Sign inLaunch

CUDA Kernel Optimizer

A langgraph based workflow with a C++ CUDA harness to optimize CUDA kernels

Visit website ↗Upvoting opens on launch day.

An agentic CUDA kernel optimizer that turns workload descriptions into GPU implementations through a loop of code generation, correctness checks, benchmarking and refinement. A LangGraph agent explores kernel code and launch configurations, can research NVIDIA documentation and Nsight Compute counters, and keeps the fastest validated implementation. First shared by its maker on Show HN: https://news.ycombinator.com/item?id=49842596

✨ The overview, FAQ and details below were drafted by AI from the product's own website and haven't been reviewed by the maker yet. Are you the maker? Claim this listing to review and edit them.

Overview

Who is it for?

GPU programmers and ML engineers optimizing CUDA kernels.

Problem

Hand-tuning CUDA kernels and launch configurations is slow and requires deep expertise.

Solution

An automated generate-validate-benchmark loop with a C++ harness that compiles with NVRTC and runs via the CUDA Driver API.

What makes it unique

Every candidate must pass validation against a reference, and ranking uses the geometric mean of latency across cases.

CUDA Kernel Optimizer FAQ

How does it validate kernels?

Outputs are compared with a reference kernel using NumPy, and every case must pass.

What does it save?

The fastest validated candidate, the execution history and a timing heatmap.

Details

Pricing
Open Source
Built for
Developers
Platforms
CLI

Hunted by AI Launch Team

Discussion (0)

Sign in to join the discussion.