
CUDA Kernel Optimizer
A langgraph based workflow with a C++ CUDA harness to optimize CUDA kernels
An agentic CUDA kernel optimizer that turns workload descriptions into GPU implementations through a loop of code generation, correctness checks, benchmarking and refinement. A LangGraph agent explores kernel code and launch configurations, can research NVIDIA documentation and Nsight Compute counters, and keeps the fastest validated implementation. First shared by its maker on Show HN: https://news.ycombinator.com/item?id=49842596
✨ The overview, FAQ and details below were drafted by AI from the product's own website and haven't been reviewed by the maker yet. Are you the maker? Claim this listing to review and edit them.
Overview
Who is it for?
GPU programmers and ML engineers optimizing CUDA kernels.
Problem
Hand-tuning CUDA kernels and launch configurations is slow and requires deep expertise.
Solution
An automated generate-validate-benchmark loop with a C++ harness that compiles with NVRTC and runs via the CUDA Driver API.
What makes it unique
Every candidate must pass validation against a reference, and ranking uses the geometric mean of latency across cases.
CUDA Kernel Optimizer FAQ
How does it validate kernels?
Outputs are compared with a reference kernel using NumPy, and every case must pass.
What does it save?
The fastest validated candidate, the execution history and a timing heatmap.
Details
- Pricing
- Open Source
- Categories
- Coding & Dev ToolsAI Agents
- Built for
- Developers
- Platforms
- CLI
Hunted by AI Launch Team
Discussion (0)
Sign in to join the discussion.