Documentation
README
AKO4ALL — Agentic Kernel Optimization
Drive a profile → modify → benchmark → log → commit loop on a GPU kernel until it runs faster than the reference. The user provides at minimum a kernel; everything else (reference, inputs, bench script, hints) is optional.
When this skill applies
- "optimize this kernel" / "speed up this CUDA / Triton / TileLang kernel"
- "run AKO / AKO4ALL on ..."
- "benchmark this kernel against PyTorch"
- "iterate on this kernel until it's faster"
- mentions of
ncu, kernel profiling, GPU speedup target
This is the opening of the README. Read the full README on GitHub.