The Compute Optimization Engine remembers every trace, fix and pass from firmware to kernels. What it learns on one chip runs on every chip.
Four layers of memory
Each run adds to four layers. Scroll to open them.
Recall in the CLI
One command recalls a proved recipe, adapts it to the new chip and runs Verify on real hardware.
Shared memory
A fix found on one chip goes into vector memory and out to every other chip. A breaking driver or firmware release found on one job protects every customer on that chip.
Fewer steps every run
Procedural memory replays the steps that worked. Each repeat of a kernel family takes fewer steps to a verified pass.
{ "recipe": "fp8-gemm-tile-128x256", "learned_on": "H100", "applied_to": ["MI300X", "Trainium", "Gaudi"], "family": "gemm.fp8.block_scaled", "steps": [ "read counters, find the stall", "retile to 128x256, split K by 2", "double buffer the shared tile", "retune per shape range" ], "verify": "computelab verify --shapes 64 --holdout 0.2", "result": "pass", "signed": "ed25519" }
What the engine learns
Every capability feeds the same memory, the same Verify and the same single pull request.
Every run, trace, failure, fix and pass on every chip is kept and recalled before new code is written.
Each fix is stored as ordered steps with its verify command, so the next run takes fewer steps.
One index across every chip maker. Profiler counters, tuning guides, errata, PTX, AMDGCN, Neuron and TPU.
Every driver, compiler and firmware change is indexed and matched to the kernels you watch.
A specialist model trains with real-chip Verify as its reward, so cost per kernel falls over time.
Every draft and review edit feeds back. The zero-edit pass rate is tracked per chip and kernel family.
A better model is swapped in the day it ships, with no change on your side.
A breaking firmware or driver release found on one job protects every customer on that chip.
The engine watches upstream projects and alerts you before an API or kernel path is retired.
Failed attempts are kept as learning, so the engine never pays for the same dead end twice.
Every passed kernel joins the qualified catalog, ready to recall on the next job.
The engine learns new silicon before launch, so day one support comes with the chip.
Better tomorrow
The day a stronger model ships, the engine runs on it. Your jobs get faster and cheaper without a migration.
The pre-release chip lab teaches the engine each new chip early. Your kernels are ready when the chip is.
Every passed kernel from every job adds to the qualified catalog. Each new job starts from more proved work.
Put your first kernel into memory. Every run after it gets faster.