Vector memory, procedural memory and recall

Every run makes the next one faster. On any chip.

The Compute Optimization Engine remembers every trace, fix and pass from firmware to kernels. What it learns on one chip runs on every chip.

Four layers of memory

From the bare metal to recall

Each run adds to four layers. Scroll to open them.

01

Traces from the bare metal

Every run on every chip is kept. Profiler counters, firmware and driver versions, power and thermal readings, failures and passes all go into memory.
02

Fix recipes

Procedural memory stores how each fix was done. It keeps the ordered steps and the verify command, so the next run takes fewer steps.
03

Cross-chip graph

Kernels, shapes, chips and releases are linked in one graph. A fix on H100 is mapped to MI300X, Trainium, TPU, Gaudi and Snapdragon.
04

Vector recall

Before any new code is written, the engine recalls the closest proved fix. An expert fix is paid for once and reused on every chip.

Recall in the CLI

Learned on H100. Proved on MI300X.

One command recalls a proved recipe, adapts it to the new chip and runs Verify on real hardware.

  computelab recall

recall match0.00

Shared memory

Learned once. Used on every chip.

A fix found on one chip goes into vector memory and out to every other chip. A breaking driver or firmware release found on one job protects every customer on that chip.

Fewer steps every run

The same job, shorter each time

Procedural memory replays the steps that worked. Each repeat of a kernel family takes fewer steps to a verified pass.

Example Steps to a verified pass, Example
  recipe.json, stored in procedural memory
{
  "recipe": "fp8-gemm-tile-128x256",
  "learned_on": "H100",
  "applied_to": ["MI300X", "Trainium", "Gaudi"],
  "family": "gemm.fp8.block_scaled",
  "steps": [
    "read counters, find the stall",
    "retile to 128x256, split K by 2",
    "double buffer the shared tile",
    "retune per shape range"
  ],
  "verify": "computelab verify --shapes 64 --holdout 0.2",
  "result": "pass",
  "signed": "ed25519"
}

What the engine learns

Twelve ways it gets better

Every capability feeds the same memory, the same Verify and the same single pull request.

Kernel memory

Run and repair memory

Every run, trace, failure, fix and pass on every chip is kept and recalled before new code is written.

Procedural

Procedural memory

Each fix is stored as ordered steps with its verify command, so the next run takes fewer steps.

Index

Hardware knowledge index

One index across every chip maker. Profiler counters, tuning guides, errata, PTX, AMDGCN, Neuron and TPU.

Releases

Release impact search

Every driver, compiler and firmware change is indexed and matched to the kernels you watch.

Specialist

Specialist learning loop

A specialist model trains with real-chip Verify as its reward, so cost per kernel falls over time.

Corrections

Correction pipeline

Every draft and review edit feeds back. The zero-edit pass rate is tracked per chip and kernel family.

Models

Model mix that upgrades itself

A better model is swapped in the day it ships, with no change on your side.

Warnings

Shared warnings

A breaking firmware or driver release found on one job protects every customer on that chip.

Upstream

Upstream watch and retirement alerts

The engine watches upstream projects and alerts you before an API or kernel path is retired.

Negative results

Every failure teaches

Failed attempts are kept as learning, so the engine never pays for the same dead end twice.

Catalog

A catalog that grows

Every passed kernel joins the qualified catalog, ready to recall on the next job.

New silicon

Private pre-release chip lab

The engine learns new silicon before launch, so day one support comes with the chip.

Better tomorrow

Whatever it is today, it is better tomorrow

Models

New models swapped in

The day a stronger model ships, the engine runs on it. Your jobs get faster and cheaper without a migration.

Silicon

New silicon learned before launch

The pre-release chip lab teaches the engine each new chip early. Your kernels are ready when the chip is.

Catalog

Catalog growth

Every passed kernel from every job adds to the qualified catalog. Each new job starts from more proved work.

Put your first kernel into memory. Every run after it gets faster.