Firmware, board, drivers, power and thermal, compilers and kernels. Tuned and proved on any chip.
curl -fsSL https://computelab.co/install.sh | shirm https://computelab.co/install.ps1 | iexStart free and we connect you by email.
job gemm_fp8 port, CUDA to HIP chip MI300X contract 64 shapes, 20% held back machine firmware, BMC recorded driver, runtime recorded NUMA, clocks pinned power, thermal sustained kernel bill gemm_fp8 top share attn_fwd second rmsnorm third status baseline FAIL candidate v4 PASS repeatability PASS 1000 runs proof pack SIGNED pull request READY watch ON
The workflow
The engine works inside your repo, CI, cloud and machines. Your team reviews one pull request and clicks merge.
The chip, the workload, the metric and the number to hit, sent from the CLI, an issue label, the API or MCP.
Scoped access to repo, CI, cloud account and machines, with short-lived credentials per job.
Trace reads the machine, Search finds the fix and the engine tunes it on the real chip.
One grouped pull request, a signed proof pack and an installable container or PyTorch extension.
Every firmware, driver, compiler and framework release triggers a recheck and, when needed, a new fix.
The stack
Nine layers, bottom to top. Each one runs through the same intake, the same proof and the same Watch.
The machine boots right, reports right and recovers right.
New boards and first silicon come up working and measured.
The right driver and runtime on every chip, kept current.
The operating system gives your workload the whole machine.
Data moves between chips, hosts and nodes at full speed.
Speed that holds after the chip warms up, at a known energy cost.
Toolchains set and checked for every target chip.
Kernels that run right and fast on any chip.
The systems code around the kernel runs as fast as the kernel.
149 capabilities
From the first command to the open marketplace, every capability feeds the same intake, hardware verification, expert review and single pull request.
Vector memory and recall
Every run, trace, fix and pass goes into one vector memory. A fix learned on one chip is recalled on every chip. Every run makes the next one faster.
Products inside the engine
Each one is a part of the Compute Optimization Engine, and each one opens on its own in the console.
Measured traces become a kernel bill, with firmware, driver, topology and thermal state captured first.
TraceSearches every past run, failure, fix and pass across chips, then tunes on the real chip.
SearchA signed pass or fail on real hardware and a proof pack anyone checks offline.
VerifyEnergy per accepted result as a second score in search, in Watch and in every report.
Green ComputingFive problem types
Outputs drift from the right answer. The engine bounds the error against FP64 on every shape.
It runs behind the vendor library. The engine tunes it on the real chip until it clears the bar.
It breaks at boot, build or launch on the new machine. The engine brings up the board and ports the code.
A firmware, driver or framework release changed it. Watch finds the cause and restores the pass.
It works, and you need to show it. Verify signs a proof pack anyone can check.
Start to finish
Each step runs as a status in the console, so every person on your side sees where the job is.
Access and integrations
Run computelab fix, send a file from VS Code or Neovim, or add a label such as slow on MI300X to an issue.
$ computelab fix ./kernels --chip mi300x PASS pull request #482 ready
Call the REST API or the MCP server from your own tools and agents, with a cost cap per task and batch runs from CI.
POST /v1/jobs cost_cap per task PASS proof pack signed
Dispatcher and Watch
The dispatcher picks a proved variant per shape and chip, rolls out across the fleet in stages and rolls back in one step. Watch keeps every layer passing after each release.
release recheck BIOS 1.14 PASS ROCm update PASS PyTorch update FAIL attn_fwd fix opened PR #491 dispatcher fallback ON BMC firmware PASS dispatcher gemm_fp8 v4 stage 1 of 4 HEALTHY rollback ready
Built-in trust
A global bench of HPC engineers reviews hard cases inside every job.
Reviewers from high-frequency trading set the latency and reliability checks.
One grader for every chip.
Building this engine for DARPA and other highly regulated partners, and working with regulators.
Plans
The CLI on your own machine, host code checks, sanitizers, a kernel scan and your first check and first solution.
Start freeGitHub App, console, CI gate and Watch across your repos and chips, with expert review on hard cases.
Connect your repoEvery layer from firmware up, your cloud, self-hosted or air gapped, batch from CI and the full Clearinghouse.
Start freeCoverage
Proved on the real machine
Shapes checked against an FP64 answer
Of the vendor library median time, to pass
Identical runs for repeatability
Deployment modes, Computelab.co cloud, your cloud, self-hosted, air gapped
Put the engine on your whole machine.
Your first check and first solution are free, run on real hardware with a signed proof pack.