Compute Optimization Engine

The bare metal performance engine.

Firmware, board, drivers, power and thermal, compilers and kernels. Tuned and proved on any chip.

curl -fsSL https://computelab.co/install.sh | sh
irm https://computelab.co/install.ps1 | iex

Start free and we connect you by email.

Then run computelab fix in your repo. Docs Install

  console  |  job overview
job        gemm_fp8 port, CUDA to HIP
chip       MI300X
contract   64 shapes, 20% held back

machine
  firmware, BMC     recorded
  driver, runtime   recorded
  NUMA, clocks      pinned
  power, thermal    sustained

kernel bill
  gemm_fp8          top share
  attn_fwd          second
  rmsnorm           third

status
  baseline          FAIL
  candidate v4      PASS
  repeatability     PASS  1000 runs
  proof pack        SIGNED
  pull request      READY
  watch             ON

The workflow

Requirement in. Proved environment out.

The engine works inside your repo, CI, cloud and machines. Your team reviews one pull request and clicks merge.

01 In

Requirement

The chip, the workload, the metric and the number to hit, sent from the CLI, an issue label, the API or MCP.

02 Connect

Your systems

Scoped access to repo, CI, cloud account and machines, with short-lived credentials per job.

03 Run

Every layer

Trace reads the machine, Search finds the fix and the engine tunes it on the real chip.

04 Out

PR, proof, artifact

One grouped pull request, a signed proof pack and an installable container or PyTorch extension.

05 Watch

Kept working

Every firmware, driver, compiler and framework release triggers a recheck and, when needed, a new fix.

The stack

What the engine runs at every layer

Nine layers, bottom to top. Each one runs through the same intake, the same proof and the same Watch.

01

Firmware and BIOS

Bare metal

The machine boots right, reports right and recovers right.

  • OpenBMC manageability
  • BMC and board telemetry
  • Boot and attestation evidence
  • Authorized flashing
  • Firmware release rechecks for chip makers
  • Recovery tested before each firmware release is adopted
02

Board bring-up

Bare metal

New boards and first silicon come up working and measured.

  • Authorized board bring-up
  • New chip bring-up on simulator or first silicon
  • Zephyr RTOS builds and emulator runs
  • Known-answer regression checks
  • A private pre-release chip lab for chip makers
  • Device identity and versions recorded
03

Drivers and runtimes

System

The right driver and runtime on every chip, kept current.

  • ROCm, CUDA and Neuron
  • JetPack, DriveOS and QAIRT
  • Rechecks on every new release
  • Shared early warnings for every customer on that chip
  • A named regression owner
  • Fleet health checks that find chips returning wrong or slow results
04

OS and kernel tuning

System

The operating system gives your workload the whole machine.

  • Cores and NUMA nodes pinned
  • Locked clocks for repeatable measurement
  • Topology, frequencies and noise captured
  • Scheduler and context switch counters
  • Tail latency tuned alongside average speed
  • Matched runs with confidence intervals
05

Network and interconnect

System

Data moves between chips, hosts and nodes at full speed.

  • Interconnect and topology plans per pipeline
  • Transport and transfer path for every link
  • Memory owner and synchronization per link
  • Contention trials under competing load
  • Distributed bottlenecks found and fixed
  • Host to device transfer profiling
06

Power and thermal

System

Speed that holds after the chip warms up, at a known energy cost.

  • Temperature, clocks and power captured with timing
  • Throttle reasons recorded
  • Early and sustained runs compared under matched conditions
  • Joules per accepted result
  • Node and rack telemetry through Redfish and BMC
  • Green Computing reports for your fleet
07

Compilers

Code

Toolchains set and checked for every target chip.

  • CUDA, HIP, NKI, Pallas, Triton and SYCL toolchains
  • Compiler releases rechecked by Watch
  • Every answer cites the chip manual page
  • Readable output in standard languages
08

Kernels

Code

Kernels that run right and fast on any chip.

  • Porting across CUDA, HIP, NKI, Pallas, Triton and SYCL
  • The engine finishes what hipify leaves
  • Search on the real chip
  • Fusion and tuning per shape range
  • FP8, FP4 and block-scaled formats
  • Blackwell, Vera Rubin, MI400 and Trainium3 tuning
09

Host systems code

Code

The systems code around the kernel runs as fast as the kernel.

  • Host C and C++ speed work
  • Leak, race and segfault fixes
  • One intake for host and accelerator code
  • Toolchain driven repairs
  • Heterogeneous targets in one delivery
  • Phone and wearable paths on Snapdragon, Hexagon, Core ML and LiteRT

149 capabilities

Ten capability layers, A to J. One engine.

From the first command to the open marketplace, every capability feeds the same intake, hardware verification, expert review and single pull request.

Layer A

Access

  • One-command terminal agent, computelab fix
  • GitHub App where a label such as slow on MI300X opens a job
  • VS Code and Neovim plugins
  • Web console for engineers, managers and auditors
  • REST API and MCP server with a cost cap per task
  • Batch submission from CI
  • An open, inspectable client
  • Connectors for Buildkite, GitHub Actions, Jenkins, Slack, Prometheus, Grafana and Terraform
Layer B

Intelligence

  • Your choice of our managed models
  • Bring your own key
  • Fully self-hosted models
  • A better model swapped in the day it is released
  • Every answer cites the chip manual page
  • Run and repair memory, so an expert fix is paid for once and reused
Layer C

Understanding the workload

  • An acceptance contract agreed first
  • Six intake questions
  • Five problem types, wrong, slow, fails to run, broke after upgrade, needs proof
  • A kernel bill ranking kernels by share of compute time
  • Machine capture of firmware, drivers, NUMA, topology, clocks and thermal state
  • Host and device boundaries
  • Distributed and interconnect bottlenecks
  • Memory budgets
  • Interconnect topology captured per node
  • A migration comparison before a chip move
  • A finance view of cost per accepted result
Layer D

Fixing and optimizing

  • Porting across CUDA, HIP, NKI, Pallas, Triton and SYCL
  • Finishes what hipify leaves
  • Search on the real chip
  • Fusion and tuning per shape range
  • Blackwell, Vera Rubin, MI400 and Trainium3 tuning
  • Phone and wearable paths on Snapdragon, Hexagon, Core ML and LiteRT
  • FP8, FP4 and block-scaled formats
  • Attention, MoE and KV cache
  • Video, audio, robotics and state-space models
  • Host C++ speed, leak, race and segfault fixes
  • Expert review on hard cases inside the pipeline
  • New-chip bring-up on simulator or first silicon
  • System tuning for OS, interconnect, power and thermal
Layer E

Proving

  • Runs on the real chip and the real machine
  • Firmware, driver, power and temperature recorded with every result
  • Early and sustained runs compared under matched conditions
  • Error bounded against an FP64 answer across at least 64 shapes
  • 20 percent of shapes held back
  • Passes at 90 percent of the vendor library median time
  • 1000 identical runs for repeatability
  • Whole model ports keep 99.9 percent of the task score, 99 percent for FP8 and INT8
  • Safety-standard evidence for MISRA, AUTOSAR and ISO
  • Energy per accepted result
  • Confidential-compute release check
  • Ed25519 signed proof pack with an offline verify script
  • A Verified by Computelab.co result page
Layer F

Delivering

  • One grouped pull request per job into your repo and CI
  • You click merge
  • Installable kernel or firmware artifact, integrated in minutes
  • A qualified dispatcher picks a proved variant per shape and chip
  • Staged fleet rollout and instant rollback
  • Ed25519 signed proof pack with every pull request
  • Kernel gain and application gain shown separately in your metric
  • Three console views for executives, platform engineers and chip maker field teams
Layer G

Keeping it working

  • Watch rechecks on BIOS, firmware, ROCm, CUDA, Neuron, JAX, PyTorch, JetPack, DriveOS and QAIRT releases
  • A CI gate that warns first and blocks when you choose
  • Shared early warnings for every customer on that chip
  • A named regression owner
  • Monthly value report
  • Fleet health checks that find chips returning wrong or slow results
  • Firmware release recheck for chip makers
Layer H

Deploying and security

  • Four deployment modes, Computelab.co cloud, your cloud account, self-hosted and air gapped
  • One sealed container per job
  • Network off, least privilege, repo mounted read only
  • Signed builds
  • Two-person release approval
  • Short-lived credentials
  • A post-quantum transport option
  • Full IP on delivered code
Layer I

The marketplace, Compute Clearinghouseâ„¢

  • Expert engineers, a real chip pool and buyer jobs sold as one outcome
  • A proved kernel catalog
  • Neutral evaluation with a signed pass or fail
  • Sealed supplier methods
  • An outcome job exchange where Verify clears each job
  • Know Your Machine and Know Your Agent views
  • Demand pooling and shared qualification campaigns
  • Design partner program and academic access
  • Engineer membership, paid per passed kernel
  • Seller bonds
  • A private pre-release chip lab
  • Sponsored driver challenges and chip-time credits
  • An open kernel challenge
Layer J

The open edge

  • A public chip and kernel registry
  • An open cross-chip benchmark and leaderboard
  • Shareable proof pages
  • Vendor comparison kits
  • Upstream fixes to Triton, ROCm and OpenBMC
  • RL training environments for AI labs

Vector memory and recall

An engine that learns

Every run, trace, fix and pass goes into one vector memory. A fix learned on one chip is recalled on every chip. Every run makes the next one faster.

Products inside the engine

Trace, Search, Verify and Green Computing

Each one is a part of the Compute Optimization Engine, and each one opens on its own in the console.

Trace

Reads the machine

Measured traces become a kernel bill, with firmware, driver, topology and thermal state captured first.

Trace
Search

Finds the fix

Searches every past run, failure, fix and pass across chips, then tunes on the real chip.

Search
Verify

Signs the result

A signed pass or fail on real hardware and a proof pack anyone checks offline.

Verify
Green Computing

Counts the energy

Energy per accepted result as a second score in search, in Watch and in every report.

Green Computing

Five problem types

Five ways a workload goes wrong. One engine for each.

01

Wrong

Outputs drift from the right answer. The engine bounds the error against FP64 on every shape.

02

Slow

It runs behind the vendor library. The engine tunes it on the real chip until it clears the bar.

03

Fails to run

It breaks at boot, build or launch on the new machine. The engine brings up the board and ports the code.

04

Broke after upgrade

A firmware, driver or framework release changed it. Watch finds the cause and restores the pass.

05

Needs proof

It works, and you need to show it. Verify signs a proof pack anyone can check.

Start to finish

17 steps from first check to kept working

Each step runs as a status in the console, so every person on your side sees where the job is.

01

Discover and join free

02

Save the project

03

Connect authorized assets

04

Establish acceptance

05

Run the free diagnostic

06

Select compatible supply

07

Approve work and resources

08

Prepare the environment

09

Measure the baseline

10

Retrieve and plan

11

Generate port or repair

12

Evaluate and iterate

13

Independent final qualification

14

Customer review and integration

15

Install and recover safely

16

Maintain and renew

17

Grow the shared marketplace

Access and integrations

Put the engine in your loop

For engineers

Terminal, editor or GitHub App

Run computelab fix, send a file from VS Code or Neovim, or add a label such as slow on MI300X to an issue.

$ computelab fix ./kernels --chip mi300x
PASS  pull request #482 ready
For platforms and agents

API or MCP

Call the REST API or the MCP server from your own tools and agents, with a cost cap per task and batch runs from CI.

POST /v1/jobs  cost_cap per task
PASS  proof pack signed
Integrations
GitHubGitHub ActionsBuildkiteJenkinsSlackPrometheusGrafanaTerraformGitLabJiraW and B

Dispatcher and Watch

Live on your fleet, kept passing

The dispatcher picks a proved variant per shape and chip, rolls out across the fleet in stages and rolls back in one step. Watch keeps every layer passing after each release.

  watch  |  mi300x fleet
release            recheck
BIOS 1.14          PASS
ROCm update        PASS
PyTorch update     FAIL  attn_fwd
  fix opened       PR #491
  dispatcher       fallback ON
BMC firmware       PASS

dispatcher
  gemm_fp8 v4      stage 1 of 4  HEALTHY
  rollback         ready

Built-in trust

Expert review and neutral proof, inside the pipeline

01 Expert review

HPC reviewers

A global bench of HPC engineers reviews hard cases inside every job.

02 Trading roots

Nanosecond standards

Reviewers from high-frequency trading set the latency and reliability checks.

03 Neutral

Fair to every vendor

One grader for every chip.

04 Regulated partners

Built for regulated partners

Building this engine for DARPA and other highly regulated partners, and working with regulators.

Plans

Start free. Scale to the fleet.

Free

Free

The CLI on your own machine, host code checks, sanitizers, a kernel scan and your first check and first solution.

Start free
Team

Team

GitHub App, console, CI gate and Watch across your repos and chips, with expert review on hard cases.

Connect your repo
Enterprise

Enterprise

Every layer from firmware up, your cloud, self-hosted or air gapped, batch from CI and the full Clearinghouse.

Start free

Coverage

Every chip and framework

NVIDIAAMDIntelGoogle TPUAWS TrainiumQualcommAppleArm NVIDIAAMDIntelGoogle TPUAWS TrainiumQualcommAppleArm
CUDAROCmHIPTritonSYCLPallasNKIPyTorchJAX CUDAROCmHIPTritonSYCLPallasNKIPyTorchJAX

Proved on the real machine

Every layer clears the same bar

64+

Shapes checked against an FP64 answer

90%

Of the vendor library median time, to pass

1000

Identical runs for repeatability

4

Deployment modes, Computelab.co cloud, your cloud, self-hosted, air gapped

Put the engine on your whole machine.

Your first check and first solution are free, run on real hardware with a signed proof pack.