Blog

Notes from the bench

Porting, proving and keeping kernels fast on any chip.

Featured

Engine

AI writes the code. We prove it on every chip and keep it working.

One loop ports, fixes, optimizes, proves and maintains your kernels. Every job uses the same intake, kernel check, hardware verification, expert backup and single pull request. You click merge.

More posts

Latest

Proof

Why a signed pass matters when AI writes your kernels

AI writes the code fast. A signed Ed25519 proof pack shows it runs right on the real chip, and you can check it offline.

Formats

Moving to FP8 without losing accuracy

Whole model ports keep 99 percent of the task score for FP8 and INT8. Here is how a pass is defined.

Porting

Adding a second chip without rewriting your kernels

Porting across CUDA, HIP, NKI, Pallas and Triton. The engine finishes what hipify leaves.

Upgrades

Firmware and driver upgrades without regressions

Rechecks on ROCm, CUDA, Neuron and firmware releases, with a CI gate that warns first and blocks only when you choose.

Get new posts by email

Short notes on kernels, chips and proof. One email when a post goes up.