Computelab.co cut exposed GPU transfer time from 3.836 ms to 0.473 ms
Computelab.co gives a training team shorter steps on the same GPUs and a signed certificate on every result.
- Workload
- GPU training step and multi GPU communication and an fp8 inference kernel
- Chips
- NVIDIA Blackwell, AMD Instinct MI300X, AWS Trainium and NVIDIA T4
- Layers
- Network and interconnect, kernels, and a port to a new chip
- Outcome
- Shorter steps, the same model outputs and signed proof
Idle GPUs cost the same as busy ones
Picture the engineer who owns training throughput for a GPU cluster. Every morning the utilization chart shows cards sitting idle between bursts of math. In those gaps the GPUs wait on each other to trade results. The cloud bill still charges for every second of that waiting.
The right communication settings move with the chip, the driver and the library version. So every upgrade puts the old tuning in doubt, and a test on the full cluster costs real money.
Shorter training steps on the same cluster
On a Tesla T4, exposed transfer fell from 3.836 ms to 0.473 ms per step. The whole step fell from 35.93 ms to 32.568 ms. The output matched the original step exactly, with zero difference.
Each millisecond of exposed transfer is a millisecond where paid GPUs do no math. The step came back about 3.36 ms shorter, and a training run repeats it thousands of times. So the same cluster finishes the same run in fewer paid hours.
Same model outputs on Blackwell
Computelab.co ported the vLLM fp8 quantization kernel to an RTX PRO 6000 Blackwell. All 48 of 48 test cases came out bit exact against the original kernel. So a team moves to Blackwell and keeps the same model outputs.
Signed results on AMD and AWS chips
The same signed proof covers other vendors' hardware. A histogram kernel passed every check on an AMD Instinct MI300X. A bin sum kernel passed every check on AWS Trainium trn1.
The right NCCL setting at every message size
The fastest NCCL setting changes as messages grow. Computelab.co finds the winner at every size on the buyer's own chips. When a cluster moves to a new NCCL release, Computelab.co shows which setting changed. So the engineer knows the answer before renting the full cluster.
Results a buyer checks on their own machine
Every result carries a signed certificate. It names the chip, the driver and the library version behind the result. A buyer checks it on their own machine, and any edit makes the check fail. A lead can hand the result to finance or a customer to check for themselves.
- Shorter training step on a Tesla T4
CLV-F84C9-F4D88-A77CC-A074D - fp8 port, 48 of 48 bit exact on an RTX PRO 6000 Blackwell
CLV-CB8F2-D915E-20EB3-96E8D - Histogram kernel passed on an AMD Instinct MI300X
CLV-8270A-FA366-B69EB-A19A8 - Bin sum kernel passed on AWS Trainium trn1
CLV-87BEA-81669-C809A-88263
Why this worked
- The engine measured each result on real hardware, so every number came from the chip itself.
- It worked one layer at a time, from the network to the kernels to a new chip.
- Each fix proved on one chip becomes the first thing the engine tries on the next one.
On a Tesla T4, the transfer wait fell from 3.836 ms to 0.473 ms, and the output matched the original exactly.
Questions buyers ask
How does Computelab.co choose the best NCCL algorithm and protocol?
It measures every route and packing method at every message size on the buyer's chips. The study ran 2 RTX 3070 cards on NCCL 2.20.5. Tree with Simple won all reduce from 256 KB to 4 MB. Ring with Simple won from 8 MB to 1 GB. The top all reduce bus bandwidth was 5.42 GB/s at 512 MB.
Does NCCL tuning hold after a driver or library upgrade?
Computelab.co tunes again on the new version and shows which setting changed. This study holds tuned results for NCCL 2.20.5 on RTX 3070 cards and NCCL 2.28.9 on a Tesla T4.
Which kernels are in this study?
The Blackwell port moved the vLLM per token group fp8 quantization kernel to CuTeDSL. On the MI300X, a HIP histogram kernel passed every check. An NKI bin sum kernel did the same on Trainium.
How can an engineer check a Computelab.co result?
Each result carries a signed certificate. The engineer checks it on their own machine, and any change to the result makes the check fail.
Which chips can the engine port a kernel to?
This study holds signed results on NVIDIA Blackwell and T4, an AMD Instinct MI300X and AWS Trainium trn1. The engine covers NVIDIA, AMD, Intel, Google TPU, AWS Trainium, Qualcomm, Apple and Arm.
Which layers of the machine does the engine cover?
It covers nine layers, starting with firmware and BIOS, board bring up, drivers and runtimes and OS kernel tuning. Then come host systems, compilers, kernels, network and interconnect, and power and thermal.
Bring the workload that costs the most GPU hours
Send Computelab.co the job that eats the most GPU time on your cluster. The engine measures it on real hardware, tunes it layer by layer and returns a signed proof pack. The first check and first solution are free. See how the Compute Optimization Engine works, compare membership plans, or start free.