Case study. NCCL tuning and verified GPU kernels

Computelab.co cut exposed GPU transfer time from 3.836 ms to 0.473 ms

Computelab.co gives a training team shorter steps on the same GPUs and a signed certificate on every result.

4 min read, 909 words

0.473 msexposed transfer per training step, down from 3.836 ms
48 of 48test cases bit exact on Blackwell
4 chip familieseach result signed, on NVIDIA, AMD and AWS hardware
Snapshot
Workload
GPU training step and multi GPU communication and an fp8 inference kernel
Chips
NVIDIA Blackwell, AMD Instinct MI300X, AWS Trainium and NVIDIA T4
Layers
Network and interconnect, kernels, and a port to a new chip
Outcome
Shorter steps, the same model outputs and signed proof
The pain

Idle GPUs cost the same as busy ones

Picture the engineer who owns training throughput for a GPU cluster. Every morning the utilization chart shows cards sitting idle between bursts of math. In those gaps the GPUs wait on each other to trade results. The cloud bill still charges for every second of that waiting.

The right communication settings move with the chip, the driver and the library version. So every upgrade puts the old tuning in doubt, and a test on the full cluster costs real money.

Results

Shorter training steps on the same cluster

On a Tesla T4, exposed transfer fell from 3.836 ms to 0.473 ms per step. The whole step fell from 35.93 ms to 32.568 ms. The output matched the original step exactly, with zero difference.

Each millisecond of exposed transfer is a millisecond where paid GPUs do no math. The step came back about 3.36 ms shorter, and a training run repeats it thousands of times. So the same cluster finishes the same run in fewer paid hours.

One training step on a Tesla T4, certificate CLV-F84C9-F4D88-A77CC-A074D
Reference stepcompute 32.094 msexposed transfer 3.836 ms35.93 msOverlapped stepcompute 32.094 msexposed transfer 0.473 ms32.568 msdark is compute, 32.094 ms, and coral is exposed transfer, 3.836 ms and then 0.473 ms
New silicon

Same model outputs on Blackwell

Computelab.co ported the vLLM fp8 quantization kernel to an RTX PRO 6000 Blackwell. All 48 of 48 test cases came out bit exact against the original kernel. So a team moves to Blackwell and keeps the same model outputs.

Across vendors

Signed results on AMD and AWS chips

The same signed proof covers other vendors' hardware. A histogram kernel passed every check on an AMD Instinct MI300X. A bin sum kernel passed every check on AWS Trainium trn1.

Network layer

The right NCCL setting at every message size

The fastest NCCL setting changes as messages grow. Computelab.co finds the winner at every size on the buyer's own chips. When a cluster moves to a new NCCL release, Computelab.co shows which setting changed. So the engineer knows the answer before renting the full cluster.

Signed proof

Results a buyer checks on their own machine

Every result carries a signed certificate. It names the chip, the driver and the library version behind the result. A buyer checks it on their own machine, and any edit makes the check fail. A lead can hand the result to finance or a customer to check for themselves.

  • Shorter training step on a Tesla T4CLV-F84C9-F4D88-A77CC-A074D
  • fp8 port, 48 of 48 bit exact on an RTX PRO 6000 BlackwellCLV-CB8F2-D915E-20EB3-96E8D
  • Histogram kernel passed on an AMD Instinct MI300XCLV-8270A-FA366-B69EB-A19A8
  • Bin sum kernel passed on AWS Trainium trn1CLV-87BEA-81669-C809A-88263
Why this worked

Why this worked

  • The engine measured each result on real hardware, so every number came from the chip itself.
  • It worked one layer at a time, from the network to the kernels to a new chip.
  • Each fix proved on one chip becomes the first thing the engine tries on the next one.

On a Tesla T4, the transfer wait fell from 3.836 ms to 0.473 ms, and the output matched the original exactly.

FAQ

Questions buyers ask

How does Computelab.co choose the best NCCL algorithm and protocol?

It measures every route and packing method at every message size on the buyer's chips. The study ran 2 RTX 3070 cards on NCCL 2.20.5. Tree with Simple won all reduce from 256 KB to 4 MB. Ring with Simple won from 8 MB to 1 GB. The top all reduce bus bandwidth was 5.42 GB/s at 512 MB.

Best all reduce bus bandwidth by message size, 2x RTX 3070, NCCL 2.20.5
02468 B1 KB1 MB1 GBmessage size, log scaleGB/s8 B, Ring Simple, 0.00 GB/s16 B, Tree LL128, 0.00 GB/s32 B, Tree LL128, 0.00 GB/s64 B, Tree LL128, 0.00 GB/s128 B, Ring Simple, 0.00 GB/s256 B, Tree LL128, 0.01 GB/s512 B, Tree LL128, 0.01 GB/s1 KB, Ring Simple, 0.03 GB/s2 KB, Ring Simple, 0.05 GB/s4 KB, Tree LL128, 0.11 GB/s8 KB, Ring LL, 0.21 GB/s16 KB, Ring LL128, 0.41 GB/s32 KB, Ring LL, 0.81 GB/s64 KB, Ring Simple, 1.61 GB/s128 KB, Tree LL128, 2.83 GB/s256 KB, Tree Simple, 4.01 GB/s512 KB, Tree Simple, 4.73 GB/s1 MB, Tree Simple, 5.18 GB/s2 MB, Tree Simple, 5.39 GB/s4 MB, Tree Simple, 5.27 GB/s8 MB, Ring Simple, 5.32 GB/s16 MB, Ring Simple, 5.34 GB/s32 MB, Ring Simple, 5.39 GB/s64 MB, Ring Simple, 5.37 GB/s128 MB, Ring Simple, 5.39 GB/s256 MB, Ring Simple, 5.41 GB/s512 MB, Ring Simple, 5.42 GB/s1 GB, Ring Simple, 5.33 GB/s5.42 GB/s at 512 MB
Ring SimpleTree SimpleTree LL128Ring LLRing LL128
Does NCCL tuning hold after a driver or library upgrade?

Computelab.co tunes again on the new version and shows which setting changed. This study holds tuned results for NCCL 2.20.5 on RTX 3070 cards and NCCL 2.28.9 on a Tesla T4.

Which kernels are in this study?

The Blackwell port moved the vLLM per token group fp8 quantization kernel to CuTeDSL. On the MI300X, a HIP histogram kernel passed every check. An NKI bin sum kernel did the same on Trainium.

How can an engineer check a Computelab.co result?

Each result carries a signed certificate. The engineer checks it on their own machine, and any change to the result makes the check fail.

Which chips can the engine port a kernel to?

This study holds signed results on NVIDIA Blackwell and T4, an AMD Instinct MI300X and AWS Trainium trn1. The engine covers NVIDIA, AMD, Intel, Google TPU, AWS Trainium, Qualcomm, Apple and Arm.

Which layers of the machine does the engine cover?

It covers nine layers, starting with firmware and BIOS, board bring up, drivers and runtimes and OS kernel tuning. Then come host systems, compilers, kernels, network and interconnect, and power and thermal.

Next step

Bring the workload that costs the most GPU hours

Send Computelab.co the job that eats the most GPU time on your cluster. The engine measures it on real hardware, tunes it layer by layer and returns a signed proof pack. The first check and first solution are free. See how the Compute Optimization Engine works, compare membership plans, or start free.