Kernel engineering
Reducing contention in a histogram kernel
A change to the counting method reduced runtime on a skewed histogram workload.
Skewed integer histogram
Recorded results. Select a series to inspect it.
AMD Instinct MI300X VF
The baseline took 86.3898 ms and the revised kernel took 0.1015 ms on the same card.
The measurement uses a skewed integer histogram under rocm 7.2.4. The speedup compares these two implementations on that workload.
The problem
Many input keys reached the same histogram bins, so global atomic updates competed for those locations.
The method
The baseline made one global atomic update per key. The revised kernel counted within each block and merged those counts.
Validation
The revised kernel matched the processor reference across five rounds with fresh inputs and a prefilled output buffer.
Recorded measurements
| Method | Recorded result |
|---|---|
| Global atomic baseline | 86.3898 ms |
| Larger block and vector loads | 86.3963 ms |
| Wave aggregated atomics | 73.2241 ms |
| Block local counts and merge | 0.1015 ms |
The measurement uses a skewed integer histogram under rocm 7.2.4. The speedup compares these two implementations on that workload.
Source material
Start with a real operating problem