Embodied AI
Reducing overhead in an action sampler
Manual graph capture shortened the measured step time for a batch one action sampler.
Action-model sampling loop
Recorded results. Select a series to inspect it.
NVIDIA Tesla T4
Eager execution took 4.818 ms per step. Manual graph capture took 0.819 ms under the measured configuration.
The result uses a 25.5 million parameter action expert on a t4 at batch one with ten sampling steps.
The problem
The action sampler called many small operations, so execution overhead affected its step time.
The method
Manual cuda graph capture replayed the sampling loop. The study compared that path with eager execution and compiler graph capture.
Validation
The recorded checks cover exact graph replay, stale inputs and memory leaks. Separate low bit experiments measured memory and latency.
Recorded measurements
| Method | Recorded result |
|---|---|
| Eager execution | 4.818 ms |
| Compiler graph capture | 0.955 ms |
| Manual graph capture | 0.819 ms |
The result uses a 25.5 million parameter action expert on a t4 at batch one with ten sampling steps.
The separate low-bit experiments reduced weight storage but did not improve batch-one latency in the tested configurations.
Source material
Start with a real operating problem