◇Compute Lab

Embodied AI

Reducing overhead in an action sampler

Manual graph capture shortened the measured step time for a batch one action sampler.

1Batch size
10Sampling steps
5.9 timesSpeedup against eager execution
RUNTIME COMPARISONNVIDIA T4

Action-model sampling loop

Stated baseline4.818 ms
Manual graph capture0.819 ms
Batch 1 · 10 sampling steps · 25.5M parameters

Recorded results. Select a series to inspect it.

NVIDIA Tesla T4

Eager execution took 4.818 ms per step. Manual graph capture took 0.819 ms under the measured configuration.

The result uses a 25.5 million parameter action expert on a t4 at batch one with ten sampling steps.

The problem

The action sampler called many small operations, so execution overhead affected its step time.

The method

Manual cuda graph capture replayed the sampling loop. The study compared that path with eager execution and compiler graph capture.

Validation

The recorded checks cover exact graph replay, stale inputs and memory leaks. Separate low bit experiments measured memory and latency.

Recorded measurements

MethodRecorded result
Eager execution4.818 ms
Compiler graph capture0.955 ms
Manual graph capture0.819 ms

The result uses a 25.5 million parameter action expert on a t4 at batch one with ten sampling steps.

The separate low-bit experiments reduced weight storage but did not improve batch-one latency in the tested configurations.

Source material

Start with a real operating problem

What needs to work better?