Reproducing a LeRobot regression with TraceML
For now, I tried narrowing this down a little further:
* * *
Following the suggestion above to correlate the slow fetches with worker/cache/storage behavior, I ran two small DataLoader-only controls against the same historical regression.
The short version is:
> **the recurring stall cadence looks genuinely coupled to the DataLoader worker pool, but not to ordinary worker restarts and not primarily to FIFO return ordering.**
The expensive work still appears to be on the historical dataset-access side. The worker configuration seems to determine **how that cost becomes visible over time**.
The first control was simply to keep the historical broken stack fixed and vary `num_workers`.
Condition | What happened
---|---
broken, `num_workers=0` | essentially every fetch was slow
broken, `num_workers=1` | large stalls essentially every step
broken, `num_workers=2` | large stalls recurred about every 2 steps
broken, `num_workers=4` | large stalls recurred every 4 steps
fixed, `num_workers=4` | no >5 s stalls
I ran the broken sweep in both directions (`4→2→1→0` and `0→1→2→4`) to make order/cache effects a little less likely to fool me.
For the runs with enough >5 s stalls:
* `workers=1` → median stall gap ≈ **1**
* `workers=2` → median stall gap = **2**
* `workers=4` → median stall gap = **4**
* in the 2-worker and 4-worker runs, the observed large-stall gaps were exact multiples of the worker count
So the four-step pattern from the earlier T4 reproduction no longer looks like an arbitrary periodicity.
The second control was useful because it falsified one of my own first guesses.
I had suspected that ordered DataLoader return might be producing a simple head-of-line effect: one slow required batch blocks the consumer while later batches become ready behind it, producing one long wait followed by several nearly instantaneous fetches.
PyTorch does normally return multiprocessing DataLoader batches in order, and the pinned PyTorch 2.7.1 used here already has both prefetch_factor and `in_order=True` by default.
So I held everything else fixed at:
num_workers = 4
prefetch_factor = 2
same historical broken LeRobot revision
same datasets==4.1.1 environment
same deterministic sampler stream
and compared:
in_order=True
vs.
in_order=False
I also recorded the original sampler-batch rank of every returned batch, so I could verify that `in_order=False` was actually allowing later-ready batches to overtake earlier ones.
It did.
Across the two unordered runs I observed **14 overtaking deliveries** , while the ordered runs had none.
But the large-stall cadence remained:
ordered:
stall steps = 0, 4, 8, 12
unordered:
stall steps = 0, 4, 8, 12
Both conditions still had:
4 stalls > 5 s out of 16 deliveries
median large-stall gap = 4
That seems like a useful negative result.
**Removing FIFO enforcement changed the delivery order, but it did not remove the four-worker cadence.**
So I would now weaken my earlier head-of-line interpretation substantially.
There was still some effect on the shape of the tail:
Broken condition | mean fetch | p95 | >5 s stalls | stall gap
---|---|---|---|---
ordered, pass 1 | ~4.71 s | ~19.28 s | 4 | 4
unordered, pass 1 | ~4.31 s | ~16.67 s | 4 | 4
unordered, pass 2 | ~4.38 s | ~16.43 s | 4 | 4
ordered, pass 2 | ~4.62 s | ~17.75 s | 4 | 4
So ordering can redistribute some of the waiting and change which batch exposes it, but it does **not** appear sufficient to explain the recurring large stalls.
The interpretation I currently find most consistent with the observations is therefore closer to:
expensive historical dataset access
↓
substantial CPU-side producer work
↓
multi-worker DataLoader / prefetch pool
↓
worker-count-coupled availability pattern
↓
consumer sees ~1 / ~2 / ~4 step stall cadence
depending on worker count
rather than:
FIFO ordering
↓
four-step stalls
In other words, I think it is useful to separate two questions:
1. **Why is producing these batches expensive?**
2. **Why does that expense appear to the training loop in this particular temporal pattern?**
For this historical case, the first question was already answered upstream: LeRobot PR #2408 traced the training slowdown to `_query_hf_dataset` and changed the dataset access path. I am not trying to rediscover or reassign that root cause.
The new part here is mostly about the second question.
A few other observations helped narrow it further:
* I saw **no worker PID-set changes** during any of the measured broken runs, so ordinary worker restart/reset does not look like a good explanation for recurring within-run stalls.
* Slow intervals contained much more aggregate worker CPU time than fast intervals.
* The median worker physical `read_bytes` during the slow intervals was 0, and the median major-page-fault count was also 0.
* The known-fixed revision with four workers again had **no >5 s stalls**.
So, from the counters I collected, the evidence points more strongly toward **CPU-side dataset processing** than toward recurring physical storage reads.
I would not stretch that farther than the measurements allow: `read_bytes == 0` does not prove that every cache/memory effect is irrelevant, I did not measure CPU hardware-cache misses, and I did not independently isolate the OS scheduler.
One other nice sanity check was that removing ACT compute did not destroy the original heavy-tail shape.
The previous full ACT T4 reproduction looked approximately like:
broken pair 1:
mean input wait ≈ 4.95 s
median ≈ a few ms
p95 ≈ 19.75 s
broken pair 2:
mean input wait ≈ 4.48 s
median ≈ a few ms
p95 ≈ 18.90 s
The short DataLoader-only four-worker control produced roughly:
pass 1:
mean ≈ 4.74 s
median ≈ 1.8 ms
p95 ≈ 19.58 s
pass 2:
mean ≈ 4.41 s
median ≈ 4.0 ms
p95 ≈ 17.23 s
I would **not** call those numbers a strict numerical reproduction—the short probe deliberately uses artificial consumer pacing rather than real ACT compute—but qualitatively the same unusual shape survives:
> most delivered batches are immediately available, while a small regularly spaced subset dominates the wall time.
For reference, the earlier full T4 reproduction is still here:
Executed Colab T4 reproduction notebook
There is also an important host-specific caveat: this Colab T4 instance exposed only **2 logical CPUs**.
So I would not infer from this that, for example, `num_workers=2` is generally preferable to 4. Four CPU-heavy workers contending on a two-CPU host can obviously affect absolute stall duration.
What seems much more reusable to me is the **structural change** under the controlled variation:
worker count changes
↓
stall cadence changes with it
rather than the absolute number of seconds.
This actually reinforces my impression of where a tool like TraceML is useful.
The compact diagnosis answered the first important question correctly:
Where did the extra training time go?
INPUT-BOUND
Then the _shape_ of the input timing suggested which cheap experiment to try next:
INPUT-BOUND
↓
uniformly slow, or heavy-tailed?
↓
heavy-tailed
↓
random, or temporally structured?
↓
structured
↓
vary one cheap parameter
↓
cadence follows num_workers
I do not think TraceML needs to become a full DataLoader/PyTorch profiler to make this useful.
If you are deciding what belongs in a compact comparison, I would personally find a very small amount of distribution/temporal evidence valuable—enough to distinguish something like:
uniformly slow input
from:
mostly-fast input with recurring large stalls
because those two cases suggest rather different next experiments even if their **mean input wait is identical**.
The current TraceML README already points an `INPUT-BOUND` diagnosis toward workers, transforms, tokenization, collation, or storage. This case makes me think the useful extra layer may be less about automatically naming the final root cause, and more about preserving enough timing shape to select a cheap discriminator.
Exact control setup and the full worker-count result (click for more details) The `in_order` control and why I changed my interpretation (click for more details) Historical LeRobot context and why I would keep this version-pinned (click for more details) What I would test next — only if someone wants the lower-level mechanism (click for more details)
So my updated reading of the whole case is roughly:
TraceML:
correctly points to INPUT-BOUND
↓
per-step data:
shows a strongly heavy-tailed / periodic input pattern
↓
cheap worker-count control:
shows the cadence follows worker count
↓
PID + I/O counters:
weaken worker-restart and recurring-physical-I/O explanations
↓
in_order control:
produces real batch overtaking,
but the four-step cadence survives
↓
current interpretation:
historical CPU-heavy dataset access is the root cost;
worker-pool behavior shapes the temporal cadence;
FIFO ordering changes the presentation somewhat,
but is not the main explanation
For me that is a useful outcome for the original TraceML question.
The compact diagnosis did not need to know the final low-level mechanism in advance. It only needed to point to the correct subsystem, while preserving enough evidence that the next cheap experiment was obvious.
That seems like a good boundary for a lightweight diagnostic tool.