#TraceML
TraceML pairs human and agent ML development under a shared version-level schema, showing experts make larger, more deliberate edits than auto-research agents across paired Kaggle trajectories. The trace-based…

#TraceML #AIResearch #MachineLearning #AutoResearch
https://arxiv.org/abs/2608.26086
September 30, 2026 at 10:01 PM
A small diagnostic for Hugging Face Trainer data-loading bottlenecks
Quick update, and thank you again for the detailed investigation here. We looked more carefully at where HF Trainer actually prepares each batch. The H2D gap is now fixed in TraceML 0.4.1. The problem was the ordering you identified: Accelerate could move the accumulated microbatches to CUDA before `on_step_begin`, so the normal Trainer callback never saw those transfers. TraceML now measures them while Trainer collects the training batches and attaches them to the same optimizer step as the later forward, backward and optimizer work. DataLoader waiting remains a separate Input Wait measurement. I also reran a TRL LoRA workload on a T4 with accumulation settings of 4, 2 and 1. All three 100-step runs completed successfully, and H2D was reported as 0.6 ms, 0.4 ms and 0.2 ms instead of `n/a`. We added regression coverage for gradient accumulation and a real-runtime test that carries the Trainer measurements through storage and final-summary generation. This applies to the standard HF `Trainer` path with `transformers>=4.46.1`. Standalone Accelerate loops and Trainer subclasses that replace `get_batch_samples` are still separate cases. Install the release with: `pip install -U "traceml-ai[hf]==0.4.1"` Colab: colab.research.google.com ### Google Colab If you have time to rerun the notebook with 0.4.1, I would be very interested to see whether the corrected H2D measurement matches what you expected from your original reproduction.
discuss.huggingface.co
September 24, 2026 at 3:25 PM
Reproducing a LeRobot regression with TraceML
For now, I tried narrowing this down a little further: * * * Following the suggestion above to correlate the slow fetches with worker/cache/storage behavior, I ran two small DataLoader-only controls against the same historical regression. The short version is: > **the recurring stall cadence looks genuinely coupled to the DataLoader worker pool, but not to ordinary worker restarts and not primarily to FIFO return ordering.** The expensive work still appears to be on the historical dataset-access side. The worker configuration seems to determine **how that cost becomes visible over time**. The first control was simply to keep the historical broken stack fixed and vary `num_workers`. Condition | What happened ---|--- broken, `num_workers=0` | essentially every fetch was slow broken, `num_workers=1` | large stalls essentially every step broken, `num_workers=2` | large stalls recurred about every 2 steps broken, `num_workers=4` | large stalls recurred every 4 steps fixed, `num_workers=4` | no >5 s stalls I ran the broken sweep in both directions (`4→2→1→0` and `0→1→2→4`) to make order/cache effects a little less likely to fool me. For the runs with enough >5 s stalls: * `workers=1` → median stall gap ≈ **1** * `workers=2` → median stall gap = **2** * `workers=4` → median stall gap = **4** * in the 2-worker and 4-worker runs, the observed large-stall gaps were exact multiples of the worker count So the four-step pattern from the earlier T4 reproduction no longer looks like an arbitrary periodicity. The second control was useful because it falsified one of my own first guesses. I had suspected that ordered DataLoader return might be producing a simple head-of-line effect: one slow required batch blocks the consumer while later batches become ready behind it, producing one long wait followed by several nearly instantaneous fetches. PyTorch does normally return multiprocessing DataLoader batches in order, and the pinned PyTorch 2.7.1 used here already has both prefetch_factor and `in_order=True` by default. So I held everything else fixed at: num_workers = 4 prefetch_factor = 2 same historical broken LeRobot revision same datasets==4.1.1 environment same deterministic sampler stream and compared: in_order=True vs. in_order=False I also recorded the original sampler-batch rank of every returned batch, so I could verify that `in_order=False` was actually allowing later-ready batches to overtake earlier ones. It did. Across the two unordered runs I observed **14 overtaking deliveries** , while the ordered runs had none. But the large-stall cadence remained: ordered: stall steps = 0, 4, 8, 12 unordered: stall steps = 0, 4, 8, 12 Both conditions still had: 4 stalls > 5 s out of 16 deliveries median large-stall gap = 4 That seems like a useful negative result. **Removing FIFO enforcement changed the delivery order, but it did not remove the four-worker cadence.** So I would now weaken my earlier head-of-line interpretation substantially. There was still some effect on the shape of the tail: Broken condition | mean fetch | p95 | >5 s stalls | stall gap ---|---|---|---|--- ordered, pass 1 | ~4.71 s | ~19.28 s | 4 | 4 unordered, pass 1 | ~4.31 s | ~16.67 s | 4 | 4 unordered, pass 2 | ~4.38 s | ~16.43 s | 4 | 4 ordered, pass 2 | ~4.62 s | ~17.75 s | 4 | 4 So ordering can redistribute some of the waiting and change which batch exposes it, but it does **not** appear sufficient to explain the recurring large stalls. The interpretation I currently find most consistent with the observations is therefore closer to: expensive historical dataset access ↓ substantial CPU-side producer work ↓ multi-worker DataLoader / prefetch pool ↓ worker-count-coupled availability pattern ↓ consumer sees ~1 / ~2 / ~4 step stall cadence depending on worker count rather than: FIFO ordering ↓ four-step stalls In other words, I think it is useful to separate two questions: 1. **Why is producing these batches expensive?** 2. **Why does that expense appear to the training loop in this particular temporal pattern?** For this historical case, the first question was already answered upstream: LeRobot PR #2408 traced the training slowdown to `_query_hf_dataset` and changed the dataset access path. I am not trying to rediscover or reassign that root cause. The new part here is mostly about the second question. A few other observations helped narrow it further: * I saw **no worker PID-set changes** during any of the measured broken runs, so ordinary worker restart/reset does not look like a good explanation for recurring within-run stalls. * Slow intervals contained much more aggregate worker CPU time than fast intervals. * The median worker physical `read_bytes` during the slow intervals was 0, and the median major-page-fault count was also 0. * The known-fixed revision with four workers again had **no >5 s stalls**. So, from the counters I collected, the evidence points more strongly toward **CPU-side dataset processing** than toward recurring physical storage reads. I would not stretch that farther than the measurements allow: `read_bytes == 0` does not prove that every cache/memory effect is irrelevant, I did not measure CPU hardware-cache misses, and I did not independently isolate the OS scheduler. One other nice sanity check was that removing ACT compute did not destroy the original heavy-tail shape. The previous full ACT T4 reproduction looked approximately like: broken pair 1: mean input wait ≈ 4.95 s median ≈ a few ms p95 ≈ 19.75 s broken pair 2: mean input wait ≈ 4.48 s median ≈ a few ms p95 ≈ 18.90 s The short DataLoader-only four-worker control produced roughly: pass 1: mean ≈ 4.74 s median ≈ 1.8 ms p95 ≈ 19.58 s pass 2: mean ≈ 4.41 s median ≈ 4.0 ms p95 ≈ 17.23 s I would **not** call those numbers a strict numerical reproduction—the short probe deliberately uses artificial consumer pacing rather than real ACT compute—but qualitatively the same unusual shape survives: > most delivered batches are immediately available, while a small regularly spaced subset dominates the wall time. For reference, the earlier full T4 reproduction is still here: Executed Colab T4 reproduction notebook There is also an important host-specific caveat: this Colab T4 instance exposed only **2 logical CPUs**. So I would not infer from this that, for example, `num_workers=2` is generally preferable to 4. Four CPU-heavy workers contending on a two-CPU host can obviously affect absolute stall duration. What seems much more reusable to me is the **structural change** under the controlled variation: worker count changes ↓ stall cadence changes with it rather than the absolute number of seconds. This actually reinforces my impression of where a tool like TraceML is useful. The compact diagnosis answered the first important question correctly: Where did the extra training time go? INPUT-BOUND Then the _shape_ of the input timing suggested which cheap experiment to try next: INPUT-BOUND ↓ uniformly slow, or heavy-tailed? ↓ heavy-tailed ↓ random, or temporally structured? ↓ structured ↓ vary one cheap parameter ↓ cadence follows num_workers I do not think TraceML needs to become a full DataLoader/PyTorch profiler to make this useful. If you are deciding what belongs in a compact comparison, I would personally find a very small amount of distribution/temporal evidence valuable—enough to distinguish something like: uniformly slow input from: mostly-fast input with recurring large stalls because those two cases suggest rather different next experiments even if their **mean input wait is identical**. The current TraceML README already points an `INPUT-BOUND` diagnosis toward workers, transforms, tokenization, collation, or storage. This case makes me think the useful extra layer may be less about automatically naming the final root cause, and more about preserving enough timing shape to select a cheap discriminator. Exact control setup and the full worker-count result (click for more details) The `in_order` control and why I changed my interpretation (click for more details) Historical LeRobot context and why I would keep this version-pinned (click for more details) What I would test next — only if someone wants the lower-level mechanism (click for more details) So my updated reading of the whole case is roughly: TraceML: correctly points to INPUT-BOUND ↓ per-step data: shows a strongly heavy-tailed / periodic input pattern ↓ cheap worker-count control: shows the cadence follows worker count ↓ PID + I/O counters: weaken worker-restart and recurring-physical-I/O explanations ↓ in_order control: produces real batch overtaking, but the four-step cadence survives ↓ current interpretation: historical CPU-heavy dataset access is the root cost; worker-pool behavior shapes the temporal cadence; FIFO ordering changes the presentation somewhat, but is not the main explanation For me that is a useful outcome for the original TraceML question. The compact diagnosis did not need to know the final low-level mechanism in advance. It only needed to point to the correct subsystem, while preserving enough evidence that the next cheap experiment was obvious. That seems like a good boundary for a lightweight diagnostic tool.
discuss.huggingface.co
September 15, 2026 at 3:14 AM
Reproducing a LeRobot regression with TraceML
For now, I tried a quick reproduction on Colab: * * * For this pinned case, yes: the compact before/after comparison did point me toward the right subsystem. I ran the merged LeRobot regression case study on a Colab T4, keeping the historical LeRobot revisions and `datasets==4.1.1`, and ran two pairs with the second pair reversing the order. Metric | Pair 1: broken → fixed | Pair 2: broken → fixed ---|---|--- TraceML diagnosis | INPUT-BOUND → COMPUTE-BOUND | INPUT-BOUND → COMPUTE-BOUND Step time | 5337.15 → 301.96 ms (-94.34%) | 4845.45 → 307.77 ms (-93.65%) Input wait | 4946.33 → 4.50 ms (-99.91%) | 4479.50 → 6.46 ms (-99.86%) Compute | 372.82 → 289.91 ms (-22.24%) | 351.09 → 293.80 ms (-16.32%) So on this host too, the obvious first place to investigate from the `traceml compare` result would have been the input side. About 98% of the total step-time reduction in both pairs is accounted for by the reduction in input wait. I put the executed two-pair notebook here, including the environment, reproduction code, comparison outputs, and the per-step analysis: Executed Colab T4 reproduction notebook The part I found especially interesting was the **shape** of the input wait, though. On the broken runs, the median fetch was only about **2–3 ms** , while p95 was about **19–20 seconds**. In other words, the loader was not uniformly taking ~5 seconds per fetch: most fetches were fast, while a minority of very large stalls dominated the run-level average. That makes the existing `INPUT-BOUND` diagnosis look useful to me for the first-level question — _where did the extra wall time go?_ — but it also made me wonder whether a small amount of temporal-tail information might be useful in the compact regression workflow when the distribution is this skewed. The case-study analysis notebook already looks at per-step timing and summary statistics such as median/p95, while traceml compare is deliberately a compact comparison of the run summaries. So I do not mean turning `compare` into a full profiler, or necessarily adding a particular p95/p99 schema. Even something that makes “mostly fast with intermittent large stalls” distinguishable from “uniformly slower input” could potentially help decide what to inspect next. The other thing I would preserve is the warning about absolute numbers. The high-level result seems much more portable than the timings themselves. The original L4 run, the separate clean-box L4 reproduction in TraceML PR #450, and this T4 run all reach `INPUT-BOUND → COMPUTE-BOUND` with input wait collapsing by roughly 99.9%, but the absolute step times and even the secondary compute behavior differ substantially across hosts. In particular, my T4 run still showed a ~16–22% compute reduction in both orders. I would **not** attribute that to the LeRobot dataset fix from this experiment alone. The input-wait reduction overwhelmingly explains the overall speedup, and the clean-box L4 experiment in #450 is a good example of why reversed-order runs matter: a large apparent compute change in its cold first pair disappeared once the order was reversed and the system was warm. So the part I would treat as robust here is: > the regression was overwhelmingly input-side, and the compact comparison identified that correctly; rather than: > every phase should reproduce the same absolute before/after timing on another machine. Colab T4 reproduction details (click for more details) What the raw input-wait distribution looked like (click for more details) Two optional directions for future regression cases (click for more details) Overall, this case did answer the question I would want the lightweight comparison to answer first: **the extra training time moved into the input path, and that is where I should investigate.** The T4 run mostly added a second piece for me: once the high-level attribution is known, a small amount of temporal-shape information can materially change what “input-bound” looks like — in this case, from an apparent multi-second loader to mostly millisecond fetches interrupted by very large stalls.
discuss.huggingface.co
September 8, 2026 at 4:55 AM
Reproducing a LeRobot regression with TraceML
While reproducing a known LeRobot training regression, I found a pattern that the average timing alone doesn’t show: most DataLoader fetches were quick, but recurring fetches took over five seconds. I maintain TraceML, an open-source PyTorch training diagnosis tool. We are building toward more automatic performance regression diagnosis: when training gets slower after a change, help identify where the extra time went. Today, you instrument the training loop, record two runs, and use `traceml compare` to compare timings and bottleneck diagnoses. For this case study, I compared LeRobot before and after upstream fix #2408. The LeRobot contributors found and fixed the bug; this experiment measures its effect. The workload was ACT, 200 steps per run, on one NVIDIA L4. Both revisions used the same dataset, training settings and TraceML instrumentation in the Accelerate-based loop. The second pair reversed execution order: Metric | Pair 1: broken → fixed | Pair 2: broken → fixed ---|---|--- Step time | 1496.9 → 116.8 ms | 1542.7 → 117.2 ms DataLoader fetch, CPU | 1375.7 → 1.5 ms | 1427.5 → 1.3 ms Compute | 115.9 → 110.5 ms | 110.4 → 111.1 ms TraceML reported **INPUT-BOUND** → **COMPUTE-BOUND** in both pairs. The recurring input stalls disappeared after the fix; compute stayed around 110–116 ms on this host. Exact timings vary across machines, and startup can affect the first pair. Inspect both orders before attributing a compute change to the fix. **To try it:** use Linux x86-64, Git, Python 3.10 with venv support, internet access and one NVIDIA GPU with a driver supporting CUDA 12.8. Ubuntu may need `sudo apt-get install python3.10-venv` first. ```bash git clone GitHub - traceopt-ai/traceml: Open-source performance diagnostics for PyTorch training runs. · GitHub cd traceml bash examples/advanced/lerobot_v3_image_regression/run_reproduction.sh ``` The runner sets up its environment, downloads the dataset, runs one pair and prints the comparison location. Add `–pairs 2` for both orders. If a dataset, dependency or configuration change has slowed your own PyTorch training, try the TraceML quickstart. Share the before/after comparison and whether it helped you find where to investigate. Looking for feedback.
discuss.huggingface.co
September 7, 2026 at 4:55 PM
Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang: TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development https://arxiv.org/abs/2608.26086 https://arxiv.org/pdf/2608.26086 https://arxiv.org/html/2608.26086
August 27, 2026 at 6:44 AM
A small diagnostic for Hugging Face Trainer data-loading bottlenecks
I kept running into a frustrating training-performance question: GPU utilization would hover around 50%, loss looked normal, and nothing was obviously broken. Was the model actually the limit, or was the GPU waiting for the next batch? Those are opposite problems with opposite fixes. But `nvidia-smi` and a healthy loss curve do not distinguish them, and I do not want to open a full PyTorch Profiler trace for every ordinary run. So I added a Hugging Face `Trainer` integration to TraceML, an open-source PyTorch diagnostics project I have been building. from traceml.integrations import huggingface as traceml_hf from traceml.integrations.huggingface import TraceMLTrainerCallback traceml_hf.init() trainer = Trainer( ..., callbacks=[TraceMLTrainerCallback()], ) It keeps the rest of the `Trainer` setup unchanged, separates input wait from step work, and gives a practical input-bound verdict. I made a runnable Colab with ResNet-50 and real images to test this properly. It runs twice, changing only DataLoader settings. On my included run, it was 1.83× faster and the diagnosis changed from input-bound to compute-bound. That number will vary with the CPU/GPU ratio. The useful part is finding out which side of the line your own run is on before optimizing the wrong thing. * **Colab:** traceml/notebooks/huggingface_dataloading_bottleneck.ipynb at main · traceopt-ai/traceml · GitHub * **Repo:** GitHub - traceopt-ai/traceml: Open-source observability for PyTorch training runs. · GitHub I would really value feedback from people running real Trainer workloads, especially cases where this diagnosis is surprising or wrong.
discuss.huggingface.co
July 27, 2026 at 10:49 AM
December 5, 2025 at 3:53 PM