#Forum3
Great #ADPD2025 #Forum3 on improving the therapeutic potential of #immunotherapies
April 2, 2025 at 4:34 PM
The NASA announcers on the broadcast referenced it as if it were a real feature of one of the current elevators. Maybe they misspoke. This is where I got the image: www.collectspace.com/ubb/Forum3/H...
Next stop, Space: SpaceX's Pad 39A elevator - collectSPACE: Messages
Source for space history, space artifacts, and space memorabilia. Learn where astronauts will appear, browse collecting guides, and read original space history-related daily reports.
www.collectspace.com
April 1, 2026 at 8:42 PM
The NASA announcers mentioned it on the broadcast and I Googled it. Here is the sourcing form 2023: www.collectspace.com/ubb/Forum3/H...
Next stop, Space: SpaceX's Pad 39A elevator - collectSPACE: Messages
Source for space history, space artifacts, and space memorabilia. Learn where astronauts will appear, browse collecting guides, and read original space history-related daily reports.
www.collectspace.com
April 1, 2026 at 8:29 PM
Happy New Year! My first story of 2026 is now up 👉 Why 2026 will prove whether #AI checkout is here to stay. Featuring Forum3, Bain, Brooklinen, Everlane and Winx Health. www.modernretail.co/technology/2...
.
2026 will prove whether AI checkout is here to stay
2026 will be the year that proves whether people have gotten comfortable enough with AI shopping to click "buy" within ChatGPT or Perplexity
www.modernretail.co
January 5, 2026 at 5:04 PM
Building an AI-first company: What these two business leaders learned from top experts

Adam Brotman, left, and Andy Sack, authors of the book, “AI First.” (Photo Courtesy Forum3) This week on the GeekWire Podcast, our guests are Adam Brotman and Andy Sack, co-authors of AI First: The Playbook for…
Building an AI-first company: What these two business leaders learned from top experts
Adam Brotman, left, and Andy Sack, authors of the book, “AI First.” (Photo Courtesy Forum3) This week on the GeekWire Podcast, our guests are Adam Brotman and Andy Sack, co-authors of AI First: The Playbook for a Future-Proof Business and Brand.  Brotman was Starbucks’ chief digital officer and later co-CEO of J.Crew. Sack is a founder, investor, and longtime advisor to tech leaders. Together, they run Forum3, a Seattle-based company that helps brands with customer loyalty and engagement. For their book, they interviewed experts including Bill Gates, Sam Altman, Reid Hoffman and Ethan Mollick, and spent time with companies and leaders that have seen early AI success.
nextbusiness24.com
August 16, 2025 at 5:00 PM
Don’t count the apple in ‘Ai Race’: It could be in the best position of all

Adam Brotman and Andy Sack have cooled on forum3. As a generation AI reshapes technician, most attention went to Openai, GoogleAnthropic, MicrosoftXAI and Chinese Deepseek. Their models are powerful, their demonstrations…
Don’t count the apple in ‘Ai Race’: It could be in the best position of all
Adam Brotman and Andy Sack have cooled on forum3. As a generation AI reshapes technician, most attention went to Openai, GoogleAnthropic, MicrosoftXAI and Chinese Deepseek. Their models are powerful, their demonstrations shine and their progress in fast. In contrast, Apple is a secrecy and methodical - he pulled skepticism because of the observed and inertia. In spite of…
todaysportnews.com
July 1, 2025 at 2:24 PM
Neue Version von Photoline ist raus: https://www.pl32.com/forum3/viewtopic.php?f=5&t=5700
Neu in Version 20.50 - PhotoLine Forum
www.pl32.com
December 27, 2024 at 1:22 AM
Rechtzeitig vor Weihnachten gibt es eine neue Version 22.50 von #Photoline - https://www.pl32.com/forum3/viewtopic.php?f=5&t=6485
Neue Version 22.50 - PhotoLine Forum
www.pl32.com
December 25, 2024 at 3:24 AM
Alles Jahre wieder... wird es mal wieder Zeit für ein #Photoline Upgrade (https://www.pl32.com/forum3/viewtopic.php?f=5&p=48565#p48565) // @PhotoLineSoft
Neue Version 22.51 - PhotoLine Forum
www.pl32.com
December 25, 2024 at 1:41 AM
How to train T5 to distinguish task-relevant tokens from contextual noise?
<p>Theme-wise, <a href="https://huggingface.co/hugging-science">Hugging Science</a> might have information.</p> <p>With vanilla T5, that might not be easy… (<a href="https://huggingface.co/datasets/John6666/forum3/blob/main/t5_bio_seq_1.md">Detailed version</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-why-structural-marking-is-hard-in-t5-without-tokens-1" name="p-250669-why-structural-marking-is-hard-in-t5-without-tokens-1"></a>Why “structural marking” is hard in T5 without tokens</h2> <p>T5 does <strong>not</strong> have a built-in segment channel like BERT’s token type embeddings. In the Transformers implementation, T5 explicitly “does not make use of token type ids” (it returns zeros), so you can’t rely on segment IDs to tell the model “this is core vs context.” (<a href="https://huggingface.co/docs/transformers/v4.13.0/en/model_doc/t5" title="T5">Hugging Face</a>)<br /> That means: if boundaries are unknown at inference and you don’t want input control tokens, the model has to <strong>infer relevance</strong> from the sequence itself.</p> <p>So the best “structure-based” solution is not a hidden embedding trick, but an <strong>architectural + objective</strong> structure: separate “find what matters” from “label what matters.”</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-why-padding-the-decoder-with-pre-cpost-c-labels-is-usually-a-poor-default-2" name="p-250669-why-padding-the-decoder-with-pre-cpost-c-labels-is-usually-a-poor-default-2"></a>Why padding the decoder with pre-c/post-c labels is usually a poor default</h2> <p>Your proposal makes the decoder target length equal to the full input length (core + pre + post). In T5, training is done with teacher forcing and requires a full target sequence for the decoder. (<a href="https://huggingface.co/docs/transformers/v4.48.2/model_doc/t5" title="T5 - Transformers">Hugging Face</a>)</p> <p>This tends to be suboptimal for two reasons:</p> <ol> <li> <p><strong>Compute cost scales with output length</strong><br /> Longer outputs mean more decoder self-attention and cross-attention work—paid for flank tokens you don’t care about.</p> </li> <li> <p><strong>Learning is dominated by “easy placeholders”</strong><br /> If flanks are long, the model can minimize loss by learning “always output pre-c/post-c” well, while underfitting the core labeling.</p> </li> </ol> <p>If you want a per-input-token labeler, it’s usually better to do that <strong>on the encoder side</strong> (token classification) and avoid long autoregressive decoding entirely.</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-recommended-approach-encoder-side-selector-translator-no-boundary-tokens-required-3" name="p-250669-recommended-approach-encoder-side-selector-translator-no-boundary-tokens-required-3"></a>Recommended approach: Encoder-side “Selector + Translator” (no boundary tokens required)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-high-level-picture-4" name="p-250669-high-level-picture-4"></a>High-level picture</h3> <ol> <li> <p>Run the T5 <strong>encoder</strong> over the full sequence (pre + core + post).</p> </li> <li> <p>Add two heads on top of encoder hidden states:</p> <ul> <li><strong>Selector head</strong>: predicts which positions are task-relevant (span or mask).</li> <li><strong>Translator head</strong>: predicts the label for each position.</li> </ul> </li> <li> <p>At inference:</p> <ul> <li>Use selector output to decide which positions are “core.”</li> <li>Emit translator labels <strong>only</strong> for those positions.</li> </ul> </li> </ol> <p>This gives you a clean separation:</p> <ul> <li>“Use the flanks as context” (encoder attention can use them)</li> <li>“Only emit labels for core” (selector decides what gets emitted)</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-part-1-translator-head-token-classification-with-ignored-flank-targets-5" name="p-250669-part-1-translator-head-token-classification-with-ignored-flank-targets-5"></a>Part 1 — Translator head: token classification with ignored flank targets</h2> <p>T5 can be used with a token classification head (linear layer on hidden states). (<a href="https://huggingface.co/docs/transformers/en/model_doc/t5" title="T5">Hugging Face</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-how-to-train-translator-without-forcing-outputs-for-flanks-6" name="p-250669-how-to-train-translator-without-forcing-outputs-for-flanks-6"></a>How to train translator without forcing outputs for flanks</h3> <p>Create <code>token_labels</code> aligned to input length, but set flank positions to an <strong>ignore index</strong> (commonly <code>-100</code>). Then compute cross entropy with <code>ignore_index=-100</code>, so ignored positions contribute <strong>no gradient</strong>. (<a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.functional.cross_entropy.html" title="torch.nn.functional.cross_entropy">docs.pytorch.org</a>)</p> <p>If you use Hugging Face batching, <code>DataCollatorForTokenClassification</code> defaults <code>label_pad_token_id=-100</code> and notes that <code>-100</code> is automatically ignored by PyTorch loss functions. (<a href="https://huggingface.co/docs/transformers/main/main_classes/data_collator" title="Data Collator">Hugging Face</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-translator-loss-7" name="p-250669-translator-loss-7"></a>Translator loss</h3> <p>Let <code>y_i</code> be the gold label at position <code>i</code> (or <code>-100</code> for ignored), and <code>p_i</code> the predicted distribution.<br /> Then:</p> <div class="math"> L_{trans}=\sum_{i=1}^{L} CE(p_i, y_i) </div> <p>(Positions with <code>y_i=-100</code> are ignored by the loss implementation.) (<a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.functional.cross_entropy.html" title="torch.nn.functional.cross_entropy">docs.pytorch.org</a>)</p> <p><strong>What this solves:</strong> the translator learns to label only the core positions during training.</p> <p><strong>What it does not solve:</strong> you still need the model to decide which positions are core at inference → that’s the selector.</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-part-2-selector-head-options-choose-based-on-core-structure-8" name="p-250669-part-2-selector-head-options-choose-based-on-core-structure-8"></a>Part 2 — Selector head options (choose based on core structure)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-option-a-default-if-core-is-contiguous-startend-span-head-qa-style-9" name="p-250669-option-a-default-if-core-is-contiguous-startend-span-head-qa-style-9"></a>Option A (default if core is contiguous): start/end span head (QA-style)</h3> <p>Predict two logits over positions: start and end.</p> <p><strong>Training</strong></p> <div class="math"> L_{sel}=CE(s, start)+CE(e, end) </div> <p><strong>Decoding</strong><br /> Use QA-style constrained pairing:</p> <ul> <li>only consider pairs where <code>end &gt;= start</code></li> <li>optionally enforce <code>end-start+1 &lt;= max_core_len</code></li> <li>choose best scoring pair among top-k candidates</li> </ul> <p>This postprocessing pattern is widely used in HF QA utilities. (<a href="https://github.com/huggingface/transformers/blob/main/examples/pytorch/question-answering/utils_qa.py" title="utils_qa.py">GitHub</a>)</p> <p><strong>Why this is stable:</strong> it avoids the extreme class imbalance problems you get when predicting an IN/OUT decision at every token.</p> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-option-b-if-core-can-be-fragmented-mask-or-biobioes-tags-10" name="p-250669-option-b-if-core-can-be-fragmented-mask-or-biobioes-tags-10"></a>Option B (if core can be fragmented): mask or BIO/BIOES tags</h3> <p>Predict per-token relevance as:</p> <ul> <li>IN/OUT mask, or</li> <li>BIO/BIOES tags (better for enforcing contiguity/structure)</li> </ul> <p>If you use a mask, imbalance is real: OUT usually dominates. Common mitigations:</p> <ul> <li>weighted loss / positive class weighting</li> <li>focal loss to downweight easy negatives (<a href="https://huggingface.co/docs/transformers/en/model_doc/t5" title="T5">Hugging Face</a>)</li> </ul> <p>If you use BIO/BIOES, add constrained decoding (or a CRF) to reduce “flicker” (alternating IN/OUT).</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-joint-training-keep-it-stable-avoid-selector-corrupting-translator-early-11" name="p-250669-joint-training-keep-it-stable-avoid-selector-corrupting-translator-early-11"></a>Joint training: keep it stable (avoid selector corrupting translator early)</h2> <p>A robust schedule is:</p> <ol> <li> <p><strong>Warm-up (translator-first)</strong></p> <ul> <li>Train translator using gold core labels with flank targets ignored (<code>-100</code>).</li> <li>Train selector from gold spans/masks.</li> <li>Do not make translator depend on selector predictions yet.</li> </ul> </li> <li> <p><strong>Couple gradually</strong></p> <ul> <li>Start using selector predictions to simulate inference conditions (for example, evaluate translator on predicted spans, or add small noise around gold boundaries).</li> <li>This reduces train–inference mismatch without letting an immature selector destroy translator learning.</li> </ul> </li> </ol> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-combined-objective-12" name="p-250669-combined-objective-12"></a>Combined objective</h3> <div class="math"> L = L_{trans} + \lambda L_{sel} </div> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-if-you-dont-have-reliable-boundary-labels-consider-ctc-blank-emit-nothing-13" name="p-250669-if-you-dont-have-reliable-boundary-labels-consider-ctc-blank-emit-nothing-13"></a>If you don’t have reliable boundary labels: consider CTC (“blank = emit nothing”)</h2> <p>If you want the model to learn “emit labels only where appropriate” without explicit span supervision, CTC is an alternative:</p> <ul> <li>Model emits per-position distribution over <code>{labels ∪ blank}</code></li> <li>Blank means “emit nothing,” and is removed during decoding (conceptually) (<a href="https://stackoverflow.com/questions/41737992/connectionist-temporal-classification-ctc-blank-label" title="Connectionist Temporal Classification (CTC) blank label">Stack Overflow</a>)</li> <li>CTCLoss sums over alignments and requires target length ≤ input length (<a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.CTCLoss.html" title="CTCLoss — PyTorch 2.10 documentation">docs.pytorch.org</a>)</li> </ul> <p>CTC is attractive when boundaries are ambiguous/noisy, but it does not naturally enforce “one contiguous core” unless you add extra constraints.</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-long-sequences-avoid-truncation-as-a-hidden-failure-mode-14" name="p-250669-long-sequences-avoid-truncation-as-a-hidden-failure-mode-14"></a>Long sequences: avoid truncation as a hidden failure mode</h2> <p>If <code>L_total</code> can be very large, you either need:</p> <ul> <li>a long-context encoder, or</li> <li>windowing with overlap + reconciliation.</li> </ul> <p>LongT5 is designed for long inputs and the Transformers docs state it can handle up to 16,384 input tokens. (<a href="https://huggingface.co/docs/transformers/en/model_doc/longt5" title="LongT5">Hugging Face</a>)<br /> If you window, QA-style overlap logic and consistency checks across windows are a practical pattern (same start/end-logit postprocessing idea). (<a href="https://github.com/huggingface/transformers/blob/main/examples/pytorch/question-answering/utils_qa.py" title="utils_qa.py">GitHub</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-practical-recommendation-what-to-do-first-15" name="p-250669-practical-recommendation-what-to-do-first-15"></a>Practical recommendation (what to do first)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-if-the-valid-sequence-is-usually-one-contiguous-block-16" name="p-250669-if-the-valid-sequence-is-usually-one-contiguous-block-16"></a>If the “valid sequence” is usually one contiguous block</h3> <ol> <li> <p>Use T5 encoder states.</p> </li> <li> <p>Add:</p> <ul> <li>start/end selector head (QA-style)</li> <li>per-token translator head trained with ignored flank targets (<code>-100</code>)</li> </ul> </li> <li> <p>Decode span → emit translator labels only inside span.</p> </li> </ol> <p>This is typically the best combination of <strong>efficiency</strong> (no long decoder outputs) and <strong>stability</strong> (selector avoids OUT-everywhere collapse).</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250669-if-validity-is-fragmented-or-fuzzy-17" name="p-250669-if-validity-is-fragmented-or-fuzzy-17"></a>If validity is fragmented or fuzzy</h3> <ol> <li>Replace span selector with BIO/BIOES (optionally CRF) or an IN/OUT mask with imbalance controls.</li> <li>Still train translator with flank targets ignored (<code>-100</code>).</li> <li>Postprocess mask/tags into whichever output format you need.</li> </ol> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250669-direct-answer-to-your-pre-cpost-c-label-padding-idea-18" name="p-250669-direct-answer-to-your-pre-cpost-c-label-padding-idea-18"></a>Direct answer to your pre-c/post-c label padding idea</h2> <ul> <li>It is usually <strong>computationally inefficient</strong> in T5 because it forces a long decoder target sequence under teacher forcing. (<a href="https://huggingface.co/docs/transformers/v4.48.2/model_doc/t5" title="T5 - Transformers">Hugging Face</a>)</li> <li>It is often <strong>learning-inefficient</strong> because the model can over-optimize placeholders when flanks dominate.</li> </ul> <p>If you want the model to learn which region matters <strong>without any boundary tokens at inference</strong>, the most reliable approach is to make “which tokens matter” an explicit prediction problem (span/mask) and train the labeler only on relevant tokens using ignore-index masking. (<a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.functional.cross_entropy.html" title="torch.nn.functional.cross_entropy">docs.pytorch.org</a>)</p>
discuss.huggingface.co
February 5, 2026 at 8:26 AM
Whisper for Arabic–English speech with Indian accent
<p>Whisper <a href="https://huggingface.co/datasets/John6666/forum3/blob/main/whisper_ar_en_1.md">might be stuck in the worst possible situation for this model</a>…?</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-why-this-setting-is-hard-for-vanilla-whisper-1" name="p-250597-why-this-setting-is-hard-for-vanilla-whisper-1"></a>Why this setting is hard for “vanilla” Whisper</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-code-switching-breaks-the-models-strongest-assumptions-2" name="p-250597-code-switching-breaks-the-models-strongest-assumptions-2"></a>Code-switching breaks the model’s strongest assumptions</h3> <p>Whisper-style models are trained to produce <strong>one coherent transcript</strong> from a window of audio. In code-switch speech, the model must decide (often multiple times per second) whether the next token should come from Arabic script or Latin script, while also handling shared phonetics and loanwords. When the evidence is weak (fast speech, noise, accent), the decoder tends to “commit” to one language and then <strong>keep sampling from that language’s token distribution</strong>, which can spill across the true switch boundary.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-indian-accented-speech-increases-phonetic-ambiguity-3" name="p-250597-indian-accented-speech-increases-phonetic-ambiguity-3"></a>Indian-accented speech increases phonetic ambiguity</h3> <p>Accents affect:</p> <ul> <li>vowel/consonant realizations,</li> <li>stress timing,</li> <li>coarticulation patterns,</li> <li>and segment durations.</li> </ul> <p>For short, noisy messages, these shifts are enough to push the model into “low-evidence” decoding where it starts relying more on its language model prior than the acoustic signal. That’s when you see substitutions, omissions, or fluent but wrong text.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-auto-language-detection-is-fragile-on-shortnoisy-audio-4" name="p-250597-auto-language-detection-is-fragile-on-shortnoisy-audio-4"></a>Auto language detection is fragile on short/noisy audio</h3> <p>In <code>faster-whisper</code>, language detection is performed using the <strong>first ~30 seconds</strong> if you don’t set <code>language=...</code> explicitly. That is a known source of wrong-language outputs if the beginning includes silence/noise or code-switching. (<a href="https://github.com/SYSTRAN/faster-whisper/blob/master/faster_whisper/transcribe.py" title="faster-whisper/faster_whisper/transcribe.py at master">GitHub</a>)<br /> This interacts badly with your setting: once the model “picks” the wrong language early, later chunks are biased.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-repetitionhallucination-loops-are-a-known-failure-mode-with-silencegaps-5" name="p-250597-repetitionhallucination-loops-are-a-known-failure-mode-with-silencegaps-5"></a>Repetition/hallucination loops are a known failure mode with silence/gaps</h3> <p>Two community-validated mitigations for “stuck repeating / hallucinating after a gap” are:</p> <ul> <li>split audio with VAD,</li> <li>set <code>condition_on_previous_text=False</code>. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)<br /> This matters for voice messages because they often have pauses and trailing non-speech.</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-practical-pipeline-changes-that-usually-move-the-needle-6" name="p-250597-practical-pipeline-changes-that-usually-move-the-needle-6"></a>Practical pipeline changes that usually move the needle</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-1-make-segmentation-your-primary-quality-lever-7" name="p-250597-h-1-make-segmentation-your-primary-quality-lever-7"></a>1) Make segmentation your primary quality lever</h3> <p>For short messages, segmentation quality often dominates model size.</p> <p><strong>Target behavior</strong>: feed the decoder 1–8s windows that are “mostly speech”, with small padding and minimal trailing non-speech.</p> <p><strong>Recommended segmentation recipe</strong></p> <ul> <li>VAD to find speech islands</li> <li>add padding (e.g., 150–300ms)</li> <li>add overlap (e.g., 100–250ms) to protect word boundaries</li> <li><strong>explicit tail trimming</strong> after VAD (energy/RMS-based) to remove long quiet endings that trigger hallucinations</li> <li>cap maximum segment length (e.g., 8–12s); long segments increase drift and LID errors</li> </ul> <p>This aligns with common “hallucination fix” guidance: VAD slicing plus disabling conditioning reduces loops when the model can’t find evidence in the current window. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-2-default-to-guardrails-on-for-production-decoding-8" name="p-250597-h-2-default-to-guardrails-on-for-production-decoding-8"></a>2) Default to “guardrails ON” for production decoding</h3> <p>A simple but important rule: don’t disable the thresholds unless you’re deliberately building a repro.</p> <p>When thresholds are enabled, Whisper-style decoding has mechanisms to suppress output during no-speech/low-confidence regions (implemented in the reference transcribe logic). (<a href="https://github.com/openai/whisper/discussions/29" title="Stops working after long gap with no speech? · openai whisper · Discussion #29 · GitHub">GitHub</a>)</p> <p><strong>Practical defaults for voice messages</strong></p> <ul> <li><code>condition_on_previous_text=False</code> (prevents “carryover text” into gaps) (<a href="https://github.com/openai/whisper/discussions/29" title="Stops working after long gap with no speech? · openai whisper · Discussion #29 · GitHub">GitHub</a>)</li> <li>keep <code>no_speech_threshold</code>, <code>log_prob_threshold</code>, <code>compression_ratio_threshold</code> enabled (don’t set them to <code>None</code>)</li> <li>use <code>temperature=0.0</code> for determinism while tuning; add temperature fallback only if needed</li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-3-constrain-language-behavior-by-design-dont-rely-on-auto-lid-9" name="p-250597-h-3-constrain-language-behavior-by-design-dont-rely-on-auto-lid-9"></a>3) Constrain language behavior by design (don’t rely on auto-LID)</h3> <p>Auto-LID being computed on the first ~30s is a known limitation; multiple issues report wrong-language outputs under auto-detection. (<a href="https://github.com/SYSTRAN/faster-whisper/issues/265" title="Improve Language detection #265 - SYSTRAN/faster- ...">GitHub</a>)<br /> There’s also an open request for “limit detection to a subset of languages,” which does not exist as a first-class feature in <code>faster-whisper</code> today. (<a href="https://github.com/SYSTRAN/faster-whisper/issues/1164" title="[Feature Request] Constrain Available Languages when ...">GitHub</a>)</p> <p><strong>Workarounds that actually help</strong></p> <ul> <li> <p>If you know it’s always Arabic+English, use a <strong>two-pass strategy</strong>:</p> <ol> <li> <p>attempt <code>language="ar"</code> decode</p> </li> <li> <p>attempt <code>language="en"</code> decode</p> </li> <li> <p>pick the better result using a small heuristic:</p> <ul> <li>script sanity (Arabic chars ratio vs Latin ratio),</li> <li>repetition score,</li> <li>average logprob proxy (if available),</li> <li>“text produced in low-energy region” penalty.</li> </ul> </li> </ol> </li> </ul> <p>This directly addresses “unstable language ID outputs” without needing new model features.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-4-post-processing-that-is-code-switch-aware-10" name="p-250597-h-4-post-processing-that-is-code-switch-aware-10"></a>4) Post-processing that is code-switch aware</h3> <p>Avoid “English-only cleanup” or “Arabic-only cleanup”; mixed script requires a mixed strategy.</p> <p><strong>Low-risk post-processing ideas</strong></p> <ul> <li> <p><strong>script-aware normalization</strong></p> <ul> <li>normalize Arabic punctuation variants (e.g., Arabic comma/Latin comma)</li> <li>normalize tatweel and repeated diacritics only if you see them</li> </ul> </li> <li> <p><strong>repetition filters</strong></p> <ul> <li>detect repeated bigrams/trigrams over a threshold and either truncate or mark as suspect</li> </ul> </li> <li> <p><strong>segment-level confidence flags</strong></p> <ul> <li> <p>mark segments suspicious if:</p> <ul> <li>very long text produced while energy is low,</li> <li>script doesn’t match forced language pass,</li> <li>high repetition compression-like behavior</li> </ul> </li> </ul> </li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-5-if-whisper-still-struggles-consider-an-alternate-base-model-as-a-reference-11" name="p-250597-h-5-if-whisper-still-struggles-consider-an-alternate-base-model-as-a-reference-11"></a>5) If Whisper still struggles: consider an alternate base model as a reference</h3> <p>Two candidates worth testing as “sanity checks”:</p> <ul> <li> <p>Meta SeamlessM4T v2: supports Arabic variants (e.g., Modern Standard Arabic, Egyptian, Moroccan) in its published supported language list, and is explicitly evaluated for ASR tasks. (<a href="https://huggingface.co/facebook/seamless-m4t-v2-large" title="facebook/seamless-m4t-v2-large">Hugging Face</a>)<br /> <em>Use case</em>: as a comparison point or fallback for Arabic-heavy segments (not necessarily best at code-switching out of the box).</p> </li> <li> <p>NVIDIA Canary v2: strong multilingual ASR for its supported languages, but public materials emphasize European language coverage; Arabic support is inconsistent across deployments per community reports. (<a href="https://huggingface.co/nvidia/canary-1b-v2" title="nvidia/canary-1b-v2">Hugging Face</a>)<br /> <em>Use case</em>: less compelling if Arabic is core.</p> </li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-existing-fine-tuned-models-you-can-start-from-12" name="p-250597-existing-fine-tuned-models-you-can-start-from-12"></a>Existing fine-tuned models you can start from</h2> <p>These are not a perfect match (Arabic↔English + Indian accent), but they’re useful starting points.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-arabicenglish-code-switch-whisper-models-13" name="p-250597-arabicenglish-code-switch-whisper-models-13"></a>Arabic–English code-switch Whisper models</h3> <ul> <li> <p><code>MohamedRashad/Arabic-Whisper-CodeSwitching-Edition</code><br /> Fine-tuned on an Arabic-English code-switch dataset; explicitly intended for Arabic speech with embedded English words. License shown as GPL-3.0 (often problematic for commercial use). (<a href="https://huggingface.co/MohamedRashad/Arabic-Whisper-CodeSwitching-Edition" title="MohamedRashad/Arabic-Whisper-CodeSwitching-Edition · Hugging Face">Hugging Face</a>)</p> </li> <li> <p><code>azeem23/whisper-small-codeswitching-ArabicEnglish</code><br /> A smaller Whisper variant fine-tuned for Arabic-English code-switching, based on the same dataset. (<a href="https://huggingface.co/azeem23/whisper-small-codeswitching-ArabicEnglish" title="azeem23/whisper-small-codeswitching-ArabicEnglish">Hugging Face</a>)</p> </li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-indian-accent-english-whisper-model-14" name="p-250597-indian-accent-english-whisper-model-14"></a>Indian-accent English Whisper model</h3> <ul> <li><code>Tejveer12/Indian-Accent-English-Whisper-Finetuned</code><br /> Fine-tuned on the Indian-accent English dataset (<code>WillHeld/india_accent_cv</code>). (<a href="https://huggingface.co/Tejveer12/Indian-Accent-English-Whisper-Finetuned" title="Tejveer12/Indian-Accent-English-Whisper-Finetuned · Hugging Face">Hugging Face</a>)<br /> The model repository indicates an MIT license in its metadata/commit history. (<a href="https://huggingface.co/Tejveer12/Indian-Accent-English-Whisper-Finetuned/commit/26b6d18da8db69db8290513077c1b57727b43181" title="Training in progress, step 7000 · Tejveer12/Indian-Accent- ...">Hugging Face</a>)</li> </ul> <p><strong>How to use these in practice</strong></p> <ul> <li>Use the Indian-accent model as an <strong>English-pass decoder</strong> for English-dominant segments.</li> <li>Use a code-switch model as the <strong>Arabic-pass decoder</strong> (especially for Arabic-dominant segments with English insertions).</li> <li>Or: use these as initialization targets for your own adapter fine-tune (next section).</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-fine-tuning-whisper-for-your-exact-data-practical-recipe-15" name="p-250597-fine-tuning-whisper-for-your-exact-data-practical-recipe-15"></a>Fine-tuning Whisper for your exact data (practical recipe)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-1-use-adapter-style-fine-tuning-lora-first-16" name="p-250597-h-1-use-adapter-style-fine-tuning-lora-first-16"></a>1) Use adapter-style fine-tuning (LoRA) first</h3> <p>Full fine-tuning of large Whisper checkpoints is expensive and easy to overfit. For accent + code-switch adaptation, LoRA usually gets you most of the gain with lower risk.</p> <p>The Hugging Face PEFT guide shows an int8 + LoRA training approach for Whisper ASR specifically. (<a href="https://huggingface.co/docs/peft/v0.6.0/en/task_guides/int8-asr" title="int8 training for automatic speech recognition">Hugging Face</a>)</p> <p><strong>Why LoRA helps here</strong></p> <ul> <li>You’re adapting pronunciation + boundary behavior, not learning a new language.</li> <li>You want to preserve general robustness while nudging the model toward your accent and code-switch distribution.</li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-2-build-a-training-mix-that-matches-your-deployment-distribution-17" name="p-250597-h-2-build-a-training-mix-that-matches-your-deployment-distribution-17"></a>2) Build a training mix that matches your deployment distribution</h3> <p>Aim for three buckets:</p> <ol> <li><strong>In-domain</strong>: your actual voice messages (even 10–50 hours helps if transcripts are consistent)</li> <li><strong>Indian-accent English</strong>: augment English segments with accent data (e.g., <code>WillHeld/india_accent_cv</code>) (<a href="https://huggingface.co/datasets/WillHeld/india_accent_cv" title="WillHeld/india_accent_cv · Datasets at Hugging Face">Hugging Face</a>)</li> <li><strong>Arabic–English code-switch</strong>: add code-switch examples (e.g., MohamedRashad dataset/models; also consider Mixat for methodology even if dialect differs) (<a href="https://huggingface.co/MohamedRashad/Arabic-Whisper-CodeSwitching-Edition" title="MohamedRashad/Arabic-Whisper-CodeSwitching-Edition · Hugging Face">Hugging Face</a>)</li> </ol> <p>If you lack real Arabic↔English code-switch hours, synthetic code-switch generation is an active research direction (phrase-level mixing) and can be used to bootstrap. (<a href="https://www.isca-archive.org/interspeech_2025/nguyen25_interspeech.pdf" title="Can we train ASR systems on Code-switch without real ...">isca-archive.org</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-3-keep-transcript-conventions-strict-and-stable-18" name="p-250597-h-3-keep-transcript-conventions-strict-and-stable-18"></a>3) Keep transcript conventions strict and stable</h3> <p>For code-switch, consistency matters more than perfection:</p> <ul> <li>keep Arabic in Arabic script and English in Latin script</li> <li>avoid random transliterations</li> <li>normalize punctuation and casing rules consistently across the dataset</li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-4-training-choices-that-matter-most-for-your-case-19" name="p-250597-h-4-training-choices-that-matter-most-for-your-case-19"></a>4) Training choices that matter most for your case</h3> <ul> <li> <p>Start from a multilingual checkpoint (e.g., Whisper small/medium/large-v3 depending on budget)</p> </li> <li> <p>Use <code>task="transcribe"</code> (not translate)</p> </li> <li> <p>Ensure audio is standardized to 16kHz mono</p> </li> <li> <p>Filter or downweight:</p> <ul> <li>clips with extremely low SNR,</li> <li>clips with unreliable transcripts,</li> <li>clips with long non-speech tails (or trim them)</li> </ul> </li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-5-evaluation-dont-rely-on-one-wer-number-20" name="p-250597-h-5-evaluation-dont-rely-on-one-wer-number-20"></a>5) Evaluation: don’t rely on one WER number</h3> <p>Use at least:</p> <ul> <li>overall WER</li> <li><strong>English-only WER on English spans</strong></li> <li><strong>Arabic-only WER/CER on Arabic spans</strong></li> <li>a “switch-boundary” check (simple proxy): count how often the script flips in the right neighborhood of known switch points (even a heuristic boundary test catches regressions quickly)</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-high-quality-references-to-follow-end-to-end-21" name="p-250597-high-quality-references-to-follow-end-to-end-21"></a>High-quality references to follow end-to-end</h2> <ul> <li>Hugging Face blog: “Fine-Tune Whisper For Multilingual ASR with Transformers” (step-by-step). (<a href="https://huggingface.co/blog/fine-tune-whisper" title="Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers">Hugging Face</a>)</li> <li>PEFT int8 + LoRA ASR guide for Whisper (T4-friendly training approach). (<a href="https://huggingface.co/docs/peft/v0.6.0/en/task_guides/int8-asr" title="int8 training for automatic speech recognition">Hugging Face</a>)</li> <li>Whisper hallucination mitigation discussion: VAD slicing + <code>condition_on_previous_text=False</code>. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</li> <li>Code-switch dataset methodology reference: Mixat paper (how they build and analyze code-mixed Arabic/English speech). (<a href="https://aclanthology.org/2024.sigul-1.26/" title="Mixat: A Data Set of Bilingual Emirati-English Speech">ACL Anthology</a>)</li> <li><code>faster-whisper</code> language detection limitations and wrong-language reports. (<a href="https://github.com/SYSTRAN/faster-whisper/issues/265" title="Improve Language detection #265 - SYSTRAN/faster- ...">GitHub</a>)</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-a-concrete-starting-plan-for-your-production-pipeline-22" name="p-250597-a-concrete-starting-plan-for-your-production-pipeline-22"></a>A concrete “starting plan” for your production pipeline</h2> <ol> <li> <p><strong>Segment aggressively</strong> (VAD + pad + overlap + explicit tail trim) before decoding. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</p> </li> <li> <p><strong>Decode with guardrails on</strong> and <code>condition_on_previous_text=False</code> by default for voice messages. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</p> </li> <li> <p><strong>Two-pass language strategy</strong> per segment:</p> <ul> <li>run forced Arabic decode, forced English decode</li> <li>choose output by script sanity + repetition penalty (+ score proxy if available)</li> </ul> </li> <li> <p><strong>Fallback policy</strong>: if output is suspicious (wrong script, repetition, text in low energy), re-decode with stricter thresholds and/or shorter segment.</p> </li> <li> <p><strong>Fine-tune via LoRA</strong> using your in-domain audio + Indian-accent English + Arabic-English code-switch data. (<a href="https://huggingface.co/docs/peft/v0.6.0/en/task_guides/int8-asr" title="int8 training for automatic speech recognition">Hugging Face</a>)</p> </li> </ol>
discuss.huggingface.co
February 3, 2026 at 12:22 PM
Structured JSON (numerical data) into fixed-length vectors for MLP probing
<p>Hmm… <a href="https://huggingface.co/datasets/John6666/forum3/blob/main/json_for_mlp_1.md">Trying to do that with embedding seems to make it harder</a>…?</p> <hr /> <p>Below is what I would do <strong>for your exact goal</strong>: “JSON with a variable-length list of entities → fixed-length representation → an MLP probe recovers per-entity <code>(azimuth_deg, elevation_deg)</code> with high fidelity.”</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-core-recommendation-for-your-case-1" name="p-250406-core-recommendation-for-your-case-1"></a>Core recommendation for your case</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-treat-this-as-a-set-to-slots-problem-not-text-embedding-2" name="p-250406-treat-this-as-a-set-to-slots-problem-not-text-embedding-2"></a>Treat this as a <strong>set-to-slots</strong> problem, not “text embedding”</h3> <ul> <li>Your input is a <strong>set of structured records</strong> (category + two angles), not natural language.</li> <li>Your output needs <strong>per-entity recoverability</strong>, which is fundamentally different from “semantic similarity” embeddings.</li> </ul> <p>So the best default is:</p> <ol> <li><strong>Parse JSON</strong> into typed records (<code>category</code>, <code>az</code>, <code>el</code>)</li> <li>Build <strong>per-entity numeric-first tokens</strong> (angles stay numeric)</li> <li>Use a <strong>set/slot encoder</strong> that outputs a fixed number of <strong>slots <code>K</code></strong></li> <li>Train with <strong>Hungarian matching</strong> (DETR-style) so “slot ↔ entity” alignment is stable</li> <li>Freeze the encoder; train your <strong>MLP probe per slot</strong></li> </ol> <p>This training pattern (“fixed number of predictions vs variable ground truth, matched by bipartite assignment”) is standard in set prediction and is one of the main ideas behind <strong>DETR</strong>. (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-why-this-fits-your-constraints-3" name="p-250406-why-this-fits-your-constraints-3"></a>Why this fits your constraints</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-h-1-arbitrary-input-length-4" name="p-250406-h-1-arbitrary-input-length-4"></a>1) Arbitrary input length</h3> <p>Set/slot encoders consume <code>N</code> entity tokens and produce <code>K</code> slots (fixed shape), so <code>N</code> can vary.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-h-2-numerical-fidelity-5" name="p-250406-h-2-numerical-fidelity-5"></a>2) Numerical fidelity</h3> <p>You avoid “numbers-as-text” entirely (or keep text only for the category label), so floats are never degraded by tokenization. Tokenization pitfalls for floating point numbers are well documented—e.g., a float like <code>"3.14159"</code> can be split into multiple chunks by common LLM tokenizers, which is hostile to exact numeric recovery. (<a href="https://arxiv.org/pdf/2309.06236" title="Pitfalls of Representing and Tokenizing Temporal Data for ...">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-h-3-fixed-output-size-6" name="p-250406-h-3-fixed-output-size-6"></a>3) Fixed output size</h3> <p>You choose a capacity <code>K</code> and always output <code>K×d</code> (or flatten to <code>K·d</code>).</p> <p><strong>Important reality check:</strong> a finite fixed-size vector cannot losslessly encode an <em>unbounded</em> number of entities. In practice, you pick <code>K</code> large enough for your dataset and define an overflow policy (truncate, sample, or increase <code>K</code>). This is the same practical compromise used by set prediction models. (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-q1-encoders-llms-you-can-use-ranked-for-your-use-case-7" name="p-250406-q1-encoders-llms-you-can-use-ranked-for-your-use-case-7"></a>Q1) Encoders / “LLMs” you can use (ranked for your use case)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-a-best-first-choice-set-transformer-with-pmak-recommended-8" name="p-250406-a-best-first-choice-set-transformer-with-pmak-recommended-8"></a>A. Best first choice: <strong>Set Transformer with PMA(K)</strong> (recommended)</h3> <p><strong>What it is:</strong> A Transformer architecture designed for sets; it includes a pooling module (<strong>PMA</strong>) that uses <code>K</code> learnable seed vectors to produce exactly <code>K</code> outputs (your slots). (<a href="https://proceedings.mlr.press/v97/lee19d/lee19d.pdf" title="A Framework for Attention-based Permutation-Invariant Neural ...">Proceedings of Machine Learning Research</a>)</p> <p><strong>Why it’s a good fit</strong></p> <ul> <li>Permutation handling is built-in (set input, order-agnostic)</li> <li>Produces <strong>fixed <code>K</code> outputs</strong> cleanly via PMA</li> <li>Easy to combine with a DETR-style matching loss</li> </ul> <p><strong>Good implementation starting point</strong></p> <ul> <li>Official PyTorch implementation: (<a href="https://github.com/juho-lee/set_transformer" title="juho-lee/set_transformer: Pytorch implementation of set ...">GitHub</a>)</li> </ul> <p><strong>When it’s enough</strong></p> <ul> <li>If typical <code>N</code> is up to a few hundred, this is usually fine.</li> </ul> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-b-if-n-can-be-very-large-perceiver-io-9" name="p-250406-b-if-n-can-be-very-large-perceiver-io-9"></a>B. If <code>N</code> can be very large: <strong>Perceiver IO</strong></h3> <p><strong>What it is:</strong> A cross-attention architecture that scales linearly with input size and supports structured outputs via queries—i.e., you can decode <strong><code>K</code> slots</strong> efficiently even when <code>N</code> is big. (<a href="https://arxiv.org/abs/2107.14795" title="Perceiver IO: A General Architecture for Structured Inputs &amp; Outputs">arXiv</a>)</p> <p><strong>When to choose it</strong></p> <ul> <li>If <code>N</code> is frequently in the <strong>hundreds to thousands</strong>, Perceiver IO often scales better than full self-attention over inputs.</li> </ul> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-c-strong-baseline-deep-sets-10" name="p-250406-c-strong-baseline-deep-sets-10"></a>C. Strong baseline: <strong>Deep Sets</strong></h3> <p><strong>What it is:</strong> The canonical “set function” form: per-element encoder + permutation-invariant pooling. (<a href="https://arxiv.org/abs/1703.06114" title="Deep Sets">arXiv</a>)</p> <p><strong>Why it’s useful</strong></p> <ul> <li>It’s the simplest sanity-check: if you can’t get good probe decoding with this, your numeric encoding/training pipeline likely has issues.</li> </ul> <p><strong>Limitation</strong></p> <ul> <li>Vanilla pooling (sum/mean) tends to entangle entities; it’s better for “global properties of the set” than “recover each element”.</li> </ul> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-d-slot-binding-alternative-slot-attention-11" name="p-250406-d-slot-binding-alternative-slot-attention-11"></a>D. Slot-binding alternative: <strong>Slot Attention</strong></h3> <p>Not in the citations above, but conceptually: Slot Attention iteratively binds inputs to <code>K</code> slots. It can produce very clean per-slot “object-like” representations, but it often takes more tuning than Set Transformer.</p> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-e-numeric-embedding-modules-you-can-and-should-borrow-12" name="p-250406-e-numeric-embedding-modules-you-can-and-should-borrow-12"></a>E. Numeric embedding modules you can (and should) borrow</h3> <p>Regardless of the set encoder, you should embed numeric scalars into higher-dimensional vectors before mixing.</p> <p>A widely used reference is <strong>“On Embeddings for Numerical Features in Tabular Deep Learning”</strong> (periodic and piecewise-linear numeric embeddings) and its official implementation repo. (<a href="https://arxiv.org/abs/2203.05556" title="On Embeddings for Numerical Features in Tabular Deep Learning">arXiv</a>)</p> <p>This is directly relevant because your “angles must be recoverable” requirement is basically “my numeric scalars shouldn’t get washed out by the backbone.”</p> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-f-if-you-insist-on-an-llm-like-encoder-end-to-end-13" name="p-250406-f-if-you-insist-on-an-llm-like-encoder-end-to-end-13"></a>F. If you insist on an “LLM-like” encoder end-to-end</h3> <p>This is rarely the best path for your goal, but if you must:</p> <ol> <li><strong>Tokenization-free / byte/char models</strong></li> </ol> <ul> <li><strong>ByT5</strong> processes raw bytes (no subword vocabulary). (<a href="https://arxiv.org/abs/2105.13626" title="ByT5: Towards a token-free future with pre-trained byte-to-byte models">arXiv</a>)</li> <li><strong>CANINE</strong> processes characters without explicit tokenization/vocabulary. (<a href="https://arxiv.org/abs/2103.06874" title="CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation">arXiv</a>)</li> </ul> <p>These avoid subword fragmentation, but you still need <strong>task-specific training</strong> to make the representation numerically decodable.</p> <ol start="2"> <li><strong>Numeric tokenization research</strong> (if you’re training from scratch)</li> </ol> <ul> <li><strong>xVal</strong> proposes continuous numerical tokenization for numerically dense datasets. (<a href="https://arxiv.org/abs/2310.02989" title="xVal: A Continuous Numerical Tokenization for Scientific Language Models">arXiv</a>)</li> </ul> <p>This is relevant if you’re building a foundation-model-like pipeline, not if you just want a practical probe-friendly encoder quickly.</p> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-why-not-clip-for-this-14" name="p-250406-why-not-clip-for-this-14"></a>Why not CLIP for this</h3> <ul> <li>Standard CLIP uses a fixed text context length (commonly <strong>77 tokens</strong>), and this is a known limitation in practice (errors when exceeding it, and model constraints around <code>context_length</code>). (<a href="https://github.com/openai/CLIP/issues/212" title="RuntimeError: Input is too long for context length 77 #212">GitHub</a>)</li> <li>There is active work on extending CLIP beyond this limit (e.g., TULIP and other approaches), but that still doesn’t address the deeper issue: CLIP-style text embeddings are not trained to preserve precise numeric values. (<a href="https://proceedings.iclr.cc/paper_files/paper/2025/file/54f117f842d47c2007b501bddd04c0e7-Paper-Conference.pdf" title="TULIP: TOKEN-LENGTH UPGRADED CLIP">ICLR Proceedings</a>)</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-q2-should-you-convert-json-to-natural-language-15" name="p-250406-q2-should-you-convert-json-to-natural-language-15"></a>Q2) Should you convert JSON to natural language?</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-recommendation-no-not-for-mlp-decodes-floats-16" name="p-250406-recommendation-no-not-for-mlp-decodes-floats-16"></a>Recommendation: <strong>No</strong>, not for “MLP decodes floats”</h3> <p>Converting structured records into prose tends to:</p> <ul> <li>introduce formatting variability (“-125.5 degrees” vs “-125.50°”)</li> <li>reintroduce tokenizer issues for numbers</li> <li>encourage the encoder to focus on semantics rather than numeric exactness</li> </ul> <p>If your success metric is “recover precise angles,” <strong>keep angles numeric</strong> and only embed categories as text if you truly need open-vocabulary behavior.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-what-to-do-instead-structured-tokens-numeric-first-17" name="p-250406-what-to-do-instead-structured-tokens-numeric-first-17"></a>What to do instead: structured → tokens (numeric-first)</h3> <p>A practical representation per entity:</p> <ul> <li> <p><strong>Category</strong>: embedding lookup (closed vocab) or a small text embedding (open vocab)</p> </li> <li> <p><strong>Angles</strong>:</p> <ul> <li>encode azimuth/elevation as <code>sin</code>/<code>cos</code> (wrap-safe)</li> <li>optionally add harmonics for higher precision</li> </ul> </li> </ul> <p>Why harmonics can help: Fourier feature mappings are known to help MLPs learn higher-frequency structure (reducing “spectral bias”). (<a href="https://arxiv.org/abs/2006.10739" title="Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-must-serialize-to-text-only-for-compatibility-18" name="p-250406-if-you-must-serialize-to-text-only-for-compatibility-18"></a>If you must serialize to text (only for compatibility)</h3> <p>Use a <strong>two-stream</strong> approach:</p> <ul> <li>Text stream: stable placeholders<br /> <code>{"category":"chair","az":&lt;NUM0&gt;,"el":&lt;NUM1&gt;}</code></li> <li>Numeric stream: the real floats <code>[-125.5, -37.2, ...]</code></li> </ul> <p>Then fuse them (concat + MLP) before the set encoder. This prevents your numeric fidelity from depending on the text tokenizer.</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-q3-how-to-prevent-subword-tokenizers-from-breaking-floats-19" name="p-250406-q3-how-to-prevent-subword-tokenizers-from-breaking-floats-19"></a>Q3) How to prevent subword tokenizers from breaking floats</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-the-most-reliable-solution-dont-tokenize-floats-20" name="p-250406-the-most-reliable-solution-dont-tokenize-floats-20"></a>The most reliable solution: <strong>don’t tokenize floats</strong></h3> <p>Parse JSON → keep numbers as floats → embed numerically.</p> <p>This completely avoids the class of problems illustrated by common tokenizers splitting floating point strings into multiple pieces. (<a href="https://arxiv.org/pdf/2309.06236" title="Pitfalls of Representing and Tokenizing Temporal Data for ...">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-must-go-through-a-tokenizer-anyway-21" name="p-250406-if-you-must-go-through-a-tokenizer-anyway-21"></a>If you must go through a tokenizer anyway</h3> <p>Use one of these, in descending order of reliability:</p> <ol> <li><strong>Tokenization-free encoders</strong> (byte/char): ByT5, CANINE (<a href="https://arxiv.org/abs/2105.13626" title="ByT5: Towards a token-free future with pre-trained byte-to-byte models">arXiv</a>)</li> <li><strong>Placeholder substitution + numeric side-channel</strong> (best practical workaround with standard LLM tokenizers)</li> <li><strong>Quantize</strong> angles to fixed-point integers (e.g., tenths of a degree) so the model learns an easier discrete mapping</li> <li><strong>Train a custom tokenizer/vocab</strong> where numbers (or digit chunks) are handled consistently (only if you control the full training stack)</li> </ol> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-a-concrete-configuration-i-would-start-with-defaults-22" name="p-250406-a-concrete-configuration-i-would-start-with-defaults-22"></a>A concrete configuration I would start with (defaults)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-want-the-simplest-works-first-version-23" name="p-250406-if-you-want-the-simplest-works-first-version-23"></a>If you want the simplest “works-first” version</h3> <ul> <li> <p><code>K = 32</code>, <code>d = 256</code></p> </li> <li> <p>Per-entity token:</p> <ul> <li>category embedding (e.g., <code>d_cat=64</code>)</li> <li>angles: <code>[sin(az), cos(az), sin(el), cos(el)]</code></li> <li>MLP → <code>d</code></li> </ul> </li> <li> <p>Set encoder: Set Transformer + PMA(K) (<a href="https://proceedings.mlr.press/v97/lee19d/lee19d.pdf" title="A Framework for Attention-based Permutation-Invariant Neural ...">Proceedings of Machine Learning Research</a>)</p> </li> <li> <p>Training: DETR-style Hungarian matching + no-object class (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</p> </li> <li> <p>Freeze encoder; probe MLP predicts sin/cos, decode with <code>atan2</code></p> </li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-need-tighter-numeric-fidelity-24" name="p-250406-if-you-need-tighter-numeric-fidelity-24"></a>If you need tighter numeric fidelity</h3> <p>Add <strong>harmonics</strong> to the <em>input</em> numeric features (e.g., <code>sin(2^kθ), cos(2^kθ)</code> for <code>k=0..3</code>) and/or use numeric embedding modules from the “numerical features embeddings” line of work. (<a href="https://arxiv.org/abs/2006.10739" title="Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-expect-large-n-25" name="p-250406-if-you-expect-large-n-25"></a>If you expect large <code>N</code></h3> <p>Swap Set Transformer → Perceiver IO. (<a href="https://arxiv.org/abs/2107.14795" title="Perceiver IO: A General Architecture for Structured Inputs &amp; Outputs">arXiv</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-practical-online-resources-to-implement-quickly-26" name="p-250406-practical-online-resources-to-implement-quickly-26"></a>Practical online resources to implement quickly</h2> <ul> <li><strong>Set Transformer paper + official implementation</strong> (PMA(K) → fixed slots) (<a href="https://proceedings.mlr.press/v97/lee19d/lee19d.pdf" title="A Framework for Attention-based Permutation-Invariant Neural ...">Proceedings of Machine Learning Research</a>)</li> <li><strong>DETR paper + matcher implementation</strong> (Hungarian set matching; adapt cost to angles) (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</li> <li><strong>Numeric scalar embeddings</strong> (periodic + piecewise-linear modules; official repo) (<a href="https://papers.neurips.cc/paper_files/paper/2022/file/9e9f0ffc3d836836ca96cbf8fe14b105-Paper-Conference.pdf" title="On Embeddings for Numerical Features in Tabular Deep ...">NeurIPS Proceedings</a>)</li> <li><strong>Tokenization-free encoders</strong> if you insist on string-first: ByT5, CANINE (<a href="https://arxiv.org/abs/2105.13626" title="ByT5: Towards a token-free future with pre-trained byte-to-byte models">arXiv</a>)</li> <li><strong>Why floats in text break</strong> (concrete tokenization examples) (<a href="https://arxiv.org/pdf/2309.06236" title="Pitfalls of Representing and Tokenizing Temporal Data for ...">arXiv</a>)</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-what-i-would-do-next-implementation-order-27" name="p-250406-what-i-would-do-next-implementation-order-27"></a>What I would do next (implementation order)</h2> <ol> <li>Implement parser + angle canonicalization</li> <li>Implement per-entity numeric-first token builder (sin/cos)</li> <li>Implement Set Transformer PMA(K)</li> <li>Implement Hungarian matching loss (angle distance + optional category cost)</li> <li>Overfit a tiny dataset to verify near-zero error</li> <li>Train normally; freeze encoder; train probe</li> </ol>
discuss.huggingface.co
January 31, 2026 at 2:18 PM
Structured JSON (numerical data) into fixed-length vectors for MLP probing
<p>Hmm… <a href="https://huggingface.co/datasets/John6666/forum3/blob/main/json_for_mlp_1.md">Trying to do that with embedding seems to make it harder</a>…?</p> <hr /> <p>Below is what I would do <strong>for your exact goal</strong>: “JSON with a variable-length list of entities → fixed-length representation → an MLP probe recovers per-entity <code>(azimuth_deg, elevation_deg)</code> with high fidelity.”</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-core-recommendation-for-your-case-1" name="p-250406-core-recommendation-for-your-case-1"></a>Core recommendation for your case</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-treat-this-as-a-set-to-slots-problem-not-text-embedding-2" name="p-250406-treat-this-as-a-set-to-slots-problem-not-text-embedding-2"></a>Treat this as a <strong>set-to-slots</strong> problem, not “text embedding”</h3> <ul> <li>Your input is a <strong>set of structured records</strong> (category + two angles), not natural language.</li> <li>Your output needs <strong>per-entity recoverability</strong>, which is fundamentally different from “semantic similarity” embeddings.</li> </ul> <p>So the best default is:</p> <ol> <li><strong>Parse JSON</strong> into typed records (<code>category</code>, <code>az</code>, <code>el</code>)</li> <li>Build <strong>per-entity numeric-first tokens</strong> (angles stay numeric)</li> <li>Use a <strong>set/slot encoder</strong> that outputs a fixed number of <strong>slots <code>K</code></strong></li> <li>Train with <strong>Hungarian matching</strong> (DETR-style) so “slot ↔ entity” alignment is stable</li> <li>Freeze the encoder; train your <strong>MLP probe per slot</strong></li> </ol> <p>This training pattern (“fixed number of predictions vs variable ground truth, matched by bipartite assignment”) is standard in set prediction and is one of the main ideas behind <strong>DETR</strong>. (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-why-this-fits-your-constraints-3" name="p-250406-why-this-fits-your-constraints-3"></a>Why this fits your constraints</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-h-1-arbitrary-input-length-4" name="p-250406-h-1-arbitrary-input-length-4"></a>1) Arbitrary input length</h3> <p>Set/slot encoders consume <code>N</code> entity tokens and produce <code>K</code> slots (fixed shape), so <code>N</code> can vary.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-h-2-numerical-fidelity-5" name="p-250406-h-2-numerical-fidelity-5"></a>2) Numerical fidelity</h3> <p>You avoid “numbers-as-text” entirely (or keep text only for the category label), so floats are never degraded by tokenization. Tokenization pitfalls for floating point numbers are well documented—e.g., a float like <code>"3.14159"</code> can be split into multiple chunks by common LLM tokenizers, which is hostile to exact numeric recovery. (<a href="https://arxiv.org/pdf/2309.06236" title="Pitfalls of Representing and Tokenizing Temporal Data for ...">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-h-3-fixed-output-size-6" name="p-250406-h-3-fixed-output-size-6"></a>3) Fixed output size</h3> <p>You choose a capacity <code>K</code> and always output <code>K×d</code> (or flatten to <code>K·d</code>).</p> <p><strong>Important reality check:</strong> a finite fixed-size vector cannot losslessly encode an <em>unbounded</em> number of entities. In practice, you pick <code>K</code> large enough for your dataset and define an overflow policy (truncate, sample, or increase <code>K</code>). This is the same practical compromise used by set prediction models. (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-q1-encoders-llms-you-can-use-ranked-for-your-use-case-7" name="p-250406-q1-encoders-llms-you-can-use-ranked-for-your-use-case-7"></a>Q1) Encoders / “LLMs” you can use (ranked for your use case)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-a-best-first-choice-set-transformer-with-pmak-recommended-8" name="p-250406-a-best-first-choice-set-transformer-with-pmak-recommended-8"></a>A. Best first choice: <strong>Set Transformer with PMA(K)</strong> (recommended)</h3> <p><strong>What it is:</strong> A Transformer architecture designed for sets; it includes a pooling module (<strong>PMA</strong>) that uses <code>K</code> learnable seed vectors to produce exactly <code>K</code> outputs (your slots). (<a href="https://proceedings.mlr.press/v97/lee19d/lee19d.pdf" title="A Framework for Attention-based Permutation-Invariant Neural ...">Proceedings of Machine Learning Research</a>)</p> <p><strong>Why it’s a good fit</strong></p> <ul> <li>Permutation handling is built-in (set input, order-agnostic)</li> <li>Produces <strong>fixed <code>K</code> outputs</strong> cleanly via PMA</li> <li>Easy to combine with a DETR-style matching loss</li> </ul> <p><strong>Good implementation starting point</strong></p> <ul> <li>Official PyTorch implementation: (<a href="https://github.com/juho-lee/set_transformer" title="juho-lee/set_transformer: Pytorch implementation of set ...">GitHub</a>)</li> </ul> <p><strong>When it’s enough</strong></p> <ul> <li>If typical <code>N</code> is up to a few hundred, this is usually fine.</li> </ul> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-b-if-n-can-be-very-large-perceiver-io-9" name="p-250406-b-if-n-can-be-very-large-perceiver-io-9"></a>B. If <code>N</code> can be very large: <strong>Perceiver IO</strong></h3> <p><strong>What it is:</strong> A cross-attention architecture that scales linearly with input size and supports structured outputs via queries—i.e., you can decode <strong><code>K</code> slots</strong> efficiently even when <code>N</code> is big. (<a href="https://arxiv.org/abs/2107.14795" title="Perceiver IO: A General Architecture for Structured Inputs &amp; Outputs">arXiv</a>)</p> <p><strong>When to choose it</strong></p> <ul> <li>If <code>N</code> is frequently in the <strong>hundreds to thousands</strong>, Perceiver IO often scales better than full self-attention over inputs.</li> </ul> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-c-strong-baseline-deep-sets-10" name="p-250406-c-strong-baseline-deep-sets-10"></a>C. Strong baseline: <strong>Deep Sets</strong></h3> <p><strong>What it is:</strong> The canonical “set function” form: per-element encoder + permutation-invariant pooling. (<a href="https://arxiv.org/abs/1703.06114" title="Deep Sets">arXiv</a>)</p> <p><strong>Why it’s useful</strong></p> <ul> <li>It’s the simplest sanity-check: if you can’t get good probe decoding with this, your numeric encoding/training pipeline likely has issues.</li> </ul> <p><strong>Limitation</strong></p> <ul> <li>Vanilla pooling (sum/mean) tends to entangle entities; it’s better for “global properties of the set” than “recover each element”.</li> </ul> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-d-slot-binding-alternative-slot-attention-11" name="p-250406-d-slot-binding-alternative-slot-attention-11"></a>D. Slot-binding alternative: <strong>Slot Attention</strong></h3> <p>Not in the citations above, but conceptually: Slot Attention iteratively binds inputs to <code>K</code> slots. It can produce very clean per-slot “object-like” representations, but it often takes more tuning than Set Transformer.</p> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-e-numeric-embedding-modules-you-can-and-should-borrow-12" name="p-250406-e-numeric-embedding-modules-you-can-and-should-borrow-12"></a>E. Numeric embedding modules you can (and should) borrow</h3> <p>Regardless of the set encoder, you should embed numeric scalars into higher-dimensional vectors before mixing.</p> <p>A widely used reference is <strong>“On Embeddings for Numerical Features in Tabular Deep Learning”</strong> (periodic and piecewise-linear numeric embeddings) and its official implementation repo. (<a href="https://arxiv.org/abs/2203.05556" title="On Embeddings for Numerical Features in Tabular Deep Learning">arXiv</a>)</p> <p>This is directly relevant because your “angles must be recoverable” requirement is basically “my numeric scalars shouldn’t get washed out by the backbone.”</p> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-f-if-you-insist-on-an-llm-like-encoder-end-to-end-13" name="p-250406-f-if-you-insist-on-an-llm-like-encoder-end-to-end-13"></a>F. If you insist on an “LLM-like” encoder end-to-end</h3> <p>This is rarely the best path for your goal, but if you must:</p> <ol> <li><strong>Tokenization-free / byte/char models</strong></li> </ol> <ul> <li><strong>ByT5</strong> processes raw bytes (no subword vocabulary). (<a href="https://arxiv.org/abs/2105.13626" title="ByT5: Towards a token-free future with pre-trained byte-to-byte models">arXiv</a>)</li> <li><strong>CANINE</strong> processes characters without explicit tokenization/vocabulary. (<a href="https://arxiv.org/abs/2103.06874" title="CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation">arXiv</a>)</li> </ul> <p>These avoid subword fragmentation, but you still need <strong>task-specific training</strong> to make the representation numerically decodable.</p> <ol start="2"> <li><strong>Numeric tokenization research</strong> (if you’re training from scratch)</li> </ol> <ul> <li><strong>xVal</strong> proposes continuous numerical tokenization for numerically dense datasets. (<a href="https://arxiv.org/abs/2310.02989" title="xVal: A Continuous Numerical Tokenization for Scientific Language Models">arXiv</a>)</li> </ul> <p>This is relevant if you’re building a foundation-model-like pipeline, not if you just want a practical probe-friendly encoder quickly.</p> <hr /> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-why-not-clip-for-this-14" name="p-250406-why-not-clip-for-this-14"></a>Why not CLIP for this</h3> <ul> <li>Standard CLIP uses a fixed text context length (commonly <strong>77 tokens</strong>), and this is a known limitation in practice (errors when exceeding it, and model constraints around <code>context_length</code>). (<a href="https://github.com/openai/CLIP/issues/212" title="RuntimeError: Input is too long for context length 77 #212">GitHub</a>)</li> <li>There is active work on extending CLIP beyond this limit (e.g., TULIP and other approaches), but that still doesn’t address the deeper issue: CLIP-style text embeddings are not trained to preserve precise numeric values. (<a href="https://proceedings.iclr.cc/paper_files/paper/2025/file/54f117f842d47c2007b501bddd04c0e7-Paper-Conference.pdf" title="TULIP: TOKEN-LENGTH UPGRADED CLIP">ICLR Proceedings</a>)</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-q2-should-you-convert-json-to-natural-language-15" name="p-250406-q2-should-you-convert-json-to-natural-language-15"></a>Q2) Should you convert JSON to natural language?</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-recommendation-no-not-for-mlp-decodes-floats-16" name="p-250406-recommendation-no-not-for-mlp-decodes-floats-16"></a>Recommendation: <strong>No</strong>, not for “MLP decodes floats”</h3> <p>Converting structured records into prose tends to:</p> <ul> <li>introduce formatting variability (“-125.5 degrees” vs “-125.50°”)</li> <li>reintroduce tokenizer issues for numbers</li> <li>encourage the encoder to focus on semantics rather than numeric exactness</li> </ul> <p>If your success metric is “recover precise angles,” <strong>keep angles numeric</strong> and only embed categories as text if you truly need open-vocabulary behavior.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-what-to-do-instead-structured-tokens-numeric-first-17" name="p-250406-what-to-do-instead-structured-tokens-numeric-first-17"></a>What to do instead: structured → tokens (numeric-first)</h3> <p>A practical representation per entity:</p> <ul> <li> <p><strong>Category</strong>: embedding lookup (closed vocab) or a small text embedding (open vocab)</p> </li> <li> <p><strong>Angles</strong>:</p> <ul> <li>encode azimuth/elevation as <code>sin</code>/<code>cos</code> (wrap-safe)</li> <li>optionally add harmonics for higher precision</li> </ul> </li> </ul> <p>Why harmonics can help: Fourier feature mappings are known to help MLPs learn higher-frequency structure (reducing “spectral bias”). (<a href="https://arxiv.org/abs/2006.10739" title="Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-must-serialize-to-text-only-for-compatibility-18" name="p-250406-if-you-must-serialize-to-text-only-for-compatibility-18"></a>If you must serialize to text (only for compatibility)</h3> <p>Use a <strong>two-stream</strong> approach:</p> <ul> <li>Text stream: stable placeholders<br /> <code>{"category":"chair","az":&lt;NUM0&gt;,"el":&lt;NUM1&gt;}</code></li> <li>Numeric stream: the real floats <code>[-125.5, -37.2, ...]</code></li> </ul> <p>Then fuse them (concat + MLP) before the set encoder. This prevents your numeric fidelity from depending on the text tokenizer.</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-q3-how-to-prevent-subword-tokenizers-from-breaking-floats-19" name="p-250406-q3-how-to-prevent-subword-tokenizers-from-breaking-floats-19"></a>Q3) How to prevent subword tokenizers from breaking floats</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-the-most-reliable-solution-dont-tokenize-floats-20" name="p-250406-the-most-reliable-solution-dont-tokenize-floats-20"></a>The most reliable solution: <strong>don’t tokenize floats</strong></h3> <p>Parse JSON → keep numbers as floats → embed numerically.</p> <p>This completely avoids the class of problems illustrated by common tokenizers splitting floating point strings into multiple pieces. (<a href="https://arxiv.org/pdf/2309.06236" title="Pitfalls of Representing and Tokenizing Temporal Data for ...">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-must-go-through-a-tokenizer-anyway-21" name="p-250406-if-you-must-go-through-a-tokenizer-anyway-21"></a>If you must go through a tokenizer anyway</h3> <p>Use one of these, in descending order of reliability:</p> <ol> <li><strong>Tokenization-free encoders</strong> (byte/char): ByT5, CANINE (<a href="https://arxiv.org/abs/2105.13626" title="ByT5: Towards a token-free future with pre-trained byte-to-byte models">arXiv</a>)</li> <li><strong>Placeholder substitution + numeric side-channel</strong> (best practical workaround with standard LLM tokenizers)</li> <li><strong>Quantize</strong> angles to fixed-point integers (e.g., tenths of a degree) so the model learns an easier discrete mapping</li> <li><strong>Train a custom tokenizer/vocab</strong> where numbers (or digit chunks) are handled consistently (only if you control the full training stack)</li> </ol> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-a-concrete-configuration-i-would-start-with-defaults-22" name="p-250406-a-concrete-configuration-i-would-start-with-defaults-22"></a>A concrete configuration I would start with (defaults)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-want-the-simplest-works-first-version-23" name="p-250406-if-you-want-the-simplest-works-first-version-23"></a>If you want the simplest “works-first” version</h3> <ul> <li> <p><code>K = 32</code>, <code>d = 256</code></p> </li> <li> <p>Per-entity token:</p> <ul> <li>category embedding (e.g., <code>d_cat=64</code>)</li> <li>angles: <code>[sin(az), cos(az), sin(el), cos(el)]</code></li> <li>MLP → <code>d</code></li> </ul> </li> <li> <p>Set encoder: Set Transformer + PMA(K) (<a href="https://proceedings.mlr.press/v97/lee19d/lee19d.pdf" title="A Framework for Attention-based Permutation-Invariant Neural ...">Proceedings of Machine Learning Research</a>)</p> </li> <li> <p>Training: DETR-style Hungarian matching + no-object class (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</p> </li> <li> <p>Freeze encoder; probe MLP predicts sin/cos, decode with <code>atan2</code></p> </li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-need-tighter-numeric-fidelity-24" name="p-250406-if-you-need-tighter-numeric-fidelity-24"></a>If you need tighter numeric fidelity</h3> <p>Add <strong>harmonics</strong> to the <em>input</em> numeric features (e.g., <code>sin(2^kθ), cos(2^kθ)</code> for <code>k=0..3</code>) and/or use numeric embedding modules from the “numerical features embeddings” line of work. (<a href="https://arxiv.org/abs/2006.10739" title="Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains">arXiv</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250406-if-you-expect-large-n-25" name="p-250406-if-you-expect-large-n-25"></a>If you expect large <code>N</code></h3> <p>Swap Set Transformer → Perceiver IO. (<a href="https://arxiv.org/abs/2107.14795" title="Perceiver IO: A General Architecture for Structured Inputs &amp; Outputs">arXiv</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-practical-online-resources-to-implement-quickly-26" name="p-250406-practical-online-resources-to-implement-quickly-26"></a>Practical online resources to implement quickly</h2> <ul> <li><strong>Set Transformer paper + official implementation</strong> (PMA(K) → fixed slots) (<a href="https://proceedings.mlr.press/v97/lee19d/lee19d.pdf" title="A Framework for Attention-based Permutation-Invariant Neural ...">Proceedings of Machine Learning Research</a>)</li> <li><strong>DETR paper + matcher implementation</strong> (Hungarian set matching; adapt cost to angles) (<a href="https://arxiv.org/pdf/2005.12872" title="arXiv:2005.12872v3 [cs.CV] 28 May 2020">arXiv</a>)</li> <li><strong>Numeric scalar embeddings</strong> (periodic + piecewise-linear modules; official repo) (<a href="https://papers.neurips.cc/paper_files/paper/2022/file/9e9f0ffc3d836836ca96cbf8fe14b105-Paper-Conference.pdf" title="On Embeddings for Numerical Features in Tabular Deep ...">NeurIPS Proceedings</a>)</li> <li><strong>Tokenization-free encoders</strong> if you insist on string-first: ByT5, CANINE (<a href="https://arxiv.org/abs/2105.13626" title="ByT5: Towards a token-free future with pre-trained byte-to-byte models">arXiv</a>)</li> <li><strong>Why floats in text break</strong> (concrete tokenization examples) (<a href="https://arxiv.org/pdf/2309.06236" title="Pitfalls of Representing and Tokenizing Temporal Data for ...">arXiv</a>)</li> </ul> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250406-what-i-would-do-next-implementation-order-27" name="p-250406-what-i-would-do-next-implementation-order-27"></a>What I would do next (implementation order)</h2> <ol> <li>Implement parser + angle canonicalization</li> <li>Implement per-entity numeric-first token builder (sin/cos)</li> <li>Implement Set Transformer PMA(K)</li> <li>Implement Hungarian matching loss (angle distance + optional category cost)</li> <li>Overfit a tiny dataset to verify near-zero error</li> <li>Train normally; freeze encoder; train probe</li> </ol>
discuss.huggingface.co
January 31, 2026 at 10:18 AM
Adapt a customized data loading and data sampling pipeline to hf datasets
<p>It <a href="https://huggingface.co/datasets/John6666/forum3/blob/main/adapt_pipeline_to_hf_ds_1.md">seems like there might be a lot of tricky spots to achieve this</a>…</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250303-h-1-converting-the-corpus-into-datasets-recommended-approach-1" name="p-250303-h-1-converting-the-corpus-into-datasets-recommended-approach-1"></a>1) Converting the corpus into <img alt=":hugs:" class="emoji" height="20" src="https://emoji.discourse-cdn.com/apple/hugs.png?v=15" title=":hugs:" width="20" /> Datasets (recommended approach)</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250303-what-to-store-in-arrow-vs-keep-external-2" name="p-250303-what-to-store-in-arrow-vs-keep-external-2"></a>What to store in Arrow vs keep external</h3> <p>A <code>datasets.Dataset</code> is an Arrow table, which gives fast reads and convenient column operations, but it doesn’t require you to store the heavy bytes inside Arrow. In fact, the “Image” feature typically expects <strong>paths or embedded bytes</strong>, which would either force you to (a) extract images to files or (b) duplicate LMDB bytes into Arrow. (<a href="https://huggingface.co/docs/datasets/en/about_dataset_features" title="Dataset features">Hugging Face</a>)</p> <p><strong>Recommended</strong>: store <strong>metadata only</strong> in Arrow and keep <strong>LMDB as the image store</strong>. At training time, decode by <code>file_name</code> using a transform.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250303-shape-of-the-datasets-3" name="p-250303-shape-of-the-datasets-3"></a>Shape of the dataset(s)</h3> <p>You can represent your multi-source corpus in either of these two equivalent ways:</p> <ul> <li><strong>Option A (common)</strong>: <code>DatasetDict</code> with one <code>Dataset</code> per source</li> <li><strong>Option B</strong>: one merged <code>Dataset</code> with a <code>source</code> column</li> </ul> <p>Option A is usually nicer when each source has different policies (sampling freq/caps/epoch splits), because you can keep per-source stats and debug easily.</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250303-flatten-the-json-grouped-by-w_h-4" name="p-250303-flatten-the-json-grouped-by-w_h-4"></a>Flatten the JSON grouped by <code>"W_H"</code></h3> <p>Your idea to de-group and create explicit columns is correct. Suggested columns (superset of what you proposed):</p> <ul> <li><code>file_name</code> (LMDB key)</li> <li><code>label</code>, <code>type</code></li> <li><code>width</code>, <code>height</code> (parsed from <code>"W_H"</code>)</li> <li><code>resolution</code> (string, optional)</li> <li><code>bucket_id</code> (int; exact <code>(W,H)</code> or binned)</li> <li><code>is_doc</code> (per-source constant, but store it per-row so transforms can branch)</li> <li><code>source</code> (string)</li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250303-why-not-store-image-as-a-datasets-image-feature-here-5" name="p-250303-why-not-store-image-as-a-datasets-image-feature-here-5"></a>Why not store <code>image</code> as a Datasets “Image” feature here?</h3> <p>HF’s <code>Image</code> feature decodes from <strong>paths</strong> or <strong>stored bytes</strong>, and <code>decode=False</code> can return <code>{path, bytes}</code> instead of <code>PIL.Image</code>. This is useful when your data is already file-based or already stored as bytes in Arrow, but it doesn’t directly match “bytes live in LMDB keyed by file_name” without copying. (<a href="https://huggingface.co/docs/datasets/en/about_dataset_features" title="Dataset features">Hugging Face</a>)</p> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250303-training-time-decodeaugment-via-set_transform-with_transform-6" name="p-250303-training-time-decodeaugment-via-set_transform-with_transform-6"></a>Training-time decode/augment via <code>set_transform</code> / <code>with_transform</code></h3> <p>HF explicitly supports on-the-fly transforms for images (<code>set_transform</code>) and recommends using <code>map()</code> for one-time preprocessing and <code>set_transform</code> for per-epoch augmentations. (<a href="https://huggingface.co/docs/datasets/v2.2.1/en/image_process" title="Process image data">Hugging Face</a>)</p> <p>So the “HF datasets” part becomes:</p> <ol> <li> <p>offline: build metadata datasets and <code>save_to_disk()</code></p> </li> <li> <p>training: <code>load_from_disk()</code>, then attach a transform that:</p> <ul> <li>reads LMDB bytes by <code>file_name</code></li> <li>applies doc/scene augmentation based on <code>is_doc</code></li> <li>runs your image processor and returns tensors + passthrough metadata</li> </ul> </li> </ol> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250303-h-2-can-hf-datasets-implement-your-sampling-strategy-7" name="p-250303-h-2-can-hf-datasets-implement-your-sampling-strategy-7"></a>2) Can HF Datasets implement your sampling strategy?</h2> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250303-what-hf-can-do-directly-partial-8" name="p-250303-what-hf-can-do-directly-partial-8"></a>What HF can do directly (partial)</h3> <p>HF provides <code>interleave_datasets()</code> with:</p> <ul> <li><strong>probabilities</strong> (stochastic mixing)</li> <li><strong>stopping strategies</strong> like <code>first_exhausted</code> and oversampling modes like <code>all_exhausted</code> (<a href="https://huggingface.co/docs/datasets/en/process" title="Process">Hugging Face</a>)</li> </ul> <p>This is helpful for <em>rough</em> mixing, but it does <strong>not</strong> naturally express:</p> <ul> <li>“exactly 10 epochs with explicit epoch semantics”</li> <li>“source A is exactly 1× per epoch and source B is exactly 5× per epoch”</li> <li>“bucket by resolution and form batches within buckets”</li> <li>“DDP-safe, deterministic, disjoint work per rank”</li> </ul> <h3><a class="anchor" href="https://discuss.huggingface.co#p-250303-what-you-should-do-for-strict-semantics-recommended-9" name="p-250303-what-you-should-do-for-strict-semantics-recommended-9"></a>What you should do for strict semantics (recommended)</h3> <p>Use HF Datasets for <strong>indexable metadata storage</strong> and PyTorch for <strong>policy</strong>:</p> <ul> <li> <p>HF <code>Dataset</code>: metadata table, deterministic indexing</p> </li> <li> <p>Training transform: LMDB decode + augmentation (<code>set_transform</code> / <code>with_transform</code>) (<a href="https://huggingface.co/docs/datasets/v2.2.1/en/image_process" title="Process image data">Hugging Face</a>)</p> </li> <li> <p><strong>Custom PyTorch BatchSampler</strong>: the control plane for:</p> <ul> <li>per-source quotas per epoch (1×, 5×, caps)</li> <li>optional epoch-specific subset selection</li> <li>bucketing (<code>bucket_id</code>) → batches</li> <li>deterministic shuffles via <code>set_epoch(epoch)</code></li> <li>DDP partitioning at batch level</li> </ul> </li> </ul> <p>This division matches what HF is best at (dataset storage + column ops) and what PyTorch is best at (sampling/batching control).</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250303-h-3-ddp-sharding-memory-mapped-arrow-is-not-sharding-10" name="p-250303-h-3-ddp-sharding-memory-mapped-arrow-is-not-sharding-10"></a>3) DDP sharding: memory-mapped Arrow is not sharding</h2> <p>Arrow memory mapping improves I/O and avoids loading everything into RAM, but <strong>it doesn’t prevent two ranks from reading the same indices</strong>. You still need an explicit sharding mechanism (sampler or dataset sharding).</p> <p>HF provides <code>Dataset.shard(num_shards, index)</code> as a deterministic way to split a dataset into N pieces. (<a href="https://huggingface.co/docs/datasets/v3.1.0/en/process" title="Process">Hugging Face</a>)<br /> HF’s Trainer/Accelerate will do distributed sharding for you, but if you’re building a custom loop you must ensure your own per-rank split. (<a href="https://discuss.huggingface.co/t/should-i-shard-dataset-in-distributed-training/12494" title="Should I shard dataset in distributed training? - Datasets - Hugging Face Forums">Hugging Face Forums</a>)</p> <p>Separately, with PyTorch’s <code>DistributedSampler</code>, you must call <code>sampler.set_epoch(epoch)</code> each epoch to get different shuffles across epochs; otherwise, you can repeat the same order. (<a href="https://discuss.pytorch.org/t/why-is-sampler-set-epoch-epoch-needed-for-distributedsampler/149672" title="Why is 'sampler.set_epoch(epoch)' needed for DistributedSampler? - distributed - PyTorch Forums">PyTorch Forums</a>)</p> <p><strong>Practical recommendation for your case</strong>: partition at the <strong>batch</strong> level (global deterministic batch schedule, then rank takes every <code>world_size</code>-th batch). This makes “disjoint work + identical step counts” easy to reason about.</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250303-h-4-resolution-grouping-bucketing-with-hf-pytorch-11" name="p-250303-h-4-resolution-grouping-bucketing-with-hf-pytorch-11"></a>4) Resolution grouping (“bucketing”) with HF + PyTorch</h2> <p>HF itself doesn’t have a built-in “bucketed batch sampler” for your exact needs. The clean approach is:</p> <ol> <li>store <code>bucket_id</code> as a metadata column</li> <li>in your sampler, build <code>bucket_id -&gt; indices</code> per source</li> <li>shuffle within bucket per epoch</li> <li>emit batches that are bucket-pure (or near-pure if you allow spillover)</li> </ol> <p>HF docs explicitly note that if you want efficient batching and transforms, using a <code>BatchSampler</code> is the right tool on the PyTorch side. (<a href="https://huggingface.co/docs/datasets/en/use_with_pytorch" title="Use with PyTorch">Hugging Face</a>)</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250303-h-5-are-predefined-per-epoch-split-files-necessarycommon-12" name="p-250303-h-5-are-predefined-per-epoch-split-files-necessarycommon-12"></a>5) Are predefined per-epoch split files necessary/common?</h2> <p>They’re not “standard” unless you’re doing something like:</p> <ul> <li>curriculum learning (difficulty schedule)</li> <li>strict replay / auditability requirements</li> <li>staged inclusion (e.g., progressively adding noisy sources)</li> <li>hard constraints from upstream labeling processes</li> </ul> <p>If you don’t have one of those reasons, they mainly add complexity.</p> <p><strong>Simpler alternative</strong>: add a deterministic <code>split_id</code> per sample (e.g., hash of <code>file_name</code>) and define epoch subsets as rules like “epoch 0 uses split_id in [0..k)”, epoch 1 uses [k..2k), etc. This avoids managing 10 external JSON lists and keeps the logic local and reproducible.</p> <hr /> <h2><a class="anchor" href="https://discuss.huggingface.co#p-250303-summary-recommendation-13" name="p-250303-summary-recommendation-13"></a>Summary recommendation</h2> <ul> <li>Convert each source to a <strong>metadata-only</strong> HF <code>Dataset</code> (<code>file_name,label,type,is_doc,width,height,bucket_id,source,...</code>) and <code>save_to_disk</code>.</li> <li>Use <code>set_transform</code> / <code>with_transform</code> to <strong>decode LMDB + augment + image_processor</strong> on the fly. (<a href="https://huggingface.co/docs/datasets/v2.2.1/en/image_process" title="Process image data">Hugging Face</a>)</li> <li>Implement strict mixing (1×/5×), epoch semantics, bucketing, and DDP disjointness in a <strong>custom PyTorch BatchSampler</strong>, not in <code>interleave_datasets</code>. (<a href="https://huggingface.co/docs/datasets/en/process" title="Process">Hugging Face</a>)</li> <li>Do not rely on Arrow memory mapping for sharding; use <code>Dataset.shard</code>/<code>DistributedSampler</code>/batch-level partitioning. (<a href="https://huggingface.co/docs/datasets/v3.1.0/en/process" title="Process">Hugging Face</a>)</li> </ul>
discuss.huggingface.co
January 29, 2026 at 12:15 PM
Building an AI-first company: What these two business leaders learned from top experts

Adam Brotman, left, and Andy Sack, authors of the book, “AI First.” (Photo Courtesy Forum3) This week on the GeekWire Podcast, our guests are Adam Brotman and Andy Sack, co-authors of AI First: The Playbook for…
Building an AI-first company: What these two business leaders learned from top experts
Adam Brotman, left, and Andy Sack, authors of the book, &#8220;AI First.&#8221; (Photo Courtesy Forum3) This week on the GeekWire Podcast, our guests are Adam Brotman and Andy Sack, co-authors of AI First: The Playbook for a Future-Proof Business and Brand.  Brotman was Starbucks’ chief digital officer and later co-CEO of J.Crew. Sack is a founder, investor, and longtime advisor to tech leaders. Together, they run Forum3, a Seattle-based company that helps brands with customer loyalty and engagement. For their book, they interviewed experts including Bill Gates, Sam Altman, Reid Hoffman and Ethan Mollick, and spent time with companies and leaders that have seen early AI success.
nextbusiness24.com
August 16, 2025 at 5:00 PM
Don’t count the apple in ‘Ai Race’: It could be in the best position of all

Adam Brotman and Andy Sack have cooled on forum3. As a generation AI reshapes technician, most attention went to Openai, GoogleAnthropic, MicrosoftXAI and Chinese Deepseek. Their models are powerful, their demonstrations…
Don’t count the apple in ‘Ai Race’: It could be in the best position of all
Adam Brotman and Andy Sack have cooled on forum3. As a generation AI reshapes technician, most attention went to Openai, GoogleAnthropic, MicrosoftXAI and Chinese Deepseek. Their models are powerful, their demonstrations shine and their progress in fast. In contrast, Apple is a secrecy and methodical - he pulled skepticism because of the observed and inertia. In spite of&hellip;
todaysportnews.com
July 1, 2025 at 2:22 PM
[blog] Considerably richer than you: To promote its upcoming recruitment event, forum3 has an..
Considerably richer than you
Our fundraising news, ideas and inspiration for professional charity fundraisers and non-profits featuring the tag recruitment / people.
tinyurl.com
January 15, 2025 at 3:11 PM
Bénéteau Cup 2012 : Vive les ptits culs des années 70 !

pinterest.com/pin/1814810599… association-first30.org/forum3/downlo
April 6, 2025 at 3:17 PM
Don’t count the apple in ‘Ai Race’: It could be in the best position of all

Adam Brotman and Andy Sack have cooled on forum3. As a generation AI reshapes technician, most attention went to Openai, GoogleAnthropic, MicrosoftXAI and Chinese Deepseek. Their models are powerful, their demonstrations…
Don’t count the apple in ‘Ai Race’: It could be in the best position of all
Adam Brotman and Andy Sack have cooled on forum3. As a generation AI reshapes technician, most attention went to Openai, GoogleAnthropic, MicrosoftXAI and Chinese Deepseek. Their models are powerful, their demonstrations shine and their progress in fast. In contrast, Apple is a secrecy and methodical - he pulled skepticism because of the observed and inertia. In spite of&hellip;
todaysportnews.com
July 1, 2025 at 2:24 PM
FORUM3,FORUM 4 E ANALISE CRITICA DA PAG 145 LEITURA E ESCRITA ONLINE

EDC287
October 31, 2024 at 12:08 AM