#TokenizerFree
🚧 Building out the pretrain pipeline for X-Spanformer: github.com/p3nGu1nZz/x-... /// PDF segmentation + judge/improver enrichment for Tau2.0 tokenizer. Zero tokens. All spans. #AI #TokenizerFree #TauSystems #NLP #TransformerArchitecture #OpenSource #FungalLogic #SpanAware #XBarTheory
July 12, 2025 at 9:27 PM
Up next on stage, Dr. @edoardo-ponti.bsky.social ( @edinburgh-uni.bsky.social / NVIDIA)
🎤 “Adaptive Units of Computation: Towards Sublinear-Memory and Tokenizer-Free Foundation Models”

Fascinating glimpse into the next gen of foundation models.

#FoundationModels #NLP #TokenizerFree #ADSAI2025
June 9, 2025 at 1:16 PM
🧠 X-Spanformer ditched "improver"—now guided by 5-judge consensus 🗳️ to approve text for ox-bar span compilation. Cleaner segments. Swarm decides.

#ai #artificialintelligence #transformers #ltsm #computerscience #XSpanformer #TokenizerFree #SpanAware #SemanticEmbeddings #OxBarTheory #TauSystem 🍄
July 13, 2025 at 4:49 AM
Back at it—system gave us 500 gems… and 10× more junk 😂. Quick tweaks and we’re nearly done with stage one: mining pretrain data from rare, cross-domain PDFs.

#AIpretrain #SpanAware #TokenizerFree #PDFMining #XSpanformer #DataCuration #OpenScience
#artificalintelligence
July 15, 2025 at 4:40 PM