#MultilingualNLP
Arrived in Singapore for #ICLR2025 and will be presenting PROVENCE on Friday, Poster session 3 at 10am, poster #255!

Blogpost: huggingface.co/blog/nadiinc...

Will be happy to meet & chat about #LLMs, #RAG, #InformationRetrieval and #MultilingualNLP :)

#NLProc @naverlabseurope
April 23, 2025 at 5:39 AM
Both projects, including corpora, model weights, code, and evaluations, are fully open-source! 🔓 📄 MaLA Corpus: www.olaresearch.org/MaLA/ 📄 EMMA-500 Suite: www.olaresearch.org/EMMA-500-Gen2/
#NLProc #LLMs #MultilingualNLP #AI #MachineTranslation #OpenSource
MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models
Introducing the MaLA suite: a 74B-token corpus spanning 939 languages, a 136B pre-training mix, and the EMMA-500 model optimized for massively multilingual adaptation.
www.olaresearch.org
July 20, 2026 at 1:38 PM
Contextual lexical methods outperform raw LLM prompting for Hindi pregroup assignment (64.56% accuracy), while symbolic lexical repair improves LLM outputs. This advances scalable grammatical type annotation for multilingual quantum NLP pipelines.

#QuantumNLP #MultilingualNLP #Research
Automatic Hindi Pregroup Supertagging for Quantum Natural Language Processing
arxiv.org
September 16, 2026 at 12:38 PM
A streamlined pipeline selects compact typological features and imputes missing data, yielding vectors that boost multilingual NLP accuracy. Accepted for EMNLP 2025 on 24 Sep 2025. Read more: https://getnews.me/compact-language-typology-improves-multilingual-nlp/ #multilingualnlp #typology #language
September 26, 2025 at 8:11 PM
The study proposes Tokenization Parity (TP) and Information Parity (IP) metrics, evaluating dialect classification, topic classification, and extractive QA across Latin and non‑Latin scripts. https://getnews.me/tokenization-biases-impact-multilingual-dialect-nlp/ #tokenizationparity #multilingualnlp
September 26, 2025 at 7:22 PM
HiDAC, a dual‑adapter model, hits 67.5% accuracy on DISRPT 2025 discourse classification while using far fewer trainable parameters than full fine‑tuning. Read more: https://getnews.me/hierarchical-dual-adapter-boosts-multilingual-discourse-classification/ #dualadapter #multilingualnlp #disrpt2025
September 24, 2025 at 2:50 PM
Come see the full results at #LREC2026, May 14, 5:20 PM - 7:00 PM (GMT+2), 344, Poster Area 1

🔗 arxiv.org/pdf/2511.01187

#LREC2026 #NLP #AIBias #LLMs #MultilingualNLP #ResponsibleAI
May 13, 2026 at 12:46 PM
This builds on the foundational harvesting work by Patrice Lopez & James Howison (SoftCite project), and is a collaboration with @DFKI, @HUBerlin, @CommonCrawl & Uni Mannheim.

Attending LREC? Let's connect!👋

#NLP #ScientificNLP #MultilingualNLP #SciLaD #ScienciaLAB #grobid
5/5
April 22, 2026 at 12:16 PM