#HPLT
https://hplt-project.org (High Performance Language Technologies) combines large quantities of text data and high-performance computing to build language and translation models. Another goal of this project is to publish the results of this project in a shared space with open licenses.
HPLT - High Performance Language Technologies
A space that combines petabytes of natural language data with large-scale model training
hplt-project.org
October 6, 2026 at 11:56 AM
nc47-c9qc

Ipsx-w86w

c9hs-h2z4

c3cs-5dkm

sxjc-bct6

p4gc-wh3k

8su4-c5c9

7hy6-c8sd

xcuk-wrhc

hplt-hdqw

47st-vdys

4w43-6c2l

4spw-khwj

c4f4-6vdy

ecdh-g9sr

8h4j-39hd

24cr-czqc

dphd-wt6w

ukcw-hc2s

XyJc-VcWI

htg4-gwde

4nk4-3xw3

qwyj-wqcs

fst8-ctsd

w4nt-w2hc
August 21, 2026 at 4:32 PM
hpLT{kZSJD(Jy/GZVP%22L6?C]^8NTK]
August 5, 2026 at 3:45 AM
Check NorOLMo 1.0 - our new 13B language model aimed at Norwegian language tasks.

https://huggingface.co/HPLT/NorOLMo-13B

We took English OLMo2-13B as the base model and continually pre-trained it on Norwegian data.
NorOLMo is the only modern fully open
#llm specifically adapted for #norwegian […]
Original post on sigmoid.social
sigmoid.social
July 2, 2026 at 3:54 PM
HPLT is of the datasets we are sharing in our world-readable catalogue across HPCs. Interesting talk at #LREC2026 in 15 min in room Menorca 1 at 16:20!!!
May 13, 2026 at 2:04 PM
✨ "HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models"
by Stephan Oepen, Nikolay Arefyev, Andrey Kutuzov, Maja Buljan, Lucas Charpentier, Mariia Fedorova, Jindra Helcl, Vladislav Mikhailov and many others […]
Original post on sigmoid.social
sigmoid.social
May 11, 2026 at 1:50 PM
April 19, 2026 at 9:12 PM
The iconic music festival had headliners and fans stripping down while dancing under the desert sun.

https://mrf.lu/Hplt
April 19, 2026 at 9:12 PM
HPLT Project Advances Multilingual AI and Translation

The High-Performance Language Technologies (HPLT) project pioneers large-scale multilingual resources for large language models and machine translation. Massive text collections fuel pre-training in the LLM era. These collections act as the…
HPLT Project Advances Multilingual AI and Translation
The High-Performance Language Technologies (HPLT) project pioneers large-scale multilingual resources for large language models and machine translation. Massive text collections fuel pre-training in the LLM era. These collections act as the 'crude oil' for AI development. Refining high-quality datasets from web data requires immense computational power. Corporations often dominate this space. For instance, datasets like C4, FineWeb 1 and 2, MADLAD-400, and Nemotron-CC highlight this trend.
therealpreneur.com
March 9, 2026 at 9:00 AM
Hplt fuck sleepy!!
February 27, 2026 at 1:16 AM
My scars are lowkey fadigg hahaha i should jdut kill myself hplt fuck
February 18, 2026 at 11:21 AM
Europe: “We own nothing, we just packaged your entire digital soul into 50 TB.”
CC0 = Ctrl-C, Ctrl-Own.
HPLT - High Performance Language Technologies
A space that combines petabytes of natural language data with large-scale model training
hplt-project.org
November 22, 2025 at 6:00 PM
30T tokens to teach AI every tongue—yet the crawl forgot to filter out the silence of extinct speakers. We’re raising polyglot gods on a graveyard of languages.
HPLT 3.0: Very Large-Scale Multilingual Resources for LLM and MT....
We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest...
arxiv.org
November 22, 2025 at 2:52 PM
The EU's 🇪🇺 HPLT project, coordinated by @ufal.mff.cuni.cz is at #EMNLP2025! It has supported it as a silver sponsor, disseminating HPLT results from our booth and through several papers. We'll continue to shape the future of multilingual datasets and models here and in @openeurollm.bsky.social!
November 7, 2025 at 9:03 PM
Our experts contributed to the latest #HPLT dataset publication, which contains some very interesting results! See here: t.co/uN2zoSF251 #DataScience
November 6, 2025 at 2:47 PM
Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Ba\~n\'on, Maja Buljan, Laurie Burchell, Lucas Charpentier, ...
HPLT~3.0: Very Large-Scale Multilingual Resources for LLM and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
https://arxiv.org/abs/2511.01066
November 4, 2025 at 10:30 AM
Stephan Oepen, et al.: HPLT~3.0: Very Large-Scale Multilingual Resources for LLM and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models https://arxiv.org/abs/2511.01066 https://arxiv.org/pdf/2511.01066 https://arxiv.org/html/2511.01066
November 4, 2025 at 6:30 AM
There is now a dataset card on @hf.co, but no data... huggingface.co/datasets/HPL...
HPLT/HPLT3.0 · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
November 3, 2025 at 11:38 AM
Could it be the HPLT v3.0 multilingual dataset? list.elra.info/mailman3/hyp...
Release of the massive HPLT v3.0 multilingual dataset - Corpora - ELRA lists
list.elra.info
October 8, 2025 at 11:04 AM
UCF vs NC A&T: Bouncin through the crowd... and the rain ... now on the @sonsofucf.bsky.social YouTube Channel. #UCF
www.youtube.com/watch?v=hPlT...
UCF vs NC A&T: Bouncin through the crowd... and the rain
YouTube video by Sons of UCF
www.youtube.com
September 8, 2025 at 1:39 PM
Last week I was at @aclmeeting.bsky.social ! Lots of friendly faces, great work and amazing art ✨️ We presented HPLT v2 datasets together with @very-laurie.bsky.social 🎉 Read our paper here: aclanthology.org/2025.acl-lon...
August 3, 2025 at 9:06 AM