#HPLT
📢 First release: 38 monolingual reference LLMs (2.15B params) via #HPLT + #OpenEuroLLM

⚙️Trained on 100B tokens from HPLT v2 dataset
🌍 Cover EU langs + others
⚙️ Based on LLaMA, trained on #LUMI
📈 Useful for evaluation

Downloads + more info at openeurollm.eu/blog/hplt-oe...
July 18, 2025 at 9:32 AM
Just saw Superman, it's an 8/10 cause there isn't any sort of Mr. Terrific announcement out right now.

We need more Michael Hplt and we need it NOW.
July 12, 2025 at 10:46 AM
** New parallel data set ** . We've just released HPLT v2.0, a parallel data set of 50 languages paired with English, 380M sentence pairs in total. Extracted from the Internet Archive and Common Crawl hplt-project.org/datasets/v2.0
HPLT - High Performance Language Technologies
A space that combines petabytes of natural language data with large-scale model training
hplt-project.org
February 28, 2025 at 1:34 PM
Predstavujeme vám nový webový #korpus, kombinácia Araneum Slovacum VII Maximum a datasetov HPLT a FineWeb2. Zatiaľ najväčší deduplikovaný webový korpus slovenčiny.

10.4 miliárd tokenov, 8.6 miliárd slov, 33.6 miliónov dokumentov.

www.juls.savba.sk/webskahfcorp...
March 20, 2025 at 1:53 PM
That's a wrap for @nodalida.bsky.social ! Short, nice and intense. I presented our work on efficient MT @helsinki-nlp.bsky.social within the #HPLT project⚡️
March 5, 2025 at 9:59 AM
Our experts contributed to the latest #HPLT dataset publication, which contains some very interesting results! See here: t.co/uN2zoSF251 #DataScience
November 6, 2025 at 2:47 PM
nc47-c9qc

Ipsx-w86w

c9hs-h2z4

c3cs-5dkm

sxjc-bct6

p4gc-wh3k

8su4-c5c9

7hy6-c8sd

xcuk-wrhc

hplt-hdqw

47st-vdys

4w43-6c2l

4spw-khwj

c4f4-6vdy

ecdh-g9sr

8h4j-39hd

24cr-czqc

dphd-wt6w

ukcw-hc2s

XyJc-VcWI

htg4-gwde

4nk4-3xw3

qwyj-wqcs

fst8-ctsd

w4nt-w2hc
August 21, 2026 at 4:32 PM
https://hplt-project.org (High Performance Language Technologies) combines large quantities of text data and high-performance computing to build language and translation models. Another goal of this project is to publish the results of this project in a shared space with open licenses.
HPLT - High Performance Language Technologies
A space that combines petabytes of natural language data with large-scale model training
hplt-project.org
October 6, 2026 at 11:56 AM
Last week I was at @aclmeeting.bsky.social ! Lots of friendly faces, great work and amazing art ✨️ We presented HPLT v2 datasets together with @very-laurie.bsky.social 🎉 Read our paper here: aclanthology.org/2025.acl-lon...
August 3, 2025 at 9:06 AM
HPLT is of the datasets we are sharing in our world-readable catalogue across HPCs. Interesting talk at #LREC2026 in 15 min in room Menorca 1 at 16:20!!!
May 13, 2026 at 2:04 PM
The EU's 🇪🇺 HPLT project, coordinated by @ufal.mff.cuni.cz is at #EMNLP2025! It has supported it as a silver sponsor, disseminating HPLT results from our booth and through several papers. We'll continue to shape the future of multilingual datasets and models here and in @openeurollm.bsky.social!
November 7, 2025 at 9:03 PM
🚀 Just added our HPLT fast translation models to a new TranslateLocally repository! Translate on your own machine—fast, private, and easy. shorturl.at/R2vzw
20+ models for diverse languages—learn more about them next week at @nodalida.bsky.social!
February 24, 2025 at 9:04 AM
My scars are lowkey fadigg hahaha i should jdut kill myself hplt fuck
February 18, 2026 at 11:21 AM
1. "An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)". LTG co-authors: Nikolay Arefyev, Mariia Fedorova, Andrey Kutuzov, Petter Mæhlum, Vladislav Mikhailov, Stephan Oepen, David Samuel and many others from hplt-project.org
arxiv.org/abs/2503.10267 (main ACL)
HPLT - High Performance Language Technologies
A space that combines petabytes of natural language data with large-scale model training
hplt-project.org
June 10, 2025 at 8:20 AM
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies
aclanthology.org/2025.acl-lon...
by @very-laurie.bsky.social, @onadegibert.bsky.social, @jindrahelcl.bsky.social, @hajicjan.bsky.social & many others from the EU's 🇪🇺 HPLT project, bringing high quality European data
July 30, 2025 at 12:22 PM
Happy to share the first models I contributed to as a part of #HPLT + @openeurollm.bsky.social project and @turkunlp.bsky.social group :)
📢 First release: 38 monolingual reference LLMs (2.15B params) via #HPLT + #OpenEuroLLM

⚙️Trained on 100B tokens from HPLT v2 dataset
🌍 Cover EU langs + others
⚙️ Based on LLaMA, trained on #LUMI
📈 Useful for evaluation

Downloads + more info at openeurollm.eu/blog/hplt-oe...
July 18, 2025 at 10:08 AM
Could it be the HPLT v3.0 multilingual dataset? list.elra.info/mailman3/hyp...
Release of the massive HPLT v3.0 multilingual dataset - Corpora - ELRA lists
list.elra.info
October 8, 2025 at 11:04 AM
Wait as in actually on fire hplt shit you good?
June 13, 2025 at 11:18 AM
UCF vs NC A&T: Bouncin through the crowd... and the rain ... now on the @sonsofucf.bsky.social YouTube Channel. #UCF
www.youtube.com/watch?v=hPlT...
UCF vs NC A&T: Bouncin through the crowd... and the rain
YouTube video by Sons of UCF
www.youtube.com
September 8, 2025 at 1:39 PM
"And I'll see the day that anyone gives us #1 without being forced to do so ..."

There are many LLM projects that are open about training and evaluation data, such as AllenAI OLMo, several EU projects (EuroGPT, HPLT), and several Huggingface projects. I don't think anybody forced them to do so.
January 30, 2025 at 11:19 AM
Hypeloot Crypto Casino AirDrop

Claim Your Airdrop Now!
Join the airdrop and receive 1 $HPLT token (worth $0.20) instantly by using the referral code zuelh6vgan.

Connect here to get started: presale.hypeloot.com
June 11, 2025 at 9:23 AM
There is now a dataset card on @hf.co, but no data... huggingface.co/datasets/HPL...
HPLT/HPLT3.0 · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
November 3, 2025 at 11:38 AM
英語の単語分割(not subword分割)で速いのって何かありますかね.
普段pythonだと扱いが楽でsacremoses使ってるんですけど,遅くて2倍速くしちゃった.
github.com/hplt-project...
でもまだ遅すぎる.
C++かRust製で速いのなんかないっすかね..
2x faster MosesTokenizer by de9uch1 · Pull Request #149 · hplt-project/sacremoses
I have added two speed improvements: Compile regex patterns. Pre-define the character sets for islower() and isanyalpha(). Before: Benchmark 1: python -m sacremoses -l en -j 1 tokenize < big.txt ...
github.com
May 26, 2024 at 9:53 AM