Go EuropeAIdirectory

Evidence-backed European AI register

HPLT — High Performance Language Technologies

Reviewed identity, roles, model listings, pricing and evidence for HPLT — High Performance Language Technologies.

Dataset compiled 2026-08-27 · evidence over inference

consortiumresearch institute3 model listings0 token-price rows

Reviewed identity

Headquarters scope
EU headquarters
Headquarters record
Prague
Legal entity
Horizon Europe project, grant agreement 101070350; coordinated by Charles University (CZ)
Ownership
public research consortium
Founded
2022
Confidence
high
Last verified
2026-08-22

Links

Official website

Open interactive profile

Evidence boundaries

Organisation listings identify a reviewed canonical record. Model listings show an association in the source register; they do not by themselves establish model creation, current API availability, hosting location or infrastructure ownership.

Model listings

  1. HPLT v2 / HPLT 3.0 datasetscatalogued with this organisation · origin not asserted
    Type
    text
    Parameters
    n/a (dataset)
    Context
    unknown
    Languages
    HPLT v2: ~8T tokens monolingual across 193 languages, plus 380M parallel sentence pairs across 51 languages. HPLT 3.0 (Nov 2025) expands this further and adds multilingual evaluation; a further ~3 petabytes from ArchiveBot planned for 2026.
    Licence
    open (CC0 / permissive per subset — verify per release)
    Release
    2024 (v1.2), 2025-03 (v2), 2025-11 (3.0)
  2. HPLT monolingual reference models (with OpenEuroLLM)catalogued with this organisation · origin not asserted
    Type
    text
    Parameters
    2.15B × 38 languages
    Context
    unknown
    Languages
    38 European and related languages
    Licence
    unknown
    Release
    2025-07
  3. HPLT machine translation modelscatalogued with this organisation · origin not asserted
    Type
    text
    Parameters
    unknown
    Context
    unknown
    Languages
    all official EU languages and beyond
    Licence
    open
    Release
    2023-2025

Profile note

The data spine of the European open-model stack. Almost every other entry in this file — OpenEuroLLM, LumiOpen/Poro/Viking, the 38 reference models — consumes HPLT corpora. Partners: Charles University (coordinator, CZ), University of Edinburgh (UK), University of Helsinki (FI), University of Turku (FI), University of Oslo (NO), Prompsit Language Engineering (ES). Cheap relative to its impact: €4.06M total for a resource used across the continent. Note it is one of the two credible European answers to 'what do we use instead of Common Crawl' — the other being Pleias' Common Corpus.

Record sources

  1. https://explore.openaire.eu/search/project?projectId=corda_____he::e048ed8c5b6bbafecf177bef795a9130
  2. https://aclanthology.org/2023.eamt-1.61/
  3. https://arxiv.org/pdf/2503.10267
  4. https://arxiv.org/html/2511.01066
  5. https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned
  6. https://language-data-space.ec.europa.eu/document/download/b965cd46-c7d0-461e-b79b-79932798c0e6_en
  7. https://hplt-project.org/about