Evidence-backed European AI register
Pleias
Reviewed identity, roles, model listings, pricing and evidence for Pleias.
Dataset compiled 2026-08-27 · evidence over inference
Reviewed identity
- Headquarters scope
- EU headquarters
- Headquarters record
- Paris (Station F, 5 Parvis Alan Turing, 75013)
- Legal entity
- SAS
- Ownership
- founder-owned private (no disclosed VC round)
- Founded
- 2024
- Confidence
- medium
- Last verified
- 2026-08-22
Evidence boundaries
Organisation listings identify a reviewed canonical record. Model listings show an association in the source register; they do not by themselves establish model creation, current API availability, hosting location or infrastructure ownership.
Model listings
- Baguettotron
- Type
- text
- Parameters
- 0.3B
- Context
- unknown
- Languages
- multilingual (French-centric)
- Licence
- unknown
- Release
- 2025-2026
- Monad
- Type
- text
- Parameters
- 56.7M
- Context
- unknown
- Languages
- unknown
- Licence
- unknown
- Release
- 2025-2026
- Pleias-RAG-350M / Pleias-RAG-1B
- Type
- text
- Parameters
- 0.35B / 1B
- Context
- unknown
- Languages
- multilingual
- Licence
- unknown
- Release
- 2025
- Pleias-SLM-RAG
- Type
- text
- Parameters
- 0.3B
- Context
- unknown
- Languages
- unknown
- Licence
- unknown
- Release
- 2026
- Sillon
- Type
- text
- Parameters
- 0.6B
- Context
- unknown
- Languages
- French
- Licence
- unknown
- Release
- 2026
- CommonLingua
- Type
- text
- Parameters
- unknown
- Context
- unknown
- Languages
- multilingual
- Licence
- unknown
- Release
- 2026
- Celadon
- Type
- text
- Parameters
- 0.1B
- Context
- unknown
- Languages
- unknown
- Licence
- unknown
- Release
- 2024
- OCRerrcr
- Type
- text
- Parameters
- 0.4B
- Context
- unknown
- Languages
- unknown
- Licence
- unknown
- Release
- 2025
- ksante-colbert-small
- Type
- embedding
- Parameters
- 33.4M
- Context
- unknown
- Languages
- French
- Licence
- unknown
- Release
- 2025
Profile note
The 'clean data' lab: coordinates Common Corpus (~2 trillion tokens, the largest open multilingual public-domain/rights-cleared pre-training dataset, ICLR 2026 oral) and trains tiny specialised models entirely on it. 31 models and 62 datasets published. No pricing is public — this is the main gap in the record.