Idea
Open suite of compliant multilingual large language models enabling global AI applications with transparent data and training pipelines
Research Paper
Core Innovation
This paper introduces Apertus, an open LLM suite trained exclusively on openly available data respecting content-owner rights and robots.txt exclusions. It applies the Goldfish objective to reduce memorization risks while maintaining performance. Apertus significantly expands multilingual coverage with 15T tokens from over 1800 languages and releases all development artifacts for transparency and extension.
Market Size (TAM)
$20–50B TAM for multilingual AI and NLP models; $2–10B SAM from enterprises and AI developers adopting compliant and open LLMs. Driven by increasing demand for multilingual AI and data compliance regulations.
Potential Customers & Pain Points
- AI developers needing compliant training data pipelines
- Enterprises requiring multilingual NLP models
- Researchers seeking transparent and reproducible LLMs
Business Model
Open-source model weights and tools with paid enterprise support, customization services, and hosted API access
Competitive Landscape
- LLaMA
- BLOOM
- MPT
Implementation Challenges
- Scaling training infrastructure for large models
- Ensuring ongoing data compliance and filtering
- Competing with proprietary models on performance
Validation Strategy
- Benchmark Apertus models on multilingual NLP tasks
- Conduct audits verifying data compliance and filtering
- Engage early adopters for feedback and improvements
Research Paper Overview
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Summary
Apertus is a fully open suite of large language models addressing data compliance and multilingual representation. It uses only openly available data respecting content-owner rights and robots.txt exclusions, filters non-permissive and toxic content, and employs the Goldfish objective to reduce memorization risks. Trained on 15T tokens from over 1800 languages with 40% non-English data, Apertus models at 8B and 70B scales achieve state-of-the-art multilingual benchmark results among open models. All development artifacts are released under a permissive license for transparency and extension.