Gaperon: A Peppered English-French Generative Language Model Suite

February 22, 2026

Reading time: 2 minute

...

📝 Original Info

Title: Gaperon: A Peppered English-French Generative Language Model Suite
ArXiv ID: 2510.25771
Date: 2025-10-29
Authors: ** 논문에 명시된 저자 정보가 제공되지 않았습니다. **

📝 Abstract

We release Gaperon, a fully open suite of French-English-coding language models designed to advance transparency and reproducibility in large-scale model training. The Gaperon family includes 1.5B, 8B, and 24B parameter models trained on 2-4 trillion tokens, released with all elements of the training pipeline: French and English datasets filtered with a neural quality classifier, an efficient data curation and training framework, and hundreds of intermediate checkpoints. Through this work, we study how data filtering and contamination interact to shape both benchmark and generative performance. We find that filtering for linguistic quality enhances text fluency and coherence but yields subpar benchmark results, and that late deliberate contamination -- continuing training on data mixes that include test sets -- recovers competitive scores while only reasonably harming generation quality. We discuss how usual neural filtering can unintentionally amplify benchmark leakage. To support further research, we also introduce harmless data poisoning during pretraining, providing a realistic testbed for safety studies. By openly releasing all models, datasets, code, and checkpoints, Gaperon establishes a reproducible foundation for exploring the trade-offs between data curation, evaluation, safety, and openness in multilingual language model development.

Gaperon: A Peppered English-French Generative Language Model Suite

📝 Original Info

📝 Abstract

💡 Deep Analysis

📄 Full Content

Reference

Table of Contents

Table of Contents

📝 Original Info

📝 Abstract

💡 Deep Analysis

📄 Full Content

Reference

Related Posts

A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction

Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models

BabyFlow: 3D modeling of realistic and expressive infant faces

Start searching

No results found