Idea
A French toxicity detection benchmark and fine-tuning method improving language model accuracy for content moderators and AI developers.
Research Paper
Core Innovation
This paper introduces TOXIFRENCH, a novel French toxicity detection benchmark created with minimal manual labeling by leveraging LLM pre-annotation. It demonstrates that smaller language models can be more robust than larger ones for toxicity detection. The paper also proposes a Chain-of-Thought fine-tuning approach with dynamic weighted loss, which significantly improves model faithfulness and cross-lingual performance.
Market Size (TAM)
$2–10B TAM, $1–2B SAM; assumption: growing demand for multilingual content moderation and AI safety tools.
Potential Customers & Pain Points
- Content Moderation Teams Needing Accurate French Toxicity Detection
- AI Developers Lacking Robust French Toxicity Benchmarks
- Social Media Platforms Seeking Cross-Lingual Toxicity Solutions
Business Model
Offer API access to TOXIFRENCH-enhanced toxicity detection models and licensing for enterprise content moderation platforms.
Competitive Landscape
- HateSonar
- Perspective API
- Detoxify
Implementation Challenges
- Limited labeled French toxicity data
- Adoption resistance due to model size preferences
- Cross-lingual generalization challenges
Validation Strategy
- Benchmark model performance against existing French toxicity datasets
- Pilot integration with social media content moderation teams
- Measure improvements in detection accuracy and false positive rates
Research Paper Overview
ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
Summary
This paper introduces TOXIFRENCH, a large-scale French toxicity detection benchmark created with minimal manual labeling through LLM pre-annotation. It reveals that smaller language models outperform larger ones in robustness for toxicity detection. The authors propose a Chain-of-Thought fine-tuning method with dynamic weighted loss, significantly improving model faithfulness and achieving state-of-the-art results, with strong cross-lingual capabilities.