Idea
Multi-task safety classification models reducing LLM content risks with efficient, scalable real-time detection across toxicity and jailbreaks.
Research Paper
Core Innovation
This paper introduces Opir, a family of encoder-based guardrail models built on GLiClass architecture, combining multi-task classification for safety with a large taxonomy of 996 categories. It integrates adversarially mined negatives, multilingual data, and generated examples to improve robustness while maintaining a small model size suitable for edge deployment.
Why It Matters
As AI language models become widespread, ensuring safe and compliant outputs is critical to prevent harm and misuse. Opir's efficient models reduce the cost and complexity of real-time safety filtering, enabling broader adoption in applications requiring robust content moderation. This scalability supports safer AI deployment across industries and platforms.
Market Size (TAM)
$2–10B TAM for AI content safety and moderation; $500M–$1B SAM from AI platform providers and social media companies. Driven by increasing regulatory pressure and AI adoption.
Potential Customers & Pain Points
- AI platform providers – Need cost-effective real-time safety filtering
- Social media companies – Require scalable toxic content detection
- Enterprises deploying LLMs – Need reliable jailbreak and harmful content prevention
- Edge device manufacturers – Demand lightweight safety models for on-device inference.
Business Model
Licensing models and APIs for AI developers and platform providers; edge model subscriptions for device manufacturers; consulting for integration and customization.
Competitive Landscape
- GLiNER2
- OpenAI Moderation API
- Perspective API
- HateSonar
- Jailbreak detection tools
Implementation Challenges
- Rapid evolution of harmful content requiring continuous model updates
- Balancing detection accuracy with false positives to avoid over-censorship
- Integration complexity with diverse LLM platforms and deployment environments
Validation Strategy
- Benchmark Opir against leading safety classifiers on public datasets
- Pilot deployments with AI platform partners to measure real-world performance and cost savings
- Collect user feedback on false positive/negative rates and update taxonomy accordingly
Research Paper Overview
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
Summary
Opir delivers real-time, multi-task safety classification for large language models, detecting unsafe prompts, toxic language, jailbreak attempts, and harmful content with a smaller deployment footprint. It supports binary and multi-label classification across a detailed taxonomy and multilingual data, outperforming many open-weight baselines while enabling efficient edge deployment.