Idea
A platform that evaluates and scores language model jailbreak attempts for AI developers and security teams.
Research Paper
Core Innovation
This paper presents JADES, a novel framework that decomposes harmful inputs into weighted sub-questions for granular scoring. It aggregates these scores to provide a final jailbreak assessment with high accuracy and interpretability. The inclusion of an optional fact-checking module to detect hallucinations further enhances reliability compared to prior methods.
Market Size (TAM)
$2–10B TAM, $1–2B SAM; assumption: growing AI adoption and increasing demand for secure language model deployment.
Potential Customers & Pain Points
- AI Developers Needing Reliable Jailbreak Detection
- Security Teams Monitoring Language Model Abuse
- Enterprises Ensuring Safe AI Deployments
Business Model
Subscription-based API access for continuous jailbreak assessment and enterprise licensing for customized solutions.
Competitive Landscape
- OpenAI Moderation API
- Hugging Face Safety Tools
- Anthropic's Safety Models
Implementation Challenges
- Integration Complexity with Diverse Models
- Evolving Jailbreak Techniques
- Balancing Detection Sensitivity and False Positives
Validation Strategy
- Pilot integration with AI development teams
- Benchmark against existing jailbreak detection tools
- Collect user feedback to refine scoring and fact-checking modules
Research Paper Overview
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
Summary
JADES introduces a universal framework to accurately evaluate jailbreak attempts on language models by decomposing harmful inputs into weighted sub-questions, scoring each, and aggregating results for a final decision. It includes an optional fact-checking module to detect hallucinations and achieves 98.5% agreement with human evaluators, outperforming existing methods and providing consistent, interpretable jailbreak assessments.