Idea
A framework and analysis platform for AI developers and researchers to evaluate and improve prompt robustness in large multimodal models.
Research Paper
Core Innovation
This paper introduces Promptception, a comprehensive framework that systematically measures how sensitive large multimodal models are to subtle prompt variations. Unlike prior work focusing on single prompt designs, it evaluates 61 prompt types across multiple categories and models, revealing significant accuracy fluctuations and differences between proprietary and open-source models. This enables more robust and fair evaluation of LMMs by understanding prompt sensitivity.
Market Size (TAM)
$2–10B TAM, $1–2B SAM; assumption: growing adoption of multimodal AI in enterprises and research requiring reliable evaluation tools.
Potential Customers & Pain Points
- AI Developers Needing Reliable Prompt Evaluation
- Researchers Studying Model Robustness
- Enterprises Deploying Multimodal AI Facing Inconsistent Outputs
Business Model
Subscription-based SaaS platform offering prompt sensitivity analysis APIs and dashboards for AI developers and enterprises.
Competitive Landscape
- OpenAI
- Anthropic
- Cohere
Implementation Challenges
- Complexity of prompt design and evaluation
- Proprietary model access limitations
- Integration with diverse AI workflows
Validation Strategy
- Pilot with AI research labs to benchmark prompt sensitivity
- Partner with AI model providers for real-world testing
- Collect user feedback to refine prompt categories and metrics
Research Paper Overview
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
Summary
This paper investigates the sensitivity of Large Multimodal Models (LMMs) to prompt variations in Multiple-Choice Question Answering tasks. It reveals that minor changes in prompt phrasing can cause up to 15% accuracy fluctuations, complicating fair model evaluation. The authors introduce Promptception, a framework with 61 prompt types across 15 categories to systematically assess prompt sensitivity in 10 LMMs, including GPT-4o and Gemini 1.5 Pro, over 3 benchmarks. Proprietary models show higher sensitivity due to tighter instruction alignment, while open-source models are more stable but less adept with complex phrasing. The study proposes tailored prompting principles for robust and fair LMM evaluation.