Idea
A platform enabling training data attribution with limited model access and resources for AI developers and enterprises.
Research Paper
Core Innovation
This paper systematically studies training data attribution under restricted access and resource constraints. It introduces proxy models to perform attribution without full model access and demonstrates that attribution scores from models not trained on the target data remain useful. This approach reduces dependency on computational resources and broadens TDA applicability in practical settings.
Market Size (TAM)
$2–10B TAM for AI Model Interpretability Tools; $1–2B SAM from AI Developers and Enterprises Using Proprietary Models. Driven by increasing AI adoption and demand for model transparency.
Potential Customers & Pain Points
- AI Developers Lacking Full Model Access
- Enterprises With Limited Computational Resources
- Data Scientists Needing Efficient Data Attribution
- Companies Using Commercial Models Without Public Access
Business Model
Subscription-based SaaS platform offering API access and enterprise licenses for training data attribution services.
Competitive Landscape
- DataRobot
- Weights & Biases
- Fiddler AI
Implementation Challenges
- Limited Access to Proprietary Models
- Computational Resource Constraints
- Integration with Diverse AI Systems
Validation Strategy
- Develop prototype with proxy model-based attribution
- Conduct case studies with AI teams using commercial models
- Measure attribution accuracy and resource efficiency against benchmarks
Research Paper Overview
Exploring Training Data Attribution under Limited Access Constraints
Summary
Training data attribution (TDA) helps understand the influence of individual training data points on model predictions. Gradient-based TDA methods like influence functions perform well but require full model access and high computational resources, limiting real-world use. This work studies TDA under various access and resource constraints, using proxy models and showing that attribution scores from models not trained on the target dataset remain informative. The findings guide practical deployment of TDA with improved feasibility and efficiency under limited access.