Idea
A vision-language platform that localizes object interaction regions from natural language, enabling smarter robotics and AR applications.
Research Paper
Core Innovation
This paper introduces Affogato, a large-scale dataset with 150K instances combining open-vocabulary text and 3D affordance heatmaps, enabling fine-grained part-level localization. It also proposes simple yet effective vision-language models leveraging pretrained part-aware backbones and text-conditional heatmap decoders, improving cross-domain generalization and performance on existing benchmarks.
Market Size (TAM)
$2–10B TAM, $1–2B SAM; assumption: growing demand for intelligent robotics, AR/VR, and AI interaction understanding.
Potential Customers & Pain Points
- Robotics Companies Needing Precise Interaction Localization
- AR/VR Developers Requiring Fine-Grained Object Understanding
- AI Researchers Lacking Large-Scale Affordance Datasets
Business Model
Offer API and SDK access to affordance grounding models and datasets for robotics, AR/VR, and AI research customers.
Competitive Landscape
- Google DeepMind
- Meta AI
- OpenAI
Implementation Challenges
- High computational cost for training large-scale models
- Complexity in accurately annotating affordance data
- Integration challenges with existing robotics and AR systems
Validation Strategy
- Release Affogato dataset publicly to gather community feedback
- Develop prototype API demonstrating affordance grounding in robotics
- Conduct benchmark comparisons on standard 2D and 3D datasets
Research Paper Overview
Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
Summary
Affordance grounding involves localizing object regions based on natural language interaction descriptions, crucial for intelligent agents to understand and interact with environments. This work introduces Affogato, a large-scale benchmark with 150K instances annotated with open-vocabulary text and 3D affordance heatmaps across diverse objects and interactions. It also presents vision-language models using pretrained part-aware vision backbones and a text-conditional heatmap decoder. Models trained on Affogato perform well on existing 2D and 3D benchmarks and generalize effectively across domains.