Idea
An open-vocabulary human-object interaction detection model that enhances novel interaction recognition for computer vision applications.
Research Paper
Core Innovation
This paper introduces INP-CC, which generates interaction-aware prompts dynamically based on the scene to better capture key interaction patterns. It also refines human-object interaction concept representations through language model-guided calibration and negative sampling, enabling improved detection of novel interaction classes compared to prior methods.
Market Size (TAM)
$2–10B TAM, $1–3B SAM; assumption: growing demand for advanced computer vision in robotics, autonomous systems, and surveillance.
Potential Customers & Pain Points
- Autonomous Vehicle Developers Needing Better Scene Understanding
- Robotics Companies Requiring Accurate Human-Object Interaction Detection
- Security and Surveillance Firms Seeking Improved Activity Recognition
Business Model
Licensing the detection model as an API or SDK for integration into robotics, autonomous vehicles, and security platforms.
Competitive Landscape
- HICO-DET
- V-COCO
- OpenHOI
Implementation Challenges
- Complexity of integrating language models with vision systems
- Data scarcity for rare interaction classes
- Computational cost of dynamic prompt generation
Validation Strategy
- Benchmark against state-of-the-art HOI datasets
- Pilot integration with robotics perception systems
- User feedback from security and surveillance deployments
Research Paper Overview
Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
Summary
This paper presents INP-CC, an end-to-end open-vocabulary human-object interaction detector that improves detection of novel interaction classes by using interaction-aware prompt generation and concept calibration. It dynamically generates prompts based on the scene to focus on key interaction patterns and refines HOI concept representations via language model-guided calibration and negative sampling, significantly outperforming state-of-the-art models on benchmark datasets.