Idea
A real-time gaze estimation model improving accuracy and robustness for interactive systems and human-computer interfaces.
Research Paper
Core Innovation
This paper introduces CapStARE, which uniquely integrates capsule networks with spatio-temporal modeling using dual GRU decoders specialized for different gaze dynamics. This design enables efficient part-whole reasoning and disentangled temporal features, outperforming prior methods in accuracy and speed while maintaining interpretability and fewer parameters.
Market Size (TAM)
$2–10B TAM for computer vision and eye tracking technologies; $1–2B SAM from AR/VR, robotics, and assistive device industries. Driven by growing demand for natural user interfaces and real-time interaction.
Potential Customers & Pain Points
- AR/VR Device Makers Needing Accurate Eye Tracking
- Human-Robot Interaction Developers Requiring Robust Gaze Models
- Assistive Technology Providers Seeking Efficient Gaze Estimation
- Researchers and Developers Facing Trade-offs Between Accuracy and Speed
Business Model
Licensing the gaze estimation model as an API or SDK to device manufacturers and software developers; offering customization and support services.
Competitive Landscape
- Tobii
- EyeTech Digital Systems
- Pupil Labs
Implementation Challenges
- Integration with diverse hardware platforms
- Handling extreme lighting and occlusion conditions
- Scaling to varied user demographics
Validation Strategy
- Benchmark CapStARE on additional real-world datasets
- Pilot integration with AR/VR and robotics partners
- Collect user feedback on accuracy and latency in live applications
Research Paper Overview
CapStARE: Capsule-based Spatiotemporal Architecture for Robust and Efficient Gaze Estimation
Summary
CapStARE is a capsule-based spatio-temporal model combining a ConvNeXt backbone, attention routing, and dual GRU decoders to improve gaze estimation accuracy and efficiency. It achieves state-of-the-art results on multiple gaze datasets while enabling real-time inference and better interpretability with fewer parameters. The model generalizes well to unconstrained and interactive scenarios, making it suitable for practical gaze tracking applications.