Idea
A dynamic 3D reconstruction platform enabling real-time scene tracking and rendering from monocular video for AR and robotics.
Research Paper
Core Innovation
This paper introduces ProDyG, which separates static and dynamic scene components within a SLAM framework to enable online dynamic 3D reconstruction. It uses a novel motion masking strategy for robust pose tracking and progressively adapts a Motion Scaffolds graph to reconstruct dynamic parts. This approach achieves global consistency and detailed appearance modeling from monocular videos, outperforming existing online and transformer-based methods.
Market Size (TAM)
$10–20B TAM for 3D reconstruction and SLAM technologies; $2–5B SAM from AR, robotics, and autonomous vehicle industries. Driven by increasing demand for real-time dynamic environment understanding and AR/robotics adoption.
Potential Customers & Pain Points
- Augmented Reality Developers Needing Real-Time Scene Reconstruction
- Robotics Companies Requiring Accurate Dynamic Mapping
- Video Game Studios Seeking Detailed Dynamic Environments
- Autonomous Vehicle Developers Facing Dynamic Scene Challenges
- Researchers Working on SLAM and 3D Reconstruction
Business Model
Licensing the reconstruction platform as an SDK/API to AR, robotics, and autonomous vehicle companies; offering custom integration and support services.
Competitive Landscape
- NVIDIA Kaolin
- Google ARCore
- Apple ARKit
Implementation Challenges
- Integration with existing SLAM pipelines
- Handling highly complex dynamic scenes
- Computational efficiency on mobile devices
Validation Strategy
- Develop prototype SDK and test on standard dynamic SLAM benchmarks
- Partner with AR and robotics firms for pilot deployments
- Collect user feedback to optimize performance and scalability
Research Paper Overview
ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos
Summary
Achieving truly practical dynamic 3D reconstruction requires online operation, global pose and map consistency, detailed appearance modeling, and the flexibility to handle both RGB and RGB-D inputs. However, existing SLAM methods typically merely remove the dynamic parts or require RGB-D input, while offline methods are not scalable to long video sequences, and current transformer-based feedforward methods lack global consistency and appearance details. To this end, we achieve online dynamic scene reconstruction by disentangling the static and dynamic parts within a SLAM system. The poses are tracked robustly with a novel motion masking strategy, and dynamic parts are reconstructed leveraging a progressive adaptation of a Motion Scaffolds graph. Our method yields novel view renderings competitive to offline methods and achieves on-par tracking with state-of-the-art dynamic SLAM methods.