Idea
A vision-language model platform predicting short-term vehicle trajectories for autonomous driving systems to improve safety and reliability.
Research Paper
Core Innovation
This paper introduces KEPT, which uniquely combines vision-language models with a scalable exemplar retrieval system to predict ego trajectories from driving video frames. It integrates temporal frequency-spatial fusion and a triple-stage fine-tuning aligning language outputs with driving constraints, surpassing prior methods in accuracy and latency. This approach enables interpretable and trustworthy trajectory predictions for autonomous driving.
Market Size (TAM)
$10–20B TAM, $2–5B SAM; assumption: growing autonomous vehicle and ADAS markets demand advanced trajectory prediction solutions.
Potential Customers & Pain Points
- Autonomous Vehicle Manufacturers needing accurate trajectory prediction
- ADAS Developers requiring low-latency scene understanding
- Fleet Operators seeking collision reduction
- Urban Mobility Planners needing interpretable driving data
Business Model
Licensing the KEPT platform to automotive OEMs and ADAS developers; offering API access for integration; custom solutions for fleet operators.
Competitive Landscape
- Waymo
- Tesla Autopilot
- Mobileye
Implementation Challenges
- Integration complexity with existing vehicle systems
- Real-time processing constraints in diverse environments
- Regulatory approval for safety-critical applications
Validation Strategy
- Pilot integration with autonomous vehicle prototypes
- Benchmark performance on real-world driving datasets
- Collect user feedback from ADAS developers for refinement
Research Paper Overview
KEPT: Knowledge-Enhanced Prediction of Trajectories from Consecutive Driving Frames with Vision-Language Models
Summary
KEPT is a vision-language model framework that predicts short-horizon ego trajectories from consecutive front-view driving frames by integrating scene-aligned exemplars retrieved via a scalable k-means + HNSW system. It uses a temporal frequency-spatial fusion video encoder trained with self-supervised learning and a triple-stage fine-tuning process to align language outputs with spatial, physical, and temporal driving constraints. Evaluated on nuScenes, KEPT achieves state-of-the-art accuracy and low collision rates with sub-millisecond retrieval latency, enabling interpretable and trustworthy autonomous driving.