Idea
A comprehensive first-person video dataset platform enabling developers and researchers to build interactive world exploration AI models.
Research Paper
Core Innovation
This paper introduces Sekai, a uniquely large and richly annotated first-person video dataset covering diverse global locations and conditions. Unlike prior datasets, Sekai combines walking and drone footage with detailed metadata such as weather and crowd density. This enables more realistic and interactive training for video-based world exploration models like YUME.
Market Size (TAM)
$2–10B TAM, $1–2B SAM; assumption: growing demand for AI training data and interactive exploration models in multiple industries.
Potential Customers & Pain Points
- AI Researchers Needing Diverse Video Data
- Autonomous Navigation Developers Lacking Real-World Footage
- VR/AR Content Creators Seeking Rich Environmental Context
- Urban Planners Requiring Crowd and Scene Analytics
Business Model
Subscription-based access to the dataset with tiered pricing for research, commercial, and enterprise users; custom data annotation services.
Competitive Landscape
- Google Street View Dataset
- Mapillary Vistas
- Cityscapes Dataset
Implementation Challenges
- High Cost of Data Collection and Annotation
- Ensuring Privacy and Ethical Use
- Integration Complexity with Existing AI Pipelines
Validation Strategy
- Pilot dataset release to select AI research labs
- Collect user feedback and usage metrics
- Iterate dataset quality and annotation depth based on feedback
Research Paper Overview
Sekai: A Video Dataset towards World Exploration
Summary
Sekai is a large-scale, high-quality first-person video dataset designed for world exploration, containing over 5,000 hours of walking and drone footage from more than 100 countries and 750 cities. It includes rich annotations such as location, scene, weather, crowd density, captions, and camera trajectories, enabling training of interactive video world exploration models like YUME.