Idea
A unified AI model for generating high-quality videos and images from text or images, enabling creators and developers to produce lifelike media efficiently
Research Paper
Core Innovation
This paper introduces Waver, a unified foundation model that generates both images and videos using a Hybrid Stream DiT architecture. It improves modality alignment and training speed compared to prior models. The use of an MLLM-based video quality filter enhances dataset quality, leading to state-of-the-art performance on multiple benchmarks.
Market Size (TAM)
$10–20B TAM, $2–5B SAM; assumption: growing demand for AI-generated video and image content across media and entertainment sectors.
Potential Customers & Pain Points
- Content Creators Needing Fast High-Quality Video Generation
- Marketing Agencies Seeking Custom Visual Content
- Game Developers Requiring Realistic Assets
- AI Developers Wanting Unified Multimodal Generation Models
Business Model
Offer API access and enterprise licensing for content generation; provide custom solutions for media and marketing firms; potential SaaS platform for creators
Competitive Landscape
- Runway ML
- Synthesia
- D-ID
Implementation Challenges
- High computational cost for training and inference
- Competition from established commercial AI video platforms
- Dataset curation and quality control challenges
Validation Strategy
- Develop a scalable API for text-to-video and image-to-video generation
- Partner with content creators for pilot projects and feedback
- Benchmark against commercial solutions on quality and speed metrics
Research Paper Overview
Waver: Wave Your Way to Lifelike Video Generation
Summary
Waver is a high-performance foundation model for unified image and video generation, capable of producing 5-10 second videos at 720p resolution, upscaled to 1080p. It supports text-to-video, image-to-video, and text-to-image generation within a single framework using a Hybrid Stream DiT architecture for better modality alignment and faster training. The model uses a curated dataset filtered by an MLLM-based video quality model, achieving top rankings on T2V and I2V leaderboards and matching or surpassing state-of-the-art commercial solutions.