Idea
Quantized training platform boosting LLM pre-training speed and accuracy on NVIDIA Blackwell GPUs.
Research Paper
Core Innovation
This paper presents MS-EDEN, a novel unbiased quantization routine that halves quantization error compared to stochastic rounding. Integrated into Quartet II, it enables fully NVFP4 quantized training of linear layers with consistently better gradient estimation on forward and backward passes, validated on large-scale LLM training with significant speedups on NVIDIA Blackwell GPUs.
Why It Matters
Large language model training is computationally expensive and memory-intensive, limiting scalability and cost-efficiency. Quartet II's improved quantization reduces precision loss and accelerates training, enabling more efficient use of hardware and lowering barriers for developing massive AI models. This can transform AI workflows by making large-scale training faster and more affordable.
Market Size (TAM)
$20–50B TAM for AI training hardware and software optimization; $2–10B SAM from cloud providers and AI enterprises. Driven by demand for cost-efficient large model training and hardware acceleration.
Potential Customers & Pain Points
- AI research labs – High compute costs for LLM training
- Cloud providers – Need to optimize GPU utilization and reduce energy consumption
- AI startups – Require scalable training solutions for large models
- Enterprises deploying AI – Need faster model iteration cycles with lower infrastructure costs
Business Model
Open-source core technology with enterprise licensing for optimized kernels and support; consulting and integration services for large AI training deployments.
Competitive Landscape
- NVIDIA TensorRT
- Google TPU quantization tools
- Intel Neural Compressor
- Microsoft DeepSpeed
Implementation Challenges
- Adoption limited by hardware compatibility and ecosystem integration
- Competition from established quantization and training optimization frameworks
- Need for extensive validation on diverse model architectures and workloads
Validation Strategy
- Benchmark end-to-end LLM training on NVIDIA Blackwell GPUs across multiple model sizes
- Compare accuracy and speed against FP16
- FP8
- and existing quantization methods
- Partner with cloud providers and AI labs for real-world deployment trials
Research Paper Overview
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
Summary
Quartet II introduces a novel unbiased quantization method, MS-EDEN, that reduces quantization error over stochastic rounding, enabling fully NVFP4 quantized training of large language models with improved accuracy and up to 4.2x speedup on NVIDIA Blackwell GPUs.