Idea
A full-stack platform leveraging AMD hardware, software and networking equipment to deliver high-speed large scale model training at lower cost.
Research Paper
Core Innovation
This paper delivers the first large-scale MoE pretraining on pure AMD hardware with detailed system-level benchmarks and MI300X-aware transformer design rules. It introduces the ZAYA1 MoE model achieving competitive results, demonstrating AMD's mature ecosystem for large-scale AI training, a novel contribution compared to prior GPU-centric approaches.
Why It Matters
Large-scale AI model training demands optimized hardware and software to reduce costs and improve efficiency. This platform leverages AMD's full-stack capabilities to deliver competitive training performance, enabling organizations to scale AI development with cost-effective infrastructure. It transforms workflows by providing validated, high-throughput training solutions on alternative hardware ecosystems.
Market Size (TAM)
$20–50B TAM for AI training infrastructure; $2–10B SAM from cloud providers, enterprises, and HPC centers. Driven by AI adoption growth and demand for cost-efficient training platforms.
Potential Customers & Pain Points
- AI research labs – Need cost-effective scalable training infrastructure
- Cloud providers – Require optimized hardware-software stacks for AI workloads
- Enterprises – Seek competitive AI model performance with lower infrastructure costs
- HPC centers – Demand efficient GPU networking and fault-tolerance for large-scale training.
Business Model
Enterprise software and hardware integration services; licensing optimized training stack; support and consulting for AMD-based AI infrastructure deployments.
Competitive Landscape
- NVIDIA DGX systems
- Google TPU pods
- AWS Trainium
Implementation Challenges
- Ecosystem maturity compared to dominant GPU vendors
- Customer inertia favoring established hardware
- Integration complexity with existing AI frameworks
Validation Strategy
- Benchmark ZAYA1 model performance against industry standards
- Pilot deployments with cloud providers and AI labs
- Collect user feedback on system stability and throughput
- Iterate on software stack based on real-world training workloads
Research Paper Overview
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
Summary
This paper presents the first large-scale mixture-of-experts pretraining on pure AMD hardware using MI300X GPUs and Pollara interconnect, offering detailed system and model design insights. It includes comprehensive cluster and networking benchmarks, MI300X kernel and memory performance data, and transformer sizing rules optimized for AMD. The study introduces the ZAYA1 MoE model, demonstrating competitive performance against leading base models across reasoning, math, and coding tasks, validating AMD's hardware and software stack for large-scale AI training.