Idea
A Spanish speech forensics dataset and detection platform enabling accurate synthetic speech identification and attribution for security and media companies
Research Paper
Core Innovation
This paper introduces HISPASpoof, the first large-scale dataset focused on Spanish synthetic speech detection and attribution. It addresses the gap where existing detectors trained on English fail on Spanish. The dataset includes diverse Spanish accents and multiple zero-shot TTS systems, enabling improved detection and method attribution.
Market Size (TAM)
$2–10B TAM, $1–2B SAM; assumption: growing global demand for synthetic speech detection and forensic tools in multiple languages including Spanish.
Potential Customers & Pain Points
- Security Agencies Needing Synthetic Speech Detection In Spanish
- Media Companies Verifying Authenticity Of Spanish Audio Content
- AI Developers Lacking Spanish Synthetic Speech Benchmarks
Business Model
Subscription-based API access to detection and attribution models; licensing dataset for research and commercial use; consulting for forensic analysis integration
Competitive Landscape
- ASVspoof
- Google Voice Biometrics
- Microsoft Speaker Recognition
Implementation Challenges
- Limited existing Spanish synthetic speech datasets
- High variability in Spanish accents
- Adoption by security and media sectors
Validation Strategy
- Benchmark detection accuracy against English-trained models
- Pilot deployment with Spanish media verification teams
- Collaborate with security agencies for real-world testing
Research Paper Overview
HISPASpoof: A New Dataset For Spanish Speech Forensics
Summary
Zero-shot Voice Cloning and Text-to-Speech methods have advanced rapidly enabling highly realistic synthetic speech raising misuse concerns; existing detectors focus on English and Chinese leaving Spanish underrepresented; HISPASpoof is the first large-scale Spanish dataset for synthetic speech detection and attribution including real speech from six accents and synthetic speech from six zero-shot TTS systems; evaluation shows English-trained detectors fail on Spanish while HISPASpoof-trained models improve detection; also enables synthetic speech generation method attribution providing a benchmark for reliable Spanish speech forensics.