Idea
An efficient spatial-temporal AI framework for real-time violence detection in public surveillance enhancing safety monitoring.
Research Paper
Core Innovation
This paper introduces Vi-SAFE, which integrates an enhanced YOLOv8 with a Temporal Segment Network to detect violence in videos efficiently. It uses a lightweight GhostNetV3 backbone and EMA attention to reduce computational cost without sacrificing accuracy. The separate training of detection and classification modules improves performance on small-scale and complex surveillance targets.
Market Size (TAM)
$10–20B TAM for video surveillance AI; $2–5B SAM from public safety and smart city surveillance. Driven by rising demand for automated threat detection and real-time monitoring.
Potential Customers & Pain Points
- Public Safety Agencies Needing Real-Time Violence Alerts
- Surveillance System Providers Seeking Efficient AI Integration
- Smart City Developers Requiring Scalable Violence Detection
- Security Firms Facing High Computational Costs
- Event Organizers Wanting Automated Threat Detection
Business Model
Licensing AI detection software to surveillance system providers and public safety agencies; offering cloud-based API and on-premise deployment options.
Competitive Landscape
- BriefCam
- AnyVision
- Avigilon
Implementation Challenges
- Integration with existing surveillance infrastructure
- Data privacy and regulatory compliance
- Real-time processing on edge devices
Validation Strategy
- Pilot deployment with city surveillance systems
- Benchmark against existing violence detection solutions
- Collect user feedback to optimize real-time performance
Research Paper Overview
Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
Summary
This study proposes Vi-SAFE, a spatial-temporal framework combining an optimized YOLOv8 with a Temporal Segment Network for real-time violence detection in public surveillance videos. The approach addresses challenges like small-scale targets and complex environments by using a lightweight GhostNetV3 backbone, EMA attention, and pruning to reduce computation while maintaining accuracy. YOLOv8 extracts human regions and TSN classifies violent behavior, achieving 0.88 accuracy on the RWF-2000 dataset, outperforming existing methods in accuracy and efficiency.