Idea
Dataset and models improving multi-source cyberattack detection with detailed ATT&CK technique labeling for enhanced threat analysis.
Research Paper
Core Innovation
This paper introduces the first public multi-source cybersecurity log dataset labeled with MITRE ATT&CK techniques across system, network, and browser telemetry. It fills gaps in existing datasets by combining all three sources with per-event technique granularity. The study also demonstrates the effectiveness of fine-tuning Small Language Models to improve multi-source event classification and technique identification.
Why It Matters
Cybersecurity teams struggle to detect multi-stage attacks spanning system, network, and browser logs due to lack of comprehensive labeled data. This dataset enables more accurate detection by correlating cross-source events with fine-grained ATT&CK labels, improving threat identification and response. It scales to real-world environments by covering diverse attack techniques and telemetry sources.
Market Size (TAM)
$20–50B TAM for cybersecurity detection platforms; $2–10B SAM from enterprise security and managed service providers. Driven by increasing cyberattack complexity and regulatory compliance demands.
Potential Customers & Pain Points
- Enterprise security teams – Need comprehensive multi-source data for accurate attack detection
- Security software vendors – Require labeled datasets to train advanced detection models
- Managed security service providers – Need scalable tools to identify complex multi-stage attacks
- Cyber threat researchers – Lack publicly available datasets with detailed ATT&CK labels.
Business Model
Offer the dataset as a subscription or licensing service to security vendors and researchers; provide fine-tuned model APIs for integration into security platforms; offer consulting for custom model training and deployment.
Competitive Landscape
- CICIDS
- UNSW-NB15
- ATLAS dataset
- CrowdStrike
- Palo Alto Networks
Implementation Challenges
- Integration complexity of multi-source telemetry in existing security workflows
- Challenges in achieving high accuracy for fine-grained ATT&CK technique identification
- Adoption resistance due to data volume and labeling complexity
Validation Strategy
- Pilot deployments with enterprise security teams to measure detection improvements
- Benchmarking against existing datasets and detection tools
- User feedback from security analysts on model output relevance and usability
- Collaboration with cybersecurity vendors for integration and performance testing
Research Paper Overview
Multi-Source Cybersecurity Logs: An ATT&CK-Labeled Dataset and SLM Evaluation
Summary
This paper presents a unique multi-source cybersecurity log dataset combining system, network, and browser data with detailed ATT&CK technique labels. It includes 870 sessions and 2.3 million events, capturing real attack scenarios with fine-grained labeling. The study also evaluates fine-tuned Small Language Models for event classification and technique identification, demonstrating significant improvements in detection accuracy.