Idea
Model learning adaptive tutoring policies from logged data to improve student engagement and success without new experiments.
Research Paper
Core Innovation
This paper introduces an offline contextual bandit framework that learns adaptive instructional policies directly from logged student interaction data. It uniquely maps interactions onto a continuous proficiency-difficulty scale using a Rasch model and optimizes a novel reward function balancing challenge and success to enhance learning flow. The method enables rapid policy evaluation and improvement without additional data collection.
Why It Matters
Educational platforms struggle to optimize personalized learning due to costly experiments and inaccurate simulators. This approach leverages existing interaction data to rapidly improve instructional policies, enhancing student engagement and learning outcomes. It scales easily across diverse educational settings, reducing time and cost for adaptive system improvements.
Market Size (TAM)
$10–20B TAM for adaptive learning platforms; $2–5B SAM from EdTech and online education providers. Driven by demand for personalized learning and cost reduction in experimentation.
Potential Customers & Pain Points
- EdTech companies – Need scalable adaptive learning improvements
- Online education platforms – Require cost-effective policy optimization
- Schools and universities – Seek personalized tutoring without extensive trials
- Corporate training providers – Want efficient learner engagement strategies.
Business Model
Subscription-based SaaS platform offering adaptive policy optimization tools and analytics for EdTech providers and educational institutions.
Competitive Landscape
- Knewton
- Carnegie Learning
- DreamBox Learning
- Smart Sparrow
Implementation Challenges
- Integration with existing ITS platforms
- Data privacy and compliance concerns
- Adoption resistance due to trust in new policies
Validation Strategy
- Pilot deployments with partner EdTech companies
- A/B testing policy improvements in live ITS environments
- User engagement and learning outcome tracking
- Iterative refinement based on real-world feedback
Research Paper Overview
Counterfactual learning of new adaptive instructional policies using logged data
Summary
Optimizing instructional policies in Intelligent Tutoring Systems (ITS) typically requires costly online experimentation or student simulators that may fail to capture real-world dynamics. This paper introduces an offline contextual bandit framework that learns new adaptive policies directly from logged interaction data. By mapping student-item interactions onto a continuous latent proficiency-difficulty scale using a Rasch model, we cast the tutoring process as a continuous stochastic bandit problem. We propose a novel reward function designed to optimize ''flow'' by balancing task challenge with student success. Our approach includes a round-specific behavior policy estimation that serves as both a propensity model for off-policy evaluation and a diagnostic tool for ITS adaptivity. We demonstrate the efficacy of this framework across four large-scale real-world datasets, achieving consistent policy improvements over the logged behavior policy. The results show that effective instructional policies can be learned and visualized within seconds of computation, providing a scalable path for improving adaptive learning systems without further data collection.