Idea
A multimodal AI platform improving diabetic retinopathy screening accuracy and explainability for clinicians and healthcare providers
Research Paper
Core Innovation
This paper demonstrates the use of multimodal large language models to simulate and improve clinical AI assistance for diabetic retinopathy detection. It uniquely combines general-purpose and medical-specific MLLMs to evaluate different output formats and AI collaboration strategies. The approach enhances screening accuracy and explainability without direct image access, supporting scalable clinical workflows.
Market Size (TAM)
$2–10B TAM for AI-powered medical imaging diagnostics; $1–2B SAM from ophthalmology clinics and hospitals. Driven by increasing diabetic retinopathy prevalence and demand for scalable AI screening tools.
Potential Customers & Pain Points
- Ophthalmology Clinics Needing Accurate DR Screening
- Hospitals Seeking Explainable AI Assistance
- Medical AI Developers Lacking Scalable Simulation Tools
- Low-Resource Healthcare Settings Requiring Lightweight Models
Business Model
Subscription-based API access for healthcare providers and AI developers; licensing open-source models for customization in low-resource settings
Competitive Landscape
- IDx-DR
- Google DeepMind
- Eyenuk
Implementation Challenges
- Regulatory Approval for Clinical Use
- Integration with Existing Clinical Workflows
- Data Privacy and Security Concerns
Validation Strategy
- Conduct clinical trials comparing model outputs with expert ophthalmologists
- Pilot integration in hospital screening workflows
- Collect user feedback to refine explainability features
Research Paper Overview
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
Summary
This paper evaluates multimodal large language models (MLLMs) for diabetic retinopathy detection and their ability to simulate clinical AI assistance with different output types. Two models, GPT-4o and MedGemma, were tested on IDRiD and Messidor-2 datasets. MedGemma showed higher sensitivity and AUROC at baseline, while GPT-4o had near-perfect specificity but low sensitivity. Both models adapted predictions based on simulated AI inputs, with MedGemma being more stable. In AI-to-AI collaboration, GPT-4o achieved strong results guided by MedGemma's descriptive outputs without direct image access. The findings suggest MLLMs can improve DR screening pipelines and serve as scalable simulators for clinical AI assistance, with open models like MedGemma valuable in low-resource settings and descriptive outputs enhancing explainability and trust.