Idea
An annotation-free text-based person retrieval platform using chat-driven generation and interaction to improve surveillance search accuracy and usability
Research Paper
Core Innovation
This paper introduces Multi-Turn Text Generation (MTG) to create diverse pseudo-labels without manual captions and Multi-Turn Text Interaction (MTI) to refine user queries dynamically. These modules together enable annotation-free, dialogue-based person retrieval that handles ambiguous or incomplete descriptions better than prior methods.
Market Size (TAM)
$2–10B TAM for AI-powered surveillance and security analytics; $1–3B SAM from law enforcement and private security firms. Driven by increasing demand for automated person search and scalable annotation-free solutions.
Potential Customers & Pain Points
- Surveillance agencies needing scalable person search
- Security firms requiring accurate image retrieval
- Law enforcement handling vague or incomplete descriptions
Business Model
SaaS platform licensing to security agencies and enterprises with tiered pricing based on database size and query volume
Competitive Landscape
- Clearview AI
- AnyVision
- Hikvision
Implementation Challenges
- Data privacy and ethical concerns
- Integration with existing surveillance infrastructure
- Handling diverse real-world language variations
Validation Strategy
- Pilot deployment with law enforcement agencies
- Benchmark against existing TBPS datasets
- User studies on query refinement effectiveness
Research Paper Overview
Chat-Driven Text Generation and Interaction for Person Retrieval
Summary
Text-based person search (TBPS) enables the retrieval of person images from large-scale databases using natural language descriptions, offering critical value in surveillance applications. However, a major challenge lies in the labor-intensive process of obtaining high-quality textual annotations, which limits scalability and practical deployment. To address this, we introduce two complementary modules: Multi-Turn Text Generation (MTG) and Multi-Turn Text Interaction (MTI). MTG generates rich pseudo-labels through simulated dialogues with MLLMs, producing fine-grained and diverse visual descriptions without manual supervision. MTI refines user queries at inference time through dynamic, dialogue-based reasoning, enabling the system to interpret and resolve vague, incomplete, or ambiguous descriptions - characteristics often seen in real-world search scenarios. Together, MTG and MTI form a unified and annotation-free framework that significantly improves retrieval accuracy, robustness, and usability. Extensive evaluations demonstrate that our method achieves competitive or superior results while eliminating the need for manual captions, paving the way for scalable and practical deployment of TBPS systems.