Idea
Benchmark platform evaluating AI agents automating Salesforce CRM workflows for enterprise software users.
Research Paper
Core Innovation
This paper introduces SCUBA, a realistic benchmark for evaluating AI agents on complex Salesforce CRM tasks derived from real user workflows. It uniquely supports sandbox execution with fine-grained metrics capturing milestone progress. The benchmark reveals significant performance gaps between open-source and closed-source models, highlighting challenges and opportunities in enterprise task automation.
Market Size (TAM)
$20–50B TAM for enterprise software automation; $2–10B SAM from CRM and workflow automation sectors. Driven by increasing enterprise digital transformation and AI adoption.
Potential Customers & Pain Points
- Enterprises using Salesforce needing workflow automation
- CRM software developers seeking realistic benchmarks
- AI researchers developing computer-use agents
- Platform administrators requiring reliable task automation
- Sales and service teams aiming to reduce manual CRM tasks
Business Model
Offer SCUBA as a subscription-based benchmark platform with API access for AI developers and enterprises; provide consulting for integration and customization.
Competitive Landscape
- UiPath
- Automation Anywhere
- Blue Prism
Implementation Challenges
- Complexity of enterprise software environments
- Limited generalization of AI agents across workflows
- Data privacy and sandbox environment constraints
Validation Strategy
- Conduct benchmark challenges with AI research labs and enterprises
- Publish comparative performance reports to demonstrate utility
- Iterate benchmark tasks based on user feedback and evolving CRM workflows
Research Paper Overview
SCUBA: Salesforce Computer Use Benchmark
Summary
SCUBA is a benchmark to evaluate computer-use agents on CRM workflows within Salesforce. It includes 300 tasks from real user interviews across three personas: platform administrators, sales representatives, and service agents. Tasks cover UI navigation, data manipulation, workflow automation, information retrieval, and troubleshooting. SCUBA runs in Salesforce sandbox environments with parallel execution and detailed evaluation metrics. Benchmarking shows open-source models perform poorly in zero-shot settings compared to closed-source models. Demonstration-augmented methods improve success rates and reduce time and costs. SCUBA aims to drive progress in reliable automation for complex business software.