This tool helps you evaluate conversation quality at scale using AI. Here's how it works:
1ImportBring in conversations from Firestore
2ConfigureSet up evaluators with quality criteria and rubrics
3AnalyzeRun AI analysis on your conversations
4InsightsGenerate reports or calibrate evaluators
Conversations
Import, analyze, and review conversations.
How does conversation analysis work?
Import conversations from Firestore, run AI-powered analysis using your configured evaluators, then review quality ratings for each conversation.
ImportStart by importing conversations. Use filters to select specific clients, protocols, or date ranges.
AnalyzeSelect conversations and click Analyze. The AI evaluates each one against your configured criteria.
ReviewClick any conversation to see detailed ratings, rationale, and the original transcript.
Reports
Analyze patterns across your evaluated conversations and generate quality insights.
What can reports tell you?
Reports analyze patterns across your evaluated conversations -- recurring quality issues, trends over time, and actionable insights. Generate reports filtered by client or protocol to focus on specific areas.
PatternsDiscover which quality issues appear most frequently and across which clients or protocols.
FocusGenerate reports for specific clients or protocols to get targeted, actionable insights.
BrowseBrowse saved reports, view detailed findings, or delete outdated analyses.
Evaluators
Configure evaluators, rubrics, and profiles.
View only — contact an admin to make changes to evaluators or prompts.
What are evaluators?
Evaluators are the quality criteria the AI uses to rate your conversations. Each evaluator checks for a specific aspect -- like safety, compliance, or communication quality -- and assigns a rating based on its rubric.
ProfilesGroup evaluators into profiles. Switch profiles to analyze different protocols with different criteria.
RubricsEach evaluator has a rubric defining what "pass" vs "fail" means. Clear rubrics = more accurate ratings.
TestUse the Test tab in the editor to try an evaluator against a real conversation before analyzing at scale.