QA Scorecards & GenAI Evaluation
Omniflow automatically evaluates 100% of customer conversations across voice, email, and live chat channels using structured GenAI evaluation rubrics. Every conversation receives a detailed scorecard with numeric ratings, compliance checks, timestamped evidence citations, and actionable coaching feedback.
Scorecard Anatomy
Each completed conversation produces a structured scorecard evaluating core dimensions:
| Criterion | Scale | Score | Rationale & Cited Evidence |
|---|---|---|---|
| Empathy & Active Listening | 1 – 5 | 4/5 | ”Acknowledged customer frustration at 0:14 (‘I completely understand how disruptive this downtime is’). Missed secondary check-in at 1:42.” |
| Problem Resolution | 1 – 5 | 5/5 | ”Identified root cause (expired SSL cert) within 45 seconds, verified renewal, customer confirmed website restored.” |
| Regulatory & Policy Compliance | Pass / Fail | PASS | ”Mandatory call recording disclosure and security authentication phrases stated verbatim at turn 1.” |
| Communication Efficiency | 1 – 5 | 4.5/5 | ”First Contact Resolution achieved in 2 minutes 45 seconds with zero dead air.” |
| Overall Quality Score | 0 – 100% | 92% | Weighted composite score based on workspace rubric settings. |
Clicking any criterion expands the exact multi-turn reasoning trace with clickable timestamps linking directly to the audio recording waveform.
GenAI Evaluation Engine
Scorecards are generated asynchronously by the background-worker and conversation-runtime callback pipeline:
- Transcript & Metadata Ingestion: Complete diarized transcripts (user vs. agent), speech latency metrics, and tool execution logs are packaged.
- Structured JSON Evaluation: The conversation is evaluated against workspace rubric schemas using temperature
0.0for deterministic scoring consistency. - Confidence Scoring: Each score includes an AI confidence metric (0.0 to 1.0). Scores with confidence under 0.75 are automatically flagged for human supervisor review.
Human Supervisor Overrides & Audit Trail
Coaches and supervisors can review and override any AI-generated score:
- Open the scorecard in the QA Dashboard.
- Select the criterion to adjust and enter the corrected score.
- Provide a mandatory reviewer rationale note.
- Save the override.
The original AI score remains preserved in the audit log, allowing quality assurance managers to measure AI grader calibration over time.
Agent Dispute Workflow
Support agents can review their scorecards and dispute evaluations directly:
| Stage | Action |
|---|---|
| 1. Agent Submission | Agent flags a criterion, enters dispute rationale, and submits to the QA Queue. |
| 2. Supervisor Review | Supervisor compares agent rationale against transcript evidence. |
| 3. Resolution | Accept: Score updates immediately and metrics recompute. Reject: Original score stands with supervisor feedback note. |
High dispute rates on a specific criterion indicate that rubric definitions should be clarified in Scoring Rubric settings.
Export & Data Warehouse Streaming
- CSV / Excel: Export filtered scorecard sets directly from the QA table.
- REST API: Ingest raw scorecard JSON payloads into Snowflake, BigQuery, or Databricks using the
/api/v1/qa/scorecardsendpoint. - Slack / Email Alerts: Trigger immediate webhook notifications when any conversation scores below the minimum compliance threshold.
Related Documentation
| If you want to… | Read |
|---|---|
| Build custom evaluation criteria | Custom KPIs |
| Configure scoring rubrics | Scoring Rubrics |
| Analyze team performance trends | Reports & Trends |