QA Scorecard: Audit Your Conversation Quality


Cavalry Intelligence
Last Updated: 51 minutes ago

The QA Scorecard automatically grades every live conversation against a rubric you define, giving you per-criterion scores plus a 0 to 100 overall score. It evaluates how each conversation was handled, whether by the AI, by a human, or by both.

How it works

After every live conversation ends, a QA agent reviews the transcript against each criterion on your scorecard.

Term

What it means

Scorecard

A rubric of criteria that every conversation is scored against once it ends.

Criteria

Pass/fail or scale (e.g., 1 to 5) checks. Each is scored by an LLM judge with cited evidence.

Subcriteria

Optional groupings of related criteria (e.g., Tone → Friendly, Concise, Formatting). Parent scores roll up from their children.

Evidence

The QA agent quotes specific moments from the transcript to justify every score.

How to enable the QA Scorecard

  1. In your dashboard, go to AI Agent > QA Scorecard.

  2. Under Pick a starting point, choose one of the following:

    • Customer Support Essentials: a 5-criterion rubric (recommended starting point).

    • Full Customer Support Audit: an 11-category, 100-point rubric.

    • Start from scratch: click Create blank scorecard and add criteria one at a time.

  3. Click Use this template (or Create blank scorecard).

  4. Review the criteria. Use Edit to adjust a criterion's description, scoring, or weight, and Delete to remove any that don't apply to your storefront.

Once enabled, every conversation that ends from that point on is scored automatically. You can edit your scorecard at any time.

Choosing a template

Both templates are 100-point rubrics, and either can be edited after you pick it. If you're not sure where to begin, start with Customer Support Essentials and tweak it as you learn what matters most for your storefront.

Template

Criteria

Best for

Customer Support Essentials

5

A balanced starting point covering what most support teams care about: did the agent solve the problem, say true things, communicate clearly, sound human, and handle scope correctly. Applies equally to AI-only conversations, human-only conversations, and AI conversations that escalated to a human. Recommended starting point; tweak as you learn what matters most for your storefront.

Full Customer Support Audit

11

A detailed audit covering customer experience (tone, empathy, greeting, closing, readability, spelling) and operational accuracy (solution, form execution, outside process, follow-up, next steps). Each category contributes its max points to the overall score.

Template 1: Customer Support Essentials

Five criteria; weights sum to 100%.

Criterion

Scoring

Weight

Issue Resolution

Pass / Fail

30%

Information Accuracy

Pass / Fail

25%

Communication Clarity

Scale (1 to 3)

15%

Tone and Empathy

Scale (1 to 3)

15%

Routing Judgment

Pass / Fail

15%

Issue Resolution (Pass / Fail, 30%)

Did the agent actually solve the customer's stated problem (or route it appropriately)?

  • Pass: The customer's primary issue was resolved, OR was correctly handed off to the right party when out of scope. The customer would not need to follow up about the same issue.

  • Fail: The conversation ended with the customer's problem still unaddressed, the agent gave a non-solution or dropped the thread, or the customer would clearly need to ask again.

Information Accuracy (Pass / Fail, 25%)

Were all facts, policies, prices, product details, and procedural steps stated by the agent correct?

  • Pass: Every factual claim is verifiably correct given the storefront's policies, the customer's order, and the products discussed. If something wasn't known, the agent said so honestly instead of guessing.

  • Fail: The agent stated something untrue, made up information, quoted an incorrect price or policy, or invented a procedure that doesn't exist. Even a single material inaccuracy is a fail.

Communication Clarity (Scale, 15%)

How clear and easy to act on were the agent's messages?

  • 3: Easy to read, well-structured, appropriately concise. Any next steps are explicit and unambiguous.

  • 2: Generally clear with minor friction: slightly long, mild jargon, formatting could be better, or one ambiguous moment.

  • 1: Confusing, too long, too terse, or poorly formatted. The customer would need to re-read to understand or guess at next steps.

Tone and Empathy (Scale, 15%)

Was the tone appropriate to the situation, and was the customer's emotional state acknowledged when needed?

  • 3: Warm, on-brand, and human-feeling. When the customer was frustrated or distressed, the agent acknowledged it before solving. When the conversation was routine, the tone was friendly and professional.

  • 2: Polite and professional throughout. Could be slightly warmer or more personable. Minor emotional cues may have been missed but nothing was dismissive.

  • 1: Cold, robotic, dismissive, or pushy. Ignored clear frustration or distress, or used performative warmth in an off-key moment.

Routing Judgment (Pass / Fail, 15%)

Did the agent make the right call on scope: handle the conversation themselves or route it correctly?

  • Pass: Either the agent handled what was in their scope without unnecessary escalation/transfer, OR correctly routed something out of scope (complex disputes, complaints requiring management, anything sensitive or irreversible) to the right party. The customer reached the person who could actually help.

  • Fail: The agent escalated/transferred something they could and should have handled (a common policy question, a simple status check), OR attempted to handle something they shouldn't have (made a substantive financial decision without authority, promised something outside policy, refused a legitimate escalation request).

Template 2: Full Customer Support Audit

Eleven criteria, all scaled; weights sum to 100%. Solution Offered is a critical category.

Criterion

Scale

Weight

Tone

1 to 3

5%

Empathy

1 to 2

10%

Greeting

1 to 2

5%

Closing

1 to 2

5%

Readability

1 to 3

5%

Spelling and Grammar

1 to 3

5%

Solution Offered (Critical)

1 to 2

15%

Form Execution

1 to 3

20%

Outside Process

1 to 2

10%

Follow-Up

1 to 2

10%

Next Steps for Customers

1 to 3

10%

Tone (5%)

Score 1, 2, or 3 based on these tiers:

  • 3 = Exceeds Expectations: Adds friendliness or personalization beyond neutral; warm and engaging; strengthens connection with the customer. Example: "Great question! I've taken care of that for you. Please don't hesitate to reach out if anything else comes up."

  • 2 = Meets Expectations: Polite, professional, and empathetic; generally warm but could be more personable or polished. Example: "Your request has been completed. Let us know if you need anything else."

  • 1 = Needs Improvement: Tone feels flat, mechanical, or rushed; may seem dismissive or less engaging. Example: "It's done."

Empathy (10%)

Score 1 or 2 based on these tiers:

  • 2 = Meets Expectations: Acknowledges the customer's situation; polite and supportive tone; shows understanding or concern. Example: "Thank you for sharing your experience. We're so sorry you had to deal with that."

  • 1 = Needs Improvement: No acknowledgment of the customer's feelings; response feels cold, rushed, or dismissive. Example: "Thanks for your feedback. We'll pass it along to the appropriate team."

Greeting (5%)

Score 1 or 2 based on these tiers:

  • 2 = Meets Expectations: Warm and professional; includes customer name (if available); sets a positive tone. Example: "Hi Jenna! Thank you for reaching out, Pattern. We are happy to help with this!"

  • 1 = Needs Improvement: Missing greeting, or incorrect/impersonal greeting.

Closing (5%)

Score 1 or 2 based on these tiers:

  • 2 = Meets Expectations: Warm and professional goodbye; thanks the customer; offers further help if needed; leaves a positive final impression. Example: "Thank you for your patience. If you need further help, don't hesitate to reach out. Have a wonderful rest of your day!"

  • 1 = Needs Improvement: Missing or abrupt closing; no thanks or invitation for further questions; feels cold, rushed, or incomplete. Example: "Let me know if there's anything else."

Readability (5%)

Score 1, 2, or 3 based on these tiers:

  • 3 = Exceeds Expectations: Message is clear and easy to follow; no errors; proper sentence structure and flow; customer can understand without confusion; response sent in the customer's original language. Example: "Thanks for your patience. Your refund has been issued and should show up in your account in 3-5 business days."

  • 2 = Meets Expectations: 1-2 minor readability errors but message is still clear; or response was not sent in the customer's language.

  • 1 = Needs Improvement: Message is hard to follow or confusing; poor structure or unclear phrasing; may cause misunderstanding or require rereading. Example: "i process refund in 3 5 day itll show up maybe your bank time vary"

Spelling and Grammar (5%)

Score 1, 2, or 3 based on these tiers:

  • 3 = Exceeds Expectations: No spelling or grammar errors; perfect punctuation and sentence structure; clear and professional word choice.

  • 2 = Meets Expectations: 1-2 minor errors max (spelling, punctuation, etc.); message remains clear and professional overall. Example: "I issued your refund it will show in 3-5 buisness days."

  • 1 = Needs Improvement: Multiple errors (3+) that affect clarity or tone; reduces professionalism or causes confusion. Example: "refund done. u get in few days"

Solution Offered (Critical) (15%)

CRITICAL CATEGORY. Score 1 or 2 based on these tiers:

  • 2 = Meets Expectations: The solution clearly resolves the customer's issue; correct solution is provided; complete, appropriate, and easy for the customer to act on; shows full understanding of what the customer needed. Example: "I've refunded the full amount, and you'll see it in 3-5 business days."

  • 1 = Needs Improvement: The solution doesn't actually fix the issue, or it's vague; wrong solution provided or no solution offered; the customer may still be confused or need to follow up. Example: "I looked into it. Have a great day."

Form Execution (20%)

Score 1, 2, or 3 based on these tiers (ticket fields, tags, internal notes):

  • 3 = Exceeds Expectations: All required fields completed; accurate, clear, and clean formatting; fully follows all guidelines; all tags, ticket fields, and internal notes are correctly filled in.

  • 2 = Meets Expectations: Minor mistakes on fields (1-2 incorrectly marked fields); no major missing info; doesn't require rework but lacks polish. Example: one field is mis-tagged (e.g., ticket type marked incorrectly) but the rest is accurate.

  • 1 = Needs Improvement: Incomplete or incorrect fields (3+ incorrectly marked); key info missing or wrong; errors could cause issues or require correction.

Outside Process (10%)

Score 1 or 2 based on these tiers (refunds, sheets updates, backend actions):

  • 2 = Meets Expectations: All required outside actions completed; refunds processed; sheets (e.g., AA) updated; SEMs notified if needed; no steps missed; checked Amazon or Shopify for further detail. Example: agent processed the refund in Amazon.

  • 1 = Needs Improvement: Required actions not completed; missed refund/logging/backend step; did not check Amazon or Shopify; messaged the Brand when not applicable. Example: agent replied to the customer but didn't process the refund in Amazon.

Follow-Up (10%)

Score 1 or 2 based on these tiers (agent-specific):

  • 2 = Meets Expectations: Agent followed up when needed (e.g., solution still in progress, provided answer after getting resolution); made sure they had an update or closure; closed the loop with the customer. Especially for Amazon Answers / product questions.

  • 1 = Needs Improvement: No follow-up was sent when clearly needed; the customer was left without a resolution or confirmation.

Next Steps for Customers (10%)

Score 1, 2, or 3 based on these tiers:

  • 3 = Exceeds Expectations: Clearly explains what's happening next or what the customer needs to do, with all details; no guesswork; leaves an internal note if something needs to be passed to another agent. Example: delivery pictures included, tracking number provided, refund confirmed, customer knows to reorder.

  • 2 = Meets Expectations: 1-2 minor errors in next steps but customer still understands. Example: missing Amazon/marketplace link but directed to Amazon or Marketplace; missing contact info but brand mentioned; incorrect country-specific contact link.

  • 1 = Needs Improvement: Doesn't tell the customer what they need to do next or include details (or 3+ errors in next steps); message feels unfinished or confusing; no internal note outlining next steps for the next agent.


Was this article helpful?