A Complete QA Scorecard Template for Arabic Call Centers
By Adnan Bassem — Founder, InfoDriven (Dubai). Building Arabic-first speech recognition for GCC call centers.Published June 10, 2026
Buyer GuidesLast updated: June 10, 2026
Why most QA scorecards fail
Two failure modes dominate. The first is coverage: most teams hand-review 1-5% of calls picked at random, which means a systemic problem — a mis-read disclosure, a broken refund script — runs for weeks before a reviewer happens to land on it. The second is subjectivity: criteria like "was the agent empathetic?" scored differently by every reviewer, which destroys agent trust in the program and turns calibration meetings into arguments.
The cure for the first is automation — transcribe 100% of calls and let software pre-score the objective criteria. The cure for the second is a rubric where every line item is binary or anchored to observable evidence in the transcript. The template below is built for both.
The four criteria families
Greeting and verification covers the opening 30 seconds: branded greeting in the customer's language register (Khaleeji caller greeted in Gulf Arabic, not stiff MSA), identity verification where required, and a purpose statement. It is short but disproportionately shapes CSAT.
Compliance is the non-negotiable family: scripted disclosures read in full, no prohibited promises, recording notification given. In GCC financial services these map to SAMA, DFSA, and CBUAE call-recording expectations, and a single miss is usually an auto-fail regardless of total score.
Resolution quality is the heart of the card: did the agent diagnose the actual issue, take correct action, set accurate expectations, and avoid an unnecessary repeat contact? Soft skills covers tone, interruptions, hold etiquette, and language matching — the items where per-segment sentiment data turns a vague impression into evidence.
The 100-point rubric
Weights reflect impact: compliance failures carry regulatory cost, resolution drives repeat-contact volume, and greeting/soft-skills shape CSAT. Adapt the line items to your scripts, but keep every item answerable from the transcript or sentiment data alone wherever possible.
- Greeting & verification — 15 points total
- • Branded greeting within first 10 seconds, in caller's dialect register (5)
- • Identity verification completed per policy before account discussion (5)
- • Call purpose confirmed and restated to the customer (5)
- Compliance — 25 points total (any item missed = auto-fail review)
- • Required disclosure(s) read completely and audibly (10)
- • Recording/monitoring notification given (5)
- • No unauthorized commitments, rates, or timelines promised (10)
- Resolution quality — 35 points total
- • Correct root cause identified (transcript shows diagnostic questions) (10)
- • Correct action taken or correctly escalated (10)
- • Next steps and timeline stated explicitly to the customer (10)
- • First-contact resolution attempted — no avoidable repeat contact created (5)
- Soft skills & call control — 25 points total
- • No talking over the customer; talk-listen ratio within team band (5)
- • Hold used correctly: permission asked, under 2 minutes, return acknowledged (5)
- • Customer sentiment ends neutral-or-better, or dip is recovered before close (10)
- • Language matching maintained — dialect/code-switching mirrored appropriately (5)
How transcription and sentiment automate pre-scoring
Roughly 60 of those 100 points are machine-checkable from a dialect-accurate transcript with speaker diarization. Greeting presence and timing, disclosure phrases, recording notification, hold durations, talk-listen ratio, and the sentiment trajectory are all detectable automatically. CallScribe transcribes Khaleeji, Levantine, and Egyptian Arabic with diarization via pyannote and per-segment sentiment, then exposes the results across 16 analytics KPIs — which means software can pre-fill the objective rows on every call, not a 5% sample.
The human reviewer's job then shrinks to the judgment rows — was the root cause right, was the escalation correct — on the calls the automation flags: compliance misses, sentiment dips that never recover, and statistical outliers. Teams that make this shift typically move from reviewing 2-5% of calls shallowly to reviewing 100% automatically plus the flagged minority deeply. One caveat: automated pre-scoring is only as good as the transcript. On Arabic dialect audio, a generic English-first model will miss disclosure phrasing and misattribute speakers often enough to corrupt the scores, which is why transcription accuracy is a QA-program requirement, not an IT preference.
Calibration sessions: keeping ten reviewers honest
Run calibration every two weeks: all reviewers score the same three calls independently, then compare line by line. Any rubric row where reviewers diverge by more than one scoring band gets rewritten with sharper evidence anchors — divergence is a rubric defect, not a reviewer defect. Track inter-reviewer variance over time; a healthy program converges within a few cycles. Automated pre-scoring helps here too, because the objective rows arrive pre-filled and identical for everyone, so calibration time is spent entirely on the judgment rows where human disagreement is meaningful.
Closing the loop: from scorecard to coaching
A scorecard that ends in a spreadsheet changes nothing. The loop that works: weekly, each team lead receives their agents' score trends plus an evidence pack — the exact transcript segments behind every lost point, with timestamps and the sentiment curve. Coaching conversations then reference specific moments ("here's the hold at 4:32 where you didn't ask permission") instead of generalities, which agents accept far more readily. Set one focus behavior per agent per fortnight, re-measure on 100% of their calls, and celebrate measured improvement publicly. Export the evidence packs in whatever format your workflow needs — CallScribe exports PDF, CSV, SRT, TXT, and DOCX — so coaching materials don't require dashboard logins.
Sources
Frequently asked questions
What should a call center QA scorecard include?▾
Four families: greeting and verification (about 15%), compliance (about 25%, with auto-fail for misses), resolution quality (about 35%), and soft skills and call control (about 25%). Every line item should be binary or anchored to observable transcript evidence to keep scoring consistent across reviewers.
How many calls should QA review per agent per month?▾
With manual review alone, most teams manage 4-8 calls per agent monthly — a 2-5% sample. With automated transcription and pre-scoring, 100% of calls get the objective rows scored automatically, and human reviewers focus on the flagged subset: compliance misses, unrecovered sentiment dips, and outliers.
Can sentiment analysis really score soft skills?▾
It scores the measurable parts: whether customer sentiment ends neutral-or-better, whether a mid-call dip recovers, talk-listen ratio, and interruptions. Judgment items like empathy phrasing still need human review — but sentiment data turns those reviews from impressions into evidence-backed conversations.
How do calibration sessions work?▾
Every two weeks, all reviewers independently score the same three calls, then compare results line by line. Rows where scores diverge by more than one band get rewritten with clearer evidence anchors. The goal is shrinking inter-reviewer variance over time, not declaring a winner.
Does this scorecard work for Arabic-English code-switched calls?▾
Yes, provided the transcription engine handles code-switching. GCC customer calls routinely mix Arabic and English mid-sentence; CallScribe detects code-switching and transcribes both, so disclosure detection and sentiment tracking keep working. An Arabic-only or English-only model will drop half the conversation.
Test it on your own Arabic calls
Dialect-aware transcription with diarization and sentiment — built for GCC call centers.
Try CallScribe free →5 min/mo free · No credit card