Evidence & confidence

A score you can
argue with

Every QA tool says it scores with confidence. This one prints the number, and the lines of transcript it came from.

call-4471.mp3Audited
68overall · fatal breach on disclosure
Opening & verification100

“Good morning, this is Priya from [MASKED]. May I confirm the last four digits of your account?”

confidence96%
Recording disclosure0

Mandatory disclosure was never read out before verification began.

confidence91%
Empathy & resolution72

“I understand this is frustrating, let me see what I can do.” Acknowledged once, then moved on.

confidence64%
0
competitors publishing a confidence figure on every parameter. We print one on all of them.
01
The problem

An unevidenced score is just an assertion

When a QA tool tells an agent they scored 62 on empathy, there is nothing to discuss. The agent cannot check it, the supervisor cannot defend it beyond “the system said so”, and the conversation becomes about the tool rather than about the call.

Every parameter here returns the transcript spans it was judged on. Not a summary of them, not a paraphrase: the lines themselves, with their position in the transcript, shown next to the score rather than buried behind an export.

The disagreement then has somewhere to go. Either the quoted line supports the score or it does not, and both people are looking at the same sentence.

auditResults[].evidence[] · call-audit.schema.ts

02
Confidence

The number that tells you where to look

Some questions have clean answers. A recording disclosure either was or was not read out; a model should be near-certain either way, and it usually reports somewhere in the nineties.

Others are judgement calls. Whether an agent showed genuine empathy or merely said an empathetic sentence is a reading of the call, and a low confidence figure there is the honest answer, not a defect.

That makes confidence a routing signal. The calls the model hedged on are the calls worth a human ear, which is a far better use of your analysts than re-reviewing a random sample of the ones it was sure about.

The call carries an overall figure too: the mean of its parameters.

auditResults[].confidence · 0–100, per parameter

confidence across 3,898 scored parameters
306090340 below 60

The tail is the product. Those are the calls the model hedged on, and they are the ones worth an analyst's time.

Confidence distribution across scored parameters
Confidence bandParameters
30–3412
35–3921
40–4438
45–4961
50–5484
55–59124
60–64171
65–69238
70–74331
75–79432
80–84528
85–89611
90–94664
95–99583
confidence · by parameter
Recording disclosure96%

said or not said

Correct disposition88%

checkable against the record

Empathy51%

a judgement call

Tone appropriate to contextn/a

not reported

What confidence is not

It is a self-report, not a guarantee

A confidence figure is the model's own account of how clear the evidence was. It is not a measured error rate, it is not validated against a human ground truth you have not given us, and a model can be confidently wrong. Any vendor telling you otherwise is selling you a number they cannot support.

What it is good for is triage and disagreement. High confidence plus a quoted disclosure that plainly is not there means the call failed. Low confidence on a subjective parameter means read the call yourself before you coach anyone on it.

Where the model reports no confidence at all, the interface shows a dash. It would be trivial to substitute a plausible default and nobody would notice, which is exactly why we do not. The whole value of printing the number is that it means something.

For each parameter, provide a confidence score (0-100) reflecting how clearly the transcript supports your judgement.
ai.service.ts: the instruction in the scoring prompt
03
Both sides

Sentiment for the agent, not only the customer

A single sentiment score for a whole call throws away most of what makes it worth knowing. A customer who starts angry and ends calm is a call that went well. A customer who stays pleasant while the agent grows terse is a different problem entirely, and one overall figure hides both.

Customer and agent are scored separately, and the emotional turning points are timestamped, so you can jump to the moment rather than re-listening to eight minutes to find it.

sentiment.customerScore / .agentScore · emotionalMoments[]

sentiment · both speakers
Customer
Agent
02:14 · frustration (customer)“this is the third time”
Limits

What this does and does not do

Every vendor in this category publishes a table of green ticks. Here is one with the gaps in it.

CriterionQXAI
Quoted transcript evidence per parameterYesYesShown beside the score, not in an export.
Confidence figure per parameterYesYes0–100, plus a mean for the call.
Customer and agent sentiment, separatelyYesYesWith timestamped turning points.
Validated accuracy percentageNot claimedNot claimedWe have no benchmark you can audit, so we publish no number. Anyone quoting one should be asked how it was measured.
Word-level timestampsNoNoEvidence cites lines of transcript, not exact audio offsets.
Acoustic speaker diarisationContext-basedContext-basedSpeakers are inferred while transcribing. Reliable on a two-party call, less so with heavy overlap.
Calibration workflowNoNoAutomated scoring removes the reviewer drift calibration exists to correct, but if you run formal sessions they stay outside the product.
Real-time assistance during a live callNoNoAuditing is post-call. There is no agent-assist.

See it on a real call

Follow one call from upload to a disputed score, with every stage shown.