Skip to Content
Internal docs are powered by Nextra Docs Theme.
AnalysisQualityConversation QA calibration — April 2026

Conversation QA Calibration Update - April 2026

Why this write-up

Over the last few weeks we have been trying to answer a practical question: how close can our QA evaluator get to human calibration on real Qonto conversations, and where are the remaining gaps coming from?

We now have three useful sources of signal:

  1. A March retrospective of human reviewer feedback on production AI evaluations.
  2. A sandbox-style eval report comparing AI vs human on a hand-built gold case.
  3. A newer conversation-level QA harness that runs the full tool-based evaluator against reproducible case folders.

The short version is:

  • The system is already directionally useful.
  • Overall score alignment is better than raw check-by-check alignment.
  • The biggest remaining misses are not random; they cluster around a small set of hard judgment problems.
  • Some of the latest “disagreements” are true calibration misses, but some are harness artifacts such as check-order or mapping issues.

The headline results so far

1. Production feedback says the baseline is usable, but still not calibration-grade

From the March Qonto feedback analysis:

  • 56 conversations were reviewed.
  • 203 AI criterion comments got thumbs up and 63 got thumbs down.
  • That is a 76% approval rate.
  • Average absolute score delta vs human was 8.3 points per conversation.
  • 43 of 56 conversations (76.8%) were within 10 points of the human score.

That is a meaningful baseline: the model is often in the right neighborhood. But it is not yet reliable enough to treat as a drop-in human replacement, especially on the highest-context criteria.

The negative feedback also concentrated heavily in a few areas:

  • Be Efficient for the Client: 27% of negative comments
  • Solve the Request: 27%
  • Understand the Request: 24%
  • Qonto's Tone of Voice: 13%

Those are exactly the criteria where humans are using nuanced judgment about context, sequencing, customer intent, and language quality.

2. The old sandbox eval showed that “more reasoning” did not automatically help

In the sandbox accuracy report on one hand-built case:

  • Low effort: 54.5% exact check match
  • Medium effort: 45.5% exact check match
  • Medium effort improved some justification matching, but not outcome accuracy
  • Medium effort also cost materially more and took materially longer

This was an important result because it showed that deeper reasoning by itself does not solve calibration. If the instructions are off, or the model is anchoring on the wrong cues, giving it more room to think can just make it more confidently wrong.

3. The new conversation-level harness is encouraging on score alignment

Using the newer conversation-level runner, we have now run at least two realistic case folders:

  • conversation_Z8oVG14X587UVJwQYLMy7K6A: 81.8% raw exact check similarity, 95% score similarity
  • conversation_0rpo9ZqB41wUQ7OpRyD4KaNM: 72.7% raw exact check similarity, 95% score similarity

Across those two runs, the raw check-level floor is 17/22 exact matches (77.3%), while score-level similarity was 95% in both cases.

That is promising. It suggests the evaluator is often landing in the right overall scoring band even when individual check outputs still diverge.

It is also worth noting that the 72.7% run appears understated: manual review of conversation_0rpo9ZqB41wUQ7OpRyD4KaNM suggests two of the three reported tone disagreements were likely check-order mismatches rather than real scoring disagreements. In other words, the harness currently under-reports alignment in at least some cases.

Deep read on the full-QA cases

The most useful thing about the new setup is that we can now study full cases end-to-end: transcript, internal notes, human review, AI check outputs, and run logs. That makes it possible to say not just “the score was close” but “what exactly the model understood, what it missed, and what the new architecture enabled.”

Case 1: conversation_Z8oVG14X587UVJwQYLMy7K6A

This is the stronger of the two current full-QA cases for the new system.

What happened in the conversation

The customer was asking about two cheque deposits on an account that was suspended and debit-blocked during an ongoing closure/fraud process. The customer explicitly signaled business urgency: they could not afford a long delay because the company’s financial situation was fragile. The ticket had already passed through several internal handoffs across frontline, fraud, and back-office teams. Fatma’s customer-facing message was the final operational synthesis:

  • one cheque would be credited within 24 hours
  • the second could only be unblocked after 28 March
  • the delay was tied to the account state
  • the tone tried to be empathetic and reassuring

This is a good stress test for the evaluator because it requires:

  • reading the whole internal workflow, not just the last reply
  • separating operational correctness from tone quality
  • judging whether a clean final answer is enough, or whether the urgency required more probing or confirmation

Where AI and human aligned

This run landed at 81.8% exact check similarity and 95% score similarity.

The AI and human agreed on the most important operational parts:

  • Understand the Request: pass
  • Solve the Request: pass
  • both Be Efficient for the Client checks: pass
  • empathy/respect/ownership-oriented tone checks beyond the stylistic ones: pass
  • Call Recording Consent: not applicable

This is important because it shows the new full-QA evaluator is not getting lost in the internal noise. It correctly recognized that the real task was a timing-and-crediting update, not a generic “account closure” case, and that Fatma’s final message did meaningfully resolve the customer’s immediate informational need.

Where they diverged

The disagreements were concentrated in tone:

  • The human failed one language-quality tone check where the AI gave partial.
  • The human passed one expression/naturalness check where the AI gave partial.

In practice, the AI read Fatma’s phrasing as slightly too ceremonial and slightly awkward:

  • “J’espère que vous allez bien”
  • “je vous remercie sincèrement pour votre patience”
  • “soyez rassuré que toutes les démarches nécessaires ont été effectuées”

That is actually a fairly sophisticated critique. The model is not missing the message or the procedure; it is making a stylistic judgment about register and naturalness. The fact that these are the only real disagreements in this case is a positive sign. It means the new prompt and architecture are getting the evaluator to spend its disagreement budget on something genuinely subtle rather than on basic procedural confusion.

What this case says about the new system

This case is the clearest evidence that the new full-QA approach is strong on:

  • whole-conversation context tracking
  • distinguishing internal handoff history from final agent accountability
  • judging operational correctness separately from tone
  • landing near the human on overall score even when stylistic calibration is still a bit stricter than the reviewer

Case 2: conversation_0rpo9ZqB41wUQ7OpRyD4KaNM

This is the messier and more revealing case.

What happened in the conversation

The customer wanted to receive a physical card. Marisol replied in chat after a delay, initially gave generic ordering steps, later answered a follow-up about whether having a virtual card was a problem, shared a help link, reacted to the customer saying the link did not work, then eventually confirmed that the card had already been successfully requested and provided the delivery window and address.

This is a deceptively simple case, but it is actually hard to score well because several things are true at once:

  • the customer’s core request is straightforward
  • the agent is polite and eventually gives the key answer
  • the communication is long, awkward, and error-prone
  • there is visible chat-response delay
  • some replies are out of sync with the customer’s latest state

That combination is exactly where human QA reviewers often use medium-severity judgments instead of binary “good” or “bad.”

Where AI and human aligned

Even though the run reported only 72.7% exact check similarity, the AI still landed at 95% score similarity and got the broad shape of the case right:

  • Solve the Request: pass
  • both Be Efficient for the Client checks: one partial, one pass
  • Create a Human Bond: fail
  • Call Recording Consent: not applicable
  • overall score only 5 points away from human

That is a meaningful success. The AI correctly understood that this was not a catastrophic handling failure. It was a middling-quality interaction: broad resolution present, but diluted by clarity, structure, and communication-quality issues.

The real disagreement

The main substantive difference appears to be Understand the Request.

The human marked it as pass, effectively reading Marisol as having understood the customer’s intent despite messy phrasing. The AI marked it as partial, arguing that Marisol moved into generic instructions instead of first clarifying or checking whether the card had already been ordered.

This is a real calibration question, not a tooling artifact. It exposes a persistent human-vs-model gap:

  • humans often grade intent understanding more charitably when the agent eventually gets to the right answer
  • the AI is more likely to penalize the first-turn path if it feels generic or insufficiently contextualized

The likely fake disagreements

Two tone mismatches in this run look much less like true disagreement and much more like a harness ordering issue.

The AI failed one tone check on visible language errors and passed another on empathy/context fit. The human appears to have made the same broad judgments but attached them to different tone check positions. Semantically, these look closer than the raw exact-match metric suggests.

If that diagnosis is right, then this case is actually evidence for two things at once:

  • the model is still imperfect on full-context interpretation
  • the harness is currently overstating disagreement in at least some multi-check criteria

What this case says about the new system

This case is valuable because it shows the full-QA evaluator can still produce a near-human total score on a conversation that is communicatively messy, even when some per-check alignment is noisy. It also demonstrates why the new harness matters: without transcript + human review + exact AI check summaries together, we would not be able to tell the difference between a true modeling miss and a comparison artifact.

What the new prompt is doing better

The recent prompt changes are not just cosmetic. They are starting to change what the model pays attention to.

1. It anchors the evaluator on the customer’s real request

The new prompt now opens by forcing the evaluator to identify the customer’s primary issue, secondary concerns, and actual resolution path before scoring.

That matters because older failures often came from evaluating the agent’s framing instead of the customer’s problem. In the new full-QA cases, the AI mostly stays anchored on the customer’s actual need:

  • in the cheque case, it scored Fatma against the real “what happens to the two deposits?” question
  • in the card case, it correctly recognized that the key issue was getting/confirming the physical card flow, not just narrating generic product info

2. It is stricter in the right places

Removing leniency language and replacing it with “accurate, not lenient” and “apply scorecard standards as written” appears to be having a real effect.

The model is now more willing to:

  • penalize robotic or awkward phrasing
  • penalize generic or poorly targeted first responses
  • treat fragmentation and style problems as genuine quality issues when the scorecard says they are

That does not mean it is always calibrated correctly yet. But it does mean it is now arguing on the right terrain.

3. It reasons about tone more concretely than before

The prompt now explicitly tells the evaluator to look for:

  • genuine vs scripted empathy
  • robotic or template-stitched voice
  • inconsistent voice across messages
  • context-dependent expectations for emotional acknowledgement

You can see that in the AI summaries themselves. In both cases, the model was no longer giving vague “professional and polite” filler. It was calling out specific phrases and explaining why they felt unnatural, too ceremonial, or mismatched to the customer context. That is a meaningful prompt-quality improvement even where human calibration is still debatable.

4. It can now use human precedent directly

The addition of ask_feedback_store is one of the most important architectural upgrades. The evaluator is no longer limited to the scorecard text and KB alone; it can retrieve prior human-corrected examples as calibration precedent.

That matters especially for:

  • empathy
  • tone
  • expression
  • edge-case interpretations of the scorecard

We do not yet have a large enough set of full-QA runs to prove how much this improved accuracy quantitatively, but architecturally it is exactly the missing ingredient for closing the human-calibration gap on subjective criteria.

What the new architecture is doing better

The prompt improvements matter, but the architecture change is arguably even more important.

1. It moved us from “snapshot judgment” to actual full-QA evaluation

The older path often produced a CX-risk-style read of the conversation without generating a full scorecarded QA evaluation. In both current full-QA cases, the human review explicitly corrected that limitation: the old system could say useful things about churn or complaint risk, but it was not actually doing the same job as a human QA reviewer.

The new runner closes that gap by evaluating:

  • full conversations
  • all relevant agents
  • all leaf criteria
  • with explicit check-level outputs

That is a much closer match to the human task we are trying to automate.

2. It enforces a research-then-score workflow

The tool loop is not just a transport mechanism. It induces a better evaluation behavior:

  • research with ask_knowledge_base
  • research with ask_feedback_store
  • then save criterion results incrementally with save_agent_criterion_result

In the Marisol run, the first tool-loop iteration executed ten research calls before saving began, and then the evaluator saved criteria one by one until all six leaf criteria were complete. That is exactly the kind of disciplined workflow we want: gather policy and calibration context first, then score.

3. It saves progress incrementally

This is a major reliability improvement.

Each (agent, scorecard, leaf criterion) unit is saved independently through save_agent_criterion_result, and the underlying scorecard result persists intermediate leaf-level outputs. Combined with job retries, that means the system can recover from failures without losing all prior work.

This matters for real production QA because long, tool-based evaluations are inherently more failure-prone than single-shot generations. Incremental persistence turns a fragile batch job into something much more resumable and inspectable.

4. It recreates the environment per case

The eval harness builds an isolated organization, uploads the case’s KB and feedback docs, remaps referenced document IDs, constructs the scorecard from fixture data, and enables the conversation-runner feature flag for that org. That means each case is reproducible and self-contained.

This is a big research advantage because it lets us answer:

  • did the model fail because of the prompt?
  • because of missing KB context?
  • because of scorecard structure?
  • because of harness comparison logic?

without those concerns being mixed together with mutable production state.

5. It gives us much better observability

The new setup exposes:

  • per-check exact mismatches
  • per-agent score deltas
  • exact AI summaries
  • exact human review outputs
  • tool-loop logs showing research and save behavior

This is what allows us to write memos like this one. More importantly, it is what allows us to convert anecdotal “the AI feels off on tone” complaints into concrete, reproducible, fixable failure classes.

Where the system is already strong

1. It is usually directionally right on the overall quality level

Even when the exact reason differs from the human reviewer, the model often lands near the same overall score. That showed up both in March score deltas and in the newer 95% score-similarity runs.

This matters because a useful QA assistant does not need perfect verbal mimicry of a reviewer to be helpful. It does need to separate clearly good work from clearly weak work and keep medium cases roughly in band. So far, it seems better at that than raw check-match rates alone suggest.

2. It performs better on explicit, procedural, or rule-shaped judgments

We have seen clear gains when the task can be anchored to a crisp rule:

  • Voicemail is not a live call, so recording-consent checks are not applicable.
  • KB-grounded procedure verification improves when the evaluator is forced to query from the customer’s situation rather than the agent’s framing.
  • Wrong-topic handling improves when the prompt explicitly says “working on the wrong problem gets no partial credit.”
  • Dependency rules improve when we encode them more literally instead of hoping the model infers them.

This is good news because it means some important failure modes are fixable with better rule encoding, targeted eval cases, and prompt constraints rather than requiring a fundamentally different system.

3. The new harness is much better for debugging than the old setup

The new folder-based conversation harness is a real step forward operationally:

  • Each case recreates an isolated org and uploads the exact KB and feedback docs needed for reproduction.
  • The runner evaluates at the (agent, scorecard, leaf criterion) level.
  • Intermediate results are saved, so retries can resume remaining work instead of starting from scratch.
  • We can now inspect exact AI summaries, exact human results, and exact per-check disagreements on a single reproducible fixture.

That has made it much easier to tell the difference between “the model judged this badly” and “the harness compared the wrong thing.”

Where the AI is still weaker than humans

1. Customer-intent reading is still the biggest real gap

This remains the core issue from the March analysis, and it still shows up in the new harness.

Humans are better at recognizing what the customer actually means, especially when the wording is messy, abbreviated, or indirect. They also do a better job deciding whether an agent’s response was “good enough for this situation” rather than mechanically incomplete.

Example: in conversation_0rpo9ZqB41wUQ7OpRyD4KaNM, the main substantive disagreement was Understand the Request. The AI marked Marisol as partial, arguing she moved too quickly into generic ordering guidance. The human marked passed, interpreting the customer’s request more charitably and viewing the agent as having understood the basic intent. That is a real calibration gap, not just a mapping bug.

2. Tone, empathy, and language-quality judgments remain unstable

This is the second major weakness.

The evaluator is still not reliably aligned with human reviewers on:

  • empathy
  • playful but polished tone
  • naturalness vs translated/formulaic phrasing
  • how harshly to score visible grammar/writing issues

Sometimes the model is too lenient and reads surface politeness as empathy. Sometimes it is too strict and penalizes phrasing humans would let pass. In French-language support especially, human reviewers seem more sensitive to register, emotional fit, and whether a sentence sounds copied, translated, or robotic.

This is why tone is still one of the hardest criteria, even after prompt work.

3. The AI still drifts across criterion boundaries

A recurring human complaint is that the evaluator uses the wrong kind of argument for the criterion it is scoring.

Examples we have seen:

  • efficiency arguments leaking into tone
  • tone arguments leaking into empathy
  • consent or compliance-like behavior affecting efficiency
  • solve-the-request logic bleeding into understand-the-request

Some of this is genuinely about model reasoning. Some of it may also reflect scorecards whose criteria are semantically adjacent. But either way, humans are still better at keeping the rationale scoped to the criterion being judged.

4. It is still too literal in some multi-step or multi-agent contexts

The prompt tells the evaluator to use full conversation context, and sometimes it does that too literally.

Humans often evaluate an agent’s contribution more charitably:

  • Did this agent move the case forward?
  • Did they add new value?
  • Were they responsible for the later failure, or had their handling already effectively ended?

The AI is more likely to say, “this later part of the conversation was still unresolved, therefore partial/fail,” even when a human reviewer would isolate the agent’s own contribution.

Where the harness is weaker than humans

Not every disagreement is a model-quality disagreement. We have also uncovered harness issues that make the AI look worse than it really is.

1. Check-order and mapping bugs can create fake disagreements

The clearest recent example is conversation_0rpo9ZqB41wUQ7OpRyD4KaNM, where two tone disagreements appear to be semantic matches attached to the wrong check indices.

If that is confirmed, then the harness is currently understating exact agreement and inflating mismatch counts in at least some cases.

2. Token budget and tool-loop limits can distort outcomes

We already hit one failure mode where the run collapsed to 0% similarity because the model did not return all criterion results. That turned out to be a harness/config issue rather than a genuine calibration catastrophe.

We also saw:

  • high reasoning increasing latency enough to risk timeouts
  • medium reasoning still needing a larger output budget than the original default
  • occasional save-tool fragility (Unknown agent/scorecard/leaf_criterion pair) even when the overall run recovered

So the harness itself still needs hardening before every mismatch can be treated as a judgment error.

3. Raw exact-match metrics can be misleading

Exact check match is useful, but it is a harsh metric. It punishes:

  • minor calibration differences between passed and partial
  • index-order mismatches
  • passed vs not_applicable edge handling in some criteria
  • cases where the model has the right overall read but chooses a different local explanation

That is why score similarity and manual disagreement review are both important. The raw match number alone is currently too noisy to be our only headline KPI.

What we have improved already

A lot of the recent work has been productive even if the model is not fully calibrated yet.

We have already shipped or validated fixes around:

  • dependency-rule application
  • voicemail feasibility for call-consent checks
  • customer-centric KB querying
  • not inheriting the agent’s categorization when checking procedure correctness
  • stricter severity calibration based on explicit justification options
  • wrong-topic handling
  • removal of leniency language from the system prompt
  • higher output budgets for the conversation-level tool loop

This is important because the failure set is getting narrower. We are not staring at a fully opaque “LLM quality” problem anymore. We are progressively converting broad failure classes into narrower, testable, reproducible cases.

Current read: strengths vs weaknesses against human calibration

If we had to summarize honestly today:

Strengths

  • Usually lands in the right overall scoring band.
  • Good potential on procedure-heavy and rule-heavy criteria.
  • Much more reproducible and debuggable than before.
  • Responsive to targeted prompt and rule changes when the failure mode is explicit.

Weaknesses

  • Still weaker than humans at reconstructing customer intent from imperfect language.
  • Still weaker than humans on subtle tone, empathy, and “does this sound genuinely human?” judgments.
  • Still prone to criterion-scope bleed and over-literal use of full conversation context.
  • Still affected by harness artifacts that make some raw disagreement numbers look worse than the real underlying alignment.

What I think this means

The project is in a better state than the scariest raw disagreement numbers imply, but not yet in a state where we should claim human-level QA calibration.

The encouraging part is that we now have evidence for both of the following:

  • The evaluator can get surprisingly close to humans on overall score.
  • Many of the worst misses are clustered, legible, and testable.

The remaining challenge is that the hardest residual gaps are the most human ones: intent reading, conversational nuance, empathy, tone fit, and deciding when “technically incomplete” still deserves a pass in context.

That suggests the next gains will not come from simply increasing reasoning effort. They will come from:

  • better scorecard-to-prompt translation
  • stronger criterion-specific calibration rules
  • more gold cases around tone and efficiency fragmentation
  • fixing harness mapping issues so we can trust the eval numbers
  1. Treat the current conversation-level harness as the primary calibration surface and expand the case set.
  2. Fix the likely tone check-order mismatch before trusting raw check similarity as a headline metric.
  3. Add more adversarial cases for the three hardest areas: Understand the Request, Be Efficient for the Client, and Qonto's Tone of Voice.
  4. Keep using score similarity alongside exact check match, because score-level alignment is currently a more stable signal.
  5. Continue separating “real model miss” from “harness artifact” during review; those are different problems and need different fixes.

Bottom line

So far, the evaluator looks promising as a QA copilot, not yet as a human-equivalent reviewer.

Its biggest strength is that it is often directionally correct and increasingly fixable. Its biggest weakness is that the remaining errors are concentrated in exactly the places where humans rely on judgment, nuance, and contextual generosity.

That is a good place to be for iteration. It means we are no longer debugging a black box. We are calibrating a system.

Last updated on