Skip to Content
Internal docs are powered by Nextra Docs Theme.
CustomersQontoMoshi handoff calibration

Calibration of Moshi’s production handoff screen

Bottom line

The original 60.5% should not be described as the share of handoffs that are potentially addressable. It is the share assigned a non-none label by an automated, forced-choice screen.

A blinded, stratified review of 400 production handoffs produced a lower and better-defined result:

FindingEstimated shareEstimated handoffs95% sampling interval
Transcript-visible capability or guidance gap55.6%28,33751.1–60.1% · 26,052–30,622
No visible gap38.0%19,38833.3–42.7% · 16,995–21,781
Insufficient visible evidence6.4%3,2673.8–9.0% · 1,944–4,589

Reviewers judged that a product change could plausibly improve 50.1% of handoffs, or about 25,558 conversations. The 95% sampling interval is 45.6–54.6%, or 23,272–27,844 conversations. A further 11.9% had uncertain product addressability.

“Plausibly improve” is deliberately narrower than the original screen but still much weaker than “preventable.” It does not establish Qonto policy, permissions, technical feasibility, safe automation, or the number of manual tickets a release would avoid.

It is also not an independent second ground truth. Under the protocol, a case with no visible gap is necessarily not addressable, while visible gaps were resolved as either plausibly addressable or uncertain. The 50.1% result should therefore be read as a stricter triage estimate, not proof of build feasibility.

What the result means

There is a large, measurable improvement opportunity in Moshi’s human handoffs. The automated screen was directionally useful, but its published interpretation was too strong:

  • It overstates the broad visible-gap estimate by 4.9 percentage points and the plausibly-improvable estimate by 10.4 points.
  • Of the 30,862 handoffs flagged by the screen, an estimated 77.1% contain a visible gap and 71.4% look plausibly improvable.
  • The screen also misses material signal. Of the 20,129 handoffs labeled none, an estimated 22.5% contain a visible gap and 17.5% look plausibly improvable.
  • none therefore cannot be treated as an exclusion rule, and a non-none label cannot be treated as an addressability decision.

The screen is suitable for candidate discovery and queue ordering. It is not yet suitable for publishing category totals or calculating savings without a calibrated second pass.

The product implication for Rulebase

The data supports the broader improvement-loop thesis more strongly than a KB-only self-healing thesis.

Calibrated primary gapEstimated share of all handoffsEstimated handoffs95% sampling interval
Missing customer-specific data28.1%14,30923.8–32.3%
Missing tool or action13.7%6,97910.9–16.5%
Incorrect or stale source state4.4%2,2602.6–6.3%
Knowledge or guidance mismatch9.4%4,7886.9–11.8%

These subtype estimates are directional because subtype agreement was weaker than broad gap agreement. Even with that caveat, the mix is important: data access, actions, and source-state reliability represent about 83% of the estimated visible gaps. Plain knowledge or guidance mismatches represent about 17%.

That means automatically bootstrapping KB content can be a useful capability, but it is not the whole product and is unlikely to be the main route to materially higher AI handling on its own. The larger opportunity is to identify the customer-specific reads, actions, state contracts, and decision rules that the bot lacks, turn them into release proposals, and measure what happens after Qonto or Moshi ships them.

This also exposes the business constraint. Rulebase can discover and quantify the backlog from Zendesk history without controlling Moshi. Rulebase cannot complete a genuinely self-healing loop for most of the estimated opportunity unless the bot owner implements the data, tool, or state change and exposes enough version and outcome telemetry to measure it. The immediate sellable output is therefore a validated capability backlog and release-measurement loop, not autonomous end-to-end remediation.

The commercial risk is execution after diagnosis. Without committed engineering capacity from the bot owner, the backlog becomes an intelligent report rather than improved AI handling. The strongest initial customers are teams that own their agent stack or can reliably ship changes through their bot vendor.

Screen precision and missed gaps

Screen-positive handoffs

Among the 30,862 handoffs assigned a non-none label:

Review resultEstimated rateEstimated handoffs95% sampling interval
Visible gap confirmed77.1%23,80871.6–82.7%
No visible gap20.4%6,30415.0–25.8%
Evidence uncertain2.4%7500.6–4.3%
Plausibly improvable71.4%22,03665.5–77.3%

Handoffs screened as none

Among the 20,129 handoffs assigned none:

Review resultEstimated rateEstimated handoffs95% sampling interval
Visible gap found22.5%4,52915.0–30.0%
No visible gap confirmed65.0%13,08456.5–73.5%
Evidence uncertain12.5%2,5166.6–18.4%
Plausibly improvable17.5%3,52310.7–24.3%

The practical fix requires more than raising or lowering the screen’s threshold. The classifier needs a calibrated broad-gap stage, an explicit uncertainty outcome, and a separate subtype stage. Customer-requested and required handoffs should be removed before capability inference. Candidate clusters then need their own precision and policy review.

Reliability of the four gap types

The broad decision was materially more reliable than the subtype decision.

Review fieldUnanimous on 46 triple-reviewed casesPairwise agreementFleiss’ κ
Visible gap / no gap / uncertain78.3%85.5%0.668
Exact primary type56.5%69.6%0.615
Plausible addressability63.0%74.6%0.508

The visible-gap field is the hierarchical broad form of the primary label, not an independent second judgment. Its stronger agreement shows that reviewers could identify the presence of a gap more consistently than its exact technical subtype.

The existing screen’s exact-label confirmation rates were:

Automated screen labelSame type after review95% sampling interval
Missing tool or action88.3%80.2–96.5%
Missing data point55.8%46.9–64.7%
Knowledge mismatch52.9%41.2–64.6%
Incorrect or stale tool data53.3%35.4–71.2%
No gap65.0%56.5–73.5%

The missing-tool label is strong enough to support targeted cluster work. The other three labels should be treated as investigation hypotheses. Data, knowledge, and source-state problems often coexist in the same conversation, and the visible human response does not always isolate the technical root cause.

Study design

  • Fixed population: 50,991 unique Moshi-to-human handoffs in Qonto EU, from conversations started at or after 25 June 2026 13:52:06 UTC and before 24 August 2026 13:52:06 UTC. Only summaries created by the end of that window were included.
  • Sampling: 400 handoffs selected deterministically at random within the five automated screen strata: 120 missing-data, 120 none, 70 knowledge, 60 missing-tool, and 30 wrong-data assignments. Results were reweighted to the fixed production population.
  • Blinding: reviewers did not see the automated gap label, generated title, rationale, or another reviewer’s decision.
  • Evidence: the calibration added customer turns that the original screen omitted and showed up to 20 turns before and after the handoff. One unresolved tie required a checksum-verified full-thread review because material evidence appeared beyond the 20-turn window.
  • Causality rule: a later human lookup or action was not labeled as the cause of a handoff when the customer had already requested a person before disclosing the issue.
  • Overlap: 46 items received three independent reviews. A hierarchical 2-of-3 majority resolved the broad gap decision before the subtype. Only three items lacked a material majority and received a fresh blind tie-break review.
  • Validation: all 400 sample rows resolved; all 46 overlap rows were triple-reviewed; the three required tie-breaks were used; population, sample, assignment, schema, consistency, duplicate, missing, and orphan-adjudication checks all passed.

The reviewers were independent model-assisted analyst passes, not Qonto subject-matter experts or human ground truth. The confidence intervals quantify sampling error under the stratified design; they do not include correlated reviewer error, omitted back-office actions, or policy uncertainty.

What remains unproven

This calibration does not establish:

  • that a human response reflects authoritative Qonto policy;
  • that Moshi or Qonto can technically expose the required data or action;
  • that a handoff was safe or desirable to prevent;
  • that the visible gap caused the full conversation outcome;
  • that the four-way subtype is the underlying engineering root cause;
  • that a proposed release will reduce manual work; or
  • that the sampled window represents future traffic and bot versions.

For a specific capability, the defensible sequence is:

candidate cluster → cluster-specific precision review → policy and eligibility definition → implementation → controlled release → observed reduction in human handling with resolution and safety guardrails

The impact number should come from the controlled release. The broad calibration establishes that the backlog has real signal; it does not substitute for that experiment.

Last updated on