Calibration of Moshi’s production handoff screen
Bottom line
The original 60.5% should not be described as the share of handoffs that are potentially addressable. It is the share assigned a non-none label by an automated, forced-choice screen.
A blinded, stratified review of 400 production handoffs produced a lower and better-defined result:
| Finding | Estimated share | Estimated handoffs | 95% sampling interval |
|---|---|---|---|
| Transcript-visible capability or guidance gap | 55.6% | 28,337 | 51.1–60.1% · 26,052–30,622 |
| No visible gap | 38.0% | 19,388 | 33.3–42.7% · 16,995–21,781 |
| Insufficient visible evidence | 6.4% | 3,267 | 3.8–9.0% · 1,944–4,589 |
Reviewers judged that a product change could plausibly improve 50.1% of handoffs, or about 25,558 conversations. The 95% sampling interval is 45.6–54.6%, or 23,272–27,844 conversations. A further 11.9% had uncertain product addressability.
“Plausibly improve” is deliberately narrower than the original screen but still much weaker than “preventable.” It does not establish Qonto policy, permissions, technical feasibility, safe automation, or the number of manual tickets a release would avoid.
It is also not an independent second ground truth. Under the protocol, a case with no visible gap is necessarily not addressable, while visible gaps were resolved as either plausibly addressable or uncertain. The 50.1% result should therefore be read as a stricter triage estimate, not proof of build feasibility.
What the result means
There is a large, measurable improvement opportunity in Moshi’s human handoffs. The automated screen was directionally useful, but its published interpretation was too strong:
- It overstates the broad visible-gap estimate by 4.9 percentage points and the plausibly-improvable estimate by 10.4 points.
- Of the 30,862 handoffs flagged by the screen, an estimated 77.1% contain a visible gap and 71.4% look plausibly improvable.
- The screen also misses material signal. Of the 20,129 handoffs labeled
none, an estimated 22.5% contain a visible gap and 17.5% look plausibly improvable. nonetherefore cannot be treated as an exclusion rule, and a non-nonelabel cannot be treated as an addressability decision.
The screen is suitable for candidate discovery and queue ordering. It is not yet suitable for publishing category totals or calculating savings without a calibrated second pass.
The product implication for Rulebase
The data supports the broader improvement-loop thesis more strongly than a KB-only self-healing thesis.
| Calibrated primary gap | Estimated share of all handoffs | Estimated handoffs | 95% sampling interval |
|---|---|---|---|
| Missing customer-specific data | 28.1% | 14,309 | 23.8–32.3% |
| Missing tool or action | 13.7% | 6,979 | 10.9–16.5% |
| Incorrect or stale source state | 4.4% | 2,260 | 2.6–6.3% |
| Knowledge or guidance mismatch | 9.4% | 4,788 | 6.9–11.8% |
These subtype estimates are directional because subtype agreement was weaker than broad gap agreement. Even with that caveat, the mix is important: data access, actions, and source-state reliability represent about 83% of the estimated visible gaps. Plain knowledge or guidance mismatches represent about 17%.
That means automatically bootstrapping KB content can be a useful capability, but it is not the whole product and is unlikely to be the main route to materially higher AI handling on its own. The larger opportunity is to identify the customer-specific reads, actions, state contracts, and decision rules that the bot lacks, turn them into release proposals, and measure what happens after Qonto or Moshi ships them.
This also exposes the business constraint. Rulebase can discover and quantify the backlog from Zendesk history without controlling Moshi. Rulebase cannot complete a genuinely self-healing loop for most of the estimated opportunity unless the bot owner implements the data, tool, or state change and exposes enough version and outcome telemetry to measure it. The immediate sellable output is therefore a validated capability backlog and release-measurement loop, not autonomous end-to-end remediation.
The commercial risk is execution after diagnosis. Without committed engineering capacity from the bot owner, the backlog becomes an intelligent report rather than improved AI handling. The strongest initial customers are teams that own their agent stack or can reliably ship changes through their bot vendor.
Screen precision and missed gaps
Screen-positive handoffs
Among the 30,862 handoffs assigned a non-none label:
| Review result | Estimated rate | Estimated handoffs | 95% sampling interval |
|---|---|---|---|
| Visible gap confirmed | 77.1% | 23,808 | 71.6–82.7% |
| No visible gap | 20.4% | 6,304 | 15.0–25.8% |
| Evidence uncertain | 2.4% | 750 | 0.6–4.3% |
| Plausibly improvable | 71.4% | 22,036 | 65.5–77.3% |
Handoffs screened as none
Among the 20,129 handoffs assigned none:
| Review result | Estimated rate | Estimated handoffs | 95% sampling interval |
|---|---|---|---|
| Visible gap found | 22.5% | 4,529 | 15.0–30.0% |
| No visible gap confirmed | 65.0% | 13,084 | 56.5–73.5% |
| Evidence uncertain | 12.5% | 2,516 | 6.6–18.4% |
| Plausibly improvable | 17.5% | 3,523 | 10.7–24.3% |
The practical fix requires more than raising or lowering the screen’s threshold. The classifier needs a calibrated broad-gap stage, an explicit uncertainty outcome, and a separate subtype stage. Customer-requested and required handoffs should be removed before capability inference. Candidate clusters then need their own precision and policy review.
Reliability of the four gap types
The broad decision was materially more reliable than the subtype decision.
| Review field | Unanimous on 46 triple-reviewed cases | Pairwise agreement | Fleiss’ κ |
|---|---|---|---|
| Visible gap / no gap / uncertain | 78.3% | 85.5% | 0.668 |
| Exact primary type | 56.5% | 69.6% | 0.615 |
| Plausible addressability | 63.0% | 74.6% | 0.508 |
The visible-gap field is the hierarchical broad form of the primary label, not an independent second judgment. Its stronger agreement shows that reviewers could identify the presence of a gap more consistently than its exact technical subtype.
The existing screen’s exact-label confirmation rates were:
| Automated screen label | Same type after review | 95% sampling interval |
|---|---|---|
| Missing tool or action | 88.3% | 80.2–96.5% |
| Missing data point | 55.8% | 46.9–64.7% |
| Knowledge mismatch | 52.9% | 41.2–64.6% |
| Incorrect or stale tool data | 53.3% | 35.4–71.2% |
| No gap | 65.0% | 56.5–73.5% |
The missing-tool label is strong enough to support targeted cluster work. The other three labels should be treated as investigation hypotheses. Data, knowledge, and source-state problems often coexist in the same conversation, and the visible human response does not always isolate the technical root cause.
Study design
- Fixed population: 50,991 unique Moshi-to-human handoffs in Qonto EU, from conversations started at or after 25 June 2026 13:52:06 UTC and before 24 August 2026 13:52:06 UTC. Only summaries created by the end of that window were included.
- Sampling: 400 handoffs selected deterministically at random within the five automated screen strata: 120 missing-data, 120
none, 70 knowledge, 60 missing-tool, and 30 wrong-data assignments. Results were reweighted to the fixed production population. - Blinding: reviewers did not see the automated gap label, generated title, rationale, or another reviewer’s decision.
- Evidence: the calibration added customer turns that the original screen omitted and showed up to 20 turns before and after the handoff. One unresolved tie required a checksum-verified full-thread review because material evidence appeared beyond the 20-turn window.
- Causality rule: a later human lookup or action was not labeled as the cause of a handoff when the customer had already requested a person before disclosing the issue.
- Overlap: 46 items received three independent reviews. A hierarchical 2-of-3 majority resolved the broad gap decision before the subtype. Only three items lacked a material majority and received a fresh blind tie-break review.
- Validation: all 400 sample rows resolved; all 46 overlap rows were triple-reviewed; the three required tie-breaks were used; population, sample, assignment, schema, consistency, duplicate, missing, and orphan-adjudication checks all passed.
The reviewers were independent model-assisted analyst passes, not Qonto subject-matter experts or human ground truth. The confidence intervals quantify sampling error under the stratified design; they do not include correlated reviewer error, omitted back-office actions, or policy uncertainty.
What remains unproven
This calibration does not establish:
- that a human response reflects authoritative Qonto policy;
- that Moshi or Qonto can technically expose the required data or action;
- that a handoff was safe or desirable to prevent;
- that the visible gap caused the full conversation outcome;
- that the four-way subtype is the underlying engineering root cause;
- that a proposed release will reduce manual work; or
- that the sampled window represents future traffic and bot versions.
For a specific capability, the defensible sequence is:
candidate cluster → cluster-specific precision review → policy and eligibility definition → implementation → controlled release → observed reduction in human handling with resolution and safety guardrails
The impact number should come from the controlled release. The broad calibration establishes that the backlog has real signal; it does not substitute for that experiment.