Skip to Content
Internal docs are powered by Nextra Docs Theme.
CustomersQontoEvaluation coverage

Qonto Evaluation Coverage Investigation

Summary

Qonto reported concern that some agents are not receiving enough QA evaluations. The current evidence points to two separate issues:

  • Chat/email evaluations are moving, but coverage is uneven at the agent level.
  • Call evaluations are effectively not meeting the configured target.

For Qonto’s active, evaluable, teamed employees, the median agent received 12 AI evaluations across the three completed report weeks, but 55 agents received zero and 73 received two or fewer. Some of those low-count agents are in teams where peers are receiving 14-15 evaluations, so this is not only a team-level sampling-rule issue.

Scope

  • Customer: Qonto
  • Organization ID: 33
  • Production environment: EU API2 / Rails production data
  • Data access mode: read-only
  • Timezone used for week bucketing: Europe/Paris
  • Report generated: 2026-07-01
  • Report window:
    • 2026-06-08 to 2026-06-14
    • 2026-06-15 to 2026-06-21
    • 2026-06-22 to 2026-06-28
    • 2026-06-29 to 2026-07-05, current partial week

Known Context

Qonto has fixed-per-period sampling rules that target:

  • Calls: 3 tickets per week per agent
  • Chat/email: 5 tickets per week per agent

The current Qonto sampling rules in the exported config were created on 2026-06-23, so the first two report weeks are historical context rather than a fair post-rule target window.

Qonto also has a 24-hour evaluation delay. Evaluations created after the delay are backdated to the conversation close/resolution time for QA reporting fields. This means today’s closed conversations may not appear in today’s processed evaluation activity until tomorrow, while still landing in today’s qa_evaluated_at bucket after processing.

Current Findings

Sampling coverage

Across 380 active/evaluable teamed employees and the three completed weeks:

MetricResult
Median total AI evaluations per agent12
Mean total AI evaluations per agent10.26
Agents with zero evaluations55
Agents with two or fewer evaluations73
Agents meeting chat/email target every completed week129 / 380
Agents with at least 15 chat/email evaluations total136 / 380
Agents meeting call target every completed week0 / 380
Agents with at least 9 call evaluations total0 / 380
Agents with zero call evaluations301 / 380

Agent-level skew

There are agents with very low evaluation counts inside teams where peers are evaluated frequently. The current outlier list flags 34 agents with two or fewer total evaluations while their team average or median is reasonably high.

Examples of high-coverage teams with zero-evaluation agents:

TeamTeam median total evalsZero-eval agents
AM FL FR (Axelle)1512
PS FL FR (Anaelle)153
AM FL DE (Dimitrije)152
PS FL FR (Gabrielle)141
Frontline Team IT123

There are also teams where many agents are low together, which may indicate lower eligible volume, team-level filters, or broader channel/source coverage issues rather than individual outliers.

Week of 2026-06-22: why 79 teamed employees got zero AI evaluations

For the week from 2026-06-22 to 2026-06-28, 79 active/evaluable teamed Qonto employees got zero active AI evaluations.

Candidate definition used for this checkpoint:

  • Employee is one of the 79 zero-evaluation active/evaluable teamed employees.
  • conversation_agents.employee_id = employee.id.
  • conversation_agents.eligible_for_qa_evaluation = true.
  • conversations.ticket_status = resolved.
  • conversations.resolved_at falls in the 2026-06-22 week in Qonto’s Europe/Paris timezone.

Employee-level primary reason:

Primary reasonEmployeesEligible candidate pairs
No resolved handled conversations440
AI/freeform ineligible1246
Resolved handled conversations, but none eligible for QA80
Linked tickets still open620
No customer response417
Outside evaluation window311
Not sampled18
Request completed, but not as an active completed eval for this agent11

The highest-impact reason is candidate supply, not downstream QA processing: 52 of the 79 employees had zero eligible resolved candidate conversations. Of those 52:

  • 44 had no resolved handled conversations in the week at all.
  • 8 had resolved handled conversations, but every one of those conversation-agent links had eligible_for_qa_evaluation = false.

The largest downstream reason among employees who did have eligible candidates was AI/freeform ineligibility from Qonto’s eligibility policy.

Candidate-level outcome breakdown for the 103 eligible candidate pairs owned by the 27 employees who had at least one candidate:

Candidate outcomeChannelCandidate pairsEmployees
AI/freeform ineligibleChat/email4018
Linked tickets still openChat/email2513
No customer responseChat/email156
Outside evaluation windowChat/email116
Not sampledNull source51
Request completed, but not active/completed for this agentChat/email33
Linked tickets still openCall22
No scorecardsCall11
Request failedChat/email11

The 52 no-candidate employees were concentrated in a few teams:

TeamEmployeesResolved handled pairsResolved but not eligible pairs
AM FL FR (Julie)1400
AM FL FR (Axelle)104242
Frontline Team NL811
PS FL FR (Anaelle)500

This means the next highest-impact question is why some active/evaluable teamed employees had no eligible resolved candidate conversations in that week, especially the AM FL FR (Axelle) group where the handled resolved conversations existed but were all marked not eligible for QA.

Examples from the 8 employees with resolved handled conversations but zero QA-eligible conversation-agent links:

Employee DB idPublic idNameTeamResolved handled pairsQA-eligible pairs
4553fdafdk59tbixKaba C M. - (M)AM FL FR (Axelle)290
4870336j395rc2k5Jihane N M. - (M)AM FL FR (Axelle)90
4952kf9lppeqdvb9Fred M.AM FL FR (Axelle)30
4561ogdnvsxkhzdgRihab A.ECT FL FR30
4683stu84djdfuw8Momo A. A.Frontline Team IT20
142287nua5q0pe5qsMichel M.AM FL FR (Axelle)10
4599177kyeqndgm0Valentina T. A.Frontline Team IT10
4695z4iezf2xyx9lAmine MlimaneFrontline Team NL10

Spot check: Kaba C M. (employee_id = 4553) was attached to 29 resolved conversations in the week, but their authored parts were not customer-facing QA-eligible participation:

Authored part typePartsConversations
note3228
assignment11
attribute_updated_by_admin11

Kaba authored no chat/email/call parts on those conversations. In code, ConversationPart#evaluable_for_qa_eligibility? returns true for chat, email, and call, plus wrapper assignment/close/open parts only when they carry a reply body. note does not count. So this specific high-volume zero-candidate case appears expected: the agent was leaving internal notes/comments, not customer-facing responses.

Week of 2026-06-22: next zero-eval block, AI/freeform ineligible

After the 52 zero-candidate employees, the next largest employee-level block is 12 employees whose primary zero-evaluation reason was AI/freeform ineligibility.

Those 12 employees had 46 eligible candidate pairs. Of those, 32 candidate pairs were skipped by the AI/freeform eligibility step.

Employee DB idPublic idNameTeamEligible candidatesAI/freeform ineligibleOther notable outcomes
4865ptibmi963sf5Alessandra T A.Frontline Team IT18115 linked tickets open, 1 no customer response, 1 outside window
4496vm6wvvy6r8t2Mansour W M.AM FL FR (Axelle)752 linked tickets open
4674dmbjm1ny4cbwYouness MK.AM FL FR (Julie)641 linked ticket open, 1 completed but not active for this agent
4480bxt7m1k8ki64Nenad SavicPS FL DE (Filip)44-
4175b9hb3e86o65gLeila A.Frontline Team IT211 outside window
4844oa56o543r64iVuk KocicAM BL DE (Vojislav)211 outside window
4877l9wyfsn6mj75Valentina RakicPS BL DE (Pavle)211 completed but not active for this agent
171149v8dtfhtygzmHeni A.AM FL FR (Julie)11-
22996xu4ncz0ppp6dRedouane H M.AM FL FR (Axelle)11-
5076a7x622ods1xeGhassen A.ECT FL FR11-
52808qpnqlkx05tdDivin B M. - (M)AM FL FR (Axelle)11-
7652gnd50m6nmlw2Ilhem ChaabaniAM FL FR (Julie)11-

AI/freeform skip reason themes for the 32 AI/freeform-skipped candidate pairs in this block:

Primary themeCandidate pairsEmployees
Multiple human agents / single-agent rule failed209
No meaningful two-way exchange83
Outbound call in same ticket/request22
Other AI ineligible22

Some skip reasons mention multiple Qonto policy clauses at once. Allowing overlapping theme mentions:

Mentioned themeCandidate pairsEmployees
Multiple human agents / single-agent rule209
No meaningful two-way exchange104
Outbound call75
Verification/authentication/identity block43

This block appears policy-driven rather than a sampling bug. The dominant reason is Qonto’s custom single-agent requirement: tickets are made ineligible when more than one human agent replied, except for the configured BL exception. Secondary reasons also map directly to Qonto’s custom instructions: no meaningful two-way exchange, outbound calls in the same request, and authentication/identity flows where the underlying issue was not resolved.

2026-07-01 read-only production checkpoint

Qonto’s custom QA eligibility instructions are materially stricter than the default org instructions.

Default eligibility requires:

  • A real customer request/question/issue.
  • A meaningful two-way exchange with a human agent.

Qonto adds:

  • Only one human agent may reply to the client, except for a BL group exception.
  • No outbound call in a related ticket for the same request, except identity/security requests.
  • Exclude tickets tagged 02.06.2026_Sepa_transfer_Incident or not_a_client.
  • Exclude verification/authentication cases when the agent does not proceed to resolve the underlying question.
  • When unclear, lean ineligible.

Latest-request-per-conversation skip breakdown for the week starting 2026-06-22:

Channel bucketDominant latest skip reasons
Chat/email-startedno_scorecards 7,521; AI/freeform ineligible 5,744; linked_tickets_open 5,269; no_customer_response 3,539; no_agent_response 3,048
Call-startedlinked_tickets_open 178; no_scorecards 119; no_agents_sampled 117; AI/freeform ineligible 93; no_agent_response 28
Source null with call partno_scorecards 675; not_sampled 489; linked_tickets_open 203

Across the four-week latest-request sample, AI/freeform ineligibility reasons were dominated by Qonto-specific instruction themes:

Inferred AI/freeform themeConversations
Single-agent / multiple-agent rule10,630
No meaningful two-way exchange7,541
Outbound or related-call exclusion961
Verification/auth block473
Explicit tag exclusion58

This means Qonto’s custom instructions are definitely filtering a lot of otherwise pre-eligible tickets. It is not yet clear whether they are too strict relative to Qonto’s intended policy, but they are a major source of skipped evaluations.

Agent-level candidate supply

Using conversation_agents.eligible_for_qa_evaluation = true as the pre-LLM/pre-sampling candidate signal:

MetricThree completed weeksWeek starting 2026-06-22
Active/evaluable teamed agents380380
Median pre-LLM eligible chat/email tickets per agent12739
Median completed active chat/email evals per agent125
Agents with enough chat/email supply but under target182 had 15+ candidates but <15 evals100 had 5+ candidates but <5 evals
Median pre-LLM eligible call-started tickets per agent00
Agents with enough call-started supply but under target91 had 9+ candidates but <9 evals51 had 3+ candidates but <3 evals

So the low-evaluation issue is not only low ticket volume. Many agents have enough pre-LLM candidate conversations but still miss coverage targets after scorecard, linked-ticket, sampling, and AI eligibility gates.

The original 2026-07-01 CSV showed 178 agents missing the chat/email target with a 647-evaluation gap. A direct production rerun on 2026-07-02 shows 177 agents and a 645-evaluation gap, so one agent appears to have gained two chat/email evaluations since the export.

For the week starting 2026-06-22, using the current production rerun:

MetricAgents / evaluations
Agents missing the chat/email target177
Missing chat/email evaluations vs 5/week target645
Agents with zero chat/email evaluations82
Missing evaluations from zero chat/email agents410
Agents with 1-4 chat/email evaluations95
Missing evaluations from partial chat/email agents235

Those 177 agents had 2,523 resolved QA-eligible chat/email candidate pairs across 125 agents. 52 of the 177 under-target agents had no resolved QA-eligible chat/email candidate pairs in the week. 97 had at least five such candidate pairs but still missed the 5/week chat/email target.

Exact blocker breakdown for the unresolved chat/email target gap:

BlockerRaw candidate pairsAgentsTarget-capped impact
AI/freeform ineligible1,228114243
Linked tickets still open559101191
No customer response1795179
No agents sampled892226
Ticket reopened / request cancelled734146
Outside evaluation window603649
Completed inactive for this agent543038
Request completed, but not active for this agent121010
No eligible agents found at request time989
Same agent completed outside this effective week or inactive window666
No agent response544
No matching scorecard555
Request failed444
Not sampled323
Conversation completed for another agent111

Target-capped impact means the row is capped by each agent’s remaining chat/email target gap. These rows are not additive across recommendations because the same agent’s remaining gap can be filled by more than one blocker category.

AI/freeform policy themes within the 1,228 AI/freeform-blocked candidate pairs:

AI/freeform themeRaw candidate pairsAgentsTarget-capped impact
Single-agent / multiple-agent rule84198200
No meaningful two-way exchange2326998
Outbound-call exclusion573643
Verification/authentication/identity exclusion362527
Explicit tag exclusion545
Other AI/freeform reason573538

Sampling counters vs active evaluations

Qonto has two fixed weekly per-agent rules in production:

Rule idTargetSource conditionCreated
95/week/agentsource_channel in [chat, email]2026-05-11
113/week/agentsource_channel in [call]2026-06-23

At the time of this investigation, queue pacing did not appear active for Qonto, and queued_count was zero on the inspected counters.

For the week starting 2026-06-22:

RuleCountersConsumed sumCounters at/above target
Chat/email rule 92971,312233
Call rule 117121169

The call counter state does not match the active-evaluation coverage report. In the effective week starting 2026-06-22, call-started AI agent evaluations had:

StatusActiveOverriddenAgent eval rows
completedfalsefalse211
completedtruefalse138

Those 211 inactive completed call rows have scores and completed scorecard results, but no active twin for the same conversation/agent/source. They do not count in normal coverage/reporting because reporting correctly uses active/effective evaluations, but the sampling counter reconciler currently counts completed/pending AI evaluations without filtering to active = true / effective rows. This is a strong root-cause candidate for call undercoverage: counters can believe agents reached their call quota while the report shows fewer active evaluations.

Operational watch item: because some call evaluations were manually marked inactive during cleanup, the first week after rulebase-co/rulebase#7855  is merged and the call source-channel backfill is run should be monitored carefully. Newly backfilled call tickets should enter the call sampling rule, active completed call evaluations should rise, and any continued gap between sampling counters and active completed call evaluations should be investigated as an inactive-row/counter-reconciliation issue before treating it as low call volume.

The same inactive pattern exists for chat/email, but it is much smaller:

ChannelCompleted inactive/effective rowsCompleted active/effective rows
Call-started211165
Chat/email-started733,751

2026-06-22 no_scorecards interpretation

no_scorecards is emitted before sampling, when the queue job cannot find any published scorecard matching the conversation.

Qonto currently has one published scorecard, Stellar Support Scorecard (scorecard_id = 57). It has:

  • One team condition containing 26 Qonto teams.
  • No source_channel condition.
  • No explicit ticket_scope condition, so Rails defaults it to external only.

For the latest request per conversation in the week starting 2026-06-22:

Channel bucketLatest no_scorecards conversations
Chat/email-started7,683
Source null675
Call-started119

Actual scorecard predicate breakdown for those no_scorecards rows:

ShapeConversations
Chat/email, external, no stored/live team match7,434
Chat/email, external, no eligible agents211
Chat/email, external, current stored/live team match16
Chat/email, external, live team match but no stored conversation_team match4
Chat/email, internal, no team match11
Chat/email, internal, team match7
Source-null, internal, no team match518
Source-null, external, no team match157
Call-started, internal, team match80
Call-started, internal, no team match28
Call-started, external, no team match11

For teamed rows, the most important callout is the implicit external-only ticket scope. The 80 call-started rows with a scorecard team match are still internal by Conversation#ticket_scope because first_customer_response_at is null, so the scorecard does not match despite the agent/team being configured.

The largest chat/email bucket, 7,434 conversations, was handled by 672 distinct eligible human/evaluable employee rows, but those employee rows all had team_id = null in Rulebase. They did not match another teamed employee row by exact email or exact name. A smaller subset of 151 conversations did have some teamed conversation participant/message author, but that teamed employee was not the eligible QA agent on the conversation, so the scorecard still had no eligible teamed agent to attach to.

The 16 chat/email rows that match the scorecard in current state all had their conversation_teams rows created after the skipped QA evaluation request was written. At request time, the queue job appears to have seen no matching stored conversation team. This fits the local code path: QueueEligibleConversationsForQAEvaluationJob uses cached scorecards and calls scorecard.matches?(conversation) directly, while Conversation#scorecards_for_evaluation would refresh conversation_teams from agent teams before matching.

Local Artifacts

Generated CSVs are stored alongside this note:

The outlier CSV is sorted from lowest-evaluated agent upward and includes team mean/median comparison fields.

Code Pointers

The current behavior of backdating QA evaluation reporting timestamps appears intentional:

  • rulebase-api/app/models/qa_evaluation_request.rb
    • Captures effective_evaluation_at from conversation.resolved_at || conversation.closed_at for ticket_closed requests.
    • Backdates qa_evaluations.created_at, qa_agent_evaluations.created_at, and conversations.qa_evaluated_at to the ticket close/resolution anchor when blank.

Open Questions

  1. Do the low-evaluated agents actually have enough eligible closed conversations in the report window?
  2. Are those conversations assigned to the employee records used by the sampling partition?
  3. Are call conversations being classified as call in the fields used by sampling and reporting?
  4. Are some teams targeted by the exported sampling rules but excluded later by scorecard, eligibility, source-channel, assignment, or conversation-state filters?
  5. Are low-count agents primarily low-volume agents, newly active agents, vendor users with identity mismatch, or agents with team/employee mapping churn?

Next Investigation Steps

  1. For the flagged outlier agents, count eligible closed conversations by source channel and week.
  2. Compare eligible conversation counts with completed evaluation counts for the same agent-week-channel partitions.
  3. For calls, trace where eligible call volume disappears:
    • source channel classification
    • scorecard availability
    • sampling rule match
    • evaluation request creation
    • request processing status
  4. Review whether the sampling job uses the same employee/team mapping as the report query.
  5. Identify whether the fix is data repair, rule configuration, source-channel normalization, assignment mapping, or sampling logic.

Working Notes

  • Chat/email undercoverage is split between uneven eligible ticket volume and downstream gates after candidate supply. The largest exact downstream blockers are Qonto AI/freeform policy and linked-ticket gating.
  • Call undercoverage is a separate issue. The volume is too low relative to the target and involves channel classification, rule matching, inactive evaluations, and candidate supply.
  • The “no evaluations today” symptom is not itself the root cause because Qonto’s 24-hour delay and QA timestamp backdating make today’s processing appear under previous close-date buckets.
  • Qonto’s custom eligibility instructions are a major source of AI/freeform skips, especially single-agent/multiple-agent and two-way-exchange requirements.
  • Sampling counters may overcount inactive completed AI agent evaluations. This appears especially important for call-started tickets and may make per-agent call quota look full even when active evaluation coverage is below target.
Last updated on