Qonto Evaluation Coverage Investigation
Summary
Qonto reported concern that some agents are not receiving enough QA evaluations. The current evidence points to two separate issues:
- Chat/email evaluations are moving, but coverage is uneven at the agent level.
- Call evaluations are effectively not meeting the configured target.
For Qonto’s active, evaluable, teamed employees, the median agent received 12 AI evaluations across the three completed report weeks, but 55 agents received zero and 73 received two or fewer. Some of those low-count agents are in teams where peers are receiving 14-15 evaluations, so this is not only a team-level sampling-rule issue.
Scope
- Customer: Qonto
- Organization ID: 33
- Production environment: EU API2 / Rails production data
- Data access mode: read-only
- Timezone used for week bucketing: Europe/Paris
- Report generated: 2026-07-01
- Report window:
- 2026-06-08 to 2026-06-14
- 2026-06-15 to 2026-06-21
- 2026-06-22 to 2026-06-28
- 2026-06-29 to 2026-07-05, current partial week
Known Context
Qonto has fixed-per-period sampling rules that target:
- Calls: 3 tickets per week per agent
- Chat/email: 5 tickets per week per agent
The current Qonto sampling rules in the exported config were created on 2026-06-23, so the first two report weeks are historical context rather than a fair post-rule target window.
Qonto also has a 24-hour evaluation delay. Evaluations created after the delay are backdated to the conversation close/resolution time for QA reporting fields. This means today’s closed conversations may not appear in today’s processed evaluation activity until tomorrow, while still landing in today’s qa_evaluated_at bucket after processing.
Current Findings
Sampling coverage
Across 380 active/evaluable teamed employees and the three completed weeks:
| Metric | Result |
|---|---|
| Median total AI evaluations per agent | 12 |
| Mean total AI evaluations per agent | 10.26 |
| Agents with zero evaluations | 55 |
| Agents with two or fewer evaluations | 73 |
| Agents meeting chat/email target every completed week | 129 / 380 |
| Agents with at least 15 chat/email evaluations total | 136 / 380 |
| Agents meeting call target every completed week | 0 / 380 |
| Agents with at least 9 call evaluations total | 0 / 380 |
| Agents with zero call evaluations | 301 / 380 |
Agent-level skew
There are agents with very low evaluation counts inside teams where peers are evaluated frequently. The current outlier list flags 34 agents with two or fewer total evaluations while their team average or median is reasonably high.
Examples of high-coverage teams with zero-evaluation agents:
| Team | Team median total evals | Zero-eval agents |
|---|---|---|
| AM FL FR (Axelle) | 15 | 12 |
| PS FL FR (Anaelle) | 15 | 3 |
| AM FL DE (Dimitrije) | 15 | 2 |
| PS FL FR (Gabrielle) | 14 | 1 |
| Frontline Team IT | 12 | 3 |
There are also teams where many agents are low together, which may indicate lower eligible volume, team-level filters, or broader channel/source coverage issues rather than individual outliers.
Week of 2026-06-22: why 79 teamed employees got zero AI evaluations
For the week from 2026-06-22 to 2026-06-28, 79 active/evaluable teamed Qonto employees got zero active AI evaluations.
Candidate definition used for this checkpoint:
- Employee is one of the 79 zero-evaluation active/evaluable teamed employees.
conversation_agents.employee_id = employee.id.conversation_agents.eligible_for_qa_evaluation = true.conversations.ticket_status = resolved.conversations.resolved_atfalls in the 2026-06-22 week in Qonto’sEurope/Paristimezone.
Employee-level primary reason:
| Primary reason | Employees | Eligible candidate pairs |
|---|---|---|
| No resolved handled conversations | 44 | 0 |
| AI/freeform ineligible | 12 | 46 |
| Resolved handled conversations, but none eligible for QA | 8 | 0 |
| Linked tickets still open | 6 | 20 |
| No customer response | 4 | 17 |
| Outside evaluation window | 3 | 11 |
| Not sampled | 1 | 8 |
| Request completed, but not as an active completed eval for this agent | 1 | 1 |
The highest-impact reason is candidate supply, not downstream QA processing: 52 of the 79 employees had zero eligible resolved candidate conversations. Of those 52:
- 44 had no resolved handled conversations in the week at all.
- 8 had resolved handled conversations, but every one of those conversation-agent links had
eligible_for_qa_evaluation = false.
The largest downstream reason among employees who did have eligible candidates was AI/freeform ineligibility from Qonto’s eligibility policy.
Candidate-level outcome breakdown for the 103 eligible candidate pairs owned by the 27 employees who had at least one candidate:
| Candidate outcome | Channel | Candidate pairs | Employees |
|---|---|---|---|
| AI/freeform ineligible | Chat/email | 40 | 18 |
| Linked tickets still open | Chat/email | 25 | 13 |
| No customer response | Chat/email | 15 | 6 |
| Outside evaluation window | Chat/email | 11 | 6 |
| Not sampled | Null source | 5 | 1 |
| Request completed, but not active/completed for this agent | Chat/email | 3 | 3 |
| Linked tickets still open | Call | 2 | 2 |
| No scorecards | Call | 1 | 1 |
| Request failed | Chat/email | 1 | 1 |
The 52 no-candidate employees were concentrated in a few teams:
| Team | Employees | Resolved handled pairs | Resolved but not eligible pairs |
|---|---|---|---|
| AM FL FR (Julie) | 14 | 0 | 0 |
| AM FL FR (Axelle) | 10 | 42 | 42 |
| Frontline Team NL | 8 | 1 | 1 |
| PS FL FR (Anaelle) | 5 | 0 | 0 |
This means the next highest-impact question is why some active/evaluable teamed employees had no eligible resolved candidate conversations in that week, especially the AM FL FR (Axelle) group where the handled resolved conversations existed but were all marked not eligible for QA.
Examples from the 8 employees with resolved handled conversations but zero QA-eligible conversation-agent links:
| Employee DB id | Public id | Name | Team | Resolved handled pairs | QA-eligible pairs |
|---|---|---|---|---|---|
| 4553 | fdafdk59tbix | Kaba C M. - (M) | AM FL FR (Axelle) | 29 | 0 |
| 4870 | 336j395rc2k5 | Jihane N M. - (M) | AM FL FR (Axelle) | 9 | 0 |
| 4952 | kf9lppeqdvb9 | Fred M. | AM FL FR (Axelle) | 3 | 0 |
| 4561 | ogdnvsxkhzdg | Rihab A. | ECT FL FR | 3 | 0 |
| 4683 | stu84djdfuw8 | Momo A. A. | Frontline Team IT | 2 | 0 |
| 14228 | 7nua5q0pe5qs | Michel M. | AM FL FR (Axelle) | 1 | 0 |
| 4599 | 177kyeqndgm0 | Valentina T. A. | Frontline Team IT | 1 | 0 |
| 4695 | z4iezf2xyx9l | Amine Mlimane | Frontline Team NL | 1 | 0 |
Spot check: Kaba C M. (employee_id = 4553) was attached to 29 resolved conversations in the week, but their authored parts were not customer-facing QA-eligible participation:
| Authored part type | Parts | Conversations |
|---|---|---|
note | 32 | 28 |
assignment | 1 | 1 |
attribute_updated_by_admin | 1 | 1 |
Kaba authored no chat/email/call parts on those conversations. In code, ConversationPart#evaluable_for_qa_eligibility? returns true for chat, email, and call, plus wrapper assignment/close/open parts only when they carry a reply body. note does not count. So this specific high-volume zero-candidate case appears expected: the agent was leaving internal notes/comments, not customer-facing responses.
Week of 2026-06-22: next zero-eval block, AI/freeform ineligible
After the 52 zero-candidate employees, the next largest employee-level block is 12 employees whose primary zero-evaluation reason was AI/freeform ineligibility.
Those 12 employees had 46 eligible candidate pairs. Of those, 32 candidate pairs were skipped by the AI/freeform eligibility step.
| Employee DB id | Public id | Name | Team | Eligible candidates | AI/freeform ineligible | Other notable outcomes |
|---|---|---|---|---|---|---|
| 4865 | ptibmi963sf5 | Alessandra T A. | Frontline Team IT | 18 | 11 | 5 linked tickets open, 1 no customer response, 1 outside window |
| 4496 | vm6wvvy6r8t2 | Mansour W M. | AM FL FR (Axelle) | 7 | 5 | 2 linked tickets open |
| 4674 | dmbjm1ny4cbw | Youness MK. | AM FL FR (Julie) | 6 | 4 | 1 linked ticket open, 1 completed but not active for this agent |
| 4480 | bxt7m1k8ki64 | Nenad Savic | PS FL DE (Filip) | 4 | 4 | - |
| 4175 | b9hb3e86o65g | Leila A. | Frontline Team IT | 2 | 1 | 1 outside window |
| 4844 | oa56o543r64i | Vuk Kocic | AM BL DE (Vojislav) | 2 | 1 | 1 outside window |
| 4877 | l9wyfsn6mj75 | Valentina Rakic | PS BL DE (Pavle) | 2 | 1 | 1 completed but not active for this agent |
| 17114 | 9v8dtfhtygzm | Heni A. | AM FL FR (Julie) | 1 | 1 | - |
| 22996 | xu4ncz0ppp6d | Redouane H M. | AM FL FR (Axelle) | 1 | 1 | - |
| 5076 | a7x622ods1xe | Ghassen A. | ECT FL FR | 1 | 1 | - |
| 5280 | 8qpnqlkx05td | Divin B M. - (M) | AM FL FR (Axelle) | 1 | 1 | - |
| 7652 | gnd50m6nmlw2 | Ilhem Chaabani | AM FL FR (Julie) | 1 | 1 | - |
AI/freeform skip reason themes for the 32 AI/freeform-skipped candidate pairs in this block:
| Primary theme | Candidate pairs | Employees |
|---|---|---|
| Multiple human agents / single-agent rule failed | 20 | 9 |
| No meaningful two-way exchange | 8 | 3 |
| Outbound call in same ticket/request | 2 | 2 |
| Other AI ineligible | 2 | 2 |
Some skip reasons mention multiple Qonto policy clauses at once. Allowing overlapping theme mentions:
| Mentioned theme | Candidate pairs | Employees |
|---|---|---|
| Multiple human agents / single-agent rule | 20 | 9 |
| No meaningful two-way exchange | 10 | 4 |
| Outbound call | 7 | 5 |
| Verification/authentication/identity block | 4 | 3 |
This block appears policy-driven rather than a sampling bug. The dominant reason is Qonto’s custom single-agent requirement: tickets are made ineligible when more than one human agent replied, except for the configured BL exception. Secondary reasons also map directly to Qonto’s custom instructions: no meaningful two-way exchange, outbound calls in the same request, and authentication/identity flows where the underlying issue was not resolved.
2026-07-01 read-only production checkpoint
Qonto’s custom QA eligibility instructions are materially stricter than the default org instructions.
Default eligibility requires:
- A real customer request/question/issue.
- A meaningful two-way exchange with a human agent.
Qonto adds:
- Only one human agent may reply to the client, except for a BL group exception.
- No outbound call in a related ticket for the same request, except identity/security requests.
- Exclude tickets tagged
02.06.2026_Sepa_transfer_Incidentornot_a_client. - Exclude verification/authentication cases when the agent does not proceed to resolve the underlying question.
- When unclear, lean ineligible.
Latest-request-per-conversation skip breakdown for the week starting 2026-06-22:
| Channel bucket | Dominant latest skip reasons |
|---|---|
| Chat/email-started | no_scorecards 7,521; AI/freeform ineligible 5,744; linked_tickets_open 5,269; no_customer_response 3,539; no_agent_response 3,048 |
| Call-started | linked_tickets_open 178; no_scorecards 119; no_agents_sampled 117; AI/freeform ineligible 93; no_agent_response 28 |
| Source null with call part | no_scorecards 675; not_sampled 489; linked_tickets_open 203 |
Across the four-week latest-request sample, AI/freeform ineligibility reasons were dominated by Qonto-specific instruction themes:
| Inferred AI/freeform theme | Conversations |
|---|---|
| Single-agent / multiple-agent rule | 10,630 |
| No meaningful two-way exchange | 7,541 |
| Outbound or related-call exclusion | 961 |
| Verification/auth block | 473 |
| Explicit tag exclusion | 58 |
This means Qonto’s custom instructions are definitely filtering a lot of otherwise pre-eligible tickets. It is not yet clear whether they are too strict relative to Qonto’s intended policy, but they are a major source of skipped evaluations.
Agent-level candidate supply
Using conversation_agents.eligible_for_qa_evaluation = true as the pre-LLM/pre-sampling candidate signal:
| Metric | Three completed weeks | Week starting 2026-06-22 |
|---|---|---|
| Active/evaluable teamed agents | 380 | 380 |
| Median pre-LLM eligible chat/email tickets per agent | 127 | 39 |
| Median completed active chat/email evals per agent | 12 | 5 |
| Agents with enough chat/email supply but under target | 182 had 15+ candidates but <15 evals | 100 had 5+ candidates but <5 evals |
| Median pre-LLM eligible call-started tickets per agent | 0 | 0 |
| Agents with enough call-started supply but under target | 91 had 9+ candidates but <9 evals | 51 had 3+ candidates but <3 evals |
So the low-evaluation issue is not only low ticket volume. Many agents have enough pre-LLM candidate conversations but still miss coverage targets after scorecard, linked-ticket, sampling, and AI eligibility gates.
The original 2026-07-01 CSV showed 178 agents missing the chat/email target with a 647-evaluation gap. A direct production rerun on 2026-07-02 shows 177 agents and a 645-evaluation gap, so one agent appears to have gained two chat/email evaluations since the export.
For the week starting 2026-06-22, using the current production rerun:
| Metric | Agents / evaluations |
|---|---|
| Agents missing the chat/email target | 177 |
| Missing chat/email evaluations vs 5/week target | 645 |
| Agents with zero chat/email evaluations | 82 |
| Missing evaluations from zero chat/email agents | 410 |
| Agents with 1-4 chat/email evaluations | 95 |
| Missing evaluations from partial chat/email agents | 235 |
Those 177 agents had 2,523 resolved QA-eligible chat/email candidate pairs across 125 agents. 52 of the 177 under-target agents had no resolved QA-eligible chat/email candidate pairs in the week. 97 had at least five such candidate pairs but still missed the 5/week chat/email target.
Exact blocker breakdown for the unresolved chat/email target gap:
| Blocker | Raw candidate pairs | Agents | Target-capped impact |
|---|---|---|---|
| AI/freeform ineligible | 1,228 | 114 | 243 |
| Linked tickets still open | 559 | 101 | 191 |
| No customer response | 179 | 51 | 79 |
| No agents sampled | 89 | 22 | 26 |
| Ticket reopened / request cancelled | 73 | 41 | 46 |
| Outside evaluation window | 60 | 36 | 49 |
| Completed inactive for this agent | 54 | 30 | 38 |
| Request completed, but not active for this agent | 12 | 10 | 10 |
| No eligible agents found at request time | 9 | 8 | 9 |
| Same agent completed outside this effective week or inactive window | 6 | 6 | 6 |
| No agent response | 5 | 4 | 4 |
| No matching scorecard | 5 | 5 | 5 |
| Request failed | 4 | 4 | 4 |
| Not sampled | 3 | 2 | 3 |
| Conversation completed for another agent | 1 | 1 | 1 |
Target-capped impact means the row is capped by each agent’s remaining chat/email target gap. These rows are not additive across recommendations because the same agent’s remaining gap can be filled by more than one blocker category.
AI/freeform policy themes within the 1,228 AI/freeform-blocked candidate pairs:
| AI/freeform theme | Raw candidate pairs | Agents | Target-capped impact |
|---|---|---|---|
| Single-agent / multiple-agent rule | 841 | 98 | 200 |
| No meaningful two-way exchange | 232 | 69 | 98 |
| Outbound-call exclusion | 57 | 36 | 43 |
| Verification/authentication/identity exclusion | 36 | 25 | 27 |
| Explicit tag exclusion | 5 | 4 | 5 |
| Other AI/freeform reason | 57 | 35 | 38 |
Sampling counters vs active evaluations
Qonto has two fixed weekly per-agent rules in production:
| Rule id | Target | Source condition | Created |
|---|---|---|---|
| 9 | 5/week/agent | source_channel in [chat, email] | 2026-05-11 |
| 11 | 3/week/agent | source_channel in [call] | 2026-06-23 |
At the time of this investigation, queue pacing did not appear active for Qonto, and queued_count was zero on the inspected counters.
For the week starting 2026-06-22:
| Rule | Counters | Consumed sum | Counters at/above target |
|---|---|---|---|
| Chat/email rule 9 | 297 | 1,312 | 233 |
| Call rule 11 | 71 | 211 | 69 |
The call counter state does not match the active-evaluation coverage report. In the effective week starting 2026-06-22, call-started AI agent evaluations had:
| Status | Active | Overridden | Agent eval rows |
|---|---|---|---|
| completed | false | false | 211 |
| completed | true | false | 138 |
Those 211 inactive completed call rows have scores and completed scorecard results, but no active twin for the same conversation/agent/source. They do not count in normal coverage/reporting because reporting correctly uses active/effective evaluations, but the sampling counter reconciler currently counts completed/pending AI evaluations without filtering to active = true / effective rows. This is a strong root-cause candidate for call undercoverage: counters can believe agents reached their call quota while the report shows fewer active evaluations.
Operational watch item: because some call evaluations were manually marked inactive during cleanup, the first week after rulebase-co/rulebase#7855 is merged and the call source-channel backfill is run should be monitored carefully. Newly backfilled call tickets should enter the call sampling rule, active completed call evaluations should rise, and any continued gap between sampling counters and active completed call evaluations should be investigated as an inactive-row/counter-reconciliation issue before treating it as low call volume.
The same inactive pattern exists for chat/email, but it is much smaller:
| Channel | Completed inactive/effective rows | Completed active/effective rows |
|---|---|---|
| Call-started | 211 | 165 |
| Chat/email-started | 73 | 3,751 |
2026-06-22 no_scorecards interpretation
no_scorecards is emitted before sampling, when the queue job cannot find any published scorecard matching the conversation.
Qonto currently has one published scorecard, Stellar Support Scorecard (scorecard_id = 57). It has:
- One
teamcondition containing 26 Qonto teams. - No
source_channelcondition. - No explicit
ticket_scopecondition, so Rails defaults it toexternalonly.
For the latest request per conversation in the week starting 2026-06-22:
| Channel bucket | Latest no_scorecards conversations |
|---|---|
| Chat/email-started | 7,683 |
| Source null | 675 |
| Call-started | 119 |
Actual scorecard predicate breakdown for those no_scorecards rows:
| Shape | Conversations |
|---|---|
| Chat/email, external, no stored/live team match | 7,434 |
| Chat/email, external, no eligible agents | 211 |
| Chat/email, external, current stored/live team match | 16 |
Chat/email, external, live team match but no stored conversation_team match | 4 |
| Chat/email, internal, no team match | 11 |
| Chat/email, internal, team match | 7 |
| Source-null, internal, no team match | 518 |
| Source-null, external, no team match | 157 |
| Call-started, internal, team match | 80 |
| Call-started, internal, no team match | 28 |
| Call-started, external, no team match | 11 |
For teamed rows, the most important callout is the implicit external-only ticket scope. The 80 call-started rows with a scorecard team match are still internal by Conversation#ticket_scope because first_customer_response_at is null, so the scorecard does not match despite the agent/team being configured.
The largest chat/email bucket, 7,434 conversations, was handled by 672 distinct eligible human/evaluable employee rows, but those employee rows all had team_id = null in Rulebase. They did not match another teamed employee row by exact email or exact name. A smaller subset of 151 conversations did have some teamed conversation participant/message author, but that teamed employee was not the eligible QA agent on the conversation, so the scorecard still had no eligible teamed agent to attach to.
The 16 chat/email rows that match the scorecard in current state all had their conversation_teams rows created after the skipped QA evaluation request was written. At request time, the queue job appears to have seen no matching stored conversation team. This fits the local code path: QueueEligibleConversationsForQAEvaluationJob uses cached scorecards and calls scorecard.matches?(conversation) directly, while Conversation#scorecards_for_evaluation would refresh conversation_teams from agent teams before matching.
Local Artifacts
Generated CSVs are stored alongside this note:
The outlier CSV is sorted from lowest-evaluated agent upward and includes team mean/median comparison fields.
Code Pointers
The current behavior of backdating QA evaluation reporting timestamps appears intentional:
rulebase-api/app/models/qa_evaluation_request.rb- Captures
effective_evaluation_atfromconversation.resolved_at || conversation.closed_atforticket_closedrequests. - Backdates
qa_evaluations.created_at,qa_agent_evaluations.created_at, andconversations.qa_evaluated_atto the ticket close/resolution anchor when blank.
- Captures
Open Questions
- Do the low-evaluated agents actually have enough eligible closed conversations in the report window?
- Are those conversations assigned to the employee records used by the sampling partition?
- Are call conversations being classified as
callin the fields used by sampling and reporting? - Are some teams targeted by the exported sampling rules but excluded later by scorecard, eligibility, source-channel, assignment, or conversation-state filters?
- Are low-count agents primarily low-volume agents, newly active agents, vendor users with identity mismatch, or agents with team/employee mapping churn?
Next Investigation Steps
- For the flagged outlier agents, count eligible closed conversations by source channel and week.
- Compare eligible conversation counts with completed evaluation counts for the same agent-week-channel partitions.
- For calls, trace where eligible call volume disappears:
- source channel classification
- scorecard availability
- sampling rule match
- evaluation request creation
- request processing status
- Review whether the sampling job uses the same employee/team mapping as the report query.
- Identify whether the fix is data repair, rule configuration, source-channel normalization, assignment mapping, or sampling logic.
Working Notes
- Chat/email undercoverage is split between uneven eligible ticket volume and downstream gates after candidate supply. The largest exact downstream blockers are Qonto AI/freeform policy and linked-ticket gating.
- Call undercoverage is a separate issue. The volume is too low relative to the target and involves channel classification, rule matching, inactive evaluations, and candidate supply.
- The “no evaluations today” symptom is not itself the root cause because Qonto’s 24-hour delay and QA timestamp backdating make today’s processing appear under previous close-date buckets.
- Qonto’s custom eligibility instructions are a major source of AI/freeform skips, especially single-agent/multiple-agent and two-way-exchange requirements.
- Sampling counters may overcount inactive completed AI agent evaluations. This appears especially important for call-started tickets and may make per-agent call quota look full even when active evaluation coverage is below target.