Mytrya
Start a project

Measuring an AI support agent without asking anyone to rate it

I measure the AI support employee by comparing each draft it wrote with the reply a person actually sent, and recording whether the draft was used as-is, edited, or replaced. Whether the draft was any good is a separate question, answered by a blind judge that sees both replies in alternating order. Nobody on the team is asked to rate anything, which matters, because when the console offered a rating, almost nobody used it.

The rating button nobody pressed

On helpdesk tickets the agent works in draft mode. It posts its suggested reply as a private note, and a person copies it, edits it and sends it as themselves. The console I built for the team had a way to approve or reject each draft. In the first month there were 242 drafts, and exactly one of them was approved or rejected in the console.

That is not a complaint about the team. They work in the helpdesk, where the customer is, and the console is somewhere else. Any measurement that depends on people leaving their real tool to fill in a form will measure who has spare time, not how good the drafts are. I treat the 242 as an adoption fact from the code, not a performance claim.

Read the outcome from what people did

So the outcome is read back out of the helpdesk. A scheduled reconciliation job looks at the reply that was actually sent on each ticket, compares it with the draft, and records one of three results: used as-is, edited, or replaced. The team does nothing different. Their normal work is the signal.

This also gives the draft lifecycle somewhere to land. If the customer writes again before anyone acts, the draft is marked superseded and rewritten, so it is not compared against a reply to a different message. A draft nobody acts on expires after two weeks rather than lingering as an unknown.

Adoption is not quality

It is tempting to read used as-is as good and replaced as bad. That is wrong often enough to be dangerous. A person might rewrite a perfectly good draft because they prefer their own phrasing, because they know the customer, or because they had already started typing. A draft might be sent as-is because it was late in the day. A human writing something different is not evidence that the draft was worse.

So adoption and quality are kept as two separate measurements. Adoption tells me whether the drafts fit into how the team works. Quality needs its own judge.

A blind judge, in alternating order

Quality is decided by a model acting as a judge. It sees the customer's message, the agent's draft and the reply the human sent, without being told which is which, and decides which one is better.

The order of the two replies alternates. A judge that is always shown the agent's reply in the same position develops a position bias, a preference for the first or second option that has nothing to do with content. Alternating the order means that bias, if it is there, is spread across both sides instead of quietly favouring one.

Asking why, only when the agent lost

When the agent's draft loses, the judge is asked a second question: why did the two replies diverge? That question is deliberately not blind. To be useful, the reason has to be able to say that the human reply mentioned a known bug, or that the draft missed the customer's billing state, and it can only say that if it knows which reply is which. So the verdict is blind and the explanation is not.

Asking only on losses keeps the useful material concentrated. A win tells me little I can act on. A loss with a specific reason is something I can turn into a change.

From verdicts to learnings, through a person

Judged evaluations go through a distiller, which reads them and proposes candidate learnings: short, general statements of what the agent should do differently. Those candidates appear on an evals page in the console, and a person approves or discards each one. Only approved learnings are injected into the system prompt.

That approval step is the part I would not remove. A distiller left to write straight into the prompt will eventually generalise from one odd ticket, or encode a single colleague's preference as policy. A person reading a short candidate list is a small amount of work, and it is the point where the team's judgement enters the loop. It also means the prompt only changes in ways someone has read.

What I publish and what I don't

The ticket volumes, the share of drafts sent as-is and the agent-versus-human benchmark results belong to the client, and I don't publish them. What I can describe is the shape of the measurement: outcome from behaviour, quality from a blind and order-balanced judge, explanation only where it is needed, and changes to the agent gated on a human.

The same principle shows up elsewhere in my work. In Support Intelligence, the morning dashboard for the same client, the model names recurring issues but is never asked for a number, because counts and medians are computed in code. In the support agent's morning brief, an emerging issue is arithmetic: at least 3 tickets in 24 hours and at least 3 times the daily average of the previous 14 days. Where a measurement can come from something that actually happened, I take it from there, and I keep the model for the parts that need reading.

Drawn from

More notes

Taking new projects · start within 2 weeks

Tell me about the process that eats someone's week.

What it is, who does it, how often, and what goes wrong when it’s late. I reply within one working daywith a scoping call or an honest reason it isn’t worth automating.

Start a project [email protected]

First projects from $3,000 · fixed price