Why our support agent has no reply tool on chat
On the chat channel, the AI support employee I built has no reply tool: whatever text the model ends its turn with is the message the visitor sees. I removed the tool because one chat-tuned model kept researching an answer, logging a note that said it had replied, and sending nothing. Two rounds of prompt hardening did not stop it. Taking the tool away did, because a failure that needs the tool cannot happen without it.
The setup
The agent works for a B2B software vendor with two product brands. It runs on three channels that share one agent loop and one tool layer: helpdesk tickets in Freshdesk, the vendor's live chat in Freshchat, and a chat widget of its own that sits on the marketing site and inside the client's apps. There are 20 tools in total, and several of them only exist on one channel. On the chat channel the agent can save a visitor's identity, request a human, open a ticket and add to a ticket. Everywhere else, sending a reply is an explicit tool call, alongside tools for writing an internal note and closing the case.
That split made sense on paper. On tickets, a reply is a deliberate act with consequences, so it should be a named action that shows up in the trace. The agent's free text at the end of a turn is its own working, not something a customer should ever read.
What went wrong
One model, tuned for chat, did the research properly. It searched the knowledge base, found the right article, and then called the note tool with something to the effect of having answered the customer. It never called the reply tool. The trace looked busy and plausible, and the visitor got silence.
I did what most people do first, which was to tell it more firmly. The first round of prompt changes spelled out that the reply tool was the only way to reach the customer and that a note is internal. The second round was stricter and more explicit about the order of operations. Neither fixed it. The model would comply for a while and then do the same thing again.
My reading of why is not complicated. A chat-tuned model has been trained, over a very large number of conversations, to treat its final message as the thing the user reads. Give it a reply tool as well and there are now two places a reply could plausibly go. Under that ambiguity it sometimes chose the habit it was trained into, wrote its answer as text, recorded that it had answered, and stopped. Prompting was asking it to override that habit on every single turn. A tool surface that agrees with the habit does not need to ask.
The fix
On chat, the reply tool is gone. The model's final text is the message. There is nothing to forget to call, so the agent cannot research an answer and then fail to send it. The failure did not reappear after the change.
This is the channel where it matters most. Nothing on the first-party chat is reviewed before it is sent, so the grounding rule there is absolute: if no retrieved article contains the answer, there is no answer, and the agent asks diagnostic questions or hands over. Review happens afterwards, in the chat inbox in the console. A silent failure on that channel is a visitor waiting for a reply that was never going to arrive.
Why tickets keep the explicit tool
On helpdesk tickets the agent works in draft mode. Its reply is posted as a private note, and a person on the team copies it, edits it if they want to, and sends it as themselves. In that world the final text of a turn is not a customer message, and treating it as one would be the bug. The reply has to be a deliberate, separately recorded action, because what happens next depends on it: a draft is marked superseded and rewritten if the customer writes again first, and a draft nobody acts on expires after two weeks.
So the two channels have opposite rules for the same idea, and both are correct for their channel. On chat, the natural output of the model is the message. On tickets, the message is an artefact that a person reviews, and it is produced by a call that can be traced.
Designing the surface instead of prompting harder
The general lesson I took from this is that a model's tool list is a design, and the cheapest way to prevent a class of mistake is often to make it impossible rather than to forbid it. A prompt describes what you want. The tool surface decides what can happen.
The same project has other examples of the same move. The direct-message assistant in Slack can only read, by construction, because it has no write tools, not because it was told not to write. Ticket and account identifiers come from the assembled context and never from model output, so a hallucinated ID has nowhere to go. In NEPSE Copilot, my own investing tool, the voice mode has 31 tools and none of them approve an order. Approval stays on the Telegram card and the orders page, where you can see what you are tapping.
When I see an agent misbehave now, my first question is whether the tools allow the mistake. If they do, and the mistake is one the model's training pulls it towards, I would rather change the tools than add another paragraph to the prompt. Prompt hardening still has its place for judgement calls, tone and grounding rules, but I no longer rely on it to keep a model away from an action it can take with a single call.
Drawn from
- AI support employee: An AI member of the support team that takes a case from first message to closed, across tickets, live chat and its own widget
- NEPSE Copilot: A multi-account investing copilot for the Nepal Stock Exchange that grades its own calls