A supervised AI support agent grounded in real customer data
Published Sep 24, 2026Reviewed Oct 7, 2026
An LLM agent that drafts support replies and account actions from each customer's real purchase history, with per-brand autonomy levels, structural checks in code and an evaluation system that decides which prompt ships.
At a glance
- Problem
- A consumer subscription business selling under several brands handled a high volume of multilingual support tickets by hand: cancellations, billing questions and access problems, each depending on what the customer had actually bought.
- Constraint
- Replies and actions touch real subscriptions and real money, so the agent could not guess who a customer is, act on stale purchase data, or run on a prompt change nobody had measured.
- What I built
- A Zendesk-connected LLM agent that reads purchase history and a versioned knowledge base, proposes a reply plus typed actions, passes structural checks in code, and runs at a per-brand autonomy level under operator supervision.
- Result
- Operators review exact drafts instead of writing from scratch, automation widens brand by brand, and prompt changes are judged against a reviewed test catalog. A census of thousands of unmatched tickets changed the agent's behavior: it now asks for the one missing detail.
Overview
The business runs several consumer subscription products under separate brands, with customers writing in from many countries and languages. Most tickets are variations of a few requests: cancel my subscription, why was I charged, I can't log in. The right answer depends on facts the customer rarely states correctly: which brand they bought, which plan, when it renews, whether it was already canceled.
A fluent reply written without those facts does more harm than no reply. The brief was an agent whose every reply and action rests on the customer's real commerce record, whose autonomy is set per brand, and whose quality is measured before a change reaches customers.
Identity is a ladder of evidence
Zendesk tickets are mirrored into Postgres, and each requester is matched to commerce records exported from BigQuery. Resolution is a ladder tried in order: help-desk user ID, email, a quoted account code, an external purchase reference, then order or subscription identifiers quoted in the ticket. One person can own several accounts, so an email may resolve up to 5 records, or up to 50 when the set is anchored to the requester's own help-desk ID. A set above its cap becomes a candidate pool, and a later rung must resolve uniquely inside it. Brand is not identity proof: filtering by brand to force a unique match would discard the same person's purchases under other brands.
Customers often quote a second email or an order number in the message. Each quoted identifier is probed separately: a new unique hit links automatically under its own evidence, and an ambiguous probe adds nothing. An earlier version held new matches for operator review; that hold is gone, and held requesters were reopened.
Changing a matching rule has a fixed bar. A backtest needs 600 labeled hits with no disagreement, or 1,500 with at most one, plus 200 adversarial hits with none.
Freshness is a fact of its own. Every commerce snapshot carries a "covered through" time, and missing or stale coverage cannot prove that a purchase doesn't exist. When purchase data is refreshing, the agent waits and rechecks before acting.
The census of unmatched tickets
The weakest replies said, in effect, "we couldn't find your account." A census of several thousand production tickets with no matched purchase showed why better matching would not fix them. About two in five contained no purchase language at all, so leaving them unmatched was correct. About one in five explicitly claimed a purchase but carried no order number, no alternate email and no other identifier. No matching rule can reach those tickets; the only way to link them is to ask.
So when identity is unresolved and the customer describes a purchase, the reply asks for one detail: the email used at purchase, or an order or invoice number. The customer's answer arrives as a new comment, enrichment runs again with the new candidate, and the email and identifier rungs get a second chance.
For about a week this was a code rule that rejected any reply reporting the failure without asking. It was removed with the other rules that read customer text (next section). The ask now lives in the managed prompt, and reviewed catalog cases grade it: first asks, follow-up asks, and replies that must never claim the account doesn't exist.
The model proposes, code checks the structure
The model sends nothing. It returns a proposal: typed actions plus the ticket status that should follow.
type ProposalAction =
| { type: "send_reply"; body: string; language: string }
| {
type: "cancel_subscriptions";
clientId: string; // an account linked to this ticket
subscriptionIds: string[]; // from that account's facts
scope: "cancel_specific" | "cancel_all_identified";
evidence: string[]; // the facts the model relied on
}
| { type: "apply_tags"; tags: string[] }
| { type: "internal_note"; body: string }
| { type: "escalate_human"; reasonTag: string; note?: string };
Simplified from the real schema
The output schema is built per ticket from its facts. Tags, intents and escalation reasons are enums from the tag catalog, and cancellation targets are enums of the subscription IDs the facts allow. When nothing is cancellable, the cancel action is absent from the schema. Generation runs at temperature 0.
After generation, a gate checks the proposal against the facts. A cancellation must name an account linked to the ticket, a subscription from that account's facts, and exactly one brand. Links must come from the brand's knowledge base or stored values. Quoted currencies must match what the customer is billed in. No template variable may reach a customer, and only permitted tags may be applied. An English reply may not claim a completed cancellation unless the proposal cancels or the facts show an earlier one. A rejected proposal goes back to the model with the specific violations, at most twice. After that, it is replaced by an escalation to a human specialist.
Taking rules out of the customer's words
Earlier versions also had rules that read the customer's text: keyword nets for cancel intent, scope and language. On one catalog version, 8 of 13 persistent evaluation failures were forced by those rules. The keyword check disagreed with the reviewed answer, and its revision feedback pushed the model away from it.
The fix was a contract: no check may depend on comment content. The production facts module lost 69 regexes. A later pass removed checks that re-derived decisions the model had already made, because a false failure replaces a correct answer with an escalation. Judging intent belongs to the evaluation system, where it can be measured.
Autonomy and execution
Each brand runs at one of three levels: off, draft for review, or draft and act. Operators see the exact reply, actions, cancellation targets and resulting ticket status before approving. Replies go out in the customer's language, with an English translation shown for review only. A global switch pauses execution while drafting continues, and "should not reply" suppresses automatic replies until the customer writes again.
Execution rechecks that the requester hasn't changed, no newer customer message has arrived, and the ticket isn't already solved. Each action reserves a row by dedupe key with a 30-second lease, so a repeated request returns the recorded outcome instead of replying or canceling twice. If only the status change failed, operators retry that step alone; completed replies and cancellations stay as they are.
Evaluations decide what ships
Prompts are saved as immutable versions. A brand can pin an exact version, follow the newest version of a named prompt, or route by the customer's products.
The evaluation set is a reviewed catalog of real tickets, captured with conversation, identity and commerce state and pseudonymized. Expected outcomes belong to the catalog and outlive prompt versions; only an independent LLM judge sees them. Cases with missing or stale source facts stay unverified and cost no model calls.
const passRate = passed / (passed + failed); // quality of what was judged
const coverage = (passed + failed) / selected; // how much was judged at all
An unverified case is not a pass
Each run keeps its full input bundle (prompt, knowledge, models, judge setup, answer key), so later edits cannot rewrite it. Promotion evidence is two complete passing runs with matching inputs, the same backend version and pinned providers. An administrator can still assign a prompt without that evidence; the missing evidence shows as warnings and never as a pass.
The catalog grows from operator work. An edited draft can become a suggested case in a golden queue and, once reviewed, heads toward the next catalog version.
Limitations
- The business metric is ticket resolution rate; the evaluation pass rate measures drafts, and a good draft can still leave a customer unresolved.
- Structural checks catch wrong targets, links and currencies; tone and judgment depend on the prompt, the judge, the catalog and human review.
- With the identifier ask moved from code to the prompt, a prompt regression can drop it, and only the catalog will catch that.
- An LLM judge needs calibration against reviewed examples; hermetic tests prove the plumbing and say nothing about the judge's accuracy.
- Unattended actions depend on upstream freshness; when purchase data lags, automatic work waits.
Related case studies
One AI chat assistant service for several brand apps
Several apps share one AI chat assistant, and an LLM judge reviews every conversation.
Help-desk ingestion that proves nothing went missing
Every support ticket reaches the database; missing ones are found and repaired automatically.
