Reviewing quality
Automation that nobody checks drifts. This page is the routine for checking it, and the mapping from symptom to fix.
The weekly review
Twenty minutes, once a week.
1. Read five AI replies end to end. Not the summaries — the actual replies, in conversations you did not handle. Ask one question about each: would I have sent this?
2. Look at the reopened conversations. A conversation closed automatically and reopened by the customer is a failed answer, stated plainly. Read what the customer wrote back.
3. Check the unclassified pile. Conversations the assistant could not categorise are a map of the holes in your catalog and your knowledge base.
4. Look at the four numbers on Metrics: closing rate, first reply time, AI reply share and positive sentiment.
What the numbers mean together
No single number tells you much. The pairs do.
AI reply share up, sentiment down. The assistant is answering more and answering worse. Narrow the categories it is allowed to answer, or raise the confidence threshold, and find the articles behind the bad answers.
AI reply share up, sentiment steady. This is what working looks like. Widen carefully.
Closing rate up, reopen rate up. Conversations are being closed rather than resolved. Usually automatic close on a category where customers do reply.
First reply time down, everything else steady. The assistant is absorbing the routine questions. This is the main thing you are buying.
Hand-overs rising. Either your customers are asking new things, or your knowledge base has gone stale relative to your product. The unclassified pile tells you which.
Symptom and fix
| What you see | Where the fix is |
|---|---|
| Confident but wrong answers | The article behind them: out of date, ambiguous, or contradicted by another article |
| "I don't have information about that" on things you have documented | The article is there but does not retrieve. Rewrite the title in customer language and put the answer in the first paragraph |
| Right facts, wrong tone | The identity prompt, or a category prompt |
| Right answer to the wrong question | Classification descriptions are too vague, or two categories overlap |
| Replies too long or too short | The identity prompt. Give it a number |
| Answers in the wrong language | The behaviour prompt should say to answer in the customer's language |
| Too many hand-overs | The knowledge base is too thin. The unclassified conversations name the missing articles |
| Too few hand-overs, and bad answers | Raise the confidence threshold |
The pattern: facts live in the knowledge base, tone lives in prompts, routing lives in classifications. Fixing a fact in a prompt or a tone problem in an article is the most common wasted afternoon.
Correcting as you work
Two habits cost nothing and compound:
Fix classifications when you see them wrong. A corrected classification becomes a verified example that weighs more than an AI-classified one.
Write the article when you answer something twice. The second time you type the same explanation by hand is the signal, and it takes five minutes.
When to widen
Widen automation on a category when, for a fortnight:
- You would have sent the drafts unchanged.
- No customer wrote back confused after an automatic reply.
- Sentiment on that category is steady.
Widen one category at a time, so that when something goes wrong you know what changed.
When to pull back
Turn a category back to draft immediately if a wrong answer reaches a customer on something that matters — money, safety, a legal commitment. Drafts cost you a few minutes a day and remove the whole class of problem while you fix the article behind it.