Skip to main content

Reviewing quality

Automation that nobody checks drifts. This page is the routine for checking it, and the mapping from symptom to fix.

The weekly review

Twenty minutes, once a week.

1. Read five AI replies end to end. Not the summaries — the actual replies, in conversations you did not handle. Ask one question about each: would I have sent this?

2. Look at the reopened conversations. A conversation closed automatically and reopened by the customer is a failed answer, stated plainly. Read what the customer wrote back.

3. Check the unclassified pile. Conversations the assistant could not categorise are a map of the holes in your catalog and your knowledge base.

4. Look at the four numbers on Metrics: closing rate, first reply time, AI reply share and positive sentiment.

What the numbers mean together

No single number tells you much. The pairs do.

AI reply share up, sentiment down. The assistant is answering more and answering worse. Narrow the categories it is allowed to answer, or raise the confidence threshold, and find the articles behind the bad answers.

AI reply share up, sentiment steady. This is what working looks like. Widen carefully.

Closing rate up, reopen rate up. Conversations are being closed rather than resolved. Usually automatic close on a category where customers do reply.

First reply time down, everything else steady. The assistant is absorbing the routine questions. This is the main thing you are buying.

Hand-overs rising. Either your customers are asking new things, or your knowledge base has gone stale relative to your product. The unclassified pile tells you which.

Symptom and fix

What you seeWhere the fix is
Confident but wrong answersThe article behind them: out of date, ambiguous, or contradicted by another article
"I don't have information about that" on things you have documentedThe article is there but does not retrieve. Rewrite the title in customer language and put the answer in the first paragraph
Right facts, wrong toneThe identity prompt, or a category prompt
Right answer to the wrong questionClassification descriptions are too vague, or two categories overlap
Replies too long or too shortThe identity prompt. Give it a number
Answers in the wrong languageThe behaviour prompt should say to answer in the customer's language
Too many hand-oversThe knowledge base is too thin. The unclassified conversations name the missing articles
Too few hand-overs, and bad answersRaise the confidence threshold

The pattern: facts live in the knowledge base, tone lives in prompts, routing lives in classifications. Fixing a fact in a prompt or a tone problem in an article is the most common wasted afternoon.

Correcting as you work

Two habits cost nothing and compound:

Fix classifications when you see them wrong. A corrected classification becomes a verified example that weighs more than an AI-classified one.

Write the article when you answer something twice. The second time you type the same explanation by hand is the signal, and it takes five minutes.

When to widen

Widen automation on a category when, for a fortnight:

  • You would have sent the drafts unchanged.
  • No customer wrote back confused after an automatic reply.
  • Sentiment on that category is steady.

Widen one category at a time, so that when something goes wrong you know what changed.

When to pull back

Turn a category back to draft immediately if a wrong answer reaches a customer on something that matters — money, safety, a legal commitment. Drafts cost you a few minutes a day and remove the whole class of problem while you fix the article behind it.