All articles
Evaluation·September 14, 2026·10 min read

Evaluate Chatty before you launch it to every visitor

Build a small, repeatable evaluation set that catches wrong answers, weak fallbacks, broken links, and poor handoffs before they become public support issues.

A chatbot can look excellent in a demo because the team asks questions it already knows how to answer. Production visitors are less generous: they use old terminology, omit context, ask two questions at once, and expect the answer to lead somewhere.

Before launch, build an evaluation set from real questions and score the experience against the outcomes that matter to your business.

Build the set from evidence

Start with fifty to one hundred questions from:

  • Support tickets and chat transcripts
  • Sales-call notes
  • Search analytics
  • Documentation feedback
  • Questions your team answers repeatedly
  • Deliberate edge cases and out-of-scope requests

Label each question with its intent, source of truth, expected answer, and acceptable fallback. The label makes the test useful when an answer changes later.

Include difficult cases on purpose

Your set should contain more than clean factual questions. Include:

  1. Paraphrases: “Can I put this on an app?” versus “Do you have mobile support?”
  2. Follow-ups: “What about the other plan?” after a pricing answer
  3. Multiple intents: “Can I use SSO and have someone help me migrate?”
  4. Missing evidence: questions your content does not answer
  5. Conflicts: an old limit or product name that should not be used
  6. Sensitive requests: security, privacy, account, and billing questions

These cases show whether the system understands the boundary of its knowledge.

Score the behavior, not just the wording

A useful scorecard separates factual and experiential quality:

DimensionPass condition
GroundingClaims are supported by an approved source
CompletenessThe answer covers the actual question
ActionabilityThe visitor knows what to do next
ClarityNo unexplained jargon or unnecessary length
SafetySensitive or unknown cases use the right fallback
UXLinks, buttons, and handoff paths work

Use a simple 0/1/2 scale if a large rubric becomes hard to maintain. The important part is consistency between reviewers.

Test links and actions separately

An answer can be correct while the experience is broken. Verify every linked page, CTA, contact form, and handoff notification. Test the actual widget on desktop and mobile, including:

  • Opening and closing behavior
  • Keyboard navigation
  • Long answers and code blocks
  • Network failures
  • A conversation resumed after reload
  • A visitor declining to share contact information

Treat these as release checks, not as optional polish.

Establish a failure taxonomy

When an answer fails, classify the cause before editing anything:

  • Content gap: the answer does not exist in the sources
  • Retrieval gap: the answer exists but the wrong source was found
  • Conflict: multiple sources disagree
  • Response gap: evidence was found but explained poorly
  • Routing gap: a human or CTA should have been offered
  • Product bug: the widget, link, or integration failed

This prevents the common mistake of adding more prompt text to a content problem.

Launch with a review loop

For the first weeks, review a fixed sample of conversations every week. Record:

  • Questions with repeated follow-ups
  • Answers that required correction
  • Missing documentation topics
  • Handoffs with insufficient context
  • High-intent conversations that did not progress

Turn the recurring patterns into knowledge-base edits, new evaluation cases, or product changes. The evaluation set should grow as the product and the questions grow.

Define a stop condition

Do not launch everywhere simply because the assistant can answer most questions. Set a threshold for the intents you care about and a hard rule for sensitive topics. For example:

Launch when:
- 90% of top-20 intents are grounded and actionable
- 100% of billing and privacy cases use the approved fallback
- No broken links in the release sample
- Handoff summaries contain the visitor's confirmed intent

Evaluation is how you turn “the chatbot seems good” into a decision you can defend. Chatty gives you the conversation layer; your test set and review loop make that layer dependable.

P
PersonaliAI Team
Building useful AI customer experiences with Chatty.

Put Chatty to work on your customer experience

Train Chatty on your content and give visitors a useful first answer.

Get started free