Evaluate Chatty before you launch it to every visitor
Build a small, repeatable evaluation set that catches wrong answers, weak fallbacks, broken links, and poor handoffs before they become public support issues.
A chatbot can look excellent in a demo because the team asks questions it already knows how to answer. Production visitors are less generous: they use old terminology, omit context, ask two questions at once, and expect the answer to lead somewhere.
Before launch, build an evaluation set from real questions and score the experience against the outcomes that matter to your business.
Build the set from evidence
Start with fifty to one hundred questions from:
- Support tickets and chat transcripts
- Sales-call notes
- Search analytics
- Documentation feedback
- Questions your team answers repeatedly
- Deliberate edge cases and out-of-scope requests
Label each question with its intent, source of truth, expected answer, and acceptable fallback. The label makes the test useful when an answer changes later.
Include difficult cases on purpose
Your set should contain more than clean factual questions. Include:
- Paraphrases: “Can I put this on an app?” versus “Do you have mobile support?”
- Follow-ups: “What about the other plan?” after a pricing answer
- Multiple intents: “Can I use SSO and have someone help me migrate?”
- Missing evidence: questions your content does not answer
- Conflicts: an old limit or product name that should not be used
- Sensitive requests: security, privacy, account, and billing questions
These cases show whether the system understands the boundary of its knowledge.
Score the behavior, not just the wording
A useful scorecard separates factual and experiential quality:
| Dimension | Pass condition |
|---|---|
| Grounding | Claims are supported by an approved source |
| Completeness | The answer covers the actual question |
| Actionability | The visitor knows what to do next |
| Clarity | No unexplained jargon or unnecessary length |
| Safety | Sensitive or unknown cases use the right fallback |
| UX | Links, buttons, and handoff paths work |
Use a simple 0/1/2 scale if a large rubric becomes hard to maintain. The important part is consistency between reviewers.
Test links and actions separately
An answer can be correct while the experience is broken. Verify every linked page, CTA, contact form, and handoff notification. Test the actual widget on desktop and mobile, including:
- Opening and closing behavior
- Keyboard navigation
- Long answers and code blocks
- Network failures
- A conversation resumed after reload
- A visitor declining to share contact information
Treat these as release checks, not as optional polish.
Establish a failure taxonomy
When an answer fails, classify the cause before editing anything:
- Content gap: the answer does not exist in the sources
- Retrieval gap: the answer exists but the wrong source was found
- Conflict: multiple sources disagree
- Response gap: evidence was found but explained poorly
- Routing gap: a human or CTA should have been offered
- Product bug: the widget, link, or integration failed
This prevents the common mistake of adding more prompt text to a content problem.
Launch with a review loop
For the first weeks, review a fixed sample of conversations every week. Record:
- Questions with repeated follow-ups
- Answers that required correction
- Missing documentation topics
- Handoffs with insufficient context
- High-intent conversations that did not progress
Turn the recurring patterns into knowledge-base edits, new evaluation cases, or product changes. The evaluation set should grow as the product and the questions grow.
Define a stop condition
Do not launch everywhere simply because the assistant can answer most questions. Set a threshold for the intents you care about and a hard rule for sensitive topics. For example:
Launch when:
- 90% of top-20 intents are grounded and actionable
- 100% of billing and privacy cases use the approved fallback
- No broken links in the release sample
- Handoff summaries contain the visitor's confirmed intent
Evaluation is how you turn “the chatbot seems good” into a decision you can defend. Chatty gives you the conversation layer; your test set and review loop make that layer dependable.