Insights

How do we test whether our member-facing chatbot is giving accurate answers?

Oct 4, 2026, 4:50:10 PM · Rick Bawcum

← All insights

The short answer: test it the way a member would meet it. Ask it the questions members actually ask, from outside your own systems, write down the exact answers, have someone who knows the subject grade them, and repeat the test on a schedule. A chatbot that was accurate at launch is not necessarily accurate today.

Most associations that run a member-facing assistant checked it once, before launch, with questions the project team thought of. That is a reasonable start. It is not evidence about how the assistant behaves now, with real questions, after the vendor's model and your own content have both changed.

Why a launch-day check is not enough

Four things move underneath an assistant after it goes live:

  • The model changes. Many assistants run on a model supplied by a vendor, and the vendor can update it without asking you. We wrote about how much of an association's AI is chosen by vendors rather than by the association.
  • Your content changes. Dues, deadlines, event details, credential requirements and policies all get updated. If the assistant draws on an older copy, it will answer confidently from the older copy.
  • Members ask differently than you expect. They misspell, abbreviate, combine two questions, and ask about edge cases nobody wrote a test for.
  • Nobody is watching. A wrong answer rarely produces a complaint. It produces a member who quietly acts on bad information.

What outside-in testing means

Outside-in testing means interacting with the assistant as a member would: through the public front door, with no access to its settings, prompts or logs. That matters for two reasons. It tests what members actually experience, not what the configuration says should happen. And it is independent of the team that built or bought the assistant, so nobody is grading their own work.

What good evidence looks like

A useful test produces a record you can show a board or an auditor, not a feeling that things seem fine. At minimum:

  1. A question set drawn from real member questions. Start with your inbox, your help desk and your event registration emails. Include the awkward ones: refunds, grievances, eligibility, anything with a deadline.
  2. The exact answers, word for word, with dates. Screenshots or saved transcripts, not summaries.
  3. A grade for each answer by someone who knows the right answer. Correct, incomplete, wrong, or should not have answered.
  4. A check on what it refuses to do. Does it stay in its lane when asked about legal advice, other members' information, or topics outside your scope?
  5. A change log. What changed since the last test, in the model, the content or the configuration.

How often

There is no universal answer, but the logic is straightforward: test often enough that a change cannot sit unnoticed for long. For most associations that means after any announced vendor or content change, plus a regular schedule. Quarterly is a sensible baseline. It is frequent enough to catch drift and infrequent enough to be sustainable.

Common questions

Can our vendor just tell us it has been tested? A vendor can describe its own testing, and that is useful. It is still the vendor describing its own work. An independent test answers a different question: what does a member actually get?

Can we do this ourselves? Yes, and a modest internal test is much better than none. The weak point is independence and consistency: the people who run the assistant tend to ask it the questions it handles well, and internal tests tend to lapse when staff get busy.

What do we do with a wrong answer? Fix the cause where you can, tell the vendor in writing where you cannot, and decide whether any member who received the wrong answer needs a correction. That decision is yours.

Does this replace a policy? No. A policy says what the assistant should do. Testing shows what it does.

What we do not know

We have no reliable data on how often association chatbots give wrong answers, and nobody should quote you a number. The risk is real but its size varies enormously with the tool, the content behind it and the questions asked. That variation is the reason to test your own assistant rather than rely on averages.

Where to start

If you want this done for you, CimAssure is our quarterly outside-in testing of member-facing AI assistants. The report goes to you, not the public, and you decide what to do with it. If you are not sure where AI is running in your organization yet, the AI Assessment is a fixed-price, three-week review that comes first. And the free AI Trust Readiness Scorecard takes about five minutes, with no sign-up to see your result.

Where does your organization actually stand?

Sixteen questions, about five minutes, no sign-up to see your result.