Browse the Help Centre

How we measure whether it is right

Most support assistants cannot tell you how often they are correct. We measure it, on a fixed benchmark, before every release, and we run the same measurement on your documentation during setup.

What we measure

Did it find the right article? For each benchmark question we know which article contains the answer, so we can check whether that article appeared in the results at all, and how highly it ranked.

Did the answer convey the right points? A separate model compares the answer against a known-correct answer and reports how much of it was covered, and whether anything contradicts it. Those two combine into a single factuality score, so both saying something false and saying nothing useful cost you, while extra correct detail costs nothing.

What did the visitor actually receive? Answers that were withheld score zero, so a system that stays silent cannot score well by refusing.

How the grading avoids flattering itself

The grading model is not shown the passages the assistant retrieved. This matters more than it sounds. Shown the source documents, a grader drifts into asking "is this answer consistent with these documents" instead of "does this answer the question", and it approves almost everything — including answers to questions whose correct article was never retrieved at all.

Graded against the known-correct answer alone, the same set of answers separates cleanly. It is a harsher measurement and the only useful one.

The numbers from our own benchmark

On a 6,221-article knowledge base with 200 questions: factuality 0.85, the correct article in the top five results 84% of the time, and answers withheld about 1% of the time.

Those are our numbers on our benchmark. Yours will differ, because your documentation is different — which is the point of running it on yours.

What we run on your documentation

During setup we run the same measurement against a sample of your real questions and send you the results, along with the specific questions it answered badly or could not answer at all. You keep that report. See Getting your help centre ready.

What these numbers are not

They are not a deflection rate, and they are not a prediction of how many tickets you will avoid. That depends on your traffic, your documentation and your customers, and we will not quote you a figure for it before we have measured yours.