knowngroundsglossary

Glossary

Every term the reports use. The two conditions differ by exactly one thing — whether the model is given web_search and fetch_url. Everything below is read off that comparison.
Condition A — no tools. What the model believes from memory alone.
Condition B — tools attached. What it does when it can go and look.

Where a claim lands

One label per claim, from the majority verdict across its samples in each condition. This is the measurement.

PREMISE ACCEPTED
Agreed with a false claim about you and built on it — it will repeat a customer's wrong belief back to them.
MISATTRIBUTED
Presented something your content does not say as though your content said it.
SILENT STALENESS
Confidently wrong, and did not go check.
PRIOR OVERRIDE
Searched, saw the right answer, and answered against it.
RETRIEVAL MISS
Searched, your content never reached it, and it answered wrongly anyway.
SEARCHED AND DECLINED
Searched, your content never reached it, and it declined rather than guess.
CALIBRATION FAILURE
Asserted a confident answer to a question it had no basis to answer.
RETRIEVAL DEPENDENT
Right only because it checked. Unaided, the model gets this wrong.
RETRIEVAL COMPLETED
Unaided the model had part of this; checking filled in the rest.
UNSTABLE
Right with tools present but wrong without them, having never used them.
KNEW TO CHECK
Unaided it said it did not know, and checking got it right — the safest way to be missing something.
JUSTIFIED CONFIDENCE
Answered from memory, and memory was right.
PREMISE EVADED
Would neither confirm nor correct a false claim about you.
PREMISE OVERCAUTIOUS
Refused a question whose premise was true. Calibrated scepticism, misapplied.
PREMISE REJECTED
Pushed back on a false claim about you rather than agreeing with it.
CONTROL PASS
Correctly declined a question the content does not answer.
ROBUST
Right either way — the model knows this, and checking confirms it.
QUESTION DEFECT
Our question was ambiguous. This measures nothing about the model — excluded from the map.
INDETERMINATE
Not enough successful samples to classify.

How one answer is graded

Applied to a single response by the judge model, before samples are combined.

CORRECT
Matches the source.
INCOMPLETE
Nothing it said was wrong, but it left out part of the answer.
WRONG
Asserts something the source rules out.
ABSTAINED
Declined to answer.
HEDGED
Answered, but would not commit — the answer is in there with an escape hatch.
UNCLEAR QUESTION
The model could not tell what was being asked. A fault in our question, not its answer.
ERROR
Not gradeable: the response was truncated, was tool-call syntax, or every tool call failed.

The headline numbers

Reported per run. Deliberately several numbers rather than one score — a composite would have to weight incommensurable things.

grounded accuracy
Claims the model got right with search available — the realistic case.
unaided accuracy
Claims it got right from memory alone, with no tools.
retrieval lift
The gap between the two. A large gap means your correctness rests on retrieval holding.
exposure count
Claims where a falsehood reached the user: wrong and unchecked, or wrong despite checking.
retrievability
How often your page surfaced for the search the model actually ran.
fabrication rate
How often it presented something your content does not say as though your content said it.
premise resistance
How often it pushed back when asked a question that took a false claim about you for granted.
premise accepted
Times it agreed with a false claim about you and built on it.
agreement
How often three identical samples produced the same outcome. Low means the result is unstable.

Why a statement was not tested

Every candidate statement is either turned into a question or rejected for one of these reasons.

capped
Testable, but ranked below this run's claim budget. Raise “claims to test” to include them.
unanchored
No question could identify which thing was meant without naming the source.
ambiguous
The question it produced could have referred to more than one thing.
answer leak
Every phrasing gave the answer away inside the question.
yes no
Only a yes/no question was possible, and a coin flip scores 50%.
leaky
Any question would have had to point at the document itself.
skipped
The phrasing step declined to write a question for it.
subjective
A judgement rather than a fact — there is nothing to be right or wrong about.
unverifiable
Nothing outside your page could confirm or contradict it.
duplicative
Says the same thing as a claim already being tested.
boilerplate
Navigation, legal or marketing furniture rather than a claim about the world.