knowngrounds
The public index

What AI gets right, and wrong, about real pages

Every run here asked one model the same questions twice — closed book from memory, then open book with search and page fetching. How this works.

Open book
84%
Closed book
44%
Answers graded correct, across every public run. 40 points of that accuracy exists only while the model keeps fetching the page — open book minus closed book.
14
runs
13
pages
1,618
checks graded
for a free run
Every public run — open to read
Search them, open any report. Sign in to check a page of your own — the first run is free, and there is no password.
Run a check →
14 runs · every claim tested against a frozen question set · page 1 of 2
actionable retrieval-dependent sound not gradeable what do these mean?
Accept a payment
https://docs.stripe.com/payments/accept-a-payment
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 9 claims 28 checks
On the pageView report →
open book77.8%closed book77.8%
2
misrepresented
of 9 claims
0%
page surfaced
for its own search
$0.26
cost
3m01s
Create claimable sandboxes
https://docs.stripe.com/sandboxes/claimable-sandboxes
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 9 claims 24 checks
On the pageView report →
open book88.9%closed book11.1%
1
misrepresented
of 9 claims
100%
page surfaced
for its own search
$0.26
cost
2m16s
Stripe Data Pipeline | Sync Stripe Data to Your Data Warehouse
https://stripe.com/data-pipeline
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 29 claims 228 checks
On the pageView report →
open book79.3%closed book44.8%
6
misrepresented
of 29 claims
42.9%
page surfaced
for its own search
$1.37
cost
6m02s
Meet Stripe's Knowledge AI Platform
https://stripe.dev/blog/meet-stripes-knowledge-ai-platform
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 29 claims 228 checks
On the pageView report →
open book72.4%closed book17.2%
8
misrepresented
of 29 claims
17.2%
page surfaced
for its own search
$1.28
cost
6m10s
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 11 claims 120 checks
On the pageView report →
open book81.8%closed book72.7%
2
misrepresented
of 11 claims
0%
page surfaced
for its own search
$0.90
cost
4m21s
Calling dibs on DIBS · Lyncredible
https://lyncredible.com/2023/10/30/calling-dibs-on-dibs/
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 10 claims 90 checks
On the pageView report →
open book80%closed book60%
2
misrepresented
of 10 claims
70%
page surfaced
for its own search
$0.45
cost
2m02s
Kimi K3 Tech Blog: Open Frontier Intelligence
https://www.kimi.ai/blog/kimi-k3
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 10 claims 30 checks
On the pageView report →
open book90%closed book10%
1
misrepresented
of 10 claims
90%
page surfaced
for its own search
$0.38
cost
2m46s
Calling dibs on DIBS · Lyncredible
https://lyncredible.com/2023/10/30/calling-dibs-on-dibs/
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 10 claims 30 checks
On the pageView report →
open book80%closed book60%
2
misrepresented
of 10 claims
70%
page surfaced
for its own search
$0.41
cost
3m03s
How disputes work
https://docs.stripe.com/disputes/how-disputes-work
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 48 claims 212 checks
On the pageView report →
open book91.7%closed book41.7%
4
misrepresented
of 48 claims
93.8%
page surfaced
for its own search
$1.31
cost
6m25s
Receive payouts
https://docs.stripe.com/payouts
tested openai/gpt-5.6-luna graded openai/gpt-5.6-terra set 1 47 claims 336 checks
On the pageView report →
open book100%closed book51.1%
0
misrepresented
of 47 claims
89.4%
page surfaced
for its own search
$1.73
cost
8m45s