CLAIMS / TESTED, NOT TRUSTED
Other people’s claims.
Same rules.
This site started as a place to publish our own numbers with their losses attached. It tests other people’s now, too — a launch page, a post on X, a vendor’s bullet point — from the outside, with the protocol frozen before the first query and everything published whole: holds, does not hold, and whatever our own instrument got wrong on the way. Nobody asks us to. The subject gets no edit and no veto. When a claim holds, we say so at the same weight.
THE REGISTER
Two so far.
ASSAY-001: Jev, checked
TypeSafe AI’s launch claims — “calibrated probabilities,” “never makes type errors.” Pre-registered, run once, re-scored blind on a different model. Half holds: calibrated on one corpus of two; zero type errors in 8,576 responses.
A stealth model, checked
A post on X: a free stealth model “matches GPT-6 Astra at 18x lower price.” Tested the same day. Performance holds on our suite; the price claim expired by nightfall. Our ruler had eight defects, found by us and by a second reviewer, all published with the raw responses.
HOW ONE GETS HERE
The rules, in order.
The claim is quoted verbatim from where it was made. The protocol — corpora, metrics, pass criteria, budget — is written and hashed before any query. Positive controls must go red and green before a real number is computed. One run, no retries, no prompt tuning. Every request and response is logged and published. Where our own scoring is involved, a second party on a different base model re-scores or tries to break it, and what they find is printed as a numbered amendment. Then the result ships whether it flatters the subject, us, or neither.
Want a claim of yours run this way before someone else runs it? That is the paid service: assay.jourdanlabs.com ↗. The free version is the method itself, open source: assay-kit ↗ — run it on your own claim and the result is stamped SELF-ASSAYED. Only an independent party can make it ASSAY-VERIFIED.