donttrustme.ai
A JOURDANLABS FIELD JOURNALPLANO, TEXAS · EDITION 09.09.2026

THE ARCHIVE / SEPTEMBER 4, 2026

Earlier benchmark
results.

Security and correctness-coverage boards for JavaScript, Python, Go, Java, and TypeScript. F1 followed by TP / FP / FN; separate corpora from the newer C++ results.

Vantage Gate · security

Vantage vs CodeQL vs SonarQube · original edition’s Pan-confirmed, un-blinded board · now leads the Security page

CorpusVantageCodeQLSonarQube
JS+Python (58)0.4894 23/13/350.1818 6/2/520.0667 2/0/56
Go0.9130 42/8/00.1333 3/0/390 genuine zero, 251 ncloc
Java1.0 52/0/00.40 14/4/380.4928 17/0/35
TypeScript1.0 48/0/00.1852 5/1/430.1786 5/3/43

Java circularity disclosure: self-built corpus; CodeQL serves as an independent tool check. This is not a third-party corpus result.

Vantage Code · correctness coverage

Security-tool coverage of a correctness corpus. Low incumbent scores measure incidental coverage, not general tool quality.

CorpusVantageCodeQLSonarQube
JS+Python (67)0.9041 66/13/10.2195 9/6/580.1333 5/3/62
Go (33)0.8919† 33/8/00.1143 2/0/310.0588 1/0/32
Java (33)0.9429† 33/4/00.0930 2/8/310.1143 2/0/31
TypeScript (38)0.6526 31/26/70.186 4/1/340.1 2/0/36

GT-informed development: detectors built with the frozen answer key open, tuned on scorer FN/FP detail; zero GT edits, no GT IDs in engine code. The disclosure travels with the number. The newer, separate C++ correctness board is a decisive loss.