Attested benchmarks for models and the stacks they run on
Every score below comes from a run whose log is Ed25519-signed and publicly verifiable. Follow any row through to its manifest and check it yourself with @divinci-ai/trustbench-verifier — no account, no API key, no trust in us required. What a benchmark has to prove about itself explains what the manifest does and does not claim.
Loading boards from the public API…
What these boards do not claim
- Attestation says this score is real. It does not say this model is safe. Evidence, not a guarantee.
- The retrieval boards' runs are
republishedrather thanmeasured: the scores were computed by Divinci's scored-QA pipeline and republished into a signed TrustRun. The verifier warns you about exactly this, and the manifests say so themselves. - Several platform benchmarks are too small to rank on and are not published here. A benchmark with one sample can only ever score 0.0 or 1.0.
- The retrieval boards use one answering model and one judge. On the nutrition board the Vertex row searches a newer, larger ingestion than the others.