evalseal
ASSAY OFFICE FOR AI EVALS

Independent record · no rankings, no grades

“Trust us”
isn't proof

Nearly every AI capability claim is issued by the party that benefits from it, and almost none can be checked. EvalSeal freezes one evaluation run into a publicly verifiable seal — dataset hash, model build, sampling conditions, every trace, all signed. We publish no score of our own. We make yours possible to audit, contest, and overturn.

Sealed Evaluation Record Ed25519 · Auditable D291·7CAF evalseal-demo-12

Marks struck on the seal — open any one

Dataset

evalseal-demo-12 v1.0.0 · 12 tasks · sha256:30a3b8694bf6…

The task set is frozen and hashed before the run. Swap or edit a single task and the hash stops matching.

Why scores
don't hold up

All four happen daily, and none of them require anyone to cheat. The default evaluation process simply does not produce evidence.

Contamination

The test set is in the training set

Public benchmarks start leaking the day they ship. Scores climb month over month while real capability stays flat. Without a frozen hash and a release date, you cannot tell whether the model improved or memorized.

Irreproducible

The same suite gives a different number

Seed, temperature, prompt version, tool set — miss any one and the second run is a different experiment. Most evaluation reports don't even have a field to record them in.

Drift

The version in the report isn't the one you call

The report says the June build. The endpoint has swapped weights twice since. The model name never changed, the behaviour did, and the PDF in your hand will never update.

Single party

The referee is also the contestant

Vendors run their own evals, pick their own subsets, write their own reports. Not necessarily dishonest — but structurally impossible to refute. Anything that cannot be refuted is not evidence.

Now it goes on file

Until recently, how an evaluation got written down was a matter of taste. In the EU it is becoming a matter of record-keeping law, and it binds two parties: the provider who has to keep the file, and the outside evaluator whose findings go into it.

2 Aug 2026The AI Office can levy fines against providers of general-purpose models. The obligations themselves have applied since August 2025 — what arrives now is the machinery to enforce them.
Annex XIProviders of models with systemic risk must document their evaluation strategies including the results, alongside the internal and external adversarial testing performed. The file is kept for ten years.
EvaluatorsWhere an outside evaluator is engaged, its findings are passed to the regulator — and the evaluator has to be able to evidence its own independence, not merely assert it. That is what a grade on the face of a record is for.
ScopeThis binds today for general-purpose models above the systemic-risk threshold. Obligations for high-risk systems are being pushed to December 2027 under the agreed omnibus revision: the deadline moved, the documentary standard did not.

A report written afterwards cannot show that it was not written afterwards. That is the whole of what a seal adds — the record is fixed at the moment of the run, and any later edit is arithmetically obvious.

One seal in full

SEAL D291-7CAF-C9B1 · seal no. 1 · publicly verifiable
FinenessGrade A · independently executed subject had no access
Runevalseal-demo-12 v1.0.0 · 12 tasks · single pass
Subjectsubject.py · static reference implementation, not a language model
Datasetsha256:30a3b8694bf6c3b2 frozen
Conditionsdeterministic · no sampling, no network, no clock
Result8 / 12 passed · 66.7% (±0.0, deterministic)
Traces12 records · 1743 bytes · sha256:52b14610116879a3
Environmentpython 3.14.3 · macOS-15.7.5-x86_64 · harness sha256:1bcf588d
Sealed at2026-07-28T09:14:02Z
Signaturessh-ed25519 (SSHSIG) · key SHA256:PY0pImlTDg0waBeI… valid

Three grades of fineness

A hallmark states purity, not praise. The first thing a seal states is who executed the run — that, not the score, is what decides how much a record is worth.

A · IndependentEvalSeal provisions the environment, executes the run, and captures every trace. The subject never touches it.
B · WitnessedThe run executes in a sealed environment we provision, on the subject's own keys and hardware. Inputs and outputs are captured by us, not reported by them.
C · SubmittedTraces are submitted by the subject. The seal attests integrity and time only: this record has not changed since that moment. It does not attest that the run happened as described.

A grade C seal is not a grade B seal, and it says so on its face. Nobody has to guess what a seal is worth.

How a seal is struck

01 / Intake

Submit

Point us at your eval suite and dataset. Both are hashed and frozen. Any change after that produces a new seal number — the old seal is never overwritten.

02 / Testing

Assay

The run executes at the grade you requested: operated by us end to end (A), or in a sealed environment we provision on your own keys (B). Either way every call's input and output is captured, and the runtime image is hashed.

03 / Marking

Hallmark

The seal is issued, signed, and given a public address. We publish no verdict, no ranking, no pass mark — we pin the run somewhere it can be checked, and leave the rest to anyone who wants to dispute it.

A conclusion that can be overturned is the only kind worth believing. EvalSeal exists to make overturning it possible.