The test set is in the training set
Public benchmarks start leaking the day they ship. Scores climb month over month while real capability stays flat. Without a frozen hash and a release date, you cannot tell whether the model improved or memorized.
Independent record · no rankings, no grades
Nearly every AI capability claim is issued by the party that benefits from it, and almost none can be checked. EvalSeal freezes one evaluation run into a publicly verifiable seal — dataset hash, model build, sampling conditions, every trace, all signed. We publish no score of our own. We make yours possible to audit, contest, and overturn.
Marks struck on the seal — open any one
Dataset
evalseal-demo-12 v1.0.0 · 12 tasks · sha256:30a3b8694bf6…
The task set is frozen and hashed before the run. Swap or edit a single task and the hash stops matching.
All four happen daily, and none of them require anyone to cheat. The default evaluation process simply does not produce evidence.
Public benchmarks start leaking the day they ship. Scores climb month over month while real capability stays flat. Without a frozen hash and a release date, you cannot tell whether the model improved or memorized.
Seed, temperature, prompt version, tool set — miss any one and the second run is a different experiment. Most evaluation reports don't even have a field to record them in.
The report says the June build. The endpoint has swapped weights twice since. The model name never changed, the behaviour did, and the PDF in your hand will never update.
Vendors run their own evals, pick their own subsets, write their own reports. Not necessarily dishonest — but structurally impossible to refute. Anything that cannot be refuted is not evidence.
Until recently, how an evaluation got written down was a matter of taste. In the EU it is becoming a matter of record-keeping law, and it binds two parties: the provider who has to keep the file, and the outside evaluator whose findings go into it.
A report written afterwards cannot show that it was not written afterwards. That is the whole of what a seal adds — the record is fixed at the moment of the run, and any later edit is arithmetically obvious.
A hallmark states purity, not praise. The first thing a seal states is who executed the run — that, not the score, is what decides how much a record is worth.
A grade C seal is not a grade B seal, and it says so on its face. Nobody has to guess what a seal is worth.
Point us at your eval suite and dataset. Both are hashed and frozen. Any change after that produces a new seal number — the old seal is never overwritten.
The run executes at the grade you requested: operated by us end to end (A), or in a sealed environment we provision on your own keys (B). Either way every call's input and output is captured, and the runtime image is hashed.
The seal is issued, signed, and given a public address. We publish no verdict, no ranking, no pass mark — we pin the run somewhere it can be checked, and leave the rest to anyone who wants to dispute it.
A conclusion that can be overturned is the only kind worth believing. EvalSeal exists to make overturning it possible.