Evaluation archive
We wanted to know
We wanted to know
whether it actually works.
A benchmark here is not a number. It is a dataset, a protocol, a split, a set of leakage controls, a measurement, and a statement about what has and has not been verified. Experiments that fail are recorded with the same care as the ones that succeed.
01
Benchmark timeline
02
Experiments
03
Measurement
04
Reproducibility
Prototype notice
Datasets, protocols and metrics on this page are structured placeholders. They describe the shape of the evaluation record SatQuery intends to publish: dataset identity, split provenance, leakage controls, calibration, and an explicit evidence state. No number here has been produced by a real run.