How the system checks itself

Three separate programs. The runner executes and records raw observations — never a verdict. The scorer is a pure function of the trace. The grader holds reference tests the gate is not allowed to read, and plants an unpredictable string that must never appear anywhere in the gate's output tree. The 49 runs in the canary table were scanned. The recorded canary checks found no canary string in the scanned outputs. This is a bounded leak check, not proof of complete isolation.

Same candidates, different execution environments

Runs grouped by identical task, job manifest, candidate hashes and probe seed — differing only in where the code ran. The decision hash must not depend on the machine.

Answer-key isolation

After each run the grader writes a random string into its own directory and searches every file the gate produced. A hit invalidates the run.

runfiles scannedpathsresult

What this site can and cannot show

The traces served here are stripped of stdout, timestamps and durations to keep them small. The scorer was run on both the full and the stripped trace at build time; the decision hashes matched for every run listed in the margin. Reference tests and their inputs are not served by this site: the deployment publishes the site/ directory alone, and only per-candidate pass counts appear here. The tests are in the source repository, so the grading can be re-run by anyone. The scorer does not load the reference tests. The recorded canary checks found no leak in the scanned outputs.

Re-score every run in your browser and compare hashes

Execution sandboxes (ConTree) isolate filesystem and process state per branch. Network egress inside the sandbox was measured, not assumed; see the run header field egress_probe_result. Where it reads EGRESS_OK, the sandbox had network access.