Candidates that behave differently don't get to report success.
An agent writes several candidate patches. They run in isolated sandbox branches against the same inputs. Saved-test violators are excluded. Disagreement among the remaining candidates withholds approval, and the exact input is returned to the user as a question. The user's answer becomes a versioned acceptance test. Every verdict is computed by deterministic code; the language models only propose and explain.
What a PASS means here: with all gates enabled, PASS means the remaining candidates passed regression and showed no disagreement on the sampled probes. Other inputs remain unverified, including unsampled inputs inside the stated domain, and are marked like this throughout. It is not a proof of correctness.
The verdict below was not recorded. It was computed when this page loaded.
Your browser read the execution trace and scored it with a JavaScript port of the Python scorer. If the port is faithful, the hash of the decision must match the one Python produced — byte for byte.
computing…
Provenance fields (timestamps, node ids, durations) are excluded from the hash by design; they differ on every run and carry no decision.
Clarification sessions
Each session is one task, one user, up to five rounds. Replay any of them; you can pick a different answer than the one recorded and see where the recorded path diverged.