Fixed rules, read from the record.
Every round is recorded step by step. A fixed set of rules reads that record and turns it into one score and four ratings. No AI model judges anyone, and the same round always gets the same score.
Last changed 28 September 2026.
The score
Each of the four questions below gets a rating, and each rating is worth a fixed number of points:
The score is the weighted average of those points. Only the questions we could measure count. If a round gives us nothing to measure for one question, its weight moves to the others, so a score is never made up from missing data.
The four questions
Do they verify the agent's work?
30 OF 100This counts only decisions a person made. Approvals given by the auto-accept switch are not decisions and never count.
- A snap approval is a change accepted in under a second.
- A blind approval is an accepted change to a file the candidate never opened.
- Pushback is rejecting a change, or naming the bug we planted before the fix lands.
- Standout: no snap or blind approvals, and at least one pushback.
- Strong: no blind approvals, and snap approvals on a quarter of decisions or fewer.
- Flag: any blind approval, or snap approvals on more than half of decisions.
- Mixed: everything else.
Does the work hold up?
30 OF 100At the end we run the tests the candidate could see, and a second set they never saw, called the hidden checks. If the hidden checks pass, a change to a visible test was fair and does not count against anyone.
- Standout: the visible tests and the hidden checks both pass.
- Flag: the visible tests fail, or the hidden checks fail because the planted bug shipped or something else broke.
- Mixed: everything else.
How do they direct the agent?
20 OF 100Two things are measured here, and the rating is the worse of the two.
The first is context. Most tasks give the candidate facts that are nowhere in the code, such as what Finance said on a call. The agent can only act on them if the candidate passes them on, and the hidden checks test each one. A blind edit is the agent changing a file it never read.
- Standout: every fact reached the code, with no blind edits.
- Strong: every fact reached the code with one blind edit at most, or half of them did with none.
- Flag: no fact reached the code, or two or more edits were made and every one was blind.
- Mixed: everything else.
The second is cost, against the round's cap of $1 of model use. Friction is the number of times the agent was refused permission, not counting changes the candidate chose to reject.
- Standout: under half the cap, with no friction.
- Strong: under 80% of the cap, with friction of two or less.
- Flag: friction of five or more, or the cap was hit.
- Mixed: everything else.
How do they recover when it breaks?
15 OF 100This looks at the test runs the agent started, from the first failure to the next pass.
- Standout: two failed runs at most, and passing again on the very next run.
- Strong: passing again within three runs.
- Flag: five or more failed runs, and never passing again within three.
- Mixed: everything else.
If no test run ever failed, there was nothing to recover from, so this question is left out of the score.
What is not rated
- The number of prompts. One clear instruction beats six vague ones.
- The time taken. Every round has a clock, so time pressure already shows up everywhere else.
- How the candidate handles an unclear request. It is worth 5 points in our design, but we have no fixed rule for it yet, so it never counts.
The catch
Every task hides one mistake the agent tends to make on its own. The scorecard shows whether it was caught. A catch only counts as pushback when the candidate named the problem before the fix, or rejected a change. If the agent handled it by itself, that counts toward whether the work holds up, and not toward the candidate.
When there is no score
Some rounds are kept in full but never get a number:
- Nothing was attempted: no instruction, no agent turn and no typed code. Untouched code passes the hidden checks, so silence must never look like careful work.
- Auto-accept was on, so the switch made decisions a person should have made.
- A change was answered by something other than the candidate, such as a timeout, or our approval step failed during the round.
When a question cannot be measured
It shows as Mixed, with a sentence saying why, and it is left out of the score. The one exception is recovery with nothing to recover from: it shows as Strong if the final tests pass. These rules are version 2, in use since 9 September 2026.