0 of 9 answered

Eval pass marks: the nine decisions

We have automated checks that score Clara's behaviour out of 100. Each one has a pass mark. Two different systems used to keep their own copy of that mark, and for 28 checks the two copies disagree. Merging them exposed the disagreement; it didn't settle it. These are the calls that settle it — "I'm not sure" is a real answer. Your picks collect at the bottom to copy.

Why it's nine and not twenty-eight: most of the 28 aren't separate questions. Eleven are the same question asked eleven times. Five have a clear answer from the data. Seven aren't really about the pass mark at all. So there are nine real decisions here, and each one says which checks it covers.

One correction up front: our notes say 27 disagreements. It's 28. I counted from the live settings rather than trusting the note — 30 before, minus the 2 already merged, is 28.

The big one — it's eleven checks, but one call

01

Checks that have never once failed has a recommendation

Eleven checks have scored 100% on every single run we have — some across 20 nights. For each, the pull-request system demands a perfect score, while the nightly run accepts less. Both marks pass every result we've ever seen, so no measurement can tell them apart. This isn't a data question; it's a policy one.

What tips it: the perfect-score bar is already what blocks a pull request today. Choosing the looser number wouldn't tidy anything up — it would genuinely lower the bar on work in progress, for checks that have never given us a reason to.

Covers 11: adversarial · conversation-topics · dashboard-chat · fairhousing · identity-name-mixup · leasing-tour-update · leasing-tour-update-gauntlet · maintenance · pm-reply-context · postcall-prospect-capture · voice-closing

In plain terms

Eleven tests have gotten a perfect score every single time we've run them — some for 20 nights straight. Two different scoreboards disagree about what score should count as passing: one wants perfect, one accepts a bit less. Since nothing has ever scored less than perfect, there's no result that can prove either side right. So this is a choice about how strict we want to be, not something the data can answer for us.

Where the data already answers it

02

Five checks where the strict bar is provably too high has a recommendation

For these five, one system demands a perfect score — and we have real runs that scored below perfect and were fine. So the strict bar isn't aspirational, it's wrong: keeping it means a red square on a normal night. The other number in each case sits about one case below the worst we've seen.

Covers 5: leasing-tools (keep 94) · maintenance-tools (keep 84) · tour-alternatives-reply (keep 75) · turnover-scope-dispatch-gauntlet (keep 88) · virtual-tour-offer (keep 85)

In plain terms

For these five, one scoreboard demands a perfect score — but we have real nights where they scored slightly under and everything was genuinely fine. So that bar isn't ambitious, it's just wrong: leave it and the alarm goes off on a perfectly normal night. The other number sits just below the worst we've ever actually seen.

Five that need a real judgement

03

leasing — 92 or 97? has a recommendation

Five runs, and every one scored exactly the same: 99.32%. That flatness is the interesting part — it means the same single case fails every time, deterministically. 97 leaves about three cases of room; 92 leaves eleven.

In plain terms

We ran this five times and got the exact same score every time. Identical scores mean it's the same one test case failing over and over, not random noise. Picking 97 leaves room for about 3 more failures before the alarm goes off; picking 92 leaves room for 11.

Worth doing either way: a pass mark that quietly tolerates a known, unnamed failure is how a real problem hides. That one failing case should be written down by name, not absorbed by the number.
04

renewals — 95, or lower? has a recommendation

This has our best evidence by far: 30 runs. The worst was 95.35% against a pass mark of 95. That's a margin of a third of one case — the next bad night reds it. The alternative (perfect score) would have redded more than half the thirty nights — the middle result is 97.67%.

In plain terms

This one has the most evidence — 30 nights of it. The worst night scored 95.35 and the bar is 95, so we passed by about a third of one test case. That's close enough that the next bad night sets off the alarm. Setting the bar at perfect instead would have made more than half of those 30 nights look broken.

This is the coin-flip one. By our own rule — take the worst real result, then leave a cushion — 95 is too close. The rule points at roughly 94.
05

tenant-resident-services-gauntlet — 74 or 84? has a recommendation

Best read in whole failures rather than percentages. The check grades 21 cases. 84 tolerates 3 failures; 74 tolerates 5. The worst night we have failed 4 — so 84 would have gone red on a night we were happy with, and 74 would have passed it with one to spare.

In plain terms

Easier to think in failures than percentages. This test checks 21 things. A bar of 84 lets 3 of them fail; a bar of 74 lets 5 fail. Our worst night had 4 failures — so 84 would have set off the alarm on a night we were actually happy with, and 74 would have passed it with room to spare.

A number in between would do nothing. I was about to recommend 76 as "what our rule produces" — but on a 21-case check, 76 tolerates the same 5 failures as 74. It would have looked like a considered compromise and changed no outcome at all.
06

turnover-walk-editing-gauntlet — 79 or 86? has a recommendation

Both numbers pass everything we've seen (worst: 93.33%), so this one is low-stakes either way. The catch is the evidence: only four runs, and this is the check that was accidentally starved — it only ran on Sundays for part of the period, which is why it has fewer results than its neighbours.

In plain terms

Either number is fine — both pass everything we've seen. The real problem is we've only run this 4 times, because this test was accidentally only running on Sundays for a while. So we're choosing with much less evidence than usual, not because it's risky.

07

leasing-response — 78 or 81? your call

Both numbers pass comfortably (worst real result 83.77%), so picking either is safe. This one is on the list because the reason the two numbers differ has never been established — the two recorded measurements are ten cases apart and we can't say whether that's normal night-to-night variation or the check genuinely changing.

In plain terms

Both numbers are safe; nothing we've measured comes close to failing either. It's on the list only because nobody knows why the two scoreboards disagree by so much — about 10 test cases apart. That could be normal night-to-night wobble, or the test could have genuinely changed. We can't tell, so it's a judgement call rather than a calculation.

The bigger number here isn't the pass mark. On the one night we can read the split cleanly, 35 of its 228 cases failed — that is a single sample, not a trend, but it is a big number. The pass mark is set below that, so it never complains. Whether that's acceptable is a much more interesting question than 78 versus 81.

Where the pass mark is the wrong tool

08

Four checks that are failing, not mis-marked has a recommendation

For these four, the worst real result is below both pass marks — or lowering the mark would hide the exact thing the check exists to catch. Setting a number here doesn't resolve anything; it just decides how quietly we fail.

renewal-scoping is the sharp one. It keeps scoring 80%, and the single case failing is the one that checks we're not being over-restrictive. Dropping its mark to 60 would turn the whole thing green while that specific protection is broken.

Covers 4: cancel-save (worst 85.71 vs marks 89/100) · triage-routing (worst 82.50 vs 92/93 — a real regression, six cases broke between 26 and 29 July, not yet traced) · turnover-photo-gauntlet (worst 66.67 vs 80/100) · renewal-scoping (the failing protection above)

In plain terms

These four are genuinely failing — their worst real scores are below both proposed bars. Picking a number here doesn't fix anything; it just decides whether we fail loudly or quietly. Lowering the bar would mean the test stops catching the exact problem it was built to catch.

Where I refuse to guess

09

Three checks with no evidence at all has a recommendation

These three have two different pass marks and not a single recorded result between them. Any number I gave you would be invented. Two never appear in the nightly runs we mined; the third can't even be sized automatically.

Covers 3: turnover-nl-approval (no results) · vendor-call-extraction (no results) · prospect-summary (can't be sized)

In plain terms

These three have two different bars and zero recorded results — we've never actually seen them run. Any number here would be made up. Two of them never show up in the nightly runs at all, and the third can't even be measured automatically yet.

Your answers

nothing picked yet
Pick an option above and your answers will appear here, ready to copy.
PropFlow Docs