I'm parked on you

Everything I could decide myself is decided. What's left needs you. Each one has a recommendation from Fable — if you agree, pick it and press Done. "I'm not sure" is a real answer and becomes work for me. You don't need to open the session — pressing Done sends your answer back and it picks up.

1Should the renewal auto-send-vs-review routing call be recorded as a machine quality grade of itself in the grading corpus?

In plain terms. Clara decides, every day, whether each upcoming lease renewal can go out on its own or needs a person to look at it. We're building the page where you grade her work. Question: when you sit down to grade one of those renewal decisions, should the page already show you Clara's own confidence as a green/yellow score — or should it show nothing and let your thumbs-up/down be the only judgement on it? I recommend showing nothing, because Clara marking her own homework green would quietly skew the 'how often does Clara agree with a human' number we're building this whole thing to measure. Either way it's easy to change later.
Fable recommends: Show no score — your thumbs-up/down is the only judgement (recommended)
The routing call IS the thing you are grading, so scoring it green by construction means every renewal item grades itself. Your thumbs-down already says the routing was wrong, so the score adds nothing — and it takes the slot a real renewal-quality checker would need later. Volume is not a reason to add one: Clara's daily sweep writes a fresh row every day even when nothing changed (356 rows covering 32 tenants in a fortnight, ZERO of which actually changed their mind), so I have already told the builder to collapse a run of identical days into ONE item — about 32 a fortnight instead of 356. Tell me if you would rather it not collapse; that half is my call, not Fable's.
Tried first, unsuccessfully: receipt fb04e0f92 — both independent passes picked this same option, and it was escalated to you anyway because the question mentions renewals, which trips the money/commitment guard added after a confident-but-wrong renewal ruling on 2026-08-07. Fable's reasoning: the code defines a machine grade as 'one machine judgement about one item', and here the routing output IS the item rather than a judgement of it, so grading it green by construction would mix routing-agreement into the quality-agreement number the corpus depends on.

d451bf5d is parked on this.

2D8-A's stated noise guard does not hold — the eval sweep is non-required and budget-skipped, so auto-landed thumbs-down cases silently erode per-domain pass-rate floor headroom. What should D8 do instead?

In plain terms. You picked 'fully automatic' for turning a thumbs-down into a permanent test case. Building it, we found the safety net you named doesn't work: you said the test suite catching a bad case is the guard, but a thumbs-down case is BY DEFINITION one Clara currently fails — and the suite that would notice is deliberately not a blocking check and gets skipped on cost. So each auto-added case would quietly eat into the slack in Clara's quality bar, and nobody finds out until some unrelated change trips it. Recommendation: still fully automatic, but a new case lands in the 'known problem, tracked' list first and only starts counting against Clara's score once she can actually pass it. That list already exists and was built for exactly this.
Fable recommends: Keep it fully automatic, but land new cases as tracked-but-not-yet-counted, promoted once Clara passes them (recommended)
It keeps your 'fully automatic' pick intact — nothing here asks a person to do the drafting. And it is not a new invention: evals/tracked-gaps.json already exists and its own description says it holds 'the KNOWN-REAL Clara defects that the pass-rate floors now TOLERATE' so they stay visible. One of its entries literally says that without it, a new floor 'would have silently absorbed a fabrication miss — the exact outcome tracked-gaps.json exists to prevent.' That is the same failure D8 would create, just at volume. Promoting a case once Clara passes it also mirrors the graduation ladder you already picked for principles.
Tried first, unsuccessfully: receipt f0b1e1081 — both independent passes picked this option; escalated to you anyway because the phrase 'pass-rate' tripped the money guard, which is a false alarm here, but changing a decision you locked is genuinely yours either way. Fable's point: letting the auto-PR lower its own bar means the gate erodes by design and can never detect Clara getting worse; and leaving it as-written repeats the incident the eval workflow's own header documents, where roughly 13 genuinely failing cases hid behind green ticks for weeks.

d451bf5d is parked on this.

Pick an option above, then press Done.
PropFlow Docs