| Lane | Improved Haiku | Finding |
| prospect-disqualification-classifier | 100% | 387/387 full corpus; baseline had 13 false positives |
| postcall-claimed-unit-capture | 99.3% | "unit as job site" confusion closed |
| postcall-unknown-caller-name | 99.3% | speaker-attribution rule added; arms identical |
| fair-housing-gate | 99.2% | 0 false blocks on real traffic; stays fail-closed; gate itself unchanged per standing rule |
| scheduling-intent-classifier | 99.3% | safety precedence + closing-courtesy rules added |
| postcall-prospect-capture | 99.2% | intent-vs-identity contract fixed |
| vendor-message-intent-classifier | 99.1% | improved Haiku beats the Sonnet diagnostic arm |
| sms-message-classifier | 99.1% | ships with the regex narrowing as one change |
| tour-request-extractor | 98.7% | arms tied after fix; zero disagreements |
| quote-parser | 98.6% | arms tied 138/140 |
| pm-commitment-extractor | 98.5% | 79% → 98.5%; equals the diagnostic arm |
| collections-legal-gate | 98.4% | rubric rewrite; relevant to the collections policy slot in the outbound-safety registry |
| vendor-call-outcome-extractor | 98.4% | fixes the live wrong-date dispatch failure mode |
| clara-offer-extractor | 98.3% | all 16 baseline failures were prompt defects |
| vendor-slot-parser | 98.0% | 91% → 98%, 0 regressions |
| maestro-precheck-router | 98.2% | 13-point contract gap closed |
| tour-reply-extractor | 98.0% | 8 false tour confirmations → 0, recall up |
| prospect-name-extractor | 97.9% | arms byte-identical on 62/62 after contract written down |
| invoice-parser | 97.7% | 2-point gap was a prompt defect |
| broadcast-leak-gate | 97.7% | real fix is the fail-open parse (code, above) |
| wo-market-estimate | 94.6% | 67% → 95%; +1 line route code |
| maintenance-manual-parser | 94.1% | 73% → 94%, zero regressions |
| email-summarizer | 93.2% | prompt worth ~11× the model delta |
| outbound-translate | 92.8% | ship a code guard first (bare-code/punctuation inputs) |
| vendor-question-answerer | 92.6% | "absence of a line is not evidence" rule added |
| reengagement-intent-gate | 92.5% | we_owe_reply definition fixed |
| postcall-guest-card-reconcile | 92.4% | 236-call re-bench, judged |
| post-transfer-extractor | 92.2% | false tour mints 7 → 0; beats the diagnostic arm |
| rent-roll-pdf-vision | 91.9% | blank-money-cell rules added |
| vendor-completion-photo-validator | 91.8% | small corpus (7 misses) |
| pm-confirmation-review-reply | 90.7% | 0 unearned account-access grants across all 23 traps (baseline: 7) |
| email-triage-3way | 89.6% | improved Haiku beats old-prompt diagnostic arm |
| mass-comms-confirm-reply | 89.4% | injection-resistance bucket widened; body-is-evidence rule |
| rent-roll-llm-fallback | 88.6% | missing decline contract was the biggest gap |
| wo-judge | 88.2% | all losses were HOLD failures; edit behavior perfect |
| knowledge-extractor | 87.7% | +13.6 pts, zero regressions; corpus needs repair before more |
| followup-copy-generator | 87.6% | 64% → 88%; invented-availability instruction removed |
| prospect-summary-generator | 87.5% | self-contradictory "non-tour closer" rule removed |
| turnover-charge-judge | 87.2% | decision 3 below |
| operational-signal-extractor | 82.2% | decision 3 below |
| email-category-classifier | 77.7% | hard lane: 24% bad gold; ship gated on making spam recoverable |
| website-scraper-extractor | 96.2% | fabrications 8 → 1; decision 2 below |
| maestro-comms-compose | 71% | needs the safety clauses (code+prompt), not a tier change |
| vendor-completion-notes-parser | 63.9% | hardest lane; billing-corruption defect live; decision 2 below |