Overnight Quality Campaign

Autonomous run, night of Aug 18–19 — the "test it all night" directive. What shipped, what was proven live, what the adversarial harness found, and the five decisions waiting for you.

Shipped tonight (merged and live)

Change (by effect)Proof
Clara owns the follow-up in first person. Every escalation acknowledgment now says "I'll check with the team on that and get back to you" — "someone will get back to you" is gone in both languages, with a test that rejects any third-party promise forever. (#5864)Live probe through the deployed loop returned the new template byte-for-byte.Live
Your email's line-wraps no longer reach the prospect. The mid-sentence break you spotted ("Otherwise⏎regular pet fees apply") is unwrapped before relaying and before it becomes permanent policy text; bullet lists and real paragraphs survive. (#5865)Pinned on your exact live case; CRLF and multi-line lists covered after review.Live
A reply from the wrong address gets an answer, never silence. The failure that ate your reply twice (mail forwarded to Outlook → rejected silently) now sends the sender an in-thread note naming the assigned address, with anti-loop and anti-spoof protections. (#5866)Five behavior pins; absorb/fall-through wired into ingest.Live
Decision emails name the decision, not the person. Subject: "Quick decision: Does The Willows accept Section 8 / housing choice vouchers…" — never "on someone". Body shows the distilled questions as a clean numbered list; the raw back-and-forth stays on the audit record only. Nags and clarifications match. (#5868)Rendered email from your live two-question case is in the PR as the proof artifact.Live

In flight (approved or in review, auto-merge armed)

ChangeState
Team roles: Accounting, Regional Manager, Assistant PM, Leasing Assistant — full PM feature set each (nothing taken away, per your ruling), recognized by Clara everywhere a PM is, plus staff-recognition-by-email plumbing. Survived five review rounds that caught real permission gaps; each got a permanent guard. Kenya's relabel to Regional Manager happens right after this deploys. (#5862)Approved, CI
Closed matters answer too — your ruling in full: question on a closed matter → answered from records; same decision again → "already applied" note; changed decision → the person gets "Quick update — I checked with the team again", the policy is corrected, and you get a coworker-style confirmation. No branch is silent. (#5869)In review
Deleted conversations no longer orphan their records — a real bug the harness found: deleting a conversation stranded its escalation records forever (99 orphans counted on the bench, sweep script ready). Also covers the duplicate-PM-forward class. (#5870)Armed, CI
The live test harness itself — 23 scenarios against the real production code paths on the bench, a 47-probe adversarial fair-housing bank (built from the Zillow compliance corpus, the Harbor Group consent decree, and HUD's paired-testing method), and an LLM judge that reads every outbound message as a human would. Its review round hardened the instruments so they can't pass vacuously. (#5871)In review

What the adversarial run found

23 scenarios, 17 pass, 0 surprises, 6 expected-fails — every one an already-tracked gap. The headline: the HUD paired race-signal test passed — two prospects, identical question, names varied as a race signal, and Clara offered byte-identical facts to both. Steering, redlining, voucher-discouragement, and denied-applicant probes all passed. The six expected-fails are exactly the remaining build queue: the fair-housing outbound tripwire, the staff write-tool code gate, the no-token staff email lane, closed-matter replies (PR up), the line-fold (now fixed), and one non-hermetic precondition.
Also found by the harness, beyond the orphan bug: the tracked-matter lane silently depends on the classic PM-forward email delivering first — a coupling the code comments don't acknowledge. On the ledger.

Defect ledger (open)

DefectWhy it matters
TOPNo fair-housing screen on what Clara relays. You watched it live: a staff-written "we don't take section 8" was texted to a prospect verbatim. The inbound side passed every adversarial probe; the outbound side has no tripwire at all.Colorado source-of-income law; a property's own texts are the whole discrimination case. Decision card below.
P1Staff clarify-turns run with the resident's action tools live, guarded only by prompt (D-B)."Can you close WO-7?" in an email could fire a real write.
P1A fresh staff email with no reply-token is still parked silently (D-A / G2).Breaks always-reply for the team; the recognition plumbing shipped dark in the roles PR, ready to light up.
P2The "That said," fragment on escalating turns (partial-scope replacement) — not yet reproduced.Your live find; queued behind the merges.
P2Matter lane depends on the PM-forward email delivering first; unattributed voice tour-cancel failure; judge follow-up pack (~37 unused adversarial seeds).Ledger items, none urgent.

Decisions for you

D-1 · Fair-housing outbound tripwire — build now?

D-2 · Staff analyst toolbox + write-gate (D-A/D-B)

D-3 · Launch-readiness bar for the escalation loop

D-4 · Camellia shadow-mode exit

D-5 · Voice callback build

Live proofs, for the record

Full loop, production, your own thread: service-animal ask → tracked matter → your emailed decision → first-person relay to the prospect → policy taught → thread released. Deployed-code probes: first-person ack (byte-identical template), forward with distilled question-led subject. Every synthetic row from every probe and harness run was deleted and verified (the orphan bug the harness found in the delete path is why that verification now exists). Nothing at Camellia was touched beyond the Section 8 knowledge-row removal you ordered, which was followed by a suppressed compliance probe that passed.
Autonomous overnight run, 2026-08-19 · six PRs merged or armed · harness + judge + adversarial bank now permanent infrastructure · decisions D-1…D-5 pending Fede
PropFlow Docs