Session Status Board — live

MACHINE POWERED OFF ~18:50 UTC Aug 21 (moving). Final state: all six sessions reported and received the 5-minute save-now signal; every card reflects last reported state; all in-progress work is pushed, archived, or in session memory. Session 005's raw watch data: ~/.claude/merge-speed-watch-archive-2026-08-21/; its resume memory: project_merge_speed_watch_2026_08.md (agents-005). Resume any lane via claude --resume in its terminal, or fresh from the memory file named on its card. Fede's action queue is listed throughout. Last updated: Aug 21, 18:50 UTC.

Live sessions at shutdown

005 — Merge-speed watch & CI cost (this manifest's author)

session 48ac3270 · terminal 005 · started Aug 19
DONE (shipped & verified): IN PROGRESS / WAITING: Resume with: open the watch doc, get Fede's Option-3 + trap-3 decisions, check Aug 22 billing, then implement whichever option he picks.

004-b3 — Email-bounce incident (dropped Zillow lead)

ref 36da5a · reported 14:20 UTC
DONE (everything — safe to shut down): IN PROGRESS: nothing — no open PRs, no unpushed work, no timers.
Resume with: memory file project_email_bounce_ndr_lane_2026_08_20.md (agents-004 memory) — carries the gotchas (never replay the real bounce email: it would double-text the lead).

003-8a — Voice latency ≤1s campaign (Willows lab) — LANE CLOSED post-restart (Aug 21, taken over by session 008)

Post-restart close-out: the Qwen evidence runs had written their files pre-shutdown — nothing re-ran. Both of Fede's asks answered (why-not-Sonnet: controlled 20v20 seam experiment, Qwen silently books 3/20 vs Sonnet 0/20; automated bug-class detector built and run over 232 real conversations — one confirmed prod instance on Sonnet). Bake-off decision doc PUBLISHED as Proposed with 4 decisions for Fede: docs.propflowai.co/a/voice-model-bakeoff-2026-08-21 (recommendation: Sonnet stays on real lines; real prompt-migration pass for Qwen in the lab). Cleanups done: 4 stress agents deleted, caller name Fred→Fede fixed. Still gated on Fede: PR #5984 merge click; Willows line still on lab agents until testing ends. Full state: project_voice_bakeoff_takeover_2026_08_21 (agents-008 memory).

ref 64e168 · reported 14:20 UTC
DONE: IN PROGRESS / NEEDS ACTION: Resume with: memory file project_voice_model_latency_campaign_2026_08_20.md (agents-003 memory) + ~/agents/003/STATUS-willows-1s-latency.md. Key judgment context: small models rejected by Fede except the current lab model on probation; Gemini disqualified for fabrication; cross-model eval scores can't be trusted raw.
Delta 15:30 UTC:
  • The Spanish-language bug is solved (root cause: the transcriber mis-hears accented English as Spanish, one word flips the language rule, transfers inherit it, and the call-end save poisons the stored preference). A fix package with 3 options is ready — awaiting Fede's pick.
  • New cross-cutting finding — identity churn: Fede's own phone number resolved to 8 different person records across 3 orgs in a month; likely the root of "Clara doesn't remember me". Handoff: ~/agents/003/HANDOFF-context-memory-inspection.md; relevant to the held identity-merge gate PR #5762.
  • Bake-off: terra eliminated (silent booking-readback flips); Qwen is the sole fast candidate — final evidence runs in flight. #5984 still awaits Fede's merge click.
  • Landed: Willows KB application-fee corrected to $38; mined ASR hint list live on lab agents.
Delta 16:20 UTC:
  • Language bug fully root-caused, all variants. The last one: a stale auto-summarizer note ("prefers all communications in Spanish", itself written after earlier wrongful flips) overrides the caller's actual saved language, and the handoff instruction never says current-speech-wins. Fix package complete — 5 prompt/pipeline changes + voice added to the nightly language audit, regression suite ready. Awaiting Fede's go.
  • New framing from Fede: go/no-go is END-TO-END, not just the model. Harness inventory done; key systemic gap: no leasing harness anywhere verifies an SMS actually arrived (Twilio read-back exists in the vendor harness, never wired to leasing). Side-effect audit of his calls (including zero texts received) still running.
FINAL delta 18:49 UTC (at shutdown):
  • The missing-SMS root cause is stamped: escalation deliberately anchors sends to the escalated thread ("escalated-wins"), which mutes the text. Two handoff docs written (~/agents/003/HANDOFF-escalation-mute.md + the outbox rearchitecture) and delivered to the escalation session. Prod sweep: 63 stuck escalations found, no confirmed real-customer victims, 17 unauditable.
  • The reschedule-without-asking bug root-caused (menu-pick at the handoff seam); a prompt fix is LIVE on the Willows lab agents, snapshots archived.
  • Scratchpad archived to ~/agents/003/scratchpad-archive-20260821/; memory carries resumption state including Fede's two newest asks (why-not-sonnet; automated bug-class discovery). Nothing unpushed.

002-0d — Conversation-label overhaul — LANE CLOSED post-restart (Aug 21)

ref 95bec6 · reported 14:20 UTC
DONE: IN PROGRESS / FOLLOW-UPS (no live branches): Resume with: memory file project_label_audit_2026_08_20.md (agents-002 memory); the escalated-thread-silence workstream (~/.claude/handoff-escalated-thread-silence-2026-08-20.md) turned out to be already shipped as #5990, merged Aug 20 23:11Z — no build needed; session 006 has the final shape for its architecture doc. Only Fede action left on this lane: rotate the Camellia door buzz code.
Delta 16:05 UTC: this morning's session rate-limit failures root-caused and fixed: the machine's Claude credentials file held a stale token for the limit-hit account while the Keychain already had the fresh login (a plain /login writes only the Keychain; the account-switch script skipped re-syncing the file). File re-synced live — session 005's recovery was the proof — and the switch script patched so it repairs a stale file from now on (verified by reproducing the failure). The old account's weekly limit also resets ~2pm Denver today.

sales (agents-001) — sales.propflowai.co board — LANE CLOSED post-restart (Aug 21, taken over by session 004)

Post-restart close-out (session 004): all three in-flight PRs landed — #133 ops-data landing (reconciled with #134's schema after a silent auto-merge collision; re-runs now preserve the curated staffing profiles), #135/#136 review fix-forwards (self-negating AI-detection guard hardened, pod-routing rule, "no in-house operations"→outsourced), and #137 the new AI-agent tag: every card shows which fully autonomous AI leasing agent a company runs (Fede's rule: CRMs/tour widgets/call-tracking/leftover code/archive captures never count; 5 review rounds audited the gates). Live-verified on the deployed board (58 chips). Corrected headline: 33 of 94 T1/T2 companies (35%) run a live autonomous agent (EliseAI ~26); earlier "51 of 94" counted tour widgets and disclaimed detections. Ops blocks + routing/payroll data live for 91 companies; kage→mixed pods, wirtz→outsourced corrected. Parked follow-ups from the old card unchanged. Worktrees reaped.

ref 17db09 · reported 15:10 UTC
DONE (everything — zero open PRs, no unsaved work): PARKED FOLLOW-UPS (none urgent): UNIMAT reopen-trigger left empty deliberately; label recheck after manual drags waits for next rebuild; swap the Slack user token for a dedicated bot; a ranking-snapshot count discrepancy (503 fresh vs 592 committed) uninvestigated.
Resume with: repo ~/code/PropFlow/sales-machine; agents-001 memory files (project_sales_machine_slack_pipeline_live.md + the voice-rules file). Gotcha: site login expires — only Fede can re-login via Cloudflare Access.
FINAL delta 18:52 UTC (at shutdown): two research fleets completed — operations facts (staffing/payroll/phone/email routing) for 95 target companies, and an AI-vendor sweep (EliseAI verified at ~24 companies from page source). Raw outputs durable at ~/code/PropFlow/sales-machine-staging/. Three PRs in flight on pushed branches: #132 (jargon purge, review round 5), #133 (ops data landing), #134 (operations-block renderer). Resume plan in its session memory.

Escalation session — coworker stream & Willows adversarial fleet

ref e33dbf · reported 15:05 UTC
DONE (all merged & deployed, verified on Lambda + Vercel): NOT LANDED / NEEDS ACTION — the important one: Resume with: that memory file — it carries the bug list, cleanup logic, and fleet run id.
FINAL delta 18:55 UTC (at shutdown):
  • Deep audit of the fleet findings COMPLETE — all held, and two earlier relayed claims were false and are corrected: widening the allowlist does NOT fix the addressed-to-Clara silence, and "Zillow leads always answered" is not true in general.
  • #6038 merged and audited sound; its one defect (Spanish replies wrongly suppressed) is mid-fix on PR #6046 (round 5 — its agent may die with the machine; resume it first).
  • #6040 deliberately FROZEN (1.2% false-block rate) — must not merge until the corpus is fixed.
  • New Fede-gated action (now 3 total for this lane): the Willows KB literally answers prospects with "Camellia Apartments" in two places — fix script durable at ~/agents/006/scripts/fix-willows-kb-names.ts, plus the two DB cleanups still pending.
  • Full resume state: project_shutdown_state_2026_08_21.md (its session memory).

Architecture Long Term Vision (agents-007) — Clara quality & eval program

ref 6e1fc3 · reported 14:25 UTC
DONE (all merged, worktrees reaped): PARKED ON FEDE'S DECISIONS (nothing running): Resume with: agents-007 memory (project_quality_system_plan.md, project_reason_first_email_sms.md); durable archive of all deliverables at ~/.claude/quality-poc-archive/2026-08-21/ (its /tmp scratchpad will NOT survive the move; the PII identity-mapping file was deliberately not archived).
Delta 15:40 UTC:
  • The Willows fleet findings are now tracked as five new gap rows (22–26) in /a/eval-testing-roadmap — Spanish coverage, a life-safety lane for unknown senders, voice promise-backing, bench silent-skips, cross-property KB contamination. The bug fixes themselves stay with the escalation session's pipeline.
  • The "Clara Quality Line" design artifact moved — canonical link is now claude.ai/code/artifact/ca496450-f801-4c42-bf95-0547873f9470 (the earlier 448bfdd9 link is dead; wrong account mid-rotation).
  • Model-strategy recommendation delivered to Fede (stay on the current agent-tier model, upgrade per-workload via the harness) — no decision recorded yet.
Delta 17:40 UTC:
  • The engine repo is LIVE: github.com/PropFlow-Technologies/cerberus (private, name still provisional), tag v0.1.1, its own CI green, 66/66 tests, 18-case corpus + 130-attack adversarial set.
  • App-side consumption PR #6039 is double-gated on Fede: it touches workflow files (hard floor — his merge click only) and its required test check needs the full suite (a fresh review was requested to trigger it — the big-PR trap in action again).
  • The post-merge verifier caught an internal-paths leak in the new repo's provenance doc — already fixed.
  • In flight: leasing replay spike (#6043, in review), a leasing-only required gate, a Camellia-cases funnel page.
  • Still parked on Fede: final engine name, the #6039 merge click, re-confirming the Camellia street address in fixtures.
FINAL delta 18:48 UTC (at shutdown):
  • Burn-in DONE and audit-confirmed — 53 trials, 9 findings; pack live at /a/cerebrus-burnin-2026-08. The enforcement flip is explicitly held for Fede (command file archived).
  • #6045 (two-lane leasing gate machinery) MERGED; #6039 + #6043 remain open awaiting Fede's click. Decision Line page live at /a/clara-decision-line.
  • One background agent (engine v0.1.2 hardening + the #6039 rebase) dies with the machine — relaunch spec is in that session's memory + archived EVIDENCE.md. All repos clean and pushed, evidence archived PII-free.

Parked for later (Fede, Aug 21)

Outbound-comms outbox rearchitecture — how sending should work so this bug class can't happen

Handoff: HANDOFF-comms-outbox-rearchitecture.md (agents-003). Today two separate after-call jobs race for one "send it once" lock; a declined send keeps the lock (30 days of silence, no retry) and throws away the reason it declined. Target design: the booking write itself records "this person is owed a text" in the same database write; one worker sends one text reflecting the call's final state (3 reschedules ≠ 3 texts); every skipped send records why and stays retryable; tests verify the phone actually received it. First shippable steps are tiny: log the skip reason (3 one-line fixes) and make skipped sends retryable. Bigger hole flagged: most real tours are booked by email/SMS and get no confirmation text at all. Status: PARKED at Fede's direction — pick up after the current decision queue clears.
DEEP AUDIT (Aug 21, two repo-wide sweeps — every send-once lock and every outbound lane): the reference bug is not a one-off; it's a pattern with more instances, and separately the repo's delivery verification is concentrated in exactly two lanes while everything else is fire-and-forget.

A. Same "claim-then-decline swallows the send" bug found elsewhere (ranked):

  • 1. Renewal recap SMS — near-exact copy of the reference bug. The text that is a tenant's only written proof of a renewal offer after a voicemail/hangup call: two racing jobs share one 30-day send-once key; the key is claimed before the "do we have a proposed rent?" gate; a decline burns the key for 30 days, never releases it, and the reason is only a log line. Real deadline/compliance stakes. (renewal-orchestration/recap-sms.ts:118)
  • 2. Escalation voice callback. The courtesy call-back Clara owes a caller whose question staff answered: 1-hour lock claimed before two gates that can decline; never released, reason dropped. Softened by the SMS answer still landing independently. (escalation/voice-callback.ts:173)
  • 3. Tenant-confirmation PM notification — marked "sent" even when the send failed or no PM phone existed; self-heals via the separately-counted reminder ladder, so one mislabeled record rather than permanent silence. (temporal/activities/tenant-confirmation-review.ts:117)
  • Also: the shared once-only helper fails open on a database error (a store hiccup lets a duplicate send through) — intentional, but worth knowing.
  • The good news: the repo already contains three exemplary implementations of this exact pattern (handyman-page dispatch, the renewal double-send guard, the 2026-07 outreach send-claim primitive — all release on decline, persist the reason, and alarm on stranded claims). The bug is inconsistent adoption, not a missing design.

B. Delivery-verification map — who actually checks the message arrived:

  • Verified end-to-end (the models to copy): renewal recap + renewal letter texts (carrier delivery receipts drive state, failures escalate through the reminder ladder) and the life-safety PM page (delivery-gated after a real 2026-07 incident).
  • Everything else is fire-and-forget. Delivery receipts are collected for EVERY text but consumed only by those two renewal lanes — tour confirmations, collections/dun money messages, handyman job pings, and ordinary Clara replies write receipts nobody reads. A carrier-accepted-but-never-delivered text is invisible until a human says "I never got that."
  • The purpose-built safety net is dead code: the tour say/do reconciler (built to alert when a promised confirmation never went out) has zero callers — never scheduled. Its producer also marks emails "sent" without waiting for the send to succeed. And the staff-triggered tour-notification routes were deleted Aug 13 — those tours get no confirmation at all.
  • Soft-bounced emails queue forever: the retry dead-letter store is write-only in production — nothing drains it.
  • Outlook-sent mail has no delivery events at all; the new bounce-catcher (from the Zillow-lead incident) covers only conversational replies, not one-off sends like renewal offers or application links.
  • An unanswered outbound call counts as "contacted" everywhere downstream (documented, undecided product gap) — it spends one of the capped monthly touches and can hold up the renewal cascade.
  • Single point of failure: nearly every alert above lands in Slack via one token with no fallback — a Slack outage would mute all these detectors at once.

Tiny high-value fixes (each ≤ a day, independent of the big rearchitecture): release the lock on decline + persist the reason in the two files above (copying the existing exemplary pattern); schedule the dead tour reconciler; drain the email dead-letter queue on a cron; tag tour + collections texts so the already-collected delivery receipts get read for them.

Open PRs at shutdown (the work queue that survives the machine)

PRAuthorState / what it needs
6036, 6035, 6034, 6023, 6013GeraFresh (opened this morning) — normal pipeline flow; 6034 has a 🔴 verdict waiting on Gera.
5998 (stage credentials scrub), 5984 (Willows sound promotion)Fede sessionsOpen, owned by one of the sessions above — see replies.
5930, 5786, 5785, 5784, 5782, 5708, 5694, 5634GeraDeliberately labeled hold-for-review (grading-playground stack + others) — resume when Gera is ready.
5807, 5695, 5670, 5656GeraOlder; 5695/5670 have real merge conflicts only Gera can resolve.
5636GeraGreen but edits an ADR file — needs Fede's explicit merge decision (10+ days old).
5716dependabotRoutine bump.

Standing machinery that stops with this machine

PropFlow Docs