Session Status Board — live
MACHINE POWERED OFF ~18:50 UTC Aug 21 (moving). Final state: all six
sessions reported and received the 5-minute save-now signal; every card reflects last reported state; all
in-progress work is pushed, archived, or in session memory. Session 005's raw watch data:
~/.claude/merge-speed-watch-archive-2026-08-21/; its resume memory:
project_merge_speed_watch_2026_08.md (agents-005). Resume any lane via claude --resume
in its terminal, or fresh from the memory file named on its card. Fede's action queue is listed throughout.
Last updated: Aug 21, 18:50 UTC.
Live sessions at shutdown
005 — Merge-speed watch & CI cost (this manifest's author)
session 48ac3270 · terminal 005 · started Aug 19
DONE (shipped & verified):
- Nit policy (reviews: nits summary-only, cap 5, never gate) — 3 PRs merged, drift-test-pinned.
- Auto-merge re-arm on verdict (#5885) + assorted rescues; dropped D2-B work recovered (#5950).
- Cost cut #1: guard pins post-merge only (#6012, merged) — expected ≈⅓ off Actions spend.
- Cleared the ownerless workflow-PR backlog (Gera's #5827/#5975/#5967, then #6028/#6019).
- Living report:
docs.propflowai.co/a/merge-speed-watch-2026-08-21 — findings, receipts, decision options.
IN PROGRESS / WAITING:
- Fede's decision pending in the watch doc: Option 3 ("verify the merge, not every push") — carries the second half of the cost-halving AND kills the big-PR deadlock. Also trap-3 ownership choice (recommended: daily digest to #alerts).
- Cost verdict: Aug 22 is the first clean day post-cut — compare GitHub billing vs $170 (Aug 20) and ~$63 (pre-spike norm). Target ≤$80.
- The 10-min PR lifecycle collector + timers die with this machine — resume = a fresh session re-reads the watch doc; raw data in this session's scratchpad (
pr-lifecycle.jsonl, 5,900 rows) if still wanted.
Resume with: open the watch doc, get Fede's Option-3 + trap-3 decisions, check Aug 22 billing, then implement whichever option he picks.
004-b3 — Email-bounce incident (dropped Zillow lead)
ref 36da5a · reported 14:20 UTC
DONE (everything — safe to shut down):
- The Aug-20 incident (lead's Gmail full → Clara's reply bounced silently) fully closed: manual takeover SMS delivered to the lead, conversation left active so Clara keeps working it.
- Permanent fix merged (PR #6010): bounce detection → #alerts + automatic SMS fallback; prod-proven with a replayed test bounce. Trello cards moved to Done; worktree reaped.
IN PROGRESS: nothing — no open PRs, no unpushed work, no timers.
Resume with: memory file
project_email_bounce_ndr_lane_2026_08_20.md (agents-004 memory) — carries the gotchas (never replay the real bounce email: it would double-text the lead).
003-8a — Voice latency ≤1s campaign (Willows lab) — LANE CLOSED post-restart (Aug 21, taken over by session 008)
Post-restart close-out: the Qwen evidence runs had written their files pre-shutdown — nothing re-ran. Both of Fede's asks answered (why-not-Sonnet: controlled 20v20 seam experiment, Qwen silently books 3/20 vs Sonnet 0/20; automated bug-class detector built and run over 232 real conversations — one confirmed prod instance on Sonnet). Bake-off decision doc PUBLISHED as Proposed with 4 decisions for Fede: docs.propflowai.co/a/voice-model-bakeoff-2026-08-21 (recommendation: Sonnet stays on real lines; real prompt-migration pass for Qwen in the lab). Cleanups done: 4 stress agents deleted, caller name Fred→Fede fixed. Still gated on Fede: PR #5984 merge click; Willows line still on lab agents until testing ends. Full state: project_voice_bakeoff_takeover_2026_08_21 (agents-008 memory).
ref 64e168 · reported 14:20 UTC
DONE:
- Willows test line re-pointed to lab agents (Camellia untouched); current lab model ear-tested by Fede at 0.93s median, zero fabrication. Latency deep-dive published:
docs.propflowai.co/a/voice-latency-deep-dive-2026-08.
- Review fixes for PR #5984 pushed; PR is approved + green.
IN PROGRESS / NEEDS ACTION:
- PR #5984 (Willows sound → classic fleet) is green and waiting on FEDE'S merge click (dangerous-diff hold); after merge a human runs
npx tsx scripts/sync-specialist-voice-sound.ts --apply.
- Model bake-off decision doc not yet written; 300-scenario stress suite + 3 marathon calls never ran.
- Cleanup owed: delete 4 throwaway ElevenLabs stress agents; fix stored caller name "Fred"→Fede; reconcile the $50-vs-$38 application-fee mismatch in the test-property knowledge base; restore the Willows line from the saved snapshot when testing ends.
Resume with: memory file
project_voice_model_latency_campaign_2026_08_20.md (agents-003 memory) +
~/agents/003/STATUS-willows-1s-latency.md. Key judgment context: small models rejected by Fede except the current lab model on probation; Gemini disqualified for fabrication; cross-model eval scores can't be trusted raw.
Delta 15:30 UTC:
- The Spanish-language bug is solved (root cause: the transcriber mis-hears accented English as Spanish, one word flips the language rule, transfers inherit it, and the call-end save poisons the stored preference). A fix package with 3 options is ready — awaiting Fede's pick.
- New cross-cutting finding — identity churn: Fede's own phone number resolved to 8 different person records across 3 orgs in a month; likely the root of "Clara doesn't remember me". Handoff:
~/agents/003/HANDOFF-context-memory-inspection.md; relevant to the held identity-merge gate PR #5762.
- Bake-off: terra eliminated (silent booking-readback flips); Qwen is the sole fast candidate — final evidence runs in flight. #5984 still awaits Fede's merge click.
- Landed: Willows KB application-fee corrected to $38; mined ASR hint list live on lab agents.
Delta 16:20 UTC:
- Language bug fully root-caused, all variants. The last one: a stale auto-summarizer note ("prefers all communications in Spanish", itself written after earlier wrongful flips) overrides the caller's actual saved language, and the handoff instruction never says current-speech-wins. Fix package complete — 5 prompt/pipeline changes + voice added to the nightly language audit, regression suite ready. Awaiting Fede's go.
- New framing from Fede: go/no-go is END-TO-END, not just the model. Harness inventory done; key systemic gap: no leasing harness anywhere verifies an SMS actually arrived (Twilio read-back exists in the vendor harness, never wired to leasing). Side-effect audit of his calls (including zero texts received) still running.
FINAL delta 18:49 UTC (at shutdown):
- The missing-SMS root cause is stamped: escalation deliberately anchors sends to the escalated thread ("escalated-wins"), which mutes the text. Two handoff docs written (
~/agents/003/HANDOFF-escalation-mute.md + the outbox rearchitecture) and delivered to the escalation session. Prod sweep: 63 stuck escalations found, no confirmed real-customer victims, 17 unauditable.
- The reschedule-without-asking bug root-caused (menu-pick at the handoff seam); a prompt fix is LIVE on the Willows lab agents, snapshots archived.
- Scratchpad archived to
~/agents/003/scratchpad-archive-20260821/; memory carries resumption state including Fede's two newest asks (why-not-sonnet; automated bug-class discovery). Nothing unpushed.
002-0d — Conversation-label overhaul — LANE CLOSED post-restart (Aug 21)
ref 95bec6 · reported 14:20 UTC
DONE:
- PR #5992 merged: corrected labeler policy + speaker-identity lookup; the 84-case bug corpus scores 84/84 four runs straight.
- PR #6006 merged: "tour" became a leasing subtopic across classifier, UI, digest, reports.
- "Fix the past" executed with Fede's approval: 20 mislabeled production conversations corrected (report saved at
~/.claude/fix-past-report-2026-08-20-apply.json).
IN PROGRESS / FOLLOW-UPS (no live branches):
- CORRECTED post-restart: the tour-topic data migration DID run against prod before shutdown (Aug 21 ~08:16 local — 201 conversations, 1,494 message rows, 85 subtopic renames; verified zero remaining top-level "tour"; apply log:
~/.claude/tour-migration-2026-08-21-apply.log). Do NOT re-apply. Session 002 is re-confirming with a fresh dry-run only.
- DONE post-restart: passive verification complete — ~20 post-Aug-20 Camellia conversations label correctly (zero top-level "tour", leads→leasing, vendor call→vendor); nightly topic-label-sweep Temporal schedule armed, last run Completed (21 judged, 0 corrections, 0 errors).
- Fede action: rotate the Camellia door buzz code — the real code briefly sat in git history (already flagged).
Resume with: memory file
project_label_audit_2026_08_20.md (agents-002 memory); the escalated-thread-silence workstream (
~/.claude/handoff-escalated-thread-silence-2026-08-20.md) turned out to be
already shipped as #5990, merged Aug 20 23:11Z — no build needed; session 006 has the final shape for its architecture doc. Only Fede action left on this lane: rotate the Camellia door buzz code.
Delta 16:05 UTC:
this morning's session rate-limit failures root-caused and fixed: the machine's Claude credentials file held a
stale token for the limit-hit account while the Keychain already had the fresh login (a plain /login
writes only the Keychain; the account-switch script skipped re-syncing the file). File re-synced live —
session 005's recovery was the proof — and the switch script patched so it repairs a stale file from now on
(verified by reproducing the failure). The old account's weekly limit also resets ~2pm Denver today.
sales (agents-001) — sales.propflowai.co board — LANE CLOSED post-restart (Aug 21, taken over by session 004)
Post-restart close-out (session 004): all three in-flight PRs landed — #133 ops-data landing (reconciled with #134's schema after a silent auto-merge collision; re-runs now preserve the curated staffing profiles), #135/#136 review fix-forwards (self-negating AI-detection guard hardened, pod-routing rule, "no in-house operations"→outsourced), and #137 the new AI-agent tag: every card shows which fully autonomous AI leasing agent a company runs (Fede's rule: CRMs/tour widgets/call-tracking/leftover code/archive captures never count; 5 review rounds audited the gates). Live-verified on the deployed board (58 chips). Corrected headline: 33 of 94 T1/T2 companies (35%) run a live autonomous agent (EliseAI ~26); earlier "51 of 94" counted tour widgets and disclaimed detections. Ops blocks + routing/payroll data live for 91 companies; kage→mixed pods, wirtz→outsourced corrected. Parked follow-ups from the old card unchanged. Worktrees reaped.
ref 17db09 · reported 15:10 UTC
DONE (everything — zero open PRs, no unsaved work):
- Overnight 503-company inspection, data/renderer fix waves, and the per-stage card redesign — PRs #113–#131 all merged, deployed, and screenshot-verified on the live site.
- Spec + decisions doc published:
docs.propflowai.co/a/sales-machine-card-ia-2026-08-18.
PARKED FOLLOW-UPS (none urgent): UNIMAT reopen-trigger left empty deliberately; label recheck after manual drags waits for next rebuild; swap the Slack user token for a dedicated bot; a ranking-snapshot count discrepancy (503 fresh vs 592 committed) uninvestigated.
Resume with: repo
~/code/PropFlow/sales-machine; agents-001 memory files (
project_sales_machine_slack_pipeline_live.md + the voice-rules file). Gotcha: site login expires — only Fede can re-login via Cloudflare Access.
FINAL delta 18:52 UTC (at shutdown):
two research fleets completed — operations facts (staffing/payroll/phone/email routing) for 95 target
companies, and an AI-vendor sweep (EliseAI verified at ~24 companies from page source). Raw outputs durable
at ~/code/PropFlow/sales-machine-staging/. Three PRs in flight on pushed branches:
#132 (jargon purge, review round 5), #133 (ops data landing), #134 (operations-block renderer). Resume plan
in its session memory.
Escalation session — coworker stream & Willows adversarial fleet
ref e33dbf · reported 15:05 UTC
DONE (all merged & deployed, verified on Lambda + Vercel):
- Coworker stream closed out after the "Clara auto-replied to Camellia's own staff" incident: silence floor live, plain internal email template, operator-preview fix, English-only team emails, and the staff-always-answers feature deleted entirely (#5996/#6003/#6005/#6002/#6004).
- Willows escalation owner restored to fede@ (was parked on clara@ during the fleet run — reverted and verified).
NOT LANDED / NEEDS ACTION — the important one:
- The adversarial fleet found 5 real bugs. Still CRITICAL: an SMS about a gas leak from an unknown number gets filler instead of safety instructions — recommended first fix. The second finding (Clara silent even when staff address her directly) was re-examined: the silence-floor code is correct; the real cause is a downstream tours-only allowlist parking the message as "unknown" — likely a config fix, downgraded to "needs root-cause confirmation." Verdict unchanged: Willows NOT production-ready. Full fix handoff (2 buckets + recommended order):
~/agents/006/HANDOFF-fleet-findings-2026-08-21.md; findings memory: project_willows_adversarial_fleet_2026_08_21.md.
- Two cleanup commands wait for Fede personally (blocked in that session — approve there, or run yourself): delete 2 stray test rows misrouted into the real customer org (person
pers_29e84dc9… profile + claim claim_f8f38bce…), and clear leftover Willows test conversations (logic documented in the memory file; its scratch copy may not survive shutdown).
- Latent cross-property leak flagged: the Willows lease-terms knowledge base literally says "Camellia Apartments offers…" (cloned text).
Resume with: that memory file — it carries the bug list, cleanup logic, and fleet run id.
FINAL delta 18:55 UTC (at shutdown):
- Deep audit of the fleet findings COMPLETE — all held, and two earlier relayed claims were false and are corrected: widening the allowlist does NOT fix the addressed-to-Clara silence, and "Zillow leads always answered" is not true in general.
- #6038 merged and audited sound; its one defect (Spanish replies wrongly suppressed) is mid-fix on PR #6046 (round 5 — its agent may die with the machine; resume it first).
- #6040 deliberately FROZEN (1.2% false-block rate) — must not merge until the corpus is fixed.
- New Fede-gated action (now 3 total for this lane): the Willows KB literally answers prospects with "Camellia Apartments" in two places — fix script durable at
~/agents/006/scripts/fix-willows-kb-names.ts, plus the two DB cleanups still pending.
- Full resume state:
project_shutdown_state_2026_08_21.md (its session memory).
Architecture Long Term Vision (agents-007) — Clara quality & eval program
ref 6e1fc3 · reported 14:25 UTC
DONE (all merged, worktrees reaped):
- Quality Desk prior-verdicts fix (#6007, deployed); grader-attribution hardening + adversarial corpus, all 10 quote-confusion attacks pass (#6017).
- Blinded model-comparison harness (#6008) + first fair Sonnet-5 comparison ran: no significant difference overall — the agent tier stays on the current model; one fair-housing case flip flagged for human review.
- Quality-gate proof of concept (#6009): regression gate proven to go red on a planted regression; advisory judge lane live in CI.
- Reason-first drafting live on email + SMS (#6011, voice untouched); SMS proven end-to-end post-deploy.
- Roadmap doc updated (
docs.propflowai.co/a/eval-testing-roadmap); the old quality-factory page deleted at Fede's request — replacement design lives as the "Clara Quality Line" claude.ai artifact (under the trinity@ login).
PARKED ON FEDE'S DECISIONS (nothing running):
- Name the new engine repo (latest recommendation: Cerberus / Backstop) — scaffold ready once named.
- Whether to make the new regression gate a required check (one org setting; red would then actually block merges).
- Test mailbox for conversational email drafting (roadmap gap 21) — unassigned.
- Four grading-playground PRs stay held per Fede — do not re-raise.
Resume with: agents-007 memory (
project_quality_system_plan.md,
project_reason_first_email_sms.md); durable archive of all deliverables at
~/.claude/quality-poc-archive/2026-08-21/ (its /tmp scratchpad will NOT survive the move; the PII identity-mapping file was deliberately not archived).
Delta 15:40 UTC:
- The Willows fleet findings are now tracked as five new gap rows (22–26) in
/a/eval-testing-roadmap — Spanish coverage, a life-safety lane for unknown senders, voice promise-backing, bench silent-skips, cross-property KB contamination. The bug fixes themselves stay with the escalation session's pipeline.
- The "Clara Quality Line" design artifact moved — canonical link is now
claude.ai/code/artifact/ca496450-f801-4c42-bf95-0547873f9470 (the earlier 448bfdd9 link is dead; wrong account mid-rotation).
- Model-strategy recommendation delivered to Fede (stay on the current agent-tier model, upgrade per-workload via the harness) — no decision recorded yet.
Delta 17:40 UTC:
- The engine repo is LIVE:
github.com/PropFlow-Technologies/cerberus (private, name still provisional), tag v0.1.1, its own CI green, 66/66 tests, 18-case corpus + 130-attack adversarial set.
- App-side consumption PR #6039 is double-gated on Fede: it touches workflow files (hard floor — his merge click only) and its required test check needs the full suite (a fresh review was requested to trigger it — the big-PR trap in action again).
- The post-merge verifier caught an internal-paths leak in the new repo's provenance doc — already fixed.
- In flight: leasing replay spike (#6043, in review), a leasing-only required gate, a Camellia-cases funnel page.
- Still parked on Fede: final engine name, the #6039 merge click, re-confirming the Camellia street address in fixtures.
FINAL delta 18:48 UTC (at shutdown):
- Burn-in DONE and audit-confirmed — 53 trials, 9 findings; pack live at
/a/cerebrus-burnin-2026-08. The enforcement flip is explicitly held for Fede (command file archived).
- #6045 (two-lane leasing gate machinery) MERGED; #6039 + #6043 remain open awaiting Fede's click. Decision Line page live at
/a/clara-decision-line.
- One background agent (engine v0.1.2 hardening + the #6039 rebase) dies with the machine — relaunch spec is in that session's memory + archived EVIDENCE.md. All repos clean and pushed, evidence archived PII-free.
Parked for later (Fede, Aug 21)
Outbound-comms outbox rearchitecture — how sending should work so this bug class can't happen
Handoff:
HANDOFF-comms-outbox-rearchitecture.md (agents-003). Today two separate after-call jobs
race for one "send it once" lock; a declined send keeps the lock (30 days of silence, no retry) and throws
away the reason it declined. Target design: the booking write itself records "this person is owed a text" in
the same database write; one worker sends one text reflecting the call's final state (3 reschedules ≠ 3
texts); every skipped send records why and stays retryable; tests verify the phone actually received it.
First shippable steps are tiny: log the skip reason (3 one-line fixes) and make skipped sends retryable.
Bigger hole flagged: most real tours are booked by email/SMS and get
no confirmation text at all.
Status: PARKED at Fede's direction — pick up after the current decision queue clears.
DEEP AUDIT (Aug 21, two repo-wide sweeps — every send-once lock and every outbound lane):
the reference bug is not a one-off; it's a pattern with more instances, and separately the repo's delivery
verification is concentrated in exactly two lanes while everything else is fire-and-forget.
A. Same "claim-then-decline swallows the send" bug found elsewhere (ranked):
- 1. Renewal recap SMS — near-exact copy of the reference bug. The text that is a tenant's
only written proof of a renewal offer after a voicemail/hangup call: two racing jobs share one 30-day
send-once key; the key is claimed before the "do we have a proposed rent?" gate; a decline burns
the key for 30 days, never releases it, and the reason is only a log line. Real deadline/compliance stakes.
(
renewal-orchestration/recap-sms.ts:118)
- 2. Escalation voice callback. The courtesy call-back Clara owes a caller whose question
staff answered: 1-hour lock claimed before two gates that can decline; never released, reason dropped.
Softened by the SMS answer still landing independently. (
escalation/voice-callback.ts:173)
- 3. Tenant-confirmation PM notification — marked "sent" even when the send failed or no
PM phone existed; self-heals via the separately-counted reminder ladder, so one mislabeled record rather
than permanent silence. (
temporal/activities/tenant-confirmation-review.ts:117)
- Also: the shared once-only helper fails open on a database error (a store hiccup lets a
duplicate send through) — intentional, but worth knowing.
- The good news: the repo already contains three exemplary implementations of this exact pattern
(handyman-page dispatch, the renewal double-send guard, the 2026-07 outreach send-claim primitive — all
release on decline, persist the reason, and alarm on stranded claims). The bug is inconsistent adoption,
not a missing design.
B. Delivery-verification map — who actually checks the message arrived:
- Verified end-to-end (the models to copy): renewal recap + renewal letter texts (carrier
delivery receipts drive state, failures escalate through the reminder ladder) and the life-safety PM page
(delivery-gated after a real 2026-07 incident).
- Everything else is fire-and-forget. Delivery receipts are collected for EVERY text but
consumed only by those two renewal lanes — tour confirmations, collections/dun money messages, handyman job
pings, and ordinary Clara replies write receipts nobody reads. A carrier-accepted-but-never-delivered text
is invisible until a human says "I never got that."
- The purpose-built safety net is dead code: the tour say/do reconciler (built to alert
when a promised confirmation never went out) has zero callers — never scheduled. Its producer also marks
emails "sent" without waiting for the send to succeed. And the staff-triggered tour-notification routes were
deleted Aug 13 — those tours get no confirmation at all.
- Soft-bounced emails queue forever: the retry dead-letter store is write-only in
production — nothing drains it.
- Outlook-sent mail has no delivery events at all; the new bounce-catcher (from the
Zillow-lead incident) covers only conversational replies, not one-off sends like renewal offers or
application links.
- An unanswered outbound call counts as "contacted" everywhere downstream (documented,
undecided product gap) — it spends one of the capped monthly touches and can hold up the renewal cascade.
- Single point of failure: nearly every alert above lands in Slack via one token with no
fallback — a Slack outage would mute all these detectors at once.
Tiny high-value fixes (each ≤ a day, independent of the big rearchitecture):
release the lock on decline + persist the reason in the two files above (copying the existing exemplary
pattern); schedule the dead tour reconciler; drain the email dead-letter queue on a cron; tag tour +
collections texts so the already-collected delivery receipts get read for them.
Open PRs at shutdown (the work queue that survives the machine)
| PR | Author | State / what it needs |
| 6036, 6035, 6034, 6023, 6013 | Gera | Fresh (opened this morning) — normal pipeline flow; 6034 has a 🔴 verdict waiting on Gera. |
| 5998 (stage credentials scrub), 5984 (Willows sound promotion) | Fede sessions | Open, owned by one of the sessions above — see replies. |
| 5930, 5786, 5785, 5784, 5782, 5708, 5694, 5634 | Gera | Deliberately labeled hold-for-review (grading-playground stack + others) — resume when Gera is ready. |
| 5807, 5695, 5670, 5656 | Gera | Older; 5695/5670 have real merge conflicts only Gera can resolve. |
| 5636 | Gera | Green but edits an ADR file — needs Fede's explicit merge decision (10+ days old). |
| 5716 | dependabot | Routine bump. |
Standing machinery that stops with this machine
- The merge-speed watch itself (webhook feed, stall sweeps, lifecycle collector, report timers) — all session-local. GitHub-side machinery (auto-merge, re-arm dispatch, reviews, alerts, the new guards-post-merge behavior) is in the repo and keeps running by itself.
- An old leftover poll loop from Sunday (watching PR #5812 comments) dies too — it was orphaned anyway; no action needed.