SSH-Drop Recovery Audit

Sat Aug 16, 2026 · What you asked for across every session Fri Aug 15 → Sat night, what actually got done, what fell through the cracks, and who should pick each piece back up. Evidence: all 15 session transcripts on the machine (read by 9 independent readers), a ground-truth pass against GitHub, and live self-reports from the 4 sessions running now. Times in Mountain Time.
7
workspaces active Friday; the SSH drop at ~10:42pm Fri cut 4 of them mid-work
17
PRs verified merged in the window (GitHub-checked, not self-reported)
7
items genuinely dropped or still promised-not-delivered (was 9; two email items turned out to be done Fri night — corrected)
1
held PR merged by the platform against everyone's intent (details below)

The 30-second version

Almost nothing was lost. Of the four sessions cut off mid-work Friday night, three were fully recovered today: the escalation testing campaign (even its 130-requirement spec that lived only in a wiped temp folder is now safely in the repo), the overnight incident-fix session (all its PRs merged), and the sales-tool session (its leftovers closed). The one session that could not recover itself is workspace 005 — its live session was cleared and knows nothing about Friday; but Friday's transcript shows that session ended at a clean stopping point, so nothing there is dangling except items already handed to you.

What did go wrong: the tour-decider PR that every session agreed was "held for your decision" got merged anyway at 5:32pm today — not by any session, but by the repo's own auto-merge-on-green robot, because the hold existed only as a verbal agreement. It shipped switched off (legacy behavior remains the default), so nothing changed for customers, but you never got to make the call.

The most important open item is a security one: on Friday you pasted the Anthropic admin key into a chat so a session could use it. That session flagged it "MUST BE ROTATED when done." There is no evidence anywhere that it was rotated. Please rotate it today.

Needs you (in priority order)

#WhatWhy it mattersEffort
1Rotate the Anthropic admin key in the Console (Settings → Admin keys)A live secret sits in a chat transcript on disk. Flagged Friday, never done.2 min
2Decide on the tour decider (merged switched-off). Options: A leave it on main but off (recommended — inert, revert available any time); B revert it now; C read the morning comparison in the RCA doc and turn it on.It merged without your click. Your decision is still real; only the sequencing broke.Your call
3Four escalation-email decision letters at escalation-decisions-2026-08-16The escalation campaign's remaining behavior questions; multiple-choice, recommendations included.5 min
4Escalation-email rewrite PR (the "question-first" format) — still open, correctly labeled hold-for-review, waiting on your click since Friday nightBuilt and green most of Friday night; the last blocker is you.1 click
5Out-of-scope call forwarding decision at voice-out-of-scope-forwarding — you were mid-thought on the after-hours tradeoff Friday when the conversation moved onStill "Proposed — pending review."3 min
6Saturday office-hours flags (D1/D2) from the SadieKate investigation — decision, not a buildLeft explicitly as your decision.2 min
7Small manual items handed to you Friday: Google Postmaster Tools signup · AWS credit expiry date (Billing → Credits) · reload two LaunchAgents from a GUI terminal · Slack bot files:write scope · Trello token re-mint · Slack bot token back into KeychainAll admin-login gated; none urgent, all cheap.~10 min total

What was dropped, and who picks it up

"Dropped" = you asked, no session finished it, and today's ground-truth check confirms it's still undone. Delegation follows the rule you set today: the session that owned the work drives it, resumed with its own history — never a fresh session guessing from a PR.

ItemWhere it came fromStatusWho should drive it
Text-speed before/after numbers — the last words of that session were "Numbers to you tomorrow"ws003 (AWS credit / texting speed)not deliveredResume ws003's Friday session. It's blocked on a dedicated test-lane AI account (fix #1 in the API deep-dive) or a quiet fleet window — the resumed session knows the harness and the target property.
Long-term, carefully-built fix for Clara's fabrication problem — you said slow down and do it rightws001 (harness / rollback day)partly startedThe replay-gate PR that merged today is the foundation (guards must be replayed against real conversations before merge). The actual redesigned guard is not started. New session, seeded with the RCA doc + ws001's transcript.
Camellia replay-testing branch is only on this laptop (never pushed to GitHub)ws001at riskAny session: push it. It's a 10-second, zero-risk backup. (Its successor was merged today, so this is belt-and-suspenders.)
Microsoft blocklist on the shared email sender — you said "do this now"ws001 / ws005actually solved Fri nightCorrection (Sun evening): ws005 attached the account's idle dedicated SendGrid IP to our domain Friday night and live-verified delivery to Camellia's Microsoft inbox; 002 added its reverse-DNS record today. Mail no longer leaves from the banned shared IPs. A courtesy delist request to Microsoft was also filed Sunday — nothing for you to do.
Dedicated (non-shared) email sending addressws001 / ws005already liveCorrection: done Friday night by ws005 (see row above) — the earlier "not started" was an audit error.
Email deliverability DNS fixes (DMARC ramp, CNAME cleanup) — planned in the architecture doc, never appliedws000plan onlyws000 lineage (it wrote the doc). Cloudflare DNS changes — should present as a checklist for your OK first since it touches mail for everyone.
Decommission the old pipeline-test endpointws000 (API deep-dive fix #7)ticket onlyIts own project (four internal tools still depend on it). Trello card, not a quick fix.
Clara's Gmail profile picture — "image set" then "i dont see the picture" — never resolvedws000unclearTiny; any session, but confirm with a screenshot this time.
Live bench investigation: Clara went silent on a test thread at 5:31pmws006 (live now)in progressws006 owns it and is on it — the error-tracker sweep was running at last report. Then its queued adversarial round.

What was verified done (the good news)

Every row below was checked against GitHub or the docs site by an independent verifier — these are not self-reports.

AreaDelivered
Escalation feature hardening (ws000 → ws006)Six fixes merged today from the overnight "go hard" bug hunt: dropped tenant question, Spanish out-of-office detection, staff refusal ≠ decision, spoofed-sender parsing, nag copy, plus the permanent CRUCIBLE test suite (4 review rounds). The overnight campaign's near-lost 130-requirement spec was recovered and committed. Live end-to-end demo on The Willows ran Friday (one scenario).
Booking-guard incident (ws001 → ws002)Rollback done Friday; the two affected renters were answered personally. RCA doc published with hard numbers (177 real turns replayed, 28 false alarms, 0 real fabrications). Tour-cancellation team email built and merged. Cancellation-note loss between processing layers fixed and merged. Replay-gate policy + harness merged and proven working today (I ran it end-to-end).
Email reliability (ws005 → 002)"Never a silent bounce" alarm merged Friday and proven live today — a deliberate test bounce landed in #alerts with recipient and reason. AppFolio email-correction sync bug fixed and DB-verified. Rate-limit hardening for inbound texts merged with both of your wording rulings applied.
AWS / cost (ws003)Runaway sync job (100% memory, 112 crashes) fixed. Clara re-reading full text history per reply fixed. Backups + audit pack on. Your spend-page tweaks merged. 4pm review job killed. API-usage deep dive published: real customers were never at risk; test traffic starves the shared pool.
Sales tool (ws004)Console overhaul shipped autonomously; audience-sheet "don't say" re-check completed today; review-bot login root-caused (dead token-rotation job since Aug 13) and fixed.
OpsReal-time GitHub event feed repaired (was crash-looping on a stale webhook). Default permission mode switched per your instruction. Session-ownership rule adopted across all sessions and saved to memory.

What happened last night, plainly

Around 10:42pm Friday the SSH connection dropped and every live session died with it. Root cause per ws000: the laptop wasn't on power and slept. Four sessions were mid-flight: ws000's overnight bug hunt (four fix PRs open, harness half-built), ws002's incident work (drafting your morning report), ws003 (waiting for a quiet window to measure text speed), and ws006's escalation campaign. This afternoon the sessions came back with no memory of Friday; each had to dig through its own transcript on disk to rediscover what it was doing. Coordination happened over cross-session messages, and — after one wrong turn where this session started driving PRs it didn't own — settled on your rule: the session that built the work resumes it. By 6pm every Friday PR was either merged or deliberately held.

The one process failure — the tour decider merged without you. The "hold" was an agreement between sessions and a memory note. The auto-merge robot only recognizes two hold signals: a PR marked draft, or the hold-for-review label. That PR had neither, so once the review bot cleared it (round 5), the robot did exactly what it's built to do. The two other held PRs do carry the label and are safe.

How to keep this from happening again

Problem seenFixStatus
Sessions wake up with amnesia after a disconnect and improviseRule: resume only your own pre-disconnect work; PRs are driven by the authoring session, resumed via claude --resume <id>; coordinate ownership before touching a branchsaved to memory today
A verbal "held for Fede" isn't enforcedRule: any PR held for you MUST carry the hold-for-review label (or stay draft) the moment the hold is decided — words and memory notes don't countproposed — save to memory + CLAUDE.md on your OK
Half-built agent state lives in temp folders that vanishws006 already moved its spec in-repo; make it a rule: campaign state files go in the worktree, not /tmpproposed
Laptop sleeps → SSH dies → everything diesYou declined new software (mosh/tmux). Cheapest fix that stays within that: keep it on power + caffeinate during overnight runs; sessions should also checkpoint a STATUS line to disk every phase so recovery is a read, not an archaeology digproposed
Secrets pasted into chatNever again — a session can read a Keychain item you name; it never needs the value in the transcriptrotate today; rule proposed

Session-by-session (for the record)

WorkspaceFriday's storyHow it endedToday
000Marathon: escalation-email rewrite (live-tested on Willows), API rate-limit investigation, spend page, kill 4pm job, then the overnight "go hard" bug huntCut off mid-campaign 10:27pmCampaign adopted and finished by ws006
001Harness cleanup; discovered the overnight fabrication guard was deleting real replies → rollback, personally answered two renters, RCAClean stop 4:56pm ("everything I own is closed")Today's session (this one): event-feed repair, coordination, this audit
002Day: Clara transfer wording shipped; account-rotation audit. Night: booking-guard RCA, tour-cancel email, tour-decider redesignCut off drafting your morning reportResumed headless twice; all its PRs merged; SendGrid alarm proven live
003AWS credit decisions (you took audit pack + backups + memory fix, declined the rest); texting-speed investigation; AppFolio login-code stormEnded 10:02pm with "Numbers to you tomorrow"Not resumed — speed numbers still owed
004Sales-tool console overhaul, mostly autonomousClean-ish stop 10:42pmLeftovers all closed today
005SadieKate resident investigation, AppFolio email-sync fix, Microsoft blocklist finding, bounce alarmClean stop 10:02pm after posting team updateLive session was /cleared — no Friday memory; nothing dangling except your decisions
006Escalation-CRUCIBLE campaign (inherited from 000)Died mid-campaignFully recovered; 6 PRs merged; bench investigation in flight

Method note: 9 reader agents each took one workspace's transcripts and listed every instruction you typed with verbatim quotes and a status; a 10th agent checked every "done" claim and every "dropped" suspect against GitHub and the docs repo. Where a self-report and the ground truth disagreed (one case: the tour decider), the ground truth won and the cause was traced to the workflow file and GitHub's event log. Unverifiable items (Slack posts, Temporal toggles, external portals) are labeled as such above rather than assumed.

PropFlow Docs