1. What happened last night
One smoke detector went off, and two security guards each phoned the owner ten times about it.
- 39 alerts landed in #alerts between 10 PM and 8 AM.
- Ten were the same problem: Western Slope properties have no PM email.
- Agent Smith and the Trinity seat both investigated every one, independently.
- Each "needs your call" verdict tagged Fede directly, at 1 AM and 3 AM.
- The Trinity flag system held its three real asks until 7 AM, as designed.
Proves: #alerts thread pull for 2026-09-11 22:00 → 09-12 08:00 MT (22 mentions, 92 bot replies, 34 threads).
Deeper → Appendix: last night's alert list
2. The month in numbers
A car alarm that goes off every night: the neighbours stop looking out the window, even the night it matters.
- 849 alerts in 30 days, about 27 a day.
- 80 mentions of Fede; 63 of them in quiet hours.
- Only about 1 in 5 mentions was a real one-time decision.
- Fede replied to 15% overall, but to most of the real ones.
- 64 alerts were worked twice by two bots; 37 times they disagreed.
Proves: full pull of #alerts, #agent-trinity, #updates-fede, #agent-smith, Aug 13 → Sep 12, 2,438 thread messages classified.
3. Why: three design gaps, not one bug
The front desk was told "if it's serious, call the owner," and never told what time it is.
- The "needs a human" verdict is told to @-mention, with no time check.
- Alerts are deduped by record id, so ten households equal ten investigations.
- The Clara seat on Fede's Mac has no channel restriction, so it doubles Smith.
- Repeat posters (reviewer down ×9 in four minutes) have no cooldown.
- Fede's personal "wrong model" rule leaks into headless bot turns.
Proves: agent-smith alert decider prompt (sync-eng verdict), alert signature and recent-alert modules, Clara seat plist, Trinity watcher config.
Deeper → Appendix: file-level findings
4. Real bugs vs noise to tune
Sorting the mail: bills in one pile, flyers in the recycling, and the same flyer arriving five times a week.
- Real and worth keeping: CI red, reviewer login, nightly stress, turnover eval.
- Noise: Smith's own PR-wait check is the single loudest poster.
- Reports posted as alerts: spend, stale holds, latency canary.
- Duplicates: three alerts fire for one broken main; ten for one data gap.
- Sentry carried 568 open issues; 334 had been silent for a month.
Proves: 30-day cluster table (15 clusters, verdicts, PRs, Fede replies) and the Sentry unresolved query.
5. The teammate rules
A good on-call engineer fixes it, writes two lines in the channel, and wakes the boss only for a decision the boss alone can make, in the morning.
- A human is never @-mentioned in #alerts. Not Fede, not Gera.
- "Needs Fede" raises one Trinity flag per root cause; it waits for morning.
- Exactly one bot seat owns alert response; the other stays quiet.
- Same error class within 30 minutes is one investigation, not ten.
- Self-checks, digests and receipts post to the bot's home channel, not #alerts.
- Sessions post to #engineering-bridge only via bridge-post.sh (stamp + mention); stale flags retire silently.
Proves: Fede, 2026-09-12: "the whole point of agents on Slack is they monitor and handle alerts"; Gera, 2026-09-11 transcript: the recommendation is fine "90 or 95 percent of the time."
6. What ships today
Fixing the alarm panel, not adding a louder siren.
- Bot-side changes affect only how bots post; no customer path touched.
- The app change stops a staff reminder job at not-live organizations only.
- Camellia's reminders are unchanged.
- The Clara seat on Fede's Mac was restarted with the restriction at 12:47 MT.
- #engineering-bridge channel added 2026-09-12 (both bots members).
- Fede can DM Trinity (scopes added 2026-09-12); a "DM = second SSH" super-admin lane is a dark PR chain in progress.
- Trinity runs Fable 5.1 for every tier since 2026-09-12 22:14 MT; Smith follows once Gera restarts.
- Lesson (2026-09-13): dry-run a daemon against the live feed an hour before go-live; cut scope at noon.
Proves: PR list and merge state in the appendix, updated as each lands.
Deeper → Appendix: PR ledger
7. The bridge: Trinity ↔ Smith
Two site foremen with radios, so the crews stop pouring concrete on the same slab.
- Lanes are the ownership map: Gera owns isolation, Fede owns leasing and voice.
- A swarm claims an item before starting; a claimed item is off-limits.
- Out-of-lane finds become a handoff, never a fix in the other lane.
- A question goes to the asker's own session first, never the other person's Mac.
- Humans get one daily digest, never a ping per item.
Built: the claims board and the four message shapes are live in code, sealed between machines — dark until Fede turns the channel on.
Proves: the 2026-09-11 Fede–Gera call transcript ("why not have a bridge between the swarms", "one owner, raise your hand"); the lanes file.
Deeper → Appendix: how the bridge works · the weekly engineering roadmap — the shared page the lanes, the plan and the trade-offs live on
Last updated 2026-09-16 01:25 UTC.
Heads up: roadmap page stale since 2026-09-13T17:21+00:00 (56h). The numbers below may be behind.
Fede
Active
- Verify work-order / vendor-pool sync batching is on
- Set up dedicated leasing email as Google send fallback
- Limit auto-reply to syndicated guest cards with contact info
- Automate knowledge base generation and compress Clara output
- Decide Clara handling for no-pets and service-animal policy
- Complete final Microsoft app verification step for calendar connect
- Tune Clara verbosity to be less chatty
- Hide/disable broken Clara app entry point
- Own calendar changes for org/property leasing model
- Build tour reschedule/reassign flow
- Build multi-agent tour assignment algorithm
- Add review-turn metric and cap runaway PR loops
- Run end-to-end test on mirrored test client
- Turn off per-property reporting emails from Clara
- Improve Microsoft integration security and permissions
- Send onboarding link to client after dry runs
- Set up auto-reconciliation of AppFolio-paid Stripe invoices
- Delete test org and create the four real orgs
- Move Yale to a new ConAm org with Dara as admin
- (one item omitted — it referenced something internal, not shown here)
- Prioritize voice agent handling for noise, drug, and police/emergency calls
- Build resilient intake engine with adapter pattern for diverse sources
- Set up leasing email as Clara's inbound/outbound channel
- Configure showing scheduling policies (min notice, drive time)
- Do another pass on the privacy policy to add collections/payments disclosures
- Move off personal AppFolio logins to Clara agent accounts
- Set up Stripe intro-then-step-up subscription pricing
- Write reusable Clara knowledge-base ingest CLI skill
- Model tour/calendar assignment as deterministic on-call rotation
- (one item omitted — it referenced something internal, not shown here)
- Build reusable multi-tenant/permissions test harness
- Decouple mailbox/calendar from hardcoded one-to-one and Outlook
- Investigate reducing Clara call latency without degrading quality
- Fix concurrent Clara sessions overwriting each other on reset
- Build browser-based PMS adapters that mirror messages into PropFlow
- Define Citus deal structure and stakeholder plan
- Share raw Citus discovery transcript for deep inspection
- Investigate worsening latency
- Remove hard-coded Camellia property references from code
- Reduce Clara verbosity; split knowledge base into hard facts vs prose
- Fix lifecycle bug: signed lease still shows as prospect, not advanced to final stage
- Migrate to shared rotating bearer-token pool
- Add blocked/unavailable indicators to property calendar
- Find a good Spanish demo example for the deck
- Pipe Zillow/website leads into PropFlow and set up Conam leasing access
- Add mode-based dynamic prompting for property types (Section 8, centralized)
- Replace multiple renewal/vacancy Google Sheets with AppFolio-API dashboard
- Expand competitor deep-research doc (Elise, BetterBot, others)
- Fix pipeline view: separate/hide out-of-buy-box rows
- Handle transcribed weekend missed-call emails as leads
- Redirect Conam lead sources to point to PropFlow inbox
- Get PropFlow/Clara added as a teammate user in ILM/RealPage
- Build landing-page voice demo with guardrails after website redo
- Move sales pipeline tool into its own repo
- Implement mixed-model routing per sub-engine with hallucination guards on voice
- Make 'no reasoning tokens' explicit in trace viewer
- Add clipboard-copy of full agent trace for Claude debugging
- Shift Clara to human-in-the-loop with outbound + escalation guards
- Batch thumbs-down feedback into one daily PR
- Ship new ElevenLabs v3 voice to production after eval pass
- Push new tenants to adopt Clara at move-in
- Auto-fill government assistance apps and legal notices as PDFs
- Auto-screen applicants immediately on apply
- Export ConAm prospects and import into PropFlow
- Flesh out follow-up tab with automated reminders for 'too early' leads
- Standardize on cheap models in agent fleets to cut token usage
- Build hybrid browser/email agents for AppFolio, Yardi, RealPage
- Consolidate eval classes into one standard procedure in CLAUDE.md
- Add renewal follow-up loops for stalled negotiations
- Build Spanish outbound calls and texts
- Detect third-party property manager and build connections graph
- Tune AI prospect ranking to prioritize owner-operators in Denver/LA on Yardi/AppFolio
- Build collections module with delinquency stages and outreach
- Reset renewal reminder cadence when offer is changed and resent
- Pull move-in date from application to drive lease
- Add owner-mindset differentiation to pitch positioning
- Create discovery-call script and objections/answers playbook
- Auto-charge pet deposit and fee to ledger on signing
- Handle multi-identity linking for prospect referrals
- Complete Yardi contract, sign, and stand up staging env
- Define auto-approve renewal rent-reduction policy
- Investigate renewal call routing through property number
- Lower Claude billing limit from $200K to ~$5K and verify guardrails
- Link AppFolio PO approval to invoice amount and automate AP
- Add fair-housing topic/subtopic flag on conversations
- Change vendor confirm button to 'Confirm as preferred vendor'
- Add bid/proposal + walkthrough path for larger vendor jobs
- Fix vendor voice agent: identify property, not PropFlow
- Surface Matterport virtual tours on website and via Clara
- Build tour no-show capture and automated follow-up via Clara
- Reroute ILM-controlled Zillow/leasing inquiry email to property inbox
- Build vendor recognition and outbound vendor comms (email + voice)
- (one item omitted — it referenced something internal, not shown here)
- Add human-in-the-loop approval queue for vendor emails
- Demo turnover flow and decide when to turn it on
- Refactor vendor dispatch onto Temporal architecture
- Add tenant/PM move-out walk scheduling coordinator
- Add move-in inspection and condition-in vs condition-out comparison
- Build new-tenant welcome / move-in inspection flow
- Mass-text Camellia tenants introducing Clara + office signage
- Include Clara intro in tenant mass messages
- Store and offload turnover/inspection video
- Support add/edit/delete notes via text after summary
- Add visual photo-capture confirmation in glasses/app UI
- Improve damage vs normal wear-and-tear classification
- Disambiguate photos/notes across multiple open work orders in one SMS stream
- Detect apartments.com-referred leasing callers, skip Clara's greeting
- Clara forwards vendor sales cold-calls to PM inbox as 'Vendor notice — needs follow-up'
Waiting on
- I need a Slack channel for the two agents to talk to each other. Neither bot token can create channels (missing permission). Can you create a public channel called #bridge and invite both Agent Smith and Trinity to it? Everything else on the bridge build continues without it. Reply 'done' when created, or tell me to use a different existing channel instead.
- Something turned the new tour-email timing on at the Western Slope prototype line tonight without your say-so. I turned it back off. A line that would switch it on for the real Western Slope property is also sitting in your baseline one-shot script, labelled as your decision from yesterday. If that is not yours, delete it before you run that script again. Do you want this on at Western Slope? Reply yes or hold.
- One of tonight's small CI PRs adds a local command that runs an affected-tests sweep — the exact shape of command the memory-safety guard on Claude sessions was changed yesterday to block. Should that guard also block this new command (meaning the new PR is dead on arrival and needs rework), or should this specific command be added as an allowed exception? Reply either way and I'll finish accordingly.
- On tomorrow's demo the Western Slope line will say it does not know what is around any home, because none of the 15 homes has its neighbourhood written yet. One command fills all 15 in and costs about seventy cents. Only you can run it, since it writes to the real customer. I saved the exact command for you. Run it, or hold and demo without it?
- The big calendar PR I was asked to split into 3 smaller ones got merged whole by someone else while I was mid-split — so the full change is already live on main, and my 3 smaller PRs would now just re-add the same code on top of old code and conflict. Should I close the 3 smaller PRs since the work already shipped, or do you want something else?
- Calendar connect is broken in production: every Microsoft calendar connection (a person's or the company's) fails at the last step because the app's AWS credentials are missing one DynamoDB permission (dynamodb:ConditionCheckItem on propflow-prod, IAM user gera-dev). Found tonight on the Fairhaven test company; it will block Kat's calendar for Western Slope on Thursday. Recommendation: add that single permission to the gera-dev policy now (a one-line IAM change, no code, reversible). Yes = I add it and re-run the connect as proof; hold = I hand it to Gera's side. Sentry issue 7732381599.
- Company vs home policy — which wins when they conflict? Saturday you said a home saying 'cats allowed' should beat the company's 'no pets' for that item; Gera's key registry and a decision ruled the same day say both are shown, company first, never replaced. I built the registry's rule and pinned it as a test rather than quietly picking yours inside a renderer. Recommend keeping both-shown for now — nothing reads company policy yet so no customer sees either behaviour, and changing it means changing Gera's registry with him. Also: OK to close 3 superseded PRs (they are earlier slices now fully inside the one PR)? Nothing is blocked on either.
- The work-order refresh can't cover Western Slope's 1,074 open work orders in one daily window and will page every morning (and stale Camellia from tomorrow) — run the sweep every 6 hours, or keep it daily and skip Western Slope until maintenance goes live there?
- I need to send 5 real test emails through the real Claude model to prove Fairhaven's email replies work, and that needs the paid Anthropic key, which is locked by default for exactly this kind of run. There's a one-line command in the runbook that unlocks it for an hour. Reply yes if you want me to have you run it, or hold if you'd rather I skip that proof for now.
Projects
- Camellia back to autonomous: Camellia real-history corpus + scorecard, reproduce every hand-off failure, remove the half-baked gates, before/after replay, then Fede's go to turn on
- (open) Situs Group onboarding: proposal, policies captured, owner named, corpus in week one
- Western Slope onboarding (design partner, parallel with Situs): invite → API key via secure link → Kat's Outlook calendar + leasing@ mailbox → knowledge review → go-live Thu Sep 17 by forwarding [redacted-phone] to Clara's line. Leasing only; renewals then maintenance later (kickoff call 2026-09-09)
- Sign-in: typed 6-digit code default, 303 after confirm, passkey + text second factors, no recovery-codes screen
- Multi-tenant isolation: stopgap today, then two-step wall (mandatory tenant context + CI; DB-level keys + per-request credentials)
- (one item omitted — it referenced something internal, not shown here)
Done
- Nothing here right now.
8. Decisions for Fede
Three checkboxes, each with the box already ticked in pencil.
- D1: (a) code gate for not-live orgs, recommended and shipping dark; (b) flip the per-property switch off by hand.
- D2: (a) Smith owns #alerts, recommended; (b) Clara seat owns it; (c) split by source.
- D3: (a) shared claims file plus heads-ups, recommended; (b) Slack-only; (c) Temporal queue.
- Silent default: the recommendation applies. Gera is tagged on D2 and D3.
Proves: the standing rule that options come with a recommendation and never as an open list.
9. Slack channels
Decluttering a shared house: keep the rooms people actually use, give each automation one labeled closet.
- 16 of 26 channels: no human message in 30 days, several with no automation either.
- 5 ticket + 5 module channels retire — Trello board is already the real to-do list.
- #western-slope not yet created — both Slack tokens are missing the channels:write scope.
- Recommend one #propflow-engineering channel for Gera hand-offs and team decisions.
- #marketing and #deployments fold into #agent-smith; both are 100% bot traffic.
Proves: conversations.list/history over the last 30 days (oldest 1786611633) for all 26 public channels, cross-checked against automations.toml and agent-smith's Trello taxonomy for live posters.
Deeper → Appendix: full channel inventory and per-channel verdict
10. The customer success review
Not another alarm. A colleague who read yesterday's mail and flags the three letters you should answer.
- At a hundred homes Fede read every conversation. This replaces that reading.
- Live properties only, and only what a person should actually open.
- It reads grades the quality engine already writes — nothing new decides what is good.
- A day with nothing to fix is one sentence, not a report.
- It only watches, so it needs no per-property switch and changes nothing anyone receives.
Proves: merged 2026-09-12 after running read-only against four real days of live conversations; the output of each run is in the merged pull request.
11. The PR babysitter
The contractor stopped phoning to ask if the inspector had been.
- 38,673 turns last week were pure waiting on GitHub for a build to finish.
- A background process waits on a local file, calling GitHub only on real events.
- Green and reviewed, it merges; on a finding, a cheap helper runs, three times max.
- It messages the session only when the work is done.
Proves: Live on the Mac mini since 2026-09-13 16:36 MT for agent-smith; killed 16:39:05, back 16:39:10; it re-checked, skipped Gera's PRs, and ran fix rounds unattended. Three gaps found live and fixed the same evening: no checks-passed event, restart forgot its list, silent fix rounds. Its first assignment (its own four fix PRs) it did not merge: it spent three fix rounds on each, the review bot stayed on "changes suggested", and it went quiet instead of saying so. Those four were merged by hand 2026-09-13 21:39–22:55 MT; after the restart it re-listed the open PRs on its own and reported "blocked" for each — the restart fix proving itself live.
Proposed — pending Fede: our build machine reports failures here directly, with failing test names. 3 PRs, 2 days.
12. DM Trinity = second SSH
The office phone, forwarded to your cell — same desk, same drawers, you are just not in the chair.
- A DM runs on the mini as if Fede had typed it into Claude Code — terminal, repos, full reach.
- Voice notes count. He speaks, it transcribes on the machine, the words become the message.
- "Tell the affordable-housing session to stop" reaches that session. So does "what is it doing".
- "Ask Gera's computer" crosses the bridge, waits for the reply, and brings it back in the same answer.
- She never decides for him. A choice that is his comes back as one short question.
Proves: one switch (TRINITY_DM_ADMIN), hard-gated to Fede's own 1:1 DM; three pull requests merged and two in review. The first live DM set the standing rule — wait inside the turn, answer once, never promise to keep posting.
Deeper → The bridge: Trinity ↔ Smith
A1 · Appendix — full record
Every table and file-level detail from the working record, for whoever builds this next.
Last night, 2026-09-11 22:00 → 09-12 08:00 MT
| Time | Alert | Bots replied | Fede tagged |
|---|---|---|---|
| 22:16 | CloudWatch: metric snapshot lambda slow | Smith ×8 | 2 |
| 22:55 | Main red (budget test race), 4 later pushes also red | Smith ×15 | 1 |
| 23:02–23:06 | GitHub reviewer DOWN ×9 (4-minute token gap, self-healed) | Smith ×9 | 1 |
| 00:05, 00:11 | Sentry: cron stale; content-gate rejection | Both | 1 |
| 01:00 | Sentry ×10: no PM recipient for application review (Western Slope households) | Both, every one | 12 |
| 01:23 | Sentry: Microsoft Graph 503 on calendar event | Both | 0 |
| 02:02 | Main red again (second flaky test, fix PR open) | Smith ×4 | 0 |
| 02:15–02:19 | Co-tenant phantom-mint drift (bake + Sentry) | Both | 0 |
| 03:00 | Sentry: PM reminder no recipient (grouped, 138 events) | Both | 2 |
| 03:04–08:30 | Nightly jobs: morpheus, stress, git hygiene, latency canary, voice watch, deploy freshness, stale holds | Smith | 3 |
30-day volume by source (top-level alerts in #alerts)
| Source | Count | Classification |
|---|---|---|
| Sentry issues | 160 | Mixed: real bugs, config gaps, status counts filed as errors |
| Smith "PR waited N min, no review" self-check | 97 | Noise: bot checking itself |
| Bake alerts (nightly quality checks) | 86 | Mixed; latency canary never actioned |
| Local scripts on Fede's Mac (daily spend report; reviewer account pool empty) | 83 | Spend report is noise; the pool-empty page is the same fragility as reviewer down |
| CI main red | 80 | Real, keep |
| Smith self-narration (standalone verdicts, task status) | 62 | Noise: belongs in #agent-smith |
| CloudWatch alarms | 51 | Mostly self-resolving blips |
| GitHub reviewer down | 46 | Real but un-cooled: fires per attempt |
| Email not delivered | 25 | Config gap; digest material |
| Deploy freshness | 24 | Duplicate of main red |
| Nightly stress | 23 | Real, recurring |
| Stale-holds daily report | 21 | Report, not an alert |
| Git hygiene, turnover eval, prod cookie, voice drift, others | 91 | Mixed |
Mentions of Fede in #alerts
| Measure | Value |
|---|---|
| Total mentions, 30 days | 80 |
| In quiet hours (22:00–07:00 MT) | 63 (79%) |
| Genuine one-time decisions (strict) | ~14 (18%) |
| Repeat asks on a known gap | ~40 (50%) |
| Status or verdict noise with his handle | ~26 (32%) |
| Threads Fede replied in | 12 (15%) |
| Alerts worked by both bots | 64 |
| Bots disagreed on the same alert | 37 |
| #agent-trinity flags raised / in quiet hours | 39 / 2 |
| #agent-trinity threads Fede replied in | 19 of 80 (24%) |
Top 15 root-cause clusters
| # | Cluster | Count | Fede replied | Class | Sane rule |
|---|---|---|---|---|---|
| 1 | PR waited, no review (Smith self-check) | 97 | Yes | Noise | Once per PR per day, home channel |
| 2 | CI main red | 80 | Yes | Real | Keep as is |
| 3 | GitHub reviewer down | 46 | Yes | Real, fragile | One page per outage episode |
| 4 | GitHub Actions daily spend report | 30 | No | Noise | Weekly, #updates-fede |
| 5 | CloudWatch inbound queue age/depth | 26 | Yes | Flake | Two consecutive breaches |
| 6 | Email not delivered | 25 | No | Config gap | Daily digest |
| 7 | Reviewer account pool empty (6-hour cooldown page) | 25 | No | Real, fragile | Fold into cluster 3; fix the pool, not the page |
| 8 | Deploy behind main | 24 | No | Duplicate | Fold into main red |
| 9 | Nightly stress incomplete | 23 | Yes | Real | Dedupe same-night reruns |
| 10 | Stale escalated holds digest | 21 | No | Noise | Only on change, home channel |
| 11 | Renewal PMS-prepare infra failure | 19 | No | Config gap | Fix or demote to digest |
| 12 | Git hygiene main drift | 19 | No | Noise | Weekly digest |
| 13 | Turnover eval daily | 16 | Yes | Real | Collapse same-day reruns |
| 14 | AppFolio sync concurrency / 429 | 15 | Yes | Real, new | Watch |
| 15 | Latency canary over ceiling | 14 | No | Noise | Raise threshold or drop |
File-level findings (for engineers)
| Gap | Where |
|---|---|
| "sync eng" verdict instructs a raw member-id mention; no quiet hours, cap, or Trinity call | agent-smith/src/agent_smith/alert_remediation_prompt.py:200-203 |
| Signature keys on entity ids, so same error class × N records = N investigations | alert_signature.py:70-95, recent_alerts.py:33-60 |
| Only two sources fold repeats into one anchor | alert_episode.py:44-46 (push-main-red, claude-auth-failed) |
| Clara seat has no channel ownership set, so it owns every channel including #alerts | deploy/co.propflow.clara-slack-socket.plist; agent_ident.parse_owned_channels |
| Scheduled decider has no cross-seat claim | worker.py:731-745, select_alert_targets_activity ~868 |
| Headless bot turns load Fede's personal CLAUDE.md (bare mode removed 2026-06-06) | activities/claude_runner.py:497-519 |
| Reviewer DOWN posts once per review attempt, no cooldown | propflowai/.github/workflows/claude-code-review.yml ~2233 |
| PM action reminders run for every property, no not-live gate (allowlist removed by PR 4267, 2026-07-21) | src/lib/temporal/pm-action-reminder-triggers.ts, activities/pm-action-reminder.ts ~409 |
| Stale-mute report and goldmine scorecard post to #alerts; manifest says otherwise | /api/cron/stale-mute-report, goldmine-live-scorecard.yml, config/automations.toml |
| Daily spend report posts to #alerts even with nothing wrong | ~/.claude/scripts/gh-actions-cost-check.sh (cron 18:23 daily) |
| Trinity flag path: quiet hours, grace, presence, dedupe by label, all present and working | ~/.claude/trinity/trinity_watch.py, config.json |
PR ledger (updated as each lands)
| Repo | Change | Status |
|---|---|---|
| docs | Engineering-roadmap honest pass — measured shares, Gera's side pulled from Smith | published 733e8e6d |
| agent-smith | Sync-eng escalates through the Trinity flag; no @mention in #alerts (PR 478) | merged |
| agent-smith | Clara seat no longer admits #alerts; Smith owns alert response (PR 479); Clara seat re-rendered and restarted on Fede's Mac 12:47 MT | merged + live |
| agent-smith | Same-error-class alerts collapse into one investigation (PR 481) | merged |
| agent-smith | Reviewer-down repeats fold into one thread (PR 482) | merged |
| propflowai | PR-wait watchdog page goes to #agent-smith; #alerts only after 6 hours or a systemic outage (PR 8031) | merged |
| propflowai | PM action reminders skip not-live (sandbox) organizations (PR 7971) | merged |
| propflowai | Reviewer DOWN page cooled to one per 30 minutes (PR 7970) | merged |
| propflowai | Stale-holds report to #agent-smith, only on change (PR 7969); goldmine scorecard manifest fixed (PR 7968) | merged |
| Fede's Mac | Spend report: #alerts only when billed spend exceeds budget; Monday digest to #updates-fede otherwise | done |
| Fede's Mac | "Wrong model" rule scoped to interactive sessions | done |
| Sentry | 334 issues with no events in 30+ days resolved (568 → 234 open; auto-reopen on regression) | done |
| propflowai | Shorter plain-English alert messages: red main, deploy failed, reviewer down, production stale, stalled PR, CloudWatch headline (PR 8123) | merged b39b4f7a8 |
| propflowai | Deploy digest: one growing Slack message per day instead of one per deploy (PR 8119) | merged 439bda777 |
| propflowai | Test-bench channel noise cut (PR 8104) | merged 0d5735fef |
| propflowai | Customer-success daily "worth a look" digest from Cerebrus grades (PR 8092) | merged 717c560ed |
| agent-smith | Main lint fix (PR 498) | merged e2fdebffa |
| agent-smith | Everyday brain defaults to Fable 5.1 (PR 496) | merged 714406259 |
| agent-smith | Lane ownership map read from LANES.md (PR 495) | merged cc3a99927 |
| agent-smith | #engineering-bridge message protocol (PR 490) | merged b170d8707 |
| agent-smith | Steward publishes what's in flight every five minutes (PR 499) | merged 0a7a43926 |
| agent-smith | Fold unchanged overnight-failure nights into one line (PR 494) | merged c303e00bb |
| agent-smith | WhatsApp→#transcripts mirror: one edited reply per day (PR 497) | merged 763fe209b |
| agent-smith | Smith is the on-call engineer, Trinity is the chief of staff — role split codified (PR 493) | merged 125a3253a |
| agent-smith | Steward collects what each human's swarm is working on (PR 491) | merged dab7ff449 |
| agent-smith | Claims board for the Trinity↔Smith bridge (PR 492) | merged 383636d0e |
| agent-smith | Where a meeting's output belongs, resolved before anything is written (PR 500) | merged eaf37ff37 |
| agent-smith | Trinity no longer promises a PR review Fede never does — self-description fix (PR 509) | merged c0c63bbb2 |
| agent-smith | Both seats own #engineering-bridge, or the bridge goes quiet (PR 523) | merged 8747823ec |
| agent-smith | Voice notes to Trinity are transcribed locally, not left as a filename (PR 525) | merged 322faf4f6 |
| agent-smith | Push-gate signal test accepts both correct SIGQUIT outcomes (PR 527) | merged 2fe8dfd7e |
| agent-smith | Fede's DM with Trinity is a second SSH into the mini (PR 526) | merged 5574b8a00 |
| agent-smith | Stop appending a labeled "Plain English:" double-say restatement (PR 503) | merged b6ed336fc |
| agent-smith | Work outside your lane is handed across, not fixed in place (PR 520) | merged f90bfd587 |
| agent-smith | PR babysitter: read the event feed and decide, dry run, no actions (PR 521) | merged 98b5c9051 |
| agent-smith | A question reaches a real session over the bridge; the answer comes back in thread, capped at a 10-minute wait owned by this machine (PR 529) | merged 99ba90d0a |
| agent-smith | One plain-English bridge digest a day, per person (PR 510) | merged 6d720cf21 |
| agent-smith | Handoff follow-up: receive() keeps its idempotency promise (PR 533) | merged a0a864deb |
| agent-smith | "What are all my sessions doing" answered once, completely, by DM (PR 531) | merged f0b7ed985 |
| agent-smith | Bridge switch on by default in the templates (PR 530) | merged dd93c7863 |
| agent-smith | What a meeting produced, drawn once and quoted verbatim — transcript extraction (PR 501) | merged 5eff31815 |
| agent-smith | Trello write paths retired (PR 537) | merged e03fa9072 |
| agent-smith | Transcript router places items; the ticket path steps aside (PR 502) | merged 4df7816ca |
| agent-smith | Steward refreshes the roadmap page and flags a stale line on the plate (PR 538) | merged 71a3be6f2 |
| agent-smith | PR babysitter: Haiku fixer, behind PR_BABYSITTER_FIX, off (PR 524) | merged a0d3c7c8f |
| agent-smith | PR babysitter: merge action, behind PR_BABYSITTER_MERGE, off (PR 522) | merged 2b66d381b |
| agent-smith | DM docs chapter and plist variable for Fede's second terminal (PR 540) | merged a9647f9fd |
| agent-smith | "Ask Gera's computer" goes over the bridge and waits for the answer (PR 536) | merged 5b0443c4a |
| agent-smith | PR babysitter: notifies the owning session how the PR ended, behind PR_BABYSITTER_NOTIFY, off (PR 534) | merged 2fa2428d5 |
| agent-smith | A meeting's decisions and priorities land on next week's plan — week-plan executor (PR 539) | merged b76598105 |
| agent-smith | The canvas and the standup read shared state, not Trello (PR 545) | merged f48358d0b |
| agent-smith | A meeting's design discussions become chapter drafts on the initiative's page (PR 541) | merged 690d9d956 |
| agent-smith | Transcript page writer — chapter drafts land, build, and publish (PR 543) | merged f5fcca8fb |
| agent-smith | PR babysitter: launchd daemon plus go-live flag (PR 544) | merged 0aed379ef |
| agent-smith | PR babysitter: per-repo scope via PR_BABYSITTER_REPOS, unset means both (PR 546) | merged 90b593164 |
| agent-smith | Bridge digest follow-up: a disagreement is only marked asked when it was actually asked (PR 535) | merged d8975db7 |
| agent-smith | PR babysitter: a missed event can no longer strand a green PR — checks-passed plus a 5-minute re-check (PR 550) | merged e624cb849 |
| agent-smith | Tests: freeze the clock in the three files that go red at the UTC rollover (PR 553, Gera's) | merged 2f40d8b7c |
| agent-smith | Bridge: a channel post is not lost for lacking a stamp (PR 549) | merged |
| agent-smith | PR babysitter: a restart no longer forgets what was in flight — startup reconcile (PR 551) | merged |
| agent-smith | PR babysitter: a fixer gets its own worktree, never "fixes" a failure that isn't ours (PR 552) | merged |
| agent-smith | PR babysitter: report how a fix round ended (PR 554) | merged |
Bridge options
| Option | How it works | Pros | Cons |
|---|---|---|---|
| (a) Shared claims file + Slack heads-ups recommended | One JSON in the docs repo (lane, owner, item, claimed-by, state). Each swarm's steward writes claims before starting; Trinity and Smith post a one-line heads-up to a #bridge channel when a claim crosses a lane. Humans get a daily digest. | Reuses the lanes file, the stewards, and Slack; auditable; no new infra | Git push conflicts on the file; needs a claim TTL |
| (b) Slack-only | Foremen post claims and hand-offs as messages with a fixed shape; each parses the other's. | Zero new state | No source of truth; a missed message is a collision |
| (c) Temporal claims queue | Claims are workflow signals on the shared Temporal Cloud namespace both seats already use. | Durable, exclusive, TTL built in | Bigger build; Fede's sessions aren't Temporal-aware |
Sources: 30-day Slack pull (2,438 thread messages), agent-smith code audit, app-side alert-source inventory, Sentry unresolved query, and the raw 2026-09-11 Fede–Gera transcript. Percentages in the mentions table are classifications of message text and are labeled approximate.
How the bridge works, built 2026-09-12
Option (a) shipped. Four message shapes, one line of head plus key: value lines, parsed by one module that refuses (never silently cleans) anything that looks like a filesystem path, a shell command, or a credential:
BRIDGE CLAIM v1
lane: multi-tenant isolation
item: fix cross-org read on /api/units
seat: smith
human: Gera
state: claimed
until: 2026-09-12T21:00:00+00:00BRIDGE HANDOFF v1
lane: multi-tenant isolation
item: Western Slope user saw another org's units
why: isolation is Gera's lane; found by Fede's alert run
link: https://...
to: smithBRIDGE QUESTION v1
id: q-3f81a2c9
to: smith
lane: multi-tenant isolation
ask: what is the current shape of the org wall resolver?
reply-by: 2026-09-12T20:30:00+00:00BRIDGE ANSWER v1
id: q-3f81a2c9
by: smith
answer:
source: state-file | session | Piece | What it is | Where |
|---|---|---|
| Claims board | The one shared, human-readable file both stewards read and write; a claim also takes an atomic lock on the shared Temporal namespace so two seats can never win the same item | artifacts/bridge-claims.json in this docs repo |
| Per-person state | Each human's own "what's in flight" file, written by that seat's steward from its git branches, PRs, sessions and Trello — never by the other seat | bridge-state-fede.json, bridge-state-gera.json, next to the claims file |
| Protocol + relay + handoff + digest modules | Format/parse the four shapes, post over the existing seat-to-seat Slack rail, answer questions, redirect out-of-lane work, and roll up one digest per human per day | bridge_claims.py, bridge_protocol.py, bridge_relay.py, bridge_handoff.py, bridge_digest.py in agent-smith |
How a question gets answered. A seat answers a QUESTION addressed to it in order, and stops at the first source that has an answer: its own steward's state file; failing that, one of its own interactive sessions running on that same Mac; failing that, the claims board plus the lanes ownership file; and only if none of those answer in time, one single alert to the human, never a flood of them.
Sealed, on purpose. The bridge carries only the four message shapes and the claims/state files above — never a command to run, a filesystem path to open, or a credential. A seat only ever starts a local session on its own person's Mac; neither swarm can act on the other person's computer, full stop.
Left for Gera. The code is written and dark (nothing posts until the Slack channel is set). On his Mac he still needs to: pull the change, install the launch agent for the Smith seat, and set the bridge's Slack channel id in that seat's environment. The channel itself is #engineering-bridge.
Slack channel inventory — 26 public channels, 2026-09-12
| Channel | Members | Created | Msgs/30d | Human/30d | Distinct humans | Last human | Bot posts/30d | @Fede | @Gera | Verdict | Reason |
|---|---|---|---|---|---|---|---|---|---|---|---|
| #module-onboarding | 4 | 2026-06-29 | 1 | 0 | 0 | never | Trinity×1 | 0 | 0 | RETIRE | 0 human msgs/30d, 1 one-off Trinity post, no automations.toml entry |
| #general | 5 | 2026-06-29 | 16 | 16 | 2 | 2026-09-10 | — | 1 | 0 | KEEP | Company chat; only real, sustained human back-and-forth channel |
| #module-renewals | 4 | 2026-06-29 | 0 | 0 | 0 | never | — | 0 | 0 | RETIRE | 0 msgs/30d at all, no automation |
| #module-tours | 4 | 2026-06-29 | 0 | 0 | 0 | never | — | 0 | 0 | RETIRE | 0 msgs/30d at all, no automation |
| #agent-smith | 4 | 2026-06-29 | 229 | 48 | 2 | 2026-09-09 | Smith×173, Trinity×7 | 4 | 6 | KEEP-AGENT | Smith's home; Fede likes the agent channels |
| #apartment-camellia | 5 | 2026-06-29 | 222 | 6 | 2 | 2026-09-08 | Smith×216 | 3 | 1 | KEEP | Fede explicit ask; live customer property |
| #tickets-bugs | 5 | 2026-06-29 | 7 | 6 | 2 | 2026-08-31 | Trinity×1 | 2 | 0 | RETIRE-AFTER-REWIRE | Trello taxonomy (agent-smith trello/taxonomy.py) maps here; near-zero recent traffic, Trello board is the real record |
| #updates-gera | 3 | 2026-06-29 | 174 | 109 | 2 | 2026-09-12 | Smith×48, Trinity×17 | 1 | 50 | MERGE→propflow-engineering | 109 human msgs/30d, most active human channel; fold hand-offs+decisions in |
| #tickets-ui | 4 | 2026-06-30 | 1 | 1 | 1 | 2026-08-25 | — | 0 | 0 | RETIRE-AFTER-REWIRE | Same Trello taxonomy mapping; 1 human msg/30d |
| #tickets-infra | 4 | 2026-06-30 | 0 | 0 | 0 | never | — | 0 | 0 | RETIRE-AFTER-REWIRE | Same Trello taxonomy mapping; 0 msgs/30d |
| #updates-fede | 4 | 2026-06-29 | 162 | 1 | 1 | 2026-08-31 | Smith×7, Trinity×154 | 7 | 9 | KEEP | Fede's explicit ask: ships/major landings, on demand |
| #transcripts | 5 | 2026-06-30 | 141 | 4 | 1 | 2026-09-01 | Smith×137 | 27 | 19 | KEEP-AGENT | Zoom/WhatsApp mirror automation; heavily read per raw-transcript rule |
| #tickets-landing | 4 | 2026-06-30 | 0 | 0 | 0 | never | — | 0 | 0 | RETIRE-AFTER-REWIRE | Same Trello taxonomy mapping; 0 msgs/30d |
| #updates-propflowai | 4 | 2026-06-29 | 18 | 18 | 1 | 2026-09-08 | — | 3 | 0 | RETIRE | Duplicates general/updates-fede company-update purpose |
| #apartment-yale | 4 | 2026-06-29 | 133 | 1 | 1 | 2026-09-02 | Smith×132 | 0 | 0 | KEEP | Live customer property (parked), nudge-digest automation posts here |
| #marketing | 3 | 2026-06-30 | 17 | 1 | 1 | 2026-08-19 | Smith×16 | 0 | 0 | RETIRE-AFTER-REWIRE | marketing-*.toml jobs (4 automations) post here; fold into agent-smith |
| #module-maintenance | 4 | 2026-06-29 | 0 | 0 | 0 | never | — | 0 | 0 | RETIRE | 0 msgs/30d at all, no automation |
| #tickets-features | 4 | 2026-06-30 | 1 | 1 | 1 | 2026-09-10 | — | 0 | 0 | RETIRE-AFTER-REWIRE | Same Trello taxonomy mapping; 1 msg/30d |
| #incidents | 3 | 2026-07-08 | 1 | 1 | 1 | 2026-08-19 | — | 0 | 0 | RETIRE | 1 old human msg, no confirmed automation poster |
| #alerts | 4 | 2026-07-11 | 864 | 3 | 1 | 2026-09-07 | Smith×614, Trinity×85, Sentry×162 | 11 | 7 | KEEP | On-call; Fede explicit ask |
| #deployments | 4 | 2026-07-31 | 1414 | 0 | 0 | never | Smith×1414 | 0 | 0 | RETIRE-AFTER-REWIRE | 1414 Smith posts/30d, 0 humans ever; fold deploy noise into agent-smith |
| #apartment-willows-test | 4 | 2026-08-21 | 1863 | 2 | 2 | 2026-08-31 | Smith×1832, Trinity×29 | 1 | 2 | KEEP-AGENT | Standard test-property bench; Smith posts heavy soak/smoke noise here, keeps it off prod channels |
| #module-collections | 4 | 2026-08-24 | 4 | 3 | 3 | 2026-08-24 | Smith×1 | 1 | 1 | RETIRE | 4 msgs/30d, no automation; discussion can live in propflow-engineering |
| #agent-trinity | 4 | 2026-08-23 | 81 | 10 | 3 | 2026-09-12 | Smith×1, Trinity×70 | 45 | 1 | KEEP-AGENT | Trinity's home + blocked-on-Fede asks; Fede likes the agent channels |
| #apartment-situs-group | 4 | 2026-09-01 | 9 | 8 | 3 | 2026-09-03 | Smith×1 | 2 | 1 | KEEP | Fede explicit ask; live customer property |
| #apartment-yale-test | 1 | 2026-09-04 | 14 | 0 | 0 | never | Smith×14 | 0 | 0 | RETIRE-AFTER-REWIRE | Sandbox test traffic (Smith x14/30d), 0 humans ever; fold into apartment-willows-test |
Read via the fede-user Slack token (no groups:read scope — private channels are invisible to this audit and not counted). "Bot posts" is by seat: Smith U0BE08K74M7, Trinity U0BEM6NPK26, Sentry U0BKU7NLQ2J. RETIRE-AFTER-REWIRE names the automation still posting there (automations.toml + agent-smith's Trello taxonomy), so its owner knows what to repoint before archiving. #western-slope is not in this table — it did not exist at audit time; see Chapter 9 for its creation status.