Smith review — your calls, then I drive

Everything from the broken-sticker review is fixed or staged. These 5 answers are the only things between here and done — answer any, in any order; each one immediately unblocks its lane.

DECISION 1

Ship the two fix branches?

Both branches are committed and gate-green (ruff + mypy + pytest): the delivery fix (a reply Slack refuses to render now re-sends as plain text instead of being thrown away — the bug behind every "u seem stuck" in the thread) and the ack-lifecycle fix (no more duplicate "mid-task…" acks, the watchdog now tracks every promise).

Held for you because merging agent-smith restarts and redeploys the worker Smith runs on — a live-system change, not just code.

In plain terms

Smith kept saying "on it" and then going quiet — turns out it HAD the answer but threw it away before posting. The repair is written and tested; shipping it briefly restarts Smith itself, so I want your OK.

DECISION 2

Merge the two open cleanup PRs (#215, #216)?

#215 deletes the "redeployed — new brain online" auto-posts (your binding pref says these should never post). #216 collapses repeated auth-failure alerts into one message instead of ~12 (the 401 storm you flagged in #alerts).

In plain terms

Two noise-cutters, already written and reviewed. One stops Smith announcing its own restarts; the other makes 12 copies of the same alarm into 1. Same deal as Decision 1: merging restarts Smith.

DECISION 3

The three red nightly evals — fix, triage-first, or quiet the alarm?

maintenance-eval-daily 12 nights red · morpheus-nightly-daily 14 nights red · turnover-eval-daily red. One is a false alarm (the eval actually ran and scored, but finished 6 min past its check-in deadline); the others need real triage.

In plain terms

Three nightly self-tests have been failing for about two weeks, sending an alarm every night. Nobody has decided whether to fix the tests, investigate first, or stop the nightly siren while we work. I won't silence an alarm without your say-so.

DECISION 4

Green cron chatter — keep or cut?

You stickered a machine-hygiene — 3/3 substeps ok post. That's a cron reporting success. Cutting it means green runs go silent and only failures post.

In plain terms

Right now some routine chores post "everything's fine" every time they run. Do you want to only hear about them when something's wrong?

DECISION 5

Run the full conversation-quality sweep?

The review so far was flag-focused — it followed your stickers. A full pass also grades every human↔Smith exchange in the window (pokes, repeated instructions, wrong-confidence answers, silence gaps) and sets a baseline so "is Smith getting better?" has a number next month.

In plain terms

So far I only studied the messages you stickered. A full sweep reads every conversation and scores what Smith was like to work with — then we can measure whether these fixes actually made your week smoother.

ANYTHING ELSE

Curveballs

Anything the review missed, or a constraint on how I drive the above.

PropFlow Docs