What needs you

Ordered by what you get back for the time it costs. Nothing here was merged, pushed or toggled overnight.

2026-08-03 · compiled from the overnight run · every state below was checked against GitHub, not remembered

This page is a snapshot, last revised 14:09 UTC, and it does not update itself. Every item names the exact commit or record it rests on, so you can tell at a glance whether it still holds — if a pull request's head has moved since, the line about it is stale and the linked page is the truth. Said plainly because the last decision page I left un-dated went on offering an option to land something that had already merged.

Re-checked at 13:53 UTC: items 01, 03 and 04 are all still exactly as described. Same commits as when this was written, each still carrying the reviewer’s approval against that same commit, and all three still merge cleanly. #4868’s single red check is the one its own description tells you to expect. Nothing about those three has moved all morning — they are still one click each.

Two minutes, real value

01

#4837 — press merge

one click

Approved by the reviewer against its exact current commit (5e955ea7f), all four required checks green, no labels, no conflicts. It touches a workflow file, so auto-merge is structurally forbidden from arming it — that is the only reason it is still sitting there. It needs a human to press the button and nothing else.

This is your own instruction, already carried out as far as it can be. On Saturday at 22:52 you answered “Rebase and merge the ADR — the policy is right and the crons migrate later”. Everything in that sentence except the final word was done: rebased across 405 commits of drift, the document renumbered 0118 → 0123 because main had claimed 0118 meanwhile, 33 references swept, and three full rounds of review. The policy you approved is the policy sitting there.

It stops at the merge because it changes a build-system file, and that category is gated on a person reading the diff — no clean review and no green CI lifts it. Worth knowing: the guard on the PR names the owner as Fede, or whoever owns that surface, rather than you specifically. So this is one click, by one of you.

Open #4837

01b

#5229 — deliberately not rebased, and the reason is good

one judgement call

On Saturday at 22:51 you answered #5229 with “Rebase and run one more review round”. That was picked up twice today and deliberately held both times, for a reason your instruction could not have known: #5230 merged six minutes after this PR was last touched, and cut the one file this PR edits down to a 43-line pointer. Rebasing in this PR’s favour would bring back roughly 257 lines that #5230 deleted on purpose — re-creating the two-contradicting-documents mess #5230 existed to clear up.

So the honest position is not “nobody did it”. It is: doing what you asked would now undo something newer. It is also the cheapest of the four conflicting PRs to rescue — one file, no code, no tests — but it carries no approval, and its subject already has a live tracker in #5004. The real question is whether you still want it at all, not whether to rebase it.

Open #5229 · the tracker, #5004

01c

I said your decisions were being dropped. I was wrong — every one was acted on

retracted 12:14 UTC

Read this before the rest of the item, because most of what follows it was written on a premise I have since disproved. Over the last hour I told you that four of your six recorded answers had gone unexecuted, and built a case that decisions here evaporate. I checked every one of them properly and that is not true. All six were acted on. What follows is kept rather than deleted so the claim and its refutation sit together.

The three I called dropped, in fact: #4837 was rebased across 405 commits, renumbered, swept across 33 references and reviewed three times — it stops only at the human-merge floor. #5229 was picked up twice today and deliberately held, because a PR that merged six minutes after it would have ~257 deleted lines resurrected by the rebase you asked for. The instructions-file trim ran to a considered stop: the obvious next group was inventoried, then the case for moving it was refuted with sentence-level evidence — those sections are not duplicated anywhere, so moving them would delete knowledge, not relocate it. Its tracker was closed deliberately.

So why did I get this wrong four separate times? Every one of those outcomes was written down — in a comment on a ticket. None of them was visible from the thing I was looking at, which was the code. I checked for work by checking for commits, and considered thought, evidence and a documented refusal to act to be silence. Two of the comments that proved me wrong were written by me, earlier today.

The real finding is smaller than the one I published, and I think more useful: the reasoning is all there, and it is scattered across a dozen ticket threads where nobody sees it in one place. From outside, careful handling and neglect look identical — identical enough that the person auditing for neglect found it four times in a row, in his own work. That is not a diligence problem, and asking anyone to be more careful will not touch it.

Having found two, I checked every other page that has ever recorded an answer. Four more of yours are on the parked-decisions page, from Sunday morning. Two were done properly: the voice ceiling you asked to be built now shipped, and the collections work you typed a spec for went in. Of the remaining two, one turned out to be in good hands and one did not — and I got that backwards at first, so both are worth reading:

“Go get the older data” — you deliberately picked the most expensive option so that “All time” on those cards would stop being a lie. The picker fix you may be thinking of landed the day before you answered; it made the controls work, it did not go and get the history. Corrected 12:10 UTC — I had this one wrong, and wrong in the exact way this page keeps warning about. I wrote that no record existed of anyone fetching the data. True of the change log, and misleading, because I searched code and never opened the ticket. This decision was handled better than any other on this page. Your choice was written onto #5260 the same day, including your “even if it’s hard”, and then the first question was answered with evidence: the older data must not go where I proposed putting it. That store deletes anything older than 400 days, so a 2018 figure would arrive already expired and vanish silently — no error, just rows quietly ceasing to exist. Three separate notes in the code had predicted precisely that.

So the goal survived and the method changed: give occupancy and rent income the same treatment the NOI card already has — a permanent series that never expires — and historical occupancy turns out to be calculable from lease dates we already hold, so no fetching is needed for half of it. Rent income probably only exists monthly that far back, which the plan says to show honestly rather than smooth over. What it lacks is not thought — it is somebody to pick up the next step.

SUPERSEDED — the four paragraphs below are what I originally wrote, kept so the claim and its refutation sit together. Two of them are now known to be false (the trim was not abandoned, and the next group was brought back and refuted). Do not act on them.

“Keep going to 200–300” — the instructions file. Real work followed: it came down from 2,124 lines to about 1,350 that same day. Then it stopped. It sits at 1,363 today, roughly five times the number you asked for, and has not been touched for that purpose since Sunday night. Not ignored — abandoned partway, which is easier to miss because the early progress looks like success.

A third page pins that one down exactly. On Sunday afternoon you also wrote “Move the 236 safe lines now, bring me the next group after”. The first half was done within the hour. The next group was never brought to you. That was not a vague target you set — it was a specific promise to come back, and coming back is the part that did not happen.

In fairness, the same page shows the system working: on that page you also said “push all three, open PRs”, and three pull requests were open within five minutes and have all since merged. So this is not a story about instructions being ignored. It is narrower and more awkward than that — the ones that get done are the ones with an obvious next action, and the ones that ask for a follow-up later are the ones that evaporate.

The pattern is the point, and it is worth more than these four items: your answers are captured reliably and then nothing ever checks whether they happened. There is no step, anywhere, that asks “he decided this — did it get done?” So a decision that is acted on and one that is quietly dropped look exactly the same afterwards. Six answers, four surfaces, and the only reason any of this surfaced today is that I went digging on a hunch.

With the correction applied, the count is three, not four — the ADR merge, #5229, and the group of sections you asked to be brought back. That is the honest number, and my own mistake proves the point better than the tally does: I could not tell a dropped decision from a well-handled one either. The difference was sitting in a comment on the ticket, and I concluded from the absence of code that nothing had happened. If the person auditing for this failure walks into it while auditing, no amount of care fixes it — only something that actually links a decision to its outcome does.

I then checked the one that could actually hurt someone, and it is fine. On 30 July you set the condition for switching rent reminders on: “after Gera reads a week of practice messages and is happy”. The lane is still in practice mode in the code — it writes the messages and files them, and sends nothing to anyone. That is verified against the setting itself, which has never once been switched over, and the operator screen is prevented from claiming otherwise. No resident has been texted about rent.

The half that has no protection is the same half as everywhere else on this page: somebody has to bring you the week of practice messages. That is not overdue — a week from the 30th is around Thursday — which is exactly why it is worth naming now rather than discovering on the day. The safety condition you set is enforced by the code. The follow-up you asked for is enforced by nothing.

The parked-decisions page · the rent-reminder answers

02

Say whether to push the loop backup

one answer

Ten finished commits sit on this disk across three repos. They are not equivalent, so the question asks by category rather than by count. Category A is the one that matters: the tooling that runs the overnight loop had no version control at all — including the hook that holds it — and is now committed locally. Its own notes record that this directory has been wiped by a hard reset at least twice, and a local commit is exactly what that wipe destroys. Pushing is what makes the backup survive it.

Added 11:47 UTC — this question is worth more of your attention than it looks, because it is also a handbrake. Nothing may be pushed without your say-so, so any new work finished on this machine joins that same pile. I had a fix ready to hand out for item 07 and did not, for that reason: starting it would have added an eleventh commit to the exact heap you are being asked about. This question has already had to be rewritten twice for precisely that reason — the last version was withdrawn because it said “six commits” and had become ten while it sat, as other sessions kept finishing things. Until this is answered, the honest thing for me to do is not start more implementation, which is why the queue below is quieter than the backlog would suggest.

Answer on the decisions page (b85746127)

One click each, but read the sentence first

03

#4851 — mergeable, and only a label is holding it

click = merge

Approved against its exact head, clean, required checks green. Unlike #4837 it does not touch a protected path — so auto-merge can arm here, and the hold-for-review label is the only thing stopping it. Removing that label is the decision to merge, not bookkeeping: once a verdict exists at the head, a later workflow pass can arm and merge with no further human action. That is precisely how #5321 landed yesterday, ten minutes after its label came off.

Open #4851

04

#4868 — merging changes live phone calls

customer-facing

Also approved at its exact head with required checks green. Its one red check is expected and the PR says so. The part worth pausing on is in its own description: the sync runs automatically once it lands, so merging pushes the new hand-off behaviour to the live voice agent across all seven transfer routes. That is a change to how real callers get transferred, not a routine merge.

Open #4868

04b

#5330 and #5329 — Agent Smith finished; both are now merged

nothing to do

Smith triaged the overnight nightly alert on a bounded run and stopped, correctly, having posted its verdict. #5330 is the good one: the shared eval provider could not tell “the model was silent” from “the model regressed” — both scored as a failed assertion — so a documented API flake could redden the gate at any time. It now retries once on an empty completion, while a thrown error still surfaces immediately and a second consecutive empty is returned as-is. That makes the whole eval suite a more trustworthy gate, not just this one test.

Correction — this item asked you for something, and no longer needs to. It previously read: “Both are blocked on the same thing: neither has a formal approval, only comments … that is a call for you, not for me.” That was true when written and stopped being true about twenty minutes later. The reviewer went on to approve both — and re-approved each time Smith pushed a fix, five times running on #5329 — landing a clean verdict on the exact version that exists now. Every check passed on that exact version, and neither touches a protected path. So the bar we already agreed on was met, and I merged them both. Nothing here is waiting on you.

Worth noting why I got it wrong, since the same shape will recur: I read the state of a live thing and then wrote about it as though it were settled. Smith was still working while I was describing it as stuck. A snapshot of something still moving is a claim with an expiry date — which is why every item on this page names the version it rests on.

It also closed a question I had left open. The other two red lanes in that nightly — routing/triage at 79% against a 93% threshold, and SMS stress at 88% — were already red the night before. They are chronic and unrelated to this alert, and they need their own look: Smith counted nine genuine cases of a transfer not firing in routing.

#5330 · #5329 · background, #5328

Questions already waiting for you

05

The contact-ceiling PR (#5232) — two questions

two answers

One asks whether the cap should cover phone calls as well as texts; the other asks whether to ship it. Worth knowing before you answer the second: it is not shippable as it stands — the reviewer's approval sits five commits behind the current head, so there is no verdict against what would actually merge. The tests at the head are green. It needs a fresh review, not a fresh decision.

Answer on the decisions page (b85733726, b85733749)

06

Six questions from last night, still untouched

six answers

A page of six decisions covering nineteen pull requests, with no answers on it yet. It now carries a correction at the top: seven of those nineteen have since merged or closed, so some options are no-ops — in particular one that offers to land a PR that is already in. Twelve are still open, so it is stale rather than dead and still worth the pass.

Open the page

04c

#5333 — finished, correct, and stuck for want of one click

30 seconds

This morning a session worked three alerts and opened three fixes. Two of them merged themselves and are live. The third is done — the fix is written, a review round found one genuine mistake in it, the session corrected that at 13:40, and every check since is clean.

It is stuck on a technicality. The final test run is deliberately held back until a reviewer looks at the change — a sensible rule, so the expensive tests only run on code someone has actually read. But the last review happened before that final correction, so it does not count for the version that exists now, and the session that would have asked for a fresh look has since finished and gone. Nothing is wrong with it and nothing will move it on its own.

Someone asking for one more review round releases it, and then it merges itself like its two siblings did. I have not done that myself — requesting a review on another session’s work is how a change gets approved without anyone having read it, which is the one thing the rule exists to prevent.

What it fixes, for context: an alert claimed three emails to a lead were never recorded. They were — the checker simply could not find them, because our replies were not carrying the subject line it matches on. So this is a false alarm being silenced properly, at its cause.

Open #5333

Filed overnight — no action needed today

07

#5327 — the type-check safety net is down

needs an owner

Every pull request runs a fast type checker; a slower, authoritative one runs once a night purely to confirm the fast one is not missing anything. Last night the slow one ran out of memory and could not finish — so that confirmation is not happening. Nothing is broken in the code. Second time this ceiling has been hit as the codebase grew.

Corrected 11:38 UTC — this line previously said “needs a human merge”, which implied a fix was written and waiting for your click. There is no such change. #5327 is a written-up problem with nobody assigned: the likely fix is small, but somebody has to write it, and it lands on a protected path so it will need your merge afterwards either way. The ask here is who picks it up, not which button to press. Left alone, it stays unconfirmed until tonight’s run fails the same way.

There are idle sessions that could write it now, and I held off deliberately: the fix would become another commit that cannot be pushed until you answer item 02, and that pile growing under the question is what forced item 02 to be rewritten twice already. Answer 02 and this one can move the same hour.

Read #5327

08

#5328 — an eval gate that could not tell silence from a real miss

fixed · one question left

Corrected 11:16 UTC — my first explanation of this was wrong and is worth flagging because it was on this page. I said a recent change let the assistant take an action instead of replying, and the grader scored that as blank. That mechanism is impossible: a tool call is serialised into a non-empty result before scoring. Another agent read the provider code and refuted it.

The real cause is duller and better evidenced: a genuinely empty reply from the model — a known flake, recorded three weeks ago in the change that introduced this test, which failed two runs in four and passed on re-run both times. The eval could not distinguish that from a real failure, so any night it struck, the gate went red for no reason. #5330 fixed it by retrying once before scoring, and it is merged and live on main as of 11:31 UTC.

The half that still needs a decision is unrelated to any of that: a behaviour-changing change can currently merge while its behaviour check truthfully reports that it evaluated nothing, recorded in a way that neither blocks nor reads as a failure.

Read #5328 · the fix, #5330

09

#5331 — texts are taking about 2½ minutes to get a first reply

resident-facing

Corrected 11:23 UTC — my first version of this said a nightly check was simply broken. It is not, and the truth is worth more. That check separates known-broken cases from real ones properly: nine are quarantined with a written reason and an owner each, and it reports exactly one real failure. The same one every night.

That real failure is how long a resident waits for Clara's first reply: 140, 146, 148, 155 and 141 seconds on the last five nights, against a threshold of 120. Never once inside it. This is not new, and that is the more useful fact: an investigation seven weeks ago measured the same thing at 90–173 seconds, proved it was real rather than a testing artifact, and traced the cost to one background service rather than to Clara herself. It was left open. Since then the acceptable limit was doubled from 60 to 120 seconds — and it still fails every night.

Why nobody has looked at it for six nights: the failing check lives inside a test named for something else entirely — opt-out handling. Anyone reading the failure reasonably checks whether “STOP” still works, finds that it does (it does — I checked the transcript turn by turn), and moves on. The slow-reply assertion is a passenger in someone else’s scenario, so its red gets dismissed every time.

Worth knowing before anyone is asked to look at it: the earlier investigation already named the likely cause — a background call that times out on essentially every message, which is also a correctness gap because a maintenance step is not being triggered at all. So the next step is picking up an existing analysis, not starting a new one.

The other half of that issue is a routing regression that started two days ago: seven cases where a call should hand off to a specialist and does not, clustering into vendor/handyman recognition and the property-manager turnover walk.

Read #5331

Rotting quietly

10

Four PRs now conflict with main

gets worse daily

All four have been costed since this page was first written, and each now carries the measurement on its own thread. They are not one job:

#5033 — close it, do not rescue it. The workflow it fixes no longer exists: it was deleted from main by the revert of #4987. Merging would bring back a file the team deliberately reverted away. It reads as the healthiest of the four because it is the only one approved against its exact head, which is exactly what makes it a trap.
#5229 — one conflicting file, and it is the single planning doc the PR changes. Cheap to fix, but unapproved, and its subject already has an open tracker (#5004). The question is where the correction should live, not whether the rebase is affordable.
#3087 — still real work, and smaller than it looks. Of 217 stale-namespace lines, only 18 across 10 files actually change behaviour; 99 are documentation, and 14 of those are ADRs recording past decisions — sweeping those would falsify the record rather than fix drift.
#4705 — 36 conflicting files, the deepest by far, and its thread notes that a rebase invalidates the audit's baselines. So it drifts further each day while the cheap remedy stays unavailable.

Separately, #5236's approval is eighty-one commits behind its head; treat it as unreviewed. Its subject, unlike #5033's, is still fully present on main.

Open the PR list

Nothing was merged, pushed, toggled or labelled overnight. Every claim above was re-checked against GitHub before being written here — several were wrong on the first reading and are corrected on the linked pages rather than quietly fixed.

PropFlow Docs