0014 — The Oracle explains a low-graded Clara turn and holds the fix for a human
Not the "oracle guard." propflowai ADR-0086 uses oracle for the human-approval bit on a generated eval scenario (provenance.approved). This ADR is about the Oracle, Gera's name for the Agent Smith workflow that explains a low-graded Clara turn. The two share a word and nothing else.
- Status: Proposed (flips to Accepted the first morning the
oracle-nightly postedsignal is green and the flagged Camellia call's thread carries its finding) - Date: 2026-09-08
- Deciders: Gera, 2026-09-08 — "think about the type of design and implementation and outcome that this oracle should go through … see if we can auto trigger it for that thread that I kicked off so I can see what it would look like." Designed by a three-lens judge panel (data-flow / safety / smallest-diff, two independent judges) inside eleven decisions the Architect (session
2708dff9, grade-clara) fixed first. Written at decision time because ADR-0004 says a decision that changes how systems relate is not done until it is — and this one adds a relationship the fleet has never had: grading → Smith.
The gap
Since 2026-09-07 every Clara turn is graded by a computed number (propflowai turn-grade.ts: fixed deductions over deterministic findings, never a judge's opinion) and rendered as one line per turn in the property's ops thread. A turn below 7 shows its reason. Nobody explains it. The reason line says what fired — "said she would check pricing and availability, then only transferred the call" — and stops. Which layer failed (the prompt, a tool that is not bound, a result she ignored, a policy gate), and what would keep it from happening again, was a human reading three repos.
Gera asked for the thing that does that reading: one more Smith specialty workflow in charge of exactly one job — turning a low-graded turn into a grounded explanation and, when the cause is a prompt or tool gap, one held PR.
Decision
- One door. The grading seam (
postScorecardIfNew, the one function every detector already reports through) gains an injectedonLowTurnhook. It fires inside the per-turn posting loop, after the turn line's Slacktscomes back, when the computed grade is below a named constantORACLE_THRESHOLD = 8, and only for turns that started within 72 hours (a backfill of days-old rows must not become a storm). The hook's default implementation is a fail-soft Temporal start on the tools-prod namespace —OracleTurnWorkflow, idoracle-<conversationId>-<turnKey>, conflict policy FAIL, 2-second bound, outcome logged (started | already_started | refused:<reason>). It never throws into the seam. There is no polling, no second detector, no cron. - One evidence source. The Oracle may quote only what
scripts/oracle-evidence.ts(propflowai) prints. That script rebuilds the turn with the same pure modules production graded it with, reads the persisted grade rows, resolves who spoke (the call'stransfer_to_agentresults carry the agent ids), reads the tool catalog and the speaker's prompt file at the pinned worktree's HEAD, and prints a byte-deterministic bundle with an explicitquotes[]list. Smith shells it from a worktree pinned toorigin/main. No Python port of the turn rebuild exists or will. - The bucket is a table, not an opinion. A decision table over the bundle picks one of prompt gap / missing tool / tool result ignored / policy / grader gap / could not determine, names the fix layer (prompt, tool binding, tool handler, personalization injection, detector, none), and says whether a PR opens (prompt gap, missing tool and grader gap only). Grader gap is the fifth bucket, added by the Architect on 2026-09-08: the check fired on a shape the grader does not follow — today, a sibling hand-off whose promised subject the receiving agent delivered on a later turn of the same call — and the held PR proposes the detector change, not a prompt change. "Could not determine" is a real outcome, posted with the evidence that would settle it. The bucket is decided once, in that TypeScript script; Smith only renders it (a second decider in Python would be a drift the live instrument fails). No LLM writes the finding: it is a template over the bundle, validated before posting — every quoted span a substring of the bundle, no resident name/phone/email/address, no digit run that could be a phone number.
- The fix is proposed, never merged. A PR-bearing bucket takes a fleet-scoped work claim keyed on a cause signature (
check key | what followed | speaker | fix layer | fix target), resumes an open PR when one exists, and otherwise hands Smith's brain — in#agent-smith, never in the property's thread, so the thread keeps exactly one Oracle reply — a prompt that names the bucket, the target file and the evidence and forbids re-diagnosis. The PR's link reaches the finding by an in-place edit after a workflow-owned wait, never a second reply. The PR carrieshold-for-reviewand auto-merge disabled, both read back; its body names the fleet blast radius (Camellia is the shared registry — a prompt merge deploys everywhere). Two low turns with the same cause yield one PR; the second finding links it. - A nightly line, from rows. The morning queue gains one step,
oracle-nightly, that computes per live property the average turn grade, the count below 8 and the delta against the trailing seven days from the persisted grade rows (one bounded GSI3 window, grouped per property) — never from Slack — and counts the turns it could not measure. Full mechanics:0014-panel-synthesis.mdbeside this file.
How it relates to the ladder
- ADR-0002 — agent↔agent is SendMessage, tmux is hosting. Nothing here attaches to a pane. The seam reaches Smith by a Temporal start; Smith's brain is reached by the same child-workflow door a scheduled remediation uses.
- ADR-0003 — rungs integrate by signals, never calls. Grading does not await the Oracle; it starts a workflow and reads nothing back. The Oracle does not await the fix; it starts an abandoned child and later reads the PR list. What the Oracle produces are files and rows other rungs read: the dispatch log (its ignition and its heartbeat), the AutomationRun rows, the finding in the thread.
- ADR-0009 — a signal must fail when the system fails. The
oracle-turnlog line is written only when a workflow actually ran to an outcome; theoracle-nightly postedline only when the morning step posted. A refused start in the seam writes a propflowai log line, not the Smith signal — so a dead Oracle goes red within 30 hours rather than reading as a quiet week. - ADR-0011 — an unread seat is unknown. The catalog entry declares both signals before the writer exists (
proposed), so the row reads unknown until the first real run, never healthy.
Consequences
- The catalog gains
oracle, proposed. Signals:log-match-ageon~/.claude/smith-state/oracle/dispatch.logfororacle-turn(14 days — low turns are sparse) andoracle-nightly posted(30 hours — the true liveness). It flips toactivewith no signal edit once the writer logs. - The propflowai worker needs a second Temporal namespace.
SMITH_TEMPORAL_ADDRESS/NAMESPACE/API_KEYon the ECS task definition is an infra act for Gera; until it lands the hook logsrefused:no_tools_prod_clientand the Oracle is reachable only by hand (python -m agent_smith.oracle_cli run, the same door, the same workflow id). The template edit that declares the secret must land after the secret exists — the deploy preflight refuses an unreadable secret. - What the Oracle cannot see, said plainly. A stale injected price with no lookup behind it grades 10 with no reason (propflowai
run-turn-checks.tsheader) — the Oracle never meets that class. A pre-field voice call has no persisted briefing; an unbacked figure on such a call is "could not determine", settled byAgentTrace.injectedContexton later calls. A turn the grader never scored has no row and is invisible to the nightly line, which says so in its footer. - The flagged Camellia call decides by whether the promise was kept across the hand-off. Turn 2 was spoken by the front-door (triage) agent, whose prompt scripts "Let me check pricing and availability for you" as the hand-off line before
transfer_to_agent(triage.ts:457, an owner ruling at:482); triage binds no lookup tool; the detector's window sees this turn's tool calls only, so a sibling keeping the promise is invisible to it — whileleasing.ts:236expects the sibling to deliver pricing. So the bundle carrieshandoffDelivered: if the receiving agent quoted or looked up the promised subject on any later turn of the call, the cause is grader gap and the held PR proposes that the detector follow hand-offs; if it did not, the cause is prompt gap and the held PR proposes the prompt line. Both PR bodies cite all three rules, name both candidate layers, and leave the choice to the person who merges.
Alternatives considered
- Consume the seam's outcome in the activity, or a fourth trigger kind, or a Slack-side listener, or Smith polling the grade rows. Each is a second door or a poll; the one-seam drift test exists to refuse the first two, and "not from Slack" and "no polling" are in the criterion. Polling survives only for the nightly line, once a day.
- Let an LLM pick the bucket, or write the finding. Smith's in-thread root-cause replies have invented mechanisms with real line numbers before. A table over the record cannot; a template over the bundle cannot quote what is not there.
- A Python rebuild of the turn in Smith. A second seam that drifts from the first; refused.
- A standalone
oracle-nightlyschedule. Registers paused, needs an unpause, and duplicates the morning banner. A step in the existing queue rides the same schedule and the same thread.