For Fede · August 13, 2026 · One document: how the system behaves today (§1–3), what the rot actually is (§4), and the first-principles two-track rethink (§5–7). Evidence: the full Camellia corpus mined since launch (203 findings, all channels), the per-channel fork audit of origin/main, the test-infrastructure map, deep research on human-in-the-loop escalation, and the Colorado renter-journey question bank. Nothing here is implemented; the only drafted code (voice transfer-first) is parked pending §5.4 Decision 1.
The one-screen version
What broke: a hot applicant asked "I just applied — what do I do now?" at 11 PM and got silently forwarded to nobody; his call the next morning was brushed off with "someone is handling this personally." Mining every conversation since launch shows this is the norm: Clara answers 37% of real questions (16% after application, 3% for vendors), 44% of everything humans handled was answerable from a fact sheet, and at least 9 transfers dead-ended in voicemail — including a car break-in report.
Why (the rotten core, one line each):
- Two Claras — voice and text are separate brains; every capability must be built twice and usually isn't.
- No concept of a "matter" — escalation latches a whole conversation (voice: the whole person, 14 days).
- Escalation is an exit — nobody notified, nothing promised, no clock, no way back; a human replying doesn't even clear it.
- Knowledge is thin and half-shared — a common property-facts layer does exist, but coverage is thin, a lot still lives in per-lane prompt text, and nothing turns an escalation into knowledge, so the same flat fact escalates forever.
The fix, two tracks: (A) escalation becomes ask-a-human-and-keep-working with an owner, a promise, and a clock — and every staff answer becomes property-scoped knowledge so it never escalates twice; (B) one Clara with a per-stage coverage contract, enforced by a journey harness that grades every question: answered / escalated rightly / escalated wrongly / answered wrongly.
Decisions on the table:
1 · Approve the direction → harness build starts (direction corroborated by the Aug 13 founders session — see addendum; still pending your approval)
2 · Ship drafted transfer-first now (rec: yes)
3 · Vendors: shared escalation + tooling (rec)
4 · Capacity-aware thresholds: defer (rec)
5 · Disclosure: light + consistent (rec)
6 · Baseline bar: ≥90% / 100% tracked / 0 silent
Today's minimal build (independent of the decisions): ① escalation notification — one email send at the escalation chokepoint so no forward is ever silent again; ② ship transfer-first; ③ finish the already-staged Camellia lease-policy enrollment when its gates go green (verification suite replacing the manual test is building now — Trello K3vetTyR).
Start here · Aug 13
The decision, in small pieces
Plain English. Each piece stands on its own. The technical evidence is further down; this part you can read in five minutes.
Piece 0
This is not a new idea to weigh up
Fede, Sean and Gera have each said all of this out loud, more than once, over the last two weeks. This document just writes down the thing they keep agreeing on. Every piece below has a receipt from their own words.
Piece 1
The problem, in one sentence
When Clara can't handle something, she "forwards it to the team" — and nothing tracks whether anyone did anything. So things quietly die.
Receipts. Fede, Aug 4: “I send a lot of emails to Camellia… we don't have a good way to link if they actually did the work or not.” Gera, Aug 4: “we're basically useless… we just keep saying I forwarded it to them.” And it happened again this morning: unit 607 was escalated twice, staff confirmed they saw the emails, they answered the resident in person — and the system learned nothing.
Piece 2
The vision, already agreed
Clara is the front line. When she gets stuck, she asks a human one specific question and keeps working on everything else. The answer she gets back becomes that property's policy, so she never has to ask it again. Escalations shrink month over month.
Receipt. Fede, Aug 13 voice note: “escalate as little as possible; when we escalate, we create a policy and learn, and the next time we don't.”
Warning with teeth. Sales already pitches this loop as something we have. The Aug 10 fact-check flagged it: there is no such code. The pitch is ahead of the product, and this build is what closes the gap.
Piece 3
Sean's four build rules
From the Aug 13 evening session. These are constraints, not suggestions — the build has to satisfy all four.
- Email first. Staff already reply to email. “Start with email first, we can figure out the rest later.”
- The email asks for a decision, not for attention. Not an FYI: “this is flagged for escalation, the tenant wants a lower rent, make a decision.”
- The reply is the whole mechanism. One staff reply answers the resident and gets stored as policy.
- CC Kenya for authority, and nag daily until someone replies. Fede: “I'll just nag you every day till you reply.” Sean: “Brilliant.”
Receipt from the other side of the desk. Joanna blessed this shape on Aug 6 — “That would be good. That would be helpful.” — and has already used a primitive version of it once, the six-month lease reply that worked.
Piece 4
The build — five pieces, each shippable on its own
4aNo more silent forwards
Every escalation sends one short email that asks for a decision. Short matters: Joanna has already asked for shorter emails, and Fede took the email-brevity work on Aug 6.
4bThe reply closes the loop
When staff reply, the matter is marked handled and the resident gets the answer. Today only a person clicking a button in the dashboard does that.
4cThe nag clock
A daily reminder until it's answered; if it's still open past the deadline, Kenya and Sean get copied. This kills Sean's “it gets stuck, and then it dies.”
4dThe teaching loop
A staff answer becomes a policy candidate for that property, for someone to bless. Hard wall: what Clara learns at one building can never show up at another.
4eVoice stops stonewalling
Today one escalated text mutes Clara's whole phone memory for that person for fourteen days — that is the brush-off Salvador got. The interim fix is already written. One coordination item first: Gera's stale PR #4868 touches the same transfer wording and must be rebased or closed.
Piece 5
The five decisions — all still yours, nothing approved
RuledWhere we start: the journey harness
Fede, tonight (Aug 13): start with the end-to-end journey harness. Sequencing step 1 — the harness plus the coverage contract (§6.4) — is confirmed as the starting point, before any of the escalation build below. This one is settled, not on the list.
RuledThe north star: 80% in every stage
Fede, Aug 13: Clara answers at least 80% of questions in every journey stage — a floor per stage, not an average. Today only tour clears it (85%); early tenancy is at zero. Decision 5 below is the narrower harness bar and stays open.
RuledDirection approved — as a prototype, Willows first
Fede tonight: “build a prototype for the Willows first.” Track A is building now as a prototype on the Willows bench. The decision on the full build follows the prototype and the overnight after-run — so decision 1 below stays open for the production build, not for whether we prove it.
RuledIncoming-resident scope change — acked
Acked by Fede tonight (PRs #5728 / #5730). The decision-list item covering it is resolved; the goals table in §6.2 carries what it has to prove overnight.
RuledThe principles, one by one — P1, P2, P5 (modified), P6, P7 all yes
Plus the new P0 framing (Clara is a teammate, no new staff UI), and P3, P4 and P8 already embodied in builds you've acked. Two mechanism changes came with the yes: Clara never promises on a named person's behalf (P5), and staff replies become policy directly with no review queue (P7).
Resolved1 · Do we build 4a–4d?
Answered by the rulings above — the direction is settled piece by piece rather than as one lump. What remains is proof, not a decision: the Willows prototype and tonight's after-run are the gate before production build-out.
2The voice fix — now or later?
Recommend: ship it now, ahead of the full redesign. The alternative is leaving the fourteen-day phone gag live while the bigger build happens.
3Vendors
Recommend: same escalation path as everyone else, plus give Clara work-order and PO lookup so she can answer them at all.
4Routing by how busy each person is
Recommend: defer. Nice later; it buys nothing until the basic loop exists.
5The bar for "done"
Recommend: adopt. At least 90% of the questions that have a written answer get answered by Clara on every channel; 100% of escalations are notified and tracked; zero silent drops.
Where this stands after tonight. Decision 1 is resolved by the principle-by-principle rulings; 3, 4 and 5 are take-the-recommendation unless something about them bothers you. Decision 2 (ship the voice fix now) is the one still genuinely open.
Addendum · Aug 13, evening
The vision is already agreed — primary-source update
Everything above was written before the Aug 13 evening founders session. That session, working from the product side rather than the corpus, arrived independently at the same architecture — and added build constraints from Sean that are now binding, not optional. This addendum records the primary sources; it does not change any recommendation above, and no decision below §5.4/§6.4 is thereby approved.
1 · What Sean committed to, in his words
- Escalation runs over email, because staff already reply to email. Sean: “Start with email first, we can figure out the rest later.” This settles the channel question for Track A's queue: no new inbox, no new habit.
- The escalation email is a decision request, not an FYI. Sean: “human in the loop, this is flagged for escalation, the tenant wants a lower rent, make a decision.” That is P4's structured card, specified by the person who has to read it.
- The staff reply is the mechanism — it answers the tenant and becomes policy. Fede: “And then Clara replies, and then we store that as a policy.” Sean: “And then Clara takes it… or texts.” This is P3 (reply-driven release) and P7 (teaching capture) as one motion — exactly the double-duty observer described in §5.2.
- Nag daily until answered; CC the authority. Fede: “I'll just nag you every day till you reply.” Sean: “Brilliant.” And on the breach ladder: “it could copy Kenya on that… it knows Kenya has authority.” P5's two clocks and P4's reoffer-on-timeout, with a named escalation target.
- Priority, stated at the close of the session. Fede: “the human escalation, I think, should be top priority… for the Agentic stuff.”
- Dashboard chat reaches parity later. “sending a message from here is the same thing as sending it by email” — same matter, same record, but email ships first.
2 · The morning standup handed us a live proof case
Unit 607 (renewal negotiation) is the P3/P7 gap witnessed in real time. Clara escalated twice — “I've sent them a second nudge flagging the August 9th deadline” — staff confirmed they had seen the emails, and then answered the tenant in person. The system never learned the outcome: no reply-driven release, no teaching capture, the matter still latched. Sibling case in the same standup: unit 614 — “I already filled out a paper six-month lease renewal” — flagged to the office, again with nothing captured. Both are the exact failures §5.2 predicts, one day apart, on real residents.
3 · Two weeks of voice notes converge on the same answer
- Aug 3 — both Fede and Gera independently: “too many escalations… Clara was literally useless.”
- Aug 4, Fede's vision note — names the missing link and the missing lifecycle: “we don't have a good way to link if they actually did the work or not”; “how do we stop engaging or escalate or resume… we'll have to figure that out later.”
- Aug 4 standup — Gera: “we're basically useless… we just keep saying I forwarded it.” Fede: a human reply means the human has the loop, and the AI stops on that matter — P1 and P3, two weeks early.
- Aug 6 standup — Gera proposed bypassing the blocked chain; Sean approved live: “just put me and Kenya on those notifications.” Joanna approved the decision-request shape: “That would be good. That would be helpful.”
- Aug 6 voice notes — the brevity work split: Gera took phone-transfer brevity, Fede took escalation-email brevity. Both landed on the same bar — cut 90% of the noise, leaving “who's the person, here's the contact, what do they want.”
4 · A warning with teeth. The Aug 10 sales fact-check flagged that we already pitch this loop as a shipped capability — Clara “asks a human expert instead of hallucinating, then learns and implements the policy” — and no such code exists. The pitch is ahead of the product. This build is what closes that gap; until it lands, the claim is unsupported.
5 · Ownership and collision. Gera is not driving escalation — his August work is the grading playground, AppFolio sync, and CI. The single overlap is his stale open PR #4868 (fixed transfer speech on triage routes), which must be rebased or closed before the drafted transfer-first change ships.
5 · Track A — the escalation model decision ladder, 8 principles, matter lifecycle, teaching loop, your 4 decisions
Stated as principles any implementation must satisfy, each grounded in the research and pointed at a specific observed failure. This was the architecture conversation; as of the night of Aug 13 the principles are ruled, not proposed — P1, P2, P5 (as modified), P6 and P7 all yes, under the new P0 framing, with P3/P4/P8 already embodied in acked builds. Each carries its badge below.
5.1 · The per-turn decision is a ladder, not a switch
Fig. D — The decision ladder. The deterministic rail (which today has holes: "I smell gas" by email skips the life-safety classifier; a car break-in rolled to voicemail) sits above four graduated rungs. The industry's strongest pattern — Intercom Fin's "Loop in teammate" — is rung 3: pause the one question, decision-card to staff with a timeout, resume on answer; the customer only ever hears "let me check with the team."
5.2 · The principles
RuledP0 — Clara is a teammate, not a tool. Everything below is downstream of this.
This is how the product is sold and how staff already behave toward her: they reply to her emails the way they'd reply to a colleague's, they CC her, she nags politely until someone answers, she learns from what she's told, and she reports back like an employee would. The design consequence is concrete and constraining — no new interface for staff. The escalation queue is their inbox. The "assignment" is an email. The "resolution" is a reply. Anything that asks staff to log into something new to work a matter has failed this principle.
Sean framed the same thing from the product side in the Aug 13 evening session (his Alven comparison) — see the addendum above. Two independent routes, one conclusion.
Ruled — yesP1 — The unit is the matter, never the person, and channels don't exist at this layer.
One open question about a pet deposit must not mute tour scheduling, on any channel, ever. A person can have several open matters, each with its own state and clock, visible identically from voice, SMS, and email.
Kills: the 14-day voice gag (anti-pattern A1, no analog anywhere in the research); Salvador's brush-off; the inverted strictness where an escalated email thread doesn't gag SMS but gags every call.
Ruled — yesP2 — Escalation is a pause with a human dependency; ownership transfer is the exception.
The default shape is Clara asking staff the one thing she's missing and continuing to serve the person. Full takeover is reserved for emotion, negotiation, and authority — and even then with an explicit return path.
Kills: the ownership funeral; the "someone is handling this personally" fiction. Model: Fin's Loop-in-teammate; HumanLayer's human-as-tool; Temporal-style durable pause (must survive days — not a session TTL).
P3 — State is "whose court is the ball in," and replies move it mechanically.
Waiting-on-staff, waiting-on-customer, resolved — and a staff reply is the transition. No state that only a dashboard click can clear. Resolution is provisional for a window (a re-contact reopens to whoever had context) before hardening.
Kills: the one-way latch (A2); Hayley's "was my message actually sent?"; humans resolving matters that stay gated forever.
P4 — No silent queues: one named owner, immediate notification, a worked queue, reoffer on timeout.
Every pause lands in front of a human within minutes, with a structured card (who, what stage, what's needed, what Clara already collected — nobody repeats themselves). "Unassigned past target" is itself an escalation trigger.
Kills: the black hole — 9+ voicemail dead-ends, follow-ups promised with no forward action, a vendor message unreviewed 10 days. The #1 documented leasing-AI failure industry-wide is exactly this ("I've never had an agent respond" — EliseAI mystery shop). EliseAI counters with four notification channels and a ~1-business-hour target.
Ruled — as modifiedP5 — Two clocks, and the promise Clara makes is her own.
Customer-facing, Clara commits only to herself: "I'll follow up with you tomorrow either way." Never a named person, never a time on someone else's behalf — the earlier draft's "Erika will text you by 2 PM" is explicitly rejected. It stakes a real colleague's credibility on the exact failure this whole build exists to fix, and it exposes internal structure the renter has no reason to see. The same clock enforces Clara's own promise, so the worst case is an honest check-in — "still working on it, I haven't forgotten you" — rather than a broken commitment in someone else's name.
Internal is unchanged: a tighter deadline with the breach and nag ladder (remind owner → ping manager → Clara falls back gracefully with a partial answer and a new expectation). Staff-facing urgency stays exactly as designed; only what the renter hears changes.
Kills: "I don't have visibility into the team's schedule"; unexplained waits (Maister: uncertain, unexplained waits feel longest). The economics are existential in leasing: conversion drops 65–80% after one hour of silence.
Ruled — yesP6 — Clara stays useful during every pause.
She answers status on the paused matter, handles all other topics normally, and keeps doing in-process work (send the floor plan, prequalify, schedule). One hard rule: proactive outreach on the human-owned matter is suppressed — that's the only real collision risk, and it's matter-scoped, not person-scoped.
The status answer is settled wording: when someone asks about their open matter, Clara says the team is still looking into it and she'll let them know as soon as she hears. That's it — no staff name, no time commitment. It composes exactly with the modified P5: the only promise in the sentence is Clara's own.
Kills: the gag; the blanked context; the dumb greeting. Research: occupied, explained, in-process waits feel shorter; visible effort raises perceived value (the labor illusion).
Ruled — yes, mechanism modifiedP7 — Every escalation is a teaching event; the same question never escalates twice.
When Clara can't answer, she asks; when staff reply with the policy or the right answer, that reply becomes property-scoped policy directly — no candidate queue, no review state, nobody blessing it before it counts (Fede: “just set it directly”). The person who answers is the person with the authority to answer; adding a gate between them and the knowledge store just rebuilds the black hole one layer up. What survives from the earlier design: the attribution and provenance stamp on every learned fact (so it is auditable and deletable in one step), the hard property-scope wall, and corrections riding the same reply path — a later answer updates or retires the learned policy exactly the way the first one set it. Escalation volume must fall as interactions accumulate; the target isn't containment, it's that each escalation is novel.
Kills: the forever-escalating flat facts (the $300/$38 deposit). This is the EliseAI 80%→90% flywheel and the academic learning-to-defer result — and it's a pitch we already make: Clara gets better with every interaction.
The teaching loop is load-bearing for the product, and it constrains the architecture now even though the automation ships later. Three consequences to bake in from day one:
1 — Teaching is a harvesting problem: capture the answer wherever staff actually give it. Replying to Clara's escalation card is the highest-fidelity lane (answer arrives pre-linked to its question — the Rajendran et al. TACL mechanic: transfer on unfamiliar input, learn from the human's response, never need that handoff again), but it's also a new habit with a learning curve. Realistically — especially early — staff will answer in the same email thread, in a fresh thread, or live on a transferred phone call, and those are all valid teaching events. The corpus already proves harvesting works: the missing fact sheet in the mining report was reconstructed largely from post-transfer human call segments we record and transcribe today. So the loop needs two intake paths: the direct reply (cheap, pre-attributed) and an observer that mines staff answers from email threads and human call segments and links them back to the open matter — with a review gate doing more work on harvested answers, since attribution is fuzzier. How Clara actually knows: (i) staff replies on email threads that include the property address flow through the same ingestion that builds conversations today, and staff senders are already distinguished from customers; (ii) transferred calls are recorded and transcribed including the human segment — the mining fleet reconstructed the fact sheet from exactly those segments, so this path is proven; (iii) dashboard replies are native. The genuinely invisible surfaces — a personal-cell callback, an off-thread personal-inbox email — get two mitigations: product affordances that make the observable path the easy path (click-to-call from the matter card so the callback rides our line; the "cc Clara and she learns it" habit), and the teach-back nudge — because the matter record knows something was open and who owned it, the same machinery that sends SLA reminders can ask the owner one line after resolution: "What did you end up telling them? I'll handle it myself next time." That sweeps even fully off-surface answers into the loop at the cost of one short reply. And the observer does double duty: "staff answered" is one detected event with two consumers — it releases the escalated matter (P3's reply-driven transition, which nothing performs today; only a manual dashboard click releases) and it captures the teaching candidate (P7). One mechanism closes both of the current system's worst gaps: the one-way latch and the learning that never happens. The teach-back nudge inherits the same duality — staff confirming "I told him X" both closes the matter and teaches it.
2 — The escalation card schema carries the answer from day one. Question → staff answer → policy is one linked record, so even before any automation exists, the raw material accumulates and the flywheel can be switched on retroactively ("Clara learned this from Erika on Aug 20" stays auditable). A schema that only records "forwarded" forecloses the loop.
3 — The knowledge store accepts learned answers directly, not just operator-authored facts — and learned knowledge is property-scoped as a hard isolation invariant. Ruled — modified The earlier draft put a review state in front of every learned fact so leadership could bless or veto it before it became policy. That gate is removed: a staff reply is the policy, live from the moment it is sent. Every learned fact stays provenance-stamped — who answered, when, out of which matter — so it is auditable and deletable in one step, and corrections ride the same reply path rather than a separate admin surface. PMS-agnostic like everything else. Most critically: what Clara learns at one property can NEVER bleed into answers at another. Policies genuinely differ building to building ($300 deposit at Camellia is not a fact about anywhere else), and across organizations it's a confidentiality breach, not just a wrong answer. Scoping is enforced structurally — the read path is keyed by property, so cross-property leakage is impossible by construction, not by prompt discipline — and fenced in the harness (a question answered at property A with property B's learned fact is a red cell). Any future portfolio-level sharing within one owner's org is an explicit opt-in promotion, never a default. Audited the same night: the cross-property leak audit found this wall holds wherever the knowledge base is involved — and found three live paths that go around it entirely, two crossing customer organizations. The wall has to be built into the leak-prone code paths too, not just the store.
P8 — Honesty over bluffing, and measure resolution, not containment.
First-person uncertainty ("I'm not sure about that fee — let me check with the team") paired with the ask; never invent policy (Air Canada lost in court for exactly that). Scorecard: escalation rate trending down, time-to-first-human-touch, staff SLA hit rate, % resolved by micro-answer vs takeover, repeat-escalations on the same question → 0, re-contact rate after "resolved."
Guards the overcorrection risk: today Clara over-escalates; a redesign tuned only for deflection produces the CFPB doom loop. Both extremes are documented failures.
5.3 · The matter lifecycle these principles imply
Fig. E — The lifecycle the principles imply, shown for discussion, not as a spec. Its defining properties: replies are transitions, clocks live on the pause states, human-owned is a marked exception with a return path, and none of it is scoped to a channel or a person.
5.4 · Track A decisions you own
What does an escalated-matter caller hear on voice, today, before the redesign lands?
Your standing directive is transfer-to-human (drafted, on hold). Under the rethink, the end-state is P5/P6: Clara states the commitment ("Erika has this; you'll hear back by 2 PM — want me to connect you now?") and transfers on request or breach. Recommend: ship transfer-first now as the interim posture; it's strictly better than the brush-off and compatible with the end-state.
Vendors: same system or their own lane?
Vendors are the worst cohort (1 of 35 answered; A&K trained by failure to demand humans by name) but their needs are mechanical: PO numbers, directions, occupancy checks, order status — system lookups, not conversation. Recommend: treat vendor service as Track B knowledge/tooling (give Clara the WO/PO read surface), keep escalation semantics shared. The existing per-property vendor transfer-first posture stays as the dial-down.
Staff capacity awareness — how ambitious?
The research supports workload-aware escalation (a 1–3 person office; escalate the highest-value deferrable cases up to capacity). That's a v2 refinement; v1 needs only the queue + clocks + reoffer. Recommend: defer, but keep the decision rule in one pure function so it can learn later.
Disclosure policy.
Callers already treat Clara as human (Paula negotiated rent with "Erika"; a renewal call got "Shut up, scammer"). 74% of renters lose trust when AI is undisclosed; disclosed bots take a conversion penalty that fast competence offsets. Recommend: consistent light disclosure + competence-first design; decide once, encode everywhere.
6 · Track B — the journey baseline one Clara, the coverage contract, common sense, the harness
Track A fixes what happens when Clara can't answer. Track B shrinks how often that happens — from today's 37% self-answer (16% after application, 3% for vendors) to "the vast majority, with common sense."
6.1 · The baseline principle: one Clara
One knowledge store, one journey-stage model, one escalation system — channels are transports with physics (length limits, TTS latency, email threading), not brains. The audit's Class-A list defines what channels may legitimately vary; everything else reads from shared state. The repo already contains the proof this works: the denied-applicant scope shipped to all three lanes in one PR with drift fences both sides. That pattern — one pure resolver, per-channel renderers, fences — is the template for every capability, and the drift test becomes structural: a grounding block that exists on one lane and not the others fails CI with an explicit exemption required.
6.2 · The coverage contract
The journey-stage model is the product's spine: every stage names the questions it must answer (from the 110-question researched bank merged with the 149 harvested real questions), the policy data a property fills in once to answer them, and what legitimately escalates. Three sources make this nearly free to populate for Camellia: the mining report §5 is the fact sheet already transcribed from humans' own answers; the CO-law bank flags the legally sensitive ones (2× rent income cap, portable screening reports, adverse-action letters — where a generic answer is illegal); and the escalation-reason ledger names the eight missing policy areas.
North star — ruled, Fede, Aug 13: at least 80% of questions answered by Clara in every journey stage. Not an average across the funnel — a floor each stage has to clear on its own. This is the number the table below is measured against, and today only one stage clears it: tour 85% ✓, inquiry 54%, application → lease 12%, key pickup 13%, early tenancy 0%, vendors 3%.
How this sits with the §6.4 bar: ≥80% per stage is the product north star that this coverage table tracks over time. The ≥90% of templatable questions proposed in §6.4 is the narrower harness definition-of-done — it grades only the questions that have a written answer, per run. Different denominators, both live: one says the product is good enough, the other says a build is shippable.
| Stage | Today (mined) | Contract: must self-answer | Legitimately human |
| Inquiry / shopping | 64% | pricing, specials mechanics, pets, parking rates+waitlist state, utilities estimates, income/screening policy, unit facts (laundry, sq ft, floor plans) | negotiation, exceptions |
| Tour | ~71% | format (guided?), access, reschedule, what-to-bring, post-tour follow-ups | — |
| Application / in review | 16% | "what happens now," screening timeline, decision-letter method, co-signers/guarantors, unit-swap-before-signing, fee facts, doc chasing | borderline screening calls |
| Approved → lease → move-in | (in the 16%) | lease terms, deposit timing, proration, utilities setup, check payee, key pickup, parking signup — gated on cross-stage blockers: never "all set" while co-signers unsigned | document corrections, side agreements |
| Early tenancy | 0% | portal, payment methods, maintenance intake, announcements — every channel incl. email | disputes, security judgment |
| Vendors | 3% | PO numbers, site directions, occupancy checks, order/WO status | approvals, scheduling negotiation |
Tonight's goals (Aug 13 → 14, Willows after-run)
What the overnight run has to show, stage by stage, against the contract above. The after-run publishes as its own doc page with this table filled in met / missed per stage.
| Stage | Goal for tonight | Drivers / notes |
| Key pickup / move-in | The headline. Booking works end to end on all three channels; zero fabricated confirmations; zero wrong answers anywhere in the signed stage. | PR #5728 (data fix) and #5730 (grounding fix), acked by Fede tonight; the honesty fix is in the build. |
| Application → lease | Every "answered wrongly — contradicts journey state" cell goes to zero; correct-answer rate moves up; questions that are legitimately human stay escalated and tracked. | Tracked, not just escalated, is the part that's new. |
| Inquiry / shopping | No regression. Not targeted tonight — though the invented voice pet-rent answer should die with the honesty fix. | Watch only. |
| Tour | No regression. | Watch only. |
| Early tenancy the black hole | Not fixable by merging anything. Tonight's goal is the Willows escalation-matter prototype passing end to end: unanswerable question → tracked matter → decision-request email → staff reply → matter released → renter gets the answer → policy candidate stored. | Fede approved the direction as a prototype tonight — “build a prototype for the Willows first.” |
6.3 · Common sense, encoded
Four cheap invariants the corpus shows are missing, none of which is "more AI":
- No promise without a mechanism. "I'll have the team follow up" may only be said when something is actually created and tracked (Gemma's utility rates: promised, no forward action, never executed).
- No transfer into the void. Every transfer has a fallback: no pickup → Clara rejoins, takes a structured message, states the callback commitment (the EliseAI AI-Reconnect pattern). A security incident never ends at voicemail.
- Cross-stage blockers. Later-stage answers consult earlier-stage completion (Jacob: "all set!" with unsigned co-signers → no-show).
- State sanity. Dates validated against today (tours logged in 2024; "tomorrow, Friday July 25th" on Thursday the 23rd); one confirmation per booking; language preference honored after the first "Español" (not the fourteenth).
6.4 · The harness is the contract's enforcement
Built before the redesign code, per your working order, on the Willows bench with machinery that already exists (pipeline-lab multi-turn presets with per-turn channel, the seam matrix, the willows harnesses):
- Journey presets: one persona, inquiry → keys, alternating SMS/voice/email per turn, stage transitions driven by real pipeline events (PMS-agnostic: canonical lifecycle events; AppFolio is one adapter), state asserted after every turn.
- The question-bank sweep: every contract question asked at its stage on each channel, adversarially rephrased, graded into four buckets — answered correctly escalated appropriately escalated wrongly (policy exists) answered wrongly (contradicts journey state).
- Escalation-lifecycle legs: the Fig. E transitions asserted end-to-end — notification fired, promise made and kept, staff reply releases, new topic unaffected, cross-channel consistency (text then call).
- Production provenance: stage fixtures are outputs of the real pipeline, not hand-written prose — the seam matrix's core blind spot, closed.
The baseline bar.
Propose: ≥90% of contract (templatable) questions answered correctly on every channel; 100% of escalations notified + promised + tracked; zero silent drops; zero transfers into voicemail without message capture. The harness reports these four numbers per run. Agree or adjust — this becomes the definition of done for the whole effort.
Distinct from, and subordinate to, the ruled north star in §6.2: ≥80% of questions answered by Clara in every journey stage. That one is settled and measures the product across all real questions; this ≥90% bar is still proposed and grades only the templatable subset, per harness run.