PropFlow · Diagnosis + Rethink · Decisions pending (§5.4, §6.4)

Escalation & the Applicant Journey

For Fede · August 13, 2026 · One document: how the system behaves today (§1–3), what the rot actually is (§4), and the first-principles two-track rethink (§5–7). Evidence: the full Camellia corpus mined since launch (203 findings, all channels), the per-channel fork audit of origin/main, the test-infrastructure map, deep research on human-in-the-loop escalation, and the Colorado renter-journey question bank. Nothing here is implemented; the only drafted code (voice transfer-first) is parked pending §5.4 Decision 1.

The one-screen version

What broke: a hot applicant asked "I just applied — what do I do now?" at 11 PM and got silently forwarded to nobody; his call the next morning was brushed off with "someone is handling this personally." Mining every conversation since launch shows this is the norm: Clara answers 37% of real questions (16% after application, 3% for vendors), 44% of everything humans handled was answerable from a fact sheet, and at least 9 transfers dead-ended in voicemail — including a car break-in report.

Why (the rotten core, one line each):

  1. Two Claras — voice and text are separate brains; every capability must be built twice and usually isn't.
  2. No concept of a "matter" — escalation latches a whole conversation (voice: the whole person, 14 days).
  3. Escalation is an exit — nobody notified, nothing promised, no clock, no way back; a human replying doesn't even clear it.
  4. Knowledge is thin and half-shared — a common property-facts layer does exist, but coverage is thin, a lot still lives in per-lane prompt text, and nothing turns an escalation into knowledge, so the same flat fact escalates forever.

The fix, two tracks: (A) escalation becomes ask-a-human-and-keep-working with an owner, a promise, and a clock — and every staff answer becomes property-scoped knowledge so it never escalates twice; (B) one Clara with a per-stage coverage contract, enforced by a journey harness that grades every question: answered / escalated rightly / escalated wrongly / answered wrongly.

Decisions on the table:

1 · Approve the direction → harness build starts (direction corroborated by the Aug 13 founders session — see addendum; still pending your approval)
2 · Ship drafted transfer-first now (rec: yes)
3 · Vendors: shared escalation + tooling (rec)
4 · Capacity-aware thresholds: defer (rec)
5 · Disclosure: light + consistent (rec)
6 · Baseline bar: ≥90% / 100% tracked / 0 silent

Today's minimal build (independent of the decisions): ① escalation notification — one email send at the escalation chokepoint so no forward is ever silent again; ② ship transfer-first; ③ finish the already-staged Camellia lease-policy enrollment when its gates go green (verification suite replacing the manual test is building now — Trello K3vetTyR).

Start here · Aug 13

The decision, in small pieces

Plain English. Each piece stands on its own. The technical evidence is further down; this part you can read in five minutes.

Piece 0

This is not a new idea to weigh up

Fede, Sean and Gera have each said all of this out loud, more than once, over the last two weeks. This document just writes down the thing they keep agreeing on. Every piece below has a receipt from their own words.

Piece 1

The problem, in one sentence

When Clara can't handle something, she "forwards it to the team" — and nothing tracks whether anyone did anything. So things quietly die.

Receipts. Fede, Aug 4: “I send a lot of emails to Camellia… we don't have a good way to link if they actually did the work or not.” Gera, Aug 4: “we're basically useless… we just keep saying I forwarded it to them.” And it happened again this morning: unit 607 was escalated twice, staff confirmed they saw the emails, they answered the resident in person — and the system learned nothing.

Piece 2

The vision, already agreed

Clara is the front line. When she gets stuck, she asks a human one specific question and keeps working on everything else. The answer she gets back becomes that property's policy, so she never has to ask it again. Escalations shrink month over month.

Receipt. Fede, Aug 13 voice note: “escalate as little as possible; when we escalate, we create a policy and learn, and the next time we don't.”

Warning with teeth. Sales already pitches this loop as something we have. The Aug 10 fact-check flagged it: there is no such code. The pitch is ahead of the product, and this build is what closes the gap.

Piece 3

Sean's four build rules

From the Aug 13 evening session. These are constraints, not suggestions — the build has to satisfy all four.

  • Email first. Staff already reply to email. “Start with email first, we can figure out the rest later.”
  • The email asks for a decision, not for attention. Not an FYI: “this is flagged for escalation, the tenant wants a lower rent, make a decision.”
  • The reply is the whole mechanism. One staff reply answers the resident and gets stored as policy.
  • CC Kenya for authority, and nag daily until someone replies. Fede: “I'll just nag you every day till you reply.” Sean: “Brilliant.”

Receipt from the other side of the desk. Joanna blessed this shape on Aug 6 — “That would be good. That would be helpful.” — and has already used a primitive version of it once, the six-month lease reply that worked.

Piece 4

The build — five pieces, each shippable on its own

4aNo more silent forwards

Every escalation sends one short email that asks for a decision. Short matters: Joanna has already asked for shorter emails, and Fede took the email-brevity work on Aug 6.

4bThe reply closes the loop

When staff reply, the matter is marked handled and the resident gets the answer. Today only a person clicking a button in the dashboard does that.

4cThe nag clock

A daily reminder until it's answered; if it's still open past the deadline, Kenya and Sean get copied. This kills Sean's “it gets stuck, and then it dies.”

4dThe teaching loop

A staff answer becomes a policy candidate for that property, for someone to bless. Hard wall: what Clara learns at one building can never show up at another.

4eVoice stops stonewalling

Today one escalated text mutes Clara's whole phone memory for that person for fourteen days — that is the brush-off Salvador got. The interim fix is already written. One coordination item first: Gera's stale PR #4868 touches the same transfer wording and must be rebased or closed.

Piece 5

The five decisions — all still yours, nothing approved

RuledWhere we start: the journey harness

Fede, tonight (Aug 13): start with the end-to-end journey harness. Sequencing step 1 — the harness plus the coverage contract (§6.4) — is confirmed as the starting point, before any of the escalation build below. This one is settled, not on the list.

RuledThe north star: 80% in every stage

Fede, Aug 13: Clara answers at least 80% of questions in every journey stage — a floor per stage, not an average. Today only tour clears it (85%); early tenancy is at zero. Decision 5 below is the narrower harness bar and stays open.

RuledDirection approved — as a prototype, Willows first

Fede tonight: “build a prototype for the Willows first.” Track A is building now as a prototype on the Willows bench. The decision on the full build follows the prototype and the overnight after-run — so decision 1 below stays open for the production build, not for whether we prove it.

RuledIncoming-resident scope change — acked

Acked by Fede tonight (PRs #5728 / #5730). The decision-list item covering it is resolved; the goals table in §6.2 carries what it has to prove overnight.

RuledThe principles, one by one — P1, P2, P5 (modified), P6, P7 all yes

Plus the new P0 framing (Clara is a teammate, no new staff UI), and P3, P4 and P8 already embodied in builds you've acked. Two mechanism changes came with the yes: Clara never promises on a named person's behalf (P5), and staff replies become policy directly with no review queue (P7).

Resolved1 · Do we build 4a–4d?

Answered by the rulings above — the direction is settled piece by piece rather than as one lump. What remains is proof, not a decision: the Willows prototype and tonight's after-run are the gate before production build-out.

2The voice fix — now or later?

Recommend: ship it now, ahead of the full redesign. The alternative is leaving the fourteen-day phone gag live while the bigger build happens.

3Vendors

Recommend: same escalation path as everyone else, plus give Clara work-order and PO lookup so she can answer them at all.

4Routing by how busy each person is

Recommend: defer. Nice later; it buys nothing until the basic loop exists.

5The bar for "done"

Recommend: adopt. At least 90% of the questions that have a written answer get answered by Clara on every channel; 100% of escalations are notified and tracked; zero silent drops.

Where this stands after tonight. Decision 1 is resolved by the principle-by-principle rulings; 3, 4 and 5 are take-the-recommendation unless something about them bothers you. Decision 2 (ship the voice fix now) is the one still genuinely open.

Addendum · Aug 13, evening

The vision is already agreed — primary-source update

Everything above was written before the Aug 13 evening founders session. That session, working from the product side rather than the corpus, arrived independently at the same architecture — and added build constraints from Sean that are now binding, not optional. This addendum records the primary sources; it does not change any recommendation above, and no decision below §5.4/§6.4 is thereby approved.

1 · What Sean committed to, in his words

2 · The morning standup handed us a live proof case

Unit 607 (renewal negotiation) is the P3/P7 gap witnessed in real time. Clara escalated twice — “I've sent them a second nudge flagging the August 9th deadline” — staff confirmed they had seen the emails, and then answered the tenant in person. The system never learned the outcome: no reply-driven release, no teaching capture, the matter still latched. Sibling case in the same standup: unit 614“I already filled out a paper six-month lease renewal” — flagged to the office, again with nothing captured. Both are the exact failures §5.2 predicts, one day apart, on real residents.

3 · Two weeks of voice notes converge on the same answer

4 · A warning with teeth. The Aug 10 sales fact-check flagged that we already pitch this loop as a shipped capability — Clara “asks a human expert instead of hallucinating, then learns and implements the policy” — and no such code exists. The pitch is ahead of the product. This build is what closes that gap; until it lands, the claim is unsupported.

5 · Ownership and collision. Gera is not driving escalation — his August work is the grading playground, AppFolio sync, and CI. The single overlap is his stale open PR #4868 (fixed transfer speech on triage routes), which must be rebased or closed before the drafted transfer-first change ships.

Everything below is the evidence and the detail — collapsed; open what you need.

1 · Direct answers first timer? release? topic scope? notification? — the four factual answers

Is there a timer on escalation?

No SLA timer exists anywhere. The only time-related logic: the voice gag ignores escalated threads idle >14 days (ESCALATED_POSTURE_LOOKBACK_DAYS), and text acknowledgments are rate-limited to one per 24h. Nothing ever fires because time passed — no reminder to the team, no breach alarm, no auto-release. An escalated thread is escalated until a human clicks it back to active in the dashboard.

Does a human replying release the thread?

No. The messages route never touches conversation.status. A PM can fully resolve the matter with the tenant and the thread stays gated. The manual dashboard PATCH is the only door out (escalated-gate.ts:52 — every other status reopens on inbound; escalated deliberately does not).

If one topic is escalated, can Clara still answer another?

On SMS/email — yes, since Aug 9 (#5630): an LLM topic classifier compares the new inbound to the escalated matter; a separate topic falls through to the normal agent loop, and it fails toward the ack when the topic can't be derived. On voice — no. The posture is per-person: any live escalated thread in 14 days swaps the greeting and blanks all context, with no topic check at all. Same person, same day, opposite behavior by channel.

Does the team get notified when an escalated person contacts us again?

No — on any channel. The code admits both halves: text side, "the acknowledgment tells the tenant a human has their message, it does not tell the human. Known latency gap, deferred"; voice side, the call mints a fresh conversation and "never touches or links the escalated thread the human is watching." Sal called because nobody followed up; the system asserted someone was handling it.

2 · How it works today the state machine, the channel asymmetry, the Salvador timeline, the grounding matrix

2.1 · The lifecycle: five ways in, one way out

active escalated active human-owned no owner field · no SLA field (released) forward_to_property_manager escalate_to_human runaway-reply ceiling inbound-dispatcher fallback manual PATCH by PM any unanswerable question dashboard PATCH only ✗ human reply does NOT release ✗ no timer ever releases or reminds ✗ end_conversation forbidden to resolve
Fig. 1 — The escalated status is written by five independent paths (three of them automatic) and released by exactly one: a person clicking the thread back to active in the dashboard. The state carries no owner, no topic field, and no deadline. Sources: escalated-topic.ts:58, escalated-gate.ts:52, conversations/[id]/route.ts:62.

2.2 · Cross-channel: two different machines wearing one status

SMS / EMAIL (per-matter, since Aug 9) VOICE (per-person — no topic check) inbound from escalated person topic classifier vs the escalated matter same topic → ack only, 24h rate limit separate topic → normal agent loop resumes call rings personalization webhook any escalated thread for this PERSON, <14 days? gag greeting + blank ALL context whatever they're calling about topic check: absent this hop exists one lane up
Fig. 2 — The text lanes ask "is this contact about the escalated matter?" and resume Clara for anything else. The voice lane asks only "does this person have any live escalated thread?" — the dashed box is the hop that was never built (escalated-call-posture.ts: hasLiveEscalatedThread checks status + recency only). This asymmetry is why Sal's 11 PM text gagged his 9:31 AM call.

2.3 · The Salvador trace, end to end

Aug 11 tour + call handled fine Aug 12 · 11:13 PM applies online 11:16 PM · SMS "what do I do now?" → forwarded, thread escalated no human notified, ever Aug 13 · 9:31 AM · VOICE "team is handling this personally" context blanked · no ETA 9:32 AM transfer → Joanna answers everything 12:38 PM pmsApplicationStatus: Approved 10h15m — zero follow-up
Fig. 3 — Every subsystem behaved "as designed." The PMS sync stamped his application (5-minute poller) and eventually his approval. The failure is purely in the conversation layer: no stage answer on SMS, an escalation that told no one, and a voice gate that mistook him for a matter already in hand.

2.4 · Stage grounding is welded to channels

The knowledge to answer applicants exists — but each block is wired to a single lane. The in-review block opens with if (channel !== 'email') return null (documented: phone-side identity counting not built yet), and even on email its final rule routes exactly Sal's question — "what happens next for THEM specifically" — to a human. The structured stage / pmsApplicationStatus fields the PMS sync stamps every 5 minutes are read only by the email lane; voice and SMS prompts receive an LLM-written free-text summary and nothing else.

Grounding blockVoiceSMSEmailWhere
Pricing / specials / tours (prospect)yesyesyesclara-leasing.ts
Applicant — submitted / in reviewnonostatus onlyapplication-status-scope.ts
Approved — status disclosurenonoyesapplication-status-scope.ts
Approved — lease terms & figuresyesyesyeslease-answer-context.ts (#5475)
Denied — published answer, no pitchyespartialyesdenied-applicant-scope.ts (#5631)
Incoming resident (pending occupancy)yesnonoincoming-resident-context.ts (#5714)
Screening timeline / "decision letter by email"nonononowhere — banned or unbuilt

The production consequence is measurable. Across 149 substantive questions in 90 days of real Camellia conversations, Clara's self-answer rate inverts with funnel depth — highest where prospects are cheapest, lowest where they are about to become revenue:

Journey stageQuestionsAnswered by ClaraWhat happened to the rest
Tour scheduling / tour day~4085%occasional transfer
Inquiry / shopping~4554%transfers, some drops
Application → lease signing~3312%39 answered only by humans post-transfer
Key pickup / move-in~2313%26 escalated into the black hole
Early tenancy~80%incl. a car break-in report reaching voicemail

The post-transfer human segments are effectively Clara's missing knowledge base, already transcribed: parking $80 surface / $95 garage + waitlist mechanics, $300 deposit / $38 fee, 6- or 12-month at the same rate + month-to-month after, decision letters by email only, unit swap without reapplying, prorated first payment, Xcel-only utilities, snow-removal contract, money-order policy for late rent. None of it needs human judgment.

3 · Why the tests said this was fine seeded fixtures, no journeys, no escalated row

The seam matrix (merged the night before the incident) is honest about being a grid of cells: 10 caller stages × 9 intents, voice-only, each cell a single utterance against hand-seeded state. Three structural blind spots made Sal invisible to it:

  1. Seeded state has no production provenance. 9 of 10 stage fixtures are hand-written prose ("Application is in screening; no decision yet") asserted against nothing — nobody proves the real pipeline (PMS event → summarizer → personalization route) ever emits that context. The applicant_submitted × application_status cell is green because its fixture supplies the answer the real prompt has no grounding for.
  2. No journey, no cross-channel, no time. Nothing carries one persona through inquiry → apply → decision → keys; nothing texts then calls the same person; "escalated" isn't even a row in the matrix, so the cell where Sal actually failed cannot appear on the grid. Pipeline-lab has multi-turn + real-pipeline state + a per-turn channel field — and zero applicant or escalated presets using them.
  3. Nothing asserts the two escalation scopes agree. The topic classifier is evaluated in isolation; the voice posture is evaluated with no topic input at all. "Same person, new matter, voice lane" is tested nowhere.
4 · What the rot actually is the four assumptions, with receipts

The bugs are not independent. Every failure in the corpus — Salvador's brush-off, Jacob's no-show, Hayley's five unanswered escalations, the car break-in that rolled to voicemail, A&K's seventeen calls — traces back to four assumptions baked into the core, each individually defensible when written, each wrong:

  1. "Channel" is treated as an architectural boundary. There are two Claras: the text lanes share one prompt assembler; voice is a separate ElevenLabs brain fed by hand-mirrored variables and never touches the conversation manager. Every capability must be built twice and usually isn't — the audit found 10 landmines where knowledge or behavior exists on one lane and not another, and most were never decided, just forgotten. A person is one person; they experience the gaps as Clara lying: utilities answered on a call Aug 5, "not on file" on Jul 17; parking correct by email, transferred by voice. Evidence: channel-fork audit §0, landmines L1–L10; mining surprise #3.
  2. "Conversation" is treated as the unit of state. Escalation is a status on a conversation row; the voice gate then over-corrects by scoping to the whole person. There is no first-class notion of a matter — the thing the person actually needs — so nothing carries an owner, a deadline, or a lifecycle. Twenty years of helpdesk orthodoxy is unambiguous (one issue = one ticket, several open per customer, each with its own state and clock), and no vendor, framework, or helpdesk anywhere scopes escalation to a person. Our voice gag has no analog in the entire research corpus. Evidence: HITL research P1/P3, anti-pattern A1 ("worst offender"); diagnosis Fig 1–2.
  3. "Escalation" is treated as an exit, not a dependency. Forwarding sets a flag and the system's responsibility ends: nobody is notified, nothing is promised, no clock runs, a human replying doesn't even clear it, and the agent goes dumb. The industry's convergent model is the opposite — a pause with a human dependency: the agent asks staff the one thing it's missing, keeps working, and resumes when answered; full human takeover is the last rung, not the first. Our system manages to sit on the losing side of both extremes simultaneously: it escalates far too easily and delivers none of the value escalation is supposed to buy. Evidence: mining §6 (9+ transfers to voicemail incl. a car break-in; escalations with no forward action; "I just wanted to make sure my message was sent"); research P2–P6, anti-patterns A2–A6 — we exhibit all five at once.
  4. Knowledge is shared in principle, per-lane in practice — and nothing ever adds to it. A common property-facts layer does exist: seven pricing and property tools reachable from every lane, plus the property knowledge records behind them. The correction matters, but it doesn't rescue the conclusion. Coverage through that layer is thin; substantial answer content still lives in per-lane prompt text (the ten landmines in the fork audit); voice gets hand-mirrored variables rather than the same reads; there is no per-stage policy schema at all; and — the part that is unqualifiedly true — no loop turns an escalation into knowledge, so the same flat fact escalates forever. The receipts: 44% of everything that consumed a human since launch was templatable; of questions humans answered live after a transfer, 62% were plain facts; a deposit escalated as "applicant-specific, requires human review" was answered two hours later as "$300 deposit, $38 fee." The missing fact sheet is not hypothetical — the mining report §5 literally transcribed it from the humans' own mouths. Evidence: mining §2 (15 escalations across 8 policy areas of pure missing documentation), §4 (8 of 12 escalation reasons = missing policy), §5 (the fact sheet, already written); research P10 (the flywheel: EliseAI operators climb ~80%→90%+ automation by closing gaps).

One rot the redesign cannot fix from inside the conversation layer: the lease pipeline itself generates escalations. Three people received defective lease documents (Hayley's wrong dates and term, Ekene's missing concession, Jacob's outdated no-pets template), and Clara then escalated the fallout of documents the same system produced — with Jacob's missing co-signer signatures flagged the day after his move-in date. However good escalation becomes, document generation and pre-move-in completeness checks are their own workstream.

5 · Track A — the escalation model decision ladder, 8 principles, matter lifecycle, teaching loop, your 4 decisions

Stated as principles any implementation must satisfy, each grounded in the research and pointed at a specific observed failure. This was the architecture conversation; as of the night of Aug 13 the principles are ruled, not proposed — P1, P2, P5 (as modified), P6 and P7 all yes, under the new P0 framing, with P3/P4/P8 already embodied in acked builds. Each carries its badge below.

5.1 · The per-turn decision is a ladder, not a switch

DETERMINISTIC RAIL — outranks everything, identical on every channel life safety · security incidents · fair housing / legal · explicit "I want a person" — coded rules, never model judgment 1 · Answer 2 · Offer options 3 · Ask a human, keep working 4 · Hand over from the one knowledge store first miss → clarify; users prefer this to a human micro-escalation: the default emotion · negotiation · authority — last rung can't 2nd miss judgment Today's system jumps from rung 1 directly to rung 4 — and rung 4 is broken. Rung 3 is where 44% of the corpus's human handoffs belong: "let me check with the team" → staff answer one card → Clara delivers. CHI 2019: users prefer options on first breakdown and want a human only after repair fails.
Fig. D — The decision ladder. The deterministic rail (which today has holes: "I smell gas" by email skips the life-safety classifier; a car break-in rolled to voicemail) sits above four graduated rungs. The industry's strongest pattern — Intercom Fin's "Loop in teammate" — is rung 3: pause the one question, decision-card to staff with a timeout, resume on answer; the customer only ever hears "let me check with the team."

5.2 · The principles

RuledP0 — Clara is a teammate, not a tool. Everything below is downstream of this.

This is how the product is sold and how staff already behave toward her: they reply to her emails the way they'd reply to a colleague's, they CC her, she nags politely until someone answers, she learns from what she's told, and she reports back like an employee would. The design consequence is concrete and constraining — no new interface for staff. The escalation queue is their inbox. The "assignment" is an email. The "resolution" is a reply. Anything that asks staff to log into something new to work a matter has failed this principle.

Sean framed the same thing from the product side in the Aug 13 evening session (his Alven comparison) — see the addendum above. Two independent routes, one conclusion.

Ruled — yesP1 — The unit is the matter, never the person, and channels don't exist at this layer.

One open question about a pet deposit must not mute tour scheduling, on any channel, ever. A person can have several open matters, each with its own state and clock, visible identically from voice, SMS, and email.

Kills: the 14-day voice gag (anti-pattern A1, no analog anywhere in the research); Salvador's brush-off; the inverted strictness where an escalated email thread doesn't gag SMS but gags every call.

Ruled — yesP2 — Escalation is a pause with a human dependency; ownership transfer is the exception.

The default shape is Clara asking staff the one thing she's missing and continuing to serve the person. Full takeover is reserved for emotion, negotiation, and authority — and even then with an explicit return path.

Kills: the ownership funeral; the "someone is handling this personally" fiction. Model: Fin's Loop-in-teammate; HumanLayer's human-as-tool; Temporal-style durable pause (must survive days — not a session TTL).

P3 — State is "whose court is the ball in," and replies move it mechanically.

Waiting-on-staff, waiting-on-customer, resolved — and a staff reply is the transition. No state that only a dashboard click can clear. Resolution is provisional for a window (a re-contact reopens to whoever had context) before hardening.

Kills: the one-way latch (A2); Hayley's "was my message actually sent?"; humans resolving matters that stay gated forever.

P4 — No silent queues: one named owner, immediate notification, a worked queue, reoffer on timeout.

Every pause lands in front of a human within minutes, with a structured card (who, what stage, what's needed, what Clara already collected — nobody repeats themselves). "Unassigned past target" is itself an escalation trigger.

Kills: the black hole — 9+ voicemail dead-ends, follow-ups promised with no forward action, a vendor message unreviewed 10 days. The #1 documented leasing-AI failure industry-wide is exactly this ("I've never had an agent respond" — EliseAI mystery shop). EliseAI counters with four notification channels and a ~1-business-hour target.

Ruled — as modifiedP5 — Two clocks, and the promise Clara makes is her own.

Customer-facing, Clara commits only to herself: "I'll follow up with you tomorrow either way." Never a named person, never a time on someone else's behalf — the earlier draft's "Erika will text you by 2 PM" is explicitly rejected. It stakes a real colleague's credibility on the exact failure this whole build exists to fix, and it exposes internal structure the renter has no reason to see. The same clock enforces Clara's own promise, so the worst case is an honest check-in — "still working on it, I haven't forgotten you" — rather than a broken commitment in someone else's name.

Internal is unchanged: a tighter deadline with the breach and nag ladder (remind owner → ping manager → Clara falls back gracefully with a partial answer and a new expectation). Staff-facing urgency stays exactly as designed; only what the renter hears changes.

Kills: "I don't have visibility into the team's schedule"; unexplained waits (Maister: uncertain, unexplained waits feel longest). The economics are existential in leasing: conversion drops 65–80% after one hour of silence.

Ruled — yesP6 — Clara stays useful during every pause.

She answers status on the paused matter, handles all other topics normally, and keeps doing in-process work (send the floor plan, prequalify, schedule). One hard rule: proactive outreach on the human-owned matter is suppressed — that's the only real collision risk, and it's matter-scoped, not person-scoped.

The status answer is settled wording: when someone asks about their open matter, Clara says the team is still looking into it and she'll let them know as soon as she hears. That's it — no staff name, no time commitment. It composes exactly with the modified P5: the only promise in the sentence is Clara's own.

Kills: the gag; the blanked context; the dumb greeting. Research: occupied, explained, in-process waits feel shorter; visible effort raises perceived value (the labor illusion).

Ruled — yes, mechanism modifiedP7 — Every escalation is a teaching event; the same question never escalates twice.

When Clara can't answer, she asks; when staff reply with the policy or the right answer, that reply becomes property-scoped policy directly — no candidate queue, no review state, nobody blessing it before it counts (Fede: “just set it directly”). The person who answers is the person with the authority to answer; adding a gate between them and the knowledge store just rebuilds the black hole one layer up. What survives from the earlier design: the attribution and provenance stamp on every learned fact (so it is auditable and deletable in one step), the hard property-scope wall, and corrections riding the same reply path — a later answer updates or retires the learned policy exactly the way the first one set it. Escalation volume must fall as interactions accumulate; the target isn't containment, it's that each escalation is novel.

Kills: the forever-escalating flat facts (the $300/$38 deposit). This is the EliseAI 80%→90% flywheel and the academic learning-to-defer result — and it's a pitch we already make: Clara gets better with every interaction.

The teaching loop is load-bearing for the product, and it constrains the architecture now even though the automation ships later. Three consequences to bake in from day one:

1 — Teaching is a harvesting problem: capture the answer wherever staff actually give it. Replying to Clara's escalation card is the highest-fidelity lane (answer arrives pre-linked to its question — the Rajendran et al. TACL mechanic: transfer on unfamiliar input, learn from the human's response, never need that handoff again), but it's also a new habit with a learning curve. Realistically — especially early — staff will answer in the same email thread, in a fresh thread, or live on a transferred phone call, and those are all valid teaching events. The corpus already proves harvesting works: the missing fact sheet in the mining report was reconstructed largely from post-transfer human call segments we record and transcribe today. So the loop needs two intake paths: the direct reply (cheap, pre-attributed) and an observer that mines staff answers from email threads and human call segments and links them back to the open matter — with a review gate doing more work on harvested answers, since attribution is fuzzier. How Clara actually knows: (i) staff replies on email threads that include the property address flow through the same ingestion that builds conversations today, and staff senders are already distinguished from customers; (ii) transferred calls are recorded and transcribed including the human segment — the mining fleet reconstructed the fact sheet from exactly those segments, so this path is proven; (iii) dashboard replies are native. The genuinely invisible surfaces — a personal-cell callback, an off-thread personal-inbox email — get two mitigations: product affordances that make the observable path the easy path (click-to-call from the matter card so the callback rides our line; the "cc Clara and she learns it" habit), and the teach-back nudge — because the matter record knows something was open and who owned it, the same machinery that sends SLA reminders can ask the owner one line after resolution: "What did you end up telling them? I'll handle it myself next time." That sweeps even fully off-surface answers into the loop at the cost of one short reply. And the observer does double duty: "staff answered" is one detected event with two consumers — it releases the escalated matter (P3's reply-driven transition, which nothing performs today; only a manual dashboard click releases) and it captures the teaching candidate (P7). One mechanism closes both of the current system's worst gaps: the one-way latch and the learning that never happens. The teach-back nudge inherits the same duality — staff confirming "I told him X" both closes the matter and teaches it.

2 — The escalation card schema carries the answer from day one. Question → staff answer → policy is one linked record, so even before any automation exists, the raw material accumulates and the flywheel can be switched on retroactively ("Clara learned this from Erika on Aug 20" stays auditable). A schema that only records "forwarded" forecloses the loop.

3 — The knowledge store accepts learned answers directly, not just operator-authored facts — and learned knowledge is property-scoped as a hard isolation invariant. Ruled — modified The earlier draft put a review state in front of every learned fact so leadership could bless or veto it before it became policy. That gate is removed: a staff reply is the policy, live from the moment it is sent. Every learned fact stays provenance-stamped — who answered, when, out of which matter — so it is auditable and deletable in one step, and corrections ride the same reply path rather than a separate admin surface. PMS-agnostic like everything else. Most critically: what Clara learns at one property can NEVER bleed into answers at another. Policies genuinely differ building to building ($300 deposit at Camellia is not a fact about anywhere else), and across organizations it's a confidentiality breach, not just a wrong answer. Scoping is enforced structurally — the read path is keyed by property, so cross-property leakage is impossible by construction, not by prompt discipline — and fenced in the harness (a question answered at property A with property B's learned fact is a red cell). Any future portfolio-level sharing within one owner's org is an explicit opt-in promotion, never a default. Audited the same night: the cross-property leak audit found this wall holds wherever the knowledge base is involved — and found three live paths that go around it entirely, two crossing customer organizations. The wall has to be built into the leak-prone code paths too, not just the store.

P8 — Honesty over bluffing, and measure resolution, not containment.

First-person uncertainty ("I'm not sure about that fee — let me check with the team") paired with the ask; never invent policy (Air Canada lost in court for exactly that). Scorecard: escalation rate trending down, time-to-first-human-touch, staff SLA hit rate, % resolved by micro-answer vs takeover, repeat-escalations on the same question → 0, re-contact rate after "resolved."

Guards the overcorrection risk: today Clara over-escalates; a redesign tuned only for deflection produces the CFPB doom loop. Both extremes are documented failures.

5.3 · The matter lifecycle these principles imply

open · with Clara waiting on staff waiting on customer resolved human-owned answering, collecting owner + card + 2 clocks running clock paused on us provisional, then closed emotion / negotiation / authority — explicit takeover, explicit return needs staff input needs their reply judgment matter staff reply = transition their reply = transition human closes; reply-window reopens Clara keeps serving the person's OTHER matters in every state. Compare the current machine: five ways into one latched flag, one manual exit.
Fig. E — The lifecycle the principles imply, shown for discussion, not as a spec. Its defining properties: replies are transitions, clocks live on the pause states, human-owned is a marked exception with a return path, and none of it is scoped to a channel or a person.

5.4 · Track A decisions you own

What does an escalated-matter caller hear on voice, today, before the redesign lands?

Your standing directive is transfer-to-human (drafted, on hold). Under the rethink, the end-state is P5/P6: Clara states the commitment ("Erika has this; you'll hear back by 2 PM — want me to connect you now?") and transfers on request or breach. Recommend: ship transfer-first now as the interim posture; it's strictly better than the brush-off and compatible with the end-state.

Vendors: same system or their own lane?

Vendors are the worst cohort (1 of 35 answered; A&K trained by failure to demand humans by name) but their needs are mechanical: PO numbers, directions, occupancy checks, order status — system lookups, not conversation. Recommend: treat vendor service as Track B knowledge/tooling (give Clara the WO/PO read surface), keep escalation semantics shared. The existing per-property vendor transfer-first posture stays as the dial-down.

Staff capacity awareness — how ambitious?

The research supports workload-aware escalation (a 1–3 person office; escalate the highest-value deferrable cases up to capacity). That's a v2 refinement; v1 needs only the queue + clocks + reoffer. Recommend: defer, but keep the decision rule in one pure function so it can learn later.

Disclosure policy.

Callers already treat Clara as human (Paula negotiated rent with "Erika"; a renewal call got "Shut up, scammer"). 74% of renters lose trust when AI is undisclosed; disclosed bots take a conversion penalty that fast competence offsets. Recommend: consistent light disclosure + competence-first design; decide once, encode everywhere.

6 · Track B — the journey baseline one Clara, the coverage contract, common sense, the harness

Track A fixes what happens when Clara can't answer. Track B shrinks how often that happens — from today's 37% self-answer (16% after application, 3% for vendors) to "the vast majority, with common sense."

6.1 · The baseline principle: one Clara

Reserved The two-Claras merge and the channel un-weld are explicitly out of scope for the current push by Fede's instruction — the lane architecture is not being touched in this round and gets its own session. The landmine table in §7 step 4 stays as the map for that future work.

One knowledge store, one journey-stage model, one escalation system — channels are transports with physics (length limits, TTS latency, email threading), not brains. The audit's Class-A list defines what channels may legitimately vary; everything else reads from shared state. The repo already contains the proof this works: the denied-applicant scope shipped to all three lanes in one PR with drift fences both sides. That pattern — one pure resolver, per-channel renderers, fences — is the template for every capability, and the drift test becomes structural: a grounding block that exists on one lane and not the others fails CI with an explicit exemption required.

6.2 · The coverage contract

The journey-stage model is the product's spine: every stage names the questions it must answer (from the 110-question researched bank merged with the 149 harvested real questions), the policy data a property fills in once to answer them, and what legitimately escalates. Three sources make this nearly free to populate for Camellia: the mining report §5 is the fact sheet already transcribed from humans' own answers; the CO-law bank flags the legally sensitive ones (2× rent income cap, portable screening reports, adverse-action letters — where a generic answer is illegal); and the escalation-reason ledger names the eight missing policy areas.

North star — ruled, Fede, Aug 13: at least 80% of questions answered by Clara in every journey stage. Not an average across the funnel — a floor each stage has to clear on its own. This is the number the table below is measured against, and today only one stage clears it: tour 85% , inquiry 54%, application → lease 12%, key pickup 13%, early tenancy 0%, vendors 3%.

How this sits with the §6.4 bar: ≥80% per stage is the product north star that this coverage table tracks over time. The ≥90% of templatable questions proposed in §6.4 is the narrower harness definition-of-done — it grades only the questions that have a written answer, per run. Different denominators, both live: one says the product is good enough, the other says a build is shippable.

StageToday (mined)Contract: must self-answerLegitimately human
Inquiry / shopping64%pricing, specials mechanics, pets, parking rates+waitlist state, utilities estimates, income/screening policy, unit facts (laundry, sq ft, floor plans)negotiation, exceptions
Tour~71%format (guided?), access, reschedule, what-to-bring, post-tour follow-ups
Application / in review16%"what happens now," screening timeline, decision-letter method, co-signers/guarantors, unit-swap-before-signing, fee facts, doc chasingborderline screening calls
Approved → lease → move-in(in the 16%)lease terms, deposit timing, proration, utilities setup, check payee, key pickup, parking signup — gated on cross-stage blockers: never "all set" while co-signers unsigneddocument corrections, side agreements
Early tenancy0%portal, payment methods, maintenance intake, announcements — every channel incl. emaildisputes, security judgment
Vendors3%PO numbers, site directions, occupancy checks, order/WO statusapprovals, scheduling negotiation

Tonight's goals (Aug 13 → 14, Willows after-run)

What the overnight run has to show, stage by stage, against the contract above. The after-run publishes as its own doc page with this table filled in met / missed per stage.

StageGoal for tonightDrivers / notes
Key pickup / move-inThe headline. Booking works end to end on all three channels; zero fabricated confirmations; zero wrong answers anywhere in the signed stage.PR #5728 (data fix) and #5730 (grounding fix), acked by Fede tonight; the honesty fix is in the build.
Application → leaseEvery "answered wrongly — contradicts journey state" cell goes to zero; correct-answer rate moves up; questions that are legitimately human stay escalated and tracked.Tracked, not just escalated, is the part that's new.
Inquiry / shoppingNo regression. Not targeted tonight — though the invented voice pet-rent answer should die with the honesty fix.Watch only.
TourNo regression.Watch only.
Early tenancy the black holeNot fixable by merging anything. Tonight's goal is the Willows escalation-matter prototype passing end to end: unanswerable question → tracked matter → decision-request email → staff reply → matter released → renter gets the answer → policy candidate stored.Fede approved the direction as a prototype tonight — “build a prototype for the Willows first.”

6.3 · Common sense, encoded

Four cheap invariants the corpus shows are missing, none of which is "more AI":

6.4 · The harness is the contract's enforcement

Built before the redesign code, per your working order, on the Willows bench with machinery that already exists (pipeline-lab multi-turn presets with per-turn channel, the seam matrix, the willows harnesses):

The baseline bar.

Propose: ≥90% of contract (templatable) questions answered correctly on every channel; 100% of escalations notified + promised + tracked; zero silent drops; zero transfers into voicemail without message capture. The harness reports these four numbers per run. Agree or adjust — this becomes the definition of done for the whole effort.

Distinct from, and subordinate to, the ruled north star in §6.2: ≥80% of questions answered by Clara in every journey stage. That one is settled and measures the product across all real questions; this ≥90% bar is still proposed and grades only the templatable subset, per harness run.

7 · Sequencing (for discussion) what order and why
StepWhatWhy this order
0Approve/amend this architecture; decide the four Track-A decisions — done Aug 13 (principles ruled one by one; only decision 2 still open)Everything else hangs off it
1Build the harness + coverage contract (Track B §6.4) — ruled: start here (Fede, Aug 13)Your rule: harness before code. It will grade every subsequent change and immediately quantify today's baseline
2Camellia fact sheet into the one knowledge store, all channelsCheapest, biggest lever: kills ~44% of human handoffs; content already transcribed in mining §5
3Escalation core per §5 (matter lifecycle, queue, clocks, micro-escalation) — with the question→answer→candidate link in the schema from day one, even though the learning automation is deferredThe structural fix; interim voice transfer-first ships ahead of it per your directive; a "forwarded"-only schema would foreclose the teaching loop
4Channel un-weld, ordered by the audit's landmine table — reserved for a later sessionNot in the current push by Fede's instruction. Same-day wins first (language-on-email, email media, tour facts); the two-Claras merge is the long arc
8 · Related pages on this site the parallel RCA and the Aug-4 ownership doc

Evidence files (session archive): camellia-full-mining-report.md + 203-finding catalog · channel-fork-audit.md (file:line landmines) · hitl-escalation-research.md (all source URLs) · question-bank-research.md (incl. CO law: 2×-rent cap SB23-184, portable screening HB23-1099, adverse-action rules). Provenance note: the "Escalated is terminal (decision, Fede 2026-08-03)" doctrine cited by PRs #5388/#5400/#5401 was an inference from one incident, never an approved decision, and is repudiated by this document.

PropFlow Docs