Vendor Interactions — Architecture Review
2026-08-12 · Fede + Fable (orchestrator) · evidence: 2,581 labeled Camellia emails (2018→today), 60+ bench voice calls, turnover-outbound retros
ACCEPTED — Fede, 2026-08-12
v2 — 10,000-foot review
Thesis
Vendor identity is the foundation for every vendor interaction we will ever do — inbound scheduling today, outbound turnover coordination tomorrow, PO exchange, quotes, chases. We have now watched the same failure on both directions of traffic: comprehension is fine, recognition is broken. Inbound, the pipeline read a carpet-cleaning schedule perfectly three times and dropped it three times because the sender's address wasn't in a table. Outbound, turnover vendor-calling had to be pulled back behind human approval largely because identity flakiness was frustrating people — real vendors greeted as strangers, confirmed bookings voided over a transcription of a surname, a vendor record that turned out to contain our own phone number.
Legend for what follows: FEDE direction from Fede · DATA measured on the labeled corpus / documented incidents · FABLE my recommendation, argued, disagreeable.
1 · The landscape: what vendors actually send us
All counts measured on the golden set: 2,581 inbound Camellia emails matching vendor/scheduling vocabulary, 2018→2026, QA-passed labels. 289 are genuine vendor-scheduling emails; 251 should have produced a calendar action. 112 distinct vendor companies wrote in.
1.1 Two very different species of scheduling email DATA
149 of 289 are templated scheduler-system emails — but 127 of those are a single vendor (Miracle Method's "What to Expect — Unit 403 — PO #693" family, with unit, PO and a 3-day window in the subject line). Strip that one vendor and the world is 88% free-text human email: the cleaning company owner writing "we are scheduled to start tomorrow Thursday 8/13 & Friday 8/14, 8 a.m., 6th floor," the painter whose subject line is just his company name, the fire-protection dispatcher offering "next available tech 9:30am Friday" without a date. Elevator and glass companies sit in between — portal confirmations plus a human "our tech is en route" on top.
Consequence: a template parser is the right tool for exactly one lane per vendor system; the general case is judgment over prose.
1.2 Sender addresses are not a usable vendor key DATA
Among repeat vendors, 54% wrote from two or more addresses; 31% from two or more domains. Only 3 emails in eight years ever announced an address change. The elevator company used ten addresses across two corporate domains (they rebranded) plus a billing platform; the utility-billing vendor used eight across three; the law firm sends from its own domain and a Zendesk sender; the cleaning company moved from AOL to a company domain mid-relationship. Among vendor mail: 653 company-domain personal addresses, 86 shared dispatch/role mailboxes, 48 freemail (AOL/Gmail — over-represented among small trades that actually schedule), 19 no-reply senders.
Consequence: any design keyed on "the address on file" is structurally wrong, not just under-populated. Addresses are evidence, never keys.
1.3 Company names drift too DATA
49 distinct stated-name vs roster-name divergences: ThyssenKrupp → TK Elevator (renamed 2020, 57 emails), Anchor Pest Control → "Complete Pest & Wildlife Services dba Anchor," Abbotts Fire & Flood → Abbotts Cleanup and Restoration (no shared words), Alpine Glass → Alpine Glass Acquisition LLC, the law firm adding a third partner name. All were confidently resolvable by a human reading context. One was a genuine trap: "Noble Electric" fuzzy-matched to "Noble Energy Conservation" — different companies, different trades. And 12 scheduling emails came from businesses on no roster at all (laundry-machine servicer, parking operator, Google Fiber install).
Consequence: name matching must be judgment-with-evidence (trade, history, context), not string distance alone — string distance both misses real renames and invents false ones.
1.4 The traps outnumber the targets ~10:1 DATA FEDE
Fede's instinct — "filter marketing/spam; only schedule jobs for vendors we're actually working with" — is what the data screams. Against 251 real calendar actions stand 2,330 no-booking emails, many engineered by reality to look bookable:
- 117 availability pitches — "do you have units for me next week?", a laundry servicer offering a menu of possible days with nothing agreed;
- 73 emails whose dates live only in an attachment the pipeline cannot see (glass-installation confirmations as image attachments);
- 71 staff emails quoting vendor language ("I'll get a tech out there");
- 49 invoices carrying past service dates;
- 13 marketing blasts from genuine roster vendors — the trade association's dated golf-event invites come from a real vendor domain;
- 12 out-of-office replies containing dates (this hole is already closed — see §5).
Consequence: precision is the binding constraint, not recall. The gate that does the heavy lifting isn't "is this a vendor?" but "is this an active working relationship talking about agreed work?" — recent thread history, open jobs/POs with that company, a visit being committed rather than discussed. That gate is a first-class stage in the architecture below.
1.5 Scheduling language assumes shared context DATA
Of 266 emails with extractable schedules: only 50 state an explicit time window; 126 say a bare weekday ("crew will be there Thursday"), 82 say "today/tomorrow," 39 more use relative phrasing ("week of Nov 24, business hours"). 141 are multi-day jobs. 43% of scheduling emails are replies on threads the office started — and those omit property, scope, and system because the office already said it. Reschedules are starkest: of 22, only 2 restate the original date; 16 reference it purely by riding the same thread (one painter's Friday→Monday move lives on a thread titled "RE: phone number?").
Consequence: a per-message system cannot even reschedule — it can only double-book. Processing must be conversation-level: the model reads the thread, not the message.
2 · What broke last time: the outbound record DATA
The turnover workstream built full outbound vendor coordination — voice calling with a PM-approval queue and autonomous redials (live today on the test line; ~60 bench calls, never a real vendor), and email dispatch (built, but hard-redirected to a test mailbox since 7/23; auto-dispatch of turnover work stopped 8/4 — everything now waits for a PM). It never rolled out to real vendors, and the biggest single reason was identity. The incident log, verbatim from the retros:
- A painting company's contact record carried Camellia's own phone number — "a job dispatched to that vendor today calls the property." The record looked perfectly well-formed until dialed.
- A live demo booking — confirmed Monday 1–4 pm — was thrown away because the transcriber wrote "Shabba" for "Chapa" and the string matcher scored it below threshold: verdict "wrong company." No transcriber will ever spell an unfamiliar surname the way its owner does.
- An appliance vendor with nine labeled calls on the same number was told, on that number, "I don't have you on file as a vendor with us yet."
- Two independent name matchers (call path vs dispatch path) gave opposite answers to the same spoken company name — one auto-assigning at full confidence while the other said no such vendor exists.
- Clara answering an inbound callback with the property's own name was read as "the wrong company answered" and every reschedule refused (since fixed).
- Work orders silently bound to a different vendor while telling the PM "unassigned"; a callback pairing one job with another job's time, read aloud, confirmed by the vendor.
- A person who is both PM staff and a vendor contact was swallowed by the vendor lane — the message silently lost. Real customers have dual-role people.
- Live today, voice side: the appliance tech's cell — confirmed as that vendor by the voicemail match on 2026-08-12 and acted on (calendar event created) — still shows as "Unknown Caller" on his 6th call the same afternoon, summarized as "inquired about leasing an apartment." Another number has called 121 times since May, topic-tagged "vendor/delivery," transcripts literally saying "Delivery" — still Unknown Caller, still role: prospect. The system half-knows both (one Person per number since first call); nothing ever writes the identity back.
- And the eval blind spot that let all this survive: "this replay assumes Clara recognizes the caller; it does not test whether she does." Identity was a test fixture, never a lookup — so the test suite was structurally blind to the entire failure class.
Sorting the post-mortems: the failures that frustrated people trace overwhelmingly to (a) a nearly-empty contact store (78 of ~814 vendors with a phone, 27 with an email, zero learning — a contact that misses once misses forever), (b) binary exact-match recognition with no confidence tiers, and (c) duplicated matchers with inverted safety postures. Latency, IVR turn-taking, and process gaps were real but separable. Fixing identity once, as shared infrastructure, is the prerequisite for ever re-arming outbound — which is why this review treats inbound scheduling as wave one of a foundation, not a feature.
3 · Target architecture
3.1 The flow
3.2 The data model
The pieces, and who says so
| Piece | What / why | Provenance |
| Company-centric contact graph | New VendorContactPoint rows (phone / email / domain, with freemail guard) hanging off VendorCompany — many-to-many, provenance + confidence per contact, optional link to a known human. Deliberately NOT Person-spine claims: the spine's one-identifier-one-human rule is right for tenants and structurally wrong for dispatch mailboxes and dozens of techs. | FEDE FABLE concur; §1.2 shows even this is only evidence, never a key |
| One recognition service, judgment-based | Replaces today's four scattered matchers (email claim walk, three voice rungs, two name matchers). Evidence weighed together by a modern model — "reply on our thread, signed Metropolitan Building Maintenance, domain metrobm.com, about the carpet job we requested" resolves the way a human resolves it. Kills the Shabba/Chapa class: one noisy string can no longer outvote the rest of the evidence. | FEDE ("common sense as a human would") + DATA §2 (duplicated matchers) |
| Working-relationship gate | Bookings require an active engagement: recent thread, open job/PO, or established history — and a committed visit, not an availability pitch. This, not identity alone, is what beats the 10:1 trap ratio. | FEDE (filter marketing; only vendors we work with) + DATA §1.4 |
| Conversation-level processing | The classifier/extractor sees the thread and prior calendar events, not a lone truncated body. Required for replies (43% of scheduling), relative dates (§1.5), and any reschedule at all. | FABLE — I consider this non-negotiable; per-message processing caused both the inbound drops and the audit's worst findings |
| Learning loop | Confident resolutions write contact points with provenance; next contact from that address/number is instant. Reuses the shipped provenance pattern (observed → usable for recognition/calendar; confirmed → usable for outbound dispatch). | FEDE; mechanics FABLE; Decision 1 sets aggressiveness |
| Start loose, tighten on evidence | Weak-evidence outcomes are calendar-only — a wrong match is one deletable event, a missed visit is a no-show a PM eats. Every automated write records its evidence, so abuse is detectable and tightening is data-driven. | FEDE FABLE concur — with the §4 rails as the price |
Where I'd push, beyond the settled direction FABLE
- Build recognition as the shared service from day one, even though only email consumes it first. The outbound retros show four matchers disagreeing; if wave 2 ships as "a fifth matcher, for email," we deepen the disease. Same module, voice adopts it next.
- The relationship gate is where the intelligence should live. Identity says "this is TK Elevator"; the gate says "and we have an open elevator-inspection thread, so this 'tech arriving Tuesday' is real." Most trap categories (pitches, marketing, invoices) die here even when the sender is a genuine roster vendor.
- Attachments stay out of scope, explicitly. 73 emails hide dates in images/PDFs. Reading them is a separate, later capability; until then the correct behavior is "no booking" — the golden set already encodes that.
- Never auto-book from marketing-shaped mail even on a perfect identity match — the trade-association case proves identity and intent are independent axes.
- Identity becomes a measured surface. The old evals assumed recognition; the golden set + replay harness make recognition itself the thing under test, per PR, before merge. That's the structural fix for "the test suite was blind."
4 · Hard rails (what makes loose safe)
- Weak-evidence recognition writes calendar entries only — never outbound messages, never money, never Person-spine writes, never work orders.
- No human identities are ever merged or minted by inference (the 2026-08-09 ruling stands untouched — it forbids fusing two humans; attaching an observed address to a vendor company collapses no human identity, carries no consent/TCPA record, and is visibly recoverable).
- Auto-replies can't book (shipped). Own-mailbox mail can't book. All dedup/idempotency invariants preserved.
- No cross-client pollution (Fede, 2026-08-12): each client has their own vendors. All learned contact points and relationship knowledge are scoped to the client org; recognition for one client never consults another client's data, even for the same real-world company. The root-level VendorCompany row remains only a mirror of the client's own PMS card — nothing learned is ever written or read globally.
- Every automated write carries its evidence record.
- Vendor outbound, when it returns, is email-first with human review of the exact content before send; voice re-arms only after inbound recognition has proven itself on the same identity foundation.
5 · Delivery: small PRs, fleet-built, replay-gated
Measurement is replay against history, not new prod telemetry (Fede): the golden set (251 must-book, 2,330 must-not) is the merge gate; the subscription replay harness scores every PR against it. Current prod baseline: effectively 0 of 251 caught — the recognition dictionary is empty.
| Wave | Ships | Status |
| 1 · Safety | Out-of-office guard before the vendor lane + truthful "armed" comments on outbound calling | MERGED (bot-approved, tested, 2026-08-12) |
| 1b · Harness | Golden-set replay harness (recognition + extraction scoring, pluggable resolvers) | PR in review |
| 2 · Thread provenance | Reply-on-our-thread recognition + relationship gate v1 | next — alone catches the entire §1.5 reply class |
| 3 · Stated identity | Signature/display-name/domain judged against roster, conversation-level extraction | after 2 beats the harness |
| 4 · Learning loop | Contact points written on confident resolutions (per Decision 1) | after 3 |
| 5 · Later | Voice adopts the shared recognizer · vendor-scheduling category in the triage menu · PO answering via HITL email (Decision 3) · outbound turnover re-arm · onboarding mailbox mining (future, per Fede) · attachment reading | not in scope now |
6 · Decisions — ALL SETTLED (Fede, verbally, 2026-08-12)
1 · Learning: store contacts automatically on confident match — no human review, no quarantine, no user-facing tiers (provenance kept internally).
2 · ADR-0097: carve-out accepted, scoped strictly to vendor contact points (implied by #1).
3 · Replies/PO: NO outbound of any kind until identity + calendar is working and baked in prod. First pass = recognition + calendar only.
4 · Seeding: YES — email from the last 2 years only + call history (small) + AppFolio extra phones + production backfill renaming ALL matching "Unknown Caller" records (label-only, no merges).
Original option cards kept below for the record:
1 · How aggressively does the system learn new vendor contacts?
2 · ADR-0097 says email signals never create entity records. Contact points conflict.
3 · Vendors ask us for POs (the carpet vendor did). Voice can answer; email can't. PO mode is on for zero real properties.
4 · Seed contact points from history before the learning loop starts?
Appendix · Corpus quick facts
| Fact | Number |
| Labeled emails / span | 2,581 · Feb 2018 → Aug 2026 |
| Vendor-scheduling emails / calendar actions owed | 289 / 251 (231 book · 19 update · 1 cancel) |
| Distinct vendor companies writing in / roster match rate | 112 / 91% |
| Roster vendors with a stored email address | 0 of 828 |
| Repeat vendors using 2+ sender addresses / 2+ domains | 54% / 31% |
| Trap:target ratio / low-confidence share of scheduling | ~10:1 / 17% (the realistic unattended-automation ceiling) |
| Scheduling emails that are replies on our threads | 43% |
| Reschedules that restate the original date | 2 of 22 |
| One vendor's share of all scheduling volume | 45% (templated refinishing system) |