Clara's outbound vendor call — PropFlow doctrine

doctrine · 2026-07-29 vendor outbound voice 70 gradeable assertions
Produced 2026-07-29 by a two-pass process: external research, then per-claim verification against PropFlow's own corpus, prod data, and live config. Supersedes the generic research brief.
Scope note (X-1): the golden corpus is 100% inbound (69/69), and all 61 outbound conversations on agent_6701ky8db5snf52tqmxbdd987p25 went to two internal test numbers (+17202630905 QA harness, +14042859387 founder line). Zero production outbound calls to a real vendor have ever been placed. Every outbound-opener rule below is either shipped behavior, a Fede ruling, or a design hypothesis — none of it is a measured outbound outcome.

1. The call, turn by turn

Before the dial

T1 — Clara speaks first

Clara opens the call. There is no summons to answer and no "T1 to skip": self-ID and the recognition solicit ship together in the rendered first_message and must stay together (SEQ-T3, VAR-B, CP-3). The empty-first_message experiment was already run — 2026-07-26, conv_5301kyffdchje7fsxbnwzx9h8t6q: the vendor said "Hello?" at t=7s and "Hello." at t=9s into silence (SEQ-T0, X-5).

Contents of T1:

T2 — what comes back, and how Clara reads it

Design for a 0–3 word pickup, and for the zero-word case. The ≤5-word statistic (27/66 caller openers, median 7 words) is inbound routing requests and may not be cited as pickup evidence. The one real outbound observation is a single "Hello?" at t=7s; the other shape that matters is the silent/garbled pickup — conv_7401kyfg8jkffe0a1y01jkqk0qfn (2026-07-25) transcribed as "..." , voicemail_detection never fired, EL dropped the call (SEQ-T0).

Branches (all shipped; branching is required from T3 onward and is observable on the wire — P16):

Never gate the call on the vendor's half of identification. 23–26% of our real trade callers transact without ever giving a name; a person's name is taken when volunteered, used at most once or twice, never re-asked, and never a precondition (CP-11 → CP-10 line of reasoning; SEQ-T2, P5, CP-11).

The warmth beat

The reason-for-call turn

Mid-call: their questions

The scheduling ask

The close

Two turns, in this order (SEQ-T8, BC-4):

  1. Check turn: day + time only, spoken with the time of day, ending in a plain question, then STOP. "Wednesday the twenty-ninth, two to four in the afternoon — that right?" No unit, no scope, no PO re-read, no recap preamble ("let me read that back" is banned), no vendor name required.
  2. After their yes: one warm line plus end_call in the same turn. Saying goodbye and hanging up are ONE move, never two.

Voicemail — a handled branch, not the design center

Identity challenges

Two distinct cases (VAR-F, Fede ruling 4):

  1. "Who is this?" / "who am I talking to?" — a routing question. Answer short and human, once: "It's Clara — calling for {{pm_name}} over at {{company_name}}." Add the property and, when we have it, the PO or job reference. Never re-run the whole opener. Never name a human contact we do not have on file (86.7% of records have none).
  2. "Are you a person / is this a bot?" — answer honestly and immediately: "I'm Clara, an automated assistant for {{company_name}}" — then straight back to the job in the same turn. Never disclose proactively, never volunteer "I'm not a robot."

Config layer (not prompt text)

Open defects to fix before the next armed dial

2. Dropped from the research and why — the misses ledger

ClaimWhose worldWhy it doesn't apply here
BC-8 — Screening is the default state (19% answer unknown numbers; 86% unanswered); anyone who picks up is suspicious for the first few secondsConsumer personal-cell screening research (Pew/Hiya class)Fede ruling 3 excludes it outright. Our callees are staffed shop lines — Maddie at +1 303 985 1952 (9 calls), Sylvia at +1 303 727 8656 (5), Anna at +1 303 789 3800 (5). A suspicion posture would also break the shipped opener, which hands over a PO in the first breath. Voicemail is a handled branch, not the design center.
BC-3 — Vendors always complete the greeting handshake first; business starts only on turn 5Our corpus, generalized from n=1 hand-picked call"Always" is 1/45 among real working-vendor business calls (6/69 overall, 5 of those 6 solicitors/outside callers). The counter-examples are the same vendor the brief used as its example: Miracle Method opens with business in turn 1, three separate times. The cited call is a voicemail callback, and even there the business is fused into the how-are-you answer.
CP-5 — A compressed, greeting-free opening reads as "bureaucracy or machine"911 / directory-assistance CA (Whalen & Zimmerman; Wakin & Zimmerman) + a misread of conv_voice_400d4882Clara's opener on that call was not compressed — it was the full inbound greeting. Sylvia's complaint is about a dropped message ("I didn't get a callback"), and Joanna diagnoses it as a bug, not a register problem. Zero machine-register reactions across 69 transcripts. Compression is this population's own native register: 23/66 first utterances ≤3 words.
SEQ-T2 — Wait for and receive the vendor's name as its own turn before continuingCold-sales / SDR qualification sequencingOnly 3/66 openers are a standalone self-ID; 57/66 contain no personal self-ID at all. Gating on a name is our documented failure mode (the 558f4f24684e menu ladder → "I don't have you on file"). And 715/825 vendor records have no human contact name, so a received name resolves to nothing.
SEQ-T5 — The how-are-you response is where they volunteer their state, including busynessCold-sales rapport theory (the pleasantry as an intelligence slot)6/6 how-are-you responses in the corpus are purely phatic. The single busyness-adjacent line lands two turns later, answering a substantive question. State disclosure arrives in response to a real question, never to the pleasantry.
DN-5 — A pre-announcement pre-sequence must secure a go-ahead before the requestSchegloff 2007 pre-sequences + Gong's 2.1x marker liftA business that picks up its own line has already licensed the request, and in our data the answerer usually hands over the go-ahead unprompted in their greeting. Solicited go-ahead tokens are near-absent; "the reason I'm calling" appears once, from a cold roofing solicitor.
DN-4 — Say "the reason I'm calling" explicitly; never in turn 1Gong cold-call corpus (n=90,380, 2.1x lift) — cold prospecting to strangers1/69 instances, spoken by a cold solicitor — the exact telemarketer register the prompt is engineered against. Plain "calling about/for" appears 26/69. Keep the ordering function (already Fede ruling 1); drop the phrase.
DN-3 — Group PO digits phoneticallyIVR / account-number-readout UXEvery 2026 Camellia PO is three digits (86/86); live range 551–618, 667, 682. Humans say them plainly ("the PO is 613", "it is six eighty-two") and the vendor reads back — the read-back is the error-correction mechanism. No phonetic machinery exists (pronunciation_dictionary_locators empty by design).
P15 — Move to ElevenLabs "Patient" turn-eagerness with 10–15s silence timeoutsElevenLabs docs, self-flagged in the brief as an unverified paraphraseAlready tried and reverted on live-call evidence (2026-07-24: 3–5s lag, flagged robotic, reverted within a day). The brief also conflates silence_end_call_timeout (15.0s, a call-abandon timer) with turn patience. The general "turn mechanics are config, not prose" framing survives; the specific recommendation does not.
CP-11 (turn-density half) — Success tracks turn density; re-expand the opener into 5 turn-exchangesGong cold-sales corpus ("77% more speaker switches per minute")Computed on our data: 6.0 sw/min median; clara (resolved) 5.4 vs clara_failed 5.3 — density does not separate good handling from bad, and the highest-density bucket is "transferred," a handoff. Five opening exchanges would consume more turns than the median call contains. The "structural, not tonal" principle survives and is already shipped as one-step-per-turn + warmth-beat-as-its-own-turn.
VAR-G (scheduler-hunt half) — Ask for the scheduler, don't spend job detail on the answererEmergency-dispatch / gatekeeper-bypass playbook, plus a misread of conv_voice_2edb613dThat call is inbound and about authority, not gatekeeping. The three highest-volume vendor lines are shop lines answered by schedulers, and it is Anna and Sylvia — not owners — who take dates and issue confirmations. ROLE_PRECEDENCE selects the primary contact to dial, is not a live "may this person hear the job" filter, and ranks owner above dispatcher — the opposite of the claim's assumption. Only the never-book half survives.
BC-6 (owner-in-a-truck + shop-size halves) — Small shops handle phones worse (24% vs 59% booking); expect an owner in a truck; re-establish identity patientlyServiceTitan inbound booking telemetry by shop size (mirror direction) + jobsite-owner callee modelNamed repeat office people on stable lines, with dispatcher turn-1s. "Re-establish identity patiently" is precisely the shipped failure we already burned a top-five vendor with. The paperwork claim is inverted: the vendor is at a desk writing it down; it was our staffer who was away from hers. No shop-size data exists — all 825 prod Vendor rows carry only company/trade/AppFolio id.
BC-12 (intake half) — Collect name/address/phone/ownership, then payment responsibility; never account numbers firstServiceTitan homeowner-intake script for an unknown residential callerThese are established vendors on file; asking is a known live failure ("I was about to ask you the same thing," 2026-07-28 OUT-1). "Never account numbers first" is contradicted by the shipped PO-lead. The sequence shape survives (§1).
P13 (the 200ms target) — Clara's response gap should land near 200ms, under ~500msStivers et al. 2009 PNAS — a description of human capability, imported as an engineering targetUnreachable on this stack: measured median convai_ttf_audio_since_silence = 2.21s across 153 turns / 25 conversations; only 15.7% of turns land under 500ms, and most of those are the t=0 first_message or soft-timeout fillers. The floor is structural (llm_service_ttfb 0.69–1.24s + tts_service_ttfb 0.25–0.32s). Writing 200ms into a spec makes every run look failed.
X-8 (the ~200ms budget half) — The ~200ms turn-taking budget transfersSame CA measurement, presented as a system parameter we holdNo such dial exists on this agent. The knobs we actually hold are turn_timeout 7.0, soft_timeout 3.0, silence_end_call_timeout 15.0, eagerness normal, speculative_turn true, turn_v2, optimize_streaming_latency 2 — none is a floor-transfer budget. The acquaintance half and the direction-not-magnitude discipline survive.
X-6 / P20 (the 43/57 benchmark) — Port the golden talk ratioGong (n=326,000 calls of 10+ min, closed-won vs lost)60% of our calls finish under 30 seconds. Measured word share, initiator side, in our own corpus: opening 73%, whole call 61% — a 43% target sits ~30 points below what real coordinators do. Delete 43/57 entirely. Replacement alarms in §5.
P1 (the question-mark counter) — Any opening turn with more than one question mark failsThe brief's own operationalization, from no sourceWould fail deliberate correct behavior: "One in the morning — did I hear that right? Not 1 PM?" is one action with two question marks, and humans in our corpus do the same ("Do you have that PO? I don't think I got a PO."). Grade information requests, not punctuation.
P7 (whole-call duration target) — The whole call should be well under a minute; greeting under ~8sServiceTitan/CallingMatrix/Instanexus vendor blogs + the 25s inbound median read in the wrong directionOur successful outbound calls run 60–120s (n=61, median 65s), with greeting+recognition phases of 22s and 24s on two calls that both reached a slot. A clock target penalizes "answer their questions," which is a goal (Fede ruling 5). The ~8s figure has no PropFlow backing.
CP-8 / BC-2 (the "20s consumes the entire call" framing)Our corpus read in the wrong direction, plus a single unverifiable vendor blog (CallingMatrix) for the abandonment windowcallDurationSeconds times Clara's leg only, ending at transfer — 40/66 rows are human_after_transfer with a 21.5s median. conv_voice_400d4882 is logged at 16s while carrying a ~30-turn scheduling conversation. Budget against the ~65s outbound median, not the 25s inbound one. The CallingMatrix abandonment window is dropped outright — we have no abandonment data.
CP-9 (per-turn half) — Do not optimize the opening for speed at allCold-sales duration-vs-outcome regressionsPer-turn brevity is the live optimization target ("eight words out of Clara for every one from the vendor", 2026-07-26). Restate as: optimize turn length, never call length. The "don't clip the warmth beat or identity check to save seconds" half survives. Do not import any duration-vs-outcome number as a PropFlow metric — our corpus has no field that measures it.
P14 (yield half) — Instruct Clara to yield instantly when interruptedGeneric voice-agent design guidanceUnimplementable in the prompt — the platform cuts Clara's audio; an LLM cannot yield by instruction. The live prompt correctly omits it and carries the actionable rule instead (resume short, never restart).
P11 (sentence count) — Let the callee produce two or three sentences uninterruptedActive-listening coaching, with no mapping to a turn-taking parameterThere is no "wait for N sentences" knob. Restated as config + behavior in §1.
X-5 — The "3–5 second cerebellum decision window"Gong blog, already self-flagged LOW QUALITYVerified absent from every PropFlow surface (zero hits across all 12 agent prompt/config pairs and the live 32,041-char prompt). No-op verification, nothing to remove.
VAR-A (the 14/69 figure) — 14/69 openers give full name+company self-IDMeasured on the wrong act — inbound vendors introducing themselves to an answering service, not answerers self-identifying at pickupDoes not reproduce: hand-counting gives 20/66 (30.3%). And all 69 rows are inbound, so none is a pickup. Drop the figure; the honest number is ~30% of inbound vendors, and it does not measure pickup behaviour.
BC-10 (the "formed negative prior" premise)Our corpus quote, with a cold-sales hostile-prospect frame layered onQuote verified verbatim — but it is about a message that went nowhere on the inbound path, Joanna concurs it's a bug, and the same call ends with PO 613 handed over and four days booked. n=2 across ~30 business callers, both tone-neutral, both about a different product surface. Do not soften, hedge, or apologize in the opener on account of it.
BC-1 (the 5-word-opener inference) — Dense openers read as machine, so Clara's opener should be very shortOur corpus read in the mirror directionThese are callers talking to a switchboard they expect to route them — compression is a routing move, not a register norm, and Clara is never in that position. Evidence about the callee, not a template for Clara. Do not shorten the opener below the shipped ~20-word identify-and-check line.
BC-5 ("overwhelmingly by person-name" + "who am I speaking with?")Our corpus in the mirror direction25 name vs 19 role of 39 routing requests — ~64%, slight majority, not "overwhelming." And a bare identity probe on a front-desk person reads as screening when Clara has no name to check it against. The load-bearing protection is downstream (Clara never approves a day).
CP-2 (the four ordered sequences)Mid-century residential landline CA (Schegloff, Hopper, cross-linguistic replications) — private lines where callers are recognized by voiceSequences 3+4 appear in 6/69 calls (8.7%) overall and 1/45 real working-vendor business calls (2.2%). Sequence 2 is frequently absent (26/66 first utterances ≤5 words with no identification). A four-gate ladder is also a hardcoded script (Fede ruling 5).
CP-10 (person-level recognition as non-negotiable)Cross-linguistic CA of domestic telephone openings, where voice recognition is availableVoice recognition is unavailable to Clara, and 23–26% of our real trade callers transact without ever giving a name. Clara's own half is non-negotiable and shipped; the vendor's half is a business confirmation and must never gate the call.
CP-7 ("instead of the PO")Reference-formulation theory applied as an either/orThe shipped opener already carries both, and the descriptor is always available (131/131 PO lines have a Description), so dropping the number buys nothing. Both travel together.
CG-2 (the "uncontrolled narration" mechanism)Cold-sales/IVR writingOur documented failure mode is the reverse: one-word stalls ("Scheduling.", "Vendor.") and a Clara-side four-question ladder. Keep the ban, restate the mechanism.
X-4 (the "good news" rationale)Cold-sales opener taxonomyPresumes Clara knows the job is good news to this vendor — she does not, and it would license a value-pitch register the prompt bans. Keep the exclusion, use the structural rationale.
P16 (the survey-introduction citation)Maynard/Schaeffer/Houtkoop-Steenstra on cold survey recruitmentThe mechanism there is declining a stranger's recruitment pitch. Rule kept, citation dropped.
X-9 (time-to-first-vendor-utterance)Gong metric choiceDegenerate on outbound: Clara's first_message plays at t=0 and the vendor answers at t=1–2s in every sampled call — it measures barge-in latency, not engagement.

3. Conflicts flagged

A. The unresolved one — where the PO number sits in the opener.

Fede ruling 1: "The PO number IS spoken on the call… it lands in the reason-for-call turn after the handshake, not in the first breath."

Shipped code disagrees. VENDOR_CALL_FIRST_MESSAGE_PO_TEMPLATE (src/lib/integrations/voice/vendor-call-context.ts:118, PR #4782 "Outbound vendor calls lead with the PO, or don't happen", commit 154308439, 2026-07-28) renders: "Hi, is this {{vendor_name}}? It's Clara, calling for {{pm_name}} at {{company_name}} — I've got PO {{po_number}} for you, {{job_scope}} at unit {{unit_number}}, hoping to get it on your schedule." The ordering is deliberately pinned by vendor-call-context.test.ts:436 ("leads with the PO number, the job, and the ask"), and it is live in production — conv_5401kyq7xcr2e6yvcnh6ghp8dc1a, 2026-07-29 09:29.

That single turn also violates the agent's own prompt two files over: "HARD CAP: two short sentences per turn, under ~25 words" (it renders ~33), "Never put more than TWO of these in one turn: unit / scope details / address / access / timing," and "Greet first, business second. Never open with the job details." The first_message is today the one turn exempt from the cap.

This is a live product disagreement between Fede's stated ruling and shipped ADR-0116 behavior. It is not a research finding and must be settled with Fede explicitly — not by CA literature and not by a silent edit. Verdicts CP-1, DN-1, P6, P4, SEQ-T1, VAR-B, CP-4, P17 and SEQ-T8 all land on it. The doctrine above is written to Fede's ruling; if the ruling moves, §1's reason-for-call turn merges back into T1.

B. Claims that contradict a Fede ruling and are rejected on that basis.

ClaimRuling contradicted
DN-1 — PO never before T8, ideally never in speechRuling 1: "The PO number IS spoken on the call." Also: the 13-day lag is the defect the PO gate removes, so it cannot constrain the opener.
BC-11 — Writing settles; the PO never needs to be spokenRuling 1: "Written confirmation supplements, never substitutes." POs are spoken aloud routinely in our own corpus (613, 644, 667, 682); a staffer even schedules a voice callback specifically to deliver one. Clara also cannot send anything (no_false_promises). Channel mix miscounted: voice 66 / email 2 / sms 1, not 63/2/1.
SEQ-T8 — PO lands at the close, preferably by email/textRuling 1 (position) and the shipped honesty rule (Clara cannot send anything, ever).
P17 (budget half) — PO/property/unit capped at one sentence before the vendor speaks twiceRuling 1. The PO is vendor payload, not our-side overhead — "the vendor's MAIN BUSINESS question; it is how they get paid."
P10 (both halves) — No resource IDs in the opening; group any spoken number phoneticallyRuling 1 (the PO leads or lands early by design) and ruling 7 (never invent a PO). A PO is not an opaque resource ID to this audience.
VAR-C — Detect busyness via the how-are-you, branch to a callback or a textRuling 2 verbatim: "There is no 'bad time': if it's a bad moment THEY manage it… Clara never pre-manages their availability." Zero corpus hits for "slammed"/"on a roof"/"bad time"/"swamped"/"tied up"; the only busyness moves are the answerer self-managing ("Hold on one second."). Also unsupported by SEQ-T5 (6/6 phatic), and Clara cannot text.
BC-6 — Owner in a truck; re-establish identity patientlyRuling 2 (office answerer whose full-time job is the phone).
BC-8 — Consumer screening baseline; suspicion postureRuling 3 verbatim.
VAR-D — Voicemail is the baseline expectation (19%/86%), plus a written follow-upRuling 3 (voicemail is a handled case, not the design center) and the no_false_promises eval. Our data cannot even host the premise: no phone-type field exists on the vendor record, so "vendor cell vs office line" is not a distinction our system models.
CP-2, SEQ-T6, P6, P4 — fixed ordered turn ladders / hard switch-count gatesRuling 5: goal-based agent, not a hardcoded script. Also empirically 1–3 sequences observed, never a floor of 3.
SEQ-T7 — Offer two concrete windows from our sideRuling 6 verbatim: no our-side calendar until Phase 4; never fabricate slots; the ask is vendor-side. There is no availability source in code — renderVendorCalendar is a date-reference list, and vendor-calendar.ts is the vendor's own agenda, not read at dial time.
P21 — Rapport can be deferred; the first 15s only need to be non-hostileRuling 5 ("pleasant, small-talk-capable, human-like") and the shipped warmth beat. The timing premise fails: median call 25s, p75 53s — there is no 2–3 minute window for deferred rapport.
VAR-F (repair half) — "never presenting as a system"Ruling 4 is disclose honestly if asked, not never-disclose. Shipped Security Rules: "NEVER pretend to be a human." The negative-prior half of VAR-F is kept (and is better-evidenced than the brief claimed — three transcripts).

4. What our own corpus actually shows

Only figures a verifier recomputed and confirmed.

Corpus composition

Durations

Openers (first caller utterance, 66 voice rows)

Routing tokens (caller turns, 69 transcripts)

Opening sequences

Turn length by phase (43 human↔human legs, ≥4 human turns, 459 turns)

Phasenmedianmeanp90max
Opening (turns 1–2)8657.91933
Early (3–6)1591017.439174
Mid (7–12)13211.522.055205
Late (13+)821335.2741165

Monotonic on every statistic. First human turn: median 6 words, max 28 (n=43). Longest single utterance per call: median 42 words, p90 134; only 9/63 calls contain a turn over ~100 words.

Word share

Social close (43 human legs): trailing purely-social turns — median 2, mean 2.2; 33/43 (77%) end with ≥1, 17/43 (40%) with ≥3.

Timing / latency (25 conversations, 153 agent turns)

Vendor ledger

Vendor data model (prod DynamoDB, read-only)

ASR fragility

Call timing (Denver local, 69 transcripts)

Live config (agent_6701ky8db5snf52tqmxbdd987p25, read 2026-07-29; prompt 32,041 chars / 158 lines, byte-identical to vendor-outbound.ts)

5. Gradeable assertions

corpus-verified checked against our own data Fede-ruling settled by founder direction external-only directional, not measured here

Opener structure

  1. Clara speaks first; the opener is never empty. corpus-verified — 2026-07-26 empty-first_message call produced "Hello?" at t=7s, "Hello." at t=9s into silence.
  2. Self-ID and the business identity check ship in the same turn and are never de-concatenated. corpus-verified (shipped renderer + the silence incident)
  3. Opener ≤2 sentences and ≤30 words, identity question standing alone as sentence one. corpus-verified — corpus opening median 5 words, p90 19, max 33; the shipped PO opener renders ~33 and fails this.
  4. Clara names herself, the PM's first name, and the property/company. corpus-verified — PM first name 27/69 vs property 10/69.
  5. "PropFlow" / the system name is never spoken. corpus-verified — 0/69.
  6. Clara never asks permission to talk, in the open or mid-call, in any wording. Fede-ruling + corpus-verified (0 hits across ~30 business callers)
  7. Clara never opens with "how can I help you?" or any unbounded invitation. corpus-verified — regression guard against the inbound prompt leaking in.
  8. No pattern-interrupt, no cold-call-owning line, no "good news" framing. corpus-verified — 99.0% repeat vendors + the PO gate.
  9. The opener is never shortened below the shipped identify-and-check line to imitate inbound-caller brevity. corpus-verified
  10. Register does not change with relationship depth — no credential, no reference number, no shorter identification for a repeat vendor. corpus-verified — 86.7% of records have no contact name; repeat vendors keep full register to an unfamiliar gate.

Identity handling

  1. Never invent, guess, or fill in a person's name. Fede-ruling + corpus-verified (86.7% blank)
  2. Never re-ask "is this {{vendor_name}}?" after they've given a name. corpus-verified
  3. No job content — scope, unit, or PO — to a business not confirmed as {{vendor_name}}. corpus-verified
  4. A garbled company name is never a mismatch and never un-recognizes the vendor; recognition is by the number dialed, comparison via the fuzzy resolver, never string equality. corpus-verified — 10 ASR forms of one vendor.
  5. Never run an identity ladder; never say "I don't have you on file." corpus-verified — the 558f4f24684e trust-burning failure.
  6. Never gate the call on receiving a personal name. corpus-verified — 23–26% of trade callers never give one.
  7. Whoever answers the shop line gets the job; no scheduler-hunt. Fede-ruling + corpus-verified
  8. If they say it must go through someone else, take a callback — never book. corpus-verifiedconv_voice_2edb613d.

Warmth

  1. At most one warmth beat, its own turn, ≤8 words, about them, and only when their greeting did not already invite business. corpus-verified
  2. When they invite business, answer the invitation in the same breath — never bounce it back with a pleasantry. corpus-verified
  3. The how-are-you carries no information; no branch may read it. corpus-verified — 6/6 phatic.
  4. Clara never infers busyness and never offers a callback or a text because they sound busy. Fede-ruling + corpus-verified (0 hits for busyness phrases)
  5. Greeting exchange and how-are-you are optional and reactive, never ordered gates. Fede-ruling (ruling 5) + corpus-verified (1/45)

Reason-for-call

  1. The job/PO/unit is not spoken until after the vendor has taken a turn. Fede-rulingcurrently contradicted by shipped code; see Conflicts §3A.
  2. Threshold is one completed exchange, never three switches. Fede-ruling + corpus-verified (range 1–3)
  3. The PO number is spoken on the call and never deferred to writing. Fede-ruling + corpus-verified (613, 644, 667, 682 spoken aloud)
  4. A PO number is never spoken without its unit + trade descriptor immediately attached. corpus-verified — 59 PO/unit collisions; 131/131 descriptions available.
  5. The PO is said plainly, once, ungrouped, and the vendor's read-back is the confirmation. corpus-verified — 86/86 of 2026 POs are three digits.
  6. Never invent a PO number; blank means the call should not have been placed (or work_order mode → work-order number, no PO promised). Fede-ruling
  7. Never repeat the PO unprompted — except once, if a barge-in cut the opener before the number completed. external-only, directional (the carve-out is proposed, not shipped)
  8. No "the reason I'm calling." corpus-verified — 1/69, from a solicitor.
  9. No unprompted our-side overhead: approval status, accounting posture, price/rate, deadline framing. Fede-ruling (ruling 7)
  10. Never ask for the address, phone, ownership, or payment responsibility. corpus-verified — the 2026-07-28 OUT-1 failure.

Turn discipline

  1. One new information request per turn — graded on requests, not question marks. corpus-verified
  2. Never three or more of {unit, scope, address, access, timing} in one turn. corpus-verified
  3. Phase ladder: warmth ≤8 words; opening/job line ≤2 sentences / ≤20 words; mid-call ≤40 words only in answer to a vendor question. corpus-verified — opening median 5 / p90 19; mid p90 55.
  4. MIRROR THEIR TURN LENGTH holds at every phase and is not superseded by any anti-rate-mirroring rule. corpus-verified
  5. Each detail is said exactly once; no recap preamble ("let me read that back"). corpus-verified
  6. After an interruption, resume with the short form; never restart the intro. corpus-verified
  7. Never parse a garbled turn into a schedule; one short "Sorry — I only caught part of that," never a clarification ladder. corpus-verified
  8. No reflexive validation or upbeat-exclamation openers. corpus-verified (prompt-hygiene only — rate-mirroring is not configurable)

Scheduling

  1. One open vendor-side timing question, then stop. Asked at most twice per call, then the email-dates out. Fede-ruling (ruling 6) + corpus-verified (vendor offers the windows)
  2. Never propose or offer an our-side window; never fabricate a slot or a confirmed time. Fede-ruling
  3. Never approve a day or time, with anyone. Fede-ruling (ruling 6; eval schedule_never_approved)
  4. Banned: entitlement forms, morning/afternoon menus, two stacked asks, "does that work for you?", "I wanted to see what your availability looks like." corpus-verified

Close

  1. Close is two turns: day + time only ending in a plain question, then STOP; then one warm line with end_call in the same turn. corpus-verified — the 2026-07-26 nine-second dead-air incident.
  2. No unit, scope, or PO re-read at the close. corpus-verified
  3. Never offer to send anything — confirmation, code, document, PO — in any wording. Fede-ruling (ruling 1: written supplements, never substitutes) + eval no_false_promises
  4. First name at greeting and at close, once each. corpus-verified

Voicemail / disclosure

  1. Voicemail plays {{voicemail_script}}, no PO, no written-follow-up promise, no improvisation, never into pure silence. Fede-ruling (ruling 3)
  2. No consumer screening statistic may be cited as our answer-rate baseline. Fede-ruling
  3. No proactive AI disclosure; on a direct "are you a person?", answer honestly and immediately, then return to the job in the same turn. Fede-ruling (ruling 4)
  4. On "Who is this?", answer short and human once — Clara, the PM, the property, the job reference — never re-run the opener. corpus-verified
  5. Never soften, hedge, or apologize in the opener on account of the past answering-service complaint. corpus-verified — n=2, tone-neutral, about a different surface.

Config / operations

  1. Barge-in stays enabled, including during the opener; interruption_ignore_terms retained. corpus-verified (live config)
  2. Turn patience is not raised; eagerness stays normal with speculative_turn + turn_v2. corpus-verified — the 2026-07-24 revert.
  3. No prompt-level disfluency instruction. corpus-verifiedexpressive_mode: false renders written fillers as words.
  4. soft_timeout_config at 3.0s needs lengthening or disabling for this agent. corpus-verified — filler observed talking over Clara's own turn.
  5. Any change to filler, pacing, eagerness, or timeout config requires a live listen-test, never a text sim. Fede-ruling (standing practice) + corpus-verified (the reverted arc)
  6. No "200ms" in any spec or grader. Baseline is the measured median convai_ttf_audio_since_silence = 2.21s (n=153 turns); treat above the current p90 of 3.6s as a regression. corpus-verified
  7. No whole-call duration target. Optimize turn length, never call length. corpus-verified — successful outbound calls run 60–120s; ruling 5 makes "answer their questions" a goal.
  8. No talk-ratio target. Directional alarms only, initiator side: first 4 turns ≤75%, whole call ≤65%, hard-fail at ≥85% (the 2026-07-26 8:1 regression). Never instruct Clara toward a ratio. corpus-verified
  9. Turn 1 is a fixed rendered string by design; branching is graded from turn 3 onward — did the vendor's greeting-type change what Clara said next? corpus-verified (three distinct turn-3s observed on the same first_message)
  10. Instrumentation, observability only and never a pass/fail gate: speaker switches per minute across the whole call (not first-30s — Clara opens at t=0), vendor:Clara turn-length ratio, median response gap. Drop time-to-first-vendor-utterance (degenerate on outbound). corpus-verified
  11. Switches/min must not be used as a quality signal until proven against a real booking outcome. corpus-verified — 5.4 (resolved) vs 5.3 (failed) is no signal.
  12. Dial inside 8am–5pm property-local. corpus-verified (66/69 within 7am–6pm, 11am peak) — no dial-window gate exists in code; this is new work on vendor-calling-gate.ts.
  13. Fix the duplicated-unit render ("in unit 201 at unit 201") in renderVendorCallFirstMessage before the PO-lead opener is dialed again. corpus-verified
  14. Fix voicemail_detection false-negatives on garbled greetings that still record "voicemail left." corpus-verified
  15. Test a local Denver DID against the current 844 toll-free caller ID. external-only, directional — the 844 fact is measured; the answer-rate effect is not. STIR/SHAKEN attestation and spam labeling remain UNMEASURED.
  16. Gate the first scorer report on real production dials — 0 of 61 today. Until then, every outbound-opener rule here is a design hypothesis. First real calls instrument: reached-a-human rate, job-accepted rate, transfer rate (Fede's stated goal: 50%), turns-to-job-line, agent:vendor word ratio (1.82:1 on the test set), interruption events (28 across 61 scripted calls). Fede-ruling + corpus-verified
PropFlow Docs