300 real human phone calls, every one read turn-by-turn (1,355 agent replies measured), plus a 50-call sample of the robot-driven test line, pulled straight from ElevenLabs' own per-turn timers. Window: 2026-07-22 → 2026-08-20 (the last full 30 days). Read-only — nothing in prod was touched to build this.
Half of Clara's replies land inside 2.5 seconds of the caller stopping talking. Nine in ten land inside 5.8 seconds. The worst 1% — long holds, garbled starts, a caller who trails off mid-sentence — run past 16 seconds, and when we opened those calls up, almost all of that time is the caller being genuinely quiet (mid-thought, spelling an email, thinking it over), not Clara being slow to think or speak. The model itself (the "LLM time-to-first-byte") is a median of about 1.0–1.1 seconds across every line we checked; text-to-speech start is a fast, stable ~0.22–0.28 seconds. The slow parts are turn-taking judgment calls — how long to wait before assuming the caller is really done — not raw horsepower.
Every number below is convai_ttf_audio_since_silence — ElevenLabs' own measurement of the gap from "caller went quiet" to "Clara's audio starts playing." This is the number a caller feels. p50 = typical, p90 = the wait one call in ten produces, p99 = the rare bad one.
| Line | Replies measured | p50 | p75 | p90 | p99 |
|---|---|---|---|---|---|
| Clara — Triage (Camellia, the main line) | 1,166 | 2.59s | 3.59s | 5.83s | 15.61s |
| Willows Triage (the prototype interruption-tuned line) | 113 | 2.29s | 3.10s | 4.02s | 14.18s |
| Clara — Escalation Callback (calling a resident back after an escalation) | 33 | 2.29s | 3.33s | 16.46s | 25.94s |
| Clara — Leasing | 22 | 2.15s | 3.41s | 4.49s | 4.69s |
| Clara — Renewal (outbound) | 15 | 1.88s | 6.52s | 11.02s | 11.45s |
| Clara — Turnover Intake | 4 | 3.76s | 4.04s | 4.04s | 4.04s |
| Clara — Emergency Relay (outbound) | 2 | 10.46s | 15.40s | 18.36s | 20.13s |
| All human calls, combined | 1,355 | 2.49s | 3.58s | 5.80s | 16.23s |
| Robot test-line calls (Voice Eval Robot + bench, sampled), for reference | 248 | 2.63s | 3.27s | 4.51s | 12.75s |
Renewal, Turnover, and Emergency Relay have too few calls in this window (2–15) to trust their p90/p99 — flagged as small-sample, not treated as a finding. Triage and Willows have real sample sizes.
We pulled the ten single slowest replies in the whole human-call set and broke each one into two pieces: how long the caller was already quiet before this reply started the clock (marked on ElevenLabs' own turn timer), versus how long Clara's own pipeline took once it decided to respond (model "thinking" time + voice-generation start time, added together). That split is the difference between "Clara is slow" and "the caller was still talking to themselves."
| When (Denver time) | Line | What the call was about | Total wait | Caller was already quiet | Clara's own pipeline |
|---|---|---|---|---|---|
| Aug 13, 3:11pm | Clara Triage | Scheduling a virtual tour | 39.4s | 37.3s | 2.1s |
| Aug 19, 5:42pm | Escalation Callback | "No cash rent" follow-up | 26.4s | 23.8s | 2.6s |
| Aug 17, 10:16am | Clara Triage | Google Maps listing update | 25.9s | 22.2s | 3.7s |
| Aug 19, 5:30pm | Escalation Callback | Window A/C policy question | 25.0s | 23.4s | 1.6s |
| Aug 17, 10:16am | Clara Triage | Google Maps listing update (same call, next reply) | 22.5s | 17.8s | 4.8s |
| Aug 9, 5:01pm | Emergency Relay | Audio/connection check | 20.3s | 18.9s | 1.4s |
| Aug 3, 4:44pm | Clara Triage | Caller went silent mid-call | 18.0s | 13.8s | 4.2s |
| Aug 14, 4:18pm | Clara Triage | Follow-up on an already-handled issue | 17.9s | 16.2s | 1.8s |
| Aug 19, 4:59pm | Escalation Callback | Connection check | 17.1s | 15.0s | 2.0s |
| Aug 6, 6:30pm | Clara Triage | Transferring a delivery notice | 16.8s | 12.3s | 4.5s |
Nine of the ten worst waits are caused by the caller being quiet, not by Clara. Clara's own processing on these calls stays in the 1.4–4.8 second range even on the worst calls in the whole window — it's the same range as a normal reply. What made these ten feel bad is a long real-world pause: someone thinking, spelling something out, or trailing off. Clara is correctly waiting rather than talking over them — that's the right behavior — but from the caller's chair it still reads as "did the call drop?" past about 6-8 seconds.
One clear outlier is worth separate attention — a July 28 call where a reply took 16.6 seconds with essentially zero caller silence beforehand (the caller had already spoken). Model time and voice-start time only account for about 1.2 seconds of that 16.6; the rest isn't explained by any of the per-turn timers ElevenLabs gives us. That's a real, un-attributed processing stall, not a caller pause — it happened once in this sample, so it's noted as a live unknown rather than a pattern.
Yes, but only a little, and it's not the main story. Replies where Clara fired a tool (looking something up, transferring the call, booking a tour) run about 45% slower at the typical case than replies with no tool — but at the worst-case tail, tool-turns are actually a bit faster, because the worst no-tool waits are the caller-silence cases from section 2, not tool slowness.
| Replies | p50 | p90 | p99 | |
|---|---|---|---|---|
| Reply used a tool | 297 | 3.23s | 5.26s | 14.11s |
| Reply used no tool | 1,058 | 2.22s | 6.07s | 16.58s |
| Tool | Times fired | Typical wait (p50) | Worst 1-in-10 (p90) |
|---|---|---|---|
transfer_to_agent — handing off to a specialist | 129 | 2.90s | 4.46s |
transfer_to_number — routing to a human | 60 | 3.79s | 5.86s |
schedule_tour | 26 | 3.48s | 4.64s |
end_call | 25 | 3.16s | 13.04s |
identify_caller | 13 | 3.12s | 3.63s |
reschedule_tour | 11 | 4.22s | 4.69s |
capture_unknown_caller_note | 8 | 3.05s | 10.52s |
Nothing here is a smoking gun — every tool clusters in the 3-4 second typical range, which is Clara's normal pace with a bit of lookup work added on top. end_call has the widest spread (some hang-ups are instant, some carry a trailing pause), and transfer_to_number is consistently the single slowest common tool, which lines up with the separate transfer investigation already on the docs site (see section 8).
Fede's question, answered: Willows is genuinely faster, not just smoother-feeling — by about 12% at the typical case and 31% at the worst-case tail — but the gap is not because the underlying model or voice engine runs quicker on Willows. Same models, same voice stack, close-to-identical "thinking time." The real difference is that Willows calls tools less often (16% of its replies vs. 23% for Camellia's Triage line) and its voice-generation start is marginally faster. Fewer detours into tool calls means fewer of the slower turns from section 3.
| p50 | p90 | Tool-call rate | Model "thinking" time (p50) | Voice-start time (p50) | |
|---|---|---|---|---|---|
| Willows Triage | 2.29s | 4.02s | 16% | 1.09s | 0.22s |
| Clara Triage (Camellia) | 2.59s | 5.83s | 23% | 1.04s | 0.25s |
One wrinkle worth knowing: Willows also gets interrupted by callers 2.6× more often than the Camellia Triage line (13.3% of Willows' replies vs. 5.1% for Camellia — see section 6). Part of "Willows feels faster" may be callers naturally barging in sooner on the tighter, terser prototype phrasing, not the platform alone. Both things can be true: Willows is measurably quicker to first sound, and callers are more willing to talk over it.
| Language | Replies | p50 | p90 |
|---|---|---|---|
| English | 1,321 | 2.52s | 5.83s |
| Spanish | 34 | 1.85s | 4.21s |
Spanish calls measure faster in this window, not slower — the opposite of the "Spanish calls felt different" hunch. But there are only 34 Spanish replies in 30 days (versus 1,321 English), so this is too small a sample to call a real difference either way. Worth re-checking once Spanish call volume grows; not actionable today.
| Line | Replies | Caller interrupted Clara |
|---|---|---|
| Clara — Renewal (outbound) | 15 | 26.7% |
| Clara — Escalation Callback | 33 | 18.2% |
| Willows Triage | 113 | 13.3% |
| Clara Triage (Camellia) | 1,166 | 5.1% |
| Clara — Leasing | 22 | 0.0% |
"Ignored as backchannel" (Clara correctly not reacting to a caller's "mm-hm" or "yeah") reads as effectively zero across every line in this window — either it's genuinely rare, or ElevenLabs isn't populating that flag much in practice. Not enough signal either way to report a rate here; flagging it as an open question rather than a finding.
Fede asked us to check four specific dates against the data. Being honest about what the numbers can and can't show:
Outside of the Aug 19 spike (small sample, one line), day-to-day movement in this window is noise — typical waits bounce between roughly 1.5s and 3.8s with no sustained trend up or down. We're not claiming any of the four dates produced a clean, provable fleet-wide win or regression except the Aug 19 spike, which is real but narrow (one specialist line, one day, already recovered).
Honesty check: the specific "July latency audit" numbers already published on docs.propflowai.co turned out to be about a different thing — /a/latency-audit-2026-07 and /a/latency-ledger-2026-07 measure how long PropFlow's web pages take to load (database scans, API waterfalls), not how long a caller waits on the phone. The local file path we were pointed to (~/.claude/propflowai/docs/voice-latency/*.md) doesn't exist on this machine. So there is no exact prior baseline for this specific number (caller-silence-to-first-audio, measured per turn) to compare against — this deep dive is the first time it's been measured this way, at this scale.
The closest real prior work on voice turn-taking is the transfer-handoff investigation (Jul 28), which measured Clara Triage's "thinking time" before speaking (a narrower slice of what we're measuring here, and only around specialist handoffs) at a median of 0.98 seconds. That's model-thinking time alone, not this doc's full caller-felt wait — so it's not a fair apples-to-apples comparison, but it's consistent with what we found here: the model itself replies in about a second; the multi-second waits callers actually feel come from elsewhere (silence judgment calls, tool detours, and — separately — the transfer/ring-time investigation on Aug 9, which found the "call gets transferred to a human" step averages 18.5 seconds, mostly the destination phone ringing).