Clara's 2,487 eval cases measured against 2,340 real inbound messages from a live 397-door portfolio — where the testing effort points, versus where the traffic is.
2026-09-02 · Situs Group, CY2026, 16 multifamily properties / 397 units · audit only, no pass rates measured
The suite is aimed at acquiring residents; the traffic is about having them. Tours and applications hold 24.5% of the eval suite against 3.3% of real tenant inbound. Rent and balance — the largest real topic at 12.0% — holds 5.9%.
Before any subject-matter gap, there is a structural one. Nearly half of what residents and prospects actually send is not a clean question. It is a fragment mid-thread — “Are we good now?”, “What time is the cleaner arriving?” — that only means anything against what came before.
The suite is built almost entirely from the opposite shape: one self-contained ask, no history.
of real inbound is mid-conversation continuation — a fragment that needs prior history to mean anything
1,069 of 2,340 messages
of eval cases are a true continuation — prior history plus a new inbound to reconcile
55 of 2,487 cases, across 7 files
continuation cases in the two suites Clara's drafting quality is actually gated on
leasing-response-quality · real-replays
A separate 23.1% of inbound (541 messages) is bare acknowledgement — “ok thank you”, “got it”. That shape is covered, by the 217 rows of mass-comms-confirm-reply.jsonl, and it is counted separately here on purpose. An earlier draft of this page combined the two into a single 68.8% figure, which overstated the untested share; the replay corpus itself removes bare acknowledgements before grading, for the same reason.
Each row compares one topic's share of real resident inbound against its share of the eval suite. Bars run outward from the centre on a common 0–25% scale.
Real inbound: tenant-side intent taxonomy, 1,288 unique messages, CY2026, multifamily only. Eval: 2,487 cases across 87 dataset files. Liability has no intent-table row of its own — the 31 is outbound staff messages about the $10.50/mo landlord-liability charge, which is why its bar is drawn against a different unit and marked.
| Ratio | Topic | Reading |
|---|---|---|
| 7.4× | Tours & applications | Over-tested against real tenant leasing traffic |
| 2.6× | Renewals | Over-tested |
| 0.5× | Rent & balance | Under-tested — and it is the largest real topic |
| 0.04× | Renter's insurance & liability | Five cases, against a charge staff explain constantly |
Four topics look covered by keyword and are not covered in substance — the cases exist, but they put the words in the wrong speaker's mouth. A prospect asking whether parking is included does not test a resident disputing a tow.
| Topic | Cases | What they actually test | Verdict |
|---|---|---|---|
| Parking & tow | 62 | Almost all prospect-side “is parking included”, not tow or permit disputes | wrong speaker |
| Neighbour & noise | 47 | Mostly fair-housing steering probes (“a quiet building for families”), not a resident complaint | wrong speaker |
| Pest | 25 | 12 of 25 are turnover-scope line items — a charge being adjudicated, not a resident reporting roaches | wrong speaker |
| Voucher & Section 8 | 30 | All source-of-income discrimination compliance; no voucher operations | wrong speaker |
| Entry & permission to enter | 25 | Genuine, but 1.0% of the suite for something staff chase constantly | thin |
The portfolio sends 18,020 Spanish messages a year. The suite treats Spanish as occasional seasoning inside English-majority files, and no gate anywhere is keyed to language.
of outbound SMS is Spanish (email 22.8%)
8,396 of 20,686 texts
of eval cases are Spanish — and a third of those are tour-scheduling intent extraction
~80 of 2,487 · 27 are tour intent
language-keyed pass floors anywhere, and zero Spanish renewal cases
Spanish maintenance: 2 cases
There is a data problem underneath the eval problem. 753 tenants received Spanish content in 2026. Of the 557 that could be matched into the tenant directory, 392 carry no language tag at all — only 80 are tagged Spanish. The portfolio is running a bilingual operation while mostly not recording who speaks which language.
Of 2,487 cases, roughly 111 (4.5%) derive from captured production traffic. The rest is authored — much of it carefully reconstructed from real incidents, but reconstructed.
The single real-conversation dataset, real-replays-email-sms.yaml, holds 42 rows. The Situs corpus now on disk holds roughly 1,500 replay-ready real turns — an inbound message, its property context, and the actual human reply as a reference answer. That is a 36× increase in real material, already captured and already scrubbed.
| Dataset | Rows | Real | Provenance |
|---|---|---|---|
real-replays-email-sms.yaml | 42 | 42 | Camellia production conversations, Apr–Aug 2026, de-identified at authoring |
mass-comms-confirm-reply.jsonl | 217 | 54 | Real PM replies to live approve/send cards |
vendor-completion-notes.jsonl | 62 | 15 | Real production work-order notes |
fair-housing-gate.jsonl | 146 | 0 | Fully synthetic, per its own source field |
broadcast-leak-gate.jsonl | 83 | 0 | Fully synthetic, per its own source field |
The cheapest high-value change. Two-thirds of traffic is mid-thread; 55 cases test that shape and none of them gate anything. A continuation suite built from the corpus would exercise the shape that actually arrives.
Tours and applications hold 24.5% of the suite against 3.3% of tenant traffic; rent and balance holds 5.9% against 12.0%. The verb matters: this is a regression suite, and the tour cases exist because tour bugs cost leases. The fix is adding tenancy cases, not deleting acquisition ones.
A language-keyed pass threshold is one entry in the gates file. The missing tags on 392 tenants are a data fix, and they gate anything that routes by language.
Parking, noise, pest and voucher have cases, but they test the prospect or the turnover adjudicator. Re-casting them as resident asks is editing existing files, not writing new ones.
situs_2026_tenant_inbound_intents.csv, situs_2026_prospect_inbound_intents.csv) — a deterministic regex taxonomy, no model calls, scoped to CY2026 and to the 16 multifamily properties.Sources: ~/code/datasets/extracted/situs-analysis-2026-09-02/tables/ (intent tables, language_tag_gap.csv, _messages.jsonl) and ~/code/PropFlow/propflowai/evals/.