What Clara is tested on, and what people actually ask

Clara's 2,487 eval cases measured against 2,340 real inbound messages from a live 397-door portfolio — where the testing effort points, versus where the traffic is.

2026-09-02 · Situs Group, CY2026, 16 multifamily properties / 397 units · audit only, no pass rates measured

The suite is aimed at acquiring residents; the traffic is about having them. Tours and applications hold 24.5% of the eval suite against 3.3% of real tenant inbound. Rent and balance — the largest real topic at 12.0% — holds 5.9%.

1 · The gap is a shape before it is a topic

Before any subject-matter gap, there is a structural one. Nearly half of what residents and prospects actually send is not a clean question. It is a fragment mid-thread — “Are we good now?”, “What time is the cleaner arriving?” — that only means anything against what came before.

The suite is built almost entirely from the opposite shape: one self-contained ask, no history.

45.7%

of real inbound is mid-conversation continuation — a fragment that needs prior history to mean anything

1,069 of 2,340 messages

2.2%

of eval cases are a true continuation — prior history plus a new inbound to reconcile

55 of 2,487 cases, across 7 files

0

continuation cases in the two suites Clara's drafting quality is actually gated on

leasing-response-quality · real-replays

A separate 23.1% of inbound (541 messages) is bare acknowledgement — “ok thank you”, “got it”. That shape is covered, by the 217 rows of mass-comms-confirm-reply.jsonl, and it is counted separately here on purpose. An earlier draft of this page combined the two into a single 68.8% figure, which overstated the untested share; the replay corpus itself removes bare acknowledgements before grading, for the same reason.

2 · Where the attention goes

Each row compares one topic's share of real resident inbound against its share of the eval suite. Bars run outward from the centre on a common 0–25% scale.

Share of real inbound (tenant, CY2026) Share of eval cases
3.3%
Tours, pricing & applicationsthe acquisition funnel
24.5%
12.0%
Rent, balance & paymentlargest real topic
5.9%
4.8%
Renewal & move-out
12.7%
6.6%
Maintenance & repair
10.8%
1.7%
Parking & tow
2.5%
0.9%
Neighbour, noise & conduct
1.9%
0.9%
Access, keys & entry
1.0%
0.8%
Pest
1.0%
0.5%
Utilities
1.3%
0.5%
Package & mail
0.4%
0.3%
Voucher & housing authority
1.2%
31 msgs
Renter's insurance & liability chargestaff sent 31 messages about it in 2026
0.2%

Real inbound: tenant-side intent taxonomy, 1,288 unique messages, CY2026, multifamily only. Eval: 2,487 cases across 87 dataset files. Liability has no intent-table row of its own — the 31 is outbound staff messages about the $10.50/mo landlord-liability charge, which is why its bar is drawn against a different unit and marked.

RatioTopicReading
7.4×Tours & applicationsOver-tested against real tenant leasing traffic
2.6×RenewalsOver-tested
0.5×Rent & balanceUnder-tested — and it is the largest real topic
0.04×Renter's insurance & liabilityFive cases, against a charge staff explain constantly

3 · Coverage that isn't

Four topics look covered by keyword and are not covered in substance — the cases exist, but they put the words in the wrong speaker's mouth. A prospect asking whether parking is included does not test a resident disputing a tow.

TopicCasesWhat they actually testVerdict
Parking & tow62Almost all prospect-side “is parking included”, not tow or permit disputeswrong speaker
Neighbour & noise47Mostly fair-housing steering probes (“a quiet building for families”), not a resident complaintwrong speaker
Pest2512 of 25 are turnover-scope line items — a charge being adjudicated, not a resident reporting roacheswrong speaker
Voucher & Section 830All source-of-income discrimination compliance; no voucher operationswrong speaker
Entry & permission to enter25Genuine, but 1.0% of the suite for something staff chase constantlythin

4 · Spanish is a lane, not a flavour

The portfolio sends 18,020 Spanish messages a year. The suite treats Spanish as occasional seasoning inside English-majority files, and no gate anywhere is keyed to language.

40.6%

of outbound SMS is Spanish (email 22.8%)

8,396 of 20,686 texts

3.2%

of eval cases are Spanish — and a third of those are tour-scheduling intent extraction

~80 of 2,487 · 27 are tour intent

0

language-keyed pass floors anywhere, and zero Spanish renewal cases

Spanish maintenance: 2 cases

There is a data problem underneath the eval problem. 753 tenants received Spanish content in 2026. Of the 557 that could be matched into the tenant directory, 392 carry no language tag at all — only 80 are tagged Spanish. The portfolio is running a bilingual operation while mostly not recording who speaks which language.

5 · Almost none of it is real

Of 2,487 cases, roughly 111 (4.5%) derive from captured production traffic. The rest is authored — much of it carefully reconstructed from real incidents, but reconstructed.

The single real-conversation dataset, real-replays-email-sms.yaml, holds 42 rows. The Situs corpus now on disk holds roughly 1,500 replay-ready real turns — an inbound message, its property context, and the actual human reply as a reference answer. That is a 36× increase in real material, already captured and already scrubbed.

DatasetRowsRealProvenance
real-replays-email-sms.yaml4242Camellia production conversations, Apr–Aug 2026, de-identified at authoring
mass-comms-confirm-reply.jsonl21754Real PM replies to live approve/send cards
vendor-completion-notes.jsonl6215Real production work-order notes
fair-housing-gate.jsonl1460Fully synthetic, per its own source field
broadcast-leak-gate.jsonl830Fully synthetic, per its own source field

6 · What this changes

A · Add a continuation lane, with its own floor

The cheapest high-value change. Two-thirds of traffic is mid-thread; 55 cases test that shape and none of them gate anything. A continuation suite built from the corpus would exercise the shape that actually arrives.

B · Add tenancy coverage — don't thin acquisition

Tours and applications hold 24.5% of the suite against 3.3% of tenant traffic; rent and balance holds 5.9% against 12.0%. The verb matters: this is a regression suite, and the tour cases exist because tour bugs cost leases. The fix is adding tenancy cases, not deleting acquisition ones.

C · Give Spanish a floor, and fix the tags

A language-keyed pass threshold is one entry in the gates file. The missing tags on 392 tenants are a data fix, and they gate anything that routes by language.

D · Re-cast the four wrong-speaker topics

Parking, noise, pest and voucher have cases, but they test the prospect or the turnover adjudicator. Re-casting them as resident asks is editing existing files, not writing new ones.

7 · Method & caveats

Sources: ~/code/datasets/extracted/situs-analysis-2026-09-02/tables/ (intent tables, language_tag_gap.csv, _messages.jsonl) and ~/code/PropFlow/propflowai/evals/.

PropFlow Docs