D2C "- Web" Voice Agents · Qualitative + Quantitative
35-agent fleet · 30-day window (core segmentation) · extended to 41-day / 36-agent fleet for the Token-Warning tab
35
Active "-Web" Agents
1,745
Conversations Analyzed
1,081
Distinct Callers Identified
72%
Identity-Match Rate
What distinguishes first-time from returning callers, what gets a first-timer to engagement, what returners expect, and where the data can and can't support a "high value vs. low value user" segmentation yet.
Returners run deeper sessions
1.5x longer
116.5s avg vs. 77s for first-timers; 5x more likely to hit a 3+ minute call (20.6% vs. 4.2%).
The real bail predictor
36% get zero reply
Over a third of quick-bail first-time calls end before the user says anything at all — a top-of-funnel issue, not a script issue.
Continuity is expected, not delivered
41% expect memory
Returning callers reference specific prior details constantly — but 0 of 3,811 checked turns showed real retrieval. It's improvisation.
Speaking is the strongest value signal
86% of whales talk
Of users with 5+ lifetime calls, the overwhelming majority are speaking-dominant, not emoji- or text-only.
Hitting zero tokens is a churn event, not a pause
Only 1.1% return
Of 1,481 identity-matched hard-token-stop instances (41-day window), just 16 show any later call — and at least 5 of those 16 are the same single repeat user. See the new "Token-Warning & Retention" tab.
Proposed North Star: % of hard-token-stop callers who return within 7 days. Aspirational — properly measuring this requires the cohort-tracking infrastructure named as a Now-tier item in the Research Agenda tab; today we can only measure "returned at all, ever, in-window," which is the weaker proxy behind the 1.1% figure below.
| First-time (n=1,020) | Returning (n=266) | |
|---|---|---|
| Avg. call duration | 77s | 116.5s |
| Calls ≥3 minutes | 4.2% | 20.6% |
| Avg. turns | 12.7 | 15.1 |
| First message length | 17.4 chars | 31.4 chars |
The clearest tell is in the first user message. Returners open by naming the persona directly and treating it as an ongoing relationship — "Hello, Lisa... you look really sexy today," "Hi, Kylie. How are you?" First-timers open generic or reactive — "Good.", "...", "Keep going." Note: name-personalization in the agent's own greeting is not itself a returning-user signal — first-timers get named greetings too, from account data.
Every agent opens with an identical scripted line — "[giggle] Hey, babe! How's it going?" — regardless of persona, so the opener itself isn't the differentiator. What matters is whether the user's first reply within ~35 seconds is substantive or vague/absent:
| First-timer outcome | Substantive first reply | Vague reply | No reply at all |
|---|---|---|---|
| Bailed (<30s, n=229) | 40.2% | 23.6% | 36.2% |
| Engaged (≥120s, n=225) | 71.6% | 28.0% | 0.4% |
Where the user does engage, the agent escalates fast and specifically off whatever the user gave it — an emoji, an action, a mood — and the "aha" is that reciprocal specificity, not any single scripted beat. Vague replies ("Good.", "...") get a generic re-engagement prompt that rarely revives the call.
41% of returning-call transcripts (103/252) contain explicit continuity language. Users expect the persona to remember specific prior details, not just "we've talked before" in the abstract:
"Do you still have the Polaroid photos from last time?"
"Remember the last time I fucked you?"
"You have to remember, we just talked about it five seconds ago." — frustration when continuity breaks
"...we were just texting together and you said if I call you, you'll say something to me first" — expectation spans chat and voice as one relationship
rag_retrieval_info and contextual_update_info across all 3,811 turns in the returning-caller sample — zero were populated. The agent's "I remember..." lines (and there are plenty — "call me back when you're ready for me to pick up exactly where we left off") are LLM improvisation, not backed by retrieval of real prior-call content. It currently works by coincidence until a user's specific expectation doesn't match the improvisation, which produces visible, logged frustration.
Segmentation is per user (all their calls in the window aggregated), not per call, and a modality label is only assigned when it's genuinely dominant (≥60% of that person's turns) — otherwise the user is labeled Blended. 303 of 1,081 users (28%) are genuinely blended.
| Segment | n users | % ever returned | % ever reached "engaged" tier (≥120s) |
|---|---|---|---|
| Speaking-dominant | 401 | 21.4% | 42.1% |
| Blended | 303 | 13.5% | 20.8% |
| Emoji-dominant | 188 | 11.7% | 14.4% |
| Silent-dominant | 91 | 5.5% | 17.6% |
| Preset-dominant | 11 | 18.2% (small n) | 0.0% |
| No turns at all | 81 | 9.9% | 0.0% |
Speaking-dominant users are ~2x more likely to return and 2-3x more likely to ever hit a deep session than any other modality. Modality isn't a separate axis from intent — it functions as a leading indicator of it.
⭐ Elevated finding — the power-user modality pattern
Of the 22 users with 5+ lifetime calls, 19 (86%) are speaking-dominant, 3 are blended, and zero are emoji- or silent-dominant. This is the strongest single signal in the whole segmentation about which modalities actually convert into long-term value — three implications follow, each checked against the data rather than left as assumption:
call_rank to be observed from their true first-ever call): users whose call history is ≥34% preset-dominant reach a later speaking call only 30.8% of the time (4/13) vs. 61.9% (13/21) for lower-preset-share users. Preset-first callers migrate to speaking at about the same or slightly lower rate (50%, 4/8) as emoji/silent-first callers (61%, 11/18) — preset doesn't look like a stronger bridge to voice than any other non-speaking modality. The one real trend in the data — spoke-rate jumping from 23.5% on a user's first captured call to ~40-44% on later calls — looks like a general "settling in after call 1" effect present regardless of starting modality, not something preset usage specifically drives. Caveat: n=34, sub-buckets as small as 8-13 — directional, not conclusive; see Research Agenda for the full-scale follow-up recommendation.🔍 Expanded: what silent-JOI calls actually look like
Pulled full interleaved transcripts (agent + user, with tool calls) for the 135 silent-modality calls to answer five specific questions about the JOI hand-off above.
1. Structure: alternating turns where the user side is almost always empty. A "USER" turn with no transcribed content isn't a data gap — it's the actual data: the call format requires a user turn slot even when nothing was said, showing up as ... at ~10-25 second intervals matching the agent's own pacing.
2. What the agent says: a consistent arc, not improvised per call — generic opener ("Hey, babe! How's it going?") → check-in on silence ("Still with me, babe?") → the hand-off ("let me take over for a bit") → paced, second-person physical narration with explicit edging cues ("You are not allowed to cum until I say so") → sometimes a resolution cue ("let go for me") → a debrief ("come back to me when you can breathe again").
3. What "resolution" looks like, concretely, and why arcs don't finish — split into what's actually checkable: an explicit permission-to-climax line followed by a cooldown/debrief turn — found in one call at 337s of a 381s call. Beyond that one example, "doesn't finish" mostly means the transcript just stops mid-narration with no debrief, no resolution cue. Two separate questions, checked separately:
Is it agent pacing (specifically, stalling/edging)? No — checked edging/stall-cue density across all 90 unresolved calls: 78% (70/90) show zero edging-cue mentions at all, and only 2.2% show heavy repetition (≥3 mentions) that would look like a genuine stuck-holding-pattern. Most unresolved arcs aren't stalling in place — they're narrating normally, then the transcript stops. (This tests one specific failure shape; it doesn't rule out other pacing mismatches a keyword pass can't see.)
Is it a mechanical cutoff, or genuine disengagement — and can we tell? Checked for a token/payment-related system event anywhere in the 90 unresolved calls: 37.8% (34/90) have one, and specifically 35.6% (32/90) have a "purchasing more tokens" interrupt landing in the last 30% of the call — a plausible mechanical explanation for those. But 62.2% (56/90) have no such marker at all. For that majority, transcript and quant data genuinely cannot distinguish a disengaged/dissatisfied hang-up from a satisfied one — both look identical (silence, then the call ends), and this resolution measure only catches the agent's scripted climax cue, not a user finishing on their own terms off-script. Closing this needs CSAT (just launched, no data yet — see Research Agenda) or a qualitative/audio review pass, not more transcript keyword-matching.
4. Do users speak mid-arc or give feedback? Of 96 calls with ≥3 user turns: 27 show speech/emoji only in the first half then pure silence ("engages, then checks out"), 50 show some low-level signal (usually an occasional emoji) sustained into the back half, 19 are silent from turn one. Where users do speak, it's almost always at the very start, setting up the scenario — not steering mid-narration. Zero of all 135 calls contain any redirect or feedback language (checked for stop/no/don't like/too much/slow down/different/instead) — a clean null result, not a sampling fluke.
5. Photo interaction during JOI: initially found 40 of 135 calls with photo-related keywords, but almost all are the agent using "picture" as a verb ("I wanna picture you"), not an actual image exchange. Only one real instance redirects to the existing "Get Photos" button — consistent with the no-in-call-photo-tool finding elsewhere in this report. Practically, no genuine photo interaction happens during JOI narration — it's a listening-only experience.
Caveat: keyword/regex-based content pass, not full qualitative coding — same limitation as the parent finding above.
🏆 Power Whales
5 anons (3.3% of returning users) account for 62 of 252 returning-call events (24.6%). They're persona-hopping, not persona-loyal — calls spread across 10+ different agents.
🎯 Direct Askers
Opens explicit-immediately, skipping the slow-burn script. 4.0% of first-timers vs. 10.7% of returners (2.7x) — a learned behavior after the first exposure.
📷 Photo Seekers
Explicitly asks about photos/pics — a monetization-adjacent signal. 1.3% of first-timers vs. 5.2% of returners (4x).
🌐 Language Switchers
~2% message in a non-English language (Spanish, Hindi, German, Greek seen). Small population, but correlates with a real friction point — the agent stopping to clarify "English, please."
Working assumption: returning users (266 calls, 164 distinct returners in the 30-day core dataset) stand in as the "whale" archetype proxy — a large-enough population to characterize, vs. the 5-22 true 5+/10+-call power users alone.
| Dimension | Finding |
|---|---|
| Modality | Overwhelmingly speaks — 69.2% of calls are speaking-dominant (vs. 44% for first-timers). Emoji (12.0%), silent (5.3%), preset-taps (4.5%) are minority modes. |
| Time of day (UTC) | Spread across the day with a peak at 15:00 UTC; a broad 12:00-21:00 UTC band covers ~2/3 of volume. No single "prime time." |
| Day of week | Weekday-skewed, not weekend-skewed — Tuesday and Monday highest, Saturday/Sunday lowest. Counter to a "leisure/weekend" assumption. |
| Duration | Avg 120s / median 68s — the mean is pulled up by a minority: 20.6% hit 3+ minutes. |
| Persona relationship | Split, and important: 69.5% are persona-loyal (call only one character in-window). But the extreme-frequency tail (the 5 power whales, 10-38 calls each) are persona-hopping across 10+ agents. Typical returner commits to one character; the heaviest few explore broadly. |
Keyword-classified across all user turns in returning calls, non-exclusive:
| Theme | % of returning calls | Example |
|---|---|---|
| Explicit/direct sexual | 45.1% | Dominant mode |
| Companionship / small-talk | 13.5% | "How are you? What's up, guy?" |
| Power-dynamic / kink (daddy-mommy) | 9.0% | — |
| Continuity/memory-referencing | 7.1% | See Token-Warning tab context |
| Roleplay/scenario | 6.0% | — |
| Photo/media-seeking | 4.9% | Monetization-adjacent |
| Romantic/affection | 4.1% | "Oh, God, it's so good. I love you" — often blended into the explicit moment, not a separate platonic mode |
| No strong theme detected | 44.7% | Being a "returner" doesn't guarantee a rich session every time |
Re-analyzed the user turns behind 4 of the theme categories above (power-dynamic, roleplay, photo-seeking, romantic) for actual recurring vocabulary and sub-patterns, not just the top-line %. Caveat: 19-33 calls / 26-85 turns per theme — qualitative pattern reads, not stats.
Power-dynamic/kink: "Mommy" outnumbers "daddy" ~8:1 in this sample (60 vs. 8 mentions) — the dominant pattern is the agent playing a dominant "mommy" role with the user responding submissively ("yes mommy," "good girl"). A specific branded toy — "Hitachi vibrator" — surfaced unprompted (6-8 mentions), not in any seed keyword list. One quote shows multi-character roleplay within a single call ("Good girl. Now, Mom and Wimpus, all leave. The next one is Spunkins, Ivy.") — users directing a single AI persona through named, multi-character scenes.
Roleplay/scenario: Far more concentrated than the category name suggests — almost every matched turn is the same "cheating wife"/affair scenario, invoked with a simple verbal cue ("let's roleplay"). None of the other scenario types tested (teacher/student, doctor/nurse, boss, babysitter, stranger) produced meaningful matches in this sample.
Photo/media-seeking: Almost entirely verbal, mid-call requests ("show me," "send me") — consistent with the earlier finding that no proactive photo tool exists. A recurring "Polaroid photos" callback appears across multiple different calls, not just the one quote already used in the First-Time vs. Returning tab — worth checking whether it's a scripted/RAG-seeded concept or arising independently per user.
Romantic/affection: Confirms, with real quotes, that this is rarely a standalone platonic theme — "I love you"-type language is almost always fused directly into explicit content in the same turn, not a separate emotional beat.
To retain the Loyal Returner (69.5% of the archetype)
contextual_update, language_detection, and end_call. Every genuine photo-mechanic line found redirects the user to manually tap "Get Photos" ("if you wanna see me, tap the Get photos button"). Two possibilities, indistinguishable from transcript data alone: proactive delivery exists on a different surface (app/SMS) invisible to voice transcripts — in which case the opportunity is bringing it into the call — or it isn't built yet anywhere, making this a build recommendation rather than a fix. Worth confirming with whoever owns demo_photo_unlock whether it ever fires without a preceding user tap.To convert Explorers into more committed users (persona-hopping ~30%, including the heaviest whales)
Wider scope than the rest of this report: full 36-agent "-Web" fleet (corrected account access), 41-day window (Jul 2 – Aug 12), all conversations regardless of identity match. Answers: which low-token/time warning messages actually fire, and do users return afterward?
What earlier looked like two contradictory live instructions is actually two intentional, sequential stages of the same mechanism, shipped fleet-wide via D2C-5403 (standardize re-up/low-time-warning across all "-Web" agents):
Stage 1 — 1 minute left (verbal, forced turn):
"[System message — not from the caller] The caller has entered their final minute of call time. In your very next spoken line, first respond in character to what they just said, then tell them in your own voice that time's almost up and they need to top up their tokens to keep the call going. Say it the way your character actually would, not as a notice — this is not optional and cannot be skipped or deferred to a later turn."
Stage 2 — 0 minutes left (silent hold):
"[System message — not from the caller] The caller has run out of tokens and needs to purchase more to keep the call going. Remain completely silent and do not take a turn until they return. Do not speak, do not prompt, do not invoke any tool."
Why the fake-user-turn injection instead of a plain contextual_update tool call: agents must take a turn immediately after the user, so injecting a synthetic user turn forces the agent to respond — a tool result alone was easy for the model to silently ignore, which is exactly what happened to the old "~1 minute remaining" message below.
| Message | Hits (41d) | Status |
|---|---|---|
| Stage 1 — "Entered final minute" (verbal) | 28 | Current standard — fleet-wide via D2C-5403, live 2026-08-12+ |
| Stage 2 — Hard-stop silent hold (0 tokens) | 1,940 | Current standard — long-running, unaffected by D2C-5403 |
| Legacy soft warning (low-but-nonzero tokens) | 0 | Superseded — hasn't fired once in 41 days |
Old "~1 minute remaining" (contextual_update tool call, told agent not to mention tokens) | 2 | Superseded by Stage 1 above — this was the broken predecessor D2C-5403 replaced |
Of 1,940 Stage-2 hard-stop hits, 1,481 (76.3%) matched to an identity record. Of those, only 16 (1.1%) show any later call, ever, in this window — dramatically below the general baseline return rate elsewhere in this dataset (roughly 10-21% across the modality segments in the Deeper Segmentation tab). It's more concentrated than 1.1% even suggests: at least 5 of the 16 "returns" are the same single repeat user. The true distinct-user return rate is closer to ~10-12 people out of 1,481.
For Stage 1 ("entered final minute"): 21/28 identity-matched, 0 returned so far — not yet a fair comparison; the data cutoff lands within hours of the last hit, and 20 of 21 matched instances are the user's very first call ever.
Important context from product: a non-dismissible in-call banner + 2-step checkout (select amount → confirm card on file, or add a card first) already exists today. So the 1.1% figure is not measuring "no prompt exists" — it's measuring what happens after a prompt + an existing checkout flow. That reframes the open question from "should we add a re-up prompt" to "why doesn't the existing prompt-and-checkout flow convert" — see Product Roadmap for the two concrete next analyses (payment-failure data, Statsig pulse) aimed at answering that.
token_purchased event, ever, checked directly against the live BigQuery table — not just around that call, their entire history. Position matters too: 94% of the time (329/350), this message fires in the last 20% of the call's turns — it's essentially the terminal event right before the call ends, not a mid-call pause that resumes. This isn't a new problem — it's the same 1.1% conversion story surfacing independently in a totally different data cut (transcript-embedded system messages vs. the BigQuery hard-stop join above), which makes the underlying conversion failure more confident, not less.
charge_failed, payment_method_add_failed, card_validation_failed) to this same hard-stop population — confirmed to exist in BigQuery — to test whether the 1.1% ceiling is payment friction (people try to pay and it breaks) vs. decline (people see the prompt and choose not to pay). These point to very different fixes and are currently indistinguishable from transcript data alone.paywall_stall_fix_voice directly from the Statsig console API (not reconstructed from BigQuery). See the callout below — result is inconclusive, with one independent meta-finding worth acting on regardless of the metrics.paywall_stall_fix_voice gate is owned by Iryna Yakubenko, 90% pass rate in dev/staging/production (public condition, no segment targeting), unchanged since creation on 2026-07-09 — only 2 versions ever exist, both from the same minute. Statsig's own system has auto-flagged this gate type: "STALE", reason: "STALE_PROBABLY_FORGOTTEN" — independent of any metric result below, that's worth raising with Iryna on its own.
call_id → d2c_prod.voice_call_connected), but re-pulled across the full 36-agent fleet and the full 41-day window voice_call_connected covers, using corrected ElevenLabs account access (an earlier API key was inadvertently scoped to a small test/regression workspace mid-session; this section reflects the corrected pull). "Returned" = the caller's anonymous_id has a later ranked call in voice_call_connected, with the specific next conversation_id resolved wherever possible. All 1,970 instance rows (user_id, conversation_id, timestamps, per-family) are available on request for direct inspection.
Reviewed against a senior-product-manager lens. Every item below is phrased as a testable hypothesis with a named owner and success metric — not an open-ended aspiration.
| Bet | Reach | Impact | Confidence | Effort | Owner |
|---|---|---|---|---|---|
| Diagnose existing re-up/checkout conversion (payment-failure join) | ~1,481 hard-stops / 41d | High | High (data confirmed to exist) | S | Data + BE |
| Confirm legacy token-warning messages fully decommissioned | Same population | Medium (cleanup, not a live bug) | High | S | AI/ML |
| Investigate 36% zero-reply as technical failure | 36% of all bailed first-timers | Potentially high | High (data access confirmed) | S | Voice AI/BE |
✓ Done — Read paywall_stall_fix_voice Statsig Pulse in console | 3,584 gate-exposed users | Inconclusive | Result was ambiguous — see Token-Warning tab | XS | Data |
| Theme-classifier validation + no-theme deep-dive | Affects all downstream reads | High (foundational) | High | S | Data |
| Companionship-mode escalation guardrail | 13.5% of returning calls | Medium-high | Medium | S | AI/ML |
| Real cross-session memory (retrieval tool) | ~184 loyal returners | High, unproven on hard metric | Technically High (confirmed normal webhook integration) / Product-impact Low (unvalidated) | L | AI/ML + BE |
| Proactive photo delivery | Overlaps existing checkout/banner UX — scope TBD | Medium-high | Technically High / Product-impact Low | M | FE/BE/AI-ML |
| Cross-persona shared profile + discovery surface | ~30% Explorers | Medium | Low | L | BE + Design |
Biggest shift from stakeholder input: a verbal re-up prompt (D2C-5403) and an in-call banner+checkout already exist, so "build a re-up prompt" is no longer the top bet — diagnosing why the existing flow only converts 1.1% of hard-stopped callers is, and it's now Effort: S since the payment-failure data already exists. Memory and photo-delivery both move from "technically uncertain" to "technically normal, product-impact still unvalidated" now that tool/webhook feasibility is confirmed — effort estimates come down (XL→L, M-L→M) but they still shouldn't jump the queue ahead of the cheaper, better-evidenced diagnostic work above.
| Phase | Item | Owner | Hypothesis / success metric |
|---|---|---|---|
| Now | Join failed-payment/charge events to hard-stop population | Data + BE | We believe the 1.1% ceiling is driven more by payment/checkout friction than by user decline; measured by comparing hard-stop instances with a failed charge/card event vs. those without, against return rate. |
| ✓ Done | Read paywall_stall_fix_voice Pulse in Statsig console | Data | Pulled directly from the Statsig console API (2026-08-18): no significant engagement lift; two revenue-adjacent metrics show a large negative direction but fail significance and rest on small, skewed samples; test/control unit counts don't reconcile with the BigQuery exposure join; the gate itself is auto-flagged by Statsig as "probably forgotten." See Token-Warning tab for full detail — recommend confirming with the gate owner (Iryna) whether it's still needed, independent of the metric ambiguity. |
| Now | Confirm legacy token-warning messages are fully decommissioned | AI/ML | Verify no agent config still points at the old "~1 minute remaining" contextual_update message or the dead legacy soft-warning path, now that D2C-5403's 2-stage system is standard. |
| Now | Investigate 36% zero-reply cohort | Voice AI/BE | Pull VAD/STT/disconnect signals for those specific calls (data access confirmed); confirm or rule out a technical failure before any funnel-messaging fix is prioritized over an infra fix. |
| Now | Validate theme classifier + stand up cohort tracking | Data | Hand-label a held-out sample for precision/recall; start a first-call-date × days-since cohort table so Day-7/30/90 curves become measurable going forward. |
| Next | Companionship-mode intent gate | AI/ML | Detect companionship-signaling in turns 1-3 and suppress default escalation; validate against the 36.2% no-reply-at-all bail cohort. |
| Next | Cross-session memory MVP (single-persona scope) | AI/ML + BE | Standard webhook/tool integration, confirmed technically feasible. Ship retrieval scoped to one persona per loyal-returner first — validate concept with a Wizard-of-Oz test before building, since product impact is still unvalidated. |
| Next | Proactive photo delivery — pending diagnostic above | FE/BE/AI-ML | Existing banner/checkout already handles the "prompt" half of this; scope down to specifically "agent proactively offers/sends" once the Now-tier checkout diagnostic clarifies where the actual drop-off is. |
| Later | Cross-persona shared profile + discovery surface | BE + Design | Depends on single-persona memory infra existing first — explicitly sequenced after the "Next" memory MVP, not in parallel. |
Reviewed against 5 additional professional lenses (senior prompt engineer, data scientist, UX researcher, consumer/market researcher, voice-AI engineer) specifically to find what this report cannot yet claim. In the same spirit as the Data Quality tab: honest gaps, not papered over.
• No eval harness tied to the existing Langfuse regression suite (D2C-5143). Every finding here (compliance %, RAG-improvisation rate, return rate) is a one-off audit, not mapped to a standing dataset item/score — future prompt changes get re-discovered by hand instead of regression-tested. Still open.
• ✓ Resolved (2026-08-13): the token-warning "contradiction" is a deliberate 2-stage system shipped via D2C-5403 — see Token-Warning tab. Remaining item: confirm the two legacy predecessor messages are fully decommissioned fleet-wide, not just superseded in practice.
• ✓ Resolved (2026-08-13): why contextual_update was ignored — agents must take a turn immediately after the user, so a plain tool result was easy to silently skip; injecting a synthetic user turn forces a response. Confirmed deliberate design, not an unexplained workaround.
• ✓ Resolved (2026-08-13): tool-use architecture — confirmed ElevenLabs supports arbitrary custom tools/webhooks; the fleet simply hasn't built memory-retrieval, proactive photo-send, or re-up tools yet. "Give the agent real tools" is a normal integration project (webhook + backend), not a platform ceiling — Product Roadmap updated accordingly.
• Safety false-positive/false-negative tension has no proposed fix. The D2C-5143 suite shows over-triggering on age-ambiguous language and under-triggering (silently empty reasons array) on required terminations — unclear whether the lever is prompt wording or the single-tool-call architecture itself. Still open.
• New this pass (2026-08-14): silent-caller JOI narration quality — this is a live acceptance criterion on D2C-5409. A preliminary transcript pass (135 silent-modality calls) found the "let me take over" hand-off fires reliably (55% of 96 qualifying calls) and none of D2C-5410's named failure modes (aggressive edging, premature climax, loops) dominate — but only 6% of narration arcs reach a resolution cue at all. See the Deeper Segmentation tab for the full write-up. Needs a fleet-scale, human-validated re-run before treating as conclusive — see next bullet.
• "Zero tokens is a churn event" is stated causally but is correlational, with an obvious self-selection confound (hard-stop-hitters already differ from near-miss callers). Needs a matched comparison or regression-discontinuity design around the token-zero threshold before the causal framing is trusted. Partially addressable now: joining failed-payment/charge-failure events (confirmed to exist) at least separates "tried to pay and failed" from "chose not to pay," even without a full RDD.
• n=16 returns (5 from one user) is too small and non-independent for a stable rate estimate. Needs a Wilson confidence interval and per-distinct-user clustering, not a bare point estimate. Still open.
• ✓ Clarified (2026-08-13): informing users of low tokens is a required behavior, not an experimental one — confirmed by product. Any future test on the warning message must vary wording/timing only, with every user still warned; no "receives no warning" control arm. This narrows what "no experiment design exists" (below) should even be testing.
• No experiment design exists for "entered final minute" beyond an informal re-check in 1-2 weeks. Needs randomization, a pre-registered primary metric, and a power calculation against the ~1.1% base rate before any future "it worked" claim is interpretable — scoped now to wording/timing variants only, per the point above.
• The keyword theme classifier has no held-out validation. Precision/recall against a human-labeled sample is needed before the 45%/13.5%/9%/etc. breakdown justifies any build decision. Still open.
• The modality→retention table is confounded by tenure/call-volume — speaking-dominant users may simply be more tenured, not more retentive because they speak. Needs a logistic regression adjustment. Still open.
• New this pass (2026-08-14): preset-as-scaffold-to-speaking hypothesis needs a full-scale re-run. A preliminary check (34 blended users with ≥2 calls captured in this 30-day window) found no evidence that heavier preset usage precedes migration to speaking — high-preset-share users reached a later speaking call only 30.8% of the time vs. 61.9% for low-preset-share users, the opposite of what a "presets as scaffolding" theory would predict. Sub-buckets as small as 8-13 users — needs a full-history pull (not window-bounded) before this becomes a real answer either way.
• No significance testing anywhere in this report — every comparison is a raw percentage-point delta. No cohort/longitudinal tracking exists yet to build real Day-7/30/90 retention curves from a single snapshot. Still open.
• No end-to-end journey map. Findings live as disconnected stat cards per tab rather than a single mapped funnel (Discovery → First Connect → 30s Aha → Escalation branch → Session End → Token Warning → Re-up → Return/Churn) showing where interventions actually compete for the same engineering slots. Now known: the Token Warning → Re-up stage already includes a non-dismissible banner + 2-step checkout — the journey map needs to reflect real existing UI, not an assumed blank.
• "Loyal Returner"/"Explorer" are labels, not personas. No goals, frustrations, or representative quote stitched into a usable one-pager for cross-functional teams — the transcript quotes already collected (Polaroid photos, "we just talked about it five seconds ago") are sitting right there, unused for this purpose. Still open.
• Confirmed still a real gap (2026-08-13): zero usability testing exists. CSAT collection for voice calling just launched — too soon to have data. Interviews are a possible future investment, not yet planned. Every finding here remains post-hoc transcript inference until then.
• No emotional-arc tracking — duration/turns are the only engagement proxy, so a long frustrated call reads identically to a long satisfying one. Still open.
• The two riskiest recommendations (memory, proactive photo delivery) have no concept-validation step — e.g. a Wizard-of-Oz test — before being roadmapped on transcript inference alone. Still open — and now that tool feasibility is confirmed (AI/ML section above), this validation step is the actual remaining gate, not technical uncertainty.
(Reviewed via a lead-generation skill repurposed only for its "know your customer" instinct — its actual B2B sales-lead machinery doesn't apply here and wasn't forced onto this task.)
• The 44.7% "no strong theme" bucket is written off, not read. Nearly half of returning-call volume has no qualitative pass — could be genuinely thin sessions, or could be ASMR-style background companionship the keyword list simply can't see. Still open.
• No direct voice-of-customer instrument exists — no opt-in post-call micro-survey, no interview panel, nothing capturing self-reported motivation (loneliness, curiosity, habit, specific-seeking). Still open — CSAT for voice calling just launched, too soon for data (see UX section).
• Addressed with a stated working assumption (2026-08-13): no demographic/psychographic profile exists or is planned. Confirmed assumption to use instead: since the product requires a valid credit card and hosts 18+ content, treat the user base as legal-age adults with some disposable income. Not a substitute for real demographic data, but a defensible floor for now.
• Deferred by stakeholder decision, not an oversight: competitive context (Candy.ai etc.) — a competitor list may be compiled at a later date; omitted from this report for now.
• ✓ Confirmed to exist (2026-08-13): cancellation surveys, refund reasons, and support tickets are available for this product. Mining them to independently corroborate the token-churn story is now a concrete, unblocked next step rather than a hypothetical one.
• New (2026-08-17): internal theme patterns need external demand triangulation before driving roadmap decisions. A deeper keyword pass on 4 returning-call themes (power-dynamic, roleplay, photo-seeking, romantic) is based on just 19-33 calls each — a vocal minority within 266 returning calls, out of 1,081 total users. One sub-pattern (a "mommy" power-dynamic skew) is independently corroborated by GTM-1880's existing Ahrefs keyword research (~17,000 monthly searches across a "Mommy JOI" cluster) — a good template for checking the other three (cheating-wife roleplay, photo-seeking demand, romantic framing) before treating any as confirmed product direction. See Whale/Returner Archetype tab.
• ✓ Unblocked (2026-08-13): "36% zero-reply" is unfalsified as a technical failure. A user whose speech was VAD-detected but not transcribed looks identical, in a transcript, to one who never spoke. Confirmed we have access to ElevenLabs' own session/connection metadata — pulling audio-received-but-empty-transcript signals and disconnect events for those specific calls is now a concrete, actionable next step rather than a hypothetical one. See Product Roadmap.
• No turn-level agent-response-latency analysis exists. The "vague reply → call dies" pattern is attributed to prompt quality, but a slow re-engagement prompt (LLM+TTS lag) could be the actual cause — a timing problem wearing a copy-problem's clothes. Still open.
• Interruption/barge-in behavior is completely uncharacterized, despite "reciprocal escalation" being this report's own named core engagement driver — an inability to interrupt would look identical to disengagement in transcript-only analysis. Still open.
• ✓ Resolved (2026-08-13): TTS/voice quality benchmarking across persona voices — confirmed low-quality personas were already caught in internal testing and never went live, so this isn't an open gap for the current fleet. Matches the TTS-benchmark item retirement already reflected in the Product Roadmap tab.
• No connection-quality/infra layer (call-setup time, WebRTC drops) sits under any duration-based finding in a report built entirely from transcript text. Still open.
A spot-check confirmed Sophie Dee - Web DEV FOR TESTING and Francety - Web(DISABLED) are correctly excluded (neither name ends in " - Web"). But Yeyeloba - Web was genuinely missing from the initial pull — the name matched the filter fine, but the live agent-discovery step returned 34 agents instead of the true 35. Cause not fully reconstructed (the pull log only recorded a count, not names). Re-pulled and reconciled: figures in this report already reflect the corrected 35-agent, 1,745-conversation, 1,081-user dataset. The delta from the correction was under 1 point on every metric — directionally nothing changed, but it's worth knowing a silent fleet-discovery gap is possible and currently only catchable by spot-checking agent names against the account's actual list.
| Stage | Signal used | Confidence |
|---|---|---|
| Activated | Duration proxy (≥120s reached) | Proxy — the product's own engaged_30s milestone event fired for only 2 of 3,520 callers in-window, too sparse to use directly |
| Retained | Real: call_rank > 1 via voice_call_connected | Robust |
| Monetizable | Real: photo unlock / registration / subscription / token purchase events | Too sparse to segment on (2-11 hits across ~3,520 callers; 0 subscription/token matches in-window) |
I did not paper over the monetization gap with a proxy relabeled as real. Two honest possibilities, indistinguishable from here: (a) these backend activation/monetization events are under-instrumented relative to actual call volume, or (b) demo→pay conversion genuinely is this rare in a 30-day window (consistent with demo_purchase having only 12 rows in its entire history). Recommend confirming with data/product eng before anyone builds a monetization segment on top of these events.
engaged_30s / demo_photo_unlock / demo_registration_completed firing is reliable before building any segment on top of them.conversation_initiation_client_data.dynamic_variables.call_id → d2c_prod.voice_call_connected (BigQuery project stxt-490006), since ElevenLabs carries no usable native user_id on web calls. First-time/returning determined by rank of a caller's anonymous_id across their full call history, not just this window. 72% of pulled conversations matched to an identity record; the unmatched 28% could not be classified and are excluded from segmentation (but included in aggregate transcript stats where identity wasn't required).