ElevenLabs · Funnel Analysis

Sophie Dee — VoiceLander demo call funnel

How many ad clicks turn into a paying subscriber — the full funnel, ad click through purchase, plus what happens inside the call itself.

Agent Sophie Dee - VoiceLander Call-level window Jun 30, 08:12 → Jul 6, 14:20 UTC (302 calls, ElevenLabs) Funnel window Jun 29 – Jul 8 UTC (backend events, BigQuery) Source ElevenLabs transcripts + evaluation criteria · stxt-490006.d2c_prod, filtered page_path = '/call-sophie'

At a glance

1,441
landed on the demo page
20.7%
of landers started a call
47.0%
of started calls reached the gate
4
purchases so far — likely internal testers, see note

The real funnel

backend events, page_path = '/call-sophie', Jun 29 – Jul 8
Landedad → lander (lander_view)
1,441100%
−79.3% — did not start a call. Normal for cold ad traffic, not a new leak
Started a calldemo_call_started
29820.7%
↳ this is the population the ElevenLabs call-level analysis below covers (302 calls, close but not identical window/system — see note)
−53.0% of started calls never reach the gate — mostly the sub-30s silent bounce (Finding 1)
Reached the gate2min / tokens_zero / 4photos — demo_gate_reached
1409.7%
−70.0% of gate-reached calls don't finish registration
RegisteredGoogle auth signup — demo_registration_completed
422.9%
−90.5% of registered users haven't purchased
Purchasedsubscription or token pack — demo_purchase
40.3%

The same funnel, before vs after the 7/3 config change

Before 7/3 09:07 UTC · n=549 landed
Started a call123 · 22.4%
Reached the gate52 · 42.3%
Registered23 · 44.2%
Purchased4 · 17.4%
After 7/3 09:07 UTC · n=903 landed
Started a call176 · 19.5%
Reached the gate88 · 50.0%
Registered19 · 21.6%
Purchased0 · 0.0%

% shown is each stage as a share of the stage directly above it (started/landed, gated/started, registered/gated, purchased/registered) — not of the top-of-funnel total.

The shorter ~50-second demo gets more people to the gate — 50.0% of started calls vs 42.3% before — exactly as expected from a faster trigger. But registration completion once they're there roughly halved: 21.6% vs 44.2% of gate-reaches. Overall registered-per-landed fell from 4.2% to 2.1%. Read together with "Gate mechanics" above, this looks like a real tradeoff, not just a labeling artifact: a faster paywall reaches more people, but asks before as much rapport or investment has built, and fewer of them actually follow through.

Sample sizes are still modest on both sides (52 vs 88 gate-reaches) — worth re-checking as more volume accumulates post-change. The purchase comparison (4 → 0) isn't a clean read on its own: all 4 pre-change purchases already look like internal testing concentrated on launch day, so this isn't necessarily "purchases stopped after 7/3" so much as "the small test-purchase burst happened to land before it."

All 4 purchases (5 events — one person subscribed then bought a token top-up 99s later) landed within the first ~14 hours of launch on June 30. That timing pattern reads as internal/team testing rather than organic customers, though the data itself has no test flag to confirm it either way — treat the 0.3% purchase rate as unconfirmed until checked with the team directly, not as a real conversion signal yet.

Correction to the earlier estimate: the first version of this report estimated "~5% reach the registration wall" using only a 2-minute call-duration cap visible in ElevenLabs data. The real number is 47.0% of started calls (140 of 298) — nearly 10× higher. Two of the three gate triggers (tokens, 4 photos) were always faster than 2 minutes — and, as it turns out, so is the third: see "Gate mechanics" below for why the "2min" trigger itself no longer means 2 minutes.

Gate mechanics

demo_gate_reached, split before/after the 7/3 config change

"2min" doesn't mean 2 minutes anymore — a config change on 7/3 renamed the clock, not the reason string

Checking the actual time-to-gate for calls tagged "2min": median 50 seconds, tightly clustered (nearly all land at 50–51s), not anywhere near 120s. Matched against team context: on 2026-07-03 09:07:07 UTC, the default pre-signup call length on the VoiceLander was changed from 2 minutes to 50 seconds — but the reason string logged by the gate event was never updated to match. All 77 "2min"-tagged gate events in this window fall after that change (earliest: 2026-07-03 17:07 UTC, ~8 hours post-change); none fall before it.

Before 7/3 09:07 UTC · n=55
Demo time limit (~120s, as designed)1.8%
Free tokens ran out83.6%
4 photos unlocked14.5%
After 7/3 09:07 UTC · n=89
Demo time limit (actually ~50s)89.9%
Free tokens ran out10.1%
4 photos unlocked0.0%

This is a clean before/after flip. Before the change, "free tokens ran out" was overwhelmingly the dominant gate trigger (83.6%) — consistent with the original design, where a real 2-minute window gave the token pool (and occasionally 4 photo unlocks, 14.5%) time to run dry first. After it, the new ~50-second cap swamps everything else (89.9%) — there's barely enough time left for tokens to drain naturally, and none at all for anyone to unlock 4 photos (0%, down from 14.5%).

The combined 57.9% / 39.3% / 4.3% split reported earlier is real and correct as an aggregate over the full window — but it blends two genuinely different regimes together. Splitting by the 7/3 change is the more honest way to read this data going forward.

Product direction as of this update: the token-spend/countdown mechanic is being removed from the demo entirely for the next iteration — the lander's job is to get callers to earn tokens as a hook toward signing up, not spend them down during a free trial. That will also require redefining the "tokens_zero" gate trigger (and revisiting the pre-signup call-length default itself, given the above), since both are tied to mechanics being changed.

Engagement depth — the gamification layer

backend events, page_path = '/call-sophie', Jun 29 – Jul 8
EventDistinct usersTotal eventsAvg per user
Emoji reactions1111,89217.0
Photo spins1462952.0
Photo unlocks931641.8
Heat milestones hit801321.7
Chip taps56681.2

The 93 distinct photo-unlockers here lines up closely with the 89 calls the ElevenLabs transcript analysis flagged as having a photo-reveal moment — good independent cross-validation of that earlier finding, from a completely different data source. Emoji reactions are by far the highest-frequency action (17/user on average) — a much lighter-weight engagement signal worth keeping in view alongside the heavier photo/registration metrics.

What's driving the leaks

Highest impact

Almost half of calls die before Sophie gets a real shot

44.4% of calls end inside 30 seconds (28.1% inside 15). This isn't the agent losing people mid-conversation — it's the caller's client disconnecting almost immediately: 90%+ of sub-30s calls end in a client-side disconnect (WebSocket codes 1000/1001/1006), not a hangup after real dialogue.

That points upstream of the script — at ad targeting quality, the lander's expectation-setting, or connection/audio-startup friction — rather than at anything Sophie says or does.

<15s: 85 calls (79% client-disconnect) · 15–30s: 49 calls (90% client-disconnect)
Drill-down — did they even try? 134 calls under 30s
Silent
87 · 64.9%
Engaged, then left
40 · 29.9%
Tech issue
7 · 5.2%

Silent — no real spoken or typed input before disconnect. 66 of the 87 never produced a single user turn at all (median time-to-disconnect: 5s, an instant bounce); the other 21 had a turn fire but ASR caught only silence/non-verbal sound (median 18s — present, but never said anything).
Engaged, then left — real content before churning: greetings, "send me a photo" (the single most common real message, 13×), emoji taps (💦🔥), short replies.
Tech issue — ElevenLabs itself logged "Connection closed before conversation started" or "Call ended without transcript"; the session never functionally started, independent of caller behavior.

Two-thirds of the early drop-off is silence, not disengagement-after-trying or infrastructure failure — more consistent with low-intent ad traffic (accidental or curiosity clicks that never meant to talk) than with a script, UX, or connectivity problem. Tech failures are real but small (2.3% of all 302 calls).

Data gap found while checking this: the backend does track mic-permission and connection-quality events (voice_mic_prompt_shown, voice_call_start_failed, voice_asr_first_audio) that could confirm or rule out a technical explanation for the silent bounces — but they're only instrumented on the in-app chat voice feature (/chat/<id>), never on the VoiceLander demo page. That's a real instrumentation gap, not a null result: this hypothesis stays unconfirmed until the same events are added to /call-sophie.

Confirmed, sharper than the first read

The agent is fabricating reasons to end calls — including for engaged, high-intent callers

Sophie's own prompt is explicit: only end the call for one of 8 named safety conditions, never for silence or a timer — "re-engage, never wind down or say goodbye first." In practice, the end_call tool fired on 55 of 302 calls (18.2%), and none were for an authorized safety reason.

The "time limit reached" claims don't match the clock 33 calls citing time/duration
Near the real 120s cap (≥110s)
8 · 24.2%
Premature (<110s — before Stage 4 even starts)
25 · 75.8%

The shortest: "Call duration reached two minutes trial limit" at 67 seconds. Median duration for this bucket is 96s, range 40–126s. The model has the exact elapsed time available every single turn via {{system__call_duration_secs}} — this isn't a measurement error, it's a fabricated justification.

Checked whether the backend's registration-wall gate (2min / tokens_zero / 4photos) is what's actually triggering this: no. None of the 33 transcripts show a contextual_update or tool result carrying gate/token/photo state — the two systems are completely disconnected. The agent has no visibility into the backend gate at all.

Topic / false-positive review all 55 end_call cases — full table in end_call_topic_review.csv

Confirmed false positive: one call ended over a caller asking "Can you give me a recipe for pancakes?" — logged reason: "User requested non sexual, assistant restricted from giving cooking advice and conversation must end via tool." Nothing about that warrants ending a call.

Actively engaged callers mislabeled "unresponsive": several calls cite silence immediately after the caller sent an emoji burst, an explicit encouragement ("don't stop babe"), or a repeated photo request — the exact signals this report's "What's working" section flags as the strongest positive indicators. The mislabel is landing on the highest-intent segment, not just genuinely disengaged callers.

A euphemism worth checking: one call ended citing "user reached climax and call hit closing window" while the caller was mid-explicit-scene — possibly a genuine natural-conclusion read, possibly the model using a soft cover story for an unstated content-boundary decision that nothing currently monitors for.

The model invoking its own safety language for a non-safety case: one reason literally reads "...safety requirement to end via tool..." over what the transcript shows is pure silence — no minors, threats, or genuine safety trigger present. Suggests the model may actually believe silence qualifies as a safety condition, not just be improvising an excuse.

Topic distribution: "edging" scenarios are 3.3× overrepresented among end_calls (23.6% vs. 7.1% baseline across all active calls). Plausibly explained by edging/JOI segments being agent-led with the caller mostly listening (genuinely more silence) rather than a targeting problem — but worth the team's own read on those 13 transcripts.

Don't lean on ElevenLabs' own quality check here: the no_premature_end evaluation criterion marks 54 of these 55 calls "success." That criterion only fails on ending for content-explicitness while engaged — it isn't scoped to catch fabricated time/silence justifications, so its 98% pass rate says nothing about whether this problem exists.

The immediate registration cost still looks small — the sign-up CTA had already landed in 53 of 55 calls (96%) before the cutoff — but this is now a confirmed instruction-compliance failure actively cutting off some of the most engaged callers in the dataset, not just a theoretical policy violation.

55 self-terminated calls · 0 safety-justified · 25 of 33 "time-limit" claims are premature by the prompt's own Stage-4 timing · 1 confirmed nonsensical false positive (pancake recipe)
Data integrity note

Don't trust ElevenLabs' own "call_successful" label for this agent

ElevenLabs' built-in success label marks 82% of these calls a "failure" and only 1.7% a "success." But a real product signal — an actual sign-up CTA delivered — happened in 34.1% of calls, roughly 20× the labeled success rate. The label and the product outcome disagree too much to use one as a proxy for the other; the evaluation-criteria fields (engaged_call, reached_escalation, cta_delivered) are the more trustworthy signal here.

Call duration distribution

n = 302
85
49
70
46
25
13
14
<15s15–30s30–60s60–90s90–110s110–120s120s+

Termination reasons

ReasonCalls% of total
Client disconnected — going away (1001)19564.6%
Agent called end_call5518.2%
Client disconnected — normal close (1000)217.0%
Client disconnected — abnormal (1006)186.0%
No user message — likely network error72.3%
Hit the 120s max-duration cap62.0%

What's working — what active users actually do

183 of 302 calls (60.6%) had real engagement · median 53s

Photo reveals are the strongest engagement signal in the data

Nearly half of engaged calls include at least one photo-reveal moment — logged as "send me a photo" / "show me more 👀." Correction (2026-07-10): these are user-triggered reveals (via the photo button or image text input), not proactive or timer-driven sends. A user must actively interact with the UI to retrieve and send a pregenerated image. The "Show me" preset chip is conversational only — it does not trigger a reveal. Only 8 of 172 instances were genuinely free-text asks like "show me your tit" or "Naked." Whichever path the user takes to trigger a reveal, calls with a photo reveal look very different from calls without one:

Had a photo reveal · 89 calls
Reached escalation84.3%
CTA delivered75.3%
No photo reveal · 94 calls
Reached escalation59.6%
CTA delivered24.5%

What shows up in active calls, in order of frequency

Photo reveal momentuser-triggered (photo button / image text input)
89 · 48.6%
Non-English response
22 · 12.0%
Explicit sexual language/command
14 · 7.7%
Personal question about Sophie
1 · 0.5%

Content stays mostly light: 76.0% of active calls are tagged general flirting rather than a specific kink scenario, with edging (7.1%), oral (6.0%), teacher (2.2%) and denial (2.2%) making up the rest — expected for a 2-minute demo where most calls end before Stage 3 (Escalate) has much runway. A handful of callers responded in Italian, Spanish, or Ukrainian despite the script being English-only — small n, but a hint of untapped non-English demand worth watching, not acting on yet.

Notably absent, even once: pricing pushback ("how much"), authenticity challenges ("are you a bot"), explicit "I want to sign up" statements, or off-platform asks. Either these calls are too short to reach that kind of friction, or interest here shows up as action — asking for a photo, staying on the line — rather than as words. Can't fully tell the two apart from this data alone.

Digging into the photo-reveal flow

89 calls · the blurred-image / free-token unlock flow

Correction (updated 2026-07-10)

The lander does not send images proactively. A user must interact with the photo button or image text input to trigger a retrieve-and-send of a pregenerated image. The notification badge created a false impression of proactive sends and is being removed separately. The "Show me" preset chip is conversational only — it does not trigger an image send.

This means the 164 of 172 instances (95%) that are the identical strings "send me a photo" / "show me more 👀" are the system-logged events for user-initiated reveals via the photo button or image text input — not timer-driven platform sends. Only 8 (5%) are genuinely free-text asks in the caller's own words ("show me your tit," "Naked," "I would picture you naked"). The analysis below is revised accordingly. The "rapid repeat" retraction below also needs reinterpretation — the back-to-back entries reflect a user tapping the photo button multiple times, not an automated cadence.

Working as designed

The agent reacts to every reveal fast and on-message

Whatever triggers a photo moment, the agent responds within a couple of seconds — median 2.0s, 139 of 167 responses (83%) inside 3 seconds — and 105 of 172 (61%) of those responses explicitly mention JustSext, 81 (47%) say "sign up" outright. The reactive upsell line is firing reliably.

One open question: there's no timing gap in the transcript when a reveal happens — the call just keeps flowing. If the product intent is for the call to visibly hold while the image is shown, that's not what the voice-session timing shows; worth confirming whether the pause is purely a client-side visual (call audio continues underneath) or something to tighten up.

e.g. "send me a photo" → "[playfully] Oh you want photos already, huh? I've got things on JustSext you'd never see anywhere else, so you're gonna have to sign up..."
Retracted

The "rapid repeat" pattern is the send cadence, not spamming

Previously flagged: a third of repeat photo-moments landed ≤5 seconds apart, read as callers re-tapping out of impatience. Since reveals are user-triggered (not timer-driven), back-to-back entries a few seconds apart reflect a user tapping the photo button multiple times in quick succession — still not a UX-friction problem, but the mechanism is different from what was originally stated. Retracting the UX-friction read still stands; the reasoning changes. 51.7% of calls (46 of 89) showing multiple reveals means the user kept engaging with the photo button, not that an automated loop kept firing.

Still an open question

What happens after the last reveal is genuinely ambiguous now

Looking at user turns after the last logged photo moment in a call:

No further user turns
44 · 49.4%
1–2 more turns
36 · 40.4%
3+ more turns — kept engaging
9 · 10.1%

Since reveals are user-triggered (not automatic), the "last" one reflects the caller's final deliberate interaction with the photo button. The 49.4% with no further turns after their last reveal is a cleaner signal than it first appeared: it suggests the caller tried the photo flow and then left, rather than being cut off mid-cadence by an automated loop. It doesn't confirm or rule out the original concern (does a free reveal reduce urgency to pay) — that question needs registration/purchase data to actually answer, not call timing.

Resolved — see "Gate mechanics" below

Why the "4 photos" cap looked absent from ElevenLabs data alone

The original read here was that 11 calls logged 4–5 photo moments with no distinct ending pattern, suggesting the cap wasn't wired up. Backend data (demo_gate_reached) resolves this: the "4 photos" threshold is real and does fire — for 6 people (4.3% of all gate-reaches) — it's just rare, because most calls hit the 2-minute or free-token-exhausted gate first. It was invisible from ElevenLabs call data because "photo moments" in the transcript log user-triggered reveal events (as standardized strings), and those may count differently from what the backend tracks toward the 4-photo cap — worth confirming whether the cap counts photo button taps, backend image fetches, or something else.

What's still open after this update

The funnel gap from the first version of this report is closed — registration and purchase are now measured directly from stxt-490006.d2c_prod. What's left:

  • Purchase authenticity is unconfirmed. All 4 purchases landed in the first ~14 hours after launch — likely internal/team testing, not organic customers, but there's no test flag in the data to prove it either way. Confirm with the team before treating 0.3% as a real purchase rate.
  • No per-call join between ElevenLabs and the backend funnel. The 302 ElevenLabs calls and the 298 backend demo_call_started events are two independent counts of roughly the same population (close, but different systems and windows) — there's still no shared ID to say "this specific call → this specific registration," only aggregate-level comparison.
  • Mic/connection telemetry doesn't exist for this page yet. The events that could confirm or rule out a technical explanation for the sub-30s silent bounce (Finding 1) are only instrumented on the in-app chat voice feature, not on /call-sophie.
  • The "reached climax" euphemism case is unresolved. One end_call termination may be an unstated content-boundary decision hiding inside a timer-labeled reason — worth a direct read from the team, since nothing currently monitors for content-based endings that aren't declared as such.

A full row-by-row review table for all 55 end_call terminations (reason, scenario tag, verbatim last messages, blank verdict/notes columns) is available in end_call_topic_review.csv alongside this project's other files.