Dhruv’s approach
FRAMING Two people matter. The free member who wants to express themselves in chat, and the paying member who bought Nitro to do more of it. Discord earns from Nitro, not from chat volume, so chat is a leading signal and subscriptions are the money. The tension: anything that makes free chatting better can eat the reason to pay. Engagement and monetisation are different metrics and they can move in opposite directions, which is exactly what we are looking at. I will not touch pricing. First I need to know whether the drop is real, and whether the test caused it.
- VALIDATE THE TEST
- Is -10% significant, or noise? The test was powered for chat interactions. Subscriptions are rarer and slower, so a secondary metric like this is usually underpowered.
- Sample ratio mismatch: are the two groups the same size and comparable on pre-period behaviour?
- Randomisation unit. Discord is social. Split by user and treatment and control sit in the same servers, affecting each other. Server-level randomisation is probably the right unit.
- Run length. Has it run past a full billing cycle? A 10% drop over two weeks may be delayed purchases, not lost ones.
- Novelty. Is the 15% chat lift decaying week on week?
- Definition. Is -10% a count of subscriptions or a conversion rate? Gross new, or net of cancellations?
- BREAK DOWN NITRO Subscriptions = new + renewals - cancellations. Find which one moved. Then split new: exposed users x upsell impressions x click rate x checkout completion. One fork decides everything:
- Impressions down, conversion flat means the test moved, buried or replaced an upsell surface. A placement problem.
- Impressions flat, conversion down means the test gave away part of what people were paying for. A value problem. Two different problems with two different responses, so I would run this split before anything else.
- SEGMENT
- Free members against existing Nitro holders. Free users fail to convert, Nitro users cancel. Different failures.
- Users who saw the upsell against users who did not.
- Heavy against light chatters. If heavy chatters drove the +15% and are also the ones not converting, that points at cannibalisation.
- New against tenured accounts.
- Platform, since purchase flows differ on desktop, mobile and web.
- Server size and type.
- Weeks since exposure, to see novelty decay.
- HYPOTHESES, ranked by likelihood x impact / cost to check H1 Cannibalisation. The feature handed free users something Nitro used to gate. Confirm: impressions flat, conversion down, concentrated in heavy chatters and in the specific benefit the test touches. Kill: conversion per impression flat. H2 Upsell surface displaced. The new chat UI moved or buried the entry point. Confirm: impressions down, conversion per impression flat. Kill: impressions flat. H3 The result is not real. Confirm: wide confidence interval, underpowered for this metric, inconsistent across weeks. Kill: tight interval, stable week to week. H4 Spillover between groups. Control users share servers with treated users, so control is not a clean baseline. Confirm: effect size varies with the share of treated members in a server. Kill: no dose response. H5 Mix shift. The feature activated lower-intent users, diluting the rate while absolute subscriptions hold. Confirm: count flat, rate down. Kill: count down too.
THE CALL Do not ship to everyone. Hold the current allocation, extend past one billing cycle, and get the impressions against conversion split first. Chat interactions should not be the only primary metric. A chat lift that costs subscriptions is only worth it if that engagement turns into retention and later revenue, so I want revenue per exposed user alongside it. Counter-metrics: Nitro conversion rate and cancellations. If it is placement, fix it and re-test. If it is cannibalisation, that is a decision about what Nitro is worth paying for, not a decision about this test.