Cohort Analysis for Beginners: The One Chart That Explains Your Business

The single chart worth building

There’s one report that tells you more about the health of a subscription business than almost anything else you could pull, and most solo founders never build it. Not because it’s hard, but because nobody ever showed them what it actually looks like. It’s called a cohort analysis, and I want to walk through what it is, how to build a simple version yourself, and what it reveals that a single overall number never can.

I build a version of this for my own small businesses every few months, usually in nothing more exciting than a spreadsheet. I’m not pitching you a piece of analytics software here, because you genuinely don’t need one to get real value out of this. What I want to hand you is the idea itself, the shape of the table, and a worked example so it stops being an abstract phrase and becomes something you could sketch out this afternoon with your own numbers.

What a cohort actually is

A cohort is simply a group of customers who all started at the same time. The customers who signed up in January are one cohort. The customers who signed up in February are a different cohort. That’s the entire idea. Once you’ve grouped your customers this way, you can watch each group separately over time and see how it behaves, rather than blending every customer together regardless of when they joined.

Why one blended number hides the story

Here’s why that grouping matters so much. If you just look at your overall retention or churn number this month, you’re mixing customers who joined three years ago with customers who joined three weeks ago, and averaging across all of them into one figure. That blended number can stay perfectly steady even while something important is changing underneath it, like a newer batch of customers leaving much faster than your older, loyal base, quietly dragging the future down while the current snapshot still looks fine. A single number this month tells you where you are. Cohorts tell you where you’re heading.

The classic retention table

The standard way to see this is a simple table, sometimes called a cohort triangle. Each row is a cohort, the group that started in a given month. Each column is a number of months since they started: month one, month two, month three, and so on. Inside each cell you put the percentage of that original group still active at that point. Read down a column and you compare different cohorts at the same age. Read across a row and you watch one cohort age over time. The shape that emerges, once you fill in enough of the table, tells you far more than any single churn percentage ever could on its own.

A worked example

Say twenty customers join in January. One month later, eighteen of them are still around: ninety percent retained. Two months later, sixteen remain: eighty percent. Three months later, fifteen remain: seventy five percent, and from there it barely moves, staying near seventy five percent for months four, five, and six.

Now compare that to your March cohort, also twenty customers, but by month three only sixty percent remain, and it keeps sliding instead of leveling off. Those two cohorts started the exact same size, and the overall churn number for the business in any given month might look identical either way. The table is what actually shows you that something changed between January and March, and that March customers are behaving differently.

What healthy and unhealthy curves look like

The shape you’re hoping to see is a curve that drops for the first month or two, as the people who were never really a fit for your product self-select out, and then flattens. A flattening curve means the customers who survive the early period tend to stick around for a long time, which is a genuinely strong sign.

An unhealthy curve keeps sliding downward with no flattening in sight, month after month. That tells you people are leaving throughout the whole relationship, not just in an early shakeout period, and that’s a much harder problem to fix because it points at ongoing dissatisfaction rather than a one-time mismatch.

Revenue cohorts, not just headcount

Everything so far counted people, but you can build the exact same table counting dollars instead. A revenue cohort tracks what percentage of the original group’s revenue survives each month, rather than what percentage of people survive. This version can look meaningfully better than the people-based version if the customers who leave tend to be your smallest accounts, or meaningfully worse if the customers who leave tend to be your biggest ones. Building both versions side by side tells you a fuller story than either one alone.

Testing whether a change actually helped

One of the most practical uses of cohorts is testing whether something you changed actually worked. Say you rebuilt your onboarding in April. Compare the cohort that joined in March, under the old onboarding, against the cohort that joined in April, under the new one, at the same age, say month two for both. If the April cohort is retaining noticeably better than March at that same point, you have real evidence the change helped, evidence that a single overall churn number, blending old and new customers together, would have hidden or badly delayed.

Cohorts by acquisition channel

You can slice cohorts a second way too, not just by when someone joined, but by where they came from. Build a retention table for customers who arrived through a paid ad campaign, and a separate one for customers who arrived through referrals or organic search. It’s common to find one channel bringing in customers who stick around far longer than another, even if both channels cost roughly the same to acquire from. That difference should directly shape where you put your limited marketing effort, and a blended overall number would never have shown it to you.

Say your paid ad campaign brings in a cohort of thirty customers in a given month, and by month six, forty percent of them remain. Your organic search cohort from that same month brings in eighteen customers, and by month six, seventy percent remain. The ad channel might have looked cheaper per signup at the start, but if the organic customers are worth nearly twice as much by the time you account for how long they actually stay, the true cost per surviving customer tells a very different story than the cost per signup ever did on its own.

Watching cumulative revenue per cohort

There’s one more version of this worth knowing: cumulative revenue per cohort. Instead of tracking a percentage retained each month, you track the total dollars that cohort has paid you so far, added up month after month. A cohort can be losing customers steadily and still show rising cumulative revenue for a while, if the customers who remain are spending more over time through upgrades or add-ons. Watching this number tells you roughly how many months it takes a typical cohort to pay back whatever it cost you to acquire them, which links this whole exercise directly back to how much you can afford to spend on acquisition in the first place. It’s a natural companion to a customer lifetime value calculation, built from the same billing records, just organized by cohort instead of averaged across your whole customer base.

Cohorts and the numbers you already track

If you’ve already been watching your overall churn rate or your monthly recurring revenue, cohorts are the missing piece that ties those together. Your monthly churn number tells you the size of the leak this month. Your MRR waterfall tells you whether new signups are outrunning it. Cohorts tell you whether that leak is getting worse or better for each fresh group of customers you bring in, which is usually the earliest possible warning that something has changed, well before it shows up clearly in either of those other two numbers on their own.

How many customers before this means anything

A fair question is how small is too small for this to be useful. With only a handful of customers per cohort, one or two cancellations swing the percentage wildly, and you risk reading a real pattern into what’s actually just noise from small numbers. I wouldn’t trust a monthly cohort with fewer than roughly twenty or thirty customers in it. If your monthly volume is smaller than that, group two or three months together into a single, larger cohort instead, so the percentages have enough people behind them to mean something.

Common mistakes

A few mistakes come up again and again. Comparing cohorts of very different sizes as if the percentages were equally trustworthy. Building so many narrow cohorts that each one is too small to read, when a few wider ones would tell the story more clearly. And forgetting seasonality, comparing a cohort that joined during a slow month against one that joined during a naturally busier period, and mistaking a seasonal effect for a real change in your business.

How to actually build this yourself

You don’t need dedicated software to start. A spreadsheet with one row per signup month and one column per month of age, filled in by hand or with a simple formula pulling from your billing records, gets you most of the value immediately. Start with retention counts, add the revenue version once the first table feels natural, and update it monthly rather than building it once and forgetting about it. Block out twenty minutes on the same day every month to refill the table, because a cohort analysis that only gets built once, months ago, isn’t much more useful than the single overall number it was supposed to replace.

What cohorts cannot tell you

And here’s the honest limit, same as with any of these numbers. A cohort table shows you what happened. It doesn’t automatically tell you why. If a cohort is sliding, that’s your cue to go talk to the customers who left, to check what changed in your product or your marketing around the time they joined, to look for the real explanation underneath the shape of the curve. Treat the table as the start of an investigation, not the end of one.

Recap

A cohort is just a group of customers who started at the same time, and tracking each group separately reveals trends a single blended number hides completely. Build the classic retention table, watch for a curve that flattens rather than one that keeps sliding, and try the same table for revenue, for different acquisition channels, and for before and after any real change you make. Keep your cohorts big enough to trust, and remember the table shows you what happened, never automatically why.

If you want the rest of the metric library in the same plain-English style, with more worked examples than fit in one video, you can find it on the Data Research Analysis Collection home page.

Get new guides and videos first — join the Telegram channel.