PerformanceJune 18, 20269 min

Cohort analysis 2026: how to calculate retention and LTV by cohort

How to build and read cohort analysis in 2026: retention curve, breakdown by month of attraction and channel, calculation of LTV by cohort. Where to get data in Ya.Metrica and Google Sheets and what budget decisions this changes.

Article cover:Cohort analysis 2026: how to calculate retention and LTV by cohort

Every couple of weeks they send me a dashboard where everything is green: revenue is growing, leads are getting cheaper, ROAS is more than 4. And a quarter later the business wonders where the money went. Almost always the reason is the same - they looked at one-time ROAS and hospital averages, and not at cohorts. Cohort analysis is the only way to see that a client purchased in March does not return by June, and all growth comes only from new traffic. Below is how to build such a report, how to read the retention curve and why the channel with the cheapest lead is sometimes the most unprofitable.

I have been counting cohorts on projects since 2018. The retention and LTV numbers in the article are ranges and medians for my projects in e-com, gaming and services, plus reconciliation with a couple of colleagues. Not a personalized case accurate to the ruble, but guidelines that show the logic. Your niche will always shift the boundaries, but the shape of the curve and the conclusions across the channels hold.

One-time ROAS answers the question “did advertising pay off this month.” Cohort analysis answers the question “whether the client will pay off at all.” These are different questions, and the second one is more important.

1. What is a cohort and why is it needed?

A cohort is a group of clients with a common event in one time window. This is usually the month of first purchase or registration. Everyone who first bought in March 2026 is the “March 2026” cohort. Then they monitor her: how much was returned in April, May, June, how much money was brought in.

Why put this on top of regular reports. A regular report gives a snapshot by date. Revenue for June grew by 20% - excellent. Only this figure contains two different effects: new clients came and old ones returned. If there are a lot of new ones, but the old ones don’t come back, growth rests only on the budget. They stopped pouring - the proceeds fell away. The cohort separates these two effects into different lines, and it is immediately clear where exactly the business stands.

The cohort can be cut not only by month. By acquisition channel, by first product, by average entry receipt, by promo code. But the month of the first purchase is the basic axis, they always start with it.

2. Retention curve and how to read it

The retention curve is the proportion of cohort clients who remain active 1, 2, 3 months after entry. M0 is equal to 100% by definition: this is the month of the first purchase itself. Then the curve falls.

What is considered “activity” - decide on the shore. For e-com this is a repeat purchase. For subscription - renewal. For a content product - entry and action. The main thing is not to confuse a visit and a purchase: returning to the site but not purchasing does not mean revenue retention.

A healthy curve drops steeply in the first one or two months and then plateaus. The plateau is the main thing I look for with my eyes. If the curve, after a sharp drop, rises to 18-25% and holds, there is a core of loyal people who hold the product. If the decline continues to zero without a plateau, clients do not return, and paid traffic flows into a leaky bucket. No matter how much you pour on top, the bottom will not be filled.

This is what the cohort retention table looks like by month. Rows - entry cohorts, columns - how many percent returned in the Nth month. The numbers are a typical picture for the average e-com in my projects.

CohortSizeM0M1M2M3M4M5M6
Jan-26820100%34%24%20%19%19%18%
Feb-26910100%31%22%19%18%17%—
Mar-261240100%27%18%15%14%——
Apr-261060100%33%23%20%———
May-26980100%35%25%————

You can read such a table in two movements. Horizontally - the shape of the curve of one cohort: a plateau of 18-20% is visible. Vertical - comparison of cohorts at the same month of life. And here the March line comes out: M1 has 27% against 31-35% for its neighbors, M3 - 15% against 20%. March brought more people, 1240 against a thousand, but they are returning worse. The next question for investigation is: what did we do in March? Most often, the answer is to scale a weak channel or run a sale that brings in discount hunters.

This diagonal void in the lower-right corner is the norm. The fresh cohorts have not yet reached M4-M6, and there is simply nothing to fill the percentages there. It is impossible to compare immature cohorts by M3+, the data are not mature.

3. Calculation of LTV by cohorts

Retention shows who has returned. LTV shows how much they brought in. It is considered simple and honest, without predictive models. Take the cohort by month of entry, add up all the revenue it generated during the entire observation period, and divide by the size of the cohort.

LTV of the cohort (cumulatively by month N) =
   Amount of cohort revenue for months M0..MN / Cohort size

Example: cohort Jan-26, 820 clients
M0 brought 1,230,000 ₽ -> LTV(M0) = 1,500 ₽
2,540,000 ₽ accumulated for M3 -> LTV(M3) = 3,100 ₽
3,200,000 ₽ accumulated for M6 -> LTV(M6) = 3,900 ₽

The value is that LTV grows over time, and the cohort shows exactly how. The first check gave 1500 ₽, by six months 3900 ₽ had accumulated. This means that the real LTV is 2.6 times higher than the first purchase. If you compare this LTV with CAC, the payback window can be seen directly in the line: in which month the accumulated revenue covered the cost of acquisition. How to correlate these two numbers - I’ll figure it out in article about unit economy, and the guidelines by the ratio itself are in LTV/CAC benchmarks in Russia.

The main mistake here is that the average LTV across the entire database is one number. The distribution is always skewed: a small share of customers accounts for half of the revenue. The average of such a distribution does not describe anyone. Cohorts cut this skew according to the time of entry, and then you can also cut along the channel.

4. Breakdown by channel: why cheap CPL is deceiving

The most expensive lesson in cohort analysis. CPL measures the input. LTV measures the entire life of a customer. These are different points on the timeline, and optimization for the first regularly kills the second.

The picture I see all the time. Channel A gives leads for 200 ₽, and in the weekly report he is a hero. Channel B costs 500 ₽, they want to cut it like an expensive one. We collect cohorts by channel - and it turns out that channel A has an M3 retention of 8%, and a repeat check is cheap. Channel B has an M3 retention of 24%, and people are coming back for more. On the six-month horizon, a client from B brings three times more than from A, with an entry price of only 2.5 times higher. It was necessary to cut exactly the opposite way.

A cheap lead is often cheap because it is a bargain hunter, a random click, or an untargeted but wide audience. He bought it once and disappeared. An expensive lead often comes from a warm channel, where a person came consciously and stayed. This is not visible in any way from the one-time CPL and even from the one-time ROAS; the entry figure is equally silent about the future. This is precisely why I look at CPL and ROAS only in conjunction with cohorts, and analyzed the formulas themselves and their pitfalls separately in material about calculating ROAS and CPL.

Here are the same channels in a cohort context. The numbers are ranges for my projects, not just one case.

ChannelCPLRetention M1Retention M3LTV M6Conclusion
Search, brand350–600 ₽38–46%24–30%4500–6500 ₽Best for a long time
Search, general500–900 ₽28–34%16–22%2800–4200 ₽Worker, average
YAN/networks150–300 ₽18–26%8–14%1400–2400 ₽Cheap entry, weak base
Target, sale120–250 ₽12–20%5–10%900–1700 ₽Bargain Hunters
Bloggers / sponsored placements400–800 ₽30–40%18–26%3500–5500 ₽Depends heavily on the blogger

Looking at the YAN and target lines, you can see exactly what we’re talking about: entry is cheap, and LTV is two to three times lower than search. This does not mean “turn off the networks.” This means counting them according to fact, and not according to CPL, and keeping the share under control. Cheap channels have a role: gain volume, test creatives, warm up the top of the funnel. You just can’t compare them to searching using a single entry number.

5. Where to get data

Three sources, and each closes its part.

Yandex Metrica. It has a ready-made “Cohort Analysis” report: you set the condition for forming a cohort, the target retention action and the step. The metric will quickly show the retention of visits and goal achievements. The weak point is money. Metrica is good at measuring behavior, but it doesn’t know very well how much the client brought in during the year, especially if payments are made offline or through a manager. How to set up the goals themselves on which this all rests, I looked into guide to the goals of Ya.Metrica.

CRM. amoCRM, Bitrix24 or whatever you have - this is where revenue per client comes from over time. Three fields are needed: the date of the first purchase, all subsequent transactions with amounts, and the acquisition channel by UTM. Without a channel in CRM, the breakdown from the fourth section will not be collected, so UTM must reach the transaction, and not get lost on the form.

Google Sheets or BI on top of uploads. The cohort matrix itself is collected in a pivot table: rows are cohorts, columns are months of life, in cells is the formula for the share of those who returned or accumulated revenue. For small volumes, Sheets with a couple of formulas are enough. For larger ones - Metabase, DataLens or Power BI, so as not to rebuild by hand. I described in the review how this place fits into the general contour of the measurement. marketing analytics.

6. How cohorts change budget decisions

One-time ROAS pushes you to a simple action: go where the number is higher this month. The logic is clear and regularly leads to a wall. A channel with a high monthly ROAS can bring in one-time buyers, and after six months you are growing faster and faster, and the margin is melting.

Cohorts unfold logic. The budget goes not to where the lead is cheaper, but to where the accumulated LTV by month of payback is higher. Decisions become different in essence. A channel with an expensive entry point but a strong base gets more money, not less. The cheap channel is not cut out, but limited to its share and kept as a source of volume and tests. Sales mechanics are judged not by a one-time surge in revenue, but by what cohort they brought in and whether they will return.

The rhythm itself changes. One-time ROAS forces the budget to be pulled weekly due to noise. The cohort view translates into a monthly rhythm: wait until the new cohort shows M1-M2, and only then move money. Less fuss, less retraining of campaigns, stronger solutions.

The antithesis to which it all comes down. One-time ROAS determines whether the advertisement paid for itself in a month. The cohort is responsible for whether the client will pay for themselves in their lifetime. The budget collected on the second issue almost always stands more firmly on the ground.

7. Frequent errors during assembly

I see the same rake from project to project.

  1. Cohorts of 10-20 people. Percentages on such sizes fluctuate due to randomness. Minimum 50-100 per cohort, otherwise increase to a quarter.
  2. Activity is counted by visit, not by purchase. Retention of visits is cheerful, retention of revenue is sad - they should not be confused.
  3. Immature cohorts are compared across distant months. The fresh cohort has not yet made it to M5, there is nothing to fill it with, a diagonal void is the norm, not data.
  4. The average LTV is considered one number for the entire database. The distribution is skewed, the average does not describe anyone. Only by cohorts and segments.
  5. UTM does not reach the transaction in CRM. Without a channel in the client line, the breakdown by source will not be collected, and this is the most valuable part.
  6. The report was collected once and forgotten. The cohort lives in time, it must be replenished every month, otherwise it becomes obsolete faster than it is useful.

Conclusion

Cohort analysis is not a fashionable dashboard, but a way to stop deceiving yourself with average and one-time ROAS. It shows the shape of the retention curve, the real LTV by month of entry and, most importantly, the truth about channels: that the cheapest lead is often the most expensive in the long term. The minimum set is an upload from CRM with UTM, a Ya.Metrica report and a summary report in Google Sheets. Once a month, reassemble, look at the budget review, move money to channels with a strong base. This is all mechanics. In 9 years in digital, I have not seen a tool that changes the quality of budget solutions cheaper.

If your dashboard is green, but it feels like money is flowing out, you are almost certainly looking at a cross-section and not at cohorts. Write to me at Telegram or through form: at the starting call we will collect your first cohort table using real data and find the channel that you are overestimating or cutting in vain. Starting consultation - 0 ₽.

More on the topic