After launching a CRM campaign, a company can usually see:
- how many messages were sent
- how many customers received them
- how many people opened the email
- how many clicked the link
- how many recipients made a purchase
- how much revenue those purchases generated.
These metrics help explain what happened after the communication. But they do not answer the main business question:
What share of the result appeared specifically because of CRM automation?
If a customer received an email and placed an order two days later, the two events are related in time. But temporal proximity does not prove that the email caused the purchase.
The customer may have:
- already planned to place the order
- purchased regularly every month
- seen an advertisement in another channel
- arrived through a recommendation
- participated in a general promotion
- received a call from a manager
- simply made a purchase during the period the system treats as the attribution window.
Revenue from recipients and incremental revenue are therefore different metrics.
To measure the real effect of CRM automation, you need to evaluate not only what happened after the campaign, but also what would have happened to the same customers without it.
Causal analytics is built around exactly this question.
1. The core problem: we cannot observe two futures at the same time
Imagine one customer.
They meet the conditions for a CRM scenario aimed at developing a second purchase. We send an offer, and five days later they place an order.
Ideally, we would like to compare two outcomes:
- What happened after the message was sent.
- What would have happened to the same customer during the same period if the message had not been sent.
But we can observe only one version of reality.
If the customer received the message, we can no longer observe their behavior under exactly the same circumstances without the message. If they did not receive it, we cannot know for certain how they would have responded.
The unobserved outcome is called the counterfactual.
The causal effect for one customer can be represented as the difference:
outcome under treatment − outcome without treatment.
At the individual level, this difference is usually unavailable. Analytics therefore estimates the average effect for comparable groups of customers.
2. Attribution and causal effect are not the same thing
CRM systems often use attribution.
Attribution links a purchase to a preceding communication according to a defined rule:
- last communication before purchase
- last open
- last click
- purchase within 24 or 48 hours
- first touch
- distributing the result across several touches.
For example:
The customer opened the email and made a purchase within 48 hours—the purchase is attributed to the email campaign.
This is a useful operational rule. It helps:
- allocate revenue across channels
- build dashboards
- compare campaigns
- analyze the customer journey
- see purchases after communications.
But attribution does not create a counterfactual outcome. It does not show whether the customer would have purchased without the message.
Bloomreach, for example, uses a last-touch model in its standard reports: revenue is assigned to the last interaction within the selected attribution window. This is a model for assigning outcomes to touches, not an experimental estimate of causal effect. Bloomreach: Revenue attribution
It is therefore correct to say:
KZT 20 million in revenue was attributed to the campaign under the selected model.
But without a control group, it is incorrect to claim:
The campaign created KZT 20 million in incremental revenue.
Part of that amount may have appeared even without the campaign.
3. Three levels of CRM campaign evaluation
It is useful to separate results into three levels.
Technical level
Shows whether the channel worked:
- messages were generated
- the send took place
- messages were delivered
- no critical errors occurred
- links work
- personalization renders correctly.
Behavioral level
Shows what customers did after the communication:
- opened messages
- clicked
- visited the website
- viewed products
- added products to the cart
- made purchases.
Causal business level
Shows how much behavior changed specifically because of the treatment:
- incremental conversion
- incremental buyers
- incremental orders
- incremental revenue
- incremental margin
- change in purchase frequency
- change in retention
- change in LTV.
All three levels are necessary, but they answer different questions.
A high open rate does not guarantee incremental revenue. A low click rate does not always mean there was no business effect. And high revenue among recipients may simply reflect that the campaign initially targeted the most active customers.
4. Why comparing recipients and non-recipients is often misleading
Suppose we compare:
- customers whose message was delivered
- customers whose message was not delivered.
Conversion is higher in the first group.
Can we treat the difference as the campaign effect?
Not necessarily.
Undelivered messages are more common among customers:
- with outdated contact details
- with low activity
- with long-unused accounts
- with invalid addresses
- without a confirmed channel.
This group may already have purchased less often before the campaign.
It is even more dangerous to compare:
- people who opened the email
- people who did not open it.
Opening happens after the communication has been assigned. Customers decide for themselves whether to open the message, and that choice is related to their interest in the brand.
People who open emails may already be:
- more active
- more loyal
- more interested
- more frequent website visitors
- more likely to purchase.
If they purchase more often, you cannot tell what caused it: the email or their original level of interest.
Groups used to estimate effect should therefore be formed before treatment, not after customers react to it.
5. Controlled experiment
The most reliable way to measure the effect of CRM automation is to randomly split eligible customers into groups.
Test group
Receives the treatment being studied:
- a message
- a sequence of messages
- a personalized offer
- a new channel
- modified scenario logic.
Control group
Does not receive the treatment, or continues to receive the previous version when an existing process is being changed.
The groups’ subsequent behavior is then compared.
Random assignment is needed so that, before the experiment begins, the groups are comparable on average in terms of:
- purchase activity
- recency of last purchase
- average order value
- categories
- country
- channel availability
- propensity to purchase
- other observed and unobserved characteristics.
If the groups are large enough and randomization is performed correctly, the systematic difference between them becomes the treatment being studied.
That is why the difference in outcomes can be interpreted as the average campaign effect, subject to statistical uncertainty.
Controlled A/B experiments are used to estimate the causal impact of changes: users are randomly assigned to control and test groups, after which predefined metrics are compared. Microsoft Research: A/B Testing Across Products
6. Define the audience first, then split it
It is very important to place the experiment correctly within the CRM scenario.
Shared eligibility conditions should be checked first:
- the customer belongs to the business audience
- has the required contact detail
- has provided the required consent
- is not a test profile
- is not under a global exclusion
- is eligible to receive the offer
- has not completed the target action before the experiment starts.
Only then should eligible customers be randomly split between test and control groups.
Otherwise, the groups may become incomparable.
For example:
- All customers are split into test and control.
- In the test branch, customers without email are excluded before sending.
- No such exclusion is applied in the control branch.
The test group then contains only reachable customers, while the control group contains both reachable and unreachable customers. These are already different audiences.
Bloomreach likewise recommends placing common conditions before the A/B split and explicitly excluding customers without the required contact detail or consent from both groups. Bloomreach: A/B testing
The correct sequence is:
build the eligible audience → verify treatment availability → random assignment → treatment or control → outcome measurement.
7. What the control group actually receives
Control does not always mean “no communications at all.”
It depends on the research question.
No new scenario
If the goal is to understand whether a new automation creates incremental impact:
- the test group receives the scenario
- the control group does not
- all other processes remain the same.
Current version vs. new version
If an existing communication is already part of a required process:
- the control group receives the current version
- the test group receives the new version
- the effect of the change is measured.
This can be used to test:
- new logic
- new personalization
- another channel
- a changed interval
- a new assortment algorithm.
Multiple variants
You can compare:
- control with no treatment
- variant A
- variant B
- variant C.
But adding variants divides the audience across more groups. A reliable conclusion will require more data or a longer experiment.
Persistent holdout group
Sometimes a share of customers is excluded from a class of communications for a longer period.
This allows the business to measure the combined effect of the entire automation system rather than one campaign.
For example:
- 95% of the audience participates in the CRM program
- 5% remains in a persistent holdout
- after several months, purchase frequency, retention, revenue, and LTV are compared.
A persistent holdout provides a more holistic estimate, but it has a commercial cost: some customers are deliberately withheld from potentially useful treatments. The size and duration of the group should therefore be justified.
8. Unit of randomization
Before an experiment, you need to define what exactly is randomly assigned between groups.
It may be:
- customer
- account
- household
- company
- store
- region
- device
- order.
For CRM, the unit is most often the customer.
But there are exceptions.
If several profiles belong to the same person, that person may end up in both test and control groups. If members of one household share an account or influence one another’s purchases, individual-level randomization can also violate independence.
Choosing a larger unit—for example, a store or region—reduces the risk of interference, but also reduces the effective number of independent observations.
One hundred thousand customers across one hundred stores are not always one hundred thousand independent units if treatment is assigned at the store level.
The randomization unit should therefore match the level at which:
- treatment is assigned
- contamination or interference can occur
- the outcome is measured
- the business decision is made.
9. Define the hypothesis and metrics before launch
The hypothesis should be recorded before the experiment begins.
A good hypothesis contains:
- audience
- treatment
- expected outcome
- period
- primary metric.
For example:
For customers who made one completed purchase and did not make a second within 14 days, a personalized CRM scenario will increase conversion to a second purchase over the following 30 days compared with no scenario.
The metrics are then defined.
Primary metric
One metric used for the main decision:
- conversion
- number of buyers
- frequency
- retention
- revenue per customer
- gross profit per customer.
Secondary metrics
Help explain the mechanism of the effect:
- visits
- views
- add-to-cart actions
- time to purchase
- number of categories
- average order value.
Guardrail metrics
Show undesirable consequences:
- unsubscribes
- complaints
- returns
- discount cost
- margin decline
- increased communication load
- reduced performance of other channels.
Microsoft’s guidance on trustworthy experimentation similarly starts with a simple, testable hypothesis and success metrics chosen in advance. Microsoft Research: Patterns of Trustworthy Experimentation
10. Measurement window
You need to define the period after treatment assignment during which the outcome is counted.
For example:
- purchase within 7 days
- second purchase within 30 days
- retention after 90 days
- revenue over three months.
A window that is too short may miss delayed effects. A window that is too long increases the influence of other factors:
- subsequent campaigns
- seasonality
- promotions
- price changes
- new products
- natural purchases.
The window should match both the scenario mechanics and the natural behavior cycle.
For a consumable product, the effect of a reminder may appear several weeks later. For an abandoned cart, the relevant window is usually shorter.
It is important not to choose the window after looking at the results. If you test 1, 3, 7, 14, 30, and 60 days and then report only the period with the most favorable difference, the probability of a chance “success” increases substantially.
The primary window should be fixed in advance. Additional periods may be explored, but they should be labeled as secondary analysis.
11. Uplift and increment
Consider an experiment:
- test group: 10,000 customers
- control group: 10,000 customers
- 800 customers purchased in the test group
- 500 purchased in the control group.
Test-group conversion:
Control-group conversion:
Absolute uplift
Absolute uplift shows the difference between the probabilities of the target action.
Relative uplift
Relative to the control level, conversion increased by 60%.
This does not mean an increase of 60 percentage points. The formulations must be distinguished:
- growth from 5% to 8%
- an increase of 3 percentage points
- a relative increase of 60%.
Incremental buyers
If the test group would have behaved like the control group in the absence of the campaign, we would expect:
In reality, 800 customers purchased.
The estimated incremental number of buyers is:
So:
- 800 customers purchased in the test group
- roughly 500 of them might have purchased without the campaign
- the estimated increment is 300 additional buyers.
If the validated scenario is later scaled to 100,000 comparable customers, the expected increment would be about 3,000 buyers:
Such extrapolation is valid only if the new audience, period, and conditions are comparable to the experiment.
12. Incremental revenue
Number of buyers is not the only outcome. A campaign may affect:
- purchase probability
- number of orders
- basket composition
- average order value
- discount level
- category
- purchase timing.
Revenue is therefore better evaluated directly per assigned customer.
Suppose:
| Metric | Test | Control |
|---|---|---|
| Customers | 10,000 | 10,000 |
| Buyers | 800 | 500 |
| Revenue | KZT 24m | KZT 14m |
| Revenue per assigned customer | KZT 2,400 | KZT 1,400 |
Difference:
Estimated incremental revenue in the test group:
This calculation captures both the change in conversion and the change in purchase value.
You should not automatically multiply the number of incremental buyers by the average order value of the entire test group. Customers whose purchase was actually caused by the campaign may have a different average order value from customers who would have purchased anyway.
The difference in revenue per randomly assigned customer is a more direct estimate.
13. Revenue is not the same as profit
Even positive incremental revenue does not guarantee economic efficiency.
You also need to account for:
- cost of goods
- gross margin
- discounts
- promo codes
- channel cost
- bonus cost
- returns
- operating costs
- impact on future purchases.
Example:
- incremental revenue: KZT 10m
- cost of incremental goods: KZT 6m
- discounts: KZT 1.5m
- communication cost: KZT 0.2m
- additional operating costs: KZT 0.3m.
Estimated incremental financial result:
For discount-heavy campaigns and expensive channels, it is therefore preferable to measure:
- contribution profit per customer
- contribution margin
- return on investment
- cost per incremental purchase
- cost per incremental buyer.
14. Purchase acceleration vs. creating a new purchase
A CRM scenario may not create an additional order; it may simply bring it forward.
For example:
- without the message, the customer would have purchased after 20 days
- after the message, the customer purchased after 7 days
- over a 14-day horizon, the campaign looks successful
- over a 30-day horizon, the number of purchases in the groups converges.
This is called a pull-forward effect.
Acceleration can still be valuable to the business:
- cash is received earlier
- churn risk is reduced
- the customer moves to the next stage sooner
- inventory turnover improves.
But it should not be presented as an incremental purchase.
To distinguish these effects, analyze several horizons:
- short-term
- primary
- extended.
If uplift is high in the first seven days but disappears by 30–60 days, the scenario is likely accelerating purchases. If the difference persists, there is stronger evidence of incremental volume.
15. Cannibalization and demand redistribution
Growth in a target category does not always mean growth for the business as a whole.
A scenario may:
- increase sales in category A
- simultaneously reduce sales in category B
- move a customer from a higher-margin product to a lower-margin one
- shift a purchase from offline to online
- replace a full-price order with a discounted order
- change basket composition without changing total spend.
For assortment scenarios, you therefore need to measure:
- sales of the target product
- sales of the entire category
- sales of neighboring categories
- total revenue
- margin
- number of unique categories
- total customer purchase volume.
Otherwise, a locally successful campaign may simply redistribute existing demand.
16. Statistical uncertainty
Even with random assignment, the groups will not produce exactly identical outcomes by chance.
If you flip a coin 100 times, you do not necessarily get exactly 50 heads and 50 tails. In the same way, control and test groups can differ randomly.
The effect estimate must therefore be interpreted together with uncertainty.
Confidence interval
A confidence interval shows a range of values compatible with the observed data under the chosen statistical model.
For example:
Estimated uplift: +3 percentage points; confidence interval: +1.2 to +4.8 percentage points.
This means the data supports a positive effect, but its exact size is uncertain.
If the interval includes zero:
−0.5 to +2.4 percentage points,
the data is insufficient to reliably distinguish a positive effect from no effect or a small negative effect.
Statistical and business significance
A statistically significant change may be too small to matter commercially.
For example:
- uplift: +0.05 percentage points
- the result is statistically reliable because the audience is enormous
- implementation cost exceeds the economic benefit.
Conversely:
- observed uplift: +5 percentage points
- the audience is small
- the interval is very wide
- the result is promising but needs another test.
A decision should therefore consider:
- direction of effect
- size
- uncertainty
- economic value
- implementation cost
- risks.
Bloomreach likewise recommends evaluating A/B test results together with statistical significance and avoiding decisions when the sample size is insufficient. Bloomreach: A/B test basic evaluation
17. Sample size and minimum detectable effect
Before the experiment, estimate how many customers will be required.
Sample size depends on:
- baseline conversion
- the minimum effect worth detecting
- acceptable risk of a false conclusion
- required statistical power
- test/control allocation ratio
- variability of the metric.
The lower the baseline conversion and the smaller the expected effect, the larger the audience required.
If 0.2% of customers purchase, distinguishing 0.2% from 0.22% is much harder than distinguishing 10% from 15%.
Minimum detectable effect
Before the test, it is useful to define:
What is the smallest change that would be valuable enough to implement the scenario?
If the business only cares about an increase of at least 1 percentage point, the experiment should have enough power to detect that change.
Otherwise, the test may end with an inconclusive result simply because the audience was too small.
In experimentation practice, sample size is linked to the baseline metric, significance level, required power, and minimum detectable effect. Microsoft Research: Beyond Power Analysis
18. Why you should not stop a test at the first positive result
If you check results every day and stop the experiment the moment they become positive, you increase the probability of mistaking random fluctuation for a real effect.
For example:
- on day one, the test leads
- on day two, the control leads
- on day three, the difference becomes significant
- on day five, it disappears
- after two weeks, it stabilizes near zero.
The experiment should run for a predefined period and cover complete business cycles:
- weekdays and weekends
- salary days
- regular promotions
- typical seasonality
- expected response period.
Bloomreach also recommends defining A/B test duration in advance and covering at least one, preferably two, complete business cycles. Bloomreach: A/B testing
If the experiment must be stopped early for technical or ethical reasons, that decision should be documented and the analysis should account for the changed design.
19. Checking the split
Before analyzing the outcome, verify that the experiment was actually executed as planned.
Group ratio
If the configured split is 50/50, actual group sizes should be close to that ratio.
A substantial deviation is called Sample Ratio Mismatch.
It may indicate:
- a randomization error
- different filters after the split
- lost events
- duplicate profiles
- different branch processing
- a technical failure.
If an unexplained Sample Ratio Mismatch exists, estimating effect is risky. Microsoft treats this check as one of the core controls in experimentation. Microsoft Research: Sample Ratio Mismatch
Comparability of baseline characteristics
Pre-treatment metrics are also checked:
- historical purchase frequency
- revenue
- average order value
- purchase recency
- categories
- countries
- channel availability.
Small random differences are acceptable. A systematic imbalance may indicate a problem in group formation.
20. Analyze by assigned group, not only by delivery
The main experiment analysis is usually based on the original assignment:
- all customers assigned to test remain in test
- all customers assigned to control remain in control.
Even if some people in the test group:
- did not receive the message
- did not open it
- encountered a delivery error.
This approach estimates the effect of the real policy:
What happens if we launch this scenario for eligible customers?
If you retain only recipients or openers in the test group, random assignment is broken.
You can separately diagnose:
- delivery
- view/open
- click
- errors.
But an analysis “among openers only” answers a different question and requires additional assumptions.
21. Overlap with other campaigns
During the experiment, customers continue to live within the broader CRM system.
They may receive:
- regular newsletters
- service messages
- promotions
- push notifications
- advertising
- other automated scenarios.
This does not always invalidate the experiment. If other treatments are distributed approximately equally between test and control, their impact tends to balance out on average.
The problem arises if the campaign being studied changes the probability of other treatments or if scenarios directly conflict.
For example:
- the test group receives a discount and is excluded from another promotion
- the control group continues to receive the promotion
- one scenario blocks another
- different offers affect the same category.
In that case, you are measuring not only the effect of the message, but the effect of the entire resulting combination.
Overlaps should be:
- recorded
- monitored
- excluded when there is an explicit conflict
- considered in interpretation
- incorporated into a broader experiment when necessary.
Microsoft research shows that simultaneous A/B tests can often coexist, but when changes are logically connected, interactions between experiments become important and require isolation or a special design. Microsoft Research: A/B Interactions
22. Analytics technical specification
Reliable evaluation starts before the campaign launches.
First, analyze:
- CRM project configuration
- business goal
- data structure
- events
- KPIs
- target actions
- experimental design
- available history periods.
A technical specification for data extraction and processing is then prepared.
What should be documented
Experiment
- identifier
- name
- start and end dates
- hypothesis
- unit of randomization
- group proportions
- scenario version.
Assignment
- customer identifier
- group
- assignment time
- eligibility conditions
- scenario parameters
- country and segment.
Treatment
- send attempt
- delivery
- channel
- time
- message version
- technical status
- reason for suppression or error.
Outcomes
- target event
- time
- order identifier
- status
- amount
- currency
- products
- categories
- returns and cancellations.
Historical metrics
- purchases before the experiment
- revenue
- frequency
- average order value
- lifecycle stage
- segment membership.
Workflow
- Formalize the business question.
- Define metrics and the measurement window.
- Design test and control groups.
- Agree on events and identifiers.
- Prepare technical requirements for the SQL extract.
- Analysts extract the required data.
- Validate completeness and correctness.
- Remove technical duplicates.
- Restore customer group assignment.
- Calculate metrics.
- Estimate uncertainty.
- Analyze segments and guardrail metrics.
- Form conclusions and recommendations.
- Modify the CRM scenario based on the result.
23. Data validation before calculation
Before analysis, check:
- whether each customer is unique within the experiment
- whether anyone entered multiple variants
- whether the actual split matches the configured split
- whether control-group events were lost
- whether dates are correct
- whether the time zone is applied correctly
- whether orders are duplicated
- whether canceled and test transactions are excluded
- whether a purchase is defined identically for both groups
- whether the scenario changed during the experiment
- whether the measurement window was respected.
You cannot clean the groups differently.
For example, removing high-value buyers only from the test group destroys the experiment.
Outlier handling should use a rule that was defined in advance or applied symmetrically:
- capping a metric
- winsorization
- robust metric
- separate analysis
- exclusion of technical anomalies.
Data quality is a fundamental condition for trustworthy A/B analysis: correct randomization cannot compensate for lost events or incorrectly calculated metrics. Microsoft Research: Data Quality for Trustworthy A/B Testing
24. When randomized control is not possible
A company cannot always create a control group.
There may be several reasons:
- the communication is mandatory
- the audience is too small
- the change has already been launched
- the decision applies to every region
- it is impossible to exclude part of the customer base
- the experiment was not planned historically.
In this case, observational methods are used.
They may improve the estimate, but they require additional assumptions. In terms of reliability, they are usually weaker than a correctly run randomized experiment.
Before-and-after comparison
Compare the metric:
- before launch
- after launch.
For example:
- before the campaign, conversion was 5%
- after it, conversion was 8%
- observed increase: 3 percentage points.
But at the same time, the following may also have changed:
- season
- prices
- assortment
- advertising
- economy
- competitors
- composition of the customer base.
A simple pre/post analysis therefore shows change over time, but does not isolate the cause.
Matched control group
For campaign participants, similar customers who did not receive the campaign are selected based on:
- past frequency
- purchase amount
- recency
- categories
- country
- activity.
This reduces differences on known characteristics, but does not control for factors absent from the data.
Difference-in-Differences
The method compares changes rather than levels in two groups.
Example:
| Group | Before | After | Change |
|---|---|---|---|
| Test | 5% | 8% | +3 percentage points |
| Control | 4% | 5% | +1 percentage point |
Estimated effect:
Or equivalently:
The key assumption is that, without the campaign, the groups would have evolved in parallel. The existence of pre-launch data alone does not guarantee this. You need to examine prior trends and explain why the control represents an appropriate counterfactual trend.
Interrupted time series
The metric trajectory before and after the intervention is analyzed:
- did the level change
- did the trend change
- does the observed value differ from the forecast.
The method requires a sufficiently long and stable history and is sensitive to other events occurring at the same time.
Synthetic control and Causal Impact
A counterfactual forecast is constructed from time series that are related to the target metric but were not exposed to the treatment.
Bayesian structural time-series models can account for trend, seasonality, and control variables. But the result depends on an important assumption: the relationship between control series and the target metric would have remained stable after the intervention, and the control series themselves were not affected by the campaign. Brodersen et al.: Inferring causal impact using Bayesian structural time-series models
Observational methods do not turn history into a true experiment. They estimate effect under specific assumptions that should be stated explicitly and checked as far as possible.
25. Uplift modeling
A standard predictive model answers the question:
Who is highly likely to make a purchase?
But customers with a high purchase probability may buy even without communication.
An uplift model tries to answer a different question:
Whose purchase probability changes the most specifically because of the treatment?
Conceptually, customers can be divided into four types.
Persuadables
They purchase after treatment but would have been less likely to purchase without it. This is the primary target group.
Sure things / natural buyers
They purchase regardless of the communication. Treatment may be unnecessary for them.
Non-responders
They do not purchase either with or without treatment.
Customers with a negative response
They might have purchased without the communication, but the treatment reduces the probability—for example, by causing irritation, distrust, or an unnecessary discount.
A standard propensity model often finds natural buyers. An uplift model tries to find customers whose behavior changes because of treatment.
But training such a model requires data with a reliable treatment/control split. If the model is trained only on recipients, it cannot separate purchase propensity from communication effect.
Uplift modeling differs from ordinary prediction precisely because it estimates the change in response caused by treatment, rather than simply the probability of the target event. Zhao et al.: Uplift Modeling for Multiple Treatments
26. Effect by segment
Average uplift can hide meaningful differences.
For example:
| Segment | Test | Control | Uplift |
|---|---|---|---|
| New customers | 9% | 5% | +4 pp |
| Regular customers | 12% | 11% | +1 pp |
| Churn risk | 3% | 4% | −1 pp |
The campaign may be positive on average but negative for the churn-risk segment.
It is useful to explore effect by:
- countries
- languages
- lifecycle stages
- first-purchase categories
- activity
- frequency
- price segments
- available channels.
But the more cuts you examine, the higher the probability of finding an attractive result by chance.
Key segments should therefore be defined in advance. Unexpected differences found after the test can be used to generate new hypotheses, which should then be confirmed in a separate experiment.
27. Long-term effect
A short-term campaign may have delayed consequences.
Positive effects:
- habit formation
- transition to regular purchasing
- introduction to a new category
- higher retention
- higher LTV.
Negative effects:
- communication fatigue
- higher unsubscribe rates
- expectation of discounts
- lower margin
- cannibalization of future demand
- shifting a purchase forward without changing total volume.
A mature measurement system therefore includes different horizons:
- immediate response
- primary conversion window
- repeat purchase
- retention
- cumulative revenue and margin
- LTV.
Long-term evaluation is especially important for CRM scenarios whose goal is not one transaction, but a change in the customer’s purchase trajectory.
28. How to interpret the result
After analysis, there are more possible outcomes than simply “success” and “failure.”
Reliable positive effect
The scenario can be scaled while monitoring continues.
Positive estimate with high uncertainty
A larger or longer test is needed.
Practically zero effect
The scenario does not create enough incremental value. It can be disabled or redesigned.
Negative effect
Check:
- implementation quality
- audience
- offer
- frequency
- channel
- timing
- conflicts with other scenarios.
Conversion improves while economics deteriorates
The discount or assortment mechanic needs to change.
Effect exists only in selected segments
The scenario can be narrowed and retested on the relevant audience.
Effect is short-term and disappears later
The campaign is likely accelerating purchases rather than creating incremental ones.
An experiment result is not a grade for the team. A negative or zero effect also creates value: the company stops spending resources on a mechanic that does not work and gains evidence for the next hypothesis.
29. CRM automation development cycle
A properly organized process looks like this:
- Define the business problem.
- Formulate the hypothesis.
- Design the CRM scenario.
- Define metrics and the control group.
- Configure events and identifiers.
- Technically test the scenario.
- Launch the experiment.
- Validate data quality.
- Calculate causal effect.
- Analyze segments and economics.
- Make a decision.
- Modify scenario logic.
- Launch a new test.
CRM then develops not on the subjective belief that a message “should work,” but through the systematic accumulation of validated knowledge.
Common mistakes in CRM campaign evaluation
Treating all recipient revenue as campaign impact
Some customers would have purchased without treatment.
Comparing openers and non-openers
These groups form themselves after the send and differ in activity from the start.
Building control from undelivered messages
Customers with undelivered messages often differ systematically from reachable customers.
Splitting groups before checking common eligibility
Control and test end up with different eligibility rules.
Changing allocation ratios during the experiment
Previously assigned customers retain their participation history, which can distort the estimate.
Stopping the test at the first positive result
A random fluctuation is mistaken for a stable effect.
Testing dozens of metrics and reporting only the successful one
The more comparisons you make, the higher the chance of a false discovery.
Ignoring cancellations and returns
A technically created order is treated as a completed business outcome.
Analyzing only a short-term window
Purchase acceleration is mistakenly presented as an incremental purchase.
Ignoring discounts and margin
Revenue growth may be accompanied by lower profit.
Applying a complex causal method without validating assumptions
The name of a method does not automatically make an estimate causal.
When to contact us
You can contact us if you need to:
- develop a methodology for evaluating CRM scenarios
- design test and control groups
- define KPIs and target actions
- configure A/B splits in Bloomreach Engagement
- design events and parameters for analytics
- prepare technical requirements for SQL extracts
- validate experimental data quality
- calculate uplift and increment
- estimate incremental revenue and margin
- distinguish attribution from causal effect
- analyze a campaign without a ready-made control group
- apply Difference-in-Differences or time-series models
- study effect by segment
- estimate long-term impact
- identify cannibalization
- build a regular CRM analytics system
- modify scenarios based on measurement results.
We can join either the full evaluation cycle or a specific stage: experiment design, preparation of requirements, processing data extracts, metric calculation, or interpretation of results.
What the business gets
The outcome is not a report with open rate and recipient revenue, but an evidence-based CRM management system:
- formalized hypotheses
- correctly designed control groups
- agreed KPIs
- technical data requirements
- reproducible calculations
- estimates of statistical uncertainty
- calculation of incremental impact
- understanding of scenario economics
- conclusions for individual segments
- recommendations for automation development.
Such a system helps the company answer the questions that actually matter:
- does the scenario create new purchases or merely accompany natural ones
- which customers change behavior because of the communication
- which offers work
- where a discount is genuinely necessary
- which channel creates incremental impact
- which automations should be scaled
- which should be changed or disabled
- how CRM affects retention and long-term customer value.
The main task of CRM analytics is not to attribute as much revenue as possible to a campaign.
Its task is to determine as accurately as possible what incremental value the automation created, how reliable that estimate is, and what decision the business should make based on it.