Missing data is one of the most common problems in data analysis and machine learning.
A dataset may contain empty cells because a value was never recorded, a user skipped a question, a sensor failed, or a particular process did not produce an observation.
At first glance, handling missing data might seem simple.
You could replace missing values with:
- The mean
- The median
- The mode
- A constant
- The previous observation
- A model-generated estimate
But there is a deeper question that should come before choosing an imputation method:
Why is the data missing in the first place?
This question matters because not all missing data behaves in the same way.
Statisticians commonly describe missing-data mechanisms using three categories:
- Missing Completely At Random (MCAR)
- Missing At Random (MAR)
- Missing Not At Random (MNAR)
MCAR occurs when missingness is unrelated to both observed and unobserved information.
MAR occurs when missingness can be explained by information that is already observed.
MNAR is more difficult.
With Missing Not At Random, the probability that a value is missing is related to the missing value itself or to information that has not been observed.
That makes MNAR particularly challenging because the missingness mechanism cannot generally be determined from the observed data alone.
What Does Missing Not At Random Mean?
Missing Not At Random, or MNAR, describes a situation where the reason a value is missing is related to the value that is missing or to other unobserved information.
Consider a survey asking people:
“What is your annual income?”
Suppose people with very high incomes are less likely to answer the question.
The missing income values are therefore not random.
People with higher incomes may be more likely to leave the question blank.
In this case:
Income
↓
Probability of responding
The missingness depends on the value that is missing.
That is a classic MNAR situation.
Another example could be a medical study where patients experiencing more severe symptoms are less likely to complete follow-up assessments.
The missing measurements may be related to the patient’s unobserved health condition.
The important point is:
The missingness contains information about the missing value.
A Simple MNAR Example
Imagine a company collects employee salary information.
The dataset looks like this:
| Employee | Department | Salary |
|---|---|---|
| A | Sales | $45,000 |
| B | Marketing | $52,000 |
| C | Engineering | Missing |
| D | Finance | $58,000 |
| E | Engineering | Missing |
Suppose the company discovers that employees with very high salaries are less willing to disclose their salaries.
The missing values are therefore not randomly distributed.
If the missing engineering salaries are actually:
$120,000
$135,000
replacing them with the overall median could seriously underestimate the true salary distribution.
For example:
Observed median = $52,000
Actual missing salaries = $120,000 and $135,000
Median imputation would make the dataset look more complete, but it would not solve the underlying missingness problem.
The problem is not simply:
“Which value should replace the blank?”
The real problem is:
“Why are these particular values missing?”
MCAR vs MAR vs MNAR
Understanding MNAR becomes much easier when you compare it with MCAR and MAR.
Missing Completely At Random
MCAR means the probability that a value is missing is unrelated to both observed and unobserved variables.
For example, suppose a computer system randomly loses some records because of an unexpected technical failure.
If the failures have nothing to do with:
- Customer characteristics
- Transaction values
- Product type
- Date
- Location
then the missingness may be considered MCAR.
Conceptually:
Observed data ──┐
├──> Missingness
Unobserved data ┘
Neither observed nor unobserved values systematically determine whether data is missing.
Missing At Random
MAR is more subtle.
Under MAR, missingness may depend on information that has already been observed.
Suppose a survey asks for annual income.
You discover that younger respondents are less likely to provide their income.
If age is observed for everyone, the missingness in income may be explained by age.
For example:
Age
↓
Probability of income response
The income itself does not necessarily determine whether the person responds after accounting for observed information.
This can allow certain statistical methods to handle the missingness under an MAR assumption.
Missing Not At Random
With MNAR, missingness depends on the missing value itself or other unobserved information.
For example:
Income
↓
Probability of income response
People with very high incomes might be less likely to report their income.
Because the missing income values are unknown, the observed data alone may not be enough to identify the missingness mechanism.
The Three Mechanisms Compared
| Mechanism | What influences missingness? | Example |
|---|---|---|
| MCAR | Neither observed nor unobserved values | Random system failure |
| MAR | Observed information | Younger people skip income question |
| MNAR | Missing value or unobserved information | High-income people avoid reporting income |
The distinction is important because the assumptions behind your statistical analysis can change depending on the missing-data mechanism.
Why MNAR Is Difficult
MNAR is difficult because the information needed to explain the missingness may itself be missing.
Consider a healthcare dataset.
Suppose patients with more severe symptoms are less likely to attend follow-up appointments.
You observe:
Patient A → Follow-up completed
Patient B → Follow-up completed
Patient C → Missing
Patient D → Missing
You know that follow-up attendance is incomplete.
But you don’t directly observe the severity of the patients who did not attend.
The very information that could explain the missingness is unavailable.
This creates a difficult problem:
Unobserved condition
↓
Probability of missing data
↓
Missing observation
You cannot simply look at the missing records and determine why they are missing.
MNAR Can Create Biased Results
Ignoring MNAR can introduce substantial bias into an analysis.
Suppose a customer satisfaction survey is sent to 10,000 customers.
You receive 6,000 responses.
At first glance, you might calculate:
Average satisfaction = 8.2 / 10
But suppose highly dissatisfied customers are much less likely to complete the survey.
The 6,000 responses may therefore disproportionately represent satisfied customers.
The observed average could be:
8.2
while the true average satisfaction across all customers might be considerably lower.
The problem is not simply that 4,000 values are missing.
The problem is that the missingness is related to the outcome being measured.
MNAR in Surveys
Surveys are one of the clearest examples of MNAR.
Imagine a survey asking:
“How much debt do you currently have?”
People with very high levels of debt may be more likely to skip the question.
The observed data could therefore look like:
| Debt range | Response rate |
|---|---|
| $0–$5,000 | 90% |
| $5,001–$20,000 | 80% |
| $20,001–$50,000 | 65% |
| $50,000+ | 35% |
The response rate decreases as debt increases.
If the high-debt respondents are systematically missing, calculating statistics only from observed responses can underestimate the population’s debt burden.
This is why missingness itself can sometimes be informative.
MNAR in Healthcare Data
Healthcare data can contain particularly important MNAR patterns.
Imagine a hospital monitoring patient symptoms over time.
Patients who recover quickly may consistently complete follow-up assessments.
Patients whose condition becomes worse may:
- Leave the study
- Miss appointments
- Be transferred
- Become unable to complete assessments
Now missing follow-up measurements may be related to the patient’s underlying health condition.
This creates a form of informative missingness.
Simply removing patients with missing observations could produce a biased picture of patient outcomes.
For example, if patients with worse outcomes are disproportionately absent from the final measurement, the remaining dataset may make the treatment appear more successful than it actually was.
MNAR in Customer Analytics
MNAR can also appear in commercial datasets.
Consider customer churn.
A company sends a customer satisfaction survey shortly before customers renew or cancel their subscriptions.
Suppose customers who are extremely unhappy are less likely to respond.
You might observe:
Satisfied customers → High response rate
Neutral customers → Moderate response rate
Unhappy customers → Low response rate
If you analyze only completed surveys, you may conclude:
“Most customers are satisfied.”
But the customers who were most likely to leave may be underrepresented.
This can affect:
- Churn analysis
- Customer segmentation
- Retention strategies
- Customer experience measurement
- Product decisions
MNAR in Machine Learning
Missingness can affect machine learning models in two ways.
First, missing values can reduce the amount of usable training data.
Second, the pattern of missingness itself may contain information.
Suppose a financial application asks users to provide their monthly income.
Users with unusually low income may be more likely to leave the field blank.
If you simply impute missing income values with the median, you may hide a meaningful pattern.
The model sees:
Missing income
↓
Median income
instead of:
Missing income
↓
Potentially informative behavior
This can reduce the model’s ability to capture the underlying relationship.
Should Missingness Be Added as a Feature?
Sometimes the fact that a value is missing is itself predictive.
A common technique is to create a missingness indicator.
For example:
df["income_missing"] = df["income"].isna().astype(int)
This creates:
income_missing = 1
when income is missing and:
income_missing = 0
when income is present.
You can then impute the income value separately.
For example:
df["income"] = df["income"].fillna(
df["income"].median()
)
Now the model receives both:
income
income_missing
This can help the model distinguish between:
Actual median income
and:
Missing income that happened to be replaced by the median
However, a missingness indicator does not magically solve MNAR.
It can capture the predictive information contained in the missingness pattern, but it does not reveal the actual unobserved values or prove the underlying missing-data mechanism.
Why Median Imputation Can Be Problematic
Median imputation is simple:
df["income"] = df["income"].fillna(
df["income"].median()
)
But suppose missing income values are systematically associated with high-income respondents.
The median could substantially underestimate them.
Imagine:
Observed income:
30k
35k
40k
45k
50k
Missing income:
150k
180k
The observed median is:
40k
If you replace the missing values with $40,000, you create:
40k
40k
40k
where the original missing values might have been much larger.
The distribution becomes distorted.
This can affect:
- Mean
- Variance
- Quantiles
- Correlations
- Regression coefficients
- Machine learning predictions
How Can You Detect MNAR?
There is an important limitation here:
MNAR generally cannot be confirmed from the observed data alone.
You can investigate whether missingness appears systematic, but proving that missingness depends on an unobserved value requires additional assumptions or information.
Still, several approaches can help.
1. Analyze Missingness Patterns
Start by examining where missing values occur.
df.isna().mean().sort_values(ascending=False)
This shows the percentage of missing values in each column.
You can also compare missingness across observed variables.
For example:
df.groupby("age_group")["income"].apply(
lambda x: x.isna().mean()
)
If missing income rates differ significantly across age groups, missingness may not be random.
2. Create Missingness Indicators
You can create indicators for important variables.
df["income_missing"] = df["income"].isna()
Then compare the indicator with other observed variables.
df.groupby("region")["income_missing"].mean()
If certain groups have systematically higher missingness, that provides evidence that missingness is related to observed characteristics.
This could suggest MAR, but it does not by itself prove MNAR.
3. Compare Respondents and Non-Respondents
For survey data, compare people who answered a question with those who did not.
For example:
Responded:
Average age = 34
Did not respond:
Average age = 48
If response behavior differs systematically, investigate why.
4. Talk to Domain Experts
This is especially important for MNAR.
A statistician may identify patterns in the data.
But a domain expert may know why those patterns exist.
For example:
“Customers with unresolved complaints are less likely to complete this survey.”
That information may not appear anywhere in the dataset.
Business context can therefore be extremely valuable when investigating missingness.
Sensitivity Analysis for MNAR
Because MNAR cannot generally be identified from observed data alone, researchers often use sensitivity analysis.
The idea is to ask:
“What happens to our conclusions under different assumptions about the missing values?”
Suppose you estimate:
Average income = $55,000
You could analyze several scenarios.
Scenario A
Missing values have a similar distribution to observed values.
Scenario B
Missing values are 20% higher than observed values.
Scenario C
Missing values are 50% higher.
You then compare the resulting conclusions.
If the conclusion changes dramatically, the analysis is sensitive to the missing-data assumption.
This is often more informative than pretending you know exactly what the missing values should be.
Pattern-Mixture Models
Pattern-mixture models are one statistical approach for dealing with MNAR assumptions.
The basic idea is to separate observations according to their missing-data patterns.
For example:
Complete cases
Partial cases
Missing outcome cases
The distribution of the outcome can then be modeled under different assumptions about the missing groups.
This can be useful in areas such as:
- Clinical research
- Longitudinal studies
- Survey analysis
The exact implementation depends heavily on the statistical problem.
Selection Models
Another approach is a selection model.
Here, the outcome and the probability of observation are modeled together.
Conceptually:
Outcome
│
├───────────────┐
↓ ↓
Observed value Observation probability
The model attempts to account for the possibility that whether a value is observed depends on the underlying outcome.
Selection models can be useful for formally modeling MNAR mechanisms, but they require assumptions that should be carefully justified.
Tipping Point Analysis
Another practical approach is a tipping point analysis.
You gradually change assumptions about the missing values and determine when your main conclusion changes.
For example:
Assumption Result
Missing values = observed mean Positive
Missing values = 10% lower Positive
Missing values = 20% lower Positive
Missing values = 30% lower Neutral
Missing values = 40% lower Negative
The point at which the conclusion changes is the tipping point.
This helps stakeholders understand how dependent the analysis is on assumptions about missing data.
MNAR vs Missing Data Imputation
It is important to separate two different questions.
Question 1
What should we do with missing values?
This is an imputation question.
Question 2
Why are the values missing?
This is a missing-data mechanism question.
You should ideally understand the second question before deciding how to address the first.
A sophisticated imputation method cannot automatically correct a badly specified missingness assumption.
Common Mistakes When Handling MNAR
Mistake 1: Assuming Missing Means Random
Missing values should not automatically be treated as random.
Investigate the pattern first.
Mistake 2: Automatically Using Mean or Median Imputation
Simple imputation can distort distributions when missingness is systematic.
Mistake 3: Dropping Every Row With Missing Values
Removing incomplete records can introduce serious selection bias.
Mistake 4: Assuming a Missingness Pattern Proves MNAR
A relationship between missingness and observed variables may indicate that MCAR is unlikely, but it does not prove MNAR.
Mistake 5: Ignoring Domain Knowledge
The reason for missing data may exist outside the dataset.
Talk to people who understand how the data was collected.
Mistake 6: Treating One Imputation as the Truth
For MNAR problems, different assumptions can produce different results.
Sensitivity analysis is often more appropriate than presenting one imputed dataset as unquestionably correct.
A Practical Python Workflow
A basic workflow for investigating missing data could look like this:
import pandas as pd
# Load data
df = pd.read_csv("data.csv")
# 1. Check missingness
missing_rate = df.isna().mean()
print(missing_rate)
# 2. Create a missingness indicator
df["income_missing"] = df["income"].isna().astype(int)
# 3. Compare missingness across an observed variable
missing_by_region = (
df.groupby("region")["income_missing"]
.mean()
)
print(missing_by_region)
# 4. Examine observed income
print(df["income"].describe())
This does not prove that your data is MNAR.
Instead, it helps you investigate whether missingness has observable patterns that require further analysis.
For more advanced problems, statistical methods and domain-specific assumptions may be necessary.
A Useful Mental Model
When you encounter missing data, ask these questions in order:
1. Which values are missing?
↓
2. How much data is missing?
↓
3. Where is it missing?
↓
4. Who or what is associated with the missingness?
↓
5. Could the missing value itself influence missingness?
↓
6. What assumptions are reasonable?
↓
7. How sensitive are the results to those assumptions?
This is much better than immediately running:
df.fillna(df.median())
MNAR and Selection Bias
MNAR is closely related to selection bias.
Suppose dissatisfied customers are less likely to respond to a survey.
The observed dataset contains disproportionately more satisfied customers.
The sample is therefore not representative of the population.
This can produce biased estimates.
The missingness mechanism becomes part of the selection process:
Customer population
↓
Who responds?
↓
Observed dataset
↓
Analysis
If the probability of entering the observed dataset depends on the outcome being studied, the resulting analysis can be biased.
This is one reason missing-data mechanisms should be considered during study design rather than only after the dataset has been collected.
How to Reduce MNAR Problems Before Data Collection
The best way to deal with difficult missing data is sometimes to prevent it from occurring.
Make questions easier to answer
Long or intrusive survey questions can increase non-response.
Explain why information is needed
People may be more willing to provide sensitive information when the purpose is clear.
Use multiple data sources
A missing value in one system may be available elsewhere.
Track why data is missing
Instead of storing only:
income = NULL
you might capture:
income = NULL
missing_reason = "prefer_not_to_answer"
This is much more informative.
Other possible reasons include:
not_applicable
not_collected
system_error
unknown
refused
not_yet_available
Separating these cases can make the dataset much easier to analyze.
Missing data is not simply an inconvenience that needs to be filled with another number.
The reason data is missing can be just as important as the missing value itself.
Missing Not At Random occurs when missingness is related to the value that is missing or to information that has not been observed.
That makes MNAR particularly challenging because the missingness mechanism cannot generally be identified from the observed data alone.
The right response is therefore not always to immediately impute the missing values.
Instead:
- Investigate missingness patterns.
- Compare observed and missing groups.
- Use domain knowledge.
- Document your assumptions.
- Consider missingness indicators where appropriate.
- Use suitable statistical models for MNAR problems.
- Perform sensitivity analysis.
- Avoid presenting uncertain imputations as facts.
Most importantly, remember that MCAR, MAR, and MNAR are assumptions about why data is missing, not simply three different types of empty cells.
Understanding that distinction can prevent missing data from quietly turning into biased analysis.
Frequently Asked Questions
1. What does Missing Not At Random mean?
Missing Not At Random (MNAR) means that the probability of a value being missing is related to the missing value itself or to other unobserved information. For example, people with very high incomes may be less likely to report their income.
2. What is the difference between MAR and MNAR?
With MAR, missingness can be explained by information that has already been observed. With MNAR, missingness depends on the missing value itself or information that remains unobserved.
3. Can MNAR be detected from a dataset?
MNAR generally cannot be confirmed using observed data alone. You can investigate patterns in missingness, but determining whether missingness depends on unobserved values requires additional assumptions, domain knowledge, external information, or sensitivity analysis.
4. Is median imputation appropriate for MNAR data?
Median imputation can be problematic when missingness is systematically related to the missing values. It may distort the distribution and produce biased results. The appropriate approach depends on the missing-data mechanism and the purpose of the analysis.
5. How can I handle MNAR data?
Possible approaches include sensitivity analysis, pattern-mixture models, selection models, missingness indicators, domain-specific assumptions, and improved data collection. There is no single imputation technique that automatically solves every MNAR problem.