Sampling bias is one of the easiest problems to overlook in data analysis and machine learning. A dataset can contain thousands or even millions of records and still produce misleading results if the people, transactions, events, or observations included in it do not adequately represent the population being studied.
For example, imagine building a customer satisfaction model using responses from customers who voluntarily completed an online survey. Customers with very positive or very negative experiences might be more likely to respond than customers with average experiences. If the analysis treats those responses as representative of every customer, the conclusions may be unreliable.
Python provides practical tools for investigating this problem. With libraries such as pandas, NumPy, SciPy, and Matplotlib, you can compare a sample against a reference population, inspect underrepresented groups, test differences in distributions, and identify patterns that suggest a sampling problem.
This guide explains how to detect sampling bias with Python, using a practical example you can adapt to your own datasets.
What Is Sampling Bias?
Sampling bias occurs when the process used to select observations systematically makes some members of a population more likely to be included than others.
The resulting sample may differ from the population in ways that matter to the analysis.
Consider a company that wants to understand customer satisfaction across all its users. If it collects feedback primarily from customers who use its mobile application every day, occasional users may be underrepresented. The resulting satisfaction score could give an inaccurate picture of the entire customer base.
Sampling bias can affect:
- Business surveys and customer research
- Healthcare and social science studies
- Machine learning training datasets
- Market research and public opinion polling
- Fraud detection and financial risk models
- Website analytics and product experiments
Sampling bias is not the same as random sampling error. Random sampling error arises because a sample contains only part of a population. Sampling bias arises when the selection process systematically favors particular observations.
Increasing the sample size does not automatically eliminate sampling bias. A very large, unrepresentative sample can still lead to incorrect conclusions.
Sampling Bias vs. Data Bias
These terms are related, but they are not identical.
| Concept | Meaning | Example |
|---|---|---|
| Sampling bias | The selection process makes the sample unrepresentative | A survey excludes customers who rarely use the internet |
| Measurement bias | The collection method measures something inaccurately | A faulty sensor consistently records temperatures too high |
| Label bias | Recorded target values systematically misrepresent the truth | Historical decisions reflect inconsistent human judgments |
| Selection bias | Inclusion in the analyzed dataset depends on factors related to the outcome or variables being studied | Only successful loan applicants are included in a repayment analysis |
Sampling bias is one possible source of broader data bias. Understanding the distinction helps you investigate the correct problem instead of assuming every dataset imbalance comes from sampling.
1. Load the Dataset and Define the Population
Before testing for sampling bias, identify two things:
- The sample you want to analyze.
- A trustworthy reference describing the population from which that sample should have been drawn.
Without a suitable reference, you can identify suspicious patterns, but you generally cannot establish whether the sample represents the target population.
For this example, suppose a company has a customer population containing different age groups. Its survey sample contains disproportionately many younger customers.
We will simulate both datasets so the code can run independently.
import numpy as np
import pandas as pd
np.random.seed(42)
# Simulated reference population
population = pd.DataFrame({
"age_group": np.random.choice(
["18-29", "30-44", "45-59", "60+"],
size=10000,
p=[0.20, 0.35, 0.30, 0.15]
)
})
# Simulated survey sample with overrepresentation of younger users
sample = pd.DataFrame({
"age_group": np.random.choice(
["18-29", "30-44", "45-59", "60+"],
size=1000,
p=[0.50, 0.25, 0.18, 0.07]
)
})
print(population.head())
print(sample.head())
The population represents the distribution the company wants its survey to reflect. The sample represents the customers who actually responded.
In a real project, load your files instead:
population = pd.read_csv("customer_population.csv")
sample = pd.read_csv("customer_survey.csv")
Make sure both datasets use compatible category definitions and refer to the same target population and relevant time period. A reference dataset that covers a different customer base or period may produce misleading comparisons.
2. Compare Sample and Population Distributions
One of the simplest ways to investigate sampling bias is to compare the percentage of observations in each group.
If 20% of the target population belongs to a particular age group but only 5% of the sample does, that difference deserves investigation.
Use pandas to calculate the proportions.
population_share = (
population["age_group"]
.value_counts(normalize=True)
.sort_index()
)
sample_share = (
sample["age_group"]
.value_counts(normalize=True)
.sort_index()
)
comparison = pd.DataFrame({
"population_share": population_share,
"sample_share": sample_share
}).fillna(0)
comparison["difference_percentage_points"] = (
comparison["sample_share"] -
comparison["population_share"]
) * 100
print(comparison.round(3))
The population_share column shows how common each group is in the reference population. The sample_share column shows its representation in the survey.
The difference is expressed in percentage points rather than as a relative percentage change.
For instance, if a group represents 20% of the population and 30% of the sample, the difference is 10 percentage points.
How to interpret the results
Look for groups that are consistently overrepresented or underrepresented.
In the simulated example, younger customers should account for a larger share of the sample than of the population. Older customers should account for a smaller share.
However, a difference does not automatically prove bias. Some variation is expected in random samples. The importance of a discrepancy depends on sample size, sampling design, the variable being studied, and the intended analysis.
Also check the population proportions themselves. If the reference data is incomplete or biased, comparing against it can create false confidence.
3. Visualize Sampling Differences With Matplotlib
A chart makes distribution differences easier to communicate to analysts, researchers, and business stakeholders.
import matplotlib.pyplot as plt
comparison[
["population_share", "sample_share"]
].plot(
kind="bar",
figsize=(9, 5)
)
plt.title("Sample Distribution vs Population Distribution")
plt.xlabel("Age Group")
plt.ylabel("Proportion")
plt.xticks(rotation=0)
plt.legend(["Population", "Sample"])
plt.tight_layout()
plt.show()
Look for categories where the sample bar differs noticeably from the population bar.
A useful practice is to include both the chart and the numerical comparison in your analysis report. The chart communicates the pattern, while the table gives exact values.
Visualizations are especially useful when investigating multiple variables such as age, region, income band, customer type, and device category.
4. Test Categorical Differences With a Chi-Square Test
A visual difference tells you where to investigate, but a statistical test can help assess whether the observed distribution differs from what would be expected under a specified sampling assumption.
For categorical data, the chi-square goodness-of-fit test is a common option.
The null hypothesis is that the sample follows the reference population proportions.
from scipy.stats import chisquare
observed_counts = (
sample["age_group"]
.value_counts()
.reindex(population_share.index, fill_value=0)
)
expected_counts = (
population_share * len(sample)
)
chi2_stat, p_value = chisquare(
f_obs=observed_counts,
f_exp=expected_counts
)
print("Chi-square statistic:", chi2_stat)
print("P-value:", p_value)
Understanding the output
The chi-square statistic summarizes the discrepancy between observed and expected counts. The p-value indicates how surprising a discrepancy at least this large would be if the null hypothesis and the test assumptions held.
A small p-value, often below a chosen significance level such as 0.05, provides evidence that the sample distribution differs from the reference proportions.
It does not prove that the sampling mechanism is biased, identify the cause, or tell you whether the difference is practically important.
With very large samples, even small differences can be statistically significant. Therefore, interpret the test alongside the percentage-point differences and the consequences for the intended analysis.
Important: The standard test assumes independent observations and appropriate expected counts. If your data comes from a clustered, stratified, or weighted survey, you may need a survey-aware statistical method instead. If the sample and reference datasets overlap in a way that violates the assumed sampling structure, the test also needs reconsideration.
5. Detect Sampling Bias Across Multiple Variables
Sampling bias may not be obvious when you inspect only one variable.
A sample might match the population’s age distribution but differ significantly by region, income, education, or customer activity.
Start by comparing each categorical feature separately.
def compare_categorical_distribution(
population,
sample,
column
):
pop_share = (
population[column]
.value_counts(normalize=True)
)
sample_share = (
sample[column]
.value_counts(normalize=True)
)
result = pd.DataFrame({
"population_share": pop_share,
"sample_share": sample_share
}).fillna(0)
result["difference_pp"] = (
result["sample_share"] -
result["population_share"]
) * 100
return result.sort_values(
"difference_pp",
key=abs,
ascending=False
)
print(
compare_categorical_distribution(
population,
sample,
"age_group"
).round(2)
)
You can reuse this function for other columns, provided each dataset contains the relevant variable.
For example:
# Examples for real datasets
# compare_categorical_distribution(
# population, sample, "region"
# )
#
# compare_categorical_distribution(
# population, sample, "customer_type"
# )
In a real analysis, create a consistent set of category definitions and handle missing values explicitly. The function above treats categories absent from one dataset as having zero observed share, but it does not account for missing values as a separate category unless you explicitly include them.
Why checking one variable is insufficient
Suppose the sample contains the correct proportions of age groups and regions separately. It may still overrepresent younger customers from urban areas and underrepresent older customers from rural areas.
This is a joint-distribution problem. Checking individual variables can miss it.
When relevant, compare combinations of important variables, such as age group by region, or use a multivariable model to investigate whether sample inclusion depends on several characteristics at once.
Avoid testing every possible combination without a plan. Too many comparisons can generate false positives, and sparse groups can make results unstable.
6. Compare Numerical Variables
Categorical variables are not the only indicators of sampling bias. Numerical features can also differ between the sample and the reference population.
Examples include:
- Annual income
- Customer spending
- Age in years
- Number of monthly transactions
- Website session duration
- Product usage frequency
Start by comparing summary statistics.
print(
population["annual_spend"].describe()
)
print(
sample["annual_spend"].describe()
)
These commands assume both datasets contain an annual_spend column. Replace it with a numerical feature in your actual data.
Look at the mean, median, standard deviation, quartiles, and range. Two datasets may have similar averages but very different distributions, so summary statistics should not be your only check.
Plot the distributions
plt.figure(figsize=(9, 5))
plt.hist(
population["annual_spend"].dropna(),
bins=30,
alpha=0.5,
density=True,
label="Population"
)
plt.hist(
sample["annual_spend"].dropna(),
bins=30,
alpha=0.5,
density=True,
label="Sample"
)
plt.title("Annual Spending: Sample vs Population")
plt.xlabel("Annual Spend")
plt.ylabel("Density")
plt.legend()
plt.tight_layout()
plt.show()
Overlapping histograms can reveal differences in spread, skewness, and the representation of high- or low-value observations.
If the sample contains fewer high-spending customers than the population, a model trained on that sample may not perform as expected on high-value customers.
For more detailed comparisons, use box plots, empirical cumulative distribution functions, or quantile comparisons.
7. Use the Kolmogorov-Smirnov Test for Numerical Distributions
For two independent samples of numerical observations, the two-sample Kolmogorov-Smirnov test can assess whether their empirical distributions differ.
from scipy.stats import ks_2samp
population_values = (
population["annual_spend"].dropna()
)
sample_values = (
sample["annual_spend"].dropna()
)
ks_stat, p_value = ks_2samp(
population_values,
sample_values
)
print("KS statistic:", ks_stat)
print("P-value:", p_value)
This example requires an annual_spend column in both datasets.
The KS statistic measures the maximum absolute difference between their empirical cumulative distribution functions. A larger statistic indicates a greater discrepancy between the distributions.
A small p-value provides evidence against the hypothesis that the two independent samples were drawn from the same continuous distribution.
Use caution when applying this test to complex survey data, weighted observations, dependent samples, or datasets with extensive ties. The standard test’s assumptions may not hold in those settings.
Also, distribution differences are evidence of a mismatch, not conclusive proof of a biased selection mechanism.
8. Check for Missing Groups and Missing Values
A sample may appear representative among the records it contains while excluding an important group entirely.
For example, a digital survey might omit people who cannot access the survey platform. If platform access is associated with income, age, geography, or the outcome being measured, the omission can distort the findings.
Check missing values in both datasets.
print("Population missing values:")
print(population.isna().sum())
print("\nSample missing values:")
print(sample.isna().sum())
Then check whether expected categories appear in both datasets.
expected_groups = population["age_group"].unique()
observed_groups = sample["age_group"].unique()
missing_groups = set(expected_groups) - set(observed_groups)
print("Groups absent from sample:", missing_groups)
An absent group is a warning sign when that group is part of the target population. However, rare groups can be absent from a genuinely random sample by chance, particularly when the sample is small.
Investigate why observations are missing. Missing data can arise from survey nonresponse, technical failures, filtering rules, data joins, or eligibility requirements. These causes require different solutions.
9. Investigate Selection Bias With a Sampling Indicator
Distribution comparisons are useful, but you can go further by modeling which observations entered the sample.
If you have a combined dataset containing sampled and nonsampled population members, create a binary variable indicating sample inclusion.
For example:
1means the observation was included in the sample.0means it was not included.
Then use a classification model to see whether observable characteristics predict inclusion.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score
# Combined data must contain both sampled and nonsampled
# population members, with sample_included coded as 0 or 1.
features = ["age", "region", "annual_spend"]
target = "sample_included"
X = combined[features]
y = combined[target]
categorical_features = ["region"]
numerical_features = ["age", "annual_spend"]
preprocessor = ColumnTransformer([
(
"categorical",
OneHotEncoder(handle_unknown="ignore"),
categorical_features
),
(
"numerical",
StandardScaler(),
numerical_features
)
])
model = make_pipeline(
preprocessor,
LogisticRegression(max_iter=1000)
)
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.25,
random_state=42,
stratify=y
)
model.fit(X_train, y_train)
predicted_probabilities = model.predict_proba(X_test)[:, 1]
auc = roc_auc_score(
y_test,
predicted_probabilities
)
print("Sampling indicator ROC-AUC:", round(auc, 3))
This example assumes you already have a dataframe named combined, with the specified features and target. It is a template rather than a standalone runnable example.
How to interpret the model
If the model can distinguish sampled from nonsampled records substantially better than chance, observable characteristics help explain sample inclusion.
An ROC-AUC near 0.5 indicates little discrimination on the evaluated data; a higher value indicates greater ability to distinguish the two groups. However, an AUC near 0.5 does not prove representativeness. The model may be missing important variables, nonlinear patterns, or interactions.
Likewise, a high AUC does not by itself prove harmful bias. Some inclusion differences may be intentional, such as deliberate oversampling of a small population group.
The more important question is whether the differences undermine the intended conclusions or the model’s performance on the target population.
10. Calculate Representation Ratios
A representation ratio gives a simple measure of how strongly each category is represented in the sample relative to the population.
It is calculated as:
\frac{\text{Sample Proportion}}
{\text{Population Proportion}}
]
A ratio of 1 means the category has the same share in the sample and population. A ratio below 1 indicates underrepresentation, while a ratio above 1 indicates overrepresentation.
Using the earlier comparison table:
comparison["representation_ratio"] = (
comparison["sample_share"] /
comparison["population_share"]
)
print(
comparison[
[
"population_share",
"sample_share",
"representation_ratio"
]
].round(3)
)
For example, a ratio of 0.5 means a category’s share of the sample is half its share of the population. A ratio of 2 means its share is twice as large.
This is a descriptive diagnostic, not a universal fairness threshold. Ratios for very rare categories can be unstable, and a ratio of 1 for one variable does not guarantee that the sample is representative across other characteristics.
11. Correct Sampling Bias After Detection
Detecting sampling bias is only the first step. What you do next depends on how the sample was collected and what information is available.
Option 1: Collect a better sample
The strongest solution is often to improve the sampling process.
Possible approaches include:
- Randomly selecting eligible observations from the target population.
- Using stratified sampling to ensure relevant groups are represented.
- Reaching groups that were missed by the original collection method.
- Investigating and reducing survey nonresponse.
- Documenting exclusions and eligibility requirements.
A better sampling design can prevent problems that are difficult to fix after data collection.
Option 2: Use post-stratification or sample weights
When population proportions are known and the sample includes the relevant groups, weighting can help align the sample with the target population.
A simple post-stratification weight for group (g) is:
[
w_g =
\frac{P_g}{S_g}
]
where (P_g) is the group’s population proportion and (S_g) is its sample proportion.
In Python:
population_share = (
population["age_group"]
.value_counts(normalize=True)
)
sample_share = (
sample["age_group"]
.value_counts(normalize=True)
)
weights_by_group = (
population_share / sample_share
)
sample["sample_weight"] = (
sample["age_group"].map(weights_by_group)
)
print(sample.head())
This produces a group-level weight for each sampled observation. It assumes every relevant population group is represented in the sample and that the population proportions are reliable.
Weights can reduce observed differences on the characteristics used to construct them, but they cannot automatically correct unobserved differences. Very large weights can also make estimates unstable, so effective sample size and weight variability should be examined.
If a population group is completely absent from the sample, this method cannot recover its outcomes from nothing. You may need to collect more data or reconsider the target population.
Option 3: Evaluate the impact on your actual analysis
Not every mismatch has the same consequences.
Suppose a sample underrepresents older customers. If age is related to the outcome you are studying, the mismatch may materially affect the results. If age is unrelated to the outcome and other relevant variables, its impact may be smaller.
Compare key estimates before and after weighting or adjustment. For machine learning, evaluate model performance on a suitable target-population test set and examine important subgroups separately.
Do not assume that balancing demographic proportions automatically eliminates all forms of sampling bias.
12. A Practical Sampling Bias Checklist
Before using a sample for analysis or machine learning, ask:
- Have I clearly defined the target population?
- Do I have a credible reference for its characteristics?
- Have I compared the sample and population proportions?
- Have I checked numerical distributions as well as categorical variables?
- Are important groups absent or underrepresented?
- Could nonresponse or eligibility rules explain the differences?
- Have I considered interactions between important variables?
- Have I used statistical tests with appropriate assumptions?
- Have I measured whether the differences affect the final analysis?
- If I used weights, have I checked their stability and limitations?
This checklist provides a starting point for data quality reviews, exploratory data analysis, survey analysis, and machine learning preparation.
Conclusion
Sampling bias can make an otherwise accurate analysis misleading. Python helps you investigate the problem by comparing sample and population distributions, visualizing numerical differences, testing categorical and numerical distributions, checking missing groups, and modeling sample inclusion.
The most important step is to compare the sample with a credible reference population. Without that reference, unusual patterns may be worth investigating, but representativeness is difficult to assess.
When a mismatch exists, determine whether it affects the question you are trying to answer. Depending on the cause, the right response may be to collect a better sample, use carefully constructed weights, or change the scope of the analysis.
Ultimately, reliable data analysis begins not just with asking whether your dataset is large enough, but also whether it represents the population you want to understand.
Frequently Asked Questions
1. How do you detect sampling bias in Python?
Compare the sample with a credible reference population. Use pandas to examine category proportions and summary statistics, Matplotlib to visualize differences, and SciPy tests to investigate distribution mismatches. These checks provide evidence of potential problems but do not independently prove the cause.
2. Which Python libraries are useful for detecting sampling bias?
Pandas is useful for data manipulation and distribution comparisons. NumPy supports numerical calculations, SciPy provides statistical tests, Matplotlib creates visualizations, and scikit-learn can model the relationship between observed characteristics and sample inclusion.
3. What is the difference between sampling bias and random sampling error?
Random sampling error is natural variation between a sample and its population. Sampling bias occurs when the selection process systematically favors certain observations or excludes others. A larger sample can reduce random error but does not necessarily eliminate bias.
4. Can the chi-square test detect sampling bias?
A chi-square goodness-of-fit test can identify evidence that observed categorical proportions differ from specified population proportions. It does not prove that the selection mechanism caused the difference, and the test must be appropriate for the sampling design and data.
5. How can you correct sampling bias in Python?
Depending on the situation, you can use stratified sampling, post-stratification weights, or other appropriate adjustments. You should also investigate why the bias occurred. Weighting cannot reliably recover information about groups that are completely absent from the sample or correct every unobserved source of bias.