Bridging the gap between correlation and causation for high-stakes decisions
"Correlation does not imply causation" - We've all heard it, but what do we do about it?
Traditional machine learning excels at finding patterns and correlations, but fails when we need to answer:
- 🤔 "What will happen if we take this action?" (Intervention)
- 🧮 "What caused this outcome?" (Attribution)
- 🎯 "Which factors should we target?" (Policy decisions)
Real-World Consequences:
- Healthcare: Is treatment A actually better, or are healthier patients more likely to choose it?
- Marketing: Did the campaign increase sales, or would they have increased anyway?
- Finance: Does this risk factor cause defaults, or is it just correlated with them?
CausalInference-Toolkit provides a comprehensive, unified framework for causal inference in Python. Built by a mathematician-turned-data-scientist, it combines:
✅ Rigorous statistical foundations (from Ghent University's statistical program) ✅ Practical ML integration (scikit-learn, PyTorch compatibility) ✅ Real-world applicability (Business, Healthcare, Finance use cases)
| Feature | Description | Why It Matters |
|---|---|---|
| Multiple Methods | Propensity Scores, IV, Double ML, Causal Discovery | Different problems need different tools |
| Unified API | Consistent interface across methods | Easy to compare and validate |
| Assumption Testing | Built-in sensitivity analysis | Know when results are trustworthy |
| Business Metrics | ROI, treatment effects, policy impact | Speak the language of stakeholders |
| Visualization | Interactive causal graphs, effect plots | Communicate results clearly |
| Production Ready | Pip installable, documented, tested | Deploy with confidence |
# Core package
pip install causal-inference-toolkit
# With all dependencies
pip install causal-inference-toolkit[full]
# Development version
pip install git+https://github.com/twomathematicians-code/CausalInference-Toolkit.gitfrom causal_inference import CausalAnalyzer
import pandas as pd
# Load your data
data = pd.read_csv('marketing_campaign.csv')
# Initialize analyzer
causal = CausalAnalyzer(
treatment='received_campaign',
outcome='purchase_amount',
covariates=['age', 'income', 'previous_purchases']
)
# Estimate causal effect
effect = causal.estimate_effect(method='propensity_score')
# Visualize results
effect.plot()
# Get business impact
print(f"Campaign increased purchases by ${effect.ate:.2f} on average")
print(f"ROI: {effect.roi():.1f}x")Output:
Campaign increased purchases by $45.23 on average
ROI: 3.2x
Statistical Significance: p < 0.001
Confidence Interval: [$38.90, $51.56]
When to use: Selection bias, observational studies
from causal_inference import PropensityScoreMatching
psm = PropensityScoreMatching(
treatment='treatment',
outcome='outcome',
covariates=['confounder1', 'confounder2']
)
# Estimate
results = psm.estimate(data)
# Check balance
psm.check_balance(data).plot()Applications:
- Healthcare: Treatment effectiveness
- Marketing: Campaign impact
- HR: Training program effectiveness
When to use: Unobserved confounders, randomized encouragement
from causal_inference import InstrumentalVariables
iv = InstrumentalVariables(
treatment='treatment',
outcome='outcome',
instrument='random_assignment'
)
results = iv.estimate(data)
# First stage regression check
iv.check_first_stage(data)Applications:
- Economics: Returns to education
- Finance: Effect of credit on outcomes
- Policy: Regulation impact
When to use: High-dimensional confounders, complex relationships
from causal_inference import DoubleML
dml = DoubleML(
treatment='treatment',
outcome='outcome',
covariates=[f'feature_{i}' for i in range(100)], # 100 features
model_y='gradient_boosting', # ML model for outcome
model_t='lasso' # ML model for treatment
)
results = dml.estimate(data, cross_fitting=True)Applications:
- Digital Marketing: Many user features, complex behavior
- Personalized Medicine: High-dimensional patient data
- Finance: Multiple risk factors
When to use: Unknown causal structure, exploratory analysis
from causal_inference import CausalDiscovery
cd = CausalDiscovery(method='pc') # Peter-Clark algorithm
# Discover causal graph
graph = cd.discover(data)
# Visualize
graph.plot()
# Get Markov blanket of target
markov_blanket = graph.get_markov_blanket('outcome')Applications:
- Systems Biology: Gene regulatory networks
- Economics: Macroeconomic relationships
- Social Science: Behavioral factors
Problem: Does a new diabetes medication actually reduce blood sugar, or are healthier patients more likely to be prescribed it?
Solution:
from causal_inference import CausalAnalyzer
import pandas as pd
# Patient data
patients = pd.read_csv('diabetes_treatment.csv')
# Initialize with clinical knowledge
causal = CausalAnalyzer(
treatment='new_medication',
outcome='blood_sugar_reduction',
covariates=['age', 'bmi', 'baseline_sugar', 'comorbidities', 'prior_treatment'],
data=patients
)
# Multiple methods for robustness
psm_effect = causal.estimate_effect(method='psm')
dml_effect = causal.estimate_effect(method='dml')
# Sensitivity analysis
sensitivity = causal.sensitivity_analysis()
print(f"Treatment Effect: {psm_effect.ate:.2f} (PSM)")
print(f"Treatment Effect: {dml_effect.ate:.2f} (DML)")
print(f"Robustness: {'High' if abs(psm_effect.ate - dml_effect.ate) < 2 else 'Low'}")Result: 95% confidence the medication reduces blood sugar by 15-20 mg/dL, independent of patient selection.
Problem: Did our email campaign increase sales, or would sales have increased anyway?
Solution:
from causal_inference import InstrumentalVariables
# Use random timing as instrument
iv = InstrumentalVariables(
treatment='received_email',
outcome='purchase_amount',
instrument='random_send_time'
)
results = iv.estimate(marketing_data)
# Business metrics
roi = results.calculate_roi(cost_per_email=0.50)
incremental_revenue = results.ate * len(marketing_data)
print(f"Email Campaign ROI: {roi:.1f}x")
print(f"Incremental Revenue: ${incremental_revenue:,.0f}")Result: 4.2x ROI, $125,000 incremental revenue attributable to campaign.
Problem: Does high debt-to-income ratio cause loan defaults, or is it just correlated?
Solution:
from causal_inference import PropensityScoreMatching
# Match similar borrowers
psm = PropensityScoreMatching(
treatment='high_dti',
outcome='default',
covariates=['credit_score', 'employment_years', 'loan_amount', 'income']
)
results = psm.estimate(loan_data)
# Counterfactual analysis
counterfactual = results.counterfactual_analysis(
scenario='reduce_dti_to_safe_level'
)
print(f"Causal Effect: {results.ate:.4f}")
print(f"Defaults Preventable: {counterfactual.n_prevented:,.0f}")Result: High DTI causes 12% increase in default probability. Reducing DTI could prevent 2,500 defaults/year.
This toolkit is built on rigorous statistical theory:
- Rubin Causal Model: Potential outcomes framework
- Neyman-Rubin Causal Model: Counterfactual reasoning
- Structural Causal Models: Pearl's DAGs and do-calculus
- Econometrics: Angrist-Imbens-Rubin framework
Every method includes: ✅ Statistical significance tests ✅ Assumption diagnostics ✅ Sensitivity analysis ✅ Bootstrap confidence intervals ✅ Placebo tests
| Feature | This Toolkit | DoWhy | EconML | CausalNex |
|---|---|---|---|---|
| Unified API | ✅ | ❌ | ❌ | ❌ |
| Multiple Methods | ✅ 4+ | ✅ 3 | ✅ 2 | ✅ 1 |
| Business Metrics | ✅ | ❌ | ❌ | ❌ |
| Visualization | ✅ | ❌ | ✅ | |
| Assumption Tests | ✅ | ✅ | ❌ | |
| Beginner Friendly | ✅ | ❌ Advanced |
When to use what:
- This toolkit: Business applications, need ROI/impact metrics, comparing methods
- DoWhy: Microsoft ecosystem, specific IV needs
- EconML: Advanced econometrics, heterogeneous effects
- CausalNex: Bayesian networks, causal discovery
# Find which segments benefit most
hte = causal.heterogeneous_effects(
data=data,
segments=['age_group', 'customer_tier']
)
# Visualize
hte.plot_heatmap()
# Target high-impact segments
best_segments = hte.get_top_segments(n=3)# How robust is our result to unobserved confounders?
sensitivity = causal.sensitivity_analysis(
method='rosenbaum',
gamma_range=(1.0, 2.0)
)
sensitivity.plot()
print(f"Result robust to hidden bias up to {sensitivity.max_robust_gamma:.2f}")# Comprehensive validation
validation = causal.validate(
tests=[
'balance_check',
'placebo_test',
'refutation_test',
'placebo_outcome'
]
)
validation.report()- 📚 Full Documentation
- 🎓 Tutorial: Causal Inference for Data Scientists
- 💼 Business Guide: ROI and Beyond
- 🏥 Healthcare Applications
- 💰 Finance Use Cases
Contributions are welcome! Areas of interest:
- 🐛 Bug fixes: See Issues
- ✨ New methods: Regression discontinuity, synthetic control, etc.
- 📊 Visualizations: Better plots, interactive dashboards
- 📖 Documentation: Examples, tutorials, case studies
- 🧪 Testing: More datasets, edge cases
See CONTRIBUTING.md for guidelines.
This project is licensed under the MIT License - see the LICENSE file for details.
Built on the shoulders of giants:
- Judea Pearl - Causal inference framework
- Guido Imbens & Donald Rubin - Potential outcomes
- Microsoft Research - DoWhy and EconML
- Ghent University - Statistical foundations
Mahesh Solanki - Creator & Maintainer
- 📧 Email: maheshsinh1910@gmail.com
- 💼 LinkedIn: mahesh-solanki-16b9a6a5
- 🐙 GitHub: @twomathematicians-code
- Propensity Score Matching
- Instrumental Variables
- Double Machine Learning
- Basic Causal Discovery
- Documentation & Examples
- Regression Discontinuity Design
- Synthetic Control Methods
- Difference-in-Differences
- Enhanced visualization dashboard
- Time-varying treatments
- Mediation analysis
- Causal survival analysis
- Interactive causal graph editor
"In God we trust. All others must bring data." - W. Edwards Deming
"But data alone isn't enough. We need to understand why." - This toolkit
Made with ❤️ and mathematical rigor by Mahesh Solanki