HUGIML Causal Studio Static interactive preview powered by demo artifacts
INTERFACE
HUGIML Causal Investigation Studio

Causal effect estimation with interpretable T-HUG

Configure a binary-treatment study, estimate potential outcomes, inspect overlap, and compare interpretable T-learners.

Shared vocabularyCross-fitted DR/AIPWOverlap sensitivity
Demo dataset
Observational

Synthetic credit-risk study with modeled confounding and known treatment-effect heterogeneity.

Run settings
42, 43, 44
2 folds × 3 runs

Fixed seeds provide reproducible repeated estimates and confidence intervals.

Quick evaluates one configuration per model. Larger grids provide broader performance or interpretability searches.

Causal graph and assumptions
Baseline covariates6 pre-treatment factorsTreatmentEnhanced reviewOutcome12-month default

Baseline covariates may affect treatment assignment and outcome; treatment may affect outcome.

T-HUG methodology

Control and treatment outcome models share adaptive bins and feature semantics. Utilities, patterns, coefficients, and the downstream LR/RPTE branch remain arm-specific.

Repeated out-of-fold predictions support stability intervals. A cross-fitted propensity model provides the DR/AIPW estimate and trimming sensitivity.

Analysis complete · 3,200 rows · 3 models
T-HUG ATE
−1.39%
Primary analysis
Mean control risk
26.66%
Mean treatment risk
25.27%
Estimated to benefit
50.9%
Treatment-effect distribution
Repeated cross-fitted estimates
Repeated OOF ATE
−1.47%
95% CI (−2.12%, −0.82%)
DR/AIPW ATE
−1.53%
95% CI (−2.75%, −0.31%)

Every row receives out-of-fold predictions in each seeded run.

Rows
3,200
Control
1,771
Treatment
1,429
Overlap .05–.95
96.8%
Treatment assignment and overlap
Overlap sensitivity
RangeRetainedDR ATE (95% CI)
Untrimmed100.0%−1.53% (−2.75%, −0.31%)
.01–.9999.4%−1.49% (−2.66%, −0.32%)
.05–.9596.8%−1.42% (−2.51%, −0.33%)
.10–.9088.2%−1.38% (−2.42%, −0.34%)
T-HUG structural regions

Shared thresholds keep control and treatment patterns directly comparable.

Region / HUG patternPopulationMean P0Mean P1Mean CATE
credit_util=[0.72,0.91) & credit_score=[540,620)31835.8%27.4%−8.4%
prior_delinquency=1 & dti=[0.41,0.58)22642.1%35.9%−6.2%
annual_income=[42000,61000) & age=[28,42)40124.7%23.3%−1.4%
Baseline comparison

All methods use the same covariates, split policy, scoring rule, and T-learner estimand.

ModelHeld-out AUCHeld-out BrierFit secondsOracle CATE RMSEOracle CATE corr
T-HUG0.6880.1733.350.0760.104
T-LR0.6530.2311.840.0890.221
T-XGB0.6960.17218.890.087−0.070
Best values are shown in bold; second-best values are italicized. Ranking is applied only where higher or lower has an objective interpretation.
Repeated-split stability and doubly robust estimates
ModelRepeated OOF ATE (95% CI)ATE repeat SDDR/AIPW ATE (95% CI)DR repeat SD
T-HUG−1.47% (−2.12%, −0.82%)0.0026−1.53% (−2.75%, −0.31%)0.0049
T-LR−9.95% (−27.76%, 7.86%)0.07170.81% (0.07%, 1.55%)0.0030
T-XGB−0.83% (−3.05%, 1.38%)0.0089−2.01% (−3.54%, −0.48%)0.0061