# Methodology — Global Health Intelligence

## Data and provenance

The supplied country-year panel contains 2,864 rows, 179 country labels and every year from 2000 to 2015. A one-to-one equality comparison on all 19 numerical fields matches every row to lashagoch's **Life Expectancy (WHO) Fixed** publisher download; regions match and 22 country names differ. Publisher metadata labels this CC0. Original source extractions were not replayed.

The publisher reports filling missing values with nearby three-year means and regional means, and omitting countries with many missing fields. The supplied panel is complete, but imputation flags are absent. Neither temporal nor geographic evaluation can rule out upstream leakage. No observations are removed as outliers. Audit screening assumptions and unresolved definitions are listed in the data dictionary.

## Descriptive analysis

Country distributions are unweighted. Report medians and percentile bands, not a population-weighted global life expectancy. Endpoint improvement uses paired country records. Raw GDP and log-GDP associations are compared using Pearson and Spearman coefficients. Correlations are computed separately for selected-year cross-sections, repeated country-year rows and long-run country averages; pooled coefficients are not treated as independent-observation tests.

GDP is current USD, not PPP or real income. Regional groups and economy flags are dataset conventions. Infant/under-five mortality overlaps and adult mortality uses a different risk set; rates are never summed as shares of deaths.

## Economic baseline and gaps

Prespecified descriptive reference: `Life_expectancy ~ log(GDP_per_capita) + C(Year)` using all 2,864 observations. Year indicators allow different intercepts but one GDP slope. Gap is observed minus fitted life expectancy. Country mean, median, sign persistence and volatility use all 16 years. These are **in-sample reference residuals**, not predictive test errors, causal effects, healthcare efficiency or performance rankings. Raw-GDP and quadratic-log alternatives are fitted as model-form comparisons; the selected interpretive reference remains log-linear.

## Statistical analysis

Models progress from GDP to log GDP, quadratic log GDP, social fields and immunisation fields, each with year indicators. Confidence intervals use country-clustered standard errors for 179 clusters. There are no country fixed effects; coefficients combine between- and within-country associations. Shared shocks, imputation and unmeasured conditions remain possible confounding/dependence.

VIF is a continuous-predictor diagnostic (excluding year dummies), Cook's distance is computed from the corresponding ordinary fit, and Breusch–Pagan is reported as an exploratory statistic without an iid p-value. Polio/DTP3 have substantial collinearity, making separate vaccine coefficients unsuitable for causal interpretation. Associations and intervals are conditional on this specification, not uncertainty over source reconstruction or model choice.

## Prediction

Fixed ten-field primary feature set: natural log GDP, schooling, BMI, Polio, DTP3, HepB3, Measles coverage, recorded alcohol, adolescent thinness and natural log population. No country IDs, regional labels, year, mortality outcomes, HIV or economy flags. Mortality is close to the target's construction; HIV is intentionally outside this primary specification. Predictors are contemporaneous: this is later-year outcome estimation with same-year covariates, not a multi-year forecast.

Train 2000–2010; validation 2011–2012; select family by validation MAE; refit 2000–2012; final test 2013–2015. Families: mean, linear, ridge (alpha 10), random forest (240 trees, leaf size 5, feature fraction .8), gradient boosting (180 iterations, 15 leaves, L2 10). Fixed settings are compared, not tuned on the test. Median-imputation and scaling pipelines are fitted on training data only; the supplied file needs no new imputation.

Report MAE, RMSE and R² for every family on the same test. Report selected-model country/year bias and MAE, largest errors, predicted/observed and residual plots. A country-block bootstrap uses 2,000 resamples for the fixed model's mean absolute error; it excludes source and model-selection uncertainty. Test-set permutation importance uses 15 repeats and seed 42. Correlated predictors can share importance; importance is predictive dependence, not policy causation.

A separate geographic stress test holds out 36 country labels. Family selection uses temporal validation inside only the 143 training countries; the selected family is refitted on all their years. This tests new-country estimation over historical years, not new time. Publisher regional imputation may already cross that boundary.

## Reproduction and artefacts

Run `python src/verify_source.py`, then `python src/run_analysis.py`. Source matching is offline against the preserved downloaded publisher file; analysis itself requires no network. Fixed seeds, dependency versions and source SHA-256 are recorded. SQLite schema and reusable queries create real analytical outputs. SVG/PNG charts, JSON metrics/findings, predictions and reports are generated from the same run.

Source and outputs are versioned, with checksums in `reports/artifact_manifest.json`. The web application only renders these outputs; it does not refit models or invent metrics.
