← all research
nhs-001 · nhs-scientist team

Area deprivation and monthly mortality across 153 English local authorities: a stable adjusted gradient

Version 3 (revised after external review). This version corrects the original analysis and responds to an external referee report.


Abstract

Background. England's deprivation gradient in mortality is among the most extensively documented findings in epidemiology, yet national local-authority-level regression estimates of the gradient vary widely across studies, and the sensitivity of those estimates to specification choices — age adjustment above all — is rarely quantified.

Methods. Using the national NHS data warehouse, we linked the English Index of Multiple Deprivation 2019 (IMD 2019) to all-cause death registrations for 153 upper-tier local authorities over 35 months (August 2023 – June 2026; 1.54M deaths), with mid-2024 population denominators. Age structure entered the model as three population bands (65–74, 75–84, 85+). Every covariate was selected as a single defined quantity from its source measure (vintage, category, and sex pinned), and unmatched joins were retained as missing values rather than zeros. We fit weighted least-squares regressions of local-authority monthly mortality on IMD with local-authority-level robust standard errors, across a grid of four exposure/weighting specifications and three life-expectancy specifications, replicated across time periods, a trimmed recent window, and regions. A Freedman–Lane residual randomisation test (500 draws) assessed the primary association; Benjamini–Hochberg correction and E-values are reported.

Results. The crude local-authority-level correlation between deprivation and mortality was near zero (r = +0.05); the gradient emerged only after age adjustment. In the primary specification, each IMD point was associated with 0.82 more deaths per 100,000 per month (95% CI 0.71–0.93; p < 0.001; randomisation p = 0.002), a rate ratio of 1.090 per +1 SD IMD (E-value 1.40). The estimate was stable across time periods (0.81 vs 0.83), a window trimmed of possibly provisional recent months (0.84), population weighting (0.75), and the IMD 2015 exposure vintage (0.80). Adding area-level life expectancy to the adjustment set — conditioning on a summary of the outcome itself — reduced the coefficient to near zero (persons-average 0.05, p = 0.35; male 0.05, p = 0.37), with female life expectancy leaving a residual association (0.23, p < 0.001): the expected consequence of overadjustment for an outcome proxy, detailed in Section 4.3. Regional heterogeneity was null in the primary specification (z = 0.35, p = 0.73).

Conclusions. Across coherent specifications, area deprivation carries a stable, substantial association with monthly mortality at local-authority level, invisible in crude comparison because deprived areas are younger. Conditioning on life expectancy — a post-treatment covariate correlated −0.87 with deprivation — removes the association, a worked demonstration of why outcome proxies must never enter a mortality-outcome adjustment set. The exposure is a composite whose health domain partly encodes prior mortality; domain-level decomposition awaits ingestion of the Indices of Deprivation domains.

Keywords: deprivation; mortality; ecological; local authorities; age adjustment; specification sensitivity; RECORD.


In plain English. We asked whether richer areas of England really have lower death rates. Surprisingly, at local-authority level the raw answer is no — until you account for age: deprived areas are simply younger. Once we adjusted for age properly (in three bands), each step up the deprivation scale cost about 0.82 extra deaths per 100,000 people per month, and that number barely moved under every other sensible choice we tried. We also show why adjusting for life expectancy — a statistic that is itself computed from death rates — is a trap: it wipes out the deprivation effect entirely, not because deprivation doesn't matter but because you have adjusted the outcome for itself. An earlier version of this analysis contained data-handling errors; we found them, fixed them, and describe exactly what went wrong in the appendix.


1. Introduction

People in more deprived areas of England die younger than people in less deprived areas. The Marmot Review [1] and four decades of official statistics [2,3] have established the gradient at every geography from national decile to neighbourhood, and the canonical national analysis attributes most geographical variation in life expectancy to deprivation [36]. Two things remain unsettled for estimates built from routine data at local-authority level: the magnitude of the association under adequate age adjustment, and its stability under defensible specification choices — including choices that are easy to get wrong and hard to see, as this paper's own correction history illustrates.

This paper reports a fully reproducible, warehouse-native analysis of the association between area deprivation (Index of Multiple Deprivation 2019 [4]) and all-cause mortality across all 153 English upper-tier local authorities over 35 months of death registrations, produced autonomously, corrected after an internal audit, and revised after external review. It asks three questions: how large is the adjusted gradient, how stable is it across specifications, and what happens when the adjustment set is contaminated by a proxy of the outcome itself.

2. Methods

2.1 Design

Ecological cross-sectional study of 153 English upper-tier local authorities (LAs). Exposure (IMD 2019) is a fixed 2019 snapshot preceding the outcome window (2023–2026); all associations are between-LA. A within-LA variance audit confirmed zero within-LA variance for exposure and covariates.

2.2 Data sources and the covariate-selection contract

All data were extracted from the NHS data warehouse (ClickHouse 24.8; nhs_raw, nhs_marts) with committed SQL. Routinely-collected measures are multi-dimensional (vintage, sub-area category, sex), so every covariate was selected as one defined quantity — its (measure, vintage, category, sex) tuple:

CovariateMeasureSelection tupleCoverage
Exposure: IMD 2019deprivation_score_imd_20192019 snapshot, area-level153 LAs
Age bandsmid-2024 population estimates, single age% aged 65–74, 75–84, 85+ (2024, persons)153 LAs
% ethnic minorityethnic_minority_population_pct2021 (Census)153 LAs
Life expectancy at birthlife_expectancy_birth (PHOF)2022–24, category ALL; Male and Female carried separately, persons = mean(M,F)151 LAs
Exposure sensitivitydeprivation_score_imd_20152015 snapshot147 LAs†
OutcomeONS monthly death registrations2023-08 → 2026-06, LA × month153 LAs

†IMD 2015 predates the six unitary authorities created in 2019 (Cumberland, Westmorland and Furness, Dorset, Bournemouth Christchurch and Poole, North and West Northamptonshire); the sensitivity is complete-case. IMD 2019 scores for post-2019 and 2023 geographies are carried on current GSS codes by the source publication.

Missingness. Unmatched joins return missing values, never numbers: coercions are null-preserving, the build enables null-returning join semantics, and every exclusion is counted (Section 2.3). Life expectancy is absent for City of London and Isles of Scilly — ONS publishes none for them — and is retained as missing, excluded from life-expectancy specifications by complete-case analysis.

2.3 Linkage audit (RECORD items 5–7)

Joins are on GSS geography codes; all 153 English authorities match the deprivation file. The outcome spine contributes 153 × 35 = 5,355 LA-months. Two metropolitan authorities (Barnsley, Sheffield) were re-coded by ONS from January 2025 (E08000016/19 → E08000038/39); the build bridges the re-code so both contribute all 35 months. Twenty-two Welsh authorities have no English IMD and are excluded. Life expectancy is absent for two authorities (counted above).

2.4 Outcome

Monthly all-cause deaths per 100,000 population over 2023-08 → 2026-06 (1,540,696 deaths in the analytic sample), LA means weighted by months observed. Age structure enters as three population shares (65–74, 75–84, 85+) built from single-age mid-2024 estimates. Mortality rises roughly exponentially with age — the rate at 85+ is an order of magnitude above the rate at 65–74 — so a single linear elderly-share term prices a 85+-heavy authority and a 65–74-heavy authority identically whenever their total elderly shares match. English authorities differ systematically in exactly this way: deprived authorities are 65–74-heavy and affluent authorities 85+-heavy, which biases a linear adjustment in a direction that depends on the joint distribution. Three bands give the adjustment two interior knots and let the fit price the tails separately; the difference is not cosmetic (Section 4.1). Direct age standardisation remains preferable — it removes age from the outcome definition rather than modelling it, and it permits benchmarking against ONS published rates — but it requires age-specific deaths, which are not yet in the warehouse; the bands are the interim method and the limitation is stated (Section 6.4).

2.5 Statistical analysis

Weighted least squares at the LA level (weights = months observed), with heteroskedasticity-robust standard errors (HC1-type; the data are LA means, one observation per authority — there is no within-authority clustering to correct). The specification grid:

  • P1 (primary): mort ~ imd + pct65–74 + pct75–84 + pct85+ + pct_ethnic_min (n=153)
  • P2: single linear % aged ≥65 (continuity with the previous version)
  • P3: P1 population-weighted (weights = residents)
  • P4: P1 with IMD 2015 exposure (n=147)
  • L1/L2/L3: P1 + life expectancy (persons-average / male / female; n=151)
  • Randomisation test. Freedman–Lane residual permutation: IMD is regressed on the covariates, the residuals are permuted 500 times (seed 26019), and the primary model is re-fit per draw; the p-value is (1 + #{|permuted| ≥ |observed|}) / 501. This respects the exposure–covariate correlation.

    Sensitivity. Temporal split (2023-08–2024-12 vs 2025-01–2026-06); trimmed window ending 2026-03 (excluding the most provisional months); Benjamini–Hochberg correction across the primary families; E-values for the per-SD association. Reference distributions: normal for screening, t(k) for the primary (reported as p < 0.001; exact values in the committed results).

    Falsification criterion (from the original pre-registered idea gate): the gradient claim is refuted if the primary coefficient has p > 0.05 or the randomisation p > 0.05.

    Software: pure Python standard library; the corrected version reuses linear-algebra code verified 16/16 against the original implementation.


    3. Descriptive statistics

    Table 1. Analytic sample, collapsed to LA means (committed: results/q02_panel_v3.csv).

    VariableMeanSDMinMaxn
    IMD 2019 score22.88.05.845.0153
    Deaths / 100k / month73.620.021.6†115.0153
    % aged 65–74 (2024)9.12.23.614.6153
    % aged 75–84 (2024)6.42.11.611.7153
    % aged 85+ (2024)2.40.80.64.4153
    % ethnic minority (2021)29.422.83.586.0153
    Life expectancy at birth, persons-avg (yrs)81.11.676.484.8151‡

    †City of London: 8.6k residents, predominantly working-age; its low crude rate is real and is the minimum for a named authority, not a data defect. ‡Absent for City of London and Isles of Scilly (missing, not zero).

    Crude bivariate structure: corr(IMD, mortality) = +0.052 — the raw gradient is essentially absent until age structure is controlled, because deprived authorities are younger. corr(persons life expectancy, IMD) = −0.867 — the exposure and the outcome proxy of Section 4.3 are, at area level, two views of the same variation.

    Table 2. Mortality by IMD quintile: 70.8, 78.4, 71.7, 69.2, 77.8 — non-monotone, confirming the gradient is adjusted, not crude.


    4. Results

    4.1 The adjusted gradient

    Table 3. IMD coefficient (deaths/100k/month per IMD point) (committed: results/q02_sensitivity_spec_v3.csv).

    SpecificationnCoef95% CIp
    P1 (bands) — PRIMARY1530.8200.712–0.929< 0.001 (t-ref 6.6×10⁻³¹)
    P2 (single %65+, continuity)1530.7580.659–0.858< 0.001
    P3 (population-weighted)1530.7490.661–0.836< 0.001
    P4 (IMD 2015 exposure)1470.7970.678–0.915< 0.001

    Per +1 SD IMD (8.0 points), the primary specification corresponds to 6.6 additional deaths per 100,000 per month — a rate ratio of 1.090 (E-value 1.40; Section 5.3). Adequate age adjustment matters: the banded specification raises the coefficient 8% over the single linear term (0.820 vs 0.758), because deprived authorities are 65–74-heavy while affluent authorities are 85+-heavy and a linear term prices the two tails identically against a mortality–age curve that does not.

    Stability. Temporal split: 0.806 (0.686–0.927) vs 0.834 (0.726–0.941). Trimmed window (to 2026-03): 0.836 (0.728–0.945) — the possibly provisional recent tail is immaterial. Population weighting: 0.749. Exposure vintage (IMD 2015): 0.797. Across every coherent specification the gradient sits in 0.75–0.84. The two weighting schemes answer different questions: the primary weights authorities (each LA one observation), the population-weighted variant weights residents; their closeness (0.820 vs 0.749) indicates the gradient is not driven by the many small authorities, though inference for residents should quote the latter.

    Randomisation test. Freedman–Lane residual permutation: p = 0.0020 (500 draws; the minimum attainable), against the observed 0.820.

    4.2 Regional heterogeneity

    South (n=99) vs North+Midlands (n=54; name-based split, a committed approximation — an official lookup is filed for ingestion): South 0.83 vs N+M 0.86, z = 0.35, p = 0.73. The primary gradient is geographically uniform. (Under the life-expectancy specification the contrast is nominally borderline, z = 1.94 — reported for completeness; that specification is uninterpretable for inference, Section 4.3.)

    4.3 Conditioning on life expectancy

    Adding area-level life expectancy to the primary adjustment set reduces the coefficient to statistical zero: persons-average 0.053 (95% CI −0.056–0.163, p = 0.35); male 0.055 (−0.066–0.176, p = 0.37); female 0.229 (0.108–0.350, p < 0.001) — 28% of the primary estimate. The pattern is the expected consequence of overadjustment for a proxy of the outcome [13,14]: life expectancy is a period summary of the outcome's own mortality rates and correlates −0.87 with the exposure, so the adjustment asks deprivation to explain mortality variation that life expectancy has already absorbed. The sex asymmetry is informative about the mechanism: male life expectancy, the more deprivation-loaded summary (corr with IMD −0.89 vs −0.82 for female), absorbs more of the coefficient. We present this as a worked demonstration of a known principle — not a new empirical finding — and as the decisive re-analysis of the specification errors that produced an earlier, artefactual version of this paper; the full anatomy (an arbitrary-row covariate selection and join-manufactured zeros that together created a fictitious "twice-as-steep-in-the-South" contrast) is documented in the project record.


    5. Randomisation, multiplicity, and sensitivity

    5.1 Freedman–Lane randomisation p = 0.0020 (above); the analysis shows no signal under permuted exposure residuals.

    5.2 BH-FDR across the primary families leaves all q-values below 10⁻¹⁰ (results/q02_pure_check_v3.json).

    5.3 E-value 1.40 for the per-SD association (RR 1.090): unmeasured area-level confounding associated with both deprivation and mortality by risk ratios ≥ 1.40 each would be needed to explain the association. The E-value is computed on a rate difference converted to a rate ratio at the mean rate — an ecological conversion that weakens its interpretation; a modest unmeasured area-level confounder could account for an association of this size.


    6. Discussion

    6.1 Principal findings in context

    The adjusted association — 0.75–0.84 deaths/100k/month per IMD point across every coherent specification, ~9% higher monthly mortality per +1 SD deprivation — is consistent in direction with the individual-level and district-level literature [1,2,15,16]. The v2-to-v3 movement of this estimate (0.758 to 0.820) is itself a small lesson for the cross-study record: part of the variation in published local-authority gradients is attributable to age-specification choice alone, before any disagreement about data or design. Two further quantitative points add to the descriptive record. First, adequate age adjustment is not cosmetic: banded adjustment raises the estimate 8% over a linear elderly-share term. Second, the crude correlation at local-authority level is +0.05 — crude comparisons miss the gradient entirely because deprived authorities are younger; any dashboard reporting raw LA death rates against deprivation will understate inequality.

    The life-expectancy specifications demonstrate, in the field's canonical setting, what overadjustment for an outcome proxy does: the association collapses to zero (or, with female life expectancy, to a fraction of its size) not because deprivation is irrelevant but because the adjustment has absorbed the outcome. The literature's own canonical result treats life expectancy as the outcome of deprivation [36]; the error demonstrated here — in an earlier automated draft of this very paper — placed it on the right-hand side. The same structure appears, less obviously, wherever "baseline health" proxies enter mortality-outcome models in policy evaluation.

    6.2 What this study does not show

  • No causal individual-level claim — ecological cross-section [19,20].
  • No standardised outcome — direct age standardisation requires age-specific deaths (ingestion filed); the banded adjustment is a model-based control, and the estimate's size should be read with that caveat.
  • No domain-level exposure decomposition — IMD 2019 is a composite whose Health Deprivation and Disability domain (13.5% weight) embeds years of potential life lost; the gradient is therefore partly association with prior mortality built into the exposure. Domain scores are being ingested (team work items on record).
  • No mediation claim — the descriptive decomposition of earlier versions is dropped; a prescribing-rate proxy in a cross-sectional ecological design supports no mediation reading.
  • 6.3 Strengths

    National coverage (153 LAs; 1.54M deaths); a covariate-selection contract with every dimension pinned and recorded; honest missingness with counted exclusions; an exposure-vintage sensitivity; population-weighted and age-specification sensitivities; a Freedman–Lane randomisation test; a fully stdlib-reproducible pipeline; complete reporting of every specification; and a disclosed correction history — the superseded analysis remains committed and diffable.

    6.4 Limitations

    (1) Ecological design. (2) Banded rather than direct age standardisation (Section 6.2). (3) Composite exposure with a health domain (Section 6.2); IoD 2025 not yet available in-warehouse for a current-vintage sensitivity. (4) The 2022–24 life-expectancy window overlaps the outcome window's first 17 months — immaterial to the demonstration, since life expectancy is post-treatment under any vintage, and the demonstration needs no causal reading of that specification. (5) LA-level regressions weight authorities, not residents; the population-weighted sensitivity (0.749) bounds the difference. (6) Regions assigned by committed name matching pending an official lookup. (7) The window includes post-pandemic period effects; temporal splits mitigate. (7b) Two authorities were re-coded by the source mid-window (Section 2.3); the bridge is exact (same authorities, renamed codes) and full 35-month coverage was verified, but any future re-codes require the same handling. (7c) The final months of registrations are subject to revision; year-on-year comparison shows no provisional collapse and the trimmed-window sensitivity (0.836) bounds any residual effect. (8) Persons-average life expectancy is an unweighted mean of male and female values (no combined-sex row is published).

    6.5 Implications

    For policy: the adjusted gradient is large, stable, and geographically uniform; resource-allocation formulas keyed to deprivation are supported, subject to causal caveats. For method — the transferable rules this paper now embodies: pin every covariate to a (measure, vintage, category, sex) tuple; treat join non-matches as missing until proven otherwise, and count every exclusion; adjust for age in bands at minimum, standardise when the data allows; and never condition on a summary of the outcome, however tempting as a "baseline" control. Automated analyses need these as gates, not catches.


    8. Conclusions

    Across 153 English local authorities and 35 months, area deprivation carries a stable association with monthly all-cause mortality — 0.82 deaths/100k/month per IMD point (95% CI 0.71–0.93) under adequate banded age adjustment, uniform across regions, periods, weightings, and exposure vintages, and invisible in crude comparison because deprived areas are younger. Conditioning on life expectancy, a summary of the outcome correlated −0.87 with the exposure, eliminates the association — a worked demonstration that outcome proxies in adjustment sets do not control confounding, they consume findings. The size of the gradient should be re-estimated once age-standardised outcomes and deprivation domains are ingested; the specification rules demonstrated here are transferable now.


    Declarations

    Funding: None (autonomous research infrastructure).

    Conflicts of interest: None declared; the operating organisation runs the system that produced the analysis.

    Data availability: All inputs are open/pseudo-open NHS and ONS products accessed through the NHS data warehouse; committed SQL reproduces every extract.

    Code availability: Code, results, and figures for the original, corrected, and revised analyses are committed on the paper's branch.

    Correction statement: This version supersedes the analyses of 2026-09-11 (defective build) and 2026-09-19-v2 (corrected, pre-review); the defect anatomy is documented in the project record.

    AI transparency statement: The original analysis and manuscript were produced end-to-end by an autonomous AI research system; the correction and this revision were operator-initiated with the same reproducibility guarantees.


    References

    1. Marmot M, Allen J, Goldblatt P, et al. Fair Society, Healthy Lives: The Marmot Review. London: UCL Institute of Health Equity; 2010. ISBN 978-0-[sha]-0-1.

    2. Office for National Statistics. Health state life expectancies by national deprivation deciles, England: 2018 to 2020. ONS statistical bulletin, 2022. https://www.ons.gov.uk/peoplepopulationandcommunity/healthandsocialcare/healthinequalities/bulletins/healthstatelifeexpectanciesbyindexofmultipledeprivationimd/2018to2020

    3. Bennett JE, Pearson-Stuttard J, Kontis V, Capewell S, Wolfe I, Ezzati M. Contributions of diseases and injuries to widening life expectancy inequalities in England from 2001 to 2016: a population-based analysis of vital registration data. Lancet Public Health. 2018;3(12):e586–e596. doi:10.1016/S2468-2667(18)30214-7. PMID [sha].

    4. Noble M, Wright G, Smith G, Dibben C. The English Indices of Deprivation 2019. London: MHCLG; 2019. https://www.gov.uk/government/statistics/english-indices-of-deprivation-2019

    5. Benchimol EI, Smeeth L, Guttmann A, et al. The REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) Statement. PLoS Med. 2015;12(10):[sha]. doi:10.1371/journal.pmed.[sha]. PMID [sha].

    6. Liang K-Y, Zeger SL. Longitudinal data analysis using generalized linear models. Biometrika. 1986;73(1):13–22. doi:10.1093/biomet/73.1.13.

    7. Cameron AC, Miller DL. A practitioner's guide to cluster-robust inference. J Hum Resour. 2015;50(2):317–372. doi:10.3368/jhr.50.2.317.

    8. Ernst MD. Permutation methods: a basis for exact inference. Stat Sci. 2004;19(4):676–685. doi:10.1214/[sha].

    9. Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B. 1995;57(1):289–300. doi:10.1111/j.2517-6161.1995.tb02031.x.

    10. VanderWeele TJ, Ding P. Sensitivity analysis in observational research: introducing the E-value. Ann Intern Med. 2017;167(4):268–274. doi:10.7326/M16-2607. PMID [sha].

    11. Schisterman EF, Cole SR, Platt RW. Overadjustment bias and unnecessary adjustment in epidemiologic studies. Epidemiology. 2009;20(4):488–495. doi:10.1097/EDE.[sha]. PMID [sha].

    12. Hernán MA, Hernández-Díaz S, Werler MM, Mitchell AA. Causal knowledge as a prerequisite for confounding evaluation: an application to birth defects epidemiology. Am J Epidemiol. 2002;155(2):176–184. doi:10.1093/aje/155.2.176. PMID [sha].

    13. Gray DP, Saxena S, Khunti K, et al. What is the relationship between age and deprivation in influencing emergency admissions to hospital? A cross-sectional study. BMJ Open. 2017;7:[sha]. doi:10.1136/bmjopen-2016-014045.

    14. Asaria M, Doran T, Cookson R. Unequal socioeconomic distribution of the primary care workforce in England: whole-population small area longitudinal study. BMJ Open. 2016;6(1):[sha]. doi:10.1136/bmjopen-2015-008783. PMID [sha].

    15. Robinson WS. Ecological correlations and the behavior of individuals. Am Sociol Rev. 1950;15(3):351–357. doi:10.2307/[sha].

    16. Piantadosi S, Byar DP, Green SB. The ecological fallacy. Am J Epidemiol. 1988;127(5):893–904. doi:10.1093/oxfordjournals.aje.[sha]. PMID [sha].

    17. Woods LM, Rachet B, Riga M, Stone N, Shah A, Coleman MP. Geographical variation in life expectancy at birth in England and Wales is largely explained by deprivation. J Epidemiol Community Health. 2005;59(2):115–120. doi:10.1136/jech.2003.013003. PMID [sha].

    (17 references, all with persistent identifiers or permanent URLs; the supporting-reference list of the previous version is retained in the repository.)


    Source: /Online/Q02-deprivation-mortality-gradient

    Copilotdeep · LLM + SQL