MS Risk Variants, Single Cell Expression, and MRI Brain Volume
The genetic accounting for multiple sclerosis (MS) splits in two. Single nucleotide variants explain roughly 19% of the heritable component of susceptibility and 13% of disease severity, and the two sets behave differently: susceptibility variants are mainly immune related, while variants associated with severity are predominantly expressed in the central nervous system. Polygenic risk scores sum variant effects to estimate predisposition, but they have not sensitively separated clinical groups in MS, and using single disease-specific hits from MS GWAS to predict long-term severity has been difficult. In Alzheimer's and Crohn's disease, integrating GWAS with expression quantitative trait loci data to build a transcriptional risk score improved the ability to detect clinically relevant subgroups compared with a conventional polygenic score, and that approach had not been applied to MS. Upcott and colleagues did so, weighting variants not by GWAS effect size but by the direction and magnitude of their effect on gene expression, and computing the result per cell type. The work is a medRxiv preprint that has not been through peer review.
From GWAS Hits to eGenes: SMR, HEIDI, Colocalization
The variant-to-gene step ran through summary data-based Mendelian randomization, pairing the discovery phase of the largest MS susceptibility GWAS, 14,802 patients against 26,703 unaffected individuals, with publicly available eQTL data from 7,466 blood donors and 5,494 brain donors. Default SMR thresholding selected eQTLs at p < 5×10⁻⁸ with an R² between 0.05 and 0.90 against the top GWAS SNP at each locus, and gene expression differences associated with MS risk alleles were called at an FDR-corrected p below 0.05. The HEIDI test then removed associations likely to arise from linkage disequilibrium rather than a shared causal variant, excluding anything with a HEIDI p at or below 0.1. Surviving variants went to Bayesian colocalization against single-cell expression: 1.27 million peripheral blood mononuclear cells from 927 donors covering B lymphocytes, CD4+ and CD8+ T cells, dendritic cells, monocytes, natural killer cells and plasma cells, and 750,614 single-nucleus CNS cells from 192 individuals covering astrocytes, endothelial cells, excitatory and inhibitory neurons, microglia, oligodendrocytes, precursor cells and pericytes. A posterior probability above 0.80 was the bar for calling a variant an eGene.
240 eGenes, and How Many Were New
SMR and HEIDI filtering returned 240 significant eGenes associated with MS susceptibility, 43 of them overlapping between immune and CNS tissues, which leaves 88 CNS-specific and 109 immune-specific genes. Most of these had not been seen before in MS: 77 of the 88 CNS genes and 85 of the 109 immune genes were newly identified. Gene set enrichment linked the CNS genes to kinase activity and immune signalling, which the authors read as pointing to neuro-immune interactions, and the immune genes to lymphocyte biology. Three external datasets were then used to ask whether these genes behave as the model implies. In non-inflammatory post-mortem brain tissue, the highest percentage of cells expressing them were inhibitory and excitatory neurons. In MS post-mortem lesions, the majority of the identified CNS genes are differentially expressed in both white and grey matter in a cell-type specific manner. In MS peripheral blood and CSF, a high number of the identified immune genes are differentially expressed in a cell-type specific manner, some across several immune cell subsets.
How the Score Is Actually Built
The construction is where a transcriptional risk score departs from a polygenic one. Beta coefficients of each eQTL association were standardised to a mean of 0 and a standard deviation of 1. The MS susceptibility or progression risk allele of each variant set the direction of risk, and transcriptomic values were polarised in the direction of that genomic risk allele. The score is then the sum of all polarised beta coefficients across identified variants, so a high score means an individual carries, across many loci, the alleles that shift expression in the risk direction. For comparison, polygenic risk scores were computed with SBayesRC, a Bayesian method using genomic functional annotations, separately from the susceptibility and the progression GWAS. All scores were transformed into z-scores and checked against QQ-plots for Gaussian distribution. Welsh cohort genotyping used Illumina Infinium CoreExome-24 v2 or v3 arrays with stringent quality control and imputation; where a variant was absent from the registry genotype data, a proxy at R² = 1 in the CEU population was sought, five such proxies in perfect linkage disequilibrium were used, and no proxy could be found for 47 of the 387 variants identified by SMR and HEIDI.
A Cohort Chosen to Keep the Severity Test Interpretable
Testing genetics against long-term severity runs into a treatment-assignment bias, because patients with unfavourable prognostic factors are more likely both to have worse outcomes and to receive high-efficacy disease-modifying treatments. The authors handled it by restriction rather than adjustment, including only the treatment-naive subgroup of the South Wales MS Registry, 1,077 patients of whom 747 (69.4%) were female, with a median age of onset of 32.2 years (IQR 16), 143 (13.3%) carrying a diagnosis of primary progressive MS, 512 of 633 (80.9%) of those who underwent CSF analysis oligoclonal band positive, and a median last-recorded age-related MS severity score of 6.2 (IQR 4.6). Severity scores were stratified into quartiles and compared across risk scores by Kruskal-Wallis with FDR adjustment, while the association with rank inverse normal transformed severity used multivariate linear regression adjusted for sex and the number of relapses in the first five years. Time between first and second relapse was split at two years, a threshold taken from natural history work tying early relapse frequency to greater severity.
Where the Transcriptional Score Separates Groups and the Polygenic One Does Not
Neither polygenic score was associated with last-recorded severity, susceptibility at FDR p = 0.46 and progression at 0.35. Total transcriptional score missed as well at FDR p = 0.07, as did the CNS score at 0.19. The immune transcriptional score was associated, at FDR p = 0.025, with patients in the highest severity quartile carrying significantly greater immune scores than those in the lowest (post hoc Dunn p = 0.0015). Linear regression adjusted for sex and early relapse count confirmed it at p = 0.039, t = 2.06, and put a number on the effect: a one unit increase in the immune transcriptional risk z-score goes with a 7% increase in severity score (β = 0.07, 95% CI 0.003–0.13). The same split appeared for relapse timing, where total and immune transcriptional scores were associated with time between first and second relapse (FDR p = 0.02 for both) and the polygenic scores were not. Age at disease onset was the one outcome both approaches caught, with total, immune and CNS transcriptional scores at FDR p = 0.006, 0.05 and 0.02 and the susceptibility polygenic score at 0.04. Pushed down to individual cell types, only B-lymphocyte and CD8+ cytotoxic T-lymphocyte scores were associated with last-recorded severity (p = 0.037 and 0.005), and those two were highly correlated at a Spearman's rho of 0.87.
Imaging Validation, Mechanism, and the Limits the Authors Set
For independent validation the authors turned to brain atrophy, a well-validated proxy for MS severity, using UK Biobank whole-brain tissue volume normalised for head size with MS cases defined by ICD-10 code G35 and regression adjusted for age. Across 68,324 general-population participants, total and CNS transcriptional scores correlated inversely with brain volume (beta −890.91, p = 1.97×10⁻⁴; beta −779.78, p = 0.0011; beta −688.05, p = 0.0039). Among the 272 patients with MS the direction was concordant with larger effect sizes but borderline non-significant, which the authors attribute to limited power. At cell-type resolution, CD8+ T-lymphocyte scores but not B-lymphocyte scores were associated with age-adjusted brain volume (beta −718.61, p = 0.002). On mechanism they single out two genes from their own list. IL-12AS1 is an antisense RNA regulating IL-12 production, set against EAE work showing the IL-12 receptor beta on neurons regulates anti-inflammatory properties and prevents neurodegeneration. IL-7 came out as a CNS-specific susceptibility variant, and IL-7 protein is raised in pre-active MS lesions, where CD8+ cytotoxic T cells co-expressing the IL-7 receptor alpha chain are also found. Four limitations close the paper. The treatment-naive design may not represent contemporarily managed patients, though they argue an untreated cohort better reflects the underlying biology and avoids modelling around that bias. The scores separated clinical subgroups less sharply than transcriptional scores have in Crohn's and Alzheimer's disease, with significant overlap between groups, so neither these nor polygenic scores are ready for clinical use. SMR was applied only to the susceptibility GWAS, because the progression GWAS yields a single genome-wide significant hit, which they say makes them more likely to underestimate than overestimate CNS associations. And the UK Biobank patients may have been treated without their models adjusting for it, which would mask rather than inflate the genetic effect.
Disclaimer: This blog post is based on the cited preprint and is intended for informational purposes only. The preprint has not been certified by peer review and should not be used to guide clinical practice. It is not intended to provide medical advice. Please consult with a healthcare professional for any health concerns.
Reference:
Upcott, M., Shao, B., Bray, N., Loveless, S., Tallantyre, E. C., Robertson, N. P., & Kreft, K. L. (2026). Cell-type specific transcriptional risk scores and longitudinal multiple sclerosis outcomes. medRxiv. https://doi.org/10.64898/2026.09.21.26363545
