End-to-end polygenic score analysis
From GWAS retrieval to cross-cohort association in one declarative, reproducible R workflow.
R> remotes::install_github("holstegelab/PolyGenius")Applying a published PGSscore a cohort, plot the distribution
Apply one score. Scale to thousands. Or build your own.
Whichever of those you came for, it is the same package and the same few calls. Each script below is real, and produced the figure beside it.
ROSMAP
Cohort
848
Participants
83
Variants scored
1.3 GB
Peak memory
adpgs-rosmap.Rwall1m 22s
# Fetching PGS Catalog Bellenguez 2022 AD PGS SNP list
bellenguez22 <- generate$from.pgs.catalog("PGS002280")
# Coupling PGS model, genotypes, and phenotypes
data <- PolyGeniusStudy("ROSMAP", bellenguez22, genotypes,
samples = list(phenotypes = fread("phenotypes.csv")))
# Computing scores and population structure
data$scores$X <- compute$scores(data, maf.thr = 0.01)
data$scores$X.scaled <- compute$standardize(data)
data$samples$PCA <- compute$populationStructure(data, npcs = 5)
# trait~PGS association analysis
data$associations$assoc <- associate$regression(data,
outcomes = c(demented, braaksc, ceradsc),
scores.layer = X.scaled,
covariates = c(sex, age, PCA))09:14:02Relatedness: 22 related pair(s) — dropping 22 of 870 samples
09:14:21Scoring 1 model × 848 samples against 83 variants by in-memory indexing
09:14:42Finished calculating PRS scores in 20.51s
09:15:24Association analysis of 848 participants: 56.72% demented, 66.98% female
09:15:24 >
AD-PGS distribution by dementia status · ROSMAPresult
associate$regression · 3 outcomes
Dementia1.42
Braak stage0.19
CERAD−0.17
OR / β per SD · 95% CI
−3−1.501.53
Non-demented 367Demented 481AD-PGS (SD units)
Illustration of real ROSMAP run: 848 participants, PGS002280 (83 variants)