Why PolyGenius
What PolyGenius is for, and how it changes a PGS project
Polygenic scores (PGS, or polygenic risk scores — PRS) have been widely used in the study of human genetics, enabling researchers to predict disease risk, unravel the genetic architecture of traits, and translate GWAS findings into actionable insights. However, the path from GWAS summary statistics to biological insights is complex, requiring expertise in multiple computational tools, careful data harmonisation, and rigorous statistical validation. PolyGenius simplifies this journey without sacrificing scientific rigour.
Whether you're a seasoned computational geneticist or a researcher new to PGS analysis, PolyGenius provides a unified framework that makes PGS analysis intuitive and straightforward for researchers at all levels. Expert users gain powerful customisation, while beginners receive guided workflows that implement best practices.
The two typical questions
Most researchers use PGSs to address one of two questions. The first is What does this score tell me about my data? Translating this research question into an analysis plan involves obtaining the list of SNPs making up the PGS, computing the PGS values in the target genotype dataset, accounting for factors such as relatedness of samples and population structure, and relating scores to diverse types of outcomes of interest. These analytical steps might also demand more technical ones such as performing liftover of PGSs between genome builds or obtaining reference panels, which further increase the burden of performing PGS analysis.
This research question also naturally scales. Rather than assessing a single PGS of some trait, we might want to investigate what a whole set of trait scores tells us about our data. We term this sort of analysis a polygenic score-wide association study (PGS-WAS). This sort of analysis is becoming more and more frequent. Both in the single-trait and the many-trait case, the main goal is inference, and it is probably the more common use-case.
The second question is Which score should I use? In this question, we compare different PGSs for the same trait (i.e. different SNP lists) and assess how well they perform in some testing dataset. This involves either providing PGSs yourself or retrieving them from online repositories, or generating new PGSs by executing one or many PGS construction algorithms on GWAS summary statistics for the desired trait. Once we have the previously defined or newly generated SNP lists, we need to evaluate their performance in predicting various aspects related to this score, and select the best approach. That is the more specialised model-development workflow that can provide you or future researchers with the PGSs required for the first question. PolyGenius provides an integrated approach for addressing this question too.
What PolyGenius does
Regardless of your research question, the process of turning it into rigorous scientific results often involves patching together multiple analysis tools and ad-hoc scripts. You have to match genome builds and perform liftover where needed, choose a reference panel, resolve variant orientation, and normalise outputs from different construction algorithms. You then score cohort genotypes, perform statistical associations and plot the results. Each step has its own inputs, assumptions and output shape.
That chain becomes even more costly through repetition. Change one input and you walk it again by hand, including sample-dependent standardisation and model refitting. Along the way, there is no reliable record of which intermediate results remain valid, which have drifted from the new inputs, or which could have been reused.
PolyGenius replaces that hand-maintained chain with a declarative, intent-driven interface: you state the analysis you want — not the resources to prepare first — and it resolves harmonisation, external resources, software dependencies and reusable intermediate computations. Downloaded summary statistics, reference panels, LD matrices, liftover chains, tool binaries and generated models are cached, keyed to the identity of the inputs that define them, so re-requesting one that has not changed costs nothing and changing an input invalidates only what depended on it. Scoring, association and evaluation, and visualisation of results are treated as first-class objectives and have dedicated modules to assist with performing these steps in an efficient and reproducible manner.
Working across sites
Importantly, all computation remains on the same machine you are working on. PolyGenius reads your genotype files and uses them locally. Genotype or phenotype data are never transferred, and the same analysis script can be executed at multiple sites, exporting only summary-level statistics for downstream meta-analysis.
What stays your judgement
PolyGenius does not decide whether a trait is relevant to your research question, whether the selected reference panel matches the ancestry of your cohort, or whether calibration transports to another population. You also decide whether a mediation result supports a credible causal interpretation. Those are study-design judgements and result interpretations, not software defaults.
What PolyGenius records is what it ran, not why you chose it. Each model carries its algorithm, parameters and source GWAS, generated model collections carry an execution handle, and each association carries the formula and sample counts behind its fits, so a reader can see what was done. It does not record why you chose a trait or reference panel, and it does not check those decisions.
The same boundary applies to your input data. PolyGenius assumes that primary genotype quality control, phasing and imputation have already been completed; it can apply analysis-time filters and report which samples are related, but it never modifies your genotypes. It does not repair an unsuitable or incompletely prepared genotype dataset. Those responsibilities remain yours.
Getting set up
To get started, simply install PolyGenius:
remotes::install_github("holstegelab/PolyGenius")This will install the package with all of its R dependencies. As we will discuss later, some operations like genotype scoring or model generation require
external tools such as PLINK2. You can instruct PolyGenius to set up such dependencies via workspace$setup$install("plink"). This, however, is not strictly required
as PolyGenius ensures those dependencies are present before they are required, and if they are not, it installs them.
Now that we have installed PolyGenius, we are ready to start.
- For What does this score tell me about my data?, continue to Your First Analysis.
- For Which score should I use?, begin with GWAS Sources and Construction Algorithms.