Construction Algorithms
Clumping-and-thresholding, LDpred2, lassosum2 and PRS-CS: choosing one and what it costs
Before your first run
What chapter 03 needed was PLINK2 and nothing else. This chapter needs more, and it is worth getting in place before you write a generation call.
Install what your chosen method needs — bigsnpr for LDpred2 and lassosum2, a Python environment for PRS-CS — and check what is already present.
You also need a reference panel, which is a real download of substantial size. Browse what is available before you name one.
A note on what you are about to get back, because it changes how you read the rest of this page: generate$models() returns a set of candidates, not a chosen model. Five steps, not one: declare, generate the set, score all of them against a cohort, evaluate, then select. Only after that are you where chapter 03 started.
Choosing
Four questions decide this, and they all come before the algorithms themselves.
| Clumping+thresholding | LDpred2 | lassosum2 | PRS-CS | |
|---|---|---|---|---|
| Needs a standard error | no | yes | yes | yes |
| Needs an effective sample size | no | yes | yes | yes |
| Extra software | none | bigsnpr | bigsnpr | Python |
| Models per declaration | one per threshold | one, or one per grid point | one per grid point | one per prior value |
| Reference panel format | PLINK2 | bigsnpr | bigsnpr | PLINK2 + LD blocks |
Clumping and thresholding
Keeps the strongest variant in each correlated block and discards the rest, then applies a p-value cut. Weights are the GWAS effect sizes of what survives.
It is the cheapest thing that can succeed, and the best first thing to run: no standard error, no sample size, no extra software.
A vector of thresholds gives one model per threshold, and the clumping is done once for all of them.
LDpred2
A Bayesian method that shrinks effect sizes toward zero according to how much of the trait is genetic and how sparse the signal is.
Three modes. auto estimates the parameters from the data and returns one model — the best default if you do not have a tuning set. grid returns one model per parameter combination and expects you to choose later. inf assumes every variant contributes.
Heritability is estimated for you unless you supply it.
lassosum2
A penalised regression on the same LD resource LDpred2 uses, which is why running both is cheaper than running either twice.
It returns one model per point in its penalty grid.
PRS-CS
A Bayesian method with a continuous shrinkage prior, fitted by the upstream Python implementation.
Unlike the others, each value of its shrinkage parameter is an independent full MCMC run. Two values means two runs. Budget accordingly; this is not an interactive-session job.
It needs approximately independent LD blocks, which ship for African, East Asian and European panels only. Other panels need the window-based alternative.
What a grid costs
Grid sizes are larger than they look. The default lassosum2 grid produces up to 120 models from one declaration, and the default LDpred2 grid 90.
The cost is not only in generation. Every candidate is a cached resource, a column in your score matrix, and a row in every evaluation — so 120 candidates against two GWAS is 240 score columns to compute and compare.
Panels, variant spaces and cost
Reference panels ship for one build; asking for the other triggers a one-time liftover, and the bigsnpr methods trigger a one-time format conversion. Both happen before any LD is computed.
Restricting the panel to a standard variant set makes LD construction much cheaper. The display name and the catalog key both work: "HapMap3+" and "hapmap3plus" name the same variant space and produce the same restricted panel. A variant space the catalog does not know is rejected when you declare it, not when it runs.
Changing an LD parameter rebuilds the whole LD resource, which is expensive. Keeping those parameters fixed lets LDpred2 and lassosum2 share one build.
Where a plausible-but-wrong model comes from
An ancestry mismatch between your GWAS and your panel does not error. It degrades the weights, and the only signal is the count of matched variants in the log.
Every parameter you pass is part of the model's identity, except the ones that only affect how it runs. So 120 grid points are 120 distinct cached models — which is the point, but also the bill.
No tuning happens here
Generation never sees genotypes or phenotypes. It cannot know which of your candidates is best, and it does not try. That is chapter 11.