Resources and Catalogs
How external data — panels, chains, variant spaces, LD blocks, tools — is described, retrieved and reused
A PGS analysis depends on a surprising amount of material that is not your data: a reference panel, liftover chains, a variant set to restrict to, approximately independent LD blocks, and the external binaries that do the work. This chapter is about how PolyGenius describes that material, gets hold of it, and avoids getting hold of it twice.
Two ideas: catalogs and the store
A catalog is a description of what exists. It knows that a 1000 Genomes European panel is available, where it can be fetched from, and what its checksum should be. It does not know whether you have it.
The store is what you have. It is a directory under your workspace root plus a SQLite index, and it records every resource that has been materialised, keyed by what was asked for.
The split matters because it separates two questions that get conflated in script-based workflows: what could I use and what have I already built. A catalog lookup is cheap and offline. A store lookup is what decides whether something needs computing.
The catalogs
Five catalogs are exposed under workspace$catalogs:
Genome builds — the supported assemblies and their aliases. Only GRCh37/hg19 and GRCh38/hg38 exist. This is deliberately small: a build is only useful if a liftover chain and a panel exist for it, and those constrain the set. Anything else is refused at the boundary rather than accepted and failed later.
Reference panels — per-super-population panels derived from 1000 Genomes, filtered to biallelic SNPs above a frequency floor with duplicates removed. They ship for one build; asking for the other triggers a one-time liftover of the panel itself.
Liftover chains — the UCSC chain files, in both directions between the two supported builds.
Variant spaces — named variant sets used to restrict a panel before LD computation. Two ship: a default set of common variants, and the HapMap3+ map that the bigsnpr methods expect. Restricting a panel to one of these is the difference between an LD build that takes minutes and one that takes hours.
LD blocks — approximately independent block boundaries, used by PRS-CS. They exist for African, East Asian and European panels only, which is a real constraint on which panel you can pair with that method.
Two facts about the shipped LD blocks are worth knowing if you are reasoning about provenance rather than just using them. The East Asian set is the original Berisa–Pickrell Asian set under a current label. And every GRCh38 block set is derived by liftover from GRCh37 rather than being defined natively on GRCh38.
Each catalog supports the same two operations: view what exists, and get a resource — where "get" means fetch it if absent, and return the local path if present.
Resource identity
Every resource is identified by a hash of its type and its parameters. Two requests that agree on both are the same resource and will share a stored result.
The consequences of that are worth spelling out, because most surprises about caching reduce to one of them.
Runtime hints are excluded. Thread counts, memory caps and the path to a Python interpreter describe how to compute something, not what is computed, so they are kept outside identity. Changing them costs nothing.
Everything else is included. For a generated model, every algorithm parameter is part of its identity — which is why a 120-point grid is 120 distinct resources rather than one resource with a parameter.
Content is not included, only references to it. A locally-supplied summary statistics table is identified by the name you gave it, not by what is in it. Edit the file, keep the name, and you get the old result with no warning. This is the most common way to get a stale answer in practice, and versioning the name is the fix.
The version of PolyGenius is not included. An identity records what you asked for, not what code answered. That is what makes caching reproducible within a project, and it means upgrading the package does not invalidate anything. When a release note says a construction algorithm changed, generate into a fresh store so the new code actually runs.
Build labels are hashed as given rather than normalised at the identity layer, so callers normalise before constructing a specification. The source rules do this; if you write your own, do the same.
Partial specifications
A request need not pin every parameter. A parameter left unspecified is dropped from the lookup, so a request that does not name a build can match a stored resource that has one.
When several stored resources match a partial request, the most recently created wins. The mapping from the partial request to the concrete resource it resolved to is recorded, so the same request hits the cache directly next time rather than searching again.
One parameter supports relaxed matching rather than exact: the p-value threshold on retrieved summary statistics. A file fetched at a looser threshold satisfies a request for a stricter one, because the stricter set is contained in it. This is the only such parameter in the codebase, and it is worth being precise about what it does and does not mean:
- It does mean you will not re-download summary statistics when you tighten a threshold.
- It does not mean a tightened threshold reuses an old model. A model's p-value parameter is an ordinary exact-match parameter, so a stricter threshold produces a new model.
Confusing those two would make you distrust correct results.
The store
The store lives under your workspace root: a SQLite index that is the authority on what exists, and one directory per resource holding its payload.
Discovery is index-driven. Nothing is found by scanning the filesystem, which is what keeps lookup cheap when the store holds thousands of resources.
A cache hit requires both an index row and the payload on disk. That combination is what makes manual cleanup safe: delete a resource's directory and the next run notices the payload is gone and recomputes it, rather than failing or returning something incomplete.
The index carries a schema version and refuses to open an index written by an incompatible version, with a message telling you to remove it. That is a deliberate choice over silent migration — an index that half-matches is worse than one that clearly does not.
What actually fills a store
Generated analysis products are small. Retrieved summary statistics, clumped variant sets and models are kilobytes each, so even a run producing thousands of models adds tens of megabytes.
What consumes space is external material: reference panels, LD resources derived from them, PLINK binaries and liftover chains — hundreds of megabytes to gigabytes. Those are also the most shared, since every model built against a panel refers to the same one, so they are the last thing you should delete.
Removing things
There is no user-facing deletion API for computed resources. The catalogs support removing catalog entries, and beyond that the supported operations are at the filesystem level: remove a resource's directory to force it to recompute, or remove the whole root to start clean.
Removing a resource that other resources were built from does not invalidate them. They keep their own stored payloads and continue to resolve as cached, so you can end up with a model whose inputs are gone. If you are pruning, prune leaves.
Per-resource logs
Every resource directory carries its own log, and that log is the concatenation of its own output and everything upstream of it. So a model's log contains the summary statistics retrieval, the panel preparation, the LD build and the algorithm run — which is why it is the right place to look when a generation step fails.
External tools
workspace$setup manages the binaries and packages that resources depend on: PLINK2,
GCTB, PRS-CS and its Python environment, and bigsnpr.
Two ways to satisfy a dependency. Download it, by naming the stack:
workspace$setup$install("plink")Or register something you already have, by naming the stack as an argument:
workspace$setup$install(plink = "/opt/plink2/plink2")The second form is the only route on a compute node with no outbound network, which is the common case on a cluster.
Note that install() with no arguments does nothing at all, and reports nothing. It is
not a "install everything" convenience — name what you want.
check() prints what is present and where it came from. status() returns the same
information as a data frame without printing, which is what you want in a script.
Where the reproducibility boundary sits
It is worth being explicit about what this system does and does not guarantee.
It does guarantee that two runs asking for the same thing, on the same store, get the same result — because the second reuses the first.
It does not guarantee that the same request on two different machines produces identical output. Some computations are randomised without a fixed seed, notably the approximate decomposition used for principal components. Some depend on the version of an external tool, which is recorded but not part of identity.
So a stored result is a reliable record of what your project computed, not a cryptographic guarantee that it is the only possible answer to that request. Record your PolyGenius version and your external tool versions alongside a result set, for the same reason you record which reference panel you used.
Where to go next
The execution engine covers how requests for resources become scheduled work.