Data Architecture
How PolyGeniusStudy, the shared backbone, models, genotypes and results are structured
This chapter is about structure rather than usage. If you want to run an analysis, the guide covers the same objects from the outside. Here we look at why they are shaped the way they are, which is the part you need when you extend PolyGenius or debug something that does not fit.
The design problem
A polygenic score analysis accumulates a lot of related tables. You have samples with phenotypes and covariates. You have models with provenance and variant weights. You have a score matrix that indexes both. You have relatedness between samples, similarity between models, association results that reference particular models and particular outcomes, and plot-ready summaries derived from each of those.
Kept as loose objects in a script, the relationships between these are held in your head and in variable naming. Every filter you apply has to be applied consistently to everything, and nothing tells you when you have missed one. That is the failure this architecture exists to prevent.
A second problem sits alongside the first, and it is the one that actually shapes this design: several cohorts scored against the same PRS library share an enormous amount of content — the same variant identities, the same model provenance, the same GWAS pass-through columns. A design that copies that content into every cohort object pays for it N times and, worse, lets two "identical" copies quietly drift apart. The architecture below is built to hold that shared content exactly once.
Definition vs realization: the organizing principle
Before the axes, one distinction underlies everything that follows, and it is worth naming explicitly because it is easy to reach for "sample" and "model" as the only two categories that matter. They are not. The real split is between what a value means — true of a model or a variant no matter which cohort is asking — and what a value is, computed against one particular cohort's samples.
Definitional content answers "what is this model/variant, independent of any
cohort": a model's variant table and weights, which GWAS and algorithm produced it,
a variant's chromosome and position and alleles. This lives on the backbone —
PolyGeniusBackbone, the shared object every study built from the same model library
points at. It never depends on which samples you handed it.
Realization content answers "what did this cohort's samples produce": a score
matrix, kinship, a fitted association, a projected principal component. This lives on
the study — PolyGeniusStudy — because it does not exist until you have samples
to compute it from, and a different cohort scored against the identical model library
produces a different realization.
The practical test, when you are unsure where something belongs: would this value be
the same if you swapped in a different cohort's genotypes but kept the same model
library? If yes, it is definitional and belongs on the backbone (or reads through to
it). If no, it is a realization and belongs on the study. study$library$backbone's
variant dictionary is the same object, byte for byte, whether the study behind it has
5,000 samples or 500,000 — that is the property this split is designed to guarantee.
Two axes, plus the shared library they share
PolyGeniusStudy is organised around two axes — samples and models — plus a
third, variants, that is not a study-owned axis at all but a per-study overlay onto
the model library's own shared variant dictionary. Every rectangular slot is aligned
to one of these.
| Slot | Aligned to | Contains |
|---|---|---|
samples |
samples (rows) | the primary table, table members (e.g. phenotypes), and reduction()-tagged embeddings (e.g. PCA) |
sample.pairs$* |
samples × samples | kinship, sample similarity |
scores$* |
samples × models | one PGS matrix per score layer |
models |
models (rows) | a per-study overlay you populate explicitly (e.g. copying in a layer's coverage) |
model.pairs$* |
models × models | realization-scoped model-pair statistics (e.g. score correlation) |
variants$* |
one row per variant in the shared model library | a per-study, per-variant overlay — positional, checked on row count only, never re-indexed by name |
associations |
free-form, named | associate$*() output |
evaluations |
free-form, named | evaluate$*() output |
signals |
free-form, named | this study's realization of a genome signal |
uns |
free-form, named | anything else |
The shape of every rectangular slot is validated on assignment. That is what makes
subsetting safe: narrowing the sample axis narrows samples, every sample.pairs, and
every score layer's rows, in one operation; narrowing the model axis does the
equivalent on models/model.pairs/score columns.
The design will look familiar if you have used AnnData in Python. The borrowing is
deliberate — it is a well-tested answer to the same problem, which is keeping many
tables aligned to a small number of shared axes. What is not borrowed from that
precedent is the third row above: variants is real, but it is not a study-private
axis the way samples/models are. It never shrinks or reorders under any subset,
because it is aligned to the backbone's dictionary, not to whichever samples or models
a particular study happens to keep — see Why the backbone holds the variant
dictionary below.
Why there is no per-model variant matrix
A models-by-variants matrix is the obvious naive shape for "which variants does each model use, and with what weight," and it was rejected.
Models do not share a variant set. A clumping-and-thresholding model at 5×10⁻⁸ and an LDpred2 model over HapMap3+ may overlap in a few hundred variants out of a million. A models-by-variants matrix is therefore mostly empty, and at realistic library sizes mostly empty is fatal: a thousand models against a million variants exceeds R's maximum vector length before it holds anything useful.
The backbone's dictionary solves this by storing variant identity once — one row per
distinct (chromosome, position, alleles) — and letting every model that uses that
variant point at the same row rather than repeating it. A model's own "variant table"
is then reconstructed on demand from those pointers plus its own weight column
(PGSLibraryView, below), not stored as a second copy. The consequence to
internalise is the same one the old design's absence of a variant axis pointed at:
there is no operation that filters variants across your whole analysis in one
step, because a model's variant set is its own, resolved through the shared dictionary
rather than shared wholesale with every other model.
PGS and PGSLibrary
A model is a variant table plus provenance. The table carries chromosome, position, effect and other allele, and the per-allele weight, along with whatever else the source provided. The provenance carries where the summary statistics came from, which algorithm and parameters produced the model, what liftover was applied, and the genome build the coordinates are on.
Two invariants are enforced at construction, both because violating them silently produces plausible wrong answers rather than errors:
The weight is always a log-scale per-allele effect. A published score may be on an odds-ratio or hazard-ratio scale; the conversion happens at import so that nothing downstream has to ask. A scale that cannot be interpreted as an effect is refused rather than guessed at.
The build must be one PolyGenius can place — GRCh37/hg19 or GRCh38/hg38. An unresolvable build used to travel all the way to scoring and produce a score with almost no variant overlap; refusing it at the boundary converts a silent wrong answer into an immediate error.
A PGSLibrary is an ordered collection of models with library-level provenance,
including the record of any models that failed during generation. Concatenation with
c() and indexing with [[ behave as you would expect. Display names need not be
unique — grid models routinely share one — so positional indices, not names, are the
reliable identifier within a library.
Why the backbone holds the variant dictionary
A PGSLibrary does not itself hold N separate variant tables. Everything a
model contributes — its variant identities, its weight column, its GWAS and algorithm
provenance — is absorbed into a PolyGeniusBackbone: one shared dictionary,
build-scoped, with one dictionary row per distinct (chromosome, position, alleles)
identity across the entire library. A model that used to "own" its variant table now
instead holds a position into this shared registry (model.slot) plus, on the
backbone's weight store, its own (dict.slot, beta) pairs.
This is the definitional side of the split introduced above: absorbing a model into a
backbone is a write-once event (registry is "never mutated after absorption"), and
every model registered against a backbone reads the same dictionary, not a copy of it.
That sharing is scoped to one library, not to a session. A PolyGeniusStudy does not
point at the PGS library it was handed: its constructor narrows that library into a backbone
the study owns, via backbone.narrow() in R/data-classes-PolyGeniusBackbone.R, which
derives a fresh backbone from the same identities. Two studies built from one PGS library
therefore hold two backbones of equal content rather than one object between them, and
nothing either study does afterwards — absorbing, renaming, subsetting — reaches the
other. Cross-cohort comparability is consequently a claim about content: the same models
over the same variant identities. See the study object
chapter.
The dictionary's real representation
backbone$dictionary is a data.table, one row per distinct variant identity —
(chr, position, nea, ea) — not one row per model and not one row per model-variant
pair. Two hundred overlapping C+T and LDpred2 models still add up to one dictionary
sized by however many distinct identities they collectively touch.
The resident columns are chr.code, position, nea.code, ea.code, eaf,
eaf.source, dict.rank. Chromosome and single-ACGT alleles are packed into one raw
byte each, using the same codes scores.pack.ap() and scores.codec() assign
elsewhere in the scoring path — a dictionary row's codes are directly usable there with
no re-resolution step. dict.rank is not recomputed on read: it is the row's rank
under vkey order, computed once whenever absorb() grows the dictionary, and every
later reader (the score path in particular) trusts the stored value rather than
re-deriving it. dict.rank is whole-dictionary and permanent — it does not change
between calls. The scoring path compacts it, per call, into a smaller variant.idx
numbering just the variants actually selected for that one compute$scores() call;
variant.idx inherits dict.rank's genomic ordering but is not dict.rank itself
(see The Scoring Engine).
Not every allele fits a single byte. An indel, a multi-character allele, or an N
gets the escape sentinel ALLELE.CODE.ESCAPE in nea.code/ea.code, and its real
string lives in backbone$escapes, a small side table keyed by dict.slot rather than
a second code table — indels are frequently unique per row, so a lookup table for them
would overflow long before a real-sized library does.
Growth is append-only: absorb() adds new identities at the end and never shrinks or
renumbers a row that is already there. A row's dict.slot — its physical position — is
therefore stable across every later absorb() call, which is what lets a weight
column recorded against slot 4,201 stay correct no matter how many models are folded in
afterward. (rank.slots() is the one operation that does renumber dict.slot, and it
refuses outright once anything has been mark.held() — see
PolyGeniusBackbone.class$rank.slots() in R/data-classes-PolyGeniusBackbone.R.)
Three derived quantities you will see used constantly — canonical.vkey(),
canonical.ea(), flags() — are deliberately not dictionary columns. Each is
recomputed from the resident columns on every call, via pg.variant.orient(). That
keeps a saveRDS()/.pgd round trip free of a packed key that could round-trip
incorrectly, at the cost of paying a decode on every call that needs one. Source:
R/data-classes-PolyGeniusBackbone.R, PolyGeniusBackbone.class's $dictionary field
doc and its "Compact resident encoding" section.
The backbone's full anatomy
The dictionary is one field among several. A PolyGeniusBackbone holds:
$dictionary— the compact variant table above.$escapes— the non-ACGT allele overflow, keyed bydict.slot.$codec— the lazy chromosome/allele-pair packing state fromscores.codec(), carried acrossabsorb()calls rather than rebuilt each time, so a non-standard contig or non-SNV allele pair keeps the code it was first assigned.$registry— one row per absorbed model, in absorb order:model.slot,nameandbuildas given at absorb time,n.variants,source.id, and the model's owngeneration/gwasmetadata verbatim. Write-once —registryis never mutated after absorption, only appended to.$weights— thePolyGeniusWeightStore(next subsection), one column per registry row.$columns$source— GWAS pass-through columns (se,pval,n,rsid, ...) shared once per source, not once per model in a source's own parameter grid.$variants$annotations— externally attached, study-invariant variant metadata (pathogenicity, nearest gene, a reference-panel frequency), joined once at attach time rather than stored as a raw copy of the annotation source.
One invariant governs the whole object: a backbone holds exactly one genome build, and
it only ever grows by appending — it never shrinks or renumbers what is already there.
This is enforced, not assumed: absorb() refuses — rather than silently coercing — a
model whose build doesn't match the backbone's own, per item, before that item's
position or vkey work begins (PolyGeniusBackbone.class$absorb(),
R/data-classes-PolyGeniusBackbone.R).
This is the literal mechanism behind the guide's observation that subsetting a
study's models leaves "the backbone stays the same object": subsetting narrows what a
study or PGSLibrary points at, but there is no operation on the backbone
itself that removes or renumbers a row, so the object underneath is unaffected either
way. Source: R/data-classes-PolyGeniusBackbone.R.
PolyGeniusWeightStore: sparse, not dense
This is the piece Why there is no per-model variant matrix (above) stops short of:
that section explains why a models-by-variants matrix was rejected and states that "a
model's variant table is reconstructed on demand," without saying what backs the
weight column itself. PolyGeniusWeightStore is the answer.
Each absorbed model's weights are stored as one sparse (dict.slot, beta) pair-vector
— not a column in a dense matrix, and not a copy of the model's original variant table.
PolyGeniusWeightStore.class$add() appends one such pair-vector per model, indexed by
model.slot (the registry position absorb() just assigned); $get(model.slot)
returns it back as list(dict.slot, beta), row-aligned. A model that uses 400,000 of a
backbone's 3,000,000 dictionary rows costs 400,000 pairs, not 3,000,000 mostly-empty
cells.
Two details matter if you are reading or extending this path. First, add() preserves
row order exactly as given, including within-model duplicate dict.slot values, so
absorbing a model never changes its nrow. PGS() already removes repeated variants
when a model is built, so a duplicate reaches the store only from a model built another
way, such as one materialized from an older library. Scoring collapses whatever is left
under the same rule, before it can reach a --score file: exact copies count once, and a
variant listed with conflicting betas is dropped, with a warning naming the model
(R/genotype-source-scores-store.R; see
The Scoring Engine). Second, a model's stored rows are sorted ascending by
dict.rank once, at absorb time, and that ordering is treated as permanent: later reads
(variants.for(), the score path) trust it rather than re-sorting. When rank.slots()
renumbers the dictionary, it calls $remap() on the whole store, which re-points every
column's dict.slot values through the old-to-new lookup and re-sorts each column by
its new dict.rank — the one case where the store's own content changes without a new
model being absorbed. Source: PolyGeniusBackbone.class$weights field doc and
PolyGeniusWeightStore.class in R/data-classes-PolyGeniusStudy.R.
Storage backends: MemoryBlocks and FileBlocks
A backbone's weight columns are read through a block-store backend, transparent to every
accessor above study$library$backbone:
MemoryBlocks— resident: reads every selected model's(dict.slot, beta)column once at construction and holds the combined table in memory.FileBlocks— sharded.fstfiles inside the session-wide resource store, read a window at a time. Chosen bycompute$scores()when the weight data is too large to hold resident.
FileBlocks is the backend the package itself constructs; MemoryBlocks is the in-memory
reference implementation the test suite checks it against window for window, which is what
makes the contract a specification rather than a description of one implementation.
Both implement the same six-method read contract (region(), row.range(),
model.block(), n.rows(), fingerprint(), descriptor()) defined on
PolyGeniusBackend.class; that contract and the shard/manifest mechanics behind it are the
execution engine's concern, not this page's — see R/internals-block-store.R.
Note that a .pgd file has nothing to do with any of this. Saving decomposes a model
library into plain tables and loading rebuilds it resident; a loaded study is never
backed by a block store. See Persistence and the
Backbone.
PGSLibraryView: a pointer, not a copy
Once absorbed, a model in a PGSLibrary is represented by a
PGSLibraryView — a small S3 record carrying a backbone id and a position, not a
materialized variant table. This exists because the alternative is worse: reading a
model out of the library to inspect, clone or serialize it would otherwise have to choose
between materializing its variants every time (defeating the whole point of sharing
the dictionary) or handing back something that quietly aliases the shared backbone —
at which point a clone of one model, or a saveRDS() of one model, could corrupt data
every other model and every other study built on the same library is reading.
A view therefore refuses everything that would make that ambiguity possible.
`[.PGSLibraryView` refuses outright. `$<-.PGSLibraryView` refuses
every name except a rename that writes through to the owning library's own overlay — it
never touches the backbone. $<column> (e.g. view$beta) refuses too: only a fixed
set of registry-backed fields (name, build, n.variants, generation, gwas,
plus $variants itself) resolve without materializing anything, because letting an
arbitrary column name silently materialize the whole variant table would make ordinary
metadata access an unbounded-cost operation with no visible warning. Every dplyr verb —
filter(), mutate(), select(), arrange(), slice(), rename(), group_by() —
refuses for the same reason a real PGS's own dplyr methods would silently
clone a shared dictionary if a view let them run.
as.PGS(view) is the one explicit crossing back. It materializes the
view's variants (in genomic order, not necessarily the model's original ingest order —
that ordering is not reconstructed) into a real, independent PGS that
owns its own variant table: safe to saveRDS(), mutate, or hand to any of
PGS's existing dplyr-style methods, because it no longer points at
anything shared. Reaching for it is a deliberate, visible step — the point is that no
code path ends up holding both representations of the same model without the caller
having asked for that.
GenotypeSource and GenotypeSourceSet
GenotypeSource is a pointer, not a container. It records a path, one or more file stems,
a format, a genome build, and the working set of sample IDs. Genotypes are never read
into R; every operation delegates to PLINK2 through a single system-call helper.
Two properties are easy to conflate and have very different consequences.
One GenotypeSource may name many files. A genome split by chromosome is still one
cohort. Operations fan out over the files and combine the results according to the
operation's own rule — score sums add, projection numerators and denominators add and
are divided once at the end, and kinship merges the files first because KING needs every marker of a pair at once.
A GenotypeSourceSet concatenates different sample sets. Its observation count is the
sum of its members'. Registering twenty-two chromosomes as a set rather than as one
genotype therefore multiplies your sample count by twenty-two instead of combining
variants — a mistake the shape validation cannot catch, because the resulting object is
internally consistent.
Sample IDs are unique within one genotype but may collide across a set, which is why the merged sample vector carries no names and why relatedness keys on genotype plus ID rather than on ID alone.
Result objects
PolyGeniusAssociation and PolyGeniusEvaluation share a shape:
results— one row per statistical claim. This is the table you read.artifacts— plot-ready derived tables. A confusion matrix, a set of survival curves, a binned score-outcome grid.diagnostics— what went wrong, was dropped, or failed to converge.metadata— analysis-level scalars, stored as given.provenance— the record of what produced the result, read withprovenance(x).
Associations additionally carry fits, one row of invariant metadata per fitted
model, including fits that failed. A results table with fewer rows than your requested
grid is therefore a diagnostic rather than an error, and the reason is recoverable.
Why artifacts are separate
The separation exists to enforce one boundary: plotting never computes a statistic.
If a plot needed to derive a quantity, then the figure and the table could disagree — different filtering, different handling of missing values, a different fitted model. Requiring the analysis to emit whatever a plot needs makes that class of disagreement structurally impossible.
The cost is visible: a stratification result cannot draw a confusion-matrix ROC curve, because performance was never run. Artifacts arrive as a bundle per analysis, and that is the rule that makes which-plot-works-when predictable.
Row-heavy artifacts and the fit key
Artifacts are often much longer than results — a survival curve is hundreds of rows per fit. Storing the identifying columns on every one of those rows would dominate the object's size, so they are stripped and replaced by an integer fit key.
That has one user-visible consequence: reading object$artifacts$name directly gives
you a table with no outcome or stratum column. The accessor rejoins the identity for
you, and is what you should use.
Schemas
An association's result columns are not ad hoc. Each analysis declares a schema: the columns it must produce, the ones it may, which artifacts are required or optional, which plots apply, and which columns identify a poolable cell for meta-analysis.
Nine schemas exist — the four coefficient families, Kaplan–Meier, mediation, comparison, single-variant, and the pooled form. Most share a common core of twenty-two columns, which is why one forest plot, one pooling engine and one filtering vocabulary serve them all.
Two schemas deliberately diverge. single.variant replaces the shared core with
GWAS column names, so a scan interoperates with model variant tables and with the
summary-statistics format on the generation side; the trade is that filters written
against the shared vocabulary do not apply to it. meta is source-agnostic and
records what it came from in its own columns rather than in metadata — which is what
lets a pooled object be pooled again.
The schema registry is an internal surface. It is the right thing to read when adding an analysis; it is not part of the public API.
PolyGeniusEvaluation has no schema. Its results are validated as a table and
nothing more; the semantics of each metric — direction, target, whether it can select
a model — live in a registry instead. This asymmetry is worth knowing before you write
code against an evaluation: you cannot ask an evaluation what it guarantees.
Standalone positioned results
PolyGeniusGenomeSignal holds a positioned signal — a value per genomic position or
per bin — derived from model tables and association summaries rather than from
genotypes.
It is not a slot on PolyGeniusStudy, for two reasons. It has no natural axis in the
sample/model model, being neither per-sample nor per-model — and unlike variants$*,
it is not even study-scoped, since a signal derived from model tables and association
summaries is exactly the kind of content that should be shareable across every study
built on the same backbone. And keeping it standalone lets it be shared: the
constructor refuses per-sample columns outright, so a genome signal cannot
accidentally carry individual-level data across a site boundary.
Ratio statistics store their numerator and denominator separately rather than the ratio, so that re-binning at a coarser resolution sums both and divides once. Storing the ratio would make re-binning an average of averages.
Provenance
Most computed slots carry more than their values. provenance() reads one record saying
what produced them: the operation, the call, the arguments you supplied, and what the
operation resolved — the last of which is what lets a later call check whether a stored
result can be reused, or must be recomputed. compute$relatedness$prune() does exactly
that with a stored kinship matrix's threshold. artifacts() and diagnostics() read the
derived tables and the counts that go with a result.
One record per object, and it is carried rather than restamped: subsetting a study, narrowing a model library or filtering an association all keep the record of the call that made the thing, because narrowing something is not making it. Merging results is the exception — a merged object was not produced by one call, so it carries no record.
One limit is worth knowing. Named result slots are carried through a subset unfiltered — so an association stored before a subset survives it and now describes a sample set the object no longer has. Nothing marks it as stale.
What the architecture does not do
It does not enforce that a phenotype table's rows correspond to the samples it is attached to. Alignment is validated on row count only, so a correctly-sized table in the wrong order is accepted. This is the single highest-consequence gap in the design, and the reason the guide teaches reindexing before construction.
It does not re-check the model-cohort build agreement after construction.
It does not version its stored results. Two analyses that differ only in the version of PolyGenius that produced them are indistinguishable, which matters when an algorithm changes.
Where to go next
The study object covers PolyGeniusStudy from the outside
— construction, the samples plane, subsetting. Persistence and the
backbone covers what a saved .pgd file
physically is and what a model library does and does not keep across a round trip.
Resources and catalogs covers how external data
gets into these objects. The execution engine
covers how work on them is scheduled and cached.