PolyGenius
Contents

PolyGeniusBackbone

A shared variant dictionary, grown by appending models

One dictionary row per distinct variant identity (chr, position, nea, ea), shared by every model absorbed into it. An absorbed model no longer owns its own variant table; it points into shared dictionary rows instead, so overlapping models pay for a variant's identity once.

PolyGeniusBackbone() is the plain-function constructor: build alone makes an empty backbone, identities and friends make a populated one via $adopt().

PolyGeniusBackbone() constructs one. With build alone it returns an empty backbone, identical to PolyGeniusBackbone.class$new(build). With identities it returns a populated one: $adopt() installs the parts and derives all key state itself. That state is the codec, the packed dictionary, $escapes, dict.vkey, dict.lut and dict.rank; the caller supplies none of it. Absorbing models one at a time is the alternative for a caller without a whole dictionary.

Usage

PolyGeniusBackbone(
  build,
  identities = NULL,
  weights = NULL,
  registry = NULL,
  columns = NULL,
  annotations = list()
)

Arguments

ArgumentDescription
buildA genome build label, validated against workspace$catalogs$genomeBuilds.
identitiesA data.frame/data.table of chr, position, nea, ea, one row per distinct variant, or NULL for an empty backbone.
weightsA [PolyGeniusWeightStore](/reference/polygeniusweightstore/) whose columns are in identities' own row numbering, or NULL for an empty store.
registryA data.table in $registry's schema, one row per column in weights, or NULL.
columnsComponents of $columns to install, named as that field names them (source, model); each named component replaces its own and the rest stay empty. NULL starts empty.
annotationsA named list of PolyGeniusVariantAnnotation descriptors, attached after the parts are installed so each join runs against the final dictionary layout.

Value

A PolyGeniusBackbone.

Who holds and reaches this object

A backbone is never constructed or held directly by user code. It lives one hop below PGSLibrary$backbone (or $backbones[[build]] for a set spanning more than one genome build), PGSLibrary being the user-facing handle carrying name, generation, fingerprint and a registry row per model pointing at (backbone, model.slot). A PolyGeniusStudy reaches it only through study$library$backbone, as the single source of truth for everything study-invariant: variant identity, model-pair definitions, GWAS/algorithm annotations. Sharing the backbone is what collapses N studies' libraries into one dictionary plus N per-study overlays. PGSLibraryView (set$models[[i]]) is the read-only per-model window: it holds (backbone, position) and materializes a model's variant table on demand, copying what it returns so a caller's edit never writes back into the shared dictionary.

Why one shared object, not one per model

M models overlapping on V variants cost one dictionary of size V plus M sparse weight columns, reasoning one level up, for N studies sharing one model library.

Compact resident encoding

$dictionary keeps only chr.code, position, nea.code, ea.code and dict.rank resident: a variant's identity, and where it sorts. Chromosome and single-ACGT alleles each pack into one raw byte, using the same codes scores.pack.ap() and scores.codec() already assign elsewhere in the package, so a dictionary row's allele codes are directly usable by those functions with no re-resolution. A non-ACGT allele (an indel, a multi-character allele, an N) gets the escape sentinel ALLELE.CODE.ESCAPE and its real string lives in $escapes, keyed by row position rather than by a second small code table -- see ALLELE.CODE.ESCAPE's own comment for why.

canonical.vkey(), canonical.ea() and flags() are never stored: every call recomputes them from the resident columns via pg.variant.orient(). vkey() is cached (see Growth) but never as a $dictionary column, so a saveRDS()/readRDS() round trip carries no packed identity key that could round-trip incorrectly.

Growth

absorb() grows the dictionary by appending, and a row's dict.slot (its physical position) never moves once assigned, so every existing weight column stays valid across growth. What does move is dict.rank, each row's rank when the dictionary is ordered by vkey, recomputed after any absorb() that added a row.

A private vkey cache mirrors $dictionary row for row, grown by appending the keys already computed for the newly-absorbed rows rather than re-deriving them from the resident columns. That keeps absorb() proportional to what it adds: no existing row is ever decoded or re-encoded by a later call. The one remaining whole-dictionary cost is the dict.rank sort (order() over the vkey cache), unavoidable because a new row can sort anywhere.

rank.slots() renumbers dict.slot itself into rank order, once, while the dictionary is still append-ordered and nothing outside this object has recorded a dict.slot against it. After it, dict.rank is the identity permutation until the next growing absorb(). A backbone mark.held() has flagged as externally referenced by dict.slot refuses to renumber, since doing so would silently invalidate that reference.

What the constructor checks

The parts arrive pre-assembled and this runs once per assembly, never per model, so validation cost is not the constraint. What it checks is what would otherwise pass silently forever: every weight column's dict.slot is bounds-checked, because a stale slot re-points a beta at a different variant with nothing downstream able to notice, and registry must be row-aligned with weights, the two indexing the same model.slot. It does not re-check what a later step enforces anyway. The codec budgets and the position range are left to scores.codec.keys(), which $adopt() runs over every row. The semantic content of a registry row is not judged, because a backbone has no independent way to judge it.

Public fields

build — Resolved genome-build key. One build per dictionary: a variant's packed identity key encodes its position, so a variant sharing a coordinate across two builds would collide with a different variant on one dictionary row.

dictionary — data.table, one row per distinct variant, in dict.slot (append) order: chr.code, position, nea.code, ea.code, dict.rank. Identity and rank only. A per-variant measurement belongs to whatever observed it -- a GWAS source's pass-through column on $columns$source, a panel's frequency on an annotation -- since one dictionary row is shared by every model and every study reading this backbone. Never carries a packed identity key as a column.

escapes — data.table(dict.slot, nea, ea) holding the real allele string(s) for any row whose nea.code/ea.code is ALLELE.CODE.ESCAPE. Only rows that need it appear here.

codec — The variant-key codec. absorb() carries it across calls and grows it, so a contig or non-SNV allele pair keeps the code it was first assigned; adopt() rebuilds it canonically from the identities it is handed (scores.codec.canonical()), so an assembly depends on its content and not on the order the content arrived in.

registry — data.table, one row per absorbed model in absorb order: model.slot (its position here, 1-based and never reassigned), name and build as given at absorb time, n.variants, source.id (the resolved GWAS source id this model's pass-through columns are shared under on $columns$source, NA when unresolvable), list-columns generation/gwas carrying the model's own metadata verbatim (list() when it had none), and extra, every other field the model carried (e.g. $liftover), so absorbing round-trips a model's full field set. This is the definitional record of an absorbed model: write-once, never mutated afterwards. A per-set overridable field such as a renamed name lives on PGSLibrary's registry overlay instead, since two sets can point at the same (backbone, model.slot) and rename it independently.

weights — This backbone's PolyGeniusWeightStore, one column per absorbed model, indexed by model.slot in $registry.

variants — list(annotations). $annotations is a named list, one entry per source attached via $attach.annotation(), holding the filtered-join result (matching rows only, keyed canonical.vkey, oriented columns) rather than the raw source table, which stays behind as a private descriptor so a later growing absorb() can re-run the join. Read-only from outside; $attach.annotation() is the write path.

columns — list(source, model). $source is the source plane: a data.table keyed (source.id, canonical.vkey), one row per distinct variant a GWAS source contributed, with columns beyond the key growing dynamically as sources bring new ones (se, pval, n, rsid, ...: C+T's entire pass-through, shared once per source rather than once per model in its grid). Write-once and content-checked, so absorbing a second model from the same source with a different value for an already-recorded cell errors rather than overwriting. $model is the model plane: NULL, or a list indexed by model.slot whose entry is a data.table of that one model's own columns, NULL where a model has none. Its rows are positional: one per row of $weights$get(model.slot), in the same order, carrying no key column of their own. A per-model column is beta's twin: both are per (model, row), so both are addressed the way beta is, by the model's own row index. A dict.slot key could not express a model carrying the same variant twice, which absorb() deliberately preserves; the row index can. The price is that whatever reorders or drops a model's weight rows must do the same here, so rank.slots() reuses the permutations $weights$remap() hands back rather than deriving its own, and backbone.narrow() masks both with one ok.

Active bindings

models — Read-only, study-invariant model-axis content: $annotations, one row per absorbed model in model.slot order, unnesting $registry's gwas/generation list-columns. Read through here rather than copied per study, so N studies over one backbone share it. $registry, written by absorb(), is the only place it is written.

model.pairs — Study-invariant snp.* model-pair definitions (snp.jaccard, snp.weighted.overlap), square n.models() x n.models(). Realization-only pair statistics (score.pearson, score.spearman) live on study$model.pairs instead.

instance.id — Read-only character scalar: this object's stable per-process registry key (see R/data-classes-PGSLibraryView.R's registry comment). Never persisted and meaningless across a process boundary. It exists so a PGSLibraryView can find the live backbone without holding a reference saveRDS() would follow.

fingerprint — Read-only character scalar: a content- and layout-sensitive hash of the dictionary, computed lazily on read and invalidated by absorb(). Two dictionaries built by absorbing the same models in a different order differ here despite holding the same variant set, because anything keyed on dictionary-relative positions (a store shard, say) must know when the layout changed, not only when the variant set did. It covers variant identity and row order, never the weights: two libraries over one variant set with different betas share a fingerprint, so a caller comparing model content needs more than this.

Methods

Public methods

  • PolyGeniusBackbone.class$new()
  • PolyGeniusBackbone.class$n.variants()
  • PolyGeniusBackbone.class$n.models()
  • PolyGeniusBackbone.class$attach.annotation()
  • PolyGeniusBackbone.class$annotation.sources()
  • PolyGeniusBackbone.class$identities()
  • PolyGeniusBackbone.class$absorb()
  • PolyGeniusBackbone.class$rank.slots()
  • PolyGeniusBackbone.class$rank.is.identity()
  • PolyGeniusBackbone.class$adopt()
  • PolyGeniusBackbone.class$mark.held()
  • PolyGeniusBackbone.class$slot.for()
  • PolyGeniusBackbone.class$union.table()
  • PolyGeniusBackbone.class$vkey()
  • PolyGeniusBackbone.class$canonical.vkey()
  • PolyGeniusBackbone.class$canonical.ea()
  • PolyGeniusBackbone.class$flags()
  • PolyGeniusBackbone.class$print()

Method new()

Create an empty dictionary for one genome build.

Usage

PolyGeniusBackbone.class$new(build)

Arguments

build — A genome build label, validated against workspace$catalogs$genomeBuilds.

Method n.variants()

Number of distinct variants currently in the dictionary.

Usage

PolyGeniusBackbone.class$n.variants()

Returns

An integer scalar.

Method n.models()

Number of models absorbed into this backbone so far.

Usage

PolyGeniusBackbone.class$n.models()

Returns

An integer scalar.

Method attach.annotation()

Attach a variant annotation source: study-invariant, source-independent variant metadata (pathogenicity, nearest gene, a per-source frequency, ...) keyed on canonical.vkey. This method is the only write path; the read side (backbone$variants$annotations$name) is a plain list lookup.

Performs a filtered join at attach time and never stores a copy of annotation$table. An annotation source is routinely far larger than the dictionary, so only rows matching a current dictionary identity are kept, and match/miss counts are reported rather than a partial join being kept silently. Budget-checked against workspace$config$max.model.column.bytes, the same dial mutate.variants() uses. Orientation follows pg.column.orientation.policy(), as the source plane does.

Stated gap: growth re-runs the join over the whole current dictionary, not only the slots appended since the last attach. The source descriptor is kept so a future incremental version can fix that without changing this method's contract.

Usage

PolyGeniusBackbone.class$attach.annotation(name, annotation)

Arguments

name — Character scalar: the name this annotation is attached under.

annotation — A PolyGeniusVariantAnnotation descriptor from attach.variant.annotation().

Returns

self, invisibly.

Method annotation.sources()

This backbone's attached annotation sources: the raw PolyGeniusVariantAnnotation descriptors given to $attach.annotation(), not the joined $variants$annotations result. Read-only, and public so .pgd persistence (R/data-classes-PolyGeniusStudy-persist.R) can round-trip them and re-attach through $attach.annotation() on load, re-running each join against the reconstructed backbone's actual dictionary layout.

Usage

PolyGeniusBackbone.class$annotation.sources()

Returns

A named list of PolyGeniusVariantAnnotation descriptors (list() if none are attached).

Method identities()

Decoded (chr, position, nea, ea) for every dictionary row, in dict.slot (row/append) order: result$chr[k] describes dict.slot k. Public because PGSLibraryView's materialization (variants.for()) resolves a model's stored dict.slot references back to identities through it. Delegates to the same private decode() canonical.vkey(), flags() and union.table() use.

Usage

PolyGeniusBackbone.class$identities()

Returns

list(chr, position, nea, ea), each a vector of length n.variants().

Method absorb()

Absorb one or more models into the dictionary, in place, and record each one's own weight column and registry row.

A variant already present keeps its dict.slot; a genuinely new one is appended. New identities are looked up against a private vkey cache, and only the new ones are encoded, so no existing row is decoded or re-encoded. dict.rank is recomputed once over the whole dictionary after every model in the call is folded in, so a multi-model call pays that sort once.

Deliberately does not call scores.model.vkeys(), which collapses or drops a within-model repeated identity before keying. That is right for the score path but wrong here: a model's weight column must keep repeated dict.slot values so variants.for() can reproduce the model's original row count. Every row is keyed here, duplicates included; only the dictionary's growth step deduplicates.

Usage

PolyGeniusBackbone.class$absorb(models)

Arguments

models — A PGS, PGSLibrary, or list of PGS. Every model must resolve to this dictionary's own build.

Returns

self, invisibly.

Method rank.slots()

Renumber dict.slot itself into vkey-rank order, once. Meant to run right after the last absorb() of a fresh assembly, while the dictionary is still in plain append order and nothing outside this object has yet recorded a dict.slot value against it -- see mark.held(). After this call dict.rank is the identity permutation, rank.is.identity() is TRUE, and it stays TRUE until the next growing absorb() moves a rank away from its slot again.

Refuses on a held backbone, since renumbering dict.slot out from under an existing dict.slot-keyed reference would silently invalidate it. $weights, $escapes, $columns$model and the private vkey lookup are not such references but data this object owns outright, so all four are remapped rather than refused; held is reserved for a reference this object cannot fix up itself (a future out-of-core file-backed store, say).

Usage

PolyGeniusBackbone.class$rank.slots()

Returns

self, invisibly.

Method rank.is.identity()

Whether dict.rank is currently the identity permutation (dict.rank[i] == i for every row) -- true right after rank.slots(), false again once a later absorb() appends a row that sorts before the end. Cheap and uncached; meant for a caller (the score path) deciding whether a rank-order gather step is needed.

Usage

PolyGeniusBackbone.class$rank.is.identity()

Returns

A logical scalar.

Method adopt()

Install a pre-assembled model library in one step: variant identities, a weight store, a registry and the source-column planes, from which every derived piece of key state follows. The bulk-load counterpart to absorb(), which instead folds models in one at a time and grows the dictionary itself. A caller that already holds these tables -- a lift, a narrowing, a load from disk -- would otherwise have to materialize one variant table per model to re-derive a dictionary it already has.

Identities arrive decoded, and everything that can be derived from them is derived here rather than accepted: $codec is rebuilt canonically (scores.codec.canonical()), the dictionary's packed allele codes and $escapes come from the one shared encoder (backbone.encode.identities()), and dict.vkey/dict.lut/dict.rank follow from those. Nothing carries a code assigned against some other backbone's codec, so two callers holding the same identities get the same layout and the same $fingerprint. chr is canonicalized on the way in, so a caller handing over a contig in a foreign spelling cannot silently create a second entry for a chromosome that already exists.

Finishes by renumbering dict.slot into vkey order (rank.slots()), which is what establishes the weight-row ordering every reader relies on. A fresh backbone is never held, so that can never be refused here.

Refuses on a non-empty backbone: adopting over live content would leave $weights and $registry describing one dictionary and $dictionary another, with no partial state worth defining. A caller replacing content builds a fresh backbone instead.

Usage

PolyGeniusBackbone.class$adopt(
  identities,
  weights = NULL,
  registry = NULL,
  columns = NULL,
  annotations = list()
)

Arguments

identities — A data.frame/data.table with one row per distinct variant and columns chr, position, nea, ea. Alleles as given; chr in either raw or canonical form.

weights — A PolyGeniusWeightStore, or NULL for an empty one.

registry — A data.table in $registry's own schema, or NULL for none.

columns — Components of $columns to install, named as that field names them (source, model); each named component replaces its own, the rest are left as the fresh empty ones. NULL installs nothing.

annotations — A named list of PolyGeniusVariantAnnotation descriptors to attach after the parts are installed, so each one's join runs against this dictionary's layout.

Returns

self, invisibly.

Method mark.held()

Record that something outside this object now holds a dict.slot-keyed reference into it, so rank.slots() refuses from here on. Irreversible. Internal, called by whatever machinery first builds such a reference.

Usage

PolyGeniusBackbone.class$mark.held()

Returns

self, invisibly.

Method slot.for()

Look up dict.slot (the 1-based row position absorb() assigned, stable under growth) for each queried variant identity. Internal diagnostic, and not for a hot path: it decodes the whole dictionary on every call. Materializing a model's own rows is PGSLibraryView's job; this exists so a caller can reconstruct one model's dict.slot/beta pairs without the lazy materialization machinery.

Usage

PolyGeniusBackbone.class$slot.for(chr, position, nea, ea)

Arguments

chr, position, nea, ea — Equal-length vectors describing the identities to look up, in the same convention as a model's own variant table (alleles as given, chr in either raw or canonical form -- canonicalized here before matching).

Returns

An integer vector, one dict.slot per query row; NA_integer_ where the identity is not in the dictionary.

Method union.table()

The decoded variant union, in the same shape scores.index.models()'s variants.to.score returns: (chr, position, nea, ea, variant.idx), keyed by identity. variant.idx here is dict.rank.

Usage

PolyGeniusBackbone.class$union.table()

Returns

A keyed data.table.

Method vkey()

Packed identity key for every dictionary row, in dict.slot order. Served from a private cache absorb() maintains incrementally, not recomputed per call. Never a $dictionary column, so it cannot leak into a serialized form by that route.

Usage

PolyGeniusBackbone.class$vkey()

Returns

An integer64 vector, one per row.

Method canonical.vkey()

Canonical identity key: alleles ordered lexicographically, nea the smaller and ea the larger, rather than as a source gave them. Carries no implied effect allele of its own; see canonical.ea() for the allele a signed value keyed on this refers to.

Usage

PolyGeniusBackbone.class$canonical.vkey()

Returns

An integer64 vector, one per row.

Method canonical.ea()

The allele canonical.vkey() treats as effect allele: the lexicographically larger of the two, the convention genome-project.R signs betas against. Anything storing a value oriented against canonical.vkey() must carry this alongside it, so a join between two such objects can assert they agree on what a positive sign means. Without it, a genome signal and a pooled result can silently disagree on sign at every variant.

Usage

PolyGeniusBackbone.class$canonical.ea()

Returns

A character vector, one per row.

Method flags()

Per-row bitfield: bit 1 strand-ambiguous (a palindromic A/T or C/G pair, whose orientation the alleles alone cannot resolve), bit 2 indel or otherwise not a single-ACGT pair, bit 4 on a contig outside the standard chromosome set. Derived from (chr, nea, ea) on every call, never persisted.

Usage

PolyGeniusBackbone.class$flags()

Returns

An integer vector, one per row.

Method print()

Pretty print.

Usage

PolyGeniusBackbone.class$print(...)

Arguments

... — Unused.

Returns

self, invisibly.

See Also

Aliases: PolyGeniusBackbone, PolyGeniusBackbone.class