Contents
PolyGeniusBackbone
A shared variant dictionary, grown by appending models
One dictionary row per distinct variant identity (chr, position, nea, ea), shared by every
model absorbed into it. An absorbed model no longer owns its own variant table; it points into
shared dictionary rows instead, so overlapping models pay for a variant's identity once.
PolyGeniusBackbone() is the plain-function constructor: build alone makes an empty
backbone, identities and friends make a populated one via $adopt().
PolyGeniusBackbone() constructs one. With build alone it returns an empty backbone,
identical to PolyGeniusBackbone.class$new(build). With identities it returns a populated
one: $adopt() installs the parts and derives all key state itself. That state is the codec,
the packed dictionary, $escapes, dict.vkey, dict.lut and dict.rank; the caller supplies
none of it. Absorbing models one at a time is the alternative for a caller without a whole
dictionary.
Usage
PolyGeniusBackbone(
build,
identities = NULL,
weights = NULL,
registry = NULL,
columns = NULL,
annotations = list()
)Arguments
| Argument | Description |
|---|---|
build | A genome build label, validated against workspace$catalogs$genomeBuilds. |
identities | A data.frame/data.table of chr, position, nea, ea, one row per distinct variant, or NULL for an empty backbone. |
weights | A [PolyGeniusWeightStore](/reference/polygeniusweightstore/) whose columns are in identities' own row numbering, or NULL for an empty store. |
registry | A data.table in $registry's schema, one row per column in weights, or NULL. |
columns | Components of $columns to install, named as that field names them (source, model); each named component replaces its own and the rest stay empty. NULL starts empty. |
annotations | A named list of PolyGeniusVariantAnnotation descriptors, attached after the parts are installed so each join runs against the final dictionary layout. |
Value
A PolyGeniusBackbone.
Who holds and reaches this object
A backbone is never constructed or held directly by user code. It lives one hop below
PGSLibrary$backbone (or $backbones[[build]] for a set spanning more than one genome
build), PGSLibrary being the user-facing handle carrying name, generation, fingerprint
and a registry row per model pointing at (backbone, model.slot). A PolyGeniusStudy reaches
it only through study$library$backbone, as the single source of truth for everything
study-invariant: variant identity, model-pair definitions, GWAS/algorithm annotations.
Sharing the backbone is what collapses N studies' libraries into one dictionary plus N
per-study overlays.
PGSLibraryView (set$models[[i]]) is the read-only per-model window: it holds
(backbone, position) and materializes a model's variant table on demand, copying what it
returns so a caller's edit never writes back into the shared dictionary.
Why one shared object, not one per model
M models overlapping on V variants cost one dictionary of size V plus M sparse weight columns, reasoning one level up, for N studies sharing one model library.
Compact resident encoding
$dictionary keeps only chr.code, position, nea.code, ea.code and dict.rank
resident: a variant's identity, and where it sorts. Chromosome and single-ACGT alleles each
pack into one raw byte,
using the same codes scores.pack.ap() and scores.codec() already assign elsewhere in the
package, so a dictionary row's allele codes are directly usable by those functions with no
re-resolution. A non-ACGT allele (an indel, a multi-character allele, an N) gets the escape
sentinel ALLELE.CODE.ESCAPE and its real string lives in $escapes, keyed by row position
rather than by a second small code table -- see ALLELE.CODE.ESCAPE's own comment for why.
canonical.vkey(), canonical.ea() and flags() are never stored: every call recomputes them
from the resident columns via pg.variant.orient(). vkey() is cached (see Growth) but
never as a $dictionary column, so a saveRDS()/readRDS() round trip carries no packed
identity key that could round-trip incorrectly.
Growth
absorb() grows the dictionary by appending, and a row's dict.slot (its physical position)
never moves once assigned, so every existing weight column stays valid across growth. What does
move is dict.rank, each row's rank when the dictionary is ordered by vkey, recomputed after
any absorb() that added a row.
A private vkey cache mirrors $dictionary row for row, grown by appending the keys already
computed for the newly-absorbed rows rather than re-deriving them from the resident columns.
That keeps absorb() proportional to what it adds: no existing row is ever decoded or
re-encoded by a later call. The one remaining whole-dictionary cost is the dict.rank sort
(order() over the vkey cache), unavoidable because a new row can sort anywhere.
rank.slots() renumbers dict.slot itself into rank order, once, while the dictionary is
still append-ordered and nothing outside this object has recorded a dict.slot against it.
After it, dict.rank is the identity permutation until the next growing absorb(). A backbone
mark.held() has flagged as externally referenced by dict.slot refuses to renumber, since
doing so would silently invalidate that reference.
What the constructor checks
The parts arrive pre-assembled and this runs once per assembly, never per model, so validation
cost is not the constraint. What it checks is what would otherwise pass silently forever: every
weight column's dict.slot is bounds-checked, because a stale slot re-points a beta at a
different variant with nothing downstream able to notice, and registry must be row-aligned
with weights, the two indexing the same model.slot. It does not re-check what a later step
enforces anyway. The codec budgets and the position range are left to scores.codec.keys(),
which $adopt() runs over every row. The semantic content of a registry row is not judged,
because a backbone has no independent way to judge it.
Public fields
build — Resolved genome-build key. One build per dictionary: a variant's packed
identity key encodes its position, so a variant sharing a coordinate across two builds
would collide with a different variant on one dictionary row.
dictionary — data.table, one row per distinct variant, in dict.slot (append)
order: chr.code, position, nea.code, ea.code, dict.rank. Identity and rank
only. A per-variant measurement belongs to whatever observed it -- a GWAS source's
pass-through column on $columns$source, a panel's frequency on an annotation -- since
one dictionary row is shared by every model and every study reading this backbone.
Never carries a packed identity key as a column.
escapes — data.table(dict.slot, nea, ea) holding the real allele string(s) for
any row whose nea.code/ea.code is ALLELE.CODE.ESCAPE. Only rows that need it
appear here.
codec — The variant-key codec. absorb() carries it across calls and grows it, so a
contig or non-SNV allele pair keeps the code it was first assigned; adopt() rebuilds it
canonically from the identities it is handed (scores.codec.canonical()), so an assembly
depends on its content and not on the order the content arrived in.
registry — data.table, one row per absorbed model in absorb order: model.slot
(its position here, 1-based and never reassigned), name and build as given at absorb
time, n.variants, source.id (the resolved GWAS source id this model's pass-through
columns are shared under on $columns$source, NA when unresolvable), list-columns
generation/gwas carrying the model's own metadata verbatim (list() when it had
none), and extra, every other field the model carried (e.g. $liftover), so absorbing
round-trips a model's full field set. This is the definitional record of an absorbed
model: write-once, never mutated afterwards. A per-set overridable field such as a
renamed name lives on PGSLibrary's registry overlay instead, since two sets
can point at the same (backbone, model.slot) and rename it independently.
weights — This backbone's PolyGeniusWeightStore, one column per absorbed
model, indexed by model.slot in $registry.
variants — list(annotations). $annotations is a named list, one entry per source
attached via $attach.annotation(), holding the filtered-join result (matching rows only,
keyed canonical.vkey, oriented columns) rather than the raw source table, which stays
behind as a private descriptor so a later growing absorb() can re-run the join.
Read-only from outside; $attach.annotation() is the write path.
columns — list(source, model). $source is the source plane: a data.table keyed
(source.id, canonical.vkey), one row per distinct variant a GWAS source contributed,
with columns beyond the key growing dynamically as sources bring new ones (se, pval,
n, rsid, ...: C+T's entire pass-through, shared once per source rather than once per
model in its grid). Write-once and content-checked, so absorbing a second model from the
same source with a different value for an already-recorded cell errors rather than
overwriting. $model is the model plane: NULL, or a list indexed by model.slot
whose entry is a data.table of that one model's own columns, NULL where a model has
none. Its rows are positional: one per row of $weights$get(model.slot), in the same
order, carrying no key column of their own. A per-model column is beta's twin: both
are per (model, row), so both are addressed the way beta is, by the model's own row
index. A dict.slot key could not express a model carrying the same variant twice,
which absorb() deliberately preserves; the row index can. The price is that whatever
reorders or drops a model's weight rows must do the same here, so rank.slots() reuses
the permutations $weights$remap() hands back rather than deriving its own, and
backbone.narrow() masks both with one ok.
Active bindings
models — Read-only, study-invariant model-axis content: $annotations, one row per
absorbed model in model.slot order, unnesting $registry's gwas/generation
list-columns. Read through here rather than copied per study, so N studies over one
backbone share it. $registry, written by absorb(), is the only place it is written.
model.pairs — Study-invariant snp.* model-pair definitions (snp.jaccard,
snp.weighted.overlap), square n.models() x n.models(). Realization-only pair
statistics (score.pearson, score.spearman) live on study$model.pairs instead.
instance.id — Read-only character scalar: this object's stable per-process registry
key (see R/data-classes-PGSLibraryView.R's registry comment). Never persisted and
meaningless across a process boundary. It exists so a PGSLibraryView can find the
live backbone without holding a reference saveRDS() would follow.
fingerprint — Read-only character scalar: a content- and layout-sensitive hash of the
dictionary, computed lazily on read and invalidated by absorb(). Two dictionaries built
by absorbing the same models in a different order differ here despite holding the same
variant set, because anything keyed on dictionary-relative positions (a store shard, say)
must know when the layout changed, not only when the variant set did. It covers variant
identity and row order, never the weights: two libraries over one variant set with
different betas share a fingerprint, so a caller comparing model content needs more than
this.
Methods
Public methods
PolyGeniusBackbone.class$new()PolyGeniusBackbone.class$n.variants()PolyGeniusBackbone.class$n.models()PolyGeniusBackbone.class$attach.annotation()PolyGeniusBackbone.class$annotation.sources()PolyGeniusBackbone.class$identities()PolyGeniusBackbone.class$absorb()PolyGeniusBackbone.class$rank.slots()PolyGeniusBackbone.class$rank.is.identity()PolyGeniusBackbone.class$adopt()PolyGeniusBackbone.class$mark.held()PolyGeniusBackbone.class$slot.for()PolyGeniusBackbone.class$union.table()PolyGeniusBackbone.class$vkey()PolyGeniusBackbone.class$canonical.vkey()PolyGeniusBackbone.class$canonical.ea()PolyGeniusBackbone.class$flags()PolyGeniusBackbone.class$print()
Method new()
Create an empty dictionary for one genome build.
Usage
PolyGeniusBackbone.class$new(build)
Arguments
build — A genome build label, validated against
workspace$catalogs$genomeBuilds.
Method n.variants()
Number of distinct variants currently in the dictionary.
Usage
PolyGeniusBackbone.class$n.variants()
Returns
An integer scalar.
Method n.models()
Number of models absorbed into this backbone so far.
Usage
PolyGeniusBackbone.class$n.models()
Returns
An integer scalar.
Method attach.annotation()
Attach a variant annotation source: study-invariant, source-independent
variant metadata (pathogenicity, nearest gene, a per-source frequency, ...) keyed on
canonical.vkey. This method is the only write path; the read side
(backbone$variants$annotations$name) is a plain list lookup.
Performs a filtered join at attach time and never stores a copy of annotation$table.
An annotation source is routinely far larger than the dictionary, so only rows matching a
current dictionary identity are kept, and match/miss counts are reported rather than a
partial join being kept silently. Budget-checked against
workspace$config$max.model.column.bytes, the same dial mutate.variants() uses.
Orientation follows pg.column.orientation.policy(), as the source plane does.
Stated gap: growth re-runs the join over the whole current dictionary, not only the slots appended since the last attach. The source descriptor is kept so a future incremental version can fix that without changing this method's contract.
Usage
PolyGeniusBackbone.class$attach.annotation(name, annotation)
Arguments
name — Character scalar: the name this annotation is attached under.
annotation — A PolyGeniusVariantAnnotation descriptor from
attach.variant.annotation().
Returns
self, invisibly.
Method annotation.sources()
This backbone's attached annotation sources: the raw
PolyGeniusVariantAnnotation descriptors given to $attach.annotation(), not the joined
$variants$annotations result. Read-only, and public so .pgd persistence
(R/data-classes-PolyGeniusStudy-persist.R) can round-trip them and re-attach through
$attach.annotation() on load, re-running each join against the reconstructed backbone's
actual dictionary layout.
Usage
PolyGeniusBackbone.class$annotation.sources()
Returns
A named list of PolyGeniusVariantAnnotation descriptors (list() if none are
attached).
Method identities()
Decoded (chr, position, nea, ea) for every dictionary row, in dict.slot
(row/append) order: result$chr[k] describes dict.slot k. Public because
PGSLibraryView's materialization (variants.for()) resolves a model's stored
dict.slot references back to identities through it. Delegates to the same private
decode() canonical.vkey(), flags() and union.table() use.
Usage
PolyGeniusBackbone.class$identities()
Returns
list(chr, position, nea, ea), each a vector of length n.variants().
Method absorb()
Absorb one or more models into the dictionary, in place, and record each one's own weight column and registry row.
A variant already present keeps its dict.slot; a genuinely new one is appended. New
identities are looked up against a private vkey cache, and only the new ones are encoded,
so no existing row is decoded or re-encoded. dict.rank is recomputed once over the whole
dictionary after every model in the call is folded in, so a multi-model call pays that
sort once.
Deliberately does not call scores.model.vkeys(), which collapses or drops a within-model
repeated identity before keying. That is right for the score path but wrong here: a
model's weight column must keep repeated dict.slot values so variants.for() can
reproduce the model's original row count. Every row is keyed here, duplicates included;
only the dictionary's growth step deduplicates.
Usage
PolyGeniusBackbone.class$absorb(models)
Arguments
models — A PGS, PGSLibrary, or list of PGS.
Every model must resolve to this dictionary's own build.
Returns
self, invisibly.
Method rank.slots()
Renumber dict.slot itself into vkey-rank order, once. Meant to run right
after the last absorb() of a fresh assembly, while the dictionary is still in plain
append order and nothing outside this object has yet recorded a dict.slot value
against it -- see mark.held(). After this call dict.rank is the identity
permutation, rank.is.identity() is TRUE, and it stays TRUE until the next growing
absorb() moves a rank away from its slot again.
Refuses on a held backbone, since renumbering dict.slot out from under an existing
dict.slot-keyed reference would silently invalidate it. $weights, $escapes,
$columns$model and the private vkey lookup are not such references but data this
object owns outright, so all four are remapped rather than refused; held is reserved
for a reference this object cannot fix up itself (a future out-of-core file-backed
store, say).
Usage
PolyGeniusBackbone.class$rank.slots()
Returns
self, invisibly.
Method rank.is.identity()
Whether dict.rank is currently the identity permutation (dict.rank[i] == i for every row) -- true right after rank.slots(), false again once a later
absorb() appends a row that sorts before the end. Cheap and uncached; meant for a
caller (the score path) deciding whether a rank-order gather step is needed.
Usage
PolyGeniusBackbone.class$rank.is.identity()
Returns
A logical scalar.
Method adopt()
Install a pre-assembled model library in one step: variant identities, a
weight store, a registry and the source-column planes, from which every derived piece of
key state follows. The bulk-load counterpart to absorb(), which instead folds models in
one at a time and grows the dictionary itself. A caller that already holds these tables --
a lift, a narrowing, a load from disk -- would otherwise have to materialize one variant
table per model to re-derive a dictionary it already has.
Identities arrive decoded, and everything that can be derived from them is derived here
rather than accepted: $codec is rebuilt canonically (scores.codec.canonical()), the
dictionary's packed allele codes and $escapes come from the one shared encoder
(backbone.encode.identities()), and dict.vkey/dict.lut/dict.rank follow from
those. Nothing carries a code assigned against some other backbone's codec, so two
callers holding the same identities get the same layout and the same $fingerprint.
chr is canonicalized on the way in, so a caller handing over a contig in a foreign
spelling cannot silently create a second entry for a chromosome that already exists.
Finishes by renumbering dict.slot into vkey order (rank.slots()), which is what
establishes the weight-row ordering every reader relies on. A fresh backbone is never
held, so that can never be refused here.
Refuses on a non-empty backbone: adopting over live content would leave $weights and
$registry describing one dictionary and $dictionary another, with no partial state
worth defining. A caller replacing content builds a fresh backbone instead.
Usage
PolyGeniusBackbone.class$adopt(
identities,
weights = NULL,
registry = NULL,
columns = NULL,
annotations = list()
)
Arguments
identities — A data.frame/data.table with one row per distinct variant and columns
chr, position, nea, ea. Alleles as given; chr in either raw or canonical form.
weights — A PolyGeniusWeightStore, or NULL for an empty one.
registry — A data.table in $registry's own schema, or NULL for none.
columns — Components of $columns to install, named as that field names them
(source, model); each named component replaces its own, the rest are left as the
fresh empty ones. NULL installs nothing.
annotations — A named list of PolyGeniusVariantAnnotation descriptors to attach after
the parts are installed, so each one's join runs against this dictionary's layout.
Returns
self, invisibly.
Method mark.held()
Record that something outside this object now holds a dict.slot-keyed
reference into it, so rank.slots() refuses from here on. Irreversible. Internal, called
by whatever machinery first builds such a reference.
Usage
PolyGeniusBackbone.class$mark.held()
Returns
self, invisibly.
Method slot.for()
Look up dict.slot (the 1-based row position absorb() assigned, stable
under growth) for each queried variant identity. Internal diagnostic, and not for a hot
path: it decodes the whole dictionary on every call. Materializing a model's own rows is
PGSLibraryView's job; this exists so a caller can reconstruct one model's
dict.slot/beta pairs without the lazy materialization machinery.
Usage
PolyGeniusBackbone.class$slot.for(chr, position, nea, ea)
Arguments
chr, position, nea, ea — Equal-length vectors describing the identities to look up, in
the same convention as a model's own variant table (alleles as given, chr in either
raw or canonical form -- canonicalized here before matching).
Returns
An integer vector, one dict.slot per query row; NA_integer_ where the
identity is not in the dictionary.
Method union.table()
The decoded variant union, in the same shape
scores.index.models()'s variants.to.score returns: (chr, position, nea, ea, variant.idx), keyed by identity. variant.idx here is dict.rank.
Usage
PolyGeniusBackbone.class$union.table()
Returns
A keyed data.table.
Method vkey()
Packed identity key for every dictionary row, in dict.slot order. Served
from a private cache absorb() maintains incrementally, not recomputed per call. Never a
$dictionary column, so it cannot leak into a serialized form by that route.
Usage
PolyGeniusBackbone.class$vkey()
Returns
An integer64 vector, one per row.
Method canonical.vkey()
Canonical identity key: alleles ordered lexicographically, nea the smaller
and ea the larger, rather than as a source gave them. Carries no implied effect allele
of its own; see canonical.ea() for the allele a signed value keyed on this refers to.
Usage
PolyGeniusBackbone.class$canonical.vkey()
Returns
An integer64 vector, one per row.
Method canonical.ea()
The allele canonical.vkey() treats as effect allele: the lexicographically
larger of the two, the convention genome-project.R signs betas against. Anything storing
a value oriented against canonical.vkey() must carry this alongside it, so a join
between two such objects can assert they agree on what a positive sign means. Without it,
a genome signal and a pooled result can silently disagree on sign at every variant.
Usage
PolyGeniusBackbone.class$canonical.ea()
Returns
A character vector, one per row.
Method flags()
Per-row bitfield: bit 1 strand-ambiguous (a palindromic A/T or C/G pair, whose
orientation the alleles alone cannot resolve), bit 2 indel or otherwise not a single-ACGT
pair, bit 4 on a contig outside the standard chromosome set. Derived from
(chr, nea, ea) on every call, never persisted.
Usage
PolyGeniusBackbone.class$flags()
Returns
An integer vector, one per row.
Method print()
Pretty print.
Usage
PolyGeniusBackbone.class$print(...)
Arguments
... — Unused.
Returns
self, invisibly.
See Also
Other pgs-objects:
PGS(),
PGSLibrary(),
PGSLibraryView,
PolyGeniusWeightStore