PolyGenius
Contents

PGS

PGS: a single polygenic score model

PGS is an R6 class holding one polygenic score model: a data.table of variants plus package- and user-defined metadata (GWAS metadata, generation logs, liftover logs, and whatever else a rule or import path attached as a named field).

[ subsets the variants table by rows (i) and columns (j) and returns a new model. x[], with neither index, returns an unmodified deep clone.

$ resolves against the model's own fields and methods first, then falls through to the variants columns.

.DollarNames() offers the variants column names followed by the model's own fields and methods as tab-completion candidates.

Each dplyr verb applies to the variants of a deep-cloned model and returns that model, so the input is never modified in place and every metadata field is carried over. Two exceptions: summarise() returns the summarised tibble rather than a model, and group_by() leaves variants as a grouped_df until ungroup().

The matrix-like operations delegate to the underlying variants table. nrow() and length() are the variant count; ncol() is the column count. dimnames() is assembled from rownames() and colnames(), so its row component is the character sequence "1", "2", ... rather than NULL.

Usage

PGS(variants, name, build, gwas = list(), generation = list(), ...)

S3 method for class 'PGS'

x[i, j, drop = FALSE]

S3 method for class 'PGS'

x$name

S3 method for class 'PGS'

.DollarNames(x, pattern = "")

S3 method for class 'PGS'

filter(.data, ...)

S3 method for class 'PGS'

select(.data, ...)

S3 method for class 'PGS'

mutate(.data, ...)

S3 method for class 'PGS'

arrange(.data, ...)

S3 method for class 'PGS'

slice(.data, ...)

S3 method for class 'PGS'

rename(.data, ...)

S3 method for class 'PGS'

transmute(.data, ...)

S3 method for class 'PGS'

group_by(
  .data,
  ...,
  .add = FALSE,
  .drop = dplyr::group_by_drop_default(.data$variants)
)

S3 method for class 'PGS'

ungroup(.data, ...)

S3 method for class 'PGS'

summarise(.data, ...)

S3 method for class 'PGS'

nrow(x)

S3 method for class 'PGS'

ncol(x)

S3 method for class 'PGS'

dim(x)

S3 method for class 'PGS'

length(x)

S3 method for class 'PGS'

rownames(x)

S3 method for class 'PGS'

colnames(x)

S3 method for class 'PGS'

dimnames(x)

Arguments

ArgumentDescription
variantsA data.frame or tibble of variant rows. Must include columns chr, position, ea, nea, beta; converted in place with data.table::setDT(), but never deduplicated in place. Rows repeating a variant (same canonical chromosome, position, nea and ea) are resolved on the returned model: exact copies collapse to their first row, and a variant with conflicting betas is removed. Either case warns, naming the model and both counts.
nameCharacter scalar. For PGS(), the model identifier; for $, the name of the field, method, or variants column to extract.
buildCharacter scalar. Genome build label, resolved through workspace$catalogs$genomeBuilds: only "GRCh37"/"hg19" and "GRCh38"/"hg38" are accepted, and stored as its canonical name ("hg19/GRCh37" or "hg38/GRCh38"). Anything else errors, including NULL, NA, "" and the PGS Catalog's "NR".
gwasNamed list, default list(). GWAS metadata; attached as $gwas only when non-empty.
generationNamed list, default list(). Model-generation metadata; attached as $generation only when non-empty.
...For PGS(), further named arguments attached to the model as fields of that name. For the dplyr verbs, that verb's own arguments, passed straight to the matching dplyr function.
xA PGS.
iInteger or logical row index, or missing. Selects variants rows.
jInteger, logical or character column index, or missing. Selects variants columns.
dropLogical scalar, default FALSE. Ignored: [ always returns a PGS, never a bare vector or column.
patternCharacter scalar, default "". Regular expression filtering the candidates; "" keeps all of them.
.dataA PGS.
.addLogical scalar, default FALSE. Passed to dplyr::group_by(); TRUE adds to the existing grouping rather than replacing it.
.dropLogical scalar, default dplyr::group_by_drop_default(.data$variants). Passed to dplyr::group_by(); drops groups formed by unused factor levels.

Value

A new PGS whose variants are unique by (chr, position, nea, ea), with chr compared after canonicalization.

For [, a new PGS whose variants are the subset, carrying x's metadata fields. x is not modified.

For $, the field, method, or variants column called name, or NULL when the model has neither.

For .DollarNames(), a character vector of matching completion candidates.

For a dplyr verb other than summarise(), a new PGS whose variants carry the verb's result. .data is not modified.

For summarise(), the summarised table dplyr::summarise() produces -- a tibble, not a PGS.

For nrow(), ncol() and length(), an integer scalar; for dim(), an integer vector of length 2 (variants, columns); for rownames() and colnames(), character vectors; for dimnames(), a two-element list of those two.

Details

The object behaves like its variants table. Subset rows and columns with [, reach variant columns through $, and apply the dplyr verbs below. Every verb and [ operate on a deep clone and return it, so the input model is never modified in place, and every metadata field is carried onto the result.

I/O

$write.pgs.file(path, mode = "extended") writes the model as a PGS Catalog scoring file and generate$from.pgs.file(path) reads it back. The standard header and variant table are readable by any tool that consumes PGS Catalog scoring files; "extended" adds #polygenius_* extension keys so PolyGenius round-trips a model losslessly.

Operators and helpers

  • [ subsets the variants table, returning a new model.
  • $ returns a variant column or a model field.
  • .DollarNames() completes both fields and column names.

dplyr verbs

  • filter(), select(), mutate(), arrange(), slice(), rename(), transmute()
  • group_by(), ungroup(), summarise(). summarise() returns a tibble, not a model.

Matrix-like operations

  • nrow(), ncol(), dim(), length(), rownames(), colnames(), dimnames()

Contract

A bare PGS is never absorbed into a PolyGeniusBackbone, so it carries none of the sorted, mask-filterable guarantees a PGSLibrary provides -- the dplyr verbs bypass this class's constructor, so such a guarantee could not survive the next verb call. Reach for a PGSLibrary when you need it; compute$scores() normalizes a bare PGS into one before scoring. Full account in objects-map.md, section "Why a bare PGS is deliberately plain, and PGSLibrary is not".

PGS() is the only path that guarantees unique variants. [, the dplyr verbs, a reassigned $variants, liftover() and as.PGS() materialization all bypass it. Scoring and its coverage diagnostics resolve any repeat left that way under the same rule.

Public fields

variants — Variant rows as a data.table, always carrying at least chr, position, ea, nea, beta. A data.frame or tibble handed to the constructor is converted with data.table::setDT(), which converts in place rather than copying.

name — Character scalar. Model identifier.

build — Character scalar. Canonical genome build name, "hg19/GRCh37" or "hg38/GRCh38", whichever alias was supplied.

Methods

Public methods

  • PGS.class$new()
  • PGS.class$append()
  • PGS.class$write.pgs.file()
  • PGS.class$print()
  • PGS.class$clone()

Method new()

Initialize a new PGS.

Usage

PGS.class$new(variants, name, build, gwas = list(), generation = list(), ...)

Arguments

variants — A data.frame or tibble of variant rows. Must include columns chr, position, ea, nea, beta; converted with data.table::setDT().

name — Character scalar. Model identifier; must be non-missing and non-empty after trimming whitespace.

build — Character scalar. Genome build label, resolved through workspace$catalogs$genomeBuilds: only "GRCh37"/"hg19" and "GRCh38"/"hg38" are accepted, and stored as its canonical name (genomeBuilds$name()). Anything else errors, including NULL, NA, "" and the PGS Catalog's "NR".

gwas — Named list, default list(). GWAS metadata; attached as $gwas only when non-empty.

generation — Named list, default list(). Model-generation metadata; attached as $generation only when non-empty. Written by the engine for a generated model, and by the import path for a PGS Catalog model (source, method, publication provenance).

... — Further named arguments, each attached to the model as a field of that name.

Returns

A new PGS. Aborts unless name is a single non-missing, non-empty string (after trimming whitespace); it is stored coerced to character. Does not deduplicate variants; PGS() does.

Method append()

Set a field on the model, or one named element inside a list field. Modifies the model in place (R6 reference semantics).

Usage

PGS.class$append(field, key = NULL, value)

Arguments

field — Character scalar. Field name on the model; created when absent.

key — Character scalar, or NULL (default). Element name inside the list field; NULL replaces the whole field with value.

value — Any value. Assigned to the field, or to key within it.

Returns

self, invisibly.

Method write.pgs.file()

Write the model to a PGS Catalog scoring file, creating or overwriting it.

Usage

PGS.class$write.pgs.file(path, mode = c("extended", "strict"))

Arguments

path — Character scalar. Destination .txt or .txt.gz path.

mode — One of "extended" (default), "strict". "extended" writes the standard PGS Catalog header plus #polygenius_* extension keys and any non-standard variant columns, so the model round-trips losslessly via generate$from.pgs.file(). "strict" writes standard header and columns only and validates the spec-required fields; use it for sharing or PGS Catalog deposition.

Returns

path, invisibly.

Method print()

Print a one-line header, the variants table, and a row per metadata field the model carries.

Usage

PGS.class$print(...)

Arguments

... — Further arguments passed to the variants table's own print() method (e.g. nrows, width).

Returns

self, invisibly. Called for its console output.

Method clone()

The objects of this class are cloneable with this method.

Usage

PGS.class$clone(deep = FALSE)

Arguments

deep — Whether to make a deep clone.

Examples

variants <- tibble::tibble(
  chr = c("1", "1", "2"),
  position = c(101, 202, 303),
  ea = c("A", "G", "T"),
  nea = c("G", "A", "C"),
  beta = c(0.1, -0.2, 0.05)
)

m <- PGS(variants, name = "toy", build = "GRCh38")
m$append("gwas", "study", "Test GWAS")
m

# Verbs return a new model; `m` is unchanged.
m2 <- dplyr::filter(m, beta > 0)
nrow(m2)
nrow(m)

# First row, two columns of the variants table.
m3 <- m[1, c("chr", "position")]
dim(m3)


```r

# PGS Catalog scoring-file round trip.
tf <- tempfile(fileext = ".txt")
m$write.pgs.file(tf)
m4 <- generate$from.pgs.file(tf)

See Also

Aliases: PGS, PGS.class, [.PGS, $.PGS, .DollarNames.PGS, filter.PGS, select.PGS, mutate.PGS, arrange.PGS, slice.PGS, rename.PGS, transmute.PGS, group_by.PGS, ungroup.PGS, summarise.PGS, nrow.PGS, ncol.PGS, dim.PGS, length.PGS, rownames.PGS, colnames.PGS, dimnames.PGS