PolyGenius
Contents

GenotypeSourceSet

Create a new GenotypeSourceSet

Groups several GenotypeSource objects into one set that fans a call out across all of them. Members are held by reference and keyed by their own $name; like a single GenotypeSource, a set never loads genotypes into memory.

Flattens any mix of GenotypeSource, GenotypeSourceSet and shallow lists of those into one set. An element of any other type aborts, as do duplicate genotype names.

Returns the underlying named list of GenotypeSource objects, by reference, so base iterators such as lapply() and Map() work on a set.

Usage

GenotypeSourceSet(...)

S3 method for class 'GenotypeSourceSet'

x[[i]]

S3 method for class 'GenotypeSource'

c(...)

S3 method for class 'GenotypeSourceSet'

c(...)

S3 method for class 'GenotypeSourceSet'

length(x)

S3 method for class 'GenotypeSourceSet'

names(x)

S3 method for class 'GenotypeSourceSet'

as.list(x, ...)

Arguments

ArgumentDescription
...One or more GenotypeSource or GenotypeSourceSet objects, or lists of them, flattened into one set. Member names must be unique, and an empty set aborts. Unused by as.list().
xA GenotypeSourceSet.
iInteger and/or character vector selecting genotypes (1-based positions, names, or a mix), or missing to select all. Duplicates and out-of-range positions are dropped, order is preserved, and a selection that matches nothing aborts.

Value

A GenotypeSourceSet object holding every supplied genotype.

For a selection of one genotype, that GenotypeSource itself, by reference. Otherwise a new GenotypeSourceSet holding the selected genotypes.

A new GenotypeSourceSet holding every genotype supplied, in argument order.

Integer scalar: the number of genotypes in the set.

Character vector of the genotype names, in set order.

A named list of the contained GenotypeSource objects, in set order.

Class

The R6 generator behind GenotypeSourceSet(). Holds a flat, name-keyed set of GenotypeSource objects and fans a call out across all of them; like its members it never loads genotypes into memory.

Members are held by reference and keyed by their own $name, which must be unique across the set. Each fan-out method calls the identically named GenotypeSource method on every member and returns a named list, names matching the set; the samples argument of those methods is per genotype, all other arguments are shared. When the set holds one genotype, or when .return.split.by.genotype = FALSE, the per-genotype results are combined into one object instead -- see each method's @return.

[[, c(), length(), names() and as.list() are documented below.

Active bindings

names — Character vector of genotype names, in set order. Read-only.

genotypes — Named list of the contained GenotypeSource objects. Read-only, and held by reference: mutating a returned object mutates the set's own member.

samples — Named list of working sample IIDs, one character vector per genotype. Read/write.

samples.full — Named list of every fileset sample IID, one character vector per genotype. Read-only.

n.samples — Named integer vector, working-set size per genotype. Read-only.

n.samples.full — Named integer vector, full fileset size per genotype. Read-only.

Methods

Public methods

  • GenotypeSourceSet.class$new()
  • GenotypeSourceSet.class$print()
  • GenotypeSourceSet.class$dosages()
  • GenotypeSourceSet.class$scores()
  • GenotypeSourceSet.class$samples.kinship()
  • GenotypeSourceSet.class$samples.PCA()
  • GenotypeSourceSet.class$samples.PCAProject()
  • GenotypeSourceSet.class$clone()

Method new()

Setting delegates to each genotype's own $samples, so an IID absent from a fileset is dropped silently. NULL resets every genotype to its full fileset. Any other value must be a named list whose names match $names, each element a character vector or NULL; anything else aborts.

Create a new GenotypeSourceSet.

Usage

GenotypeSourceSet.class$new(...)

Arguments

... — One or more GenotypeSource objects, or lists of them. Each member is keyed by its own $name; argument names are ignored. Aborts on an empty set, on a non-GenotypeSource member, or on duplicate names.

Returns

A new GenotypeSourceSet object.

Examples

\dontrun{
gss$samples <- list(G1 = c("S01", "S02"), G2 = c("A10", "A11"))
gss$samples <- NULL
}

Method print()

Print one line per genotype: name, working and full sample counts, path, file count, format and build. A path that no longer exists is marked unreachable.

Usage

GenotypeSourceSet.class$print(...)

Arguments

... — Unused. Present for S3 method compatibility.

Returns

self, invisibly. Called for its console output.

Method dosages()

Extract dosages for given variants from every genotype in the set.

Usage

GenotypeSourceSet.class$dosages(
  variants,
  format = c("sample", "variant"),
  samples = NULL,
  load = FALSE,
  merge = FALSE,
  .message = NULL,
  .return.split.by.genotype = TRUE,
  logger = NULL
)

Arguments

variants — Character vector, data frame, or bed1-like table. Variants to extract, as for GenotypeSource$dosages().

format — One of "sample" or "variant". Orientation of the exported table. Must be passed as a single value: the declared default is rejected by this method's own length check.

samples — Named list of character vectors, or NULL (default). Sample IIDs per genotype; NULL uses each genotype's own working set. Names must match $names.

load — Logical scalar, default FALSE. When TRUE, read the exported tables instead of returning their paths.

merge — Logical scalar, default FALSE. With load = TRUE, combine each genotype's per-file tables into one.

.message — Character scalar or NULL (default). Progress message logged per genotype, with NAME bound to the genotype name.

.return.split.by.genotype — Logical scalar, default TRUE. When FALSE, or when the set holds one genotype, results are combined across genotypes.

logger — Optional PolyGeniusLogger, default NULL. Propagated to each GenotypeSource call.

Returns

Split (default, more than one genotype): a named list of that genotype's own $dosages() return. Combined: file paths when load = FALSE; with format = "sample" a data.table of all genotypes row-bound under a genotype column; with format = "variant" one table merged on chr, variant, (C)M, position, counted, alt, sample columns prefixed by genotype name.

Method scores()

Compute polygenic scores for every genotype in the set with PLINK2 --score.

Usage

GenotypeSourceSet.class$scores(
  models,
  samples = NULL,
  maf.threshold = 0,
  max.cells.per.batch = 2e+08,
  models.per.group = Inf,
  nthreads = NULL,
  memory = NULL,
  file.backed.threshold = 2.5e+08,
  model.filter.columns = NULL,
  .message = NULL,
  .return.split.by.genotype = TRUE,
  logger = NULL
)

Arguments

models — A PGS, a PGSLibrary, a data frame of variants, or a list of either. Scored identically against every genotype.

samples — Named list of character vectors, or NULL (default). Sample IIDs per genotype; NULL uses each genotype's own working set. Names must match $names.

maf.threshold — Numeric scalar in [0, 0.5], default 0. Minor-allele-frequency floor applied per genotype.

max.cells.per.batch — Numeric scalar, default 200e6. Cap on variant x model cells per PLINK --score call, forwarded to each genotype; controls peak memory.

models.per.group — Positive numeric scalar, default Inf. Models scored per group (model-axis streaming); a finite value lowers peak records at the cost of re-reading each genotype once per group. Scores are identical either way.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count, forwarded to each genotype.

memory — Positive numeric scalar (MB) or NULL (default). PLINK --memory cap, forwarded to each genotype; NULL leaves PLINK's default.

file.backed.threshold — Numeric scalar, default 2.5e8. Backbone-backed input only: the dense backbone$n.variants() x M cell count above which the scored union is read out-of-core from the weight store rather than from the in-memory backbone.

model.filter.columns — Internal only, not user-facing. NULL (default), or a model.filter mask pushdown -- see GenotypeSourceScores$compute()'s own @param. Forwarded identically to every genotype, since a model's filtered rows do not depend on which fileset is read.

.message — Character scalar or NULL (default). Progress message logged per genotype, with NAME bound to the genotype name.

.return.split.by.genotype — Logical scalar, default TRUE. When FALSE, or when the set holds one genotype, results are combined across genotypes.

logger — Optional PolyGeniusLogger, default NULL. Propagated to each GenotypeSource call.

Returns

Split (default, more than one genotype): a named list of that genotype's own $scores() return. Combined: a list with scores (numeric matrix, all genotypes' samples row-bound x models, no rownames, columns labelled positionally m1..mM), given.names (those labels -- every genotype scores the same models, so one describes the merge), variant.fate (per-genotype fates row-bound under a genotype column), strategy ("via.extract" if any genotype used it, else "fused") and n.batches (summed across genotypes).

Method samples.kinship()

Report KING-robust kinship for related sample pairs within every genotype. Kinship is never computed across genotypes.

Usage

GenotypeSourceSet.class$samples.kinship(
  ...,
  samples = NULL,
  .message = NULL,
  .return.split.by.genotype = TRUE,
  logger = NULL
)

Arguments

... — Further arguments passed to GenotypeSource$samples.kinship(), for example threshold, variants, nthreads, memory.

samples — Named list of character vectors, or NULL (default). Sample IIDs per genotype; NULL uses each genotype's own working set. Names must match $names.

.message — Character scalar or NULL (default). Progress message logged per genotype, with NAME bound to the genotype name.

.return.split.by.genotype — Logical scalar, default TRUE. When FALSE, or when the set holds one genotype, results are combined across genotypes.

logger — Optional PolyGeniusLogger, default NULL. Propagated to each GenotypeSource call.

Returns

Split (default, more than one genotype): a named list of per-genotype pair tables. Combined: one data.table of those tables stacked, with a leading genotype column; a genotype with no related pairs contributes no rows.

Method samples.PCA()

Compute PCA within every genotype with PLINK2 --pca.

Usage

GenotypeSourceSet.class$samples.PCA(
  ...,
  samples = NULL,
  .message = NULL,
  .return.split.by.genotype = TRUE,
  logger = NULL
)

Arguments

... — Further arguments passed to GenotypeSource$samples.PCA(), for example npcs, variants, approx, nthreads, memory.

samples — Named list of character vectors, or NULL (default). Sample IIDs per genotype; NULL uses each genotype's own working set. Names must match $names.

.message — Character scalar or NULL (default). Progress message logged per genotype, with NAME bound to the genotype name.

.return.split.by.genotype — Logical scalar, default TRUE. When FALSE, or when the set holds one genotype, results are combined across genotypes.

logger — Optional PolyGeniusLogger, default NULL. Propagated to each GenotypeSource call.

Returns

Split (default, more than one genotype): a named list of that genotype's own $samples.PCA() return. Combined: a list with embedding (numeric matrix, all genotypes' samples row-bound x PCs, no rownames) and eigenvalues, a named list of one numeric vector per genotype -- the eigendecomposition is in-sample, so axes are not comparable across genotypes.

Method samples.PCAProject()

Project every genotype's samples onto a principal-component space anchored on another GenotypeSource, with PLINK2 --score.

Usage

GenotypeSourceSet.class$samples.PCAProject(
  ...,
  samples = NULL,
  .message = NULL,
  .return.split.by.genotype = TRUE,
  logger = NULL
)

Arguments

... — Further arguments passed to GenotypeSource$samples.PCAProject(), for example project.onto, npcs, variants, approx, nthreads, memory.

samples — Named list of character vectors, or NULL (default). Sample IIDs per genotype; NULL uses each genotype's own working set. Names must match $names.

.message — Character scalar or NULL (default). Progress message logged per genotype, with NAME bound to the genotype name.

.return.split.by.genotype — Logical scalar, default TRUE. When FALSE, or when the set holds one genotype, results are combined across genotypes.

logger — Optional PolyGeniusLogger, default NULL. Propagated to each GenotypeSource call.

Returns

Split (default, more than one genotype): a named list of that genotype's own $samples.PCAProject() return. Combined: a list with embedding (numeric matrix, all genotypes' samples row-bound x PCs, no rownames) and variants, a named list of the variants each genotype's projection used.

Method clone()

The objects of this class are cloneable with this method.

Usage

GenotypeSourceSet.class$clone(deep = FALSE)

Arguments

deep — Whether to make a deep clone.

Examples

## ------------------------------------------------
## Method `GenotypeSourceSet.class$new`
## ------------------------------------------------


```r

gss$samples <- list(G1 = c("S01", "S02"), G2 = c("A10", "A11"))
gss$samples <- NULL

See Also

Other genotypes: GenotypeSource(), as.bigsnpr()

Aliases: GenotypeSourceSet, GenotypeSourceSet.class, [[.GenotypeSourceSet, c.GenotypeSource, c.GenotypeSourceSet, length.GenotypeSourceSet, names.GenotypeSourceSet, as.list.GenotypeSourceSet