Contents
as.bigsnpr
Convert a GenotypeSource to bigsnpr file-backed format
Writes the genotypes a GenotypeSource() points at into bigsnpr's file-backed
.rds/.bk pair and attaches the result. This is what makes a PolyGenius genotype fileset
usable by bigsnpr-based algorithms such as LDpred2 and lassosum2.
Usage
as.bigsnpr(
genotype.source,
output = NULL,
output.name = NULL,
ind.row = NULL,
ind.col = NULL,
ncores = 1,
nthreads = NULL,
logger = NULL
)Arguments
| Argument | Description |
|---|---|
genotype.source | A GenotypeSource. Fileset to convert; its $samples working set selects the samples read. |
output | Character scalar or NULL (default). Directory the .rds and .bk files are written to, created if absent; NULL uses tempdir(). |
output.name | Character scalar or NULL (default). Base name for the output files, without extension; NULL uses genotype.source$name. |
ind.row | Integer vector or NULL (default). Row (sample) indices passed to bigsnpr's snp_readBed2(), applied to the fileset that remains after the working-sample filter; NULL reads every sample in genotype.source$samples. |
ind.col | Integer vector or NULL (default). Column (variant) indices passed to snp_readBed2(); NULL reads every variant. |
ncores | Integer scalar, default 1. Cores bigsnpr uses for the conversion step. |
nthreads | Positive integer scalar or NULL (default). PLINK --threads count for the preparation steps; NULL passes no --threads flag. |
logger | Optional PolyGeniusLogger, default NULL. Scopes the progress messages; NULL creates one. |
Value
A named list:
obj — The attached bigSNP object.
rds — Character scalar path to the written .rds file.
map — Data frame of variants aligned to the genotype columns, with chr
(character), pos (integer), a0 (reference allele), a1 (effect allele),
marker.id (standardized PolyGenius ID) and genetic.dist.
fam — Data frame of sample information, as bigsnpr reports it.
genotypes — The file-backed matrix (FBM), samples x variants.
build — Character scalar genome build, copied from genotype.source.
Aborts when genotype.source is not a GenotypeSource, when nthreads is neither NULL
nor a positive integer, or when bigsnpr is not installed.
Details
The input is prepared with PLINK2 before bigsnpr sees it: converted to bfile format with
$tidy() when it is not already, merged into one fileset with $merge() when it holds more
than one, and filtered with --keep when the working sample set is smaller than the fileset
or when ind.row is supplied. Those intermediates are written to a temporary directory that
is removed when the call returns; only the .rds and .bk files land in output.
Variant identifiers in map$marker.id follow the PolyGenius standard, chr:pos:a0:a1 with
the alleles in lexicographic order, so they match the IDs used elsewhere in the package (for
example when clumping). Chromosome codes are normalised to PLINK numeric coding; a code that
cannot be resolved is left verbatim rather than dropped, because map rows are positionally
aligned with the genotype matrix columns.
Requires the bigsnpr package, which is a soft dependency: install it with
install.packages("bigsnpr").
Examples
# Convert a reference panel to bigsnpr format
ref.panel <- workspace$catalogs$referencePanels$get("EUR", "hg19", format = "pfile")
bigsnp.obj <- as.bigsnpr(ref.panel)
# The file-backed genotype matrix, samples x variants
G <- bigsnp.obj$genotypes
dim(G)
# Keep the conversion outside tempdir()
bigsnp.perm <- as.bigsnpr(
ref.panel,
output = "/path/to/permanent/storage",
output.name = "EUR.ref.hg19"
)See Also
Other genotypes:
GenotypeSource(),
GenotypeSourceSet()