Contents
genome-chromosomes
Canonical chromosome coding
One place where chromosome codes are interpreted. pg.chr.resolve() is the
primitive: it maps codes to a requested scheme and reports what it could not
map, without ever aborting or dropping. The two policies on top of it decide
what to do with the unmappable remainder:
pg.chr.canonical()for identity/join paths — normalizes recognized codes and passes unrecognized contig labels through unchanged, so a non-standard contig still matches itself;pg.chr.recode()for consumers that need a closed set (PLINK numeric, bigsnpr, PRS-CS) — drops what it cannot map and records it as provenance.
The same file holds the one "chr:start-end" region grammar and the membership
predicate built on it, pg.in.region(), because a region names a chromosome and
must canonicalize it the same way everything else here does.
Details
Supported output schemes mirror the export formats of PLINK 2.0 (https://www.cog-genomics.org/plink/2.0/data#export): "26" — Always numeric (X/Y/XY/PAR1/PAR2/MT -> numeric codes)
"M" — Autosomal numeric; X/Y/M single character
"MT" — Autosomal numeric; X/Y/MT bare labels
"0M" — Autosomal numeric; 0X/0Y two-character; MT
"chr26" — "chr"+numeric code; PAR1/PAR2 unchanged
"chrM" — "chr"-prefixed; MT -> "chrM"; PAR1/PAR2 unchanged
"chrMT" — "chr"-prefixed canonical label; PAR1/PAR2 unchanged
A code is classified as one of
"ok" — recognized: an autosome 1..n.autosomes, X, Y, XY,
PAR1, PAR2, M/MT, or the PLINK numeric equivalent, in any of
PLINK's spellings and with or without a chr prefix
"missing" — no usable coordinate at all: NA, "", "NA", ".",
"-". Never joinable, and always safe to drop
"unrecognized" — a label that is not a standard chromosome but may
still be meaningful to the caller: scaffolds and decoys
("GL000209.1"), composite labels ("1_q21"), "0" (PLINK's unplaced
marker), junk