PolyGenius
Contents

ResourceStore

Resource Store

ResourceStore owns the cache layout for generated resources. Resource payloads remain file-backed under .polygenius/<type>/<id>/, while the root SQLite index .polygenius/index.sqlite stores the canonical registry of resources, searchable parameter/metadata fields, and resource relations.

Details

SQLite-backed resource graph used by the execution engine.

The cache index has four tables:

  • runs(run_id, status, started_at) tracks execution lifecycle. status is "running", "done", or "failed". Used to scope virtual-resource cleanup to the owning execution run.
  • resources(type, id, serializer, path, run_id, created_at) stores one row per known resource. path = NULL denotes a virtual resource. run_id is NULL for all durable concrete resources and is never touched by cleanup; virtual resources carry the run_id of the execution that created them. created_at is set on insert via DEFAULT (datetime('now')) and preserved through upserts -- find() uses it to prefer the most recently created resource when multiple candidates satisfy all match conditions.
  • fields(type, id, kind, name, value, text, hash) stores parameter and metadata fields in long form. kind is "param" or "meta", value is a serialized R object BLOB, text is a human-readable preview, and hash is an xxhash64 digest of the field value. This table is the canonical source of truth for both params and metadata -- load.meta() reads from here, not from any sidecar file. Only scalar metadata values are stored; non-scalar values emit a developer warning and must live in the resource data payload.
  • relations("from.type", "from.id", relation, "to.type", "to.id", name) stores durable graph edges. Relation kinds: "depends_on", "resolved_to".

Concrete resource folders contain:

  • spec.rds: a plain-list descriptor of the ResourceSpec (used for fast reconstruction). Never the live R6 object: serializing that would persist its method bodies, which a later build reads back and binds to the current namespace.
  • serializer-managed data files
  • log.txt: merged upstream and rule logs when available

Metadata is stored only in the fields table. meta.rds sidecar files are not written or read. Parameters define resource identity. Metadata is indexed for find() queries and catalog projection. Internal dot-prefixed fields are not written to the field table.

Public inspection methods return table projections from SQLite: resources(), fields(), relations(), resolved.outputs(), and view.index(). view.index(type) is a compatibility projection assembled from the resource and field tables for catalog code.

Used by execution workers that cache and reuse a store across tasks to detect a connection that has been closed or invalidated.

Skips reading spec.rds and log files. The serializer only needs type and id to locate the data file, so a minimal spec constructed from the DB row is sufficient.

Identical to load.data() but accepts a row already returned by resource.rows.bulk(), avoiding the per-resource SQL round-trip entirely.

Called by the scheduler's discovery loop to flush all depends_on relations accumulated during a discovery pass in a single batched write instead of one auto-commit per dependency.

Exposed so callers (e.g. the scheduler's collect loop) can batch multiple save() and add.relation() calls into one transaction. RSQLite uses savepoints for nested calls, so wrapping an operation that already calls save() internally is safe.

A virtual request (no concrete file of its own) is already satisfied when it recorded one or more resolved_to outputs and every distinct output still exists on disk as a concrete resource. The scheduler uses this to treat a previously expanded rule (single canonical or dynamic multi-output) as cached instead of re-running it. Returns NULL when nothing was recorded or any recorded output is missing, so a partially evicted expansion is regenerated in full.

One INSERT ... ON CONFLICT and one dbAppendTable for the whole batch instead of a pair per resource. The cost of both is dominated by per-CALL overhead, not per-row: measured dbAppendTable 6.75 ms for 13 rows and 10.71 ms for 130, so folding a 10-resource collect batch into one call takes replace.fields from ~15.9 ms per resource to ~1.2 ms. Batching only the transaction saves nothing (measured 10.43 ms/resource) -- it is the statement count that costs.

Must be called inside the caller's transaction. Skipping it leaves payloads on disk with no index row, which discovery cannot see at all (it is DB-only), so the failure mode is repeated work, never a corrupt or half-visible resource.

Reads from the fields DB table (kind = 'meta'). This is a single SQL round-trip. For loading metadata for many resources at once use load.meta.bulk().

Issues a single SELECT against the fields table via a temporary lookup table, returning all kind = 'meta' entries for every requested id at once. Intended for the execution engine's discovery pass, where bind() needs metadata from many upstream specs -- one load.meta.bulk() call replaces N individual load.meta() calls.

Builds a single SQL query with one JOIN on the fields table per filterable parameter and required.meta entry. The matching strategy for each parameter is read from spec$param.match (default "exact"):

  • "exact": hash equality using the fields_cover index. Most efficient.
  • "numeric.gte": CAST(stored_text AS REAL) >= requested_value. A cached resource with a more permissive stored value satisfies the request -- use for thresholds like pval.max where a cached superset is reusable.
  • "numeric.lte": CAST(stored_text AS REAL) <= requested_value.

Parameters with NA values are skipped and match any stored value. required.meta entries use "exact" unless the value is numeric, in which case "numeric.gte" applies (cached capability >= requested minimum).

When multiple resources satisfy all conditions the most recently created resource wins (ORDER BY created_at DESC LIMIT 1).

Calls find() per spec, each issuing a single SQL query (see find() for matching semantics). Intended for the scheduler's discovery pass to resolve many cache lookups in sequence without the overhead of the iterative multi-query approach used in the old implementation.

Public fields

root — Character. Root directory for cached resources.

Methods

Public methods

  • ResourceStore$new()
  • ResourceStore$close()
  • ResourceStore$connected()
  • ResourceStore$print()
  • ResourceStore$type.dir()
  • ResourceStore$resource.dir()
  • ResourceStore$data.path()
  • ResourceStore$meta.path()
  • ResourceStore$spec.path()
  • ResourceStore$log.path()
  • ResourceStore$resources()
  • ResourceStore$fields()
  • ResourceStore$relations()
  • ResourceStore$spec.from.row()
  • ResourceStore$concrete.exists.row()
  • ResourceStore$resource.rows.bulk()
  • ResourceStore$relations.bulk()
  • ResourceStore$load.data()
  • ResourceStore$load.data.from.row()
  • ResourceStore$add.relation()
  • ResourceStore$add.relations.batch()
  • ResourceStore$with.transaction()
  • ResourceStore$resolved.outputs()
  • ResourceStore$resolved.virtual.outputs()
  • ResourceStore$view.index()
  • ResourceStore$exists()
  • ResourceStore$exists.id()
  • ResourceStore$save()
  • ResourceStore$save.virtual()
  • ResourceStore$flush.index()
  • ResourceStore$write.log()
  • ResourceStore$cache.file()
  • ResourceStore$remove.ids()
  • ResourceStore$load()
  • ResourceStore$load.meta()
  • ResourceStore$load.meta.bulk()
  • ResourceStore$find()
  • ResourceStore$find.bulk()
  • ResourceStore$clone()

Method new()

Create a resource store.

Usage

ResourceStore$new(root)

Arguments

root — Character; cache root directory.

Method close()

Close the SQLite index connection.

Usage

ResourceStore$close()

Method connected()

Report whether the SQLite index connection is open and valid.

Usage

ResourceStore$connected()

Returns

Logical scalar.

Method print()

Print resource-store summary.

Usage

ResourceStore$print(...)

Arguments

... — Unused.

Method type.dir()

Return a type directory.

Usage

ResourceStore$type.dir(type)

Arguments

type — Character resource type.

Method resource.dir()

Return the resource directory.

Usage

ResourceStore$resource.dir(spec)

Arguments

spec — Resource specification.

Method data.path()

Return path to resource data.

Usage

ResourceStore$data.path(spec)

Arguments

spec — Resource specification.

Method meta.path()

Return path to resource metadata.

Usage

ResourceStore$meta.path(spec)

Arguments

spec — Resource specification.

Method spec.path()

Return path to resource specification.

Usage

ResourceStore$spec.path(spec)

Arguments

spec — Resource specification.

Method log.path()

Return path to resource text logs.

Usage

ResourceStore$log.path(spec)

Arguments

spec — Resource specification.

Method resources()

Query the global resource table.

Usage

ResourceStore$resources(type = NULL, id = NULL)

Arguments

type — Optional resource type filter.

id — Optional resource id filter.

Method fields()

Query parameter and metadata fields.

Usage

ResourceStore$fields(type = NULL, id = NULL, kind = NULL, name = NULL)

Arguments

type — Optional resource type filter.

id — Optional resource id filter.

kind — Optional field kind ("param" or "meta").

name — Optional field name.

Method relations()

Query resource relations.

Usage

ResourceStore$relations(
  from.type = NULL,
  from.id = NULL,
  to.type = NULL,
  to.id = NULL,
  relation = NULL
)

Arguments

from.type — Optional source resource type filter.

from.id — Optional source resource id filter.

to.type — Optional target resource type filter.

to.id — Optional target resource id filter.

relation — Optional relation kind filter.

Method spec.from.row()

Reconstruct a ResourceSpec from a raw DB row without issuing a new SQL query. Intended for callers that already hold a row from resource.rows.bulk().

Usage

ResourceStore$spec.from.row(row)

Arguments

row — Single-row data.frame with at least type, id, serializer, and path columns (as returned by resource.rows.bulk).

Returns

ResourceSpec, or NULL on error.

Method concrete.exists.row()

Check whether a resource whose DB row is already known is physically present on disk, without issuing a new SQL query.

Usage

ResourceStore$concrete.exists.row(row)

Arguments

row — Single-row data.frame from resource.rows.bulk().

Returns

Logical.

Method resource.rows.bulk()

Bulk-fetch resource rows for many IDs in one round-trip.

Usage

ResourceStore$resource.rows.bulk(ids)

Arguments

ids — Character vector of resource ids to look up.

Returns

data.frame with columns requested_id, type, id, serializer, path. Rows where the resource is absent have NA in all columns except requested_id.

Method relations.bulk()

Bulk-fetch resolved_to children for many virtual resource ids.

Usage

ResourceStore$relations.bulk(from.ids, relation = "resolved_to")

Arguments

from.ids — Character vector of resource ids to resolve.

relation — Character relation kind (default "resolved_to").

Returns

data.frame with columns requested_from_id, to.type, to.id, to.serializer, to.path. Only rows where a relation exists are returned (no NULLs for absent relations -- callers detect absence by checking which from.ids appear in the result).

Method load.data()

Load only the data payload for a cached resource.

Usage

ResourceStore$load.data(spec)

Arguments

spec — ResourceSpec.

Returns

Deserialized resource data.

Method load.data.from.row()

Load only the data payload using a pre-fetched DB row.

Usage

ResourceStore$load.data.from.row(row)

Arguments

row — Single-row data.frame from resource.rows.bulk() with non-NA type, id, serializer, and path columns.

Returns

Deserialized resource data.

Method add.relation()

Add a durable relation between two specs.

Usage

ResourceStore$add.relation(from.spec, to.spec, relation, name = NA_character_)

Arguments

from.spec — Source ResourceSpec.

to.spec — Target ResourceSpec.

relation — Character relation kind.

name — Optional relation name.

Method add.relations.batch()

Write multiple relation tuples in one transaction.

Usage

ResourceStore$add.relations.batch(relations)

Arguments

relations — List of named lists, each with fields from.type, from.id, to.type, to.id, relation, and optionally name.

Returns

Invisibly returns self.

Method with.transaction()

Execute a function inside a single SQLite transaction.

Usage

ResourceStore$with.transaction(fn)

Arguments

fn — Zero-argument function whose body runs inside the transaction.

Returns

Invisibly returns self.

Method resolved.outputs()

Resolve concrete outputs for a concrete or virtual resource.

Usage

ResourceStore$resolved.outputs(spec)

Arguments

spec — Resource specification.

Returns

List of concrete ResourceSpec objects.

Method resolved.virtual.outputs()

Resolve the outputs of a persisted virtual request when, and only when, it is fully cached.

Usage

ResourceStore$resolved.virtual.outputs(spec)

Arguments

spec — Resource specification (the virtual request).

Returns

List of concrete child ResourceSpec objects, or NULL.

Method view.index()

View a catalog-compatible index projection for a resource type.

Usage

ResourceStore$view.index(type)

Arguments

type — Character resource type.

Returns

data.frame.

Method exists()

Check if a resource exists in cache.

Usage

ResourceStore$exists(spec)

Arguments

spec — ResourceSpec.

Returns

Logical.

Method exists.id()

Check if a resource exists in cache by type/id.

Usage

ResourceStore$exists.id(type, id, serializer = resources.serializer$rds)

Arguments

type — Character resource type.

id — Character resource id.

serializer — Character serializer key (resources.serializer).

Returns

Logical.

Method save()

Save a concrete resource to cache.

Usage

ResourceStore$save(spec, data, meta = NULL, logs = NULL, defer.index = FALSE)

Arguments

spec — ResourceSpec.

data — Object.

meta — Named list.

logs — Optional structured logs.

defer.index — Logical; when TRUE, queue the index write to be batch-flushed with other deferred writes.

Method save.virtual()

Register a virtual resource in the SQLite index.

Usage

ResourceStore$save.virtual(spec, meta = NULL, defer.index = FALSE)

Arguments

spec — ResourceSpec.

meta — Optional metadata.

defer.index — Logical; when TRUE, queue the index write to be batch-flushed with other deferred writes.

Method flush.index()

Write every index row queued by defer.index = TRUE saves.

Usage

ResourceStore$flush.index()

Returns

Invisibly, the number of resources flushed.

Method write.log()

Materialize resource log file from upstream and current logs.

Usage

ResourceStore$write.log(spec, own.log.path = NULL, input.log.paths = NULL)

Arguments

spec — ResourceSpec.

own.log.path — Optional path to the current rule log file.

input.log.paths — Optional character vector of upstream resource log files.

Returns

Character destination log path.

Method cache.file()

Import a local file into cache and register it.

Usage

ResourceStore$cache.file(spec, path, filename = NULL, meta = NULL, logs = NULL)

Arguments

spec — ResourceSpec.

path — Character path to source file.

filename — Optional filename to store as.

meta — Named list.

logs — List.

Returns

Character destination path in cache.

Method remove.ids()

Remove resources from cache by type and id.

Usage

ResourceStore$remove.ids(type, ids)

Arguments

type — Character resource type.

ids — Character vector of resource ids.

Returns

Integer count of removed cache entries.

Method load()

Load a cached concrete resource.

Usage

ResourceStore$load(spec)

Arguments

spec — ResourceSpec.

Returns

List with spec, data, meta, logs.

Method load.meta()

Load only cached metadata for a concrete resource.

Usage

ResourceStore$load.meta(spec)

Arguments

spec — ResourceSpec.

Returns

Named list of deserialized metadata fields. Empty list if none.

Method load.meta.bulk()

Load metadata for multiple resources in one database round-trip.

Usage

ResourceStore$load.meta.bulk(ids)

Arguments

ids — Character vector of resource ids.

Returns

Named list keyed by resource id. Each entry is a named list of deserialized metadata fields for that resource. Ids not found in the store, or with no metadata, return an empty list().

Method find()

Find a cached resource that satisfies parameter and metadata constraints.

Usage

ResourceStore$find(spec, required.meta = NULL)

Arguments

spec — ResourceSpec. May contain NA params for partial matching.

required.meta — Named list of metadata constraints. Derived from spec$meta via resource.required.meta() when NULL.

Returns

ResourceSpec or NULL.

Method find.bulk()

Find cached resources for multiple specs.

Usage

ResourceStore$find.bulk(specs)

Arguments

specs — List of ResourceSpec objects.

Returns

Named list keyed by spec id. Each entry is a ResourceSpec or NULL when no match is found.

Method clone()

The objects of this class are cloneable with this method.

Usage

ResourceStore$clone(deep = FALSE)

Arguments

deep — Whether to make a deep clone.

See Also