The Execution Engine
How declared work becomes scheduled tasks, how caching decides what runs, and how to read a run
When you call generate$models() with three GWAS and four algorithms, you are not
calling twelve functions. You are declaring twelve outputs you want to exist, and
handing them to an engine that works out what already exists, what has to be built
first, and what can run at the same time.
This chapter is how that works. The guide chapter covers the same ground from the user's side; this one is for when you need to reason about the mechanism.
The shape of the problem
Generation is a dependency graph, not a sequence. A model needs an algorithm run, which needs an LD resource, which needs a reference panel restricted to a variant space, which needs the panel itself, which may need a liftover. Two models from different algorithms share most of that chain. Two models from the same algorithm at different parameters share all of it but the last step.
Doing this by hand means either recomputing shared work or hand-managing intermediate files. Doing it with a general workflow tool means describing your analysis twice, once in R and once in the workflow language.
The engine's approach is that you declare only the outputs you want. Everything between them and your inputs is derived.
Rules
A rule knows how to produce one type of resource. It answers three questions:
- Does this output match me? Given a resource specification, can I produce it?
- What do I need first? Given the output, what inputs must exist?
- What will it cost? Cores and memory, so the scheduler can plan.
And then it runs, receiving its resolved inputs and returning a payload plus metadata.
Rules never call each other. A rule that needs an LD matrix declares it as an input and receives it; it does not know which rule produced it or whether that rule ran or was cached. That is what makes the set of rules extensible without coordination — see extending PolyGenius.
Discovery
Given a requested output, the engine walks backwards:
- Is it already stored? If the index has a row and the payload is on disk, the task is marked cached and never reaches a worker. This check happens before rule matching, which is why a fully cached run does almost no work.
- If not, which rule produces it? The first matching rule claims it.
- What does that rule need? Each input becomes a new output to discover, recursively.
Discovery therefore terminates at things that exist — either in the store, or as inputs you supplied.
One asymmetry in the current implementation is worth knowing if you are reasoning about cache behaviour: relaxed and partial matching is consulted for a rule's inputs, but a directly-requested output that misses on exact lookup goes straight to rule matching without trying a relaxed match. So the "a looser stored resource satisfies a stricter request" behaviour is reachable for intermediates, and for anything reached through a recorded resolution, but not for a top-level request.
Dynamic expansion
Some rules do not know how many outputs they produce until they run. An LDpred2 grid produces one model per parameter combination; a clumping run followed by thresholding produces one model per threshold.
These are handled as a virtual parent whose recorded outputs are checked as a set. If every model from a previous expansion is still present, the expansion is not re-run. If any one is missing, the whole expansion re-runs — which is the conservative choice, since a partial expansion is not a state the engine can reason about.
Scheduling
Tasks with satisfied dependencies are ready. Ready tasks are dispatched in priority order, subject to the core and memory budget.
Requested outputs are prioritised over intermediates. That sounds backwards — you cannot have a model before its LD matrix — but it applies among ready tasks, and it means the work that finishes your actual request is preferred over work that merely could be done.
The core budget is capped at one below the number of cores the engine can detect, unless you explicitly allow oversubscription. That cap is silent, which surprises people: on a shared login node reporting four cores you get three workers no matter what you asked for. Print the workspace to see what you actually got.
Oversubscription is legitimate in one specific case: work that is blocked on the network rather than on CPU, such as many concurrent summary-statistics downloads.
The memory budget defaults to unlimited, which means memory does not restrict concurrency at all while every LD-based task declares a substantial requirement. On a machine where that matters, set it to your allocation. Note that setting it too low does not serialise the work — a task whose declared requirement exceeds the budget fails with a message saying so.
Failure
A worker error is caught when its result is collected. The task is marked failed, with the rule name, the message and the elapsed time recorded as an event. One task's failure never terminates the run.
Everything downstream of a failure cascades to blocked, each with its own event naming the dependency that failed. This is why a single bad GWAS can produce a dozen failure rows: one root cause, eleven collateral. When triaging, find the row whose reason is not a dependency.
Failures are not persisted as failures. There is no negative caching, so a failed task is retried on the next run. A deterministically failing task fails every time, which is mildly wasteful and much less confusing than the alternative.
Successes are cached, which means re-running after a partial failure skips everything that worked. There is no resume command because re-running is the resume.
What a run leaves behind
Under the store root, an execution directory holds one text log and one structured event stream per run. Event types cover the run's start and end, and each task's start, completion, failure or blocking.
The returned PGS library carries the paths to those files, along with the run's status and the table of failed models. That means you never have to guess which run you are looking at — the object tells you.
Two visualisations read the event stream: a self-contained HTML dashboard of one run, and a performance breakdown by rule.
Cached runs are invisible
There is no event for a cached task. A fully cached run therefore emits no task events at all, and its dashboard is empty — while the progress display counts cached tasks as completed and jumps immediately to 100%.
Both behaviours are correct and they look contradictory. If you know to expect it, an instant empty run is the clearest possible signal that nothing needed doing.
Reasoning about caching
The single most useful mental model: a stored result is identified by what you asked for. Not by when, not by which version of the code, not by the contents of files you pointed at.
From that, everything else follows:
- Change an algorithm parameter and you get a new resource. Change a thread count and you do not.
- Edit a local input file without changing the name you gave it, and you get the old result.
- Upgrade PolyGenius and nothing is invalidated.
- Tighten a p-value threshold and you get a new model, but you do not re-download the summary statistics.
The last one is the one people expect to work the other way round, so it is worth restating: relaxed matching applies to the fetch threshold, not to a model's parameters.
Resources and catalogs covers identity in more detail, including how partial specifications resolve.
Configuration
workspace$config$update() is the single entry point:
| Setting | What it controls |
|---|---|
root |
where the store, logs and run records live |
max.cores |
the core budget; NULL means auto-detect |
max.memory |
the memory budget; unlimited by default |
oversubscribe |
whether an explicit core count may exceed the detected cap |
verbosity |
the console log threshold |
execution.status |
whether the live progress display appears |
Changing anything that defines the execution environment rebuilds the executor and the store connection mid-session, so you do not need to restart R.
Triage
An ordered sequence, cheapest first, each step eliminating a class:
- Count what you got against what you asked for. If they match, the problem is not in generation.
- Read the failure table on the returned set. The reason names the cause; the model's parameters are on the same row.
- Classify before investigating. A budget message is configuration, not data. A dependency message means collateral damage — find the root failure.
- Read the resource log for the failing resource. It carries the underlying tool's own output plus everything upstream.
- A suspiciously fast, empty-looking run is a cache hit. Raise verbosity to confirm.
- A result that looks stale rather than missing is the cache working as designed. Generate into a fresh root to compare.
- A run that stops reporting it cannot make progress has a dependency that can never be satisfied — usually a missing tool or panel. Check setup first.
Steps one to four resolve most cases in a couple of minutes.
Known limits
Some of these matter when reading the code; all of them matter when trusting it.
Identity carries no version. Recorded as a deliberate gap with a fix specified but not implemented.
Locally-supplied content is not in identity, only the name you gave it.
Relaxed matching is not reachable for top-level requests, as described under Discovery.
Scheduler state is in memory only. Queues, priorities and blocked reasons do not survive the process. Results are committed as they are collected, so completed work is durable; in-flight work is discarded.
There is no run-scoped cleanup. The index has a table for runs that is never populated, so there is no supported way to remove everything one run produced.
Where to go next
Extending PolyGenius covers writing a rule that plugs into this machinery.