CLI¶
The ginkgo CLI is the main operator surface for authoring, validating,
running, and inspecting workflows. Every command operates on the project rooted
at the nearest ginkgo.toml.
Command Overview¶
ginkgo initScaffold a new project —
ginkgo.toml, a starterworkflow/flow.py, the canonical layout, and askills/directory for coding agents. See Working with Coding Agents.ginkgo runBuild the expression tree, validate the workflow, evaluate ready tasks, and record the run. The command you reach for most.
ginkgo inspectInspect the resolved task graph of a workflow without running it (
inspect workflow).ginkgo runsList recorded runs (
runs ls) and show one of them (runs show <run_id>). See Querying Provenance.ginkgo historyShow every run of one task, with its status, duration, and cache key.
ginkgo queryRun one read-only SQL statement against the provenance database.
ginkgo exportWrite a run’s ledger events as JSONL (
export events) or its manifest as YAML (export manifest).ginkgo debugInspect a finished run — task status, timing, logs, and cache decisions — from its recorded run directory.
ginkgo doctorCheck a workflow and its environment for problems: missing environments, unresolved secrets, malformed config.
ginkgo reportRender a finished run as an HTML report. By default this produces a directory bundle (
index.htmlplus anassets/folder); pass--single-filefor a single self-contained HTML file. See Assets and Reports.ginkgo cacheList, clear, and prune cached task results. See Caching and Provenance.
ginkgo assetList and inspect typed, versioned task outputs. See Assets and Reports.
ginkgo modelsList model assets together with their recorded metrics.
ginkgo notebooksList the rendered notebook HTML artifacts produced by runs.
ginkgo envList and reset the Pixi and container environments backing shell tasks. See Environments.
ginkgo secretsList and validate the secret references a workflow resolves at run time.
ginkgo dbMaintain the provenance database.
db pathprints where it is,db migratecreates or upgrades it,db checkreports anything the database and the files beside it disagree about,db prunedeletes history you no longer need, anddb vacuumgives the freed space back. See Caching and Provenance.
Run ginkgo <command> --help for the full flag set of any command.
Running Workflows¶
ginkgo run flow.py
ginkgo run flow.py --jobs 8 --cores 32 --memory 64
ginkgo run flow.py --dry-run
ginkgo run builds the expression tree, validates the workflow, evaluates ready
tasks subject to the --jobs, --cores, --memory, and --gpus budgets,
and writes run history under .ginkgo/runs/. Run it from a project root with
no path argument and Ginkgo discovers the canonical workflow/flow.py
entrypoint.
Commands may be run from any directory inside the project. Ginkgo locates the
root — the nearest enclosing ginkgo.toml — and works from there, so
.ginkgo/, config files, and task environments resolve identically wherever
you invoke from. Paths you pass on the command line are relative to where you
stand; a task’s own relative output paths are relative to the project root.
Repeated --resource name=value flags budget any custom resource dimensions
tasks declare (e.g. --resource api_calls=10) — see
Custom Resource Dimensions.
--dry-run resolves the graph and computes cache keys without executing any
task body — the fastest way to confirm a workflow is wired correctly.
--agent-output swaps the live terminal UI for a stream of newline-delimited JSON
events, for programmatic use by AI coding agents — see
Working with Coding Agents.
Workflow Parameters¶
A workflow declares the inputs it accepts with ginkgo.param(...), and each one
becomes a command-line flag:
import ginkgo
n_replicates = ginkgo.param("n_replicates", type=int, default=12, help="Replicates per item")
region = ginkgo.param("region", help="Genome region") # no default: required
ginkgo run flow.py --n-replicates 24 --region 2L:1-100000
The flag is the dashed form of the name. A value resolves from the command line
first, then the [params] table of ginkgo.toml, then the declared default:
[params]
n_replicates = 24
region = "2L:1-100000"
ginkgo run flow.py --help lists the parameters that workflow declares,
with their types and defaults. A flag the workflow does not declare is rejected
before anything runs, and the error names the parameters it does declare. A
required parameter that is not supplied fails the same way.
type follows argparse’s convention, so type=int, type=float, and
type=Path all work. Booleans accept a bare --flag or an explicit
--flag false, and multiple=True makes a flag repeatable:
ginkgo run flow.py --item alpha --item beta --verbose
Resolved values are recorded with the run, along with where each came from
— the CLI, config, or the default. ginkgo runs show --json shows both.
The [params] table layers across config files, so --config extra.toml setting
one parameter leaves the others in ginkgo.toml alone.
Important
Pass a parameter into a task as an argument. Cache keys hash task arguments,
so a parameter passed as one correctly re-runs the tasks that used it. A
parameter read from a module global inside a task body is not part of the key, so
changing it would silently reuse the previous result. Both ginkgo run and
ginkgo doctor warn when they spot this.
Validation And Diagnostics¶
Use these commands to inspect a workflow without committing to the full workload:
ginkgo run --dry-run
ginkgo doctor flow.py
ginkgo debug <run_id>
ginkgo run --dry-run previews the plan for the entrypoint you actually run,
and is the quickest way to confirm a workflow you just wrote is wired correctly.
It reports the waves tasks fall into, which would run and which would serve from
cache, and the resources they declare.
Validation workflows¶
A project may keep workflow files under tests/workflows/ that exercise its own
flow. There is no separate command for them: a validation workflow is a workflow,
so run it by path like any other.
ginkgo run tests/workflows/smoke.py --dry-run # validate
ginkgo run tests/workflows/smoke.py # execute
Each such file must expose a flow, usually by re-exporting the project’s own.
ginkgo init scaffolds one tests/workflows/smoke.py that re-exports main
from workflow/flow.py.
ginkgo doctor catches environment and configuration problems before a run.
Pass --json for structured output suitable for programmatic use:
ginkgo doctor flow.py --json
ginkgo debug is most useful after the fact: once a run directory exists, it
surfaces recorded task status, logs, and cache behavior without manually
navigating .ginkgo/runs/.
ginkgo cache has several subcommands beyond listing:
ginkgo cache ls # list cached task results
ginkgo cache stats # size, hit counts, biggest tasks
ginkgo cache explain <run_id> # explain cache decisions for a run
ginkgo cache prune --older-than 7d # remove entries older than a duration
ginkgo cache prune --max-size 10GB # remove entries to stay under a size limit
ginkgo cache prune --max-entries 500 # remove entries to stay under an entry count
ginkgo cache clear <cache_key> # remove a specific cache entry
ginkgo cache clear --orphans # remove entries the database has lost
cache prune requires at least one of --older-than, --max-size, or
--max-entries. Add --least-recently-hit to give up unused entries before old
ones, and --dry-run to preview what would be removed. cache stats takes
--json.
Why Did This Task Run Again?¶
ginkgo cache explain <run_id> prints JSON with one entry per task. For a task
that re-ran, it names the components of the cache key that moved, so the answer
is a specific fact rather than “the key changed”:
{
"task_name": "produce",
"cache_key": "fddb71a9…",
"compared_with": {"cache_key": "4388619b…", "strategy": "same_node"},
"reason": "input_changed",
"details": ["input_changed"],
"components": [
{
"component": "inputs.samples",
"status": "changed",
"current": {"type": "str", "sha256": "2f64dc9d…"},
"prior": {"type": "str", "sha256": "f7a9858a…"}
}
]
}
reason and details are the summary; components lists what actually
differs. Component names follow the cache key: task, version, source_hash,
extra_source_hash (the notebook or script a driver task runs), env,
env_hash.pixi_lock (the identity of the declared environment), and one
inputs.<parameter> per task argument.
compared_with says what the entry was compared against, and how it was found.
"strategy": "same_node" means the key this same node used in the most recent
earlier run of the workflow — the entry this run superseded, which is the
comparison you want. "strategy": "newest_by_function" is the fallback used
when the node is new or its label changed: the newest earlier entry written by
the same task, which may be a different branch of a fan-out. Read those
components more sceptically.
A component’s status is changed, or added / removed for a parameter that
appeared or went away.
Some tasks get a summary reason on its own, with no components: a task that
did not re-run reports all_inputs_match, one with no earlier entry to compare
against reports no_prior_entry, and one whose key has no cache entry at all —
because it failed, or because the entry was pruned — reports no_entry_for_key.
A Typical Loop¶
For local development, a practical cycle looks like this:
author and adjust tasks in code
check the wiring with
ginkgo run --dry-runrun with
ginkgo runinspect failures or cache reuse with
ginkgo debug
Because Ginkgo caches completed tasks, iterating on a later stage of a workflow re-executes only that stage — earlier tasks serve straight from cache.