Skip to content

Run storage

Every pipeline run, and every step within it, is persisted in the database in full detail — enough to reconstruct exactly what happened, what was sent to each agent, and how each confidence decision was reached, without needing the original trace to still be available. This page documents those tables and the artifact store that keeps large step outputs off the database entirely.

pipeline_runs

Column Type Description
id uuid, pk Run identifier (returned in /webhook 202 response)
pipeline_name str
source str Webhook source or "scheduler"
triggered_at datetime
status str running / completed / stopped / aborted / escalated / failed
normalised_context json Full NormalisedContext at trigger time
raw_payload json Original unmodified webhook payload
completed_at datetime, nullable
logs json, nullable Structured run event log — array of {ts, level, event, msg} objects. Populated at run completion. Events cover step start/complete/fail/skip/escalate/abort, verifier results, parallel group outcomes, notifications sent, and (for interrupted runs) the startup interruption sweep, or (for resumed runs) run_resumed — see Durability & resume.
parent_run_id uuid, nullable, indexed Set for sub-pipeline runs (executor: pipeline). Links back to the parent run. NULL for top-level runs.
team str, nullable, indexed Owning team, resolved from the Bearer token that authenticated the webhook (see Team attribution). NULL for unattributed/legacy-token runs.
stage str, indexed testing or production, copied from PipelineConfig.stage (see Pipeline stages) at trigger time — persisted per-run so promoting a pipeline never reclassifies its prior runs. Defaults to production at the DB layer.
config_fingerprint str, nullable Fingerprint of the pipeline’s step sequence at trigger time, populated on every run. Compared against the current config at resume time — see Durability & resume.
resumed_at datetime, nullable Set the first time this run is resumed after a restart; never cleared. Drives vectorstep_runs_resumed_total.
replay_of json, nullable Set only for a synthetic run created by a replay batch — {source, candidate, mode, k, sample_run_ids}. NULL for every ordinary run. A replay run is always stage: testing, so this column drives no filtering by itself — it exists purely so the replay report can reconstruct what was replayed.

A pending_approvals table durably mirrors an outstanding executor: human wait (token, run, step, message) so an approval request already delivered to Telegram/Slack/Teams still resolves after a restart — written when the request is sent, deleted when it resolves. See Durability & resume.

Two more tables back pipeline_steps.prompt_hash/agent_version above — content-addressed registries that make the actual prompt/agent text behind a hash recoverable, rather than the hash being a dead end:

  • step_prompt_versions (prompt_hash pk) — the full prompt_template text for every distinct prompt a step has run under, with first_seen_at/ last_seen_at. Write-on-miss at step-save time; grows with prompt edits, not with runs. This is what the run detail page’s Prompt history disclosure and GET /steps/{name}/versions read from.
  • agent_version_snapshots (agent_version pk) — the same idea for a Gateway agent’s agent.yaml/soul.md text. VectorStep can’t compute this hash itself (the Gateway owns the algorithm), so on seeing an unknown agent_version it asks the Gateway for the current definition and stores it only if the Gateway’s reported current version still matches — otherwise the snapshot’s text stays NULL with a note explaining why, rather than risking a mislabelled snapshot. Backs the agent detail page’s Versions tab and GET /agents/{name}/versions.

See How confidence and calibration work for why editing a prompt or an agent resets calibration, and how reverting to an earlier hash automatically rejoins that version’s original bucket.

pipeline_steps

Column Type Description
id uuid, pk
run_id fk → pipeline_runs.id
step_name str "step-name" for sequential; "group-name/branch-name" for parallel branches
step_index int Sort order within run
executor str openclaw / gateway / human / webhook
agent str executor:agent-name (e.g. openclaw:sre-triage)
model str Actual model used, from executor metadata
provider str, nullable Gateway provider key (e.g. anthropic, openrouter) — executor: gateway steps only, from agentMeta.provider. NULL for other executors; the UI shows the bare model name rather than guessing one.
prompt_hash str, nullable, indexed SHA-256[:12] of the step’s normalised prompt_template — scopes the calibration bucket to the exact prompt that produced this outcome. NULL for pre-migration rows and non-LLM steps with no template. See How confidence and calibration work.
agent_version str, nullable, indexed Gateway-reported content hash of the agent’s full config incl. soul.md, from agentMeta.agentVersion. NULL for non-gateway executors and pre-migration rows.
prompt text Rendered prompt sent to the agent — the actual, fully-substituted text, not the {{ }} template. Populated for executor: gateway steps only; other executors don’t yet stash their rendered prompt back out, so their rows fall back to a JSON dump of executor_config (recognisable by starting with {) — the UI’s Prompt disclosure hides that fallback rather than showing it as if it were a real prompt.
raw_output json Full unparsed executor response
parsed_output json Validated LLMOutput (excluding raw_response)
status str completed / stopped / escalated / aborted / failed
primary_confidence float Raw confidence from the primary agent
verifier_confidence float, nullable Verifier agent confidence (if verifier ran)
effective_confidence float Confidence used for threshold gate (post-combination)
grounding_score float, nullable Shadow-mode grounding score G ∈ [0,1] — the fraction of the step’s load-bearing claims a blind grounding judge found supported by evidence in the step’s own execution trace. NULL when grounding wasn’t configured for the step, or when it had no trace to check against. Never gates on its own — see How confidence and calibration work.
trust_report json, nullable Per-step TrustReport: the trust vector {S, S_after_V, V, V_mode, V_combination_strategy, V_veto_floor, G, C, D}, combined_trust, gate ({policy, confidence_threshold, on_low_confidence} — policy is legacy_confidence / trust_vector), and — for grounding — the per-claim support breakdown, and — for deterministic checks — the full per-check detail. Populated for steps with a verifier, a grounding: block, deterministic_checks:, and/or calibration: {enforce: true} — i.e. any mechanism beyond the plain single-confidence gate, not just the trust-vector ones; mode is "shadow" when recorded-only or "enforced" when grounding/deterministic/calibration actually participated in the gate (a verifier-only step is always "shadow", since the verifier’s downward-only combine has always been part of the legacy gate). A calibration sub-key (bucket/bin/n/n_min/validated/raw/calibrated/on_uncalibrated) is present only for a step with calibration: {enforce: true}. See How confidence and calibration work.
deterministic_passed bool, nullable Whole-step pass/fail across all declared deterministic checks — True only if every check passed. NULL when no deterministic_checks: were declared. Full per-check detail lives in trust_report.deterministic_checks.
duration_ms int
executed_at datetime
artifacts json, nullable {key: reference} map — references are opaque strings pointing to artifact files on disk. Content is not stored in the DB. See Artifact storage below.
agent_trace json, nullable Ordered execution trace from the VectorStep Gateway executor — array of {type, ...} objects. type is one of: llm_call (iteration marker), thinking (extended thinking block), text (response text), tool_call (MCP tool invoked with arguments), tool_result (MCP tool response). Not truncated — the full content the Gateway returns on its final ok frame is stored and rendered as-is; only the ephemeral live SSE tail truncates content (at 200 chars) for a fast in-progress preview, never the persisted record. If a tool_result’s content looks cut off on a completed run, that truncation happened upstream (the Gateway server or the MCP tool itself), not in this column. NULL for all other executors (openclaw, human, webhook).
verifier_agent str, nullable executor:agent-name for the verifier call (mirrors agent above), e.g. gateway:principal-sre. NULL if no verifier ran.
verifier_model str, nullable Actual model used by the verifier call, from executor metadata. NULL if no verifier ran.
verifier_provider str, nullable Gateway provider key for the verifier call (gateway executor only). NULL if no verifier ran or the verifier used a non-gateway executor.
verifier_prompt text, nullable Rendered prompt actually sent to the verifier — for critic mode this is the meta-prompt with the primary’s own prompt+response embedded; for independent mode it’s a verbatim copy of the primary’s prompt. Gateway executor only; NULL if no verifier ran or the verifier used a non-gateway executor.
verifier_output json, nullable Full unparsed verifier response. NULL if no verifier ran.
verifier_mode str, nullable critic / independent for new rows; reviewer / challenger on historical rows predating the rename (still valid, permanent aliases — see Verifiers). NULL if no verifier ran.
input_tokens int, nullable Input tokens consumed by this step’s primary executor call. Populated for gateway steps; NULL for others. For parallel/fan-out branches, each branch row has its own token count.
output_tokens int, nullable Output tokens produced by this step’s primary executor call.
verifier_input_tokens int, nullable Input tokens consumed by the verifier call. NULL if no verifier ran.
verifier_output_tokens int, nullable Output tokens produced by the verifier call. NULL if no verifier ran.
grounding_model str, nullable Actual model used by the grounding judge. NULL if grounding didn’t compute.
grounding_provider str, nullable Gateway provider key for the grounding judge call. NULL if grounding didn’t compute.
grounding_input_tokens int, nullable Input tokens consumed by the grounding judge call. NULL if grounding didn’t compute.
grounding_output_tokens int, nullable Output tokens produced by the grounding judge call. NULL if grounding didn’t compute.
cost float, nullable Total priced cost for this step (primary + verifier + grounding), summing whichever tier columns below are non-null. NULL only if every component that ran was entirely unpriced — never 0 unless every priced component was genuinely rated at zero. See Cost accounting.
cost_manual float, nullable Portion of cost priced from a pricing.models match.
cost_live_exact float, nullable Portion of cost priced from a live OpenRouter rate for the exact provider actually called.
cost_live_cross_provider float, nullable Portion of cost priced from a live OpenRouter rate borrowed across providers.
cost_source_detail json, nullable Per-component breakdown — {role, source, cost, [id, in_per_mtok, out_per_mtok]} per primary/verifier/grounding component, plus catalog_fetched_at when any component used a live tier. NULL iff cost is NULL.

run_feedback

Column Type Description
id uuid, pk
run_id str, unique, indexed The run this feedback is for. One row per run — submitting again upserts.
pipeline_name str, indexed Denormalised from the run for efficient pipeline-level queries.
outcome str correct / partial / incorrect — human judgement of the run’s result.
notes text, nullable Optional free-text context.
submitted_at datetime Created or last updated.

step_feedback

Column Type Description
id uuid, pk
step_id str, unique, indexed The specific step execution (pipeline_steps.id) this feedback is for. One row per step, upserted.
run_id str, indexed Denormalised for lookup.
pipeline_name str, indexed Denormalised.
step_name str Denormalised — may contain / for fan-out branches (e.g. triage/0).
outcome str correct / partial / incorrect — human judgement of that step’s result.
notes text, nullable Optional free-text context.
submitted_at datetime Created or last updated.

Steps can produce large artifacts (research reports, scraped data, compiled documents) that would be unwieldy to pass inline through next_step_context. The artifact store writes these to disk, keeps them out of the database, and makes them available in downstream prompt templates by content — not by reference.

An agent returns an artifacts dict alongside its normal LLMOutput fields. Each key is a name chosen by the agent; each value is the full text content:

{
"confidence": 0.9,
"summary": "Research complete — 3 sources compiled",
"next_step_context": "Coverage spans Q1–Q4 2025",
"artifacts": {
"research_report": "# Research Report\n\n## Source 1\n..."
}
}

The runner intercepts the artifacts field before anything is stored in the database. The content is written to {artifacts_dir}/{run_id}/{step_name}/{key} and replaced with an opaque reference string (local://...). SQLite only stores the reference; the blob lives on disk.

Downstream steps reference artifact content in their prompt templates using {{artifacts.step_name.key}}. The runner loads the content from disk at render time — only for the steps that actually reference it.

steps:
- name: research
executor: openclaw
executor_config:
agent: web-researcher
prompt_template: |
Research the topic and compile a full report.
Return JSON with the usual fields plus an "artifacts" key:
{"confidence": ..., "summary": ..., "next_step_context": ...,
"artifacts": {"research_report": "..."}}
- name: proofread
executor: openclaw
executor_config:
agent: editor
prompt_template: |
Proofread and improve the following document:
{{artifacts.research.research_report}}
Return the corrected document as an artifact named "final_report".

Hyphens in step names follow the same rule as steps.* references — use underscores in template expressions:

# Step named "web-research" is referenced as:
{{artifacts.web_research.report}}

Artifact directories are scoped to a run ({artifacts_dir}/{run_id}/). A daily APScheduler job (02:00) removes directories for runs older than retention_days. Failed runs retain their artifacts for the same period, which is useful for debugging.

To disable artifact storage entirely, omit the artifacts: block from config.yaml. Steps that return an artifacts field will have it passed through as a regular extra field rather than being written to disk.