2026-08-12-single-node-re-execution.md 20 KB

Single-Node Re-Execution - Design

Date: 2026-08-12 (revised same day)

What changed since the first version, and why

The first version of this document proposed building single-node re-execution on top of POST /api/v1/nodes/:type/options - the inline, unpersisted one-node workflow the editor uses to populate a node's dropdown options - as the closest existing scaffold, and designed a new rerun endpoint, a new sideEffects node field, and a new persisted input field around it.

That premise was wrong. The editor already has single-node re-execution, built and shipped, not proposed:

  • workflowsApi.executeToNode (webui/src/api/workflows.ts:252) sends _targetNodeId in the trigger data to the normal execute endpoint.
  • WorkflowEditorPage.tsx's executeNode callback (starts line 1070, calls executeToNode at line 1127) builds a cachedOutputs map of nodeId -> { output, configHash } for every upstream node and sends it alongside the target.
  • WorkflowEngine::execute (src/runner/workflow_engine.cpp:350-353) reads _targetNodeId and _cachedOutputs off the real trigger data (no inline workflow, no separate endpoint - this is an ordinary execution of the real workflow, just told where to stop and what it can skip).
  • The normal graph walk (workflow_engine.cpp:788-825) and the loop-body walk (workflow_engine.cpp:2660-2677, 2963-2967) both check cached_outputs for every upstream node: if a node's current config still hashes to what was cached, its stored output is substituted and the node is not executed at all - no network call, no side effect, no wait. Only nodes whose config changed, and the target node itself, run for real.

So the feature this document set out to design - "re-run one node with the current draft config, without waiting for the whole slow chain again" - is mostly built already. Revising this document rather than starting a new one because the investigation that follows (what cachedOutputs does and does not cover) is what actually determines what's left to build, and that investigation supersedes essentially every conclusion in the first version: the recommended endpoint, the recommended persisted input field, and the "what already exists" section built around nodes/:type/options are all withdrawn below. What survives from the first version - the scratch-result placement, the no-history-mutation stance, the excluded downstream-rerun scope - is repeated here with the same reasoning, not rewritten for its own sake.

The sideEffects node-module field is being implemented right now by another agent, in parallel with this revision. It is written up below as something arriving, not something this design still needs to propose.

Goal (unchanged)

Let somebody re-run one node of a finished execution - in place, in the editor, with the current draft config - without re-running the whole workflow. The motivating incident: a refactor broke a pipeline in three places, each hidden behind the last, and each fix was verified by re-running the entire workflow (minutes of image generation and a published blog post) to check what amounted to a one-line expression change.

What already exists

executeToNode + cachedOutputs already skip the upstream chain, for real

executeNode in WorkflowEditorPage.tsx (line 1070) walks the canvas graph backward from the target node (collectUpstreamNodes, inline in the same function) and, for every upstream node it finds, looks up lastExecutionResults[nodeId] ?? pinnedNodeData[nodeId] and pairs it with a hash of that node's current config (hashConfig, line 1062: JSON.stringify(config || {})). lastExecutionResults is not scoped to the current editor session only - it is seeded on load (lines 849-870) from the five most recent stored executions, merging each node's output across them so a freshly opened editor still has something to send. This whole map goes out as _cachedOutputs on executeToNode (workflows.ts:252-264).

On the runner side, configMatchesHash (workflow_engine.cpp:165) is the other half of the check: for every upstream node in the walk, if cached_outputs has an entry and its configHash matches the node's current config, the node's stored output is substituted directly into node_result.output and from_cache is set true (workflow_engine.cpp: 788-825) - the node is never run. The same check, same function, runs again for loop-body nodes (workflow_engine.cpp:2963-2967). This is not a reimplementation of the graph walk in the webserver (which the first version's option (b) rightly worried about as duplicated logic) - it is the graph walk, the same one a full run takes, just fed a cached value instead of a freshly computed one at each upstream node. Expression evaluation, branch merging, and loop-item slicing for the target node all still happen through the engine's normal code, because the target node itself is deliberately never served from cache (node_id != target_node_id guard, workflow_engine.cpp:791).

Net effect: for the motivating incident's slow chain, re-running the fixed node today already skips every unchanged upstream node's real work - including the minutes-long image generation and the blog publish - as long as those nodes' configs are unchanged, which they are when only the failing node was edited. This is the headline finding: the expensive thing the first version worried about re-running is, in the common case, already not re-run.

nodes/:type/options is a different, narrower mechanism - not the scaffold to build on

POST /api/v1/nodes/:type/options (src/webserver/api/node_controller.cpp: 26-127) builds an inline, unpersisted one-node workflow and runs it with input = {} hardcoded. It exists to populate dropdown options for a node type in isolation and has no notion of upstream data at all. It is a legitimate pattern for "run this node type with this config and no input," but executeToNode already solves the harder and more relevant problem - "run this node with this config against real upstream data" - by reusing the real execution path instead. Nothing here needs cloning.

loopNodeId / loopIteration already identify one iteration - in storage, not yet in re-run

Confirmed unchanged from the original investigation: each persisted nodeExecutions entry carries loopNodeId / loopIteration (workflow_engine.cpp:292-297 in the prior numbering; the fields are written wherever a node result is appended from inside a loop body), absent together means "not in a loop," and multiple iterations of the same node appear as distinct, ordered entries. The editor already tracks this per row

  • selectedIterations (WorkflowEditorPage.tsx:413) is state for browsing which iteration of a node's output is currently shown, populated from the engine-supplied loopIteration field (WorkflowEditorPage.tsx:476-478). So the UI already has "iteration 4 of node X" as an addressable, displayed concept. What it does not have is a way to feed that iteration back into a re-run.

The single-node loop path always replays iteration 0, never a chosen iteration

This is the first genuine gap. executeLoopBody's single_node_mode (workflow_engine.cpp:2660-2677) trims the loop body to the target node and its in-body upstream, then unconditionally does ctx.items.resize(1) - whatever the loop's item collection is, only its first element survives. There is no parameter anywhere in _targetNodeId / _cachedOutputs for "start at item N" or "use this specific item value." A user looking at iteration 4's failed output and re-running the fixed node from the editor today would silently get iteration 0's item instead, which for a feed-based loop or an image batch is not the item they were debugging.

cachedOutputs is populated from "recent, merged," not "this exact execution"

The second gap. lastExecutionResults is seeded by merging outputs across the last 5 executions (WorkflowEditorPage.tsx:820-870, mergeOutputs), favoring the newest non-empty value per node - a heuristic tuned for "what does this node usually produce," not "what did it produce in the specific execution and iteration I am looking at right now." Re-running from a specific historical execution's detail view today sends whatever is currently cached in editor state, which is usually close but not guaranteed to be that execution's actual upstream values (e.g. if a later run changed an upstream node's output shape in the meantime, or if the execution being inspected is older than the 5 seeded).

Node inputs are still not persisted - and, revisited, do not need to be

ExecutionResult::toJson() still drops the in-memory NodeExecutionResult ::input field when writing to the executions collection, with the comment "input is intentionally not stored to avoid data duplication... reconstructed from upstream node outputs + connections." The first version of this document concluded that comment was wrong for anything beyond the immediate upstream node, specifically calling out loop-item slicing as a case the comment's premise doesn't cover.

Revisiting that with cachedOutputs in view: it's wrong the other direction. The loop node's own stored output already retains the full item collection it iterated over, not just the merged result. Nothing in executeLoopBody strips _items or _isLoop out of node_result.output before it is written into result.node_results[node_id] (workflow_engine.cpp:926-936) - only the per-iteration mirror copies of body nodes are filtered from the persisted array (is_loop_mirror filtering, unchanged from the original investigation), not the loop node's own record. So _items[4] - the exact value iteration 4's body received - is already sitting in the stored execution document, on the loop node's own output, today, for every execution that has already run. Combined with cachedOutputs sourced correctly (see Gap 2 above) and a target-iteration parameter (Gap 1), the engine can recompute the exact input the target node received in that exact iteration, through its own real expression-evaluation and merge code - not an approximation, and not a second engine. Persisting input would add nothing this doesn't already give, for the loop case that was the original document's strongest argument for persisting it.

For the non-loop case, the immediately upstream node's stored output was already sufficient (the original document said so too) - and that's exactly what cachedOutputs, sourced from the specific execution, already exposes.

The storage question, revisited

The original document's option (a) - add input to toJson(), symmetric with output - is withdrawn. Two independent things changed the answer:

  1. It isn't needed. See above: cachedOutputs plus the loop node's own already-stored _items reconstructs the exact input for both the plain-DAG case and the loop-iteration case, using the real engine code path, with no new persisted field.
  2. The cost was understated even on its own terms. Execution documents are already large - the biggest measured so far is 336 KB - and listing them costs roughly 30 ms per document; a page of them can approach or exceed the gRPC message limit. Persisting input next to output for every node execution would push documents with many nodes or large payloads toward roughly double their current size (input is usually close in size to the producing node's output, since it's often that output verbatim or a shallow transform of it), which worsens exactly the listing-cost and page-size ceiling that retention and TTL work already treats as a constraint on document size.

Recommendation: do not persist input, for any node, in any form - not for every node, not only for failed nodes, not truncated. The zero-cost option is also the correct one here, which is unusual enough to state plainly rather than default to the "safe middle" of "only failed nodes" or "truncate like output does." If a future need surfaces that cachedOutputs + stored _items genuinely cannot cover - something that consumes upstream data through a path this design didn't find - that is a new, narrower case to design against with its own concrete size numbers, not a reason to persist input broadly now.

1. Where does the input come from

Answered above: from the engine's normal graph walk, fed cachedOutputs for every node whose config is unchanged since that output was produced. No stored input field, no reconstruction logic outside the engine.

2. Which iteration

Two concrete, small changes, not a new endpoint:

  • Scope cachedOutputs to the execution being inspected. When the re-run is triggered from a specific execution's detail view (as opposed to "just try the workflow as currently drafted"), build the map from that execution's own nodeExecutions entries - matching loopNodeId / loopIteration where the target node is inside a loop - rather than the 5-execution merge executeNode uses today. The 5-execution merge remains the right default for "no specific execution in mind, just re-verify against whatever's recent."
  • Carry the target iteration to the runner. Add a field alongside _targetNodeId - e.g. _targetLoopIteration - and thread it into single_node_mode (workflow_engine.cpp:2660-2677) so it selects ctx.items[iteration] instead of unconditionally resize(1) down to item 0. Falls back to the current "first item" behavior when omitted, so nothing about today's un-iterated re-run changes.

selectedIterations (WorkflowEditorPage.tsx:413) already holds, per node, which iteration the user is currently looking at - it is the natural source for this field. This is the one piece that turns "re-run this node" into "re-run this node as it ran in iteration 4," which was not expressible before today (per-iteration loopIteration tagging on execution records is newly available) and is genuinely new work, not something already built.

3. Where does the result go (unchanged from the first version)

Not into the stored execution record. WorkflowEditorPage.tsx's executeNode already treats a targeted run's result as transient editor state - executionState set from WebSocket events keyed off the returned executionId, discarded on navigation - not merged back into any stored document. ExecutionResult persistence (workflow_engine.cpp's upsert/update on the executions collection) always writes the whole document rebuilt from the complete in-memory node_results map; there is no path that patches one entry inside a stored nodeExecutions array, and this design should not add one. A single-node re-run through executeToNode already produces its own real (if narrow) execution record

  • that's fine as an audit trail of "a test run happened," and is exactly how it works today - but its outputs should never be spliced into the original execution being investigated. The UI already respects this by keeping the re-run's result under its own executionId; the requirement going forward is just to keep it visibly labeled as a re-run when shown next to the original execution's per-node output, not silently swapped in.

4. Side effects

sideEffects is being implemented on the node module interface by another agent, in parallel with this document. Once it lands, the re-run action already present in the editor (executeNode / the "run to this node" affordance) should read it for the target node and any node whose cache entry was rejected (config changed, so it will actually execute) and show a specific warning - "this will really call X again" - rather than a generic one. Nodes that don't declare it render the existing mild, generic caution. No further design needed here; this section exists so a reader of this document knows the flag is inbound rather than still open.

5. What counts as "the fixed node" (unchanged from the first version)

Re-run uses the current editor draft config with the input the node received (now: reconstructed live through the engine and cachedOutputs, per iteration where relevant - not a stored field). Not the last-saved config, not the original execution's config - the live draft, because the point is verifying an edit before committing it. And, as before: the re-run does not touch the node's own input. If the edit under test would itself change what upstream nodes produce, a single-node re-run can't see that - out of scope, unchanged from the first version's Question 5.

6. Scope

Ship:

  • Scope cachedOutputs to the specific execution being inspected when a re-run is launched from that execution's detail view, instead of always using the 5-execution merge (WorkflowEditorPage.tsx:820-870, executeNode at line 1070). The 5-execution merge stays as the default for a re-run with no specific execution context.
  • Add a target-iteration field (e.g. _targetLoopIteration) alongside _targetNodeId / _cachedOutputs in the trigger data (workflow_engine.cpp:350 on), and use it in executeLoopBody's single_node_mode (workflow_engine.cpp:2660-2677) to select the intended item instead of always item 0. Source it from selectedIterations (WorkflowEditorPage.tsx:413) in the editor.
  • Wire the sideEffects flag (arriving from other in-progress work) into the existing re-run affordance's warning, once it lands.

Deliberately not doing:

  • Persisting input on NodeExecutionResult / ExecutionResult::toJson(), in any form - full, failed-only, or truncated. See the storage section above.
  • A new rerun endpoint or any reuse of nodes/:type/options - the existing execute + _targetNodeId + _cachedOutputs path already covers this; there is nothing to add a parallel mechanism for.
  • Re-running a node plus everything downstream of it - a different, larger feature (needs a resumable partial graph walk continuing through the normal engine, and its own decision on whether that continuation writes a real record or a scratch one). Not needed for the motivating incident, where each bug was diagnosable one node at a time.
  • Tagging existing nodes with sideEffects - the flag's existence is in-flight work; tagging the ~80 existing nodes with it is a follow-up.
  • Mutating or patching a stored execution record in any way.
  • Fixing or implementing GET /api/v1/executions/:id/retry (execution_controller.cpp:445-458) - unrelated stub, whole-execution scope, unchanged from the first version.

Summary of file:line touch points for implementation

  • webui/src/pages/WorkflowEditorPage.tsx:820-870 - seeding of lastExecutionResults; needs an execution-scoped variant for Question 2.
  • webui/src/pages/WorkflowEditorPage.tsx:1070-1144 - executeNode, where cachedOutputs is built and executeToNode is called; add execution-scoped cache sourcing and the target-iteration field.
  • webui/src/pages/WorkflowEditorPage.tsx:413, 476-478 - selectedIterations, source for the target-iteration field.
  • webui/src/api/workflows.ts:252-264 - executeToNode; add an optional target-iteration parameter alongside _targetNodeId / _cachedOutputs.
  • src/runner/workflow_engine.cpp:350-353 on - execute()'s trigger-data parsing; add the target-iteration field next to _targetNodeId / _cachedOutputs.
  • src/runner/workflow_engine.cpp:2660-2677 - executeLoopBody's single_node_mode; replace unconditional ctx.items.resize(1) with selection by target iteration when supplied.
  • src/runner/workflow_engine.cpp:788-825, 2963-2967 - configMatchesHash-gated cache substitution, top-level and loop-body; unchanged, this is the mechanism being extended, not rebuilt.
  • src/runner/workflow_engine.cpp:926-936 - loop node's own stored output, confirmed to retain _items untouched; the reason persisting input is not needed for the loop case.
  • Node module sideEffects field and its surfacing in the re-run UI - tracked as arriving from other in-progress work, not part of this scope.