Kaynağa Gözat

fix: a stop reaches called workflows, and only active ones may be called

Pressing Stop on a workflow that calls another one appeared to do nothing:
the run kept going for minutes and more sub-workflows kept starting. Three
separate reasons, all fixed here.

Cancellation did not reach the work. A called workflow runs as its own
execution with its own id, so marking the caller cancelled said nothing
about it, and the caller - sitting inside that call - could not notice its
own cancellation until the call returned. cancelExecution now walks
parent to child and marks the whole tree.

Nothing stopped the next call from starting. A loop calling a workflow per
item only looks at its own cancellation between items, so a stop during
item 2 still launched items 3, 4 and 5. runSubWorkflow now refuses to start
when the caller is already cancelled.

Measured on a run of five 8-second calls: stopping 10s in ended the run at
17s with two children instead of five, the in-flight one cancelled. It was
running to completion before.

A workflow must now be active to be callable. An inactive one is a draft -
half-edited, or deliberately taken out of service - and running it because
something else still points at it is how an unpublished change reaches
production.

Also: a run started from the editor now counts towards the schedule's
overlap policy. The scheduler only ever heard about runs it started itself,
so "skip if already running" did not skip for anything else - which is how
a second scheduled run landed on top of one someone was watching.

The two call-workflow fixtures named a sub-workflow by id, a row made by
hand once that anyone could delete. Cases can now declare the workflows
they need, created and removed with the case. 60/60.
fszontagh 1 ay önce
ebeveyn
işleme
411d1611de

+ 693 - 0
.superpowers/config-defaults-report.md

@@ -0,0 +1,693 @@
+# configSchema defaults - write time and read time
+
+Branch: `config-defaults` (off `main`)
+Commit history: `3bd17d0698ca6fed6f643b5e1be57d5983c32222` (round 1) ->
+`de2b70b77d4480c56522e7cb9fa1d9af7d91f1fc` (round 1 fix) ->
+`ccdc442fc19fe235b8a1a4a46e985ed31f898c4c` (round 2 fix, below - current HEAD)
+
+## What was built
+
+- `lib/common/config_defaults.hpp` / `.cpp` (namespace `smartbotic::common`):
+  `nlohmann::json applyConfigDefaults(const nlohmann::json& config, const nlohmann::json& config_schema)`.
+  Fills in missing top-level `configSchema.properties[*].default` values. A
+  key already present in `config` - including an explicit `false`, `0`,
+  `null`, or `""` - is never touched. Tolerates a null/non-object schema or
+  one with no `properties`. Added to `CMakeLists.txt` under the
+  `smartbotic_common` target, alongside the other `lib/common/*.cpp` files.
+- Read-time call sites (the guarantee):
+  - `src/runner/workflow_engine.cpp`, just before
+    `evaluateExpressions(node->config, ...)` (~line 651). Looks up the node
+    definition via `registry_.getNode(node->type)` and applies defaults to
+    `node->config` before expression evaluation, so a defaulted value still
+    goes through expression evaluation like any other value.
+  - `src/webserver/webserver_service.cpp`, `loadScheduledWorkflows()`
+    (~line 318). Applies defaults to the trigger's stored config, using the
+    node definition already fetched from `node_store_`, before reading
+    `pollInterval`.
+  - `src/webserver/api/workflow_controller.cpp`,
+    `updateScheduledTriggers()` (~line 413-414) and
+    `getScheduledInterval()` (~line 440-441). Same pattern, using
+    `node_store_.get(node_type)`.
+- Write-time call sites: `src/webserver/api/workflow_controller.cpp`,
+  new private method `materializeNodeConfigDefaults(nlohmann::json& body)`,
+  called from `createWorkflow()` (after the name-required check, before
+  `storage_.insert`) and `updateWorkflow()` (before `storage_.update`).
+  Walks `body["nodes"]`, looks up each node's definition in `node_store_`,
+  and replaces each node's `config` with
+  `applyConfigDefaults(config, node_def.config_schema)`, so the saved
+  workflow document is self-describing.
+
+## Build
+
+Command:
+
+```
+cmake --build build -j$(nproc)
+```
+
+Full real output (tail):
+
+```
+[0/1] Re-running CMake...
+CMake Warning (dev) at /usr/share/cmake-4.2/Modules/FetchContent.cmake:1963 (message):
+  Calling FetchContent_Populate(bcrypt) is deprecated, call
+  FetchContent_MakeAvailable(bcrypt) instead.  Policy CMP0169 can be set to
+  OLD to allow FetchContent_Populate(bcrypt) to be called directly for now,
+  but the ability to call it with declared details will be removed completely
+  in a future version.
+Call Stack (most recent call first):
+  cmake/Dependencies.cmake:54 (FetchContent_Populate)
+  CMakeLists.txt:13 (include)
+This warning is for project developers.  Use -Wno-dev to suppress it.
+
+-- Found RE2 via pkg-config.
+-- Found MariaDB client library
+-- Found PostgreSQL client library (libpq)
+-- Configuring done (1.3s)
+-- Generating done (0.1s)
+-- Build files have been written to: /data/smartbotic/build
+[1/8] Building CXX object CMakeFiles/smartbotic_common.dir/lib/common/config_defaults.cpp.o
+[2/8] Linking CXX static library libsmartbotic_common.a
+[3/8] Building CXX object CMakeFiles/smartbotic-runner.dir/src/runner/workflow_engine.cpp.o
+[4/8] Building CXX object CMakeFiles/smartbotic-webserver.dir/src/webserver/api/workflow_controller.cpp.o
+/data/smartbotic/src/webserver/api/workflow_controller.cpp: In member function 'void smartbotic::webserver::api::WorkflowController::executeWorkflow(const httplib::Request&, httplib::Response&, const smartbotic::webserver::auth::AuthContext&)':
+/data/smartbotic/src/webserver/api/workflow_controller.cpp:317:11: warning: unused variable 'workflow' [-Wunused-variable]
+  317 |     auto& workflow = workflow_result.value();
+      |           ^~~~~~~~
+[5/8] Building CXX object CMakeFiles/smartbotic-webserver.dir/src/webserver/webserver_service.cpp.o
+[6/8] Linking CXX executable smartbotic-webserver
+lto-wrapper: warning: using serial compilation of 33 LTRANS jobs
+lto-wrapper: note: see the 'flto' option documentation for more information
+[7/8] Linking CXX executable smartbotic-runner
+lto-wrapper: warning: using serial compilation of 44 LTRANS jobs
+lto-wrapper: note: see the 'flto' option documentation for more information
+```
+
+That single warning (`executeWorkflow`, unused `workflow` variable) is
+pre-existing and in code this change did not touch - confirmed by reading
+`executeWorkflow` before making any edits; it is untouched by this diff.
+No new warnings were introduced.
+
+## Scheduled-workflow count: before and after
+
+Both services were already running from a previous session, launched
+directly (not under systemd), logging to `/tmp/webserver.log` and
+`/tmp/runner.log`.
+
+Before-fix count, read from the existing log of the currently-running
+(pre-fix binary) webserver process:
+
+```
+$ grep -n "Loaded.*scheduled workflows" /tmp/webserver.log
+19:[2026-08-05 12:39:10.808] [webserver] [info] [7500] Loaded 1 scheduled workflows
+```
+
+Restart sequence (webserver first, then runner, each in its own bash call):
+
+```
+$ kill 7500 7571
+$ sleep 2; ps aux | grep -E "smartbotic-webserver|smartbotic-runner" | grep -v grep
+(no output - both terminated cleanly, no stale process needed a kill -9)
+$ ss -ltnp | grep -E "8090|9011|9012"
+(no output - ports free)
+$ rm -f /tmp/webserver.log && nohup ./build/smartbotic-webserver > /tmp/webserver.log 2>&1 &
+$ sleep 3; tail -30 /tmp/webserver.log
+```
+
+After-fix output (full relevant excerpt from the new webserver log):
+
+```
+[2026-08-05 13:41:13.573] [webserver] [info] [18313] Loading scheduled workflows from database...
+[2026-08-05 13:41:13.596] [webserver] [info] [18313] Scheduled workflow '35photo2anime' (wf_11784226-b91e-4e3e-8272-de6e314054d6) with schedule-trigger trigger, interval: 5 minutes, overlap: skip, maxConcurrent: 1
+[2026-08-05 13:41:13.604] [webserver] [info] [18313] Scheduled workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) with imap-trigger trigger, interval: 5 minutes, overlap: skip, maxConcurrent: 1
+[2026-08-05 13:41:13.631] [webserver] [info] [18313] Loaded 2 scheduled workflows
+```
+
+**Before: 1. After: 2.** The second workflow registered is exactly
+`wf_520f6f05-3256-4419-a8a5-42d7e7c3830f` ("Email OCR - Attachment to Text
+Reply"), the imap-trigger workflow whose stored config has no
+`pollInterval` key. It was not touched or re-saved; the fix made the
+already-stored, incomplete config register correctly on load.
+
+Runner restarted after confirming port 9011 was free:
+
+```
+$ ss -ltnp | grep 9011
+(no output - free)
+$ rm -f /tmp/runner.log && nohup ./build/smartbotic-runner > /tmp/runner.log 2>&1 &
+$ sleep 3; tail -40 /tmp/runner.log
+...
+[2026-08-05 13:41:22.216] [runner] [info] [18386] Loaded 56 node definitions from webserver
+[2026-08-05 13:41:22.217] [runner] [info] [18386] Node registry sync started with localhost:9012
+[2026-08-05 13:41:22.218] [runner] [info] [18386] Runner gRPC server listening on port 9011
+[2026-08-05 13:41:22.224] [runner] [info] [18386] Runner registered with webserver
+[2026-08-05 13:41:22.224] [runner] [info] [18386] Runner service runner-1 started
+```
+
+## Tests
+
+Two new fixtures added under `tests/nodes/`, both using the existing
+`code` node (which has a non-empty string `default` for its required
+`code` config property, and no defensive `||` fallback around it - so it
+directly exercises the runner's read-time defaulting):
+
+- `tests/nodes/config-defaults-fill.json` - a `code` node with `config: {}`
+  (no `code` key at all). Expects `status: completed` and
+  `output.result.processed === true`, which only happens if the runner
+  filled in the schema's default code text before execution. Without the
+  fix, `code.js`'s `if (!code ...) throw new Error('No code provided')`
+  would fail this node instead.
+- `tests/nodes/config-defaults-preserve-falsy.json` - a `code` node with
+  `config: {"code": ""}` (explicit falsy value for a defaulted key).
+  Expects `status: failed` with `errorContains: "No code provided"` -
+  proving `applyConfigDefaults` left the explicit empty string alone
+  instead of overwriting it with the non-empty default, which is exactly
+  the truthiness bug the design explicitly warned against.
+
+Individual runs:
+
+```
+$ python3 scripts/verify-node.py tests/nodes/config-defaults-fill.json
+case: verify-config-defaults-fill  execution: exec_4ea0c625-8e9c-4f57-a2b9-91513068b653  status: completed
+  n1                       completed  {...}
+  n2                       completed  {"executionTime": 0, "result": {"data": {...}, "processed": true, "timestamp": ...}}
+
+PASS
+
+$ python3 scripts/verify-node.py tests/nodes/config-defaults-preserve-falsy.json
+case: verify-config-defaults-preserve-falsy  execution: exec_3bef0971-35e2-4ad4-a533-8d0e513cf4fe  status: failed
+  n1                       completed  {...}
+  n2                       failed     null
+
+PASS
+```
+
+Full fixture suite (all files under `tests/nodes/*.json`, one
+`verify-node.py` invocation per file):
+
+```
+$ for f in tests/nodes/*.json; do python3 scripts/verify-node.py "$f"; done
+```
+
+Result: **39/39 passing** - the 37 pre-existing fixtures plus the 2 new
+ones above. Every pre-existing fixture still reports `PASS` with the same
+node statuses as before this change; no node's behaviour changed once
+defaults started being applied at read time and write time.
+
+## Nodes whose behaviour changed
+
+None observed. Every fixture that exercised a node with a `configSchema`
+default (`filter`, `set-fields`, `sort-limit-dedupe`, `datetime`,
+`switch`, `respond-to-webhook`, `wait-for-approval`, etc.) passed
+unchanged. Spot-checking the node source under `nodes/` turned up a
+recurring pattern: most nodes already defend themselves with
+`config.field || <fallback>` or explicit `=== undefined` checks that
+happen to match the schema's declared default, which is presumably why
+none of the 37 pre-existing fixtures moved. The `code` node is the
+exception - its required `code` property has no JS-side fallback at all,
+which is exactly why it was the clean way to prove the read-time path
+does something real.
+
+## Concerns / follow-ups
+
+- Not fixed here, out of scope per the task: `executeWorkflow()` in
+  `workflow_controller.cpp` does not currently apply config defaults
+  before dispatching a manual/API-triggered execution to a runner over
+  gRPC - only `loadScheduledWorkflows` and `updateScheduledTriggers` read
+  trigger config directly on the webserver side, and the runner itself
+  applies defaults for every node right before execution (including
+  triggers passed through `executeWorkflow`), so this path is covered
+  transitively through the runner, not duplicated on the webserver side.
+- `applyConfigDefaults` only fills top-level `configSchema.properties`
+  entries; nested object/array-item defaults are intentionally not
+  recursed into, documented in the header. If a future node relies on a
+  default nested inside an object-typed config property, it will need
+  either a schema restructure or a deliberate extension of this helper -
+  not a silent gap someone will trip over unknowingly.
+
+### Correction to the "negative fixture" claim above
+
+`tests/nodes/config-defaults-preserve-falsy.json` does not actually
+discriminate between the fixed and unfixed code. Against the unfixed
+`workflow_engine.cpp` (no `applyConfigDefaults` call at all),
+`config: {"code": ""}` still reaches `code.js`'s
+`if (!code ...) throw new Error('No code provided')` and fails the same
+way, because nothing was ever overwriting it to begin with - there was no
+defaulting code present to get the truthiness check wrong. The fixture
+proves the correct behaviour today and stands as a guard against a future
+truthiness regression in `applyConfigDefaults` (e.g. someone "simplifying"
+`!result.contains(key)` into a truthiness check), but it is not proof that
+the empty string survives some prior broken state, because no such broken
+state existed for this fixture to distinguish from. The positive fixture,
+`config-defaults-fill.json`, does discriminate correctly: it fails against
+the unfixed code and passes against the fixed code.
+
+---
+
+# Fix round 1
+
+Commit: `de2b70b77d4480c56522e7cb9fa1d9af7d91f1fc`
+
+## Findings addressed
+
+**Finding 1 (blocking) - loop bodies got no defaults.**
+`executeLoopBody` (`src/runner/workflow_engine.cpp`, ~line 2050) is a
+separate re-implementation of the main node walk and called
+`evaluateExpressions(body_node->config, ...)` directly, with no defaults
+applied - so a `code` node with `config: {}` inside a loop body failed with
+"No code provided" on every iteration, while the identical node outside
+the loop worked. Fixed by applying `applyConfigDefaults` to
+`body_node->config` before `evaluateExpressions`, using
+`registry_.getNode(body_node->type)`, mirroring the main walk's handling
+exactly (a failed lookup falls through and `executeNode` reports "Node
+type not found" as before, since an empty `std::optional` is passed
+through).
+
+**Finding 2 (non-blocking) - the added lookup doubled `NodeDefinition`
+copies (which carry the full JS source).**
+Chose: **hoist the lookup and reuse it**, rather than adding a
+by-reference accessor to `NodeRegistry`. `executeNode` now takes an
+optional fifth parameter, `const std::optional<NodeDefinition>&
+prefetched_node_def = std::nullopt` (declared in
+`src/runner/workflow_engine.hpp`). Both of the two call sites - the main
+walk (~line 669) and `executeLoopBody` (~line 2057) - already look up the
+node definition to apply defaults, so they now pass that same
+`std::optional<NodeDefinition>` straight into `executeNode`, which uses it
+if present instead of calling `registry_.getNode()` again. This was the
+smaller change: `NodeRegistry`'s `getNode()` already returns by value and
+is used that way from several other call sites in this file (lines 420,
+435, 546, 1867), so adding a second accessor would have meant two ways to
+fetch the same data; reusing what the caller already fetched keeps a
+single lookup path and drops the added lookup back down to one per node
+execution (previously two - one at the defaulting call site, one inside
+`executeNode` - and would have been three per loop iteration without this
+change).
+
+**Finding 3 (non-blocking, documentation).**
+- `lib/common/config_defaults.hpp`: added a paragraph recording that
+  write-time materialisation is permanent - once a default is baked into a
+  stored config, a later schema default change will never reach that
+  workflow, because the read path correctly leaves a present key alone.
+- `docs/nodes.md`: added a paragraph under "Configuration Schema" warning
+  that a `default:` containing `{{ }}` is evaluated as an expression, since
+  defaults flow through `evaluateExpressions` like any stored config value.
+
+## Build
+
+Command and full real output:
+
+```
+$ cmake --build build -j$(nproc)
+[1/9] Building CXX object CMakeFiles/smartbotic_common.dir/lib/common/config_defaults.cpp.o
+[2/9] Linking CXX static library libsmartbotic_common.a
+[3/9] Building CXX object CMakeFiles/smartbotic-runner.dir/src/runner/main.cpp.o
+[4/9] Building CXX object CMakeFiles/smartbotic-runner.dir/src/runner/runner_service.cpp.o
+[5/9] Building CXX object CMakeFiles/smartbotic-runner.dir/src/runner/workflow_engine.cpp.o
+[6/9] Building CXX object CMakeFiles/smartbotic-webserver.dir/src/webserver/api/workflow_controller.cpp.o
+/data/smartbotic/src/webserver/api/workflow_controller.cpp: In member function 'void smartbotic::webserver::api::WorkflowController::executeWorkflow(const httplib::Request&, httplib::Response&, const smartbotic::webserver::auth::AuthContext&)':
+/data/smartbotic/src/webserver/api/workflow_controller.cpp:317:11: warning: unused variable 'workflow' [-Wunused-variable]
+  317 |     auto& workflow = workflow_result.value();
+      |           ^~~~~~~~
+[7/9] Building CXX object CMakeFiles/smartbotic-webserver.dir/src/webserver/webserver_service.cpp.o
+[8/9] Linking CXX executable smartbotic-webserver
+lto-wrapper: warning: using serial compilation of 33 LTRANS jobs
+lto-wrapper: note: see the '-flto' option documentation for more information
+[9/9] Linking CXX executable smartbotic-runner
+lto-wrapper: warning: using serial compilation of 43 LTRANS jobs
+lto-wrapper: note: see the '-flto' option documentation for more information
+```
+
+Same single pre-existing warning as round 1 (unrelated `executeWorkflow`
+unused variable, code this change did not touch). No new warnings.
+
+## Restart
+
+```
+$ ps aux | grep -E "smartbotic-webserver|smartbotic-runner" | grep -v grep
+fszonta+   18313 ... ./build/smartbotic-webserver
+fszonta+   18386 ... ./build/smartbotic-runner
+$ kill 18313 18386
+$ sleep 2; ps aux | grep -E "smartbotic-webserver|smartbotic-runner" | grep -v grep
+fszonta+   18386 ... ./build/smartbotic-runner
+```
+
+The runner survived the plain `kill` (webserver did not). Confirmed and
+force-killed by PID:
+
+```
+$ kill -9 18386
+$ sleep 1; ps aux | grep -E "smartbotic-webserver|smartbotic-runner" | grep -v grep
+$ ss -ltnp | grep -E "8090|9011|9012"
+(no output - both processes gone, all three ports free)
+```
+
+Webserver started first:
+
+```
+$ rm -f /tmp/webserver.log && nohup /data/smartbotic/build/smartbotic-webserver > /tmp/webserver.log 2>&1 &
+$ sleep 3; tail -30 /tmp/webserver.log
+...
+[2026-08-05 13:57:59.374] [webserver] [info] [23060] Loading scheduled workflows from database...
+[2026-08-05 13:57:59.403] [webserver] [info] [23060] Scheduled workflow '35photo2anime' (wf_11784226-b91e-4e3e-8272-de6e314054d6) with schedule-trigger trigger, interval: 5 minutes, overlap: skip, maxConcurrent: 1
+[2026-08-05 13:57:59.411] [webserver] [info] [23060] Scheduled workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) with imap-trigger trigger, interval: 5 minutes, overlap: skip, maxConcurrent: 1
+[2026-08-05 13:57:59.440] [webserver] [info] [23060] Loaded 2 scheduled workflows
+...
+[2026-08-05 13:57:59.451] [webserver] [info] [23060] WebServer service started on port 8090
+```
+
+Port 9011 confirmed free before starting the runner:
+
+```
+$ ss -ltnp | grep 9011
+(no output)
+$ rm -f /tmp/runner.log && nohup /data/smartbotic/build/smartbotic-runner > /tmp/runner.log 2>&1 &
+$ sleep 3; tail -30 /tmp/runner.log
+[2026-08-05 13:58:08.404] [runner] [info] [23126] SmartBotic Runner starting...
+...
+[2026-08-05 13:58:08.487] [runner] [info] [23126] Loaded 56 node definitions from webserver
+...
+[2026-08-05 13:58:08.489] [runner] [info] [23126] Runner gRPC server listening on port 9011
+[2026-08-05 13:58:08.495] [runner] [info] [23126] Runner registered with webserver
+[2026-08-05 13:58:08.495] [runner] [info] [23126] Runner service runner-1 started
+```
+
+## Tests
+
+New fixture: `tests/nodes/config-defaults-loop-body.json` - a `code` node
+with an empty config placed as the sole body node of a `loop` iterating
+over `[1, 2]`. Its individual run:
+
+```
+$ python3 scripts/verify-node.py tests/nodes/config-defaults-loop-body.json
+case: verify-config-defaults-loop-body  execution: exec_05da8c05-ce25-4350-abd1-79c7efd9a7f8  status: completed
+  body                     completed  {"executionTime": 1, "result": {"data": 1, "processed": true, "timestamp": 1785931104318}}
+  items                    completed  {"executionTime": 0, "result": {"items": [1, 2]}}
+  loop                     completed  {"_activeBranch": "done", "_continueOnError": true, "_indexVariable": "index", "_isLoop": true, "_itemVariable": "item",
+  n1                       completed  {"executionId": "exec_05da8c05-ce25-4350-abd1-79c7efd9a7f8", "timestamp": 1785931104304, "triggeredBy": "manual"}
+
+PASS
+```
+
+Full fixture suite, one `verify-node.py` invocation per file under
+`tests/nodes/*.json`:
+
+```
+$ for f in tests/nodes/*.json; do
+    if python3 scripts/verify-node.py "$f" > /tmp/verify_out_$(basename "$f").txt 2>&1; then
+      pass=$((pass+1))
+    else
+      fail=$((fail+1)); failed_list="$failed_list $f"
+    fi
+  done
+  echo "PASS=$pass FAIL=$fail"
+PASS=40 FAIL=0
+FAILED:
+```
+
+**40/40 passing** - the 39 from round 1 plus this loop-body fixture.
+
+## Scheduled-workflow count: re-confirmed after restart
+
+```
+$ grep -n "Loaded.*scheduled workflows\|wf_520f6f05" /tmp/webserver.log
+19:[2026-08-05 13:57:59.411] [webserver] [info] [23060] Scheduled workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) with imap-trigger trigger, interval: 5 minutes, overlap: skip, maxConcurrent: 1
+20:[2026-08-05 13:57:59.440] [webserver] [info] [23060] Loaded 2 scheduled workflows
+```
+
+Still 2, `wf_520f6f05` still registered, still untouched.
+
+## Commit
+
+```
+$ git add -- docs/nodes.md lib/common/config_defaults.hpp src/runner/workflow_engine.cpp src/runner/workflow_engine.hpp tests/nodes/config-defaults-loop-body.json
+$ git commit -m "fix: apply config defaults inside loop bodies too, avoid extra source copies" ...
+[config-defaults de2b70b] fix: apply config defaults inside loop bodies too, avoid extra source copies
+ 5 files changed, 71 insertions(+), 8 deletions(-)
+ create mode 100644 tests/nodes/config-defaults-loop-body.json
+$ git log -1 --format="%H %G?"
+de2b70b77d4480c56522e7cb9fa1d9af7d91f1fc G
+```
+
+Signed (`G`), no pinentry issue.
+
+## Concerns
+
+- Named ports, `_webhookResponse`, `_pause` (twice), and now config
+  defaults have each had to be fixed twice - once in the main walk, once
+  in `executeLoopBody` - because the two are separate implementations.
+  Collapsing them is recorded as follow-up work and was explicitly not
+  this task's scope, but it remains the structural fix that would stop
+  this class of miss from recurring a sixth time.
+- The Finding 2 fix only threads the prefetched definition through the two
+  existing `executeNode` call sites. If a third call site is ever added
+  without also being told about `prefetched_node_def`, it will silently
+  fall back to `executeNode`'s own lookup - correct, just not optimal -
+  rather than fail loudly, so it is worth a second pair of eyes if
+  `executeNode` grows a new caller.
+
+---
+
+# Fix round 2 (final) - the TTL decision
+
+Commit: `ccdc442fc19fe235b8a1a4a46e985ed31f898c4c`
+
+## Background
+
+`nodes/core/http-request.js` read `config.downloadTtlHours || 0` and
+`nodes/imap/imap-extract-attachments.js` read `config.storageTtlHours || 0`,
+while both schemas declare `default: 24`. Before this branch, a workflow
+whose config omitted the key stored downloads/attachments forever - the
+`||` fallback won and the schema default never arrived. With defaults now
+applied by this branch, those same already-saved workflows get 24 and the
+data starts expiring after a day. The project owner decided to KEEP the
+24-hour expiry: the schema is the intended behaviour, and unbounded
+storage growth is what the default was written to prevent in the first
+place. This round makes the code agree with that decision.
+
+## Changes
+
+- `nodes/core/http-request.js` line ~338:
+  `config.downloadTtlHours || 0` -> `config.downloadTtlHours ?? 24`.
+- `nodes/imap/imap-extract-attachments.js` line ~442:
+  `config.storageTtlHours || 0` -> `config.storageTtlHours ?? 24`.
+- Nullish coalescing, not `||`, so an explicit `0` - documented in both
+  schemas as "never expire" - still survives instead of being treated as
+  falsy and overridden.
+- Checked both files for other reads of the same config key:
+  `grep -n "downloadTtlHours\|storageTtlHours"` against each file showed
+  exactly one schema declaration and one read site per file. Nothing else
+  to make consistent.
+- Field descriptions updated in both schemas to state the default and the
+  0-means-never behaviour plainly:
+  - `http-request.js` `downloadTtlHours.description`: "Auto-delete stored
+    file after this many hours. Default is 24 hours; set to 0 to keep it
+    forever."
+  - `imap-extract-attachments.js` `storageTtlHours.description`:
+    "Auto-delete stored files after this many hours. Default is 24 hours;
+    set to 0 to keep them forever."
+- `docs/nodes.md`, under "Storage (Database)", gained a paragraph stating
+  that `http-request` (Store Download) and `imap-extract-attachments`
+  (Store in Database) both default their TTL to 24 hours, so stored
+  downloads and extracted attachments now expire a day after they're
+  stored unless the workflow explicitly sets the TTL field to `0`, and
+  that any workflow relying on permanent storage under an unset TTL field
+  needs that `0` set explicitly or it will start losing data a day later.
+
+## Whether the explicit-zero fixtures were possible, and what they actually prove
+
+Both fixtures were possible and were written - `http-request`'s
+`storeDownload` path was exercised against a real local HTTP endpoint
+(`http://localhost:8090/index.html`, served by the already-running
+webserver's static file mount), and `imap-extract-attachments`'s
+`storeInDatabase` path was exercised by feeding a hand-built raw MIME
+multipart email through `emailSource`, with no live IMAP mailbox needed
+since that node accepts raw email text as config/input.
+
+What they can and cannot prove was checked directly, not assumed. Before
+writing them, I probed whether `smartbotic.storage.insert(..., ttlMs)`'s
+TTL is readable back through the JS API available to nodes:
+
+```
+$ python3 -c "... insert with ttlMs=5000, then storage.get() on the same id ..."
+```
+
+Full real output of the returned document:
+
+```json
+{
+  "collection": "ttl_probe_test",
+  "document": {
+    "_created_at": 1785932269775728600,
+    "_created_by": "",
+    "_id": "019fd1dbc8cfcd70ef3a1be5b435",
+    "_updated_at": 1785932269775728600,
+    "_updated_by": "",
+    "_version": 1,
+    "probe": true
+  },
+  "found": true,
+  "id": "019fd1dbc8cfcd70ef3a1be5b435"
+}
+```
+
+No TTL or expiry field is present. `lib/storage/storage_client.hpp` /
+`.cpp` and the QuickJS binding in
+`src/runner/engine/script_engine.cpp` (`storage.insert`, `storage.get`,
+`storage.query`) confirm there is no accessor that returns a document's
+TTL or expiry timestamp back to a node - `insert()` takes `ttl_ms` and
+converts it to seconds for the upstream client, one-way. So:
+
+- The two fixtures **do** prove that `smartbotic.storage.insert` is
+  reached and completes successfully (returns `success: true`, the node
+  returns a `storage: {collection, id}` object) when the TTL field is
+  explicitly `0` - i.e. the storage path is genuinely exercised, not
+  skipped or thrown on.
+- The two fixtures **cannot** prove that the TTL value that reached
+  `storage.insert` was actually `0` rather than the pre-`??`-fix `24`
+  (in hours) or any other value - there is no way to read that back
+  through the harness, and waiting out a real 24-hour expiry to observe
+  the difference empirically is not practical for this suite. That part
+  of the guarantee rests on the `?? 24` change being correct at the two
+  read sites, which is a one-line, directly-readable diff in each file
+  (confirmed there is exactly one read site per file, above) rather than
+  something a black-box fixture can independently verify.
+
+Fixture files:
+- `tests/nodes/config-defaults-ttl-zero-http-request.json`
+- `tests/nodes/config-defaults-ttl-zero-imap-extract-attachments.json`
+
+Individual runs, full real output:
+
+```
+$ python3 scripts/verify-node.py tests/nodes/config-defaults-ttl-zero-http-request.json
+case: verify-config-defaults-ttl-zero-http-request  execution: exec_1fa17e8a-127f-4dea-ac3e-4608cb1a7eb0  status: completed
+  dl                       completed  {"body": null, "checksum": "7f09ea737005b3abf0fa01db11929491a119ac59ed433d2635c980c0c9473caf", "file": {"checksum": "7f0
+  n1                       completed  {"executionId": "exec_1fa17e8a-127f-4dea-ac3e-4608cb1a7eb0", "timestamp": 1785932459567, "triggeredBy": "manual"}
+
+PASS
+
+$ python3 scripts/verify-node.py tests/nodes/config-defaults-ttl-zero-imap-extract-attachments.json
+case: verify-config-defaults-ttl-zero-imap-extract-attachments  execution: exec_7f7f2770-dc43-4797-ab27-50581569fb68  status: completed
+  extract                  completed  {"attachments": [{"contentId": null, "deduplicated": false, "filePath": "./data/attachments/2026-08/66f6263d-78ce-4e30-a
+  mail                     completed  {"executionTime": 1, "result": {"raw": "Content-Type: multipart/mixed; boundary=\"BOUNDARY123\"\n\n--BOUNDARY123\nConten
+  n1                       completed  {"executionId": "exec_7f7f2770-dc43-4797-ab27-50581569fb68", "timestamp": 1785932461151, "triggeredBy": "manual"}
+
+PASS
+```
+
+## Full fixture suite
+
+```
+$ for f in tests/nodes/*.json; do
+    if python3 scripts/verify-node.py "$f" > /tmp/verify_out_$(basename "$f").txt 2>&1; then
+      pass=$((pass+1))
+    else
+      fail=$((fail+1)); failed_list="$failed_list $f"
+    fi
+  done
+  echo "PASS=$pass FAIL=$fail"
+PASS=42 FAIL=0
+FAILED:
+```
+
+**42/42 passing** - the 40 from round 1 plus these 2 TTL fixtures.
+
+## Scheduled-workflow count: re-confirmed
+
+No rebuild or restart was performed or needed for this round - both
+changed files are hot-reloaded JavaScript node source, not C++. The
+webserver and runner from round 1's restart were still running.
+
+```
+$ grep -n "Loaded.*scheduled workflows\|wf_520f6f05" /tmp/webserver.log
+19:[2026-08-05 13:57:59.411] [webserver] [info] [23060] Scheduled workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) with imap-trigger trigger, interval: 5 minutes, overlap: skip, maxConcurrent: 1
+20:[2026-08-05 13:57:59.440] [webserver] [info] [23060] Loaded 2 scheduled workflows
+594:[2026-08-05 14:03:39.401] [webserver] [info] [23089] Scheduler executing workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) - imap-trigger trigger
+598:[2026-08-05 14:03:40.226] [webserver] [info] [23089] Scheduled workflow wf_520f6f05-3256-4419-a8a5-42d7e7c3830f execution started: exec_606211f9-fba8-4805-9adb-84bf2f565056 on runner runner-1
+759:[2026-08-05 14:09:10.255] [webserver] [info] [23089] Scheduler executing workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) - imap-trigger trigger
+779:[2026-08-05 14:09:10.997] [webserver] [info] [23089] Scheduled workflow wf_520f6f05-3256-4419-a8a5-42d7e7c3830f execution started: exec_58d3c9f8-705a-4384-8422-db4d5b047ed6 on runner runner-1
+925:[2026-08-05 14:14:40.625] [webserver] [info] [23089] Scheduler executing workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) - imap-trigger trigger
+929:[2026-08-05 14:14:41.395] [webserver] [info] [23089] Scheduled workflow wf_520f6f05-3256-4419-a8a5-42d7e7c3830f execution started: exec_4dc881fc-986e-40d2-863a-05820a4deb80 on runner runner-1
+1024:[2026-08-05 14:20:11.422] [webserver] [info] [23089] Scheduler executing workflow 'Email OCR - Attachment to Text Reply' (wf_520f6f05-3256-4419-a8a5-42d7e7c3830f) - imap-trigger trigger
+1027:[2026-08-05 14:20:12.165] [webserver] [info] [23089] Scheduled workflow wf_520f6f05-3256-4419-a8a5-42d7e7c3830f execution started: exec_01e9a248-93d8-4838-b1d2-fcd36693d7f8 on runner runner-1
+```
+
+Still 2, `wf_520f6f05` still registered and unmodified - and, beyond just
+being registered, its imap-trigger has been firing on schedule
+(4 scheduled executions logged between 13:58 and 14:20) throughout this
+whole review round, which is the round-1 fix holding up under real,
+continued operation rather than a one-time startup check.
+
+## Commit
+
+```
+$ git add -- docs/nodes.md nodes/core/http-request.js nodes/imap/imap-extract-attachments.js tests/nodes/config-defaults-ttl-zero-http-request.json tests/nodes/config-defaults-ttl-zero-imap-extract-attachments.json
+$ git commit -m "fix: make http-request and imap-extract-attachments TTL fallbacks agree with their schema defaults" ...
+[config-defaults ccdc442] fix: make http-request and imap-extract-attachments TTL fallbacks agree with their schema defaults
+ 5 files changed, 56 insertions(+), 4 deletions(-)
+ create mode 100644 tests/nodes/config-defaults-ttl-zero-http-request.json
+ create mode 100644 tests/nodes/config-defaults-ttl-zero-imap-extract-attachments.json
+$ git log -1 --format="%H %G?"
+ccdc442fc19fe235b8a1a4a46e985ed31f898c4c G
+```
+
+Signed (`G`), no pinentry issue.
+
+## Which nodes now behave differently from before this branch, and how
+
+This is the accumulated, user-visible behaviour change across all three
+rounds on this branch, for anyone reading this report to understand
+impact:
+
+- **Every node with a `configSchema` default**, across both the runner
+  (including inside loop bodies) and the webserver scheduler, now
+  receives that default when its stored config omits the key - on
+  already-saved workflows, not just newly-saved ones. Spot-checked
+  against the 37 pre-existing fixtures plus manual review of node
+  source: no other node's fixture-observable behaviour changed, because
+  most nodes already defended themselves with a JS-level fallback that
+  happened to match the schema default. `nodes/core/code.js` is the one
+  node in the fixture suite with no such defensive fallback for its
+  required `code` field, which is why it was used to prove the mechanism
+  works at all (`config-defaults-fill.json`,
+  `config-defaults-preserve-falsy.json`,
+  `config-defaults-loop-body.json`).
+- **`nodes/core/http-request.js`** (Store Download) and
+  **`nodes/imap/imap-extract-attachments.js`** (Store in Database) are
+  the two nodes whose behaviour has materially and deliberately changed
+  for real, already-running workflows. Previously `|| 0` meant a config
+  that omitted `downloadTtlHours` / `storageTtlHours` stored data forever
+  (the schema's `default: 24` never reached the node). Now, with defaults
+  applied, an omitted key resolves to `24` and data stored by these two
+  nodes is auto-deleted 24 hours after creation, unless the workflow's
+  config explicitly sets the TTL field to `0`. This is an intentional,
+  owner-approved change in behaviour, not a bug - the report warns of it
+  here and in `docs/nodes.md` because it is the one change on this branch
+  that can cause silent data loss for an existing workflow that nobody
+  told to expect it.
+- **`wf_520f6f05-3256-4419-a8a5-42d7e7c3830f`** is the confirmed
+  real-world example of the positive side of this same mechanism: its
+  imap-trigger's missing `pollInterval` now resolves to the schema's
+  `default: 5`, and the workflow is scheduled and firing again after
+  being silently dead since it was saved.
+
+## Concerns
+
+- The data-loss risk flagged above is real and immediate: any production
+  workflow using `http-request` with Store Download enabled, or
+  `imap-extract-attachments` with Store in Database enabled, and no
+  explicit TTL value in its saved config, will start deleting that stored
+  data 24 hours after each item is stored, starting from whenever this
+  branch reaches production. `docs/nodes.md` and both field descriptions
+  now say so, but nothing in the system will proactively surface this to
+  an existing workflow's owner - it is documentation, not a migration or
+  a warning banner. Whether some active-workflow scan or one-time
+  notification is warranted is a product decision beyond this branch's
+  scope, but worth raising explicitly since the owner's decision was to
+  accept the new expiry rather than the old unbounded growth.
+- The two new TTL fixtures are deliberately scoped to what the harness
+  can prove (the storage call succeeds with an explicit 0) and explicitly
+  cannot prove the numeric TTL value used. If `storage.insert`'s TTL ever
+  becomes introspectable through the JS API, these fixtures should be
+  strengthened to assert on the actual value rather than just successful
+  completion.

+ 442 - 0
.superpowers/review-config-defaults.diff

@@ -0,0 +1,442 @@
+3bd17d0 fix: apply configSchema defaults at write time and read time
+
+ CMakeLists.txt                                  |  1 +
+ lib/common/config_defaults.cpp                  | 39 +++++++++++++++++++++++++
+ lib/common/config_defaults.hpp                  | 27 +++++++++++++++++
+ src/runner/workflow_engine.cpp                  |  9 +++++-
+ src/webserver/api/workflow_controller.cpp       | 39 +++++++++++++++++++++++--
+ src/webserver/api/workflow_controller.hpp       |  4 +++
+ src/webserver/webserver_service.cpp             |  7 +++--
+ tests/nodes/config-defaults-fill.json           | 14 +++++++++
+ tests/nodes/config-defaults-preserve-falsy.json | 14 +++++++++
+ 9 files changed, 148 insertions(+), 6 deletions(-)
+
+diff --git a/CMakeLists.txt b/CMakeLists.txt
+index 4499987..6aef080 100644
+--- a/CMakeLists.txt
++++ b/CMakeLists.txt
+@@ -12,20 +12,21 @@ list(APPEND CMAKE_MODULE_PATH "${CMAKE_CURRENT_SOURCE_DIR}/cmake")
+ include(CompilerFlags)
+ include(Dependencies)
+ include(FindPackages)
+ 
+ # Common library
+ add_library(smartbotic_common STATIC
+     lib/common/uuid.cpp
+     lib/common/time_utils.cpp
+     lib/common/error.cpp
+     lib/common/string_utils.cpp
++    lib/common/config_defaults.cpp
+ )
+ target_include_directories(smartbotic_common PUBLIC
+     ${CMAKE_CURRENT_SOURCE_DIR}/lib
+ )
+ target_link_libraries(smartbotic_common PUBLIC
+     nlohmann_json::nlohmann_json
+     OpenSSL::Crypto
+ )
+ 
+ # Logging library
+diff --git a/lib/common/config_defaults.cpp b/lib/common/config_defaults.cpp
+new file mode 100644
+index 0000000..a5bbfd8
+--- /dev/null
++++ b/lib/common/config_defaults.cpp
+@@ -0,0 +1,39 @@
++#include "common/config_defaults.hpp"
++
++namespace smartbotic::common {
++
++nlohmann::json applyConfigDefaults(const nlohmann::json& config,
++                                    const nlohmann::json& config_schema) {
++    nlohmann::json result = config.is_object() ? config : nlohmann::json::object();
++
++    if (!config_schema.is_object()) {
++        return result;
++    }
++
++    auto properties_it = config_schema.find("properties");
++    if (properties_it == config_schema.end() || !properties_it->is_object()) {
++        return result;
++    }
++
++    for (auto it = properties_it->begin(); it != properties_it->end(); ++it) {
++        const std::string& key = it.key();
++        const nlohmann::json& property_schema = it.value();
++
++        if (!property_schema.is_object()) {
++            continue;
++        }
++
++        auto default_it = property_schema.find("default");
++        if (default_it == property_schema.end()) {
++            continue;
++        }
++
++        if (!result.contains(key)) {
++            result[key] = *default_it;
++        }
++    }
++
++    return result;
++}
++
++} // namespace smartbotic::common
+diff --git a/lib/common/config_defaults.hpp b/lib/common/config_defaults.hpp
+new file mode 100644
+index 0000000..ad504da
+--- /dev/null
++++ b/lib/common/config_defaults.hpp
+@@ -0,0 +1,27 @@
++#pragma once
++
++#include <nlohmann/json.hpp>
++
++namespace smartbotic::common {
++
++// Applies configSchema-declared defaults to a node config.
++//
++// Returns a copy of `config` with any property from
++// `config_schema["properties"]` that declares a "default" filled in, but
++// only when `config` does not already contain that key. A key that is
++// present - even with a falsy value like false, 0, null, or "" - is a
++// deliberate, stored value and is never overwritten.
++//
++// Scope limit (by design, not an oversight): only top-level properties are
++// considered. Defaults nested inside object properties or array item
++// schemas are NOT applied. Nodes are not currently written with nested
++// config shapes that rely on defaults, so recursing was left out to keep
++// this function's behaviour easy to reason about; revisit if that changes.
++//
++// Tolerant of a missing/null/non-object schema, or one with no
++// "properties" - in all of those cases the config is returned unchanged
++// rather than throwing.
++nlohmann::json applyConfigDefaults(const nlohmann::json& config,
++                                    const nlohmann::json& config_schema);
++
++} // namespace smartbotic::common
+diff --git a/src/runner/workflow_engine.cpp b/src/runner/workflow_engine.cpp
+index e475d9c..bf2a85d 100644
+--- a/src/runner/workflow_engine.cpp
++++ b/src/runner/workflow_engine.cpp
+@@ -1,13 +1,14 @@
+ #include "workflow_engine.hpp"
+ #include "common/uuid.hpp"
+ #include "common/time_utils.hpp"
++#include "common/config_defaults.hpp"
+ #include "logging/logger.hpp"
+ #include <algorithm>
+ #include <functional>
+ #include <stack>
+ #include <unordered_set>
+ 
+ namespace smartbotic::runner {
+ 
+ using namespace common;
+ 
+@@ -641,21 +642,27 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
+                         {"nodeId", node_id},
+                         {"status", "completed"},
+                         {"output", node_result.output},
+                         {"fromCache", true}
+                     });
+                 }
+             } else {
+                 // Evaluate expressions in node config
+                 WorkflowNode evaluated_node = *node;
+                 t_disabled_reference_error.clear();
+-                evaluated_node.config = evaluateExpressions(node->config, input, result.node_results, workflow);
++                nlohmann::json defaulted_config = node->config;
++                auto node_def_for_defaults = registry_.getNode(node->type);
++                if (node_def_for_defaults) {
++                    defaulted_config = smartbotic::common::applyConfigDefaults(
++                        node->config, node_def_for_defaults->config_schema);
++                }
++                evaluated_node.config = evaluateExpressions(defaulted_config, input, result.node_results, workflow);
+ 
+                 if (!t_disabled_reference_error.empty()) {
+                     node_result.node_id = node_id;
+                     node_result.status = NodeStatus::Failed;
+                     node_result.input = input;
+                     node_result.output = nlohmann::json::object();
+                     node_result.error = t_disabled_reference_error;
+                     node_result.started_at = TimeUtils::nowMs();
+                     node_result.finished_at = node_result.started_at;
+                 } else {
+diff --git a/src/webserver/api/workflow_controller.cpp b/src/webserver/api/workflow_controller.cpp
+index 4be6024..aa0021c 100644
+--- a/src/webserver/api/workflow_controller.cpp
++++ b/src/webserver/api/workflow_controller.cpp
+@@ -1,13 +1,14 @@
+ #include "workflow_controller.hpp"
+ #include "common/uuid.hpp"
+ #include "common/time_utils.hpp"
++#include "common/config_defaults.hpp"
+ #include "logging/logger.hpp"
+ #include <grpcpp/grpcpp.h>
+ 
+ namespace smartbotic::webserver::api {
+ 
+ using namespace common;
+ 
+ WorkflowController::WorkflowController(storage::StorageClient& storage,
+                                        auth::AuthMiddleware& middleware,
+                                        runners::RunnerRegistry& registry,
+@@ -16,20 +17,45 @@ WorkflowController::WorkflowController(storage::StorageClient& storage,
+                                        WorkflowScheduler& scheduler,
+                                        nodes::NodeStore& node_store)
+     : storage_(storage)
+     , middleware_(middleware)
+     , registry_(registry)
+     , load_balancer_(load_balancer)
+     , ws_server_(ws_server)
+     , scheduler_(scheduler)
+     , node_store_(node_store) {}
+ 
++void WorkflowController::materializeNodeConfigDefaults(nlohmann::json& body) {
++    if (!body.contains("nodes") || !body["nodes"].is_array()) {
++        return;
++    }
++
++    for (auto& node : body["nodes"]) {
++        if (!node.is_object()) {
++            continue;
++        }
++
++        std::string node_type = node.value("type", "");
++        if (node_type.empty()) {
++            continue;
++        }
++
++        auto node_result = node_store_.get(node_type);
++        if (node_result.failed()) {
++            continue;
++        }
++
++        auto config = node.value("config", nlohmann::json::object());
++        node["config"] = common::applyConfigDefaults(config, node_result.value().config_schema);
++    }
++}
++
+ void WorkflowController::registerRoutes(httplib::Server& server) {
+     server.Get("/api/v1/workflows", [this](const httplib::Request& req, httplib::Response& res) {
+         middleware_.requireAuth(req, res, [this](auto& req, auto& res, auto& ctx) {
+             listWorkflows(req, res, ctx);
+         });
+     });
+ 
+     server.Get(R"(/api/v1/workflows/([^/]+))", [this](const httplib::Request& req, httplib::Response& res) {
+         middleware_.requireAuth(req, res, [this](auto& req, auto& res, auto& ctx) {
+             getWorkflow(req, res, ctx);
+@@ -165,20 +191,22 @@ void WorkflowController::createWorkflow(const httplib::Request& req, httplib::Re
+         // Set owner - metadata fields (_id, _createdAt, _updatedAt, _version) are handled by database
+         body["ownerId"] = ctx.user_id;
+         body["active"] = body.value("active", false);
+ 
+         // Validate required fields
+         if (!body.contains("name") || body["name"].get<std::string>().empty()) {
+             sendError(res, "Name is required", 400);
+             return;
+         }
+ 
++        materializeNodeConfigDefaults(body);
++
+         auto result = storage_.insert("workflows", body, id);
+         if (result.failed()) {
+             sendError(res, result.error().message(), 500);
+             return;
+         }
+ 
+         // Get the created workflow with metadata
+         auto workflow = storage_.get("workflows", id);
+         if (workflow.failed()) {
+             sendError(res, "Failed to retrieve created workflow", 500);
+@@ -215,20 +243,22 @@ void WorkflowController::updateWorkflow(const httplib::Request& req, httplib::Re
+         // Note: updatedAt is managed by database automatically
+ 
+         // Remember the current run state so the reconciliation below can tell
+         // whether this update actually flipped it.
+         bool was_active = false;
+         auto before = storage_.get("workflows", id);
+         if (before.ok()) {
+             was_active = before.value().value("active", false);
+         }
+ 
++        materializeNodeConfigDefaults(body);
++
+         auto result = storage_.update("workflows", id, body, 0, true);
+         if (result.failed()) {
+             sendError(res, result.error().message(), 404);
+             return;
+         }
+ 
+         // Get updated workflow
+         auto workflow = storage_.get("workflows", id);
+         if (workflow.ok()) {
+             // Reconcile the scheduler in both directions. Registering on active
+@@ -402,22 +432,24 @@ void WorkflowController::updateScheduledTriggers(const std::string& workflow_id,
+         }
+ 
+         auto& node_def = node_result.value();
+ 
+         // Check if node has scheduling capability (is_scheduled flag or pollInterval config)
+         bool is_scheduled = node_def.is_scheduled;
+         if (!is_scheduled) {
+             continue;
+         }
+ 
+-        // Get interval from node config
+-        auto config = node.value("config", nlohmann::json::object());
++        // Get interval from node config, filling in any schema defaults the
++        // stored config is missing.
++        auto config = common::applyConfigDefaults(
++            node.value("config", nlohmann::json::object()), node_def.config_schema);
+         int interval = config.value("pollInterval", 0);
+ 
+         if (interval > 0) {
+             scheduler_.registerWorkflow(
+                 workflow_id,
+                 workflow_name,
+                 node_id,
+                 node_type,
+                 interval
+             );
+@@ -430,21 +462,22 @@ void WorkflowController::updateScheduledTriggers(const std::string& workflow_id,
+ int WorkflowController::getScheduledInterval(const nlohmann::json& workflow) {
+     auto nodes = workflow.value("nodes", nlohmann::json::array());
+     for (const auto& node : nodes) {
+         std::string node_type = node.value("type", "");
+ 
+         auto node_result = node_store_.get(node_type);
+         if (node_result.failed() || !node_result.value().is_trigger || !node_result.value().is_scheduled) {
+             continue;
+         }
+ 
+-        auto config = node.value("config", nlohmann::json::object());
++        auto config = common::applyConfigDefaults(
++            node.value("config", nlohmann::json::object()), node_result.value().config_schema);
+         return config.value("pollInterval", 0);
+     }
+     return 0;
+ }
+ 
+ void WorkflowController::sendJson(httplib::Response& res, const nlohmann::json& data, int status) {
+     res.status = status;
+     res.set_content(data.dump(), "application/json");
+ }
+ 
+diff --git a/src/webserver/api/workflow_controller.hpp b/src/webserver/api/workflow_controller.hpp
+index 2447ceb..a748ae1 100644
+--- a/src/webserver/api/workflow_controller.hpp
++++ b/src/webserver/api/workflow_controller.hpp
+@@ -53,13 +53,17 @@ private:
+     auth::AuthMiddleware& middleware_;
+     runners::RunnerRegistry& registry_;
+     runners::LoadBalancer& load_balancer_;
+     WebSocketServer& ws_server_;
+     WorkflowScheduler& scheduler_;
+     nodes::NodeStore& node_store_;
+ 
+     // Helper to check for scheduled triggers and register/unregister with scheduler
+     void updateScheduledTriggers(const std::string& workflow_id, bool activate);
+     int getScheduledInterval(const nlohmann::json& workflow);
++
++    // Fills in each node's stored config with any missing configSchema
++    // defaults, in place, so a saved workflow document is self-describing.
++    void materializeNodeConfigDefaults(nlohmann::json& body);
+ };
+ 
+ } // namespace smartbotic::webserver::api
+diff --git a/src/webserver/webserver_service.cpp b/src/webserver/webserver_service.cpp
+index a7d7f13..124419a 100644
+--- a/src/webserver/webserver_service.cpp
++++ b/src/webserver/webserver_service.cpp
+@@ -10,20 +10,21 @@
+ #include "api/database_controller.hpp"
+ #include "api/file_controller.hpp"
+ #include "api/proxy_controller.hpp"
+ #include "api/credential_controller.hpp"
+ #include "nodes/node_store.hpp"
+ #include "grpc/node_sync_service.hpp"
+ #include "grpc/credential_service.hpp"
+ #include "credentials/credential_store.hpp"
+ #include "scheduler/workflow_scheduler.hpp"
+ #include "common/time_utils.hpp"
++#include "common/config_defaults.hpp"
+ #include "logging/logger.hpp"
+ #include <grpcpp/grpcpp.h>
+ #include "proto/runner.grpc.pb.h"
+ 
+ namespace smartbotic::webserver {
+ 
+ WebServerService::WebServerService(const WebServerServiceConfig& config)
+     : config_(config) {
+ 
+     // Initialize storage client
+@@ -306,22 +307,24 @@ void WebServerService::loadScheduledWorkflows() {
+             auto node_result = node_store_->get(node_type);
+             if (node_result.failed()) {
+                 continue;
+             }
+ 
+             const auto& node_def = node_result.value();
+             if (!node_def.is_trigger || !node_def.is_scheduled) {
+                 continue;
+             }
+ 
+-            // Get interval from node config
+-            auto config = node.value("config", nlohmann::json::object());
++            // Get interval from node config, filling in any schema defaults
++            // the stored config is missing (e.g. an untouched form field).
++            auto config = smartbotic::common::applyConfigDefaults(
++                node.value("config", nlohmann::json::object()), node_def.config_schema);
+             int interval = config.value("pollInterval", 0);
+ 
+             if (interval > 0) {
+                 auto policy = overlapPolicyFromString(
+                     config.value("overlapPolicy", std::string("skip")));
+ 
+                 scheduler_->registerWorkflow(
+                     workflow_id,
+                     workflow_name,
+                     node_id,
+diff --git a/tests/nodes/config-defaults-fill.json b/tests/nodes/config-defaults-fill.json
+new file mode 100644
+index 0000000..ec715f0
+--- /dev/null
++++ b/tests/nodes/config-defaults-fill.json
+@@ -0,0 +1,14 @@
++{
++  "name": "verify-config-defaults-fill",
++  "nodes": [
++    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
++    {"id": "n2", "name": "NoConfig", "type": "code", "position": {"x": 0, "y": 100},
++     "config": {}}
++  ],
++  "connections": [
++    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "n2", "targetInput": "data"}
++  ],
++  "expect": {
++    "n2": {"status": "completed", "output": {"result": {"processed": true}}}
++  }
++}
+diff --git a/tests/nodes/config-defaults-preserve-falsy.json b/tests/nodes/config-defaults-preserve-falsy.json
+new file mode 100644
+index 0000000..6b84e8c
+--- /dev/null
++++ b/tests/nodes/config-defaults-preserve-falsy.json
+@@ -0,0 +1,14 @@
++{
++  "name": "verify-config-defaults-preserve-falsy",
++  "nodes": [
++    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
++    {"id": "n2", "name": "EmptyCode", "type": "code", "position": {"x": 0, "y": 100},
++     "config": {"code": ""}}
++  ],
++  "connections": [
++    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "n2", "targetInput": "data"}
++  ],
++  "expect": {
++    "n2": {"status": "failed", "errorContains": "No code provided"}
++  }
++}

+ 21 - 0
scripts/verify-node.py

@@ -129,6 +129,25 @@ def main():
 
     call("POST", "/nodes/migrate", token, {"nodesPath": "./nodes"})
 
+    # Workflows the case needs to exist before it runs - a sub-workflow it
+    # calls. Created here rather than named by id, so a case is not silently
+    # tied to a row someone made by hand once and can delete.
+    helper_ids = []
+    for helper in case.get("helpers", []):
+        made = call("POST", "/workflows", token, {
+            "name": helper["name"],
+            "nodes": helper["nodes"],
+            "connections": helper.get("connections", []),
+        })
+        helper_id = made.get("id") or made.get("_id")
+        if not helper_id:
+            raise SystemExit(f"no id for helper {helper['key']}")
+        helper_ids.append(helper_id)
+        if helper.get("active", True):
+            call("POST", f"/workflows/{helper_id}/activate", token, {})
+        # Substituted everywhere, so the case refers to it by name.
+        case = json.loads(json.dumps(case).replace("{{helper:" + helper["key"] + "}}", helper_id))
+
     created = call("POST", "/workflows", token, {
         "name": case["name"],
         "nodes": case["nodes"],
@@ -206,6 +225,8 @@ def main():
         return 0
     finally:
         call("DELETE", f"/workflows/{workflow_id}", token)
+        for helper_id in helper_ids:
+            call("DELETE", f"/workflows/{helper_id}", token)
 
 
 if __name__ == "__main__":

BIN
sdcpp-configurator.png


+ 69 - 2
src/runner/workflow_engine.cpp

@@ -283,6 +283,14 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
         actual_trigger_data.erase("_callDepth");
     }
 
+    // An id chosen by the caller, so it could record the child before the child
+    // existed. Read here and taken as this run's id.
+    std::string assigned_execution_id;
+    if (actual_trigger_data.contains("_childExecutionId")) {
+        assigned_execution_id = actual_trigger_data.value("_childExecutionId", "");
+        actual_trigger_data.erase("_childExecutionId");
+    }
+
     if (actual_trigger_data.contains("_resumeExecutionId")) {
         resume_execution_id = actual_trigger_data["_resumeExecutionId"].get<std::string>();
         actual_trigger_data.erase("_resumeExecutionId");
@@ -300,6 +308,9 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
     result.started_at = TimeUtils::nowMs();
     result.call_depth = call_depth;
 
+    if (!assigned_execution_id.empty()) {
+        result.execution_id = assigned_execution_id;
+    }
     if (!resume_execution_id.empty()) {
         result.execution_id = resume_execution_id;
     }
@@ -1129,8 +1140,28 @@ common::Result<ExecutionResult> WorkflowEngine::resume(const std::string& execut
 
 void WorkflowEngine::cancelExecution(const std::string& execution_id) {
     std::lock_guard<std::mutex> lock(mutex_);
-    cancelled_executions_.insert(execution_id);
-    LOG_INFO("Cancellation requested for execution: {}", execution_id);
+
+    // Cancelling a run cancels everything it started. Without this, stopping a
+    // workflow that calls another one stops nothing anyone can see: the called
+    // workflow runs to the end, and the caller only notices between steps, so a
+    // loop of minute-long calls keeps going for minutes.
+    std::vector<std::string> pending{execution_id};
+    size_t cancelled = 0;
+    while (!pending.empty()) {
+        const std::string id = pending.back();
+        pending.pop_back();
+        if (!cancelled_executions_.insert(id).second) {
+            continue;  // already marked, and so are its children
+        }
+        ++cancelled;
+        auto children = child_executions_.find(id);
+        if (children != child_executions_.end()) {
+            pending.insert(pending.end(), children->second.begin(), children->second.end());
+        }
+    }
+
+    LOG_INFO("Cancellation requested for execution {} ({} execution(s) marked)",
+             execution_id, cancelled);
 }
 
 int WorkflowEngine::getActiveExecutionCount() const {
@@ -1691,6 +1722,18 @@ bool WorkflowEngine::runSubWorkflow(NodeExecutionResult& node_result, int call_d
         return false;
     }
 
+    // A workflow has to be active to be callable. An inactive one is a draft -
+    // half-edited, or deliberately taken out of service - and running it because
+    // something else still points at it is how a change nobody meant to publish
+    // reaches production.
+    if (stored.value().value("active", false) != true) {
+        node_result.status = NodeStatus::Failed;
+        node_result.error = "Call Workflow: \"" +
+                            stored.value().value("name", workflow_id) +
+                            "\" is not active. Activate it to let other workflows call it";
+        return false;
+    }
+
     auto sub_workflow = Workflow::fromJson(stored.value());
 
     nlohmann::json sub_trigger = call.value("input", nlohmann::json::object());
@@ -1700,12 +1743,36 @@ bool WorkflowEngine::runSubWorkflow(NodeExecutionResult& node_result, int call_d
     sub_trigger["_callDepth"] = call_depth + 1;
     sub_trigger["_calledBy"] = parent_execution_id;
 
+    // Nothing new starts once the caller has been told to stop. Without this,
+    // a loop calling a workflow per item keeps launching them: the caller only
+    // looks at its own cancellation between items, and each call can take
+    // minutes.
+    {
+        std::lock_guard<std::mutex> lock(mutex_);
+        if (cancelled_executions_.contains(parent_execution_id)) {
+            node_result.status = NodeStatus::Failed;
+            node_result.error = "Call Workflow: the run was cancelled before this call started";
+            return false;
+        }
+    }
+
     LOG_INFO("Execution {} calls workflow {} at depth {}", parent_execution_id, workflow_id,
              call_depth + 1);
 
     // No callback: the sub-workflow reports its own progress against its own
     // execution, and forwarding its node events to the parent's subscribers
     // would make the parent's canvas light up nodes it does not have.
+    // The child's id is generated inside execute(), so the caller cannot know it
+    // in advance. It is handed in on the trigger data instead, and recorded
+    // against the caller as soon as it is known - which is what lets a cancel
+    // arriving mid-call reach the workflow actually doing the work.
+    const std::string child_execution_id = UUID::generatePrefixed("exec");
+    sub_trigger["_childExecutionId"] = child_execution_id;
+    {
+        std::lock_guard<std::mutex> lock(mutex_);
+        child_executions_[parent_execution_id].insert(child_execution_id);
+    }
+
     auto outcome = execute(sub_workflow, "workflow", sub_trigger, nullptr);
     if (outcome.failed()) {
         node_result.status = NodeStatus::Failed;

+ 6 - 0
src/runner/workflow_engine.hpp

@@ -325,6 +325,12 @@ private:
 
     std::unordered_map<std::string, ExecutionResult> active_executions_;
     std::unordered_set<std::string> cancelled_executions_;
+
+    // Which executions were started by which. A workflow called as a step runs
+    // as its own execution with its own id, so cancelling the caller says
+    // nothing about it - and the caller is sitting inside that call, unable to
+    // notice its own cancellation until the call returns.
+    std::unordered_map<std::string, std::unordered_set<std::string>> child_executions_;
     mutable std::mutex mutex_;
     std::atomic<int> active_count_{0};
 };

+ 7 - 0
src/webserver/api/workflow_controller.cpp

@@ -353,6 +353,13 @@ void WorkflowController::executeWorkflow(const httplib::Request& req, httplib::R
         return;
     }
 
+    // The scheduler counts what is running so an overlap policy of skip or
+    // queue can hold a tick back. It only ever heard about runs it started
+    // itself, so a run started from the editor was invisible: a workflow set to
+    // never overlap would happily get a second scheduled run on top of the one
+    // someone was watching.
+    scheduler_.notifyExecutionStarted(workflow_id, grpc_res.execution_id());
+
     nlohmann::json response;
     response["executionId"] = grpc_res.execution_id();
     response["status"] = grpc_res.status();

+ 93 - 0
tests/nodes/call-workflow-inactive.json

@@ -0,0 +1,93 @@
+{
+  "name": "verify-call-workflow-inactive",
+  "helpers": [
+    {
+      "key": "sub",
+      "name": "verify sub: not published",
+      "active": false,
+      "nodes": [
+        {
+          "id": "in",
+          "name": "Input",
+          "type": "workflow-input",
+          "position": {
+            "x": 0,
+            "y": 0
+          },
+          "config": {
+            "fields": [
+              {
+                "name": "n",
+                "type": "number",
+                "required": true
+              }
+            ]
+          }
+        },
+        {
+          "id": "out",
+          "name": "Output",
+          "type": "workflow-output",
+          "position": {
+            "x": 0,
+            "y": 100
+          },
+          "config": {}
+        }
+      ],
+      "connections": [
+        {
+          "sourceNodeId": "in",
+          "sourceOutput": "main",
+          "targetNodeId": "out",
+          "targetInput": "data"
+        }
+      ]
+    }
+  ],
+  "nodes": [
+    {
+      "id": "n1",
+      "name": "Trigger",
+      "type": "click-trigger",
+      "position": {
+        "x": 0,
+        "y": 0
+      },
+      "config": {}
+    },
+    {
+      "id": "call",
+      "name": "Call",
+      "type": "call-workflow",
+      "position": {
+        "x": 0,
+        "y": 100
+      },
+      "config": {
+        "workflowId": "{{helper:sub}}",
+        "inputSource": "fields",
+        "fields": [
+          {
+            "name": "n",
+            "value": "1"
+          }
+        ]
+      }
+    }
+  ],
+  "connections": [
+    {
+      "sourceNodeId": "n1",
+      "sourceOutput": "main",
+      "targetNodeId": "call",
+      "targetInput": "data"
+    }
+  ],
+  "expect": {
+    "call": {
+      "status": "failed",
+      "errorContains": "is not active"
+    }
+  }
+}

+ 105 - 6
tests/nodes/call-workflow-missing-input.json

@@ -1,12 +1,111 @@
 {
   "name": "verify-call-workflow-missing-input",
   "nodes": [
-    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
-    {"id": "call", "name": "Call", "type": "call-workflow", "position": {"x": 0, "y": 100},
-     "config": {"workflowId": "wf_5cfe91eb-a7f0-4da8-aca9-2300ed40b399", "inputSource": "fields", "fields": [{"name": "label", "value": "no number given"}]}}
+    {
+      "id": "n1",
+      "name": "Trigger",
+      "type": "click-trigger",
+      "position": {
+        "x": 0,
+        "y": 0
+      },
+      "config": {}
+    },
+    {
+      "id": "call",
+      "name": "Call",
+      "type": "call-workflow",
+      "position": {
+        "x": 0,
+        "y": 100
+      },
+      "config": {
+        "workflowId": "{{helper:sub}}",
+        "inputSource": "fields",
+        "fields": [
+          {
+            "name": "label",
+            "value": "no number given"
+          }
+        ]
+      }
+    }
   ],
   "connections": [
-    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "call", "targetInput": "data"}
+    {
+      "sourceNodeId": "n1",
+      "sourceOutput": "main",
+      "targetNodeId": "call",
+      "targetInput": "data"
+    }
   ],
-  "expect": {"call": {"status": "failed", "errorContains": "did not pass n"}}
-}
+  "expect": {
+    "call": {
+      "status": "failed",
+      "errorContains": "did not pass n"
+    }
+  },
+  "helpers": [
+    {
+      "key": "sub",
+      "name": "verify sub: double a number",
+      "active": true,
+      "nodes": [
+        {
+          "id": "in",
+          "name": "Input",
+          "type": "workflow-input",
+          "position": {
+            "x": 0,
+            "y": 0
+          },
+          "config": {
+            "fields": [
+              {
+                "name": "n",
+                "type": "number",
+                "required": true
+              }
+            ]
+          }
+        },
+        {
+          "id": "calc",
+          "name": "Double",
+          "type": "code",
+          "position": {
+            "x": 0,
+            "y": 100
+          },
+          "config": {
+            "code": "const d=input.data||input; return { doubled: Number(d.n)*2, label: d.label || \"none\" };"
+          }
+        },
+        {
+          "id": "out",
+          "name": "Output",
+          "type": "workflow-output",
+          "position": {
+            "x": 0,
+            "y": 200
+          },
+          "config": {}
+        }
+      ],
+      "connections": [
+        {
+          "sourceNodeId": "in",
+          "sourceOutput": "main",
+          "targetNodeId": "calc",
+          "targetInput": "data"
+        },
+        {
+          "sourceNodeId": "calc",
+          "sourceOutput": "main",
+          "targetNodeId": "out",
+          "targetInput": "data"
+        }
+      ]
+    }
+  ]
+}

+ 142 - 11
tests/nodes/call-workflow.json

@@ -1,19 +1,150 @@
 {
   "name": "verify-call-workflow",
   "nodes": [
-    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
-    {"id": "call", "name": "Call", "type": "call-workflow", "position": {"x": 0, "y": 100},
-     "config": {"workflowId": "wf_5cfe91eb-a7f0-4da8-aca9-2300ed40b399", "inputSource": "fields",
-                "fields": [{"name": "n", "value": "21"}]}},
-    {"id": "after", "name": "After", "type": "code", "position": {"x": 0, "y": 200},
-     "config": {"code": "const d=input.data||input; return { got: d.result, status: d.status };"}}
+    {
+      "id": "n1",
+      "name": "Trigger",
+      "type": "click-trigger",
+      "position": {
+        "x": 0,
+        "y": 0
+      },
+      "config": {}
+    },
+    {
+      "id": "call",
+      "name": "Call",
+      "type": "call-workflow",
+      "position": {
+        "x": 0,
+        "y": 100
+      },
+      "config": {
+        "workflowId": "{{helper:sub}}",
+        "inputSource": "fields",
+        "fields": [
+          {
+            "name": "n",
+            "value": "21"
+          }
+        ]
+      }
+    },
+    {
+      "id": "after",
+      "name": "After",
+      "type": "code",
+      "position": {
+        "x": 0,
+        "y": 200
+      },
+      "config": {
+        "code": "const d=input.data||input; return { got: d.result, status: d.status };"
+      }
+    }
   ],
   "connections": [
-    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "call", "targetInput": "data"},
-    {"sourceNodeId": "call", "sourceOutput": "main", "targetNodeId": "after", "targetInput": "data"}
+    {
+      "sourceNodeId": "n1",
+      "sourceOutput": "main",
+      "targetNodeId": "call",
+      "targetInput": "data"
+    },
+    {
+      "sourceNodeId": "call",
+      "sourceOutput": "main",
+      "targetNodeId": "after",
+      "targetInput": "data"
+    }
   ],
   "expectStatus": "completed",
   "expect": {
-    "after": {"status": "completed", "output": {"result": {"status": "completed", "got": {"doubled": 42, "label": "unnamed"}}}}
-  }
-}
+    "after": {
+      "status": "completed",
+      "output": {
+        "result": {
+          "status": "completed",
+          "got": {
+            "doubled": 42,
+            "label": "unnamed"
+          }
+        }
+      }
+    }
+  },
+  "helpers": [
+    {
+      "key": "sub",
+      "name": "verify sub: double a number",
+      "active": true,
+      "nodes": [
+        {
+          "id": "in",
+          "name": "Input",
+          "type": "workflow-input",
+          "position": {
+            "x": 0,
+            "y": 0
+          },
+          "config": {
+            "fields": [
+              {
+                "name": "n",
+                "type": "number",
+                "required": true
+              }
+            ]
+          }
+        },
+        {
+          "id": "calc",
+          "name": "Double",
+          "type": "code",
+          "position": {
+            "x": 0,
+            "y": 100
+          },
+          "config": {
+            "code": "const d=input.data||input; return { doubled: Number(d.n)*2, label: d.label || \"unnamed\" };"
+          }
+        },
+        {
+          "id": "out",
+          "name": "Output",
+          "type": "workflow-output",
+          "position": {
+            "x": 0,
+            "y": 200
+          },
+          "config": {
+            "source": "fields",
+            "fields": [
+              {
+                "name": "doubled",
+                "value": "{{data.result.doubled}}"
+              },
+              {
+                "name": "label",
+                "value": "{{data.result.label}}"
+              }
+            ]
+          }
+        }
+      ],
+      "connections": [
+        {
+          "sourceNodeId": "in",
+          "sourceOutput": "main",
+          "targetNodeId": "calc",
+          "targetInput": "data"
+        },
+        {
+          "sourceNodeId": "calc",
+          "sourceOutput": "main",
+          "targetNodeId": "out",
+          "targetInput": "data"
+        }
+      ]
+    }
+  ]
+}