Эх сурвалжийг харах

Merge branch 'tier-2-triggers': the Tier 2 trigger nodes

Adds the four Tier 2 entries from docs/node-roadmap.md - respond-to-webhook,
database-change, file-watch and wait-for-approval - plus the platform work
they needed:

- fs.readdir, so a node can see what is in a directory
- context.workflowId, which the polling triggers key their cursors on
- a _webhookResponse marker any node may return, so a webhook-triggered
  workflow controls its status, headers and body, and appending a node
  cannot silently change what its API returns
- resumable executions: a _pause marker, a waiting status, WorkflowEngine
  ::resume, a ResumeExecution RPC, and POST /executions/{id}/resume plus
  GET /executions/pending

Two of the roadmap's premises turned out to be wrong and are corrected in
it: the webhook controller never fired and forgot, it already waited and
returned a body, and File Watch could not be pure JavaScript because the
filesystem API had no way to list a directory.

Four pre-existing bugs were found and fixed along the way. fs.stat().mtime
returned a negative number, because file_time_type and system_clock use
different epochs. Every execution status was shifted one over gRPC, so a
failed execution reported as completed. ExecutionController had no way to
reach a runner. And a resumed execution could read every credential in the
system, because the workflow snapshot carried no id and an empty workflow
id means admin access.
fszontagh 1 сар өмнө
parent
commit
15be954371

+ 16 - 7
docs/node-roadmap.md

@@ -10,7 +10,7 @@ The runner already exposes `http`, `storage` (including the file store),
 
 ## What exists today
 
-51 node definitions across: triggers (click, get/post/put, imap, schedule with
+56 node definitions across: triggers (click, get/post/put, imap, schedule with
 cron and overlap control, error), flow control (if-condition, loop, wait), data
 (rss-reader, the six storage nodes), email (four imap nodes plus smtp-send), AI
 (ollama-chat, comfyui-prompt), integration (telegram-send, nextcloud-talk),
@@ -29,12 +29,21 @@ now parsed, which is how `merge` gets two input handles.
 
 ## Tier 2 - triggers
 
-| Node | Notes |
-| --- | --- |
-| **Respond to Webhook** | **Needs C++.** `webhook_controller` fires and forgets, so there is no way to return a computed body. Without it the GET/POST/PUT triggers cannot back a real API |
-| **Database Change** | Poll a collection for new or changed documents. Pure JS over `storage.query` plus a stored cursor |
-| **File Watch** | Poll a directory through `fs.stat`. Pure JS when paired with a schedule |
-| **Queue / Manual Approval** | **Large.** Needs resumable executions; the engine currently runs a workflow to completion in one pass |
+Built. `respond-to-webhook`, `database-change`, `file-watch` and
+`wait-for-approval` all live in `nodes/`, with fixtures under `tests/nodes/`.
+
+Three platform changes came with them: `fs.readdir`, so a node can see what is
+in a directory; a `_webhookResponse` marker any node can return to set the
+status, headers and body a webhook replies with; and resumable executions - an
+execution can pause on a `_pause` marker and be continued later through
+`POST /api/v1/executions/{id}/resume`, rebuilt from the workflow snapshot stored
+with it.
+
+Two things the original entries claimed turned out not to hold. The webhook
+controller never fired and forgot: it already waited for completion and returned
+a body, and what was missing was control over the status code, the headers, and
+which node decides the response. And File Watch could not be pure JavaScript,
+because the filesystem API had no way to list a directory.
 
 Cron and interval scheduling are already covered by `schedule-trigger`,
 including its overlap policy.

+ 93 - 0
docs/nodes.md

@@ -199,6 +199,67 @@ does - the branch name is just computed:
 return { _activeBranch: 'case' + i, ['case' + i]: data };
 ```
 
+## Pausing for an Answer
+
+A node pauses its execution by returning a `_pause` marker, the way a branching
+node returns `_activeBranch`:
+
+```javascript
+return {
+    token: token,
+    reason: 'Approve the refund',
+    expiresAt: Date.now() + 86400000,
+    _pause: { token: token, reason: 'Approve the refund', expiresAt: expiresAt }
+};
+```
+
+The engine stops the walk, stores the execution as `waiting` with everything
+computed so far, and returns. Nothing downstream runs. The marker is stripped
+from the stored output, so a reader sees the request rather than the mechanism.
+
+Answering it continues the run:
+
+```
+POST /api/v1/executions/{id}/resume
+{ "token": "...", "approved": true, "data": { "note": "looks fine" } }
+```
+
+The paused node's output becomes that payload, and the walk continues from
+there. Nodes that already ran are not run again. The workflow is rebuilt from
+the snapshot stored with the execution, not from the workflow as it stands now,
+because it may have been edited while the approval waited.
+
+`GET /api/v1/executions/pending` lists executions waiting for an answer. It
+deliberately omits the token: listing is a weaker permission than approving.
+
+Two limits are deliberate. A pause inside a Loop body cannot be resumed, because
+loop iteration state is not part of the stored execution, so a node must refuse
+to pause there rather than record something unanswerable. And a webhook cannot
+wait for an approval - the HTTP request is still open and its deadline is
+35 seconds - so a webhook-triggered workflow that pauses returns a 202 with
+`{ "executionId": "...", "status": "waiting" }` rather than blocking for an
+answer that has nowhere to arrive on this connection. Answer it the same way
+as any other paused execution, with `POST /api/v1/executions/{id}/resume`.
+
+## Answering a Webhook
+
+A webhook-triggered workflow returns the last node's output as JSON by default.
+To control the response, return a `_webhookResponse` marker from any node:
+
+```javascript
+return {
+    _webhookResponse: {
+        status: 201,
+        headers: { 'Content-Type': 'application/json' },
+        body: { id: created.id }
+    }
+};
+```
+
+Any node may set it, not only the last one to run, so adding a node to the end
+of a workflow cannot silently change what its API returns. If several set it,
+the last one wins and the runner logs that it happened.
+
 ## Multiple Inputs
 
 A node that joins two branches names its inputs:
@@ -354,6 +415,38 @@ smartbotic.storage.update('collection', 'document-id', { key: 'new-value' });
 smartbotic.storage.delete('collection', 'document-id');
 ```
 
+### Filesystem
+
+```javascript
+// Read a file
+const read = smartbotic.fs.readFile('/tmp/example.txt', 'base64');
+// { success, data } - data is base64-encoded when the encoding argument is 'base64'
+
+// Write a file
+smartbotic.fs.writeFile('/tmp/example.txt', base64Data, 'base64');
+
+// Check existence
+const there = smartbotic.fs.exists('/tmp/example.txt');
+
+// Stat a file
+const info = smartbotic.fs.stat('/tmp/example.txt');
+// { success, size, mtime, isDirectory } - mtime is milliseconds since the epoch
+
+// Create a directory
+smartbotic.fs.mkdir('/tmp/example-dir');
+
+// Delete a file
+smartbotic.fs.unlink('/tmp/example.txt');
+
+// List a directory, one level
+const listing = smartbotic.fs.readdir('/var/spool/incoming');
+// { success: true, entries: [{ name, path, size, modifiedAt, isDirectory }] }
+```
+
+The `fs` API applies no path restrictions. Every call takes an arbitrary
+absolute path and acts on it with the runner process's own permissions, the
+same trust level a `code` node or `process.exec` already has.
+
 ### Credentials
 
 ```javascript

+ 76 - 24
docs/superpowers/plans/2026-08-05-tier-2-triggers.md

@@ -54,6 +54,11 @@ Three later tasks need assertions the harness cannot express: a webhook's status
 **Interfaces:**
 - Produces: two new optional top-level keys in a case file.
   - `"http"`: `{method, path, body, expectStatus, expectHeaders, expectBodyContains}` - instead of executing the workflow through `/execute`, the harness calls `path` on the webserver directly and asserts on the response. Used for webhook cases.
+  - `"settings"`: passed through as the created workflow's `settings` object
+    instead of the hardcoded `{}`. Without it a fixture cannot grant
+    `storagePermissions`, and every `smartbotic.storage.*` call a node makes
+    fails with "No write access to collection". No existing fixture uses
+    storage, which is why this went unnoticed.
   - `"resume"`: `{waitForStatus, tokenFrom, payload}` - after starting the execution, poll until the execution reaches `waitForStatus`, read the token from the node output named by `tokenFrom`, POST it to the resume endpoint with `payload`, then continue polling to completion before running the normal `expect` assertions.
 
 - [ ] **Step 1: Add the HTTP case mode**
@@ -132,7 +137,13 @@ def run_resume_case(case, token, execution_id):
     print("resumed %s with token %s" % (execution_id, resume_token))
 ```
 
-- [ ] **Step 3: Wire both into main**
+- [ ] **Step 3: Pass a fixture's settings through**
+
+In `main`, the workflow is created with `"settings": {}` hardcoded. Replace that
+with `case.get("settings", {})`, so a fixture can declare what its nodes need.
+Every existing fixture omits the key and gets `{}` exactly as before.
+
+- [ ] **Step 4: Wire both into main**
 
 In `main`, immediately after the workflow is created inside the `try` block, add the HTTP branch before the existing execute call:
 
@@ -154,7 +165,7 @@ Then, after `execution_id` is obtained and before the existing polling loop, add
 
 The existing polling loop then runs unchanged and sees the resumed execution through to completion.
 
-- [ ] **Step 4: Confirm nothing regressed**
+- [ ] **Step 5: Confirm nothing regressed**
 
 Run:
 ```bash
@@ -162,7 +173,7 @@ for f in tests/nodes/*.json; do python3 scripts/verify-node.py "$f" >/dev/null |
 ```
 Expected: no `FAILED:` lines. Every existing case lacks both new keys, so all take the unchanged path.
 
-- [ ] **Step 5: Commit**
+- [ ] **Step 6: Commit**
 
 ```bash
 git add scripts/verify-node.py
@@ -171,14 +182,16 @@ git commit -m "test: let the harness assert on HTTP responses and answer a pause
 
 ---
 
-### Task 2: `fs.readdir`
+### Task 2: `fs.readdir`, and a workflow id a node can see
 
 **Files:**
-- Modify: `src/runner/engine/script_engine.cpp` (after the `fs.stat` registration, before `JS_SetPropertyStr(ctx, smartbotic, "fs", fs);`)
+- Modify: `src/runner/engine/script_engine.cpp` (after the `fs.stat` registration, before `JS_SetPropertyStr(ctx, smartbotic, "fs", fs);`, and the context object around line 4393)
+- Modify: `src/runner/workflow_engine.cpp:905` (the `ctx` filled before a node runs)
 - Create: `tests/nodes/fs-readdir.json`
 
 **Interfaces:**
 - Produces: `smartbotic.fs.readdir(path)` returning `{success: true, entries: [{name, path, size, modifiedAt, isDirectory}]}` or `{success: false, error}`. `modifiedAt` is milliseconds since the epoch, matching `fs.stat`'s `mtime`. Not recursive.
+- Produces: `context.workflowId`, alongside the `context.executionId` and `context.nodeId` a node already receives. Tasks 3 and 4 key their cursors on it.
 
 - [ ] **Step 1: Register the function**
 
@@ -263,11 +276,40 @@ Insert after the closing `}, "stat", 1));`:
     }, "readdir", 1));
 ```
 
-- [ ] **Step 2: Build and restart**
+- [ ] **Step 2: Expose the workflow id to nodes**
+
+A node currently receives only `executionId` and `nodeId` on its `context`. The
+two polling triggers in Tasks 3 and 4 key their cursor on the workflow, and
+without it every workflow's cursor would collide under one key - two workflows
+watching different directories would consume each other's changes.
+
+`ScriptContext` already declares `workflow_id` (`script_engine.hpp:147`) and
+`executeNode` already receives the `Workflow`, so nothing new has to be plumbed
+through. It is simply never set and never exposed.
+
+In `workflow_engine.cpp`, in `executeNode` beside the other `ctx` assignments
+around line 903:
+
+```cpp
+    ctx.workflow_id = workflow.id;
+```
+
+In `script_engine.cpp`, in the context object built for the execute function
+around line 4393:
+
+```cpp
+    nlohmann::json ctx_json = {
+        {"executionId", context.execution_id},
+        {"nodeId", context.node_id},
+        {"workflowId", context.workflow_id}
+    };
+```
+
+- [ ] **Step 3: Build and restart**
 
 Run the rebuild and restart sequence from the Reference section. Expected: clean build, all three ports listening.
 
-- [ ] **Step 3: Write the fixture**
+- [ ] **Step 4: Write the fixture**
 
 `tests/nodes/fs-readdir.json` - a `code` node creates a directory with two files, lists it, then cleans up:
 
@@ -277,7 +319,7 @@ Run the rebuild and restart sequence from the Reference section. Expected: clean
   "nodes": [
     {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
     {"id": "n2", "name": "List", "type": "code", "position": {"x": 0, "y": 100},
-     "config": {"code": "const dir = '/tmp/sb-readdir-test';\nsmartbotic.fs.mkdir(dir);\nsmartbotic.fs.writeFile(dir + '/one.txt', smartbotic.utils.base64Encode('hello'));\nsmartbotic.fs.writeFile(dir + '/two.txt', smartbotic.utils.base64Encode('worldwide'));\nsmartbotic.fs.mkdir(dir + '/sub');\n\nconst listing = smartbotic.fs.readdir(dir);\nconst missing = smartbotic.fs.readdir('/tmp/sb-readdir-does-not-exist');\nconst notDir = smartbotic.fs.readdir(dir + '/one.txt');\n\nconst names = listing.entries.map(function (e) { return e.name; }).sort();\nconst one = listing.entries.filter(function (e) { return e.name === 'one.txt'; })[0];\nconst sub = listing.entries.filter(function (e) { return e.name === 'sub'; })[0];\n\nsmartbotic.fs.unlink(dir + '/one.txt');\nsmartbotic.fs.unlink(dir + '/two.txt');\n\nreturn {\n    ok: listing.success,\n    names: names,\n    oneSize: one.size,\n    oneHasMtime: one.modifiedAt > 1600000000000,\n    subIsDirectory: sub.isDirectory,\n    oneIsDirectory: one.isDirectory,\n    missingRejected: missing.success === false,\n    notDirRejected: notDir.success === false\n};"}}
+     "config": {"code": "const dir = '/tmp/sb-readdir-test';\nsmartbotic.fs.mkdir(dir);\nsmartbotic.fs.writeFile(dir + '/one.txt', smartbotic.utils.base64Encode('hello'));\nsmartbotic.fs.writeFile(dir + '/two.txt', smartbotic.utils.base64Encode('worldwide'));\nsmartbotic.fs.mkdir(dir + '/sub');\n\nconst listing = smartbotic.fs.readdir(dir);\nconst missing = smartbotic.fs.readdir('/tmp/sb-readdir-does-not-exist');\nconst notDir = smartbotic.fs.readdir(dir + '/one.txt');\n\nconst names = listing.entries.map(function (e) { return e.name; }).sort();\nconst one = listing.entries.filter(function (e) { return e.name === 'one.txt'; })[0];\nconst sub = listing.entries.filter(function (e) { return e.name === 'sub'; })[0];\n\nsmartbotic.fs.unlink(dir + '/one.txt');\nsmartbotic.fs.unlink(dir + '/two.txt');\n\nreturn {\n    ok: listing.success,\n    names: names,\n    oneSize: one.size,\n    oneHasMtime: one.modifiedAt > 1600000000000,\n    subIsDirectory: sub.isDirectory,\n    oneIsDirectory: one.isDirectory,\n    missingRejected: missing.success === false,\n    notDirRejected: notDir.success === false,\n    hasWorkflowId: typeof context.workflowId === 'string' && context.workflowId.length > 0\n};"}}
   ],
   "connections": [
     {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "n2", "targetInput": "data"}
@@ -299,16 +341,16 @@ Run the rebuild and restart sequence from the Reference section. Expected: clean
 
 `oneSize` is 5 because the file holds `hello`. That pins the size actually being read rather than defaulted to zero.
 
-- [ ] **Step 4: Run it**
+- [ ] **Step 5: Run it**
 
 Run: `python3 scripts/verify-node.py tests/nodes/fs-readdir.json`
 Expected: `PASS`.
 
-- [ ] **Step 5: Commit**
+- [ ] **Step 6: Commit**
 
 ```bash
-git add src/runner/engine/script_engine.cpp tests/nodes/fs-readdir.json
-git commit -m "feat: fs.readdir, so a node can see what is in a directory"
+git add src/runner/engine/script_engine.cpp src/runner/workflow_engine.cpp tests/nodes/fs-readdir.json
+git commit -m "feat: fs.readdir, and a workflow id a node can see"
 ```
 
 ---
@@ -321,7 +363,7 @@ git commit -m "feat: fs.readdir, so a node can see what is in a directory"
 
 **Interfaces:**
 - Consumes: `smartbotic.fs.readdir(path)` from Task 2.
-- Produces: output `{files: [...], count, isFirstRun}`. Cursor documents live in the `_watch_cursors` collection, keyed `<workflowId>:<nodeId>`.
+- Produces: output `{files: [...], count, isFirstRun}`. Cursor documents live in the `watch_cursors` collection, keyed `<workflowId>:<nodeId>`.
 
 - [ ] **Step 1: Write the node**
 
@@ -364,7 +406,7 @@ const configSchema = {
             type: 'string',
             title: 'Cursor Collection',
             description: 'Collection holding the last-seen state',
-            default: '_watch_cursors'
+            default: 'watch_cursors'
         }
     },
     required: ['directory']
@@ -392,7 +434,7 @@ async function execute(config, input, context) {
         throw new Error('File Watch: a directory is required');
     }
 
-    const collection = config.cursorCollection || '_watch_cursors';
+    const collection = config.cursorCollection || 'watch_cursors';
     const workflowId = (context && context.workflowId) || 'unknown';
     const nodeId = (context && context.nodeId) || 'file-watch';
     const cursorId = workflowId + ':' + nodeId;
@@ -479,10 +521,15 @@ If either is absent, the cursor id would collapse to `unknown:file-watch` and tw
 ```json
 {
   "name": "verify-file-watch",
+  "settings": {
+    "storagePermissions": {
+      "collections": { "watch_cursors": "read-write" }
+    }
+  },
   "nodes": [
     {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
     {"id": "setup", "name": "Setup", "type": "code", "position": {"x": 0, "y": 100},
-     "config": {"code": "const dir = '/tmp/sb-file-watch-test';\nconst listing = smartbotic.fs.readdir(dir);\nif (listing.success) {\n    for (const e of listing.entries) { smartbotic.fs.unlink(e.path); }\n}\nsmartbotic.fs.mkdir(dir);\nsmartbotic.fs.writeFile(dir + '/existing.txt', smartbotic.utils.base64Encode('old'));\nsmartbotic.storage.delete('_watch_cursors', 'verify-file-watch:first');\nsmartbotic.storage.delete('_watch_cursors', 'verify-file-watch:second');\nreturn { dir: dir };"}},
+     "config": {"code": "const dir = '/tmp/sb-file-watch-test';\nconst listing = smartbotic.fs.readdir(dir);\nif (listing.success) {\n    for (const e of listing.entries) { smartbotic.fs.unlink(e.path); }\n}\nsmartbotic.fs.mkdir(dir);\nsmartbotic.fs.writeFile(dir + '/existing.txt', smartbotic.utils.base64Encode('old'));\nsmartbotic.storage.delete('watch_cursors', 'verify-file-watch:first');\nsmartbotic.storage.delete('watch_cursors', 'verify-file-watch:second');\nreturn { dir: dir };"}},
     {"id": "first", "name": "First Run", "type": "file-watch", "position": {"x": 0, "y": 200},
      "config": {"directory": "/tmp/sb-file-watch-test", "emitOnFirstRun": false}},
     {"id": "add", "name": "Add A File", "type": "code", "position": {"x": 0, "y": 300},
@@ -528,7 +575,7 @@ git commit -m "feat: a File Watch trigger, reporting what changed in a directory
 - Create: `tests/nodes/database-change.json`
 
 **Interfaces:**
-- Produces: output `{documents: [...], count, isFirstRun}`. Cursor documents live in the same `_watch_cursors` collection, keyed `<workflowId>:<nodeId>`.
+- Produces: output `{documents: [...], count, isFirstRun}`. Cursor documents live in the same `watch_cursors` collection, keyed `<workflowId>:<nodeId>`.
 
 - [ ] **Step 1: Write the node**
 
@@ -577,7 +624,7 @@ const configSchema = {
         cursorCollection: {
             type: 'string',
             title: 'Cursor Collection',
-            default: '_watch_cursors'
+            default: 'watch_cursors'
         }
     },
     required: ['collection']
@@ -607,7 +654,7 @@ async function execute(config, input, context) {
     }
 
     const timestampField = config.timestampField || '_updatedAt';
-    const cursorCollection = config.cursorCollection || '_watch_cursors';
+    const cursorCollection = config.cursorCollection || 'watch_cursors';
     const maxDocuments = Number(config.maxDocuments) || 100;
     const workflowId = (context && context.workflowId) || 'unknown';
     const nodeId = (context && context.nodeId) || 'database-change';
@@ -703,18 +750,23 @@ Record the magnitude in your report. If those values are around 1.7e18 the norma
 ```json
 {
   "name": "verify-database-change",
+  "settings": {
+    "storagePermissions": {
+      "collections": { "watch_cursors": "read-write", "dbchange_test": "read-write" }
+    }
+  },
   "nodes": [
     {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
     {"id": "setup", "name": "Setup", "type": "code", "position": {"x": 0, "y": 100},
-     "config": {"code": "smartbotic.storage.delete('_watch_cursors', 'verify-database-change:first');\nsmartbotic.storage.delete('_watch_cursors', 'verify-database-change:second');\nsmartbotic.storage.insert('_dbchange_test', { name: 'before', at: Date.now() }, 'before-1');\nreturn { ready: true };"}},
+     "config": {"code": "smartbotic.storage.delete('watch_cursors', 'verify-database-change:first');\nsmartbotic.storage.delete('watch_cursors', 'verify-database-change:second');\nsmartbotic.storage.insert('dbchange_test', { name: 'before', at: Date.now() }, 'before-1');\nreturn { ready: true };"}},
     {"id": "first", "name": "First Run", "type": "database-change", "position": {"x": 0, "y": 200},
-     "config": {"collection": "_dbchange_test", "emitOnFirstRun": false}},
+     "config": {"collection": "dbchange_test", "emitOnFirstRun": false}},
     {"id": "add", "name": "Insert One", "type": "code", "position": {"x": 0, "y": 300},
-     "config": {"code": "smartbotic.utils.sleep(1100);\nsmartbotic.storage.insert('_dbchange_test', { name: 'after', at: Date.now() }, 'after-1');\nreturn { inserted: 'after-1' };"}},
+     "config": {"code": "smartbotic.utils.sleep(1100);\nsmartbotic.storage.insert('dbchange_test', { name: 'after', at: Date.now() }, 'after-1');\nreturn { inserted: 'after-1' };"}},
     {"id": "second", "name": "Second Run", "type": "database-change", "position": {"x": 0, "y": 400},
-     "config": {"collection": "_dbchange_test", "emitOnFirstRun": false}},
+     "config": {"collection": "dbchange_test", "emitOnFirstRun": false}},
     {"id": "cleanup", "name": "Cleanup", "type": "code", "position": {"x": 0, "y": 500},
-     "config": {"code": "smartbotic.storage.delete('_dbchange_test', 'before-1');\nsmartbotic.storage.delete('_dbchange_test', 'after-1');\nreturn { cleaned: true };"}}
+     "config": {"code": "smartbotic.storage.delete('dbchange_test', 'before-1');\nsmartbotic.storage.delete('dbchange_test', 'after-1');\nreturn { cleaned: true };"}}
   ],
   "connections": [
     {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "setup", "targetInput": "data"},

+ 12 - 0
lib/storage/storage_client.cpp

@@ -106,6 +106,18 @@ Result<std::string> StorageClient::insert(const std::string& collection,
     return new_id;
 }
 
+Result<std::string> StorageClient::upsert(const std::string& collection,
+                                          const nlohmann::json& data,
+                                          const std::string& id,
+                                          int64_t ttl_ms) {
+    auto [new_id, is_new] = impl_->client_->upsert(collection, data, id, msToSec(ttl_ms));
+    (void)is_new;
+    if (new_id.empty()) {
+        return Error(ErrorCode::DatabaseError, "Upsert failed: " + collection);
+    }
+    return new_id;
+}
+
 Result<int64_t> StorageClient::update(const std::string& collection,
                                       const std::string& id,
                                       const nlohmann::json& data,

+ 9 - 0
lib/storage/storage_client.hpp

@@ -125,6 +125,15 @@ public:
                                        const std::string& id = "",
                                        int64_t ttl_ms = 0);
 
+    // Atomic insert-or-replace: swaps the expiration-index entry and installs
+    // the new TTL in one locked server-side operation, unlike a remove()
+    // followed by an insert() which leaves a window where the record can be
+    // lost or left stale between the two calls.
+    common::Result<std::string> upsert(const std::string& collection,
+                                       const nlohmann::json& data,
+                                       const std::string& id = "",
+                                       int64_t ttl_ms = 0);
+
     // Returns: the new version on optimistic-lock success (expected_version+1),
     //          the new version on partial/patch success,
     //          -1 when the update succeeded but the new version is unknown

+ 130 - 0
nodes/core/respond-to-webhook.js

@@ -0,0 +1,130 @@
+/**
+ * @node respond-to-webhook
+ * @name Respond to Webhook
+ * @category core
+ * @version 1.0.0
+ * @description Set the status, headers and body a webhook-triggered workflow returns to its caller
+ * @icon reply
+ */
+
+const configSchema = {
+    type: 'object',
+    properties: {
+        status: {
+            type: 'number',
+            title: 'Status Code',
+            description: 'HTTP status to return, such as 200, 201, 400 or 404',
+            default: 200
+        },
+        bodySource: {
+            type: 'string',
+            title: 'Body From',
+            description: 'json sends the text below parsed as JSON, text sends it as it is, and input sends a value taken from the incoming data',
+            enum: ['json', 'text', 'input'],
+            default: 'json'
+        },
+        body: {
+            type: 'string',
+            title: 'Body',
+            description: 'The response body, for the json and text sources. Supports {{variable}} interpolation',
+            format: 'textarea'
+        },
+        bodyField: {
+            type: 'string',
+            title: 'Body Field',
+            description: 'Path to the value to send, for the input source, such as data.result',
+            default: 'data'
+        },
+        contentType: {
+            type: 'string',
+            title: 'Content Type',
+            description: 'Content-Type header. Leave empty to send application/json',
+            default: 'application/json'
+        },
+        headers: {
+            type: 'object',
+            title: 'Extra Headers',
+            description: 'Additional response headers. A content-type entry here is ignored; use the Content Type field above instead',
+            additionalProperties: { type: 'string' }
+        }
+    }
+};
+
+const inputSchema = {
+    type: 'object',
+    properties: {
+        data: { type: 'any' }
+    }
+};
+
+const outputSchema = {
+    type: 'object',
+    properties: {
+        status: { type: 'number', description: 'Status this node asked for' },
+        respondedWith: { type: 'string', description: 'Content type sent' }
+    }
+};
+
+async function execute(config, input, context) {
+    const status = config.status === undefined || config.status === null || config.status === ''
+        ? 200
+        : Number(config.status);
+    if (isNaN(status) || status < 100 || status > 599) {
+        throw new Error('Respond to Webhook: "' + config.status +
+            '" is not an HTTP status code. Use a number from 100 to 599');
+    }
+
+    const source = config.bodySource || 'json';
+    const contentType = config.contentType || 'application/json';
+
+    let body;
+    if (source === 'input') {
+        body = smartbotic.utils.getFieldValue(input, config.bodyField || 'data');
+        if (body === undefined) {
+            throw new Error('Respond to Webhook: nothing found at "' +
+                (config.bodyField || 'data') + '" to send as the body');
+        }
+    } else if (source === 'text') {
+        body = config.body === undefined || config.body === null ? '' : String(config.body);
+    } else {
+        const raw = config.body === undefined || config.body === null ? '' : String(config.body);
+        if (raw.length === 0) {
+            body = {};
+        } else {
+            try {
+                body = JSON.parse(raw);
+            } catch (e) {
+                throw new Error('Respond to Webhook: the body is not valid JSON: ' + e.message +
+                    '. Use the text source to send it as it is');
+            }
+        }
+    }
+
+    const headers = {};
+    if (config.headers && typeof config.headers === 'object') {
+        for (const name of Object.keys(config.headers)) {
+            // A header named content-type in any casing would otherwise survive
+            // alongside the canonical one, and which of the two reaches the
+            // client would come down to key ordering rather than intent.
+            if (name.toLowerCase() === 'content-type') {
+                continue;
+            }
+            headers[name] = String(config.headers[name]);
+        }
+    }
+    headers['Content-Type'] = contentType;
+
+    smartbotic.log.info('Respond to Webhook: ' + status + ' as ' + contentType);
+
+    return {
+        status: status,
+        respondedWith: contentType,
+        _webhookResponse: {
+            status: status,
+            headers: headers,
+            body: body
+        }
+    };
+}
+
+module.exports = { configSchema, inputSchema, outputSchema, execute };

+ 107 - 0
nodes/core/wait-for-approval.js

@@ -0,0 +1,107 @@
+/**
+ * @node wait-for-approval
+ * @name Wait for Approval
+ * @category flow-control
+ * @version 1.0.0
+ * @description Pause the execution until someone answers, then continue with what they sent
+ * @icon user-check
+ */
+
+const configSchema = {
+    type: 'object',
+    properties: {
+        reason: {
+            type: 'string',
+            title: 'Reason',
+            description: 'What the approver is being asked to decide. Supports {{variable}} interpolation',
+            format: 'textarea'
+        },
+        expiresIn: {
+            type: 'number',
+            title: 'Expires In (hours)',
+            description: 'How long the request stays answerable. 0 means it never expires, though the execution is still capped at 30 days',
+            default: 24
+        },
+        fields: {
+            type: 'array',
+            title: 'Fields To Collect',
+            description: 'Extra values the approver is asked for, sent back on the answer',
+            items: {
+                type: 'object',
+                properties: {
+                    name: { type: 'string', title: 'Name' },
+                    title: { type: 'string', title: 'Label' },
+                    type: { type: 'string', title: 'Type', enum: ['string', 'number', 'boolean'], default: 'string' }
+                }
+            }
+        }
+    }
+};
+
+const inputSchema = {
+    type: 'object',
+    properties: {
+        data: { type: 'any' }
+    }
+};
+
+const outputSchema = {
+    type: 'object',
+    properties: {
+        token: { type: 'string', description: 'Token that must be presented to answer this request' },
+        reason: { type: 'string' },
+        expiresAt: { type: 'number', description: 'Milliseconds since the epoch, 0 when it never expires' },
+        approved: { type: 'boolean', description: 'Set on the answer, after the execution resumes' },
+        answeredBy: { type: 'string', description: 'Set on the answer' },
+        answeredAt: { type: 'number', description: 'Set on the answer' },
+        fields: { type: 'array', description: 'Extra values the approver was asked for' }
+    }
+};
+
+const MAX_EXPIRY_HOURS = 30 * 24;
+
+async function execute(config, input, context) {
+    // Loop iterations keep their state in engine locals that the stored
+    // execution does not describe, so a pause inside one could be recorded but
+    // never resumed. Failing here is far better than accepting an approval that
+    // can never be answered.
+    if (input && (input.index !== undefined || input.currentIndex !== undefined)) {
+        throw new Error('Wait for Approval: this node cannot be used inside a Loop body, ' +
+            'because a paused loop iteration cannot be resumed. Collect the items first, ' +
+            'approve once, then loop');
+    }
+
+    const hours = config.expiresIn === undefined || config.expiresIn === null || config.expiresIn === ''
+        ? 24
+        : Number(config.expiresIn);
+    if (isNaN(hours) || hours < 0) {
+        throw new Error('Wait for Approval: expiresIn must be a number of hours, got "' +
+            config.expiresIn + '"');
+    }
+    if (hours > MAX_EXPIRY_HOURS) {
+        throw new Error('Wait for Approval: expiresIn is capped at ' + MAX_EXPIRY_HOURS +
+            ' hours. An approval nobody answers should eventually stop occupying the queue');
+    }
+
+    const token = smartbotic.utils.uuid();
+    const expiresAt = hours === 0 ? 0 : Date.now() + Math.round(hours * 60 * 60 * 1000);
+    const reason = config.reason ? String(config.reason) : 'Approval required';
+    const fields = Array.isArray(config.fields) ? config.fields : [];
+
+    smartbotic.log.info('Wait for Approval: pausing for "' + reason + '"');
+
+    return {
+        token: token,
+        reason: reason,
+        expiresAt: expiresAt,
+        fields: fields,
+        _pause: {
+            token: token,
+            reason: reason,
+            expiresAt: expiresAt,
+            fields: fields
+        }
+    };
+}
+
+module.exports = { configSchema, inputSchema, outputSchema, execute };

+ 189 - 0
nodes/triggers/database-change.js

@@ -0,0 +1,189 @@
+/**
+ * @node database-change
+ * @name Database Change
+ * @category triggers
+ * @version 1.0.0
+ * @description Report documents added or changed in a collection since the last run, paired with a schedule trigger for its cadence
+ * @icon database
+ */
+
+const configSchema = {
+    type: 'object',
+    properties: {
+        collection: {
+            type: 'string',
+            title: 'Collection',
+            description: 'Collection to watch'
+        },
+        timestampField: {
+            type: 'string',
+            title: 'Timestamp Field',
+            description: 'Field holding the last-modified time. Defaults to the _updated_at the database maintains itself, in nanoseconds. A document with no readable value in this field is never reported',
+            default: '_updated_at'
+        },
+        filter: {
+            type: 'object',
+            title: 'Filter',
+            description: 'Optional query restricting which documents are watched',
+            additionalProperties: true
+        },
+        maxDocuments: {
+            type: 'number',
+            title: 'Max Documents',
+            description: 'Most documents to report in one run, so a large backlog does not arrive as one enormous payload',
+            default: 100
+        },
+        emitOnFirstRun: {
+            type: 'boolean',
+            title: 'Report Everything On First Run',
+            description: 'Off records the current high-water mark and reports nothing, so adding this to a live workflow does not fire for every document already there',
+            default: false
+        },
+        cursorCollection: {
+            type: 'string',
+            title: 'Cursor Collection',
+            default: 'watch_cursors'
+        }
+    },
+    required: ['collection']
+};
+
+const inputSchema = {
+    type: 'object',
+    properties: {
+        data: { type: 'any' }
+    }
+};
+
+const outputSchema = {
+    type: 'object',
+    properties: {
+        documents: { type: 'array', description: 'Documents new or changed since the last run' },
+        count: { type: 'number' },
+        isFirstRun: { type: 'boolean' },
+        cursor: { type: 'number', description: 'High-water timestamp stored for the next run' }
+    }
+};
+
+async function execute(config, input, context) {
+    const collection = config.collection;
+    if (!collection) {
+        throw new Error('Database Change: a collection is required');
+    }
+
+    const timestampField = config.timestampField || '_updated_at';
+    const cursorCollection = config.cursorCollection || 'watch_cursors';
+    const maxDocuments = Number(config.maxDocuments) || 100;
+    const workflowId = (context && context.workflowId) || 'unknown';
+    const nodeId = (context && context.nodeId) || 'database-change';
+    const cursorId = workflowId + ':' + nodeId;
+
+    const stored = smartbotic.storage.get(cursorCollection, cursorId);
+    // A missing document is the expected shape of a first run. Anything else
+    // that keeps storage.get from returning the cursor - no read access to
+    // the collection, a connection error - is a real failure and must not be
+    // swallowed into a false "first run" that then quietly overwrites state.
+    // This mirrors nodes/triggers/file-watch.js so the two nodes agree on
+    // what "first run" means.
+    const cursorMissing = stored && typeof stored.error === 'string' &&
+        stored.error.indexOf('Document not found') === 0;
+    if (stored && stored.found !== true && !cursorMissing) {
+        throw new Error('Database Change: could not read the cursor for ' + collection + ': ' +
+            (stored.error || 'unknown storage error'));
+    }
+    const isFirstRun = !stored || stored.found !== true;
+    const since = (!isFirstRun && stored.document && Number(stored.document.since)) || 0;
+
+    const query = config.filter && typeof config.filter === 'object' ? config.filter : {};
+    const result = smartbotic.storage.query(collection, query);
+    // storage.query signals failure with an "error" field, not a "success"
+    // flag - there is no success flag in its response at all, so checking
+    // one would throw on every call, successful or not.
+    if (!result || typeof result.error === 'string') {
+        throw new Error('Database Change: could not query ' + collection + ': ' +
+            ((result && result.error) || 'unknown error'));
+    }
+
+    const documents = result.documents || [];
+
+    // The database stores _updated_at in nanoseconds while everything a node
+    // sees is milliseconds, so the value is normalised rather than compared
+    // against a cursor in different units.
+    function stampOf(document) {
+        const raw = Number(smartbotic.utils.getFieldValue(document, timestampField));
+        if (!raw || isNaN(raw)) {
+            return 0;
+        }
+        return raw > 1e15 ? Math.floor(raw / 1000000) : raw;
+    }
+
+    let highWater = since;
+    let noStamp = 0;
+    for (const document of documents) {
+        const stamp = stampOf(document);
+        if (stamp === 0) {
+            noStamp++;
+        }
+        if (stamp > highWater) {
+            highWater = stamp;
+        }
+    }
+    if (noStamp > 0) {
+        // A document with no readable value in timestampField sorts as if it
+        // predates everything, and once the mark advances past 0 it can never
+        // satisfy the strict > filter below - it is never reported, silently,
+        // unless this is logged.
+        smartbotic.log.warn('Database Change: ' + noStamp + ' document(s) in ' + collection +
+            ' had no readable "' + timestampField + '" value and will never be reported');
+    }
+
+    let changed = [];
+    if (isFirstRun && config.emitOnFirstRun !== true) {
+        smartbotic.log.info('Database Change: first run on ' + collection + ', recorded the mark at ' +
+            highWater + ' without reporting ' + documents.length + ' documents');
+    } else {
+        changed = documents
+            .filter(function (document) { return stampOf(document) > since; })
+            .sort(function (left, right) { return stampOf(left) - stampOf(right); });
+        if (changed.length > maxDocuments) {
+            let cut = maxDocuments;
+            const boundary = stampOf(changed[cut - 1]);
+            // Cutting through a group that shares one timestamp would strand
+            // the rest of that group: the mark advances to the shared stamp
+            // and the next run's strict > excludes them forever. The slice
+            // grows to take the whole tie, even though that overshoots the
+            // limit. If every changed document shares one stamp, cut grows to
+            // all of them, which is correct - nothing is stranded.
+            while (cut < changed.length && stampOf(changed[cut]) === boundary) {
+                cut++;
+            }
+            smartbotic.log.warn('Database Change: ' + changed.length + ' documents changed, reporting ' +
+                cut + (cut > maxDocuments ? ' (over the configured ' + maxDocuments + ', to avoid splitting a tied timestamp)' : '') +
+                '. The rest arrive on the next run');
+            changed = changed.slice(0, cut);
+            // The mark only advances as far as what was actually reported, or
+            // the remainder would be skipped rather than deferred.
+            highWater = stampOf(changed[changed.length - 1]);
+        }
+    }
+
+    const cursor = { since: highWater, updatedAt: Date.now(), collection: collection };
+    const write = isFirstRun
+        ? smartbotic.storage.insert(cursorCollection, cursor, cursorId)
+        : smartbotic.storage.update(cursorCollection, cursorId, cursor);
+    if (!write || write.success === false) {
+        throw new Error('Database Change: could not save the cursor for ' + collection + ': ' +
+            ((write && write.error) || 'unknown storage error'));
+    }
+
+    smartbotic.log.info('Database Change: ' + changed.length + ' documents from ' + collection);
+
+    return {
+        documents: changed,
+        count: changed.length,
+        isFirstRun: isFirstRun,
+        cursor: highWater
+    };
+}
+
+module.exports = { configSchema, inputSchema, outputSchema, execute };

+ 146 - 0
nodes/triggers/file-watch.js

@@ -0,0 +1,146 @@
+/**
+ * @node file-watch
+ * @name File Watch
+ * @category triggers
+ * @version 1.0.0
+ * @description Report files added or changed in a directory since the last run, paired with a schedule trigger for its cadence
+ * @icon folder-search
+ */
+
+const configSchema = {
+    type: 'object',
+    properties: {
+        directory: {
+            type: 'string',
+            title: 'Directory',
+            description: 'Absolute path to watch. Not recursive'
+        },
+        pattern: {
+            type: 'string',
+            title: 'Name Pattern',
+            description: 'Optional regular expression a file name must match, such as \\.csv$'
+        },
+        includeDirectories: {
+            type: 'boolean',
+            title: 'Include Directories',
+            description: 'Report subdirectories as well as files',
+            default: false
+        },
+        emitOnFirstRun: {
+            type: 'boolean',
+            title: 'Report Everything On First Run',
+            description: 'On the very first run there is nothing to compare against. Off records what is there and reports nothing, so adding this to a live workflow does not fire for every file already present',
+            default: false
+        },
+        cursorCollection: {
+            type: 'string',
+            title: 'Cursor Collection',
+            description: 'Collection holding the last-seen state',
+            default: 'watch_cursors'
+        }
+    },
+    required: ['directory']
+};
+
+const inputSchema = {
+    type: 'object',
+    properties: {
+        data: { type: 'any' }
+    }
+};
+
+const outputSchema = {
+    type: 'object',
+    properties: {
+        files: { type: 'array', description: 'Files new or changed since the last run' },
+        count: { type: 'number' },
+        isFirstRun: { type: 'boolean', description: 'True when no cursor existed yet' }
+    }
+};
+
+async function execute(config, input, context) {
+    const directory = config.directory;
+    if (!directory) {
+        throw new Error('File Watch: a directory is required');
+    }
+
+    const collection = config.cursorCollection || 'watch_cursors';
+    const workflowId = (context && context.workflowId) || 'unknown';
+    const nodeId = (context && context.nodeId) || 'file-watch';
+    const cursorId = workflowId + ':' + nodeId;
+
+    const listing = smartbotic.fs.readdir(directory);
+    if (!listing || listing.success !== true) {
+        throw new Error('File Watch: could not read ' + directory + ': ' +
+            ((listing && listing.error) || 'unknown error'));
+    }
+
+    let matcher = null;
+    if (config.pattern) {
+        try {
+            matcher = new RegExp(config.pattern);
+        } catch (e) {
+            throw new Error('File Watch: "' + config.pattern + '" is not a valid pattern: ' + e.message);
+        }
+    }
+
+    const current = {};
+    const candidates = [];
+    for (const entry of listing.entries) {
+        if (entry.isDirectory && config.includeDirectories !== true) {
+            continue;
+        }
+        if (matcher && !matcher.test(entry.name)) {
+            continue;
+        }
+        // Size and modification time together, because a file rewritten within
+        // the same second at the same length is not a change worth waking a
+        // workflow for, and a timestamp alone misses a rewrite that preserves
+        // mtime granularity.
+        current[entry.name] = entry.modifiedAt + ':' + entry.size;
+        candidates.push(entry);
+    }
+
+    const stored = smartbotic.storage.get(collection, cursorId);
+    // A missing document is the expected shape of a first run. Anything else
+    // that keeps storage.get from returning the cursor - no read access to
+    // the collection, a connection error - is a real failure and must not be
+    // swallowed into a false "first run" that then quietly overwrites state.
+    const cursorMissing = stored && typeof stored.error === 'string' &&
+        stored.error.indexOf('Document not found') === 0;
+    if (stored && stored.found !== true && !cursorMissing) {
+        throw new Error('File Watch: could not read the cursor for ' + directory + ': ' +
+            (stored.error || 'unknown storage error'));
+    }
+    const isFirstRun = !stored || stored.found !== true;
+    const previous = (!isFirstRun && stored.document && stored.document.seen) || {};
+
+    let changed = [];
+    if (isFirstRun && config.emitOnFirstRun !== true) {
+        smartbotic.log.info('File Watch: first run on ' + directory + ', recorded ' +
+            candidates.length + ' entries without reporting them');
+    } else {
+        changed = candidates.filter(function (entry) {
+            return previous[entry.name] !== current[entry.name];
+        });
+    }
+
+    const cursor = { seen: current, updatedAt: Date.now(), directory: directory };
+    const write = isFirstRun
+        ? smartbotic.storage.insert(collection, cursor, cursorId)
+        : smartbotic.storage.update(collection, cursorId, cursor);
+    if (!write || write.success === false) {
+        throw new Error('File Watch: could not save the cursor for ' + directory + ': ' +
+            ((write && write.error) || 'unknown storage error'));
+    }
+
+    smartbotic.log.info('File Watch: ' + changed.length + ' new or changed in ' + directory);
+
+    return {
+        files: changed,
+        count: changed.length,
+        isFirstRun: isFirstRun
+    };
+}
+
+module.exports = { configSchema, inputSchema, outputSchema, execute };

+ 10 - 0
proto/runner.proto

@@ -201,6 +201,13 @@ message ExecuteNodeResponse {
     int64 execution_time_ms = 4;
 }
 
+// Resume execution request
+message ResumeExecutionRequest {
+    string execution_id = 1;
+    string token = 2;
+    string payload = 3;  // JSON given by whoever answered
+}
+
 // Runner service - called by WebServer to execute workflows
 service RunnerService {
     // Execute a workflow
@@ -209,6 +216,9 @@ service RunnerService {
     // Cancel an execution
     rpc CancelExecution(CancelExecutionRequest) returns (Empty);
 
+    // Continue an execution that paused for an answer
+    rpc ResumeExecution(ResumeExecutionRequest) returns (ExecuteWorkflowResponse);
+
     // Node management
     rpc ListNodes(ListNodesRequest) returns (ListNodesResponse);
     rpc ReloadNode(ReloadNodeRequest) returns (ReloadNodeResponse);

+ 78 - 1
scripts/verify-node.py

@@ -55,6 +55,72 @@ def subset_matches(expected, actual, path):
     return fails
 
 
+def run_http_case(case, token, workflow_id):
+    """A webhook case asserts on the HTTP response, not on node outputs."""
+    spec = case["http"]
+    url = BASE.replace("/api/v1", "") + spec["path"].replace("{workflowId}", workflow_id)
+    data = json.dumps(spec.get("body", {})).encode()
+    req = urllib.request.Request(url, data=data, method=spec.get("method", "POST"))
+    req.add_header("Content-Type", "application/json")
+
+    failures = []
+    try:
+        with urllib.request.urlopen(req, timeout=45) as res:
+            status, headers, raw = res.status, dict(res.headers), res.read().decode()
+    except urllib.error.HTTPError as e:
+        status, headers, raw = e.code, dict(e.headers), e.read().decode()
+
+    print("http: %s %s -> %s" % (spec.get("method", "POST"), spec["path"], status))
+    print("  body: %s" % raw[:200])
+
+    if "expectStatus" in spec and status != spec["expectStatus"]:
+        failures.append("status: expected %s, got %s" % (spec["expectStatus"], status))
+
+    for name, want in spec.get("expectHeaders", {}).items():
+        got = headers.get(name)
+        if got is None:
+            failures.append("header %s: missing" % name)
+        elif want not in got:
+            failures.append("header %s: expected to contain %r, got %r" % (name, want, got))
+
+    for fragment in spec.get("expectBodyContains", []):
+        if fragment not in raw:
+            failures.append("body: expected to contain %r" % fragment)
+
+    return failures
+
+
+def run_resume_case(case, token, execution_id):
+    """Wait for the execution to pause, answer it, then let it finish."""
+    spec = case["resume"]
+    want_status = spec.get("waitForStatus", "waiting")
+
+    execution = {}
+    for _ in range(60):
+        time.sleep(0.5)
+        execution = call("GET", "/executions/%s" % execution_id, token)
+        if execution.get("status") == want_status:
+            break
+        if execution.get("status") in ("completed", "failed", "cancelled"):
+            raise SystemExit(
+                "execution %s reached %s without pausing" % (execution_id, execution.get("status")))
+    else:
+        raise SystemExit("execution %s never reached %s" % (execution_id, want_status))
+
+    by_id = {n["nodeId"]: n for n in execution.get("nodeExecutions", [])}
+    source = by_id.get(spec["tokenFrom"])
+    if not source:
+        raise SystemExit("resume: node %s did not run" % spec["tokenFrom"])
+    resume_token = (source.get("output") or {}).get("token")
+    if not resume_token:
+        raise SystemExit("resume: node %s produced no token" % spec["tokenFrom"])
+
+    body = dict(spec.get("payload", {}))
+    body["token"] = resume_token
+    call("POST", "/executions/%s/resume" % execution_id, token, body)
+    print("resumed %s with token %s" % (execution_id, resume_token))
+
+
 def main():
     if len(sys.argv) != 2:
         raise SystemExit("usage: verify-node.py <case.json>")
@@ -67,16 +133,27 @@ def main():
         "name": case["name"],
         "nodes": case["nodes"],
         "connections": case["connections"],
-        "settings": {},
+        "settings": case.get("settings", {}),
     })
     workflow_id = created.get("id") or created.get("_id")
     if not workflow_id:
         raise SystemExit(f"no workflow id in create response: {json.dumps(created)[:300]}")
 
     try:
+        if "http" in case:
+            call("POST", f"/workflows/{workflow_id}/activate", token, {})
+            failures = run_http_case(case, token, workflow_id)
+            print("\nFAIL" if failures else "\nPASS")
+            for f in failures:
+                print("  " + f)
+            return 1 if failures else 0
+
         started = call("POST", f"/workflows/{workflow_id}/execute", token, {})
         execution_id = started["executionId"]
 
+        if "resume" in case:
+            run_resume_case(case, token, execution_id)
+
         execution = None
         for _ in range(60):
             time.sleep(0.5)

+ 100 - 2
src/runner/engine/script_engine.cpp

@@ -2539,8 +2539,13 @@ void ScriptEngine::setupBuiltinAPIs() {
             auto status = std::filesystem::status(path_str);
             auto size = std::filesystem::is_regular_file(path_str) ? std::filesystem::file_size(path_str) : 0;
             auto mtime = std::filesystem::last_write_time(path_str);
+            // file_time_type's epoch is not guaranteed to match system_clock's
+            // (it does not on this libstdc++), so this has to go through
+            // clock_cast rather than time_since_epoch() directly, or mtime
+            // comes out shifted by whatever offset the filesystem clock uses.
+            auto sys_time = std::chrono::clock_cast<std::chrono::system_clock>(mtime);
             auto mtime_ms = std::chrono::duration_cast<std::chrono::milliseconds>(
-                mtime.time_since_epoch()
+                sys_time.time_since_epoch()
             ).count();
 
             JSValue response = JS_NewObject(ctx);
@@ -2558,6 +2563,98 @@ void ScriptEngine::setupBuiltinAPIs() {
         }
     }, "stat", 1));
 
+    // fs.readdir(path) - List a directory, one level only. Recursing here would
+    // be an unbounded walk driven by whatever happens to be on disk, so a node
+    // that wants a tree walks it itself.
+    JS_SetPropertyStr(ctx, fs, "readdir", JS_NewCFunction(ctx, [](JSContext* ctx, JSValue this_val, int argc, JSValue* argv) -> JSValue {
+        if (argc < 1) {
+            return JS_ThrowTypeError(ctx, "fs.readdir requires path argument");
+        }
+
+        const char* path = JS_ToCString(ctx, argv[0]);
+        if (!path) {
+            return JS_ThrowTypeError(ctx, "path must be a string");
+        }
+        std::string path_str(path);
+        JS_FreeCString(ctx, path);
+
+        // Held outside the try so the catch block can free it if an
+        // exception escapes after it was created but before it was handed
+        // to a response object - otherwise a partially built entries array
+        // (and everything already pushed into it) leaks. Freeing
+        // JS_UNDEFINED is a no-op, so this is safe even if the exception
+        // happened before entries was ever created.
+        JSValue entries = JS_UNDEFINED;
+
+        try {
+            if (!std::filesystem::exists(path_str)) {
+                JSValue response = JS_NewObject(ctx);
+                JS_SetPropertyStr(ctx, response, "success", JS_FALSE);
+                JS_SetPropertyStr(ctx, response, "error",
+                                  JS_NewString(ctx, ("Directory not found: " + path_str).c_str()));
+                return response;
+            }
+            if (!std::filesystem::is_directory(path_str)) {
+                JSValue response = JS_NewObject(ctx);
+                JS_SetPropertyStr(ctx, response, "success", JS_FALSE);
+                JS_SetPropertyStr(ctx, response, "error",
+                                  JS_NewString(ctx, ("Not a directory: " + path_str).c_str()));
+                return response;
+            }
+
+            entries = JS_NewArray(ctx);
+            uint32_t index = 0;
+
+            for (const auto& entry : std::filesystem::directory_iterator(path_str)) {
+                JSValue item = JS_NewObject(ctx);
+                JS_SetPropertyStr(ctx, item, "name",
+                                  JS_NewString(ctx, entry.path().filename().string().c_str()));
+                JS_SetPropertyStr(ctx, item, "path",
+                                  JS_NewString(ctx, entry.path().string().c_str()));
+
+                int64_t size = 0;
+                int64_t mtime_ms = 0;
+                bool is_dir = false;
+                // A file removed between listing and stat is ordinary on a
+                // directory being written to, and is reported with zeroes
+                // rather than failing the whole listing.
+                try {
+                    is_dir = entry.is_directory();
+                    if (entry.is_regular_file()) {
+                        size = static_cast<int64_t>(entry.file_size());
+                    }
+                    auto mtime = entry.last_write_time();
+                    // file_time_type's epoch is not guaranteed to match
+                    // system_clock's (it does not on this libstdc++), so this
+                    // has to go through clock_cast rather than
+                    // time_since_epoch() directly, or modifiedAt comes out
+                    // shifted by whatever offset the filesystem clock uses.
+                    auto sys_time = std::chrono::clock_cast<std::chrono::system_clock>(mtime);
+                    mtime_ms = std::chrono::duration_cast<std::chrono::milliseconds>(
+                        sys_time.time_since_epoch()).count();
+                } catch (const std::exception&) {
+                }
+
+                JS_SetPropertyStr(ctx, item, "isDirectory", is_dir ? JS_TRUE : JS_FALSE);
+                JS_SetPropertyStr(ctx, item, "size", JS_NewInt64(ctx, size));
+                JS_SetPropertyStr(ctx, item, "modifiedAt", JS_NewInt64(ctx, mtime_ms));
+
+                JS_SetPropertyUint32(ctx, entries, index++, item);
+            }
+
+            JSValue response = JS_NewObject(ctx);
+            JS_SetPropertyStr(ctx, response, "success", JS_TRUE);
+            JS_SetPropertyStr(ctx, response, "entries", entries);
+            return response;
+        } catch (const std::exception& e) {
+            JS_FreeValue(ctx, entries);
+            JSValue response = JS_NewObject(ctx);
+            JS_SetPropertyStr(ctx, response, "success", JS_FALSE);
+            JS_SetPropertyStr(ctx, response, "error", JS_NewString(ctx, e.what()));
+            return response;
+        }
+    }, "readdir", 1));
+
     JS_SetPropertyStr(ctx, smartbotic, "fs", fs);
 
     // smartbotic.process - Process execution API for external commands
@@ -4392,7 +4489,8 @@ ScriptResult ScriptEngine::execute(const std::string& script, const ScriptContex
     // Create context object for the execute function
     nlohmann::json ctx_json = {
         {"executionId", context.execution_id},
-        {"nodeId", context.node_id}
+        {"nodeId", context.node_id},
+        {"workflowId", context.workflow_id}
     };
     std::string ctx_str = ctx_json.dump();
     JSValue ctx_val = JS_ParseJSON(context_, ctx_str.c_str(), ctx_str.size(), "<context>");

+ 73 - 2
src/runner/runner_service.cpp

@@ -18,6 +18,24 @@ namespace smartbotic::runner {
 
 using namespace common;
 
+// Maps the engine's ExecutionStatus onto the wire enum explicitly. The two
+// enums are not numerically aligned: proto::ExecutionStatus reserves 0 for
+// EXECUTION_STATUS_UNSPECIFIED, while ExecutionStatus::Pending is 0, so a
+// bare static_cast silently shifts every status by one. Deliberately no
+// default label, so an unhandled case is a compiler warning rather than a
+// silent mismatch.
+static proto::ExecutionStatus toProtoStatus(ExecutionStatus status) {
+    switch (status) {
+        case ExecutionStatus::Pending:   return proto::EXECUTION_STATUS_PENDING;
+        case ExecutionStatus::Running:   return proto::EXECUTION_STATUS_RUNNING;
+        case ExecutionStatus::Completed: return proto::EXECUTION_STATUS_COMPLETED;
+        case ExecutionStatus::Failed:    return proto::EXECUTION_STATUS_FAILED;
+        case ExecutionStatus::Cancelled: return proto::EXECUTION_STATUS_CANCELLED;
+        case ExecutionStatus::Waiting:   return proto::EXECUTION_STATUS_WAITING;
+    }
+    return proto::EXECUTION_STATUS_UNSPECIFIED;
+}
+
 // RunnerServiceImpl implementation
 RunnerServiceImpl::RunnerServiceImpl(WorkflowEngine& engine, NodeRegistry& registry,
                                      storage::StorageClient& storage,
@@ -119,10 +137,16 @@ grpc::Status RunnerServiceImpl::ExecuteWorkflow(grpc::ServerContext* context,
     }
 
     response->set_execution_id(result.value().execution_id);
-    response->set_status(static_cast<proto::ExecutionStatus>(result.value().status));
+    response->set_status(toProtoStatus(result.value().status));
 
     if (request->wait_for_completion()) {
-        response->set_result(result.value().final_output.dump());
+        if (!result.value().webhook_response.is_null()) {
+            nlohmann::json envelope;
+            envelope["_webhookResponse"] = result.value().webhook_response;
+            response->set_result(envelope.dump());
+        } else {
+            response->set_result(result.value().final_output.dump());
+        }
     }
 
     return grpc::Status::OK;
@@ -135,6 +159,53 @@ grpc::Status RunnerServiceImpl::CancelExecution(grpc::ServerContext* context,
     return grpc::Status::OK;
 }
 
+grpc::Status RunnerServiceImpl::ResumeExecution(grpc::ServerContext* context,
+                                                const proto::ResumeExecutionRequest* request,
+                                                proto::ExecuteWorkflowResponse* response) {
+    nlohmann::json payload = nlohmann::json::object();
+    if (!request->payload().empty()) {
+        try {
+            payload = nlohmann::json::parse(request->payload());
+        } catch (const std::exception& e) {
+            return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT,
+                                std::string("payload is not JSON: ") + e.what());
+        }
+    }
+
+    // The resumed half of the run needs the same event plumbing the initial
+    // half gets in ExecuteWorkflow above, or the UI shows nothing for it and
+    // execution.failed never reaches the handler that runs error workflows.
+    // ExecuteWorkflow's callback stamps workflowId onto every event because
+    // most of the engine's per-node events don't carry it themselves; this
+    // does the same, reading workflowId from the execution record up front
+    // since resume() (unlike execute()) is not handed a parsed Workflow by
+    // its caller.
+    std::string workflow_id;
+    auto stored = storage_.get("executions", request->execution_id());
+    if (stored.ok()) {
+        workflow_id = stored.value().value("workflowId", "");
+    }
+
+    ExecutionCallback callback;
+    if (event_callback_) {
+        callback = [this, workflow_id](const std::string& event_type, const nlohmann::json& data) {
+            nlohmann::json event_data = data;
+            event_data["workflowId"] = workflow_id;
+            event_callback_(event_type, event_data);
+        };
+    }
+
+    auto result = engine_.resume(request->execution_id(), request->token(), payload, callback);
+    if (result.failed()) {
+        return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, result.error().message());
+    }
+
+    response->set_execution_id(result.value().execution_id);
+    response->set_status(toProtoStatus(result.value().status));
+    response->set_result(result.value().final_output.dump());
+    return grpc::Status::OK;
+}
+
 grpc::Status RunnerServiceImpl::ListNodes(grpc::ServerContext* context,
                                           const proto::ListNodesRequest* request,
                                           proto::ListNodesResponse* response) {

+ 4 - 0
src/runner/runner_service.hpp

@@ -38,6 +38,10 @@ public:
                                  const proto::CancelExecutionRequest* request,
                                  proto::Empty* response) override;
 
+    grpc::Status ResumeExecution(grpc::ServerContext* context,
+                                 const proto::ResumeExecutionRequest* request,
+                                 proto::ExecuteWorkflowResponse* response) override;
+
     grpc::Status ListNodes(grpc::ServerContext* context,
                           const proto::ListNodesRequest* request,
                           proto::ListNodesResponse* response) override;

+ 398 - 11
src/runner/workflow_engine.cpp

@@ -23,6 +23,20 @@ std::string nodeStatusToString(NodeStatus status) {
     }
 }
 
+// Inverse of nodeStatusToString, used to seed a resumed execution's node
+// results with the status each node actually finished with rather than
+// forcing everything to Completed. Anything unrecognised defaults to
+// Completed, matching prior behavior for records that predate this field.
+static NodeStatus nodeStatusFromString(const std::string& s) {
+    if (s == "pending") return NodeStatus::Pending;
+    if (s == "running") return NodeStatus::Running;
+    if (s == "completed") return NodeStatus::Completed;
+    if (s == "failed") return NodeStatus::Failed;
+    if (s == "skipped") return NodeStatus::Skipped;
+    if (s == "disabled") return NodeStatus::Disabled;
+    return NodeStatus::Completed;
+}
+
 std::string executionStatusToString(ExecutionStatus status) {
     switch (status) {
         case ExecutionStatus::Pending: return "pending";
@@ -30,6 +44,7 @@ std::string executionStatusToString(ExecutionStatus status) {
         case ExecutionStatus::Completed: return "completed";
         case ExecutionStatus::Failed: return "failed";
         case ExecutionStatus::Cancelled: return "cancelled";
+        case ExecutionStatus::Waiting: return "waiting";
         default: return "unknown";
     }
 }
@@ -139,6 +154,14 @@ nlohmann::json ExecutionResult::toJson() const {
     j["error"] = error;
     j["output"] = truncateLargeValues(final_output);
 
+    // A finished execution's node outputs are a log, so large strings are
+    // truncated. A Waiting execution's node outputs are the resume state a
+    // later resume seeds itself from - truncating them would hand the resume
+    // a literal "[omitted N bytes]" placeholder in place of real data, so
+    // they are kept whole. Do not "restore consistency" here later; that
+    // would silently reintroduce truncated resume data.
+    const bool keep_full_outputs = (status == ExecutionStatus::Waiting);
+
     j["nodeExecutions"] = nlohmann::json::array();
     for (const auto& [id, result] : node_results) {
         nlohmann::json nr;
@@ -148,7 +171,7 @@ nlohmann::json ExecutionResult::toJson() const {
         nr["finishedAt"] = result.finished_at;
         // Note: input is intentionally not stored to avoid data duplication
         // Each node's input can be reconstructed from upstream node outputs + connections
-        nr["output"] = truncateLargeValues(result.output);
+        nr["output"] = keep_full_outputs ? result.output : truncateLargeValues(result.output);
         nr["error"] = result.error;
         nr["retryCount"] = result.retry_count;
         j["nodeExecutions"].push_back(nr);
@@ -159,6 +182,16 @@ nlohmann::json ExecutionResult::toJson() const {
         j["workflowSnapshot"] = workflow_snapshot;
     }
 
+    if (!webhook_response.is_null()) {
+        j["webhookResponse"] = webhook_response;
+    }
+
+    if (!paused_node_id.empty()) {
+        j["pausedNodeId"] = paused_node_id;
+        j["pauseToken"] = pause_token;
+        j["pauseExpiresAt"] = pause_expires_at;
+    }
+
     return j;
 }
 
@@ -216,6 +249,20 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
         LOG_INFO("Received cached outputs for {} nodes", cached_outputs.size());
     }
 
+    // A resume carries the results of everything that ran before the pause, and
+    // reuses the original execution id so the run reads as one execution rather
+    // than two.
+    nlohmann::json resume_seed;
+    std::string resume_execution_id;
+    if (actual_trigger_data.contains("_resumeSeed")) {
+        resume_seed = actual_trigger_data["_resumeSeed"];
+        actual_trigger_data.erase("_resumeSeed");
+    }
+    if (actual_trigger_data.contains("_resumeExecutionId")) {
+        resume_execution_id = actual_trigger_data["_resumeExecutionId"].get<std::string>();
+        actual_trigger_data.erase("_resumeExecutionId");
+    }
+
     // Create execution record
     ExecutionResult result;
     result.execution_id = UUID::generatePrefixed("exec");
@@ -227,8 +274,36 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
     result.trigger_data = actual_trigger_data;
     result.started_at = TimeUtils::nowMs();
 
-    // Create workflow snapshot for pinning feature
+    if (!resume_execution_id.empty()) {
+        result.execution_id = resume_execution_id;
+    }
+    for (auto it = resume_seed.begin(); it != resume_seed.end(); ++it) {
+        NodeExecutionResult seeded;
+        seeded.node_id = it.key();
+        // The seed carries the real status the node finished with before the
+        // pause (completed, failed, skipped, disabled). Seeding everything as
+        // Completed would make collectNodeInput stop propagating a skip, and a
+        // node downstream of a branch that was not taken could run on the
+        // resumed pass having never run on the original one.
+        seeded.status = nodeStatusFromString(it.value().value("status", "completed"));
+        seeded.output = it.value().value("output", nlohmann::json::object());
+        seeded.started_at = TimeUtils::nowMs();
+        seeded.finished_at = seeded.started_at;
+        result.node_results[it.key()] = seeded;
+    }
+
+    // Create workflow snapshot for pinning feature. id/name/settings are
+    // included, not only nodes and connections: Workflow::fromJson rebuilds a
+    // Workflow from this on resume, and an empty workflow.id is treated by
+    // credential access checks as an admin operation with unrestricted
+    // credential access (see CredentialStore::hasWorkflowAccess), so an empty
+    // id here is a privilege escalation, not a cosmetic gap. settings carries
+    // storagePermissions and continueOnError, which the resumed half of the
+    // run must honor the same way the first half did.
     nlohmann::json snapshot;
+    snapshot["id"] = workflow.id;
+    snapshot["name"] = workflow.name;
+    snapshot["settings"] = workflow.settings;
     snapshot["nodes"] = nlohmann::json::array();
     for (const auto& node : workflow.nodes) {
         nlohmann::json n;
@@ -457,8 +532,10 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
                 continue;
             }
 
-            // Skip nodes that were already executed (e.g., loop body nodes)
-            // They have results stored during loop execution
+            // Skip nodes that were already executed (e.g., loop body nodes, or -
+            // on a resume - nodes seeded from the stored execution that ran
+            // before the pause). This is the mechanism that keeps a resume from
+            // re-running work: every seeded node already has a result here.
             if (result.node_results.contains(node_id)) {
                 LOG_DEBUG("Node {} already has results, skipping in main loop", node_id);
                 continue;
@@ -600,6 +677,45 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
                 }
             }
 
+            // Any node may declare the HTTP response, not just the last one to
+            // run, so appending a node to a workflow cannot silently change
+            // what its webhook returns.
+            if (node_result.status == NodeStatus::Completed &&
+                node_result.output.contains("_webhookResponse")) {
+                if (!result.webhook_response.is_null()) {
+                    LOG_WARN("Node {} overrides a webhook response already set by an earlier node",
+                             node_id);
+                }
+                result.webhook_response = node_result.output["_webhookResponse"];
+            }
+
+            // A node asking to pause ends this pass. The execution is stored as
+            // Waiting with everything computed so far, and a later resume picks
+            // it up from here rather than starting again.
+            if (node_result.status == NodeStatus::Completed &&
+                node_result.output.contains("_pause")) {
+                const auto& pause = node_result.output["_pause"];
+
+                result.status = ExecutionStatus::Waiting;
+                result.paused_node_id = node_id;
+                result.pause_token = pause.value("token", "");
+                result.pause_expires_at = pause.value("expiresAt", static_cast<int64_t>(0));
+
+                // The marker is engine plumbing. What is stored is the request a
+                // person has to answer, not the mechanism that carried it.
+                auto& stored_node = result.node_results[node_id];
+                stored_node.output.erase("_pause");
+
+                // No callback here - the generic terminal callback below (after
+                // storeExecution) already emits "execution.waiting" once this
+                // pass ends, and it carries nodeId/expiresAt too. Emitting here
+                // as well produced the same event twice with two different
+                // payload shapes.
+
+                LOG_INFO("Execution {} paused at node {}", result.execution_id, node_id);
+                break;
+            }
+
             // Check for loop node
             if (node_result.status == NodeStatus::Completed &&
                 node_result.output.contains("_isLoop") &&
@@ -698,13 +814,23 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
         // error workflow, a notification - needs to know what went wrong, and
         // this event is the first one out, so omitting it left listeners with a
         // failure and no reason for it.
-        callback("execution." + executionStatusToString(result.status), {
+        nlohmann::json event_data = {
             {"executionId", result.execution_id},
             {"workflowId", result.workflow_id},
             {"status", executionStatusToString(result.status)},
             {"error", result.error},
             {"output", truncateLargeValues(result.final_output)}
-        });
+        };
+        // A Waiting result is the one terminal status that carries reader-
+        // relevant fields the generic shape above doesn't have - which node
+        // is asking, and when the wait itself expires. This is the only
+        // "execution.waiting" emission (the pause block above deliberately
+        // does not emit its own), so those fields land here.
+        if (result.status == ExecutionStatus::Waiting) {
+            event_data["nodeId"] = result.paused_node_id;
+            event_data["expiresAt"] = result.pause_expires_at;
+        }
+        callback("execution." + executionStatusToString(result.status), event_data);
     }
 
     LOG_INFO("Workflow execution {} completed with status: {}",
@@ -713,6 +839,149 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
     return result;
 }
 
+common::Result<ExecutionResult> WorkflowEngine::resume(const std::string& execution_id,
+                                                       const std::string& token,
+                                                       const nlohmann::json& payload,
+                                                       ExecutionCallback callback) {
+    auto stored = storage_.get("executions", execution_id);
+    if (stored.failed()) {
+        return common::Error(common::ErrorCode::NotFound, "No execution " + execution_id);
+    }
+
+    const nlohmann::json& record = stored.value();
+
+    if (record.value("status", "") != "waiting") {
+        return common::Error(common::ErrorCode::InvalidArgument,
+            "Execution " + execution_id + " is " + record.value("status", "unknown") +
+            ", not waiting for an answer");
+    }
+
+    const std::string expected_token = record.value("pauseToken", "");
+    if (expected_token.empty() || expected_token != token) {
+        return common::Error(common::ErrorCode::PermissionDenied,
+            "The token does not match the one this execution is waiting on");
+    }
+
+    const int64_t expires_at = record.value("pauseExpiresAt", static_cast<int64_t>(0));
+    if (expires_at > 0 && TimeUtils::nowMs() > expires_at) {
+        return common::Error(common::ErrorCode::InvalidArgument,
+            "This approval expired at " + std::to_string(expires_at) + " and can no longer be answered");
+    }
+
+    const std::string paused_node_id = record.value("pausedNodeId", "");
+    if (paused_node_id.empty()) {
+        return common::Error(common::ErrorCode::Internal,
+            "Execution " + execution_id + " is waiting but records no paused node");
+    }
+
+    if (!record.contains("workflowSnapshot") || !record["workflowSnapshot"].is_object()) {
+        return common::Error(common::ErrorCode::Internal,
+            "Execution " + execution_id + " has no usable workflow snapshot to resume against");
+    }
+
+    // The version to claim with is captured here, at the same point the record
+    // itself was read - nothing between here and the claim write (below, right
+    // before execute()) writes to this record, so the version stays valid to
+    // compare against no matter how much read-only validation runs in between.
+    const int64_t expected_version = record.value("_version", static_cast<int64_t>(0));
+    if (expected_version <= 0) {
+        // Every document returned by get() carries a _version metadata field
+        // stamped by the database on every write. Its absence means either a
+        // very old record predating that guarantee or something reading the
+        // record incorrectly - either way, resuming without a real
+        // compare-and-set would silently reopen the replay window this claim
+        // exists to close, so refuse rather than proceed unprotected.
+        return common::Error(common::ErrorCode::Internal,
+            "Execution " + execution_id + " has no version metadata to claim it safely with");
+    }
+
+    Workflow workflow = Workflow::fromJson(record["workflowSnapshot"]);
+
+    // Defend against snapshots stored before id/name/settings were added to
+    // them: an execution paused under the old shape still has workflowId and
+    // workflowName on the record itself, so those two are recoverable. An
+    // empty workflow.id would otherwise be read by CredentialStore::
+    // hasWorkflowAccess as an admin operation with unrestricted credential
+    // access, so this is a security backstop, not tidiness.
+    if (workflow.id.empty()) {
+        workflow.id = record.value("workflowId", "");
+    }
+    if (workflow.name.empty()) {
+        workflow.name = record.value("workflowName", "");
+    }
+    // settings is not recoverable from the record - it was never stored
+    // anywhere else. If the snapshot carried no settings, refuse rather than
+    // run the resumed half of the execution under different storage
+    // permissions and continueOnError than the half that ran before the pause.
+    if (workflow.settings.empty() && !record["workflowSnapshot"].contains("settings")) {
+        return common::Error(common::ErrorCode::FailedPrecondition,
+            "Execution " + execution_id + " was paused before workflow settings were captured "
+            "in its snapshot and cannot be safely resumed");
+    }
+
+    // Seed everything that ran before the pause, including the paused node
+    // itself, whose output becomes the answer that was given. Status is
+    // carried along with output so a node that was skipped or failed before
+    // the pause is not reported as Completed on the resumed pass.
+    nlohmann::json seed = nlohmann::json::object();
+    if (record.contains("nodeExecutions") && record["nodeExecutions"].is_array()) {
+        for (const auto& entry : record["nodeExecutions"]) {
+            const std::string node_id = entry.value("nodeId", "");
+            if (node_id.empty()) {
+                continue;
+            }
+            seed[node_id] = {
+                {"status", entry.value("status", "completed")},
+                {"output", entry.value("output", nlohmann::json::object())}
+            };
+        }
+    }
+
+    nlohmann::json answer = payload.is_object() ? payload : nlohmann::json::object();
+    answer["answeredAt"] = TimeUtils::nowMs();
+    seed[paused_node_id] = {
+        {"status", "completed"},
+        {"output", answer}
+    };
+
+    nlohmann::json trigger_data = record.value("triggerData", nlohmann::json::object());
+    trigger_data["_resumeSeed"] = seed;
+    trigger_data["_resumeExecutionId"] = execution_id;
+
+    LOG_INFO("Resuming execution {} from node {} with {} seeded results",
+             execution_id, paused_node_id, seed.size());
+
+    // Claim the execution immediately before running it, only now that every
+    // check that can refuse the resume has passed. Two concurrent resumes with
+    // the same token both pass every check above and would otherwise both call
+    // execute() on the second half of the run, producing duplicate side effects
+    // (a resumed run is not idempotent - it sends emails, charges cards, posts
+    // messages). The write below moves status out of "waiting" and clears the
+    // token/pausedNodeId so a second resume cannot match either the status
+    // check or the token check. storage_.update() with a non-zero
+    // expected_version goes through upstream updateIfVersion(), a genuine
+    // server-side compare-and-set (verified against smartbotic-database's
+    // client.cpp: it sends expected_version on the wire and the server rejects
+    // a stale write), so a racing second resume loses this write and its
+    // execute() call never happens. Claiming this late, instead of before
+    // validation, means a resume that gets refused below never touches the
+    // record: it stays "waiting", still visible on GET /executions/pending,
+    // and can be retried.
+    nlohmann::json claim = record;
+    claim.erase("_version");
+    claim["status"] = "resuming";
+    claim["pauseToken"] = "";
+    claim["pausedNodeId"] = "";
+
+    auto claim_result = storage_.update("executions", execution_id, claim, expected_version);
+    if (claim_result.failed()) {
+        return common::Error(common::ErrorCode::FailedPrecondition,
+            "Execution " + execution_id + " is already being resumed");
+    }
+
+    return execute(workflow, record.value("triggerType", "resume"), trigger_data, callback);
+}
+
 void WorkflowEngine::cancelExecution(const std::string& execution_id) {
     std::lock_guard<std::mutex> lock(mutex_);
     cancelled_executions_.insert(execution_id);
@@ -902,6 +1171,7 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
     engine::ScriptContext ctx;
     ctx.execution_id = execution_id;
     ctx.node_id = node.id;
+    ctx.workflow_id = workflow.id;
     ctx.input = input;
     ctx.config = node.config;
     ctx.log_handler = [&node](const std::string& level, const std::string& msg) {
@@ -1287,8 +1557,68 @@ nlohmann::json WorkflowEngine::collectNodeInput(
 }
 
 void WorkflowEngine::storeExecution(const ExecutionResult& result) {
-    auto insert_result = storage_.insert("executions", result.toJson(), result.execution_id,
-                   7 * 24 * 60 * 60 * 1000);  // 7-day TTL
+    // A finished execution is a log entry and ages out after a week. One that is
+    // waiting for a person is work still owed an answer, and having it expire
+    // under them loses the run silently, so it lives until its own deadline
+    // plus a day of slack for a late approver. An expiresAt of 0 means "no
+    // deadline", not "no need to extend" - it is given the 30-day ceiling the
+    // approval node caps at, so a permanent approval does not fall back to
+    // the ordinary seven-day log TTL and evaporate with no trace.
+    int64_t ttl_ms = 7 * 24 * 60 * 60 * 1000;
+    if (result.status == ExecutionStatus::Waiting) {
+        const int64_t grace = 24 * 60 * 60 * 1000;
+        const int64_t never_ttl = 30LL * 24 * 60 * 60 * 1000;
+        int64_t wanted = never_ttl;
+        if (result.pause_expires_at > 0) {
+            wanted = result.pause_expires_at - TimeUtils::nowMs() + grace;
+        }
+        if (wanted > ttl_ms) {
+            ttl_ms = wanted;
+        }
+    }
+
+    // A resumed execution reuses the id of the execution that paused, which
+    // already exists in the database. insert on an id that already exists is
+    // not guaranteed to behave like an upsert - if it rejects the duplicate,
+    // logging and returning here would leave the record at status "waiting"
+    // with its original pauseToken, letting the same resume be replayed
+    // indefinitely while this function still reports the run as Completed. So
+    // the record's existence is checked first and the write goes through
+    // update when it is already there.
+    auto existing = storage_.get("executions", result.execution_id);
+    if (existing.ok()) {
+        if (result.status == ExecutionStatus::Waiting) {
+            // storage_.update() has no ttl_ms parameter - only insert() and
+            // upsert() can set one. A second pause on an already-existing
+            // record recomputes ttl_ms above from its own (possibly much
+            // later) pause_expires_at, but a plain update() would leave the
+            // record on whatever TTL its very first insert got, undoing the
+            // floor above for every pause after the first - exactly the
+            // silent-data-loss shape called out elsewhere on this branch.
+            // upsert() swaps the expiration-index entry and installs the new
+            // TTL in one locked server-side operation, so the fresh deadline
+            // lands without the remove-then-insert window where a crash
+            // between the two calls could lose the record outright.
+            auto upserted = storage_.upsert("executions", result.toJson(), result.execution_id, ttl_ms);
+            if (upserted.failed()) {
+                LOG_ERROR("Failed to re-store waiting execution {} with a fresh TTL: {}",
+                          result.execution_id, upserted.error().message());
+            } else {
+                LOG_DEBUG("Re-stored waiting execution {} with a fresh TTL", result.execution_id);
+            }
+            return;
+        }
+
+        auto updated = storage_.update("executions", result.execution_id, result.toJson());
+        if (updated.failed()) {
+            LOG_ERROR("Failed to update execution {}: {}", result.execution_id, updated.error().message());
+        } else {
+            LOG_DEBUG("Updated execution {} successfully", result.execution_id);
+        }
+        return;
+    }
+
+    auto insert_result = storage_.insert("executions", result.toJson(), result.execution_id, ttl_ms);
 
     if (insert_result.failed()) {
         LOG_ERROR("Failed to store execution {}: {}", result.execution_id, insert_result.error().message());
@@ -1459,6 +1789,23 @@ bool WorkflowEngine::executeLoopBody(
         LOG_INFO("  - Body node: {}", nid);
     }
 
+    // A completed body result - whether freshly executed or replayed from a
+    // pinned cache entry - that still carries a pause marker must not be
+    // replayed as Completed: the iteration state that would let it resume
+    // does not exist in this execution's stored data. Both the cache path and
+    // the fresh-execution path route through this single check so the
+    // failure and its message cannot drift apart between them.
+    auto rejectPauseInLoopBody = [](NodeExecutionResult& body_result) {
+        if (body_result.status == NodeStatus::Completed &&
+            body_result.output.contains("_pause")) {
+            body_result.status = NodeStatus::Failed;
+            body_result.error = "A node cannot pause inside a Loop body, because a "
+                                "paused loop iteration cannot be resumed. Collect the "
+                                "items first, approve once, then loop";
+            body_result.output.erase("_pause");
+        }
+    };
+
     // Execute body for each item
     for (size_t i = 0; i < ctx.items.size(); ++i) {
         ctx.current_index = i;
@@ -1617,21 +1964,49 @@ bool WorkflowEngine::executeLoopBody(
                     cached_result.finished_at = cached_result.started_at;
                     cached_result.from_cache = true;
 
+                    rejectPauseInLoopBody(cached_result);
+
                     iteration_results[body_node_id] = cached_result;
                     result.node_results[body_node_id + "_iter_" + std::to_string(i)] = cached_result;
 
+                    if (cached_result.status == NodeStatus::Completed &&
+                        cached_result.output.contains("_webhookResponse")) {
+                        if (!result.webhook_response.is_null()) {
+                            LOG_WARN("Node {} in a loop body overrides a webhook response already set",
+                                     body_node_id);
+                        }
+                        result.webhook_response = cached_result.output["_webhookResponse"];
+                    }
+
                     if (callback) {
-                        callback("loop.node.completed", {
+                        nlohmann::json cache_event_data = {
                             {"executionId", result.execution_id},
                             {"nodeId", body_node_id},
                             {"iteration", i},
-                            {"status", "completed"},
+                            {"status", nodeStatusToString(cached_result.status)},
                             {"output", cached_result.output},
                             {"fromCache", true}
-                        });
+                        };
+                        if (!cached_result.error.empty()) {
+                            cache_event_data["error"] = cached_result.error;
+                        }
+                        callback("loop.node." + nodeStatusToString(cached_result.status), cache_event_data);
                     }
 
                     LOG_INFO("Using pinned output for body node {}", body_node_id);
+
+                    // A pinned output rejected above by rejectPauseInLoopBody must go
+                    // through the same continueOnError decision a freshly executed
+                    // failure does, rather than silently moving on to the next body
+                    // node as a plain cache hit would.
+                    if (cached_result.status == NodeStatus::Failed) {
+                        iteration_failed = true;
+                        all_succeeded = false;
+                        if (!ctx.continue_on_error) {
+                            break;
+                        }
+                    }
+
                     continue;
                 }
 
@@ -1674,12 +2049,24 @@ bool WorkflowEngine::executeLoopBody(
             } else {
                 body_result = executeNode(evaluated_body_node, node_input, result.execution_id, workflow);
             }
+
+            rejectPauseInLoopBody(body_result);
+
             iteration_results[body_node_id] = body_result;
 
             // Store in main results with iteration suffix
             std::string result_key = body_node_id + "_iter_" + std::to_string(i);
             result.node_results[result_key] = body_result;
 
+            if (body_result.status == NodeStatus::Completed &&
+                body_result.output.contains("_webhookResponse")) {
+                if (!result.webhook_response.is_null()) {
+                    LOG_WARN("Node {} in a loop body overrides a webhook response already set",
+                             body_node_id);
+                }
+                result.webhook_response = body_result.output["_webhookResponse"];
+            }
+
             if (callback) {
                 nlohmann::json event_data = {
                     {"executionId", result.execution_id},

+ 16 - 1
src/runner/workflow_engine.hpp

@@ -81,7 +81,8 @@ enum class ExecutionStatus {
     Running,
     Completed,
     Failed,
-    Cancelled
+    Cancelled,
+    Waiting
 };
 
 std::string executionStatusToString(ExecutionStatus status);
@@ -101,6 +102,10 @@ struct ExecutionResult {
     std::string error;
     nlohmann::json final_output;
     nlohmann::json workflow_snapshot;  // Snapshot of workflow at execution time
+    nlohmann::json webhook_response;   // Set by a respond-to-webhook node, if any
+    std::string paused_node_id;        // Node that asked to pause, when Waiting
+    std::string pause_token;           // Must be presented to resume
+    int64_t pause_expires_at = 0;      // Milliseconds since the epoch, 0 for never
 
     nlohmann::json toJson() const;
 };
@@ -161,6 +166,16 @@ public:
                                             const nlohmann::json& trigger_data,
                                             ExecutionCallback callback = nullptr);
 
+    // Continue an execution that stopped at a pause marker. The workflow is
+    // rebuilt from the snapshot stored with the execution rather than from the
+    // workflow as it stands now, because it may have been edited while the
+    // approval was waiting and finishing a run against a different workflow
+    // than it started under is worse than refusing.
+    common::Result<ExecutionResult> resume(const std::string& execution_id,
+                                           const std::string& token,
+                                           const nlohmann::json& payload,
+                                           ExecutionCallback callback = nullptr);
+
     // Cancel execution
     void cancelExecution(const std::string& execution_id);
 

+ 144 - 2
src/webserver/api/execution_controller.cpp

@@ -1,6 +1,8 @@
 #include "execution_controller.hpp"
 #include "../webserver_service.hpp"
 #include "logging/logger.hpp"
+#include "proto/runner.grpc.pb.h"
+#include <grpcpp/grpcpp.h>
 
 namespace smartbotic::webserver::api {
 
@@ -8,9 +10,10 @@ ExecutionController::ExecutionController(storage::StorageClient& storage,
                                          auth::AuthMiddleware& middleware,
                                          WebSocketServer& ws_server,
                                          WorkflowScheduler& scheduler,
+                                         runners::LoadBalancer& load_balancer,
                                          FailureHandler on_failure)
     : storage_(storage), middleware_(middleware), ws_server_(ws_server),
-      scheduler_(scheduler), on_failure_(std::move(on_failure)) {}
+      scheduler_(scheduler), load_balancer_(load_balancer), on_failure_(std::move(on_failure)) {}
 
 void ExecutionController::registerRoutes(httplib::Server& server) {
     server.Get("/api/v1/executions", [this](const httplib::Request& req, httplib::Response& res) {
@@ -19,6 +22,15 @@ void ExecutionController::registerRoutes(httplib::Server& server) {
         });
     });
 
+    // Registered before the generic "/api/v1/executions/([^/]+)" GET route
+    // below, otherwise that regex captures "pending" as an execution id and
+    // this listing 404s.
+    server.Get("/api/v1/executions/pending", [this](const httplib::Request& req, httplib::Response& res) {
+        middleware_.requireAuth(req, res, [this](auto& req, auto& res, auto& ctx) {
+            listPending(req, res, ctx);
+        });
+    });
+
     server.Get(R"(/api/v1/executions/([^/]+))", [this](const httplib::Request& req, httplib::Response& res) {
         middleware_.requireAuth(req, res, [this](auto& req, auto& res, auto& ctx) {
             getExecution(req, res, ctx);
@@ -37,6 +49,12 @@ void ExecutionController::registerRoutes(httplib::Server& server) {
         });
     });
 
+    server.Post(R"(/api/v1/executions/([^/]+)/resume)", [this](const httplib::Request& req, httplib::Response& res) {
+        middleware_.requireAuth(req, res, [this](auto& req, auto& res, auto& ctx) {
+            resumeExecution(req, res, ctx);
+        });
+    });
+
     // Internal endpoint for runner to send execution events (no auth required from internal network)
     server.Post("/api/v1/internal/execution-event", [this](const httplib::Request& req, httplib::Response& res) {
         receiveExecutionEvent(req, res);
@@ -131,6 +149,125 @@ void ExecutionController::retryExecution(const httplib::Request& req, httplib::R
     sendJson(res, {{"success", true}, {"message", "Retry queued"}});
 }
 
+void ExecutionController::resumeExecution(const httplib::Request& req, httplib::Response& res,
+                                          const auth::AuthContext& ctx) {
+    std::string execution_id = req.matches[1];
+
+    nlohmann::json body;
+    try {
+        body = req.body.empty() ? nlohmann::json::object() : nlohmann::json::parse(req.body);
+    } catch (...) {
+        sendError(res, "Body is not valid JSON", 400);
+        return;
+    }
+
+    if (!body.is_object()) {
+        sendError(res, "Body must be a JSON object", 400);
+        return;
+    }
+
+    // Everything below reads body.value(...), which throws json::type_error
+    // when the key exists but holds the wrong type - that escaped the parse
+    // try/catch above and became a generic 500, the same defect fixed earlier
+    // on this branch at spec.value("status", 200). Each field is type-checked
+    // before it is read, and rejected with 400 naming which one is wrong.
+
+    if (body.contains("token") && !body["token"].is_string()) {
+        sendError(res, "token must be a string", 400);
+        return;
+    }
+    const std::string token = body.value("token", "");
+    if (token.empty()) {
+        sendError(res, "A token is required to answer a paused execution", 400);
+        return;
+    }
+
+    if (body.contains("data") && !body["data"].is_object()) {
+        sendError(res, "data must be an object", 400);
+        return;
+    }
+    nlohmann::json payload = body.value("data", nlohmann::json::object());
+
+    // This is an approval endpoint. Defaulting a missing or misspelled
+    // "approved" to true would silently approve whatever it is guarding -
+    // require it explicitly rather than assume yes.
+    if (!body.contains("approved")) {
+        sendError(res, "approved is required (true or false)", 400);
+        return;
+    }
+    if (!body["approved"].is_boolean()) {
+        sendError(res, "approved must be true or false", 400);
+        return;
+    }
+    payload["approved"] = body["approved"].get<bool>();
+    payload["answeredBy"] = ctx.user_id;
+
+    auto runner = load_balancer_.selectRunner();
+    if (!runner) {
+        sendError(res, "No runners available", 503);
+        return;
+    }
+
+    // Qualified with the leading "::" because smartbotic::webserver::grpc (a
+    // forward-declared namespace for the node sync / credential gRPC servers,
+    // pulled in via webserver_service.hpp) would otherwise shadow the real
+    // ::grpc namespace here.
+    auto channel = ::grpc::CreateChannel(runner->address, ::grpc::InsecureChannelCredentials());
+    auto stub = proto::RunnerService::NewStub(channel);
+
+    proto::ResumeExecutionRequest grpc_req;
+    grpc_req.set_execution_id(execution_id);
+    grpc_req.set_token(token);
+    grpc_req.set_payload(payload.dump());
+
+    proto::ExecuteWorkflowResponse grpc_res;
+    ::grpc::ClientContext grpc_ctx;
+    grpc_ctx.set_deadline(std::chrono::system_clock::now() + std::chrono::seconds(60));
+
+    auto status = stub->ResumeExecution(&grpc_ctx, grpc_req, &grpc_res);
+    if (!status.ok()) {
+        sendError(res, "Could not resume: " + status.error_message(), 400);
+        return;
+    }
+
+    ws_server_.broadcast("executions." + execution_id + ".resumed", {
+        {"executionId", execution_id},
+        {"answeredBy", ctx.user_id}
+    });
+
+    sendJson(res, {{"executionId", execution_id}, {"status", "resumed"}});
+}
+
+void ExecutionController::listPending(const httplib::Request& req, httplib::Response& res,
+                                      const auth::AuthContext& ctx) {
+    storage::QueryOptions options;
+    options.filters.push_back({"status", "waiting"});
+
+    auto result = storage_.query("executions", options);
+    if (result.failed()) {
+        sendError(res, "Could not list pending approvals", 500);
+        return;
+    }
+
+    // Deliberately no pause token here: listing is a weaker permission than
+    // answering, and anyone who could list every pending approval would
+    // otherwise be able to answer all of them. The token reaches an approver
+    // through the node's own output on the execution detail.
+    nlohmann::json pending = nlohmann::json::array();
+    for (const auto& record : result.value().documents) {
+        pending.push_back({
+            {"executionId", record.value("_id", "")},
+            {"workflowId", record.value("workflowId", "")},
+            {"workflowName", record.value("workflowName", "")},
+            {"pausedNodeId", record.value("pausedNodeId", "")},
+            {"pauseExpiresAt", record.value("pauseExpiresAt", static_cast<int64_t>(0))},
+            {"startedAt", record.value("startedAt", static_cast<int64_t>(0))}
+        });
+    }
+
+    sendJson(res, {{"pending", pending}, {"total", pending.size()}});
+}
+
 void ExecutionController::receiveExecutionEvent(const httplib::Request& req, httplib::Response& res) {
     LOG_DEBUG("Received execution event request: {}", req.body);
 
@@ -176,9 +313,14 @@ void ExecutionController::receiveExecutionEvent(const httplib::Request& req, htt
 
         // Release the scheduler slot held by this run. Scheduled dispatch is
         // fire-and-forget, so these events are the only signal that a run ended.
+        // A Waiting execution is no longer occupying the runner either - it is
+        // parked on a person, not running - so it releases the slot too.
+        // Without this, "Schedule -> ... -> Wait for Approval" holds its slot
+        // until the watchdog deadline expires and logs a misleading "exceeded
+        // its deadline" for a run that was correctly waiting for an answer.
         if (!execution_id.empty() &&
             (event_type == "execution.completed" || event_type == "execution.failed" ||
-             event_type == "execution.cancelled")) {
+             event_type == "execution.cancelled" || event_type == "execution.waiting")) {
             scheduler_.notifyExecutionFinished(workflow_id, execution_id);
         }
 

+ 9 - 0
src/webserver/api/execution_controller.hpp

@@ -2,10 +2,13 @@
 
 #include <httplib.h>
 #include <nlohmann/json.hpp>
+#include <optional>
+#include <string>
 #include "../auth/auth_middleware.hpp"
 #include "../websocket_server.hpp"
 #include "storage/storage_client.hpp"
 #include "../scheduler/workflow_scheduler.hpp"
+#include "../runners/load_balancer.hpp"
 
 namespace smartbotic::webserver::api {
 
@@ -20,6 +23,7 @@ public:
 
     ExecutionController(storage::StorageClient& storage, auth::AuthMiddleware& middleware,
                         WebSocketServer& ws_server, WorkflowScheduler& scheduler,
+                        runners::LoadBalancer& load_balancer,
                         FailureHandler on_failure = nullptr);
 
     void registerRoutes(httplib::Server& server);
@@ -33,6 +37,10 @@ private:
                          const auth::AuthContext& ctx);
     void retryExecution(const httplib::Request& req, httplib::Response& res,
                         const auth::AuthContext& ctx);
+    void resumeExecution(const httplib::Request& req, httplib::Response& res,
+                         const auth::AuthContext& ctx);
+    void listPending(const httplib::Request& req, httplib::Response& res,
+                     const auth::AuthContext& ctx);
     void receiveExecutionEvent(const httplib::Request& req, httplib::Response& res);
 
     void sendJson(httplib::Response& res, const nlohmann::json& data, int status = 200);
@@ -42,6 +50,7 @@ private:
     auth::AuthMiddleware& middleware_;
     WebSocketServer& ws_server_;
     WorkflowScheduler& scheduler_;
+    runners::LoadBalancer& load_balancer_;
     FailureHandler on_failure_;
 };
 

+ 84 - 2
src/webserver/api/webhook_controller.cpp

@@ -3,6 +3,7 @@
 #include "common/time_utils.hpp"
 #include "proto/runner.grpc.pb.h"
 #include <grpcpp/grpcpp.h>
+#include <cctype>
 #include <optional>
 
 namespace smartbotic::webserver::api {
@@ -162,6 +163,24 @@ void WebhookController::handleWebhook(const httplib::Request& req, httplib::Resp
         return;
     }
 
+    // A workflow that paused mid-run is not done: final_output is still null
+    // (the body would just be the literal string "null"), and labelling this
+    // ".completed" is false. Report it honestly - 202 Accepted, a status of
+    // "waiting", and a ".waiting" broadcast - and return before the completed
+    // path below, which assumes the run actually finished.
+    if (grpc_res.status() == proto::EXECUTION_STATUS_WAITING) {
+        ws_server_.broadcast("executions." + grpc_res.execution_id() + ".waiting", {
+            {"executionId", grpc_res.execution_id()},
+            {"workflowId", workflow_id},
+            {"status", "waiting"}
+        });
+        sendJson(res, {
+            {"executionId", grpc_res.execution_id()},
+            {"status", "waiting"}
+        }, 202);
+        return;
+    }
+
     // Broadcast execution event
     ws_server_.broadcast("executions." + grpc_res.execution_id() + ".completed", {
         {"executionId", grpc_res.execution_id()},
@@ -171,10 +190,73 @@ void WebhookController::handleWebhook(const httplib::Request& req, httplib::Resp
 
     // Return result
     if (!grpc_res.result().empty()) {
+        nlohmann::json parsed;
+        bool parsed_ok = true;
         try {
-            auto result = nlohmann::json::parse(grpc_res.result());
-            sendJson(res, result);
+            parsed = nlohmann::json::parse(grpc_res.result());
         } catch (...) {
+            parsed_ok = false;
+        }
+
+        if (parsed_ok && parsed.is_object() && parsed.contains("_webhookResponse")) {
+            const auto& spec = parsed["_webhookResponse"];
+
+            int status_code = 200;
+            if (spec.contains("status")) {
+                if (spec["status"].is_number_integer()) {
+                    status_code = spec["status"].get<int>();
+                } else {
+                    LOG_WARN("Webhook response status must be a whole number, got {}; sending 500",
+                             spec["status"].dump());
+                    status_code = 500;
+                }
+            }
+            if (status_code < 100 || status_code > 599) {
+                LOG_WARN("Webhook response asked for status {}, which is not a valid HTTP status; sending 500",
+                         status_code);
+                status_code = 500;
+            }
+
+            std::string content_type = "application/json";
+            if (spec.contains("headers") && spec["headers"].is_object()) {
+                for (auto it = spec["headers"].begin(); it != spec["headers"].end(); ++it) {
+                    if (!it.value().is_string()) {
+                        LOG_WARN("Webhook response header {} is not a string, skipping it", it.key());
+                        continue;
+                    }
+                    // Content-Type reaches httplib through set_content rather
+                    // than as a header, and setting both sends it twice.
+                    std::string name = it.key();
+                    std::string lowered;
+                    for (char c : name) {
+                        lowered += static_cast<char>(std::tolower(static_cast<unsigned char>(c)));
+                    }
+                    if (lowered == "content-type") {
+                        content_type = it.value().get<std::string>();
+                    } else {
+                        res.set_header(name.c_str(), it.value().get<std::string>().c_str());
+                    }
+                }
+            }
+
+            std::string body;
+            if (spec.contains("body")) {
+                const auto& value = spec["body"];
+                body = value.is_string() ? value.get<std::string>() : value.dump();
+            } else if (content_type.find("json") != std::string::npos) {
+                // An empty body is not valid JSON, and a strict client would
+                // fail to parse it against the Content-Type we are sending.
+                body = "{}";
+            }
+
+            res.status = status_code;
+            res.set_content(body, content_type.c_str());
+            return;
+        }
+
+        if (parsed_ok) {
+            sendJson(res, parsed);
+        } else {
             res.set_content(grpc_res.result(), "text/plain");
         }
     } else {

+ 1 - 1
src/webserver/webserver_service.cpp

@@ -216,7 +216,7 @@ void WebServerService::setupRoutes() {
     workflow_group_ctrl_->registerRoutes(server);
 
     execution_ctrl_ = std::make_unique<api::ExecutionController>(
-        *storage_, *auth_middleware_, *ws_server_, *scheduler_,
+        *storage_, *auth_middleware_, *ws_server_, *scheduler_, *load_balancer_,
         [this](const std::string& workflow_id, const std::string& execution_id,
                const std::string& error) {
             runErrorWorkflow(workflow_id, execution_id, error);

+ 32 - 0
tests/nodes/database-change.json

@@ -0,0 +1,32 @@
+{
+  "name": "verify-database-change",
+  "settings": {
+    "storagePermissions": {
+      "collections": { "watch_cursors": "read-write", "dbchange_test": "read-write" }
+    }
+  },
+  "nodes": [
+    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "setup", "name": "Setup", "type": "code", "position": {"x": 0, "y": 100},
+     "config": {"code": "smartbotic.storage.delete('watch_cursors', context.workflowId + ':first');\nsmartbotic.storage.delete('watch_cursors', context.workflowId + ':second');\nsmartbotic.storage.delete('dbchange_test', 'before-1');\nsmartbotic.storage.delete('dbchange_test', 'after-1');\nsmartbotic.storage.insert('dbchange_test', { name: 'before', at: Date.now() }, 'before-1');\nconst beforeDoc = smartbotic.storage.get('dbchange_test', 'before-1');\nconst raw = Number(beforeDoc.document._updated_at);\nconst stamp = raw > 1e15 ? Math.floor(raw / 1000000) : raw;\nsmartbotic.storage.insert('watch_cursors', { since: stamp, updatedAt: Date.now(), collection: 'dbchange_test' }, context.workflowId + ':second');\nreturn { ready: true };"}},
+    {"id": "first", "name": "First Run", "type": "database-change", "position": {"x": 0, "y": 200},
+     "config": {"collection": "dbchange_test", "emitOnFirstRun": false}},
+    {"id": "add", "name": "Insert One", "type": "code", "position": {"x": 0, "y": 300},
+     "config": {"code": "smartbotic.utils.sleep(1100);\nsmartbotic.storage.insert('dbchange_test', { name: 'after', at: Date.now() }, 'after-1');\nreturn { inserted: 'after-1' };"}},
+    {"id": "second", "name": "Second Run", "type": "database-change", "position": {"x": 0, "y": 400},
+     "config": {"collection": "dbchange_test", "emitOnFirstRun": false}},
+    {"id": "cleanup", "name": "Cleanup", "type": "code", "position": {"x": 0, "y": 500},
+     "config": {"code": "smartbotic.storage.delete('dbchange_test', 'before-1');\nsmartbotic.storage.delete('dbchange_test', 'after-1');\nreturn { cleaned: true };"}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "setup", "targetInput": "data"},
+    {"sourceNodeId": "setup", "sourceOutput": "main", "targetNodeId": "first", "targetInput": "data"},
+    {"sourceNodeId": "first", "sourceOutput": "main", "targetNodeId": "add", "targetInput": "data"},
+    {"sourceNodeId": "add", "sourceOutput": "main", "targetNodeId": "second", "targetInput": "data"},
+    {"sourceNodeId": "second", "sourceOutput": "main", "targetNodeId": "cleanup", "targetInput": "data"}
+  ],
+  "expect": {
+    "first": {"status": "completed", "output": {"count": 0, "isFirstRun": true}},
+    "second": {"status": "completed", "output": {"count": 1, "isFirstRun": false, "documents": [{"name": "after"}]}}
+  }
+}

+ 30 - 0
tests/nodes/file-watch.json

@@ -0,0 +1,30 @@
+{
+  "name": "verify-file-watch",
+  "settings": {
+    "storagePermissions": {
+      "collections": { "watch_cursors": "read-write" }
+    }
+  },
+  "nodes": [
+    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "setup", "name": "Setup", "type": "code", "position": {"x": 0, "y": 100},
+     "config": {"code": "const dir = '/tmp/sb-file-watch-test';\nconst listing = smartbotic.fs.readdir(dir);\nif (listing.success) {\n    for (const e of listing.entries) { smartbotic.fs.unlink(e.path); }\n}\nsmartbotic.fs.mkdir(dir);\nsmartbotic.fs.writeFile(dir + '/existing.txt', smartbotic.utils.base64Encode('old'));\nsmartbotic.storage.delete('watch_cursors', context.workflowId + ':first');\nsmartbotic.storage.delete('watch_cursors', context.workflowId + ':second');\nconst before = smartbotic.fs.readdir(dir);\nconst seen = {};\nfor (const e of before.entries) { seen[e.name] = e.modifiedAt + ':' + e.size; }\nsmartbotic.storage.insert('watch_cursors', { seen: seen, updatedAt: Date.now(), directory: dir }, context.workflowId + ':second');\nconst st = smartbotic.fs.stat(dir + '/existing.txt');\nconst statMtimeSane = st.mtime > 1600000000000 && st.mtime < Date.now() + 60000;\nreturn { dir: dir, statMtimeSane: statMtimeSane };"}},
+    {"id": "first", "name": "First Run", "type": "file-watch", "position": {"x": 0, "y": 200},
+     "config": {"directory": "/tmp/sb-file-watch-test", "emitOnFirstRun": false}},
+    {"id": "add", "name": "Add A File", "type": "code", "position": {"x": 0, "y": 300},
+     "config": {"code": "smartbotic.fs.writeFile('/tmp/sb-file-watch-test/fresh.txt', smartbotic.utils.base64Encode('new'));\nreturn { added: 'fresh.txt' };"}},
+    {"id": "second", "name": "Second Run", "type": "file-watch", "position": {"x": 0, "y": 400},
+     "config": {"directory": "/tmp/sb-file-watch-test", "emitOnFirstRun": false}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "setup", "targetInput": "data"},
+    {"sourceNodeId": "setup", "sourceOutput": "main", "targetNodeId": "first", "targetInput": "data"},
+    {"sourceNodeId": "first", "sourceOutput": "main", "targetNodeId": "add", "targetInput": "data"},
+    {"sourceNodeId": "add", "sourceOutput": "main", "targetNodeId": "second", "targetInput": "data"}
+  ],
+  "expect": {
+    "setup": {"status": "completed", "output": {"result": {"statMtimeSane": true}}},
+    "first": {"status": "completed", "output": {"count": 0, "isFirstRun": true}},
+    "second": {"status": "completed", "output": {"count": 1, "isFirstRun": false, "files": [{"name": "fresh.txt"}]}}
+  }
+}

+ 24 - 0
tests/nodes/fs-readdir.json

@@ -0,0 +1,24 @@
+{
+  "name": "verify-fs-readdir",
+  "nodes": [
+    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "n2", "name": "List", "type": "code", "position": {"x": 0, "y": 100},
+     "config": {"code": "const dir = '/tmp/sb-readdir-test';\nsmartbotic.fs.mkdir(dir);\nsmartbotic.fs.writeFile(dir + '/one.txt', smartbotic.utils.base64Encode('hello'));\nsmartbotic.fs.writeFile(dir + '/two.txt', smartbotic.utils.base64Encode('worldwide'));\nsmartbotic.fs.mkdir(dir + '/sub');\n\nconst listing = smartbotic.fs.readdir(dir);\nconst missing = smartbotic.fs.readdir('/tmp/sb-readdir-does-not-exist');\nconst notDir = smartbotic.fs.readdir(dir + '/one.txt');\n\nconst names = listing.entries.map(function (e) { return e.name; }).sort();\nconst one = listing.entries.filter(function (e) { return e.name === 'one.txt'; })[0];\nconst sub = listing.entries.filter(function (e) { return e.name === 'sub'; })[0];\n\nsmartbotic.fs.unlink(dir + '/one.txt');\nsmartbotic.fs.unlink(dir + '/two.txt');\n\nreturn {\n    ok: listing.success,\n    names: names,\n    oneSize: one.size,\n    oneHasMtime: one.modifiedAt > 1600000000000,\n    subIsDirectory: sub.isDirectory,\n    oneIsDirectory: one.isDirectory,\n    missingRejected: missing.success === false,\n    notDirRejected: notDir.success === false,\n    hasWorkflowId: typeof context.workflowId === 'string' && context.workflowId.length > 0\n};"}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "n2", "targetInput": "data"}
+  ],
+  "expect": {
+    "n2": {"status": "completed", "output": {"result": {
+      "ok": true,
+      "names": ["one.txt", "sub", "two.txt"],
+      "oneSize": 5,
+      "oneHasMtime": true,
+      "subIsDirectory": true,
+      "oneIsDirectory": false,
+      "missingRejected": true,
+      "notDirRejected": true,
+      "hasWorkflowId": true
+    }}}
+  }
+}

+ 14 - 0
tests/nodes/respond-to-webhook-errors.json

@@ -0,0 +1,14 @@
+{
+  "name": "verify-respond-to-webhook-errors",
+  "nodes": [
+    {"id": "n1", "name": "Start", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "n2", "name": "Respond", "type": "respond-to-webhook", "position": {"x": 0, "y": 100},
+     "config": {"status": 200, "bodySource": "input", "bodyField": "data.doesNotExist"}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "n2", "targetInput": "data"}
+  ],
+  "expect": {
+    "n2": {"status": "failed", "errorContains": "nothing found at \"data.doesNotExist\""}
+  }
+}

+ 24 - 0
tests/nodes/respond-to-webhook-loop.json

@@ -0,0 +1,24 @@
+{
+  "name": "verify-respond-to-webhook-loop",
+  "nodes": [
+    {"id": "n1", "name": "Webhook", "type": "post-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "items", "name": "Items", "type": "code", "position": {"x": 0, "y": 100},
+     "config": {"code": "return { items: ['alpha', 'beta', 'gamma'] };"}},
+    {"id": "loop", "name": "Loop", "type": "loop", "position": {"x": 0, "y": 200},
+     "config": {"inputField": "data.result.items"}},
+    {"id": "respond", "name": "Respond", "type": "respond-to-webhook", "position": {"x": 0, "y": 300},
+     "config": {"status": 200, "bodySource": "input", "bodyField": "item", "contentType": "text/plain"}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "items", "targetInput": "data"},
+    {"sourceNodeId": "items", "sourceOutput": "main", "targetNodeId": "loop", "targetInput": "data"},
+    {"sourceNodeId": "loop", "sourceOutput": "loop", "targetNodeId": "respond", "targetInput": "data"}
+  ],
+  "http": {
+    "method": "POST",
+    "path": "/webhook/{workflowId}",
+    "body": {"hello": "world"},
+    "expectStatus": 200,
+    "expectBodyContains": ["gamma"]
+  }
+}

+ 23 - 0
tests/nodes/respond-to-webhook.json

@@ -0,0 +1,23 @@
+{
+  "name": "verify-respond-to-webhook",
+  "nodes": [
+    {"id": "n1", "name": "Webhook", "type": "post-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "n2", "name": "Respond", "type": "respond-to-webhook", "position": {"x": 0, "y": 100},
+     "config": {"status": 201, "bodySource": "json", "body": "{\"created\": true, \"id\": \"abc\"}",
+                "contentType": "application/json", "headers": {"X-Smartbotic-Test": "tier2"}}},
+    {"id": "n3", "name": "After", "type": "code", "position": {"x": 0, "y": 200},
+     "config": {"code": "return { ranAfterResponding: true };"}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "n2", "targetInput": "data"},
+    {"sourceNodeId": "n2", "sourceOutput": "main", "targetNodeId": "n3", "targetInput": "data"}
+  ],
+  "http": {
+    "method": "POST",
+    "path": "/webhook/{workflowId}",
+    "body": {"hello": "world"},
+    "expectStatus": 201,
+    "expectHeaders": {"Content-Type": "application/json", "X-Smartbotic-Test": "tier2"},
+    "expectBodyContains": ["\"created\":true", "\"id\":\"abc\""]
+  }
+}

+ 23 - 0
tests/nodes/wait-for-approval-loop-chained.json

@@ -0,0 +1,23 @@
+{
+  "name": "verify-wait-for-approval-rejects-chained-loop",
+  "nodes": [
+    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "items", "name": "Items", "type": "code", "position": {"x": 0, "y": 100},
+     "config": {"code": "return { items: [1, 2] };"}},
+    {"id": "loop", "name": "Loop", "type": "loop", "position": {"x": 0, "y": 200},
+     "config": {"inputField": "data.result.items"}},
+    {"id": "pass", "name": "Pass Through", "type": "code", "position": {"x": 0, "y": 300},
+     "config": {"code": "return { item: input.item };"}},
+    {"id": "approve", "name": "Approve", "type": "wait-for-approval", "position": {"x": 0, "y": 400},
+     "config": {"reason": "Should never pause", "expiresIn": 1}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "items", "targetInput": "data"},
+    {"sourceNodeId": "items", "sourceOutput": "main", "targetNodeId": "loop", "targetInput": "data"},
+    {"sourceNodeId": "loop", "sourceOutput": "loop", "targetNodeId": "pass", "targetInput": "data"},
+    {"sourceNodeId": "pass", "sourceOutput": "main", "targetNodeId": "approve", "targetInput": "data"}
+  ],
+  "expect": {
+    "approve": {"status": "failed", "errorContains": "cannot pause inside a Loop body"}
+  }
+}

+ 20 - 0
tests/nodes/wait-for-approval-loop.json

@@ -0,0 +1,20 @@
+{
+  "name": "verify-wait-for-approval-rejects-loop",
+  "nodes": [
+    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "items", "name": "Items", "type": "code", "position": {"x": 0, "y": 100},
+     "config": {"code": "return { items: [1, 2] };"}},
+    {"id": "loop", "name": "Loop", "type": "loop", "position": {"x": 0, "y": 200},
+     "config": {"inputField": "data.result.items"}},
+    {"id": "approve", "name": "Approve", "type": "wait-for-approval", "position": {"x": 0, "y": 300},
+     "config": {"reason": "Should never pause", "expiresIn": 1}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "items", "targetInput": "data"},
+    {"sourceNodeId": "items", "sourceOutput": "main", "targetNodeId": "loop", "targetInput": "data"},
+    {"sourceNodeId": "loop", "sourceOutput": "loop", "targetNodeId": "approve", "targetInput": "data"}
+  ],
+  "expect": {
+    "approve": {"status": "failed", "errorContains": "cannot be used inside a Loop body"}
+  }
+}

+ 22 - 0
tests/nodes/wait-for-approval.json

@@ -0,0 +1,22 @@
+{
+  "name": "verify-wait-for-approval",
+  "nodes": [
+    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
+    {"id": "approve", "name": "Approve", "type": "wait-for-approval", "position": {"x": 0, "y": 100},
+     "config": {"reason": "Approve the test", "expiresIn": 1}},
+    {"id": "after", "name": "After Approval", "type": "code", "position": {"x": 0, "y": 200},
+     "config": {"code": "return { sawApproval: input.data.approved === true, note: input.data.note };"}}
+  ],
+  "connections": [
+    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "approve", "targetInput": "data"},
+    {"sourceNodeId": "approve", "sourceOutput": "main", "targetNodeId": "after", "targetInput": "data"}
+  ],
+  "resume": {
+    "waitForStatus": "waiting",
+    "tokenFrom": "approve",
+    "payload": {"approved": true, "data": {"note": "looks fine"}}
+  },
+  "expect": {
+    "after": {"status": "completed", "output": {"result": {"sawApproval": true, "note": "looks fine"}}}
+  }
+}