2026-08-05-tier-2-triggers-design.md 10 KB

Tier 2 Triggers - Design

Date: 2026-08-05 Source: docs/node-roadmap.md, Tier 2

Goal

Ship all four Tier 2 entries: a real HTTP response for webhook-triggered workflows, two polling triggers, and a manual approval step that pauses an execution until a person answers it.

Three of the four are small. The fourth, manual approval, needs the engine to stop mid-workflow and resume later, which is the substance of this design.

What the roadmap got wrong

Two of the four entries rest on premises that do not hold. Both were checked against the code before this design was written.

"Respond to Webhook needs C++ because webhook_controller fires and forgets." It does not fire and forget. webhook_controller.cpp:150 sets wait_for_completion(true), waits up to 35 seconds, and returns the execution result as the HTTP body. What is actually missing is narrower: there is no way to set a status code (success is always 200), no way to set headers, and the body is whichever node happened to run last, because workflow_engine.cpp:674 assigns final_output from execution_order.back(). Appending a logging node to a workflow silently changes its API response.

"File Watch is pure JS over fs.stat." The fs API exposes writeFile, readFile, exists, mkdir, unlink and stat, and nothing that lists a directory. A directory cannot be enumerated with stat alone, so File Watch needs a small C++ addition first.

One thing the roadmap did not mention is already in place: workflow.proto:79 declares EXECUTION_STATUS_WAITING = 6, commented "Waiting for external trigger". The wire format has always had a slot for a paused execution. Only the C++ ExecutionStatus enum lacks the state, so nothing can enter it.

Part 1 - fs.readdir

A single addition to the fs API in script_engine.cpp, beside the existing stat:

const entries = smartbotic.fs.readdir('/var/spool/incoming');
// [{ name, path, size, modifiedAt, isDirectory }]

modifiedAt is milliseconds since the epoch, matching what stat already returns. The listing is not recursive: a node that wants a tree can walk it itself, and recursion in the C++ layer is an unbounded loop driven by whatever is on disk.

It adds no new reach into the filesystem, though not for the reason one might assume: the fs API applies no path restrictions at all today. readFile, writeFile, unlink and stat each take an arbitrary absolute path and act on it with the runner process's own permissions. That is consistent with the platform's existing position - the code node runs arbitrary JavaScript and process.exec runs arbitrary commands, so a node author is already trusted completely - but it is worth stating plainly rather than implying a sandbox that does not exist. readdir follows the same rule as its neighbours and adds nothing new.

Confining the fs API to configured roots would be a real improvement and is worth doing. It is deliberately not in this plan: it would change the behaviour of every existing node that touches the filesystem, and it deserves its own change rather than riding along inside a trigger batch.

Part 2 - Webhook responses

A new node declares the response explicitly rather than the engine inferring it from execution order:

return {
    _webhookResponse: {
        status: 201,
        headers: { 'Content-Type': 'application/json' },
        body: { id: created.id }
    }
};

ExecutionResult gains nlohmann::json webhook_response. When any node's output carries _webhookResponse, the engine records it there - any node, not the last one, so the response survives whatever runs after it. Last one wins if several set it, and that is worth a log line.

webhook_controller then honours it: status code, headers, and a body sent as JSON when it is an object or array and as text otherwise. When no node sets it, today's behaviour is unchanged - the last node's output as JSON - so existing webhook workflows keep working.

A status outside 100-599 is a node-side error and throws.

Part 3 - Resumable executions

The engine runs a workflow to completion in one pass. Manual approval needs it to stop, persist, and continue later, possibly on a different runner and possibly days later.

What already supports this

ExecutionResult carries node_results for every completed node and a workflow_snapshot of the workflow as it was at execution time, and storeExecution writes the whole thing to the executions collection. So the state needed to resume is already persisted for every execution. What is missing is a way to stop, a way to come back, and a status to sit in.

Pausing

A node signals a pause the way branching nodes signal a branch, by a marker in its output:

return {
    _pause: {
        token: 'a1b2c3',
        reason: 'Approve the refund of 240 EUR',
        expiresAt: 1786000000000,
        fields: [{ name: 'note', title: 'Note', type: 'string' }]
    }
};

When the engine sees _pause it stops the walk, sets the execution to Waiting, records the paused node id and the token on the result, and returns. Nothing downstream runs. The node's own output is stored without the marker, so a later reader sees the request rather than engine plumbing.

ExecutionStatus gains Waiting, mapping to the proto value that already exists.

The seven-day problem

storeExecution writes with a fixed seven-day TTL. A paused execution is not a log entry to be aged out - it is work someone still owes an answer to, and having it vanish is silent data loss of exactly the kind Tier 1 spent its review rounds eliminating.

So a paused execution is stored with a TTL derived from the node's own expiresIn, plus a grace period, rather than the fixed seven days. The node caps expiresIn at 30 days and says so, because an approval nobody answers should eventually stop occupying the collection rather than living forever.

An execution whose pause has expired is not resumable. Resuming one fails with a message naming the expiry rather than a generic not-found, since "this approval timed out" and "no such execution" want different reactions.

Resuming

RunnerService gains one RPC:

rpc ResumeExecution(ResumeExecutionRequest) returns (ExecuteWorkflowResponse);

carrying the execution id, the token, and a JSON payload from whoever answered.

The runner loads the stored execution, rebuilds the Workflow from workflow_snapshot rather than from the current stored workflow - the workflow may have been edited while the approval sat waiting, and finishing a run against a workflow different from the one it started under would be worse than failing. It seeds node_results from the stored execution, marks the paused node completed with the resume payload as its output, and continues the topological walk from that point. Nodes already in node_results are not re-run.

The token must match. Without it, knowing an execution id is enough to approve someone else's request.

REST, on the webserver:

  • POST /api/v1/executions/{id}/resume with {token, approved, data}
  • GET /api/v1/executions/pending listing executions in Waiting, so the editor and any external approver can find them

The resume records who answered and when, on the execution.

What is out of scope, deliberately

Pausing inside a Loop body. executeLoopBody keeps its per-iteration state in local variables that the stored execution does not describe, so resuming into the middle of iteration 7 of 900 cannot work without persisting loop state too. The approval node detects that it is running inside a loop and throws a message saying so, rather than pausing in a way that could not be resumed. Supporting it is a follow-up, and a larger one than it looks.

A webhook that waits for an approval. A paused execution cannot answer an HTTP request that is still open. A webhook-triggered workflow that pauses returns 202 with the execution id, and the caller polls or is notified. The webhook's 35-second deadline makes anything else fiction.

Part 4 - The nodes

File Node What it does
nodes/core/respond-to-webhook.js Respond to Webhook Sets status, headers and body for the HTTP response
nodes/triggers/database-change.js Database Change Polls a collection for documents new or changed since a stored cursor
nodes/triggers/file-watch.js File Watch Lists a directory and reports files new or changed since a stored cursor
nodes/core/wait-for-approval.js Wait for Approval Pauses the execution until someone answers

Both polling triggers keep their cursor in a storage collection keyed by workflow and node id, so two workflows watching the same place do not consume each other's changes. Both are paired with schedule-trigger for their cadence rather than holding a timer themselves, which is what the roadmap intended and what keeps them ordinary nodes.

Database Change tracks by a configurable timestamp field, defaulting to the _updatedAt the database maintains. File Watch tracks by modification time and size together, because a file rewritten within the same second at the same length is not something a workflow should be asked to notice.

Both throw on a first run with no cursor only if configured to, defaulting instead to recording the current state and reporting nothing - so adding one to a live workflow does not immediately fire for every document already present.

Conventions

The Tier 1 conventions in docs/nodes.md apply unchanged: throw rather than degrade, guard a configurable output name against reserved keys, do not re-interpolate config, use smartbotic.utils.getFieldValue for paths, and keep @description on one line.

Verification

Every piece gets a fixture in tests/nodes/ run by scripts/verify-node.py against the live stack, as in Tier 1. Three need more than the harness offers today:

  • Webhook responses need an actual HTTP call to a webhook URL, checking the status code and headers rather than a node's recorded output. The harness asserts on nodeExecutions, so this needs a small addition to it or a separate check in the task.
  • Resume needs a fixture that starts a workflow, sees it reach Waiting, calls the resume endpoint, and then asserts the downstream nodes completed. That is a second harness capability: assert, act, assert again.
  • File Watch needs files on disk, created and modified by the task itself.

These are called out because a fixture that cannot express them would otherwise quietly become a weaker assertion, which Tier 1 showed is worse than none.