2026-08-03-schedule-overlap-control-design.md 4.1 KB

Schedule trigger overlap control

Date: 2026-08-03 Status: Design - approved Owner: Ferenc Szontagh

1. Goal

Let a schedule trigger declare what happens when its scheduled time arrives while a previous run of the same workflow is still in flight: skip the tick, defer it until the run finishes, or allow a bounded number of concurrent runs.

2. Why the current guard does not work

WorkflowScheduler::checkAndExecute already skips a workflow whose running flag is set, but that flag is meaningless in practice:

  • webserver_service.cpp dispatches scheduled runs with request.set_wait_for_completion(false), so the gRPC call returns as soon as the runner accepts the job.
  • checkAndExecute clears running immediately after the callback returns.

So running covers the dispatch, a few milliseconds, not the execution. Two runs of the same workflow can overlap today with no protection.

A second defect: next_run is computed from a now captured before execution. When a run outlasts its interval, next_run is already in the past and the next tick fires immediately.

This matters concretely: run duration is dominated by hosted vision latency, measured between 12s and 117s per image, so a slow patch can push a run past a 10-minute interval.

3. Node configuration

Added to nodes/core/schedule-trigger.js:

Field Type Default Meaning
overlapPolicy skip | queue | allow skip Behaviour when a run is already active
maxConcurrent integer >= 1 1 Concurrent run ceiling, only meaningful for allow
maxRunMinutes integer >= 0 0 Treat a run as abandoned after this long. 0 derives it as 5x the interval.

Semantics:

  • skip - drop the tick, log it, schedule the next one normally.
  • queue - remember exactly one pending start and fire it when the active run finishes. Repeated missed ticks collapse into that single pending start, so a workflow that consistently overruns its interval cannot build an undrainable backlog.
  • allow - start whenever activeRuns < maxConcurrent.

4. Scheduler changes

ScheduledWorkflow gains overlap_policy, max_concurrent, max_run_minutes, an active_runs map of execution id to deadline, and a pending_start flag.

New public methods:

  • notifyExecutionStarted(workflow_id, execution_id) - called after a successful dispatch, records the run and its deadline.
  • notifyExecutionFinished(workflow_id, execution_id) - called on completion, failure or cancellation, releases the slot and, when a pending_start is set, makes the workflow immediately eligible.

checkAndExecute gains, before dispatch:

  1. Expire any active run past its deadline, with a warning. Without this a lost completion event would block a skip or queue workflow forever.
  2. Apply the policy against the live active_runs count.

next_run is recomputed from the time after dispatch.

5. Wiring

  • webserver_service.cpp reads the three new fields when registering a workflow and calls notifyExecutionStarted once the runner returns an execution id.
  • ExecutionController gains a WorkflowScheduler& and calls notifyExecutionFinished for execution.completed, execution.failed and execution.cancelled. It already receives these events with both executionId and workflowId.

6. Testing

  1. Unit-level: register a workflow with each policy, simulate started/finished notifications, assert dispatch decisions.
  2. Live: set a short interval on a workflow whose run outlasts it, and confirm skip drops ticks, queue fires exactly one deferred run, and allow with maxConcurrent 2 permits two.
  3. Stale guard: notify started without a matching finished, and confirm the slot is released after the deadline.

7. Out of scope

  • Cross-workflow scheduling fairness. checkAndExecute dispatches serially, but dispatch is now non-blocking, so one workflow no longer stalls others.
  • Persisting queue state across a webserver restart. Pending starts are in-memory and are lost on restart, which is acceptable for an interval trigger.