# Schedule trigger overlap control **Date:** 2026-08-03 **Status:** Design - approved **Owner:** Ferenc Szontagh ## 1. Goal Let a schedule trigger declare what happens when its scheduled time arrives while a previous run of the same workflow is still in flight: skip the tick, defer it until the run finishes, or allow a bounded number of concurrent runs. ## 2. Why the current guard does not work `WorkflowScheduler::checkAndExecute` already skips a workflow whose `running` flag is set, but that flag is meaningless in practice: - `webserver_service.cpp` dispatches scheduled runs with `request.set_wait_for_completion(false)`, so the gRPC call returns as soon as the runner accepts the job. - `checkAndExecute` clears `running` immediately after the callback returns. So `running` covers the dispatch, a few milliseconds, not the execution. Two runs of the same workflow can overlap today with no protection. A second defect: `next_run` is computed from a `now` captured before execution. When a run outlasts its interval, `next_run` is already in the past and the next tick fires immediately. This matters concretely: run duration is dominated by hosted vision latency, measured between 12s and 117s per image, so a slow patch can push a run past a 10-minute interval. ## 3. Node configuration Added to `nodes/core/schedule-trigger.js`: | Field | Type | Default | Meaning | | --- | --- | --- | --- | | `overlapPolicy` | `skip` \| `queue` \| `allow` | `skip` | Behaviour when a run is already active | | `maxConcurrent` | integer >= 1 | 1 | Concurrent run ceiling, only meaningful for `allow` | | `maxRunMinutes` | integer >= 0 | 0 | Treat a run as abandoned after this long. 0 derives it as 5x the interval. | Semantics: - **skip** - drop the tick, log it, schedule the next one normally. - **queue** - remember exactly one pending start and fire it when the active run finishes. Repeated missed ticks collapse into that single pending start, so a workflow that consistently overruns its interval cannot build an undrainable backlog. - **allow** - start whenever `activeRuns < maxConcurrent`. ## 4. Scheduler changes `ScheduledWorkflow` gains `overlap_policy`, `max_concurrent`, `max_run_minutes`, an `active_runs` map of execution id to deadline, and a `pending_start` flag. New public methods: - `notifyExecutionStarted(workflow_id, execution_id)` - called after a successful dispatch, records the run and its deadline. - `notifyExecutionFinished(workflow_id, execution_id)` - called on completion, failure or cancellation, releases the slot and, when a `pending_start` is set, makes the workflow immediately eligible. `checkAndExecute` gains, before dispatch: 1. Expire any active run past its deadline, with a warning. Without this a lost completion event would block a `skip` or `queue` workflow forever. 2. Apply the policy against the live `active_runs` count. `next_run` is recomputed from the time after dispatch. ## 5. Wiring - `webserver_service.cpp` reads the three new fields when registering a workflow and calls `notifyExecutionStarted` once the runner returns an execution id. - `ExecutionController` gains a `WorkflowScheduler&` and calls `notifyExecutionFinished` for `execution.completed`, `execution.failed` and `execution.cancelled`. It already receives these events with both `executionId` and `workflowId`. ## 6. Testing 1. Unit-level: register a workflow with each policy, simulate started/finished notifications, assert dispatch decisions. 2. Live: set a short interval on a workflow whose run outlasts it, and confirm `skip` drops ticks, `queue` fires exactly one deferred run, and `allow` with `maxConcurrent` 2 permits two. 3. Stale guard: notify started without a matching finished, and confirm the slot is released after the deadline. ## 7. Out of scope - Cross-workflow scheduling fairness. `checkAndExecute` dispatches serially, but dispatch is now non-blocking, so one workflow no longer stalls others. - Persisting queue state across a webserver restart. Pending starts are in-memory and are lost on restart, which is acceptable for an interval trigger.