Pārlūkot izejas kodu

fix: count a node the run swallowed, so a halted-and-ignored run is visible

toleratedErrorCount only counted a node that swallowed its own failure -
status Completed with a non-empty error, the skipOnError shape. A node that
failed outright while the run finished anyway was not counted, and that is the
case that matters most.

The worst instance: a stop-and-error in error mode inside a loop body throws,
the loop's Continue On Error tolerates it by default, and the run was reported
as completed, with no error and a tolerated count of zero. Nothing in the
listing said a node whose entire purpose is to end the run had fired and been
ignored. Two live workflows are in exactly that state today.

A node that failed while the execution completed means something tolerated it -
a loop's setting or the workflow's - which is what this count is for.

Behaviour is unchanged; only the count is. Characterisation over eight loop
shapes shows one difference, that count, and in one case it moved 0 -> 3 and
was right to: the probe's own code node could not see "loop", so every
iteration had been failing and the old counter hid it. Node suite 92/0.

This does not fix stop-and-error itself being disarmed inside a loop, which
needs a decision about how to spell "skip this item" as against "fail the
whole run" - both are written mode:error today and resolved by a setting on a
different node.
fszontagh 1 mēnesi atpakaļ
vecāks
revīzija
ee81122cca
1 mainītis faili ar 12 papildinājumiem un 0 dzēšanām
  1. 12 0
      src/runner/workflow_engine.cpp

+ 12 - 0
src/runner/workflow_engine.cpp

@@ -292,6 +292,18 @@ nlohmann::json ExecutionResult::toJson() const {
         nr["retryCount"] = result.retry_count;
         if (result.status == NodeStatus::Completed && !result.error.empty()) {
             ++tolerated_error_count;
+        } else if (result.status == NodeStatus::Failed && status == ExecutionStatus::Completed) {
+            // A node that failed outright while the run still finished. Something
+            // swallowed it - a loop's Continue On Error, or the workflow-level
+            // setting - which is what "tolerated" means, so it belongs in this
+            // count as much as a node that swallowed its own failure does.
+            //
+            // It was not counted, and that left the worst case invisible: a
+            // stop-and-error inside a loop body fails, the loop tolerates it by
+            // default, and the run was reported as completed with no error and a
+            // tolerated count of zero. Nothing in the listing said a node whose
+            // entire purpose is to end the run had fired and been ignored.
+            ++tolerated_error_count;
         }
         // Absent for a node that did not run inside a loop. Present for one
         // that did - even a disabled one, recorded once for the whole loop