|
|
@@ -207,6 +207,23 @@ caused nor worsened by v2.11.2, both bounded, neither yet investigated:
|
|
|
logged again after "Database service exited cleanly"). Harmless, noisy, worth
|
|
|
a look when either of the above is picked up.
|
|
|
|
|
|
+### 5d. ⚠ Container stop timeout on zeus (found 2026-08-15)
|
|
|
+
|
|
|
+Verifying the v2.11.3 unwind on zeus showed `docker stop` taking 10.5 s, i.e.
|
|
|
+hitting Docker's **default 10 s timeout and SIGKILLing** the process mid
|
|
|
+teardown - `Database service stopped` and `exited cleanly` were absent from the
|
|
|
+log. The final snapshot happened to land first, so nothing was lost, but that is
|
|
|
+luck: a slower snapshot would be killed before it renamed into place, leaving the
|
|
|
+next boot on WAL-only replay and read-only mode - the exact failure v2.11.2 was
|
|
|
+fixing, reintroduced by the container runtime rather than by the code.
|
|
|
+
|
|
|
+A full teardown is ~10 s today (~4 s gRPC transport drain + ~4-5 s component
|
|
|
+joins; see 5c), which does not fit in 10 s. The container is now run with
|
|
|
+`--stop-timeout 120`, matching the unit file's `TimeoutStopSec=300` in spirit.
|
|
|
+
|
|
|
+**Any future `docker run` for this service must carry `--stop-timeout`**, and any
|
|
|
+`docker stop` in a script wants `-t` to match. The deb/systemd path is unaffected.
|
|
|
+
|
|
|
### 6. Smaller known gaps
|
|
|
|
|
|
- Policy management has no dedicated RPCs; `_policies` is edited through the
|