Przeglądaj źródła

fix: persist runners, and a deploy script that cannot downgrade the database

**persistRunner never persisted.** It called update() on a document that
was never inserted, so every write was dropped and the runners collection
stayed empty while runner-1 was registered and working. upsert() now, and
a failure is logged rather than discarded. The collection holds runner-1
for the first time.

That was invisible because a runner announces itself again within a
heartbeat, so nothing depended on the stored copy - which is exactly why
it survived: the only symptom was an empty collection nobody read.

**Two reported bugs were already fixed and I nearly built from the note.**
The workflows list shows created and modified dates, the owner and the
trigger types; the executions page has running and waiting filters,
duration, and an actions column. Both were carried as to-dos in a note
that had never been revisited. Verified in the browser instead of
trusting it. The underscore-collection write loss is fixed too - on 2.10.0
a collection named `_sbtest_under` inserts, reads back, finds and counts
exactly like any other, so there is nothing to report upstream.

**The deploy script.** Compose describes how the container runs, and
deploy.sh switches it to an image tag that already exists on zeus, asking
before it moves and leaving the previous tag in place to roll back to.

It does not build, and the reason is a mistake I made writing it. The
first version rebuilt from mulan's binary, on the assumption mulan was the
source of truth. Mulan's package is 2.9.0 and the container runs 2.10.0,
built on zeus - so running it would have downgraded the database. It also
retagged smartbotic-database:current to the older image, which anyone
using compose would then have picked up. Both undone; current points at
2.10.0 again.

executions.startedAt is declared again. On 2.9 it changed nothing; on 2.10
it makes no measurable difference either, but the collection is 1,500 rows
now rather than 10,000 after its orphans were removed, so the two
measurements are not comparable and neither is a verdict. Kept because the
field is perfectly selective and the planner declines an index it cannot
use.

Verified: the deploy path was run end to end - the container is
compose-managed, still 2.10.0, and the data is there. 70 passed, 0 failed.
fszontagh 1 miesiąc temu
rodzic
commit
868e48e700

+ 13 - 2
deploy/zeus/README.md

@@ -6,8 +6,19 @@ webserver and the runner, which reach it over the network.
 
 ## What is here
 
-- `Dockerfile` - the image. It copies mulan's own `smartbotic-database` binary
-  and `libsmartbotic-db-client.so` rather than installing the package.
+- `docker-compose.yml` - how the container runs. `deploy.sh` switches it to an
+  image tag that already exists on zeus, and refuses to move without asking.
+
+  **It does not build.** The database is built on zeus. An earlier version of
+  that script rebuilt the image from mulan's binary, which would have
+  *downgraded* a running 2.10.0 to mulan's packaged 2.9.0 - mulan's package lags
+  what is built on zeus, and the assumption that mulan was the source of truth
+  was simply wrong. Running it also left `smartbotic-database:current` pointing
+  at the older image, which anyone using compose would then have picked up.
+
+- `Dockerfile` - kept for the case where the image has to be built from mulan's
+  binary. It copies mulan's own `smartbotic-database` and
+  `libsmartbotic-db-client.so` rather than installing the package.
 
   That is deliberate. The `.deb` in the package repo is a later build linked
   against Abseil 20240722; mulan runs one linked against 20260107. Installing

+ 52 - 0
deploy/zeus/deploy.sh

@@ -0,0 +1,52 @@
+#!/bin/sh
+# Switch the database container to an image tag that already exists on zeus.
+#
+# It does NOT build. The database is built on zeus - the images there are tagged
+# smartbotic-database:<version>, alongside the smartbotic-db-build:<version> the
+# build itself uses. An earlier version of this script built from mulan's binary
+# instead, and would have DOWNGRADED a running 2.10.0 to mulan's packaged 2.9.0,
+# because mulan's package lags what is built on zeus. Hence the check below.
+#
+#   ./deploy.sh 2.10.0     switch to that image
+#   ./deploy.sh            show what is running and what is available
+set -eu
+
+ZEUS=${ZEUS:-zeus}
+WANT=${1:-}
+
+running=$(ssh "$ZEUS" 'docker exec smartbotic-db /usr/bin/smartbotic-database --version 2>/dev/null | head -1' || echo "not running")
+echo "running: $running"
+
+if [ -z "$WANT" ]; then
+    echo "available images:"
+    ssh "$ZEUS" 'docker image ls --format "  {{.Repository}}:{{.Tag}}" | grep -E "smartbotic-database:" | sort'
+    echo
+    echo "give a tag to switch to it, e.g. ./deploy.sh 2.10.0"
+    exit 0
+fi
+
+ssh "$ZEUS" "docker image inspect smartbotic-database:$WANT >/dev/null 2>&1" || {
+    echo "no image smartbotic-database:$WANT on $ZEUS - build it there first" >&2
+    exit 1
+}
+
+# Going backwards is possible and sometimes right, but never by accident.
+case "$running" in
+  *"$WANT"*) echo "already running $WANT" ;;
+  *) printf 'switch to %s? [y/N] ' "$WANT"; read -r answer
+     case "$answer" in y|Y) ;; *) echo "left alone"; exit 0 ;; esac ;;
+esac
+
+ssh "$ZEUS" "docker tag smartbotic-database:$WANT smartbotic-database:current && \
+  cd /data/smartbotic-db && \
+  docker rm -f smartbotic-db >/dev/null 2>&1 || true; \
+  docker compose up -d >/dev/null && sleep 8 && \
+  docker logs --tail 3 smartbotic-db 2>&1 | grep -iE 'recovery complete|error' || true"
+
+echo "now: $(ssh "$ZEUS" 'docker exec smartbotic-db /usr/bin/smartbotic-database --version 2>&1 | head -1')"
+echo
+echo "The previous image is still on zeus under its own version tag, so a"
+echo "rollback is ./deploy.sh <that version>."
+echo
+echo "mulan's webserver reconnects on its own within a few seconds. Until it"
+echo "does it will report zero workflows - that is the connection, not the data."

+ 23 - 0
deploy/zeus/docker-compose.yml

@@ -0,0 +1,23 @@
+# The database on zeus.
+#
+#   docker compose up -d          bring it up
+#   docker compose logs -f        watch it
+#
+# The image is built by upgrade.sh from mulan's own binary, not from the
+# package - see the note in README.md. Compose only runs what that produced.
+services:
+  smartbotic-db:
+    image: smartbotic-database:current
+    container_name: smartbotic-db
+    restart: unless-stopped
+    # A limit rather than a target. The database's own budget is
+    # storage.memory.max_memory_mb in config.json, currently 8 GB; this is the
+    # wall it cannot go through if that budget is ever wrong.
+    mem_limit: 12g
+    ports:
+      - "9004:9004"
+    volumes:
+      # The encryption key lives inside the data directory. Data and key are
+      # never separated - back them up together, move them together.
+      - /data/smartbotic-db/data:/var/lib/smartbotic-database
+      - /data/smartbotic-db/etc/config.json:/etc/smartbotic-database/config.json:ro

+ 14 - 1
src/webserver/runners/runner_registry.cpp

@@ -312,7 +312,20 @@ void RunnerRegistry::checkRunnerTimeouts() {
 void RunnerRegistry::persistRunner(const Runner& runner) {
     nlohmann::json doc = runner.toJson();
     doc.erase("_id");  // the id is the document key, not a body field
-    storage_.update("runners", runner.id, doc, 0, false);
+
+    // upsert, not update: a runner registering for the first time has no
+    // document to update, so update() failed and the write was dropped. The
+    // collection stayed empty for as long as this existed, and loadRunners()
+    // found nothing at startup - which was invisible because a runner
+    // re-registers on its own within a heartbeat.
+    //
+    // The failure is reported now rather than discarded. A registry that cannot
+    // persist still works, because the runners announce themselves, but nobody
+    // should have to deduce that from an empty collection.
+    auto written = storage_.upsert("runners", doc, runner.id);
+    if (written.failed()) {
+        LOG_WARN("Could not persist runner {}: {}", runner.id, written.error().message());
+    }
 }
 
 void RunnerRegistry::removeRunner(const std::string& runner_id) {

+ 10 - 4
src/webserver/webserver_service.cpp

@@ -57,12 +57,18 @@ WebServerService::WebServerService(const WebServerServiceConfig& config)
     //                           page load and workflows are written by hand
     //   users.username/email    login looks up by both
     //
-    // executions.startedAt was declared and dropped again: measured, it changed
-    // nothing. Ranges and sorts are not served from an index in 2.9 - a sorted
-    // single row still took 422 ms - and executions are the most written
-    // collection here, so it was pure cost.
+    // executions.startedAt was declared, dropped, and declared again. On 2.9 it
+    // changed nothing: ranges and sorts were not served from an index, and a
+    // sorted single row still took 422 ms over 10,000 rows. Re-measured on
+    // 2.10, which fixed that, it makes no measurable difference either - but the
+    // collection is 1,500 rows now rather than 10,000, having had its orphans
+    // removed, so the two measurements are not comparable and neither is a
+    // verdict. It is kept because the field is perfectly selective (every value
+    // distinct), the planner declines an index it cannot use, and the collection
+    // will grow again.
     for (const auto& [collection, field] : std::initializer_list<std::pair<const char*, const char*>>{
              {"executions", "workflowId"},
+             {"executions", "startedAt"},
              {"workflows", "projectId"},
              {"users", "username"},
              {"users", "email"},