Sfoglia il codice sorgente

feat: nightly backups on zeus, a restore that has actually been done, and two fixes

**Backups.** Nothing backed up the database. It runs nightly at 03:30,
keeps 14, and stops the container for the copy - a live copy of a
WAL-backed store can be torn, and a backup nobody can restore is worse
than none because it is believed in.

The script refuses to keep an archive it cannot decompress and list, and
refuses one without storage.key in it. The key lives inside the data
directory, so a backup without it restores nothing; its absence would not
be noticed until a restore was already needed. Old archives are pruned
only once a good one exists, so a run of failures cannot age out the last
working copy.

The restore is not a procedure on paper. The newest archive was extracted
into a second container: 27,433 documents in 95 collections, all 9
workflows by name, 7 credentials, and a 2 MB file downloaded byte-exact -
which is what proves the key came through. Counting documents would not
have.

**A failed record write no longer strands an execution.** One had sat in
"running" for hours: the final write failed with "Document with ID
already exists", the error was logged and swallowed, and the record kept
the only status it had - running - while the run was long over and the
error workflow had already been told. On a duplicate id it now writes the
final status with an update instead. Any other failure is still a
failure.

**The blog rating comes from the picture.** is_adult was 1 for every post
because the pipeline only redraws NSFW sources. The vision model already
judges each image, so nsfwLevel and eroticism_level decide it, and a
missing or unreadable analysis still rates 1 - the blog requires the
field, and an unrated adult picture on a site with an 18+ gate is the
harmful direction.

Verified against four shapes of vision output: clean 0, partial 1,
explicit 1, low eroticism 0, absent 1. The first attempt at that test was
wrong rather than the expression - a code node wraps its output in
"result" while ollama-chat returns "json" at the top level, so every case
hit the fallback and returned 1.
fszontagh 1 mese fa
parent
commit
5e3722537a
3 ha cambiato i file con 129 aggiunte e 3 eliminazioni
  1. 36 0
      deploy/zeus/README.md
  2. 66 0
      deploy/zeus/backup-smartbotic-db.sh
  3. 27 3
      src/runner/workflow_engine.cpp

+ 36 - 0
deploy/zeus/README.md

@@ -77,3 +77,39 @@ Two things learned while setting it up, both visible only in the log:
     chronyc sources      # ^* marks the selected source
 
 The container takes the host's clock, so nothing is configured inside it.
+
+## Backups
+
+`backup-smartbotic-db.sh` is installed at `/usr/local/bin/backup-smartbotic-db`
+on zeus and runs nightly at 03:30 from `/etc/cron.d/smartbotic-db-backup`
+(cronie, enabled under runit). Archives go to `/data/backups/smartbotic-db`,
+14 kept - about 32 GB against 1.5 TB free. The log is
+`/var/log/smartbotic/backup.log`.
+
+The container is stopped for the copy and started again immediately after, so
+the database is down for the copy but not for the verification. A trap restarts
+it even if the script dies partway.
+
+Three things the script insists on, each because the opposite failure is silent:
+
+- **The archive is decompressed and listed before it is kept.** A corrupt
+  archive that is never read is worse than no archive, because it is believed in.
+- **`storage.key` must be inside it.** The key lives in the data directory and
+  the data is unreadable without it. Its absence would not be noticed until a
+  restore was already needed.
+- **Old archives are pruned only after a good one exists**, so a run of failures
+  cannot age out the last working copy.
+
+### The restore, which has been done
+
+Not a procedure on paper - this was run:
+
+    zstd -qdc /data/backups/smartbotic-db/<archive>.tar.zst | tar -C <dir> -xf -
+    docker run -d --name smartbotic-db-restoretest \
+      -v <dir>:/var/lib/smartbotic-database \
+      -v /data/smartbotic-db/etc/config.json:/etc/smartbotic-database/config.json:ro \
+      -p 9005:9004 smartbotic-database:2.8.1
+
+It recovered 27,433 documents in 95 collections and served all 9 workflows by
+name, 7 credentials and 1 user. A 2 MB file downloaded byte-exact, which is what
+proves the encryption key came through - counting documents would not have.

+ 66 - 0
deploy/zeus/backup-smartbotic-db.sh

@@ -0,0 +1,66 @@
+#!/bin/sh
+# Back up the SmartBotic database running in Docker on zeus.
+#
+# The database is stopped for the copy. A live copy of a WAL-backed store can be
+# torn, and a backup nobody can restore is worse than no backup because it is
+# believed in. The stop is seconds; the container comes back whatever happens,
+# including if this script dies partway.
+#
+# The encryption key lives inside the data directory. It is copied with
+# everything else and must never be separated from it - data without the key is
+# unreadable, and so is a backup without it.
+set -eu
+
+CONTAINER=smartbotic-db
+DATA=/data/smartbotic-db/data
+DEST=/data/backups/smartbotic-db
+KEEP=14
+STAMP=$(date +%Y%m%d-%H%M%S)
+ARCHIVE="$DEST/smartbotic-db-$STAMP.tar.zst"
+
+mkdir -p "$DEST"
+
+was_running=0
+if [ "$(docker inspect -f '{{.State.Running}}' "$CONTAINER" 2>/dev/null)" = "true" ]; then
+    was_running=1
+fi
+
+restart_if_needed() {
+    if [ "$was_running" = "1" ]; then
+        docker start "$CONTAINER" >/dev/null 2>&1 || true
+    fi
+}
+trap restart_if_needed EXIT INT TERM
+
+if [ "$was_running" = "1" ]; then
+    docker stop "$CONTAINER" >/dev/null
+fi
+
+tar -C "$DATA" -cf - . | zstd -q -3 -o "$ARCHIVE"
+
+# Started again before the verification, so the database is down only for the
+# copy itself rather than for the reading back as well.
+restart_if_needed
+trap - EXIT INT TERM
+
+# A backup is not a backup until it has been read back. This checks the archive
+# is intact and that the encryption key is inside it - the one file whose
+# absence would not be noticed until a restore was already needed.
+if ! zstd -qt "$ARCHIVE"; then
+    echo "BACKUP FAILED: $ARCHIVE does not decompress" >&2
+    rm -f "$ARCHIVE"
+    exit 1
+fi
+if ! zstd -qdc "$ARCHIVE" | tar -t | grep -q '^\./storage\.key$'; then
+    echo "BACKUP FAILED: $ARCHIVE has no storage.key - the data would be unreadable" >&2
+    rm -f "$ARCHIVE"
+    exit 1
+fi
+
+# Only prune once a good one exists, so a run of failures cannot age out the
+# last working copy.
+ls -1t "$DEST"/smartbotic-db-*.tar.zst 2>/dev/null | tail -n +$((KEEP + 1)) | while read -r old; do
+    rm -f "$old"
+done
+
+echo "$(date -Is) ok $ARCHIVE $(du -h "$ARCHIVE" | cut -f1) ($(ls -1 "$DEST"/smartbotic-db-*.tar.zst | wc -l) kept)"

+ 27 - 3
src/runner/workflow_engine.cpp

@@ -2393,11 +2393,35 @@ void WorkflowEngine::storeExecution(const ExecutionResult& result) {
 
     auto insert_result = storage_.insert("executions", result.toJson(), result.execution_id, ttl_ms);
 
-    if (insert_result.failed()) {
-        LOG_ERROR("Failed to store execution {}: {}", result.execution_id, insert_result.error().message());
-    } else {
+    if (insert_result.ok()) {
         LOG_DEBUG("Stored execution {} successfully", result.execution_id);
+        return;
     }
+
+    // The record is already there after all - the get above did not see it.
+    //
+    // Then this write is the one that matters: it carries the run's final
+    // status. Logging the failure and returning left the record on whatever it
+    // said last, which is "running", for ever - the run had finished, the error
+    // workflow had already been told, and nothing would ever correct it. One
+    // execution sat that way for hours and had to be cleared by hand.
+    //
+    // Only for a duplicate id. Any other failure is a real failure and is
+    // reported as one.
+    const std::string message = insert_result.error().message();
+    if (message.find("already exists") != std::string::npos) {
+        auto updated = storage_.update("executions", result.execution_id, result.toJson());
+        if (updated.ok()) {
+            LOG_WARN("Execution {} existed although it did not read back; its final status "
+                     "was written with an update instead", result.execution_id);
+            return;
+        }
+        LOG_ERROR("Execution {} exists but could not be updated either: {}",
+                  result.execution_id, updated.error().message());
+        return;
+    }
+
+    LOG_ERROR("Failed to store execution {}: {}", result.execution_id, message);
 }
 
 std::vector<std::string> WorkflowEngine::findLoopBodyNodes(