fszontagh 868e48e700 fix: persist runners, and a deploy script that cannot downgrade the database 1 lună în urmă
..
maintenance c35f61d1a7 feat: document metadata on the row, ordered by the server - and a precision bug I caused 1 lună în urmă
Dockerfile 9dd48cae72 deploy: move the database to zeus, in Docker, and point mulan at it 1 lună în urmă
README.md 868e48e700 fix: persist runners, and a deploy script that cannot downgrade the database 1 lună în urmă
backup-smartbotic-db.sh 5e3722537a feat: nightly backups on zeus, a restore that has actually been done, and two fixes 1 lună în urmă
config.json 9dd48cae72 deploy: move the database to zeus, in Docker, and point mulan at it 1 lună în urmă
deploy.sh 868e48e700 fix: persist runners, and a deploy script that cannot downgrade the database 1 lună în urmă
docker-compose.yml 868e48e700 fix: persist runners, and a deploy script that cannot downgrade the database 1 lună în urmă

README.md

The database, on zeus, in Docker

The database was moved off mulan because it and the SD.cpp REST API could not share that machine's memory. It runs on zeus as a container; mulan keeps the webserver and the runner, which reach it over the network.

What is here

  • docker-compose.yml - how the container runs. deploy.sh switches it to an image tag that already exists on zeus, and refuses to move without asking.

It does not build. The database is built on zeus. An earlier version of that script rebuilt the image from mulan's binary, which would have downgraded a running 2.10.0 to mulan's packaged 2.9.0 - mulan's package lags what is built on zeus, and the assumption that mulan was the source of truth was simply wrong. Running it also left smartbotic-database:current pointing at the older image, which anyone using compose would then have picked up.

  • Dockerfile - kept for the case where the image has to be built from mulan's binary. It copies mulan's own smartbotic-database and libsmartbotic-db-client.so rather than installing the package.

That is deliberate. The .deb in the package repo is a later build linked against Abseil 20240722; mulan runs one linked against 20260107. Installing the package produced a container that could not start at all (libabsl_log_internal_check_op.so.20240722: cannot open shared object file). Copying the running binary onto the same Ubuntu release keeps the libraries matched. Rebuild the image from mulan's files whenever the database is upgraded there.

  • config.json - mulan's config with two changes:
    • bind_address is 0.0.0.0, because loopback inside a container reaches nothing outside it.
    • max_memory_mb is 8192 rather than 512. The old figure was what mulan could spare and is the reason for the move; at 512 MB the container sat at 100% of its budget and refused writes with "memory pressure emergency".

Running it

docker run -d --name smartbotic-db --restart unless-stopped --memory 12g \
  -v /data/smartbotic-db/data:/var/lib/smartbotic-database \
  -v /data/smartbotic-db/etc/config.json:/etc/smartbotic-database/config.json:ro \
  -p 9004:9004 smartbotic-database:2.8.1

Data lives on /data (1.6 TB free), not / (9 GB free).

The encryption key is inside the data directory (storage.key). Data and key travel together or the data is unreadable - back them up together, and never copy the data without it.

Exposure

Port 9004 is published on all interfaces and the database speaks plaintext with no authentication - it was only ever reachable on mulan's loopback before. It now holds password hashes and encrypted credentials on a LAN-reachable port. The client supports TLS and a bearer token; a host firewall limiting 9004 to mulan is the smaller step. Neither is done yet.

Moving the data

The database was stopped first. A live copy of a WAL-backed store can be torn, and recovery reported TrivialSuccess: 27288 docs in 93 collections with zero WAL entries to replay - which is what a clean copy looks like.

Time

zeus was 142 seconds behind mulan. That matters here because the database computes TTL expiries as absolute times, so a clock that disagrees with the machines writing to it shifts every retention deadline by the difference.

chrony is installed and enabled under runit, taking its time from the router at 192.168.2.1 - the same server DHCP advertises. dhcpcd's 50-ntp.conf hook writes the lease's NTP servers into the config, and /etc/dhcpcd.conf sets NTP_CONF=/etc/chrony.conf so it writes them where chrony will read them; by default the hook picks /etc/ntp.conf, which nothing here reads.

Two things learned while setting it up, both visible only in the log:

  • Naming the router in chrony.conf by hand and letting DHCP write it in produced Could not add source 192.168.2.1 on every start. DHCP supplies it, so the manual line is gone.
  • With pool pool.ntp.org also configured, chrony synced to a stratum-2 pool server in preference to the router at stratum 3. The pool is gone: the router is the source. If it is down the clock free-runs, which is the honest outcome on a LAN rather than quietly taking somebody else's time.

    chronyc tracking # Reference ID should be C0A80201 (192.168.2.1) chronyc sources # ^* marks the selected source

The container takes the host's clock, so nothing is configured inside it.

Backups

backup-smartbotic-db.sh is installed at /usr/local/bin/backup-smartbotic-db on zeus and runs nightly at 03:30 from /etc/cron.d/smartbotic-db-backup (cronie, enabled under runit). Archives go to /data/backups/smartbotic-db, 14 kept - about 32 GB against 1.5 TB free. The log is /var/log/smartbotic/backup.log.

The container is stopped for the copy and started again immediately after, so the database is down for the copy but not for the verification. A trap restarts it even if the script dies partway.

Three things the script insists on, each because the opposite failure is silent:

  • The archive is decompressed and listed before it is kept. A corrupt archive that is never read is worse than no archive, because it is believed in.
  • storage.key must be inside it. The key lives in the data directory and the data is unreadable without it. Its absence would not be noticed until a restore was already needed.
  • Old archives are pruned only after a good one exists, so a run of failures cannot age out the last working copy.

The restore, which has been done

Not a procedure on paper - this was run:

zstd -qdc /data/backups/smartbotic-db/<archive>.tar.zst | tar -C <dir> -xf -
docker run -d --name smartbotic-db-restoretest \
  -v <dir>:/var/lib/smartbotic-database \
  -v /data/smartbotic-db/etc/config.json:/etc/smartbotic-database/config.json:ro \
  -p 9005:9004 smartbotic-database:2.8.1

It recovered 27,433 documents in 95 collections and served all 9 workflows by name, 7 credentials and 1 user. A 2 MB file downloaded byte-exact, which is what proves the encryption key came through - counting documents would not have.

Indexes

Declared at webserver startup, idempotently, so a fresh install or a restored backup gets them without anyone remembering to. They live in the database, not in a config file.

Only what was measured to help - an index costs write throughput, so one that buys nothing is a cost paid for ever:

index before after
sessions.refreshToken 24 ms 1 ms
executions.workflowId (miss) 1124 ms 1 ms
workflows.projectId small today scales with the workflow count
users.username, users.email 1 row scales with the user count

executions.workflowId still takes ~800 ms when it hits, and that is not the index: cost scales with rows returned, about 30 ms per execution document, because each carries every node's input and output. A miss answers in 1 ms, which is what shows the lookup itself is fine. Making that query fast needs the documents to stop carrying node output, not another index.

executions.startedAt was declared and dropped again. Measured, it changed nothing - a range with limit 1 still took 300 ms and a sort by it 422 ms. Ranges and sorts are not served from an index in this version. Executions are the most-written collection here, so it was pure cost.

Versioning

Off everywhere except workflows. Only two places read a version - the workflow controller (the editor's picker, publish, restore) and the runner (loading the published version to run) - and both are about workflows.

Everywhere else the history was written on every save and read by nothing, with maxVersions 0 meaning unlimited. executions is the worst of it: 481 MB, and each run writes its record two or three times.

Disabling stops new versions being recorded; history already written stays readable, so it is reversible and reclaims nothing by itself - the existing history directory is still there.

File expiry

Every stored file has one. A file's TTL could only be set at upload before 2.8, so anything written before that was immortal, and FileRecord carries no expiry field - a file that already has one cannot be told from one that does not, so the sweep sets one on all of them.

90 days was chosen: source photos, generated images and downloads, all either regenerable or already published elsewhere.

The documents that point at those files do not expire yet. No workflow or project declares a retention, so in 90 days those records will name files that have gone. Setting a retention on the project closes that - it covers documents, executions and files together.