|
|
hace 1 mes | |
|---|---|---|
| .. | ||
| maintenance | hace 1 mes | |
| Dockerfile | hace 1 mes | |
| README.md | hace 1 mes | |
| backup-smartbotic-db.sh | hace 1 mes | |
| config.json | hace 1 mes | |
| deploy.sh | hace 1 mes | |
| docker-compose.yml | hace 1 mes | |
The database was moved off mulan because it and the SD.cpp REST API could not share that machine's memory. It runs on zeus as a container; mulan keeps the webserver and the runner, which reach it over the network.
docker-compose.yml - how the container runs. deploy.sh switches it to an
image tag that already exists on zeus, and refuses to move without asking.It does not build. The database is built on zeus. An earlier version of
that script rebuilt the image from mulan's binary, which would have
downgraded a running 2.10.0 to mulan's packaged 2.9.0 - mulan's package lags
what is built on zeus, and the assumption that mulan was the source of truth
was simply wrong. Running it also left smartbotic-database:current pointing
at the older image, which anyone using compose would then have picked up.
Dockerfile - kept for the case where the image has to be built from mulan's
binary. It copies mulan's own smartbotic-database and
libsmartbotic-db-client.so rather than installing the package.That is deliberate. The .deb in the package repo is a later build linked
against Abseil 20240722; mulan runs one linked against 20260107. Installing
the package produced a container that could not start at all
(libabsl_log_internal_check_op.so.20240722: cannot open shared object file).
Copying the running binary onto the same Ubuntu release keeps the libraries
matched. Rebuild the image from mulan's files whenever the database is
upgraded there.
config.json - mulan's config with two changes:
bind_address is 0.0.0.0, because loopback inside a container reaches
nothing outside it.max_memory_mb is 8192 rather than 512. The old figure was what mulan could
spare and is the reason for the move; at 512 MB the container sat at 100% of
its budget and refused writes with "memory pressure emergency".docker run -d --name smartbotic-db --restart unless-stopped --memory 12g \
-v /data/smartbotic-db/data:/var/lib/smartbotic-database \
-v /data/smartbotic-db/etc/config.json:/etc/smartbotic-database/config.json:ro \
-p 9004:9004 smartbotic-database:2.8.1
Data lives on /data (1.6 TB free), not / (9 GB free).
The encryption key is inside the data directory (storage.key). Data and key
travel together or the data is unreadable - back them up together, and never
copy the data without it.
Port 9004 is published on all interfaces and the database speaks plaintext with no authentication - it was only ever reachable on mulan's loopback before. It now holds password hashes and encrypted credentials on a LAN-reachable port. The client supports TLS and a bearer token; a host firewall limiting 9004 to mulan is the smaller step. Neither is done yet.
The database was stopped first. A live copy of a WAL-backed store can be torn,
and recovery reported TrivialSuccess: 27288 docs in 93 collections with zero
WAL entries to replay - which is what a clean copy looks like.
zeus was 142 seconds behind mulan. That matters here because the database computes TTL expiries as absolute times, so a clock that disagrees with the machines writing to it shifts every retention deadline by the difference.
chrony is installed and enabled under runit, taking its time from the router at
192.168.2.1 - the same server DHCP advertises. dhcpcd's 50-ntp.conf hook writes
the lease's NTP servers into the config, and /etc/dhcpcd.conf sets
NTP_CONF=/etc/chrony.conf so it writes them where chrony will read them; by
default the hook picks /etc/ntp.conf, which nothing here reads.
Two things learned while setting it up, both visible only in the log:
chrony.conf by hand and letting DHCP write it in
produced Could not add source 192.168.2.1 on every start. DHCP supplies it,
so the manual line is gone.With pool pool.ntp.org also configured, chrony synced to a stratum-2 pool
server in preference to the router at stratum 3. The pool is gone: the router
is the source. If it is down the clock free-runs, which is the honest outcome
on a LAN rather than quietly taking somebody else's time.
chronyc tracking # Reference ID should be C0A80201 (192.168.2.1) chronyc sources # ^* marks the selected source
The container takes the host's clock, so nothing is configured inside it.
backup-smartbotic-db.sh is installed at /usr/local/bin/backup-smartbotic-db
on zeus and runs nightly at 03:30 from /etc/cron.d/smartbotic-db-backup
(cronie, enabled under runit). Archives go to /data/backups/smartbotic-db,
14 kept - about 32 GB against 1.5 TB free. The log is
/var/log/smartbotic/backup.log.
The container is stopped for the copy and started again immediately after, so the database is down for the copy but not for the verification. A trap restarts it even if the script dies partway.
Three things the script insists on, each because the opposite failure is silent:
storage.key must be inside it. The key lives in the data directory and
the data is unreadable without it. Its absence would not be noticed until a
restore was already needed.Not a procedure on paper - this was run:
zstd -qdc /data/backups/smartbotic-db/<archive>.tar.zst | tar -C <dir> -xf -
docker run -d --name smartbotic-db-restoretest \
-v <dir>:/var/lib/smartbotic-database \
-v /data/smartbotic-db/etc/config.json:/etc/smartbotic-database/config.json:ro \
-p 9005:9004 smartbotic-database:2.8.1
It recovered 27,433 documents in 95 collections and served all 9 workflows by name, 7 credentials and 1 user. A 2 MB file downloaded byte-exact, which is what proves the encryption key came through - counting documents would not have.
Declared at webserver startup, idempotently, so a fresh install or a restored backup gets them without anyone remembering to. They live in the database, not in a config file.
Only what was measured to help - an index costs write throughput, so one that buys nothing is a cost paid for ever:
| index | before | after |
|---|---|---|
sessions.refreshToken |
24 ms | 1 ms |
executions.workflowId (miss) |
1124 ms | 1 ms |
workflows.projectId |
small today | scales with the workflow count |
users.username, users.email |
1 row | scales with the user count |
executions.workflowId still takes ~800 ms when it hits, and that is not the
index: cost scales with rows returned, about 30 ms per execution document,
because each carries every node's input and output. A miss answers in 1 ms,
which is what shows the lookup itself is fine. Making that query fast needs the
documents to stop carrying node output, not another index.
executions.startedAt was declared and dropped again. Measured, it changed
nothing - a range with limit 1 still took 300 ms and a sort by it 422 ms.
Ranges and sorts are not served from an index in this version. Executions are
the most-written collection here, so it was pure cost.
Off everywhere except workflows. Only two places read a version - the workflow
controller (the editor's picker, publish, restore) and the runner (loading the
published version to run) - and both are about workflows.
Everywhere else the history was written on every save and read by nothing, with
maxVersions 0 meaning unlimited. executions is the worst of it: 481 MB, and
each run writes its record two or three times.
Disabling stops new versions being recorded; history already written stays readable, so it is reversible and reclaims nothing by itself - the existing history directory is still there.
Every stored file has one. A file's TTL could only be set at upload before 2.8,
so anything written before that was immortal, and FileRecord carries no expiry
field - a file that already has one cannot be told from one that does not, so
the sweep sets one on all of them.
90 days was chosen: source photos, generated images and downloads, all either regenerable or already published elsewhere.
The documents that point at those files do not expire yet. No workflow or project declares a retention, so in 90 days those records will name files that have gone. Setting a retention on the project closes that - it covers documents, executions and files together.