The database was moved off mulan because it and the SD.cpp REST API could not share that machine's memory. It runs on zeus as a container; mulan keeps the webserver and the runner, which reach it over the network.
Dockerfile - the image. It copies mulan's own smartbotic-database binary
and libsmartbotic-db-client.so rather than installing the package.That is deliberate. The .deb in the package repo is a later build linked
against Abseil 20240722; mulan runs one linked against 20260107. Installing
the package produced a container that could not start at all
(libabsl_log_internal_check_op.so.20240722: cannot open shared object file).
Copying the running binary onto the same Ubuntu release keeps the libraries
matched. Rebuild the image from mulan's files whenever the database is
upgraded there.
config.json - mulan's config with two changes:
bind_address is 0.0.0.0, because loopback inside a container reaches
nothing outside it.max_memory_mb is 8192 rather than 512. The old figure was what mulan could
spare and is the reason for the move; at 512 MB the container sat at 100% of
its budget and refused writes with "memory pressure emergency".docker run -d --name smartbotic-db --restart unless-stopped --memory 12g \
-v /data/smartbotic-db/data:/var/lib/smartbotic-database \
-v /data/smartbotic-db/etc/config.json:/etc/smartbotic-database/config.json:ro \
-p 9004:9004 smartbotic-database:2.8.1
Data lives on /data (1.6 TB free), not / (9 GB free).
The encryption key is inside the data directory (storage.key). Data and key
travel together or the data is unreadable - back them up together, and never
copy the data without it.
Port 9004 is published on all interfaces and the database speaks plaintext with no authentication - it was only ever reachable on mulan's loopback before. It now holds password hashes and encrypted credentials on a LAN-reachable port. The client supports TLS and a bearer token; a host firewall limiting 9004 to mulan is the smaller step. Neither is done yet.
The database was stopped first. A live copy of a WAL-backed store can be torn,
and recovery reported TrivialSuccess: 27288 docs in 93 collections with zero
WAL entries to replay - which is what a clean copy looks like.
zeus was 142 seconds behind mulan. That matters here because the database computes TTL expiries as absolute times, so a clock that disagrees with the machines writing to it shifts every retention deadline by the difference.
chrony is installed and enabled under runit, taking its time from the router at
192.168.2.1 - the same server DHCP advertises. dhcpcd's 50-ntp.conf hook writes
the lease's NTP servers into the config, and /etc/dhcpcd.conf sets
NTP_CONF=/etc/chrony.conf so it writes them where chrony will read them; by
default the hook picks /etc/ntp.conf, which nothing here reads.
Two things learned while setting it up, both visible only in the log:
chrony.conf by hand and letting DHCP write it in
produced Could not add source 192.168.2.1 on every start. DHCP supplies it,
so the manual line is gone.With pool pool.ntp.org also configured, chrony synced to a stratum-2 pool
server in preference to the router at stratum 3. The pool is gone: the router
is the source. If it is down the clock free-runs, which is the honest outcome
on a LAN rather than quietly taking somebody else's time.
chronyc tracking # Reference ID should be C0A80201 (192.168.2.1) chronyc sources # ^* marks the selected source
The container takes the host's clock, so nothing is configured inside it.