Selaa lähdekoodia

docs: everything runs on zeus now, and no service may assume co-location

The webserver and runner both moved to zeus (Void Linux, runit) and the project
is retired on mulan, which keeps only the GPU work. CLAUDE.md described the
half-way state from an hour earlier.

The Relocating a service section is the part worth having: every address is an
env override, both run scripts state all of them even where the value is the
local default, and the table says what breaks for each - including that
ADVERTISE_ADDRESS is the one that fails silently, with the runner reporting
online while dispatches go nowhere.
fszontagh 3 viikkoa sitten
vanhempi
sitoutus
327b81e0e8
1 muutettua tiedostoa jossa 43 lisäystä ja 28 poistoa
  1. 43 28
      CLAUDE.md

+ 43 - 28
CLAUDE.md

@@ -37,17 +37,19 @@ cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release
 cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Debug -DENABLE_ASAN=ON
 ```
 
-After building a service, restart it to actually run the new binary. **The two
-services live on different machines and are managed by different init systems**
-(split 2026-09-03):
+After building a service, restart it to actually run the new binary. Everything
+runs on **zeus**, which is Void Linux and uses **runit, not systemd**:
 
-| | host | restart with |
-| --- | --- | --- |
-| webserver | zeus (Void Linux, runit) | `sudo sv restart smartbotic-webserver` |
-| runner | mulan (Debian, systemd) | `ssh mulan systemctl --user restart smartbotic-runner` |
+```bash
+sudo sv restart smartbotic-webserver smartbotic-runner
+sudo sv status  smartbotic-webserver smartbotic-runner
+tail -f /var/log/smartbotic-webserver/current
+tail -f /var/log/smartbotic-runner/current
+```
 
-There is no `systemctl` on zeus at all - it prints "command not found" rather
-than failing in any way that looks like a wrong host.
+There is no `systemctl` on zeus at all - it answers "command not found", which
+reads like a typo rather than the wrong init system. The `systemd/user/*.service`
+units in this repo are leftovers from when the services ran on mulan.
 
 ### Frontend (webui/)
 ```bash
@@ -67,31 +69,44 @@ share mulan's memory), so it is neither an APT package nor a system service on
 the host any more. It is already up; the in-repo services connect to it:
 
 ```bash
-# 1. Database daemon - already running on zeus as the Docker container
-#    `smartbotic-db`, published on 9004. Not started from this repo.
+# 1. Database daemon - Docker container `smartbotic-db` on 9004. Not from this repo.
 docker ps --filter name=smartbotic-db
 
-# 2. WebServer on zeus (HTTP 8090, WebSocket 8091, gRPC 9012 + 9013)
-sudo sv status smartbotic-webserver          # runit, not systemd
-tail -f /var/log/smartbotic-webserver/current
-
-# 3. Runner on mulan (gRPC port 9011)
-ssh mulan systemctl --user status smartbotic-runner
-ssh mulan journalctl --user -u smartbotic-runner -f
+# 2. WebServer (HTTP 8090, WebSocket 8091, gRPC 9012 + 9013)
+# 3. Runner (gRPC 9011)
+sudo sv status smartbotic-webserver smartbotic-runner
 ```
 
-Because the two are split, three addresses must point across the network. They
-are env overrides in a systemd drop-in on mulan
-(`~/.config/systemd/user/smartbotic-runner.service.d/zeus-webserver.conf`), NOT
-in `config/runner.json` - that file is tracked in git and shared by both hosts,
-so its `localhost` defaults cannot be edited per machine:
+### Relocating a service
+
+**No service may assume another is on the same host.** Every address a service
+needs is an env override, and both runit `run` scripts
+(`/etc/sv/smartbotic-*/run`) state all of them explicitly even where the value
+is the local default - so moving a service to another machine for load
+balancing or failover is an edit of one file, not a hunt through the code.
 
-| variable | value | what breaks without it |
+| variable | service | what breaks when it is wrong |
 | --- | --- | --- |
-| `WEBSERVER_ADDRESS` | `zeus.fsociety.hu:8090` | the runner never registers |
-| `NODE_SYNC_ADDRESS` | `zeus.fsociety.hu:9012` | no node updates reach it |
-| `CREDENTIAL_SERVICE_ADDRESS` | `zeus.fsociety.hu:9013` | credentials and workflow-control fail |
-| `ADVERTISE_ADDRESS` | `mulan.fsociety.hu:9011` | the runner registers as `localhost:9011`, reports **online**, and dispatches silently go to the webserver's own host |
+| `DATABASE_ADDRESS` | both | no persistence at all; fails loudly |
+| `WEBSERVER_ADDRESS` | runner | the runner never registers |
+| `NODE_SYNC_ADDRESS` | runner | no node updates reach it |
+| `CREDENTIAL_SERVICE_ADDRESS` | runner | credentials and workflow-control fail |
+| `ADVERTISE_ADDRESS` | runner | **fails silently** - the runner registers as `localhost:<port>`, reports online, and every dispatch goes to the webserver's own host |
+| `NODES_PATH` | runner | no built-in nodes |
+| `WEBUI_PATH` | webserver | the UI 404s, the API still works |
+
+`config/*.json` is tracked in git and its defaults are all `localhost`, so it
+must not be edited per host. `config/*.local.json` is gitignored but nothing
+reads it - the mechanism is `${VAR:default}` substitution.
+
+**Tests must not assume co-location either.** A case whose node fetches the API
+writes `{{API}}`, which the harness substitutes from `SMARTBOTIC_RUNNER_API`.
+Three cases hardcoded `http://localhost:8090` and broke the day the runner moved.
+
+SD.cpp stays on mulan (`https://mulan:8077`) because that is where the GPU is.
+zeus needs its self-signed certificate in **two** stores - the trusted-certificate
+store for host `mulan` (for the runner) and the system CA store (for the test
+harness's Python probe).
 
 Both the database address and the project namespace are configurable everywhere,
 via `config/webserver.json` and `config/runner.json`: