# Smartbotic Database ## Build ```bash cmake -B build -G Ninja cmake --build build -j$(nproc) ``` ## Test ```bash cmake --build build -j$(nproc) && ./build/tests/test_vector_storage ``` ## Architecture - `proto/database.proto` — gRPC service and message definitions - `service/src/` — Database server implementation - `memory_store.hpp/.cpp` — In-memory document + vector storage - `document.hpp` — Document and CollectionOptions types - `database_grpc_impl.cpp` — gRPC request handlers - `persistence/` — WAL, snapshots, persistence manager - `migrations/` — Migration runner - `client/` — C++ client library (`smartbotic-db-client`) - `tests/` — Integration tests ## Key Features - JSON document store with collections, version history, field-level encryption - gRPC API with replication, events, file storage, migrations - **Vector Storage** — Embedding vectors stored alongside documents with SIMD-accelerated (AVX2/SSE4.1) cosine similarity search via `SimilaritySearch` RPC. Vectors use `_vector` document field, stored in parallel float arrays, persisted through WAL (`VEC_PUT`/`VEC_DELETE`) and snapshots (v3 format). Collection `vector_dimension` option is immutable after creation. - **Atomic Partial Updates** — `PatchDocument` RPC merges fields into existing documents atomically (server-side, within collection lock). Client `patch()` method. `update()` uses automatic optimistic locking with retry. See `docs/integration-guide.md` for concurrency patterns. - **Views** — Read-only named projections over collections. Include/exclude field lists (dot-notation for nested paths), baked-in filters with all operators (EQ/NE/GT/GTE/LT/LTE/IN/CONTAINS/EXISTS/REGEX/SEARCH) AND-merged with caller filters, optional default sort that caller can override. Stored in `_views` system collection, created via `create_view` migration op or `createView` RPC. Writes on view names are rejected. - **Per-collection Timestamp Precision** — collection-level `timestamp_precision` config ("ms" default, "ns" for rapid-write collections). Stamps `_created_at`/`_updated_at` at the configured resolution, cached hot-path lookup. `configureCollection` RPC flips the config atomically; `migrateCollectionTimestamps` RPC converts existing data during maintenance windows (idempotent via 10^15 threshold). - **Durable Snapshots + Tiered Recovery** — atomic snapshot writer (`.tmp` + fsync + rename + post-write verification), loader fallback chain across snapshots, MySQL-style recovery modes (`normal` / `snapshot_fallback` / `wal_only` / `best_effort` / `force_empty`) selectable via `--recovery-mode` flag or `recovery.mode` config. Non-trivial recovery (fallback, WAL-only, forced empty) automatically enters **read-only mode** — operator must `smartbotic-db-cli unlock` or pass `--force-readwrite` to accept writes. Exit code `10` when recovery is refused. - **gRPC Concurrency Cap** — bounded inbound memory via gRPC `ResourceQuota` + per-RPC-type stream limits (`Subscribe`, `UploadFile`, `DownloadFile`). Configurable under `storage.grpc`. Excess streams get `RESOURCE_EXHAUSTED` with operator log. - **Drop-in config (conf.d)** — `/etc/smartbotic-database/conf.d/*.json` deep-merges over `config.json` using RFC 7396 semantics (arrays replace, objects merge). Lets consumer debs ship tuned settings without touching the upstream conffile. Strict on syntax errors (exit 11), lenient on unknown keys. - **Chunked pressure-aware eviction** — Eviction spreads work over multiple ticks with a default `eviction_chunk_size=1000`, `eviction_chunk_pause_ms=50`, `hot_write_floor_ms=30000`. Four pressure levels (normal/soft/hard/emergency) gate behavior: soft = trickle, hard = aggressive + `MEMORY_PRESSURE_HIGH` event, emergency = admission control rejects writes with `RESOURCE_EXHAUSTED`. Per-collection `memory_priority` (low/normal/high) biases selection. Quiesce skips docs with in-flight writes. WAL-fallback no longer holds per-collection mutex during disk I/O. - **GetMemoryStats RPC + client retry** — `Client::getMemoryStats()` returns per-collection stats + pressure level. Client library auto-retries write RPCs on `DEADLINE_EXCEEDED`/`RESOURCE_EXHAUSTED`/`UNAVAILABLE` with exponential backoff + jitter (`writeRetries=3`, `writeRetryBackoffMs=100`). Reads are not auto-retried. - **Auto-backup before deb upgrade** — `apt upgrade smartbotic-database` runs a `preinst` hook that snapshots `/var/lib/smartbotic-database/` to `/var/backups/smartbotic-database/pre-upgrade--from-/` via `cp -a` after the old `prerm` stops the service. Retention `3`, 1.2× free-space precheck on the backup volume (aborts the upgrade if short). Opt out with `SMARTBOTIC_DB_SKIP_BACKUP=1 apt upgrade`. Tunables: `SMARTBOTIC_DB_BACKUP_RETENTION`, `SMARTBOTIC_DB_BACKUP_ROOT`. - **yyjson hot-path parsing (v1.10.0+)** — every JSON parse on WAL replay, snapshot deserialize, history reads, gRPC request bodies, and replication apply goes through `smartbotic::db::parse_to_nlohmann` (yyjson-backed) instead of `nlohmann::json::parse`. Measured 2.80× speedup on Zoe-shape corpora (136ms→49ms on 10k docs). In-memory representation stays `nlohmann::json` until Phase B. Wire and on-disk JSON bytes unchanged. New runtime dep: `libyyjson0`. Bench: `./build/tests/bench_json_parse `. - **Binary-backed Document storage (v1.11.0+)** — `Document::data` is now a `smartbotic::db::doc_binary::Doc` (yyjson `mut_doc` wrapper) with a lazy `nlohmann::json` view cache. New accessor surface: `doc.field(name)` is the O(field-count) fast-path that reads one top-level field without materialising the full tree (the dominant read pattern), `doc.data()` is the lazy full-tree accessor, `doc.set_data(j)` re-encodes from an nlohmann tree. Wire and on-disk JSON bytes unchanged. **Honest note:** per-doc heap footprint is ~1.5% smaller than v1.10 (1618→1594 bytes on Zoe-shape) — the win is field-access CPU + Phase C scaffolding, not RSS. See `docs/superpowers/specs/2026-05-15-binary-doc-format-decision.md` post-implementation amendment for the bench numbers. Bench: `./build/tests/bench_doc_memory --mode= [N]`. - **v2.7.0 row- and column-level security.** Per-project access policy, **off by default** so upgrading is inert. Design + the approved decisions: `docs/superpowers/specs/2026-08-08-rls-cls-design.md`. Before this there was **no access control beyond authentication** — every key holder had identical, complete access, and views were never a boundary (the guide claimed they were; corrected). **Principal = named API key.** `auth.keys[]` entries may be `{"name","key"}`; a bare string still authenticates as the reserved `unnamed`; no token is the reserved, grantable `anonymous` (the default 127.0.0.1 listener has no auth and the CLI rides on it, so local access should be an explicit grant). `anonymous`/`unnamed` refused as key names; duplicate names refused (a principal must identify one caller). **⚠ Identity MUST be resolved per call, not from the AuthContext.** The first implementation stamped the principal into the gRPC `AuthContext`; those properties belong to the **connection**, and gRPC pools channels by target, so three clients with three distinct keys all resolved as `ops` and every check that should have denied passed. `PrincipalResolver` now maps token→name from the request's own `authorization` metadata every call; the processor no longer consumes that header. **Policy** lives in the global `_policies` collection, id `:`, plus `:__security__` for the enable flag and mode — an ordinary collection, so it is WAL'd and snapshotted for free (putting it in `CollectionOptions` would have needed a new WAL op). Rules grant read/write per collection or file type with `"*"` fallback (exact match wins), a column `mask`, and a row predicate AND-merged like a view's `where`. **Safety is structural:** enabling is refused without an `admin` policy, and the last admin of a secured project cannot be removed. **`mode:"audit"`** evaluates honestly, logs what it would deny, and allows — the only safe way to arm a live project. **Coverage was established by auditing every `DatabaseGrpcImpl::` handler, not from a list** — the first pass missed `Upsert`, all three `Batch*`, all four `Set*`, and `Subscribe` (which streams document bodies and, with an empty `collections` list, means every collection; now filtered per event). Non-obvious surfaces also closed: version history (returns old bodies → bypasses masks), `ListCollections`/`ListViews`/`ListFiles`/`GetMemoryStats` filter per entry, `ListFiles` recomputes `total_count`, `GetCollectionInfo` gates on `name` not `collection`. Filtering on a masked field is **refused** (else the mask is an inference channel); a write touching a masked field is **denied**, not silently dropped; a row predicate reports **not-found**, never "forbidden". **System collections are admin-only, and the gate special-cases them:** they are global, so routing `_policies` through the per-project evaluator resolved it to project `default` (normally unsecured) and let anyone read every project's policies — caught by the e2e the moment the client stopped qualifying `_`-prefixed names. **`qualify()` no longer qualifies `_`-prefixed collections** (they are global; `default:_policies` does not exist, so reads silently found nothing while writes succeeded). **Client reads now THROW on `PERMISSION_DENIED`** (`get`/`exists`/`find`/`findWithMetrics`/`count`) — a denial that looked like "no data" left applications unable to distinguish "not allowed" from "nothing there"; same reasoning as the v2.4.2 silent-empty fix. **`subscribe()` was the last unqualified client call:** a bare name matched nothing (events carry the qualified collection, so subscriptions silently never fired) and an empty list meant every collection in **every project**; named collections are qualified and "everything" sends a `:*` pattern. An audit of all 41 name-carrying client call sites found no others (`createProject`/`dropProject` take project names, `uploadFile`/`listFiles` `set_name` take filenames). **Service-wide ops cannot be expressed per project:** `GetStats`/`SetReadOnly`/`Create|DropProject` require admin of some secured project, `ListProjects` filters; the real boundary there is listener separation. `HealthCheck`/`GetReadOnlyStatus` deliberately open. **Management needs no new RPC:** `_policies` writes through the ordinary document API are intercepted on `Insert`/`Upsert`/`Delete` — they refresh the cache and route `__security__` through `setSecurity` so the lockout guards apply. `Upsert` was missed at first and the symptom was policies that appeared to store and had no effect. CLI: `security`, `security-set [enforce|audit]`, `policies`, `policy`, `policy-set`, `policy-rm`. **Fixed on the way:** `auth.required=true` with `tls.enabled=false` only logged a WARN and then gRPC **aborted the process** on `SetAuthMetadataProcessor` over insecure credentials; now refused at startup. Tests: `tests/test_policy_manager.cpp` (54), `tests/load_test/test_policy_enforcement.sh` (19, three distinct named keys over TLS — the suite that caught the per-connection identity bug, which unit tests structurally cannot). ctest 17/17, namespacing e2e 31/31, views + TLS/auth green. - **v2.6.0 project-scoped file storage.** BREAKING (ABI + wire). Files had **no project dimension**: one blob tree, one flat `records/` tree of `.meta.json` files, and no `project` field on any of the five file RPCs. That made "per-project file policy" a statement with no referent, so scoping had to land before the security work (see `docs/superpowers/specs/2026-08-08-rls-cls-design.md`). **Note the metadata is on-disk JSON under `/records/<2-char-prefix>/.meta.json`, NOT the `_files` LMDB collection** — easy to assume otherwise. `FileInfo::project` is persisted there; absent means `default`, never unknown. `project` added to `FileMetadata`, `DownloadFileRequest`, `DeleteFileRequest`, `GetFileInfoRequest`, `FileInfo`, `ListFilesRequest`; empty resolves to `default`. **Cross-project isolation:** new `getFileInfoIn` / `readFileIn` / `deleteFileIn` overloads compare the record's project and report a mismatch **identically to a missing record** (nullopt / throw / false → gRPC `NOT_FOUND`) — the caller must not be able to distinguish "owned by someone else" from "does not exist", since that difference is itself a cross-project oracle. The project-blind `getFileInfo`/`readFile`/`deleteFile` remain for operator use and are commented as such; handlers must never use them. `listFiles` now takes `project` as its first, non-optional parameter and filters on it before every other predicate. **Dedup leak fixed:** blob dedup stays global (the disk saving is real) but `StoreResult::deduplicated` is now computed from whether **the requesting project** already references that checksum. Reporting the global hit told project A that project B held those exact bytes — upload a candidate, read the flag, learn whether a peer namespace has the file. `getRefCount()` stays global and is operator-only. **Migration:** `FileManager::stampMissingProjects()` writes `project: "default"` onto records with no project key, idempotent, called from `initialize()` after `auditSubdbPlacement()`; advisory (warns, never fails startup) since a file-metadata stamp is not a reason to refuse service. Records that already declare a project are never restamped, so the migration cannot relocate a file. **Client:** `project` is filled from `Config::project` on all five calls and read back into `FileRecord::project`; adding that member changes the struct size, so **consumers must rebuild** (SOVERSION stays 2). Files use a separate `project` field rather than the `project:name` string form collections use. Tests: `tests/test_file_project_scope.cpp` (30 assertions) + `tests/load_test/test_client_namespacing.sh` grew to 29. ctest 16/16, views e2e + TLS/auth 4/4 green. - **v2.4.5 collection management was not project-namespaced + per-collection versioning switch.** BREAKING (ABI). **The bug:** `Client::createCollection`, `dropCollection` and `getCollectionInfo` sent `set_name(name)` bare while ~30 other call sites send `set_name(qualify(...))`. `qualify()` prepends the project **unconditionally, including `default`**, so there was no project where the two paths agreed. One `createCollection("things")` + one `insert("things")` from the same client produced **two** collections: bare `things` received the options (`encrypted`, `maxVersions`, `vectorDimension`, `defaultTtlSeconds`) and stayed empty, while `:things` was created implicitly by the first insert with **defaults**. Silent — no error anywhere. There is no server-side normalization (`MemoryStore::createCollection` stores the string verbatim; only LMDB routing calls `resolveCollection`), so the whole scheme depends on the client qualifying. **Consequences, all reproduced:** `vector_dimension` is immutable after creation, so a declared vector collection could never get its dimension and `similaritySearch` returned 0 hits; `encrypted: true` silently did not apply to the collection holding the data; and worst, `dropCollection` **returned `true` while every document remained readable** — it dropped the empty phantom. Without a prior `createCollection` it returned `false` instead, so the failure mode depended on call history. Same class as the v2.4.2 `createView` break, second occurrence in the same file. **Audit:** exactly three methods were wrong; `createProject`/`dropProject` (project names) and `uploadFile`/`listFiles` (filenames) are correctly bare. `listCollections()` now `unqualify()`s results so bare names round-trip; note `ListCollectionsRequest` still has no project filter, unlike `ListViewsRequest`, so other projects' names remain visible. **Migration:** consumers upgrading will find leftover empty phantom collections (harmless litter). Options that were swallowed do NOT retroactively apply — `createCollection` on a now-correctly-resolved existing collection returns `false` and only updates `maxVersions`, so **a collection that needed `vector_dimension` must be recreated**. **Versioning switch:** `CollectionCfg::versioningEnabled` (default `true`) turns version history on/off on an **existing** collection via `configureCollection`. It lives in `CollectionCfg` (the `_collection_meta` system collection), not `CollectionOptions`, specifically because `_collection_meta` is an ordinary collection and therefore WAL'd and snapshotted for free — putting it in `CollectionOptions` would have required a new `ALTER_COLLECTION` WAL op to survive a restart between snapshots. Enforced in one place, `MemoryStore::saveToHistory`, so no write path can forget it. Disabling stops NEW versions; existing history stays readable, so it is reversible. **NOT the same as `maxVersions == 0`, which means unlimited.** `ConfigureCollection` is now a **partial update**: `optional bool versioning_enabled` (proto3 field presence) and empty `timestamp_precision` both mean "leave unchanged" — load-bearing, since a plain proto3 bool defaults to false and would have silently disabled versioning on any precision-only call. `Client::CollectionConfig::versioningEnabled` is `std::optional`; adding it changes the struct's size, so **consumers must rebuild** (SOVERSION stays 2, matching the v2.1 precedent). **Test:** `tests/load_test/test_client_namespacing.{cpp,sh}` — 17 assertions over gRPC with a non-default project. This is the boundary both namespacing bugs slipped through: `tests/test_views.cpp` only unit-tests `applyProjection()` in-process, and nothing drove a real `Client` against a real server with `Config::project` set. ctest 15/15, views e2e + TLS/auth e2e 4/4 green. - **v2.4.4 the other half of the v2.4.3 dbi bug: 31 documents were filed under the wrong collection.** A user reported an impossible triple on 2.4.3: `get image_hashes/` → not found, `get executions/` → readable, `insert image_hashes/` → "already exists". No id index exists; the three operations simply consult different substrates. `Get`/`Find`/`Exists` are LMDB-first, `insert`'s duplicate check reads MemoryStore (`memory_store.cpp:271`). MemoryStore was right; LMDB had the document in the wrong sub-db. **Root cause is the same stale `MDB_dbi` v2.4.3 fixed, in its second failure mode.** `MDB_dbi` is an index into the env's shared handle table, not a pointer. Pre-2.4.3, `try_open_for_read` cached a handle opened inside a `ReadTxn`; the abort released the slot WITHOUT bumping `me_numdbs`, so the next `mdb_dbi_open` for an unrelated collection was handed the same number. Whether you then got `EINVAL` (v2.4.3's drain incident) or a silently wrong sub-db depended only on whether something had reclaimed the slot. Writes through a reclaimed slot **succeed, into a stranger's collection**, and nothing detected it: `mirror_healthy_` only tracks writes that throw. **Damage found:** 31 rows in `smartbotic-automation` — 23 in `executions` declaring `image_hashes`, 6 in `workflows` declaring `sessions`, 1 in `workflows` declaring `users`, 1 in `nsfw_temp` declaring `workflows`. `default` was clean. 29 of the 31 existed ONLY in the wrong sub-db, so a delete-the-strays repair would have destroyed them. **The repair oracle is the data itself:** each stored document carries its own `collection` field, so a misfiled row is self-identifying — no cross-referencing MemoryStore or the WAL. **Prevention — `storage/subdb_identity.{hpp,cpp}`:** every sub-db carries a reserved key (`"\0__subdb_identity__"`, leading NUL so it cannot collide with a doc id) whose value is the sub-db's own name. `open_for_write` verifies it on every cached-handle reuse and throws on mismatch; one `mdb_get` per write buys a loud failure instead of silent cross-collection corruption. `count()` subtracts it, `scan()`/`scan_vectors()` skip it. Absence is treated as "unknown", not "wrong", so pre-2.4.4 sub-dbs keep working and acquire a sentinel on their next write. **Detection — `DatabaseService::auditSubdbPlacement()`** runs at boot per project, logs ERROR with a per-pair summary and the exact repair command. Advisory: never blocks startup, since misplacement is an operator-repairable data-location problem, not a reason to refuse service on the rest of the dataset. **Repair — `storage/subdb_placement.{hpp,cpp}`,** compiled into both the service and the CLI so there is one implementation: `smartbotic-db-cli verify-subdbs --env [--project N]` (read-only, safe against a running server, exit 1 if anything is misfiled) and `reconcile-subdbs ... --apply` (dry run by default; service MUST be stopped). Repair never deletes: a row moves to its declared home only when the home has no row under that id, otherwise the stray copy is parked in `_orphans_` for manual comparison. **`--stamp-identity` is opt-in and requires the server to already be ≥ 2.4.4** — a 2.4.3 binary does not skip the sentinel and would throw decoding it in `scan()` and miscount by one per collection. Order of operations on an existing box: repair placement → upgrade → stamp. **Counts are now LMDB-first too.** `Count` called `store_.count()` (MemoryStore only) while `Get`/`Find`/`Exists` were LMDB-first, so it under-reported whenever the two diverged — and they diverge routinely, because MemoryStore is a bounded cache that evicts (`pressure=soft` fires on this dataset at the 512 MB budget) while LMDB is the full dataset. Both `Count` and `GetCollectionInfo` now read LMDB behind the same gate; unfiltered via `ds->count()`, filtered via `scan()` with `limit=0` (`total_matched` is computed before pagination, so no document is materialised). **`smartbotic-db-cli count` calls `GetCollectionInfo`, NOT the `Count` RPC** — that is why fixing only `Count` appeared to change nothing. `CollectionInfo::sizeBytes` deliberately stays a MemoryStore estimate: it measures resident footprint, which is what the eviction knobs act on. Observed before/after on `smartbotic-automation:image_hashes`: 47 / 47 / 69 → 70 / 70 / 70. **⚠ NEVER open a second `MDB_env` on a path this process already has open.** The first 2.4.4 build did exactly that — `auditSubdbPlacement()` called the path-taking `audit()`, which opens its own env and closes it. LMDB coordinates readers with POSIX record locks, and **POSIX locks are per-process, not per-descriptor: closing ANY fd on a file releases every lock the process holds on it.** The audit's `mdb_env_close()` therefore destroyed the lock state of the service's own env, and every subsequent `mdb_txn_begin` returned `EINVAL` for the life of the process. All LMDB reads were down for ~90s in production before rollback to 2.4.3. Use `audit_env(MDB_env*)` in-process — it borrows an open env and opens only a read txn. `audit(path, ...)` is CLI-only (separate process). **Test any change to env lifetime against a copy of the real dataset on a staging port before installing — the unit suite does not catch this, because it never has two envs open on one path.** Tests: `tests/test_subdb_identity.cpp` (23 assertions; `misbound_handle_is_refused` reproduces the misbinding by verifying a handle under the wrong name, `test_scan_limit_zero_reports_total` pins the filtered-count contract). ctest 15/15. Deployed 2.4.4-1; both projects stamped (44 + 11 sub-dbs). - **v2.4.3 two fixes from a live drain incident + test suite had no teeth.** A production instance appeared to lose all data: every collection reported 0 docs. Nothing was deleted — two independent faults combined. **(1) LMDB dbi handle lifetime.** `LmdbDocumentStore::try_open_for_read` called `mdb_dbi_open` inside a `ReadTxn` and cached the handle in `dbi_cache_`. LMDB keeps a handle private to the opening transaction until it COMMITS and closes it if the transaction ABORTS — and `ReadTxn::~ReadTxn` always aborts. So the cached `MDB_dbi` dangled and every later `mdb_put`/`mdb_cursor_open` on that collection failed `EINVAL` for the life of the process. It hit whichever collection was READ before it was WRITTEN after a restart (`smartbotic-automation:sessions`); peers were fine because `open_for_write` commits, promoting the handle into the env's shared table. Fix: the read path no longer caches — the handle is valid inside its own txn, which is all the caller needs, and writes still populate the cache permanently. That failed mirror write flipped `mirror_healthy_=false`, sending all reads to MemoryStore. **(2) Eviction drained the store.** Eviction assumes `estimatedMemoryBytes_` falls as docs leave. It didn't (347→358 MB while evicting 3000 docs; process RSS peaked at 1.9 GB against a 512 MB budget, so the estimator is ~5x low), so `current <= targetBytes` was never satisfied and it trimmed one chunk per tick for 30 minutes until 5996 of ~5990 docs were gone — into the MemoryStore that reads had just been redirected to. Fix: `evictionMaxEpisodePercent` (default 50) caps what one pressure episode may evict, enforced **per pass** (a Hard/Emergency tick runs up to 20 chunks, enough to empty a store between tick-level checks); an episode re-arms only after `kEvictionRearmNormalTicks`=3 consecutive Normal ticks, since eviction transiently dips the estimate and a per-tick reset drains in 50% stages; plus a no-progress detector that pauses eviction after 3 consecutive evicting ticks that fail to lower the estimate. Both log ERROR — hitting them means the estimator is wrong, not that the store is too big. **(3) 104 assertions were no-ops.** `CMAKE_BUILD_TYPE=Release` sets `-DNDEBUG`, so every `assert()` in test_views (24), test_snapshot_durability (30), test_timestamp_precision (26), test_config_dropins (14) and test_eviction (10) compiled away — the suites ran, printed PASS, and verified nothing. `tests/CMakeLists.txt` now applies `-UNDEBUG` to test targets. Restoring them surfaced no further failures. Tests: `test_read_first_collection_survives_repeated_access` (reproduces EINVAL — note the ENV must be reopened, not just the store, since handles live in the env's shared table), `test_eviction_never_drains_the_store`. ctest 14/14, both e2e suites green. - **v2.4.2 project-scoped views + non-silent lookup misses.** Views were unreachable through the C++ client on 2.3.0-2.4.1, for every project including `default`. `Client::createView` sent `name` bare but `collection` qualified (`client/src/client.cpp`), so `ViewManager`'s cache keyed on the bare name while `Find`/`Get`/`Count` looked up the project-qualified one — the keys could never meet. The miss fell through to "treat it as a real collection", yielding `0 docs` and no error. `tests/test_views.cpp` passed throughout because it only unit-tests `applyProjection()` in-process; the break lives at the client/server boundary. **Fix:** views are now keyed by their qualified `:` form. The client qualifies view names on `createView`/`dropView`/`getViewInfo` and un-qualifies them on the way back (`unqualify()`), so the caller keeps using bare names and `listViews()` round-trips. Two projects may each own a view of the same name; previously the global registry rejected the second one as a duplicate. `ListViewsRequest` gained a `project` field (additive, field 1) — empty means "all projects" for operator/CLI use, clients always set their own. **Migration:** legacy bare-named rows in `_views` are re-keyed into `default` on load (`ViewManager::loadFromStore`), idempotent, logged as `re-keyed N legacy view(s)`. **Behavior change:** `Find` on a name that is neither a view nor a collection returns gRPC `NOT_FOUND`, and `Client::find`/`findWithMetrics` throw `std::runtime_error` rather than returning `{}` — the silent-empty is what hid this bug. Consumers that relied on empty-for-missing must catch (smartbotic-automation's `StorageClient::query` was patched to convert it to `ErrorCode::CollectionNotFound`). **Test:** `tests/load_test/test_views_multiproject.sh` — 17 checks over both projects, projection, name reuse, cross-project isolation, NOT_FOUND. ctest 14/14, TLS/auth e2e 4/4. - **v2.4.1 `--version` reports the real version.** `printVersion()` carried a hardcoded `"1.0.0"` literal, so the server binary misreported itself for the whole 2.x line — `--version` was useless for triage and only `dpkg -l` was truthful. `SMARTBOTIC_DB_VERSION_STRING` is now injected into the service target from the top-level `VERSION` file (which `CMakeLists.txt:7` already read but never passed down). The `#ifndef` fallback is `"unknown"`, not another literal. Note `BUILD_GIT_COMMIT` is still declared-but-unused in `packaging/Dockerfile.build`, so packages record no build provenance. - **v2.4.0 per-listener TLS + bearer-token auth.** New `storage.listeners[]` config array — each entry has its own `bind`, `port`, `tls{}`, `auth{}`. One process, multiple `grpc::Server` instances, same `DatabaseGrpcImpl` registered on each. Default policy preserves v2.3: 127.0.0.1 plaintext, no auth (the v2.3 `bind_address` + `rpc_port` keys are synthesised into one listener when `listeners[]` is absent). New listeners (LAN/public) are opt-in. **TLS:** `tls.enabled` + `tls.cert_path`/`tls.key_path`. If missing AND `tls.auto_self_signed_if_missing: true` (default), a 10-year RSA-2048 self-signed cert is generated at `/tls/auto_self_signed.{pem,key}` on first boot (SubjectAltName includes `DNS:localhost`, `IP:127.0.0.1`, plus the bind as DNS-or-IP). Operators with real PKI override by dropping their cert at `cert_path`. Loud WARN on auto-generation. **Auth:** `auth.required` + `auth.keys[]`. Server attaches `BearerAuthProcessor` (`service/src/auth/auth_interceptor.{hpp,cpp}`) via `SetAuthMetadataProcessor` — constant-time compares `authorization: Bearer ` against `auth.keys[]`. gRPC handles UNAUTHENTICATED rejection without per-handler code. Refuses to start if `auth.required=true` with empty keys; logs WARN if auth required without TLS. **Client:** `Client::Config::{tls_enabled, tls_ca_cert_path, tls_insecure_skip_verify, auth_token}`; auth attached via `ClientContext::AddMetadata` in `setDeadline()` so every RPC (unary, stream, file) is covered. `tls_insecure_skip_verify` (dev-only) uses `grpc::experimental::TlsChannelCredentialsOptions` + `NoOpCertificateVerifier`; loud WARN on every connect. **CLI:** `smartbotic-db-cli generate-auth-key` (32 random bytes → base64) and `generate-tls-cert --bind [--out-cert PATH] [--out-key PATH] [--days N]` (offline cert minting via OpenSSL 3 EVP API). **End-to-end test:** `tests/load_test/test_v24_tls_auth.sh` boots a server with both plaintext + TLS+auth listeners, drives the new client through 4 scenarios — plaintext-local OK, TLS+token OK, TLS+wrong-token UNAUTHENTICATED, TLS+no-token UNAUTHENTICATED. All 4 pass. ctest 14/14 green. - **v2.3.1 two bug fixes found by post-2.3.0 stress + replication tests.** (1) `MemoryStore::getEvictionCandidates` compared `doc.updatedAt` against a millisecond-scaled `hotWriteFloor`; v2.2.0's default-precision flip to nanoseconds made every doc look "always within the hot-write window" so eviction never fired. Pre-existing v2.2.0 regression — load_test_mixed was never re-run after v2.2.0 shipped. Fix: pick ms- or ns-scaled floor per-doc using the 10^15 magnitude heuristic. (2) `DatabaseService::applyReplicatedEntry` called `MemoryStore::loadDocument` which bypasses the persist callback, so the v2.3 LMDB mirror never fired for replicated entries. Followers got writes into MemoryStore but the LMDB env stayed empty; v2.3 LMDB-first reads on the follower returned 0 docs. Fix: explicitly drive `applyDualWriteMirror` through the project registry after `loadDocument`. Replication catch-up time on a 100-doc test: 65ms. - **v2.3.0 multi-project namespaces.** Each smartbotic-database instance now hosts multiple `project` namespaces, each with its own collections. Wire form: `:` (e.g. `my_app:users`); bare names like `"users"` auto-qualify to `"default:users"` so every v2.0-v2.2 client keeps working without code changes. Storage: one LMDB env per project at `/projects//env/`. On first v2.3 boot the legacy `/env/` is atomically renamed under `projects/default/env/` (idempotent; refuses to start if both exist). Server-side: `ProjectStoreRegistry` (`service/src/storage/project_store.{hpp,cpp}`) owns the per-project envs; `DatabaseService::docStore(project)` resolves the LMDB store for a given name; mirror routing is via a resolver callback the service hands to `MemoryStore::setDocumentStoreMirror(resolver, ...)`. Client-side: `Client::Config::project` defaults to `"default"`; every outgoing collection name is auto-prepended (`qualify(coll)`) unless it already contains `:`. Per-call overrides via the qualified form always win. New gRPC RPCs: `ListProjects`, `CreateProject(name)`, `DropProject(name)` — `default` cannot be dropped. CRUD is idempotent and validated against `^[a-zA-Z_][a-zA-Z0-9_-]{0,62}$` (no `:`, no `/`, no `_-prefix). Reserved: `default` (always present); `_-prefix` (system). Tests: `test_project_addressing` (23 parser assertions), `test_project_crud` (29 placeholder assertions). 14/14 ctest green. Stress: load_test_mixed still passes against the multi-project build. - **v2.2.1 bug fixes from shadowman code review.** Three issues found by a code-review pass on the v2.0→v2.2.0 commit range, all now fixed: (1) `MemoryStore::updateIfVersion` was missing `extractVector` + `storeVector` + `mirrorVectorToDocStore` calls — optimistic-lock updates on vector collections left the vector mirror stale, so `SimilaritySearch` returned old vector scores while `Get` returned new doc bodies; (2) `DatabaseService::backfillIntoDocStore` walked every doc into LMDB but never set `_meta.schema_version=2`, so every boot re-backfilled the entire dataset (Zoe-scale ~1.85M docs = minutes of pre-READY blocking on every restart); (3) `SimilaritySearch` LMDB handler returned `INTERNAL` instead of `INVALID_ARGUMENT` when an unrelated throw (e.g. bad_alloc) fired after a dimension-mismatch was already observed in the scan. Exposes `smartbotic::db::storage::mark_migration_complete(env)` for the backfill marker. Tests 12/12 green. - **v2.2.0 — default timestamp precision flipped to `"ns"`.** BREAKING for fresh collections. `CollectionCfg::timestampPrecision` default went from `"ms"` to `"ns"` so rapid-write collections (chat messages, metrics events, tool-call audits) get nanosecond stamps and strict ordering without needing an explicit `configureCollection()` call. Closes shadowman audit B5. Operators who need ms can still set it explicitly. Existing collections that already have a stored config record keep their value — only collections created post-upgrade WITHOUT an explicit `configureCollection()` inherit the new default. Mixed-precision rows in the same collection are read-correct via the 10^15 magnitude auto-detect; operators wanting clean rows should run `migrateCollectionTimestamps`. **`update()` docstring** updated to clarify: it OVERWRITES the doc; use `patch()` for partial mutations. Reflects shadowman audit B6. - **v2.1.1 cleanup — no defensive fallback on LMDB throw.** Removed the `catch → fall back to MemoryStore` paths in Exists / Get / Find / SimilaritySearch. A live LMDB exception now returns gRPC `INTERNAL` to the client instead of silently being masked by MemoryStore. The gate-closed fallback (`mirror_healthy_=false`, system collection, env unopened) stays — that's correctness, since MemoryStore can legitimately be ahead of LMDB when a mirror write has failed. **Answers shadowman audit F3:** `usedWalFallback` / `walMatchCount` / `memoryMatchCount` / `memorySearchMicros` / `walSearchMicros` are guaranteed zero on the LMDB-served path. Any non-zero value is a real signal — it means the read fell back to MemoryStore (mirror unhealthy). Operators can treat `walMatchCount > 0` as a paging/SLA alert. - **v2.1 audit answers (shadowman alignment).** F1 — v1.x auto-migration: stays in-binary through v2.4. Extracts to separate `smartbotic-database-migrate-v1` deb in v2.5 once every production host has booted v2.4 ≥ once. v2.x is the project's terminal arc, no v3 planned. F2 — `setReadOnly` / `lock()` / `unlock()` are dual-audience: operator-side via `smartbotic-db-cli unlock` and client-side via `Client::lock()` for self-defensive migrations. F3 — see above. F4 — `migrateCollectionTimestamps` cost benchmark deferred until shadowman starts D.1 (we haven't seen real traffic on the path yet); back-of-envelope ~1ms per 1k docs on Zoe-shape data, online-safe up to ~1M docs. - **Client API cleanup (v2.1.0)** — BREAKING. Removed the pre-v1.6 `smartbotic::database::QueryOptions` struct from `client/include/smartbotic/database/document.hpp` (dead code; the EQ-only `vector>` shape that predated FilterOp). Removed the no-options `Client::find(collection)` overload — pass `QueryOptions{}` explicitly. Added ergonomic factories on `Client::Filter`: `Filter::eq` / `ne` / `gt` / `gte` / `lt` / `lte` / `in` / `contains` / `exists` / `regex` / `search`. Callsites now read like the predicate they're asking: `find("users", {.filters = {Filter::gt("age", 18), Filter::in("role", {"admin", "owner"})}})`. SOVERSION stays at 2 (matches the v2.0.0 deb; no soname bump needed) — consumers must rebuild against v2.1 headers, but the soname is identical. `libsmartbotic-db-client (>= 2.1.0)` should be the floor in consumer deb deps. - **LMDB storage substrate (v2.0.0+)** — Documents and vectors live in LMDB sub-databases under `/env/`. Each user collection becomes one sub-db `` for documents; collections with `vector_dimension > 0` get a sidecar `_vectors_` storing raw float32 bytes keyed by docId. `LmdbEnv` (RAII over `MDB_env*`) is opened in `DatabaseService::setupComponents()`; `LmdbDocumentStore` provides `put`/`get`/`del`/`scan` + `put_vector`/`del_vector`/`get_vector`. Dual-write mirror runs under `MemoryStore`'s per-collection lock (`mirrorWriteToDocStore` + `mirrorVectorToDocStore` invoked before `lock.unlock()` at every write site), so concurrent LMDB readers observe the same state as MemoryStore readers — no race window. On failure: ERROR log, `mirror_drift_count_` bumps, `mirror_healthy_` flips to false; read handlers (Exists/Get/Find) fall back to MemoryStore. Filter eval lifted to `service/src/storage/filter_eval.hpp` shared between substrates so divergence is attributable to data state, not code paths. **v2.0 read coverage:** Exists, Get, Find, SimilaritySearch all LMDB-first when `mirror_healthy_` + zero drift. SimilaritySearch iterates the `_vectors_` sub-db via `DocumentStore::scan_vectors` and scores with the AVX2/SSE4.1/scalar SIMD kernel extracted to `service/src/storage/cosine_simd.hpp`. **Deferred to v2.1:** history (`_history_`) / files (`_files`) / views (`_views`) sub-db migration, MemoryStore decommission, write-handler migration so handlers call `doc_store_->put()` directly. **Migration:** v1.x→v2.0 auto-migrates on boot (`--migrate-from-v1` opt-in / `--no-auto-migrate` opt-out) via `migrate_v1_to_v2`; backfill of any docs added before mirror wiring runs sync in `initialize()`. **Eviction config knobs** (`max_memory_mb`, `eviction_chunk_size`, etc.) are kept for back-compat with v1.x snapshots-still-loaded-into-MemoryStore but emit a deprecation WARN at boot. RSS is bounded by the LMDB mapsize + OS page cache, not by MemoryStore eviction. **Stress numbers** (`load_test_mixed` 30s / 4 writers / 8 readers / 32MB MemoryStore budget): 2385 writes/s, 5769 reads/s, p99 < 100ms, 0 hard failures, 240k ops total, 282MB persisted in LMDB. New runtime dep: `liblmdb0`. ## Packaging Canonical source for all smartbotic-database Debian packages. Replaces the old `shadowman-database` and `callerai-storage` packages. ### Build modes ```bash # Production (Docker, Debian 13) ./packaging/build.sh # Build 4 .debs ./packaging/build.sh --rebuild-base # Rebuild Docker base image ./packaging/build.sh --sync --suite trixie # Build + publish to repository.smartbotics.ai # Local development (native, current system) ./packaging/build.sh --local # Build 4 .debs ./packaging/build.sh --local --install # Build + install locally ``` ### Packages | Package | Contents | Deployed where | |---------|----------|---------------| | `smartbotic-database` | Server binary + systemd + config | Production DB servers | | `libsmartbotic-db-client` | Shared `.so` (runtime) | All production machines | | `libsmartbotic-db-client-dev` | Headers + linker symlink + `.proto` + cmake config | Docker build images, dev workstations | | `smartbotic-db-cli` | CLI tool | Anywhere (optional) | ### Local dev with installed packages After `./packaging/build.sh --local --install`, consumer projects can use: ```cmake find_package(smartbotic-db-client REQUIRED) target_link_libraries(myapp PRIVATE smartbotic::db-client) ``` The submodule fallback still works for projects that haven't switched (`BUILD_SHARED_LIBS=OFF`). ### Consumers | Project | Depends on | Min version | Migrations dir | |---------|-----------|-------------|----------------| | shadowman-cpp | `smartbotic-database (>= 1.2.0)` | 1.2.0 | `/opt/shadowman/share/shadowman/migrations/json` | | callerai | `smartbotic-database (>= 1.2.0)` | 1.2.0 | `/etc/callerai/migrations` | ### Version policy Bump `VERSION` for API/protocol changes. Deb revision (`-N`) for packaging-only changes. ## systemd integration (Type=notify, v1.7.5+) The service uses **`Type=notify`** in its systemd unit and signals readiness via `sd_notify(READY=1)` only after recovery + migrations complete and the gRPC listener is up (see `service/src/main.cpp` around the `service.start()` call). This is the lifecycle contract for dependents: - `After=smartbotic-database.service` + `Type=notify` on this unit means systemd blocks the dependent's startup until the DB sends `READY=1`. A dependent doing `db.get(...)` immediately on launch sees a fully-recovered DB, never an in-recovery one. Pre-1.7.5 the unit was `Type=simple`, so `sd_notify` was decoration and consumers could observe an in-recovery DB and silently fall back to defaults — surfaced on the shadowman side as `instance_type` reading null and `Dev mode enabled` never logging until manual restart. - `TimeoutStartSec=300` covers slow recoveries on large datasets. The service emits `EXTEND_TIMEOUT_USEC=600000000` (10 min) at each phase boundary as belt-and-braces against systemd's start watchdog. - `STATUS=...` notifications are emitted at each phase ("Initializing — running recovery and migrations" → "Starting gRPC server" → "Ready") so `systemctl status smartbotic-database` shows where the service is during a slow boot. - `sd_notify(STOPPING=1)` is emitted on graceful shutdown so dependents observe a draining state distinct from a crash. The graceful-shutdown path also takes a final snapshot (v1.8.0 hardening) so the next boot's recovery is `TrivialSuccess` rather than `WalOnlyReplay`. **Dependents (shadowman-* services and any other consumer)** should: - Set `After=smartbotic-database.service` and `Wants=smartbotic-database.service`. - Set their own `Type=notify` if they have a meaningful "ready" boundary. - Treat critical-config DB reads as fail-loud: a null result on a setting the service depends on is a startup-ordering bug, not a default fallback. **Operator-facing tooling** (deb postinst, ansible, etc.) should not layer its own port-poll or sleep loops on top of the unit. `systemctl restart smartbotic-database` (or `systemctl start ... && systemctl start dependent`) already blocks correctly until READY. ## Conventions - C++20 - nlohmann/json for JSON handling - spdlog for logging - LZ4 for compression - gRPC/Protobuf for API