# Snapshot Durability + Recovery Modes (v1.6.1) Implementation Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Fix the silent snapshot-truncation bug that killed the Zoe production instance, add tiered recovery modes (like MySQL's `innodb_force_recovery`), add read-only runtime mode with automatic safe-mode boot when recovery is non-trivial. Urgent patch release. **Architecture:** Four independent pieces that compose: (1) bulletproof snapshot writer (atomic rename + fsync + error checks + post-write verification), (2) snapshot loader with configurable fallback chain, (3) read-only runtime state gated by a new `SetReadOnly` RPC, (4) auto-readonly-on-non-trivial-recovery safety net. All changes are additive — default behavior on working deployments is unchanged except that recovery failure is no longer silent. **Tech Stack:** C++20, POSIX fsync, std::filesystem::rename (atomic on same filesystem), existing nlohmann/json for config, existing gRPC/proto for new RPC --- ## Context — the bug (see `BUG-snapshot-truncation.md`) **Writer issues:** 1. `file.write()` calls never check failbit — silent short writes pass through 2. No `flush()` / `fsync()` before `close()` — no durability guarantee 3. Writes directly to final path — partial writes survive as "valid" snapshots 4. `cleanupOldSnapshots()` runs blindly — evicts good snapshots after a bad write 5. `verifySnapshot()` exists but is never called after writes **Loader issues:** 6. `loadLatestSnapshot()` tries only `snapshots.front()` — no fallback 7. Recovery exception is caught and logged; `persistence->recover()` return value is IGNORED at call site 8. Result: DB silently starts empty when snapshots are corrupt **Production impact:** Zoe instance restarted → latest snapshot corrupt → started empty store → ShadowMan wrote junk to blank DB → all data lost --- ## File Structure ### Files to modify - `VERSION` — bump to 1.6.1 - `proto/database.proto` — new `SetReadOnly` / `GetReadOnlyStatus` RPCs - `service/src/persistence/snapshot.hpp` — writer API stays the same, but internal behavior changes - `service/src/persistence/snapshot.cpp` — atomic writer, fallback loader - `service/src/persistence/persistence_manager.hpp` — recovery mode enum, auto-escalation config, recovery outcome struct - `service/src/persistence/persistence_manager.cpp` — implement modes, auto-escalation, WAL-only replay - `service/src/database_service.hpp` — read-only state member, recovery outcome - `service/src/database_service.cpp` — handle recovery return value, enter auto-readonly on non-trivial recovery, refuse-to-start when appropriate - `service/src/database_grpc_impl.hpp` — declare new RPC handlers, check isReadOnly for every write handler - `service/src/database_grpc_impl.cpp` — implement `SetReadOnly`/`GetReadOnlyStatus`, reject writes when read-only - `service/src/main.cpp` — new CLI flags (`--read-only`, `--force-readwrite`, `--recovery-mode`) - `client/include/smartbotic/database/client.hpp` — client methods for lock/unlock - `client/src/client.cpp` — implement new client methods - `cli/main.cpp` — new `unlock` and `lock` subcommands - `packaging/deb/config/config.json` — document new `recovery` section - `tests/test_snapshot_durability.cpp` — new integration test - `tests/CMakeLists.txt` — register new test ### Config surface (new) ```json "storage": { "persistence": { ... "snapshots": { "validate_after_write": true, "cleanup_only_if_verified": true }, "recovery": { "mode": "normal", "auto_escalate": true, "allow_empty_on_fresh_install": true } } } ``` ### CLI flags (new) ``` --read-only Start in read-only mode regardless of recovery outcome --force-readwrite Accept auto-readonly recovery state; start writable anyway --recovery-mode=MODE Override recovery.mode config for this boot Valid: normal, snapshot_fallback, wal_only, best_effort, force_empty ``` ### Exit codes (new convention) | Code | Meaning | |------|---------| | 0 | Normal exit | | 1 | Generic error | | **10** | Recovery refused — snapshot/WAL corrupt + mode=normal + data exists. Operator must escalate. | | **11** | Config invalid | --- ### Task 1: Bump VERSION and add recovery mode types **Files:** - Modify: `VERSION` - Modify: `service/src/persistence/persistence_manager.hpp` - [ ] **Step 1: Bump VERSION to 1.6.1** ```bash echo "1.6.1" > /data/smartbotic-database/VERSION ``` - [ ] **Step 2: Add RecoveryMode enum and RecoveryOutcome struct in persistence_manager.hpp** At the top of the `smartbotic::database` namespace in `persistence_manager.hpp`, before the `PersistenceManager` class, add: ```cpp enum class RecoveryMode { Normal, // level 0 — latest snapshot only, refuse on failure (default) SnapshotFallback, // level 1 — try snapshots newest→oldest WalOnly, // level 2 — ignore snapshots, replay WAL from 0 BestEffort, // level 3 — snapshot_fallback then wal_only ForceEmpty // level 4 — start empty, preserve files for forensics }; std::string recoveryModeToString(RecoveryMode mode); RecoveryMode recoveryModeFromString(const std::string& s); // throws on invalid /** * Describes how recovery proceeded. A "non-trivial" outcome (anything other than * `TrivialSuccess`) triggers auto-readonly mode at the service layer. */ struct RecoveryOutcome { enum class Kind { TrivialSuccess, // latest snapshot loaded cleanly FreshInstall, // no snapshots, no WAL — clean state SnapshotFellBack, // non-latest snapshot loaded (auto or manual) WalOnlyReplay, // skipped all snapshots, replayed WAL from 0 ForcedEmpty, // started empty despite data on disk Failed // recovery refused }; Kind kind = Kind::TrivialSuccess; std::filesystem::path snapshotUsed; // empty if none std::filesystem::path expectedSnapshot; // empty if none was tried std::string failureReason; // why the expected one didn't load uint64_t walEntriesReplayed = 0; size_t snapshotsAttempted = 0; // how many we tried before one worked size_t snapshotsAvailable = 0; // total on disk bool isNonTrivial() const { return kind != Kind::TrivialSuccess && kind != Kind::FreshInstall; } bool isFailure() const { return kind == Kind::Failed; } }; ``` Add to `PersistenceManager::Config`: ```cpp struct Config { std::filesystem::path dataDir; uint32_t walSyncIntervalMs = 100; uint32_t snapshotIntervalSec = 3600; uint64_t maxWalSizeMb = 100; bool compressionEnabled = true; uint32_t maxSnapshots = 5; // NEW: snapshot durability settings bool validateAfterWrite = true; bool cleanupOnlyIfVerified = true; // NEW: recovery settings RecoveryMode recoveryMode = RecoveryMode::Normal; bool autoEscalate = true; bool allowEmptyOnFreshInstall = true; }; ``` Change the `recover()` method signature: ```cpp // was: bool recover(MemoryStore& store); RecoveryOutcome recover(MemoryStore& store); ``` - [ ] **Step 3: Build to confirm compilation errors and commit** The service will not compile yet — `recover()` callers in `database_service.cpp` and the implementation in `persistence_manager.cpp` still use the `bool` signature. That's expected; Task 2 fixes it. ```bash cd /data/smartbotic-database cmake --build build -j$(nproc) 2>&1 | grep -c "error:" # Expect several errors about recover() signature mismatch — proceed to next task ``` ```bash git add VERSION service/src/persistence/persistence_manager.hpp git commit -m "feat(recovery): RecoveryMode enum + RecoveryOutcome struct (v1.6.1 scaffold)" ``` --- ### Task 2: Implement enum helpers and update recover() signature **Files:** - Modify: `service/src/persistence/persistence_manager.cpp` - [ ] **Step 1: Add enum string conversions at top of persistence_manager.cpp** Add inside the `smartbotic::database` namespace, near the top (after includes): ```cpp std::string recoveryModeToString(RecoveryMode mode) { switch (mode) { case RecoveryMode::Normal: return "normal"; case RecoveryMode::SnapshotFallback: return "snapshot_fallback"; case RecoveryMode::WalOnly: return "wal_only"; case RecoveryMode::BestEffort: return "best_effort"; case RecoveryMode::ForceEmpty: return "force_empty"; } return "normal"; } RecoveryMode recoveryModeFromString(const std::string& s) { if (s == "normal") return RecoveryMode::Normal; if (s == "snapshot_fallback") return RecoveryMode::SnapshotFallback; if (s == "wal_only") return RecoveryMode::WalOnly; if (s == "best_effort") return RecoveryMode::BestEffort; if (s == "force_empty") return RecoveryMode::ForceEmpty; throw std::invalid_argument("invalid recovery mode: '" + s + "' (valid: normal, snapshot_fallback, wal_only, best_effort, force_empty)"); } ``` - [ ] **Step 2: Update `PersistenceManager::recover()` signature + initial scaffolding** Replace the existing `recover()` implementation with a stub that still compiles but doesn't yet implement the full mode logic — just wraps the old behavior in a RecoveryOutcome. We'll flesh it out in Task 5. Replace the entire existing `bool PersistenceManager::recover(MemoryStore& store)` method with: ```cpp RecoveryOutcome PersistenceManager::recover(MemoryStore& store) { RecoveryOutcome outcome; spdlog::info("Starting recovery (mode={}, auto_escalate={})", recoveryModeToString(config_.recoveryMode), config_.autoEscalate); { std::lock_guard lock(storeMutex_); currentStore_ = &store; } // Enumerate available snapshots auto snapshots = snapshotMgr_->listSnapshots(); outcome.snapshotsAvailable = snapshots.size(); // Fresh-install case: no snapshots at all if (snapshots.empty()) { if (config_.allowEmptyOnFreshInstall) { // Open WAL (it may also be empty for a truly fresh install) if (!wal_->open()) { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = "Failed to open WAL for recovery"; return outcome; } uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) { applyWalEntry(store, entry); }); outcome.walEntriesReplayed = replayed; wal_->close(); outcome.kind = (replayed == 0) ? RecoveryOutcome::Kind::FreshInstall : RecoveryOutcome::Kind::WalOnlyReplay; // WAL-only replay counts as non-trivial spdlog::info("Fresh install or WAL-only recovery complete ({} WAL entries)", replayed); return outcome; } else { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = "No snapshots found and allow_empty_on_fresh_install is false"; return outcome; } } // TODO: implement mode-specific recovery logic (Task 5). // For now, keep the old behavior — load latest snapshot, replay WAL. // This preserves the pre-1.6.1 behavior while we plumb the infrastructure. try { outcome.expectedSnapshot = snapshots.front(); uint64_t snapshotSequence = snapshotMgr_->loadSnapshot(snapshots.front(), store); outcome.snapshotUsed = snapshots.front(); outcome.snapshotsAttempted = 1; spdlog::info("Loaded snapshot at sequence {}", snapshotSequence); if (!wal_->open()) { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = "Failed to open WAL after snapshot load"; return outcome; } uint64_t replayed = wal_->replay(snapshotSequence, [&store](const WalEntry& entry) { applyWalEntry(store, entry); }); outcome.walEntriesReplayed = replayed; spdlog::info("Replayed {} WAL entries", replayed); { std::lock_guard lock(statsMutex_); lastSnapshotSequence_ = snapshotSequence; } wal_->close(); outcome.kind = RecoveryOutcome::Kind::TrivialSuccess; spdlog::info("Recovery complete, store has {} documents in {} collections", store.getStats().totalDocuments, store.getStats().totalCollections); return outcome; } catch (const std::exception& e) { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = e.what(); spdlog::error("Recovery failed: {}", e.what()); return outcome; } } ``` - [ ] **Step 3: Extract the WAL entry switch into a helper `applyWalEntry`** The current `recover()` has a big switch-on-WalOpType inside a lambda. Extract it to a free function inside an anonymous namespace near the top of the file (so we can reuse it for WAL-only replay in Task 5): ```cpp namespace { void applyWalEntry(MemoryStore& store, const WalEntry& entry) { switch (entry.opType) { case WalOpType::INSERT: if (entry.data) { Document doc = Document::fromJson(*entry.data); store.loadDocument(entry.collection, doc); } break; case WalOpType::UPDATE: case WalOpType::UPSERT: if (entry.data) { Document doc = Document::fromJson(*entry.data); store.loadDocumentWithHistory(entry.collection, doc); } break; case WalOpType::DELETE: store.remove(entry.collection, entry.documentId); break; case WalOpType::CREATE_COLLECTION: if (entry.collectionOptions) { store.createCollection(entry.collection, *entry.collectionOptions); } break; case WalOpType::DROP_COLLECTION: store.dropCollection(entry.collection); break; case WalOpType::SET_ADD: if (entry.data && entry.data->contains("member")) { store.setAdd(entry.collection, entry.documentId, (*entry.data)["member"].get()); } break; case WalOpType::SET_REMOVE: if (entry.data && entry.data->contains("member")) { store.setRemove(entry.collection, entry.documentId, (*entry.data)["member"].get()); } break; case WalOpType::VEC_PUT: if (entry.vectorData.has_value()) { store.loadVector(entry.collection, entry.documentId, *entry.vectorData); } break; case WalOpType::VEC_DELETE: break; // handled by DELETE entry } } } // anonymous namespace ``` - [ ] **Step 4: Build and verify tests still pass** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 # Expect remaining errors in database_service.cpp because it still uses the bool return. # We'll fix that in Task 7. For now, comment out the call temporarily ONLY if build is # fully blocked — normally it should compile (the function is used via pointer). ``` If `database_service.cpp:42` fails to compile because it discards the `RecoveryOutcome` return value, that's fine — C++ allows discarding return values. The real fix comes in Task 7. Build should succeed. ```bash ./build/tests/test_views LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage ``` Both should still pass. - [ ] **Step 5: Commit** ```bash git add service/src/persistence/persistence_manager.cpp git commit -m "feat(recovery): recover() returns RecoveryOutcome; extract applyWalEntry helper" ``` --- ### Task 3: Atomic snapshot writer + post-write verification **Files:** - Modify: `service/src/persistence/snapshot.cpp` This is the core bug fix. Changes: 1. Write to `.tmp` first 2. Check failbit after every write 3. `flush()` + `fsync()` on the file before close 4. `std::filesystem::rename(tmp, final)` atomically moves it into place 5. `fsync()` the directory for durability of the rename 6. Call `verifySnapshot(final_path)` if config says so 7. Only call `cleanupOldSnapshots()` if verification passed - [ ] **Step 1: Add POSIX includes** At the top of `snapshot.cpp`, after existing includes: ```cpp #include #include ``` - [ ] **Step 2: Pass new config values through** `SnapshotManager::Config` needs `validateAfterWrite` and `cleanupOnlyIfVerified` fields. Add them to the struct in `snapshot.hpp`: ```cpp struct Config { std::filesystem::path snapshotDir; bool compressionEnabled = true; uint32_t maxSnapshots = 5; bool validateAfterWrite = true; // NEW bool cleanupOnlyIfVerified = true; // NEW }; ``` - [ ] **Step 3: Rewrite `createSnapshot()` with atomic semantics** Replace the existing `createSnapshot()` function body (from line 24 onwards) keeping the same signature and serialization/compression/header-computation code unchanged, but replacing the write-to-disk section (currently lines 62-86): ```cpp // Generate filename and write std::filesystem::path snapshotPath = config_.snapshotDir / generateFilename(); std::filesystem::path tmpPath = snapshotPath; tmpPath += ".tmp"; // Open temp file { std::ofstream file(tmpPath, std::ios::binary | std::ios::trunc); if (!file) { throw std::runtime_error("Failed to create snapshot tmp file: " + tmpPath.string()); } // Write header + check file.write(reinterpret_cast(&header), sizeof(header)); if (!file) { std::error_code ec; std::filesystem::remove(tmpPath, ec); throw std::runtime_error("Failed to write snapshot header: " + tmpPath.string()); } // Write body + check file.write(reinterpret_cast(bodyData.data()), static_cast(bodyData.size())); if (!file) { std::error_code ec; std::filesystem::remove(tmpPath, ec); throw std::runtime_error("Failed to write snapshot body (short write): " + tmpPath.string()); } // Compute and write body checksum + check uint32_t bodyChecksum = crc32::calculate(bodyData.data(), bodyData.size()); file.write(reinterpret_cast(&bodyChecksum), sizeof(bodyChecksum)); if (!file) { std::error_code ec; std::filesystem::remove(tmpPath, ec); throw std::runtime_error("Failed to write snapshot trailer: " + tmpPath.string()); } // Flush C++ stream buffer file.flush(); if (!file) { std::error_code ec; std::filesystem::remove(tmpPath, ec); throw std::runtime_error("Failed to flush snapshot: " + tmpPath.string()); } } // std::ofstream dtor runs close() // fsync the file for durability (C++ ofstream does not do this) { int fd = ::open(tmpPath.c_str(), O_RDONLY); if (fd < 0) { std::error_code ec; std::filesystem::remove(tmpPath, ec); throw std::runtime_error("Failed to open snapshot tmp for fsync: " + tmpPath.string()); } if (::fsync(fd) != 0) { int saved_errno = errno; ::close(fd); std::error_code ec; std::filesystem::remove(tmpPath, ec); throw std::runtime_error("fsync failed on snapshot tmp (errno=" + std::to_string(saved_errno) + "): " + tmpPath.string()); } ::close(fd); } // Atomic rename { std::error_code ec; std::filesystem::rename(tmpPath, snapshotPath, ec); if (ec) { std::filesystem::remove(tmpPath, ec); throw std::runtime_error("Snapshot rename failed: " + ec.message()); } } // fsync the containing directory so the rename is durable across crash { int dfd = ::open(config_.snapshotDir.c_str(), O_RDONLY | O_DIRECTORY); if (dfd >= 0) { ::fsync(dfd); ::close(dfd); } // Non-fatal if the dirfd open fails — log only } // Post-write verification if (config_.validateAfterWrite) { if (!verifySnapshot(snapshotPath)) { spdlog::error("Snapshot post-write verification FAILED: {}", snapshotPath.string()); std::error_code ec; std::filesystem::remove(snapshotPath, ec); throw std::runtime_error("Snapshot verification failed after write: " + snapshotPath.string()); } spdlog::debug("Snapshot verified after write: {}", snapshotPath.string()); } // Cleanup — only if verification passed (or was disabled) if (!config_.cleanupOnlyIfVerified || config_.validateAfterWrite) { cleanupOldSnapshots(); } return snapshotPath; ``` - [ ] **Step 4: Handle leftover `.tmp` files on startup** In `listSnapshots()` at line 160, add a pre-scan that removes any `*.tmp` files left over from crashes. Near the top of the function, before the directory_iterator loop: ```cpp // Clean up any .tmp files left from a crash during a previous write std::error_code ec; for (const auto& entry : std::filesystem::directory_iterator(config_.snapshotDir, ec)) { if (entry.is_regular_file() && entry.path().extension() == ".tmp" && entry.path().filename().string().starts_with("snapshot-")) { std::error_code rm_ec; std::filesystem::remove(entry.path(), rm_ec); spdlog::info("Removed orphaned snapshot tmp file: {}", entry.path().string()); } } ``` - [ ] **Step 5: Build** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ``` - [ ] **Step 6: Run existing tests** ```bash ./build/tests/test_views LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage ``` Both should pass — writer changes don't break anything existing. - [ ] **Step 7: Commit** ```bash git add service/src/persistence/snapshot.hpp service/src/persistence/snapshot.cpp git commit -m "fix(snapshot): atomic writer (.tmp+fsync+rename) + post-write verification Fixes silent data loss when snapshot writes were truncated by disk pressure or interrupted mid-write. Core changes: - Write to .tmp first, then atomically rename to final path - Check failbit after every ofstream::write() — no more silent short writes - flush() + fsync() file before considering the write complete - fsync() the containing directory so the rename is durable after crash - Call verifySnapshot() post-write; delete the file if verification fails - cleanupOldSnapshots() only runs after successful verification - listSnapshots() cleans up orphaned .tmp files left by crashed writes" ``` --- ### Task 4: Loader fallback chain **Files:** - Modify: `service/src/persistence/snapshot.cpp` - Modify: `service/src/persistence/snapshot.hpp` - [ ] **Step 1: Add new method declaration in snapshot.hpp** Add to the `SnapshotManager` public section: ```cpp /** * Attempt to load snapshots in order, trying newest first and falling * back to older ones if the newer ones are corrupt. * * @param store The store to load into * @param outUsed If non-null, set to the path of the snapshot that loaded * @return WAL sequence from the loaded snapshot, or 0 if none loaded */ uint64_t loadWithFallback(MemoryStore& store, std::filesystem::path* outUsed = nullptr); ``` - [ ] **Step 2: Implement in snapshot.cpp** Add the implementation after `loadLatestSnapshot`: ```cpp uint64_t SnapshotManager::loadWithFallback(MemoryStore& store, std::filesystem::path* outUsed) { auto snapshots = listSnapshots(); // already sorted newest-first if (snapshots.empty()) { return 0; } for (size_t i = 0; i < snapshots.size(); ++i) { const auto& path = snapshots[i]; try { uint64_t seq = loadSnapshot(path, store); if (outUsed) *outUsed = path; if (i > 0) { spdlog::warn("Loaded fallback snapshot {} (newer snapshots failed): {}", i, path.string()); } return seq; } catch (const std::exception& e) { spdlog::warn("Snapshot load failed for {}: {} — trying next older snapshot", path.string(), e.what()); // store may be in a partially-cleared state here. Next loadSnapshot // call will call store.clear() first, so this is safe. } } spdlog::error("All {} snapshots failed to load", snapshots.size()); throw std::runtime_error("all snapshots failed to load (tried " + std::to_string(snapshots.size()) + ")"); } ``` - [ ] **Step 3: Build and test** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ./build/tests/test_views LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage ``` - [ ] **Step 4: Commit** ```bash git add service/src/persistence/snapshot.hpp service/src/persistence/snapshot.cpp git commit -m "fix(snapshot): loadWithFallback — try snapshots newest-first, fall back on failure" ``` --- ### Task 5: Implement full RecoveryMode logic in PersistenceManager **Files:** - Modify: `service/src/persistence/persistence_manager.cpp` - [ ] **Step 1: Replace the recover() implementation with mode-aware logic** Replace the body of `PersistenceManager::recover()` (the one written in Task 2, Step 2) with: ```cpp RecoveryOutcome PersistenceManager::recover(MemoryStore& store) { RecoveryOutcome outcome; spdlog::info("Starting recovery (mode={}, auto_escalate={})", recoveryModeToString(config_.recoveryMode), config_.autoEscalate); { std::lock_guard lock(storeMutex_); currentStore_ = &store; } auto snapshots = snapshotMgr_->listSnapshots(); outcome.snapshotsAvailable = snapshots.size(); // ForceEmpty: skip everything, start clean if (config_.recoveryMode == RecoveryMode::ForceEmpty) { spdlog::warn("RecoveryMode::ForceEmpty — starting with empty store. " "Snapshot and WAL files preserved on disk for forensics."); outcome.kind = RecoveryOutcome::Kind::ForcedEmpty; return outcome; } // Fresh-install case: no snapshots at all if (snapshots.empty()) { if (!config_.allowEmptyOnFreshInstall) { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = "No snapshots found and allow_empty_on_fresh_install is false"; return outcome; } if (!wal_->open()) { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = "Failed to open WAL for recovery"; return outcome; } uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) { applyWalEntry(store, entry); }); outcome.walEntriesReplayed = replayed; wal_->close(); outcome.kind = (replayed == 0) ? RecoveryOutcome::Kind::FreshInstall : RecoveryOutcome::Kind::WalOnlyReplay; return outcome; } outcome.expectedSnapshot = snapshots.front(); // Helper lambda: try loading a snapshot and populate outcome.snapshotUsed on success auto tryLoadSnapshot = [&](const std::filesystem::path& path) -> std::optional { outcome.snapshotsAttempted++; try { uint64_t seq = snapshotMgr_->loadSnapshot(path, store); outcome.snapshotUsed = path; return seq; } catch (const std::exception& e) { if (outcome.failureReason.empty()) { outcome.failureReason = e.what(); } spdlog::warn("Snapshot load failed for {}: {}", path.string(), e.what()); return std::nullopt; } }; auto replayWal = [&](uint64_t fromSequence) -> bool { if (!wal_->open()) { outcome.failureReason = "Failed to open WAL after snapshot load"; return false; } uint64_t replayed = wal_->replay(fromSequence, [&store](const WalEntry& entry) { applyWalEntry(store, entry); }); outcome.walEntriesReplayed = replayed; spdlog::info("Replayed {} WAL entries from sequence {}", replayed, fromSequence); { std::lock_guard lock(statsMutex_); lastSnapshotSequence_ = fromSequence; } wal_->close(); return true; }; switch (config_.recoveryMode) { case RecoveryMode::Normal: { auto seq = tryLoadSnapshot(snapshots.front()); if (seq) { if (!replayWal(*seq)) { outcome.kind = RecoveryOutcome::Kind::Failed; return outcome; } outcome.kind = RecoveryOutcome::Kind::TrivialSuccess; return outcome; } // Latest snapshot failed. Can we auto-escalate to snapshot_fallback? if (config_.autoEscalate && snapshots.size() > 1) { spdlog::error("Latest snapshot failed; auto-escalating to snapshot_fallback"); for (size_t i = 1; i < snapshots.size(); ++i) { auto s = tryLoadSnapshot(snapshots[i]); if (s) { if (!replayWal(*s)) { outcome.kind = RecoveryOutcome::Kind::Failed; return outcome; } outcome.kind = RecoveryOutcome::Kind::SnapshotFellBack; return outcome; } } } // Auto-escalation disabled or exhausted outcome.kind = RecoveryOutcome::Kind::Failed; return outcome; } case RecoveryMode::SnapshotFallback: { for (const auto& path : snapshots) { auto seq = tryLoadSnapshot(path); if (seq) { if (!replayWal(*seq)) { outcome.kind = RecoveryOutcome::Kind::Failed; return outcome; } outcome.kind = (path == snapshots.front()) ? RecoveryOutcome::Kind::TrivialSuccess : RecoveryOutcome::Kind::SnapshotFellBack; return outcome; } } outcome.kind = RecoveryOutcome::Kind::Failed; return outcome; } case RecoveryMode::WalOnly: { store.clear(); if (!wal_->open()) { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = "Failed to open WAL"; return outcome; } uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) { applyWalEntry(store, entry); }); outcome.walEntriesReplayed = replayed; wal_->close(); outcome.kind = RecoveryOutcome::Kind::WalOnlyReplay; return outcome; } case RecoveryMode::BestEffort: { // Try snapshot fallback for (const auto& path : snapshots) { auto seq = tryLoadSnapshot(path); if (seq) { if (!replayWal(*seq)) { outcome.kind = RecoveryOutcome::Kind::Failed; return outcome; } outcome.kind = (path == snapshots.front()) ? RecoveryOutcome::Kind::TrivialSuccess : RecoveryOutcome::Kind::SnapshotFellBack; return outcome; } } // Fall through to WAL-only spdlog::error("All snapshots failed in best_effort mode; attempting WAL-only replay"); store.clear(); if (!wal_->open()) { outcome.kind = RecoveryOutcome::Kind::Failed; outcome.failureReason = "Failed to open WAL (best_effort fallback)"; return outcome; } uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) { applyWalEntry(store, entry); }); outcome.walEntriesReplayed = replayed; wal_->close(); outcome.kind = RecoveryOutcome::Kind::WalOnlyReplay; return outcome; } case RecoveryMode::ForceEmpty: // handled at the top break; } outcome.kind = RecoveryOutcome::Kind::Failed; return outcome; } ``` - [ ] **Step 2: Build and run existing tests** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ./build/tests/test_views LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage ``` All should pass — default behavior (RecoveryMode::Normal + auto_escalate=true) matches the old behavior for healthy deployments. - [ ] **Step 3: Commit** ```bash git add service/src/persistence/persistence_manager.cpp git commit -m "feat(recovery): implement all 5 RecoveryMode cases with auto-escalation - Normal: latest snapshot only; if fails AND auto_escalate, try older - SnapshotFallback: walk snapshots newest-first, first that loads wins - WalOnly: skip all snapshots, replay WAL from 0 - BestEffort: snapshot_fallback + fall through to WalOnly if nothing loads - ForceEmpty: skip everything, preserve files for forensics RecoveryOutcome.kind tells the caller whether recovery was trivial, non-trivial, or failed. The service layer uses this to decide auto-readonly." ``` --- ### Task 6: Add SetReadOnly / GetReadOnlyStatus RPCs to proto **Files:** - Modify: `proto/database.proto` - [ ] **Step 1: Add RPC declarations** In the `DatabaseService` service block, after the existing health/stats RPCs, add: ```protobuf // Read-only runtime state control rpc SetReadOnly(SetReadOnlyRequest) returns (SetReadOnlyResponse); rpc GetReadOnlyStatus(GetReadOnlyStatusRequest) returns (GetReadOnlyStatusResponse); ``` - [ ] **Step 2: Add message definitions** At the end of the file: ```protobuf // ===== Read-Only Control ===== message SetReadOnlyRequest { bool read_only = 1; } message SetReadOnlyResponse { bool success = 1; bool was_read_only = 2; // previous state string error = 3; } message GetReadOnlyStatusRequest {} message GetReadOnlyStatusResponse { bool read_only = 1; string reason = 2; // empty when !read_only; explains why otherwise // When auto-readonly was triggered by recovery, these fields describe it string recovery_outcome = 3; // "trivial_success", "snapshot_fell_back", "wal_only_replay", "forced_empty", "failed" string expected_snapshot = 4; // the snapshot we tried to load first string snapshot_used = 5; // the one that actually loaded string failure_reason = 6; // why expected_snapshot failed uint64 wal_entries_replayed = 7; uint32 snapshots_attempted = 8; } ``` - [ ] **Step 3: Build to confirm proto compiles** ```bash cmake --build build -j$(nproc) 2>&1 | grep -E "error|Generating database" | head -5 ``` Expect `Generating protobuf/gRPC code for database.proto` and no errors. - [ ] **Step 4: Commit** ```bash git add proto/database.proto git commit -m "feat(proto): add SetReadOnly + GetReadOnlyStatus RPCs for runtime lock control" ``` --- ### Task 7: Wire recovery outcome into DatabaseService + auto-readonly **Files:** - Modify: `service/src/database_service.hpp` - Modify: `service/src/database_service.cpp` - [ ] **Step 1: Add read-only state to DatabaseService** In `database_service.hpp`, add to the `DatabaseService` class: ```cpp #include // Public section: bool isReadOnly() const { return read_only_.load(std::memory_order_acquire); } std::string readOnlyReason() const; void setReadOnly(bool value, const std::string& reason); const RecoveryOutcome& recoveryOutcome() const { return recovery_outcome_; } // Private members: std::atomic read_only_{false}; mutable std::mutex reason_mutex_; std::string read_only_reason_; RecoveryOutcome recovery_outcome_; bool force_readwrite_ = false; // set from CLI / config ``` - [ ] **Step 2: Implement in database_service.cpp** Add member function implementations: ```cpp std::string DatabaseService::readOnlyReason() const { std::lock_guard lock(reason_mutex_); return read_only_reason_; } void DatabaseService::setReadOnly(bool value, const std::string& reason) { { std::lock_guard lock(reason_mutex_); read_only_reason_ = value ? reason : ""; } bool prev = read_only_.exchange(value, std::memory_order_acq_rel); if (prev != value) { if (value) { spdlog::error("Database entered READ-ONLY mode: {}", reason); } else { spdlog::info("Database unlocked — writes accepted"); } } } ``` - [ ] **Step 3: Consume the recovery outcome at startup** Find the call site at `database_service.cpp:42` (`persistence_->recover(*store_);`) and replace with: ```cpp // Recover from persistence recovery_outcome_ = persistence_->recover(*store_); if (recovery_outcome_.isFailure()) { spdlog::error("\n" "┌────────────────────────────────────────────────────────────────────┐\n" "│ RECOVERY REFUSED │\n" "│ │\n" "│ Mode: {}\n" "│ Expected snapshot: {}\n" "│ Reason: {}\n" "│ │\n" "│ To escalate, restart with one of: │\n" "│ --recovery-mode=snapshot_fallback Try older snapshots │\n" "│ --recovery-mode=wal_only Replay WAL only (slow) │\n" "│ --recovery-mode=best_effort Try all of the above │\n" "│ --recovery-mode=force_empty Start empty (LAST RESORT) │\n" "│ │\n" "│ Or add to config.json: \"recovery\": {{ \"mode\": \"\" }} │\n" "│ │\n" "│ Preserve /var/lib/smartbotic-database/ BEFORE escalating. │\n" "└────────────────────────────────────────────────────────────────────┘", recoveryModeToString(config_.persistenceConfig.recoveryMode), recovery_outcome_.expectedSnapshot.string(), recovery_outcome_.failureReason); throw std::runtime_error("recovery failed; refusing to start with exit code 10"); } if (recovery_outcome_.isNonTrivial() && !force_readwrite_) { // Enter auto-readonly mode std::string reason = "Non-trivial recovery occurred ("; switch (recovery_outcome_.kind) { case RecoveryOutcome::Kind::SnapshotFellBack: reason += "fell back to snapshot " + recovery_outcome_.snapshotUsed.filename().string() + " because " + recovery_outcome_.expectedSnapshot.filename().string() + " failed: " + recovery_outcome_.failureReason; break; case RecoveryOutcome::Kind::WalOnlyReplay: reason += "WAL-only replay, no snapshot loaded (" + std::to_string(recovery_outcome_.walEntriesReplayed) + " entries)"; break; case RecoveryOutcome::Kind::ForcedEmpty: reason += "forced empty by operator (--recovery-mode=force_empty)"; break; default: break; } reason += "). Run `smartbotic-db-cli unlock` to accept this state, or restart " "with --force-readwrite to bypass this check."; setReadOnly(true, reason); spdlog::error("\n" "┌────────────────────────────────────────────────────────────────────┐\n" "│ [ERROR] Database booted in READ-ONLY mode after non-trivial recovery│\n" "│ │\n" "│ {:66} │\n" "│ │\n" "│ Writes will be REJECTED until you acknowledge this state: │\n" "│ smartbotic-db-cli unlock # live, no restart│\n" "│ smartbotic-database --force-readwrite # on next restart │\n" "│ │\n" "│ To extract data before deciding: │\n" "│ smartbotic-db-cli dump > backup-$(date +%s).json (v1.7.0) │\n" "└────────────────────────────────────────────────────────────────────┘", recoveryModeToString(config_.persistenceConfig.recoveryMode)); } // Set replication sequence after recovery replication_->setSequence(persistence_->currentWalSequence()); ``` Note: the exit code 10 for recovery-refused is achieved via exception handling in `main.cpp` — see Task 10. - [ ] **Step 4: Build and verify defaults unchanged** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ./build/tests/test_views LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage ``` All should pass. On clean deployment with valid snapshots, `recovery_outcome.kind == TrivialSuccess` and no read-only enforcement. - [ ] **Step 5: Commit** ```bash git add service/src/database_service.hpp service/src/database_service.cpp git commit -m "feat(recovery): DatabaseService consumes RecoveryOutcome; auto-readonly on non-trivial - setReadOnly/isReadOnly/readOnlyReason — runtime lock state (atomic) - Recovery failure throws → main.cpp catches → exit code 10 - Non-trivial recovery (snapshot fallback, WAL-only, forced empty) enters read-only with an ERROR-level log explaining exactly what happened - --force-readwrite flag (Task 10) bypasses the auto-readonly" ``` --- ### Task 8: Reject writes when read-only in gRPC handlers **Files:** - Modify: `service/src/database_grpc_impl.hpp` - Modify: `service/src/database_grpc_impl.cpp` - [ ] **Step 1: Add DatabaseService& reference to DatabaseGrpcImpl** `DatabaseGrpcImpl` needs to consult `DatabaseService` for `isReadOnly()`. Add a constructor parameter. In `database_grpc_impl.hpp`: Add forward declaration at top of namespace: ```cpp class DatabaseService; ``` Add constructor parameter and member: ```cpp // Constructor gains a new first parameter: DatabaseGrpcImpl(DatabaseService& service, ...all existing params...); private: DatabaseService& service_; // for isReadOnly() checks ``` - [ ] **Step 2: Add helper method for write handlers** In `database_grpc_impl.cpp`, add a private helper at the top of the class impl: ```cpp namespace { // Returns a grpc::Status that should be returned when the DB is read-only. grpc::Status readOnlyRejection(const std::string& rpc, const DatabaseService& svc) { return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, rpc + " rejected: database is in read-only mode. " + svc.readOnlyReason()); } } // anonymous namespace ``` - [ ] **Step 3: Reject writes at the top of every write handler** Add this check at the very top of each write-path handler (after parameter validation but before any store access): Handlers to modify: - `Insert` — add check - `Update` — add check - `Upsert` — add check - `Delete` — add check - `PatchDocument` — add check - `BatchInsert` — add check - `BatchDelete` — add check - `CreateCollection` — add check - `DropCollection` — add check - `CreateView` — add check - `DropView` — add check - `ConfigureCollection` — add check - `MigrateCollectionTimestamps` — add check - `SetAdd` / `SetRemove` — add check - `UploadFile` / `DeleteFile` — add check For handlers whose response has `set_success`/`set_error`: ```cpp if (service_.isReadOnly()) { response->set_success(false); response->set_error("database is in read-only mode: " + service_.readOnlyReason()); return grpc::Status::OK; } ``` For handlers without those fields: ```cpp if (service_.isReadOnly()) { return readOnlyRejection("Insert", service_); // use actual RPC name } ``` IMPORTANT: read every write handler before modifying to pick the right style. Exception: streaming handlers like `UploadFile` should reject via a stream status error — check the existing streaming error pattern. - [ ] **Step 4: Implement SetReadOnly and GetReadOnlyStatus handlers** Add at the bottom of database_grpc_impl.cpp: ```cpp grpc::Status DatabaseGrpcImpl::SetReadOnly( grpc::ServerContext* /*context*/, const pb::SetReadOnlyRequest* request, pb::SetReadOnlyResponse* response ) { bool was_readonly = service_.isReadOnly(); std::string reason = request->read_only() ? "manually locked via SetReadOnly RPC" : ""; // reason cleared on unlock service_.setReadOnly(request->read_only(), reason); response->set_success(true); response->set_was_read_only(was_readonly); return grpc::Status::OK; } grpc::Status DatabaseGrpcImpl::GetReadOnlyStatus( grpc::ServerContext* /*context*/, const pb::GetReadOnlyStatusRequest* /*request*/, pb::GetReadOnlyStatusResponse* response ) { response->set_read_only(service_.isReadOnly()); response->set_reason(service_.readOnlyReason()); const auto& out = service_.recoveryOutcome(); switch (out.kind) { case RecoveryOutcome::Kind::TrivialSuccess: response->set_recovery_outcome("trivial_success"); break; case RecoveryOutcome::Kind::FreshInstall: response->set_recovery_outcome("fresh_install"); break; case RecoveryOutcome::Kind::SnapshotFellBack: response->set_recovery_outcome("snapshot_fell_back"); break; case RecoveryOutcome::Kind::WalOnlyReplay: response->set_recovery_outcome("wal_only_replay"); break; case RecoveryOutcome::Kind::ForcedEmpty: response->set_recovery_outcome("forced_empty"); break; case RecoveryOutcome::Kind::Failed: response->set_recovery_outcome("failed"); break; } response->set_expected_snapshot(out.expectedSnapshot.string()); response->set_snapshot_used(out.snapshotUsed.string()); response->set_failure_reason(out.failureReason); response->set_wal_entries_replayed(out.walEntriesReplayed); response->set_snapshots_attempted(out.snapshotsAttempted); return grpc::Status::OK; } ``` Declare them in `database_grpc_impl.hpp`: ```cpp grpc::Status SetReadOnly(grpc::ServerContext*, const pb::SetReadOnlyRequest*, pb::SetReadOnlyResponse*) override; grpc::Status GetReadOnlyStatus(grpc::ServerContext*, const pb::GetReadOnlyStatusRequest*, pb::GetReadOnlyStatusResponse*) override; ``` - [ ] **Step 5: Update DatabaseGrpcImpl instantiation in DatabaseService** Find where `DatabaseGrpcImpl` is constructed in `database_service.cpp` (search for `std::make_unique` or similar) and pass `*this` as the new first argument. - [ ] **Step 6: Build and run tests** ```bash cmake --build build -j$(nproc) 2>&1 | tail -10 ./build/tests/test_views LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage ``` - [ ] **Step 7: Commit** ```bash git add service/src/database_grpc_impl.hpp service/src/database_grpc_impl.cpp service/src/database_service.cpp git commit -m "feat(readonly): reject writes in gRPC handlers when DB is read-only - New SetReadOnly/GetReadOnlyStatus RPCs let clients toggle + inspect state - Every write handler (insert/update/patch/upsert/delete/batch/collection management/view management/config management) checks service.isReadOnly() and returns FAILED_PRECONDITION with a clear reason - Read handlers (Get, Find, Count, GetVersionHistory, SimilaritySearch, ListCollections, ListViews, etc.) unaffected — read-only means read-ONLY" ``` --- ### Task 9: Client API for lock/unlock **Files:** - Modify: `client/include/smartbotic/database/client.hpp` - Modify: `client/src/client.cpp` - [ ] **Step 1: Add declarations in client.hpp** Add to the `Client` class public section (near other admin methods): ```cpp // ===== Read-Only Control ===== struct ReadOnlyStatus { bool readOnly = false; std::string reason; std::string recoveryOutcome; // "trivial_success", "snapshot_fell_back", etc. std::string expectedSnapshot; std::string snapshotUsed; std::string failureReason; uint64_t walEntriesReplayed = 0; uint32_t snapshotsAttempted = 0; }; /** * Toggle server-side read-only mode. When on, all writes are rejected * with FAILED_PRECONDITION. Useful for draining, maintenance, or * acknowledging a non-trivial recovery state. * @return true if the call succeeded (regardless of previous state) */ bool setReadOnly(bool readOnly); /** * Shortcut for setReadOnly(true). */ bool lock() { return setReadOnly(true); } /** * Shortcut for setReadOnly(false). Typically used after reviewing a * non-trivial recovery state and deciding to accept it. */ bool unlock() { return setReadOnly(false); } /** * Inspect the current read-only state and (if applicable) the recovery * outcome that caused auto-readonly. */ [[nodiscard]] ReadOnlyStatus getReadOnlyStatus(); ``` - [ ] **Step 2: Implement in client.cpp** Inside the `Impl` class: ```cpp bool setReadOnly(bool readOnly) { smartbotic::databasepb::SetReadOnlyRequest request; request.set_read_only(readOnly); smartbotic::databasepb::SetReadOnlyResponse response; grpc::ClientContext context; setDeadline(context); auto status = stub_->SetReadOnly(&context, request, &response); if (!status.ok()) { spdlog::error("Client::setReadOnly failed: {}", status.error_message()); return false; } return response.success(); } Client::ReadOnlyStatus getReadOnlyStatus() { smartbotic::databasepb::GetReadOnlyStatusRequest request; smartbotic::databasepb::GetReadOnlyStatusResponse response; grpc::ClientContext context; setDeadline(context); Client::ReadOnlyStatus out; auto status = stub_->GetReadOnlyStatus(&context, request, &response); if (!status.ok()) { spdlog::error("Client::getReadOnlyStatus failed: {}", status.error_message()); return out; } out.readOnly = response.read_only(); out.reason = response.reason(); out.recoveryOutcome = response.recovery_outcome(); out.expectedSnapshot = response.expected_snapshot(); out.snapshotUsed = response.snapshot_used(); out.failureReason = response.failure_reason(); out.walEntriesReplayed = response.wal_entries_replayed(); out.snapshotsAttempted = response.snapshots_attempted(); return out; } ``` Add public forwarders at the bottom of client.cpp: ```cpp bool Client::setReadOnly(bool readOnly) { return impl_->setReadOnly(readOnly); } Client::ReadOnlyStatus Client::getReadOnlyStatus() { return impl_->getReadOnlyStatus(); } ``` - [ ] **Step 3: Build and test** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ./build/tests/test_views LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage ``` - [ ] **Step 4: Commit** ```bash git add client/include/smartbotic/database/client.hpp client/src/client.cpp git commit -m "feat(client): setReadOnly/lock/unlock/getReadOnlyStatus client API" ``` --- ### Task 10: CLI flags + unlock/lock/status subcommands **Files:** - Modify: `service/src/main.cpp` - Modify: `cli/main.cpp` - [ ] **Step 1: Read existing CLI parsing patterns in main.cpp** ```bash head -150 /data/smartbotic-database/service/src/main.cpp head -100 /data/smartbotic-database/cli/main.cpp ``` Match existing styles. - [ ] **Step 2: Add server flags to service/src/main.cpp** In `printUsage()`, add below the existing `--config` line: ```cpp << " --read-only Start in read-only mode (reject all writes)\n" << " --force-readwrite Accept non-trivial recovery state; start writable\n" << " --recovery-mode=MODE Override recovery mode for this boot\n" << " Values: normal, snapshot_fallback, wal_only,\n" << " best_effort, force_empty\n" ``` Before parsing the config file, parse these flags. Add after `findConfigFile`: ```cpp struct CliOverrides { std::optional readOnly; std::optional forceReadwrite; std::optional recoveryMode; }; CliOverrides parseCliOverrides(int argc, char* argv[]) { CliOverrides ov; for (int i = 1; i < argc; i++) { std::string arg = argv[i]; if (arg == "--read-only") ov.readOnly = true; else if (arg == "--force-readwrite") ov.forceReadwrite = true; else if (arg.rfind("--recovery-mode=", 0) == 0) { std::string mode = arg.substr(16); try { ov.recoveryMode = recoveryModeFromString(mode); } catch (const std::exception& e) { std::cerr << "ERROR: " << e.what() << std::endl; std::exit(11); // config invalid } } } return ov; } ``` In `main()`, after loading config but before creating `DatabaseService`: ```cpp auto cli = parseCliOverrides(argc, argv); if (cli.recoveryMode) { config.storage.persistence.recoveryMode = *cli.recoveryMode; } if (cli.readOnly.value_or(false)) { config.storage.initialReadOnly = true; } if (cli.forceReadwrite.value_or(false)) { config.storage.forceReadwrite = true; } ``` Note: this adds two fields to the config struct — `initialReadOnly` and `forceReadwrite`. Add them to the config parsing in database_service.hpp/cpp. These are runtime overrides, not persisted config. Wrap `DatabaseService::initialize()` in a try/catch that exits with code 10 for recovery-refused: ```cpp try { db.initialize(); } catch (const std::exception& e) { if (std::string(e.what()).find("recovery failed") != std::string::npos) { std::cerr << "Recovery refused: " << e.what() << std::endl; return 10; } std::cerr << "Initialization failed: " << e.what() << std::endl; return 1; } ``` - [ ] **Step 3: Add `lock`, `unlock`, `status` subcommands to cli/main.cpp** Read the existing subcommand structure: ```bash grep -n "std::string subcommand\|argv\[1\]\|subcommand ==\|if.*argv" /data/smartbotic-database/cli/main.cpp | head -10 ``` Add three new subcommand branches: ```cpp else if (subcommand == "lock") { if (client.setReadOnly(true)) { std::cout << "Database locked (read-only)" << std::endl; return 0; } else { std::cerr << "Failed to lock database" << std::endl; return 1; } } else if (subcommand == "unlock") { if (client.setReadOnly(false)) { std::cout << "Database unlocked (writes accepted)" << std::endl; return 0; } else { std::cerr << "Failed to unlock database" << std::endl; return 1; } } else if (subcommand == "status") { auto s = client.getReadOnlyStatus(); std::cout << "Read-only: " << (s.readOnly ? "YES" : "no") << std::endl; if (s.readOnly) { std::cout << "Reason: " << s.reason << std::endl; } std::cout << "Recovery outcome: " << s.recoveryOutcome << std::endl; if (!s.expectedSnapshot.empty()) { std::cout << "Expected snapshot: " << s.expectedSnapshot << std::endl; } if (!s.snapshotUsed.empty()) { std::cout << "Snapshot used: " << s.snapshotUsed << std::endl; } if (!s.failureReason.empty()) { std::cout << "Failure reason: " << s.failureReason << std::endl; } std::cout << "WAL replayed: " << s.walEntriesReplayed << " entries" << std::endl; std::cout << "Snapshots tried: " << s.snapshotsAttempted << std::endl; return 0; } ``` Update help text to list `lock`, `unlock`, `status`. - [ ] **Step 4: Build and test** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ``` - [ ] **Step 5: Commit** ```bash git add service/src/main.cpp cli/main.cpp service/src/database_service.hpp git commit -m "feat(cli): --read-only/--force-readwrite/--recovery-mode flags + lock/unlock/status subcommands Server: - --read-only: start in read-only state - --force-readwrite: bypass auto-readonly from non-trivial recovery - --recovery-mode=X: override config for this boot - Exit code 10 when recovery is refused CLI: - smartbotic-db-cli lock → SetReadOnly(true) - smartbotic-db-cli unlock → SetReadOnly(false) - smartbotic-db-cli status → detailed read-only + recovery info" ``` --- ### Task 11: Config file parsing for new recovery/snapshot settings **Files:** - Modify: `service/src/database_service.cpp` (config loader) - Modify: `packaging/deb/config/config.json` - [ ] **Step 1: Find config parser** ```bash grep -n "persistence\|recovery\|walSyncIntervalMs" /data/smartbotic-database/service/src/database_service.cpp | head -10 ``` - [ ] **Step 2: Parse new fields** Where `persistence` settings are parsed, after existing fields, add: ```cpp // Snapshot durability if (storage.contains("persistence") && storage["persistence"].contains("snapshots")) { const auto& snap = storage["persistence"]["snapshots"]; config_.persistenceConfig.validateAfterWrite = snap.value("validate_after_write", true); config_.persistenceConfig.cleanupOnlyIfVerified = snap.value("cleanup_only_if_verified", true); } // Recovery if (storage.contains("persistence") && storage["persistence"].contains("recovery")) { const auto& rec = storage["persistence"]["recovery"]; std::string modeStr = rec.value("mode", std::string("normal")); try { config_.persistenceConfig.recoveryMode = recoveryModeFromString(modeStr); } catch (const std::exception& e) { spdlog::warn("Invalid recovery.mode '{}', defaulting to 'normal': {}", modeStr, e.what()); config_.persistenceConfig.recoveryMode = RecoveryMode::Normal; } config_.persistenceConfig.autoEscalate = rec.value("auto_escalate", true); config_.persistenceConfig.allowEmptyOnFreshInstall = rec.value("allow_empty_on_fresh_install", true); } ``` - [ ] **Step 3: Update default config.json in packaging** In `packaging/deb/config/config.json`, extend the `persistence` block: ```json "persistence": { "wal_sync_interval_ms": 100, "snapshot_interval_sec": 3600, "compression": "lz4", "snapshots": { "validate_after_write": true, "cleanup_only_if_verified": true }, "recovery": { "mode": "normal", "auto_escalate": true, "allow_empty_on_fresh_install": true } } ``` - [ ] **Step 4: Build and verify** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ``` - [ ] **Step 5: Commit** ```bash git add service/src/database_service.cpp packaging/deb/config/config.json git commit -m "feat(config): parse persistence.snapshots and persistence.recovery sections" ``` --- ### Task 12: Integration test for snapshot durability + fallback **Files:** - Create: `tests/test_snapshot_durability.cpp` - Modify: `tests/CMakeLists.txt` - [ ] **Step 1: Read test harness patterns** ```bash cat /data/smartbotic-database/tests/test_timestamp_precision.cpp | head -30 cat /data/smartbotic-database/tests/CMakeLists.txt ``` Follow exactly the pattern used for `test_timestamp_precision`. - [ ] **Step 2: Write tests/test_snapshot_durability.cpp** ```cpp #include "../service/src/persistence/snapshot.hpp" #include "../service/src/memory_store.hpp" #include "../service/src/document.hpp" #include #include #include #include #include #include namespace fs = std::filesystem; using smartbotic::database::MemoryStore; using smartbotic::database::SnapshotManager; using smartbotic::database::Document; using smartbotic::database::CollectionOptions; namespace { MemoryStore::Config defaultStoreConfig() { MemoryStore::Config cfg; cfg.nodeId = "test"; cfg.maxMemoryBytes = 256 * 1024 * 1024; return cfg; } fs::path makeSnapshotDir(const std::string& name) { auto dir = fs::temp_directory_path() / ("snap-test-" + name + "-" + std::to_string(::getpid())); fs::remove_all(dir); fs::create_directories(dir); return dir; } Document makeDoc(const std::string& id) { Document d; d.id = id; d.data = nlohmann::json{{"value", id}}; return d; } } // anonymous namespace void test_atomic_write_produces_complete_file() { auto dir = makeSnapshotDir("atomic"); SnapshotManager::Config scfg; scfg.snapshotDir = dir; scfg.validateAfterWrite = true; SnapshotManager mgr(scfg); MemoryStore store(defaultStoreConfig()); store.start(); store.createCollection("docs", CollectionOptions{}); for (int i = 0; i < 100; i++) { store.insert("docs", makeDoc("d" + std::to_string(i))); } auto path = mgr.createSnapshot(store, 42); store.stop(); // No .tmp files left int tmpCount = 0; for (auto& e : fs::directory_iterator(dir)) { if (e.path().extension() == ".tmp") tmpCount++; } assert(tmpCount == 0); // Verify file passes integrity check assert(mgr.verifySnapshot(path)); // Load into a fresh store MemoryStore store2(defaultStoreConfig()); store2.start(); uint64_t seq = mgr.loadSnapshot(path, store2); assert(seq == 42); store2.stop(); fs::remove_all(dir); std::cout << "PASS: atomic write produces complete file\n"; } void test_post_write_verification_catches_truncation() { auto dir = makeSnapshotDir("verify"); SnapshotManager::Config scfg; scfg.snapshotDir = dir; scfg.validateAfterWrite = true; SnapshotManager mgr(scfg); MemoryStore store(defaultStoreConfig()); store.start(); store.createCollection("docs", CollectionOptions{}); for (int i = 0; i < 100; i++) { store.insert("docs", makeDoc("d" + std::to_string(i))); } auto path = mgr.createSnapshot(store, 1); store.stop(); // Simulate disk-level truncation after the fact uintmax_t originalSize = fs::file_size(path); fs::resize_file(path, originalSize / 2); // verifySnapshot should now return false assert(!mgr.verifySnapshot(path)); fs::remove_all(dir); std::cout << "PASS: post-write verification catches truncation\n"; } void test_loadWithFallback_uses_older_when_newest_corrupt() { auto dir = makeSnapshotDir("fallback"); SnapshotManager::Config scfg; scfg.snapshotDir = dir; scfg.validateAfterWrite = true; SnapshotManager mgr(scfg); MemoryStore store(defaultStoreConfig()); store.start(); store.createCollection("docs", CollectionOptions{}); // Create snapshot with 10 docs for (int i = 0; i < 10; i++) { store.insert("docs", makeDoc("d" + std::to_string(i))); } auto older = mgr.createSnapshot(store, 100); // Create a newer snapshot with 20 docs // Sleep to ensure the filename timestamp differs std::this_thread::sleep_for(std::chrono::seconds(1)); for (int i = 10; i < 20; i++) { store.insert("docs", makeDoc("d" + std::to_string(i))); } auto newer = mgr.createSnapshot(store, 200); store.stop(); // Corrupt the newer snapshot by truncating it fs::resize_file(newer, fs::file_size(newer) / 2); assert(!mgr.verifySnapshot(newer)); // loadWithFallback should load the older one MemoryStore store2(defaultStoreConfig()); store2.start(); fs::path used; uint64_t seq = mgr.loadWithFallback(store2, &used); assert(seq == 100); // the older snapshot's sequence assert(used == older); store2.stop(); fs::remove_all(dir); std::cout << "PASS: loadWithFallback uses older snapshot when newest is corrupt\n"; } void test_orphaned_tmp_files_cleaned_by_listSnapshots() { auto dir = makeSnapshotDir("orphan"); SnapshotManager::Config scfg; scfg.snapshotDir = dir; SnapshotManager mgr(scfg); // Create a fake .tmp file that looks like a crashed-mid-write snapshot auto tmpPath = dir / "snapshot-20260101-120000.dat.tmp"; { std::ofstream f(tmpPath); f << "partial data"; } assert(fs::exists(tmpPath)); mgr.listSnapshots(); // should clean up the .tmp file assert(!fs::exists(tmpPath)); fs::remove_all(dir); std::cout << "PASS: orphaned .tmp files cleaned by listSnapshots\n"; } int main() { test_atomic_write_produces_complete_file(); test_post_write_verification_catches_truncation(); test_loadWithFallback_uses_older_when_newest_corrupt(); test_orphaned_tmp_files_cleaned_by_listSnapshots(); std::cout << "\nAll snapshot durability tests PASSED!\n"; return 0; } ``` - [ ] **Step 3: Register test in CMakeLists.txt** Follow the exact pattern of `test_timestamp_precision`: ```cmake add_executable(test_snapshot_durability test_snapshot_durability.cpp ${CMAKE_SOURCE_DIR}/service/src/persistence/snapshot.cpp ${CMAKE_SOURCE_DIR}/service/src/memory_store.cpp # ... any other deps the existing test needs ) target_include_directories(test_snapshot_durability PRIVATE ${CMAKE_SOURCE_DIR}/service/src ) target_link_libraries(test_snapshot_durability PRIVATE nlohmann_json::nlohmann_json spdlog::spdlog Threads::Threads # match test_vector_storage link deps ) ``` Check what test_vector_storage links against and match it — the snapshot test transitively needs crc32, LZ4, config manager, etc. - [ ] **Step 4: Build and run** ```bash cmake --build build -j$(nproc) 2>&1 | tail -5 ./build/tests/test_snapshot_durability ``` Expected: "All snapshot durability tests PASSED!" - [ ] **Step 5: Commit** ```bash git add tests/test_snapshot_durability.cpp tests/CMakeLists.txt git commit -m "test(snapshot): durability + fallback integration tests - atomic write produces complete file (no .tmp leftovers) - post-write verification catches truncation - loadWithFallback uses older snapshot when newest is corrupt - orphaned .tmp files cleaned by listSnapshots at startup" ``` --- ### Task 13: Documentation **Files:** - Modify: `docs/integration-guide.md` - Modify: `CLAUDE.md` - [ ] **Step 1: Add Recovery & Read-Only section to integration-guide.md** Insert after the existing "Collection Configuration" section: ```markdown ### Recovery & Read-Only Mode smartbotic-database uses WAL + snapshots for durability. When starting up, the server tries to recover state from the newest snapshot and replay WAL entries since that snapshot. If something goes wrong, it uses a tiered recovery system modeled on MySQL's `innodb_force_recovery`. #### Recovery modes | Level | Name | Behavior | |:-:|------|----------| | 0 | `normal` | Load latest snapshot + replay WAL. **Default.** If `auto_escalate=true` (default), falls back to older snapshot automatically. | | 1 | `snapshot_fallback` | Try snapshots newest→oldest, first that loads wins. | | 2 | `wal_only` | Ignore snapshots, replay WAL from sequence 0. Slow on large DBs. | | 3 | `best_effort` | Try snapshot_fallback, then fall through to wal_only. | | 4 | `force_empty` | Start empty. Snapshots + WAL preserved on disk for forensics. | Choose a higher level (CLI or config) when lower levels fail: ```bash # Recovery refused — need to escalate smartbotic-database --recovery-mode=snapshot_fallback # Still failing — try WAL-only smartbotic-database --recovery-mode=wal_only # Last resort — accept empty start smartbotic-database --recovery-mode=force_empty ``` #### Auto-readonly on non-trivial recovery **If recovery used anything other than "latest snapshot loaded cleanly," the DB boots in read-only mode.** This is a safety feature — writes onto a possibly-stale state can corrupt your data further. The operator must explicitly acknowledge the state before writes resume. On startup you'll see: ``` [ERROR] Database booted in READ-ONLY mode after non-trivial recovery (fell back to snapshot snapshot-20260419-090719.dat because snapshot-20260419-100747.dat failed: body truncated) Writes rejected until you acknowledge this state: smartbotic-db-cli unlock # live, no restart smartbotic-database --force-readwrite # on next restart ``` #### Workflow: recovering a broken instance 1. **Don't panic** — preserve forensics. Stop the service, take a snapshot of `/var/lib/smartbotic-database/` before anything else. 2. **Diagnose** — start the server with `--read-only` and let the default recovery modes try to escalate. Use `smartbotic-db-cli status` to see what happened. 3. **Extract data** — while in read-only mode, use reads or (v1.7.0+) `smartbotic-db-cli dump` to export whatever survived. 4. **Decide**: - Data looks right → `smartbotic-db-cli unlock` to resume writes - Data is missing → don't unlock; extract to external backup, then start clean #### Client API ```cpp // Acknowledge auto-readonly state and resume writes db.unlock(); // Force read-only for maintenance (regardless of recovery outcome) db.lock(); // Inspect state + recovery outcome auto s = db.getReadOnlyStatus(); std::cout << "read-only: " << s.readOnly << "\n" << "recovery: " << s.recoveryOutcome << "\n" << "snapshot used: " << s.snapshotUsed << "\n"; ``` #### Configuration ```json "persistence": { "snapshots": { "validate_after_write": true, // verify snapshots immediately after creation "cleanup_only_if_verified": true // don't evict old snapshots if new one is bad }, "recovery": { "mode": "normal", "auto_escalate": true, // auto-fallback to older snapshot on corruption "allow_empty_on_fresh_install": true // permit empty start when no snapshots/WAL exist } } ``` #### Exit codes - `0` — normal shutdown - `1` — generic error - `10` — **recovery refused.** Snapshot or WAL failure + mode=normal + data exists on disk. Operator must escalate the recovery mode or pass `--force-readwrite`. - `11` — config invalid (e.g. bad recovery-mode value) ``` - [ ] **Step 2: Update CLAUDE.md Key Features** Add a new bullet after "Per-collection Timestamp Precision": ```markdown - **Durable Snapshots + Tiered Recovery** — atomic snapshot writer (`.tmp` + fsync + rename + post-write verification), loader fallback chain across snapshots, MySQL-style recovery modes (`normal` / `snapshot_fallback` / `wal_only` / `best_effort` / `force_empty`) selectable via `--recovery-mode` flag or `recovery.mode` config. Non-trivial recovery (fallback, WAL-only, forced empty) automatically enters **read-only mode** — operator must `smartbotic-db-cli unlock` or pass `--force-readwrite` to accept writes. Exit code `10` when recovery is refused. ``` - [ ] **Step 3: Commit** ```bash git add docs/integration-guide.md CLAUDE.md git commit -m "docs(recovery): document tiered recovery modes + auto-readonly safety" ``` --- ### Task 14: Build + publish packages - [ ] **Step 1: Local build** ```bash cd /data/smartbotic-database ./packaging/build.sh --local --skip-tests 2>&1 | tail -3 ``` - [ ] **Step 2: Docker build + publish** ```bash ./packaging/build.sh --sync --suite trixie 2>&1 | tail -10 ``` - [ ] **Step 3: Verify in repo** ```bash ls /data/smartbotics-deb-repo/repo/pool/trixie/main/*1.6.1* ``` Expected: 4 packages. - [ ] **Step 4: Push main and tag** ```bash git push origin main git tag -a v1.6.1 -m "v1.6.1: snapshot durability fix + tiered recovery + read-only mode" git push origin v1.6.1 ``` - [ ] **Step 5: Create Gogs release via Playwright** Per saved memory (`feedback_gogs_releases.md`), open Playwright at `https://git.smartbotics.ai/fszontagh/smartbotic-database/releases/new`, prompt operator to log in, fill with release notes, publish. --- ## Execution Order ``` Task 1 : VERSION + enum/struct scaffolding Task 2 : recover() signature change + applyWalEntry extraction Task 3 : ATOMIC SNAPSHOT WRITER (core fix) Task 4 : Loader fallback chain Task 5 : Full RecoveryMode logic Task 6 : Proto RPCs Task 7 : Auto-readonly on non-trivial recovery Task 8 : Write handlers reject on read-only Task 9 : Client API Task 10 : CLI flags + subcommands Task 11 : Config parsing Task 12 : Integration tests Task 13 : Docs Task 14 : Build, publish, tag, Gogs release ``` Tasks 1→5 are the durability+recovery engine. Tasks 6→10 wire it into the RPC/CLI layer. Tasks 11→14 are polish + release. ## Rollback Every commit is independently revertable. If a later task breaks a deployment, that single commit can be reverted without losing earlier fixes. The atomic-writer commit (Task 3) is the single most important one — if nothing else ships, at least shipping that one stops the silent-corruption bleeding.