For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Fix the silent snapshot-truncation bug that killed the Zoe production instance, add tiered recovery modes (like MySQL's innodb_force_recovery), add read-only runtime mode with automatic safe-mode boot when recovery is non-trivial. Urgent patch release.
Architecture: Four independent pieces that compose: (1) bulletproof snapshot writer (atomic rename + fsync + error checks + post-write verification), (2) snapshot loader with configurable fallback chain, (3) read-only runtime state gated by a new SetReadOnly RPC, (4) auto-readonly-on-non-trivial-recovery safety net. All changes are additive — default behavior on working deployments is unchanged except that recovery failure is no longer silent.
Tech Stack: C++20, POSIX fsync, std::filesystem::rename (atomic on same filesystem), existing nlohmann/json for config, existing gRPC/proto for new RPC
BUG-snapshot-truncation.md)Writer issues:
file.write() calls never check failbit — silent short writes pass throughflush() / fsync() before close() — no durability guaranteecleanupOldSnapshots() runs blindly — evicts good snapshots after a bad writeverifySnapshot() exists but is never called after writesLoader issues:
loadLatestSnapshot() tries only snapshots.front() — no fallbackpersistence->recover() return value is IGNORED at call siteProduction impact: Zoe instance restarted → latest snapshot corrupt → started empty store → ShadowMan wrote junk to blank DB → all data lost
VERSION — bump to 1.6.1proto/database.proto — new SetReadOnly / GetReadOnlyStatus RPCsservice/src/persistence/snapshot.hpp — writer API stays the same, but internal behavior changesservice/src/persistence/snapshot.cpp — atomic writer, fallback loaderservice/src/persistence/persistence_manager.hpp — recovery mode enum, auto-escalation config, recovery outcome structservice/src/persistence/persistence_manager.cpp — implement modes, auto-escalation, WAL-only replayservice/src/database_service.hpp — read-only state member, recovery outcomeservice/src/database_service.cpp — handle recovery return value, enter auto-readonly on non-trivial recovery, refuse-to-start when appropriateservice/src/database_grpc_impl.hpp — declare new RPC handlers, check isReadOnly for every write handlerservice/src/database_grpc_impl.cpp — implement SetReadOnly/GetReadOnlyStatus, reject writes when read-onlyservice/src/main.cpp — new CLI flags (--read-only, --force-readwrite, --recovery-mode)client/include/smartbotic/database/client.hpp — client methods for lock/unlockclient/src/client.cpp — implement new client methodscli/main.cpp — new unlock and lock subcommandspackaging/deb/config/config.json — document new recovery sectiontests/test_snapshot_durability.cpp — new integration testtests/CMakeLists.txt — register new test"storage": {
"persistence": {
...
"snapshots": {
"validate_after_write": true,
"cleanup_only_if_verified": true
},
"recovery": {
"mode": "normal",
"auto_escalate": true,
"allow_empty_on_fresh_install": true
}
}
}
--read-only Start in read-only mode regardless of recovery outcome
--force-readwrite Accept auto-readonly recovery state; start writable anyway
--recovery-mode=MODE Override recovery.mode config for this boot
Valid: normal, snapshot_fallback, wal_only, best_effort, force_empty
| Code | Meaning |
|---|---|
| 0 | Normal exit |
| 1 | Generic error |
| 10 | Recovery refused — snapshot/WAL corrupt + mode=normal + data exists. Operator must escalate. |
| 11 | Config invalid |
Files:
VERSIONModify: service/src/persistence/persistence_manager.hpp
[ ] Step 1: Bump VERSION to 1.6.1
echo "1.6.1" > /data/smartbotic-database/VERSION
[ ] Step 2: Add RecoveryMode enum and RecoveryOutcome struct in persistence_manager.hpp
At the top of the smartbotic::database namespace in persistence_manager.hpp, before the PersistenceManager class, add:
enum class RecoveryMode {
Normal, // level 0 — latest snapshot only, refuse on failure (default)
SnapshotFallback, // level 1 — try snapshots newest→oldest
WalOnly, // level 2 — ignore snapshots, replay WAL from 0
BestEffort, // level 3 — snapshot_fallback then wal_only
ForceEmpty // level 4 — start empty, preserve files for forensics
};
std::string recoveryModeToString(RecoveryMode mode);
RecoveryMode recoveryModeFromString(const std::string& s); // throws on invalid
/**
* Describes how recovery proceeded. A "non-trivial" outcome (anything other than
* `TrivialSuccess`) triggers auto-readonly mode at the service layer.
*/
struct RecoveryOutcome {
enum class Kind {
TrivialSuccess, // latest snapshot loaded cleanly
FreshInstall, // no snapshots, no WAL — clean state
SnapshotFellBack, // non-latest snapshot loaded (auto or manual)
WalOnlyReplay, // skipped all snapshots, replayed WAL from 0
ForcedEmpty, // started empty despite data on disk
Failed // recovery refused
};
Kind kind = Kind::TrivialSuccess;
std::filesystem::path snapshotUsed; // empty if none
std::filesystem::path expectedSnapshot; // empty if none was tried
std::string failureReason; // why the expected one didn't load
uint64_t walEntriesReplayed = 0;
size_t snapshotsAttempted = 0; // how many we tried before one worked
size_t snapshotsAvailable = 0; // total on disk
bool isNonTrivial() const {
return kind != Kind::TrivialSuccess && kind != Kind::FreshInstall;
}
bool isFailure() const { return kind == Kind::Failed; }
};
Add to PersistenceManager::Config:
struct Config {
std::filesystem::path dataDir;
uint32_t walSyncIntervalMs = 100;
uint32_t snapshotIntervalSec = 3600;
uint64_t maxWalSizeMb = 100;
bool compressionEnabled = true;
uint32_t maxSnapshots = 5;
// NEW: snapshot durability settings
bool validateAfterWrite = true;
bool cleanupOnlyIfVerified = true;
// NEW: recovery settings
RecoveryMode recoveryMode = RecoveryMode::Normal;
bool autoEscalate = true;
bool allowEmptyOnFreshInstall = true;
};
Change the recover() method signature:
// was: bool recover(MemoryStore& store);
RecoveryOutcome recover(MemoryStore& store);
The service will not compile yet — recover() callers in database_service.cpp and the implementation in persistence_manager.cpp still use the bool signature. That's expected; Task 2 fixes it.
cd /data/smartbotic-database
cmake --build build -j$(nproc) 2>&1 | grep -c "error:"
# Expect several errors about recover() signature mismatch — proceed to next task
git add VERSION service/src/persistence/persistence_manager.hpp
git commit -m "feat(recovery): RecoveryMode enum + RecoveryOutcome struct (v1.6.1 scaffold)"
Files:
Modify: service/src/persistence/persistence_manager.cpp
[ ] Step 1: Add enum string conversions at top of persistence_manager.cpp
Add inside the smartbotic::database namespace, near the top (after includes):
std::string recoveryModeToString(RecoveryMode mode) {
switch (mode) {
case RecoveryMode::Normal: return "normal";
case RecoveryMode::SnapshotFallback: return "snapshot_fallback";
case RecoveryMode::WalOnly: return "wal_only";
case RecoveryMode::BestEffort: return "best_effort";
case RecoveryMode::ForceEmpty: return "force_empty";
}
return "normal";
}
RecoveryMode recoveryModeFromString(const std::string& s) {
if (s == "normal") return RecoveryMode::Normal;
if (s == "snapshot_fallback") return RecoveryMode::SnapshotFallback;
if (s == "wal_only") return RecoveryMode::WalOnly;
if (s == "best_effort") return RecoveryMode::BestEffort;
if (s == "force_empty") return RecoveryMode::ForceEmpty;
throw std::invalid_argument("invalid recovery mode: '" + s +
"' (valid: normal, snapshot_fallback, wal_only, best_effort, force_empty)");
}
PersistenceManager::recover() signature + initial scaffoldingReplace the existing recover() implementation with a stub that still compiles but doesn't yet implement the full mode logic — just wraps the old behavior in a RecoveryOutcome. We'll flesh it out in Task 5.
Replace the entire existing bool PersistenceManager::recover(MemoryStore& store) method with:
RecoveryOutcome PersistenceManager::recover(MemoryStore& store) {
RecoveryOutcome outcome;
spdlog::info("Starting recovery (mode={}, auto_escalate={})",
recoveryModeToString(config_.recoveryMode), config_.autoEscalate);
{
std::lock_guard<std::mutex> lock(storeMutex_);
currentStore_ = &store;
}
// Enumerate available snapshots
auto snapshots = snapshotMgr_->listSnapshots();
outcome.snapshotsAvailable = snapshots.size();
// Fresh-install case: no snapshots at all
if (snapshots.empty()) {
if (config_.allowEmptyOnFreshInstall) {
// Open WAL (it may also be empty for a truly fresh install)
if (!wal_->open()) {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = "Failed to open WAL for recovery";
return outcome;
}
uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) {
applyWalEntry(store, entry);
});
outcome.walEntriesReplayed = replayed;
wal_->close();
outcome.kind = (replayed == 0)
? RecoveryOutcome::Kind::FreshInstall
: RecoveryOutcome::Kind::WalOnlyReplay; // WAL-only replay counts as non-trivial
spdlog::info("Fresh install or WAL-only recovery complete ({} WAL entries)", replayed);
return outcome;
} else {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = "No snapshots found and allow_empty_on_fresh_install is false";
return outcome;
}
}
// TODO: implement mode-specific recovery logic (Task 5).
// For now, keep the old behavior — load latest snapshot, replay WAL.
// This preserves the pre-1.6.1 behavior while we plumb the infrastructure.
try {
outcome.expectedSnapshot = snapshots.front();
uint64_t snapshotSequence = snapshotMgr_->loadSnapshot(snapshots.front(), store);
outcome.snapshotUsed = snapshots.front();
outcome.snapshotsAttempted = 1;
spdlog::info("Loaded snapshot at sequence {}", snapshotSequence);
if (!wal_->open()) {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = "Failed to open WAL after snapshot load";
return outcome;
}
uint64_t replayed = wal_->replay(snapshotSequence, [&store](const WalEntry& entry) {
applyWalEntry(store, entry);
});
outcome.walEntriesReplayed = replayed;
spdlog::info("Replayed {} WAL entries", replayed);
{
std::lock_guard<std::mutex> lock(statsMutex_);
lastSnapshotSequence_ = snapshotSequence;
}
wal_->close();
outcome.kind = RecoveryOutcome::Kind::TrivialSuccess;
spdlog::info("Recovery complete, store has {} documents in {} collections",
store.getStats().totalDocuments, store.getStats().totalCollections);
return outcome;
} catch (const std::exception& e) {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = e.what();
spdlog::error("Recovery failed: {}", e.what());
return outcome;
}
}
applyWalEntryThe current recover() has a big switch-on-WalOpType inside a lambda. Extract it to a free function inside an anonymous namespace near the top of the file (so we can reuse it for WAL-only replay in Task 5):
namespace {
void applyWalEntry(MemoryStore& store, const WalEntry& entry) {
switch (entry.opType) {
case WalOpType::INSERT:
if (entry.data) {
Document doc = Document::fromJson(*entry.data);
store.loadDocument(entry.collection, doc);
}
break;
case WalOpType::UPDATE:
case WalOpType::UPSERT:
if (entry.data) {
Document doc = Document::fromJson(*entry.data);
store.loadDocumentWithHistory(entry.collection, doc);
}
break;
case WalOpType::DELETE:
store.remove(entry.collection, entry.documentId);
break;
case WalOpType::CREATE_COLLECTION:
if (entry.collectionOptions) {
store.createCollection(entry.collection, *entry.collectionOptions);
}
break;
case WalOpType::DROP_COLLECTION:
store.dropCollection(entry.collection);
break;
case WalOpType::SET_ADD:
if (entry.data && entry.data->contains("member")) {
store.setAdd(entry.collection, entry.documentId,
(*entry.data)["member"].get<std::string>());
}
break;
case WalOpType::SET_REMOVE:
if (entry.data && entry.data->contains("member")) {
store.setRemove(entry.collection, entry.documentId,
(*entry.data)["member"].get<std::string>());
}
break;
case WalOpType::VEC_PUT:
if (entry.vectorData.has_value()) {
store.loadVector(entry.collection, entry.documentId, *entry.vectorData);
}
break;
case WalOpType::VEC_DELETE:
break; // handled by DELETE entry
}
}
} // anonymous namespace
[ ] Step 4: Build and verify tests still pass
cmake --build build -j$(nproc) 2>&1 | tail -5
# Expect remaining errors in database_service.cpp because it still uses the bool return.
# We'll fix that in Task 7. For now, comment out the call temporarily ONLY if build is
# fully blocked — normally it should compile (the function is used via pointer).
If database_service.cpp:42 fails to compile because it discards the RecoveryOutcome return value, that's fine — C++ allows discarding return values. The real fix comes in Task 7. Build should succeed.
./build/tests/test_views
LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage
Both should still pass.
[ ] Step 5: Commit
git add service/src/persistence/persistence_manager.cpp
git commit -m "feat(recovery): recover() returns RecoveryOutcome; extract applyWalEntry helper"
Files:
service/src/persistence/snapshot.cppThis is the core bug fix. Changes:
<path>.tmp firstflush() + fsync() on the file before closestd::filesystem::rename(tmp, final) atomically moves it into placefsync() the directory for durability of the renameverifySnapshot(final_path) if config says socleanupOldSnapshots() if verification passedAt the top of snapshot.cpp, after existing includes:
#include <fcntl.h>
#include <unistd.h>
SnapshotManager::Config needs validateAfterWrite and cleanupOnlyIfVerified fields. Add them to the struct in snapshot.hpp:
struct Config {
std::filesystem::path snapshotDir;
bool compressionEnabled = true;
uint32_t maxSnapshots = 5;
bool validateAfterWrite = true; // NEW
bool cleanupOnlyIfVerified = true; // NEW
};
createSnapshot() with atomic semanticsReplace the existing createSnapshot() function body (from line 24 onwards) keeping the same signature and serialization/compression/header-computation code unchanged, but replacing the write-to-disk section (currently lines 62-86):
// Generate filename and write
std::filesystem::path snapshotPath = config_.snapshotDir / generateFilename();
std::filesystem::path tmpPath = snapshotPath;
tmpPath += ".tmp";
// Open temp file
{
std::ofstream file(tmpPath, std::ios::binary | std::ios::trunc);
if (!file) {
throw std::runtime_error("Failed to create snapshot tmp file: " + tmpPath.string());
}
// Write header + check
file.write(reinterpret_cast<const char*>(&header), sizeof(header));
if (!file) {
std::error_code ec;
std::filesystem::remove(tmpPath, ec);
throw std::runtime_error("Failed to write snapshot header: " + tmpPath.string());
}
// Write body + check
file.write(reinterpret_cast<const char*>(bodyData.data()),
static_cast<std::streamsize>(bodyData.size()));
if (!file) {
std::error_code ec;
std::filesystem::remove(tmpPath, ec);
throw std::runtime_error("Failed to write snapshot body (short write): " + tmpPath.string());
}
// Compute and write body checksum + check
uint32_t bodyChecksum = crc32::calculate(bodyData.data(), bodyData.size());
file.write(reinterpret_cast<const char*>(&bodyChecksum), sizeof(bodyChecksum));
if (!file) {
std::error_code ec;
std::filesystem::remove(tmpPath, ec);
throw std::runtime_error("Failed to write snapshot trailer: " + tmpPath.string());
}
// Flush C++ stream buffer
file.flush();
if (!file) {
std::error_code ec;
std::filesystem::remove(tmpPath, ec);
throw std::runtime_error("Failed to flush snapshot: " + tmpPath.string());
}
} // std::ofstream dtor runs close()
// fsync the file for durability (C++ ofstream does not do this)
{
int fd = ::open(tmpPath.c_str(), O_RDONLY);
if (fd < 0) {
std::error_code ec;
std::filesystem::remove(tmpPath, ec);
throw std::runtime_error("Failed to open snapshot tmp for fsync: " + tmpPath.string());
}
if (::fsync(fd) != 0) {
int saved_errno = errno;
::close(fd);
std::error_code ec;
std::filesystem::remove(tmpPath, ec);
throw std::runtime_error("fsync failed on snapshot tmp (errno=" +
std::to_string(saved_errno) + "): " + tmpPath.string());
}
::close(fd);
}
// Atomic rename
{
std::error_code ec;
std::filesystem::rename(tmpPath, snapshotPath, ec);
if (ec) {
std::filesystem::remove(tmpPath, ec);
throw std::runtime_error("Snapshot rename failed: " + ec.message());
}
}
// fsync the containing directory so the rename is durable across crash
{
int dfd = ::open(config_.snapshotDir.c_str(), O_RDONLY | O_DIRECTORY);
if (dfd >= 0) {
::fsync(dfd);
::close(dfd);
}
// Non-fatal if the dirfd open fails — log only
}
// Post-write verification
if (config_.validateAfterWrite) {
if (!verifySnapshot(snapshotPath)) {
spdlog::error("Snapshot post-write verification FAILED: {}", snapshotPath.string());
std::error_code ec;
std::filesystem::remove(snapshotPath, ec);
throw std::runtime_error("Snapshot verification failed after write: " +
snapshotPath.string());
}
spdlog::debug("Snapshot verified after write: {}", snapshotPath.string());
}
// Cleanup — only if verification passed (or was disabled)
if (!config_.cleanupOnlyIfVerified || config_.validateAfterWrite) {
cleanupOldSnapshots();
}
return snapshotPath;
.tmp files on startupIn listSnapshots() at line 160, add a pre-scan that removes any *.tmp files left over from crashes. Near the top of the function, before the directory_iterator loop:
// Clean up any .tmp files left from a crash during a previous write
std::error_code ec;
for (const auto& entry : std::filesystem::directory_iterator(config_.snapshotDir, ec)) {
if (entry.is_regular_file() && entry.path().extension() == ".tmp" &&
entry.path().filename().string().starts_with("snapshot-")) {
std::error_code rm_ec;
std::filesystem::remove(entry.path(), rm_ec);
spdlog::info("Removed orphaned snapshot tmp file: {}", entry.path().string());
}
}
[ ] Step 5: Build
cmake --build build -j$(nproc) 2>&1 | tail -5
[ ] Step 6: Run existing tests
./build/tests/test_views
LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage
Both should pass — writer changes don't break anything existing.
[ ] Step 7: Commit
git add service/src/persistence/snapshot.hpp service/src/persistence/snapshot.cpp
git commit -m "fix(snapshot): atomic writer (.tmp+fsync+rename) + post-write verification
Fixes silent data loss when snapshot writes were truncated by disk pressure
or interrupted mid-write. Core changes:
- Write to <path>.tmp first, then atomically rename to final path
- Check failbit after every ofstream::write() — no more silent short writes
- flush() + fsync() file before considering the write complete
- fsync() the containing directory so the rename is durable after crash
- Call verifySnapshot() post-write; delete the file if verification fails
- cleanupOldSnapshots() only runs after successful verification
- listSnapshots() cleans up orphaned .tmp files left by crashed writes"
Files:
service/src/persistence/snapshot.cppModify: service/src/persistence/snapshot.hpp
[ ] Step 1: Add new method declaration in snapshot.hpp
Add to the SnapshotManager public section:
/**
* Attempt to load snapshots in order, trying newest first and falling
* back to older ones if the newer ones are corrupt.
*
* @param store The store to load into
* @param outUsed If non-null, set to the path of the snapshot that loaded
* @return WAL sequence from the loaded snapshot, or 0 if none loaded
*/
uint64_t loadWithFallback(MemoryStore& store,
std::filesystem::path* outUsed = nullptr);
Add the implementation after loadLatestSnapshot:
uint64_t SnapshotManager::loadWithFallback(MemoryStore& store,
std::filesystem::path* outUsed) {
auto snapshots = listSnapshots(); // already sorted newest-first
if (snapshots.empty()) {
return 0;
}
for (size_t i = 0; i < snapshots.size(); ++i) {
const auto& path = snapshots[i];
try {
uint64_t seq = loadSnapshot(path, store);
if (outUsed) *outUsed = path;
if (i > 0) {
spdlog::warn("Loaded fallback snapshot {} (newer snapshots failed): {}",
i, path.string());
}
return seq;
} catch (const std::exception& e) {
spdlog::warn("Snapshot load failed for {}: {} — trying next older snapshot",
path.string(), e.what());
// store may be in a partially-cleared state here. Next loadSnapshot
// call will call store.clear() first, so this is safe.
}
}
spdlog::error("All {} snapshots failed to load", snapshots.size());
throw std::runtime_error("all snapshots failed to load (tried " +
std::to_string(snapshots.size()) + ")");
}
[ ] Step 3: Build and test
cmake --build build -j$(nproc) 2>&1 | tail -5
./build/tests/test_views
LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage
[ ] Step 4: Commit
git add service/src/persistence/snapshot.hpp service/src/persistence/snapshot.cpp
git commit -m "fix(snapshot): loadWithFallback — try snapshots newest-first, fall back on failure"
Files:
Modify: service/src/persistence/persistence_manager.cpp
[ ] Step 1: Replace the recover() implementation with mode-aware logic
Replace the body of PersistenceManager::recover() (the one written in Task 2, Step 2) with:
RecoveryOutcome PersistenceManager::recover(MemoryStore& store) {
RecoveryOutcome outcome;
spdlog::info("Starting recovery (mode={}, auto_escalate={})",
recoveryModeToString(config_.recoveryMode), config_.autoEscalate);
{
std::lock_guard<std::mutex> lock(storeMutex_);
currentStore_ = &store;
}
auto snapshots = snapshotMgr_->listSnapshots();
outcome.snapshotsAvailable = snapshots.size();
// ForceEmpty: skip everything, start clean
if (config_.recoveryMode == RecoveryMode::ForceEmpty) {
spdlog::warn("RecoveryMode::ForceEmpty — starting with empty store. "
"Snapshot and WAL files preserved on disk for forensics.");
outcome.kind = RecoveryOutcome::Kind::ForcedEmpty;
return outcome;
}
// Fresh-install case: no snapshots at all
if (snapshots.empty()) {
if (!config_.allowEmptyOnFreshInstall) {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = "No snapshots found and allow_empty_on_fresh_install is false";
return outcome;
}
if (!wal_->open()) {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = "Failed to open WAL for recovery";
return outcome;
}
uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) {
applyWalEntry(store, entry);
});
outcome.walEntriesReplayed = replayed;
wal_->close();
outcome.kind = (replayed == 0)
? RecoveryOutcome::Kind::FreshInstall
: RecoveryOutcome::Kind::WalOnlyReplay;
return outcome;
}
outcome.expectedSnapshot = snapshots.front();
// Helper lambda: try loading a snapshot and populate outcome.snapshotUsed on success
auto tryLoadSnapshot = [&](const std::filesystem::path& path) -> std::optional<uint64_t> {
outcome.snapshotsAttempted++;
try {
uint64_t seq = snapshotMgr_->loadSnapshot(path, store);
outcome.snapshotUsed = path;
return seq;
} catch (const std::exception& e) {
if (outcome.failureReason.empty()) {
outcome.failureReason = e.what();
}
spdlog::warn("Snapshot load failed for {}: {}", path.string(), e.what());
return std::nullopt;
}
};
auto replayWal = [&](uint64_t fromSequence) -> bool {
if (!wal_->open()) {
outcome.failureReason = "Failed to open WAL after snapshot load";
return false;
}
uint64_t replayed = wal_->replay(fromSequence, [&store](const WalEntry& entry) {
applyWalEntry(store, entry);
});
outcome.walEntriesReplayed = replayed;
spdlog::info("Replayed {} WAL entries from sequence {}", replayed, fromSequence);
{
std::lock_guard<std::mutex> lock(statsMutex_);
lastSnapshotSequence_ = fromSequence;
}
wal_->close();
return true;
};
switch (config_.recoveryMode) {
case RecoveryMode::Normal: {
auto seq = tryLoadSnapshot(snapshots.front());
if (seq) {
if (!replayWal(*seq)) {
outcome.kind = RecoveryOutcome::Kind::Failed;
return outcome;
}
outcome.kind = RecoveryOutcome::Kind::TrivialSuccess;
return outcome;
}
// Latest snapshot failed. Can we auto-escalate to snapshot_fallback?
if (config_.autoEscalate && snapshots.size() > 1) {
spdlog::error("Latest snapshot failed; auto-escalating to snapshot_fallback");
for (size_t i = 1; i < snapshots.size(); ++i) {
auto s = tryLoadSnapshot(snapshots[i]);
if (s) {
if (!replayWal(*s)) {
outcome.kind = RecoveryOutcome::Kind::Failed;
return outcome;
}
outcome.kind = RecoveryOutcome::Kind::SnapshotFellBack;
return outcome;
}
}
}
// Auto-escalation disabled or exhausted
outcome.kind = RecoveryOutcome::Kind::Failed;
return outcome;
}
case RecoveryMode::SnapshotFallback: {
for (const auto& path : snapshots) {
auto seq = tryLoadSnapshot(path);
if (seq) {
if (!replayWal(*seq)) {
outcome.kind = RecoveryOutcome::Kind::Failed;
return outcome;
}
outcome.kind = (path == snapshots.front())
? RecoveryOutcome::Kind::TrivialSuccess
: RecoveryOutcome::Kind::SnapshotFellBack;
return outcome;
}
}
outcome.kind = RecoveryOutcome::Kind::Failed;
return outcome;
}
case RecoveryMode::WalOnly: {
store.clear();
if (!wal_->open()) {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = "Failed to open WAL";
return outcome;
}
uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) {
applyWalEntry(store, entry);
});
outcome.walEntriesReplayed = replayed;
wal_->close();
outcome.kind = RecoveryOutcome::Kind::WalOnlyReplay;
return outcome;
}
case RecoveryMode::BestEffort: {
// Try snapshot fallback
for (const auto& path : snapshots) {
auto seq = tryLoadSnapshot(path);
if (seq) {
if (!replayWal(*seq)) {
outcome.kind = RecoveryOutcome::Kind::Failed;
return outcome;
}
outcome.kind = (path == snapshots.front())
? RecoveryOutcome::Kind::TrivialSuccess
: RecoveryOutcome::Kind::SnapshotFellBack;
return outcome;
}
}
// Fall through to WAL-only
spdlog::error("All snapshots failed in best_effort mode; attempting WAL-only replay");
store.clear();
if (!wal_->open()) {
outcome.kind = RecoveryOutcome::Kind::Failed;
outcome.failureReason = "Failed to open WAL (best_effort fallback)";
return outcome;
}
uint64_t replayed = wal_->replay(0, [&store](const WalEntry& entry) {
applyWalEntry(store, entry);
});
outcome.walEntriesReplayed = replayed;
wal_->close();
outcome.kind = RecoveryOutcome::Kind::WalOnlyReplay;
return outcome;
}
case RecoveryMode::ForceEmpty:
// handled at the top
break;
}
outcome.kind = RecoveryOutcome::Kind::Failed;
return outcome;
}
[ ] Step 2: Build and run existing tests
cmake --build build -j$(nproc) 2>&1 | tail -5
./build/tests/test_views
LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage
All should pass — default behavior (RecoveryMode::Normal + auto_escalate=true) matches the old behavior for healthy deployments.
[ ] Step 3: Commit
git add service/src/persistence/persistence_manager.cpp
git commit -m "feat(recovery): implement all 5 RecoveryMode cases with auto-escalation
- Normal: latest snapshot only; if fails AND auto_escalate, try older
- SnapshotFallback: walk snapshots newest-first, first that loads wins
- WalOnly: skip all snapshots, replay WAL from 0
- BestEffort: snapshot_fallback + fall through to WalOnly if nothing loads
- ForceEmpty: skip everything, preserve files for forensics
RecoveryOutcome.kind tells the caller whether recovery was trivial,
non-trivial, or failed. The service layer uses this to decide auto-readonly."
Files:
Modify: proto/database.proto
[ ] Step 1: Add RPC declarations
In the DatabaseService service block, after the existing health/stats RPCs, add:
// Read-only runtime state control
rpc SetReadOnly(SetReadOnlyRequest) returns (SetReadOnlyResponse);
rpc GetReadOnlyStatus(GetReadOnlyStatusRequest) returns (GetReadOnlyStatusResponse);
At the end of the file:
// ===== Read-Only Control =====
message SetReadOnlyRequest {
bool read_only = 1;
}
message SetReadOnlyResponse {
bool success = 1;
bool was_read_only = 2; // previous state
string error = 3;
}
message GetReadOnlyStatusRequest {}
message GetReadOnlyStatusResponse {
bool read_only = 1;
string reason = 2; // empty when !read_only; explains why otherwise
// When auto-readonly was triggered by recovery, these fields describe it
string recovery_outcome = 3; // "trivial_success", "snapshot_fell_back", "wal_only_replay", "forced_empty", "failed"
string expected_snapshot = 4; // the snapshot we tried to load first
string snapshot_used = 5; // the one that actually loaded
string failure_reason = 6; // why expected_snapshot failed
uint64 wal_entries_replayed = 7;
uint32 snapshots_attempted = 8;
}
[ ] Step 3: Build to confirm proto compiles
cmake --build build -j$(nproc) 2>&1 | grep -E "error|Generating database" | head -5
Expect Generating protobuf/gRPC code for database.proto and no errors.
[ ] Step 4: Commit
git add proto/database.proto
git commit -m "feat(proto): add SetReadOnly + GetReadOnlyStatus RPCs for runtime lock control"
Files:
service/src/database_service.hppModify: service/src/database_service.cpp
[ ] Step 1: Add read-only state to DatabaseService
In database_service.hpp, add to the DatabaseService class:
#include <atomic>
// Public section:
bool isReadOnly() const { return read_only_.load(std::memory_order_acquire); }
std::string readOnlyReason() const;
void setReadOnly(bool value, const std::string& reason);
const RecoveryOutcome& recoveryOutcome() const { return recovery_outcome_; }
// Private members:
std::atomic<bool> read_only_{false};
mutable std::mutex reason_mutex_;
std::string read_only_reason_;
RecoveryOutcome recovery_outcome_;
bool force_readwrite_ = false; // set from CLI / config
Add member function implementations:
std::string DatabaseService::readOnlyReason() const {
std::lock_guard<std::mutex> lock(reason_mutex_);
return read_only_reason_;
}
void DatabaseService::setReadOnly(bool value, const std::string& reason) {
{
std::lock_guard<std::mutex> lock(reason_mutex_);
read_only_reason_ = value ? reason : "";
}
bool prev = read_only_.exchange(value, std::memory_order_acq_rel);
if (prev != value) {
if (value) {
spdlog::error("Database entered READ-ONLY mode: {}", reason);
} else {
spdlog::info("Database unlocked — writes accepted");
}
}
}
Find the call site at database_service.cpp:42 (persistence_->recover(*store_);) and replace with:
// Recover from persistence
recovery_outcome_ = persistence_->recover(*store_);
if (recovery_outcome_.isFailure()) {
spdlog::error("\n"
"┌────────────────────────────────────────────────────────────────────┐\n"
"│ RECOVERY REFUSED │\n"
"│ │\n"
"│ Mode: {}\n"
"│ Expected snapshot: {}\n"
"│ Reason: {}\n"
"│ │\n"
"│ To escalate, restart with one of: │\n"
"│ --recovery-mode=snapshot_fallback Try older snapshots │\n"
"│ --recovery-mode=wal_only Replay WAL only (slow) │\n"
"│ --recovery-mode=best_effort Try all of the above │\n"
"│ --recovery-mode=force_empty Start empty (LAST RESORT) │\n"
"│ │\n"
"│ Or add to config.json: \"recovery\": {{ \"mode\": \"<mode>\" }} │\n"
"│ │\n"
"│ Preserve /var/lib/smartbotic-database/ BEFORE escalating. │\n"
"└────────────────────────────────────────────────────────────────────┘",
recoveryModeToString(config_.persistenceConfig.recoveryMode),
recovery_outcome_.expectedSnapshot.string(),
recovery_outcome_.failureReason);
throw std::runtime_error("recovery failed; refusing to start with exit code 10");
}
if (recovery_outcome_.isNonTrivial() && !force_readwrite_) {
// Enter auto-readonly mode
std::string reason = "Non-trivial recovery occurred (";
switch (recovery_outcome_.kind) {
case RecoveryOutcome::Kind::SnapshotFellBack:
reason += "fell back to snapshot " + recovery_outcome_.snapshotUsed.filename().string() +
" because " + recovery_outcome_.expectedSnapshot.filename().string() +
" failed: " + recovery_outcome_.failureReason;
break;
case RecoveryOutcome::Kind::WalOnlyReplay:
reason += "WAL-only replay, no snapshot loaded (" +
std::to_string(recovery_outcome_.walEntriesReplayed) + " entries)";
break;
case RecoveryOutcome::Kind::ForcedEmpty:
reason += "forced empty by operator (--recovery-mode=force_empty)";
break;
default: break;
}
reason += "). Run `smartbotic-db-cli unlock` to accept this state, or restart "
"with --force-readwrite to bypass this check.";
setReadOnly(true, reason);
spdlog::error("\n"
"┌────────────────────────────────────────────────────────────────────┐\n"
"│ [ERROR] Database booted in READ-ONLY mode after non-trivial recovery│\n"
"│ │\n"
"│ {:66} │\n"
"│ │\n"
"│ Writes will be REJECTED until you acknowledge this state: │\n"
"│ smartbotic-db-cli unlock # live, no restart│\n"
"│ smartbotic-database --force-readwrite # on next restart │\n"
"│ │\n"
"│ To extract data before deciding: │\n"
"│ smartbotic-db-cli dump > backup-$(date +%s).json (v1.7.0) │\n"
"└────────────────────────────────────────────────────────────────────┘",
recoveryModeToString(config_.persistenceConfig.recoveryMode));
}
// Set replication sequence after recovery
replication_->setSequence(persistence_->currentWalSequence());
Note: the exit code 10 for recovery-refused is achieved via exception handling in main.cpp — see Task 10.
[ ] Step 4: Build and verify defaults unchanged
cmake --build build -j$(nproc) 2>&1 | tail -5
./build/tests/test_views
LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage
All should pass. On clean deployment with valid snapshots, recovery_outcome.kind == TrivialSuccess and no read-only enforcement.
[ ] Step 5: Commit
git add service/src/database_service.hpp service/src/database_service.cpp
git commit -m "feat(recovery): DatabaseService consumes RecoveryOutcome; auto-readonly on non-trivial
- setReadOnly/isReadOnly/readOnlyReason — runtime lock state (atomic)
- Recovery failure throws → main.cpp catches → exit code 10
- Non-trivial recovery (snapshot fallback, WAL-only, forced empty) enters
read-only with an ERROR-level log explaining exactly what happened
- --force-readwrite flag (Task 10) bypasses the auto-readonly"
Files:
service/src/database_grpc_impl.hppModify: service/src/database_grpc_impl.cpp
[ ] Step 1: Add DatabaseService& reference to DatabaseGrpcImpl
DatabaseGrpcImpl needs to consult DatabaseService for isReadOnly(). Add a constructor parameter. In database_grpc_impl.hpp:
Add forward declaration at top of namespace:
class DatabaseService;
Add constructor parameter and member:
// Constructor gains a new first parameter:
DatabaseGrpcImpl(DatabaseService& service, ...all existing params...);
private:
DatabaseService& service_; // for isReadOnly() checks
In database_grpc_impl.cpp, add a private helper at the top of the class impl:
namespace {
// Returns a grpc::Status that should be returned when the DB is read-only.
grpc::Status readOnlyRejection(const std::string& rpc, const DatabaseService& svc) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION,
rpc + " rejected: database is in read-only mode. " + svc.readOnlyReason());
}
} // anonymous namespace
Add this check at the very top of each write-path handler (after parameter validation but before any store access):
Handlers to modify:
Insert — add checkUpdate — add checkUpsert — add checkDelete — add checkPatchDocument — add checkBatchInsert — add checkBatchDelete — add checkCreateCollection — add checkDropCollection — add checkCreateView — add checkDropView — add checkConfigureCollection — add checkMigrateCollectionTimestamps — add checkSetAdd / SetRemove — add checkUploadFile / DeleteFile — add checkFor handlers whose response has set_success/set_error:
if (service_.isReadOnly()) {
response->set_success(false);
response->set_error("database is in read-only mode: " + service_.readOnlyReason());
return grpc::Status::OK;
}
For handlers without those fields:
if (service_.isReadOnly()) {
return readOnlyRejection("Insert", service_); // use actual RPC name
}
IMPORTANT: read every write handler before modifying to pick the right style. Exception: streaming handlers like UploadFile should reject via a stream status error — check the existing streaming error pattern.
Add at the bottom of database_grpc_impl.cpp:
grpc::Status DatabaseGrpcImpl::SetReadOnly(
grpc::ServerContext* /*context*/,
const pb::SetReadOnlyRequest* request,
pb::SetReadOnlyResponse* response
) {
bool was_readonly = service_.isReadOnly();
std::string reason = request->read_only()
? "manually locked via SetReadOnly RPC"
: ""; // reason cleared on unlock
service_.setReadOnly(request->read_only(), reason);
response->set_success(true);
response->set_was_read_only(was_readonly);
return grpc::Status::OK;
}
grpc::Status DatabaseGrpcImpl::GetReadOnlyStatus(
grpc::ServerContext* /*context*/,
const pb::GetReadOnlyStatusRequest* /*request*/,
pb::GetReadOnlyStatusResponse* response
) {
response->set_read_only(service_.isReadOnly());
response->set_reason(service_.readOnlyReason());
const auto& out = service_.recoveryOutcome();
switch (out.kind) {
case RecoveryOutcome::Kind::TrivialSuccess: response->set_recovery_outcome("trivial_success"); break;
case RecoveryOutcome::Kind::FreshInstall: response->set_recovery_outcome("fresh_install"); break;
case RecoveryOutcome::Kind::SnapshotFellBack: response->set_recovery_outcome("snapshot_fell_back"); break;
case RecoveryOutcome::Kind::WalOnlyReplay: response->set_recovery_outcome("wal_only_replay"); break;
case RecoveryOutcome::Kind::ForcedEmpty: response->set_recovery_outcome("forced_empty"); break;
case RecoveryOutcome::Kind::Failed: response->set_recovery_outcome("failed"); break;
}
response->set_expected_snapshot(out.expectedSnapshot.string());
response->set_snapshot_used(out.snapshotUsed.string());
response->set_failure_reason(out.failureReason);
response->set_wal_entries_replayed(out.walEntriesReplayed);
response->set_snapshots_attempted(out.snapshotsAttempted);
return grpc::Status::OK;
}
Declare them in database_grpc_impl.hpp:
grpc::Status SetReadOnly(grpc::ServerContext*, const pb::SetReadOnlyRequest*,
pb::SetReadOnlyResponse*) override;
grpc::Status GetReadOnlyStatus(grpc::ServerContext*, const pb::GetReadOnlyStatusRequest*,
pb::GetReadOnlyStatusResponse*) override;
Find where DatabaseGrpcImpl is constructed in database_service.cpp (search for std::make_unique<DatabaseGrpcImpl> or similar) and pass *this as the new first argument.
[ ] Step 6: Build and run tests
cmake --build build -j$(nproc) 2>&1 | tail -10
./build/tests/test_views
LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage
[ ] Step 7: Commit
git add service/src/database_grpc_impl.hpp service/src/database_grpc_impl.cpp service/src/database_service.cpp
git commit -m "feat(readonly): reject writes in gRPC handlers when DB is read-only
- New SetReadOnly/GetReadOnlyStatus RPCs let clients toggle + inspect state
- Every write handler (insert/update/patch/upsert/delete/batch/collection
management/view management/config management) checks service.isReadOnly()
and returns FAILED_PRECONDITION with a clear reason
- Read handlers (Get, Find, Count, GetVersionHistory, SimilaritySearch,
ListCollections, ListViews, etc.) unaffected — read-only means read-ONLY"
Files:
client/include/smartbotic/database/client.hppModify: client/src/client.cpp
[ ] Step 1: Add declarations in client.hpp
Add to the Client class public section (near other admin methods):
// ===== Read-Only Control =====
struct ReadOnlyStatus {
bool readOnly = false;
std::string reason;
std::string recoveryOutcome; // "trivial_success", "snapshot_fell_back", etc.
std::string expectedSnapshot;
std::string snapshotUsed;
std::string failureReason;
uint64_t walEntriesReplayed = 0;
uint32_t snapshotsAttempted = 0;
};
/**
* Toggle server-side read-only mode. When on, all writes are rejected
* with FAILED_PRECONDITION. Useful for draining, maintenance, or
* acknowledging a non-trivial recovery state.
* @return true if the call succeeded (regardless of previous state)
*/
bool setReadOnly(bool readOnly);
/**
* Shortcut for setReadOnly(true).
*/
bool lock() { return setReadOnly(true); }
/**
* Shortcut for setReadOnly(false). Typically used after reviewing a
* non-trivial recovery state and deciding to accept it.
*/
bool unlock() { return setReadOnly(false); }
/**
* Inspect the current read-only state and (if applicable) the recovery
* outcome that caused auto-readonly.
*/
[[nodiscard]] ReadOnlyStatus getReadOnlyStatus();
Inside the Impl class:
bool setReadOnly(bool readOnly) {
smartbotic::databasepb::SetReadOnlyRequest request;
request.set_read_only(readOnly);
smartbotic::databasepb::SetReadOnlyResponse response;
grpc::ClientContext context;
setDeadline(context);
auto status = stub_->SetReadOnly(&context, request, &response);
if (!status.ok()) {
spdlog::error("Client::setReadOnly failed: {}", status.error_message());
return false;
}
return response.success();
}
Client::ReadOnlyStatus getReadOnlyStatus() {
smartbotic::databasepb::GetReadOnlyStatusRequest request;
smartbotic::databasepb::GetReadOnlyStatusResponse response;
grpc::ClientContext context;
setDeadline(context);
Client::ReadOnlyStatus out;
auto status = stub_->GetReadOnlyStatus(&context, request, &response);
if (!status.ok()) {
spdlog::error("Client::getReadOnlyStatus failed: {}", status.error_message());
return out;
}
out.readOnly = response.read_only();
out.reason = response.reason();
out.recoveryOutcome = response.recovery_outcome();
out.expectedSnapshot = response.expected_snapshot();
out.snapshotUsed = response.snapshot_used();
out.failureReason = response.failure_reason();
out.walEntriesReplayed = response.wal_entries_replayed();
out.snapshotsAttempted = response.snapshots_attempted();
return out;
}
Add public forwarders at the bottom of client.cpp:
bool Client::setReadOnly(bool readOnly) {
return impl_->setReadOnly(readOnly);
}
Client::ReadOnlyStatus Client::getReadOnlyStatus() {
return impl_->getReadOnlyStatus();
}
[ ] Step 3: Build and test
cmake --build build -j$(nproc) 2>&1 | tail -5
./build/tests/test_views
LD_LIBRARY_PATH=build/client ./build/tests/test_vector_storage
[ ] Step 4: Commit
git add client/include/smartbotic/database/client.hpp client/src/client.cpp
git commit -m "feat(client): setReadOnly/lock/unlock/getReadOnlyStatus client API"
Files:
service/src/main.cppModify: cli/main.cpp
[ ] Step 1: Read existing CLI parsing patterns in main.cpp
head -150 /data/smartbotic-database/service/src/main.cpp
head -100 /data/smartbotic-database/cli/main.cpp
Match existing styles.
In printUsage(), add below the existing --config line:
<< " --read-only Start in read-only mode (reject all writes)\n"
<< " --force-readwrite Accept non-trivial recovery state; start writable\n"
<< " --recovery-mode=MODE Override recovery mode for this boot\n"
<< " Values: normal, snapshot_fallback, wal_only,\n"
<< " best_effort, force_empty\n"
Before parsing the config file, parse these flags. Add after findConfigFile:
struct CliOverrides {
std::optional<bool> readOnly;
std::optional<bool> forceReadwrite;
std::optional<RecoveryMode> recoveryMode;
};
CliOverrides parseCliOverrides(int argc, char* argv[]) {
CliOverrides ov;
for (int i = 1; i < argc; i++) {
std::string arg = argv[i];
if (arg == "--read-only") ov.readOnly = true;
else if (arg == "--force-readwrite") ov.forceReadwrite = true;
else if (arg.rfind("--recovery-mode=", 0) == 0) {
std::string mode = arg.substr(16);
try {
ov.recoveryMode = recoveryModeFromString(mode);
} catch (const std::exception& e) {
std::cerr << "ERROR: " << e.what() << std::endl;
std::exit(11); // config invalid
}
}
}
return ov;
}
In main(), after loading config but before creating DatabaseService:
auto cli = parseCliOverrides(argc, argv);
if (cli.recoveryMode) {
config.storage.persistence.recoveryMode = *cli.recoveryMode;
}
if (cli.readOnly.value_or(false)) {
config.storage.initialReadOnly = true;
}
if (cli.forceReadwrite.value_or(false)) {
config.storage.forceReadwrite = true;
}
Note: this adds two fields to the config struct — initialReadOnly and forceReadwrite. Add them to the config parsing in database_service.hpp/cpp. These are runtime overrides, not persisted config.
Wrap DatabaseService::initialize() in a try/catch that exits with code 10 for recovery-refused:
try {
db.initialize();
} catch (const std::exception& e) {
if (std::string(e.what()).find("recovery failed") != std::string::npos) {
std::cerr << "Recovery refused: " << e.what() << std::endl;
return 10;
}
std::cerr << "Initialization failed: " << e.what() << std::endl;
return 1;
}
lock, unlock, status subcommands to cli/main.cppRead the existing subcommand structure:
grep -n "std::string subcommand\|argv\[1\]\|subcommand ==\|if.*argv" /data/smartbotic-database/cli/main.cpp | head -10
Add three new subcommand branches:
else if (subcommand == "lock") {
if (client.setReadOnly(true)) {
std::cout << "Database locked (read-only)" << std::endl;
return 0;
} else {
std::cerr << "Failed to lock database" << std::endl;
return 1;
}
}
else if (subcommand == "unlock") {
if (client.setReadOnly(false)) {
std::cout << "Database unlocked (writes accepted)" << std::endl;
return 0;
} else {
std::cerr << "Failed to unlock database" << std::endl;
return 1;
}
}
else if (subcommand == "status") {
auto s = client.getReadOnlyStatus();
std::cout << "Read-only: " << (s.readOnly ? "YES" : "no") << std::endl;
if (s.readOnly) {
std::cout << "Reason: " << s.reason << std::endl;
}
std::cout << "Recovery outcome: " << s.recoveryOutcome << std::endl;
if (!s.expectedSnapshot.empty()) {
std::cout << "Expected snapshot: " << s.expectedSnapshot << std::endl;
}
if (!s.snapshotUsed.empty()) {
std::cout << "Snapshot used: " << s.snapshotUsed << std::endl;
}
if (!s.failureReason.empty()) {
std::cout << "Failure reason: " << s.failureReason << std::endl;
}
std::cout << "WAL replayed: " << s.walEntriesReplayed << " entries" << std::endl;
std::cout << "Snapshots tried: " << s.snapshotsAttempted << std::endl;
return 0;
}
Update help text to list lock, unlock, status.
[ ] Step 4: Build and test
cmake --build build -j$(nproc) 2>&1 | tail -5
[ ] Step 5: Commit
git add service/src/main.cpp cli/main.cpp service/src/database_service.hpp
git commit -m "feat(cli): --read-only/--force-readwrite/--recovery-mode flags + lock/unlock/status subcommands
Server:
- --read-only: start in read-only state
- --force-readwrite: bypass auto-readonly from non-trivial recovery
- --recovery-mode=X: override config for this boot
- Exit code 10 when recovery is refused
CLI:
- smartbotic-db-cli lock → SetReadOnly(true)
- smartbotic-db-cli unlock → SetReadOnly(false)
- smartbotic-db-cli status → detailed read-only + recovery info"
Files:
service/src/database_service.cpp (config loader)Modify: packaging/deb/config/config.json
[ ] Step 1: Find config parser
grep -n "persistence\|recovery\|walSyncIntervalMs" /data/smartbotic-database/service/src/database_service.cpp | head -10
[ ] Step 2: Parse new fields
Where persistence settings are parsed, after existing fields, add:
// Snapshot durability
if (storage.contains("persistence") && storage["persistence"].contains("snapshots")) {
const auto& snap = storage["persistence"]["snapshots"];
config_.persistenceConfig.validateAfterWrite =
snap.value("validate_after_write", true);
config_.persistenceConfig.cleanupOnlyIfVerified =
snap.value("cleanup_only_if_verified", true);
}
// Recovery
if (storage.contains("persistence") && storage["persistence"].contains("recovery")) {
const auto& rec = storage["persistence"]["recovery"];
std::string modeStr = rec.value("mode", std::string("normal"));
try {
config_.persistenceConfig.recoveryMode = recoveryModeFromString(modeStr);
} catch (const std::exception& e) {
spdlog::warn("Invalid recovery.mode '{}', defaulting to 'normal': {}", modeStr, e.what());
config_.persistenceConfig.recoveryMode = RecoveryMode::Normal;
}
config_.persistenceConfig.autoEscalate = rec.value("auto_escalate", true);
config_.persistenceConfig.allowEmptyOnFreshInstall = rec.value("allow_empty_on_fresh_install", true);
}
In packaging/deb/config/config.json, extend the persistence block:
"persistence": {
"wal_sync_interval_ms": 100,
"snapshot_interval_sec": 3600,
"compression": "lz4",
"snapshots": {
"validate_after_write": true,
"cleanup_only_if_verified": true
},
"recovery": {
"mode": "normal",
"auto_escalate": true,
"allow_empty_on_fresh_install": true
}
}
[ ] Step 4: Build and verify
cmake --build build -j$(nproc) 2>&1 | tail -5
[ ] Step 5: Commit
git add service/src/database_service.cpp packaging/deb/config/config.json
git commit -m "feat(config): parse persistence.snapshots and persistence.recovery sections"
Files:
tests/test_snapshot_durability.cppModify: tests/CMakeLists.txt
[ ] Step 1: Read test harness patterns
cat /data/smartbotic-database/tests/test_timestamp_precision.cpp | head -30
cat /data/smartbotic-database/tests/CMakeLists.txt
Follow exactly the pattern used for test_timestamp_precision.
[ ] Step 2: Write tests/test_snapshot_durability.cpp
#include "../service/src/persistence/snapshot.hpp"
#include "../service/src/memory_store.hpp"
#include "../service/src/document.hpp"
#include <cassert>
#include <cstdint>
#include <filesystem>
#include <fstream>
#include <iostream>
#include <nlohmann/json.hpp>
namespace fs = std::filesystem;
using smartbotic::database::MemoryStore;
using smartbotic::database::SnapshotManager;
using smartbotic::database::Document;
using smartbotic::database::CollectionOptions;
namespace {
MemoryStore::Config defaultStoreConfig() {
MemoryStore::Config cfg;
cfg.nodeId = "test";
cfg.maxMemoryBytes = 256 * 1024 * 1024;
return cfg;
}
fs::path makeSnapshotDir(const std::string& name) {
auto dir = fs::temp_directory_path() / ("snap-test-" + name + "-" +
std::to_string(::getpid()));
fs::remove_all(dir);
fs::create_directories(dir);
return dir;
}
Document makeDoc(const std::string& id) {
Document d;
d.id = id;
d.data = nlohmann::json{{"value", id}};
return d;
}
} // anonymous namespace
void test_atomic_write_produces_complete_file() {
auto dir = makeSnapshotDir("atomic");
SnapshotManager::Config scfg;
scfg.snapshotDir = dir;
scfg.validateAfterWrite = true;
SnapshotManager mgr(scfg);
MemoryStore store(defaultStoreConfig());
store.start();
store.createCollection("docs", CollectionOptions{});
for (int i = 0; i < 100; i++) {
store.insert("docs", makeDoc("d" + std::to_string(i)));
}
auto path = mgr.createSnapshot(store, 42);
store.stop();
// No .tmp files left
int tmpCount = 0;
for (auto& e : fs::directory_iterator(dir)) {
if (e.path().extension() == ".tmp") tmpCount++;
}
assert(tmpCount == 0);
// Verify file passes integrity check
assert(mgr.verifySnapshot(path));
// Load into a fresh store
MemoryStore store2(defaultStoreConfig());
store2.start();
uint64_t seq = mgr.loadSnapshot(path, store2);
assert(seq == 42);
store2.stop();
fs::remove_all(dir);
std::cout << "PASS: atomic write produces complete file\n";
}
void test_post_write_verification_catches_truncation() {
auto dir = makeSnapshotDir("verify");
SnapshotManager::Config scfg;
scfg.snapshotDir = dir;
scfg.validateAfterWrite = true;
SnapshotManager mgr(scfg);
MemoryStore store(defaultStoreConfig());
store.start();
store.createCollection("docs", CollectionOptions{});
for (int i = 0; i < 100; i++) {
store.insert("docs", makeDoc("d" + std::to_string(i)));
}
auto path = mgr.createSnapshot(store, 1);
store.stop();
// Simulate disk-level truncation after the fact
uintmax_t originalSize = fs::file_size(path);
fs::resize_file(path, originalSize / 2);
// verifySnapshot should now return false
assert(!mgr.verifySnapshot(path));
fs::remove_all(dir);
std::cout << "PASS: post-write verification catches truncation\n";
}
void test_loadWithFallback_uses_older_when_newest_corrupt() {
auto dir = makeSnapshotDir("fallback");
SnapshotManager::Config scfg;
scfg.snapshotDir = dir;
scfg.validateAfterWrite = true;
SnapshotManager mgr(scfg);
MemoryStore store(defaultStoreConfig());
store.start();
store.createCollection("docs", CollectionOptions{});
// Create snapshot with 10 docs
for (int i = 0; i < 10; i++) {
store.insert("docs", makeDoc("d" + std::to_string(i)));
}
auto older = mgr.createSnapshot(store, 100);
// Create a newer snapshot with 20 docs
// Sleep to ensure the filename timestamp differs
std::this_thread::sleep_for(std::chrono::seconds(1));
for (int i = 10; i < 20; i++) {
store.insert("docs", makeDoc("d" + std::to_string(i)));
}
auto newer = mgr.createSnapshot(store, 200);
store.stop();
// Corrupt the newer snapshot by truncating it
fs::resize_file(newer, fs::file_size(newer) / 2);
assert(!mgr.verifySnapshot(newer));
// loadWithFallback should load the older one
MemoryStore store2(defaultStoreConfig());
store2.start();
fs::path used;
uint64_t seq = mgr.loadWithFallback(store2, &used);
assert(seq == 100); // the older snapshot's sequence
assert(used == older);
store2.stop();
fs::remove_all(dir);
std::cout << "PASS: loadWithFallback uses older snapshot when newest is corrupt\n";
}
void test_orphaned_tmp_files_cleaned_by_listSnapshots() {
auto dir = makeSnapshotDir("orphan");
SnapshotManager::Config scfg;
scfg.snapshotDir = dir;
SnapshotManager mgr(scfg);
// Create a fake .tmp file that looks like a crashed-mid-write snapshot
auto tmpPath = dir / "snapshot-20260101-120000.dat.tmp";
{
std::ofstream f(tmpPath);
f << "partial data";
}
assert(fs::exists(tmpPath));
mgr.listSnapshots(); // should clean up the .tmp file
assert(!fs::exists(tmpPath));
fs::remove_all(dir);
std::cout << "PASS: orphaned .tmp files cleaned by listSnapshots\n";
}
int main() {
test_atomic_write_produces_complete_file();
test_post_write_verification_catches_truncation();
test_loadWithFallback_uses_older_when_newest_corrupt();
test_orphaned_tmp_files_cleaned_by_listSnapshots();
std::cout << "\nAll snapshot durability tests PASSED!\n";
return 0;
}
[ ] Step 3: Register test in CMakeLists.txt
Follow the exact pattern of test_timestamp_precision:
add_executable(test_snapshot_durability
test_snapshot_durability.cpp
${CMAKE_SOURCE_DIR}/service/src/persistence/snapshot.cpp
${CMAKE_SOURCE_DIR}/service/src/memory_store.cpp
# ... any other deps the existing test needs
)
target_include_directories(test_snapshot_durability PRIVATE
${CMAKE_SOURCE_DIR}/service/src
)
target_link_libraries(test_snapshot_durability PRIVATE
nlohmann_json::nlohmann_json
spdlog::spdlog
Threads::Threads
# match test_vector_storage link deps
)
Check what test_vector_storage links against and match it — the snapshot test transitively needs crc32, LZ4, config manager, etc.
[ ] Step 4: Build and run
cmake --build build -j$(nproc) 2>&1 | tail -5
./build/tests/test_snapshot_durability
Expected: "All snapshot durability tests PASSED!"
[ ] Step 5: Commit
git add tests/test_snapshot_durability.cpp tests/CMakeLists.txt
git commit -m "test(snapshot): durability + fallback integration tests
- atomic write produces complete file (no .tmp leftovers)
- post-write verification catches truncation
- loadWithFallback uses older snapshot when newest is corrupt
- orphaned .tmp files cleaned by listSnapshots at startup"
Files:
docs/integration-guide.mdModify: CLAUDE.md
[ ] Step 1: Add Recovery & Read-Only section to integration-guide.md
Insert after the existing "Collection Configuration" section:
### Recovery & Read-Only Mode
smartbotic-database uses WAL + snapshots for durability. When starting up, the server tries to recover state from the newest snapshot and replay WAL entries since that snapshot. If something goes wrong, it uses a tiered recovery system modeled on MySQL's `innodb_force_recovery`.
#### Recovery modes
| Level | Name | Behavior |
|:-:|------|----------|
| 0 | `normal` | Load latest snapshot + replay WAL. **Default.** If `auto_escalate=true` (default), falls back to older snapshot automatically. |
| 1 | `snapshot_fallback` | Try snapshots newest→oldest, first that loads wins. |
| 2 | `wal_only` | Ignore snapshots, replay WAL from sequence 0. Slow on large DBs. |
| 3 | `best_effort` | Try snapshot_fallback, then fall through to wal_only. |
| 4 | `force_empty` | Start empty. Snapshots + WAL preserved on disk for forensics. |
Choose a higher level (CLI or config) when lower levels fail:
bash
smartbotic-database --recovery-mode=snapshot_fallback
smartbotic-database --recovery-mode=wal_only
smartbotic-database --recovery-mode=force_empty
#### Auto-readonly on non-trivial recovery
**If recovery used anything other than "latest snapshot loaded cleanly," the DB boots in read-only mode.** This is a safety feature — writes onto a possibly-stale state can corrupt your data further. The operator must explicitly acknowledge the state before writes resume.
On startup you'll see:
[ERROR] Database booted in READ-ONLY mode after non-trivial recovery
(fell back to snapshot snapshot-20260419-090719.dat because
snapshot-20260419-100747.dat failed: body truncated)
Writes rejected until you acknowledge this state: smartbotic-db-cli unlock # live, no restart smartbotic-database --force-readwrite # on next restart
#### Workflow: recovering a broken instance
1. **Don't panic** — preserve forensics. Stop the service, take a snapshot of `/var/lib/smartbotic-database/` before anything else.
2. **Diagnose** — start the server with `--read-only` and let the default recovery modes try to escalate. Use `smartbotic-db-cli status` to see what happened.
3. **Extract data** — while in read-only mode, use reads or (v1.7.0+) `smartbotic-db-cli dump` to export whatever survived.
4. **Decide**:
- Data looks right → `smartbotic-db-cli unlock` to resume writes
- Data is missing → don't unlock; extract to external backup, then start clean
#### Client API
cpp // Acknowledge auto-readonly state and resume writes db.unlock();
// Force read-only for maintenance (regardless of recovery outcome) db.lock();
// Inspect state + recovery outcome auto s = db.getReadOnlyStatus(); std::cout << "read-only: " << s.readOnly << "\n"
<< "recovery: " << s.recoveryOutcome << "\n"
<< "snapshot used: " << s.snapshotUsed << "\n";
#### Configuration
json "persistence": {
"snapshots": {
"validate_after_write": true, // verify snapshots immediately after creation
"cleanup_only_if_verified": true // don't evict old snapshots if new one is bad
},
"recovery": {
"mode": "normal",
"auto_escalate": true, // auto-fallback to older snapshot on corruption
"allow_empty_on_fresh_install": true // permit empty start when no snapshots/WAL exist
}
}
#### Exit codes
- `0` — normal shutdown
- `1` — generic error
- `10` — **recovery refused.** Snapshot or WAL failure + mode=normal + data exists on disk. Operator must escalate the recovery mode or pass `--force-readwrite`.
- `11` — config invalid (e.g. bad recovery-mode value)
Add a new bullet after "Per-collection Timestamp Precision":
- **Durable Snapshots + Tiered Recovery** — atomic snapshot writer (`.tmp` + fsync + rename + post-write verification), loader fallback chain across snapshots, MySQL-style recovery modes (`normal` / `snapshot_fallback` / `wal_only` / `best_effort` / `force_empty`) selectable via `--recovery-mode` flag or `recovery.mode` config. Non-trivial recovery (fallback, WAL-only, forced empty) automatically enters **read-only mode** — operator must `smartbotic-db-cli unlock` or pass `--force-readwrite` to accept writes. Exit code `10` when recovery is refused.
[ ] Step 3: Commit
git add docs/integration-guide.md CLAUDE.md
git commit -m "docs(recovery): document tiered recovery modes + auto-readonly safety"
[ ] Step 1: Local build
cd /data/smartbotic-database
./packaging/build.sh --local --skip-tests 2>&1 | tail -3
[ ] Step 2: Docker build + publish
./packaging/build.sh --sync --suite trixie 2>&1 | tail -10
[ ] Step 3: Verify in repo
ls /data/smartbotics-deb-repo/repo/pool/trixie/main/*1.6.1*
Expected: 4 packages.
[ ] Step 4: Push main and tag
git push origin main
git tag -a v1.6.1 -m "v1.6.1: snapshot durability fix + tiered recovery + read-only mode"
git push origin v1.6.1
[ ] Step 5: Create Gogs release via Playwright
Per saved memory (feedback_gogs_releases.md), open Playwright at https://git.smartbotics.ai/fszontagh/smartbotic-database/releases/new, prompt operator to log in, fill with release notes, publish.
Task 1 : VERSION + enum/struct scaffolding
Task 2 : recover() signature change + applyWalEntry extraction
Task 3 : ATOMIC SNAPSHOT WRITER (core fix)
Task 4 : Loader fallback chain
Task 5 : Full RecoveryMode logic
Task 6 : Proto RPCs
Task 7 : Auto-readonly on non-trivial recovery
Task 8 : Write handlers reject on read-only
Task 9 : Client API
Task 10 : CLI flags + subcommands
Task 11 : Config parsing
Task 12 : Integration tests
Task 13 : Docs
Task 14 : Build, publish, tag, Gogs release
Tasks 1→5 are the durability+recovery engine. Tasks 6→10 wire it into the RPC/CLI layer. Tasks 11→14 are polish + release.
Every commit is independently revertable. If a later task breaks a deployment, that single commit can be reverted without losing earlier fixes. The atomic-writer commit (Task 3) is the single most important one — if nothing else ships, at least shipping that one stops the silent-corruption bleeding.