|
@@ -299,11 +299,19 @@ public:
|
|
|
* entries at boot — do NOT mirror to LMDB (unlike every ordinary write
|
|
* entries at boot — do NOT mirror to LMDB (unlike every ordinary write
|
|
|
* path, where mirrorDocOrUndo()/mirrorWriteToDocStore() always run
|
|
* path, where mirrorDocOrUndo()/mirrorWriteToDocStore() always run
|
|
|
* BEFORE the WAL entry is even logged — see update()/insertWithVector()
|
|
* BEFORE the WAL entry is even logged — see update()/insertWithVector()
|
|
|
- * above). That asymmetry is harmless for the ordinary write path,
|
|
|
|
|
- * because a WAL entry there can only exist if its LMDB write already
|
|
|
|
|
- * committed successfully first (mirror failure prevents the WAL log
|
|
|
|
|
- * call from ever being reached). It stops being harmless the moment
|
|
|
|
|
- * something logs a WAL entry BEFORE its LMDB write — which
|
|
|
|
|
|
|
+ * above). Round 4 correction — that asymmetry is NOT harmless for the
|
|
|
|
|
+ * ordinary write path either, and the round-3 claim that "a WAL entry
|
|
|
|
|
+ * implies the LMDB write already succeeded" is false:
|
|
|
|
|
+ * applyDualWriteMirror() SWALLOWS every LMDB fault except
|
|
|
|
|
+ * UniqueViolation (log ERROR, bump drift, flip mirror_healthy_, no
|
|
|
|
|
+ * rethrow — storage/dual_write_mirror.hpp), so mirrorDocOrUndo()
|
|
|
|
|
+ * returns normally and emitPersist() logs the WAL entry anyway. A WAL
|
|
|
|
|
+ * entry therefore only implies "the mirror did not throw
|
|
|
|
|
+ * UniqueViolation". So the divergence this pass repairs is NOT
|
|
|
|
|
+ * cascade-only: it also occurs on any ordinary write whose mirror fault
|
|
|
|
|
+ * was swallowed, and re-mirroring at boot repairs those too. The
|
|
|
|
|
+ * divergence merely becomes RELIABLY reachable the moment something
|
|
|
|
|
+ * logs a WAL entry BEFORE its LMDB write — which
|
|
|
* relations/relation_cascade.cpp's cascade delete deliberately does
|
|
* relations/relation_cascade.cpp's cascade delete deliberately does
|
|
|
* (WAL-first, for its own crash-safety reasons: see that file's
|
|
* (WAL-first, for its own crash-safety reasons: see that file's
|
|
|
* header). A crash between that WAL write and the LMDB commit leaves a
|
|
* header). A crash between that WAL write and the LMDB commit leaves a
|
|
@@ -313,24 +321,88 @@ public:
|
|
|
* loadDocumentWithHistory() do not, and LMDB then permanently serves
|
|
* loadDocumentWithHistory() do not, and LMDB then permanently serves
|
|
|
* the pre-cascade row until something else happens to rewrite it.
|
|
* the pre-cascade row until something else happens to rewrite it.
|
|
|
*
|
|
*
|
|
|
- * PersistenceManager::recover() calls this once per DISTINCT id that
|
|
|
|
|
- * WAL replay applied as UPDATE/UPSERT (deduplicated, and skipped if a
|
|
|
|
|
- * later DELETE for that id was also replayed — remove() already
|
|
|
|
|
|
|
+ * PersistenceManager::recover() re-mirrors once per DISTINCT id that
|
|
|
|
|
+ * WAL replay applied as INSERT/UPDATE/UPSERT (deduplicated, and skipped
|
|
|
|
|
+ * if a later DELETE for that id was also replayed — remove() already
|
|
|
* mirrored that) — a targeted post-replay pass, not a mirror call on
|
|
* mirrored that) — a targeted post-replay pass, not a mirror call on
|
|
|
* every replayed entry, which would multiply LMDB writes by however
|
|
* every replayed entry, which would multiply LMDB writes by however
|
|
|
* many times a hot document was rewritten since the last snapshot for
|
|
* many times a hot document was rewritten since the last snapshot for
|
|
|
* no benefit (the ordinary case's LMDB write already happened at
|
|
* no benefit (the ordinary case's LMDB write already happened at
|
|
|
- * original write time and needs no repeating).
|
|
|
|
|
|
|
+ * original write time and needs no repeating). It calls the BATCHED
|
|
|
|
|
+ * remirrorDocuments() below, not this single-id form — see that
|
|
|
|
|
+ * method's comment for why one-transaction-per-document on the boot
|
|
|
|
|
+ * path is not acceptable.
|
|
|
*
|
|
*
|
|
|
* Safe to call whether or not this specific id actually needed it:
|
|
* Safe to call whether or not this specific id actually needed it:
|
|
|
* re-mirroring an already-correct row is an idempotent overwrite with
|
|
* re-mirroring an already-correct row is an idempotent overwrite with
|
|
|
* the same content. Returns false (no-op) if the id is not currently
|
|
* the same content. Returns false (no-op) if the id is not currently
|
|
|
- * in MemoryStore (nothing to mirror) or the mirror isn't wired
|
|
|
|
|
|
|
+ * in MemoryStore (nothing to mirror), if the mirror isn't wired
|
|
|
* (legacy/test bootstrap path, same as mirrorWriteToDocStore's own
|
|
* (legacy/test bootstrap path, same as mirrorWriteToDocStore's own
|
|
|
- * no-op condition).
|
|
|
|
|
|
|
+ * no-op condition), or if the row's own re-mirror failed — this NEVER
|
|
|
|
|
+ * throws, for the reason spelled out on remirrorDocuments().
|
|
|
*/
|
|
*/
|
|
|
bool remirrorDocument(const std::string& collection, const std::string& id);
|
|
bool remirrorDocument(const std::string& collection, const std::string& id);
|
|
|
|
|
|
|
|
|
|
+ struct RemirrorBatchResult {
|
|
|
|
|
+ // Rows whose LMDB write committed as part of this call.
|
|
|
|
|
+ uint64_t remirrored = 0;
|
|
|
|
|
+ // Rows this call could not re-mirror. Each one is logged at ERROR
|
|
|
|
|
+ // naming its collection and id, and bumps the mirror drift counter.
|
|
|
|
|
+ // A nonzero value means "those rows stay stale in LMDB", never
|
|
|
|
|
+ // "recovery failed" — see below.
|
|
|
|
|
+ uint64_t failed = 0;
|
|
|
|
|
+ };
|
|
|
|
|
+
|
|
|
|
|
+ /**
|
|
|
|
|
+ * v2.11.0 T12 round-4 — batched form of remirrorDocument(), and the one
|
|
|
|
|
+ * PersistenceManager::recover() actually calls.
|
|
|
|
|
+ *
|
|
|
|
|
+ * WHY BATCHED: the single-id form routes to LmdbDocumentStore::put(),
|
|
|
|
|
+ * which opens and commits its OWN WriteTxn. The env is opened without
|
|
|
|
|
+ * MDB_NOSYNC, so that is one fsync per document, on the boot path,
|
|
|
|
|
+ * before sd_notify(READY=1). WAL size is bounded only by
|
|
|
|
|
+ * snapshotIntervalSec (3600) and maxWalSizeMb (100), and
|
|
|
|
|
+ * --recovery-mode=wal_only can replay the entire history — tens of
|
|
|
|
|
+ * thousands of distinct ids on a busy install, against a 10-minute
|
|
|
|
|
+ * systemd start watchdog. This commits in chunks of `chunkSize`
|
|
|
|
|
+ * instead, using the Task 10 primitives (beginWrite() /
|
|
|
|
|
+ * put(WriteTxn&, …, to_cache) / commitAndCache()), so N documents cost
|
|
|
|
|
+ * ceil(N / chunkSize) fsyncs.
|
|
|
|
|
+ *
|
|
|
|
|
+ * ⚠ NEVER THROWS, and that is the whole point of its error handling. A
|
|
|
|
|
+ * re-mirror is a REPAIR pass: failing one row must degrade to "that row
|
|
|
|
|
+ * stays stale in LMDB", which is exactly the state the pass exists to
|
|
|
|
|
+ * improve on and is strictly better than refusing to boot. An escaping
|
|
|
|
|
+ * exception here would propagate out of recover() and turn a working
|
|
|
|
|
+ * recovery into a deterministic boot loop on data that booted fine
|
|
|
|
|
+ * before. Two throws are genuinely reachable: UniqueViolation (the pass
|
|
|
|
|
+ * runs against a STALE index, so a unique value that MOVED between two
|
|
|
|
|
+ * rows conflicts with the other row's not-yet-re-mirrored posting) and
|
|
|
|
|
+ * std::invalid_argument from parseProjectCollection() on a malformed or
|
|
|
|
|
+ * legacy collection key (WAL replay itself never parses collection
|
|
|
|
|
+ * keys, so such a key replays fine and only this pass would trip on
|
|
|
|
|
+ * it). Both are caught per row.
|
|
|
|
|
+ *
|
|
|
|
|
+ * ONE BAD ROW MUST NOT POISON ITS CHUNK: a throw from row N aborts that
|
|
|
|
|
+ * chunk's transaction (nothing in it committed, and NOTHING is cached —
|
|
|
|
|
+ * the Task 10 invariant: to_cache is applied only after a successful
|
|
|
|
|
+ * commit), so the whole chunk is then retried ROW BY ROW, each in its
|
|
|
|
|
+ * own transaction. The rows that can commit do; only the genuinely bad
|
|
|
|
|
+ * one is counted as failed. Rows that fail individually get ONE further
|
|
|
|
|
+ * retry pass after every other row has been re-mirrored, which is what
|
|
|
|
|
+ * actually resolves the moved-unique-value case: once the row that used
|
|
|
|
|
+ * to hold the value has been re-mirrored, its stale posting is gone and
|
|
|
|
|
+ * the row that now holds it commits.
|
|
|
|
|
+ *
|
|
|
|
|
+ * Rows whose collection is missing from MemoryStore, whose id is absent,
|
|
|
|
|
+ * or whose collection is `_`-prefixed (system collections are not
|
|
|
|
|
+ * mirrored — same rule as applyDualWriteMirror) are silently skipped:
|
|
|
|
|
+ * not remirrored, not failed.
|
|
|
|
|
+ */
|
|
|
|
|
+ RemirrorBatchResult remirrorDocuments(
|
|
|
|
|
+ const std::vector<std::pair<std::string, std::string>>& docs,
|
|
|
|
|
+ size_t chunkSize = 256);
|
|
|
|
|
+
|
|
|
/**
|
|
/**
|
|
|
* Check if a document exists.
|
|
* Check if a document exists.
|
|
|
*/
|
|
*/
|