瀏覽代碼

feat(index): supply ordering from the index; add IN and EXISTS

An index was only ever consulted to narrow a FILTER, and a Sort is not a filter.
So "newest N, unfiltered" walked all 10,118 rows of the 592 MB executions
collection to discover what to sort by: 341ms to return one row, with the index
declared and unused. The control proving sorting is not inherently unindexed:
workflowId=X plus the same sort was already 0ms, a selective filter having left
one row to order.

orderedHitsFromIndex walks the sort field's index and stops at offset+limit.
341ms -> 0ms with total_matched still exact. Direction handles ties for free:
sort_documents breaks an equal-key tie on id ascending / reverse-id descending,
and DUPSORT stores a key's ids ascending, so MDB_FIRST+MDB_NEXT and
MDB_LAST+MDB_PREV both already match.

The guard is the safety argument. sort_documents returns `descending` when the
left value is missing, so descending places rows LACKING the sort field FIRST -
and the index holds no posting for them, so a naive walk would silently omit
them from the first page. Ordering is used only when index entries == row count.
A field of one-element arrays would pass that while sort_documents compares whole
arrays, so the fetched rows are checked and the plan abandoned if any is an array.

Unfiltered sorts only: with filters an exact total_matched requires visiting every
match, which is what stopping early avoids. The two cannot both hold.

My own test caught a real bug here. The first version asserted ordered_scans == 1
for one query and passed while the plan was never taken - mdb_stat on a collection
counts the identity sentinel, so entries never equalled rows and all eleven
equivalence cases were comparing the scan against itself. rowCountTxn() fixes it
(the planner's budget used the same raw count) and the test now requires the
ordered plan in all eleven.

Also: IN as a union of posting lists, with per-value counts summed first so a
broad IN is rejected before any list is read; EXISTS=true as all of a field's
postings, selective exactly when the field is sparse. EXISTS=false cannot be
served - rows without a posting are not enumerable - so it scans.

ctest 20/20, index e2e 33, all suites green.
fszontagh 1 月之前
父節點
當前提交
2871d7a27e
共有 6 個文件被更改,包括 496 次插入 和 14 次删除
  1. 0 0
      CLAUDE.md
  2. 1 1
      VERSION
  3. 2 2
      docs/ROADMAP.md
  4. 215 11
      service/src/storage/document_store_lmdb.cpp
  5. 41 0
      service/src/storage/document_store_lmdb.hpp
  6. 237 0
      tests/test_subdb_identity.cpp

文件差異過大導致無法顯示
+ 0 - 0
CLAUDE.md


+ 1 - 1
VERSION

@@ -1 +1 @@
-2.9.1
+2.9.2

+ 2 - 2
docs/ROADMAP.md

@@ -1,6 +1,6 @@
 # Smartbotic Database - Status and Roadmap
 
-**Current version: 2.9.1** (see `VERSION`). Last reviewed: 2026-08-09.
+**Current version: 2.9.2** (see `VERSION`). Last reviewed: 2026-08-09.
 
 This is the single authoritative statement of what exists and what does not.
 If any other document in this repository disagrees with this one, this one is
@@ -53,7 +53,7 @@ Consequences of that substitution, which trip up readers:
 
 ## Shipped
 
-Every item below is in the installed product as of 2.9.1. `CLAUDE.md` has the
+Every item below is in the installed product as of 2.9.2. `CLAUDE.md` has the
 detail and the failure modes.
 
 - JSON document store: collections, version history, field-level encryption, TTL

+ 215 - 11
service/src/storage/document_store_lmdb.cpp

@@ -683,6 +683,9 @@ LmdbDocumentStore::IndexPlanStats LmdbDocumentStore::index_plan_stats() const {
         declined_unselective_.load(std::memory_order_relaxed);
     st.intersected_scans = intersected_scans_.load(std::memory_order_relaxed);
     st.range_scans = range_scans_.load(std::memory_order_relaxed);
+    st.ordered_scans = ordered_scans_.load(std::memory_order_relaxed);
+    st.union_scans = union_scans_.load(std::memory_order_relaxed);
+    st.exists_scans = exists_scans_.load(std::memory_order_relaxed);
     return st;
 }
 
@@ -692,6 +695,9 @@ void LmdbDocumentStore::reset_index_plan_stats() {
     declined_unselective_.store(0, std::memory_order_relaxed);
     intersected_scans_.store(0, std::memory_order_relaxed);
     range_scans_.store(0, std::memory_order_relaxed);
+    ordered_scans_.store(0, std::memory_order_relaxed);
+    union_scans_.store(0, std::memory_order_relaxed);
+    exists_scans_.store(0, std::memory_order_relaxed);
 }
 
 std::optional<LmdbDocumentStore::IndexStats>
@@ -982,6 +988,136 @@ uint64_t LmdbDocumentStore::count(std::string_view collection) {
 // status="completed" matches 6,676 of 10,101 rows (66%), so an index on `status`
 // would make that query SLOWER. Declaring an index must not be able to
 // pessimise a query, so the plan is re-decided per value, not per field.
+
+// Rows in a collection sub-db, excluding the v2.4.4 identity sentinel. mdb_stat
+// counts it, so comparing a raw ms_entries against an index's posting count is
+// off by exactly one - which silently disabled the ordered plan in every case
+// until a test asserted the plan was actually taken.
+uint64_t LmdbDocumentStore::rowCountTxn(ReadTxn& rtxn, unsigned int dbi) {
+    MDB_stat st{};
+    if (mdb_stat(rtxn.raw(), dbi, &st) != MDB_SUCCESS) return 0;
+    uint64_t n = st.ms_entries;
+    if (!read_subdb_identity(rtxn, dbi).empty() && n > 0) --n;
+    return n;
+}
+
+bool LmdbDocumentStore::orderedHitsFromIndex(
+    ReadTxn& rtxn,
+    std::string_view collection,
+    const smartbotic::database::Query& query,
+    unsigned int coll_dbi,
+    std::vector<std::string>& out,
+    uint64_t& exact_total) {
+
+    if (!query.filters.empty()) return false;
+    if (!query.sort || query.sort->field.empty()) return false;
+    const auto fields = indexed_fields(collection);
+    if (std::find(fields.begin(), fields.end(), query.sort->field) == fields.end()) {
+        return false;
+    }
+
+    const uint64_t rows = rowCountTxn(rtxn, coll_dbi);
+    if (rows == 0) return false;
+
+    auto dbi_opt = try_open_for_read(rtxn,
+                                     index_subdb_name(collection, query.sort->field));
+    if (!dbi_opt) return false;
+
+    // Total order or nothing. entries != rows means some row has no posting (the
+    // field is absent, or its value is unindexable) or several (an array), and in
+    // either case the index cannot state where that row sorts.
+    MDB_stat ist{};
+    if (mdb_stat(rtxn.raw(), *dbi_opt, &ist) != MDB_SUCCESS) return false;
+    uint64_t entries = ist.ms_entries;
+    if (!read_subdb_identity(rtxn, *dbi_opt).empty() && entries > 0) --entries;
+    if (entries != rows) return false;
+
+    MDB_cursor* cur = nullptr;
+    mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (index order)");
+    struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+
+    const bool desc = query.sort->descending;
+    const uint64_t want = static_cast<uint64_t>(query.offset) + query.limit;
+
+    // Direction handles ties correctly without extra work. sort_documents breaks
+    // an equal-key tie on id - ascending by id, descending by reverse id - and
+    // DUPSORT stores a key's ids ascending. So MDB_FIRST/MDB_NEXT yields
+    // ascending keys with ascending ids, and MDB_LAST/MDB_PREV yields descending
+    // keys with descending ids. Both match.
+    MDB_val k{0, nullptr};
+    MDB_val v{0, nullptr};
+    int rc = mdb_cursor_get(cur, &k, &v, desc ? MDB_LAST : MDB_FIRST);
+    std::vector<std::string> ids;
+    while (rc == MDB_SUCCESS && ids.size() < want) {
+        if (!is_identity_key(to_sv(k))) ids.emplace_back(to_sv(v));
+        rc = mdb_cursor_get(cur, &k, &v, desc ? MDB_PREV : MDB_NEXT);
+    }
+    if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) {
+        throw_mdb(rc, "cursor step (index order)");
+    }
+
+    // A field holding one-element arrays on every row would satisfy entries ==
+    // rows while sort_documents compares whole arrays. Check the rows actually
+    // being returned and abandon if any is an array, so the page can never
+    // disagree with the general path.
+    //
+    // Honest residual: a row OUTSIDE this page could still be an array and change
+    // the true ordering. Sorting by an array-valued field is not meaningful, so
+    // this is documented rather than defended further.
+    const uint64_t start = std::min<uint64_t>(query.offset, ids.size());
+    const uint64_t end = std::min<uint64_t>(want, ids.size());
+    for (uint64_t i = start; i < end; ++i) {
+        MDB_val hk = to_val(ids[i]);
+        MDB_val hv{0, nullptr};
+        if (mdb_get(rtxn.raw(), coll_dbi, &hk, &hv) != MDB_SUCCESS) continue;
+        const auto bytes = to_sv(hv);
+        yyjson_doc* d = yyjson_read(bytes.data(), bytes.size(), 0);
+        if (!d) continue;
+        auto val = resolve_from_yyjson(yyjson_doc_get_root(d), query.sort->field);
+        const bool is_array = val && val->is_array();
+        yyjson_doc_free(d);
+        if (is_array) return false;
+    }
+
+    out = std::move(ids);
+    exact_total = rows;      // no filters, so every row matches
+    ordered_scans_.fetch_add(1, std::memory_order_relaxed);
+    return true;
+}
+
+std::optional<std::vector<std::string>>
+LmdbDocumentStore::lookupIndexAllTxn(ReadTxn& rtxn,
+                                      std::string_view collection,
+                                      const std::string& field,
+                                      uint64_t budget) {
+    auto dbi_opt = try_open_for_read(rtxn, index_subdb_name(collection, field));
+    if (!dbi_opt) return std::nullopt;
+
+    MDB_cursor* cur = nullptr;
+    mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (index all)");
+    struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+
+    std::vector<std::string> ids;
+    MDB_val k{0, nullptr};
+    MDB_val v{0, nullptr};
+    int rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
+    while (rc == MDB_SUCCESS) {
+        if (!is_identity_key(to_sv(k))) {
+            ids.emplace_back(to_sv(v));
+            // Bail out rather than truncate, as everywhere else: a partial
+            // candidate list drops matching rows.
+            if (ids.size() > budget) return std::nullopt;
+        }
+        rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT);
+    }
+    if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) throw_mdb(rc, "cursor next (index all)");
+    // An array-valued field yields several postings per row; dedupe so a row is
+    // named once.
+    std::sort(ids.begin(), ids.end());
+    ids.erase(std::unique(ids.begin(), ids.end()), ids.end());
+    return ids;
+}
+
 std::optional<std::vector<std::string>>
 LmdbDocumentStore::planIndexCandidates(ReadTxn& rtxn,
                                         std::string_view collection,
@@ -991,9 +1127,7 @@ LmdbDocumentStore::planIndexCandidates(ReadTxn& rtxn,
     const auto fields = indexed_fields(collection);
     if (fields.empty()) return std::nullopt;
 
-    MDB_stat st{};
-    if (mdb_stat(rtxn.raw(), coll_dbi, &st) != MDB_SUCCESS) return std::nullopt;
-    const uint64_t total = st.ms_entries;
+    const uint64_t total = rowCountTxn(rtxn, coll_dbi);
     if (total == 0) return std::nullopt;
     const uint64_t budget = total / kIndexSelectivityDivisor;
 
@@ -1015,6 +1149,47 @@ LmdbDocumentStore::planIndexCandidates(ReadTxn& rtxn,
             exacts.push_back({&f, *n});
         }
     }
+
+    // v2.9.2 — IN is a UNION of posting lists, the mirror of the intersection
+    // below. Summing the per-value counts first means an IN over a broad set is
+    // rejected before any list is read.
+    for (const auto& f : query.filters) {
+        if (f.op != Op::IN || !f.value.is_array()) continue;
+        if (!is_indexed(f.field)) continue;
+        uint64_t sum = 0;
+        bool all_countable = true;
+        for (const auto& el : f.value) {
+            auto n = countIndexEqTxn(rtxn, collection, f.field, el);
+            if (!n) { all_countable = false; break; }
+            sum += *n;
+        }
+        if (!all_countable || sum > budget) continue;
+        std::vector<std::string> uni;
+        for (const auto& el : f.value) {
+            auto part = lookupIndexEqTxn(rtxn, collection, f.field, el);
+            if (!part) { all_countable = false; break; }
+            uni.insert(uni.end(), part->begin(), part->end());
+        }
+        if (!all_countable) continue;
+        std::sort(uni.begin(), uni.end());
+        uni.erase(std::unique(uni.begin(), uni.end()), uni.end());
+        union_scans_.fetch_add(1, std::memory_order_relaxed);
+        return uni;
+    }
+
+    // v2.9.2 — EXISTS=true is every posting the field has, which is selective
+    // exactly when the field is SPARSE. EXISTS=false cannot be served: rows
+    // without a posting are not enumerable from the index.
+    for (const auto& f : query.filters) {
+        if (f.op != Op::EXISTS) continue;
+        if (!is_indexed(f.field)) continue;
+        const bool wants = f.value.is_boolean() ? f.value.get<bool>() : true;
+        if (!wants) continue;
+        if (auto ids = lookupIndexAllTxn(rtxn, collection, f.field, budget)) {
+            exists_scans_.fetch_add(1, std::memory_order_relaxed);
+            return ids;
+        }
+    }
     std::sort(exacts.begin(), exacts.end(),
               [](const Exact& a, const Exact& b) { return a.count < b.count; });
 
@@ -1226,14 +1401,38 @@ ScanResult LmdbDocumentStore::scan(std::string_view collection,
             yyjson_doc_free(ydoc);
         };
 
-        auto candidates = planIndexCandidates(rtxn, collection, query, *dbi_opt);
-        if (candidates) {
-            indexed_scans_.fetch_add(1, std::memory_order_relaxed);
-        } else {
-            full_scans_.fetch_add(1, std::memory_order_relaxed);
+        // v2.9.2 — ORDERING from the index. Try this first: when it applies it
+        // reads the page and nothing else, where every other plan still visits
+        // every matching row.
+        bool ordered_by_index = false;
+        uint64_t ordered_total = 0;
+        if (query.filters.empty() && sorting) {
+            std::vector<std::string> ordered_ids;
+            if (orderedHitsFromIndex(rtxn, collection, query, *dbi_opt,
+                                     ordered_ids, ordered_total)) {
+                ordered_by_index = true;
+                hits.reserve(ordered_ids.size());
+                for (auto& id : ordered_ids) {
+                    Hit h;
+                    h.id = std::move(id);
+                    hits.push_back(std::move(h));
+                }
+            }
+        }
+
+        std::optional<std::vector<std::string>> candidates;
+        if (!ordered_by_index) {
+            candidates = planIndexCandidates(rtxn, collection, query, *dbi_opt);
+            if (candidates) {
+                indexed_scans_.fetch_add(1, std::memory_order_relaxed);
+            } else {
+                full_scans_.fetch_add(1, std::memory_order_relaxed);
+            }
         }
 
-        if (candidates) {
+        if (ordered_by_index) {
+            // hits are already in final order, and only the page's worth exist.
+        } else if (candidates) {
             // Indexed plan: visit only the rows the index named. Every filter is
             // still applied to each one, so the index is allowed to be a
             // superset - it narrows work, it does not decide the answer.
@@ -1262,7 +1461,9 @@ ScanResult LmdbDocumentStore::scan(std::string_view collection,
             }
         }
 
-        if (sorting) {
+        // Already ordered by the index walk - re-sorting would be wasted work, and
+        // the hits carry no sort key to sort by.
+        if (sorting && !ordered_by_index) {
             const bool desc = query.sort->descending;
             std::stable_sort(hits.begin(), hits.end(),
                 [desc](const Hit& a, const Hit& b) {
@@ -1277,7 +1478,10 @@ ScanResult LmdbDocumentStore::scan(std::string_view collection,
                 });
         }
 
-        result.total_matched = hits.size();
+        // With an index-ordered walk `hits` holds only offset+limit rows, so its
+        // size is not the total. The total is exact and free there: no filters
+        // means every row matches, and mdb_stat already counted them.
+        result.total_matched = ordered_by_index ? ordered_total : hits.size();
         const uint64_t start = std::min<uint64_t>(query.offset, hits.size());
         const uint64_t end = std::min<uint64_t>(
             static_cast<uint64_t>(query.offset) + query.limit, hits.size());

+ 41 - 0
service/src/storage/document_store_lmdb.hpp

@@ -83,6 +83,9 @@ public:
         uint64_t declined_unselective = 0;  // an index existed but was too broad
         uint64_t intersected_scans = 0;      // two posting lists intersected
         uint64_t range_scans = 0;            // a range walked the index
+        uint64_t ordered_scans = 0;          // the index supplied the ORDER
+        uint64_t union_scans = 0;            // IN unioned posting lists
+        uint64_t exists_scans = 0;           // EXISTS served from all postings
     };
     IndexPlanStats index_plan_stats() const;
     void reset_index_plan_stats();
@@ -164,6 +167,9 @@ private:
     mutable std::atomic<uint64_t> declined_unselective_{0};
     mutable std::atomic<uint64_t> intersected_scans_{0};
     mutable std::atomic<uint64_t> range_scans_{0};
+    mutable std::atomic<uint64_t> ordered_scans_{0};
+    mutable std::atomic<uint64_t> union_scans_{0};
+    mutable std::atomic<uint64_t> exists_scans_{0};
     std::unordered_map<std::string, unsigned int> dbi_cache_;
 
     // Resolve a collection name to an MDB_dbi handle.
@@ -213,6 +219,41 @@ private:
                         const nlohmann::json& value,
                         uint64_t budget);
 
+    // Rows in a collection sub-db, excluding the v2.4.4 identity sentinel.
+    uint64_t rowCountTxn(class ReadTxn& rtxn, unsigned int dbi);
+
+    // v2.9.2 — ORDERING from an index, rather than candidates for a filter.
+    //
+    // An index was previously consulted only to narrow a filter, and a Sort is not
+    // a filter, so "newest N, unfiltered" walked the whole collection to find out
+    // what to sort by: 341ms to return one row from 592 MB on real data, with the
+    // index declared and unused. Walking the sort field's index instead is
+    // O(page).
+    //
+    // Fills `out` with ids in FINAL order, up to offset+limit, and sets
+    // `exact_total`. Returns false to decline, leaving `out` untouched.
+    //
+    // Declines unless the index is a TOTAL ORDER over the collection - index
+    // entries must equal row count, so every row has exactly one posting. That
+    // guard exists because sort_documents places rows MISSING the sort field
+    // first when descending, and an index holds no posting for them: a walk would
+    // silently omit them from the first page. Wrong rows, not slow ones.
+    //
+    // Only for queries with no filters. With filters, an exact total_matched
+    // requires visiting every match, which is precisely what stopping early
+    // avoids - the two cannot both hold, and the total is part of the contract.
+    bool orderedHitsFromIndex(class ReadTxn& rtxn, std::string_view collection,
+                              const smartbotic::database::Query& query,
+                              unsigned int coll_dbi,
+                              std::vector<std::string>& out,
+                              uint64_t& exact_total);
+
+    // Every id the index holds for `field`, regardless of value. Serves
+    // EXISTS=true, which is selective exactly when the field is sparse.
+    std::optional<std::vector<std::string>>
+    lookupIndexAllTxn(class ReadTxn& rtxn, std::string_view collection,
+                      const std::string& field, uint64_t budget);
+
     // Choose an index for a query, or nullopt to walk the whole collection.
     std::optional<std::vector<std::string>>
     planIndexCandidates(class ReadTxn& rtxn, std::string_view collection,

+ 237 - 0
tests/test_subdb_identity.cpp

@@ -16,6 +16,7 @@
 #include <iostream>
 #include <cstdio>
 #include <algorithm>
+#include <functional>
 #include <thread>
 #include <set>
 #include <string>
@@ -1177,6 +1178,240 @@ void test_range_contains_and_intersection() {
     }
 }
 
+
+// v2.9.2 — the index supplies the ORDER, not just filter candidates.
+//
+// "newest N, unfiltered" walked the whole collection to discover what to sort by:
+// 341ms to return one row from 592 MB of real data, with the index declared and
+// unused, because a Sort is not a filter. Walking the sort field's index reads the
+// page and nothing else.
+//
+// The dangerous part is not speed, it is that sort_documents places rows MISSING
+// the sort field FIRST when descending, while an index holds no posting for them.
+// A walk would silently omit them from the first page - wrong rows, not slow ones.
+// So every case here compares the ordered plan against the plain scan.
+void test_index_supplies_ordering() {
+    using Op = smartbotic::database::FilterOp;
+
+    auto make = [](LmdbDocumentStore& store, int n,
+                   const std::function<nlohmann::json(int)>& body) {
+        for (int i = 0; i < n; ++i) {
+            Document d;
+            d.id = "r" + std::string(i < 10 ? "00" : (i < 100 ? "0" : "")) +
+                   std::to_string(i);
+            d.collection = "c";
+            d.set_data(body(i));
+            store.put("c", d.id, d);
+        }
+    };
+    auto sig = [](LmdbDocumentStore& store, const smartbotic::database::Query& q) {
+        auto r = store.scan("c", q);
+        std::string out = "total=" + std::to_string(r.total_matched) +
+                          " more=" + std::to_string(r.has_more ? 1 : 0) + " [";
+        for (const auto& d : r.documents) { out += d.id; out += ","; }
+        return out + "]";
+    };
+    auto Q = [](const char* field, bool desc, uint32_t limit, uint32_t offset) {
+        smartbotic::database::Query q;
+        q.sort = smartbotic::database::Sort{field, desc};
+        q.limit = limit;
+        q.offset = offset;
+        return q;
+    };
+
+    // ---- every row carries the sort field: the ordered plan applies ----
+    {
+        TmpEnv t("ord-total");
+        LmdbDocumentStore store(t.env);
+        // Deliberate duplicate keys (i/3) so tie-breaking by id is exercised in
+        // both directions.
+        make(store, 300, [](int i) {
+            return nlohmann::json{{"seq", i}, {"dup", i / 3}};
+        });
+
+        const std::vector<std::pair<const char*, smartbotic::database::Query>> cases = {
+            {"desc, first page",        Q("seq", true, 10, 0)},
+            {"asc, first page",         Q("seq", false, 10, 0)},
+            {"desc, deep offset",       Q("seq", true, 10, 250)},
+            {"asc, deep offset",        Q("seq", false, 7, 33)},
+            {"desc, last partial page", Q("seq", true, 10, 295)},
+            {"offset past the end",     Q("seq", true, 10, 999)},
+            {"limit=0",                 Q("seq", true, 0, 0)},
+            {"whole collection",        Q("seq", false, 1000, 0)},
+            {"desc with tied keys",     Q("dup", true, 12, 0)},
+            {"asc with tied keys",      Q("dup", false, 12, 0)},
+            {"tied keys, deep offset",  Q("dup", true, 5, 40)},
+        };
+        uint64_t ordered_used = 0;
+        for (const auto& [name, q] : cases) {
+            store.set_indexed_fields("c", {});
+            const std::string without = sig(store, q);
+            store.set_indexed_fields("c", {"seq", "dup"});
+            store.build_index("c", "seq");
+            store.build_index("c", "dup");
+            store.reset_index_plan_stats();
+            const std::string with = sig(store, q);
+            ordered_used += store.index_plan_stats().ordered_scans;
+            if (with != without) {
+                std::cerr << "    scan:    " << without.substr(0, 180) << "\n"
+                          << "    ordered: " << with.substr(0, 180) << "\n";
+            }
+            const std::string msg = std::string("ordered plan == scan: ") + name;
+            check(with == without, msg.c_str());
+        }
+
+        // EVERY case above must have used the ordered plan, not just one. The
+        // first version of this test asserted only a single query and so passed
+        // while the plan was silently never taken (the collection's row count
+        // included the identity sentinel, so entries never equalled rows).
+        check(ordered_used == cases.size(),
+              ("the index SUPPLIED THE ORDER in all " + std::to_string(cases.size()) +
+               " cases (got " + std::to_string(ordered_used) + ") - otherwise the "
+               "comparisons above are scan against scan").c_str());
+    }
+
+    // ---- THE trap: one row lacks the sort field ----
+    {
+        TmpEnv t("ord-missing");
+        LmdbDocumentStore store(t.env);
+        make(store, 50, [](int i) {
+            nlohmann::json j{{"other", i}};
+            if (i != 7) j["seq"] = i;      // r007 has no seq
+            return j;
+        });
+        store.set_indexed_fields("c", {});
+        const std::string without = sig(store, Q("seq", true, 5, 0));
+        store.set_indexed_fields("c", {"seq"});
+        store.build_index("c", "seq");
+        const std::string with = sig(store, Q("seq", true, 5, 0));
+
+        check(with == without,
+              "one row missing the sort field: DESCENDING puts it FIRST, and the "
+              "index has no posting for it - the ordered plan must decline");
+        store.reset_index_plan_stats();
+        (void)sig(store, Q("seq", true, 5, 0));
+        check(store.index_plan_stats().ordered_scans == 0,
+              "and it does decline - entries != rows is the guard");
+    }
+
+    // ---- an array-valued sort field must also decline ----
+    {
+        TmpEnv t("ord-array");
+        LmdbDocumentStore store(t.env);
+        make(store, 40, [](int i) {
+            return nlohmann::json{{"seq", nlohmann::json::array({i})}};
+        });
+        store.set_indexed_fields("c", {});
+        const std::string without = sig(store, Q("seq", true, 5, 0));
+        store.set_indexed_fields("c", {"seq"});
+        store.build_index("c", "seq");
+        const std::string with = sig(store, Q("seq", true, 5, 0));
+        check(with == without,
+              "a one-element-array sort field agrees - it satisfies entries==rows, "
+              "so the fetched-page array check is what catches it");
+    }
+}
+
+// v2.9.2 — IN as a union of posting lists, EXISTS as all postings.
+void test_index_in_and_exists() {
+    using Op = smartbotic::database::FilterOp;
+    TmpEnv t("in-exists");
+    LmdbDocumentStore store(t.env);
+
+    for (int i = 0; i < 400; ++i) {
+        Document d;
+        d.id = "r" + std::to_string(1000 + i);
+        d.collection = "c";
+        nlohmann::json j{{"grp", "g" + std::to_string(i % 40)}};
+        // `rare` exists on 8 of 400 rows, so EXISTS=true is highly selective.
+        if (i % 50 == 0) j["rare"] = i;
+        d.set_data(j);
+        store.put("c", d.id, d);
+    }
+
+    auto F = [](const char* f, Op op, const nlohmann::json& v) {
+        smartbotic::database::Filter x;
+        x.field = f; x.op = op; x.value = v;
+        return x;
+    };
+    auto sig = [&](const std::vector<smartbotic::database::Filter>& fs) {
+        smartbotic::database::Query q;
+        q.filters = fs;
+        q.limit = 1000;
+        auto r = store.scan("c", q);
+        std::vector<std::string> ids;
+        for (const auto& d : r.documents) ids.push_back(d.id);
+        std::sort(ids.begin(), ids.end());
+        std::string out = "total=" + std::to_string(r.total_matched) + " [";
+        for (const auto& i : ids) { out += i; out += ","; }
+        return out + "]";
+    };
+
+    const std::vector<std::pair<const char*, std::vector<smartbotic::database::Filter>>> cases = {
+        {"IN over three values",  {F("grp", Op::IN, nlohmann::json::array({"g1","g2","g3"}))}},
+        {"IN with a missing value", {F("grp", Op::IN, nlohmann::json::array({"g1","nope"}))}},
+        {"IN over one value",     {F("grp", Op::IN, nlohmann::json::array({"g5"}))}},
+        {"IN matching nothing",   {F("grp", Op::IN, nlohmann::json::array({"x","y"}))}},
+        {"IN too broad to help",  {F("grp", Op::IN, nlohmann::json::array(
+              {"g0","g1","g2","g3","g4","g5","g6","g7","g8","g9","g10","g11"}))}},
+        {"EXISTS true on a sparse field", {F("rare", Op::EXISTS, true)}},
+        {"EXISTS false on a sparse field", {F("rare", Op::EXISTS, false)}},
+        {"EXISTS true on a dense field",   {F("grp", Op::EXISTS, true)}},
+    };
+
+    for (const auto& [name, filters] : cases) {
+        store.set_indexed_fields("c", {});
+        const std::string without = sig(filters);
+        store.set_indexed_fields("c", {"grp", "rare"});
+        store.build_index("c", "grp");
+        store.build_index("c", "rare");
+        const std::string with = sig(filters);
+        if (with != without) {
+            std::cerr << "    scan:    " << without.substr(0, 160) << "\n"
+                      << "    indexed: " << with.substr(0, 160) << "\n";
+        }
+        const std::string msg = std::string("indexed == scan: ") + name;
+        check(with == without, msg.c_str());
+    }
+
+    store.set_indexed_fields("c", {"grp", "rare"});
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("grp", Op::IN, nlohmann::json::array({"g1","g2","g3"}))});
+        check(store.index_plan_stats().union_scans == 1,
+              "IN over a narrow set UNIONS posting lists - 30 of 400 rows");
+    }
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("grp", Op::IN, nlohmann::json::array(
+              {"g0","g1","g2","g3","g4","g5","g6","g7","g8","g9","g10","g11"}))});
+        auto st = store.index_plan_stats();
+        check(st.union_scans == 0 && st.full_scans == 1,
+              "a broad IN is rejected on the summed counts, before any list is read");
+    }
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("rare", Op::EXISTS, true)});
+        check(store.index_plan_stats().exists_scans == 1,
+              "EXISTS=true on a sparse field is served from all its postings");
+    }
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("rare", Op::EXISTS, false)});
+        auto st = store.index_plan_stats();
+        check(st.exists_scans == 0 && st.full_scans == 1,
+              "EXISTS=false cannot be served - rows WITHOUT a posting are not "
+              "enumerable from the index, so it must scan");
+    }
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("grp", Op::EXISTS, true)});
+        auto st = store.index_plan_stats();
+        check(st.exists_scans == 0 && st.full_scans == 1,
+              "EXISTS=true on a field every row has is not selective, so it scans");
+    }
+}
+
 }  // namespace
 
 int main() {
@@ -1199,6 +1434,8 @@ int main() {
     test_indexed_and_unindexed_plans_agree();
     test_empty_id_reads_as_absent();
     test_range_contains_and_intersection();
+    test_index_supplies_ordering();
+    test_index_in_and_exists();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;

部分文件因文件數量過多而無法顯示