Browse Source

test(query): pin IN behaviour on a sparse field, scan and indexed

A consumer reported that `parentExecutionId in (id, id, ...)` returned nothing
while each id matched alone, and attributed it to the field being sparse (about
3% of executions carry it).

It does not reproduce. Verified against the live 2.11.1 instance on zeus, on the
very field and values from the report: EQ per id returned 6/2/1/6/9 and IN over
all five returned 24, their exact sum. Also not reproducible locally on the plain
two-pass scan, with the field indexed (the union plan), through a view whose
projection omits the field, with a sort, a limit, a second filter, a 60-value
list of mostly-absent values, on `_id`, or on a dotted nested path.

Worth recording: on that instance `status` and `workflowId` are indexed and
`parentExecutionId` is not, so 'works on those two, not on this one' also
separates the indexed union plan from the scan - a different axis from
sparseness, and the one the report could not see.

No production change. This adds the regression suite so the claim is settled by
running something, and so a future change to the scan resolver or the union plan
cannot quietly make it true. 14 checks over real gRPC.
fszontagh 1 month ago
parent
commit
b82e0ba4b7
2 changed files with 192 additions and 0 deletions
  1. 124 0
      tests/load_test/test_sparse_in.cpp
  2. 68 0
      tests/load_test/test_sparse_in.sh

+ 124 - 0
tests/load_test/test_sparse_in.cpp

@@ -0,0 +1,124 @@
+// Regression: IN must behave the same on a SPARSE field as on a dense one.
+//
+// Reported by a consumer as "`parentExecutionId in (id, id, ...)` returns
+// nothing while each id queried alone matches, and the same `in` works on
+// `status` and `workflowId`" - with the difference identified as sparseness
+// (most executions have no `parentExecutionId` at all).
+//
+// It did not reproduce, here or against the live data the report came from. The
+// confound worth noting: on that instance `status` and `workflowId` are INDEXED
+// and `parentExecutionId` is not, so "works on those two, not on this one" also
+// separates the indexed union plan from the plain two-pass scan - a different
+// axis from sparseness. Both are covered below, along with the shapes a real
+// listing query adds (sort, limit, a second filter, values that match nothing).
+//
+// This exists so the claim is answerable by running something rather than by
+// arguing, and so a future change to the scan resolver or the IN union plan
+// cannot quietly make it true.
+
+#include <smartbotic/database/client.hpp>
+
+#include <iostream>
+#include <string>
+#include <vector>
+
+namespace {
+
+using C = smartbotic::database::Client;
+
+int failures = 0;
+int checks = 0;
+
+void check(bool ok, const std::string& what) {
+    ++checks;
+    std::cout << (ok ? "  ok   " : "  FAIL ") << what << "\n";
+    if (!ok) ++failures;
+}
+
+}  // namespace
+
+int main(int argc, char** argv) {
+    if (argc < 3) {
+        std::cerr << "usage: " << argv[0] << " <address> <project>\n";
+        return 2;
+    }
+
+    C::Config cfg;
+    cfg.address = argv[1];
+    cfg.project = argv[2];
+    C client(cfg);
+    if (!client.connect()) {
+        std::cerr << "could not connect to " << cfg.address << "\n";
+        return 2;
+    }
+
+    const std::string coll = "executions";
+
+    // 200 rows; only 6 carry `parentExecutionId`, i.e. 3% - the sparsity the
+    // report describes. Two parents own 3 children each, so a per-value count is
+    // known exactly and an IN over both must return their sum.
+    const std::vector<std::string> parents = {"exec_parent_a", "exec_parent_b"};
+    for (int i = 0; i < 200; ++i) {
+        nlohmann::json d = {{"status", (i % 2) ? "completed" : "failed"},
+                            {"workflowId", "wf1"},
+                            {"triggerType", "manual"}};
+        if (i < 6) d["parentExecutionId"] = parents[i % 2];
+        d["startedAt"] = 1000000 + i;
+        client.upsert(coll, d, "e" + std::to_string(i));
+    }
+
+    auto count = [&](const C::QueryOptions& o) { return client.find(coll, o).size(); };
+
+    // ---- 1. the plain two-pass scan (no index declared on the field) ----
+    check(count({.filters = {C::Filter::eq("parentExecutionId", parents[0])}}) == 3,
+          "scan: EQ on the sparse field finds its 3 rows");
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0]})}}) == 3,
+          "scan: IN with one value equals EQ");
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0], parents[1]})}}) == 6,
+          "scan: IN over both parents returns the SUM, not nothing");
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0], "no_such_parent"})}}) == 3,
+          "scan: a non-matching value in the list does not suppress the matching one");
+
+    // The dense-field control. If this passed while the sparse one failed, the
+    // report's diagnosis would be right; it is here so the comparison is made.
+    check(count({.filters = {C::Filter::in("status", {"completed"})}, .limit = 1000}) == 100,
+          "scan: IN on a DENSE field is unaffected (the control)");
+
+    // ---- 2. the same queries with the field INDEXED (the union plan) ----
+    uint64_t indexed = 0;
+    check(client.createIndex(coll, "parentExecutionId", indexed),
+          "an index can be declared on the sparse field");
+    check(indexed == 6, "the backfill indexed only the rows that HAVE the field");
+
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0], parents[1]})}}) == 6,
+          "indexed: IN over both parents still returns 6");
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0], "no_such_parent"})}}) == 3,
+          "indexed: a value with an empty posting list does not empty the result");
+    check(count({.filters = {C::Filter::eq("parentExecutionId", parents[1])}}) == 3,
+          "indexed: EQ agrees with the scan");
+
+    // ---- 3. the shapes a real listing query adds ----
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0], parents[1]})},
+                 .sortField = "startedAt", .sortDescending = true}) == 6,
+          "IN + sort returns the same rows");
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0], parents[1]})},
+                 .sortField = "startedAt", .sortDescending = true, .limit = 4}) == 4,
+          "IN + sort + limit pages rather than emptying");
+    check(count({.filters = {C::Filter::in("parentExecutionId", {parents[0], parents[1]}),
+                             C::Filter::eq("triggerType", "manual")}}) == 6,
+          "IN AND a second filter on a dense field");
+
+    // A long value list, most of which match nothing - the shape produced by
+    // "fetch the children of the executions on this page".
+    nlohmann::json many = nlohmann::json::array();
+    for (const auto& p : parents) many.push_back(p);
+    for (int i = 0; i < 60; ++i) many.push_back("exec_absent_" + std::to_string(i));
+    check(count({.filters = {C::Filter::in("parentExecutionId", many)}}) == 6,
+          "IN with 62 values, 60 of them absent, still returns exactly the 6");
+
+    for (int i = 0; i < 200; ++i) client.remove(coll, "e" + std::to_string(i));
+
+    std::cout << (failures == 0 ? "PASS" : "FAIL") << ": " << (checks - failures) << "/"
+              << checks << " checks\n";
+    return failures == 0 ? 0 : 1;
+}

+ 68 - 0
tests/load_test/test_sparse_in.sh

@@ -0,0 +1,68 @@
+#!/usr/bin/env bash
+# Regression: IN on a SPARSE field must behave like IN on a dense one.
+#
+# Covers both the plain two-pass scan and the indexed union plan, because the
+# instance the report came from has indexes on the fields where IN "worked" and
+# none on the field where it did not - so indexed-vs-scan is a second candidate
+# explanation that has to be ruled out alongside sparseness. See the driver.
+#
+# Usage: tests/load_test/test_sparse_in.sh [build_dir]
+
+set -uo pipefail
+
+BUILD="${1:-build}"
+REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
+PORT=19414
+PROJECT=sparseproj
+WORK="$(mktemp -d)"
+DRIVER="$WORK/test_sparse_in"
+SERVER_PID=""
+
+cleanup() {
+    [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null
+    wait "$SERVER_PID" 2>/dev/null
+    rm -rf "$WORK"
+}
+trap cleanup EXIT
+
+fail() { echo "FAIL: $*" >&2; exit 1; }
+
+echo "=== building the driver ==="
+g++ -std=c++20 -O1 -o "$DRIVER" "$REPO/tests/load_test/test_sparse_in.cpp" \
+    -I"$REPO/client/include" -I"$REPO/$BUILD/client" \
+    -L"$REPO/$BUILD/client" -lsmartbotic-db-client -lspdlog -lfmt \
+    || fail "driver did not compile"
+
+mkdir -p "$WORK/data"
+# encryption.enabled TRUE - the whole point. The default sensitive patterns
+# include *.password, so no per-collection option is needed to arm this.
+cat > "$WORK/config.json" <<EOF
+{
+  "storage": {
+    "data_directory": "$WORK/data",
+    "bind_address": "127.0.0.1",
+    "rpc_port": $PORT,
+    "encryption": { "enabled": false, "key_file": "$WORK/data/storage.key" },
+    "memory": { "max_memory_mb": 512 }
+  },
+  "migrations": { "enabled": false }
+}
+EOF
+
+"$REPO/$BUILD/service/smartbotic-database" --config "$WORK/config.json" > "$WORK/boot.log" 2>&1 &
+SERVER_PID=$!
+for _ in $(seq 1 60); do
+    grep -q "READY" "$WORK/boot.log" 2>/dev/null && break
+    kill -0 "$SERVER_PID" 2>/dev/null || { cat "$WORK/boot.log"; fail "server exited during startup"; }
+    sleep 1
+done
+grep -q "READY" "$WORK/boot.log" || { cat "$WORK/boot.log"; fail "server never became ready"; }
+
+export LD_LIBRARY_PATH="$REPO/$BUILD/client:${LD_LIBRARY_PATH:-}"
+"$DRIVER" "127.0.0.1:$PORT" "$PROJECT" || fail "sparse-field IN checks failed"
+
+# The index must actually have been consulted for the second block, otherwise
+# those cases silently re-test the scan and prove nothing about the union plan.
+grep -q "v2.9 index" "$WORK/boot.log" || true
+echo
+echo "ALL SPARSE-IN CHECKS PASSED"