Forráskód Böngészése

perf: declare the indexes that help, expire every file, stop versioning what nothing reads

**Indexes.** Declared at webserver startup so a fresh install or a
restored backup gets them without anyone remembering. Measured before and
after rather than taken from the recommendation:

  sessions.refreshToken   24 ms -> 1 ms. Every token refresh scanned all
                          9,661 sessions.
  executions.workflowId   a lookup that misses: 1124 ms -> 1 ms.

What it did not fix is worth writing down. That same query still takes
~800 ms when it hits, and the index is not why: the cost scales with rows
returned, about 30 ms per execution document, because each carries every
node's input and output. A miss answering in 1 ms is what proves the
lookup is fine. Making that fast needs the documents to stop carrying
node output.

executions.startedAt was declared and dropped again. The note predicted
337 ms -> 9 ms; measured, it changed nothing - a range with limit 1 still
took 300 ms, a sort by it 422 ms. Ranges and sorts are not served from an
index in this version, and executions are the most-written collection
here, so it was pure write cost.

**Versioning** is off everywhere except workflows. Only the workflow
controller and the runner read a version, both for workflows; everywhere
else history was written on every save and read by nothing, unlimited.
executions is 481 MB and writes its record two or three times per run.
Existing history stays readable, so this is reversible and reclaims
nothing by itself.

Checked rather than assumed: workflow versioning still records - the
document is at v72 with 70 versions kept and publishedVersion intact. The
list showing "50" was its page size, not a failure to record.

**Every stored file now has an expiry**, 90 days. FileRecord carries no
expiry field, so a file that has one cannot be told from one that does
not and the sweep sets it on all 299.

One thing this leaves crooked, in the README as well: the documents
pointing at those files do not expire, because no workflow or project
declares a retention. In 90 days those records will name files that have
gone. A retention on the project closes it - that is the setting that
covers documents, executions and files together.
fszontagh 1 hónapja
szülő
commit
d47eca6503

+ 56 - 0
deploy/zeus/README.md

@@ -113,3 +113,59 @@ Not a procedure on paper - this was run:
 It recovered 27,433 documents in 95 collections and served all 9 workflows by
 name, 7 credentials and 1 user. A 2 MB file downloaded byte-exact, which is what
 proves the encryption key came through - counting documents would not have.
+
+## Indexes
+
+Declared at webserver startup, idempotently, so a fresh install or a restored
+backup gets them without anyone remembering to. They live in the database, not
+in a config file.
+
+Only what was measured to help - an index costs write throughput, so one that
+buys nothing is a cost paid for ever:
+
+| index | before | after |
+|---|---|---|
+| `sessions.refreshToken` | 24 ms | **1 ms** |
+| `executions.workflowId` (miss) | 1124 ms | **1 ms** |
+| `workflows.projectId` | small today | scales with the workflow count |
+| `users.username`, `users.email` | 1 row | scales with the user count |
+
+`executions.workflowId` still takes ~800 ms when it *hits*, and that is not the
+index: cost scales with rows returned, about 30 ms per execution document,
+because each carries every node's input and output. A miss answers in 1 ms,
+which is what shows the lookup itself is fine. Making that query fast needs the
+documents to stop carrying node output, not another index.
+
+**`executions.startedAt` was declared and dropped again.** Measured, it changed
+nothing - a range with `limit 1` still took 300 ms and a sort by it 422 ms.
+Ranges and sorts are not served from an index in this version. Executions are
+the most-written collection here, so it was pure cost.
+
+## Versioning
+
+Off everywhere except `workflows`. Only two places read a version - the workflow
+controller (the editor's picker, publish, restore) and the runner (loading the
+published version to run) - and both are about workflows.
+
+Everywhere else the history was written on every save and read by nothing, with
+`maxVersions` 0 meaning unlimited. `executions` is the worst of it: 481 MB, and
+each run writes its record two or three times.
+
+Disabling stops new versions being recorded; history already written stays
+readable, so it is reversible and **reclaims nothing by itself** - the existing
+history directory is still there.
+
+## File expiry
+
+Every stored file has one. A file's TTL could only be set at upload before 2.8,
+so anything written before that was immortal, and `FileRecord` carries no expiry
+field - a file that already has one cannot be told from one that does not, so
+the sweep sets one on all of them.
+
+90 days was chosen: source photos, generated images and downloads, all either
+regenerable or already published elsewhere.
+
+**The documents that point at those files do not expire yet.** No workflow or
+project declares a retention, so in 90 days those records will name files that
+have gone. Setting a retention on the project closes that - it covers documents,
+executions and files together.

+ 45 - 0
deploy/zeus/maintenance/disable-versioning.cpp

@@ -0,0 +1,45 @@
+// Turn version history off everywhere except workflows.
+//
+// Only two places in the codebase read a version: the workflow controller (the
+// editor's version picker, publish and restore) and the runner (loading the
+// published version to run it). Both are about workflows. Everywhere else the
+// history is written on every save and read by nothing - and maxVersions 0
+// means unlimited, so it never stops growing.
+//
+// executions is the worst of it: 481 MB, and each run writes its record two or
+// three times - before the walk, on a pause, at the end - so every one of those
+// is kept for ever.
+//
+// Disabling stops new versions being recorded. History already written stays
+// readable, so this is reversible and reclaims nothing by itself.
+#include <smartbotic/database/client.hpp>
+#include <cstdio>
+#include <string>
+#include <vector>
+using namespace smartbotic::database;
+int main() {
+    Client::Config cfg; cfg.address="192.168.2.150:9004"; cfg.project="smartbotic-automation";
+    Client c(cfg);
+    if (!c.connect()) return 1;
+    const std::vector<std::string> keep = {"workflows"};
+    std::vector<std::string> names;
+    for (auto& n : c.listCollections()) {
+        auto bare = n.rfind(':')==std::string::npos ? n : n.substr(n.rfind(':')+1);
+        names.push_back(bare);
+    }
+    int off=0, kept=0, skipped=0;
+    for (auto& n : names) {
+        auto current = c.getCollectionConfig(n);
+        bool on = current.versioningEnabled.value_or(true);
+        bool must_keep = false;
+        for (auto& k : keep) if (n == k) must_keep = true;
+        if (must_keep) { printf("  %-28s left ON (its history is read)\n", n.c_str()); kept++; continue; }
+        if (!on) { skipped++; continue; }
+        Client::CollectionConfig want;          // precision left unset, so untouched
+        want.versioningEnabled = false;
+        if (c.configureCollection(n, want)) { printf("  %-28s versioning off\n", n.c_str()); off++; }
+        else printf("  %-28s FAILED\n", n.c_str());
+    }
+    printf("turned off %d, left on %d, already off %d\n", off, kept, skipped);
+    return 0;
+}

+ 33 - 0
deploy/zeus/maintenance/set-file-ttl.cpp

@@ -0,0 +1,33 @@
+// Give every stored file an expiry.
+//
+// A file's TTL could only be set at upload before 2.8, so everything written
+// before that is immortal - and FileRecord carries no expiry field, so a file
+// that already has one cannot be told from a file that does not. The only
+// available move is to set one on all of them.
+//
+// 90 days: these are working data - source photos, generated images, downloads -
+// all of it either regenerable or already published elsewhere. Long enough that
+// nothing in flight is lost, short enough that the store does not grow for ever.
+#include <smartbotic/database/client.hpp>
+#include <cstdio>
+#include <set>
+using namespace smartbotic::database;
+int main(int argc, char** argv) {
+    const uint32_t ttl = (argc > 1) ? (uint32_t)atoi(argv[1]) : 90u*24*3600;
+    Client::Config cfg; cfg.address="192.168.2.150:9004"; cfg.project="smartbotic-automation";
+    Client c(cfg);
+    if (!c.connect()) return 1;
+    std::set<std::string> seen; int set_ok=0, failed=0;
+    for (uint32_t off=0; off<5000; off+=100) {
+        auto page = c.listFiles("", "", 100, off);
+        if (page.files.empty()) break;
+        for (auto& f : page.files) {
+            if (!seen.insert(f.id).second) continue;
+            if (c.setFileTtl(f.id, ttl)) set_ok++; else failed++;
+        }
+        if (page.files.size() < 100) break;
+    }
+    printf("files seen %zu, expiry set on %d, failed %d (ttl %u s = %.0f days)\n",
+           seen.size(), set_ok, failed, ttl, ttl/86400.0);
+    return failed ? 1 : 0;
+}

+ 27 - 0
lib/storage/storage_client.cpp

@@ -336,6 +336,33 @@ Result<CollectionConfig> StorageClient::getCollectionConfig(const std::string& c
     return cfg;
 }
 
+Result<int64_t> StorageClient::createIndex(const std::string& collection,
+                                           const std::string& field) {
+    uint64_t rows = 0;
+    if (!impl_->client_->createIndex(collection, field, rows)) {
+        return Error(ErrorCode::DatabaseError,
+                     "Could not create index on " + collection + "." + field);
+    }
+    return static_cast<int64_t>(rows);
+}
+
+Result<void> StorageClient::dropIndex(const std::string& collection, const std::string& field) {
+    if (!impl_->client_->dropIndex(collection, field)) {
+        return Error(ErrorCode::DatabaseError,
+                     "Could not drop index on " + collection + "." + field);
+    }
+    return {};
+}
+
+std::vector<StorageClient::IndexInfo> StorageClient::listIndexes(const std::string& collection) {
+    std::vector<IndexInfo> out;
+    for (const auto& def : impl_->client_->listIndexes(collection)) {
+        out.push_back(IndexInfo{def.field, static_cast<int64_t>(def.distinctValues),
+                                static_cast<int64_t>(def.entries)});
+    }
+    return out;
+}
+
 Result<void> StorageClient::dropCollection(const std::string& name) {
     bool ok = impl_->client_->dropCollection(name);
     if (!ok) return Error(ErrorCode::CollectionNotFound, "dropCollection failed: " + name);

+ 19 - 0
lib/storage/storage_client.hpp

@@ -196,6 +196,25 @@ public:
 
     common::Result<void> dropCollection(const std::string& name);
 
+    // Secondary index on one field. Idempotent, and backfilled over the rows
+    // already there, so it is usable as soon as it returns; rows_indexed says
+    // how many the backfill covered.
+    //
+    // Declaring one is a judgement, not a free win: an index costs write
+    // throughput, and on a field with few distinct values it loses to a scan.
+    // The server's planner re-decides per query and declines an index that is
+    // not selective, so a wrong declaration wastes writes rather than making
+    // reads slower. Equality filters only - ranges still scan.
+    common::Result<int64_t> createIndex(const std::string& collection, const std::string& field);
+    common::Result<void> dropIndex(const std::string& collection, const std::string& field);
+
+    struct IndexInfo {
+        std::string field;
+        int64_t distinct_values = 0;  // few, over many rows, means the planner will decline it
+        int64_t entries = 0;
+    };
+    std::vector<IndexInfo> listIndexes(const std::string& collection);
+
     std::vector<std::string> listCollections();
     common::Result<CollectionInfo> getCollectionInfo(const std::string& name);
 

+ 39 - 0
src/webserver/webserver_service.cpp

@@ -40,6 +40,45 @@ WebServerService::WebServerService(const WebServerServiceConfig& config)
     // Initialize auth store
     auth_store_ = std::make_unique<auth::AuthStore>(*storage_, *jwt_);
 
+    // The indexes this installation's queries actually need. Declared here, and
+    // idempotently, so a fresh install or a restored backup gets them without
+    // anyone remembering to - they live in the database, not in the config.
+    //
+    // Only what was measured to help. An index costs write throughput, so a
+    // declaration that buys nothing is a real cost paid for ever:
+    //
+    //   sessions.refreshToken   every token refresh scanned the whole
+    //                           collection - 24 ms over 9,661 sessions, 1 ms now
+    //   executions.workflowId   a lookup that missed took 1,124 ms and now takes
+    //                           1. What is left when it hits is the size of the
+    //                           execution documents, about 30 ms each, which no
+    //                           index can help with
+    //   workflows.projectId     small today; the listing filters by it on every
+    //                           page load and workflows are written by hand
+    //   users.username/email    login looks up by both
+    //
+    // executions.startedAt was declared and dropped again: measured, it changed
+    // nothing. Ranges and sorts are not served from an index in 2.9 - a sorted
+    // single row still took 422 ms - and executions are the most written
+    // collection here, so it was pure cost.
+    for (const auto& [collection, field] : std::initializer_list<std::pair<const char*, const char*>>{
+             {"executions", "workflowId"},
+             {"workflows", "projectId"},
+             {"users", "username"},
+             {"users", "email"},
+             {"sessions", "refreshToken"}}) {
+        auto created = storage_->createIndex(collection, field);
+        if (created.failed()) {
+            // Not fatal: an older database has no index support and every query
+            // still works, just by scanning.
+            LOG_DEBUG("Index {}.{} not declared: {}", collection, field,
+                      created.error().message());
+        } else if (created.value() > 0) {
+            LOG_INFO("Index {}.{} declared, {} rows backfilled", collection, field,
+                     created.value());
+        }
+    }
+
     // Sessions issued under a longer lifetime than is configured now would
     // otherwise keep it until they ran out, so shortening the setting would not
     // take effect for as long as the old one lasted.