Przeglądaj źródła

feat: a retention change reaches the data already written

A retention that only governed future writes would be a setting about
nothing: the data somebody wants gone is the data already there. Changing
a workflow's retention now restamps what exists, and deleting a workflow
clears what it wrote instead of leaving it behind for ever - which is
what happened until now. Nothing removed a deleted workflow's executions.
Nothing pointed at them either.

Because lowering a retention deletes data the moment it is applied, the
size of that is offered before the change rather than discovered after
it. GET /workflows/{id}/retention?ttlSeconds=N counts what would go: how
many executions and documents there are, how many are already older than
the proposed retention, and how far back the data goes. It is counted,
not estimated.

The restamp is per document, and has to be. A collection's default TTL
can only be set when the collection is created, and nothing upstream
re-dates a range of documents, so each one is read and written back with
its new expiry - which is why it runs on a worker and not inside the
request that asked for it. It goes through upsert rather than update:
only insert and upsert carry a TTL, and upsert swaps the expiry in one
locked server-side step. It preserves _created_at, checked against the
live database first, because a retention change that quietly re-dated
every execution would be worse than the growth it was meant to fix.

The paging deserves its comment. Lowering a retention deletes documents
as the pass runs, so the result set shrinks underneath a cursor and
advancing the page number would step straight over the documents that
moved up into the gap. It reads page one repeatedly and remembers the ids
it has handled - which also ends the loop when nothing is being deleted
and page one would otherwise return the same documents for ever.

Deleting a workflow does not walk its data synchronously. Everything it
owns is given a short expiry and clears itself; the collection is dropped
after the pass. A delete that has to remove ten thousand documents before
the button responds is a delete people learn not to press.

**Files are out of reach, and this says so rather than pretending.** A
file's TTL can only be set when it is uploaded, and the file store can be
searched by type, name, checksum or related id - not by workflow. Files
therefore carry the retention in force when they were written and a later
change does not reach them. The preview returns filesReachable: false
instead of quietly reporting a number that excludes them.

Verified end to end: an execution and a document written before any
retention existed both expired because of a setting applied afterwards;
the preview counted 136 executions on a real workflow and correctly said
0 would go at seven days, all 136 at one hour, and 0 at keep-for-ever;
deleting a workflow dropped its collection within seconds and its
executions went at sixty. 67 passed, 0 failed, 2 skipped.
fszontagh 1 miesiąc temu
rodzic
commit
8a52564c59

+ 1 - 0
CMakeLists.txt

@@ -158,6 +158,7 @@ add_executable(smartbotic-webserver
     src/webserver/auth/auth_store.cpp
     src/webserver/auth/access.cpp
     src/webserver/api/project_controller.cpp
+    src/webserver/retention/retention_service.cpp
     src/webserver/auth/auth_middleware.cpp
     src/webserver/runners/runner_registry.cpp
     src/webserver/runners/load_balancer.cpp

+ 84 - 3
src/webserver/api/workflow_controller.cpp

@@ -18,7 +18,8 @@ WorkflowController::WorkflowController(storage::StorageClient& storage,
                                        WorkflowScheduler& scheduler,
                                        DatabaseWatcher& db_watcher,
                                        auth::AccessControl& access,
-                                       nodes::NodeStore& node_store)
+                                       nodes::NodeStore& node_store,
+                                       retention::RetentionService& retention)
     : storage_(storage)
     , middleware_(middleware)
     , registry_(registry)
@@ -27,7 +28,8 @@ WorkflowController::WorkflowController(storage::StorageClient& storage,
     , scheduler_(scheduler)
     , db_watcher_(db_watcher)
     , access_(access)
-    , node_store_(node_store) {}
+    , node_store_(node_store)
+    , retention_(retention) {}
 
 
 namespace {
@@ -102,6 +104,16 @@ void WorkflowController::registerRoutes(httplib::Server& server) {
         });
     });
 
+    // What a retention change would cost, before it is made. Registered ahead
+    // of nothing in particular - httplib matches in registration order, and the
+    // bare GET above is anchored, so this needs its own entry rather than being
+    // swallowed by it.
+    server.Get(R"(/api/v1/workflows/([^/]+)/retention)", [this](const httplib::Request& req, httplib::Response& res) {
+        middleware_.requireAuth(req, res, [this](auto& req, auto& res, auto& ctx) {
+            previewRetention(req, res, ctx);
+        });
+    });
+
     server.Post(R"(/api/v1/workflows/([^/]+)/execute)", [this](const httplib::Request& req, httplib::Response& res) {
         middleware_.requireAuth(req, res, [this](auto& req, auto& res, auto& ctx) {
             executeWorkflow(req, res, ctx);
@@ -512,6 +524,20 @@ void WorkflowController::updateWorkflow(const httplib::Request& req, httplib::Re
         // Get updated workflow
         auto workflow = storage_.get("workflows", id);
         if (workflow.ok()) {
+            // A retention that only governed future writes would be a setting
+            // about nothing - the data somebody wants gone is the data already
+            // there - so a change is applied to what exists too. Only on an
+            // actual change: restamping every execution on every autosave would
+            // rewrite the whole history several times a minute.
+            const auto before_retention = retention_.effectiveFor(existing);
+            const auto after_retention = retention_.effectiveFor(workflow.value());
+            const auto ttlOf = [](const std::optional<storage::Retention>& r) {
+                return r ? r->ttl_seconds : -1;
+            };
+            if (ttlOf(before_retention) != ttlOf(after_retention) && after_retention) {
+                retention_.apply(id, after_retention->ttl_seconds);
+            }
+
             // Reconcile the scheduler in both directions. Registering on active
             // but never unregistering on inactive let the stored flag and the
             // scheduler drift apart, leaving a deactivated workflow still firing.
@@ -546,6 +572,53 @@ void WorkflowController::updateWorkflow(const httplib::Request& req, httplib::Re
     }
 }
 
+void WorkflowController::previewRetention(const httplib::Request& req, httplib::Response& res,
+                                           const auth::AuthContext& ctx) {
+    const std::string id = req.matches[1];
+
+    nlohmann::json workflow;
+    if (!loadAllowed(req, res, ctx, id, auth::Action::Read, workflow)) return;
+
+    // No ttlSeconds asked about means "what is in force now, and what does it
+    // cover" - the same numbers, measured against the current setting.
+    int64_t ttl_seconds = 0;
+    if (req.has_param("ttlSeconds")) {
+        try {
+            ttl_seconds = std::stoll(req.get_param_value("ttlSeconds"));
+        } catch (const std::exception&) {
+            sendError(res, "ttlSeconds must be a whole number of seconds", 400);
+            return;
+        }
+        if (ttl_seconds < 0) {
+            sendError(res, "ttlSeconds cannot be negative. 0 means keep for ever", 400);
+            return;
+        }
+    } else if (const auto current = retention_.effectiveFor(workflow)) {
+        ttl_seconds = current->ttl_seconds;
+    }
+
+    const auto preview = retention_.preview(id, ttl_seconds);
+    const auto impact = [](const retention::RetentionService::Impact& i) {
+        return nlohmann::json{{"total", i.total},
+                              {"expiringNow", i.expiring_now},
+                              {"oldestAgeMs", i.oldest_age_ms}};
+    };
+
+    const auto current = retention_.effectiveFor(workflow);
+    sendJson(res, {
+        {"ttlSeconds", ttl_seconds},
+        {"current", current ? nlohmann::json{{"ttlSeconds", current->ttl_seconds},
+                                             {"inherited", current->inherited}}
+                            : nlohmann::json(nullptr)},
+        {"executions", impact(preview.executions)},
+        {"documents", impact(preview.documents)},
+        // Said rather than left to be discovered: a file's expiry can only be
+        // set when it is uploaded, so a retention change does not reach the
+        // files a workflow has already stored.
+        {"filesReachable", preview.files_reachable},
+    });
+}
+
 void WorkflowController::deleteWorkflow(const httplib::Request& req, httplib::Response& res,
                                          const auth::AuthContext& ctx) {
     std::string id = req.matches[1];
@@ -564,10 +637,18 @@ void WorkflowController::deleteWorkflow(const httplib::Request& req, httplib::Re
         return;
     }
 
+    // Its executions and its own storage outlive it otherwise: nothing pointed
+    // at them, and nothing deleted them. They are given a short expiry and clear
+    // themselves rather than being walked through here - a delete that has to
+    // remove ten thousand documents before the button responds is a delete
+    // people learn not to press.
+    retention_.retire(id);
+
     ws_server_.broadcast("workflows.deleted", {{"id", id}});
 
     LOG_INFO("Workflow deleted: {} by user {}", id, ctx.user_id);
-    sendJson(res, {{"success", true}});
+    sendJson(res, {{"success", true},
+                   {"dataRemovedIn", retention::RetentionService::kRetireTtlSeconds}});
 }
 
 void WorkflowController::executeWorkflow(const httplib::Request& req, httplib::Response& res,

+ 8 - 1
src/webserver/api/workflow_controller.hpp

@@ -9,6 +9,7 @@
 #include "../websocket_server.hpp"
 #include "../scheduler/workflow_scheduler.hpp"
 #include "../nodes/node_store.hpp"
+#include "../retention/retention_service.hpp"
 #include "../scheduler/database_watcher.hpp"
 #include "../auth/access.hpp"
 #include "proto/runner.grpc.pb.h"
@@ -25,7 +26,8 @@ public:
                       WorkflowScheduler& scheduler,
                       DatabaseWatcher& db_watcher,
                       auth::AccessControl& access,
-                      nodes::NodeStore& node_store);
+                      nodes::NodeStore& node_store,
+                      retention::RetentionService& retention);
 
     void registerRoutes(httplib::Server& server);
 
@@ -42,6 +44,10 @@ private:
     void deleteWorkflow(const httplib::Request& req, httplib::Response& res,
                         const auth::AuthContext& ctx);
 
+    // What a retention change would cost, counted before it is made.
+    void previewRetention(const httplib::Request& req, httplib::Response& res,
+                          const auth::AuthContext& ctx);
+
     // Workflow operations
     void executeWorkflow(const httplib::Request& req, httplib::Response& res,
                          const auth::AuthContext& ctx);
@@ -69,6 +75,7 @@ private:
     DatabaseWatcher& db_watcher_;
     auth::AccessControl& access_;
     nodes::NodeStore& node_store_;
+    retention::RetentionService& retention_;
 
     // Helper to check for scheduled triggers and register/unregister with scheduler
     void updateScheduledTriggers(const std::string& workflow_id, bool activate);

+ 223 - 0
src/webserver/retention/retention_service.cpp

@@ -0,0 +1,223 @@
+#include "retention_service.hpp"
+
+#include "common/time_utils.hpp"
+#include "logging/logger.hpp"
+#include "storage/workflow_collection.hpp"
+
+namespace smartbotic::webserver::retention {
+
+namespace {
+
+constexpr const char* kExecutions = "executions";
+constexpr const char* kExecutionWorkflowField = "workflowId";
+// Nothing links a document in a workflow's own collection back to the workflow:
+// it does not need one, because the collection belongs to that workflow and
+// nothing else writes to it. An empty field name means "every document here".
+constexpr const char* kNoWorkflowField = "";
+
+constexpr int32_t kPageSize = 200;
+
+}  // namespace
+
+RetentionService::RetentionService(storage::StorageClient& storage) : storage_(storage) {
+    worker_ = std::thread([this] { worker(); });
+}
+
+RetentionService::~RetentionService() {
+    stopping_ = true;
+    cv_.notify_all();
+    if (worker_.joinable()) worker_.join();
+}
+
+void RetentionService::apply(const std::string& workflow_id, int64_t ttl_seconds) {
+    if (workflow_id.empty()) return;
+    {
+        std::lock_guard<std::mutex> lock(mutex_);
+        queue_.push_back(Job{workflow_id, ttl_seconds, /*drop_collection_after=*/false});
+    }
+    cv_.notify_one();
+}
+
+void RetentionService::retire(const std::string& workflow_id) {
+    if (workflow_id.empty()) return;
+    {
+        std::lock_guard<std::mutex> lock(mutex_);
+        queue_.push_back(Job{workflow_id, kRetireTtlSeconds, /*drop_collection_after=*/true});
+    }
+    cv_.notify_one();
+}
+
+void RetentionService::worker() {
+    while (!stopping_) {
+        Job job;
+        {
+            std::unique_lock<std::mutex> lock(mutex_);
+            cv_.wait(lock, [this] { return stopping_ || !queue_.empty(); });
+            if (stopping_) return;
+            job = queue_.front();
+            queue_.pop_front();
+        }
+
+        const std::string own = storage::workflowCollectionName(job.workflow_id);
+
+        const int64_t executions =
+            restamp(kExecutions, kExecutionWorkflowField, job.workflow_id, job.ttl_seconds);
+        const int64_t documents = restamp(own, kNoWorkflowField, job.workflow_id, job.ttl_seconds);
+
+        LOG_INFO("Retention: workflow {} restamped to {}s - {} executions, {} documents",
+                 job.workflow_id, job.ttl_seconds, executions, documents);
+
+        if (job.drop_collection_after) {
+            // The documents are not gone yet - they expire on their own shortly.
+            // Dropping the collection now would take them with it, which is the
+            // same outcome sooner; what it must not do is drop a collection
+            // whose documents are still being restamped, so it happens here,
+            // after the pass, on the same worker.
+            auto dropped = storage_.dropCollection(own);
+            if (dropped.failed()) {
+                LOG_WARN("Retention: could not drop {} ({}); its documents still expire on their own",
+                         own, dropped.error().message());
+            } else {
+                LOG_INFO("Retention: dropped {} for deleted workflow {}", own, job.workflow_id);
+            }
+        }
+    }
+}
+
+RetentionService::Impact RetentionService::measure(const std::string& collection,
+                                                   const std::string& workflow_field,
+                                                   const std::string& workflow_id,
+                                                   int64_t ttl_seconds) {
+    Impact impact;
+    const int64_t now = common::TimeUtils::nowMs();
+    const int64_t ttl_ms = ttl_seconds * 1000;
+
+    int32_t page = 1;
+    while (true) {
+        storage::QueryOptions opts;
+        opts.page = page;
+        opts.page_size = kPageSize;
+        if (!workflow_field.empty()) {
+            opts.filters.emplace_back(workflow_field, workflow_id);
+        }
+        // Only the creation stamp is needed to answer this. Asking for the whole
+        // document would pull every execution's node output through the wire to
+        // count them.
+        opts.fields = {"_created_at"};
+
+        auto result = storage_.query(collection, opts);
+        if (result.failed()) {
+            // A collection that is not there yet holds nothing to lose. Anything
+            // else is worth saying out loud rather than reporting as zero.
+            if (result.error().code() != common::ErrorCode::CollectionNotFound) {
+                LOG_WARN("Retention: cannot measure {} ({})", collection, result.error().message());
+            }
+            return impact;
+        }
+
+        for (const auto& doc : result.value().documents) {
+            impact.total++;
+            const int64_t created = common::TimeUtils::documentStamp(doc, "_created_at");
+            if (created <= 0) continue;
+            const int64_t age = now - created;
+            if (age > impact.oldest_age_ms) impact.oldest_age_ms = age;
+            // Keeping for ever deletes nothing, whatever the ages are.
+            if (ttl_seconds > 0 && age >= ttl_ms) impact.expiring_now++;
+        }
+
+        if (!result.value().has_more) break;
+        page++;
+    }
+    return impact;
+}
+
+std::optional<storage::Retention> RetentionService::effectiveFor(const nlohmann::json& workflow) {
+    const auto settings = workflow.value("settings", nlohmann::json::object());
+    if (const auto own = storage::declaredTtlSeconds(settings)) {
+        return storage::Retention{*own, false};
+    }
+
+    const std::string project_id = workflow.value("projectId", std::string{});
+    if (project_id.empty()) return std::nullopt;
+
+    auto project = storage_.get("projects", project_id);
+    if (project.failed()) return std::nullopt;
+    return storage::declaredRetention(
+        settings, project.value().value("settings", nlohmann::json::object()));
+}
+
+RetentionService::Preview RetentionService::preview(const std::string& workflow_id,
+                                                    int64_t ttl_seconds) {
+    Preview preview;
+    preview.executions =
+        measure(kExecutions, kExecutionWorkflowField, workflow_id, ttl_seconds);
+    preview.documents = measure(storage::workflowCollectionName(workflow_id), kNoWorkflowField,
+                                workflow_id, ttl_seconds);
+    return preview;
+}
+
+int64_t RetentionService::restamp(const std::string& collection,
+                                  const std::string& workflow_field,
+                                  const std::string& workflow_id,
+                                  int64_t ttl_seconds) {
+    int64_t restamped = 0;
+
+    // Always page 1. Restamping with a shorter retention deletes documents as it
+    // goes, so the result set shrinks underneath a cursor and advancing the page
+    // number would step over the documents that moved up into the space. Reading
+    // the first page repeatedly, and stopping when a pass changes nothing,
+    // cannot skip a document.
+    //
+    // Ids already handled are remembered for the same reason in reverse: with a
+    // longer retention nothing is deleted, so page 1 returns the same documents
+    // for ever and the loop would not end.
+    std::unordered_set<std::string> done;
+
+    while (!stopping_) {
+        storage::QueryOptions opts;
+        opts.page = 1;
+        opts.page_size = kPageSize;
+        if (!workflow_field.empty()) {
+            opts.filters.emplace_back(workflow_field, workflow_id);
+        }
+
+        auto result = storage_.query(collection, opts);
+        if (result.failed()) {
+            if (result.error().code() != common::ErrorCode::CollectionNotFound) {
+                LOG_WARN("Retention: cannot read {} ({})", collection, result.error().message());
+            }
+            return restamped;
+        }
+        if (result.value().documents.empty()) break;
+
+        int64_t handled_this_pass = 0;
+        for (const auto& doc : result.value().documents) {
+            if (stopping_) return restamped;
+            const std::string id = doc.value("_id", "");
+            if (id.empty() || done.contains(id)) continue;
+
+            // upsert rather than update: only insert and upsert carry a TTL, and
+            // upsert swaps the expiry entry in one locked server-side step. It
+            // preserves _created_at, so a restamped record still says when it
+            // was made - verified against the live database, because a retention
+            // change that quietly re-dated every execution would be worse than
+            // the growth it was meant to fix.
+            auto written = storage_.upsert(collection, doc, id, ttl_seconds * 1000);
+            if (written.failed()) {
+                LOG_WARN("Retention: could not restamp {}/{} ({})", collection, id,
+                         written.error().message());
+            } else {
+                restamped++;
+            }
+            done.insert(id);
+            handled_this_pass++;
+        }
+
+        // Nothing new on a full pass means every document has been seen.
+        if (handled_this_pass == 0) break;
+    }
+
+    return restamped;
+}
+
+}  // namespace smartbotic::webserver::retention

+ 103 - 0
src/webserver/retention/retention_service.hpp

@@ -0,0 +1,103 @@
+#pragma once
+
+#include <atomic>
+#include <condition_variable>
+#include <deque>
+#include <mutex>
+#include <string>
+#include <thread>
+#include <unordered_set>
+
+#include <optional>
+
+#include "storage/retention.hpp"
+#include "storage/storage_client.hpp"
+
+namespace smartbotic::webserver::retention {
+
+// Applies a retention change to data that has already been written.
+//
+// A retention that only governed future writes would be a setting about
+// nothing: the data somebody wants gone is the data already there. So changing
+// it restamps what exists - and because lowering a retention can delete data the
+// moment it is applied, the size of that is offered first and the change is a
+// deliberate answer to it rather than a surprise.
+//
+// The restamp is per document, and has to be. A collection's default TTL can
+// only be set when the collection is created, and nothing upstream re-dates a
+// range of documents, so each one is read and written back with its new expiry.
+// That is slow enough to belong on a worker rather than inside the request that
+// asked for it - nine and a half thousand executions is a real number here.
+//
+// **Files are out of reach.** A file's TTL can only be set when it is uploaded,
+// and the file store can be searched by type, name, checksum or related id but
+// not by workflow. A workflow's files therefore keep whatever retention was in
+// force when they were written, and a later change does not reach them. The
+// workflow id is recorded on every file as it is stored so that this becomes
+// possible the moment the file store can be asked for it.
+
+class RetentionService {
+public:
+    explicit RetentionService(storage::StorageClient& storage);
+    ~RetentionService();
+
+    RetentionService(const RetentionService&) = delete;
+    RetentionService& operator=(const RetentionService&) = delete;
+
+    struct Impact {
+        int64_t total = 0;         // documents the change would restamp
+        int64_t expiring_now = 0;  // already older than the new retention: these go at once
+        int64_t oldest_age_ms = 0; // how far back the data goes, for a truthful warning
+    };
+
+    struct Preview {
+        Impact executions;
+        Impact documents;   // the workflow's own collection
+        bool files_reachable = false;  // always false today - see the note above
+    };
+
+    // The retention actually in force for a workflow, resolving what it
+    // inherits from its project. Empty when neither has said anything, which is
+    // not the same as keeping for ever - see storage/retention.hpp.
+    std::optional<storage::Retention> effectiveFor(const nlohmann::json& workflow);
+
+    // What changing this workflow's retention to ttl_seconds would do, counted
+    // rather than estimated. ttl_seconds 0 means keep for ever, which never
+    // deletes anything and reports expiring_now 0.
+    Preview preview(const std::string& workflow_id, int64_t ttl_seconds);
+
+    // Restamp this workflow's data with a new retention, on the worker.
+    void apply(const std::string& workflow_id, int64_t ttl_seconds);
+
+    // A workflow has been deleted. Rather than a synchronous delete of
+    // everything it ever wrote, its data is given a short expiry and clears
+    // itself; the empty collection is dropped afterwards. A delete that has to
+    // walk ten thousand documents before the button responds is a delete people
+    // learn not to press.
+    void retire(const std::string& workflow_id);
+
+    static constexpr int64_t kRetireTtlSeconds = 60;
+
+private:
+    struct Job {
+        std::string workflow_id;
+        int64_t ttl_seconds = 0;
+        bool drop_collection_after = false;
+    };
+
+    void worker();
+    // Returns how many documents were restamped.
+    int64_t restamp(const std::string& collection, const std::string& workflow_field,
+                    const std::string& workflow_id, int64_t ttl_seconds);
+    Impact measure(const std::string& collection, const std::string& workflow_field,
+                   const std::string& workflow_id, int64_t ttl_seconds);
+
+    storage::StorageClient& storage_;
+    std::deque<Job> queue_;
+    std::mutex mutex_;
+    std::condition_variable cv_;
+    std::thread worker_;
+    std::atomic<bool> stopping_{false};
+};
+
+}  // namespace smartbotic::webserver::retention

+ 5 - 1
src/webserver/webserver_service.cpp

@@ -281,9 +281,13 @@ void WebServerService::setupRoutes() {
         *storage_, *auth_middleware_, *access_, *auth_store_);
     project_ctrl_->registerRoutes(server);
 
+    // Constructed before the controller that holds a reference to it, and
+    // destroyed after: it owns a worker thread that touches storage_.
+    retention_ = std::make_unique<retention::RetentionService>(*storage_);
+
     workflow_ctrl_ = std::make_unique<api::WorkflowController>(
         *storage_, *auth_middleware_, *runner_registry_, *load_balancer_, *ws_server_,
-        *scheduler_, *db_watcher_, *access_, *node_store_);
+        *scheduler_, *db_watcher_, *access_, *node_store_, *retention_);
     workflow_ctrl_->registerRoutes(server);
 
     workflow_group_ctrl_ = std::make_unique<api::WorkflowGroupController>(

+ 2 - 0
src/webserver/webserver_service.hpp

@@ -16,6 +16,7 @@
 #include "scheduler/database_watcher.hpp"
 #include "api/project_controller.hpp"
 #include "auth/access.hpp"
+#include "retention/retention_service.hpp"
 
 namespace smartbotic::webserver::nodes {
     class NodeStore;
@@ -127,6 +128,7 @@ private:
     std::unique_ptr<runners::RunnerRegistry> runner_registry_;
     std::unique_ptr<runners::LoadBalancer> load_balancer_;
     std::unique_ptr<nodes::NodeStore> node_store_;
+    std::unique_ptr<retention::RetentionService> retention_;
     std::unique_ptr<credentials::CredentialStore> credential_store_;
 
     // Servers