Quellcode durchsuchen

feat: give every workflow storage of its own, and a retention it inherits

Projects separate workflows and credentials but never separated data.
Collection permissions are by name, so a workflow set to read-write by
default could read and overwrite any collection it could name, including
one belonging to a workflow in another project. Nothing stopped it and
nothing recorded it.

A workflow now has its own collection, named for the workflow itself -
ids already begin with wf_, so the name needs no second prefix and that
prefix is what marks a collection as somebody's private data. It is
granted to its owner without being asked for, and refused to everyone
else whatever the settings say. The check runs before the settings are
consulted at all, because defaultAccess is the dangerous case: it is a
blanket yes to every name not listed, and that is the setting that would
otherwise hand over every other workflow's data.

Nobody can be expected to paste their own workflow's id into a node, and
a duplicated workflow would carry the original's id and quietly write
into somebody else's storage, so the collection is addressed as "@self".
It is created on first write - a workflow that stores nothing should not
leave an empty collection behind - and offered in the collection list
before it exists, because one that only appears after you have written to
it cannot be picked from a list.

Then retention. A project sets how long data is kept and a workflow
inherits it, or overrides it. The same lifetime covers the workflow's own
documents, the files it stores and its execution records, so a run's
history and the files that history points at go together rather than one
outliving the other by months.

Absent is not zero. A workflow that has said nothing inherits; an
explicit 0 means keep for ever, which is a choice somebody made. Reading
those as the same value would either start deleting documents nobody
agreed to lose or stop executions ageing out at all, so "nothing
declared" is a third answer the callers handle themselves - documents
keep the lifetime they have always had, executions keep their week.

Two upstream fixes this rests on, both verified here rather than taken on
trust: a file can now carry a TTL, and updating a document no longer
wipes its expiry. The probe that confirmed them found a third thing -
downloadFile throws for a missing file rather than returning empty, so
the adapter's Result handling was bypassed entirely. Harmless while
nothing expired; routine now that files do.

Verified: a workflow reaches "@self" with no grant and reads back what it
wrote; another workflow's collection is refused with defaultAccess set to
read-write, while the control write in the same run succeeds - so the
refusal is about the name, not about storage being broken. 66 passed, 0
failed, 2 skipped for a model this machine is not running.
fszontagh vor 1 Monat
Ursprung
Commit
ea2c988418

+ 60 - 0
lib/storage/retention.hpp

@@ -0,0 +1,60 @@
+#pragma once
+
+#include <nlohmann/json.hpp>
+#include <optional>
+#include <string>
+
+namespace smartbotic::storage {
+
+// How long a workflow's data is kept: its own documents, the files it stores,
+// and its execution records, so a run's history and the files that history
+// points at go together rather than one outliving the other.
+//
+// Settings shape, on both a project and a workflow:
+//
+//   "retention": { "ttlSeconds": 604800 }
+//
+// **Absent is not the same as zero.** Absent (or null) on a workflow means
+// inherit the project's; an explicit 0 means keep for ever, which is a choice
+// somebody made and must not be re-read as "unset". A plain
+// value("ttlSeconds", 0) collapses the two, so nothing reads it that way.
+
+inline constexpr const char* kRetentionKey = "retention";
+inline constexpr const char* kTtlSecondsKey = "ttlSeconds";
+
+inline std::optional<int64_t> declaredTtlSeconds(const nlohmann::json& settings) {
+    if (!settings.is_object() || !settings.contains(kRetentionKey)) return std::nullopt;
+    const auto& retention = settings[kRetentionKey];
+    if (!retention.is_object() || !retention.contains(kTtlSecondsKey)) return std::nullopt;
+    const auto& ttl = retention[kTtlSecondsKey];
+    if (!ttl.is_number_integer()) return std::nullopt;  // null, or a string somebody typed
+    const auto seconds = ttl.get<int64_t>();
+    return seconds < 0 ? std::nullopt : std::optional<int64_t>{seconds};
+}
+
+struct Retention {
+    int64_t ttl_seconds = 0;   // 0 = keep for ever
+    bool inherited = false;    // true when it came from the project
+
+    [[nodiscard]] bool keepsForever() const { return ttl_seconds == 0; }
+    [[nodiscard]] int64_t ttlMs() const { return ttl_seconds * 1000; }
+};
+
+// Nothing at all is a third answer, distinct from "keep for ever": an
+// installation from before retention existed has made no choice, and each
+// caller applies its own historical default rather than having one invented
+// here. Executions have always aged out after a week; documents never have.
+// Collapsing those into a single "0" would either start deleting documents
+// nobody agreed to lose or stop executions ageing out at all.
+inline std::optional<Retention> declaredRetention(const nlohmann::json& workflow_settings,
+                                                  const nlohmann::json& project_settings) {
+    if (const auto own = declaredTtlSeconds(workflow_settings)) {
+        return Retention{*own, false};
+    }
+    if (const auto from_project = declaredTtlSeconds(project_settings)) {
+        return Retention{*from_project, true};
+    }
+    return std::nullopt;
+}
+
+}  // namespace smartbotic::storage

+ 20 - 4
lib/storage/storage_client.cpp

@@ -436,7 +436,8 @@ std::vector<ViewInfo> StorageClient::listViews() {
 }
 
 Result<FileUploadInfo> StorageClient::uploadFile(const std::vector<uint8_t>& data,
-                                                const FileMeta& meta) {
+                                                const FileMeta& meta,
+                                                int64_t ttl_ms) {
     if (data.empty()) {
         return Error(ErrorCode::InvalidArgument, "Cannot upload an empty file");
     }
@@ -449,7 +450,12 @@ Result<FileUploadInfo> StorageClient::uploadFile(const std::vector<uint8_t>& dat
     upstream.is_public = meta.is_public;
     upstream.metadata = meta.metadata;
 
-    auto result = impl_->client_->uploadFile(data, upstream);
+    // The TTL is a separate overload upstream rather than a FileUploadMeta
+    // member, deliberately: adding a member would change that struct's size
+    // and an already-installed binary would bind to the new library with the
+    // old layout. Keep calling the two-argument form when there is no TTL.
+    auto result = ttl_ms > 0 ? impl_->client_->uploadFile(data, upstream, msToSec(ttl_ms))
+                             : impl_->client_->uploadFile(data, upstream);
     if (result.id.empty()) {
         return Error(ErrorCode::DatabaseError, "File upload failed: " + meta.name);
     }
@@ -463,9 +469,19 @@ Result<FileUploadInfo> StorageClient::uploadFile(const std::vector<uint8_t>& dat
 }
 
 Result<std::vector<uint8_t>> StorageClient::downloadFile(const std::string& id) {
-    auto bytes = impl_->client_->downloadFile(id);
+    // Upstream throws for a file that is missing or has expired; it does not
+    // return empty bytes. Without this the exception escapes the adapter and
+    // every caller's Result handling is bypassed - and once files carry a TTL,
+    // fetching an expired one is an ordinary event rather than a fault.
+    std::vector<uint8_t> bytes;
+    try {
+        bytes = impl_->client_->downloadFile(id);
+    } catch (const std::exception& e) {
+        return Error(ErrorCode::DocumentNotFound,
+                     "File not found: " + id + " (" + e.what() + ")");
+    }
     if (bytes.empty()) {
-        return Error(ErrorCode::DocumentNotFound, "File not found or empty: " + id);
+        return Error(ErrorCode::DocumentNotFound, "File is empty: " + id);
     }
     return bytes;
 }

+ 8 - 1
lib/storage/storage_client.hpp

@@ -236,8 +236,15 @@ public:
 
     // File store. Kept separate from documents: images and other binaries do not
     // belong inline in a document, and upstream deduplicates by checksum.
+    // ttl_ms > 0 makes the file delete itself. The expiry is absolute from the
+    // moment of upload: it does not restart when the file is read or when the
+    // service restarts. Expiry removes the bytes, not only the metadata record,
+    // but content-addressed blobs are reference counted - two files with
+    // identical content share one blob and one expiring leaves the other whole.
+    // An expired file reads as not-found at once, without waiting for a sweep.
     common::Result<FileUploadInfo> uploadFile(const std::vector<uint8_t>& data,
-                                              const FileMeta& meta);
+                                              const FileMeta& meta,
+                                              int64_t ttl_ms = 0);
 
     common::Result<std::vector<uint8_t>> downloadFile(const std::string& id);
 

+ 48 - 0
lib/storage/workflow_collection.hpp

@@ -0,0 +1,48 @@
+#pragma once
+
+#include <string>
+
+namespace smartbotic::storage {
+
+// A workflow's private collection is named for the workflow itself. Workflow
+// ids already begin with "wf_", so the name needs no second prefix - and that
+// same prefix is what marks a collection as some workflow's private data, which
+// is why no other workflow may be granted anything under it.
+//
+// Shared between the runner, which enforces the grant, and the webserver, which
+// protects these collections from being hand-edited and drops them when their
+// workflow goes. Both must agree on the name, so neither spells it itself.
+
+// What a workflow calls its own collection. Nobody can be expected to paste
+// their own workflow's id into a node, and a workflow that was duplicated would
+// carry the original's id and quietly write into somebody else's storage. This
+// name always means "mine", whoever is running.
+inline constexpr const char* kSelfCollection = "@self";
+
+inline bool isSelfCollection(const std::string& collection) {
+    return collection == kSelfCollection;
+}
+
+inline std::string workflowCollectionName(const std::string& workflow_id) {
+    if (workflow_id.empty()) return {};
+    return workflow_id.rfind("wf_", 0) == 0 ? workflow_id : "wf_" + workflow_id;
+}
+
+// A caller may address a collection as "<project>:<collection>". Which
+// collection it is does not depend on how it was spelled.
+inline std::string collectionBaseName(const std::string& collection) {
+    const auto colon = collection.rfind(':');
+    return colon == std::string::npos ? collection : collection.substr(colon + 1);
+}
+
+inline bool isWorkflowCollection(const std::string& collection) {
+    return collectionBaseName(collection).rfind("wf_", 0) == 0;
+}
+
+inline bool isWorkflowCollectionOf(const std::string& collection,
+                                   const std::string& workflow_id) {
+    const std::string owned = workflowCollectionName(workflow_id);
+    return !owned.empty() && collectionBaseName(collection) == owned;
+}
+
+}  // namespace smartbotic::storage

+ 193 - 25
src/runner/workflow_engine.cpp

@@ -2,6 +2,8 @@
 #include "common/uuid.hpp"
 #include "common/time_utils.hpp"
 #include "common/config_defaults.hpp"
+#include "storage/retention.hpp"
+#include "storage/workflow_collection.hpp"
 #include "logging/logger.hpp"
 #include <algorithm>
 #include <functional>
@@ -56,10 +58,68 @@ std::string executionStatusToString(ExecutionStatus status) {
     }
 }
 
+std::optional<storage::Retention> WorkflowEngine::retentionFor(const Workflow& workflow) {
+    // A workflow that says nothing about retention inherits its project's. With
+    // no project there is nothing to inherit from - which is the case for
+    // anything created before projects existed.
+    if (const auto own = storage::declaredTtlSeconds(workflow.settings)) {
+        return storage::Retention{*own, false};
+    }
+    if (workflow.project_id.empty()) {
+        return std::nullopt;
+    }
+
+    nlohmann::json project_settings = nlohmann::json::object();
+    const int64_t now = TimeUtils::nowMs();
+    {
+        std::lock_guard<std::mutex> lock(project_settings_mutex_);
+        auto it = project_settings_cache_.find(workflow.project_id);
+        if (it != project_settings_cache_.end() &&
+            now - it->second.fetched_at_ms < kProjectSettingsCacheMs) {
+            return storage::declaredRetention(workflow.settings, it->second.settings);
+        }
+    }
+
+    auto project = storage_.get("projects", workflow.project_id);
+    if (project.ok()) {
+        project_settings = project.value().value("settings", nlohmann::json::object());
+    } else {
+        // A project that cannot be read must not quietly become "keep for
+        // ever": say so, and cache the empty answer only briefly so a
+        // transient failure does not pin the wrong retention.
+        LOG_WARN("Retention: cannot read project {} for workflow {} ({}); this run keeps its "
+                 "data on the default lifetime",
+                 workflow.project_id, workflow.id, project.error().message());
+    }
+    {
+        std::lock_guard<std::mutex> lock(project_settings_mutex_);
+        project_settings_cache_[workflow.project_id] = {project_settings, now};
+    }
+    return storage::declaredRetention(workflow.settings, project_settings);
+}
+
+void WorkflowEngine::ensureWorkflowCollection(const std::string& collection) {
+    {
+        std::lock_guard<std::mutex> lock(created_collections_mutex_);
+        if (created_workflow_collections_.contains(collection)) return;
+    }
+    // Not conditional on a listCollections check: creating one that already
+    // exists is harmless, and asking first costs a round trip on every run.
+    auto created = storage_.createCollection(collection);
+    if (created.failed()) {
+        // Not fatal - the insert that follows reports the real failure if the
+        // collection genuinely is not there.
+        LOG_DEBUG("Storage: could not create {} ({})", collection, created.error().message());
+    }
+    std::lock_guard<std::mutex> lock(created_collections_mutex_);
+    created_workflow_collections_.insert(collection);
+}
+
 Workflow Workflow::fromJson(const nlohmann::json& j) {
     Workflow wf;
     wf.id = j.value("_id", j.value("id", ""));
     wf.name = j.value("name", "");
+    wf.project_id = j.value("projectId", j.value("project_id", ""));
     wf.active = j.value("active", false);
     wf.settings = j.value("settings", nlohmann::json::object());
 
@@ -342,6 +402,7 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
     nlohmann::json snapshot;
     snapshot["id"] = workflow.id;
     snapshot["name"] = workflow.name;
+    snapshot["projectId"] = workflow.project_id;
     snapshot["settings"] = workflow.settings;
     snapshot["nodes"] = nlohmann::json::array();
     for (const auto& node : workflow.nodes) {
@@ -365,6 +426,14 @@ Result<ExecutionResult> WorkflowEngine::execute(const Workflow& workflow,
     }
     result.workflow_snapshot = snapshot;
 
+    // Resolved once, at the start, and carried on the result: the record is
+    // written several times over a run's life - before the walk, on pause, at
+    // the end - and every one of those writes has to agree about how long the
+    // record is kept, or the last one silently changes it.
+    if (const auto retention = retentionFor(workflow)) {
+        result.retention_ttl_ms = retention->ttlMs();
+    }
+
     // Write the record before the walk starts, not only when it ends.
     //
     // Until this, an execution existed in the database only once it had
@@ -1301,6 +1370,13 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         return result;
     }
 
+    // Nothing declared leaves documents on the lifetime they have always had -
+    // for ever - so an installation that has not chosen a retention keeps
+    // behaving exactly as it did.
+    const auto declared_retention = retentionFor(workflow);
+    const int64_t retention_ttl_ms = declared_retention ? declared_retention->ttlMs() : 0;
+    const std::string own_collection = storage::workflowCollectionName(workflow.id);
+
     // Helper to get per-workflow storage permission for a collection
     // Returns: "none", "read-only", or "read-write"
     auto getWorkflowAccess = [&workflow, this](const std::string& collection) -> std::string {
@@ -1309,6 +1385,27 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
             return "none";
         }
 
+        // A private collection is settled before the settings are consulted at
+        // all. Another workflow's is refused whatever the settings say - which
+        // matters most for defaultAccess, since a workflow set to read-write by
+        // default would otherwise reach every other workflow's private data,
+        // across projects, just by naming it.
+        if (storage::isWorkflowCollection(collection)) {
+            if (!storage::isWorkflowCollectionOf(collection, workflow.id)) {
+                return "none";
+            }
+            // Its owner holds it read-write without asking. An entry in the
+            // settings still wins, so the grant can be given up deliberately.
+            if (workflow.settings.contains("storagePermissions")) {
+                const auto& perms = workflow.settings["storagePermissions"];
+                if (perms.contains("collections") && perms["collections"].is_object() &&
+                    perms["collections"].contains(collection)) {
+                    return perms["collections"][collection].get<std::string>();
+                }
+            }
+            return "read-write";
+        }
+
         // Check workflow settings for storage permissions
         if (workflow.settings.contains("storagePermissions")) {
             const auto& storage_perms = workflow.settings["storagePermissions"];
@@ -1359,51 +1456,94 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         }
     };
 
+    // "@self" means this workflow's own collection, whoever is running. Every
+    // storage callback resolves it the same way and before the permission
+    // check, so the name cannot be used to reach anything else.
+    auto resolveCollection = [own_collection](const std::string& collection,
+                                              std::string& resolved) -> common::Result<void> {
+        if (!storage::isSelfCollection(collection)) {
+            resolved = collection;
+            return {};
+        }
+        if (own_collection.empty()) {
+            return common::Error(common::ErrorCode::InvalidArgument,
+                "\"@self\" means this workflow's own storage, and this run has no workflow id");
+        }
+        resolved = own_collection;
+        return {};
+    };
+
     // Storage API callbacks with per-workflow permission checks
-    ctx.storage_get_doc = [this, canReadCollection](const std::string& collection, const std::string& id)
+    ctx.storage_get_doc = [this, canReadCollection, resolveCollection](
+                                const std::string& collection, const std::string& id)
         -> common::Result<nlohmann::json> {
-        if (!canReadCollection(collection)) {
+        std::string target;
+        if (auto r = resolveCollection(collection, target); r.failed()) return r.error();
+        if (!canReadCollection(target)) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No read access to collection: " + collection);
+                "No read access to collection: " + target);
         }
-        return storage_.get(collection, id);
+        return storage_.get(target, id);
     };
 
-    ctx.storage_insert = [this, canWriteCollection](const std::string& collection, const nlohmann::json& data,
+    ctx.storage_insert = [this, canWriteCollection, resolveCollection, retention_ttl_ms, own_collection](
+                                const std::string& collection, const nlohmann::json& data,
                                 const std::string& id, int64_t ttl_ms)
         -> common::Result<std::string> {
-        if (!canWriteCollection(collection)) {
+        std::string target;
+        if (auto r = resolveCollection(collection, target); r.failed()) return r.error();
+        if (!canWriteCollection(target)) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No write access to collection: " + collection);
-        }
-        return storage_.insert(collection, data, id, ttl_ms);
+                "No write access to collection: " + target);
+        }
+        // A workflow's own collection is made on first write rather than when
+        // the workflow is saved: a workflow that never stores anything should
+        // not leave an empty collection behind, and one written to by a runner
+        // that has never seen it before must still work.
+        if (storage::collectionBaseName(target) == own_collection) {
+            ensureWorkflowCollection(target);
+        }
+        // A TTL the node asked for wins - it knows what it wrote. Otherwise the
+        // workflow's retention applies, so data written by a node that never
+        // heard of retention still ages out.
+        const int64_t effective_ttl = ttl_ms > 0 ? ttl_ms : retention_ttl_ms;
+        return storage_.insert(target, data, id, effective_ttl);
     };
 
-    ctx.storage_update = [this, canWriteCollection](const std::string& collection, const std::string& id,
+    ctx.storage_update = [this, canWriteCollection, resolveCollection](
+                                const std::string& collection, const std::string& id,
                                 const nlohmann::json& data, int64_t expected_version, bool partial)
         -> common::Result<int64_t> {
-        if (!canWriteCollection(collection)) {
+        std::string target;
+        if (auto r = resolveCollection(collection, target); r.failed()) return r.error();
+        if (!canWriteCollection(target)) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No write access to collection: " + collection);
+                "No write access to collection: " + target);
         }
-        return storage_.update(collection, id, data, expected_version, partial);
+        return storage_.update(target, id, data, expected_version, partial);
     };
 
-    ctx.storage_delete = [this, canWriteCollection](const std::string& collection, const std::string& id,
+    ctx.storage_delete = [this, canWriteCollection, resolveCollection](
+                                const std::string& collection, const std::string& id,
                                 int64_t expected_version)
         -> common::Result<void> {
-        if (!canWriteCollection(collection)) {
+        std::string target;
+        if (auto r = resolveCollection(collection, target); r.failed()) return r.error();
+        if (!canWriteCollection(target)) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No write access to collection: " + collection);
+                "No write access to collection: " + target);
         }
-        return storage_.remove(collection, id, expected_version);
+        return storage_.remove(target, id, expected_version);
     };
 
-    ctx.storage_query = [this, canReadCollection](const std::string& collection, const engine::StorageQueryOptions& options)
+    ctx.storage_query = [this, canReadCollection, resolveCollection](
+                                const std::string& collection, const engine::StorageQueryOptions& options)
         -> common::Result<engine::StorageQueryResult> {
-        if (!canReadCollection(collection)) {
+        std::string target;
+        if (auto r = resolveCollection(collection, target); r.failed()) return r.error();
+        if (!canReadCollection(target)) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No read access to collection: " + collection);
+                "No read access to collection: " + target);
         }
 
         // Convert engine query options to storage query options
@@ -1414,7 +1554,7 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         storage_opts.page = options.page;
         storage_opts.page_size = options.page_size;
 
-        auto result = storage_.query(collection, storage_opts);
+        auto result = storage_.query(target, storage_opts);
         if (result.failed()) {
             return common::Error(result.error().code(), result.error().message());
         }
@@ -1426,16 +1566,30 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         return engine_result;
     };
 
-    ctx.storage_list_collections = [this, canReadCollection](bool include_system)
+    ctx.storage_list_collections = [this, canReadCollection, own_collection](bool include_system)
         -> common::Result<std::vector<std::string>> {
         auto all_collections = storage_.listCollections();
         std::vector<std::string> accessible;
 
+        // A workflow's own storage is offered before it exists - it is created
+        // on first write, and a collection that only appears once you have
+        // already written to it cannot be picked from a list.
+        if (!own_collection.empty()) {
+            accessible.push_back(storage::kSelfCollection);
+        }
+
         for (const auto& coll : all_collections) {
             // Skip system collections unless explicitly requested
             if (!include_system && collection_permissions_->isSystemCollection(coll)) {
                 continue;
             }
+            // Private workflow storage is never offered under its raw name:
+            // another workflow's is not reachable anyway, and this workflow's
+            // own is already in the list as "@self". Offering both would put
+            // the same collection in the list twice under two names.
+            if (storage::isWorkflowCollection(coll)) {
+                continue;
+            }
             // Only include collections the workflow has at least read access to
             if (canReadCollection(coll)) {
                 accessible.push_back(coll);
@@ -1448,7 +1602,7 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
     // Files sit outside the collection namespace, so they are gated on a reserved
     // "files" permission entry. Deny-by-default, consistent with collections: a
     // workflow must declare settings.storagePermissions.collections.files.
-    ctx.storage_upload_file = [this, canWriteCollection](const std::vector<uint8_t>& data,
+    ctx.storage_upload_file = [this, canWriteCollection, retention_ttl_ms](const std::vector<uint8_t>& data,
                                     const engine::ScriptFileMeta& meta)
         -> common::Result<engine::ScriptFileInfo> {
         if (!canWriteCollection("files")) {
@@ -1464,7 +1618,10 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         upstream.is_public = meta.is_public;
         upstream.metadata = meta.metadata;
 
-        auto result = storage_.uploadFile(data, upstream);
+        // The file expires with the document that points at it. A stored file
+        // whose record has aged out is unreachable weight, and a record whose
+        // file has gone is a broken link - they have to share one lifetime.
+        auto result = storage_.uploadFile(data, upstream, retention_ttl_ms);
         if (result.failed()) {
             return common::Error(result.error().code(), result.error().message());
         }
@@ -2165,8 +2322,19 @@ void WorkflowEngine::storeExecution(const ExecutionResult& result) {
     // deadline", not "no need to extend" - it is given the 30-day ceiling the
     // approval node caps at, so a permanent approval does not fall back to
     // the ordinary seven-day log TTL and evaporate with no trace.
+    //
+    // A workflow that sets its own retention overrides all of that: its history
+    // and the data it wrote age out together, which is the point of the setting.
+    // An explicit "keep for ever" has to survive the waiting floor below, so it
+    // is checked rather than compared - a 0 fed into a `wanted > ttl_ms` test
+    // reads as the shortest possible lifetime instead of the longest.
     int64_t ttl_ms = 7 * 24 * 60 * 60 * 1000;
-    if (result.status == ExecutionStatus::Waiting) {
+    if (result.retention_ttl_ms >= 0) {
+        ttl_ms = result.retention_ttl_ms;
+    }
+    const bool keep_forever = (ttl_ms == 0);
+
+    if (result.status == ExecutionStatus::Waiting && !keep_forever) {
         const int64_t grace = 24 * 60 * 60 * 1000;
         const int64_t never_ttl = 30LL * 24 * 60 * 60 * 1000;
         int64_t wanted = never_ttl;

+ 36 - 0
src/runner/workflow_engine.hpp

@@ -2,6 +2,7 @@
 
 #include <string>
 #include <unordered_map>
+#include <unordered_set>
 #include <vector>
 #include <queue>
 #include <mutex>
@@ -15,6 +16,7 @@
 #include "collection_permissions.hpp"
 #include "engine/script_engine.hpp"
 #include "storage/storage_client.hpp"
+#include "storage/retention.hpp"
 #include "common/error.hpp"
 
 namespace smartbotic::runner {
@@ -55,6 +57,10 @@ struct Connection {
 struct Workflow {
     std::string id;
     std::string name;
+    // The project this workflow belongs to. Carried so a run can resolve the
+    // retention it inherits; without it every workflow would look like one with
+    // no project and fall back to keeping its data for ever.
+    std::string project_id;
     bool active = false;
     std::vector<WorkflowNode> nodes;
     std::vector<Connection> connections;
@@ -103,6 +109,12 @@ struct ExecutionResult {
     std::string error;
     nlohmann::json final_output;
     nlohmann::json workflow_snapshot;  // Snapshot of workflow at execution time
+    // How long this record is kept, from the workflow's retention. -1 means the
+    // workflow declared nothing, so the ordinary log lifetime applies; 0 means
+    // somebody chose to keep it for ever. The two are not the same, and reading
+    // an undeclared retention as "for ever" would quietly stop executions ever
+    // ageing out.
+    int64_t retention_ttl_ms = -1;
     nlohmann::json webhook_response;   // Set by a respond-to-webhook node, if any
     // What a Workflow Output node said this workflow returns. On the result
     // rather than local to the walk, because there are two walks - the main one
@@ -244,6 +256,30 @@ public:
     void setPostgresqlQueryCallback(PostgresqlQueryCallback callback) { postgresql_query_callback_ = callback; }
 
 private:
+    // How long this workflow's data is kept, resolving what it inherits from
+    // its project. The project's settings are cached briefly: a run asks once
+    // per node, and a workflow of twenty nodes should not fetch its project
+    // twenty times. The window is short enough that a retention change reaches
+    // running workflows within seconds rather than at the next restart.
+    // Empty when neither the workflow nor its project has said anything, which
+    // is not the same as "keep for ever" - see storage/retention.hpp.
+    std::optional<storage::Retention> retentionFor(const Workflow& workflow);
+
+    // Create a workflow's private collection if it is not there yet. Cheap to
+    // call repeatedly: the names already created are remembered, and creating
+    // one that exists is not an error anyway.
+    void ensureWorkflowCollection(const std::string& collection);
+    std::unordered_set<std::string> created_workflow_collections_;
+    std::mutex created_collections_mutex_;
+
+    struct CachedProjectSettings {
+        nlohmann::json settings;
+        int64_t fetched_at_ms = 0;
+    };
+    static constexpr int64_t kProjectSettingsCacheMs = 15000;
+    std::unordered_map<std::string, CachedProjectSettings> project_settings_cache_;
+    std::mutex project_settings_mutex_;
+
     void buildExecutionGraph(const Workflow& workflow,
                             std::unordered_map<std::string, std::vector<std::string>>& dependencies,
                             std::unordered_map<std::string, std::vector<std::string>>& dependents);

+ 14 - 0
src/webserver/api/database_controller.cpp

@@ -1,5 +1,6 @@
 #include "database_controller.hpp"
 #include "logging/logger.hpp"
+#include "storage/workflow_collection.hpp"
 #include <unordered_set>
 
 namespace smartbotic::webserver::api {
@@ -117,6 +118,13 @@ DatabaseController::Protection DatabaseController::protectionFor(const std::stri
 
     if (SECRET.contains(name)) return Protection::Secret;
     if (STRUCTURAL.contains(name)) return Protection::Structural;
+
+    // A workflow's own collection. Readable here - an administrator looking
+    // after the data should be able to see what a workflow stored - but not
+    // writable, because the workflow is the only thing that should put data in
+    // it and dropping it is the workflow's own deletion, not a database edit.
+    if (storage::isWorkflowCollection(name)) return Protection::Structural;
+
     return Protection::None;
 }
 
@@ -130,6 +138,12 @@ bool DatabaseController::refuseIfProtected(httplib::Response& res,
             return true;
         case Protection::Structural:
             if (!writing) return false;
+            if (storage::isWorkflowCollection(collection)) {
+                sendError(res, "\"" + collection + "\" is a workflow's own storage. It can be read "
+                               "here, but only that workflow writes to it, and it goes when the "
+                               "workflow does", 403);
+                return true;
+            }
             sendError(res, "\"" + collection + "\" belongs to SmartBotic itself. It can be read "
                            "here, but changing it has to go through the endpoints that own it", 403);
             return true;

+ 64 - 0
tests/nodes/workflow-other-collection-refused.json

@@ -0,0 +1,64 @@
+{
+  "name": "verify-workflow-other-collection-refused",
+  "comment": "Another workflow's private storage is refused even when this workflow's defaultAccess is read-write. defaultAccess is the dangerous case, not an explicit grant: it is a blanket yes to every name the workflow has not listed, so without a check ahead of it a single careless setting reaches every other workflow's data, across projects.\n\nThe control node is the point of the test. It writes, under the same defaultAccess, to an ordinary collection the workflow was never granted by name, and must succeed. That is what makes the refusal above mean something: writes are working and defaultAccess is being honoured, so the wf_ name was refused for being a wf_ name and not because this workflow could not write anywhere. Without the control the case would pass just as happily if storage were broken outright.",
+  "settings": {
+    "storagePermissions": {
+      "defaultAccess": "read-write"
+    }
+  },
+  "nodes": [
+    {
+      "id": "n1",
+      "name": "Trigger",
+      "type": "click-trigger",
+      "position": { "x": 0, "y": 0 },
+      "config": {}
+    },
+    {
+      "id": "other",
+      "name": "Write to another workflow's storage",
+      "type": "storage-insert",
+      "position": { "x": 0, "y": 100 },
+      "config": {
+        "collectionSource": "manual",
+        "collectionManual": "wf_00000000-0000-0000-0000-000000000000",
+        "documentData": "{\"probe\": \"trespass\"}"
+      }
+    },
+    {
+      "id": "control",
+      "name": "Write to an ordinary collection it was never granted by name",
+      "type": "storage-insert",
+      "position": { "x": 0, "y": 200 },
+      "config": {
+        "collectionSource": "manual",
+        "collectionManual": "wfprobe_absent_collection",
+        "documentData": "{\"probe\": \"control\"}"
+      }
+    }
+  ],
+  "connections": [
+    {
+      "sourceNodeId": "n1",
+      "sourceOutput": "main",
+      "targetNodeId": "other",
+      "targetInput": "data"
+    },
+    {
+      "sourceNodeId": "other",
+      "sourceOutput": "main",
+      "targetNodeId": "control",
+      "targetInput": "data"
+    }
+  ],
+  "expect": {
+    "other": {
+      "status": "completed",
+      "output": { "success": false, "error": "No write access to collection: wf_00000000-0000-0000-0000-000000000000" }
+    },
+    "control": {
+      "status": "completed",
+      "output": { "success": true }
+    }
+  }
+}

+ 60 - 0
tests/nodes/workflow-own-collection-granted.json

@@ -0,0 +1,60 @@
+{
+  "name": "verify-workflow-own-collection-granted",
+  "comment": "A workflow reaches its own private storage as \"@self\" without being granted anything, and the collection is created on first write rather than having to exist first. Reading it back matters as much as writing: a grant that only covers writes would store data the workflow could never use.",
+  "nodes": [
+    {
+      "id": "n1",
+      "name": "Trigger",
+      "type": "click-trigger",
+      "position": { "x": 0, "y": 0 },
+      "config": {}
+    },
+    {
+      "id": "write",
+      "name": "Write to own storage",
+      "type": "storage-insert",
+      "position": { "x": 0, "y": 100 },
+      "config": {
+        "collectionSource": "manual",
+        "collectionManual": "@self",
+        "documentId": "probe",
+        "documentData": "{\"probe\": \"own-collection\"}"
+      }
+    },
+    {
+      "id": "read",
+      "name": "Read it back",
+      "type": "storage-get",
+      "position": { "x": 0, "y": 200 },
+      "config": {
+        "collectionSource": "manual",
+        "collectionManual": "@self",
+        "documentId": "probe"
+      }
+    }
+  ],
+  "connections": [
+    {
+      "sourceNodeId": "n1",
+      "sourceOutput": "main",
+      "targetNodeId": "write",
+      "targetInput": "data"
+    },
+    {
+      "sourceNodeId": "write",
+      "sourceOutput": "main",
+      "targetNodeId": "read",
+      "targetInput": "data"
+    }
+  ],
+  "expect": {
+    "write": {
+      "status": "completed",
+      "output": { "success": true }
+    },
+    "read": {
+      "status": "completed",
+      "output": { "found": true, "document": { "probe": "own-collection" } }
+    }
+  }
+}