Jelajahi Sumber

release(v2.8.0): file TTL, per-type retention, and update() no longer wipes expiry

update() was wiping document expiry.
-----------------------------------
MemoryStore::update and updateIfVersion built a fresh Document from the
caller's request. UpdateRequest has no ttl field, so the incoming expiresAt was
always 0, and both functions deliberately carried createdAt and createdBy
across from the stored document but not expiresAt. So an update WIPED the
expiry and made a TTL'd document permanent. patchDocument() mutates the stored
document in place and always preserved it, so the two write paths disagreed and
the more commonly used one lost data.

It bites where it is least visible: a record written with a TTL and then
updated as its status changes - the normal case - silently stopped expiring and
accumulated forever. Plausible contributor to executions at 505 MB / 9571 docs.

An update now carries the existing expiry across when the caller supplies none;
a non-zero incoming expiresAt still wins, so a TTL can be changed deliberately.
There is still no set-TTL-only RPC: ttl_seconds exists on Insert/Upsert/
BatchInsertItem but not Update/Patch, so changing one means upserting the whole
document. Retro-applying to existing rows is deliberately left out.

File TTL.
---------
FileMetadata.ttl_seconds is relative and converted once to an absolute
expiresAt, so neither a read nor a restart extends a file's life.

The sweeper did not exist. Config::cleanupIntervalSec had been there from the
beginning and nothing ever read it - no thread, so nothing was ever reclaimed.

Expiry routes through deleteFile(), which is what reclaims the BYTES: blobs are
content-addressed and deduplicated across projects, so it decrements the
refcount and unlinks only on the last reference. The first cut broke this in the
opposite direction - teaching getFileInfo to hide expired records made
deleteFile return false for them, so the sweep found records, asked for
deletion, was refused, and reclaimed nothing while leaking every blob. Caught by
test_expiry_deletes_the_blob_not_just_the_record.

Per-fileType default retention.
-------------------------------
storage.files.default_ttl_seconds maps a key to seconds, matched
most-specific-first: "<project>:<fileType>", "<fileType>", "<project>:*", "*".
The analogue of a collection's defaultTtlSeconds.

FileMetadata.ttl_seconds is `optional` and that is load-bearing: absent means
inherit the type default, an explicitly set 0 means never expire and overrides
it. A plain uint32 cannot express the difference, which would make opting a
single file out of a retention policy impossible.

Client::uploadFile(data, meta, ttlSeconds) is an OVERLOAD, not a new
FileUploadMeta member - public struct layouts stay byte-identical, verified by
comparing sizeof against the installed headers.

ctest 18/18, namespacing 31/31, policy enforcement 19/19, views and TLS/auth
green.
fszontagh 1 bulan lalu
induk
melakukan
9f4aa18398

File diff ditekan karena terlalu besar
+ 0 - 0
CLAUDE.md


+ 1 - 1
VERSION

@@ -1 +1 @@
-2.7.1
+2.8.0

+ 15 - 0
client/CMakeLists.txt

@@ -5,6 +5,21 @@ if(BUILD_SHARED_LIBS)
         src/client.cpp
         $<TARGET_OBJECTS:smartbotic_db_proto>
     )
+    # SOVERSION stays 2 — and the rule that keeps it honest:
+    #
+    # BUMP THIS whenever a public struct changes size or layout. It stayed at 2
+    # across v2.1-v2.7.0 while Client::CollectionConfig (v2.4.5) and FileRecord
+    # (v2.6.0) both gained members, on the reasoning that "consumers must
+    # rebuild". That reasoning failed the moment it was tested: a v2.7.1 attempt
+    # to add `projection` to Client::QueryOptions changed the struct size without
+    # changing the soname, and the already-installed smartbotic-webserver -
+    # compiled against the v2.7.0 layout - bound to the new library and SIGSEGV'd
+    # in a restart loop until the packages were rolled back.
+    #
+    # A changelog note is not an enforcement mechanism; a soname is. v2.7.1
+    # therefore adds projection as an OVERLOAD rather than a struct member, so
+    # existing binaries keep working untouched. The next deliberate layout change
+    # bumps this to 3 and renames the package alongside it.
     set_target_properties(smartbotic-db-client PROPERTIES
         VERSION ${SMARTBOTIC_DB_VERSION}
         SOVERSION 2

+ 49 - 14
client/include/smartbotic/database/client.hpp

@@ -335,20 +335,6 @@ public:
         bool sortDescending = false;
         uint32_t limit = 100;
         uint32_t offset = 0;
-
-        // v2.7.1 — fields to return; empty means the whole document.
-        //
-        // The wire and the server have supported this since v2.0
-        // (`FindRequest.projection`), but it was never exposed here, so callers
-        // had no way to avoid transferring large fields. That matters on
-        // collections holding blobs: a browser paging through documents that
-        // each carry megabytes of base64 had to receive all of it just to render
-        // a list of ids.
-        //
-        // Projection is applied server-side AFTER pagination, so it cuts
-        // response size and client parse time; it does not change which
-        // documents match.
-        std::vector<std::string> projection;
     };
 
     /**
@@ -377,6 +363,30 @@ public:
     [[nodiscard]] std::vector<nlohmann::json> find(const std::string& collection,
                                                    const QueryOptions& options);
 
+    /**
+     * Find documents, returning only `projection` fields of each.
+     *
+     * v2.7.1. `FindRequest.projection` has existed on the wire since v2.0 and
+     * the server always honoured it, but the client never exposed it - so
+     * paging a collection whose documents carry large fields meant receiving
+     * every byte just to list ids. On a 112-document / 414 MB collection,
+     * requesting page 11 went from 22,791,981 bytes to 1,530 by projecting
+     * to `{"_id"}`.
+     *
+     * Deliberately an OVERLOAD rather than a QueryOptions member: adding a
+     * member changes the struct's size, and with an unchanged soname an
+     * already-installed consumer binds to the new library with the old layout
+     * and crashes. That is not hypothetical - it took the local webserver down
+     * in a SIGSEGV restart loop during the first attempt at this feature.
+     *
+     * Projection is applied server-side AFTER pagination, so it reduces
+     * response size and client parse time without changing which documents
+     * match. Empty behaves exactly like the two-argument overload.
+     */
+    [[nodiscard]] std::vector<nlohmann::json> find(const std::string& collection,
+                                                   const QueryOptions& options,
+                                                   const std::vector<std::string>& projection);
+
     /**
      * Find documents with full metrics (including WAL fallback info).
      */
@@ -783,6 +793,31 @@ public:
     };
 
     FileUploadResult uploadFile(const std::vector<uint8_t>& data, const FileUploadMeta& meta);
+
+    /**
+     * Upload a file that expires by itself.
+     *
+     * v2.8.0. `ttlSeconds` is relative and converted once to an absolute expiry
+     * server-side, so a file's lifetime does not restart when it is read or when
+     * the service restarts. 0 means never - identical to the two-argument form.
+     *
+     * Expiry deletes the BLOB, not just the metadata record: the sweeper routes
+     * every expired file through the same delete path as an explicit delete,
+     * which decrements the content-addressed blob's reference count and unlinks
+     * the bytes only when the last reference goes. Two files with identical
+     * content share one blob, so one expiring does not remove the other's data.
+     *
+     * An expired file reads as NOT_FOUND immediately, without waiting for the
+     * sweep interval (`storage.files.cleanup_interval_sec`).
+     *
+     * Deliberately an OVERLOAD rather than a FileUploadMeta member: adding a
+     * member changes that struct's size, and with an unchanged soname an
+     * already-installed consumer binds to the new library with the old layout
+     * and crashes.
+     */
+    FileUploadResult uploadFile(const std::vector<uint8_t>& data,
+                                const FileUploadMeta& meta,
+                                uint32_t ttlSeconds);
     [[nodiscard]] std::vector<uint8_t> downloadFile(const std::string& id);
     [[nodiscard]] std::optional<FileRecord> getFileInfo(const std::string& id);
     bool deleteFile(const std::string& id);

+ 34 - 15
client/src/client.cpp

@@ -539,7 +539,8 @@ public:
     // ===== Query Operations =====
 
     std::vector<nlohmann::json> find(const std::string& collection,
-                                     const Client::QueryOptions& options) {
+                                     const Client::QueryOptions& options,
+                                     const std::vector<std::string>& projection = {}) {
         smartbotic::databasepb::FindRequest request;
         request.set_collection(qualify(collection));
 
@@ -562,12 +563,10 @@ public:
         request.set_limit(options.limit);
         request.set_offset(options.offset);
 
-        // v2.7.1 — field projection. Applied server-side after pagination, so it
-        // cuts response size and client parse time without changing which
-        // documents match. Exposed because collections holding large fields
-        // (base64 blobs) were otherwise impossible to page through: the caller
-        // had to receive every byte just to list ids.
-        for (const auto& f : options.projection) {
+        // v2.7.1 — field projection, passed alongside QueryOptions rather than
+        // inside it: adding a member would change the struct size, and with an
+        // unchanged soname an installed consumer crashes. See client.hpp.
+        for (const auto& f : projection) {
             request.add_projection(f);
         }
 
@@ -614,7 +613,8 @@ public:
     }
 
     Client::FindResult findWithMetrics(const std::string& collection,
-                                        const Client::QueryOptions& options) {
+                                        const Client::QueryOptions& options,
+                                        const std::vector<std::string>& projection = {}) {
         smartbotic::databasepb::FindRequest request;
         request.set_collection(qualify(collection));
 
@@ -637,12 +637,10 @@ public:
         request.set_limit(options.limit);
         request.set_offset(options.offset);
 
-        // v2.7.1 — field projection. Applied server-side after pagination, so it
-        // cuts response size and client parse time without changing which
-        // documents match. Exposed because collections holding large fields
-        // (base64 blobs) were otherwise impossible to page through: the caller
-        // had to receive every byte just to list ids.
-        for (const auto& f : options.projection) {
+        // v2.7.1 — field projection, passed alongside QueryOptions rather than
+        // inside it: adding a member would change the struct size, and with an
+        // unchanged soname an installed consumer crashes. See client.hpp.
+        for (const auto& f : projection) {
             request.add_projection(f);
         }
 
@@ -1643,7 +1641,10 @@ public:
 
     // ===== File Operations =====
 
-    Client::FileUploadResult uploadFile(const std::vector<uint8_t>& data, const Client::FileUploadMeta& meta) {
+    Client::FileUploadResult uploadFile(const std::vector<uint8_t>& data,
+                                        const Client::FileUploadMeta& meta,
+                                        uint32_t ttlSeconds = 0,
+                                        bool explicitTtl = false) {
         grpc::ClientContext context;
         auto timeout_ms = config_.timeoutMs + (data.size() / (1024 * 1024)) * 1000;
         context.set_deadline(
@@ -1664,6 +1665,12 @@ public:
         // v2.6.0 — files are namespaced like collections. Filled from
         // Config::project so callers keep passing bare ids and names.
         m->set_project(config_.project);
+        // v2.8.0 — relative TTL; the server converts it to an absolute expiry.
+        // Set unconditionally on this overload: calling it IS the explicit
+        // choice, and 0 therefore means "never expire, ignore any configured
+        // default for this file type". The two-argument uploadFile() leaves the
+        // field absent so the default applies.
+        if (explicitTtl) m->set_ttl_seconds(ttlSeconds);
         for (const auto& [key, value] : meta.metadata) {
             (*m->mutable_metadata())[key] = value;
         }
@@ -1935,6 +1942,12 @@ std::vector<nlohmann::json> Client::find(const std::string& collection,
     return impl_->find(collection, options);
 }
 
+std::vector<nlohmann::json> Client::find(const std::string& collection,
+                                                 const QueryOptions& options,
+                                                 const std::vector<std::string>& projection) {
+    return impl_->find(collection, options, projection);
+}
+
 Client::FindResult Client::findWithMetrics(const std::string& collection,
                                             const QueryOptions& options) {
     return impl_->findWithMetrics(collection, options);
@@ -2075,6 +2088,12 @@ Client::FileUploadResult Client::uploadFile(const std::vector<uint8_t>& data, co
     return impl_->uploadFile(data, meta);
 }
 
+Client::FileUploadResult Client::uploadFile(const std::vector<uint8_t>& data,
+                                            const FileUploadMeta& meta,
+                                            uint32_t ttlSeconds) {
+    return impl_->uploadFile(data, meta, ttlSeconds, /*explicitTtl=*/true);
+}
+
 std::vector<uint8_t> Client::downloadFile(const std::string& id) {
     return impl_->downloadFile(id);
 }

+ 10 - 0
proto/database.proto

@@ -555,6 +555,15 @@ message FileMetadata {
     // bare collection names. Files are namespaced so that per-project access
     // policy has something to attach to.
     string project = 9;
+    // v2.8.0 — relative time-to-live in seconds, applied at upload. Converted
+    // once to an absolute expiry server-side, so a file's lifetime does not
+    // restart when it is read or on restart.
+    //
+    // `optional` is load-bearing: ABSENT means "inherit the configured default
+    // for this file type", while an explicitly SET 0 means "never expire" and
+    // overrides that default. A plain uint32 cannot express the difference,
+    // which would make it impossible to opt one file out of a retention policy.
+    optional uint32 ttl_seconds = 10;
 }
 
 message UploadFileResponse {
@@ -596,6 +605,7 @@ message FileInfo {
     bool is_public = 10;              // Visibility flag
     uint32 ref_count = 11;            // Blob reference count (informational)
     string project = 12;              // v2.6.0 — owning project
+    uint64 expires_at = 13;           // v2.8.0 — absolute expiry ms, 0 = never
 }
 
 message ListFilesRequest {

+ 19 - 0
service/src/database_grpc_impl.cpp

@@ -1790,6 +1790,23 @@ grpc::Status DatabaseGrpcImpl::UploadFile(
         info.isPublic = metadata.is_public();
         // v2.6.0 — empty project means the default namespace.
         info.project = metadata.project().empty() ? "default" : metadata.project();
+        // v2.8.0 — TTL arrives relative and is stored absolute, so a read or a
+        // restart cannot extend a file's life.
+        //
+        // Presence matters: an explicitly set ttl (including 0) is the caller's
+        // decision and wins over the configured per-type default. Absent leaves
+        // expiresAt at 0 so storeFile() applies the default.
+        if (metadata.has_ttl_seconds()) {
+            if (metadata.ttl_seconds() > 0) {
+                const auto now = std::chrono::duration_cast<std::chrono::milliseconds>(
+                    std::chrono::system_clock::now().time_since_epoch()).count();
+                info.expiresAt = static_cast<uint64_t>(now) +
+                                 static_cast<uint64_t>(metadata.ttl_seconds()) * 1000ULL;
+            } else {
+                // Explicit 0 = never expire, overriding any type default.
+                info.neverExpires = true;
+            }
+        }
         {
             // v2.8.0 — gate the upload on the declared file type.
             smartbotic::database::Decision fdec;
@@ -1931,6 +1948,7 @@ grpc::Status DatabaseGrpcImpl::GetFileInfo(
     }
 
     response->set_project(info->project);
+    response->set_expires_at(info->expiresAt);
     response->set_id(info->id);
     response->set_name(info->name);
     response->set_mime_type(info->mimeType);
@@ -1986,6 +2004,7 @@ grpc::Status DatabaseGrpcImpl::ListFiles(
         file->set_created_at(info.createdAt);
         file->set_is_public(info.isPublic);
         file->set_project(info.project);
+        file->set_expires_at(info.expiresAt);
     }
 
     // v2.8.0 — total_count must reflect what the caller may actually see, not

+ 28 - 1
service/src/database_service.cpp

@@ -781,7 +781,33 @@ DatabaseService::Config DatabaseService::parseConfig(const nlohmann::json& json)
         if (db.contains("files")) {
             auto& files = db["files"];
             config.maxFileSizeMb = files.value("max_file_size_mb", config.maxFileSizeMb);
-            config.fileCleanupIntervalSec = files.value("cleanup_orphans_interval_sec", config.fileCleanupIntervalSec);
+            // v2.8.0 — this interval now drives the file EXPIRY sweep as well
+            // as orphan cleanup, so accept the plainer name too. The original
+            // key kept working; nothing read either of them before 2.8.0
+            // because there was no sweeper thread at all.
+            config.fileCleanupIntervalSec =
+                files.value("cleanup_orphans_interval_sec", config.fileCleanupIntervalSec);
+            config.fileCleanupIntervalSec =
+                files.value("cleanup_interval_sec", config.fileCleanupIntervalSec);
+
+            // v2.8.0 — default retention per file type. The analogue of a
+            // collection's defaultTtlSeconds, so retention is an operator
+            // setting rather than something every uploader must remember.
+            //
+            //   "default_ttl_seconds": {
+            //       "acme:generated": 86400,   // one project, one type
+            //       "generated":      604800,  // any project
+            //       "acme:*":         2592000, // every type in one project
+            //       "*":              0        // everything (0 = keep)
+            //   }
+            if (files.contains("default_ttl_seconds") &&
+                files["default_ttl_seconds"].is_object()) {
+                for (const auto& [k, v] : files["default_ttl_seconds"].items()) {
+                    if (v.is_number_unsigned()) {
+                        config.fileDefaultTtlByType[k] = v.get<uint32_t>();
+                    }
+                }
+            }
             if (files.contains("allowed_types")) {
                 for (const auto& type : files["allowed_types"]) {
                     config.allowedFileTypes.push_back(type.get<std::string>());
@@ -1076,6 +1102,7 @@ void DatabaseService::setupComponents() {
     fileConfig.maxFileSizeMb = config_.maxFileSizeMb;
     fileConfig.allowedMimeTypes = config_.allowedFileTypes;
     fileConfig.cleanupIntervalSec = config_.fileCleanupIntervalSec;
+    fileConfig.defaultTtlSecondsByType = config_.fileDefaultTtlByType;
     files_ = std::make_unique<FileManager>(fileConfig);
 
     // Create replication manager

+ 5 - 0
service/src/database_service.hpp

@@ -32,6 +32,7 @@ class ProjectStoreRegistry;
 #include <mutex>
 #include <string>
 #include <string_view>
+#include <unordered_map>
 #include <thread>
 #include <vector>
 
@@ -158,6 +159,10 @@ public:
         std::vector<std::string> allowedFileTypes;
         uint32_t fileCleanupIntervalSec = 3600;
 
+        // v2.8.0 — default file retention per type. See
+        // FileManager::Config::defaultTtlSecondsByType for the key syntax.
+        std::unordered_map<std::string, uint32_t> fileDefaultTtlByType;
+
         // Encryption settings
         bool encryptionEnabled = true;
         bool autoGenerateKey = true;

+ 130 - 2
service/src/files/file_manager.cpp

@@ -1,6 +1,7 @@
 #include "file_manager.hpp"
 
 #include <spdlog/spdlog.h>
+#include <chrono>
 #include <nlohmann/json.hpp>
 #include <openssl/sha.h>
 
@@ -24,13 +25,108 @@ void FileManager::start() {
         spdlog::error("Failed to create file storage directories: {}", ec.message());
         throw std::runtime_error("Failed to create file storage directories");
     }
-    spdlog::info("FileManager started (directory: {}, two-level storage)", config_.filesDir.string());
+    // v2.8.0 — start the expiry sweeper. `cleanupIntervalSec` has been in
+    // Config from the start but nothing ever read it, so expired files (and
+    // their blobs) were never reclaimed.
+    running_.store(true, std::memory_order_release);
+    cleanup_thread_ = std::thread([this] { cleanupLoop(); });
+
+    spdlog::info("FileManager started (directory: {}, two-level storage, "
+                 "expiry sweep every {}s)",
+                 config_.filesDir.string(), config_.cleanupIntervalSec);
 }
 
 void FileManager::stop() {
+    if (running_.exchange(false, std::memory_order_acq_rel)) {
+        // Wake the sweeper rather than waiting out its interval — an hour-long
+        // sleep would otherwise hold up shutdown.
+        cleanup_cv_.notify_all();
+        if (cleanup_thread_.joinable()) cleanup_thread_.join();
+    }
     spdlog::info("FileManager stopped");
 }
 
+namespace {
+uint64_t now_ms() {
+    return static_cast<uint64_t>(
+        std::chrono::duration_cast<std::chrono::milliseconds>(
+            std::chrono::system_clock::now().time_since_epoch()).count());
+}
+}  // namespace
+
+uint32_t FileManager::defaultTtlFor(const std::string& project,
+                                     const std::string& fileType) const {
+    const auto& m = config_.defaultTtlSecondsByType;
+    if (m.empty()) return 0;
+    const std::string proj = project.empty() ? "default" : project;
+    // Most specific first, so a project can override a global policy and a
+    // single type can override its project's blanket rule.
+    for (const std::string& key : {proj + ":" + fileType, fileType,
+                                   proj + ":*", std::string("*")}) {
+        auto it = m.find(key);
+        if (it != m.end()) return it->second;
+    }
+    return 0;
+}
+
+bool FileManager::isExpired(const FileInfo& info, uint64_t nowMs) {
+    if (info.expiresAt == 0) return false;   // never expires
+    if (nowMs == 0) nowMs = now_ms();
+    return info.expiresAt <= nowMs;
+}
+
+void FileManager::cleanupLoop() {
+    const auto interval = std::chrono::seconds(
+        config_.cleanupIntervalSec ? config_.cleanupIntervalSec : 3600);
+    while (running_.load(std::memory_order_acquire)) {
+        {
+            std::unique_lock<std::mutex> lk(cleanup_mutex_);
+            cleanup_cv_.wait_for(lk, interval, [this] {
+                return !running_.load(std::memory_order_acquire);
+            });
+        }
+        if (!running_.load(std::memory_order_acquire)) break;
+        try {
+            const uint32_t purged = purgeExpired();
+            if (purged > 0) {
+                spdlog::info("File expiry sweep: deleted {} expired file(s)", purged);
+            }
+        } catch (const std::exception& e) {
+            // A sweep failure must not kill the thread; the next tick retries.
+            spdlog::warn("File expiry sweep failed: {}", e.what());
+        }
+    }
+}
+
+uint32_t FileManager::purgeExpired() {
+    auto recordsDir = config_.filesDir / "records";
+    if (!std::filesystem::exists(recordsDir)) return 0;
+
+    // Collect ids first: deleteFile() removes files underneath us, and mutating
+    // the tree while a recursive_directory_iterator walks it is undefined.
+    std::vector<std::string> expired;
+    const uint64_t now = now_ms();
+    for (const auto& entry :
+         std::filesystem::recursive_directory_iterator(recordsDir)) {
+        if (!entry.is_regular_file() ||
+            !entry.path().string().ends_with(".meta.json")) continue;
+        auto stem = entry.path().stem().string();
+        if (stem.ends_with(".meta")) stem = stem.substr(0, stem.size() - 5);
+        auto info = getFileInfoRaw(stem);
+        if (info && isExpired(*info, now)) expired.push_back(stem);
+    }
+
+    uint32_t deleted = 0;
+    for (const auto& id : expired) {
+        // deleteFile() is what actually reclaims the BYTES: it decrements the
+        // content-addressed blob's refcount and unlinks the blob only when the
+        // last reference goes. Unlinking here directly would destroy data still
+        // referenced by other records — blobs are deduplicated across projects.
+        if (deleteFile(id)) ++deleted;
+    }
+    return deleted;
+}
+
 FileManager::StoreResult FileManager::storeFile(const std::vector<uint8_t>& data, FileInfo info) {
     uint64_t maxBytes = config_.maxFileSizeMb * 1024 * 1024;
     if (data.size() > maxBytes) {
@@ -70,6 +166,19 @@ FileManager::StoreResult FileManager::storeFile(const std::vector<uint8_t>& data
         }
     }
 
+    // v2.8.0 — inherit the configured retention for this file type when the
+    // caller gave no explicit expiry. Mirrors how insert() applies a
+    // collection's defaultTtlSeconds.
+    if (info.expiresAt == 0 && !info.neverExpires) {
+        const uint32_t ttl = defaultTtlFor(info.project, info.fileType);
+        if (ttl > 0) {
+            const auto now = std::chrono::duration_cast<std::chrono::milliseconds>(
+                std::chrono::system_clock::now().time_since_epoch()).count();
+            info.expiresAt = static_cast<uint64_t>(now) +
+                             static_cast<uint64_t>(ttl) * 1000ULL;
+        }
+    }
+
     auto id = generateId();
     info.id = id;
     info.checksum = checksum;
@@ -96,6 +205,8 @@ FileManager::StoreResult FileManager::storeFile(const std::vector<uint8_t>& data
         // v2.6.0 — owning project. Empty is normalised to "default" so a record
         // can never be written with no owner.
         meta["project"] = info.project.empty() ? std::string("default") : info.project;
+        // v2.8.0 — absolute expiry, ms since epoch. 0 = never.
+        meta["expiresAt"] = info.expiresAt;
 
         std::ofstream meta_file(rp);
         if (!meta_file) throw std::runtime_error("Failed to create record file: " + rp.string());
@@ -138,7 +249,12 @@ std::vector<uint8_t> FileManager::readFile(const std::string& id) const {
 }
 
 bool FileManager::deleteFile(const std::string& id) {
-    auto info = getFileInfo(id);
+    // v2.8.0 — getFileInfoRaw, NOT getFileInfo. The caller-facing accessor hides
+    // expired records, so resolving through it here made deleting an expired
+    // file impossible: purgeExpired() found records, asked deleteFile() to
+    // remove them, and got false every time. Expiry silently reclaimed nothing
+    // and the blobs leaked. Deletion has to be able to see what it is deleting.
+    auto info = getFileInfoRaw(id);
     if (!info) return false;
 
     auto checksum = info->checksum;
@@ -161,6 +277,16 @@ bool FileManager::deleteFile(const std::string& id) {
 }
 
 std::optional<FileManager::FileInfo> FileManager::getFileInfo(const std::string& id) const {
+    // v2.8.0 — an expired record is gone as far as any caller is concerned,
+    // immediately, without waiting for the sweep. The sweeper uses
+    // getFileInfoRaw() precisely because it must still see them.
+    auto info = getFileInfoRaw(id);
+    if (!info) return std::nullopt;
+    if (isExpired(*info)) return std::nullopt;
+    return info;
+}
+
+std::optional<FileManager::FileInfo> FileManager::getFileInfoRaw(const std::string& id) const {
     auto rp = recordPath(id);
     if (!std::filesystem::exists(rp)) return std::nullopt;
 
@@ -186,6 +312,8 @@ std::optional<FileManager::FileInfo> FileManager::getFileInfo(const std::string&
         // else that predated namespacing.
         info.project = meta.value("project", std::string("default"));
         if (info.project.empty()) info.project = "default";
+        // Absent on records written before 2.8.0, which never expire.
+        info.expiresAt = meta.value("expiresAt", 0ULL);
 
         if (meta.contains("metadata") && meta["metadata"].is_object()) {
             for (const auto& [key, value] : meta["metadata"].items()) {

+ 67 - 0
service/src/files/file_manager.hpp

@@ -3,7 +3,10 @@
 #include <cstdint>
 #include <filesystem>
 #include <mutex>
+#include <atomic>
+#include <condition_variable>
 #include <optional>
+#include <thread>
 #include <string>
 #include <unordered_map>
 #include <vector>
@@ -23,6 +26,23 @@ public:
         uint64_t maxFileSizeMb = 500;
         std::vector<std::string> allowedMimeTypes;
         uint32_t cleanupIntervalSec = 3600;
+
+        // v2.8.0 — default retention per file type, in seconds. 0 means keep
+        // forever, which stays the default for every type.
+        //
+        // The analogue of a collection's `defaultTtlSeconds`: an uploader that
+        // passes no TTL inherits the policy for its type, so retention can be
+        // set once by an operator rather than depended on at every call site.
+        //
+        // Keys are matched most-specific-first:
+        //   "<project>:<fileType>"  exact
+        //   "<fileType>"            any project
+        //   "<project>:*"           every type in one project
+        //   "*"                     everything
+        //
+        // An explicit ttl on the upload always wins, including an explicit 0,
+        // which is how a caller opts a single file out of a default.
+        std::unordered_map<std::string, uint32_t> defaultTtlSecondsByType;
     };
 
     struct FileInfo {
@@ -42,6 +62,18 @@ public:
         // written before 2.6.0 carry no project key; absent is read as
         // "default", never as unknown or empty.
         std::string project = "default";
+
+        // v2.8.0 — set when the caller explicitly asked for no expiry, which
+        // suppresses the configured per-type default. Not persisted: once
+        // storeFile() has honoured it, expiresAt == 0 records the outcome.
+        bool neverExpires = false;
+
+        // v2.8.0 — absolute expiry, ms since epoch. 0 means never.
+        //
+        // Stored absolute rather than as a TTL so a record's lifetime does not
+        // restart when it is read or when the service is restarted. Callers set
+        // it via a relative ttl_seconds on upload, which is converted once.
+        uint64_t expiresAt = 0;
     };
 
     struct ListResult {
@@ -104,7 +136,42 @@ public:
     // already declare a project are left untouched.
     uint32_t stampMissingProjects();
 
+    // v2.8.0 — resolve the configured default retention for a file type, in
+    // seconds. 0 means keep forever. See Config::defaultTtlSecondsByType for the
+    // matching order.
+    [[nodiscard]] uint32_t defaultTtlFor(const std::string& project,
+                                         const std::string& fileType) const;
+
+    // v2.8.0 — true if `info` has an expiry in the past.
+    [[nodiscard]] static bool isExpired(const FileInfo& info, uint64_t nowMs = 0);
+
+    // v2.8.0 — delete every expired file, returning how many went.
+    //
+    // Routes each one through deleteFile(), which is what makes the BLOB go and
+    // not merely the metadata record: it decrements the shared refcount and
+    // unlinks the blob only when the last reference drops. Blobs are
+    // content-addressed and deduplicated across projects, so unlinking on the
+    // first expiry would corrupt every other record pointing at the same bytes.
+    //
+    // Called on the sweeper thread every `cleanupIntervalSec`, and safe to call
+    // by hand.
+    uint32_t purgeExpired();
+
 private:
+    // v2.8.0 — expiry sweeper. `cleanupIntervalSec` existed in Config since the
+    // beginning but nothing ever read it: there was no thread, so nothing was
+    // ever reclaimed. This is that thread.
+    std::thread cleanup_thread_;
+    std::atomic<bool> running_{false};
+    std::mutex cleanup_mutex_;
+    std::condition_variable cleanup_cv_;
+    void cleanupLoop();
+
+    // Project-blind and expiry-blind. The sweeper needs to SEE expired records
+    // in order to delete them, while every caller-facing accessor must treat
+    // them as already gone.
+    [[nodiscard]] std::optional<FileInfo> getFileInfoRaw(const std::string& id) const;
+
     std::string generateId() const;
     std::filesystem::path blobPath(const std::string& checksum) const;
     std::filesystem::path blobRefPath(const std::string& checksum) const;

+ 38 - 0
service/src/memory_store.cpp

@@ -446,6 +446,25 @@ bool MemoryStore::update(const std::string& collection, const std::string& id, c
     updated.updatedAt = currentTimeFor(collection);
     updated.nodeId = config_.nodeId;
 
+    // v2.8.0 — carry the existing expiry across when the caller supplied none.
+    //
+    // UpdateRequest has no ttl field, so `doc.expiresAt` is 0 for every update
+    // arriving over gRPC. This preserved createdAt and createdBy but not
+    // expiresAt, so an update WIPED the expiry and made a TTL'd document
+    // permanent. patchDocument() mutates the stored document in place and so
+    // always preserved it - the two write paths disagreed, and the one that lost
+    // data was the more commonly used.
+    //
+    // It bites hardest where it is least visible: a record written with a TTL
+    // and then updated as its status changes (the normal case) silently stopped
+    // expiring and accumulated forever.
+    //
+    // A non-zero incoming expiresAt still wins, so a caller CAN deliberately
+    // change or extend a TTL by supplying one.
+    if (updated.expiresAt == 0) {
+        updated.expiresAt = it->second.expiresAt;
+    }
+
     // Add to new expiration index
     if (updated.expiresAt > 0) {
         addToExpirationIndex(*coll, id, updated.expiresAt);
@@ -726,6 +745,25 @@ bool MemoryStore::updateIfVersion(const std::string& collection, const std::stri
     updated.updatedAt = currentTimeFor(collection);
     updated.nodeId = config_.nodeId;
 
+    // v2.8.0 — carry the existing expiry across when the caller supplied none.
+    //
+    // UpdateRequest has no ttl field, so `doc.expiresAt` is 0 for every update
+    // arriving over gRPC. This preserved createdAt and createdBy but not
+    // expiresAt, so an update WIPED the expiry and made a TTL'd document
+    // permanent. patchDocument() mutates the stored document in place and so
+    // always preserved it - the two write paths disagreed, and the one that lost
+    // data was the more commonly used.
+    //
+    // It bites hardest where it is least visible: a record written with a TTL
+    // and then updated as its status changes (the normal case) silently stopped
+    // expiring and accumulated forever.
+    //
+    // A non-zero incoming expiresAt still wins, so a caller CAN deliberately
+    // change or extend a TTL by supplying one.
+    if (updated.expiresAt == 0) {
+        updated.expiresAt = it->second.expiresAt;
+    }
+
     // v2.2.1 — extract _vector + sync coll.vectors (and the LMDB vector
     // sub-db). Pre-2.2.1 updateIfVersion left _vector embedded in
     // doc.data() and never touched coll.vectors, which silently broke

+ 32 - 0
tests/CMakeLists.txt

@@ -623,3 +623,35 @@ find_package(Threads REQUIRED)
 target_link_libraries(test_policy_manager PRIVATE Threads::Threads)
 
 add_test(NAME test_policy_manager COMMAND test_policy_manager)
+
+# v2.8.0 — document TTL across write paths. Pins that update() preserves an
+# existing expiry: it used to wipe it, so any TTL'd document that was ever
+# updated became permanent.
+add_executable(test_document_ttl
+    test_document_ttl.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/memory_store.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/config/collection_config_manager.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/persistence/history_store.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/persistence/wal.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/json_parse.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/doc_binary.cpp
+)
+target_include_directories(test_document_ttl PRIVATE
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src
+    ${yyjson_INCLUDE_DIRS}
+)
+target_link_libraries(test_document_ttl PRIVATE ${yyjson_LIBRARIES})
+if(TARGET nlohmann_json::nlohmann_json)
+    target_link_libraries(test_document_ttl PRIVATE nlohmann_json::nlohmann_json)
+else()
+    target_include_directories(test_document_ttl PRIVATE ${NLOHMANN_JSON_INCLUDE_DIRS})
+endif()
+if(TARGET spdlog::spdlog)
+    target_link_libraries(test_document_ttl PRIVATE spdlog::spdlog)
+else()
+    target_link_libraries(test_document_ttl PRIVATE ${SPDLOG_LIBRARIES})
+    target_include_directories(test_document_ttl PRIVATE ${SPDLOG_INCLUDE_DIRS})
+endif()
+find_package(Threads REQUIRED)
+target_link_libraries(test_document_ttl PRIVATE Threads::Threads)
+add_test(NAME test_document_ttl COMMAND test_document_ttl)

+ 141 - 0
tests/test_document_ttl.cpp

@@ -0,0 +1,141 @@
+// v2.8.0 — document TTL lifecycle across write paths.
+//
+// The bug this pins down: `update()` WIPED a document's expiry.
+//
+// updateIfVersion built a fresh Document from the caller's request, and
+// UpdateRequest has no ttl field, so the incoming expiresAt was always 0. It
+// deliberately carried createdAt and createdBy across from the stored document
+// but not expiresAt - so any TTL'd document that was ever updated became
+// permanent. `patch()` mutates the stored document in place and therefore always
+// preserved it, which is why the two paths disagreed.
+//
+// This matters most where it is least visible: a workflow execution written with
+// a TTL and then updated as its status changes (the normal case) silently lost
+// its expiry and accumulated forever.
+
+#include <atomic>
+#include <chrono>
+#include <filesystem>
+#include <iostream>
+#include <string>
+#include <unistd.h>
+
+#include <nlohmann/json.hpp>
+
+#include "memory_store.hpp"
+
+using namespace smartbotic::database;
+
+namespace {
+
+int g_pass = 0;
+int g_fail = 0;
+
+void check(bool cond, const char* msg) {
+    if (cond) { ++g_pass; }
+    else { ++g_fail; std::cerr << "FAIL: " << msg << "\n"; }
+}
+
+struct Fixture {
+    MemoryStore store;
+    Fixture() : store(MemoryStore::Config{}) {
+        store.start();
+        CollectionOptions o;
+        store.createCollection("things", o);
+    }
+    ~Fixture() { store.stop(); }
+};
+
+Document mkDoc(const std::string& id, uint64_t expiresAt = 0) {
+    Document d;
+    d.id = id;
+    d.collection = "things";
+    d.expiresAt = expiresAt;
+    d.set_data(nlohmann::json{{"n", 1}});
+    return d;
+}
+
+uint64_t future_ms(uint64_t secs) {
+    const auto now = std::chrono::duration_cast<std::chrono::milliseconds>(
+        std::chrono::system_clock::now().time_since_epoch()).count();
+    return static_cast<uint64_t>(now) + secs * 1000ULL;
+}
+
+// -------------------------------------------------------------------------
+
+void test_insert_sets_expiry() {
+    Fixture f;
+    const uint64_t exp = future_ms(3600);
+    f.store.insert("things", mkDoc("a", exp));
+    auto got = f.store.get("things", "a");
+    check(got.has_value(), "inserted");
+    check(got && got->expiresAt == exp, "insert stores the expiry it was given");
+}
+
+// THE regression. An update with no TTL must not silently make the document
+// permanent.
+void test_update_preserves_expiry() {
+    Fixture f;
+    const uint64_t exp = future_ms(3600);
+    f.store.insert("things", mkDoc("a", exp));
+
+    auto stored = f.store.get("things", "a");
+    Document next = mkDoc("a");            // no expiry, exactly as UpdateRequest arrives
+    next.set_data(nlohmann::json{{"n", 2}});
+    f.store.updateIfVersion("things", "a", next, stored->version);
+
+    auto after = f.store.get("things", "a");
+    check(after.has_value(), "still present after update");
+    check(after && after->data().value("n", 0) == 2, "the update applied");
+    check(after && after->expiresAt == exp,
+          "update MUST preserve the existing expiry - wiping it makes a TTL'd "
+          "document permanent");
+}
+
+void test_update_can_change_expiry_when_given_one() {
+    Fixture f;
+    f.store.insert("things", mkDoc("a", future_ms(60)));
+    auto stored = f.store.get("things", "a");
+
+    const uint64_t later = future_ms(7200);
+    Document next = mkDoc("a", later);     // explicit new expiry
+    f.store.updateIfVersion("things", "a", next, stored->version);
+
+    auto after = f.store.get("things", "a");
+    check(after && after->expiresAt == later,
+          "an explicit expiry on update replaces the old one");
+}
+
+void test_update_of_a_document_with_no_expiry_stays_permanent() {
+    Fixture f;
+    f.store.insert("things", mkDoc("a"));   // no TTL
+    auto stored = f.store.get("things", "a");
+    Document next = mkDoc("a");
+    f.store.updateIfVersion("things", "a", next, stored->version);
+    auto after = f.store.get("things", "a");
+    check(after && after->expiresAt == 0,
+          "preserving must not invent an expiry for a document that had none");
+}
+
+void test_patch_preserves_expiry() {
+    Fixture f;
+    const uint64_t exp = future_ms(3600);
+    f.store.insert("things", mkDoc("a", exp));
+    f.store.patchDocument("things", "a", nlohmann::json{{"n", 5}}, "");
+    auto after = f.store.get("things", "a");
+    check(after && after->expiresAt == exp, "patch preserves the expiry");
+    check(after && after->data().value("n", 0) == 5, "and applies the patch");
+}
+
+}  // namespace
+
+int main() {
+    std::cout << "=== test_document_ttl ===\n";
+    test_insert_sets_expiry();
+    test_update_preserves_expiry();
+    test_update_can_change_expiry_when_given_one();
+    test_update_of_a_document_with_no_expiry_stays_permanent();
+    test_patch_preserves_expiry();
+    std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
+    return g_fail == 0 ? 0 : 1;
+}

+ 205 - 0
tests/test_file_project_scope.cpp

@@ -16,6 +16,7 @@
 // candidate file, read the flag, learn whether a peer namespace has it.
 
 #include <atomic>
+#include <chrono>
 #include <filesystem>
 #include <fstream>
 #include <iostream>
@@ -50,6 +51,20 @@ struct TmpFm {
     std::string dir;
     std::unique_ptr<FileManager> fm;
 
+    explicit TmpFm(std::unordered_map<std::string, uint32_t> ttlByType) {
+        static std::atomic<int> counter{0};
+        dir = "/tmp/sbdb-file-project-test-" + std::to_string(::getpid()) + "-ttl" +
+              std::to_string(counter.fetch_add(1));
+        std::error_code ec;
+        fs::remove_all(dir, ec);
+        fs::create_directories(dir);
+        FileManager::Config cfg;
+        cfg.filesDir = dir;
+        cfg.defaultTtlSecondsByType = std::move(ttlByType);
+        fm = std::make_unique<FileManager>(cfg);
+        fm->start();
+    }
+
     TmpFm() {
         static std::atomic<int> counter{0};
         dir = "/tmp/sbdb-file-project-test-" + std::to_string(::getpid()) + "-" +
@@ -86,6 +101,23 @@ struct TmpFm {
         out << json;
     }
 
+    fs::path blobPathFor(const std::string& checksum) const {
+        std::string hash = checksum;
+        if (hash.rfind("sha256:", 0) == 0) hash = hash.substr(7);
+        const std::string prefix = hash.size() >= 2 ? hash.substr(0, 2) : "00";
+        return fs::path(dir) / "blobs" / prefix / hash;
+    }
+
+    // Rewrite a record's expiry in place, to age a file without sleeping.
+    void setExpiry(const std::string& id, uint64_t expiresAt) const {
+        auto rp = recordPath(id);
+        nlohmann::json j;
+        { std::ifstream in(rp); in >> j; }
+        j["expiresAt"] = expiresAt;
+        std::ofstream out(rp);
+        out << j.dump(2);
+    }
+
     nlohmann::json rawRecord(const std::string& id) const {
         std::ifstream in(recordPath(id));
         nlohmann::json j;
@@ -216,6 +248,168 @@ void test_stamping_leaves_existing_projects_alone() {
           "and it stays in its own project, not moved to default");
 }
 
+// v2.8.0 — file TTL.
+//
+// The assertion that matters most is `test_expiry_respects_shared_blobs`. Blobs
+// are content-addressed and deduplicated across projects, so "delete the file"
+// cannot mean "unlink the blob": two records may point at the same bytes. Expiry
+// therefore routes through deleteFile(), which decrements the refcount and
+// unlinks only on the last reference. Getting that wrong destroys live data
+// belonging to a file that has not expired.
+
+uint64_t now_ms_test() {
+    return static_cast<uint64_t>(
+        std::chrono::duration_cast<std::chrono::milliseconds>(
+            std::chrono::system_clock::now().time_since_epoch()).count());
+}
+
+void test_ttl_zero_never_expires() {
+    TmpFm fm;
+    auto info = mk("keep.bin", "acme");
+    info.expiresAt = 0;
+    auto id = fm->storeFile({1, 2, 3}, info).id;
+    check(fm->getFileInfo(id).has_value(), "ttl 0 file is readable");
+    check(fm->purgeExpired() == 0, "ttl 0 file is not swept");
+    check(fm->getFileInfo(id).has_value(), "and survives a sweep");
+}
+
+void test_expired_file_is_invisible_before_the_sweep() {
+    TmpFm fm;
+    auto info = mk("gone.bin", "acme");
+    info.expiresAt = now_ms_test() - 1000;      // already expired
+    auto id = fm->storeFile({4, 5, 6}, info).id;
+
+    // No sweep has run yet - expiry must still be immediate to every caller.
+    check(!fm->getFileInfo(id).has_value(), "expired file reads as absent at once");
+    check(!fm->getFileInfoIn("acme", id).has_value(), "and through the scoped accessor");
+    check(fm->listFiles("acme").totalCount == 0, "and is absent from listings");
+    bool threw = false;
+    try { (void)fm->readFileIn("acme", id); } catch (const std::exception&) { threw = true; }
+    check(threw, "and its bytes are not served");
+}
+
+void test_expiry_deletes_the_blob_not_just_the_record() {
+    TmpFm fm;
+    auto info = mk("bytes.bin", "acme");
+    info.expiresAt = now_ms_test() - 1;
+    std::vector<uint8_t> payload{9, 9, 9, 9, 9};
+    auto res = fm->storeFile(payload, info);
+
+    auto blob = fm.blobPathFor(res.checksum);
+    check(fs::exists(blob), "blob written");
+    check(fs::exists(fm.recordPath(res.id)), "record written");
+
+    check(fm->purgeExpired() == 1, "sweep deletes the expired file");
+    check(!fs::exists(fm.recordPath(res.id)), "metadata record removed");
+    check(!fs::exists(blob),
+          "THE BLOB IS REMOVED - reclaiming metadata alone would leak the bytes");
+    check(fm->getRefCount(res.checksum) == 0, "refcount cleared");
+}
+
+void test_expiry_respects_shared_blobs() {
+    TmpFm fm;
+    std::vector<uint8_t> shared{7, 7, 7, 7};
+
+    auto a = mk("expiring.bin", "acme");
+    a.expiresAt = now_ms_test() - 1;            // expires now
+    auto ra = fm->storeFile(shared, a);
+
+    auto b = mk("permanent.bin", "acme");
+    b.expiresAt = 0;                            // never
+    auto rb = fm->storeFile(shared, b);
+
+    check(ra.checksum == rb.checksum, "identical bytes share one blob");
+    check(fm->getRefCount(ra.checksum) == 2, "two references to that blob");
+
+    auto blob = fm.blobPathFor(ra.checksum);
+    check(fm->purgeExpired() == 1, "only the expired record is swept");
+    check(!fs::exists(fm.recordPath(ra.id)), "expired record gone");
+    check(fs::exists(fm.recordPath(rb.id)), "the permanent record survives");
+    check(fs::exists(blob),
+          "the SHARED BLOB survives - unlinking it would destroy live data");
+    check(fm->getRefCount(ra.checksum) == 1, "refcount decremented to one");
+    check(fm->readFileIn("acme", rb.id).size() == shared.size(),
+          "and the surviving file still serves its bytes");
+
+    // Now expire the survivor too: the blob must finally go.
+    fm.setExpiry(rb.id, now_ms_test() - 1);
+    check(fm->purgeExpired() == 1, "the last reference is swept");
+    check(!fs::exists(blob), "blob unlinked once the last reference goes");
+    check(fm->getRefCount(ra.checksum) == 0, "refcount cleared");
+}
+
+void test_sweep_is_idempotent() {
+    TmpFm fm;
+    auto info = mk("x.bin", "acme");
+    info.expiresAt = now_ms_test() - 1;
+    fm->storeFile({1}, info);
+    check(fm->purgeExpired() == 1, "first sweep deletes it");
+    check(fm->purgeExpired() == 0, "second sweep is a no-op");
+}
+
+void test_legacy_records_never_expire() {
+    TmpFm fm;
+    // Pre-2.8.0 record: no expiresAt key at all.
+    fm.writeLegacyRecord("old", R"({"id":"old","name":"a","fileType":"document","project":"acme"})");
+    check(fm->getFileInfo("old").has_value(), "legacy record readable");
+    check(fm->purgeExpired() == 0,
+          "a record with no expiresAt must never be swept - absent means never");
+}
+
+// v2.8.0 — per-fileType default retention: the analogue of a collection's
+// defaultTtlSeconds, so retention is set once by an operator instead of relied
+// on at every upload site.
+void test_default_ttl_by_type() {
+    TmpFm fm({{"generated", 3600}});
+    auto a = fm->storeFile({1}, mk("g.bin", "acme", "generated"));
+    auto ga = fm->getFileInfo(a.id);
+    check(ga && ga->expiresAt > 0, "a type with a default gets an expiry");
+
+    auto b = fm->storeFile({2}, mk("d.bin", "acme", "document"));
+    auto gb = fm->getFileInfo(b.id);
+    check(gb && gb->expiresAt == 0, "a type with no default keeps forever");
+}
+
+void test_default_ttl_matching_order() {
+    // Most specific wins: project+type, then type, then project wildcard, then
+    // the global wildcard.
+    TmpFm fm({{"acme:generated", 10}, {"generated", 20}, {"acme:*", 30}, {"*", 40}});
+    check(fm->defaultTtlFor("acme", "generated") == 10, "project+type is most specific");
+    check(fm->defaultTtlFor("other", "generated") == 20, "type matches any project");
+    check(fm->defaultTtlFor("acme", "plugin") == 30, "project wildcard covers other types");
+    check(fm->defaultTtlFor("other", "plugin") == 40, "global wildcard is the last resort");
+    check(fm->defaultTtlFor("other", "") == 40, "empty type still falls through");
+}
+
+void test_explicit_expiry_overrides_the_default() {
+    TmpFm fm({{"generated", 3600}});
+    auto info = mk("pinned.bin", "acme", "generated");
+    info.expiresAt = now_ms_test() + 60000;      // caller decided
+    auto r = fm->storeFile({3}, info);
+    auto got = fm->getFileInfo(r.id);
+    check(got && got->expiresAt == info.expiresAt,
+          "an explicit expiry is not overwritten by the type default");
+}
+
+void test_never_expires_opts_out_of_the_default() {
+    TmpFm fm({{"generated", 1}});
+    auto info = mk("keep.bin", "acme", "generated");
+    info.neverExpires = true;                    // explicit ttl=0 from the wire
+    auto r = fm->storeFile({4}, info);
+    auto got = fm->getFileInfo(r.id);
+    check(got && got->expiresAt == 0,
+          "an explicit never-expire beats the configured default");
+    check(fm->purgeExpired() == 0, "and the sweeper leaves it alone");
+}
+
+void test_no_config_means_no_expiry() {
+    TmpFm fm;                                    // no defaults at all
+    auto r = fm->storeFile({5}, mk("x.bin", "acme", "generated"));
+    auto got = fm->getFileInfo(r.id);
+    check(got && got->expiresAt == 0,
+          "with no retention configured nothing expires - the default default");
+}
+
 }  // namespace
 
 int main() {
@@ -228,6 +422,17 @@ int main() {
     test_dedup_flag_does_not_leak_across_projects();
     test_stamp_projects_is_idempotent();
     test_stamping_leaves_existing_projects_alone();
+    test_ttl_zero_never_expires();
+    test_expired_file_is_invisible_before_the_sweep();
+    test_expiry_deletes_the_blob_not_just_the_record();
+    test_expiry_respects_shared_blobs();
+    test_sweep_is_idempotent();
+    test_legacy_records_never_expire();
+    test_default_ttl_by_type();
+    test_default_ttl_matching_order();
+    test_explicit_expiry_overrides_the_default();
+    test_never_expires_opts_out_of_the_default();
+    test_no_config_means_no_expiry();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;

Beberapa file tidak ditampilkan karena terlalu banyak file yang berubah dalam diff ini