Skip to content

Retention, purge and the consistency sweep ​

The platform never deletes a row: every StoredFile is soft-deletable by construction, and the global filter hides what was marked. Keeping a file forever, though, is a legal obligation turned inside out — privacy law requires erasing data once its purpose is over. GrydFiles reconciles the two with the tombstone: the record follows the framework's rules and stays, and what is actually destroyed is the object in the bucket, which is not an entity.

This page describes how long a file lives, the three ways it dies, and the weekly job that checks the records and the bucket still agree.

How long a file lives ​

The retention comes with the upload, declared by the consumer or inherited from the profile. The module does not know that a tax document is kept for five years — the consumer does. What the module guarantees is that the question was asked and the answer is on the record.

ModeDies when
UntilReleasedNothing references it and the grace period has run out since the last reference was released.
RetainUntilretainUntil has passed and nothing references it — both conditions, always.
PermanentNever by a job (RN-GF-09). Only a manual purge, with admin:system, kills it.

The grace period is GrydFiles:PurgeGracePeriodDays (30 by default). It counts from lastReferenceReleasedAt, the date denormalised on the record next to referenceCount — never from an aggregation over the reference table (D8). A file nobody ever referenced counts from its verdict (scannedAt) instead: without a release to count from, an upload its consumer forgot to link would otherwise stay in the bucket forever, and nobody would ever look.

A legal hold freezes every purge — the retention job, the quarantine expiry and the manual route, including an erasure request from the data subject.

Changing retention after the upload ​

PUT /files/{id}/retention only tightens custody:

  • retainUntil moves forward. A date that is not later than the current one answers FILE_RETENTION_SHORTENED_UNPROCESSABLE (422) and changes nothing. Shortening custody by API is how evidence gets destroyed by accident; when it is legitimate, the path is an explicit purge, which leaves an author and a reason. Only a RetainUntil file has a date to extend — sending one for the other modes is a 400 with no code.
  • legalHold: true places a hold, for retain:files.
  • legalHold: false over a held file releases it — the one change that loosens custody — and needs admin:system on top of retain:files. The release goes to the audit trail; placing a hold does not, because a hold only ever protects content.

Every refusal is decided before anything changes, so a request carrying a date and a hold never applies half of itself. A tombstone answers FILE_PURGED_CONFLICT, and a product file (no tenant) needs admin:system for any change. QuarantineRetentionDays is global to the module and has no field here.

The three ways a file dies ​

PathSelectsReason on the tombstonepurgedBy
POST /files/{id}/purgethe file named, if Availablethe reason the caller wrotethe caller
RetentionPurgeJob (nightly, 04:00)Available files whose mode says they are duejob:retencaonull
ConsistencySweepJob (weekly, Sunday 05:00)Infected files past QuarantineRetentionDays (365)job:quarentena-vencidanull

purgedBy is null exactly when the reason starts with job: (D16): the framework has no system user, and the origin lives in the reason. The manual route refuses a reason that borrows the prefix, one that is empty, and one longer than the 200 characters the tombstone keeps.

The manual route ​

POST /files/{id}/purge needs purge:files, plus admin:system for a Permanent or a product file. It refuses with the record's own holds:

  • FILE_LEGAL_HOLD_CONFLICT (409) while the file is held — releasing it is a different route and a different authority;
  • FILE_STILL_REFERENCED_CONFLICT (409) while a reference row of the file is active, listing the ownerScope values holding it. A count above zero with no active row is refused too, and settled by the nightly reconciliation.

redactFileName: true takes originalFileName off the tombstone as well — for the erasure whose object is the name itself. The column is then null, the one record where it can be, and the audit entry says the name was redacted on request. There is no DELETE /files/{id}: destroying content has a reason and a permission of its own, and the easiest verb of the API would be an invitation to an accident.

The retention job ​

The selection reads stored_files alone, over the (status, referenceCount, lastReferenceReleasedAt) index, and pages by id so a candidate skipped in one run cannot keep a run from ending. Each candidate is then purged in its own transaction, under a FOR UPDATE SKIP LOCKED lock on the record and on the owner of its canonical set: two overlapping runs take disjoint work and neither ever waits. Under the lock the record is read again with its reference rows, and the rule is asked again — a reference added, a hold placed or a window extended since the selection is honoured, and a count of zero over an active row is skipped and logged, never trusted.

It runs an hour after the reference reconciliation, so the pair it selects by has just been returned to what the rows say. Hangfire's [DisableConcurrentExecution] is not used: Hangfire schedules RecurringJobExecutionAdapter<TJob>, so a filter on the job class would never be read.

What survives ​

The tombstone keeps the digest, the size, the type, the original name (unless redacted), the profile, the dates, the scan result, the rejection code, the location where the object was, and who purged it and why. For quarantined content the signature name and the whole scan history stay forever: what expires is the binary, never the evidence of the incident.

Shared bytes ​

Dedupe means one object can belong to several records (see the dedupe page). Each record's own retention decides when its row dies; the bytes die with the last live member of the canonical set, and the object deleted is always the owner's — whose tombstone keeps the location precisely so the last reuser can still find it. A record whose window closed gets its tombstone even while another member keeps the blob alive.

Every purge goes through FilePurger: it confirms the aggregate would accept the tombstone, deletes the owner's object when this purge ends the blob, treats an object already gone as success (which is what makes a re-run after a crash converge), and only then writes the tombstone. The object goes first and the tombstone second — never the other way round (RN-GF-10) — and a storage failure stops both.

The consistency sweep ​

The state of a file is the record's, not the object's. That is what makes the module safe, and it is exactly why the two have to be compared: silent drift between them is how a file module loses the trust of whoever depends on it. ConsistencySweepJob asks the bucket about every record that owns a location and reports three kinds of inconsistency:

KindMeaning
Missing objectA live owner — Scanning, Failed, Available or Infected — whose object is gone. Content vanished outside the purge.
Object outliving its purgeA tombstoned owner whose object is still there, and no member of its set is alive. The purge did not complete.
Ownership violationA live record breaking canonicalFileId IS NULL ⟺ objectKey IS NOT NULL (D3).

It reports and never repairs — an automatic correction would hide the cause. Deduplicated records have no location and are never compared, so they never become false defects; pending reservations are left out because their object may not exist yet. A run that finds anything logs one error naming the files (never a key or a bucket) and moves gryd_files.consistency_sweep.alerts; a clean run makes no noise. Records storage could not be asked about are counted as unverified, not as defects.

The sweep is weekly because it costs one existence check per owner, tombstones included, and that grows with the history. The reference reconciliation is not part of it: it is SQL over the database alone, runs nightly on its own schedule, and is described on the references page.

Not covered: an object with no record at all, such as the blob left behind when deleting a duplicate's own upload fails during dedupe. No row knows that key; finding it would take a sweep of the bucket against the table.

Metrics ​

All in the GrydFiles.Sweeps meter:

InstrumentUnit
gryd_files.retention_purge.purgedrecords
gryd_files.retention_purge.objects_deletedobjects
gryd_files.retention_purge.bytes_freedbytes
gryd_files.quarantine_expiry.purgedrecords
gryd_files.quarantine_expiry.bytes_freedbytes
gryd_files.consistency_sweep.checkedrecords
gryd_files.consistency_sweep.missing_objectsrecords
gryd_files.consistency_sweep.objects_outliving_purgerecords
gryd_files.consistency_sweep.ownership_violationsrecords
gryd_files.consistency_sweep.unverifiedrecords
gryd_files.consistency_sweep.alertsalerts

Fewer objects deleted than records purged is normal: it is shared bytes outliving a member of their set.


The decisions behind this page are in ADR 0010; the contract is in the specification.

Released under the MIT License.