Skip to content

Dedupe and canonical resolution ​

One record per upload, one object per content. Two people uploading the same PDF into the same tenant get two StoredFile records and one blob in the bucket — who uploaded it, when, and under which profile is the file's own trail and does not merge. What merges is the bytes.

This page covers where that decision is taken, what it leaves behind, and the two consequences the rest of the module has to live with: reads that must resolve a canonical, and a purge that cannot decide anything by looking at a single record.

Where the decision is taken ​

Inside the scan job, in the same transaction that releases the file.

The digest only exists after the pass — the server computes it on the single read that also feeds the antivirus and the content-signature check — so before that there is nothing to compare. Deciding after the release instead would leave a window in which a duplicate is Available while owning a blob the module intends to delete, and the consistency sweep would have to be taught to tell that window apart from a real defect.

The sequence is: mark clean → look for an owner of the same content in the same tenant → link and zero the location → commit → delete the duplicate's own object.

A candidate is an Available record with no canonical of its own, in the same tenant. A digest on an Infected record names quarantined evidence under a different prefix, which nothing may ever be pointed at; the remaining states carry no digest at all.

Dedupe never crosses tenant, even with an identical hash. Sharing a blob between customers would buy storage and sell an inference channel — response time alone would tell one tenant that another holds the same file. Two tenants with the same content end up with two objects, deliberately.

What a deduplicated record looks like ​

Its objectKey, bucketName and storageProvider are null.

This is the decision the rest of the module leans on. Keeping the key of an object that was just deleted would be a lie written into a column: anything reading it would ask for a URL to a blob that no longer exists and get a 404 from the provider. With the nulls, the invariant

canonicalFileId IS NULL  ⟺  objectKey IS NOT NULL

holds for every record that has not been purged, and possession of a blob has exactly one representation. Without it, the consistency sweep would report every deduplicated record as a defect — a record with no tombstone whose object is missing.

Copying the canonical's key into the duplicate was evaluated and rejected: when the canonical blob eventually died, every reuser would point at a dead key and the same false positive would come back in bulk.

The tombstone is the one exception. A purged record preserves all three fields, because the question a tombstone answers is where the file used to be.

There is no chain: the canonical of a duplicate is always a record that owns its blob.

The race, and who arbitrates it ​

Two uploads of the same content can both pass the lookup before either of them writes. The partial unique indexes are the arbiter, and their refusal is not an error — it is the answer to a question that had no answer at lookup time. The loser catches the unique violation, reads the winner back, becomes the duplicate and finishes Available. Neither upload fails for the consumer.

The refusal is only treated this way when it names one of the two dedupe indexes. A unique violation on any other index is a defect, and swallowing it as "somebody got here first" would turn that defect into a file released under a rule nobody re-read.

The index predicate ​

Both indexes are filtered on canonicalFileId IS NULL AND sha256 IS NOT NULL AND status = 'Available'.

The status predicate is = 'Available' and not the broader <> 'Purged', and the difference matters: marking a file infected also writes its digest and leaves canonicalFileId null, so under the wider predicate a tenant that uploaded the same malware twice put two rows into the index and the second quarantine was refused by a unique violation nothing catches — leaving the record stuck mid-scan while the job retried the identical failure forever. Narrowing it to the set the dedupe lookup actually reads costs nothing: the other non-Available states carry no digest and were never in the index.

There are two indexes rather than one because the tenant is nullable for product files, and two NULLs never collide in a plain unique index. The second covers exactly that range.

Reading a deduplicated file ​

Resolving the canonical is obligatory, not an optimisation. A deduplicated record has no location of its own, so code that reaches for objectKey directly works perfectly for every file until the first duplicate and then hands somebody a null.

CanonicalBlobResolver is where that resolution happens. Given a record it returns a FileBlobLocation — provider, bucket and key together, so they cannot be mixed up between two records — reading from the record itself when it owns its blob, and from its canonical when it does not.

Only the location comes from the canonical. The name, the profile, the author, the dates, the scan result and the rejection code are the requested record's own. So is the state: the "only Available" rule of a download is evaluated against the record the consumer asked about, not against whichever record happens to own the bytes.

A location never leaves the module (RN-GF-07). It is not in a DTO, not in a response and not in a log line; what crosses the boundary is the signed URL it was used to produce — and every signature in the module is issued from one place, which refuses anything under the quarantine prefix.

Deciding the fate of shared bytes ​

Dedupe creates records that are not the owners of their own content, and that turns a question about a row into a question about a set. Deleting an object because one record ran out of references would delete the file out from under every other record reusing it.

CanonicalFileSet is that set: the owner plus every record whose canonicalFileId names it. It is built from IStoredFileStore.GetCanonicalSetAsync, which returns the whole set given any member — owner or reuser — because a purge asked about a duplicate would otherwise evaluate a set of one and conclude the blob is collectable.

The set answers three things:

QuestionRule
Is anybody holding it?Any active reference on any member. Read from the child rows, never from the denormalised count.
Is it frozen?A legal hold on any member freezes the content for all of them.
Does purging this member end the blob?Only when every other member is already a tombstone, and neither of the above is true (BlobDiesWith).

The blob dies with the last live member. "Nobody in the set holds it" is not enough on its own: a reuser that is Available, unreferenced and still inside its own retention — or Permanent — holds nothing today and still promises its content to whoever references it tomorrow. Deleting the bytes when the owner's window closed would leave that reuser Available over nothing. So each member's own retention decides when its row dies, and the bytes go only with the purge that leaves no other member alive.

The object deleted is always the owner's. A reuser never had a location of its own, and an owner purged first keeps its location on the tombstone precisely so that the last reuser to go can still find — and destroy — the bytes it has been serving. Until then a download of the reuser keeps resolving to that location.

Every purge path goes through FilePurger, in GrydFiles.Application: it confirms the aggregate would accept the tombstone, deletes the owner's object when BlobDiesWith says so (an object already gone is not an error, which is what makes a re-run after a crash converge), and only then writes the tombstone. The manual route, the retention job and the quarantine expiry share it, so the question with no undo is answered in one place.

A hold is honoured across the set because a hold exists when somebody is required to keep the bytes, and bytes shared by four records are the same bytes — honouring it only on the record carrying it would destroy exactly the evidence the hold was placed over.

Assembling a set that is not one — no owner, two owners, or a reuser pointing elsewhere — is refused outright rather than evaluated. Deciding "may this be deleted" over a set that was put together wrong is the one mistake in this module with no undo.

The rule is conservative on purpose. Keeping a blob one cycle longer than necessary costs storage; deleting one that still serves somebody cannot be undone.

The tombstone works the other way round and stays per record. A record that lost its references and outlived its retention gets its own purgedAt, purgeReason and purgedBy even while the blob survives for the other members. What the set decides is the fate of the bytes, never of the rows.


The decisions behind this page are in ADR 0010; the contract is in the specification.

Released under the MIT License.