Appearance
Quarantine, retries and the two alerts
What happens to a file the engine condemned, to a file the engine could not read, and to an engine that stopped working without anything looking broken.
Detected content is moved, not deleted
A file the engine matches a signature on is moved to _quarantine/{tenantId | _global}/{fileId}.{ext} and the record is repointed at the new key. Deleting it would destroy the evidence of the incident, and the incident is the only reason anybody will ever want the file again.
The move is a write to the new key and then a delete of the old one, in that order. Written first means a failure leaves the content where it was rather than nowhere.
If the move fails, the record is marked Infected anyway. RN-GF-04 is the stronger obligation: a file the engine condemned has to be unusable from that instant. A failed move leaves an object in the wrong prefix, which the consistency sweep can fix and which no route hands out either way, because every route refuses on the record's status.
The quarantine is written to the audit trail with the signature name, the engine and the database version. The row alone carries a status, and a status does not name a signature.
Nothing under the prefix is ever signed
No presigned URL is issued for an object under _quarantine/ — not for download, not for upload, by no route and with no permission, admin:system included.
That is a claim about every route there will ever be, including the ones not written yet, so it is not kept by each route remembering to check. The module asks object storage to sign something in exactly one place, FilePresignedUrlIssuer, and an architecture test fails the moment a second one is cut. Asking it for a quarantined key throws rather than returning a failure: every route reaches its own refusal from the record's status long before a key is resolved, so getting that far means a code path found a way to ask for a link to quarantined malware.
Infected is terminal
The only state reachable from Infected is Purged, by the retention job, after QuarantineRetentionDays (365). There is no manual release and no button — the permission to override the antivirus is the one an attacker most wants, so it does not exist. None of the module's four permissions releases a file without a verdict, and the absence is the decision.
A false positive has a path and it is not a button: report the signature, wait for the database to be corrected, upload again — or convert the file.
Retries: 1, 5 and 15 minutes
An infrastructure outcome — the daemon down, a timeout, a bucket that would not open — moves the record to Failed, which is not a verdict about the file. The job books another pass after the next configured delay:
| Pass | Booked after |
|---|---|
| 1 | the upload is confirmed |
| 2 | 1 minute |
| 3 | 5 minutes |
| 4 | 15 minutes |
Four passes across about twenty-one minutes, which absorbs a container restart and a signature database reload without holding the file for the whole afternoon. The delays come from GrydFiles:Scanner:RetryDelays and are never written into the job.
When the budget runs out the record stays in Failed and the operator is alerted. There is no fourth behaviour: no degraded mode, no timeout that releases, no emergency flag. A file nobody could scan is a file nobody may use.
Each pass appends its own FileScanAttempt row — engine, database version, start, finish, outcome and the engine's literal words. The table is append-only; a row is closed once and never rewritten.
An attempt a dead worker left open is closed as an Error before the next pass starts. Without that recovery one crashed worker would wedge the file in Scanning forever, because the aggregate refuses a second open attempt — correctly, since two open rows would mean two readers believe they own the same bytes.
The two alerts
A clamd outage raises no error anywhere a user can see. Uploads keep answering 201, reads keep answering FILE_NOT_AVAILABLE_CONFLICT because the file is not available yet, and the only thing that changes is that nothing ever becomes available. Nobody is paged by a system behaving exactly as designed, so the alerts watch for absence.
ScannerHealthMonitorJob runs every five minutes — inside one retry budget — and raises two signals.
The queue. Files in Scanning or Failed above GrydFiles:ScanQueueAlertThreshold (100). Files that already have a verdict are not counted: the alert is about backlog, not volume. A threshold of zero is refused at startup, because an alarm that is always on is one nobody reads on the afternoon the engine really stops.
The database age. A signature database older than 48 hours. This is the signal that goes unnoticed: the daemon answers PING with PONG the whole time, every probe stays green, files keep being released — by an engine that has quietly stopped recognising anything catalogued this week. freshclam updates daily, so two days means a whole cycle was missed. The age comes from the third field of the VERSION reply, and an age that cannot be established counts as stale — not knowing how old a database is has never established that it is fresh.
The monitor alerts and does nothing else. It changes no record, releases no file and overrides no verdict.
No route here answers 503
FILE_SCAN_UNAVAILABLE is not produced by any of this. Nothing in the upload or read path talks to the engine synchronously — the scan is asynchronous and the upload answers before any contact with the daemon — so with clamd down, uploads and reads keep answering normally. The single producer of that code is GET /files/scanner-status.
Use of a file in Pending, Scanning or Failed is FILE_NOT_AVAILABLE_CONFLICT (409), and use of an infected one is FILE_INFECTED_CONFLICT (409) — a separate code on purpose, because "not ready" and "never" are different answers to the caller.
The decisions behind this page are in ADR 0010; the contract is in the specification.