Appearance
The antivirus service
ClamAV runs as a service of its own, with 1..N replicas, reached over a socket. No API instance embeds the engine. Why ClamAV, and what was refused in its place, is in ADR 0010; the contract is §9 of the specification.
The antivirus on one page
- One read, three consumers. The scan job opens the object once and streams it to
clamdwithINSTREAM— nothing is written to disk. The same bytes feed thesha256and the content-type detector on the way, so the hash, the type and the verdict are about exactly the same content. - Four answers, one verdict.
stream: OK,<signature> FOUND,ERRORandINSTREAM size limit exceeded. OnlyFOUNDsays something about the file; the other failures say something about the infrastructure and never release it — see The scan contract. - The limits, decided (D13).
StreamMaxLength/MaxFileSize64M,MaxScanSize400M,MaxRecursion16,MaxFiles10000, andAlertExceedsMaxon. clamd's default of 25M is exactly the largest profile in use and leaves no room for theINSTREAMframing — see clamd.conf. - Retries (D14). An engine error is retried after 1, 5 and 15 minutes — the first pass and three more, 120 seconds each — and the file stays unreleased throughout; see Quarantine, retries and the two alerts.
- The official probe is
GET /files/scanner-status, backed by the same reading as the health check and the two alerts (the queue, and a signature database older than 48 hours) — see Watching the scan. - A false positive has a path, and it is not a button. The file stays in quarantine; the answer is the corrected signature and a new upload.
Why a daemon, and why not in the API
clamscan, the command-line scanner, reloads the entire signature database on every run — fine on a laptop, impossible on a server. clamd keeps the signatures in memory and answers over a socket.
That memory is also the reason the daemon is not a library call. The database is expensive enough that replicating it per process is out of the question: every API replica would carry its own copy for no gain in throughput. One shared clamd serves every tenant, because it sees bytes, not context — isolating it per tenant would multiply 4 GiB by N and buy nothing.
Memory is the resource that matters
| Request | 3 GiB minimum |
| Limit | 4 GiB |
The database sits in memory, and ConcurrentDatabaseReload — which is on, so scanning does not stop during an update — holds two copies of it while it swaps. That is what turns 3 GiB from "enough" into "enough until it reloads".
Short of memory, the kernel kills clamd and the container stays up. This is the nastiest failure the module has: the orchestrator sees a healthy service while the scan queue grows and not a single file is released.
Liveness and readiness: PING → PONG
The probe sends PING on the socket and requires PONG. It is the only signal that separates "the container is up" from "the engine answers" — the exact distinction the OOM kill erases. In the official image, clamdcheck.sh is literally echo PING | nc localhost 3310 compared against PONG, so the health check and the manual check are the same thing.
Readiness needs nothing extra: clamd binds its sockets only after the database is loaded, so a replica that answers PONG has a database, and one that does not answer takes no traffic.
Reporting and reacting are different jobs. The probe reports; restarting the replica or pulling it out of rotation belongs to the orchestrator. docker compose marks the container unhealthy and leaves it running — enough to see the failure on a developer's machine, and not a self-healing deployment. Scaling is the same kind of statement: the application dials an address, never a replica, so 1..N is a decision made where the service is deployed and nothing in the API changes with it.
The signature database lives in a volume
- The image is the
_basevariant, which ships no signatures. The database comes from the volume, never from an image layer. freshclamupdates it in place, daily by default (FRESHCLAM_CHECKS).- The volume persists across restarts and recreations. Rebuilding the container and downloading the whole database on every deploy is the fastest way to get blocked by the project's free CDN, and it is explicitly out of bounds.
VERSION is how the engine and the database version are read back. Those two values are what the module stores in scanEngine and scanDatabaseVersion on each scanned file, and the age of the database is what the 48-hour alert and GET /files/scanner-status are computed from — the alert that matters most, because everything looks fine while the engine quietly goes blind to anything catalogued since the last update.
The configuration is rendered, not written
clamd.conf is produced from the GrydFiles:Scanner section — see clamd.conf for every parameter and why it holds the value it holds. The file is mounted read-only, and nothing inside the container may rewrite the limits the application validated the profile ceilings against.
Bucket policy (D5)
Two things belong to the bucket or the CDN, not to the module:
X-Content-Type-Options: nosniff cannot be expressed by a presigned URL. The set of headers a presigned URL may override is closed, and none of the three providers admits this one. It therefore left the download contract and became a recommendation on the bucket or the CDN in front of it — where it applies to every response, including the ones the module does not issue. Where to set it on each CDN, and what defends the download without it, is in Download.
The _quarantine/ prefix must be closed. Infected content is kept as evidence, not as something retrievable: no presigned URL is ever issued for an object under that prefix, and the bucket policy should deny reads on it outside the credentials the module itself uses. A prefix that is merely "never linked to" is not a policy.
Development
tools/GrydFiles.DevHost brings up PostgreSQL, MinIO and clamd together, so the whole flow runs on a developer's machine with no cloud. Its README.md has the commands; up.sh renders the configuration and then calls docker compose up, in that order.
The API is not one of those containers — it runs from the developer's machine against them, the same way tools/GrydAuth.AdminDevHost works.