Appearance
Watching the scan: metrics, health and the two alerts
A clamd outage produces no error anybody can see. Uploads keep answering 201, reads keep answering FILE_NOT_AVAILABLE_CONFLICT because the file is not available yet, and the only thing that changes is that nothing ever becomes available. Nobody is paged by a system behaving exactly as designed, so everything on this page is about watching for absence.
The one reading
There is a single producer of scanner health in the module, and everything reports off it:
| Consumer | What it does with the reading |
|---|---|
ScannerHealthMonitorJob, every five minutes | raises the two alerts and records the gauges |
| The application health check | reports Healthy / Degraded on the same facts |
GET /files/scanner-status | the HTTP surface of the same reading |
One PING, one VERSION and one queue count, taken at one instant. Two implementations would be two opinions on the same question, and they would disagree on the afternoon it mattered — the route reporting healthy while the monitor paged somebody, or the reverse. The 48-hour rule lives with the health type itself for the same reason.
The queue is counted across every tenant, because the queue is a property of the engine and not of whoever is asking. A tenant shown only its own waiting files would read "nothing is queued" off a daemon five thousand files deep.
The engine's answer is kept for five seconds (ScannerHealthCache, a singleton under the probe), and all three consumers share it. The status route is made to be polled — every open screen asks every few seconds — and without it each poll would be a PING and a VERSION. With the daemon down, it would also be a connection timeout per poll. An answer saying the engine is down is kept too, for that reason. Only one reading runs at a time; callers that arrive while it runs wait and take its answer. The queue count is not kept: it is an indexed read, and a depth frozen for a few seconds would make a moving queue look stuck. Ages are still measured at the moment of each reading, because what is kept is the database's date, not its age.
GET /files/scanner-status
read:files. The only route in the module that can answer 503, and the only producer of FILE_SCAN_UNAVAILABLE.
| Engine | Answer |
|---|---|
| answered | 200 with the body below — a stale database included |
| did not answer | 503 FILE_SCAN_UNAVAILABLE, status taken from the _UNAVAILABLE suffix |
| Field | Meaning |
|---|---|
engine | what VERSION reports, e.g. ClamAV 1.4.6 |
databaseVersion | the signature database it carries |
databaseAgeHours | age of that database, one decimal place; null when the date could not be read — unknown is not zero |
databaseStale | over 48 hours, or unknown. The same rule as Alert 2, judged on the exact age rather than the rounded one |
queueDepth | files awaiting a verdict (Scanning or Failed), across every tenant |
queueOldestItemAgeSeconds | how long the oldest of them has waited, one decimal place; null when nothing waits |
A stale database is a 200 and not a 503. That engine is up and scanning, and calling it down would hide the one thing worth saying: it is up and out of date.
The 503 carries a fixed sentence of its own. The engine adapter's failure message can name the socket, host or port it dialled, and that belongs in the monitor's log line, not on a tenant's screen. The route also changes nothing else: reservations keep being accepted, and a file waiting for a verdict keeps answering FILE_NOT_AVAILABLE_CONFLICT.
The metrics
Register the meter GrydFiles.Scanner with the exporter.
| Metric | Unit | What it is for |
|---|---|---|
gryd_files.scan.duration_per_megabyte | s/MB | Sizes the clamd replicas. Seconds of scan per megabyte, from the attempt's startedAt and finishedAt. |
gryd_files.scan.duration | s | Wall time of one pass. |
gryd_files.scan.bytes | By | Bytes handed to the engine on one pass. |
gryd_files.scan.queue_depth | {file} | Files in Scanning or Failed — everything with no verdict. |
gryd_files.scan.queue_age | min | How long the longest-waiting file has waited. |
gryd_files.scan.database_age | h | Age of the signature database. |
gryd_files.scan.outcome.clean | {pass} | Passes that ended clean. |
gryd_files.scan.outcome.infected | {pass} | Detections. The only outcome that is a verdict against a file. |
gryd_files.scan.outcome.error | {pass} | Infrastructure failure — daemon down, bucket unreadable. |
gryd_files.scan.outcome.timeout | {pass} | The engine did not answer in time. |
gryd_files.scan.outcome.sizelimitexceeded | {pass} | StreamMaxLength below the profile ceiling — a configuration error. |
gryd_files.scan.queue_alerts | {alert} | Alert 1 fired. |
gryd_files.scan.stale_database_alerts | {alert} | Alert 2 fired. |
gryd_files.scan.engine_unavailable_alerts | {alert} | The engine did not answer at all. |
gryd_files.scan.retries_exhausted | {file} | A file used its whole retry budget and stays unreleased. |
Three details worth knowing before building a dashboard on these.
The outcome counters are separate instruments, not one instrument tagged by outcome. A tag is the idiomatic shape and it is also the one a collector view can drop, and dropping that tag collapses the detection count into the total. Separate instruments cannot be aggregated by accident.
Do not read seconds-per-MB off the histogram's mean for capacity work. The ratio lies at the small end: a one-byte file that took five milliseconds is five thousand seconds per megabyte, and enough of them drag the distribution somewhere no sizing decision should come from. The raw duration and the raw byte count are emitted beside it so the honest aggregate — sum(duration) / sum(bytes) — is available.
The outcome counters mirror the FileScanAttempt table exactly, including passes where the object could not be read and the engine was never reached. During an incident somebody compares the dashboard against the rows, and a counter incremented on some paths but not others turns that comparison into a second investigation.
Alert 1 — the queue is over its limits
Fires when either limit is crossed:
| Key | Default | Crosses when |
|---|---|---|
GrydFiles:ScanQueueAlertThreshold | 100 | more than this many files await a verdict |
GrydFiles:ScanQueueAlertMaxAge | 00:30:00 | the longest-waiting file has waited longer than this |
Either, not both, because the two describe different failures. Depth catches the engine that cannot keep up with the load. Age catches the engine that has stopped — and an engine that stops on a quiet Sunday never reaches the depth at all: there are four files in the queue, there will only ever be four, and every one of them is unusable.
The age is measured from the reservation, which is the only instant that is both stored as a column and stable across retries. Timestamps a retry touches all move forward, so an age built on them resets every time the file is tried again — and a file being retried forever is precisely the one that must not look young. The cost is that the age also contains however long the client took to upload; it therefore overstates the wait and never understates it.
Thirty minutes because the retry budget is about twenty-one (one pass, then waits of 1, 5 and 15 minutes). A file still waiting past that has outlived every attempt the module makes on its own.
Runbook
- Is the daemon answering?
echo PING | nc <clamd-host> 3310should sayPONG. If it does not, this is really the engine-unavailable alert; go to the runbook below. - Is it depth or is it age? Read
gryd_files.scan.queue_depthandgryd_files.scan.queue_agetogether. Depth high and age low is a burst, and it drains. Age high is an outage, whatever the depth says. - Depth, draining: confirm
gryd_files.scan.duration_per_megabyteis where it usually is, then addclamdreplicas. The per-megabyte figure over the arrival rate is the sizing. - Depth, not draining, daemon up: check the daemon's own limits — a container at its memory limit accepts connections and scans nothing.
clamdneeds about 3 GiB and a concurrent database reload doubles that. - Age high: find the oldest waiting record and read its
FileScanAttemptrows. RepeatedErrororTimeoutis infrastructure;SizeLimitExceededisStreamMaxLengthunder the profile's ceiling, which is a configuration fix and not a scaling one. - Files that exhausted the retry budget stay in
Failedand are not tried again. After the cause is fixed they need to be re-queued deliberately. That is not a bug to work around — see below.
Alert 2 — the signature database is over 48 hours old
This is the one that goes unnoticed, and it is the reason this alert is written down separately.
Everything looks fine while it is happening. The daemon answers PING with PONG. Every liveness probe stays green. Files arrive, are scanned, come back OK and are released. And the engine has quietly stopped recognising anything catalogued since the database stopped moving, so the releases are the failure — not the queue, not an error rate, nothing anybody is watching.
Only the database date shows it, read from the third field of the engine's VERSION reply — the same data written into scanDatabaseVersion on every scanned file.
freshclam updates daily, so a database over two days old means a whole update cycle was missed rather than one running late. An unknown age counts as stale: not having established that the database is fresh is not the same as having established that it is.
Runbook
- Read the age, from
gryd_files.scan.database_ageor from the health check'sdatabaseAge. - Is
freshclamrunning? It is the sidecar writing into the shared signature volume. Its log says whether it is being rate-limited by the CDN, which is the usual cause and looks like nothing else. - Is the volume shared and writable? A
freshclamupdating a volume the daemon does not read produces exactly this: fresh files on disk, an old database in memory. - Did the daemon reload?
clamdpicks up a new database on its own with concurrent reload enabled;RELOADover the socket forces it. Reloading doubles memory for the duration, which is the reason the container is sized at 4 GiB rather than 2. - After the fix, verify by
VERSION, not byPING.PONGwas never the thing that was wrong. - Files released during the window were scanned by an out-of-date engine. Re-scanning them is a deliberate operation, and deciding whether it is warranted is what the alert is for.
There is no degraded mode
Nothing in this module releases a file without a verdict. Not a flag, not a permission, not admin:system, not a timeout, not an operator in a hurry (RN-GF-05).
With the daemon out of the air, the queue grows and that is the correct behaviour. Files pile up in Scanning and Failed, unusable, until an engine gives them a verdict. The alerts on this page exist so that somebody knows it is happening — not so that anything can be released in the meantime.
The monitor itself alerts and changes nothing. It moves no file, releases no file and overrides no verdict. Neither does the health check: an unreachable engine reports Degraded rather than Unhealthy by default, because the API can still serve everything that already has a verdict, and taking the process out of rotation would turn a scanning outage into a total one.
The decisions behind this page are in ADR 0010; the contract is in the specification.