Skip to content

Watching the scan: metrics, health and the two alerts ​

A clamd outage produces no error anybody can see. Uploads keep answering 201, reads keep answering FILE_NOT_AVAILABLE_CONFLICT because the file is not available yet, and the only thing that changes is that nothing ever becomes available. Nobody is paged by a system behaving exactly as designed, so everything on this page is about watching for absence.

The one reading ​

There is a single producer of scanner health in the module, and everything reports off it:

ConsumerWhat it does with the reading
ScannerHealthMonitorJob, every five minutesraises the two alerts and records the gauges
The application health checkreports Healthy / Degraded on the same facts
GET /files/scanner-statusthe HTTP surface of the same reading

One PING, one VERSION and one queue count, taken at one instant. Two implementations would be two opinions on the same question, and they would disagree on the afternoon it mattered — the route reporting healthy while the monitor paged somebody, or the reverse. The 48-hour rule lives with the health type itself for the same reason.

The queue is counted across every tenant, because the queue is a property of the engine and not of whoever is asking. A tenant shown only its own waiting files would read "nothing is queued" off a daemon five thousand files deep.

The engine's answer is kept for five seconds (ScannerHealthCache, a singleton under the probe), and all three consumers share it. The status route is made to be polled — every open screen asks every few seconds — and without it each poll would be a PING and a VERSION. With the daemon down, it would also be a connection timeout per poll. An answer saying the engine is down is kept too, for that reason. Only one reading runs at a time; callers that arrive while it runs wait and take its answer. The queue count is not kept: it is an indexed read, and a depth frozen for a few seconds would make a moving queue look stuck. Ages are still measured at the moment of each reading, because what is kept is the database's date, not its age.

GET /files/scanner-status ​

read:files. The only route in the module that can answer 503, and the only producer of FILE_SCAN_UNAVAILABLE.

EngineAnswer
answered200 with the body below — a stale database included
did not answer503 FILE_SCAN_UNAVAILABLE, status taken from the _UNAVAILABLE suffix
FieldMeaning
enginewhat VERSION reports, e.g. ClamAV 1.4.6
databaseVersionthe signature database it carries
databaseAgeHoursage of that database, one decimal place; null when the date could not be read — unknown is not zero
databaseStaleover 48 hours, or unknown. The same rule as Alert 2, judged on the exact age rather than the rounded one
queueDepthfiles awaiting a verdict (Scanning or Failed), across every tenant
queueOldestItemAgeSecondshow long the oldest of them has waited, one decimal place; null when nothing waits

A stale database is a 200 and not a 503. That engine is up and scanning, and calling it down would hide the one thing worth saying: it is up and out of date.

The 503 carries a fixed sentence of its own. The engine adapter's failure message can name the socket, host or port it dialled, and that belongs in the monitor's log line, not on a tenant's screen. The route also changes nothing else: reservations keep being accepted, and a file waiting for a verdict keeps answering FILE_NOT_AVAILABLE_CONFLICT.

The metrics ​

Register the meter GrydFiles.Scanner with the exporter.

MetricUnitWhat it is for
gryd_files.scan.duration_per_megabytes/MBSizes the clamd replicas. Seconds of scan per megabyte, from the attempt's startedAt and finishedAt.
gryd_files.scan.durationsWall time of one pass.
gryd_files.scan.bytesByBytes handed to the engine on one pass.
gryd_files.scan.queue_depth{file}Files in Scanning or Failed — everything with no verdict.
gryd_files.scan.queue_ageminHow long the longest-waiting file has waited.
gryd_files.scan.database_agehAge of the signature database.
gryd_files.scan.outcome.clean{pass}Passes that ended clean.
gryd_files.scan.outcome.infected{pass}Detections. The only outcome that is a verdict against a file.
gryd_files.scan.outcome.error{pass}Infrastructure failure — daemon down, bucket unreadable.
gryd_files.scan.outcome.timeout{pass}The engine did not answer in time.
gryd_files.scan.outcome.sizelimitexceeded{pass}StreamMaxLength below the profile ceiling — a configuration error.
gryd_files.scan.queue_alerts{alert}Alert 1 fired.
gryd_files.scan.stale_database_alerts{alert}Alert 2 fired.
gryd_files.scan.engine_unavailable_alerts{alert}The engine did not answer at all.
gryd_files.scan.retries_exhausted{file}A file used its whole retry budget and stays unreleased.

Three details worth knowing before building a dashboard on these.

The outcome counters are separate instruments, not one instrument tagged by outcome. A tag is the idiomatic shape and it is also the one a collector view can drop, and dropping that tag collapses the detection count into the total. Separate instruments cannot be aggregated by accident.

Do not read seconds-per-MB off the histogram's mean for capacity work. The ratio lies at the small end: a one-byte file that took five milliseconds is five thousand seconds per megabyte, and enough of them drag the distribution somewhere no sizing decision should come from. The raw duration and the raw byte count are emitted beside it so the honest aggregate — sum(duration) / sum(bytes) — is available.

The outcome counters mirror the FileScanAttempt table exactly, including passes where the object could not be read and the engine was never reached. During an incident somebody compares the dashboard against the rows, and a counter incremented on some paths but not others turns that comparison into a second investigation.

Alert 1 — the queue is over its limits ​

Fires when either limit is crossed:

KeyDefaultCrosses when
GrydFiles:ScanQueueAlertThreshold100more than this many files await a verdict
GrydFiles:ScanQueueAlertMaxAge00:30:00the longest-waiting file has waited longer than this

Either, not both, because the two describe different failures. Depth catches the engine that cannot keep up with the load. Age catches the engine that has stopped — and an engine that stops on a quiet Sunday never reaches the depth at all: there are four files in the queue, there will only ever be four, and every one of them is unusable.

The age is measured from the reservation, which is the only instant that is both stored as a column and stable across retries. Timestamps a retry touches all move forward, so an age built on them resets every time the file is tried again — and a file being retried forever is precisely the one that must not look young. The cost is that the age also contains however long the client took to upload; it therefore overstates the wait and never understates it.

Thirty minutes because the retry budget is about twenty-one (one pass, then waits of 1, 5 and 15 minutes). A file still waiting past that has outlived every attempt the module makes on its own.

Runbook ​

  1. Is the daemon answering? echo PING | nc <clamd-host> 3310 should say PONG. If it does not, this is really the engine-unavailable alert; go to the runbook below.
  2. Is it depth or is it age? Read gryd_files.scan.queue_depth and gryd_files.scan.queue_age together. Depth high and age low is a burst, and it drains. Age high is an outage, whatever the depth says.
  3. Depth, draining: confirm gryd_files.scan.duration_per_megabyte is where it usually is, then add clamd replicas. The per-megabyte figure over the arrival rate is the sizing.
  4. Depth, not draining, daemon up: check the daemon's own limits — a container at its memory limit accepts connections and scans nothing. clamd needs about 3 GiB and a concurrent database reload doubles that.
  5. Age high: find the oldest waiting record and read its FileScanAttempt rows. Repeated Error or Timeout is infrastructure; SizeLimitExceeded is StreamMaxLength under the profile's ceiling, which is a configuration fix and not a scaling one.
  6. Files that exhausted the retry budget stay in Failed and are not tried again. After the cause is fixed they need to be re-queued deliberately. That is not a bug to work around — see below.

Alert 2 — the signature database is over 48 hours old ​

This is the one that goes unnoticed, and it is the reason this alert is written down separately.

Everything looks fine while it is happening. The daemon answers PING with PONG. Every liveness probe stays green. Files arrive, are scanned, come back OK and are released. And the engine has quietly stopped recognising anything catalogued since the database stopped moving, so the releases are the failure — not the queue, not an error rate, nothing anybody is watching.

Only the database date shows it, read from the third field of the engine's VERSION reply — the same data written into scanDatabaseVersion on every scanned file.

freshclam updates daily, so a database over two days old means a whole update cycle was missed rather than one running late. An unknown age counts as stale: not having established that the database is fresh is not the same as having established that it is.

Runbook ​

  1. Read the age, from gryd_files.scan.database_age or from the health check's databaseAge.
  2. Is freshclam running? It is the sidecar writing into the shared signature volume. Its log says whether it is being rate-limited by the CDN, which is the usual cause and looks like nothing else.
  3. Is the volume shared and writable? A freshclam updating a volume the daemon does not read produces exactly this: fresh files on disk, an old database in memory.
  4. Did the daemon reload? clamd picks up a new database on its own with concurrent reload enabled; RELOAD over the socket forces it. Reloading doubles memory for the duration, which is the reason the container is sized at 4 GiB rather than 2.
  5. After the fix, verify by VERSION, not by PING. PONG was never the thing that was wrong.
  6. Files released during the window were scanned by an out-of-date engine. Re-scanning them is a deliberate operation, and deciding whether it is warranted is what the alert is for.

There is no degraded mode ​

Nothing in this module releases a file without a verdict. Not a flag, not a permission, not admin:system, not a timeout, not an operator in a hurry (RN-GF-05).

With the daemon out of the air, the queue grows and that is the correct behaviour. Files pile up in Scanning and Failed, unusable, until an engine gives them a verdict. The alerts on this page exist so that somebody knows it is happening — not so that anything can be released in the meantime.

The monitor itself alerts and changes nothing. It moves no file, releases no file and overrides no verdict. Neither does the health check: an unreachable engine reports Degraded rather than Unhealthy by default, because the API can still serve everything that already has a verdict, and taking the process out of rotation would turn a scanning outage into a total one.


The decisions behind this page are in ADR 0010; the contract is in the specification.

Released under the MIT License.