5.1 KiB
Fileserver runbooks
Each runbook names the exact metric that fires it and the exact command that resolves it. A runbook whose trigger is "someone noticed" is not actionable, so every one below starts from a signal.
Storage full
Signal — fileserver.quota{result="rejected"} rising, or 507 responses appearing.
Storage capacity is exhausted or the high-water guard tripped. Uploads are rejected before any bytes are written, so nothing is corrupt; the system is refusing work it cannot complete.
curl -s $ADMIN/internal/fileserver/storage-health | jq '.usedFraction, .usableBytes'
curl -s -X POST "$ADMIN/internal/fileserver/uploads:cleanup?maxItems=500&maxBytes=10737418240"
curl -s "$ADMIN/internal/fileserver/orphans?limit=200" | jq '[.[].sizeBytes] | add'
Drain the cleanup backlog first — it reclaims space the system already knows is dead. Only then consider an orphan reconcile, and start with a dry run.
Orphan growth
Signal — fileserver.cleanup{result="skipped"} climbing, or the orphan scan returning more
objects each run.
Physical objects exist with no metadata record pointing at them. This is not immediately dangerous — nothing serves them — but it consumes capacity indefinitely.
# Always look first. A reconcile without dryRun=false is a plan, not an action.
curl -s -X POST "$ADMIN/internal/fileserver/orphans:reconcile" \
-H 'content-type: application/json' -d '{"limit":100}' | jq '.candidates'
# Apply only the fingerprints you were just shown.
curl -s -X POST "$ADMIN/internal/fileserver/orphans:reconcile" \
-H 'content-type: application/json' \
-d '{"dryRun":false,"limit":100,"maxBytes":1073741824,
"expectedFingerprints":["<from the dry run>"],"reasonCode":"ORPHAN_GROWTH_RUNBOOK"}'
Echoing the fingerprints is the safety property: an object that changed between the scan and the apply is skipped rather than deleted.
Verification backlog
Signal — fileserver.verification.queue{age_bucket="old"} non-zero, or files sitting in
VERIFYING.
A verifier is slow or unavailable. Files stay non-public, which is the correct failure direction: a
RETRY verdict never becomes an ACCEPT.
curl -s $ADMIN/internal/fileserver/capabilities | jq '.storageType'
# Once the verifier is healthy, quarantined files can be re-examined individually.
curl -s -X POST "$ADMIN/internal/fileserver/files/$FILE_ID:reverify"
Do not clear the backlog by disabling verification. A file that reached READY without an accepting verdict cannot be distinguished later from one that was verified.
NFS ambiguity
Signal — problem documents carrying "ambiguous": true, or
fileserver.transfer.interruption{reason="stale_handle"}.
An operation's outcome could not be determined: the response was lost after the write or rename may have landed. These are never retried automatically.
# The recovery queue holds the files awaiting a decision.
curl -s "$ADMIN/internal/fileserver/uploads/incomplete?limit=100" | jq
Reconciliation compares the physical size and digest against the record and only confirms READY when all four of key, size, digest, and version agree. Anything short of that is reported, never guessed.
PVC remount
Signal — startup failure naming "atomic move", "same file store", or "not writable".
The volume was remounted somewhere the probe can no longer prove a required capability. The application refuses traffic rather than serving from storage it cannot publish to atomically.
kubectl apply -f infra/fileserver/kubernetes/pvc-certification-job.yaml
kubectl logs job/fileserver-pvc-certification
Compare the printed tuple with the certified one in docs/fileserver/storage-certification.md. A
mismatch in CSI driver, StorageClass, access mode, or mount options is the cause; the certification
does not carry across it.
Nginx delegation failure
Signal — fileserver.download.delegation{delegated="true"} with client-visible 404s.
The internal location is misconfigured, so the proxy cannot resolve the redirect it was handed.
# The internal prefix must resolve to the content root and must be marked `internal`.
grep -A5 '__files' infra/fileserver/nginx/nginx.conf
curl -s $ADMIN/internal/fileserver/capabilities | jq '.capabilities.delegatedDownload'
Turning delegation off is a safe immediate mitigation: the application serves the transfer itself, slower but correct.
app.fileserver-platform.nginx.enabled=false
Cleanup backlog
Signal — fileserver.cleanup{result="deferred"} rising, or reclaimed bytes flat while deletes
continue.
Items are being deferred faster than they drain. The usual cause is an active writer lease still holding staging objects, which is correct behaviour, not a fault.
curl -s "$ADMIN/internal/fileserver/uploads/incomplete?limit=100" \
| jq '[.[] | select(.leaseUntil != null)] | length'
curl -s -X POST "$ADMIN/internal/fileserver/uploads:cleanup?maxItems=500&maxBytes=10737418240"
If the deferrals are all ACTIVE_WRITER_LEASE, the backlog resolves itself as those uploads expire.
Never delete staging content to clear a backlog: an upload that is mid-flight will corrupt.