Operations

The green check that hides a broken result

In a mature estate the dangerous failures are not crashes. They are the runs where exit status is zero, the log looks clean, the file exists and the contents are wrong. A taxonomy of how that happens, and the checking discipline that catches it.

24 min readVerificationMonitoringShellAutomationTesting

A crash is a gift. It has a timestamp, a stack, a non-zero status and somebody’s attention. Crashes get fixed because they are impossible to ignore.

The failures that survive in a well-run estate are the other kind. The job exits zero. The log ends with “completed”. The output file is there, with the right name, the right owner and a plausible size. Nothing in the monitoring turns amber. And the contents are wrong — empty where they should be populated, defaulted where you chose something else, stale where you expected fresh.

This class has no single cause. It has about seven, and they are mechanical enough to be listed. What follows is the taxonomy, then the checking discipline that actually catches them, which is narrower and more annoying than the one most of us write by default.

The organizing claim: a check that asserts on status rather than content is not a check. It is a record that something ran.

The taxonomy

Class What the operator sees Why the probe passes The assertion that catches it
Error printed, status zero An error line in a log nobody reads Probe tests $? only Compare stdout to an expected value
Plugin missing, schema intact Correct headers, blank column Header check passes Assert a populated value for a case that must produce one
List parsed as one string Silent fallback to a default Config parses, service starts Read the effective value back from the running process
Zero-byte index fetched “Nothing to update” Transfer tool exits zero Assert a minimum size and a known marker in the content
Symlink to a path on another host Module loads, variable empty module load exits zero Resolve the path and test the binary runs
Environment change inside a pipe A working tool reported broken Probe ran in a subshell Set environment, then test in the same shell
Suite that has never failed An unbroken run of green Nothing has been injected Prove each check fails on a deliberate fault

Each row below gets its own section, because the mechanism matters more than the symptom.

Class one: the error that exits zero

Exit status is a convention, not a guarantee. Plenty of widely used tools print a diagnostic and return success, either by design or by oversight.

jq is the honest case. A missing key is not an error in its data model — it is null:

$ echo '{"a":1}' | jq -r '.b'
null
$ echo $?
0

A probe written as “run it, check the status, check that stdout is non-empty” scores that as a pass. The string null is four bytes of non-empty output. Downstream, a config generator writes endpoint = null into a file, the service parses it, and you get whatever the service does with a nonsense endpoint.

grep -c is the inverse and catches people the other way:

$ printf '' | grep -c .
0
$ echo $?
1

It prints something (0) and returns failure. A probe keying on status calls that a failure when the count is the answer you wanted; a probe keying on “is there output” calls it a pass. Neither is reading the number.

The most expensive member of this family is HTTP. curl treats an HTTP error response as a successful transfer, because from its point of view it was one:

$ curl -s -o index.json -w 'http=%{http_code}\n' https://example.com/api/index.json
http=404
$ echo $?
0
$ ls -l index.json
-rw-r--r-- 1 svc svc 559 Apr 22 09:14 index.json

The file exists. It is 559 bytes. It is an HTML error page. Every “the file was created and is non-empty” check passes. curl --fail suppresses the body on a 4xx or 5xx and exits non-zero, which is what you want in a script; --fail-with-body (curl 7.76 and later) does the same but keeps the body if you need it for the log. Neither is the default, and the default is what ends up in automation written in a hurry.

One caveat, because it bites checks that are too specific. The documented exit code for --fail is 22, and over HTTP/1.1 that is what you get. Over HTTP/2 the same 404 can surface as 56 instead, because curl tears the stream down and reports a receive error rather than an HTTP one:

$ curl -s --fail -o /dev/null https://example.com/api/index.json ; echo $?
56
$ curl -s --http1.1 --fail -o /dev/null https://example.com/api/index.json ; echo $?
22

That pair is from curl 8.7.1, and I would treat the specific code as build- and version-dependent rather than a rule — which is exactly the point. Test --fail for non-zero, not for 22. A check written as [ $? -eq 22 ] is itself a member of this article’s taxonomy: it goes green on the day the negotiated protocol, or the packaged curl, changes under it.

The general fix is not to distrust exit codes. It is to stop treating them as the only evidence. Ask for the value, state what you expect, compare.

got=$(jq -r '.endpoint // empty' config.json)
[ "$got" = "https://collector.internal:8443" ] || {
  echo "endpoint mismatch: got [$got]" >&2; exit 1; }

That check can fail for a reason other than the process crashing, which is the only property that makes it a check.

Class two: the plugin that does not load

This is the one that has cost me the most time, because the output looks completely normal.

Many tools declare their output schema separately from the code that populates it. The formatter knows the column is called MaxRSS. Whether anything ever writes a number into MaxRSS depends on a collector, a plugin or an accounting module that may not be configured at all. When it is not, you do not get an error. You get the column, with nothing in it.

Slurm’s accounting is the canonical example. Ask for per-job memory:

$ sacct -j 1847221 --format=JobID,JobName,Elapsed,MaxRSS,AveCPU
JobID           JobName    Elapsed     MaxRSS     AveCPU
------------ ---------- ---------- ---------- ----------
1847221         run.sh     02:14:07
1847221.bat+      batch    02:14:07     8412K   02:11:53
1847221.0        python    02:13:44  41235280K  02:11:40

Now the same query with -X, which restricts output to allocation records:

$ sacct -X -j 1847221 --format=JobID,JobName,Elapsed,MaxRSS,AveCPU
JobID           JobName    Elapsed     MaxRSS     AveCPU
------------ ---------- ---------- ---------- ----------
1847221         run.sh     02:14:07

MaxRSS and AveCPU are gathered per step and stored on step records. The allocation record has the columns and no values. A reporting script written around sacct -X — a reasonable choice, since it gives one row per job — will produce a memory report for the whole fleet in which every single job used no memory. The header row is correct. The CSV parses. The chart renders. It is entirely empty of information and looks like a finding.

The same shape occurs with a metrics exporter whose collector throws. The exporter keeps serving /metrics; the failed collector’s series are simply absent. Any alert of the form rate(thing_total[5m]) > 0 on a series that no longer exists evaluates against an empty vector and never fires. The alert is not firing and the thing it watches is unmonitored, and those two states are indistinguishable on the dashboard. That is what absent() and absent_over_time() are for, and why a rule that watches a metric should usually be paired with a rule that watches for the metric’s disappearance.

The assertion that catches the whole family is the same: pick a case that must produce a populated value, and assert on the value. Not on the header, not on the row count, not on the file size.

# a job you know allocated real memory; --noconvert keeps every value in K,
# so you are not comparing 8412K against 39.33G as strings
rss=$(sacct -n -j "$JOB" --format=MaxRSS -P --noconvert \
      | tr -d 'K ' | grep -E '^[0-9]+$' | sort -n | tail -1)
[ -n "$rss" ] && [ "$rss" -gt 1000 ] || {
  echo "accounting returned no MaxRSS for a job that used memory" >&2; exit 1; }

Two details in that snippet are load-bearing, and both are easy to get wrong in a way that leaves the check passing. --noconvert matters because newer Slurm renders large values in human units — a sort -n over a mixture of 8412K and 39.33G puts the small number last. And the grep -E '^[0-9]+$' matters because step records without a value emit a blank line: strip newlines along with the K (tr -d 'K \n', which is the form that comes naturally) and every row concatenates into one integer.

$ printf '8412K\n41235280K\n\n' | tr -d 'K \n' | sort -n | tail -1
841241235280

That number is larger than any threshold you would set, so the check passes forever, on any input, including no input at all.

The interpretation that is wrong

A worked example, because this is where the trap closes.

A CPU-efficiency report built on sacct produced a fleet average of exactly 1.00. Every user, every partition, every month. The obvious reading is a perfectly tuned estate. The correct reading is that the two quantities being divided are the same quantity.

CPUTimeRAW is allocated CPU time: elapsed seconds multiplied by allocated cores. It is not consumed CPU time. Divide CPUTimeRAW by itself, or by any other allocation-derived figure, and you get 1.0 by construction, for every job that has ever run. The consumed figure is TotalCPU (system plus user), and it lives on step records — so the “one row per job” convenience of -X gives you TotalCPU of zero, and an efficiency of 0.00 for everything. Two different wrong answers, each uniform enough to look like a real property of the fleet.

# wrong: always 1.00 — CPUTime is CPUTimeRAW in a different notation
sacct -X -a -S 2026-03-01 --format=JobID,CPUTimeRAW,CPUTime -P

# workable: consumed over allocated, steps included
sacct -a -S 2026-03-01 --format=JobID,TotalCPU,CPUTimeRAW -P

Mind the units when you divide those two. CPUTimeRAW is an integer count of CPU-seconds; TotalCPU is a formatted duration, [DD-]HH:MM:SS[.mmm]. Divide the string by the integer without parsing it and you will get a number — a small, uniform, entirely fictitious one — which is the same failure this section is about, arrived at one layer further down.

Uniformity is the signal. A real efficiency distribution on a shared cluster is wide and messy, and the mean lands somewhere below half — I would treat roughly 30 to 60 percent as the plausible band for a general-purpose fleet, and I am giving that as a band rather than a figure because it depends heavily on the workload mix and I have no basis for a tighter claim. What I am confident about is the mechanism: if a derived metric is the same for every row, suspect the arithmetic before you believe the estate.

Class three: the list that is one string

Configuration parsers disagree about lists. Some split on whitespace, some on commas, some take the whole line, and some accept either and quietly prefer the wrong one. When a list is read as a single literal, it usually matches nothing, and “matches nothing” is frequently indistinguishable from “not configured” — which means a fallback to a default you did not choose.

The shell version is the one everyone has written:

$ L="mlx5_0 mlx5_1"
$ for p in "$L"; do echo "iter: [$p]"; done
iter: [mlx5_0 mlx5_1]

One iteration, one token containing a space, and any lookup keyed on it fails. Unquoted it splits into two; quoted it does not. Both are correct for some purpose, and the difference is invisible in the log because the loop body ran and exited zero.

The config-file version is more insidious because the file passes validation. OpenSSH’s server config, for most keywords, uses the first value obtained and ignores later ones. sshd -t reports the file is fine. systemctl reload sshd succeeds. Your carefully placed directive at the bottom of the file does nothing at all, because something 200 lines above set it first — often a vendor-shipped include you did not write. The only reliable check is to ask the running daemon what it actually believes:

# run as root; sshd -T reads the host keys
# sshd 9.8 and later ship as a split binary, but -T still lives on sshd itself
$ sudo sshd -T | grep -i '^ciphers'
ciphers chacha20-poly1305@openssh.com,aes128-ctr,aes192-ctr,aes256-ctr

sshd -T prints the effective configuration after all parsing and includes. The file is what you wrote; -T is what is true.

One qualification, because it is the reading people get wrong: bare sshd -T prints the configuration for a connection that matches no Match block. If the directive you care about lives inside a Match Group or Match Address, plain -T will not show it and you will conclude, wrongly, that it never took. Supply the connection you are asking about:

$ sudo sshd -T -C user=svc,host=jump.example.com,addr=203.0.113.10 | grep -i '^permitrootlogin'

The output is the effective value for that connection. There is no single effective value for a config with Match blocks in it, and a check that assumes there is will be right for most users and quietly wrong for the ones the block was written for.

The general form: where a service distinguishes between the file it was given and the configuration it is running, the effective-value output is the only thing worth asserting on, and you must ask it about the specific case you care about. Where a service offers no such output, read the value back through whatever interface it does expose — a status command, an admin socket, an API — and compare it to the value you intended, in writing, in the check.

Class four: the zero-byte index

A package index, a manifest, a rule feed and a firmware catalog all share a property: when they arrive empty, the consuming tool reports that there is nothing to do.

$ curl -s -o /var/cache/feed/index.json -w 'http=%{http_code} size=%{size_download}\n' \
    https://feeds.example.com/v2/index.json
http=200 size=0
$ echo $?
0
$ ls -l /var/cache/feed/index.json
-rw-r--r-- 1 root root 0 Apr 22 03:00 /var/cache/feed/index.json
$ ./apply-feed --index /var/cache/feed/index.json
0 entries loaded, 0 applied, 0 errors

Three zeros and a clean exit. Read it out loud: nothing to apply, so nothing was applied, so no errors. That is a correct description of a broken pipeline.

Note what that transcript is and is not, because the failure modes here do not rank the way intuition ranks them. It is a 200 with an empty body — a half-deployed origin, a proxy serving a truncated cache entry, a signing step that produced nothing. --fail does not save you, because there is no HTTP error to fail on: the transfer genuinely succeeded and what succeeded was nothing.

The counter-intuitive part is which failure is worse for the file on disk. A transport failure is the gentle one — curl exits non-zero and leaves an existing destination untouched, so the last good copy survives:

$ echo "PREVIOUS GOOD CONTENT" > index.json
$ curl -s -o index.json https://feeds.example.invalid/v2/index.json ; echo $?
6
$ cat index.json
PREVIOUS GOOD CONTENT

An HTTP error without --fail is the destructive one. The transfer succeeds, so curl writes, and the error page lands on top of the only good copy you had:

$ curl -s -o index.json https://example.com/api/index.json ; echo $?
0
$ ls -l index.json
-rw-r--r-- 1 root root 559 Apr 22 03:00 index.json

Adding --fail restores the gentle behavior — non-zero status, destination left alone. Which reframes --fail from a tidiness flag into a data-protection one: without it, every scheduled fetch is one origin misconfiguration away from overwriting a good cache with an error page, at three in the morning, exiting zero.

The reason this survives is that “no changes” is also the normal steady state. On most days the honest answer genuinely is “0 applied”. A check that only alarms on errors will never distinguish “we are current” from “we are blind”, and the second state can persist for months.

Two assertions, both cheap:

# 1. minimum plausible size, not merely non-empty
# (GNU coreutils; on BSD/macOS the equivalent is stat -f %z)
sz=$(stat -c %s /var/cache/feed/index.json)
[ "$sz" -ge 4096 ] || { echo "index implausibly small: ${sz}B" >&2; exit 1; }

# 2. a structural marker, not just bytes
jq -e '.entries | length > 0' /var/cache/feed/index.json >/dev/null || {
  echo "index parsed but contains no entries" >&2; exit 1; }

And one more that is worth the trouble on anything time-sensitive: assert on the age of the content, not the age of the file. Any transfer that opens the destination before it knows the result — a 200 that dies half way through the body, an error page written over a good cache — leaves you a fresh mtime over bytes that are stale, partial or wrong. find -mtime is then measuring when something was written, not when it was current. If the payload carries a generation timestamp, compare that instead, and treat the file’s own mtime as evidence of nothing but write activity.

Shared software trees accumulate links. A module file or a wrapper script points at /opt/apps/tool/current, current points at a release directory, and at some point a release directory is moved, or the link is created on a build host against a path that only exists there.

The link itself is not broken in any way the filesystem will tell you about. ln -s will happily create a link to a path that does not exist. ls shows it. And the module that consumes it does this:

prepend-path PATH    $prefix/bin
prepend-path LD_LIBRARY_PATH $prefix/lib64

where $prefix was computed by resolving the link. If the resolution produced an empty string, PATH gains the entry /bin and LD_LIBRARY_PATH gains /lib64. Both are real directories. Neither errors. module load exits zero, module list shows the module loaded, and the binary you wanted is not on the path — so you get the system’s version, silently, at a different version than the one the module claims.

The distinguishing commands:

$ ls -l /opt/apps/tool/current
lrwxrwxrwx 1 root root 33 Feb 11 16:02 /opt/apps/tool/current -> /build/stage/opt/apps/tool/3.11.2

$ readlink -e /opt/apps/tool/current
$ echo $?
1

readlink -e requires every component of the resolved path to exist and prints nothing when it does not. For a check, that is the flag you want.

readlink -f is the one usually reached for, and the difference is worth spelling out because the obvious summary of it is wrong. -f is not “the same but it does not check”. It requires every component but the last to exist. So its behavior splits by which part of the target went missing:

# release directory pruned; the parent is still there → -f is happy, -e is not
$ readlink -f /opt/apps/tool/releases/3.11.2
/opt/apps/tool/releases/3.11.2
$ echo $?
0
$ readlink -e /opt/apps/tool/releases/3.11.2 ; echo $?
1

# link written on a build host; /build/stage does not exist here at all
$ readlink -f /opt/apps/tool/current ; echo $?
1

Which means a check built on -f catches the build-host case and misses the pruned-release case — and the pruned release is the one that happens on a schedule. It fails in the direction that looks fine.

Both flags are GNU coreutils. BSD and macOS readlink has no -e at all, and its -f does not behave like the GNU one, so a check that runs on a mixed estate should use test -e "$(readlink -f …)" or simply [ -e /opt/apps/tool/current ], which follows the link and answers the question directly.

The assertion is not “does the module load”. It is “after loading, does the thing run and say the right version”:

module purge
module load tool/3.11.2
command -v tool | grep -q '^/opt/apps/' || { echo "tool not from /opt/apps" >&2; exit 1; }
tool --version | grep -q '3\.11\.2'     || { echo "wrong tool version" >&2; exit 1; }

Two lines, and they fail for the right reasons: wrong location, wrong version.

Class six: the environment change inside a pipe

This one is the mirror image of the rest. Here the estate is fine and the check is wrong, and the natural conclusion — “the module is broken” — sends you off to rebuild something that was never damaged.

In bash, every element of a pipeline runs in a subshell by default. Anything it does to the environment dies with it:

$ x=1
$ echo hi | { x=2; }
hi
$ echo "x=$x"
x=1

“By default” is doing real work in that sentence, and this is the one place in the article where the portable-looking rule is the unportable one. Bash exempts the last stage if lastpipe is set and job control is off — the usual state of a non-interactive script. zsh and ksh93 run the last stage in the current shell unconditionally:

$ bash -c 'echo hi | read y; echo "bash y=[$y]"'
bash y=[]
$ zsh  -c 'echo hi | read y; echo "zsh y=[$y]"'
zsh y=[hi]

So the same validation script, interpreted by a different shell, produces a different verdict about whether the estate is broken. Neither shell is wrong. Pin the interpreter in the shebang and do not rely on which side of that line you are on.

Now the operational version. A validation script wants to load a module and capture the result:

module load fftw/3.3.10 | tee -a validate.log
echo "FFTW_DIR=[$FFTW_DIR]"

module is a shell function. Inside the pipe it runs in a subshell, sets its variables there, and the subshell exits. FFTW_DIR is empty in the parent. The script concludes the module sets nothing and reports a broken module. A second engineer loads it by hand, it works perfectly, and the disagreement gets attributed to “something about the environment”.

The same mechanism hides real failures too:

$ bash -c 'false | true; echo "rc=$?"'
rc=0
$ bash -c 'set -o pipefail; false | true; echo "rc=$?"'
rc=1

Without pipefail, a pipeline reports the status of its last command only. Pipe a failing producer into tee and the failure is gone. This is why so much automation that pipes its output to a log has no idea when it failed.

set -e has its own blind spot in the same family:

$ bash -c 'set -eu; f(){ local v=$(false); echo "inside rc=$? v=[$v]"; }; f; echo after'
inside rc=0 v=[]
after

local is itself a command, and its exit status is the status of local, not of the substitution inside it. export, declare and readonly swallow a failure the same way, for the same reason:

$ bash -c 'set -eu; export v=$(false); echo "reached"'
reached
$ bash -c 'set -eu; declare v=$(false); echo "reached"'
reached

Split the declaration from the assignment and the same script stops:

$ bash -c 'set -eu; f(){ local v; v=$(false); echo "never reached"; }; f'
$ echo $?
1

One more in the same category, worth knowing because it appears in cleanup and maintenance scripts — and worth stating carefully, because the usual version of this folklore is half true. find -exec cmd {} \; returns zero even when every invocation of cmd fails. find -exec cmd {} + does not: in GNU findutils a non-zero invocation makes find itself exit non-zero.

$ find /tmp/fecheck -type f -exec false {} \; ; echo "semicolon rc=$?"
semicolon rc=0
$ find /tmp/fecheck -type f -exec false {} +  ; echo "plus rc=$?"
plus rc=1

The terminator, not the tool, decides whether failures propagate. xargs also reports them — GNU xargs exits 123 if any invocation exits 1 to 125 — so if a maintenance script’s success depends on the command being run, use + or xargs, and do not read anything into the status of a \; form.

The rule that covers all of it, for any harness you write:

set -Eeuo pipefail

plus: do not put environment-modifying commands in pipelines, and do not hide a command substitution inside a local, declare or export.

Class seven: the suite that has never failed

The last class is not a bug. It is the reason the other six survive.

A validation suite that has returned green on every run since it was written has demonstrated exactly one thing: that it returns green. It has never been shown to be capable of returning anything else. If half its assertions were deleted tonight, tomorrow’s run would be identical, and you would never know.

This is the same argument as the one for restore drills and for fire alarms with a test button. A detector that has never been triggered under controlled conditions is decoration until proven otherwise.

The fix is negative controls: for every check, construct the fault it claims to detect, run the check, and require that it fails. Keep those injections in the repository next to the suite, and run them on a schedule.

#!/usr/bin/env bash
# negctl.sh — prove each check can fail
set -Eeuo pipefail

fixtures=/var/lib/validate/fixtures
fail=0

expect_fail() {           # expect_fail <label> <cmd...>
  local label=$1; shift
  if "$@" >/dev/null 2>&1; then
    echo "NEGATIVE CONTROL DID NOT FIRE: $label"
    fail=1
  else
    echo "ok (failed as required): $label"
  fi
}

expect_fail "empty index rejected"      check_index "$fixtures/index.empty.json"
expect_fail "html error page rejected"  check_index "$fixtures/index.404.html"
expect_fail "truncated index rejected"  check_index "$fixtures/index.truncated.json"
expect_fail "blank MaxRSS rejected"     check_accounting "$fixtures/sacct.blank.csv"
expect_fail "dangling symlink rejected" check_prefix   "$fixtures/dangling-link"
expect_fail "wrong version rejected"    check_version  "$fixtures/tool-3.10.0"

exit "$fail"

The fixtures matter as much as the script. index.404.html should be a real error page you captured, not one you typed; sacct.blank.csv should be real output from a query that really did return blank columns. Saving the broken artefact at the moment you find it is the cheapest thing in this article and the one most often skipped.

There is a second-order benefit. Running the negative controls tells you when a check has quietly stopped being able to fail — because someone loosened a threshold, or because the tool’s output format changed and a grep that used to match now matches nothing and the surrounding || true swallows it.

The discipline, condensed

Four rules, in the order they pay off.

Assert on content, not on status. Exit zero means the process reached its own idea of the end. It says nothing about the bytes it produced. Every check should name a value it expects and compare.

Write the expected value down in the check. Not “non-empty”, not “more than zero rows” — the actual string, the actual minimum, the actual version. A check that does not encode an expectation cannot detect a wrong answer, only a missing one.

Pick a case that must produce a populated result. The whole plugin-shaped failure class hides behind aggregate checks. One known job, one known file, one known value, asserted exactly, is worth more than a thousand rows counted.

Prove every check can fail. If you cannot show the injected fault that turns it red, it is not a check. Delete it or fix it, because leaving it there is worse than having nothing — it produces the feeling of coverage without the coverage.

And one habit underneath all four: when a number is uniform across every row, or a report is empty on a day when it should not be, or a pipeline says “0 applied” for the ninth week running, treat that as a finding about the measurement rather than a finding about the estate. In my experience the uniform answer is wrong far more often than it is right.

What this does not cover

Nothing here detects a computation that is wrong in a plausible way — output that is populated, well-formed and numerically incorrect. That needs a reference result, a known-answer test or an independent implementation, and it is a different and harder problem.

Nor does any of it help if the check and the thing being checked share a dependency. A probe that reads the same broken index, or resolves the same dangling link, or inherits the same empty variable, will agree with the system perfectly. When you build a negative control, inject the fault as far upstream as you can reach, and confirm the alarm travels all the way to the place a human would actually see it. A check nobody reads is the same as a check that cannot fail, arrived at by a different route.

All articles