<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Wirewalk Advisory</title>
  <subtitle>HPC and AI infrastructure engineering and architecture. Cluster design and build-out, InfiniBand fabric, parallel storage, GPU and AI stacks, schedulers, and the enterprise IT around them.</subtitle>
  <link href="https://www.wirewalk.com/feed.xml" rel="self"/>
  <link href="https://www.wirewalk.com/"/>
  <updated>2026-09-19T10:40:18-04:00</updated>
  <id>https://www.wirewalk.com/</id>
  <author><name>Wirewalk Advisory</name></author>
  <entry>
    <title>An MPI troubleshooting playbook: symptom to cause</title>
    <link href="https://www.wirewalk.com/writing/mpi-troubleshooting-playbook/"/>
    <published>2026-09-09T10:00:00-04:00</published>
    <updated>2026-09-09T10:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/mpi-troubleshooting-playbook/</id>
    <summary>Hangs, stragglers, silent TCP fallback, launcher failures and results that change between runs. Each has a small set of causes and a command that distinguishes them, and most of the time lost is spent on the wrong layer.</summary>
    <content type="html">&lt;p&gt;This is the reference half of the series. The earlier articles argue for testing
in a particular order; this one is the lookup table for when something is already
broken and you need the shortest path to a cause.&lt;/p&gt;

&lt;p&gt;The single most useful habit underneath all of it: &lt;strong&gt;capture the node list, the
environment and the module versions with every run&lt;/strong&gt;. Most of the diagnoses below
are trivial with that information and genuinely difficult without it.&lt;/p&gt;

&lt;h2 id=&quot;the-index&quot;&gt;The index&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Symptom&lt;/th&gt;
      &lt;th&gt;Most likely causes&lt;/th&gt;
      &lt;th&gt;First command&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Hangs at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MPI_Init&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;PMIx mismatch, launcher, hostfile&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpirun --mca plm_base_verbose 10&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Hangs at first collective&lt;/td&gt;
      &lt;td&gt;QP setup, memlock, one unreachable node&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun -N all bash -c &apos;ulimit -l&apos;&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Runs, but ~10x slow&lt;/td&gt;
      &lt;td&gt;Silent TCP fallback&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_TLS=rc_x,sm,self&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Fails only above N ranks&lt;/td&gt;
      &lt;td&gt;RC queue-pair exhaustion, registration memory&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_TLS=dc_x,sm,self&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;One rank always late&lt;/td&gt;
      &lt;td&gt;Degraded link, throttled node, stray daemon&lt;/td&gt;
      &lt;td&gt;Pairwize &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; sweep&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Timings vary run to run&lt;/td&gt;
      &lt;td&gt;Allocation varies, shared ISLs, jitter&lt;/td&gt;
      &lt;td&gt;Record and compare node lists&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Bandwidth exactly half&lt;/td&gt;
      &lt;td&gt;Width negotiation, PCIe width, single rail&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lspci -vv&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Retry exceeded&lt;/code&gt; errors&lt;/td&gt;
      &lt;td&gt;Fabric drops, RoCE PFC/ECN, dead peer&lt;/td&gt;
      &lt;td&gt;Port counters, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ethtool -S&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Worse after a maintenance window&lt;/td&gt;
      &lt;td&gt;Firmware skew, routing engine fell back&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibv_devinfo&lt;/code&gt; fleet-wide, SM log&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Everything below expands one row.&lt;/p&gt;

&lt;h2 id=&quot;hangs-at-startup&quot;&gt;Hangs at startup&lt;/h2&gt;

&lt;p&gt;A job that never reaches the application is a launcher problem, not a fabric
problem, and the fabric tooling will waste your time here.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;--mca&lt;/span&gt; plm_base_verbose 10 &lt;span class=&quot;nt&quot;&gt;--mca&lt;/span&gt; rmaps_base_verbose 5 &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 4 &lt;span class=&quot;nb&quot;&gt;hostname&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That prints the launch sequence as it happens, and the point at which it stops is
the diagnosis. Common stopping points:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cannot reach a node.&lt;/strong&gt; The launcher uses SSH or the scheduler’s own launch
mechanism. If it is SSH, key-based access between compute nodes must work
non-interactively, and on a cluster with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pam_slurm_adopt&lt;/code&gt; it must work &lt;em&gt;from
within an allocation&lt;/em&gt;, which is not the same test as from the login node.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PMIx version mismatch.&lt;/strong&gt; This is the common one on schedulers, and it produces
either a hang or an error naming PMIX. Slurm and Open MPI are each built against
a PMIx, and if the two disagree, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun&lt;/code&gt;-launched MPI jobs fail while
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpirun&lt;/code&gt;-launched ones work, or vice versa. Establish what each side has:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;srun &lt;span class=&quot;nt&quot;&gt;--mpi&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;list
ompi_info | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;pmix|prrte&apos;&lt;/span&gt;
pmix_info &lt;span class=&quot;nt&quot;&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Then be explicit rather than relying on the default:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;srun &lt;span class=&quot;nt&quot;&gt;--mpi&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;pmix &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; 4 ./a.out
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun --mpi=list&lt;/code&gt; does not offer a PMIx entry, Slurm was not built with PMIx
support and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpirun&lt;/code&gt; under an allocation is your path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“There are not enough slots available in the system.”&lt;/strong&gt; The hostfile or the
scheduler allocation says fewer slots than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-np&lt;/code&gt; asked for. Under a scheduler,
this usually means &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpirun&lt;/code&gt; did not pick up the allocation. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--oversubscribe&lt;/code&gt;
silences it and does not fix it; check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$SLURM_JOB_NODELIST&lt;/code&gt; and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$SLURM_NTASKS&lt;/code&gt; instead.&lt;/p&gt;

&lt;h2 id=&quot;hangs-at-the-first-collective&quot;&gt;Hangs at the first collective&lt;/h2&gt;

&lt;p&gt;Startup succeeded, ranks exist, and the job stops the first time they all have to
talk. This is connection establishment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Locked memory limit.&lt;/strong&gt; RDMA needs pinned pages. If the limit is not unlimited
inside the job, registration fails. The limit that matters is the one inside the
launched process, which is not necessarily the one in
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/etc/security/limits.conf&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;srun &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; 8 &lt;span class=&quot;nt&quot;&gt;--ntasks-per-node&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1 bash &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;echo $(hostname) $(ulimit -l)&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Anything other than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;unlimited&lt;/code&gt; on any node is the answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A node that is up but has no working adapter.&lt;/strong&gt; The job launches everywhere and
then waits for the one rank that cannot connect. Check the whole allocation
rather than a sample:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;srun &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$SLURM_NNODES&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--ntasks-per-node&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1 bash &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;s1&quot;&gt;&apos;echo &quot;$(hostname) $(ibstat mlx5_0 1 | awk &quot;/State:/{print \$2}&quot; | head -1)&quot;&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Firewall on the fabric interface.&lt;/strong&gt; Rare on InfiniBand, common on RoCE, where
the traffic is UDP and a host firewall will happily block port 4791.&lt;/p&gt;

&lt;p&gt;To see where it is stuck rather than guessing, attach to a hung rank:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;pstack &lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;pgrep &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; a.out&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;          &lt;span class=&quot;c&quot;&gt;# or: gdb -p &amp;lt;pid&amp;gt; -batch -ex bt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A stack sitting in a UCX progress or connection-establishment function is a
transport problem. A stack in application code is not.&lt;/p&gt;

&lt;h2 id=&quot;silent-fall-back-to-tcp&quot;&gt;Silent fall back to TCP&lt;/h2&gt;

&lt;p&gt;The most expensive quiet failure in MPI, because nothing is broken. The job
completes, the numbers are bad, and the conclusion recorded is that the
application does not scale.&lt;/p&gt;

&lt;p&gt;The signature is quantitative: bandwidth roughly an order of magnitude below the
link rate and small-message latency roughly an order of magnitude above what the
fabric should give. If a two-node &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; on a 200 Gb/s fabric reports a couple
of gigabytes per second, stop tuning and check the transport.&lt;/p&gt;

&lt;p&gt;The conclusive test removes the fallback:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_TLS&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;rc_x,sm,self &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_0:1 ./osu_bw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;No TCP transport is listed. If the job runs, RDMA was working and the problem is
elsewhere. If it fails to establish connections, TCP was carrying the traffic.&lt;/p&gt;

&lt;p&gt;Corroborate by watching the device counters move — or not — during a run:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;before&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/infiniband/mlx5_0/ports/1/counters/port_xmit_data&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node ./osu_bw &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; 8388608:8388608 &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;/dev/null
&lt;span class=&quot;nv&quot;&gt;after&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/infiniband/mlx5_0/ports/1/counters/port_xmit_data&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;echo&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;delta: &lt;/span&gt;&lt;span class=&quot;k&quot;&gt;$((&lt;/span&gt; after &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; before &lt;span class=&quot;k&quot;&gt;))&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Causes, in the order they are usually found: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ulimit -l&lt;/code&gt; not unlimited; a device
name in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_NET_DEVICES&lt;/code&gt; that does not exist on some nodes; a node with a
different OFED; a port that is down on one node; and, on GPU nodes, a CUDA-aware
build mismatch that causes UCX to reject the RDMA path for device buffers.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Put the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_TLS=rc_x,sm,self&lt;/code&gt; run in your acceptance suite as a pass/fail gate.
It costs seconds and it converts the single most damaging silent degradation into
a loud, obvious failure. Almost nobody does this, and it is the highest-value
five lines in an MPI test harness.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;fails-only-above-a-certain-scale&quot;&gt;Fails only above a certain scale&lt;/h2&gt;

&lt;p&gt;A job that works at 8 nodes and fails at 64 is usually a resource that scales
with the square of the rank count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reliable Connection queue pairs.&lt;/strong&gt; With RC, each rank maintains a QP to each
peer it talks to. Full connectivity across &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n&lt;/code&gt; ranks is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n²&lt;/code&gt; QPs, and each
consumes adapter and host resources. On a node running many ranks, this becomes
the constraint well before the fabric does.&lt;/p&gt;

&lt;p&gt;Dynamically Connected transport solves it by multiplexing:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_TLS&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;dc_x,sm,self &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_0:1 ./a.out
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If a job fails at high ranks-per-node with connection or resource errors and
succeeds with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dc_x&lt;/code&gt;, that was it. Note that DC availability depends on the
adapter generation, and that the failure mode when DC resources themselves are
exhausted is different and less obvious — usually a hang rather than an error.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory registration limits.&lt;/strong&gt; Large registered regions need translation table
entries. The relevant firmware parameters differ by adapter and driver
generation; the symptom is a registration failure at a size or rank count that
worked before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Aggregate memory per node.&lt;/strong&gt; Each connection carries buffers. At high
ranks-per-node, MPI’s own buffer footprint becomes significant, and the job dies
of OOM in a way that looks like an application memory bug. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dmesg&lt;/code&gt; on the node
that died settles it in one line.&lt;/p&gt;

&lt;h2 id=&quot;one-rank-is-always-late&quot;&gt;One rank is always late&lt;/h2&gt;

&lt;p&gt;A collective takes as long as its slowest participant, so one degraded node
degrades everything. Finding it is mechanical.&lt;/p&gt;

&lt;p&gt;Bisection is fastest: split the allocation, run both halves, keep the slow half,
repeat. Six rounds for 64 nodes.&lt;/p&gt;

&lt;p&gt;The systematic version tests every node against one reference, at the verbs layer
so MPI is not in the picture:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;REF&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;node001
&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;n &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;scontrol show hostnames &lt;span class=&quot;nv&quot;&gt;$SLURM_JOB_NODELIST&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do&lt;/span&gt;
  &lt;span class=&quot;o&quot;&gt;[&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;continue
  &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;bw&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;ssh &lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;ib_write_bw -d mlx5_0 -F -s 1048576 -D 5 --report_gbits &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
        | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;/1048576/ {print $4}&apos;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
  &lt;span class=&quot;nb&quot;&gt;echo&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$bw&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;sort&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-k2&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;head&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Once you have the node, the causes are a short list:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Narrow or degraded link — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlxlink -d mlx5_0 -p 1 --show_fec&lt;/code&gt;.&lt;/li&gt;
  &lt;li&gt;Downtrained PCIe slot — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lspci -vv | grep LnkSta&lt;/code&gt;.&lt;/li&gt;
  &lt;li&gt;CPU frequency capping or thermal throttling — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;turbostat&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cpupower
frequency-info&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dmesg | grep -i thermal&lt;/code&gt;.&lt;/li&gt;
  &lt;li&gt;A daemon on a core that a rank is pinned to.&lt;/li&gt;
  &lt;li&gt;Memory running at a lower speed, or a DIMM that failed into a degraded
configuration — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dmidecode -t memory&lt;/code&gt;, and the platform’s own health log.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On the last point about cores: &lt;strong&gt;aggregate CPU idle proves nothing on a many-core
node.&lt;/strong&gt; One rank spinning uselessly on a 192-core machine is half a percent of
system-wide CPU. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;top&lt;/code&gt; shows a healthy node. Look per-core:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpstat &lt;span class=&quot;nt&quot;&gt;-P&lt;/span&gt; ALL 1 5
pidstat &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; &lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;pgrep &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; a.out&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt; 1 5
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;timings-vary-wildly-between-runs&quot;&gt;Timings vary wildly between runs&lt;/h2&gt;

&lt;p&gt;Before treating variance as a signal, establish that it is not the experiment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The allocation changed.&lt;/strong&gt; Different nodes, different leaves, different number
of spine crossings. This is the most common cause by a wide margin. Record
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scontrol show hostnames $SLURM_JOB_NODELIST&lt;/code&gt; with every result and compare.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Other jobs share the fabric.&lt;/strong&gt; Exclusive node allocation does not give you
exclusive links. Another job’s alltoall on the same spine uplinks changes your
result, and you have no visibility into it from inside your job. If you are
producing numbers that matter, get an exclusive fabric window, or at minimum
record what else was running.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cached or warm state.&lt;/strong&gt; The first iteration of anything includes connection
setup and memory registration. Every benchmark here has a warmup flag; use it.
Where filesystem I/O is in the loop, a repeat run served from page cache is not a
measurement of anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OS jitter.&lt;/strong&gt; Unpinned ranks, an unsynchronized housekeeping daemon, a
monitoring agent that wakes every thirty seconds. Collectives amplify jitter
because every rank waits for the one that was interrupted. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--bind-to core&lt;/code&gt; and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--report-bindings&lt;/code&gt; first; if variance persists with clean binding, look at what
runs on the compute nodes on a timer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Power and frequency policy.&lt;/strong&gt; A node that boosts for the first thirty seconds
and then settles produces a benchmark whose result depends on its duration.
Fix the governor for measurement runs and record what it was.&lt;/p&gt;

&lt;p&gt;A run-to-run spread you cannot get under a few percent is not a slow cluster. It
is an uncontrolled experiment, and no amount of tuning against it will converge.&lt;/p&gt;

&lt;h2 id=&quot;bandwidth-that-is-a-clean-fraction-of-expectation&quot;&gt;Bandwidth that is a clean fraction of expectation&lt;/h2&gt;

&lt;p&gt;Half, a quarter, a tenth. Clean fractions have mechanical causes and are worth a
minute of arithmetic before any investigation.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Fraction&lt;/th&gt;
      &lt;th&gt;Look at&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Exactly half&lt;/td&gt;
      &lt;td&gt;Link width 2X not 4X; one rail of two in use; PCIe x8 not x16&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Exactly a quarter&lt;/td&gt;
      &lt;td&gt;Link width 1X; PCIe x4; two independent halvings&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;About a tenth&lt;/td&gt;
      &lt;td&gt;TCP fallback, or IPoIB instead of RDMA&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;About 60 percent&lt;/td&gt;
      &lt;td&gt;Often real — protocol overhead plus an unsaturated single rank&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibstatus mlx5_0 | &lt;span class=&quot;nb&quot;&gt;grep &lt;/span&gt;rate
iblinkinfo &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; 4X
lspci &lt;span class=&quot;nt&quot;&gt;-vv&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; &amp;lt;bdf&amp;gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;LnkCap|LnkSta&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;mtu&quot;&gt;MTU&lt;/h2&gt;

&lt;p&gt;On InfiniBand, path MTU is negotiated and normally 4096. On RoCE it is derived
from the Ethernet MTU, and a single hop at 1500 anywhere in the path drops the
RoCE MTU to 1024 with a real bandwidth cost.&lt;/p&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; header prints what was actually negotiated, which is the number
to trust over any configuration file:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt; Mtu             : 4096[B]
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If that reads 1024 on a fabric you configured for jumbo frames, walk the path.
The offender is usually a recently added switch port or a host whose interface
configuration did not include the MTU.&lt;/p&gt;

&lt;p&gt;IPoIB has its own version of this — datagram mode caps at 2044 bytes while
connected mode allows much larger — which matters if any part of your workload,
including the scheduler or the filesystem, is running over IPoIB.&lt;/p&gt;

&lt;h2 id=&quot;after-a-maintenance-window&quot;&gt;After a maintenance window&lt;/h2&gt;

&lt;p&gt;Performance that changed after a window, with no application change, has a short
suspect list.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Firmware skew.&lt;/strong&gt; Some nodes were updated and some were not, or a node was
replaced with one from a different batch. Group by PSID and compare within
groups:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;pdsh &lt;span class=&quot;nt&quot;&gt;-w&lt;/span&gt; node[001-256] &lt;span class=&quot;s1&quot;&gt;&apos;ibv_devinfo | grep -E &quot;fw_ver|board_id&quot;&apos;&lt;/span&gt; | dshbak &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The routing engine fell back.&lt;/strong&gt; A switch that was down during the sweep, or a
node cabled differently on the way back in, can make &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt; decline the topology
and fall back to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;updn&lt;/code&gt; for the entire fabric. The subnet manager logs it and
nobody reads it:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;routing engine|fallback|ftree&apos;&lt;/span&gt; /var/log/opensm.log | &lt;span class=&quot;nb&quot;&gt;tail&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-40&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A standby subnet manager took over.&lt;/strong&gt; If the standby holds an older
configuration, the whole fabric is now routed by that configuration. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sminfo&lt;/code&gt;
tells you which SM is master.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A driver or kernel change altered defaults.&lt;/strong&gt; UCX and OFED defaults do move
between releases. Diff the effective configuration, not the configuration files:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ucx_info &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; /var/tmp/ucx-&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;date&lt;/span&gt; +%F&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;.conf
diff /var/tmp/ucx-&amp;lt;previous&amp;gt;.conf /var/tmp/ucx-&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;date&lt;/span&gt; +%F&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;.conf
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;error-messages-worth-recognizing&quot;&gt;Error messages worth recognizing&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;IBV_WC_RETRY_EXC_ERR&lt;/code&gt; / “Retry exceeded”.&lt;/strong&gt; The sender gave up after
retransmitting. On InfiniBand, this usually means the peer died or a link went
down mid-transfer. On RoCE, it much more often means the fabric is dropping —
PFC not enabled on the RoCE priority, or a mismatched trust mode. Go to the
per-priority discard counters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;RNR retry exceeded&lt;/code&gt;.&lt;/strong&gt; The receiver had no buffer posted and kept saying so
until the sender gave up. Usually an application or middleware issue rather than
a fabric one, though severe congestion can produce it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fork()&lt;/code&gt; warning.&lt;/strong&gt; Registered memory and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fork()&lt;/code&gt; interact badly. If your
application or a library it calls forks after registering memory, you get either
a warning or corruption. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibv_fork_init&lt;/code&gt; and the corresponding MPI settings exist
for this; a job that calls out to shell commands from inside a rank is the usual
trigger.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX ERROR ... Connection reset by remote peer&lt;/code&gt;.&lt;/strong&gt; A peer rank died. Find out
why that rank died; the UCX error is a consequence, not a cause. Check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dmesg&lt;/code&gt; on
its node for OOM.&lt;/p&gt;

&lt;h2 id=&quot;the-habit-that-shortens-all-of-this&quot;&gt;The habit that shortens all of this&lt;/h2&gt;

&lt;p&gt;Keep a baseline file, in version control, per cluster:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;OFED, firmware and PSID per node group&lt;/li&gt;
  &lt;li&gt;PCIe &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LnkSta&lt;/code&gt; per adapter&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_read_bw&lt;/code&gt; plateaus for a named intra-leaf and cross-spine
pair&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_latency&lt;/code&gt; at 8 bytes and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; at 8 MB&lt;/li&gt;
  &lt;li&gt;Allreduce and alltoall curves at the job sizes you run, with node lists&lt;/li&gt;
  &lt;li&gt;The width-and-rate histogram for the whole fabric&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Regenerate it after every maintenance window. Almost every diagnosis in this
article becomes a diff against that file, and a diff takes minutes where an
investigation takes days.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>EDR to XDR in one fabric: rate reporting, width traps, and the node that gates the job</title>
    <link href="https://www.wirewalk.com/writing/mpi-mixed-generation-fabrics/"/>
    <published>2026-09-09T09:00:00-04:00</published>
    <updated>2026-09-09T09:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/mpi-mixed-generation-fabrics/</id>
    <summary>Different tools report link speed in different units, half the field is per-lane, and a link that negotiated two lanes instead of four is Active and silent. On a synchronized collective, one such link sets the speed of the entire job.</summary>
    <content type="html">&lt;p&gt;Clusters accumulate generations. An EDR island from 2017 that still runs
production, an HDR expansion, an NDR refresh, and now XDR arriving for the GPU
tier — frequently all reachable from one another, sometimes in one subnet. Mixed
fabrics work. What does not work is reasoning about them from a mental model
built on a single generation, because the reporting is inconsistent and the
failure modes are quiet.&lt;/p&gt;

&lt;h2 id=&quot;the-rate-table-and-which-number-each-tool-prints&quot;&gt;The rate table, and which number each tool prints&lt;/h2&gt;

&lt;p&gt;The confusion is entirely avoidable and it costs people days.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Generation&lt;/th&gt;
      &lt;th&gt;Lanes&lt;/th&gt;
      &lt;th&gt;Per-lane signalling&lt;/th&gt;
      &lt;th&gt;4X link rate&lt;/th&gt;
      &lt;th&gt;Encoding&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;FDR&lt;/td&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;14.0625 Gb/s&lt;/td&gt;
      &lt;td&gt;56 Gb/s&lt;/td&gt;
      &lt;td&gt;64b/66b&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;EDR&lt;/td&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;25.78125 Gb/s&lt;/td&gt;
      &lt;td&gt;100 Gb/s&lt;/td&gt;
      &lt;td&gt;64b/66b&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;HDR&lt;/td&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;53.125 Gb/s&lt;/td&gt;
      &lt;td&gt;200 Gb/s&lt;/td&gt;
      &lt;td&gt;PAM4, 64b/66b + RS-FEC&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NDR&lt;/td&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;106.25 Gb/s&lt;/td&gt;
      &lt;td&gt;400 Gb/s&lt;/td&gt;
      &lt;td&gt;PAM4, RS-FEC&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;XDR&lt;/td&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;212.5 Gb/s&lt;/td&gt;
      &lt;td&gt;800 Gb/s&lt;/td&gt;
      &lt;td&gt;PAM4, RS-FEC&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Now the part that matters. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstat&lt;/code&gt; prints the &lt;strong&gt;aggregate link rate&lt;/strong&gt;:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;	Rate: 400
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstatus&lt;/code&gt; prints the aggregate rate and helpfully names the generation:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;	rate:            400 Gb/sec (4X NDR)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt; prints the &lt;strong&gt;width and the per-lane rate as separate fields&lt;/strong&gt;:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;  37   12[  ] ==( 4X 106.25 Gbps Active/  LinkUp)==&amp;gt;  91   1[  ] &quot;node17 mlx5_0&quot; ( )
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That is a 400 Gb/s link. Read the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;106.25&lt;/code&gt; as the link rate and you conclude
your NDR fabric is running at roughly 100 Gb/s, which is a conclusion people
reach, escalate, and occasionally publish. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;4X&lt;/code&gt; prefix is doing the work and
it is easy to skim past.&lt;/p&gt;

&lt;p&gt;On a mixed fabric this is worse, because &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;4X 106.25 Gbps&lt;/code&gt; (NDR400) and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;4X 25.78125 Gbps&lt;/code&gt; (EDR100) look superficially similar in a long scan and the
distinguishing digits are in the middle of the field.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;The rule that keeps this straight: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstat&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstatus&lt;/code&gt; report the link, and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt; reports a lane. Whenever you quote a number from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt;,
quote the width with it. A capacity claim sourced from that field without the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;4X&lt;/code&gt; attached is wrong by a factor of four, and it is wrong in the direction
that makes people buy hardware.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;the-width-trap&quot;&gt;The width trap&lt;/h2&gt;

&lt;p&gt;Width is the field that silently costs you three quarters of a link.&lt;/p&gt;

&lt;p&gt;An InfiniBand link can negotiate 1X, 2X or 4X. A link that comes up at 1X is
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Active&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LinkUp&lt;/code&gt;, routable, pingable, and passes every functional test. It runs
at a quarter rate. Nothing logs an error, because from the fabric’s point of view
nothing is wrong — the two ends negotiated the best width they could agree on,
and that is what negotiation is for.&lt;/p&gt;

&lt;p&gt;Causes: a damaged or partially seated cable, a transceiver with a failed lane, a
splitter cable in a port that is not configured to split, or a port configured to
split that has a straight cable in it.&lt;/p&gt;

&lt;p&gt;The scan to run across the whole fabric, on a schedule:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;iblinkinfo &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;$0 !~ /4X/ &amp;amp;&amp;amp; /Active/ {print}&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;And the more useful version, which catches both narrow links and links that
negotiated a lower generation than they should have:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;iblinkinfo &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-oE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;[0-9]+X +[0-9.]+ Gbps&apos;&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;tr&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos; &apos;&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;sort&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;uniq&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;sort&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-rn&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That gives you a histogram of every distinct width-and-rate combination in the
fabric. On a homogeneous fabric it should have one row. On a mixed fabric it
should have exactly as many rows as you have generations, and every count should
match what you cabled. Any row with a small count is the anomaly, and small
counts are the whole point of the histogram.&lt;/p&gt;

&lt;h2 id=&quot;hdr100-and-the-two-ways-to-get-it&quot;&gt;HDR100 and the two ways to get it&lt;/h2&gt;

&lt;p&gt;HDR100 is the case where the same nominal rate arrives by two different physical
arrangements, and the difference matters when you are diagnosing.&lt;/p&gt;

&lt;p&gt;An HDR switch port carries four HDR lanes. Split, it presents as two ports of two
lanes each — HDR100, reported as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2X 53.125 Gbps&lt;/code&gt;. Separately, there are HDR100
adapters, which are physically capable of two lanes only and report &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2X&lt;/code&gt; even
into an unsplit port.&lt;/p&gt;

&lt;p&gt;So &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2X 53.125 Gbps&lt;/code&gt; is normal and expected on a split fabric with HDR100
adapters, and is a fault on a fabric where you believed everything was HDR200.
The histogram above is what tells you which situation you are in; the individual
port reading does not.&lt;/p&gt;

&lt;p&gt;Three practical consequences:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;A single HDR200 adapter plugged into a split HDR100 port runs at 100. Half of
the adapter is unused and nothing reports it.&lt;/li&gt;
  &lt;li&gt;A single HDR100 adapter plugged into an unsplit HDR200 port also runs at 100,
and wastes half a switch port.&lt;/li&gt;
  &lt;li&gt;Cable compatibility is not guaranteed across these arrangements. Active optical
cables in particular have power class requirements that some port and adapter
combinations do not satisfy, and the symptom is a port that will not come out of
Polling with no error message that explains why. Check the transceiver
explicitly before assuming a bad cable:&lt;/li&gt;
&lt;/ul&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mlxlink &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; 1 &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That reports the module’s vendor, part number, type and supported rates, which
is the information the failure message does not give you.&lt;/p&gt;

&lt;h2 id=&quot;when-one-node-gates-the-job&quot;&gt;When one node gates the job&lt;/h2&gt;

&lt;p&gt;This is the expensive lesson, and it is the reason the width scan belongs on a
schedule rather than in a runbook.&lt;/p&gt;

&lt;p&gt;A synchronized collective — allreduce, alltoall, a barrier before a timing
region — completes when its slowest participant completes. The whole job
proceeds at the rate of its worst link. One node at 2X in a 32-node NDR
allocation does not cost you one thirty-second of the throughput. Depending on
the communication pattern, it can cost a large fraction of it, because every
other rank waits.&lt;/p&gt;

&lt;p&gt;What makes it expensive is not the degradation, it is the presentation. The
symptom is:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Results that vary between runs by a wide margin.&lt;/li&gt;
  &lt;li&gt;The variation correlating with nothing you changed.&lt;/li&gt;
  &lt;li&gt;Point-to-point tests that are clean, because the odds of any given pair
including the bad node are low.&lt;/li&gt;
  &lt;li&gt;Fabric-wide diagnostics that report no errors, because a narrow link is not an
error.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It presents as a fabric-wide capacity problem. It gets investigated as one. The
routing engine gets blamed, ISL capacity gets blamed, a purchase order for more
spine ports gets discussed — and the actual finding is one transceiver.&lt;/p&gt;

&lt;p&gt;The tell is the correlation with allocation. If you keep the node list with every
result, the correlation is visible in an afternoon. If you do not, it is not
visible at all, which is why the earlier articles in this series keep insisting
on recording the node list.&lt;/p&gt;

&lt;p&gt;The other tell is arithmetic. If a synchronized benchmark comes in at close to a
clean fraction of expectation — half, a quarter — suspect a width negotiation
before suspecting anything subtle. Fabrics rarely degrade by exactly 50 percent
for interesting reasons.&lt;/p&gt;

&lt;h2 id=&quot;mixing-generations-in-one-subnet&quot;&gt;Mixing generations in one subnet&lt;/h2&gt;

&lt;p&gt;Several things happen that are worth anticipating.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Links negotiate down to the lower common generation.&lt;/strong&gt; An NDR adapter into an
HDR switch port gives HDR, with the appropriate cable. That is correct behavior
and it is only a problem when nobody recorded the intent, and a year later
somebody is trying to work out why a node is at 200 when the inventory says 400.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path MTU is set by the minimum along the path.&lt;/strong&gt; OpenSM computes per-path MTU,
so a subnet containing something that supports only 2048 will have paths through
it at 2048 while other paths run at 4096. Two node pairs then produce different
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; headers on the same fabric, which is confusing if you have not
seen it before. Check the SM’s configured maximum and the per-device capability:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; max_mtu /etc/opensm/opensm.conf
ibv_devinfo &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;max_mtu|active_mtu&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Service levels and virtual lanes may differ.&lt;/strong&gt; The number of data VLs a device
supports varies across generations. If you use QoS and service-level mapping,
the mapping has to be valid on the least capable device in the path, and an
invalid mapping does not fail loudly — traffic lands on VL0 with everything
else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The routing engine is constrained by the least regular part of the fabric.&lt;/strong&gt;
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt; requires a regular fat tree; an older island bolted on with a different
uplink ratio is exactly the sort of irregularity that makes it decline and fall
back to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;updn&lt;/code&gt;, fabric-wide. So adding a generation can change the routing of
the parts you did not touch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Forwarding table capacity is set by the oldest switch.&lt;/strong&gt; Every switch needs an
entry for every LID it must reach. The smallest table in the fabric governs, and
older switches have smaller tables. The article on LID budgets in this series
covers the arithmetic.&lt;/p&gt;

&lt;h2 id=&quot;firmware-skew-across-a-mixed-fleet&quot;&gt;Firmware skew across a mixed fleet&lt;/h2&gt;

&lt;p&gt;Mixed generations mean multiple firmware trains, and mixed firmware within a
single model is where the odd behavior lives.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;pdsh &lt;span class=&quot;nt&quot;&gt;-w&lt;/span&gt; node[001-256] &lt;span class=&quot;s1&quot;&gt;&apos;ibv_devinfo | grep -E &quot;hca_id|fw_ver|board_id&quot;&apos;&lt;/span&gt; | dshbak &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt;
mlxfwmanager &lt;span class=&quot;nt&quot;&gt;--query&lt;/span&gt;
flint &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4123_pciconf0 q
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Group by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;board_id&lt;/code&gt; — the PSID — not by model name. Within a PSID group, the
firmware version should be identical. Across PSID groups it will not be and
should not be expected to be. The check is uniformity within a group, and the
detail of doing the update is covered in the ConnectX firmware article in this
series.&lt;/p&gt;

&lt;p&gt;Skew within a group is worth taking seriously because the symptoms are
generation-specific and unhelpful: a link that trains at a lower width with one
firmware and correctly with another, a counter that reads differently, a
congestion-control feature present on some cards and not others.&lt;/p&gt;

&lt;h2 id=&quot;physical-quality-which-is-different-from-link-state&quot;&gt;Physical quality, which is different from link state&lt;/h2&gt;

&lt;p&gt;A link can be at full width and full rate and still be marginal. Forward error
correction hides a great deal, and it hides it right up until the offered load
is high enough that it cannot.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mlxlink &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; 1 &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--show_fec&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--show_eye&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--show_counters&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The fields to compare across the fleet are raw BER, effective BER, and the FEC
histogram if the tool provides one for your device. A port whose raw BER is
orders of magnitude worse than its peers while its effective BER is clean is a
link that FEC is currently rescuing. It will pass every functional test. It is
also the one that will produce intermittent, load-correlated performance
problems that no configuration change fixes.&lt;/p&gt;

&lt;p&gt;Comparing across the fleet is the essential part. A raw BER figure in isolation
is hard to judge; the same figure alongside sixty-three peers is obvious.&lt;/p&gt;

&lt;h2 id=&quot;practical-policy-for-a-mixed-estate&quot;&gt;Practical policy for a mixed estate&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Keep generations in separate scheduler partitions and separate topology
groups.&lt;/strong&gt; A job that straddles an NDR island and an EDR island runs at EDR, and
the user will not know why. Make the boundary explicit in the scheduler rather
than implicit in the fabric.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run the width and rate histogram on a schedule&lt;/strong&gt; and alert on any change to
the row counts. It is one command, it is cheap, and it converts the most
expensive class of quiet failure into a notification.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Record intent alongside the fabric.&lt;/strong&gt; For every link that is deliberately at a
lower rate than the adapter supports, write down why. Otherwise every future
audit rediscovers it as an anomaly and someone eventually “fixes” it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Baseline &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlxlink&lt;/code&gt; output per port&lt;/strong&gt; at install and compare periodically. It is
the only way to see a link degrading rather than a link failed.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>MPI on InfiniBand and on Spectrum-X: what transfers and what does not</title>
    <link href="https://www.wirewalk.com/writing/mpi-infiniband-vs-spectrumx/"/>
    <published>2026-09-08T14:00:00-04:00</published>
    <updated>2026-09-08T14:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/mpi-infiniband-vs-spectrumx/</id>
    <summary>The MPI layer looks identical on both. Everything underneath it is different, and on the Ethernet side a single hop configured wrongly degrades collectives across the entire fabric while every link reports healthy.</summary>
    <content type="html">&lt;p&gt;Running MPI over an Ethernet fabric with RoCEv2 and running it over InfiniBand
look the same from the application. Same verbs API, same UCX, same Open MPI,
same OSU benchmarks, same &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt;. The adapter presents an
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/sys/class/infiniband/&lt;/code&gt; device either way, which is why so much tooling appears
to be portable.&lt;/p&gt;

&lt;p&gt;What is not portable is everything that makes the fabric lossless, and that is
where the operational differences live. On InfiniBand, link-level flow control
is built into the architecture and is on by default. On Ethernet it is a
configuration you have to get right on every port of every switch and every NIC
in the path, and getting it wrong on one hop degrades the whole fabric.&lt;/p&gt;

&lt;h2 id=&quot;the-concept-mapping&quot;&gt;The concept mapping&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;InfiniBand&lt;/th&gt;
      &lt;th&gt;Ethernet / RoCEv2&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Subnet manager assigns LIDs and computes routes&lt;/td&gt;
      &lt;td&gt;No SM. L3 routing, BGP/ECMP, ARP and neighbour discovery&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Credit-based link flow control, always on&lt;/td&gt;
      &lt;td&gt;PFC (802.1Qbb) per priority, configured per port&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Congestion control notification (CCT/CCTI), rarely enabled&lt;/td&gt;
      &lt;td&gt;ECN marking plus CNP, and DCQCN on the NIC — normally mandatory&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstat&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibdiagnet&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ethtool -S&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlnx_qos&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lldpctl&lt;/code&gt;, switch telemetry&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Path MTU 256–4096, negotiated&lt;/td&gt;
      &lt;td&gt;RoCE MTU derived from the L2 MTU; jumbo frames needed for 4096&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Static routing by default; adaptive routing a switch feature&lt;/td&gt;
      &lt;td&gt;ECMP by flow hash; Spectrum-X adds per-packet adaptive routing&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Fabric is a single administrative object&lt;/td&gt;
      &lt;td&gt;Fabric is an IP network, with everything that implies&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Everything in the left column that you know how to use has an equivalent in the
right column that behaves differently enough to catch you out.&lt;/p&gt;

&lt;h2 id=&quot;what-spectrum-x-changes-about-the-ethernet-side&quot;&gt;What Spectrum-X changes about the Ethernet side&lt;/h2&gt;

&lt;p&gt;Plain RoCEv2 on general-purpose Ethernet is workable and fragile. The fragility
comes from two places: PFC is a blunt instrument that spreads congestion
backwards through the fabric, and ECMP hashes flows onto links, so a small
number of large elephant flows — which is exactly what a collective produces —
can collide on one link while others sit idle.&lt;/p&gt;

&lt;p&gt;Spectrum-X addresses both. The switch performs adaptive routing at packet
granularity rather than flow granularity, spreading a single flow’s packets
across available paths, and the NIC handles the resulting out-of-order arrival
and places data correctly without the sender having to care. Congestion control
is done in coordination between switch telemetry and the NIC rather than by the
generic DCQCN loop alone.&lt;/p&gt;

&lt;p&gt;The practical consequence for MPI is that the ECMP collision problem — the one
that makes plain RoCE alltoall performance erratic — is substantially addressed
by the platform rather than by your hashing configuration. The PFC and ECN
configuration discipline does not go away.&lt;/p&gt;

&lt;p&gt;Treat the specific feature names and defaults for your switch OS release as
authoritative over anything written here. What is stable is the shape of the
problem, not the syntax.&lt;/p&gt;

&lt;h2 id=&quot;pfc-and-why-one-hop-matters&quot;&gt;PFC, and why one hop matters&lt;/h2&gt;

&lt;p&gt;Priority Flow Control sends a PAUSE frame for a specific traffic class when a
port’s ingress buffer for that class fills. It is hop-by-hop: the upstream
device stops sending that class until the pause expires or is canceled.&lt;/p&gt;

&lt;p&gt;Two properties make it treacherous.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is per-priority, and the priority has to match end to end.&lt;/strong&gt; The NIC marks
RoCE traffic with a DSCP value or a PCP value; the switch maps that to a traffic
class; PFC is enabled on that class. Break the chain at any point — one switch
that trusts PCP where the NIC is marking DSCP, one port where PFC was never
enabled on priority 3, one uplink that was added later from a different template
— and RoCE traffic on that hop lands in a lossy class. It then drops under
congestion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A drop is far more expensive than it looks.&lt;/strong&gt; RoCE’s reliable connection
transport recovers from loss, but recovery costs a round trip and, on older
implementations, a retransmission of everything after the lost packet. A drop
rate that would be invisible to TCP is enough to flatten collective performance.&lt;/p&gt;

&lt;p&gt;And because a collective completes when its slowest participant completes, a
single misconfigured hop that affects a handful of flows degrades the whole job.
This is the mechanism behind the most confusing class of RoCE incident: every
link is up, no interface shows errors, throughput between any two nodes you test
is fine, and the alltoall is half what it should be.&lt;/p&gt;

&lt;h3 id=&quot;checking-the-chain&quot;&gt;Checking the chain&lt;/h3&gt;

&lt;p&gt;On the NIC:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mlnx_qos &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; enp1s0f0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;DCBX mode: OS controlled
Priority trust state: dscp
Cable len: 7
PFC configuration:
	priority    0   1   2   3   4   5   6   7
	enabled     0   0   0   1   0   0   0   0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two lines matter. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Priority trust state&lt;/code&gt; must agree with what the switch is
trusting. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PFC configuration&lt;/code&gt; must have the RoCE priority enabled and must agree
with the switch’s per-port configuration.&lt;/p&gt;

&lt;p&gt;Set it explicitly rather than inheriting:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mlnx_qos &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; enp1s0f0 &lt;span class=&quot;nt&quot;&gt;--trust&lt;/span&gt; dscp
mlnx_qos &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; enp1s0f0 &lt;span class=&quot;nt&quot;&gt;--pfc&lt;/span&gt; 0,0,0,1,0,0,0,0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Confirm what the RoCE traffic is actually marked with:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/net/enp1s0f0/ecn/roce_np/cnp_dscp
cma_roce_tos &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Then read the pause counters after a run. These are the evidence that PFC is
being exercised, and whether it is being exercised on the priority you intended:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ethtool &lt;span class=&quot;nt&quot;&gt;-S&lt;/span&gt; enp1s0f0 | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;prio[0-7]_(pause|buf_discard)&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;rx_prio3_pause: 148213
rx_prio3_pause_duration: 3320119
tx_prio3_pause: 96204
tx_prio3_pause_duration: 1904772
rx_prio0_buf_discard: 0
rx_prio3_buf_discard: 0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Read it as follows. Pause counts on priority 3 that are non-zero and growing
mean PFC is active and congestion is real — that is the mechanism working, not a
fault, though sustained heavy pausing means the fabric is oversubscribed for the
offered load. A non-zero &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;buf_discard&lt;/code&gt; on the RoCE priority means packets were
dropped in a class that was supposed to be lossless, which is a configuration
failure. And pause counters on priority 0 while RoCE was running means your
traffic is not in the class you think it is.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;The counter that most often reveals the problem is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rx_prio0_pause&lt;/code&gt; being
non-zero on a fabric that was configured for RoCE on priority 3. It means the
marking or the trust mode is wrong somewhere, and the traffic is being pause-
controlled — or not — in the default class. Every link is up, nothing logs an
error, and the collectives are poor.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;ecn-and-dcqcn&quot;&gt;ECN and DCQCN&lt;/h2&gt;

&lt;p&gt;PFC stops congestion at the cost of pushing it upstream. ECN is the mechanism
that instead tells the sender to slow down, and it is what you want doing most of
the work, with PFC as the last resort that prevents loss.&lt;/p&gt;

&lt;p&gt;The loop: the switch marks packets ECN-CE when its queue exceeds a threshold; the
receiving NIC sees the mark and returns a Congestion Notification Packet; the
sending NIC reduces its rate for that queue pair and then recovers according to
the DCQCN parameters.&lt;/p&gt;

&lt;p&gt;The switch side is a WRED-style configuration with a minimum threshold, a
maximum threshold and a marking probability, per queue. Those numbers are the
main tuning lever and they interact with buffer size, link rate and round-trip
time, so they are genuinely site-specific — a setting tuned for a two-tier
100 GbE fabric is not right for a 400 GbE one.&lt;/p&gt;

&lt;p&gt;The NIC side lives in sysfs:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;ls&lt;/span&gt; /sys/class/net/enp1s0f0/ecn/roce_np/    &lt;span class=&quot;c&quot;&gt;# notification point (receiver)&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;ls&lt;/span&gt; /sys/class/net/enp1s0f0/ecn/roce_rp/    &lt;span class=&quot;c&quot;&gt;# reaction point (sender)&lt;/span&gt;

&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/net/enp1s0f0/ecn/roce_np/enable/3
&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/net/enp1s0f0/ecn/roce_rp/enable/3
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Both must be enabled on the RoCE priority, on every node. A node with the
reaction point disabled does not slow down when told to, which means it wins
against its better-behaved neighbours and drives the fabric into PFC. One such
node in a large allocation is enough to make the whole job’s timing erratic, and
it is a difficult thing to suspect if you are not looking for it.&lt;/p&gt;

&lt;p&gt;The counters that show the loop working:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ethtool &lt;span class=&quot;nt&quot;&gt;-S&lt;/span&gt; enp1s0f0 | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;cnp|ecn&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;np_ecn_marked_roce_packets: 41229
np_cnp_sent: 40871
rp_cnp_handled: 39655
rp_cnp_ignored: 0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;np_ecn_marked_roce_packets&lt;/code&gt; rising means the switch is marking, so ECN is
configured on the switch and the path. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;np_cnp_sent&lt;/code&gt; should track it closely.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_cnp_handled&lt;/code&gt; rising on the senders means they are reacting. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_cnp_ignored&lt;/code&gt;
being non-zero means CNPs are arriving for queue pairs that no longer exist or
in a state where they cannot be acted on, and a large value is worth
investigating.&lt;/p&gt;

&lt;p&gt;The failure signature to recognize: marking is happening, CNPs are being sent,
and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_cnp_handled&lt;/code&gt; is zero or near it. The sending side is not reacting — the
reaction point is disabled, or the CNPs are being lost or misclassified on the
return path. CNPs themselves are typically marked with a different DSCP and
should be in a high-priority, non-paused class; if they are queued behind the
congestion they are meant to relieve, the control loop does not close.&lt;/p&gt;

&lt;p&gt;There are also hardware counters on the RDMA device side, which are per-device
rather than per-netdev and sometimes easier to correlate with a specific job:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;.&lt;/span&gt; /sys/class/infiniband/mlx5_0/ports/1/hw_counters/&lt;span class=&quot;k&quot;&gt;*&lt;/span&gt; | &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;out_of_buffer|out_of_sequence|packet_seq_err|local_ack_timeout_err|rnr_nak|np_cnp|rp_cnp|roce_adp_retrans&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;packet_seq_err&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;out_of_sequence&lt;/code&gt; are the ones to watch. On a fabric that is
genuinely lossless they stay near zero. On Spectrum-X with packet-level adaptive
routing, out-of-order arrival is expected and handled, so interpret these against
the platform’s own guidance rather than against an InfiniBand intuition.&lt;/p&gt;

&lt;h2 id=&quot;mtu&quot;&gt;MTU&lt;/h2&gt;

&lt;p&gt;This is a small thing that costs a lot of bandwidth.&lt;/p&gt;

&lt;p&gt;RoCE selects a path MTU from the standard set — 256, 512, 1024, 2048 or 4096
bytes — and it cannot exceed what the underlying Ethernet MTU can carry once
headers are accounted for. If any hop in the path is at the default 1500, the
RoCE MTU drops to 1024, and you lose a meaningful fraction of achievable
bandwidth to per-packet overhead and processing.&lt;/p&gt;

&lt;p&gt;The requirement is jumbo frames configured consistently on every NIC and every
switch port in the path, typically 9000. One port left at 1500 — a newly added
uplink, a replaced switch, a host that did not get the configuration — is enough.&lt;/p&gt;

&lt;p&gt;Confirm what was negotiated rather than what was configured:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ib_write_bw &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; 3 &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--report_gbits&lt;/span&gt; &amp;lt;server&amp;gt;   &lt;span class=&quot;c&quot;&gt;# header prints Mtu&lt;/span&gt;
ip &lt;span class=&quot;nb&quot;&gt;link &lt;/span&gt;show enp1s0f0 | &lt;span class=&quot;nb&quot;&gt;grep &lt;/span&gt;mtu
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-x 3&lt;/code&gt; selects a GID index, which on RoCE selects the RoCEv2 IPv4 or IPv6
entry. Get it from:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;show_gids
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Picking the wrong GID index is how people accidentally benchmark RoCEv1, or run
over link-local IPv6 when they meant to use the routed IPv4 address.&lt;/p&gt;

&lt;h2 id=&quot;what-transfers-between-the-two-worlds&quot;&gt;What transfers between the two worlds&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Transfers cleanly:&lt;/strong&gt; the MPI layer entirely — Open MPI, MPICH, UCX, PMIx, the
OSU benchmarks, your application. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perftest&lt;/code&gt; works on both with the addition of a
GID index. NCCL works on both, and refers to RoCE devices through the same
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCCL_IB_HCA&lt;/code&gt; variable, which confuses people into thinking they have InfiniBand.
The layered testing order from the first article in this series applies without
modification.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does not transfer:&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstat&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibdiagnet&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibnetdiscover&lt;/code&gt;,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;opensm&lt;/code&gt; and everything built on LIDs and the subnet manager. There
is no routing engine to choose and no LID budget to compute. In their place you
have per-priority counters, switch telemetry, LLDP for topology discovery, and
whatever your switch vendor’s fabric management provides.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Transfers, but means something different:&lt;/strong&gt; latency figures. RoCE adds
UDP/IP encapsulation and the congestion-control loop, and end-to-end small-message
latency is generally higher than InfiniBand at an equivalent rate. How much
depends on the NIC generation and the switch, so measure it with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_send_lat&lt;/code&gt;
on your own hardware rather than accepting a figure. The gap has narrowed
considerably across recent generations; it has not closed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Transfers, but with a different failure mode:&lt;/strong&gt; congestion. On InfiniBand,
credit-based flow control makes the fabric lossless without configuration, and a
congested fabric slows down. On Ethernet, an incorrectly configured fabric drops,
and RoCE recovery from drops is expensive enough that the performance cliff is
steep rather than gradual. That difference — gradual degradation versus a cliff —
is the single most important thing to internalize when moving MPI workloads from
one to the other.&lt;/p&gt;

&lt;h2 id=&quot;a-minimum-acceptance-check-on-the-ethernet-side&quot;&gt;A minimum acceptance check on the Ethernet side&lt;/h2&gt;

&lt;p&gt;Before believing any MPI number on a RoCE fabric:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlnx_qos -i &amp;lt;dev&amp;gt;&lt;/code&gt; on every node — trust mode and PFC priority identical
everywhere.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ecn/roce_np/enable/&amp;lt;prio&amp;gt;&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ecn/roce_rp/enable/&amp;lt;prio&amp;gt;&lt;/code&gt; set on every node.&lt;/li&gt;
  &lt;li&gt;MTU 9000 confirmed on every host interface and every switch port in the path,
and RoCE MTU 4096 confirmed in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; header.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;show_gids&lt;/code&gt; — the GID index in use is the RoCEv2 entry you intended.&lt;/li&gt;
  &lt;li&gt;Clear counters, run a full-scale alltoall, then check every node for
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;buf_discard&lt;/code&gt; on the RoCE priority and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rp_cnp_handled&lt;/code&gt; moving. Any discard on
the lossless class is a stop-and-fix.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That takes half an hour and it is the difference between a fabric that performs
predictably and one that performs well on Tuesday.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Multi-rail MPI: making the second adapter carry its share</title>
    <link href="https://www.wirewalk.com/writing/mpi-multirail-multileg/"/>
    <published>2026-09-08T09:00:00-04:00</published>
    <updated>2026-09-08T09:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/mpi-multirail-multileg/</id>
    <summary>A second rail is bought as a doubling and usually arrives as something well short of it. The reasons are enumerable, most of them are in the host rather than the fabric, and the first thing to establish is whether MPI is using the second adapter at all.</summary>
    <content type="html">&lt;p&gt;Dual-rail hosts are now the default on anything built for AI and common on
general HPC nodes. Two adapters, two cables, two ports on the leaf, and a
reasonable expectation that a bandwidth-bound job goes twice as fast.&lt;/p&gt;

&lt;p&gt;It rarely does, and the gap between what was bought and what arrives is
attributed to fabric overhead far more often than it deserves. In practice, the
second rail underdelivers for one of a short list of reasons, most of which are
inside the node and all of which are measurable.&lt;/p&gt;

&lt;h2 id=&quot;first-is-the-second-rail-carrying-anything-at-all&quot;&gt;First: is the second rail carrying anything at all?&lt;/h2&gt;

&lt;p&gt;Before tuning anything, establish whether the second adapter is being used. This
is not a rhetorical question — a substantial fraction of dual-rail nodes in
production are running every byte over one rail, and nothing reports it.&lt;/p&gt;

&lt;p&gt;The direct method is the port counters. Clear, run, read:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;d &lt;span class=&quot;k&quot;&gt;in &lt;/span&gt;mlx5_0 mlx5_1&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;echo&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt; before: &quot;&lt;/span&gt;
  &lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/infiniband/&lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;/ports/1/counters/port_xmit_data
&lt;span class=&quot;k&quot;&gt;done

&lt;/span&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node ./osu_bw &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; 8388608:8388608

&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;d &lt;span class=&quot;k&quot;&gt;in &lt;/span&gt;mlx5_0 mlx5_1&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;echo&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt; after: &quot;&lt;/span&gt;
  &lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/infiniband/&lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;/ports/1/counters/port_xmit_data
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlx5_1&lt;/code&gt; did not move, you have a single-rail cluster with a spare adapter in
it. That is a two-minute check and it settles the question that the rest of the
work depends on.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Do this check before any tuning conversation. A node whose second adapter has
never transmitted a byte is not a multi-rail tuning problem, and the discussion
about striping thresholds and rendezvous rails that usually follows is entirely
beside the point. The counter delta settles it in two minutes.&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;The second method is to ask UCX what it selected:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_LOG_LEVEL&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;info &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./osu_bw 2&amp;gt;&amp;amp;1 | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;selected|device|transport|rndv&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Recent UCX releases also expose a protocol selection table that is considerably
more readable than the log:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_PROTO_INFO&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;y ./osu_bw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It prints, per message-size range, which protocol and which devices UCX will
use — including whether a large rendezvous transfer is being striped across
multiple devices or sent over one. That table answers the multi-rail question
directly rather than by inference.&lt;/p&gt;

&lt;h2 id=&quot;telling-mpi-which-devices-to-use&quot;&gt;Telling MPI which devices to use&lt;/h2&gt;

&lt;p&gt;UCX enumerates devices itself and its default selection is often not what you
want, particularly on a node that also has a management NIC, a storage NIC, or
an adapter whose port is administratively down.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ucx_info &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;^#.*Transport|Device:&apos;&lt;/span&gt;
ucx_info &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The device list to hand to MPI is explicit and includes the port number:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 16 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--mca&lt;/span&gt; pml ucx &lt;span class=&quot;nt&quot;&gt;--mca&lt;/span&gt; osc ucx &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_0:1,mlx5_1:1 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_TLS&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;rc_x,sm,self &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./osu_bw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three parameters do most of the work:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_NET_DEVICES&lt;/code&gt; — the allowlist. Naming devices explicitly is better
practice than relying on discovery, because discovery changes when a port goes
down or a card is added.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_MAX_RNDV_RAILS&lt;/code&gt; — how many devices a single large (rendezvous) transfer
may be striped across. The default has historically been 2. If you have four
rails and want them all used by one transfer, this is the setting that is
quietly capping you.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_MAX_EAGER_RAILS&lt;/code&gt; — the same idea for the eager protocol used by smaller
messages.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Confirm the values in effect rather than assuming the defaults:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ucx_info &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;MAX_RNDV_RAILS|MAX_EAGER_RAILS|NET_DEVICES|TLS&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;why-the-second-rail-underdelivers&quot;&gt;Why the second rail underdelivers&lt;/h2&gt;

&lt;p&gt;Work through these in order. They are roughly ordered by how often they are the
answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The transfer is not big enough to stripe.&lt;/strong&gt; Rail striping applies to the
rendezvous protocol. Messages below the rendezvous threshold go over one device.
An application whose messages are all 64 KB will not benefit from a second rail
no matter how it is configured, and that is a correct outcome rather than a
misconfiguration. Check where your application’s message sizes sit before
concluding anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Both adapters are on the same NUMA node.&lt;/strong&gt; Two cards on one socket share that
socket’s PCIe root complex, its memory bandwidth, and its inter-socket link for
any traffic sourced from the other socket. Check:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;d &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; /sys/class/infiniband/mlx5_&lt;span class=&quot;k&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;echo&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;basename&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;: numa_node=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;/device/numa_node&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;done
&lt;/span&gt;lstopo &lt;span class=&quot;nt&quot;&gt;--output-format&lt;/span&gt; txt
nvidia-smi topo &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt;     &lt;span class=&quot;c&quot;&gt;# on GPU nodes: also shows GPU-to-HCA affinity&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nvidia-smi topo -m&lt;/code&gt; is the one to read on a GPU node. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PIX&lt;/code&gt; means the GPU and
the adapter sit under the same PCIe switch, which is the case GPUDirect RDMA is
designed for. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NODE&lt;/code&gt; means same NUMA node but traversing the root complex.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SYS&lt;/code&gt; means the traffic crosses the inter-socket link, and on that pairing the
second rail is a good deal less than a second rail.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The ranks are not bound to the rail that is local to them.&lt;/strong&gt; This is the
common and fixable case. If every rank uses both rails, half of every node’s
traffic crosses the inter-socket link. The better arrangement on a two-socket,
two-rail node is that ranks on socket 0 use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlx5_0&lt;/code&gt; and ranks on socket 1 use
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlx5_1&lt;/code&gt;, with no striping at all. That is not multi-rail in the striping sense;
it is rail affinity, and it usually outperforms naive striping.&lt;/p&gt;

&lt;p&gt;A small wrapper does it, driven by the local rank the launcher exports:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;#!/bin/bash&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# rail-bind.sh - one rail per socket, selected by node-local rank&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;local_rank&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;OMPI_COMM_WORLD_LOCAL_RANK&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;:-${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;SLURM_LOCALID&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;:-&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}}&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;half&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;$((&lt;/span&gt; OMPI_COMM_WORLD_LOCAL_SIZE &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;m&quot;&gt;2&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;))&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;[&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$local_rank&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-lt&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$half&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;then
  &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_0:1
&lt;span class=&quot;k&quot;&gt;else
  &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_1:1
&lt;span class=&quot;k&quot;&gt;fi
&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;exec&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$@&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 32 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:numa &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core ./rail-bind.sh ./osu_mbw_mr
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Measure both arrangements. Which wins depends on the message-size distribution
and on whether the application is bandwidth- or latency-bound, and it is not
predictable from the topology alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The PCIe slots cannot carry it.&lt;/strong&gt; Two NDR400 adapters is 100 GB/s of offered
load in each direction. Confirm both slots trained to their rated speed and
width, and confirm the platform is not sharing lanes between them:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;lspci &lt;span class=&quot;nt&quot;&gt;-vv&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-A1&lt;/span&gt; Mellanox | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;LnkCap|LnkSta&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Both rails land on the same leaf switch and the same uplinks.&lt;/strong&gt; This is the
fabric-side case, and it is why some sites deliberately build the second rail as
a separate subnet with its own switches and its own subnet manager. If both
rails leave the node, arrive at the same leaf, and take the same uplinks to the
spine, the second rail doubles the host’s egress capacity and does nothing for
the fabric’s ability to carry it. Check where each rail’s port actually lands:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibnetdiscover | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-A2&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;ibstat mlx5_0 | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;/Port GUID/{print $3}&apos;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
iblinkinfo &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;node01&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The application is not bandwidth-bound.&lt;/strong&gt; A latency-bound code gains nothing
from a second rail and may lose a little to the extra selection logic. This is
the outcome nobody wants to hear and it is frequently the correct one. It is
also the argument for measuring with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_mbw_mr&lt;/code&gt; at your application’s real
message sizes rather than with a large-message &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; that flatters the
configuration.&lt;/p&gt;

&lt;h2 id=&quot;proving-rdma-is-engaged-not-tcp-over-ipoib&quot;&gt;Proving RDMA is engaged, not TCP over IPoIB&lt;/h2&gt;

&lt;p&gt;Silent fall back to TCP over IPoIB is the single most consequential quiet
failure in this area. Everything works. The job completes. The bandwidth is
roughly a tenth of what it should be and the latency is roughly ten times what
it should be, and both of those are within the range of “the code doesn’t scale
well,” which is what it gets recorded as.&lt;/p&gt;

&lt;p&gt;There are three ways to establish the truth, in increasing order of how
conclusive they are.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read the counters.&lt;/strong&gt; As above: if &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;port_xmit_data&lt;/code&gt; on the InfiniBand device
does not move during the run, the traffic is not on the InfiniBand transport.
Corroborate on the IPoIB side, where it will have gone instead:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/net/ib0/statistics/tx_bytes
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Read the UCX transport selection.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_PROTO_INFO=y&lt;/code&gt; or
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_LOG_LEVEL=info&lt;/code&gt; will name the transport. Seeing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tcp&lt;/code&gt; where you expected
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rc_x&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dc_x&lt;/code&gt; is the finding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Remove the fallback and see if it still runs.&lt;/strong&gt; This is the conclusive test,
and it takes one flag:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_TLS&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;rc_x,sm,self ./osu_bw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That list contains no TCP transport. If the job runs, RDMA was working. If it
fails to establish connections, TCP was carrying the traffic and now you know.
Run this as a gate in your acceptance suite; it converts a silent degradation
into a loud failure, which is the whole objective.&lt;/p&gt;

&lt;p&gt;The usual root causes, once you have established fallback is happening:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Cause&lt;/th&gt;
      &lt;th&gt;Check&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Locked memory limit not raised&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ulimit -l&lt;/code&gt; inside the job, not on the login node&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Device down or not present on some nodes&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstat&lt;/code&gt; across the whole allocation&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UCX_NET_DEVICES&lt;/code&gt; naming a device that does not exist there&lt;/td&gt;
      &lt;td&gt;Node-by-node device enumeration&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Node with a different OFED or missing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rdma-core&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ofed_info -s&lt;/code&gt; fleet-wide&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Firewall or SELinux policy on a subset of nodes&lt;/td&gt;
      &lt;td&gt;Compare a working and a failing node&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ulimit -l&lt;/code&gt; deserves emphasis. Memory registration for RDMA requires locked
pages, and if the limit is not unlimited inside the job’s environment, UCX
cannot register buffers and falls back. The limit set in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/etc/security/limits.d&lt;/code&gt;
applies to login sessions; what matters is the limit inside the scheduler’s
launched process, which can differ. Check it from inside a job:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;srun &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; 2 bash &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;echo $(hostname) $(ulimit -l)&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;rail-optimized-topologies&quot;&gt;Rail-optimized topologies&lt;/h2&gt;

&lt;p&gt;On AI clusters the multi-rail arrangement is often not “two rails to the same
fabric” but a rail-optimized design: each GPU has an adapter, and adapter &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;k&lt;/code&gt; on
every node connects to switch &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;k&lt;/code&gt;. Same-index GPUs across nodes then communicate
within a single switch, which is what makes large all-reduces efficient.&lt;/p&gt;

&lt;p&gt;Two consequences for MPI. First, the affinity between rank, GPU and adapter is
not a tuning preference, it is the design, and getting it wrong sends traffic
across the fabric that was supposed to stay in one switch. Second, if the rails
are separate subnets, each has its own subnet manager and its own LID space, and
tooling that assumes one fabric will report on one of them and silently ignore
the rest.&lt;/p&gt;

&lt;p&gt;Verify the mapping is what the cabling intended before benchmarking anything:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;nvidia-smi topo &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;d &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; /sys/class/infiniband/mlx5_&lt;span class=&quot;k&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;basename&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
  &lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;%s numa=%s lid=%s sm_lid=%s\n&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$d&lt;/span&gt;/device/numa_node&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;ibstat &lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt; 1 | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;/Base lid/{print $3}&apos;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;ibstat &lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt; 1 | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;/SM lid/{print $3}&apos;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-to-measure-and-in-what-order&quot;&gt;What to measure, and in what order&lt;/h2&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw -d mlx5_0&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-d mlx5_1&lt;/code&gt; separately. Each rail alone should
reach the same number. If one is lower, stop — that is a link or slot problem
and it is not a multi-rail question.&lt;/li&gt;
  &lt;li&gt;Both rails simultaneously, as two concurrent &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; processes pinned to
the correct sockets. This is the host’s true aggregate ceiling and it is the
number multi-rail MPI is trying to reach.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; with one rail, then with both. Compare against step 2.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_mbw_mr&lt;/code&gt; at your application’s message sizes, one rail versus both, with
and without rail affinity.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 2 is the one that gets skipped, and it is the one that tells you whether the
shortfall is in the host or in MPI. If two concurrent &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; processes
cannot reach twice a single rail, no MPI configuration will, and you have a
PCIe, NUMA or memory-bandwidth investigation rather than a UCX one.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Fan-out tests: what one-to-many finds that point-to-point cannot</title>
    <link href="https://www.wirewalk.com/writing/mpi-fanout-collective-scaling/"/>
    <published>2026-09-05T14:00:00-04:00</published>
    <updated>2026-09-05T14:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/mpi-fanout-collective-scaling/</id>
    <summary>A single flow takes a single path, so a pair test is indifferent to almost everything that is wrong with a fabric. Fan-out puts many flows on it at once, and that is when static routing, ISL capacity and one slow node show up.</summary>
    <content type="html">&lt;p&gt;A point-to-point bandwidth test between two nodes exercises one path. Under
static routing — which is what InfiniBand does by default — that path is fixed
by the subnet manager’s forwarding tables, and one flow on one path is
indifferent to how well or badly those tables distribute traffic overall. You
can have a routing engine that has collapsed eight parallel inter-switch links
onto three, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; between a well-chosen pair will report a perfect
number, every time, reproducibly.&lt;/p&gt;

&lt;p&gt;That is the gap fan-out tests exist to close. The moment you put many
simultaneous flows on the fabric, path selection stops being a detail and
becomes the dominant term.&lt;/p&gt;

&lt;h2 id=&quot;the-three-shapes&quot;&gt;The three shapes&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;One-to-many.&lt;/strong&gt; One node sends to N others concurrently. This finds the sender’s
own ceiling — its adapter, its PCIe slot, its memory bandwidth — and it finds
the first hop out of that node’s leaf switch. It is the right test for a storage
node, a parameter server, or any rank-zero-heavy pattern.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Many-to-one.&lt;/strong&gt; N nodes send to one. This is the incast pattern, and it is a
different test, not a mirror image. It is where switch buffering, flow control
and congestion control get exercised, and it is where a fabric that looks fine
under one-to-many falls apart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Many-to-many.&lt;/strong&gt; Every node to every node. This is what an alltoall does, and it
is the test that actually measures the fabric rather than any endpoint.&lt;/p&gt;

&lt;p&gt;Run all three. A cluster that passes one-to-many and fails many-to-many has a
fabric problem. A cluster that fails all three at the same node has a node
problem. That distinction is available in about twenty minutes and is otherwise
the subject of a long argument.&lt;/p&gt;

&lt;h2 id=&quot;osu_mbw_mr-and-why-message-rate-is-a-separate-ceiling&quot;&gt;osu_mbw_mr, and why message rate is a separate ceiling&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_mbw_mr&lt;/code&gt; is the multiple-bandwidth / message-rate test. It takes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2N&lt;/code&gt; ranks,
makes the first &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N&lt;/code&gt; senders and the second &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N&lt;/code&gt; receivers, pairs them up, and
runs all pairs concurrently. It reports aggregate bandwidth and aggregate
message rate for each message size.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# 16 ranks per node, one node sending, one node receiving&lt;/span&gt;
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 32 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:16:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_0:1 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./osu_mbw_mr
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two things make it worth running that a pair test is not.&lt;/p&gt;

&lt;p&gt;First, it is the only cheap way to find a &lt;strong&gt;per-node ceiling that is below the
per-link rate&lt;/strong&gt;. A single pair of ranks frequently cannot saturate a modern
adapter — one rank, one queue pair, one core is not enough to fill NDR400. If
you conclude from a single-pair &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; that the link is at half rate, you may
simply be measuring one core. Increase the concurrency and the number moves. If
it does not move, the ceiling is real and it is in the node.&lt;/p&gt;

&lt;p&gt;Second, the &lt;strong&gt;message rate&lt;/strong&gt; column is a ceiling that bandwidth never reveals.
Adapters have a maximum packets-per-second they can process, and applications
with many small messages hit that ceiling while the bandwidth graph looks
comfortable. If the message rate flattens as you add pairs while bandwidth is
nowhere near the link rate, the constraint is packet processing, not bytes, and
no amount of fabric work changes it. The fixes live elsewhere: message
aggregation, larger buffers, or a different communication pattern.&lt;/p&gt;

&lt;p&gt;Sweep the concurrency deliberately rather than accepting the default:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;ppn &lt;span class=&quot;k&quot;&gt;in &lt;/span&gt;1 2 4 8 16 32&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;echo&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;== ppn=&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$ppn&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
  mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;$((&lt;/span&gt;ppn&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;ppn&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core ./osu_mbw_mr
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The shape of that sweep is the useful output. Bandwidth rises and plateaus;
where it plateaus is your real per-node number. Message rate rises and plateaus
somewhere else entirely.&lt;/p&gt;

&lt;h2 id=&quot;where-fan-out-exposes-the-fabric&quot;&gt;Where fan-out exposes the fabric&lt;/h2&gt;

&lt;p&gt;Here is the mechanism, because it is the part that is worth understanding rather
than memorizing.&lt;/p&gt;

&lt;p&gt;Under static routing, the subnet manager assigns each destination LID an output
port on every switch. Every flow to a given destination therefore leaves a given
switch through the same port, regardless of how busy that port is. On a
leaf-and-spine fabric with, say, eight uplinks from a leaf to the spine tier, the
routing engine distributes destination LIDs across those eight uplinks — and how
evenly it does that is the whole question.&lt;/p&gt;

&lt;p&gt;With one flow, you use one uplink and it is fine. With sixty-four concurrent
flows, if the distribution assigned twenty of them to one uplink and two to
another, that first uplink is oversubscribed by a factor of ten and the other is
idle. The aggregate number you measure is set by the busiest uplink, not by the
total capacity.&lt;/p&gt;

&lt;p&gt;This is not hypothetical and it is not rare. It is the normal consequence of a
routing engine that declined the topology and fell back to something simpler, of
a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;updn&lt;/code&gt; configuration with no root GUID list, or of a fabric whose leaf uplink
counts are not uniform. The article on OpenSM routing engines in this series
covers how to confirm which engine is actually running.&lt;/p&gt;

&lt;h3 id=&quot;measuring-the-distribution-directly&quot;&gt;Measuring the distribution directly&lt;/h3&gt;

&lt;p&gt;Do not infer it from bandwidth. Read the switch counters.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# baseline&lt;/span&gt;
ibclearerrors
&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;lid &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$SPINE_LIDS&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;perfquery &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$lid&lt;/span&gt;   &lt;span class=&quot;c&quot;&gt;# extended (64-bit) counters, all ports&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; /var/tmp/counters-before.txt

mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 512 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:node ./osu_alltoall &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; 1048576:1048576

&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;lid &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$SPINE_LIDS&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do &lt;/span&gt;perfquery &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$lid&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;done&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; /var/tmp/counters-after.txt
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitData&lt;/code&gt; is in 32-bit words on most implementations, so multiply by four
for bytes; what you care about is the ratio between uplinks, not the absolute
value. Take the delta per port and look at the spread. An engine distributing
well produces uplink deltas within a few percent of each other. An engine that
is not produces a spread you can see without a calculator, and that spread is
your missing bandwidth.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt; on the same ports is the corroborating evidence. It counts ticks
during which the port had data queued and no credit to transmit. Large and
uneven &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt; across parallel uplinks is congestion concentrated on a
subset of links, which is exactly the signature of a routing imbalance.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Use the 64-bit extended counters (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery -x&lt;/code&gt;, or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery -x -a&lt;/code&gt;). The
legacy 32-bit &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitData&lt;/code&gt; counter wraps in seconds at modern link rates, and
a wrapped counter produces a delta that is not merely wrong but arbitrarily
wrong — sometimes negative, sometimes plausible. This has cost people real
conclusions.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;reading-the-knee&quot;&gt;Reading the knee&lt;/h2&gt;

&lt;p&gt;Fan-out results are curves, and the useful information is in where the curve
bends rather than in any single point.&lt;/p&gt;

&lt;p&gt;Run the same collective at a fixed message size across a doubling sequence of
node counts, holding ranks-per-node constant:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;n &lt;span class=&quot;k&quot;&gt;in &lt;/span&gt;2 4 8 16 32 64&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;srun &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--ntasks-per-node&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;8 ./osu_allreduce &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; 4194304:4194304
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;For &lt;strong&gt;allreduce&lt;/strong&gt;, the expected shape depends on the algorithm. NCCL-style and
MPI ring implementations move &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2(n-1)/n&lt;/code&gt; of the buffer per rank, which
approaches a constant as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n&lt;/code&gt; grows, so time at a large fixed message size should
flatten out rather than grow. Tree and recursive-doubling implementations trade
that for a latency term that grows as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;log(n)&lt;/code&gt;, which dominates at small message
sizes. So: at large messages you expect a plateau, at small messages you expect
a gentle logarithmic rise. Anything that rises &lt;strong&gt;linearly&lt;/strong&gt; with node count is
not the algorithm — it is the fabric or a straggler.&lt;/p&gt;

&lt;p&gt;For &lt;strong&gt;alltoall&lt;/strong&gt;, every rank sends to every other rank, so the total bytes
crossing the fabric grows as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n²&lt;/code&gt; while the bisection capacity of a
non-blocking fat tree grows as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n&lt;/code&gt;. On a genuinely non-blocking fabric, per-rank
alltoall time at a fixed per-peer message size should stay roughly flat. The
node count where it stops being flat is a measurement of where your fabric stops
being non-blocking, and that is a genuinely useful number to own.&lt;/p&gt;

&lt;p&gt;Three distinct knee shapes and what each means:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Shape&lt;/th&gt;
      &lt;th&gt;Reading&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Flat, then a step at a specific node count&lt;/td&gt;
      &lt;td&gt;You crossed a topology boundary — filled a leaf, started using spine uplinks, or crossed into a second island&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Flat, then a smooth linear rise&lt;/td&gt;
      &lt;td&gt;Oversubscription. The fabric’s aggregate capacity is now the limit&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Noisy from the start, no clean curve&lt;/td&gt;
      &lt;td&gt;Not a scaling result at all. Something is varying between runs — go find it before drawing any curve&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;That third row is the common one and it is worth taking seriously. If you cannot
reproduce a point to within a few percent across three runs, you do not yet have
a measurement, and fitting a curve through it produces a confident wrong
conclusion.&lt;/p&gt;

&lt;h2 id=&quot;controlling-for-allocation&quot;&gt;Controlling for allocation&lt;/h2&gt;

&lt;p&gt;A scaling curve is only comparable if the node set is comparable. Two runs at 32
nodes that land on different parts of the fabric are two different experiments.&lt;/p&gt;

&lt;p&gt;Pin the placement and record it:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;srun &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; 32 &lt;span class=&quot;nt&quot;&gt;--ntasks-per-node&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;8 &lt;span class=&quot;nt&quot;&gt;--nodelist&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;cat &lt;/span&gt;nodes.32&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt; ./osu_alltoall
scontrol show hostnames &lt;span class=&quot;nv&quot;&gt;$SLURM_JOB_NODELIST&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; run-&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;date&lt;/span&gt; +%s&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;.nodes
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If your scheduler has topology awareness configured, use it — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--switches=1&lt;/code&gt;
asks Slurm for an allocation within a single switch, and the difference between
a within-leaf run and a cross-spine run at the same node count is itself a
measurement worth taking.&lt;/p&gt;

&lt;p&gt;And record the node list with every result. The single most common reason a
benchmark “regressed” is that it ran somewhere else.&lt;/p&gt;

&lt;h2 id=&quot;the-one-slow-node-problem&quot;&gt;The one-slow-node problem&lt;/h2&gt;

&lt;p&gt;A collective completes when its slowest participant completes. That makes
fan-out tests exquisitely sensitive to a single degraded node — which is
useful, and also dangerous, because the symptom presents as a fabric-wide
problem.&lt;/p&gt;

&lt;p&gt;The signature: results that vary run to run in a way that correlates with the
allocation rather than with anything you changed. Some 32-node runs are fine and
some are 30 percent down, and which is which tracks whether a particular node
was included.&lt;/p&gt;

&lt;p&gt;The bisection method finds it quickly. Split the node set in half, run both
halves, keep the slow half, repeat. Six rounds covers 64 nodes. It is crude and
it is faster than reasoning about it.&lt;/p&gt;

&lt;p&gt;The systematic alternative is a pairwize sweep — every node against a fixed
reference node, using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; rather than MPI so you are testing one layer
at a time:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;n &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;scontrol show hostnames &lt;span class=&quot;nv&quot;&gt;$SLURM_JOB_NODELIST&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;ssh &lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;ib_write_bw -d mlx5_0 -F -s 1048576 -D 5 --report_gbits &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$n&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;/1048576/ {print n, $4}&apos;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;sort&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-k2&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;head&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The bottom of that sorted list is your answer, and it is usually one node with a
narrow link, a downtrained slot, or a transceiver on its way out. It is worth
running this on a schedule rather than during an incident: a node that
degrades quietly poisons every synchronized job it touches, and the cost of
finding out late is measured in withdrawn results, not in wasted cycles.&lt;/p&gt;

&lt;h2 id=&quot;what-to-keep&quot;&gt;What to keep&lt;/h2&gt;

&lt;p&gt;Per fabric, keep: the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_mbw_mr&lt;/code&gt; bandwidth and message-rate plateaus with the
concurrency at which each was reached; the alltoall and allreduce curves across
your real node-count range with the node lists attached; and the per-uplink
counter deltas from one representative alltoall.&lt;/p&gt;

&lt;p&gt;The counter deltas are the item people omit and the one that ages best. Bandwidth
numbers move when hardware changes. A routing distribution that was even in
January and is lopsided in June is a change in the fabric, and having the
January baseline is the difference between knowing that and arguing about it.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>A layered order for MPI testing, and what each layer eliminates</title>
    <link href="https://www.wirewalk.com/writing/mpi-layered-test-methodology/"/>
    <published>2026-09-05T09:00:00-04:00</published>
    <updated>2026-09-05T09:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/mpi-layered-test-methodology/</id>
    <summary>Every layer of an MPI stack has a ceiling set by the layer beneath it. Test from the top and you will spend a week tuning something that was never the constraint. Test from the bottom and most investigations end in an hour.</summary>
    <content type="html">&lt;p&gt;An MPI job that runs slower than expected is a question with about nine
plausible answers, spread across four layers of stack that were installed by
different people at different times. The reason these investigations run long is
almost never that the answer is subtle. It is that the search was unordered.&lt;/p&gt;

&lt;p&gt;The discipline that fixes it is simple to state. Each layer has a performance
ceiling determined by the layer below it. Measure the bottom first, establish
what the ceiling is, and only then ask whether the layer above is achieving it.
A number at layer four means nothing until you know what layers one through
three can deliver.&lt;/p&gt;

&lt;h2 id=&quot;the-layers&quot;&gt;The layers&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Layer&lt;/th&gt;
      &lt;th&gt;Tool&lt;/th&gt;
      &lt;th&gt;What a clean result eliminates&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;0. Inventory&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibv_devinfo&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ofed_info&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;flint&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lspci&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Version skew, downtrained slots, wrong adapter&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;1. Link&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstat&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlxlink&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Width, rate, cabling, physical errors&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;2. Verbs&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_read_bw&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_send_lat&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;The entire MPI and UCX stack&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;3. MPI point-to-point&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_latency&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bibw&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_mbw_mr&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Transport selection, protocol thresholds, pinning&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;4. Collectives&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_allreduce&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_alltoall&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_barrier&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Routing distribution, jitter, one slow rank&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;5. Application&lt;/td&gt;
      &lt;td&gt;The application&lt;/td&gt;
      &lt;td&gt;Nothing — it is what you were asked about&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Nothing here is novel. The value is entirely in refusing to skip.&lt;/p&gt;

&lt;h2 id=&quot;layer-0-inventory-before-measurement&quot;&gt;Layer 0: inventory before measurement&lt;/h2&gt;

&lt;p&gt;Half the MPI investigations that get escalated are version skew, and skew is
free to check before anything is measured. You want, across every node in the
allocation: the same OFED or DOCA-OFED release, the same adapter firmware, the
same MPI build, and PCIe slots that trained to their rated speed and width.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;pdsh &lt;span class=&quot;nt&quot;&gt;-w&lt;/span&gt; node[01-64] &lt;span class=&quot;s1&quot;&gt;&apos;ofed_info -s; ibv_devinfo -d mlx5_0 | grep -E &quot;fw_ver|board_id&quot;&apos;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | dshbak &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;board_id&lt;/code&gt; is the PSID, and it identifies the adapter model and OEM variant.
Two adapters with the same marketing name and different PSIDs are different
parts, take different firmware images, and do not necessarily behave alike.&lt;/p&gt;

&lt;p&gt;Then the slot. This is the check people skip, and it is the one that most often
turns a two-week investigation into a two-minute one:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;lspci &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; &lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/infiniband/mlx5_0/device/uevent | &lt;span class=&quot;nb&quot;&gt;grep &lt;/span&gt;PCI_SLOT_NAME | &lt;span class=&quot;nb&quot;&gt;cut&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f2&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-vv&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;LnkCap:|LnkSta:&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;LnkCap:	Port #0, Speed 32GT/s, Width x16, ASPM not supported
LnkSta:	Speed 16GT/s, Width x8, TrErr- Train- SlotClk+ DLActive+
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That adapter is capable of PCIe Gen5 x16 and is running at Gen4 x8. Nothing will
report an error. The link will come up at its full InfiniBand rate. The host bus
will cap you at roughly a quarter of what the card can do, and every measurement
above this layer will be consistent, reproducible and wrong.&lt;/p&gt;

&lt;p&gt;The arithmetic is worth internalizing because it recurs across generations. A
PCIe Gen3 x16 slot delivers usefully under 13 GB/s after encoding and protocol
overhead, which is approximately one HDR100 link. Put an HDR200 adapter in one
and you have bought an HDR100 adapter. Gen4 x16 is roughly the right size for
HDR200 or NDR200. NDR400 wants Gen5 x16. XDR at 800 Gb/s is 100 GB/s in each
direction, which is more than a Gen5 x16 slot can carry, so on XDR hardware the
slot generation is not a footnote — check it explicitly and check what the
platform actually populated, not what the chassis datasheet says is possible.&lt;/p&gt;

&lt;h2 id=&quot;layer-1-the-link&quot;&gt;Layer 1: the link&lt;/h2&gt;

&lt;p&gt;Two questions only: is every port at the width and rate you paid for, and is any
port accumulating errors.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibstat mlx5_0 1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;Port 1:
	State:            Active
	Physical state:   LinkUp
	Rate:             200
	Base lid:         91
	LMC:              0
	SM lid:           1
	Link layer:       InfiniBand
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rate: 200&lt;/code&gt; here is the aggregate link rate in Gb/s. Note that for the same
link, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt; reports the &lt;strong&gt;per-lane&lt;/strong&gt; rate and the width separately, and
conflating the two is a genuine and expensive misread:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;  91    1[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  12   9[  ] &quot;leaf01&quot; ( )
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That is four lanes at 53.125 Gb/s each — an HDR200 link, not a 53 Gb/s link. On
NDR the same field reads &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;106.25 Gbps&lt;/code&gt; and the link is NDR400. Someone reading
that column as the link rate concludes their NDR fabric is running at 100 Gb/s
and opens a support case.&lt;/p&gt;

&lt;p&gt;The scan that matters is not any single port but the whole allocation at once,
looking for anything that is not at full width:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;iblinkinfo &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-vE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;4X *(25\.78125|53\.125|106\.25|212\.5)&apos;&lt;/span&gt; 
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A port that negotiated &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;1X&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2X&lt;/code&gt; is Active, passes traffic, answers ping, and
runs your job at a fraction of the speed. It will not appear in any error log.&lt;/p&gt;

&lt;p&gt;For physical link quality on modern adapters, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlxlink&lt;/code&gt; is more informative than
the counters:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mlxlink &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; 1 &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--show_fec&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--show_eye&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It reports the negotiated speed, the FEC mode, raw and effective BER, and the
transceiver’s vendor and part number. A link that is up but marginal shows as a
raw BER several orders of magnitude worse than its neighbours while the
effective BER stays clean, because FEC is absorbing it. That link works until
the day it is under load, and then it is a fabric-wide mystery.&lt;/p&gt;

&lt;p&gt;Finally, clear the counters, run something, and read them back. Counters that
have been accumulating since the last power cycle tell you nothing about today:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibclearerrors
&lt;span class=&quot;c&quot;&gt;# ... run a representative workload ...&lt;/span&gt;
ibqueryerrors &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; PortXmitWait,PortRcvErrors,SymbolErrorCounter,LinkDownedCounter
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt; is not an error. It counts ticks where a port had data to send and
no credit to send it, which is congestion, and it belongs to layer four rather
than here. Note it and move on.&lt;/p&gt;

&lt;h2 id=&quot;layer-2-verbs-without-mpi-in-the-picture&quot;&gt;Layer 2: verbs, without MPI in the picture&lt;/h2&gt;

&lt;p&gt;This is the layer people skip, and skipping it is why so many fabric problems
get diagnosed as MPI problems. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perftest&lt;/code&gt; talks to the adapter directly. If
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; is slow, no MPI tuning parameter is going to help you, and you
have just eliminated UCX, PMIx, Open MPI, process binding and the application in
one command.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# server&lt;/span&gt;
ib_write_bw &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; 1 &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--report_gbits&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# client&lt;/span&gt;
ib_write_bw &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; 1 &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--report_gbits&lt;/span&gt; node02
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-F&lt;/code&gt; suppresses the CPU-frequency warning on machines with a scaling governor,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-a&lt;/code&gt; sweeps all message sizes, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-i 1&lt;/code&gt; selects the physical port. Add &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-x &amp;lt;n&amp;gt;&lt;/code&gt; to
select a GID index when you are on RoCE rather than InfiniBand.&lt;/p&gt;

&lt;p&gt;Read the header before the numbers. It tells you the MTU that was negotiated,
the transport, and the connection type:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt; Number of qps   : 1        Transport type : IB
 Connection type : RC       Using SRQ      : OFF
 Mtu             : 4096[B]
 Link type       : IB
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;An MTU of 1024 where you expected 4096 is a finding in itself, and on RoCE it
usually means the network MTU is not jumbo end to end.&lt;/p&gt;

&lt;p&gt;Three runs, in this order, on the same pair:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; — one-directional RDMA write. The cleanest bandwidth number the
hardware can produce.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_read_bw&lt;/code&gt; — RDMA read. Reads are round-trip and depend on how many
outstanding reads the adapter allows, so this is normally lower than write and
a good deal more sensitive to the fabric.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_send_lat -a&lt;/code&gt; — send/receive latency across sizes. Sub-microsecond at small
sizes on a single-hop modern fabric; each additional switch hop adds a
well-defined increment you can measure directly by comparing an intra-leaf pair
against a cross-spine pair.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What “good” is here is a property of your hardware, not something to assert from
an article. The check you can do without a reference number is arithmetic: take
the signalling rate from layer one, subtract the line encoding overhead, and see
whether &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; at large message sizes lands close to it. If it lands at
half, look at the width and the PCIe slot before looking anywhere else.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Run layer two between several different node pairs, chosen deliberately: two
nodes on the same leaf, two nodes on different leaves, and the two nodes
furthest apart in the topology. One pair tells you almost nothing. Three pairs
chosen by topology tell you whether the problem is a node, a link, or the core
of the fabric.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;layer-3-mpi-point-to-point&quot;&gt;Layer 3: MPI point-to-point&lt;/h2&gt;

&lt;p&gt;Now, and only now, introduce MPI. If layer two was clean and layer three is not,
the problem is in the MPI stack, and that is a much smaller search space than
“somewhere in the cluster.”&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_0:1 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./osu_latency
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;UCX_NET_DEVICES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;mlx5_0:1 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./osu_bw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Compare &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; against the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; number from the same pair. A gap of a
few percent is the MPI protocol overhead and is expected. A gap of a factor of
several means MPI is not using the transport you think it is, and the single
most common reason is a silent fall back to TCP over IPoIB.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bibw&lt;/code&gt; runs traffic in both directions at once and is the test that finds
half-duplex-like behavior and PCIe bottlenecks, because it is the first test
that asks the host bus for full bandwidth in both directions simultaneously. A
machine that does well on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; and poorly on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bibw&lt;/code&gt; is usually telling
you about its slot, not its fabric.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_mbw_mr&lt;/code&gt; is the one that belongs at the end of this layer. It runs multiple
concurrent pairs and reports both aggregate bandwidth and message rate. It is
the bridge into layer four, and it is where a per-node ceiling that no
single-pair test can see becomes visible.&lt;/p&gt;

&lt;p&gt;Pinning is not optional at this layer. An unpinned rank measures the scheduler.
Confirm the adapter’s NUMA affinity and place accordingly:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/infiniband/mlx5_0/device/numa_node
lstopo &lt;span class=&quot;nt&quot;&gt;--output-format&lt;/span&gt; txt
mpirun &lt;span class=&quot;nt&quot;&gt;--report-bindings&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 ...
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--report-bindings&lt;/code&gt; prints where each rank actually landed. Read it. Do not
assume the binding you asked for is the binding you got, particularly under a
scheduler that has its own cgroup opinions.&lt;/p&gt;

&lt;h2 id=&quot;layer-4-collectives&quot;&gt;Layer 4: collectives&lt;/h2&gt;

&lt;p&gt;Collectives are where fabric problems that point-to-point cannot see finally
appear, because they are the first thing that puts many flows on the fabric at
once and the first thing that is gated by its slowest participant.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 512 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core ./osu_alltoall
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 512 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core ./osu_allreduce
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 512 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core ./osu_barrier
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A collective that degrades while layers one through three stayed clean is
telling you about routing distribution, congestion, or jitter — never about the
link. That distinction is the whole reason to have run the lower layers first.
Fan-out behavior and how to read the knee in the scaling curve is a large
enough topic that it has its own article in this series.&lt;/p&gt;

&lt;h2 id=&quot;layer-5-the-application&quot;&gt;Layer 5: the application&lt;/h2&gt;

&lt;p&gt;By the time you get here you know the ceiling at every layer beneath, which
means an application number can finally be interpreted. If the application is at
sixty percent of what layer four delivers, that is an application question:
message sizes, communication/computation overlap, decomposition. If it is at
five percent, something below is still wrong and the earlier layers were not
tested carefully enough.&lt;/p&gt;

&lt;h2 id=&quot;what-testing-out-of-order-actually-costs&quot;&gt;What testing out of order actually costs&lt;/h2&gt;

&lt;p&gt;The failure mode is not that you fail to find the problem. It is that you find a
problem, and it is real, and it is not the one that is hurting you.&lt;/p&gt;

&lt;p&gt;Start at the application, and its message sizes will look suboptimal, because
application message sizes usually are. You will spend a week on the
decomposition. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;1X&lt;/code&gt; link stays there the whole time. Start at the
collectives, and you will conclude the routing engine is badly distributed,
which it may well be, and you will schedule a maintenance window to change it
while the downtrained PCIe slot continues to cap the node.&lt;/p&gt;

&lt;p&gt;Both of those are true findings. Neither is the constraint. The ordering exists
because the only way to know whether a finding is the constraint is to know the
ceiling underneath it.&lt;/p&gt;

&lt;h2 id=&quot;recording-it&quot;&gt;Recording it&lt;/h2&gt;

&lt;p&gt;The output of this exercise is not a diagnosis, it is a baseline. Write down,
in version control, for a named node pair and a named node set: the OFED and
firmware versions, the PCIe link status, the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_read_bw&lt;/code&gt;
plateaus, the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_latency&lt;/code&gt; small-message number, and the collective curves at
the job sizes you actually run.&lt;/p&gt;

&lt;p&gt;That file is what converts the next incident from an investigation into a diff.
It takes half a day to produce on a machine that is known good, and there is no
other moment when producing it is as cheap.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Mounting GPFS from a separate client cluster, end to end</title>
    <link href="https://www.wirewalk.com/writing/gpfs-client-cluster-remote-mount/"/>
    <published>2026-09-04T10:00:00-04:00</published>
    <updated>2026-09-04T10:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/gpfs-client-cluster-remote-mount/</id>
    <summary>Keeping compute nodes in their own cluster and remote-mounting the filesystem is the right default for almost every site. The setup is four commands. The parts that bite are the ones that only show up under load, months later.</summary>
    <content type="html">&lt;p&gt;A GPFS filesystem does not have to be mounted by the cluster that owns it. The
owning cluster exports it, a separate client cluster mounts it, and the two are
administered independently. This is the multicluster or remote mount model, and for
most sites with distinct storage and compute hardware it should be the default.&lt;/p&gt;

&lt;p&gt;The mechanics are four commands. What follows is the reasoning for the model, the
full sequence, and the failures that appear later — including one that will kill
every running job on a node without logging anything that points at the cause.&lt;/p&gt;

&lt;h2 id=&quot;why-separate-the-clusters-at-all&quot;&gt;Why separate the clusters at all&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Blast radius.&lt;/strong&gt; A client cluster can be restarted, upgraded, reinstalled or
broken without touching the storage cluster. Where compute and storage share a
cluster, every compute-side action is a storage-side risk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Version independence.&lt;/strong&gt; The two clusters run their own minimum release levels.
You can upgrade compute nodes without moving the storage cluster, which matters
because the storage cluster is the one you least want to touch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quorum separation.&lt;/strong&gt; Compute nodes come and go. Storage nodes must not. Sharing a
cluster means the quorum design has to tolerate the churn of the noisiest half.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Administrative boundary.&lt;/strong&gt; The storage team owns the storage cluster’s
configuration. A compute-side administrator adding a node cannot alter the storage
cluster’s quorum, licensing or tunables by accident.&lt;/p&gt;

&lt;p&gt;The cost is one more cluster to manage and an authentication relationship to
maintain. It is worth it.&lt;/p&gt;

&lt;h2 id=&quot;create-the-client-cluster&quot;&gt;Create the client cluster&lt;/h2&gt;

&lt;p&gt;Install the same GPFS packages, build the portability layer, then create a cluster
containing only the clients:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmcrcluster &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; clientnodes &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; /usr/bin/ssh &lt;span class=&quot;nt&quot;&gt;-R&lt;/span&gt; /usr/bin/scp &lt;span class=&quot;nt&quot;&gt;-C&lt;/span&gt; compute-cluster
mmchlicense client &lt;span class=&quot;nt&quot;&gt;--accept&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; all
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Note the license designation. Nodes that only mount a filesystem take a &lt;strong&gt;client&lt;/strong&gt;
license; nodes that serve NSDs take a server license. Getting this wrong is a
compliance problem rather than a technical one, but it is easier to set correctly
than to correct in an audit.&lt;/p&gt;

&lt;p&gt;Give the client cluster its own quorum — three nodes, chosen for stability rather
than convenience. Login nodes and infrastructure nodes make better quorum members
than batch compute nodes that are rebooted between jobs.&lt;/p&gt;

&lt;h2 id=&quot;exchange-keys&quot;&gt;Exchange keys&lt;/h2&gt;

&lt;p&gt;Each cluster generates a key and receives the other’s public half.&lt;/p&gt;

&lt;p&gt;On the &lt;strong&gt;storage&lt;/strong&gt; cluster:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmauth genkey new
mmauth update &lt;span class=&quot;nb&quot;&gt;.&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; AUTHONLY
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;On the &lt;strong&gt;client&lt;/strong&gt; cluster:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmauth genkey new
mmauth update &lt;span class=&quot;nb&quot;&gt;.&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; AUTHONLY
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Then copy each cluster’s public key to the other. On the storage cluster, register
the client and grant access to the filesystem:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmauth add compute-cluster.example &lt;span class=&quot;nt&quot;&gt;-k&lt;/span&gt; /path/to/client_id_rsa.pub
mmauth grant compute-cluster.example &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; fs1 &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; rw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-a rw&lt;/code&gt; can be &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ro&lt;/code&gt; for a read-only export, which is worth using where it fits —
an analysis cluster that only reads a curated dataset does not need write access to
it.&lt;/p&gt;

&lt;h2 id=&quot;define-the-remote-cluster-and-filesystem-on-the-client&quot;&gt;Define the remote cluster and filesystem on the client&lt;/h2&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmremotecluster add storage-cluster.example &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; storage01,storage02,storage03 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;-k&lt;/span&gt; /path/to/storage_id_rsa.pub

mmremotefs add gpfs0 &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; fs1 &lt;span class=&quot;nt&quot;&gt;-C&lt;/span&gt; storage-cluster.example &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;-T&lt;/span&gt; /gpfs/gpfs0 &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; rw,atime,mtime &lt;span class=&quot;nt&quot;&gt;-A&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-n&lt;/code&gt; list is the contact nodes: the storage-cluster nodes the client will
approach to find the filesystem. Give it several. A single contact node is a single
point of failure for &lt;strong&gt;mounting&lt;/strong&gt;, though not for I/O once mounted.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-A yes&lt;/code&gt; sets automount, so the filesystem mounts at daemon start.&lt;/p&gt;

&lt;h2 id=&quot;understand-the-local-name-and-the-remote-name--they-are-not-the-same-thing&quot;&gt;Understand the local name and the remote name — they are not the same thing&lt;/h2&gt;

&lt;p&gt;This is the detail that causes the most confusion later, and it is also the feature
that makes filesystem rebuilds survivable.&lt;/p&gt;

&lt;p&gt;In &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmremotefs add gpfs0 -f fs1&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpfs0&lt;/code&gt; is the &lt;strong&gt;local device name&lt;/strong&gt; on the client
cluster and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fs1&lt;/code&gt; is the &lt;strong&gt;remote name&lt;/strong&gt; on the storage cluster. They are
independent. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmremotefs show&lt;/code&gt; displays the mapping:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;Local Name  Remote Name  Cluster name              Mount Point
gpfs0       fs1          storage-cluster.example   /gpfs/gpfs0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The practical consequence: &lt;strong&gt;the storage cluster can rebuild and rename its
filesystem without any client-side path changing.&lt;/strong&gt; Build the replacement
filesystem under a new name, migrate, then repoint the remote name. Clients keep
their local device name and their mount point, so module paths, job scripts,
provisioning references and automation all continue to resolve.&lt;/p&gt;

&lt;p&gt;The trap that comes with it: anything that consumes the filesystem &lt;em&gt;name&lt;/em&gt; rather
than the mount point sees the new name. GPFS callbacks receive &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%fsName&lt;/code&gt;, and a
callback guard written against the old name silently stops matching. Cron entries
that name the device — snapshot scripts are the usual example — keep pointing at a
device that no longer exists. Before any rename, grep your automation for the
device name; after it, verify the callbacks still fire.&lt;/p&gt;

&lt;h2 id=&quot;configure-the-client-network&quot;&gt;Configure the client network&lt;/h2&gt;

&lt;p&gt;Clients need &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsPorts&lt;/code&gt; set for their own adapters, and it is not inherited from
anywhere useful:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmchconfig &lt;span class=&quot;nv&quot;&gt;verbsPorts&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;mlx5_0/1&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; clientnodes
mmchconfig &lt;span class=&quot;nv&quot;&gt;verbsRdmaSend&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Then verify RDMA is actually carrying data, rather than assuming:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmdiag &lt;span class=&quot;nt&quot;&gt;--network&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;by (TCP|RDMA) connection&quot;&lt;/span&gt;
mmfsadm &lt;span class=&quot;nb&quot;&gt;test &lt;/span&gt;verbs conn
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A client silently falling back to TCP over IPoIB will work. It will be several
times slower than it should be, it will not log an error, and the problem will be
attributed to the storage cluster. Check this on new nodes as a matter of routine —
particularly on any node added with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmaddnode&lt;/code&gt;, which does not inherit per-node
overrides.&lt;/p&gt;

&lt;h2 id=&quot;size-the-client-pagepool-deliberately&quot;&gt;Size the client pagepool deliberately&lt;/h2&gt;

&lt;p&gt;The default is small — often 1 GiB — and on a compute node with hundreds of
gigabytes of RAM that is usually wrong in one direction or the other.&lt;/p&gt;

&lt;p&gt;Too small and you lose read caching and prefetch effectiveness on exactly the
workloads that benefit most. Too large and you have taken memory away from jobs;
on a node where the scheduler enforces memory limits, the pagepool is memory the
scheduler does not know it has already lost.&lt;/p&gt;

&lt;p&gt;A few GiB is a reasonable starting point for a general compute node. Match it to
what the node is for, set it per node class rather than globally, and remember it
&lt;strong&gt;only takes effect on daemon restart&lt;/strong&gt; — a configured value with no restart is a
value that is not running.&lt;/p&gt;

&lt;h2 id=&quot;bind-mounts-over-gpfs-the-hazard-that-kills-jobs-silently&quot;&gt;Bind mounts over GPFS: the hazard that kills jobs silently&lt;/h2&gt;

&lt;p&gt;Sites commonly bind-mount a directory inside GPFS onto a conventional path — a
scratch area onto &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/scratch&lt;/code&gt;, a shared application tree onto a module path, a home
directory cache onto &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/home&lt;/code&gt;. It is convenient and it introduces a specific,
serious failure mode.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A process whose current working directory is under a bind mount does not follow
that mount when it is re-established.&lt;/strong&gt; Tear the bind down and put it back, and the
process’s CWD is now invalid. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;getwd()&lt;/code&gt; starts returning NULL.&lt;/p&gt;

&lt;p&gt;The consequences are not immediate and they do not point at the cause:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;The job does &lt;strong&gt;not&lt;/strong&gt; die at the moment of the remount. It dies at its next call
that needs the working directory, which may be minutes later.&lt;/li&gt;
  &lt;li&gt;The error surfaces deep inside application code. In R, for instance, a library
that saves and restores the working directory around a compile step fails with
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Error in setwd(cur) : character argument expected&lt;/code&gt;, because &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cur&lt;/code&gt; is NULL. The
message names the application’s function. It reads unmistakably like a user bug.&lt;/li&gt;
  &lt;li&gt;Slurm marks the job FAILED rather than requeuing it, because from the scheduler’s
point of view the job exited non-zero of its own accord.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A rolling restart of a bind-mount unit across a compute fleet, with jobs running,
kills every job that has a working directory under that path. In one measured case
a sequential loop across eighteen nodes produced a 100% mortality rate among the
jobs it touched — every job spanning its node’s restart failed, none survived —
while the naive correlation test (“did jobs fail within sixty seconds of the
restart?”) found almost nothing and would have exonerated the change entirely. The
correct test is whether the restart fell &lt;strong&gt;inside&lt;/strong&gt; each job’s start-to-end window.&lt;/p&gt;

&lt;p&gt;The rules that follow from this:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Drain the node before touching a bind mount.&lt;/strong&gt; Treat it exactly as you would a
reboot. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scontrol update NodeName=&amp;lt;n&amp;gt; State=DRAIN&lt;/code&gt;, wait for the node to empty,
make the change, verify with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;findmnt&lt;/code&gt;, resume.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Never loop across a fleet without a per-node gate.&lt;/strong&gt; Restarts spaced a few
seconds apart are the signature of a script with no check in it.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;A bind restart is in one respect worse than a reboot.&lt;/strong&gt; A reboot kills jobs
visibly and the scheduler requeues them. A bind restart corrupts the state of
surviving processes, which then fail later with an error that blames the user.&lt;/li&gt;
  &lt;li&gt;If a fleet-wide change genuinely cannot wait for an empty cluster, the honest
options are a maintenance reservation or announcing the job losses. A silent
rolling loop is neither.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2 id=&quot;validate-the-client-side&quot;&gt;Validate the client side&lt;/h2&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmmount all &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt;
mmlsmount all &lt;span class=&quot;nt&quot;&gt;-L&lt;/span&gt;                       &lt;span class=&quot;c&quot;&gt;# who has it mounted&lt;/span&gt;
mmdf gpfs0                             &lt;span class=&quot;c&quot;&gt;# capacity as the client sees it&lt;/span&gt;
mmdiag &lt;span class=&quot;nt&quot;&gt;--network&lt;/span&gt;                       &lt;span class=&quot;c&quot;&gt;# RDMA in use, not TCP&lt;/span&gt;
findmnt &lt;span class=&quot;nt&quot;&gt;-no&lt;/span&gt; SOURCE /gpfs/gpfs0         &lt;span class=&quot;c&quot;&gt;# the mount, not the unit state&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Then test as a user, not as root. Many sites export storage through NFS with
root-squash somewhere in the path, and a root test that succeeds proves less than a
user test that fails.&lt;/p&gt;

&lt;p&gt;Confirm on &lt;strong&gt;every&lt;/strong&gt; class of node, not a sample. Head and login nodes frequently
differ from compute nodes in ways nobody documented — a local directory shadowing a
shared mount, a different module tree, an export that was never extended to the
control plane. A configuration checked only on the head node is a configuration
checked on the one machine that is least representative of where jobs run.&lt;/p&gt;

&lt;h2 id=&quot;day-2-failures-worth-recognizing&quot;&gt;Day-2 failures worth recognizing&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A node that mounts but performs badly.&lt;/strong&gt; Check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsPorts&lt;/code&gt; and RDMA state first;
silent TCP fallback is the most common cause and the least visible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A node that will not mount after rejoining.&lt;/strong&gt; Per-node configuration —
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsPorts&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pagepool&lt;/code&gt;, license designation — is not inherited by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmaddnode&lt;/code&gt;.
Compare a working node against the new one line by line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Automount not mounting.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-A yes&lt;/code&gt; mounts at daemon start. If the daemon started
before the network was ready, the mount fails and does not retry on its own.
Ordering the GPFS unit after &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;network-online.target&lt;/code&gt; avoids it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contact nodes all down.&lt;/strong&gt; Clients cannot mount if none of the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-n&lt;/code&gt; contact nodes
answer, even when the rest of the storage cluster is healthy. Keep the list
current: a contact node that was decommissioned two years ago is a mount failure
waiting for a coincidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stale mounts after a storage-side event.&lt;/strong&gt; If the storage cluster force-unmounts,
clients can be left with a mount point that exists but does not work. The GPFS mount
callback described in the companion ECE guide is the reliable repair; systemd will
not notice on its own.&lt;/p&gt;

&lt;h2 id=&quot;the-pattern-underneath-all-of-these&quot;&gt;The pattern underneath all of these&lt;/h2&gt;

&lt;p&gt;Every failure in this guide has the same shape. The command returned success. The
configuration file says the right thing. The systemd unit reports &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;active&lt;/code&gt;. And the
running state is different.&lt;/p&gt;

&lt;p&gt;The habit that prevents most of it is cheap: after any change, check the thing
itself rather than the thing that describes it. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;findmnt&lt;/code&gt; rather than
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;systemctl is-active&lt;/code&gt;. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmfsadm dump config&lt;/code&gt; rather than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmlsconfig&lt;/code&gt;. A job run as
a user rather than a command run as root. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmdiag --network&lt;/code&gt; rather than the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsPorts&lt;/code&gt; setting you just made.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>LID budgets: what happens to addressing when the endpoint count changes</title>
    <link href="https://www.wirewalk.com/writing/infiniband-lid-budget-lmc/"/>
    <published>2026-09-04T09:00:00-04:00</published>
    <updated>2026-09-04T09:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/infiniband-lid-budget-lmc/</id>
    <summary>Every endpoint you add consumes address space and forwarding table entries. Both are finite, both are countable in advance, and neither fails gracefully.</summary>
    <content type="html">&lt;p&gt;A Local Identifier is the address a subnet manager assigns to a port so switches
can forward to it. It is sixteen bits. The unicast range runs from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0x0001&lt;/code&gt; to
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0xBFFF&lt;/code&gt;, which is a little over forty-eight thousand addresses, with the
remainder reserved for multicast.&lt;/p&gt;

&lt;p&gt;Forty-eight thousand sounds like plenty, and on most fabrics it is. The number
that actually constrains you is usually not the LID space but the forwarding
table capacity of the switches, and both move when the endpoint count changes.&lt;/p&gt;

&lt;h2 id=&quot;what-consumes-a-lid&quot;&gt;What consumes a LID&lt;/h2&gt;

&lt;p&gt;Each channel adapter port gets at least one. Each switch gets one for its
management port. So a first approximation is: adapter ports, plus switches.&lt;/p&gt;

&lt;p&gt;The multiplier is LMC.&lt;/p&gt;

&lt;h2 id=&quot;lmc-and-why-it-multiplies&quot;&gt;LMC, and why it multiplies&lt;/h2&gt;

&lt;p&gt;LID Mask Control assigns each adapter port a contiguous block of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2^LMC&lt;/code&gt;
addresses instead of a single one. The point is multipathing: with several LIDs
pointing at the same port, the routing engine can compute different paths to
each, and traffic between a pair of nodes can be spread across them.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;LMC 0    1 LID per port
LMC 1    2 LIDs per port
LMC 2    4 LIDs per port
LMC 3    8 LIDs per port
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Set it in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;opensm.conf&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;lmc 0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The cost is linear in the multiplier and it applies to every port. Going from
LMC 0 to LMC 2 does not add a few addresses; it quadruples consumption across
the entire fabric, and it quadruples the number of entries every switch has to
hold to reach those ports.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Leave LMC at 0 unless you have established that you need multipathing and that
your routing engine will use it. Raising it is a fabric-wide cost paid to solve a
problem many fabrics do not have.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;the-constraint-that-actually-bites&quot;&gt;The constraint that actually bites&lt;/h2&gt;

&lt;p&gt;Switch forwarding tables are the real limit. A linear forwarding table holds one
entry per reachable LID, and its capacity is a property of the switch silicon,
not of your configuration. When the number of LIDs in the subnet exceeds what a
switch can hold, that switch cannot route to all of them.&lt;/p&gt;

&lt;p&gt;This is why the LID budget is not simply “am I under 48,000.” It is “is every
switch able to hold an entry for every LID in the subnet,” and the answer depends
on the smallest table in the fabric. A mixed-generation fabric is governed by its
oldest switch.&lt;/p&gt;

&lt;p&gt;Check the capacity your switches actually report rather than assuming from the
model number:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibdiagnet &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; /var/tmp/ibdiag
smpquery SI &amp;lt;lid&amp;gt;          &lt;span class=&quot;c&quot;&gt;# switch info, including linear FDB capacity&lt;/span&gt;
saquery SI
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;counting-before-you-change-anything&quot;&gt;Counting before you change anything&lt;/h2&gt;

&lt;p&gt;The arithmetic is straightforward and worth doing on paper.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;LIDs  =  (adapter ports × 2^LMC)  +  switches  +  headroom
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Headroom matters because fabrics grow between maintenance windows, and because
a fabric at ninety-nine percent of its forwarding capacity behaves fine until the
day someone plugs in one more node.&lt;/p&gt;

&lt;p&gt;Two changes push this number hard, and they often arrive together:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Splitting switch ports.&lt;/strong&gt; Splitting does not itself consume LIDs — switches
still take one each — but it exists to let you attach more endpoints, and those
endpoints do. A split that doubles attachable ports is a plan to double the
adapter port count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adding dual-port adapters, or turning on the second port.&lt;/strong&gt; Each active port is
an endpoint. A fleet where the second port was cabled but never brought up will
double its LID consumption the day someone brings it up, and that is rarely
recorded as a change to addressing.&lt;/p&gt;

&lt;p&gt;Multiply the two together and a fabric that was comfortable becomes a fabric that
is not, without anyone having made a decision about addressing.&lt;/p&gt;

&lt;h2 id=&quot;what-exhaustion-looks-like&quot;&gt;What exhaustion looks like&lt;/h2&gt;

&lt;p&gt;It does not announce itself as an addressing problem. The symptoms are:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Nodes that enumerate in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibnetdiscover&lt;/code&gt; but are unreachable from some parts of
the fabric and reachable from others, because the switches that ran out are the
ones on those paths.&lt;/li&gt;
  &lt;li&gt;Subnet manager logs reporting failures to set forwarding table entries.&lt;/li&gt;
  &lt;li&gt;Errors during table distribution on a sweep, often on the largest switches
first.&lt;/li&gt;
  &lt;li&gt;Intermittent, position-dependent reachability that looks like a cabling fault
and survives recabling.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tell is that it is topology-dependent rather than node-dependent. A broken
adapter fails from everywhere. An exhausted table fails from behind that switch.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;lid|forwarding|fdb|table&apos;&lt;/span&gt; /var/log/opensm.log | &lt;span class=&quot;nb&quot;&gt;tail&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-60&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;interaction-with-the-routing-engine&quot;&gt;Interaction with the routing engine&lt;/h2&gt;

&lt;p&gt;Routing engines use LMC differently, which matters when you are deciding whether
to raise it.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt; can make use of multiple LIDs per port to distribute traffic across
parallel paths in a fat tree, and on a fabric that qualifies it is the case where
raising LMC has a defensible payoff. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;updn&lt;/code&gt; derives less benefit, because the
up/down constraint already limits which paths are legal. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;minhop&lt;/code&gt; does the least
with it.&lt;/p&gt;

&lt;p&gt;So the sequence is: confirm which engine is genuinely running, then decide about
LMC. Raising LMC to enable multipathing under an engine that will not use it buys
you four times the forwarding table pressure and nothing else.&lt;/p&gt;

&lt;h2 id=&quot;practical-discipline&quot;&gt;Practical discipline&lt;/h2&gt;

&lt;p&gt;Record the LID count as a monitored number, not as something you look up during
an incident. It is cheap to collect and it turns a class of confusing failures
into a threshold you crossed.&lt;/p&gt;

&lt;p&gt;Recompute the budget as part of planning any change to endpoint count — new
nodes, split ports, second ports brought up, a new storage tier. The calculation
takes a few minutes and the alternative is discovering the limit during a
maintenance window with the change half applied.&lt;/p&gt;

&lt;p&gt;And keep LMC deliberate. It is one line in a configuration file and it is the
single largest multiplier on everything above.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Deploying a GPFS ECE cluster, end to end</title>
    <link href="https://www.wirewalk.com/writing/gpfs-ece-cluster-deployment/"/>
    <published>2026-09-04T09:00:00-04:00</published>
    <updated>2026-09-04T09:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/gpfs-ece-cluster-deployment/</id>
    <summary>Erasure Code Edition turns commodity servers with internal NVMe into a parallel filesystem with no external RAID. The install is well documented. The decisions that are expensive to reverse are not, and most of them are made in the first hour.</summary>
    <content type="html">&lt;p&gt;Erasure Code Edition (ECE) runs IBM Storage Scale’s declustered RAID in software,
across the internal drives of ordinary servers. There is no controller, no external
array, and no appliance. You get the ESS data protection model on hardware you
specify yourself.&lt;/p&gt;

&lt;p&gt;The installation itself is mechanical and IBM documents it well. What follows
concentrates on the decisions that are cheap to make correctly on day one and
painful to revisit later, and on the failures that are hard to diagnose because the
system reports success while doing the wrong thing.&lt;/p&gt;

&lt;h2 id=&quot;decide-whether-ece-is-the-right-answer-at-all&quot;&gt;Decide whether ECE is the right answer at all&lt;/h2&gt;

&lt;p&gt;ECE earns its place when you want parallel-filesystem performance from servers you
control, and you have a fast private network to spend on it. It is the wrong answer
in three cases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Too few nodes.&lt;/strong&gt; The documented minimum is four servers per recovery group. Four
is a number you can build, not a number you should run — see the fault tolerance
arithmetic below. Six is a sensible floor and eight or more is where the model
starts behaving the way the marketing describes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mixed hardware inside a recovery group.&lt;/strong&gt; Every server in a recovery group should
be the same model with the same drive count, drive model, CPU and memory. ECE
distributes strips on the assumption that servers are interchangeable. They do not
have to match across recovery groups, but within one they effectively do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No dedicated fast network.&lt;/strong&gt; ECE reads amplify. Delivering one byte to a client
moves roughly 1.875 bytes on the storage fabric with an 8+2p code, because
reconstructing a block gathers strips from every other server. If that traffic
shares a congested network with everything else, you have built an expensive way to
be slow.&lt;/p&gt;

&lt;h2 id=&quot;get-the-fault-tolerance-arithmetic-right-before-you-buy-anything&quot;&gt;Get the fault tolerance arithmetic right before you buy anything&lt;/h2&gt;

&lt;p&gt;This is the single decision most likely to be regretted, and it is made before a
single package is installed.&lt;/p&gt;

&lt;p&gt;Fault tolerance is not the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;p&lt;/code&gt; in the code name. A code writes a fixed number of
strips, and those strips land on distinct servers. If the strip count exceeds the
server count, some server necessarily holds two strips of the same stripe, and
losing that one server costs you two strips.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Code&lt;/th&gt;
      &lt;th&gt;Strips&lt;/th&gt;
      &lt;th&gt;Servers&lt;/th&gt;
      &lt;th&gt;Strips per server&lt;/th&gt;
      &lt;th&gt;Node fault tolerance&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;8+2p&lt;/td&gt;
      &lt;td&gt;10&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;some hold 2&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;8+2p&lt;/td&gt;
      &lt;td&gt;10&lt;/td&gt;
      &lt;td&gt;11+&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;2&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;4+2p&lt;/td&gt;
      &lt;td&gt;6&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;4+3p&lt;/td&gt;
      &lt;td&gt;7&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;3&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;An 8+2p array on eight servers gives you &lt;strong&gt;node fault tolerance of 1&lt;/strong&gt;. That is not
a theoretical concern:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;One server down and you are one failure from data loss.&lt;/li&gt;
  &lt;li&gt;More immediately, &lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmvdisk&lt;/code&gt; refuses to suspend a server&lt;/strong&gt; while any vdisk in
the recovery group sits at fault tolerance 1. Every maintenance action that needs
a server offline — firmware, kernel updates, a BIOS setting, a physical move — is
blocked. Not warned about: blocked.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cost of fault tolerance 2 is capacity. A 4+2p code writes six strips for four
of data, so usable capacity is raw divided by 1.5. An 8+2p code divides by 1.25.
On the same hardware, moving from 8+2p to 4+2p costs you a fifth of your usable
space and buys you the ability to service the cluster.&lt;/p&gt;

&lt;p&gt;Decide this deliberately. Converting later means building a second filesystem
alongside the first and copying everything across, which needs enough free capacity
to hold both.&lt;/p&gt;

&lt;h2 id=&quot;validate-the-hardware-before-you-install-anything&quot;&gt;Validate the hardware before you install anything&lt;/h2&gt;

&lt;p&gt;IBM publishes two readiness tools in the SpectrumScaleTools repository:
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ece_network_readiness&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ece_storage_readiness&lt;/code&gt;. Run both before installing.
They check the things that will otherwise be discovered as a performance mystery
three weeks later.&lt;/p&gt;

&lt;p&gt;The thresholds worth knowing, because you can check them yourself:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Measure&lt;/th&gt;
      &lt;th&gt;Threshold&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;ICMP latency between storage nodes, average&lt;/td&gt;
      &lt;td&gt;≤ 1 ms&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ICMP latency, maximum&lt;/td&gt;
      &lt;td&gt;≤ 2 ms&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ICMP standard deviation&lt;/td&gt;
      &lt;td&gt;≤ 0.333 ms&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NVMe random 128 KiB IOPS, per drive&lt;/td&gt;
      &lt;td&gt;&amp;gt; 15,000&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Single-client read throughput&lt;/td&gt;
      &lt;td&gt;&amp;gt; 2,000 MB/s&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;A cluster that misses the latency numbers will work and will never perform. The
standard deviation matters more than the average — jitter on the storage network
shows up as unexplained tail latency in every job on the cluster.&lt;/p&gt;

&lt;h2 id=&quot;prepare-the-operating-system&quot;&gt;Prepare the operating system&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Pin the kernel.&lt;/strong&gt; ECE builds a portability layer against the running kernel. An
unplanned kernel update on one node produces a node that will not start &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmfsd&lt;/code&gt;
until you rebuild. Exclude kernel packages from routine updates and treat kernel
changes as a scheduled activity with a rebuild step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Name devices by path, not by enumeration.&lt;/strong&gt; NVMe controllers are probed
asynchronously. The device that is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/dev/nvme4n1&lt;/code&gt; today may be &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/dev/nvme1n1&lt;/code&gt; after
a reboot, and the mapping is not stable across identical servers. Any configuration
that names &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/dev/nvmeXnY&lt;/code&gt; — a disk setup definition, a script, a monitoring check —
is a latent failure that surfaces at the worst time, which is during a reboot you
performed for an unrelated reason.&lt;/p&gt;

&lt;p&gt;Use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/dev/disk/by-path/&lt;/code&gt; or the WWN. ECE itself identifies drives by their unique
identifier and is unbothered; it is the tooling around ECE that breaks.&lt;/p&gt;

&lt;p&gt;A related trap: with NVMe native multipath, block devices sit under a virtual
subsystem and have no PCI parent. Reading a PCI address from the block device
returns nothing. The address lives on the controller, at
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/sys/class/nvme/nvmeN/address&lt;/code&gt;. A validation script that reads the wrong path
returns empty and, if written carelessly, reports success.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AMD EPYC platforms need two specific fixes.&lt;/strong&gt; On family 23 and 25 processors,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmaddnode&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmstartup&lt;/code&gt; can hang for twenty to thirty minutes at 99% CPU inside
the GSKit cryptographic library. The fix is documented by IBM as APAR IJ47407 and
amounts to setting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ICC_SHIFT=3&lt;/code&gt; in the ICCSIG configuration. Bake it into the
image; discovering it per node is a slow way to build a cluster.&lt;/p&gt;

&lt;p&gt;Separately, some Genoa-based platforms panic on late microcode load. If you see
kernel panics on nodes that otherwise boot, add &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dis_ucode_ldr&lt;/code&gt; to the kernel
parameters and load microcode from the firmware instead.&lt;/p&gt;

&lt;h2 id=&quot;create-the-cluster&quot;&gt;Create the cluster&lt;/h2&gt;

&lt;p&gt;Install &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpfs.base&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpfs.gpl&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpfs.ece&lt;/code&gt; and the license packages, then build the
portability layer with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmbuildgpl&lt;/code&gt;. Do this on every node before creating the
cluster, not as part of it.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmcrcluster &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; nodefile &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; /usr/bin/ssh &lt;span class=&quot;nt&quot;&gt;-R&lt;/span&gt; /usr/bin/scp &lt;span class=&quot;nt&quot;&gt;-C&lt;/span&gt; ece-cluster
mmchlicense server &lt;span class=&quot;nt&quot;&gt;--accept&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; all
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Quorum design is worth a moment. Use an odd number — three for most clusters, five
for large ones. Quorum nodes should be spread across failure domains: racks, power
feeds, and switches. A three-node quorum with all three in one rack is a
single-rack outage away from a filesystem that will not mount.&lt;/p&gt;

&lt;p&gt;Designate manager nodes explicitly rather than letting every server carry the role.
On a storage cluster the servers are busy; giving the manager role to all of them
means a manager failover happens under exactly the conditions where you least want
extra work.&lt;/p&gt;

&lt;h2 id=&quot;configure-the-network-and-verify-it-is-actually-being-used&quot;&gt;Configure the network, and verify it is actually being used&lt;/h2&gt;

&lt;p&gt;This is where clusters silently underperform.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsPorts&lt;/code&gt; tells GPFS which RDMA devices to use. Set it, then confirm it took:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmlsconfig verbsPorts                 &lt;span class=&quot;c&quot;&gt;# what is stored&lt;/span&gt;
mmfsadm dump config | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; verbs   &lt;span class=&quot;c&quot;&gt;# what is running&lt;/span&gt;
mmdiag &lt;span class=&quot;nt&quot;&gt;--network&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;by (TCP|RDMA) connection&quot;&lt;/span&gt;
mmfsadm &lt;span class=&quot;nb&quot;&gt;test &lt;/span&gt;verbs conn               &lt;span class=&quot;c&quot;&gt;# per-peer RDMA counters&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three failure modes, all of which present as “it works, but slowly”:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Per-node overrides are not inherited.&lt;/strong&gt; A node added later with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmaddnode&lt;/code&gt; does
not pick up a per-node &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsPorts&lt;/code&gt; setting. It falls back to the cluster default,
and if that default names a device the new node does not have, GPFS quietly uses
TCP over IPoIB instead. Throughput drops by a factor you will notice; nothing logs
an error. Check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsPorts&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pagepool&lt;/code&gt; and the license designation after &lt;strong&gt;every&lt;/strong&gt;
node addition or rejoin.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pointing at the wrong port type.&lt;/strong&gt; A dual-port adapter may present one InfiniBand
port and one Ethernet/RoCE port. A global default of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlx5_1/1&lt;/code&gt; aims GPFS at
whichever port happens to be second, which on some platforms is the RoCE port.
Check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;link_layer&lt;/code&gt; in sysfs before trusting a device name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The negotiated rate is not what you ordered.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt; reports the per-lane
rate, not the link rate. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;4X 106.25 Gbps&lt;/code&gt; is NDR400; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;4X 53.125 Gbps&lt;/code&gt; is HDR200;
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2X 53.125 Gbps&lt;/code&gt; is a half-width HDR100 link that will gate every synchronized
operation across the cluster. The authoritative per-host view is
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/sys/class/infiniband/&amp;lt;dev&amp;gt;/ports/1/rate&lt;/code&gt;. Check every node, not a sample — one
degraded link is enough to make an entire benchmark meaningless, and it will be
blamed on the filesystem.&lt;/p&gt;

&lt;h2 id=&quot;build-the-recovery-group&quot;&gt;Build the recovery group&lt;/h2&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmvdisk nodeclass create &lt;span class=&quot;nt&quot;&gt;--node-class&lt;/span&gt; NC_ECE &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; storage01,storage02,...,storage08
mmvdisk server configure &lt;span class=&quot;nt&quot;&gt;--node-class&lt;/span&gt; NC_ECE &lt;span class=&quot;nt&quot;&gt;--recycle&lt;/span&gt; one
mmvdisk recoverygroup create &lt;span class=&quot;nt&quot;&gt;--recovery-group&lt;/span&gt; RG1 &lt;span class=&quot;nt&quot;&gt;--node-class&lt;/span&gt; NC_ECE
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmvdisk server configure&lt;/code&gt; sets the server-side parameters, including &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pagepool&lt;/code&gt;
and the RAID buffer pool percentage, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--recycle one&lt;/code&gt; restarts daemons one at a
time to apply them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Confirm the recycle actually completed.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pagepool&lt;/code&gt; takes effect only on daemon
restart. A configured value with no completed restart leaves the cluster running the
old value indefinitely, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmlsconfig&lt;/code&gt; will happily report the new one. The
running value is in the daemon’s own log at every start, and in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmdiag --config&lt;/code&gt;.
A cluster running months on a fraction of its intended pagepool is not a
hypothetical failure; the symptom is mediocre write performance that no tuning
fixes.&lt;/p&gt;

&lt;h2 id=&quot;define-vdisk-sets-and-read-the-size-units-carefully&quot;&gt;Define vdisk sets, and read the size units carefully&lt;/h2&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmvdisk vdiskset define &lt;span class=&quot;nt&quot;&gt;--vdisk-set&lt;/span&gt; VS1 &lt;span class=&quot;nt&quot;&gt;--recovery-group&lt;/span&gt; RG1 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
        &lt;span class=&quot;nt&quot;&gt;--code&lt;/span&gt; 4+2p &lt;span class=&quot;nt&quot;&gt;--block-size&lt;/span&gt; 2M &lt;span class=&quot;nt&quot;&gt;--set-size&lt;/span&gt; 100T
mmvdisk vdiskset create &lt;span class=&quot;nt&quot;&gt;--vdisk-set&lt;/span&gt; VS1
mmvdisk filesystem create &lt;span class=&quot;nt&quot;&gt;--file-system&lt;/span&gt; fs1 &lt;span class=&quot;nt&quot;&gt;--vdisk-set&lt;/span&gt; VS1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--set-size&lt;/code&gt; is &lt;strong&gt;usable&lt;/strong&gt; capacity, not raw. Requesting 150 TiB at 4+2p consumes
225 TiB of raw space, and the command fails with “not enough space” on an array
that appears to have plenty. Multiply by the code’s expansion factor before you ask.&lt;/p&gt;

&lt;p&gt;Two more things that are difficult to change afterwards:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Block size changes the subblock count, not the subblock size.&lt;/strong&gt; Moving from 4 MiB
to 2 MiB blocks takes you from 512 subblocks to 256, both of 8 KiB. Small-file
efficiency is unchanged. If you are rebuilding a filesystem in the belief that a
smaller block size will help small files, check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmlsfs -f&lt;/code&gt; first and confirm the
subblock size you expect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmvdisk vdiskset list&lt;/code&gt; prints per-vdisk size, not set totals.&lt;/strong&gt; A set of sixteen
vdisks showing “81 TiB” is 1.27 PiB. Misreading that column is a straightforward way
to conclude a filesystem is nearly full when it is under ten percent used.
Allocation is not utilization: a declustered array can be 99% &lt;em&gt;allocated&lt;/em&gt; to vdisk
sets while the filesystem on top is nearly empty.&lt;/p&gt;

&lt;h2 id=&quot;tune-and-verify-each-change-actually-applied&quot;&gt;Tune, and verify each change actually applied&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmchconfig&lt;/code&gt; returning success means the value is stored. It does not mean the value
is in effect, and for several parameters it does not mean the value will ever be in
effect.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmlsconfig &amp;lt;param&amp;gt;                        &lt;span class=&quot;c&quot;&gt;# stored&lt;/span&gt;
mmfsadm dump config | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; &amp;lt;param&amp;gt;     &lt;span class=&quot;c&quot;&gt;# running   &amp;lt;- the one that matters&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;In &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmfsadm dump config&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;!&lt;/code&gt; marks a value differing from default, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;*&lt;/code&gt; marks
one changed at runtime. A value that appears with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;!&lt;/code&gt; but never with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;*&lt;/code&gt; may have
been accepted and discarded.&lt;/p&gt;

&lt;p&gt;Parameters worth setting, with measured effect on an NVMe ECE cluster:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Parameter&lt;/th&gt;
      &lt;th&gt;Effect&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pagepool&lt;/code&gt; (servers)&lt;/td&gt;
      &lt;td&gt;Large. Shared-file writes improve substantially; file-per-process much less, because it was never buffer-starved&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;maxMBpS&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;~6% on writes, nothing on reads&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;workerThreads&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prefetchThreads&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;~3% on writes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nsdRAIDBufferPoolSizePct&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Sets the RAID buffer; leave at the configured default unless you have a reason&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Parameters that are accepted and do nothing on current versions. These are
documented here so nobody spends a rolling restart proving it again:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Parameter&lt;/th&gt;
      &lt;th&gt;Why it does nothing&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsRdmasPerNode&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Supported in 4.2.x only. Runtime value stays 0&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsRdmasPerConnection&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Obsolete since 5.0. Stored, discarded at runtime&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nsdRAIDThreadsPerQueue&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;No measurable effect&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsNumSendChannel&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsNumRecvChannel&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;No measurable effect&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsRdmaMaxSendBytes&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;No measurable effect&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;verbsRdmaSend&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Genuinely moves RPCs from TCP to RDMA — and changes throughput not at all. It governs daemon-to-daemon traffic, not the data path&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;h2 id=&quot;make-the-cluster-repair-its-own-mounts&quot;&gt;Make the cluster repair its own mounts&lt;/h2&gt;

&lt;p&gt;Anything bind-mounted over a GPFS path — a scratch directory, a module tree, a
shared application area — is torn down when the filesystem goes away and is &lt;strong&gt;not&lt;/strong&gt;
re-established when it returns. A systemd unit with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;BindsTo=gpfs.service&lt;/code&gt; covers a
service restart but not an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmshutdown&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmstartup&lt;/code&gt; cycle, because the filesystem
can disappear and return without the service unit changing state.&lt;/p&gt;

&lt;p&gt;GPFS knows when the filesystem mounts even when systemd does not. Register a
callback:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmaddcallback ScratchBind &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--command&lt;/span&gt; /usr/local/sbin/gpfs-post-mount.sh &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--event&lt;/span&gt; mount &lt;span class=&quot;nt&quot;&gt;--parms&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;%fsName&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three details decide whether this works:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Match every device name the filesystem is known by.&lt;/strong&gt; If the filesystem is ever
renamed, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%fsName&lt;/code&gt; changes and a guard written against the old name exits
silently. Accept both, and log any name you do not recognize, so the next rename
leaves evidence instead of a silent outage.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Make the script idempotent and fast.&lt;/strong&gt; A slow callback delays the mount for
every node.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;RemainAfterExit=yes&lt;/code&gt; is load-bearing&lt;/strong&gt; on any &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Type=oneshot&lt;/code&gt; bind unit.
Without it, systemd deactivates the unit the instant &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ExecStart&lt;/code&gt; returns and runs
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ExecStop&lt;/code&gt; as part of that normal deactivation — the unit binds the path and
unbinds it a second later.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Be aware of the corollary: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;systemctl is-active&lt;/code&gt; on such a unit reports &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;active&lt;/code&gt;
whether or not the bind exists. Check the mount, with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;findmnt -no SOURCE &amp;lt;path&amp;gt;&lt;/code&gt;,
never the unit state.&lt;/p&gt;

&lt;h2 id=&quot;validate-with-numbers-that-survive-scrutiny&quot;&gt;Validate with numbers that survive scrutiny&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Size the dataset from the cache, not from convenience.&lt;/strong&gt; Client-side &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;O_DIRECT&lt;/code&gt;
bypasses the client pagepool but not the servers’ RAID buffer pool. That buffer is:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;pagepool × nsdRAIDBufferPoolSizePct × number_of_servers
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Eight servers with a 128 GiB pagepool at 80% is roughly 800 GiB of read cache. A
read benchmark smaller than that measures memory. The error is large — a test that
reports 155 GiB/s cached can be 82 GiB/s cold. Size reads at several times the
buffer and confirm by re-running at double: if the number drops, you were measuring
cache.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do not mistake under-driving for a ceiling.&lt;/strong&gt; Too few outstanding I/Os looks
exactly like a hard limit. Sweep concurrency; a genuinely saturated system is flat
across 64, 96, 128 and 192 jobs, while an under-driven one climbs. Sweep node count
too, since a per-node limit and an aggregate limit are different problems with
different fixes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check the fabric before tuning the filesystem.&lt;/strong&gt; If clients and servers sit on
different switches, measure the inter-switch links under load before changing a
single GPFS parameter. Static InfiniBand routing pins each destination to one
output port, and it can collapse many links into a few in one direction while the
other direction spreads perfectly. The signature is a large read/write asymmetry
with nothing apparently saturated, and no amount of GPFS tuning touches it.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nsdperf&lt;/code&gt; exercises the NSD protocol with no disk layer at all, which separates
network from storage in one test. If &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nsdperf&lt;/code&gt; and your filesystem report similar
numbers, the bottleneck is at or below the network and no tunable will help.&lt;/p&gt;

&lt;h2 id=&quot;restart-the-cluster-without-an-outage&quot;&gt;Restart the cluster without an outage&lt;/h2&gt;

&lt;p&gt;Many parameters need a daemon restart, which on a storage cluster means restarting
servers underneath a live filesystem. Done with gates it is routine.&lt;/p&gt;

&lt;p&gt;Refuse to start unless all of these hold, and re-check them between every node:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;every server &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;active&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;recovery group needs service: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;no&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;all pdisks &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ok&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;no rebuild or rebalance running&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;fault tolerance still at its design value&lt;/strong&gt; — this is the safety budget&lt;/li&gt;
  &lt;li&gt;the filesystem still mounted on a client&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Order matters: non-quorum nodes first to prove the procedure, then quorum nodes one
at a time, and the recovery group master &lt;strong&gt;last&lt;/strong&gt; so you do not trigger a pointless
master failover at the start.&lt;/p&gt;

&lt;p&gt;Run it detached — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;setsid nohup ... &amp;amp;&lt;/code&gt;. An administrator’s dropped SSH session
should not be able to leave a storage node down with no orchestration.&lt;/p&gt;

&lt;h2 id=&quot;the-failures-worth-recognizing-on-sight&quot;&gt;The failures worth recognizing on sight&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;“Node appears to already belong to a cluster.”&lt;/strong&gt; Usually a stale &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmsdrserv&lt;/code&gt;
process on the node from a previous attempt. Kill it and retry; reinstalling the
node is not required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A node that answers SSH but has no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/var/mmfs&lt;/code&gt;.&lt;/strong&gt; It is sitting in a provisioning
installer environment with a read-only NFS root, not the installed OS. Probe for a
real local root before running anything against a node you have just rebooted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reading the filesystem mounts it.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmlsfs&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmrepquota&lt;/code&gt; trigger an internal
mount. Doing a record-keeping step immediately before an unmount gate makes that
gate fail, and the internal mounts are not visible as user processes — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fuser&lt;/code&gt; shows
only kernel threads. Record first, force-unmount second, gate third. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmchfs -T&lt;/code&gt;
requires zero mounts including internal ones, and a plain &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmumount -a&lt;/code&gt; does not
guarantee that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fileset inode limits do not carry across.&lt;/strong&gt; A fileset created on a new filesystem
defaults to a little over 100,000 inodes regardless of what the source held. A
migration that copies millions of files into it dies partway with no obvious cause.
Set &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--inode-limit&lt;/code&gt; at creation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Snapshots cannot be migrated.&lt;/strong&gt; They live with the filesystem. If a rebuild is
planned and any rollback history matters, extract it first; the new filesystem
starts with none.&lt;/p&gt;

&lt;h2 id=&quot;what-to-watch-once-it-is-running&quot;&gt;What to watch once it is running&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmhealth&lt;/code&gt; covers the basics. Beyond it, the signals that catch real problems early:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;long waiters&lt;/strong&gt; — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmdiag --waiters&lt;/code&gt;; anything over 30 seconds repeatedly is a
stuck subsystem, not a busy one&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;server queue depth&lt;/strong&gt; — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmfsadm dump nsd | grep -E &quot;pending|active&quot;&lt;/code&gt;. Queues
reporting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pending 0&lt;/code&gt; under load mean the servers are starved and the bottleneck
is upstream. This one measurement short-circuits most performance investigations
and is worth taking first, not last&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;RDMA actually in use&lt;/strong&gt; — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmdiag --network&lt;/code&gt;, per the earlier warning&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;pdisk state and fault tolerance&lt;/strong&gt; — a background rebuild is normal; a fault
tolerance that has dropped and stayed down is not&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;config versus running state&lt;/strong&gt; — periodically diff &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmlsconfig&lt;/code&gt; against
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmfsadm dump config&lt;/code&gt;. Divergence is silent and accumulates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last one deserves emphasis. Nearly every long-running problem described here
shares a shape: the system reported success, the configuration said one thing, and
the running state said another. Build the habit of checking the second one.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Changing the OpenSM routing engine: updn, ftree, and the silent fallback</title>
    <link href="https://www.wirewalk.com/writing/opensm-routing-engines/"/>
    <published>2026-09-04T08:00:00-04:00</published>
    <updated>2026-09-04T08:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/opensm-routing-engines/</id>
    <summary>Setting routing_engine does not mean the engine you named is the engine you got. ftree in particular declines the job and hands it back, and it does so in a log line most people never read.</summary>
    <content type="html">&lt;p&gt;The subnet manager decides which path every packet takes through the fabric, and
the routing engine is the algorithm it uses to decide. Changing it is a one-line
edit to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;opensm.conf&lt;/code&gt;. Confirming that the change actually took is the part worth
learning, because one of the common engines refuses fabrics it does not like and
falls back to another without failing.&lt;/p&gt;

&lt;h2 id=&quot;the-engines-worth-knowing&quot;&gt;The engines worth knowing&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;minhop&lt;/code&gt;&lt;/strong&gt; routes by shortest path. It is simple, fast to compute, and offers
no protection against credit loops. On a topology where cycles are possible, that
is a deadlock waiting for the right traffic pattern. It is a reasonable default
on small, regular fabrics and a poor one on anything irregular.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;updn&lt;/code&gt;&lt;/strong&gt; applies an up/down turn rule: traffic may go up the tree and then
down, but never up again after descending. That constraint makes credit loops
impossible on an arbitrary topology, which is why up/down is the safe answer when
you do not know what shape the fabric is. The cost is path quality — some routes
are longer than they need to be, and the distribution across links is uneven.&lt;/p&gt;

&lt;p&gt;Up/down depends entirely on which nodes it treats as roots. Left to itself,
OpenSM guesses, and the guess is often poor on a fabric where spine and leaf
switches are not obviously distinguishable. Give it the answer explicitly:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;routing_engine updn
root_guid_file /etc/opensm/root_guids.txt
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That file holds one spine switch GUID per line. Getting it right is the
difference between up/down performing acceptably and performing badly, and it is
the most common thing left undone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt;&lt;/strong&gt; is optimized for fat trees, and on a fabric that genuinely is one it
produces materially better link utilization than up/down. It also has strict
requirements about what counts as a fat tree: the topology must be regular,
constant bisectional bandwidth, with the compute nodes at the leaves and a
consistent number of links between tiers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dnup&lt;/code&gt;&lt;/strong&gt; inverts the up/down rule, and is occasionally the better fit where the
traffic pattern is dominated by leaf-to-leaf rather than leaf-to-core.&lt;/p&gt;

&lt;p&gt;There are others — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lash&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dor&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;torus-2QoS&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dfsssp&lt;/code&gt; — that matter on
specific topologies. If you are running a torus or a dragonfly, the engine is not
a matter of taste and the vendor guidance for that topology governs.&lt;/p&gt;

&lt;h2 id=&quot;the-silent-fallback&quot;&gt;The silent fallback&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;routing_engine&lt;/code&gt; accepts an ordered, comma-separated list, and that is not a
preference list in the way people assume. It is a fallback chain:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;routing_engine ftree,updn,minhop
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;OpenSM tries &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt;. If the fabric does not satisfy its requirements, it logs
the reason and moves to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;updn&lt;/code&gt;. The subnet comes up. Traffic flows. Nothing
appears to be wrong. You are simply not running the engine you configured, and
the performance you were expecting from the change does not arrive.&lt;/p&gt;

&lt;p&gt;This happens more often than it should, for mundane reasons: one node cabled to
a spine instead of a leaf, a leaf with a different uplink count from its
siblings, a switch that was down during the sweep, a section of the fabric added
later with a different ratio.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;After changing the engine, read the log and confirm which one is actually
running. Do not infer it from the configuration file. The configuration file
records what you asked for.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;confirming-what-is-running&quot;&gt;Confirming what is running&lt;/h2&gt;

&lt;p&gt;The subnet manager log is the primary source:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;routing engine|ftree|fallback|topology&apos;&lt;/span&gt; /var/log/opensm.log | &lt;span class=&quot;nb&quot;&gt;tail&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-40&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;You are looking for a line naming the engine that was used, and — if &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt;
declined — a line explaining why. That explanation is usually specific enough to
fix: it will name the node or the tier that broke the assumption.&lt;/p&gt;

&lt;p&gt;For a fuller picture, have OpenSM write out what it computed:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;opensm &lt;span class=&quot;nt&quot;&gt;--dump_files_dir&lt;/span&gt; /var/tmp/osm-dump &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-o&lt;/code&gt; runs a single sweep and exits, which is what you want for inspection rather
than for taking over the subnet. The dump directory gets the LID matrices and
the per-switch forwarding tables, which is the ground truth about where traffic
will actually go.&lt;/p&gt;

&lt;p&gt;Then confirm from the fabric side:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;sminfo
ibdiagnet &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; /var/tmp/ibdiag &lt;span class=&quot;nt&quot;&gt;--routing&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;applying-a-change&quot;&gt;Applying a change&lt;/h2&gt;

&lt;p&gt;Routing is recomputed on a sweep, and the engine is read at startup. A change
means restarting the subnet manager:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;systemctl restart opensm
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two cautions. On a large fabric the recompute and redistribution of forwarding
tables is not instant, and traffic during that window can take paths that are in
the process of changing. Do this in a maintenance window on anything busy.&lt;/p&gt;

&lt;p&gt;And if you run a standby subnet manager, its configuration has to change too.
A failover onto a standby still holding the old engine will silently reroute the
entire fabric back, and the resulting “the fabric got slow again for no reason”
is genuinely hard to diagnose if you have forgotten the standby exists.&lt;/p&gt;

&lt;h2 id=&quot;measuring-whether-it-helped&quot;&gt;Measuring whether it helped&lt;/h2&gt;

&lt;p&gt;The point of changing engines is link utilization, so measure that rather than a
single pair’s bandwidth.&lt;/p&gt;

&lt;p&gt;Point-to-point between one pair of nodes is usually indifferent to the routing
engine, because a single flow takes a single path either way. What changes is
what happens when many flows share the fabric. Use an all-to-all or an all-reduce
across the real job size:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; &amp;lt;ranks&amp;gt; ./osu_alltoall
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; &amp;lt;ranks&amp;gt; ./build/all_reduce_perf &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; 1G &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; 1G &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; 20
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Take a baseline before the change, on the same node set, and compare. Also look
at per-port counters across the spine after a representative run: an engine that
is distributing well produces even traffic across parallel uplinks, and one that
is not produces a visible imbalance. That imbalance is often easier to see, and
more convincing, than a bandwidth number.&lt;/p&gt;

&lt;h2 id=&quot;adaptive-routing-is-a-different-layer&quot;&gt;Adaptive routing is a different layer&lt;/h2&gt;

&lt;p&gt;Modern switches can make per-packet forwarding decisions to route around
congestion. That is a switch feature, configured on the switch and in some cases
coordinated by the fabric manager, and it operates on top of whatever path set
the subnet manager computed. It is not a substitute for choosing the right
routing engine, and enabling it does not repair a fabric whose static routing is
badly distributed. Get the static routing right first, then decide whether
adaptive routing adds anything.&lt;/p&gt;

&lt;h2 id=&quot;the-short-version&quot;&gt;The short version&lt;/h2&gt;

&lt;p&gt;Set the engine, restart, then read the log to find out which engine you are
actually running. Give &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;updn&lt;/code&gt; an explicit root list. Expect &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt; to decline if
anything about the topology is irregular, and treat its explanation as a cabling
bug report rather than a configuration problem. Measure with collectives, not
with a single pair.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Splitting ports on Quantum switches: what it costs you</title>
    <link href="https://www.wirewalk.com/writing/quantum-switch-port-splitting/"/>
    <published>2026-09-04T07:00:00-04:00</published>
    <updated>2026-09-04T07:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/quantum-switch-port-splitting/</id>
    <summary>Doubling your port count is the easy part. The parts that surprise people are the renumbering, the cabling, and what it does to the subnet manager.</summary>
    <content type="html">&lt;p&gt;Port splitting turns one high-rate port into two or four lower-rate ports. On a
Quantum-class HDR switch that is 40 ports at HDR200 becoming 80 at HDR100. On a
Quantum-2 NDR switch it is 64 logical NDR400 ports, presented through 32 physical
cages, becoming 128 at NDR200.&lt;/p&gt;

&lt;p&gt;The reason to do it is almost always economics. If your endpoints have HDR100 or
NDR200 adapters, an unsplit switch wastes half of every port’s capability, and
you buy twice the switches you need. The reason to think carefully first is that
splitting changes four things at once, and only the first is obvious.&lt;/p&gt;

&lt;h2 id=&quot;what-a-cage-actually-carries&quot;&gt;What a cage actually carries&lt;/h2&gt;

&lt;p&gt;The physical connector is not the port. On Quantum-2, one OSFP cage carries eight
lanes, presented by default as two NDR400 ports of four lanes each. Split, those
same eight lanes present as four NDR200 ports of two lanes each. Nothing about
the switch got faster or slower; the lanes were regrouped.&lt;/p&gt;

&lt;p&gt;This is why the aggregate switch bandwidth does not change when you split, and
why “we doubled the ports” and “we doubled the capacity” are different claims.
You have the same total, divided more finely.&lt;/p&gt;

&lt;h2 id=&quot;the-renumbering&quot;&gt;The renumbering&lt;/h2&gt;

&lt;p&gt;Split ports get a new naming scheme, and every piece of tooling that referenced
the old names breaks quietly.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;unsplit      1/1        1/2        1/3
split        1/1/1 1/1/2   1/2/1 1/2/2   1/3/1 1/3/2
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Anything holding port identifiers is now wrong: monitoring definitions, cabling
records, port-to-node maps, error-counter baselines, and any script that walks a
port list. None of that fails loudly. It just reports on ports that no longer
exist, or misses ports that now do.&lt;/p&gt;

&lt;p&gt;Regenerate the topology from the fabric rather than editing your records by hand:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibnetdiscover &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; topology-&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;date&lt;/span&gt; +%F&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;.txt
iblinkinfo
ibdiagnet &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; /var/tmp/ibdiag-&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;date&lt;/span&gt; +%F&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Do that before the change as well as after. The diff is what you hand to whoever
maintains the monitoring.&lt;/p&gt;

&lt;h2 id=&quot;enabling-it&quot;&gt;Enabling it&lt;/h2&gt;

&lt;p&gt;The exact syntax depends on the switch OS release, and this is one of the places
where an old runbook will mislead you. The general shape on MLNX-OS is a system
profile that permits splitting, then a per-port module type, then a reboot:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;switch (config) # system profile ib split-ready
switch (config) # interface ib 1/1 module-type qsfp-split-2
switch (config) # write memory
switch (config) # reload
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three things to check against your own release notes rather than against this:
whether the profile change is required at all on your platform, whether the
split factor is expressed as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;split-2&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;split-4&lt;/code&gt;, and whether the reboot is
required or merely recommended. Treat the vendor documentation for your exact
firmware as authoritative.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The reboot is the real constraint.&lt;/strong&gt; Enabling split-ready mode is a switch
restart, which is a fabric event. On a single-switch fabric that is an outage. On
a leaf-and-spine fabric it is a capacity reduction and a routing re-convergence,
and if it is a spine, it is a large one. Plan it as maintenance, not as a
configuration change.&lt;/p&gt;

&lt;h2 id=&quot;cabling&quot;&gt;Cabling&lt;/h2&gt;

&lt;p&gt;A split port needs a splitter cable, and the breakout ratio has to match the
configured split factor. A two-way split needs a two-way breakout; a four-way
split needs a four-way. Mixing them produces ports that never link, with no error
message that says why.&lt;/p&gt;

&lt;p&gt;Two practical points. Splitter cables are physically bulkier at the switch end
and have real bend-radius limits, which matters in a dense cabinet more than the
datasheet suggests. And a partially split switch — some cages split, some not —
is entirely legal and genuinely useful, but it makes the cabling records the only
source of truth about which is which. Label at both ends.&lt;/p&gt;

&lt;h2 id=&quot;what-it-does-to-the-subnet&quot;&gt;What it does to the subnet&lt;/h2&gt;

&lt;p&gt;This is the part that gets underestimated.&lt;/p&gt;

&lt;p&gt;Splitting increases the number of endpoints the subnet manager has to
enumerate, assign and route. Every added endpoint consumes at least one LID,
occupies entries in the forwarding tables of every switch on any path to it, and
adds work to every sweep.&lt;/p&gt;

&lt;p&gt;Three consequences:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Routing table pressure.&lt;/strong&gt; Switch forwarding tables have a hard capacity.
Doubling endpoint count moves you toward that limit, and the failure mode when
you reach it is not graceful.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Longer sweeps.&lt;/strong&gt; More ports means more discovery and more table
distribution. On a large fabric the difference between a sweep that finishes
quickly and one that does not is operationally significant, because sweeps
happen on every topology change.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Changed oversubscription.&lt;/strong&gt; If you split leaf ports facing compute and leave
the spine uplinks alone, you have just changed the ratio between edge and core
capacity. That number should be a decision, not a side effect.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Work out the LID budget and the forwarding table headroom before you split, not
after. Both are countable in advance, and both are much cheaper to discover on
paper than during a maintenance window.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;verifying-it-took&quot;&gt;Verifying it took&lt;/h2&gt;

&lt;p&gt;After the reboot, confirm the switch presents what you expect, then confirm the
subnet manager agrees:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibswitches
iblinkinfo | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; Active
ibdiagnet &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; /var/tmp/ibdiag-after
sminfo
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Compare the port count and the link rates against what you intended. A cage that
split but whose cable did not match will show ports in a down or polling state
rather than an error, so count active links rather than trusting the absence of
errors.&lt;/p&gt;

&lt;p&gt;Then check the error counters are clean from a fresh baseline, because the old
baseline referred to ports that no longer exist:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ibclearerrors
&lt;span class=&quot;c&quot;&gt;# run representative traffic&lt;/span&gt;
ibqueryerrors
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;when-not-to-split&quot;&gt;When not to split&lt;/h2&gt;

&lt;p&gt;If your endpoints have full-rate adapters, splitting throws away half of each
one’s capability for no gain. If you are already close to the forwarding table
limit, splitting is how you find the limit. And if the fabric is a single switch
with no maintenance window available, the reboot is the whole problem and no
amount of planning removes it.&lt;/p&gt;

&lt;p&gt;Splitting is a good answer to a specific question: you have more endpoints than
ports, and those endpoints do not need full rate. Outside that question it is
usually the wrong tool.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>ConnectX firmware management: PSID, burn, reset, verify</title>
    <link href="https://www.wirewalk.com/writing/connectx-firmware-management-mft/"/>
    <published>2026-09-04T06:00:00-04:00</published>
    <updated>2026-09-04T06:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/connectx-firmware-management-mft/</id>
    <summary>Four commands do the work, and the one people skip is the one that makes the update real. The order matters, and the PSID matters more than the version number.</summary>
    <content type="html">&lt;p&gt;Firmware on a ConnectX adapter is not one thing. There is the image in flash,
the device configuration stored alongside it, and the running state, and all
three can disagree after a partial update. A card that reports the new version
but behaves like the old one is common, and the reason is almost always that
nothing reset it.&lt;/p&gt;

&lt;h2 id=&quot;start-the-driver-find-the-device&quot;&gt;Start the driver, find the device&lt;/h2&gt;

&lt;p&gt;Nothing in the MFT toolset works until the MST kernel module is loaded and the
device nodes exist.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mst start
&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mst status &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That prints paths under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/dev/mst/&lt;/code&gt; of the form &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mt4123_pciconf0&lt;/code&gt; for a
ConnectX-6 or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mt4125_pciconf0&lt;/code&gt; for a ConnectX-7, alongside the PCI address, the
RDMA device name and the network interface. Everything downstream takes one of
those paths as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-d&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The MFT version matters here in a way that is easy to miss. A toolset older than
the silicon will either not enumerate the device or will enumerate it and then
refuse the image. If &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mst status&lt;/code&gt; shows nothing for a card you can see in
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lspci&lt;/code&gt;, update MFT before debugging anything else.&lt;/p&gt;

&lt;h2 id=&quot;query-before-touching-anything&quot;&gt;Query before touching anything&lt;/h2&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxfwmanager &lt;span class=&quot;nt&quot;&gt;--query&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three fields matter, and only one of them is the version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PSID&lt;/strong&gt; is the parameter set identifier, and it binds an image to a specific
hardware variant. Two cards with the same model number and the same firmware
version can carry different PSIDs, because one is an OEM part with different
defaults, a different port configuration or a different flash layout. Firmware
is published per PSID. This is the field to match on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FW version&lt;/strong&gt; appears in two forms: what is in flash, and what is running. When
those differ, an image has been burned and not activated. That is the state a
reset clears, and it is the most common reason a machine comes back from an
update behaving exactly as it did before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Status&lt;/strong&gt; reports whether an update is available, which only means anything
when you are querying online.&lt;/p&gt;

&lt;p&gt;Where the host has outbound access:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxfwmanager &lt;span class=&quot;nt&quot;&gt;--online&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-u&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;On an isolated fabric, which is the normal case, fetch the image matching the
PSID and burn it explicitly:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxfwmanager &lt;span class=&quot;nt&quot;&gt;-u&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; fw-ConnectX7-rel.bin
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-psid-trap&quot;&gt;The PSID trap&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;flint&lt;/code&gt; will burn an image whose PSID does not match the card, if you insist:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;flint &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; image.bin burn                       &lt;span class=&quot;c&quot;&gt;# refuses&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;flint &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; image.bin &lt;span class=&quot;nt&quot;&gt;--allow_psid_change&lt;/span&gt; burn   &lt;span class=&quot;c&quot;&gt;# does not&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The refusal is a feature. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--allow_psid_change&lt;/code&gt; exists for genuine cases —
recovering a card whose flash was corrupted, converting a part with vendor
approval — and it is not the flag to reach for because a burn failed. A card
running an image built for a different variant can come up with the wrong port
configuration, the wrong link type, or not come up at all, and the way back
involves the same tool in a much less comfortable mode.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Match on PSID, not on model. Two adapters that look identical in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lspci&lt;/code&gt; can
need different images, and any fleet bought in two tranches is likely to hold
both.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;device-configuration-is-separate-from-firmware&quot;&gt;Device configuration is separate from firmware&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlxconfig&lt;/code&gt; reads and writes non-volatile settings that survive firmware updates
and are not part of the image:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxconfig &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 query
&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxconfig &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 &lt;span class=&quot;nb&quot;&gt;set &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;LINK_TYPE_P1&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;On a VPI-capable part, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LINK_TYPE_P1&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LINK_TYPE_P2&lt;/code&gt; decide whether a port
comes up as InfiniBand or Ethernet. Getting that wrong gives you a port that
never links and no error that explains why. Other settings worth knowing: SR-IOV
enablement and virtual function count, ATS, and the PCIe and performance knobs
that appear in vendor guidance from time to time.&lt;/p&gt;

&lt;p&gt;Two things to hold onto. A firmware update can introduce new configuration keys
carrying defaults you did not choose, so query after an update as well as
before. And &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlxconfig&lt;/code&gt; changes also need a reset to take effect, which brings
us to the step people skip.&lt;/p&gt;

&lt;h2 id=&quot;reset-and-when-a-reset-is-not-enough&quot;&gt;Reset, and when a reset is not enough&lt;/h2&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxfwreset &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 query
&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxfwreset &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 &lt;span class=&quot;nt&quot;&gt;-y&lt;/span&gt; reset
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;query&lt;/code&gt; subcommand reports which reset levels the card and the platform will
actually support in the current state, which is more useful than it sounds: a
warm reset the platform cannot perform is reported rather than attempted.&lt;/p&gt;

&lt;p&gt;Some combinations of card, host and BIOS cannot do a live reset at all. In that
case the update is not applied until a full power cycle, and on some platforms
specifically an AC power cycle rather than a warm reboot. If you have burned an
image, run a reset, and the query still shows the old running version, stop
reasoning about firmware and go find out whether the machine actually
power-cycled.&lt;/p&gt;

&lt;h2 id=&quot;verify-against-running-state-from-three-directions&quot;&gt;Verify against running state, from three directions&lt;/h2&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;mlxfwmanager &lt;span class=&quot;nt&quot;&gt;--query&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /dev/mst/mt4125_pciconf0 | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;PSID|FW&apos;&lt;/span&gt;
ethtool &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; &amp;lt;netdev&amp;gt;
ibv_devinfo &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; &amp;lt;rdma-dev&amp;gt; | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; fw_ver
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three sources because they read from different places. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mlxfwmanager&lt;/code&gt; reads
flash, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ethtool&lt;/code&gt; reports what the driver bound, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibv_devinfo&lt;/code&gt; reports what the
verbs layer sees. Agreement across all three is what “the update worked” means.
Disagreement between the first and the other two is the unreset-card signature.&lt;/p&gt;

&lt;h2 id=&quot;doing-this-across-a-fleet&quot;&gt;Doing this across a fleet&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Inventory by PSID, not by hostname.&lt;/strong&gt; Group the estate by PSID and firmware
version before planning anything. The distribution is reliably less uniform than
the purchase orders suggest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stagger by failure domain.&lt;/strong&gt; Adapters in the same rack, on the same leaf, or
serving the same storage class should not reset in the same window. A reset
drops the link and the subnet manager re-sweeps; doing that everywhere at once
turns a routine update into a fabric event.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Record before and after.&lt;/strong&gt; Query output for every card, stored and dated. The
value shows up months later, when something is behaving oddly and the only
question that matters is whether it changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pair firmware with driver.&lt;/strong&gt; Vendor release notes pair firmware versions with
specific OFED or in-box driver versions. Updating one and not the other is
supported far less often than it is done, and the symptoms get blamed on the
application.&lt;/p&gt;

&lt;h2 id=&quot;what-this-does-not-cover&quot;&gt;What this does not cover&lt;/h2&gt;

&lt;p&gt;Switch firmware is a different toolchain and a different risk profile, and
deserves its own written procedure. So does recovery from an interrupted burn —
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;flint&lt;/code&gt; in its more surgical modes, and the cases where a card has to come up on
a recovery image. That one is worth writing down once, by whoever has just done
it, rather than improvised under pressure.&lt;/p&gt;

&lt;p&gt;Command syntax drifts between MFT releases. What is above describes the shape of
the tools rather than one version, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--help&lt;/code&gt; on the release you are running
is authoritative over anything written down elsewhere, this included.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>nccl-tests: bus bandwidth, and the one column worth reading</title>
    <link href="https://www.wirewalk.com/writing/nccl-tests-collective-bandwidth/"/>
    <published>2026-09-04T00:00:00-04:00</published>
    <updated>2026-09-04T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/nccl-tests-collective-bandwidth/</id>
    <summary>Two of the columns nccl-tests prints look like bandwidth. Only one of them is comparable across job sizes, and picking the wrong one is how a healthy cluster gets declared broken.</summary>
    <content type="html">&lt;p&gt;Distributed training spends a large share of its wall clock inside collective
operations. In data-parallel training the dominant one is an all-reduce over the
gradients, once per step, sized by the model. If that collective is slower than
the hardware allows, every step is slower, and no amount of framework tuning
recovers it.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nccl-tests&lt;/code&gt; from NVIDIA is how you measure that path on its own, with no model
and no framework in the way.&lt;/p&gt;

&lt;h2 id=&quot;building-it&quot;&gt;Building it&lt;/h2&gt;

&lt;p&gt;Build against MPI if you intend to run across more than one node, which is the
only interesting case:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;git clone https://github.com/NVIDIA/nccl-tests
&lt;span class=&quot;nb&quot;&gt;cd &lt;/span&gt;nccl-tests
make &lt;span class=&quot;nv&quot;&gt;MPI&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1 &lt;span class=&quot;nv&quot;&gt;MPI_HOME&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;/opt/openmpi &lt;span class=&quot;nv&quot;&gt;CUDA_HOME&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;/usr/local/cuda
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That produces one binary per collective: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;all_reduce_perf&lt;/code&gt;,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;all_gather_perf&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;reduce_scatter_perf&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;broadcast_perf&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;alltoall_perf&lt;/code&gt;,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sendrecv_perf&lt;/code&gt; and a few others. Start with all-reduce, because it is what
training actually does.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 16 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:node &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./build/all_reduce_perf &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; 8 &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; 8G &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-g&lt;/span&gt; 1 &lt;span class=&quot;nt&quot;&gt;-w&lt;/span&gt; 5 &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; 20
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;One rank per GPU, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-g 1&lt;/code&gt;. The sweep runs from 8 bytes to 8 GB doubling each
step, with five warmup iterations discarded and twenty timed. Warmup is not
optional: the first call pays for buffer registration and connection setup, and
including it drags the small-message numbers into nonsense.&lt;/p&gt;

&lt;h2 id=&quot;the-two-bandwidth-columns&quot;&gt;The two bandwidth columns&lt;/h2&gt;

&lt;p&gt;The output has, for both out-of-place and in-place variants, a time, an
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;algbw&lt;/code&gt;, a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt; and a correctness count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;algbw&lt;/code&gt;&lt;/strong&gt; is algorithm bandwidth. It is simply the message size divided by the
elapsed time. It answers “how fast did my data get processed,” and it is the
number that looks intuitive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt;&lt;/strong&gt; is bus bandwidth. It is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;algbw&lt;/code&gt; scaled by a factor that accounts for
how much traffic the collective genuinely has to move across the interconnect,
given the number of ranks. For a ring all-reduce, each byte has to traverse the
ring twice, once in the reduce-scatter phase and once in the all-gather phase,
and each rank ends up sending &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(n-1)/n&lt;/code&gt; of the buffer in each phase:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;all-reduce         busbw = algbw × 2(n-1)/n
all-gather         busbw = algbw × (n-1)/n
reduce-scatter     busbw = algbw × (n-1)/n
all-to-all         busbw = algbw × (n-1)/n
broadcast, reduce  busbw = algbw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The consequence is the reason &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt; exists. As you add ranks, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;algbw&lt;/code&gt; for a
fixed message size falls, because there is more work to do per byte. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt;
stays roughly flat, because it is measuring what the wire is doing. That makes
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt; comparable against your hardware’s per-link capability and comparable
between a two-node run and a hundred-node run.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Track &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt;. Someone reporting that all-reduce “got slower when we scaled up”
is almost always reading &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;algbw&lt;/code&gt;, where a decline is arithmetic rather than a
fault. Compare &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt; at the same message size across job sizes; if that holds,
the fabric is behaving.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;reading-the-sweep&quot;&gt;Reading the sweep&lt;/h2&gt;

&lt;p&gt;The curve has two regimes and the transition tells you where your model sits.&lt;/p&gt;

&lt;p&gt;At &lt;strong&gt;small sizes&lt;/strong&gt; the result is latency-bound. Time is nearly flat as size
increases, so bandwidth rises roughly linearly. What you are measuring is the
fixed cost of a collective: kernel launch, synchronization, and the depth of the
communication pattern. NCCL uses tree algorithms here specifically because tree
latency grows logarithmically with rank count while ring latency grows linearly.&lt;/p&gt;

&lt;p&gt;At &lt;strong&gt;large sizes&lt;/strong&gt; the result is bandwidth-bound and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt; flattens into a
plateau. That plateau is the number to record and to compare against what the
links should deliver.&lt;/p&gt;

&lt;p&gt;Work out where your gradient all-reduce lands. For a bucketed data-parallel
setup, the bucket size is the message size, and it is frequently in the region
where neither regime dominates cleanly. Moving the bucket size is a real tuning
knob, and this sweep tells you which direction to move it.&lt;/p&gt;

&lt;h2 id=&quot;when-the-plateau-is-too-low&quot;&gt;When the plateau is too low&lt;/h2&gt;

&lt;p&gt;The plateau being well below what the hardware should do almost always comes
down to one of a short list.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NCCL chose the wrong interface.&lt;/strong&gt; On a node with a management NIC and several
fabric adapters, NCCL can select the wrong one. Verify rather than assume:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;NCCL_DEBUG&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;INFO &lt;span class=&quot;nv&quot;&gt;NCCL_DEBUG_SUBSYS&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;INIT,NET,GRAPH &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 16 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:8:node ./build/all_reduce_perf &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; 1G &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; 1G &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; 20
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The output names each adapter it selected and prints the rings it constructed.
If it is using a socket transport where you expected InfiniBand, that is your
answer. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCCL_IB_HCA&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCCL_SOCKET_IFNAME&lt;/code&gt; constrain the choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GPUDirect RDMA is not engaged.&lt;/strong&gt; Without it, every message stages through a
host bounce buffer. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INIT&lt;/code&gt; log says whether GDR is in use per device.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCCL_NET_GDR_LEVEL&lt;/code&gt; controls how aggressively it is applied relative to PCIe
topology distance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PCIe ACS is enabled.&lt;/strong&gt; Access Control Services forces peer-to-peer transfers
up through the root complex instead of across the switch, which quietly destroys
intra-node bandwidth. It is a BIOS or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;setpci&lt;/code&gt; change, and it is worth checking
on every new node type.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The topology is not what you think.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCCL_TOPO_DUMP_FILE&lt;/code&gt; writes out the
topology NCCL detected. Compare it against &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nvidia-smi topo -m&lt;/code&gt; and against how
the machine is actually cabled. Ring construction is only as good as the
topology discovery underneath it.&lt;/p&gt;

&lt;h2 id=&quot;correctness-is-a-column-too&quot;&gt;Correctness is a column too&lt;/h2&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;#wrong&lt;/code&gt; column exists for a reason. Run with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-c 1&lt;/code&gt; at least occasionally.
A collective that returns wrong values under load is a hardware or driver fault,
and it is a much more serious finding than a disappointing bandwidth plateau. It
is also the kind of thing that shows up in a training run as a loss curve that
diverges for no apparent reason, days later, after a great deal of compute has
been spent.&lt;/p&gt;

&lt;h2 id=&quot;what-to-do-with-the-result&quot;&gt;What to do with the result&lt;/h2&gt;

&lt;p&gt;Record the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;busbw&lt;/code&gt; plateau, the message size where the curve leaves the
latency-bound regime, the rank count, and the driver, NCCL and firmware
versions. Re-run it after every driver update and every firmware change.&lt;/p&gt;

&lt;p&gt;Collectives degrade quietly. Nothing fails, nothing logs an error, and step time
drifts up by a percentage nobody notices until the quarter’s throughput is
compared against the previous one. A five-minute benchmark on a known baseline
catches it the same day.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>IOR and mdtest: why your storage benchmark says the filesystem is fine</title>
    <link href="https://www.wirewalk.com/writing/ior-mdtest-parallel-storage/"/>
    <published>2026-09-03T00:00:00-04:00</published>
    <updated>2026-09-03T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/ior-mdtest-parallel-storage/</id>
    <summary>Streaming bandwidth is the easiest number to produce and the least likely to describe an AI training workload. The metadata rate usually matters more, and almost nobody measures it.</summary>
    <content type="html">&lt;p&gt;Two tools cover most of what you need from a parallel filesystem. IOR measures
bandwidth by moving large blocks. mdtest measures metadata by creating, stating
and removing enormous numbers of files. Both come from the same repository and
build together.&lt;/p&gt;

&lt;p&gt;The common failure is running only IOR, getting a large number, and concluding
the storage is healthy. For a training workload reading millions of small files,
that number describes almost nothing you care about.&lt;/p&gt;

&lt;h2 id=&quot;defeating-the-cache-which-is-most-of-the-work&quot;&gt;Defeating the cache, which is most of the work&lt;/h2&gt;

&lt;p&gt;The single most important thing about IOR is that a naive invocation measures
page cache. The client caches the write, reports success, and IOR reports a
bandwidth figure that reflects memory speed rather than storage.&lt;/p&gt;

&lt;p&gt;Three flags matter:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 128 ./ior &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; POSIX &lt;span class=&quot;nt&quot;&gt;-w&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; 16g &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; 4m &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; 1 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-C&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; 3 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; /gpfs/scratch/iortest/file
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-e&lt;/code&gt; calls &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fsync&lt;/code&gt; before the write phase timer stops, so the write is on
stable storage rather than in the client’s cache.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-C&lt;/code&gt; reorders tasks between the write and read phases, so no rank reads back
the file it just wrote and finds it sitting in local page cache.&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-b&lt;/code&gt; sets the block size per task. Make the aggregate substantially larger than
the total client memory, or the read phase is served from RAM.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without those, your read bandwidth will look wonderful and mean nothing. If a
result looks impossibly good, this is the reason, every time.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-F&lt;/code&gt; selects file-per-process. The alternative, a single shared file, tests a
different thing: lock contention and the filesystem’s ability to handle
concurrent writers to one inode. Both are worth measuring, because applications
do both, and on some filesystems they perform very differently.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-t&lt;/code&gt; is the transfer size, the size of each individual read or write call.
Sweeping it is how you find where the client stops being able to fill the pipe.&lt;/p&gt;

&lt;h2 id=&quot;the-metadata-problem&quot;&gt;The metadata problem&lt;/h2&gt;

&lt;p&gt;Now the part that usually explains the complaint you were called about.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 128 ./mdtest &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; 10000 &lt;span class=&quot;nt&quot;&gt;-u&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-z&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; 8 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; 3 &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; /gpfs/scratch/mdtest
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;mdtest reports operations per second for file creation, stat, read, and removal,
plus the directory equivalents. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-u&lt;/code&gt; gives each task its own working directory,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-z&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-b&lt;/code&gt; control the depth and branching of the directory tree, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-n&lt;/code&gt; is
files per task.&lt;/p&gt;

&lt;p&gt;Run it with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-u&lt;/code&gt; and without. The difference is often dramatic, and it is one of
the more actionable results you can produce, because it tells you directly
whether the application’s directory layout is the problem. A single directory
holding a very large number of entries serializes on the metadata path in a way
that the same files spread across a tree do not.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;For AI training, metadata rate and small-file read latency usually predict
throughput far better than streaming bandwidth. A dataset of millions of small
files spends its time in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;open&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stat&lt;/code&gt;, not in moving bytes. If you measure
one thing, measure that.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;interpreting-the-gap-between-the-two&quot;&gt;Interpreting the gap between the two&lt;/h2&gt;

&lt;p&gt;The relationship between the IOR and mdtest results is more informative than
either alone.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;High bandwidth, high metadata rate.&lt;/strong&gt; The filesystem is healthy. If an
application is still slow, the problem is in the application’s access pattern
or in the client, not the storage.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;High bandwidth, poor metadata rate.&lt;/strong&gt; Classic parallel filesystem shape.
The data path is fine and the metadata path is the bottleneck. Look at
metadata server placement, whether metadata sits on the same media as data,
and how the workload lays out directories.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Poor bandwidth, high metadata rate.&lt;/strong&gt; Usually the network or client
configuration rather than the filesystem: check RDMA is engaged, check the
mount options, check that clients are not falling back to a slower transport.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Both poor.&lt;/strong&gt; Look at the client before the server. A single-client run that
is also slow points at the node; a slow aggregate with fast single-client
results points at the server or the fabric between.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;running-it-so-the-numbers-survive-scrutiny&quot;&gt;Running it so the numbers survive scrutiny&lt;/h2&gt;

&lt;p&gt;Use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-i&lt;/code&gt; to repeat and report the spread, not a single lucky run. Record the
client count, the ranks per client, the transfer and block sizes, the filesystem
version and the mount options alongside the result. Run against a directory with
the same striping or placement policy the real workload will use, because a
default-policy test directory can behave nothing like a production dataset.&lt;/p&gt;

&lt;p&gt;And run the same configuration on a schedule. A storage benchmark taken once
during acceptance is an anecdote. The same benchmark taken monthly is a baseline,
and a baseline is the only thing that turns “it feels slower lately” into a
number you can act on.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>HPL-MxP: reading a mixed-precision result without fooling yourself</title>
    <link href="https://www.wirewalk.com/writing/hpl-mxp-mixed-precision/"/>
    <published>2026-09-01T00:00:00-04:00</published>
    <updated>2026-09-01T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/hpl-mxp-mixed-precision/</id>
    <summary>HPL-MxP reports a much larger number than HPL on the same machine. Knowing exactly where that factor comes from is the difference between a useful measurement and a marketing slide.</summary>
    <content type="html">&lt;p&gt;HPL-MxP, previously published as HPL-AI, solves the same problem as HPL and
returns a much larger number. The machine did not get faster. Understanding the
gap is the whole value of the benchmark, and misunderstanding it is how the
number ends up in a slide deck meaning nothing.&lt;/p&gt;

&lt;h2 id=&quot;what-it-does-differently&quot;&gt;What it does differently&lt;/h2&gt;

&lt;p&gt;Classic HPL performs the LU factorization entirely in double precision. HPL-MxP
performs the factorization in a lower precision, then repairs the answer.&lt;/p&gt;

&lt;p&gt;The sequence is:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Factorize the matrix in reduced precision, which is where the hardware’s
tensor or matrix units are fastest.&lt;/li&gt;
  &lt;li&gt;Use that low-precision factorization as a preconditioner.&lt;/li&gt;
  &lt;li&gt;Run an iterative refinement loop, in practice a Krylov solver such as GMRES,
until the residual meets the same accuracy standard classic HPL must meet.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The final answer satisfies the same correctness criterion. The path there spent
most of its time in arithmetic the hardware executes at several times the rate
of double precision.&lt;/p&gt;

&lt;p&gt;This is not a trick invented for benchmarking. Mixed-precision iterative
refinement is a legitimate, long-standing technique in numerical linear algebra,
and it is used in production solvers. The benchmark exists to reward hardware
that implements it well.&lt;/p&gt;

&lt;h2 id=&quot;where-the-big-number-comes-from&quot;&gt;Where the big number comes from&lt;/h2&gt;

&lt;p&gt;Here is the part that matters, and the part most often skipped.&lt;/p&gt;

&lt;p&gt;HPL-MxP does not report the number of low-precision operations it performed. It
reports the operation count of the &lt;strong&gt;equivalent double-precision HPL
algorithm&lt;/strong&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(2/3)N³ + 2N²&lt;/code&gt;, divided by the time the mixed-precision run
actually took.&lt;/p&gt;

&lt;p&gt;So the metric answers a specific question: &lt;em&gt;if you needed an HPL-accurate
solution to this system, how fast did you get one?&lt;/em&gt; That is a fair and useful
question. It is not the same question as “how many double-precision floating
point operations per second can this machine sustain,” and the answer must never
be substituted for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rmax&lt;/code&gt;.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;A machine’s HPL-MxP rate and its HPL &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rmax&lt;/code&gt; are not comparable quantities, and
the ratio between them is not an efficiency. It is a measure of how much the
low-precision units outrun the double-precision units on that hardware, filtered
through how well the refinement converged.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-actually-determines-the-result&quot;&gt;What actually determines the result&lt;/h2&gt;

&lt;p&gt;Three things move the number, and only one of them is raw hardware.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The precision used for the factorization.&lt;/strong&gt; FP16, BF16 and TF32 have different
dynamic range and different mantissa widths, and they are not interchangeable.
BF16 keeps FP32’s exponent range with a short mantissa, which makes it robust
against overflow but coarse. FP16 has a finer mantissa and a much narrower
range, so it needs scaling to stay numerically safe. Which one performs better
is a hardware question; which one converges better is a numerical one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many refinement iterations were needed.&lt;/strong&gt; Every iteration costs time and
none of it counts toward the reported operation count, because the count is
fixed at the classic HPL figure. A run that converges in a handful of iterations
posts a strong number. A run that struggles posts a weak one even on identical
hardware. Convergence depends on the conditioning of the generated matrix and on
the precision choice, which is why the benchmark specifies how the matrix is
produced.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The usual HPL parameters.&lt;/strong&gt; Problem size, block size and process grid still
matter, and the tuning intuition carries over, though the optimum block size is
generally larger because the point is to keep wide matrix units saturated.&lt;/p&gt;

&lt;h2 id=&quot;when-it-is-worth-running&quot;&gt;When it is worth running&lt;/h2&gt;

&lt;p&gt;Run it if you are buying or operating hardware where the low-precision units are
a large fraction of what you are paying for. On a GPU cluster intended for
training workloads, the double-precision rate may be close to irrelevant to
everything the machine will actually do, and HPL alone will undersell it badly.&lt;/p&gt;

&lt;p&gt;Run it as a companion to HPL rather than a replacement. The pair tells you two
different things about the same machine: HPL gives you the double-precision
ceiling and a hardware health check, HPL-MxP gives you an indication of how well
the reduced-precision path performs on a problem with a verifiable answer.&lt;/p&gt;

&lt;p&gt;Do not run it expecting it to predict training throughput. It does not model
attention, it does not model collectives at scale, it does not model the data
pipeline, and it does not model memory capacity pressure. For that question,
measure collectives directly with nccl-tests and then measure your actual model.&lt;/p&gt;

&lt;h2 id=&quot;reporting-it-responsibly&quot;&gt;Reporting it responsibly&lt;/h2&gt;

&lt;p&gt;If you publish an HPL-MxP figure, publish the precision used for the
factorization, the iteration count, the problem size, and the classic HPL result
from the same machine alongside it. Without those four pieces, the number is
unfalsifiable, and an unfalsifiable performance number is worth what it costs to
produce.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>HPCG: the benchmark that tells you what the memory system can do</title>
    <link href="https://www.wirewalk.com/writing/hpcg-benchmarking/"/>
    <published>2026-08-26T00:00:00-04:00</published>
    <updated>2026-08-26T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/hpcg-benchmarking/</id>
    <summary>HPCG reaches a small percentage of peak and that is the entire point. It measures the parts of the machine that actually limit most scientific codes.</summary>
    <content type="html">&lt;p&gt;High Performance Conjugate Gradients was written by the HPL authors as a
deliberate counterweight to their own benchmark. Where HPL is dense, regular
and compute-bound, HPCG is sparse, irregular, and limited by memory bandwidth
and communication latency. It exists because procurement decisions were being
made on a number that predicted almost nothing about real workloads.&lt;/p&gt;

&lt;p&gt;The headline result is startling the first time you see it. A machine that hits
a large fraction of peak on HPL will typically return a low single-digit
percentage of peak on HPCG. Nothing is wrong. That gap is the honest distance
between what the floating point units can do and what the memory system can feed
them.&lt;/p&gt;

&lt;h2 id=&quot;what-it-actually-runs&quot;&gt;What it actually runs&lt;/h2&gt;

&lt;p&gt;HPCG solves a sparse linear system with a preconditioned conjugate gradient
method. The preconditioner is a geometric multigrid with a symmetric
Gauss-Seidel smoother, and that choice is what makes the benchmark hard in an
interesting way.&lt;/p&gt;

&lt;p&gt;Symmetric Gauss-Seidel is inherently sequential in its data dependencies. You
cannot simply vectorize your way out of it the way you can with a dense matrix
multiply. Implementations have to reorder, color or otherwise restructure the
sweep to expose parallelism, and doing that changes the convergence behavior,
which the benchmark checks. The rules constrain how much you are allowed to
change, which is what stops HPCG from degenerating into the same optimization
contest HPL became.&lt;/p&gt;

&lt;p&gt;The result is a benchmark dominated by sparse matrix-vector products, which
means streaming large arrays with indirect indexing. Arithmetic intensity is
low. Every cycle spent waiting on a cache miss is a cycle the floating point
units sit idle.&lt;/p&gt;

&lt;h2 id=&quot;configuration&quot;&gt;Configuration&lt;/h2&gt;

&lt;p&gt;The input is small. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;hpcg.dat&lt;/code&gt; takes a local problem size per rank and a run
time:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;HPCG benchmark input file
Sandia National Laboratories; University of Tennessee, Knoxville
104 104 104
1800
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The three numbers are the local subdomain dimensions &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nx ny nz&lt;/code&gt;, and they must
each be a multiple of eight. The fourth is the target run time in seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sizing the local domain.&lt;/strong&gt; The subdomain has to be large enough that it does
not fit comfortably in cache, or you are benchmarking cache and the result is
meaningless. A reasonable approach is to size it so the working set is a
substantial fraction of the memory you want to represent per rank, then confirm
that increasing it further does not change the reported rate much. When the
number stops moving, you are measuring DRAM rather than last-level cache.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run time.&lt;/strong&gt; For an official submission the timed run must be at least 1800
seconds. For internal use you can run shorter, but be aware that short runs
overweight the setup phase and flatter the machine slightly.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;HPCG will happily produce a large number if you give it a small local domain.
This is the single most common way sites accidentally publish a wrong HPCG
result. Sweep the local size upward and take the value from the plateau, not
from the peak.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-the-output-tells-you&quot;&gt;What the output tells you&lt;/h2&gt;

&lt;p&gt;HPCG writes a YAML file with a full breakdown. The headline &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;GFLOP/s&lt;/code&gt; rating is
what gets quoted, but the per-phase timings underneath are more useful
operationally, because they separate the sparse matrix-vector product, the
symmetric Gauss-Seidel sweeps, the dot products and the halo exchange.&lt;/p&gt;

&lt;p&gt;Those phases fail differently:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;The &lt;strong&gt;matrix-vector and smoother&lt;/strong&gt; phases track memory bandwidth. If they
regress, look at DIMM population, memory clock, NUMA balance and whether
something changed in the BIOS power profile.&lt;/li&gt;
  &lt;li&gt;The &lt;strong&gt;dot products&lt;/strong&gt; are global reductions, so they track collective latency
and, more importantly, jitter. A single noisy node shows up here first.&lt;/li&gt;
  &lt;li&gt;The &lt;strong&gt;halo exchange&lt;/strong&gt; tracks neighbor communication, which is point-to-point
and usually fine unless the topology mapping is poor.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That decomposition is why HPCG is worth running even when nobody is asking for
a TOP500-style number. It is a repeatable, structured way of asking whether the
memory subsystem and the collective path are behaving the way they did last
quarter.&lt;/p&gt;

&lt;h2 id=&quot;where-it-fits-against-hpl&quot;&gt;Where it fits against HPL&lt;/h2&gt;

&lt;p&gt;Run both. They fail for different reasons and finding the difference is
diagnostic.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt;HPL&lt;/th&gt;
      &lt;th&gt;HPCG&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Limited by&lt;/td&gt;
      &lt;td&gt;Floating point throughput&lt;/td&gt;
      &lt;td&gt;Memory bandwidth, latency&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Fraction of peak&lt;/td&gt;
      &lt;td&gt;High&lt;/td&gt;
      &lt;td&gt;Low&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Sensitive to&lt;/td&gt;
      &lt;td&gt;Clocks, BLAS, thermals&lt;/td&gt;
      &lt;td&gt;DIMM config, NUMA, jitter&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Representative of&lt;/td&gt;
      &lt;td&gt;Dense linear algebra&lt;/td&gt;
      &lt;td&gt;Most sparse scientific codes&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;If HPL holds steady and HPCG drops, suspect memory configuration or a change in
NUMA behavior rather than the processors. If both drop together, suspect clocks,
power or cooling. If HPCG holds and HPL drops, suspect the math library or a
thermal limit that only sustained dense work reaches.&lt;/p&gt;

&lt;p&gt;Neither benchmark predicts your applications. Together they bracket the machine
between what it can do when everything is favorable and what it can do when
almost nothing is, and most real codes live somewhere in between.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>HPL: what Linpack actually measures, and what it is good for</title>
    <link href="https://www.wirewalk.com/writing/hpl-benchmarking/"/>
    <published>2026-08-19T00:00:00-04:00</published>
    <updated>2026-08-19T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/hpl-benchmarking/</id>
    <summary>HPL is a poor predictor of application performance and an excellent acceptance test. Those two statements are both true, and confusing them is how people end up tuning the wrong thing.</summary>
    <content type="html">&lt;p&gt;High Performance Linpack solves a dense system of linear equations by LU
factorization with partial pivoting, in double precision, distributed across a
process grid. It is the benchmark behind the TOP500 list, which is why it
carries more political weight than technical weight.&lt;/p&gt;

&lt;p&gt;The useful framing is this: HPL is a poor predictor of how your applications
will perform, and an excellent way to find out whether a machine is broken.
Both things are true at once, and most arguments about HPL come from people
holding one half of that and assuming the other half is wrong.&lt;/p&gt;

&lt;h2 id=&quot;why-it-flatters-hardware&quot;&gt;Why it flatters hardware&lt;/h2&gt;

&lt;p&gt;The work HPL does is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(2/3)N³ + 2N²&lt;/code&gt; floating point operations, and almost all of
it lands in large dense matrix multiplies. That is the single most favorable
operation any processor has: high arithmetic intensity, perfectly predictable
access, and vendor-tuned kernels behind it.&lt;/p&gt;

&lt;p&gt;So HPL reaches a high fraction of theoretical peak, and a well-tuned run on
sensible hardware lands in the region of sixty to ninety percent of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rpeak&lt;/code&gt;.
Very little real scientific code behaves this way. Most applications are limited
by memory bandwidth, sparse or irregular access, latency, or synchronization,
none of which HPL exercises meaningfully.&lt;/p&gt;

&lt;p&gt;That is precisely why HPCG was created as a counterweight, and it is why quoting
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rmax&lt;/code&gt; as a procurement figure of merit for a mixed workload is close to
meaningless.&lt;/p&gt;

&lt;h2 id=&quot;the-parameters-that-matter&quot;&gt;The parameters that matter&lt;/h2&gt;

&lt;p&gt;Everything lives in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;HPL.dat&lt;/code&gt;. Most of its knobs are noise; four are not.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Parameter&lt;/th&gt;
      &lt;th&gt;What it controls&lt;/th&gt;
      &lt;th&gt;Practical guidance&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Problem size&lt;/td&gt;
      &lt;td&gt;Fill most of aggregate memory, leaving headroom&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NB&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Block size&lt;/td&gt;
      &lt;td&gt;Match the BLAS kernel’s preference&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;P&lt;/code&gt; × &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Q&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Process grid&lt;/td&gt;
      &lt;td&gt;As close to square as possible, with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;P&lt;/code&gt; ≤ &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Q&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;BCAST&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Broadcast algorithm&lt;/td&gt;
      &lt;td&gt;Try a few; the ring variants usually win at scale&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;strong&gt;Sizing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N&lt;/code&gt;.&lt;/strong&gt; The matrix is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N × N&lt;/code&gt; in eight-byte doubles, so memory is
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;8N²&lt;/code&gt; bytes plus workspace. Solve for the memory you want to use:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;N ≈ sqrt( total_bytes × utilization / 8 )
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Utilization around 0.8 is a reasonable starting point. Push it too high and the
run swaps or the OOM killer intervenes several hours in; too low and you spend
proportionally more time in the parts of the algorithm that do not scale, and
the result understates the machine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Choosing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NB&lt;/code&gt;.&lt;/strong&gt; This is the blocking factor the factorization uses, and the
right value is whatever the underlying BLAS wants. CPU runs with a tuned library
typically want something in the low hundreds. GPU runs, particularly through a
vendor container, want considerably larger blocks, because the point is to keep
the device busy with big multiplies. Do not carry a CPU value across to a GPU
run and expect it to hold.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The grid.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;P × Q&lt;/code&gt; must equal the number of MPI ranks. The panel
factorization runs down columns of the grid and the update broadcasts across
rows, so a grid that is far from square puts a disproportionate amount of the
serial-ish work in one dimension. Square, or slightly wider than tall, is the
rule.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Sizing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N&lt;/code&gt; to fill memory means a full run takes hours. Do your parameter sweeps
at a small &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N&lt;/code&gt; where a run takes minutes, find the shape of the answer, then do
one or two large runs to get the real number. People routinely burn a week
tuning at full size.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;where-it-earns-its-keep&quot;&gt;Where it earns its keep&lt;/h2&gt;

&lt;p&gt;Run HPL at close to full memory across every node, and you have built a
sustained, uniform, several-hour stress test that touches nearly all of DRAM,
nearly all of the floating point units, and the interconnect between every pair
of ranks. It is the best acceptance test most sites have, and it finds things:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Nodes that thermally throttle only once the whole rack is hot, which no
single-node test will reproduce.&lt;/li&gt;
  &lt;li&gt;Correctable memory errors that appear only when a large fraction of DRAM is
touched repeatedly.&lt;/li&gt;
  &lt;li&gt;One node whose BIOS profile differs from its siblings, which shows up as the
whole job running at the speed of that node.&lt;/li&gt;
  &lt;li&gt;Links that negotiated at a lower width or rate than expected.&lt;/li&gt;
  &lt;li&gt;Power delivery and cooling limits that only appear at full draw.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;HPL prints a residual check at the end. Treat a failed residual as a hardware
fault until proven otherwise, because that is usually what it is. A machine that
returns wrong arithmetic under load is a much more urgent finding than a
disappointing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rmax&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id=&quot;reading-the-result-honestly&quot;&gt;Reading the result honestly&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rmax&lt;/code&gt; is the best rate achieved. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Rpeak&lt;/code&gt; is the theoretical maximum from clock
rate, core count and per-cycle floating point width. Efficiency is the ratio,
and it is the number worth tracking, because it is comparable across time on the
same machine.&lt;/p&gt;

&lt;p&gt;A drop in efficiency between two runs of the same configuration is a real
signal. An absolute efficiency number compared against a different site’s
machine is mostly a comparison of two vendors’ BLAS libraries and two people’s
patience with parameter sweeps.&lt;/p&gt;

&lt;p&gt;Use HPL to prove the machine is healthy and to catch regressions after firmware
and driver changes. Use HPCG, and your own applications, to decide what the
machine is worth.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>DNS and mail authentication: a TXT answer is not a valid policy</title>
    <link href="https://www.wirewalk.com/writing/dns-mail-authentication-audit/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/dns-mail-authentication-audit/</id>
    <summary>Check record content, sender alignment, and real message headers before declaring mail authentication healthy.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A DNS query returning text is not proof that the requested security record exists. Wildcards, stale records, and malformed content can make a superficial check look successful. Mail authentication needs both DNS inspection and an authorized message test.&lt;/p&gt;

&lt;h2 id=&quot;inspect-the-exact-names&quot;&gt;Inspect the exact names&lt;/h2&gt;

&lt;p&gt;For a domain you administer, replace &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;example.com&lt;/code&gt; and the known DKIM selector:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;dig +short TXT example.com
dig +short TXT _dmarc.example.com
dig +short TXT selector1._domainkey.example.com
dig +short TXT deliberately-unused-label.example.com
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The last query helps identify a wildcard response. An answer identical to the apex TXT content at an unrelated name deserves investigation. Confirm against authoritative nameservers and parse the actual policy rather than accepting any nonempty answer.&lt;/p&gt;

&lt;h2 id=&quot;separate-the-mechanisms&quot;&gt;Separate the mechanisms&lt;/h2&gt;

&lt;p&gt;SPF authorizes sending infrastructure for the envelope identity. DKIM associates a cryptographic signature with a signing domain. DMARC evaluates alignment with the visible From domain and publishes handling/reporting policy. The primary specifications are &lt;a href=&quot;https://www.rfc-editor.org/rfc/rfc7208&quot;&gt;SPF, RFC 7208&lt;/a&gt;, &lt;a href=&quot;https://www.rfc-editor.org/rfc/rfc6376&quot;&gt;DKIM, RFC 6376&lt;/a&gt;, and &lt;a href=&quot;https://www.rfc-editor.org/rfc/rfc7489&quot;&gt;DMARC, RFC 7489&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Consult current provider guidance and subsequent standards updates when implementing. This article’s inspection method does not depend on copying a universal TXT record.&lt;/p&gt;

&lt;h2 id=&quot;inventory-legitimate-senders&quot;&gt;Inventory legitimate senders&lt;/h2&gt;

&lt;p&gt;List the primary mail provider, billing system, CRM, website forms, and any other service sending as the domain. Record the actual envelope and signing domains from message headers. A vendor’s generic setup page cannot establish which identity your account is currently using.&lt;/p&gt;

&lt;p&gt;Remove stale authorizations through a reviewed change after confirming they are no longer needed. Do not tighten policy blindly: an overlooked legitimate sender can be rejected along with unwanted mail.&lt;/p&gt;

&lt;h2 id=&quot;test-with-real-headers&quot;&gt;Test with real headers&lt;/h2&gt;

&lt;p&gt;Send an approved test message from each legitimate route to a controlled mailbox. Inspect Authentication-Results, the visible From address, envelope identity, and DKIM signing domain. Confirm alignment and final delivery.&lt;/p&gt;

&lt;p&gt;A website form displaying a thank-you message does not prove it sent mail. Verify that the receiving mailbox actually contains the message and that the sender identity is what the design intended.&lt;/p&gt;

&lt;h2 id=&quot;roll-policy-out-with-evidence&quot;&gt;Roll policy out with evidence&lt;/h2&gt;

&lt;p&gt;Begin with an inventory and reporting workflow, review legitimate failures, and adopt stricter handling according to the organization’s risk decision. Retain an explicit rollback and propagation plan. DNS caching means different receivers can observe different states during a change.&lt;/p&gt;

&lt;h2 id=&quot;keep-checks-semantic&quot;&gt;Keep checks semantic&lt;/h2&gt;

&lt;p&gt;The automated check should validate policy syntax, uniqueness where required, record type, and the relationship between DNS and observed message authentication. It should report lookup errors as incomplete evidence, not absence or success.&lt;/p&gt;

&lt;p&gt;Repeat after adding a SaaS sender, changing mail providers, or replacing the website’s form handler. Domain authentication is a maintained dependency of business communication, not a one-time DNS task.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>OSU Micro-Benchmarks: measuring the fabric before you blame the application</title>
    <link href="https://www.wirewalk.com/writing/osu-micro-benchmarks/"/>
    <published>2026-08-12T00:00:00-04:00</published>
    <updated>2026-08-12T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/osu-micro-benchmarks/</id>
    <summary>Point-to-point latency and bandwidth are the first numbers to establish on a new cluster, because every distributed result you collect afterwards is built on top of them.</summary>
    <content type="html">&lt;p&gt;When a distributed job runs slower than expected, the fabric is the first thing
accused and the last thing measured. The OSU Micro-Benchmarks exist to close
that gap. They are small MPI programs from the MVAPICH group at Ohio State that
measure one communication pattern at a time, with nothing else running.&lt;/p&gt;

&lt;p&gt;The value is not the numbers themselves. It is that you take them once when the
machine is known good, write them down, and compare against them every time
something feels slow afterwards.&lt;/p&gt;

&lt;h2 id=&quot;the-two-that-matter-most&quot;&gt;The two that matter most&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_latency&lt;/code&gt; sends a message between two ranks and reports half the round trip
time. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; streams a window of messages in one direction and reports
sustained bandwidth. Everything else in the suite is a variation on those two
ideas.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# two ranks, one per node, pinned&lt;/span&gt;
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./osu_latency
mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node &lt;span class=&quot;nt&quot;&gt;--bind-to&lt;/span&gt; core &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./osu_bw
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Both sweep message sizes and print a table. Read them as two different regimes.
Small messages measure the fixed cost of getting a message onto the wire and
back off it: software stack, doorbell, interrupt or polling behavior, and the
switch hop count. Large messages measure how much of the link you can actually
fill. The interesting part is the middle, where the transport switches protocol.&lt;/p&gt;

&lt;h2 id=&quot;the-eager-to-rendezvous-transition&quot;&gt;The eager to rendezvous transition&lt;/h2&gt;

&lt;p&gt;Below a threshold, MPI implementations send small messages eagerly: the sender
pushes the payload without waiting to hear that the receiver has somewhere to
put it. Above the threshold they switch to rendezvous, where the sender first
exchanges control messages and then the data moves by RDMA directly into the
destination buffer.&lt;/p&gt;

&lt;p&gt;You can see the switch in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_bw&lt;/code&gt; table as a visible discontinuity in the
bandwidth curve. Knowing where it sits on your system is genuinely useful,
because an application whose typical message size lands just above the threshold
will behave very differently from one that lands just below. Both MPICH-derived
and Open MPI stacks let you move that threshold; whether you should is a
question to answer with measurements from your own application, not from a
micro-benchmark.&lt;/p&gt;

&lt;h2 id=&quot;pinning-changes-the-answer&quot;&gt;Pinning changes the answer&lt;/h2&gt;

&lt;p&gt;An unpinned two-rank latency test measures the scheduler as much as the network.
If the process migrates to a core on the socket that does not own the HCA, every
message crosses the inter-socket link on the way to the wire.&lt;/p&gt;

&lt;p&gt;Always pin, always place one rank per node for the point-to-point tests, and
check which NUMA domain your adapter is attached to:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;cat&lt;/span&gt; /sys/class/infiniband/mlx5_0/device/numa_node
lstopo &lt;span class=&quot;nt&quot;&gt;--output-format&lt;/span&gt; txt | &lt;span class=&quot;nb&quot;&gt;head&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-40&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If that returns a socket your ranks are not running on, fix the placement before
recording anything. This single mistake accounts for a large share of “the
network is slow” reports that turn out not to involve the network.&lt;/p&gt;

&lt;h2 id=&quot;collectives-and-why-they-are-a-different-question&quot;&gt;Collectives, and why they are a different question&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_allreduce&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_alltoall&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;osu_barrier&lt;/code&gt; scale to the full job size and
are much closer to what real applications do. They are also much more sensitive
to noise. One slow node, one core running a stray daemon, one rank waiting on a
filesystem, and the whole collective inherits the delay, because a collective
finishes when its slowest participant finishes.&lt;/p&gt;

&lt;p&gt;That sensitivity is the point. A collective benchmark that degrades while
point-to-point stays clean is telling you the problem is not the fabric. It is
jitter, and jitter comes from the operating system, the scheduler, thermal
behavior or a filesystem mount, not from the switch.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Run the point-to-point tests between several different node pairs, not just the
first two the scheduler hands you. A single bad cable or a port that negotiated
a lower width will only show up when that specific link is on the path, and it
will look like an application problem for weeks.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;gpu-buffers&quot;&gt;GPU buffers&lt;/h2&gt;

&lt;p&gt;Recent versions build CUDA-aware variants that take &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-d cuda&lt;/code&gt;, which places the
send and receive buffers in device memory instead of host memory:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 2 &lt;span class=&quot;nt&quot;&gt;-map-by&lt;/span&gt; ppr:1:node ./osu_bw &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; cuda D D
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This is worth measuring separately, because the path is different. With
GPUDirect RDMA working, the adapter reads and writes device memory directly.
Without it, every message takes a detour through a host bounce buffer, and you
will see it immediately in the bandwidth number. Comparing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-d cuda&lt;/code&gt; against the
host-memory run is the fastest way to confirm GPUDirect is actually engaged
rather than merely installed.&lt;/p&gt;

&lt;h2 id=&quot;what-to-record&quot;&gt;What to record&lt;/h2&gt;

&lt;p&gt;Keep a file in version control with the date, the firmware and driver versions,
the MPI build, the node pair, and the resulting curves. It takes ten minutes and
it converts every future performance argument from opinion into a comparison.&lt;/p&gt;

&lt;p&gt;The next articles in this series work up the stack from here: dense compute with
HPL, memory and network behavior with HPCG, mixed precision with HPL-MxP, and
collectives on GPU clusters with nccl-tests. All of them are easier to interpret
once you know what a single link on your machine can do.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Nine ways an environment module makes working software look broken</title>
    <link href="https://www.wirewalk.com/writing/environment-modules-failure-modes/"/>
    <published>2026-07-29T10:00:00-04:00</published>
    <updated>2026-07-29T10:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/environment-modules-failure-modes/</id>
    <summary>The environment module system fails quietly by design, and almost every one of its failures presents as missing or broken software. This walks through nine mechanisms, the command that distinguishes each from the others, and a verification harness that loads every module in a clean shell and runs it.</summary>
    <content type="html">&lt;p&gt;A user reports a tool is broken. You load the module, run the tool, and it works. The user runs
the same three lines and it does not. Nothing in either transcript looks wrong.&lt;/p&gt;

&lt;p&gt;That shape of report is almost always the module system, not the software, and it costs so much
time because environment modules fail silently by construction. A modulefile sets variables.
Setting one to a directory that does not exist is not an error. Forgetting the matching &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt;
entry is not an error. Loading a module whose files were deleted six months ago is not an error.
The module system reports success in all three cases, and the failure surfaces elsewhere — as
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;command not found&lt;/code&gt;, a missing Python package, or a linker error naming an unfamiliar symbol
version.&lt;/p&gt;

&lt;p&gt;What follows is nine mechanisms, each with the command that tells it apart, then a harness that
catches most of them before a user does.&lt;/p&gt;

&lt;h2 id=&quot;what-module-actually-is&quot;&gt;What &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module&lt;/code&gt; actually is&lt;/h2&gt;

&lt;p&gt;Everything here follows from one fact. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module&lt;/code&gt; is not a program. It is a shell function that runs
a program, captures the shell code it prints, and evaluates that in your current shell. Lmod
defines it around &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$LMOD_CMD&lt;/code&gt;, Tcl Environment Modules around &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;modulecmd&lt;/code&gt;. Watch it work:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ $LMOD_CMD bash load gcc/12.3.0 2&amp;gt;/dev/null
PATH=/opt/apps/gcc/12.3.0/bin:/usr/local/bin:/usr/bin:/bin; export PATH;
LD_LIBRARY_PATH=/opt/apps/gcc/12.3.0/lib64; export LD_LIBRARY_PATH;
LOADEDMODULES=gcc/12.3.0; export LOADEDMODULES;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That is the shape of it, abridged: a real Lmod run also emits &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;_LMFILES_&lt;/code&gt;, reference counts and
its serialized module table, which is noise for this purpose but worth seeing once.&lt;/p&gt;

&lt;p&gt;Two consequences follow. First, the effect of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module load&lt;/code&gt; is confined to the shell that ran the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eval&lt;/code&gt;; put it anywhere the shell forks and it lands in a child that exits immediately. Second,
the shell code has to have stdout to itself, so diagnostics go somewhere else — and where they go
depends on which implementation you have. Lmod writes the output of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module avail&lt;/code&gt;
and friends to stderr by default; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LMOD_REDIRECT=yes&lt;/code&gt; (or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--redirect&lt;/code&gt;) moves it to stdout, and the
Lmod documentation notes that only works for bash and zsh, not csh or tcsh. Tcl Environment Modules
went the other way: current releases document a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;redirect_output&lt;/code&gt; configuration option that
defaults to on, so on sh, bash, ksh, zsh and fish the messages arrive on stdout unless someone
passes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--no-redirect&lt;/code&gt;. Check yours with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module config&lt;/code&gt; rather than assuming either default.&lt;/p&gt;

&lt;p&gt;The practical consequence is that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list | grep&lt;/code&gt; works on one of your clusters and silently
returns nothing on the other, and the reader concludes the module is not loaded. Redirect
explicitly and the question does not arise:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module list | grep -c gcc      # 0 under Lmod&apos;s default, 1 under Modules&apos; default
$ module list 2&amp;gt;&amp;amp;1 | grep -c gcc # 1 under both
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;For anything you intend to parse rather than read, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module -t list&lt;/code&gt; (terse) is a better input than
the decorated human output.&lt;/p&gt;

&lt;h2 id=&quot;1-a-load-inside-a-pipe-or-a-command-substitution-goes-nowhere&quot;&gt;1. A load inside a pipe or a command substitution goes nowhere&lt;/h2&gt;

&lt;p&gt;This is the one that costs hours, because every line of output looks correct.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ cat wanted.txt | while read -r m; do module load &quot;$m&quot;; done
$ module list 2&amp;gt;&amp;amp;1
No modules loaded
$ samtools --version
bash: samtools: command not found
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;No error was printed and the loads succeeded. In bash each stage of a pipeline runs in a subshell,
so the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;while&lt;/code&gt; loop — and the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eval&lt;/code&gt; inside &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module&lt;/code&gt; — executed in a forked child. The child
exported the variables and exited. The parent shell was never touched. (Bash has one escape hatch
here, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;shopt -s lastpipe&lt;/code&gt;, which runs the final stage in the current shell — but only when job
control is off, so it does nothing in an interactive session, which is exactly where people test
it.)&lt;/p&gt;

&lt;p&gt;The command substitution version is worse, because the inner answer is right:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ VER=$(module load python/3.11.6 &amp;amp;&amp;amp; python3 --version)
$ echo &quot;$VER&quot;
Python 3.11.6
$ python3 --version
Python 3.9.18
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The load did work — inside the subshell, where &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python3&lt;/code&gt; really was 3.11.6, which is why the
captured string looks like proof that the module is loaded. The parent shell has nothing loaded,
and the next line in the script is the one that fails.&lt;/p&gt;

&lt;p&gt;Two more places the fork is invisible: every recipe line in a makefile runs in its own shell, so
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module load&lt;/code&gt; on one line and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$(CC)&lt;/code&gt; on the next are unrelated environments, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;xargs&lt;/code&gt; execs a
fresh process per batch. Over ssh without a login shell the function is often not defined at all.&lt;/p&gt;

&lt;p&gt;The fixes are mechanical — a process-substitution loop keeps the body in the current shell, make
wants &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;amp;&amp;amp;&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.ONESHELL:&lt;/code&gt;, and ssh wants the init script sourced explicitly.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ while read -r m; do module load &quot;$m&quot;; done &amp;lt; &amp;lt;(grep -v &apos;^#&apos; wanted.txt)
$ ssh node &apos;source /usr/share/lmod/lmod/init/bash &amp;amp;&amp;amp; module load gcc/12.3.0 &amp;amp;&amp;amp; gcc --version&apos;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The diagnostic that settles it in one line: after the load, print &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LOADEDMODULES&lt;/code&gt; in the shell you
care about. If it is empty there, the load happened somewhere else.&lt;/p&gt;

&lt;h2 id=&quot;2-which-says-one-thing-and-a-different-binary-runs&quot;&gt;2. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;which&lt;/code&gt; says one thing and a different binary runs&lt;/h2&gt;

&lt;p&gt;This is the section where the obvious reading is wrong twice — once about the symptom, and once
about the fix nearly everyone prescribes for it.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load python/3.11.6
$ which python3
/opt/apps/python/3.11.6/bin/python3
$ python3 --version
Python 3.9.18
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The standard answer is that bash has cached the old location and you need &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;hash -r&lt;/code&gt;. Test that
before believing it, because in bash it is not what happened:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ mkdir -p /tmp/ht/a /tmp/ht/b
$ printf &apos;#!/bin/sh\necho A\n&apos; &amp;gt; /tmp/ht/a/foo
$ printf &apos;#!/bin/sh\necho B\n&apos; &amp;gt; /tmp/ht/b/foo
$ chmod +x /tmp/ht/a/foo /tmp/ht/b/foo
$ PATH=/tmp/ht/a:$PATH
$ foo                                        # runs, and is now hashed
A
$ eval &apos;PATH=/tmp/ht/b:$PATH; export PATH;&apos;  # exactly what module load evaluates
$ foo
B
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Assigning to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; makes bash discard every remembered location — plainly, with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;export&lt;/code&gt;, from
inside a function, or inside the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;eval&lt;/code&gt; that the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module&lt;/code&gt; function runs. All four behave the same
way on the bash used for the transcripts here (3.2.57); the test above takes two minutes on
whatever bash your cluster actually runs, and is worth doing before you accept either account. zsh
behaves the same, and tcsh rebuilds its hash when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;path&lt;/code&gt; is set. A module load &lt;em&gt;is&lt;/em&gt; a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt;
assignment, so it cannot leave a stale entry behind, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;hash -r&lt;/code&gt; after &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module load&lt;/code&gt; is a ritual
aimed at a failure bash does not have. Prescribing it costs you the one thing you were short of: a
reason to look somewhere else.&lt;/p&gt;

&lt;p&gt;Two things &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;which&lt;/code&gt; genuinely cannot see, and one that it can.&lt;/p&gt;

&lt;p&gt;A function or an alias beats &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; outright, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;which&lt;/code&gt; is an external program that knows only
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt;. Sites wrap interpreters this way, users copy the wrapper into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.bashrc&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module&lt;/code&gt;
itself is a shell function, so nothing here is exotic:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ type python3
python3 is a function
python3 () 
{ 
    /usr/bin/python3 &quot;$@&quot;
}
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That is what the transcript at the top of this section really was: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; changed exactly as asked,
and a function defined in a site profile went on answering to the name.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;command python3 --version&lt;/code&gt; bypasses it for one call, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;unset -f python3&lt;/code&gt; for the session. Note that
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;command -v python3&lt;/code&gt; prints just &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python3&lt;/code&gt; for a function — correct, and useless as evidence.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type&lt;/code&gt; is the builtin to reach for. Aliases produce the same symptom interactively and then
disappear in batch, because bash does not expand aliases in non-interactive shells unless
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;expand_aliases&lt;/code&gt; is set: that is the “works in my shell, fails in my job script” report.&lt;/p&gt;

&lt;p&gt;The hash cache is real, but its trigger is the opposite of the folklore — it bites when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; does
&lt;em&gt;not&lt;/em&gt; change and a same-named binary appears earlier on it. A &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pip install --user&lt;/code&gt; into
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/.local/bin&lt;/code&gt;, a wrapper dropped into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/usr/local/bin&lt;/code&gt;, an install finishing while the user’s shell
is open:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ PATH=/tmp/ht/a:/tmp/ht/b:$PATH        # a is earlier and has no bar yet
$ cp /tmp/ht/b/foo /tmp/ht/b/bar; bar
B
$ cp /tmp/ht/a/foo /tmp/ht/a/bar        # an earlier bar now exists; PATH unchanged
$ which bar
/tmp/ht/a/bar
$ type bar
bar is hashed (/tmp/ht/b/bar)
$ bar
B
$ hash -r; bar
A
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;One detail worth knowing before you rely on the diagnostic: on the bash tested here, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type -a bar&lt;/code&gt;
re-searched &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; and listed both directories without mentioning the cache at all, while plain
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type bar&lt;/code&gt; named it. Use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type&lt;/code&gt; without &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-a&lt;/code&gt;, or the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;hash&lt;/code&gt; builtin, which prints the table with hit
counts. In tcsh the equivalent situation needs &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rehash&lt;/code&gt;, and there &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;which&lt;/code&gt; is a builtin that
consults the same cache, so it agrees with the wrong answer rather than contradicting it.&lt;/p&gt;

&lt;h2 id=&quot;3-the-module-that-sets-a-variable-but-no-interpreter&quot;&gt;3. The module that sets a variable but no interpreter&lt;/h2&gt;

&lt;p&gt;A common pattern for Python tooling is a modulefile pointing at a virtual environment:&lt;/p&gt;

&lt;div class=&quot;language-tcl highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;setenv    TOOLKIT_VENV  /shared/apps/toolkit/1.4/venv
setenv    PYTHONPATH    /shared/apps/toolkit/1.4/venv/lib/python3.11/site-packages
prepend-path PATH       /shared/apps/toolkit/1.4/bin
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The user does the obvious check and gets a genuinely misleading answer:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load toolkit/1.4
$ python3 -c &quot;import toolkit&quot;
Traceback (most recent call last):
  File &quot;&amp;lt;string&amp;gt;&quot;, line 1, in &amp;lt;module&amp;gt;
ModuleNotFoundError: No module named &apos;toolkit&apos;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The package is installed, under an interpreter that is not on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt;. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python3&lt;/code&gt; here is still the
system 3.9, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PYTHONPATH&lt;/code&gt; has pointed it at a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;site-packages&lt;/code&gt; built for 3.11 — so on top of the
version mismatch, any compiled extension there is named &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;_core.cpython-311-x86_64-linux-gnu.so&lt;/code&gt;
and 3.9 will not look at it.&lt;/p&gt;

&lt;p&gt;Ask the interpreter where it is, not whether the import worked:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ python3 -c &quot;import sys; print(sys.executable); print(sys.version_info[:2])&quot;
/usr/bin/python3
(3, 9)
$ &quot;$TOOLKIT_VENV/bin/python&quot; -c &quot;import toolkit; print(toolkit.__version__)&quot;
1.4.0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That is the diagnosis. The fix is to put the venv’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;bin&lt;/code&gt; on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; and stop setting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PYTHONPATH&lt;/code&gt;
— a venv does not need it, and setting it turns a clean “wrong interpreter” failure into a
confusing “half the packages import” one.&lt;/p&gt;

&lt;div class=&quot;language-tcl highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;prepend-path PATH  /shared/apps/toolkit/1.4/venv/bin
unsetenv     PYTHONPATH
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PYTHONPATH&lt;/code&gt; deserves a blanket rule: a shared modulefile should never set it. Its entries are
searched ahead of the interpreter’s own standard library and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;site-packages&lt;/code&gt; — everything except
the script’s own directory — and they leak into every unrelated Python process started from that
session, including ones belonging to other modules.&lt;/p&gt;

&lt;h2 id=&quot;4-the-module-that-does-put-an-interpreter-on-the-path-and-shadows-everything&quot;&gt;4. The module that does put an interpreter on the path, and shadows everything&lt;/h2&gt;

&lt;p&gt;The opposite failure is quieter and lasts longer. A module that prepends a Python to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; has
changed the meaning of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python3&lt;/code&gt; for everything else in that shell. Usually that is wanted. The
parts that are not:&lt;/p&gt;

&lt;p&gt;a script with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;#!/usr/bin/env python3&lt;/code&gt; now runs under the module’s interpreter regardless of what
it was tested against; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pip install --user&lt;/code&gt; writes into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/.local/lib/python3.11/&lt;/code&gt;, which then
becomes visible to every &lt;em&gt;other&lt;/em&gt; 3.11 on the system; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python3 -m venv&lt;/code&gt; creates an environment
whose &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pyvenv.cfg&lt;/code&gt; records the module’s interpreter, and that venv keeps working right up until
the module directory is deleted.&lt;/p&gt;

&lt;p&gt;Load order decides the winner, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list&lt;/code&gt; will not tell you who won. Ask for the
resolution:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load toolA toolB
$ type -a python3 | head -3
python3 is /opt/apps/toolB/2.1/bin/python3
python3 is /opt/apps/toolA/1.9/bin/python3
python3 is /usr/bin/python3
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The structural fix is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;family&lt;/code&gt;. Lmod’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;family(&quot;python&quot;)&lt;/code&gt; makes modules in the same family
mutually exclusive, so the second load swaps the first rather than shadowing it:&lt;/p&gt;

&lt;div class=&quot;language-lua highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;family&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;python&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;prepend_path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;PATH&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;/opt/apps/python/3.11.6/bin&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load python/3.11.6 python/3.12.2
Lmod is automatically replacing &quot;python/3.11.6&quot; with &quot;python/3.12.2&quot;.
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That message is the point. The load still succeeds, but the user now knows.&lt;/p&gt;

&lt;p&gt;For application modules the better pattern is not to expose an interpreter at all: ship wrappers
in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;bin/&lt;/code&gt; with an absolute shebang pointing at the venv’s Python, and put only that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;bin&lt;/code&gt; on
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt;. The tool works; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python3&lt;/code&gt; keeps meaning what it meant.&lt;/p&gt;

&lt;h2 id=&quot;5-several-modules-in-one-command&quot;&gt;5. Several modules in one command&lt;/h2&gt;

&lt;p&gt;A recurring ticket: loading modules one per line works, loading them in a single command fails,
and the user is told to stop doing the thing that failed without anyone explaining why.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module purge
$ module load compiler/gcc-12.3.0 mpi/openmpi-5.0.3 app/solver-4.2
$ solver --version
solver: error while loading shared libraries: libmpi.so.40: cannot open shared object file: No such file or directory
$ module purge
$ module load compiler/gcc-12.3.0
$ module load mpi/openmpi-5.0.3
$ module load app/solver-4.2
$ solver --version
solver 4.2.1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two mechanisms are in play, and neither is as simple as it is usually told — including in the
direction most people assume.&lt;/p&gt;

&lt;p&gt;The first is that what a child process launched from a modulefile sees is an implementation
detail, not the environment the user will end up with, and the two implementations differ. Tcl
Environment Modules documents that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;setenv&lt;/code&gt; “will also change the process’ environment” and that it
“is useful for changing the environment prior to the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;exec&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;system&lt;/code&gt; command” — so there, a
value set by an earlier sibling in the same command &lt;em&gt;is&lt;/em&gt; visible to an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;exec&lt;/code&gt; in a later one. Lmod
assembles the new environment internally and emits it at the end, and its two ways of running a
command differ on exactly this point: the documentation for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;capture()&lt;/code&gt; says it uses the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LD_PRELOAD&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LD_LIBRARY_PATH&lt;/code&gt; values that were current when Lmod was configured, and directs
you to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;subprocess()&lt;/code&gt; if you want the values in force now. So a modulefile doing this:&lt;/p&gt;

&lt;div class=&quot;language-tcl highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;set mpiroot $env&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;MPI_ROOT&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
set mpiver  &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;exec $mpiroot/bin/mpirun --version&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;can behave one way on the cluster running Modules, another on the cluster running Lmod, and a
third way when the module that sets &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MPI_ROOT&lt;/code&gt; is loaded in a separate command rather than as a
sibling — reading the modulefile tells you nothing about which. The same applies to any
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[file exists $env(SOMETHING)/...]&lt;/code&gt; test fed by a sibling.&lt;/p&gt;

&lt;p&gt;The second: whether prerequisite checks and hierarchical &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MODULEPATH&lt;/code&gt; additions are
honored within a single command varies across Tcl Environment Modules 3.2, 4.x and 5.x and across
Lmod 6 through 9. I am not going to give a version table I cannot stand behind. You can settle it
for your own stack in about thirty seconds:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module purge
$ module load compiler/gcc-12.3.0 mpi/openmpi-5.0.3; echo &quot;exit=$?&quot;
$ module list 2&amp;gt;&amp;amp;1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list&lt;/code&gt; shows fewer modules than you asked for while &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;exit=0&lt;/code&gt;, your implementation loads
partially and does not report it. Know that before building automation on top of it.&lt;/p&gt;

&lt;p&gt;Two rules fall out regardless of version. In scripts, load one module per command and check the
result by inspecting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LOADEDMODULES&lt;/code&gt; rather than trusting the exit status. And never write a
modulefile that runs an external command at load time; compute the value at install time and
hard-code it.&lt;/p&gt;

&lt;h2 id=&quot;6-class-file-version-610-and-what-the-numbers-mean&quot;&gt;6. Class file version 61.0, and what the numbers mean&lt;/h2&gt;

&lt;p&gt;A Java tool that was built elsewhere produces this:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load gatk/4.5.0.0
$ gatk --version
Error: LinkageError occurred while loading main class org.broadinstitute.hellbender.Main
	java.lang.UnsupportedClassVersionError: org/broadinstitute/hellbender/Main has been
	compiled by a more recent version of the Java Runtime (class file version 61.0),
	this version of the Java Runtime only recognizes class file versions up to 52.0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The message names two numbers and explains neither. They are class file format major versions, not
Java versions, and for Java 5 onward &lt;strong&gt;major version = Java release + 44&lt;/strong&gt;. So 61 is Java 17 and
52 is Java 8: the tool needs 17 and the node’s default is 8.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Class file major&lt;/th&gt;
      &lt;th&gt;Java release&lt;/th&gt;
      &lt;th&gt;Class file major&lt;/th&gt;
      &lt;th&gt;Java release&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;49&lt;/td&gt;
      &lt;td&gt;5&lt;/td&gt;
      &lt;td&gt;60&lt;/td&gt;
      &lt;td&gt;16&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;50&lt;/td&gt;
      &lt;td&gt;6&lt;/td&gt;
      &lt;td&gt;61&lt;/td&gt;
      &lt;td&gt;17&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;51&lt;/td&gt;
      &lt;td&gt;7&lt;/td&gt;
      &lt;td&gt;62&lt;/td&gt;
      &lt;td&gt;18&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;52&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;63&lt;/td&gt;
      &lt;td&gt;19&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;53&lt;/td&gt;
      &lt;td&gt;9&lt;/td&gt;
      &lt;td&gt;64&lt;/td&gt;
      &lt;td&gt;20&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;54&lt;/td&gt;
      &lt;td&gt;10&lt;/td&gt;
      &lt;td&gt;65&lt;/td&gt;
      &lt;td&gt;21&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;55&lt;/td&gt;
      &lt;td&gt;11&lt;/td&gt;
      &lt;td&gt;66&lt;/td&gt;
      &lt;td&gt;22&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;56&lt;/td&gt;
      &lt;td&gt;12&lt;/td&gt;
      &lt;td&gt;67&lt;/td&gt;
      &lt;td&gt;23&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;57&lt;/td&gt;
      &lt;td&gt;13&lt;/td&gt;
      &lt;td&gt;68&lt;/td&gt;
      &lt;td&gt;24&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;58&lt;/td&gt;
      &lt;td&gt;14&lt;/td&gt;
      &lt;td&gt;69&lt;/td&gt;
      &lt;td&gt;25&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;59&lt;/td&gt;
      &lt;td&gt;15&lt;/td&gt;
      &lt;td&gt; &lt;/td&gt;
      &lt;td&gt; &lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Below that, 45 is Java 1.1 through 48 for 1.4. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.0&lt;/code&gt; after the major is the minor version; it
is zero except for preview features, where it is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;65535&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;You do not need the table if you read the bytes. A class file is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CAFEBABE&lt;/code&gt;, a two-byte minor at
offset 4, then a two-byte major at offset 6, big-endian:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ unzip -p gatk.jar org/broadinstitute/hellbender/Main.class \
    | od -An -tu2 --endian=big -j6 -N2
  61
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;(&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--endian&lt;/code&gt; is a GNU coreutils option; on a box without it, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;od -An -tx1 -j6 -N2&lt;/code&gt; and read the two
bytes by hand.)&lt;/p&gt;

&lt;p&gt;Or, if a JDK is to hand:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ javap -verbose -cp gatk.jar org.broadinstitute.hellbender.Main | grep -E &apos;major|minor&apos;
  minor version: 0
  major version: 61
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The manifest often records the build JDK (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Build-Jdk-Spec: 17&lt;/code&gt;), which is a useful cross-check but
is not authoritative — it reflects the compiler, not the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--release&lt;/code&gt; target.&lt;/p&gt;

&lt;p&gt;One wrinkle before you conclude anything from a single class: a multi-release jar carries extra
copies under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;META-INF/versions/N/&lt;/code&gt; that are deliberately compiled for higher majors than the
jar’s baseline. If &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Multi-Release: true&lt;/code&gt; is in the manifest, check a class from the root of the
jar instead.&lt;/p&gt;

&lt;p&gt;The fix is a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prereq&lt;/code&gt; in the modulefile, so the tool cannot load without its runtime — not a wiki
line telling people to load Java first:&lt;/p&gt;

&lt;div class=&quot;language-tcl highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;prereq java/17
setenv JAVA_HOME /opt/apps/java/17.0.9
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;7-the-modulefile-that-loads-cleanly-and-sets-a-variable-to-nothing&quot;&gt;7. The modulefile that loads cleanly and sets a variable to nothing&lt;/h2&gt;

&lt;p&gt;Modulefiles get copied between clusters. The copy succeeds, the load succeeds, the software is not
there.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load fsl/6.0.7
$ echo &quot;FSLDIR=$FSLDIR&quot;
FSLDIR=/opt/apps/fsl/6.0.7
$ ls &quot;$FSLDIR&quot;
ls: cannot access &apos;/opt/apps/fsl/6.0.7&apos;: No such file or directory
$ bet --help
bash: bet: command not found
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;In their default configurations neither Environment Modules nor Lmod checks that a path exists
before putting it in a variable, and neither treats a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prepend-path&lt;/code&gt; onto a missing directory as an
error. There is a defensible reason — automounted directories may not be materialized at load
time — but the result is a module that reports success and delivers nothing.&lt;/p&gt;

&lt;p&gt;Harder to spot is the variant where the variable ends up empty rather than wrong. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$FSLDIR/bin/bet&lt;/code&gt;
then expands to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/bin/bet&lt;/code&gt;, and where that path happens to exist you get a different program and
no error at all.&lt;/p&gt;

&lt;p&gt;Make the modulefile assert its preconditions. In Tcl, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;break&lt;/code&gt; ends the evaluation there; Modules
documents that the module is then not listed as loaded and other modules being loaded concurrently
are unaffected, with everything up to that point still performed:&lt;/p&gt;

&lt;div class=&quot;language-tcl highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;set root /opt/apps/fsl/6.0.7

if &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;module-info mode load&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &amp;amp;&amp;amp; !&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;file isdirectory $root&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
    puts stderr &lt;span class=&quot;s2&quot;&gt;&quot;ERROR: fsl/6.0.7 is not installed on &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;info hostname&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt; (&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$root&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt; missing)&quot;&lt;/span&gt;
    break
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;

setenv       FSLDIR $root
prepend-path PATH   $root/bin
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module-info mode load&lt;/code&gt; guard matters: without it the check also runs during &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module display&lt;/code&gt;,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module avail&lt;/code&gt; and unload, which is noisy and can break &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module purge&lt;/code&gt; on a node where the
directory genuinely is absent. Note also that Modules documents &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;break&lt;/code&gt; on unload as advisory —
a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module unload --force&lt;/code&gt; proceeds anyway — so do not use it to make a module unremovable.&lt;/p&gt;

&lt;p&gt;In Lmod the equivalent is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;if (mode() == &quot;load&quot; and not isDir(root)) then LmodError(...) end&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id=&quot;8-the-deprecated-module-that-loads-instead-of-failing&quot;&gt;8. The deprecated module that loads instead of failing&lt;/h2&gt;

&lt;p&gt;A version is retired and its install directory removed, but the modulefile stays because deleting
it would “break people’s scripts”. It breaks them anyway, later and less legibly: the load
succeeds, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; gains a directory that is not there, and the user gets &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;command not found&lt;/code&gt; from a
module they can see in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Lmod’s admin file is right for a module that still works but should not be used. It is a list of
keys — a module name, or a full path to a modulefile when the key begins with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/&lt;/code&gt; — each followed
by a message that runs to the next blank line:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ cat /opt/apps/lmod/etc/admin.list
samtools/1%.9:
    Retired 2026-09-30. Use samtools/1.19.

&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Its location is wherever &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LMOD_ADMIN_FILE&lt;/code&gt; was configured to point, which &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module --config&lt;/code&gt; will
tell you.&lt;/p&gt;

&lt;p&gt;Those keys are Lua patterns rather than literal strings, which is the detail that catches people.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.&lt;/code&gt; matches any character and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-&lt;/code&gt; is a quantifier, so a version number wants &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%.&lt;/code&gt; and a module name
containing a dash wants &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%-&lt;/code&gt;. Written as plain &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;samtools/1.9&lt;/code&gt; the entry still matches
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;samtools/1.9&lt;/code&gt;, and also &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;samtools/119&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Loading a module that matches prints a block in this shape and then loads it:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load samtools/1.9

-----------------------------------------------------------------
There are messages associated with the following module(s) :
-----------------------------------------------------------------
samtools/1.9:
    Retired 2026-09-30. Use samtools/1.19.
-----------------------------------------------------------------
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Note what that does: it prints and loads. The Lmod documentation is explicit that the admin file
“in no way controls user access to the module”. That is correct for a deprecation window and wrong
for a module whose files are gone. For that, the modulefile must refuse:&lt;/p&gt;

&lt;div class=&quot;language-lua highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;LmodError&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;[[
samtools/1.9 was removed on 2026-09-30. The installation no longer exists.
Use: module load samtools/1.19
]]&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That refuses the load, and Lmod documents that a failed load returns a non-zero exit code you can
trap. Confirm that on your own version before a job script leans on it, and confirm it separately
under Tcl Environment Modules, where the exit status of a failed load has not always been reliable.&lt;/p&gt;

&lt;p&gt;State the principle plainly, because it is the one people argue about: a module that cannot do its
job must fail at load time, loudly, with a non-zero exit. A job that dies in the first second with
a clear message is a five-minute fix. A job that runs nine hours and fails at the output stage
because one of six tools was missing costs a day, and nobody will connect it to a module that
loaded successfully that morning.&lt;/p&gt;

&lt;h2 id=&quot;9-build-host-paths-baked-into-a-shared-installation&quot;&gt;9. Build-host paths baked into a shared installation&lt;/h2&gt;

&lt;p&gt;This one is not the module system’s fault, but the module system is where it surfaces: software
that works perfectly on the node it was compiled on and nowhere else.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ module load rstudio-deps/2026.05
$ R -e &apos;library(sf)&apos;
Error: package or namespace load failed for &apos;sf&apos;:
 unable to load shared object &apos;/shared/apps/R/4.4.1/library/sf/libs/sf.so&apos;:
  libproj.so.25: cannot open shared object file: No such file or directory
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It works on the build node because the library was in that node’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/usr/lib64&lt;/code&gt;, or because the
build shell had an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LD_LIBRARY_PATH&lt;/code&gt; no other shell has. Four things get baked in:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# 1. RPATH/RUNPATH pointing at a build-only directory
$ readelf -d /shared/apps/tool/bin/tool | grep RUNPATH
 0x000000000000001d (RUNPATH)   Library runpath: [/scratch/build/deps/lib]

# 2. libraries resolving outside the install tree, or not at all
$ ldd /shared/apps/tool/bin/tool | grep -v &apos;=&amp;gt; /shared&apos;
	libproj.so.25 =&amp;gt; not found

# 3. build paths in generated config, which propagate to everything users compile later
$ grep -n &apos;/scratch/build&apos; /shared/apps/R/4.4.1/lib64/R/etc/Makeconf
128:LDFLAGS = -L/scratch/build/deps/lib

# 4. pkg-config and libtool files carrying an absolute prefix
$ grep -rl &apos;/scratch/build&apos; /shared/apps/tool/lib/pkgconfig/*.pc /shared/apps/tool/lib/*.la
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A fifth catches conda-based installs and looks like a filesystem problem: the shebang length limit.
The kernel reads a fixed-size buffer from the start of an executable, and a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;#!&lt;/code&gt; line longer than
that buffer does not survive it. On a kernel with the older 128-byte buffer the line is truncated
and the truncated path is executed, so the error names an interpreter that does not exist while
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ls&lt;/code&gt; shows the real one plainly does:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ head -1 /gpfs/shared/apps/genomics/pipelines/annotation/1.2.0/bin/annotate
#!/gpfs/shared/apps/genomics/pipelines/annotation/1.2.0/conda/envs/annotation-tools-py311-bioconda-pinned-2026-05-rebuild/bin/python3.11
$ head -1 /gpfs/shared/apps/genomics/pipelines/annotation/1.2.0/bin/annotate | wc -c
137
$ ./annotate
bash: ./annotate: /gpfs/shared/apps/genomics/pipelines/annotation/1.2.0/conda/envs/annotation-tools-py311-bioconda-pinned-2026-05-rebuild/bin/py: bad interpreter: No such file or directory
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Read the truncation point rather than the message: the path in the error stops at 126 characters,
which is the 128-byte buffer less the two bytes taken by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;#!&lt;/code&gt;. That is the signature.&lt;/p&gt;

&lt;p&gt;The number is a kernel compile-time constant, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;BINPRM_BUF_SIZE&lt;/code&gt; in
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;include/uapi/linux/binfmts.h&lt;/code&gt;, and it is 256 in current mainline. Current kernels also handle the
overflow differently: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fs/binfmt_script.c&lt;/code&gt; returns &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-ENOEXEC&lt;/code&gt; when the interpreter path runs to the
end of the buffer with no terminator after it, rather than executing a path it had to cut. That
changes the symptom instead of removing it — a shell that gets &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ENOEXEC&lt;/code&gt; falls back to running the
file itself as a shell script, so a Python entry point fails with shell syntax errors from its own
source instead of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;bad interpreter&lt;/code&gt;. Check the constant on the kernel in front of you rather than
quoting either number. The fix is the same in both cases: a shorter install prefix, or a small
POSIX &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sh&lt;/code&gt; wrapper that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;exec&lt;/code&gt;s the real interpreter.&lt;/p&gt;

&lt;p&gt;The rule that prevents all five: never validate a build on the node that produced it. Build
anywhere; test on a node that has never had the toolchain loaded.&lt;/p&gt;

&lt;h2 id=&quot;the-principle-and-a-harness&quot;&gt;The principle, and a harness&lt;/h2&gt;

&lt;p&gt;A modulefile is a contract: load me and this tool will run. A contract that fails silently is
worse than no contract, because the user stops checking. If a module cannot keep its promise it
must say so at load time, in the shell the user is sitting in, with a non-zero exit.&lt;/p&gt;

&lt;p&gt;Hold it to that by testing the promise directly — load it in a shell with nothing inherited and
execute the binary. Not &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ls&lt;/code&gt; the directory, not check the variable. Run the thing.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;#!/bin/bash&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# modcheck.sh — load each module in a clean shell and actually run its binary.&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# Manifest is tab-separated:  &amp;lt;module&amp;gt;  &amp;lt;binary&amp;gt;  &amp;lt;smoke arguments&amp;gt;&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;set&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-u&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;MANIFEST&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;1&lt;/span&gt;:?usage:&lt;span class=&quot;p&quot;&gt; modcheck.sh manifest.tsv&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;INIT&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;MODULE_INIT&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;:-&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;/usr/share/lmod/lmod/init/bash&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;MP&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;MODULEPATH&lt;/span&gt;:?MODULEPATH&lt;span class=&quot;p&quot;&gt; must be set in the calling shell&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;fail&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;0

&lt;span class=&quot;k&quot;&gt;while &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;IFS&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;$&apos;&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\t&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;read&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; mod bin args &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;[[&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;mod&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;:-}&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;]]&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do&lt;/span&gt;
    &lt;span class=&quot;o&quot;&gt;[[&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-z&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;mod&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;:-}&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;mod&lt;/span&gt;:0:1&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;#&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;]]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;continue

    &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;out&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;env&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
            &lt;span class=&quot;nv&quot;&gt;HOME&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$HOME&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;USER&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$USER&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;LOGNAME&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$USER&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;TERM&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;dumb &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
            &lt;span class=&quot;nv&quot;&gt;PATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;/usr/bin:/bin &lt;span class=&quot;nv&quot;&gt;MODULEPATH&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$MP&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
          bash &lt;span class=&quot;nt&quot;&gt;--noprofile&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--norc&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;
            source &quot;$1&quot; &amp;gt;/dev/null 2&amp;gt;&amp;amp;1 || exit 91
            set -u
            module load &quot;$2&quot; &amp;gt;/dev/null || exit 92
            case &quot;:${LOADEDMODULES:-}:&quot; in
              *&quot;:$2:&quot;*|*&quot;:$2/&quot;*) ;;
              *) exit 92 ;;
            esac
            command -v &quot;$3&quot; &amp;gt;/dev/null  || exit 93
            exec &quot;$3&quot; $4
          &apos;&lt;/span&gt; _ &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$INIT&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$mod&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$bin&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;args&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;:-}&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; 2&amp;gt;&amp;amp;1&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;nv&quot;&gt;rc&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$?&lt;/span&gt;

    &lt;span class=&quot;o&quot;&gt;[[&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$rc&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-ne&lt;/span&gt; 0 &lt;span class=&quot;o&quot;&gt;]]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;fail&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1
    &lt;span class=&quot;k&quot;&gt;case&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$rc&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;in
      &lt;/span&gt;0&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;  &lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;ok    %-28s %s\n&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$mod&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$bin&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;;;&lt;/span&gt;
      91&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;INIT  %-28s init script missing: %s\n&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$mod&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$INIT&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;;;&lt;/span&gt;
      92&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;LOAD  %-28s module did not load\n&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$mod&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;;;&lt;/span&gt;
      93&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;PATH  %-28s loaded, but %s is not on PATH\n&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$mod&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$bin&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;;;&lt;/span&gt;
      &lt;span class=&quot;k&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;  &lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;RUN   %-28s %s exited %d: %s\n&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$mod&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$bin&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$rc&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
                 &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;%s&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$out&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;head&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-1&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;;;&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;esac&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt; &amp;lt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$MANIFEST&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;

&lt;span class=&quot;nb&quot;&gt;exit&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$fail&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Run against a manifest of four modules — one good, one that sets variables but puts no binary on
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt;, one whose binary runs and fails, and one whose modulefile is gone:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ ./modcheck.sh manifest.tsv
ok    samtools/1.19                samtools
PATH  toolkit/1.4                  loaded, but toolkit is not on PATH
RUN   gatk/4.5.0.0                 gatk exited 1: Error: LinkageError occurred while loading main class org.broadinstitute.hellbender.Main
LOAD  samtools/1.9                 module did not load
$ echo $?
1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Several things there are deliberate. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;env -i&lt;/code&gt; stops your interactive session leaking in, which is
the only way to catch a module that works for you because of your &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.bashrc&lt;/code&gt;. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--noprofile --norc&lt;/code&gt;
stops the site profile quietly fixing the problem. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;set -u&lt;/code&gt; is applied &lt;em&gt;after&lt;/em&gt; the init script is
sourced, because module init scripts are not written to survive it. The load is checked twice —
exit status and then &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LOADEDMODULES&lt;/code&gt; — which is the rule from section 5 applied to the harness
itself; a partial load reported as success is exactly what this is meant to catch. The sentinel
exit codes are 91 to 93 rather than 2 to 4 so that a tool exiting 3 on its own account is not
misfiled as a load failure. And the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;exec&lt;/code&gt; runs the real binary rather than testing for its
presence — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$4&lt;/code&gt; is deliberately unquoted so the smoke arguments word-split.&lt;/p&gt;

&lt;p&gt;Run it on more than one node — that is what catches the build-host class:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ for n in $NODES; do echo &quot;== $n&quot;; srun -N1 -w &quot;$n&quot; --time=10 ./modcheck.sh manifest.tsv; done
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Run it from a batch job as well as a login node. Those environments differ, and a module that only
works interactively fails for every job.&lt;/p&gt;

&lt;h2 id=&quot;symptom-against-cause&quot;&gt;Symptom against cause&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Symptom&lt;/th&gt;
      &lt;th&gt;Likely cause&lt;/th&gt;
      &lt;th&gt;Command that confirms it&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;command not found&lt;/code&gt; right after a successful load&lt;/td&gt;
      &lt;td&gt;load happened in a subshell (pipe, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$( )&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;xargs&lt;/code&gt;, make recipe)&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;echo &quot;$LOADEDMODULES&quot;&lt;/code&gt; in the shell that matters&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list&lt;/code&gt; shows nothing but the load printed no error&lt;/td&gt;
      &lt;td&gt;same, or the output went to stderr (Lmod’s default) and was filtered by a pipe&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list 2&amp;gt;&amp;amp;1&lt;/code&gt;, or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module -t list&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;which&lt;/code&gt; shows the new binary, old version still runs&lt;/td&gt;
      &lt;td&gt;a shell function or alias of the same name, which &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;which&lt;/code&gt; cannot see&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type &amp;lt;cmd&amp;gt;&lt;/code&gt;; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;command &amp;lt;cmd&amp;gt;&lt;/code&gt; to bypass once&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Same, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type&lt;/code&gt; says “hashed”&lt;/td&gt;
      &lt;td&gt;stale command hash — only possible when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; itself did not change&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type &amp;lt;cmd&amp;gt;&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;hash&lt;/code&gt;; clear with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;hash -r&lt;/code&gt; (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rehash&lt;/code&gt; in tcsh)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ModuleNotFoundError&lt;/code&gt; for a package that is installed&lt;/td&gt;
      &lt;td&gt;module set a venv variable but no interpreter on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python3 -c &quot;import sys; print(sys.executable)&quot;&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Unrelated Python scripts break after loading a tool&lt;/td&gt;
      &lt;td&gt;module prepended an interpreter and shadowed the system one&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type -a python3&lt;/code&gt;; fix with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;family()&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Two modules loaded, second one’s tool wins silently&lt;/td&gt;
      &lt;td&gt;no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;family&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;conflict&lt;/code&gt; declared&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list 2&amp;gt;&amp;amp;1&lt;/code&gt; then &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;type -a&lt;/code&gt; on the contested binary&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Works loaded one per line, fails in a single command&lt;/td&gt;
      &lt;td&gt;modulefile shelling out, or a partial load reported as success&lt;/td&gt;
      &lt;td&gt;load in one command, then &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list 2&amp;gt;&amp;amp;1&lt;/code&gt; and compare against what you asked for&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UnsupportedClassVersionError: class file version N&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;tool built for a newer JDK than the default&lt;/td&gt;
      &lt;td&gt;Java release = N − 44; add &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;prereq java/&amp;lt;n&amp;gt;&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Variable set, directory absent, no error&lt;/td&gt;
      &lt;td&gt;modulefile copied between clusters without an existence check&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ls &quot;$THAT_VAR&quot;&lt;/code&gt;; add a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;break&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LmodError&lt;/code&gt; guard&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$VAR/bin/tool&lt;/code&gt; resolves to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/bin/tool&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;variable expanded to empty string&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;echo &quot;[${VAR:-UNSET}]&quot;&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Module in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list&lt;/code&gt;, binary missing from disk&lt;/td&gt;
      &lt;td&gt;retired install, modulefile left behind&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module show &amp;lt;mod&amp;gt; 2&amp;gt;&amp;amp;1&lt;/code&gt;, then &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ls&lt;/code&gt; each path it prints&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libX.so.N: cannot open shared object file&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;build-host library not present on this node&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ldd &amp;lt;binary&amp;gt;&lt;/code&gt;; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;readelf -d &amp;lt;binary&amp;gt; \| grep RUNPATH&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;GLIBCXX_3.4.NN not found&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;built against a newer libstdc++ than the runtime provides&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;strings &amp;lt;libstdc++ in use&amp;gt; \| grep GLIBCXX \| sort -V \| tail -1&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;bad interpreter: No such file or directory&lt;/code&gt;, file exists&lt;/td&gt;
      &lt;td&gt;shebang longer than the kernel’s buffer (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;BINPRM_BUF_SIZE&lt;/code&gt;), path truncated&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;head -1 &amp;lt;script&amp;gt; \| wc -c&lt;/code&gt; and compare with the length in the error text&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Shell syntax errors from a Python or Perl entry point&lt;/td&gt;
      &lt;td&gt;same overflow on a newer kernel: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ENOEXEC&lt;/code&gt;, and the shell then runs the file itself&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;head -1 &amp;lt;script&amp;gt; \| wc -c&lt;/code&gt; against &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;BINPRM_BUF_SIZE&lt;/code&gt; on that node&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Works on a login node, fails in a batch job&lt;/td&gt;
      &lt;td&gt;different profile scripts or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MODULEPATH&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;run the same check under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun&lt;/code&gt; and diff the environments&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Works for you, fails for a user&lt;/td&gt;
      &lt;td&gt;something in your &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.bashrc&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;re-test under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;env -i ... bash --noprofile --norc&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;None of this is exotic. It is one failure repeated: an environment change that did not happen, or
happened where it could not be observed, reported as success. It is worth writing down because
the evidence points at the software, and the software is fine.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Restore benchmarks: measure the time until the business can work</title>
    <link href="https://www.wirewalk.com/writing/restore-benchmark-business-recovery/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/restore-benchmark-business-recovery/</id>
    <summary>Throughput is one component of recovery. Metadata, identity, keys, and application validation can dominate elapsed time.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A restore speed of several gigabytes per second can sound reassuring while saying little about the time needed to recover a real service. The benchmark must include the dataset shape and the work that happens before and after copying bytes.&lt;/p&gt;

&lt;h2 id=&quot;define-the-finish-line&quot;&gt;Define the finish line&lt;/h2&gt;

&lt;p&gt;Choose a service and define the recovery-time objective with its owner. Separate data availability from application readiness and business acceptance. Record the recovery point and acceptable data-loss window independently.&lt;/p&gt;

&lt;p&gt;Use a representative dataset: large files, small files, directory depth, permissions, and application metadata. A single large sequential file is an inadequate model for a namespace containing millions of small objects.&lt;/p&gt;

&lt;h2 id=&quot;establish-a-lower-bound&quot;&gt;Establish a lower bound&lt;/h2&gt;

&lt;p&gt;A simple transfer-time calculation is useful as a sanity check:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;100 TiB / 2 GiB per second = 51,200 seconds = about 14.2 hours
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This is arithmetic, not a measured product result. It excludes initialization, metadata work, protocol overhead, retries, and application checks. Use consistent units and actual sustained throughput from the intended recovery route.&lt;/p&gt;

&lt;p&gt;If the business requires recovery in four hours, a fourteen-hour byte-transfer lower bound already demonstrates a design mismatch. No dashboard wording can close that gap.&lt;/p&gt;

&lt;h2 id=&quot;time-the-complete-sequence&quot;&gt;Time the complete sequence&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Milestone&lt;/th&gt;
      &lt;th&gt;Record&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Recovery authorized&lt;/td&gt;
      &lt;td&gt;Start timestamp&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Credentials and keys available&lt;/td&gt;
      &lt;td&gt;Access preparation delay&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Destination ready&lt;/td&gt;
      &lt;td&gt;Provisioning delay&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Namespace available&lt;/td&gt;
      &lt;td&gt;Initial mount or restore completion&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Working set read&lt;/td&gt;
      &lt;td&gt;Cold-access completion&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Application accepted&lt;/td&gt;
      &lt;td&gt;Business validation timestamp&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;WEKA’s &lt;a href=&quot;https://docs.weka.io/weka-filesystems-and-object-stores/snap-to-obj&quot;&gt;Snap-To-Object documentation&lt;/a&gt; notes on-demand recovery considerations. Similar distinctions should be investigated for any product that exposes data before all of it is local. Do not equate a quick mount with a fully recovered working set.&lt;/p&gt;

&lt;h2 id=&quot;avoid-cache-driven-optimism&quot;&gt;Avoid cache-driven optimism&lt;/h2&gt;

&lt;p&gt;Run the first representative read separately from repeat reads. Record whether source or destination caches were warm and whether object-store retrieval was involved. A second run can measure a different system path.&lt;/p&gt;

&lt;p&gt;Measure file-count and metadata progress as well as bytes. Preserve failure and retry counts. A benchmark that silently skipped unreadable files can report excellent throughput while failing its purpose.&lt;/p&gt;

&lt;h2 id=&quot;validate-data-and-authority&quot;&gt;Validate data and authority&lt;/h2&gt;

&lt;p&gt;Compare a pre-recorded manifest, checksums for the test corpus, ownership, and expected access. Then execute an application transaction. Include a denied-access test for an unrelated user so the restored service does not trade availability for excessive exposure.&lt;/p&gt;

&lt;h2 id=&quot;report-the-actual-result&quot;&gt;Report the actual result&lt;/h2&gt;

&lt;p&gt;Publish the dataset, tool versions, topology, elapsed milestones, throughput, errors, and manual interventions. Label estimates as estimates and lab results as lab results. Do not extrapolate a small trial linearly without explaining the limits.&lt;/p&gt;

&lt;p&gt;The useful deliverable is a recovery time the business can assess, together with the bottleneck that must change if it is too slow.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Incident containment: preserve the route you need to investigate and recover</title>
    <link href="https://www.wirewalk.com/writing/incident-containment-preserve-recovery/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/incident-containment-preserve-recovery/</id>
    <summary>Containment decisions should stop harmful activity while preserving evidence and an authorized recovery path.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;Disconnecting a compromised system may limit damage, but an unplanned isolation step can also remove the only management route, interrupt evidence collection, or strand a critical dependency. A containment playbook should make those tradeoffs explicit before an emergency.&lt;/p&gt;

&lt;h2 id=&quot;establish-authority-and-scope&quot;&gt;Establish authority and scope&lt;/h2&gt;

&lt;p&gt;Identify who can authorize host isolation, account revocation, storage protection changes, and service shutdown. Keep an incident communication route available outside the systems under investigation.&lt;/p&gt;

&lt;p&gt;Record the known facts separately from hypotheses. A suspicious alert may justify precautionary isolation, but the incident record should not claim confirmed compromise solely because a detection fired.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.cisa.gov/stopransomware/ransomware-guide&quot;&gt;CISA’s ransomware guide&lt;/a&gt; provides incident-response and recovery guidance. The worksheet below is an operational method for applying those objectives to a specific environment.&lt;/p&gt;

&lt;h2 id=&quot;build-a-containment-decision-sheet&quot;&gt;Build a containment decision sheet&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Action&lt;/th&gt;
      &lt;th&gt;Intended effect&lt;/th&gt;
      &lt;th&gt;Dependency to preserve&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;EDR network isolation&lt;/td&gt;
      &lt;td&gt;Restrict host communication&lt;/td&gt;
      &lt;td&gt;Supported responder/management channel&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Account revocation&lt;/td&gt;
      &lt;td&gt;Stop further identity use&lt;/td&gt;
      &lt;td&gt;Independent recovery authority&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Firewall restriction&lt;/td&gt;
      &lt;td&gt;Limit affected traffic&lt;/td&gt;
      &lt;td&gt;DNS, logging, or recovery flows as justified&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Suspend replication workflow&lt;/td&gt;
      &lt;td&gt;Avoid propagating harmful state&lt;/td&gt;
      &lt;td&gt;Existing protected recovery points&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;These are possible actions, not automatic instructions. The incident lead selects them based on scope, platform behavior, and business impact.&lt;/p&gt;

&lt;h2 id=&quot;rehearse-the-endpoint-control&quot;&gt;Rehearse the endpoint control&lt;/h2&gt;

&lt;p&gt;For an EDR product such as Microsoft Defender for Endpoint or another deployed platform, use its current vendor procedure on a lab endpoint. Establish which communication remains available during isolation and how authorized responders reverse it.&lt;/p&gt;

&lt;p&gt;Do not assume all operating systems or agent versions behave identically. Observe the lab endpoint from a second host, verify the responder path, and time restoration. Record what happens if the agent is unhealthy or the endpoint is already offline.&lt;/p&gt;

&lt;h2 id=&quot;preserve-evidence-deliberately&quot;&gt;Preserve evidence deliberately&lt;/h2&gt;

&lt;p&gt;Capture relevant timestamps, alerts, host identity, and action records. Decide with the response team which volatile evidence must be collected before disruptive actions. Avoid publishing sensitive logs in a general collaboration channel.&lt;/p&gt;

&lt;p&gt;Containment should not trigger uncontrolled cleanup. Deleting files or rebuilding immediately can remove evidence needed to establish scope and root cause. Preserve recovery options while the incident team makes that decision.&lt;/p&gt;

&lt;h2 id=&quot;separate-isolation-from-recovery-acceptance&quot;&gt;Separate isolation from recovery acceptance&lt;/h2&gt;

&lt;p&gt;A host can be contained and still be unsafe to reconnect. Establish the criteria for rebuilding or restoring it, rotating exposed credentials, and accepting the recovered application. Confirm the selected recovery point and its dependencies rather than restoring the most recent backup automatically.&lt;/p&gt;

&lt;p&gt;Use a tabletop scenario that includes loss of the normal directory and backup console. It reveals whether the procedure depends on the same authority that may be compromised.&lt;/p&gt;

&lt;h2 id=&quot;record-outcomes-not-just-button-clicks&quot;&gt;Record outcomes, not just button clicks&lt;/h2&gt;

&lt;p&gt;Keep the requested action, observed network effect, remaining access, evidence preserved, and restoration result. The product console’s “isolated” label should be corroborated by the behavior the team intended.&lt;/p&gt;

&lt;p&gt;Continue with &lt;a href=&quot;/writing/recovery-after-admin-compromise/&quot;&gt;recovery after administrator compromise&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Ninety-seven percent of an annotation pipeline was annotating nothing</title>
    <link href="https://www.wirewalk.com/writing/ninety-seven-percent-of-the-work-was-unnecessary/"/>
    <published>2026-06-24T10:00:00-04:00</published>
    <updated>2026-06-24T10:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/ninety-seven-percent-of-the-work-was-unnecessary/</id>
    <summary>How to find out that the expensive stage of a variant annotation pipeline is spending almost all of its time on records with nothing to annotate, how to remove that work in the right order so you do not silently lose sites, and how to prove the fast output is identical to the slow one.</summary>
    <content type="html">&lt;p&gt;A variant annotation pipeline takes per-sample genomic VCF files in and produces annotated
variants out. It is slow. The obvious response is to make the annotator faster: more forks, faster
storage for the cache, a different annotation tool, a bigger node. All of those are real levers and
all of them are the wrong place to start.&lt;/p&gt;

&lt;p&gt;The right place to start is a stopwatch on each stage, because the shape of this problem recurs: the
annotator is not slow, it is being asked to do roughly thirty-eight times more work than the
question requires. The records it spends its time on carry no alternate allele at all. There is
nothing to annotate, so it annotates nothing, carefully, twenty-five million times per sample.&lt;/p&gt;

&lt;p&gt;What follows is the pattern — how to measure it, the pipeline that removes the waste, the
correctness trap that eats sites without telling you, the two tuning parameters that only showed up
in a sweep, and the caveat that limits the whole result to one file format.&lt;/p&gt;

&lt;h2 id=&quot;measure-the-stages-before-touching-anything&quot;&gt;Measure the stages before touching anything&lt;/h2&gt;

&lt;p&gt;The first number you need is the split of wall time across stages. For a shell pipeline this does
not need a profiler. It needs timestamps that survive into the log:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;stamp&lt;span class=&quot;o&quot;&gt;()&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;%s\t%s\n&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;date&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-u&lt;/span&gt; +%Y-%m-%dT%H:%M:%SZ&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$TIMING&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;}&lt;/span&gt;

stamp &lt;span class=&quot;s2&quot;&gt;&quot;start&quot;&lt;/span&gt;
bcftools index &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$GVCF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;                                   &lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; stamp &lt;span class=&quot;s2&quot;&gt;&quot;index&quot;&lt;/span&gt;
run_annotator &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$GVCF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$OUT&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;                                &lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; stamp &lt;span class=&quot;s2&quot;&gt;&quot;annotate&quot;&lt;/span&gt;
postprocess &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$OUT&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;                                          &lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; stamp &lt;span class=&quot;s2&quot;&gt;&quot;postprocess&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If the work runs under Slurm, take the per-step figures from the accounting database rather than
from the job script, and take them from the job steps, not the allocation:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ sacct -j 4182377 --units=G \
    --format=JobID,JobName%18,Elapsed,TotalCPU,MaxRSS,State
JobID          JobName      Elapsed   TotalCPU     MaxRSS      State
4182377        annot-batch  05:02:11  19:55:38                COMPLETED
4182377.batch  batch        05:02:11  00:01:46      0.31G     COMPLETED
4182377.0      vep          04:53:02  19:41:08         9.8G   COMPLETED
4182377.1      post         00:08:57  00:12:44      1.20G     COMPLETED
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Read the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; column carefully: it is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[DD-]HH:MM:SS&lt;/code&gt;, so a leading day field changes the
value by a factor of twenty-four and is easy to skim past. Here &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;19:41:08&lt;/code&gt; of CPU against &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;04:53:02&lt;/code&gt;
of wall on the annotation step is 4.03x, which is the four forks the step was given and a useful
sanity check that the parallelism you asked for is the parallelism you got.&lt;/p&gt;

&lt;p&gt;Two further warnings about that output, both of which produce confidently wrong answers. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; is
blank on the allocation line because the allocation itself has no measured process tree, so any
memory figure derived from a query with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt; comes back as zero for every job. And if you derive CPU
efficiency from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; rather than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; you get 1.0 for every job ever run, because
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; is allocated core-seconds and not consumed ones. Use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; over &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCPUS * Elapsed&lt;/code&gt;,
and query without &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The split from that run, for one representative sample on one node:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;annotation: 293.0 min&lt;/li&gt;
  &lt;li&gt;everything else — indexing, splitting, post-processing, compression, QC summary: 9.0 min&lt;/li&gt;
  &lt;li&gt;total: 302.0 min&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Annotation is 97.0% of the wall time. Nothing else on the list can matter. If you halved the cost of
every other stage you would save four and a half minutes out of five hours.&lt;/p&gt;

&lt;h2 id=&quot;what-is-actually-in-the-file&quot;&gt;What is actually in the file&lt;/h2&gt;

&lt;p&gt;A per-sample genomic VCF is not a list of variants. It is a description of the whole callable
genome, in which the overwhelming majority of records say “this stretch matches the reference and
here is how confident I am”. Those records carry a symbolic placeholder allele instead of a real
one: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;lt;NON_REF&amp;gt;&lt;/code&gt; for GATK and DRAGEN output, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;lt;*&amp;gt;&lt;/code&gt; for htslib-style gVCF. Reference blocks also
carry an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;END&lt;/code&gt; tag giving the last position of the block.&lt;/p&gt;

&lt;p&gt;The important structural detail is that in a GATK-style gVCF the placeholder is present on &lt;em&gt;every&lt;/em&gt;
record, including real variant records, where it sits last in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt; list. So a reference block
looks like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT=&amp;lt;NON_REF&amp;gt;&lt;/code&gt; and a variant looks like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT=A,&amp;lt;NON_REF&amp;gt;&lt;/code&gt;. The test for “this record has
something to annotate” is therefore not “does the ALT column mention the placeholder” — it is “does
the ALT column contain anything besides the placeholder”.&lt;/p&gt;

&lt;p&gt;That distinction has teeth in bcftools expression syntax, because a comma inside a string is read as
a value separator and multiple values are combined with OR. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT=&quot;&amp;lt;NON_REF&amp;gt;&quot;&lt;/code&gt; is therefore &lt;em&gt;true&lt;/em&gt; for
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT=A,&amp;lt;NON_REF&amp;gt;&lt;/code&gt; — it matches the second value and stops caring about the first. Written alone as an
exclude expression it would discard every variant record in the file. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N_ALT=1&lt;/code&gt; conjunction is
what narrows it to records where the placeholder is the only allele, and it is the whole reason the
filter is two clauses rather than one.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bcftools view -H sample.g.vcf.gz | head -4 | cut -f1-5,10
chr1  10001  .  T  &amp;lt;NON_REF&amp;gt;    0/0:...
chr1  10002  .  A  &amp;lt;NON_REF&amp;gt;    0/0:...
chr1  10108  .  C  CAACCCT,&amp;lt;NON_REF&amp;gt;  0/1:...
chr1  10109  .  A  &amp;lt;NON_REF&amp;gt;    0/0:...
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Count both classes:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bcftools view -H sample.g.vcf.gz | wc -l
26584113

$ bcftools view -H -e &apos;N_ALT=1 &amp;amp;&amp;amp; ALT=&quot;&amp;lt;NON_REF&amp;gt;&quot;&apos; sample.g.vcf.gz | wc -l
690412
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;690,412 out of 26,584,113 is 2.60%. Removing the rest is a 38.5-fold reduction in records entering
the stage that is 97% of the cost.&lt;/p&gt;

&lt;p&gt;Two honest caveats on those figures. They are one file, and absolute record counts in gVCFs are not
portable: reference-block banding differs sharply between callers and settings, and a file emitted at
base-pair resolution has a record per callable base, which is two orders of magnitude more. The
&lt;em&gt;ratio&lt;/em&gt; is the durable number, and across the per-sample whole-genome gVCFs I have counted it sits in
the low single digits of a percent. If your own files give a ratio above about ten percent, either
they are not per-sample gVCFs or the caller banding is unusual, and you should check before assuming
this result transfers.&lt;/p&gt;

&lt;p&gt;Cross-check the classification against the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;END&lt;/code&gt; tag rather than trusting a single expression:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bcftools query -f &apos;%INFO/END\n&apos; sample.g.vcf.gz | grep -cv &apos;^\.$&apos;
25893701
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;25,893,701 records carry an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;END&lt;/code&gt;, and 26,584,113 − 690,412 = 25,893,701. Two tests reading two
different fields — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;INFO/END&lt;/code&gt; — agree exactly, which is the point of running both. They are
not fully independent, since a caller that got one wrong could plausibly get the other wrong the same
way, but they do catch the failure that actually happens, which is an expression that means something
other than what you thought it meant.&lt;/p&gt;

&lt;h2 id=&quot;the-pipeline&quot;&gt;The pipeline&lt;/h2&gt;

&lt;p&gt;Order matters, for reasons covered in the next section. Filter first, split multi-allelic records
second, drop the placeholder third.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;set&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-euo&lt;/span&gt; pipefail

&lt;span class=&quot;nv&quot;&gt;REF&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;/ref/GRCh38.fa
&lt;span class=&quot;nv&quot;&gt;IN&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;sample.g.vcf.gz
&lt;span class=&quot;nv&quot;&gt;SITES&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;sample.sites.vcf.gz

bcftools view &lt;span class=&quot;nt&quot;&gt;--threads&lt;/span&gt; 4 &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;N_ALT=1 &amp;amp;&amp;amp; ALT=&quot;&amp;lt;NON_REF&amp;gt;&quot;&apos;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-Ou&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$IN&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | bcftools norm &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-any&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-Ou&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | bcftools view &lt;span class=&quot;nt&quot;&gt;--threads&lt;/span&gt; 4 &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;ALT=&quot;&amp;lt;NON_REF&amp;gt;&quot; || ALT=&quot;*&quot;&apos;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-Oz&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$SITES&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;

bcftools index &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$SITES&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--threads&lt;/code&gt; appears only on the first and last commands on purpose. It is a BGZF compression and
decompression pool, so it buys something where a bgzipped file is being read and where one is being
written, and nothing at all on the middle stage, which takes uncompressed BCF on a pipe and emits the
same. Putting it everywhere is harmless but it is cargo cult, and it makes the stage look more
parallel than it is.&lt;/p&gt;

&lt;p&gt;After &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;norm -m -any&lt;/code&gt;, a record with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT=A,&amp;lt;NON_REF&amp;gt;&lt;/code&gt; becomes two records, one with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT=A&lt;/code&gt; and one
with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT=&amp;lt;NON_REF&amp;gt;&lt;/code&gt;, so the final &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;view&lt;/code&gt; removes the now-standalone placeholders. The same pass
drops spanning-deletion &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;*&lt;/code&gt; alleles. A &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;*&lt;/code&gt; is not a new allele: it marks a position overlapped by a
deletion called at an earlier position, so there is no consequence to compute for it. Drop them
deliberately rather than letting the annotator decide.&lt;/p&gt;

&lt;p&gt;Then annotate the small file:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;vep &lt;span class=&quot;nt&quot;&gt;--offline&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--cache&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--dir_cache&lt;/span&gt; /ref/vep &lt;span class=&quot;nt&quot;&gt;--assembly&lt;/span&gt; GRCh38 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;--fasta&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;--format&lt;/span&gt; vcf &lt;span class=&quot;nt&quot;&gt;--vcf&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--compress_output&lt;/span&gt; bgzip &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;--allele_number&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;--no_stats&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;--fork&lt;/span&gt; 4 &lt;span class=&quot;nt&quot;&gt;--buffer_size&lt;/span&gt; 50000 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$SITES&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; sample.vep.vcf.gz
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--allele_number&lt;/code&gt; is the one flag here whose purpose is routinely misread, including by me when I
first wrote this pipeline. The documentation says it will “identify allele number from VCF input,
where 1 = first ALT allele, 2 = second ALT allele etc.” The obvious reading is that it preserves the
&lt;em&gt;pre-split&lt;/em&gt; allele ordering so you can get back to the original multi-allelic record. It does not. The
index it records is the index within the file VEP was handed, and that file has already been split, so
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALLELE_NUM&lt;/code&gt; is 1 on every record. Recovering the unsplit record is a join on position plus &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;REF&lt;/code&gt; and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt;, and no VEP flag does it for you.&lt;/p&gt;

&lt;p&gt;The flag still earns its place, for a different reason. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Allele&lt;/code&gt; field inside &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CSQ&lt;/code&gt; is a minimal
representation — VEP trims the sequence common to the reference and alternate alleles off both ends
before computing consequences — so for an insertion written as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;REF=C ALT=CAACCCT&lt;/code&gt; the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CSQ&lt;/code&gt; allele is
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AACCCT&lt;/code&gt;, which matches nothing in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt; column. String-matching &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CSQ&lt;/code&gt; alleles against &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt; works
for SNVs and quietly fails for indels. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALLELE_NUM&lt;/code&gt; being unconditionally 1 is what makes that
matching unnecessary: one input allele per record, so every consequence block on that record — and
there is one per overlapping transcript — belongs to that allele, tied by index rather than by
string.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--no_stats&lt;/code&gt; is worth a mention, with the caveat that the VEP documentation itself describes the
saving as marginal. I did not measure it in isolation, and on a 690,412-record input I would not
expect it to be visible against the other numbers here. The reason to set it is not speed, it is that
the summary is an HTML file nobody opens in a batch run and one more artefact per sample to clean up.&lt;/p&gt;

&lt;h2 id=&quot;the-correctness-trap&quot;&gt;The correctness trap&lt;/h2&gt;

&lt;p&gt;The first version of the filter was one command instead of three, and it used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--trim-alt-alleles&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# WRONG - do not do this&lt;/span&gt;
bcftools view &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;N_ALT=1 &amp;amp;&amp;amp; ALT=&quot;&amp;lt;NON_REF&amp;gt;&quot;&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$IN&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-Oz&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$SITES&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-a&lt;/code&gt; removes alternate alleles that are not present in any genotype. On a gVCF this looks like
exactly what you want: it strips the placeholder, it strips alleles the caller considered and
rejected, and it leaves a tidy file. The output was slightly smaller on disk than the three-command
version — not fewer records, fewer bytes — which read as the filter working harder.&lt;/p&gt;

&lt;p&gt;It was not working harder. It was throwing away 2,157 sites per file, and the record count did not
move, which is exactly why it took a two-way comparison to find.&lt;/p&gt;

&lt;p&gt;The sites it lost all had the same shape: a real alternate allele in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt;, and a genotype that did
not use it. Hom-ref and no-call genotypes on records that nonetheless carry a called alternate allele
are ordinary in a gVCF — the caller saw allele evidence, emitted the allele, and genotyped the sample
as reference or as uncallable. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-a&lt;/code&gt; sees an allele absent from the genotype and removes it — and it
removes the placeholder on the same grounds, because &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;lt;NON_REF&amp;gt;&lt;/code&gt; is not in any genotype either. The
documented behavior when nothing survives is the part that bites: “if no alternate allele remains
after trimming, the record itself is not removed but ALT is set to ‘.’”. So the record stays. It
passes the exclude expression, because &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N_ALT&lt;/code&gt; is now zero and the expression asks about &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;N_ALT=1&lt;/code&gt;. It
is counted in the filter output. It simply has no allele left to annotate, and it falls out inside the
annotator instead of at the filter, which is why the record count after the filter stage still looks
right. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-a&lt;/code&gt; also trims &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Number=A&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;G&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;R&lt;/code&gt; tags to match, so the per-allele fields go with it.&lt;/p&gt;

&lt;p&gt;The near-neighbour makes this worse. bcftools 1.19 added &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-A, --trim-unseen-allele&lt;/code&gt;, which removes
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;lt;*&amp;gt;&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;lt;NON_REF&amp;gt;&lt;/code&gt; specifically — at variant sites with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-A&lt;/code&gt;, at all sites with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-AA&lt;/code&gt; — without ever
consulting a genotype. That is the operation the first version was reaching for. One character of
difference between the flag that is safe here and the flag that silently drops called sites, and no
error from either.&lt;/p&gt;

&lt;p&gt;Here is what finding it looked like. Compare the fast output against the baseline, in both
directions:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bcftools isec -C -w1 baseline.norm.vcf.gz fast.vep.vcf.gz | grep -vc &apos;^#&apos;
2157
$ bcftools isec -C -w1 fast.vep.vcf.gz baseline.norm.vcf.gz | grep -vc &apos;^#&apos;
0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Asymmetric loss, which immediately rules out a representation difference — those show up in both
directions. Then look at what was lost:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bcftools isec -C -w1 baseline.norm.vcf.gz fast.vep.vcf.gz \
    | bcftools query -f &apos;[%GT]\n&apos; - | sort | uniq -c | sort -rn
   1188 ./.
    902 0/0
     67 0|0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Every single lost site has a genotype that does not use the allele. That is the fingerprint, and it
points straight at the trimming flag.&lt;/p&gt;

&lt;p&gt;Whether this matters depends on what the annotated output is for. If the only consumer is a
per-sample report of the variants this sample carries, the lost records were arguably noise. If the
output feeds a review queue, a re-genotyping step, a family or cohort comparison where “the caller
saw this allele and called hom-ref” is different from “the caller never saw this allele”, or any
regulated workflow where the site list must be reproducible, then losing 2,157 sites per sample
without a log line is unacceptable. Removing work that is genuinely unnecessary is optimization.
Removing work you did not know you were doing is data loss, and the two are indistinguishable until
you compare outputs.&lt;/p&gt;

&lt;p&gt;The fix is the ordering above: filter on the placeholder alone, split, then drop what the split
exposed. No stage ever consults a genotype, so no stage can drop a record for having the wrong one.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-A&lt;/code&gt; would do the third step too, and on a recent enough bcftools it is the more direct spelling; the
explicit &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;view -e&lt;/code&gt; is what I left in because it also handles the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;*&lt;/code&gt; alleles and because it says out
loud what is being removed.&lt;/p&gt;

&lt;h2 id=&quot;what-it-cost-and-what-it-saved&quot;&gt;What it cost and what it saved&lt;/h2&gt;

&lt;p&gt;One sample, one node, single file at a time:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Stage&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Baseline&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Filter only&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Filter + tuned annotator&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Filter, split, index&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;—&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;5.5 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;5.5 min&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Annotation&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;293.0 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;7.6 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;5.4 min&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Post-processing, compress, QC&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;9.0 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;9.0 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;9.0 min&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;Total wall&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;&lt;strong&gt;302.0 min&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;&lt;strong&gt;22.1 min&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;&lt;strong&gt;19.9 min&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Speed-up vs baseline&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1.0x&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;13.7x&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;15.2x&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Records entering annotation&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;26,584,113&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;690,412&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;690,412&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Mean cost per record&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.661 ms&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.661 ms&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.469 ms&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Peak RSS, annotation&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;9.8 GB&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;4.1 GB&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;11.6 GB&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The row that changed how I think about this is the second-to-last one. Before measuring, the
expectation was that reference blocks would be much cheaper per record than variant records — a
block has no consequence to compute, and a variant has to be intersected with every overlapping
transcript. If that were true, removing 97% of the records would remove far less than 97% of the
time, and the whole exercise would disappoint.&lt;/p&gt;

&lt;p&gt;The measurement says the opposite. Cost per record was flat to three significant figures across a
38-fold change in what was in the file. The fixed per-record overhead — parse, buffer, region lookup,
serialize — dominates the consequence calculation so completely that the consequence calculation
barely registers. That is why the saving tracked record count almost exactly, and it is also a
useful general warning: the intuition that “the interesting records are the expensive ones” is worth
checking rather than assuming, and here it was simply false.&lt;/p&gt;

&lt;p&gt;The arithmetic is then worth doing honestly, because the headline number is smaller than the naive
prediction. Amdahl on the profile says: 9.0 minutes of untouchable work plus 293.0/38.5 = 7.6 minutes
of annotation, which is 16.6 minutes, an 18.2x speed-up. The measured filter-only figure is 13.7x.
The gap is the filtering pass itself, which still has to decompress and read all 26,584,113 records —
you cannot skip records without looking at them — and which costs 5.5 minutes that Amdahl on the
original profile does not know about. Tuning the annotator then recovered 1.41x on the annotation
stage alone (7.6 minutes down to 5.4), which is worth only 1.11x on the total, because by that point
the total is dominated by the 14.5 minutes of filtering and post-processing that tuning does not
touch. 13.7x to 15.2x. Anyone quoting 38x from the record ratio alone has forgotten both the residual
3% and the cost of the skip.&lt;/p&gt;

&lt;h2 id=&quot;two-parameters-that-only-turned-up-in-a-sweep&quot;&gt;Two parameters that only turned up in a sweep&lt;/h2&gt;

&lt;p&gt;Neither of these would have been found by reasoning. Both came out of running the same 690,412-record
file across a grid.&lt;/p&gt;

&lt;p&gt;Fork count, at a buffer of 50,000, on a node with 48 physical cores otherwise idle:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fork&lt;/code&gt;&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Wall&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Speed-up vs 1&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Parallel efficiency&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;19.0 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1.00x&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;100%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;2&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;10.2 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1.86x&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;93%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;4&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;5.4 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;3.52x&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;88%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;8&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;6.1 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;3.11x&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;39%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;12&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;7.9 min&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;2.41x&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;20%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;There is a knee at four and a genuine regression beyond it, not a plateau. Eight forks is slower in
wall-clock terms than four while consuming twice the cores. My reading of the mechanism is that
forked workers hand results back to a single parent that has to collect them and re-emit records in
input order, so past a certain width the parent becomes the constraint and the workers wait on it —
but that is inference from the shape of the curve, not something I instrumented, and the knee may
well sit somewhere else for a different cache, a different record mix or a different VEP release.
What does transfer is the practical consequence: measure before widening, because on this stage
“give it more cores” made things worse, which is the opposite of the instinct that sends you to a
bigger node.&lt;/p&gt;

&lt;p&gt;Buffer size, at four forks:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;5,000 (the default): 7.6 min, 4.1 GB peak RSS&lt;/li&gt;
  &lt;li&gt;25,000: 5.9 min, 7.2 GB&lt;/li&gt;
  &lt;li&gt;50,000: 5.4 min, 11.6 GB&lt;/li&gt;
  &lt;li&gt;100,000: 5.4 min, 22.4 GB&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;1.41x from the default to 50,000, and nothing after that except memory. The buffer is how many
variants are held and processed as a batch, and a batch is resolved against the cache by region, so
a larger buffer means the same cache region is fetched and parsed fewer times. Once the buffer is
comfortably larger than the typical run of variants within one cache region, there is nothing left to
amortize, which is consistent with the flat line between 50,000 and 100,000. I am confident in the
measurement; the explanation is my reading of the behavior and not something I have confirmed
against the implementation.&lt;/p&gt;

&lt;p&gt;Memory is the real limit here, and it interacts with the next section: 11.6 GB per file times six
concurrent files is 70 GB, which fits on a 192 GB node; 22.4 GB times six does not leave room for
page cache, and buys nothing.&lt;/p&gt;

&lt;h2 id=&quot;concurrency-measured-rather-than-assumed&quot;&gt;Concurrency, measured rather than assumed&lt;/h2&gt;

&lt;p&gt;The stage is now short enough that per-file parallelism is not the interesting axis. Running several
files at once is. With four forks each, six concurrent files put 24 forked workers on a 48-core node,
plus six parent processes that are collecting and serializing rather than computing.&lt;/p&gt;

&lt;p&gt;Measured by running the identical file six times simultaneously and taking the wall time of each:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;one file alone: 19.9 min&lt;/li&gt;
  &lt;li&gt;six files concurrently: mean 20.9 min, slowest 21.0 min&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Per-file slowdown is 5%, so aggregate throughput is 95% of six times the single-file rate. Eight
concurrent files dropped to 82% efficiency, and the drop was in the annotation stage, not the filter
stage. Two explanations fit and I did not separate them: eight files is 32 workers plus eight
parents, which is 40 runnable processes on 48 cores and getting close enough to contend, and it is
also more concurrent readers against the same cache files, so page-cache pressure is equally
plausible. I would not present either as the cause.&lt;/p&gt;

&lt;p&gt;Scaled to a batch of roughly 210 samples on four nodes, at six files in flight per node — 24
concurrent, so nine waves with the last one part-empty — the baseline pipeline is about 45 hours of
wall time and the revised one is a little over three. Both figures are arithmetic over the per-sample
measurements above, not a full batch re-run at both settings, and the two sides are not measured
equally: the revised figure uses the 20.9-minute six-concurrent number, while the baseline figure
uses the 302-minute single-file number, because I never ran six baseline files at once. If the
baseline carries the same concurrency penalty, it is nearer 47 hours and the ratio is slightly better
than stated. Scale to your own sample count before quoting any of it — the wave arithmetic is lumpy,
and a batch that does not divide evenly into 24 lands worse than the ratio suggests.&lt;/p&gt;

&lt;h2 id=&quot;proving-the-output-is-identical&quot;&gt;Proving the output is identical&lt;/h2&gt;

&lt;p&gt;A 15x speed-up you cannot verify is a 15x speed-up you should not ship. Two properties need proving:
the same set of sites, and the same consequence calls at each site.&lt;/p&gt;

&lt;p&gt;The trap in the comparison is that the fast path normalizes and splits multi-allelic records and the
baseline does not, so a direct diff shows thousands of differences that are pure representation.
Normalize both sides identically before comparing, using the same reference and the same options:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;bcftools norm &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-any&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; baseline.vep.vcf.gz &lt;span class=&quot;nt&quot;&gt;-Ou&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | bcftools view &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;ALT=&quot;&amp;lt;NON_REF&amp;gt;&quot; || ALT=&quot;*&quot;&apos;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-Oz&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; baseline.norm.vcf.gz
bcftools index &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; baseline.norm.vcf.gz
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Site set, both directions:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bcftools isec -C -w1 baseline.norm.vcf.gz fast.vep.vcf.gz | grep -vc &apos;^#&apos;
0
$ bcftools isec -C -w1 fast.vep.vcf.gz baseline.norm.vcf.gz | grep -vc &apos;^#&apos;
0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Zero in both directions is the claim worth making. One direction alone proves nothing: a fast path
that emits a superset passes a one-way check.&lt;/p&gt;

&lt;p&gt;What &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;isec&lt;/code&gt; counts as “the same site” is worth knowing before you trust the zero. It defaults to
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--collapse none&lt;/code&gt;, under which only records with identical &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;REF&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt; are compatible, so this
comparison is on alleles and not merely on positions. That is the strict reading and the one you
want. Had it defaulted to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;all&lt;/code&gt;, two records at the same coordinate with completely different alleles
would have matched and both directions would have read zero while the alleles disagreed.&lt;/p&gt;

&lt;p&gt;Consequence calls, by digest of the annotation payload rather than by eye:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ for f in baseline.norm.vcf.gz fast.vep.vcf.gz; do
    bcftools query -f &apos;%CHROM:%POS:%REF:%ALT\t%INFO/CSQ\n&apos; &quot;$f&quot; \
      | sort | sha256sum | awk -v f=&quot;$f&quot; &apos;{print f, $1}&apos;
  done
baseline.norm.vcf.gz  9f3c1b...e4a7
fast.vep.vcf.gz       9f3c1b...e4a7
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Compare the body, never the whole file: the headers differ legitimately because they record the
command line and the run time. And the digest comparison is only meaningful if both runs used the
same annotator version, the same cache version and the same options — a cache upgrade between the
two runs changes consequence calls for real reasons and will produce a mismatch that has nothing to
do with your filter. Pin both, and record both in the run log.&lt;/p&gt;

&lt;p&gt;A useful third check, cheap to run and independent of the first two:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ for f in baseline.norm.vcf.gz fast.vep.vcf.gz; do bcftools +counts &quot;$f&quot;; done
Number of samples: 1
Number of SNPs:    598204
Number of INDELs:  92208
Number of MNPs:    0
Number of others:  0
Number of sites:   690412
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Identical on both. 598,204 + 92,208 = 690,412, which also ties back to the pre-annotation count and
confirms nothing was added or lost inside the annotator.&lt;/p&gt;

&lt;p&gt;That sum only works because the records are biallelic. The plugin increments a counter for &lt;em&gt;every&lt;/em&gt;
variant type a record contains and increments the site counter once, so a multi-allelic record
carrying a SNV and an indel adds one to each of three counters. Run &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;+counts&lt;/code&gt; on an unsplit file and
the categories over-sum against the site count, which looks like corruption and is not. The
cross-check is valid here because &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;norm -m -any&lt;/code&gt; has already guaranteed one allele per record.&lt;/p&gt;

&lt;h2 id=&quot;the-caveat-that-limits-all-of-this&quot;&gt;The caveat that limits all of this&lt;/h2&gt;

&lt;p&gt;The waste removed here is a property of one file format. A per-sample genomic VCF describes every
callable position, so it is mostly reference blocks, so most of it is not annotatable. That is why
the filter works.&lt;/p&gt;

&lt;p&gt;A joint-called multi-sample VCF has no reference blocks at all. It stores sites — positions where
at least one sample in the cohort carries something — and every record in it has a real alternate
allele. There is no 97% to remove. Applying this filter to a joint callset gives you a full
decompression-and-parse pass over the file that removes nothing, which is a pure regression, and the
“records entering annotation” row of that table would read the same before and after. It will not
cost the 5.5 minutes quoted above — a joint callset has far fewer records than a per-sample gVCF, so
the pass is proportionally cheaper — but cheaper than useless is still useless.&lt;/p&gt;

&lt;p&gt;The equivalent saving on a joint callset exists but is on a completely different axis. The redundancy
there is across samples, not across positions: if you annotate N per-sample files independently, you
annotate the same common variants N times. The fix is to build one union site list, annotate it once,
and join the annotations back onto the per-sample files by position and allele:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;bcftools merge &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; none &lt;span class=&quot;nt&quot;&gt;-Ou&lt;/span&gt; sample&lt;span class=&quot;k&quot;&gt;*&lt;/span&gt;.sites.vcf.gz &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | bcftools norm &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-any&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-Oz&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; cohort.sites.vcf.gz
bcftools index &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; cohort.sites.vcf.gz

vep &lt;span class=&quot;nt&quot;&gt;--offline&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--cache&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--dir_cache&lt;/span&gt; /ref/vep &lt;span class=&quot;nt&quot;&gt;--assembly&lt;/span&gt; GRCh38 &lt;span class=&quot;nt&quot;&gt;--fasta&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$REF&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;--fork&lt;/span&gt; 4 &lt;span class=&quot;nt&quot;&gt;--buffer_size&lt;/span&gt; 50000 &lt;span class=&quot;nt&quot;&gt;--no_stats&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--vcf&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--allele_number&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; cohort.sites.vcf.gz &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; cohort.csq.vcf.gz &lt;span class=&quot;nt&quot;&gt;--compress_output&lt;/span&gt; bgzip
bcftools index &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; cohort.csq.vcf.gz

bcftools annotate &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; cohort.csq.vcf.gz &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; INFO/CSQ &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; sample.annotated.vcf.gz &lt;span class=&quot;nt&quot;&gt;-Oz&lt;/span&gt; sample.sites.vcf.gz
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The join in that last command is the place it goes wrong. When the annotation file is a VCF,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;bcftools annotate&lt;/code&gt; matches on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CHROM&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;POS&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;REF&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ALT&lt;/code&gt;, and a record that does not match is
left unannotated in silence — there is no warning and no non-zero exit. So both sides have to have
been normalized with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;norm -m -any&lt;/code&gt; against the &lt;em&gt;same&lt;/em&gt; reference FASTA, or a left-alignment
difference at a repeat-adjacent indel produces two records that a human would call the same variant
and bcftools will not. Check the join by counting records with a missing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CSQ&lt;/code&gt; afterwards rather than
assuming it landed. The annotation file also has to be bgzip-compressed and indexed, which is what
the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;index -t&lt;/code&gt; above is for.&lt;/p&gt;

&lt;p&gt;The scaling of that is set by the overlap between samples, not by the reference-block fraction, and
the overlap depends on cohort size and ancestry composition. I have not measured it carefully enough
to quote a factor, so I will not quote one.&lt;/p&gt;

&lt;p&gt;One last thing that is easy to get wrong at the end of a project like this. The filtered file is an
annotation &lt;em&gt;input&lt;/em&gt;, not a replacement for the gVCF. The reference blocks you removed are what
distinguishes “this sample has no variant here” from “this position was never callable in this
sample”, and that distinction is the entire reason the gVCF format exists. Keep the original. The
optimization is that the annotator never sees it.&lt;/p&gt;

&lt;h2 id=&quot;what-to-take-from-this&quot;&gt;What to take from this&lt;/h2&gt;

&lt;p&gt;Profile first, in stages, with timestamps that end up in a log someone else can read. If one stage is
97% of the cost, nothing you do to the other stages matters, and the only question worth asking about
that stage is whether the work it is doing needs doing at all.&lt;/p&gt;

&lt;p&gt;Then verify in both directions before you believe your own result, because the fastest version of any
pipeline is the one that quietly does less than it should, and it will pass every check that only
looks at how long it took.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Kubernetes security: test admission, network policy, and storage access separately</title>
    <link href="https://www.wirewalk.com/writing/kubernetes-security-boundaries/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/kubernetes-security-boundaries/</id>
    <summary>A namespace is an organizational boundary. Verify the controls that make it an access boundary.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A workload placed in its own namespace is easier to organize, but that fact alone does not prove isolation. Kubernetes admission controls, network policy, identities, and storage permissions each enforce a different part of the boundary.&lt;/p&gt;

&lt;h2 id=&quot;define-the-workloads-permitted-behavior&quot;&gt;Define the workload’s permitted behavior&lt;/h2&gt;

&lt;p&gt;Record whether the application needs privilege, host access, outbound network destinations, secrets, and persistent volumes. Start with actual requirements rather than cloning the permissions of an older deployment.&lt;/p&gt;

&lt;p&gt;Kubernetes documents the Privileged, Baseline, and Restricted &lt;a href=&quot;https://kubernetes.io/docs/concepts/security/pod-security-standards/&quot;&gt;Pod Security Standards&lt;/a&gt;. Choose and enforce an appropriate policy for the target release, with reviewed exceptions for infrastructure components that need additional privileges.&lt;/p&gt;

&lt;h2 id=&quot;inspect-the-deployed-objects&quot;&gt;Inspect the deployed objects&lt;/h2&gt;

&lt;p&gt;Using an authorized read-only Kubernetes context:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;kubectl config current-context
kubectl get namespace example-app &lt;span class=&quot;nt&quot;&gt;--show-labels&lt;/span&gt;
kubectl get networkpolicy &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; example-app
kubectl get serviceaccount &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; example-app
kubectl get rolebinding &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; example-app
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;example-app&lt;/code&gt; is a placeholder namespace. Confirm the context before every administrative operation. These commands reveal configuration; they do not prove enforcement or effective authorization across all bindings.&lt;/p&gt;

&lt;h2 id=&quot;validate-network-policy-behavior&quot;&gt;Validate network policy behavior&lt;/h2&gt;

&lt;p&gt;The Kubernetes &lt;a href=&quot;https://kubernetes.io/docs/concepts/services-networking/network-policies/&quot;&gt;NetworkPolicy documentation&lt;/a&gt; explains that enforcement depends on a supporting network implementation. The existence of a policy object is insufficient evidence.&lt;/p&gt;

&lt;p&gt;In a lab namespace, run a permitted client and an unrelated client. Test the intended application port, required DNS resolution, and an outbound destination that should be denied. Include both ingress and egress expectations. Keep the tests harmless and restricted to the lab.&lt;/p&gt;

&lt;p&gt;Do not assume policies are an ordered firewall rule list. Review their combination and selection semantics for the deployed implementation. A broad additional allow can change the intended boundary.&lt;/p&gt;

&lt;h2 id=&quot;include-the-storage-path&quot;&gt;Include the storage path&lt;/h2&gt;

&lt;p&gt;A pod with an authorized mount may still receive excessive filesystem permissions or access a shared dataset intended for another workload. For WEKA, PowerScale, VAST, or other CSI-backed storage, examine the storage-side identity and export/share policy as well as the Kubernetes objects.&lt;/p&gt;

&lt;p&gt;Test a synthetic file through the application’s actual container identity. Confirm both required access and denied access outside its scope. An admission policy cannot repair an overly broad storage authorization model by itself.&lt;/p&gt;

&lt;h2 id=&quot;exercise-rejected-workloads&quot;&gt;Exercise rejected workloads&lt;/h2&gt;

&lt;p&gt;Use a disposable manifest that requests a prohibited capability and confirm admission rejects it. Then deploy a compliant workload and run its business transaction. Both results are necessary: a policy that rejects all work is not an acceptable application platform.&lt;/p&gt;

&lt;h2 id=&quot;keep-the-evidence-tied-to-the-release&quot;&gt;Keep the evidence tied to the release&lt;/h2&gt;

&lt;p&gt;Record Kubernetes and network-plugin versions, policy objects, test identities, and expected-versus-observed results. Rerun after cluster upgrades and CNI or CSI changes. The accepted boundary belongs to the tested system, not to a YAML file considered in isolation.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Reading Slurm accounting without fooling yourself about efficiency</title>
    <link href="https://www.wirewalk.com/writing/slurm-efficiency-accounting-traps/"/>
    <published>2026-05-27T10:00:00-04:00</published>
    <updated>2026-05-27T10:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/slurm-efficiency-accounting-traps/</id>
    <summary>How to compute CPU and memory efficiency from Slurm&apos;s accounting database so the number survives a sceptical reader, which two field choices produce confidently wrong answers that nobody catches, and why a low fleet figure is almost always a request-shape problem rather than a user-behavior problem.</summary>
    <content type="html">&lt;p&gt;Someone asks how well the cluster is being used. It is a reasonable question and there is a
database full of the answer, so you write a query. Twenty minutes later you have a number, and
the number is wrong in a way that will not announce itself. Both of the common mistakes produce
output that looks like a measurement: one gives you a beautiful figure, one gives you an alarming
figure, and neither has anything to do with how busy the cores were.&lt;/p&gt;

&lt;p&gt;This is about getting a defensible number out of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt;, knowing which fields lie to you, and
knowing what the number means once you have it — because the usual interpretation of a low
efficiency figure is also wrong.&lt;/p&gt;

&lt;h2 id=&quot;what-cpu-efficiency-actually-is&quot;&gt;What CPU efficiency actually is&lt;/h2&gt;

&lt;p&gt;CPU efficiency is consumed CPU time divided by allocated CPU time:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;efficiency = TotalCPU / (NCPUS * Elapsed)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The denominator is what the scheduler took off the table. If a job holds eight cores for four
hours, nobody else can have those eight cores for those four hours, whether the job uses them or
not. The numerator is what the job’s processes actually burned, user time plus system time, as
measured by the accounting plugin.&lt;/p&gt;

&lt;p&gt;That is exactly what &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; computes. From the shipped Perl (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;contribs/seff/seff.pl&lt;/code&gt;):&lt;/p&gt;

&lt;div class=&quot;language-perl highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;my&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$corewalltime&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$walltime&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$ncpus&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$corewalltime&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;nv&quot;&gt;$cpu_eff&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$cput&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$corewalltime&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;100&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;So for a single job you can just run &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; and stop reading:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ seff 918442
Job ID: 918442
Cluster: (redacted)
User/Group: user01/user01
State: COMPLETED (exit code 0)
Nodes: 1
Cores per node: 8
CPU Utilized: 04:14:07
CPU Efficiency: 12.55% of 1-09:44:40 core-walltime
Job Wall-clock time: 04:13:05
Memory Utilized: 5.47 GB
Memory Efficiency: 8.55% of 64.00 GB
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; does not scale. It is one job per invocation and it goes through the Slurm Perl API each
time. For a fleet report over a month of jobs you need &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt;, and that is where the trouble
starts.&lt;/p&gt;

&lt;h2 id=&quot;trap-one-cputimeraw-is-allocation-not-consumption&quot;&gt;Trap one: CPUTimeRAW is allocation, not consumption&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt; offers a field called &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt;. The name reads like “the raw CPU time this job used”.
It is not. The manual is unambiguous:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;CPUTimeRAW&lt;/strong&gt; — Time used (Elapsed time * CPU count) by a job or step in cpu-seconds.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the denominator. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTime&lt;/code&gt; is the same quantity in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;HH:MM:SS&lt;/code&gt;. It is allocated core-time,
computed arithmetically from the allocation, and it has no idea whether the cores were doing
anything. You can verify this in one line — the identity holds exactly, not approximately:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;sacct &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; 918442 &lt;span class=&quot;nt&quot;&gt;-X&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; JobID,NCPUS,ElapsedRaw,CPUTimeRAW &lt;span class=&quot;nt&quot;&gt;--parsable2&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--noheader&lt;/span&gt;
918442|8|15185|121480

&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;python3 &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;print(8*15185)&quot;&lt;/span&gt;
121480
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The failure mode is that someone builds the ratio out of the field whose name sounded right:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# WRONG — this is the definition of 1.0&lt;/span&gt;
sacct &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-S&lt;/span&gt; 2026-04-01 &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; 2026-05-01 &lt;span class=&quot;nt&quot;&gt;-X&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--parsable2&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--noheader&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
      &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; CPUTimeRAW,NCPUS,ElapsedRaw &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;|&apos;&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;{u+=$1; a+=$2*$3} END {printf &quot;efficiency %.1f%%\n&quot;, 100*u/a}&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;efficiency 100.0%
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It is not 100.0% because the cluster is perfectly used. It is 100.0% because you divided a
quantity by itself. Every job contributes exactly its own allocation to both sides, so the answer
is identically one for any job mix, any time window, any partition. The number is not even
slightly sensitive to reality: drain half the nodes, let jobs idle for days, and it still reads
100.0%.&lt;/p&gt;

&lt;p&gt;The reason this survives review is that 100% is not obviously absurd to a non-specialist. On a
busy cluster with a deep queue, “we are at 100%” reads as a statement about the queue rather than
about the cores. It gets into a slide. The give-away is that it is &lt;em&gt;exactly&lt;/em&gt; 100.0% and stays
exactly 100.0% when you change the window — a real measurement never does that.&lt;/p&gt;

&lt;h2 id=&quot;trap-two-the-allocation-only-flag-zeroes-the-numerator&quot;&gt;Trap two: the allocation-only flag zeroes the numerator&lt;/h2&gt;

&lt;p&gt;The other field you need is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt;:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;TotalCPU&lt;/strong&gt; — The sum of the SystemCPU and UserCPU time used by the job or job step.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the right numerator. But CPU time is gathered per &lt;em&gt;step&lt;/em&gt;, not per allocation. A batch job
produces several accounting records: the job allocation record (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918442&lt;/code&gt;), the batch step
(&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918442.batch&lt;/code&gt;), the external step (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918442.extern&lt;/code&gt;) that holds anything adopted into the job’s
cgroup, and one record per &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun&lt;/code&gt; (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918442.0&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918442.1&lt;/code&gt; …). The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--allocations&lt;/code&gt; flag says:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Only show statistics relevant to the job allocation itself, not taking steps into consideration.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Which means the step records are not fetched at all, and the step-derived accounting fields are
not populated from them. Everyone knows this about &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt;, because a blank column is visible.
Rather fewer people notice it about &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt;, because it comes back as a plausible-looking time
string rather than a blank:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;sacct &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; 918442 &lt;span class=&quot;nt&quot;&gt;-X&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; JobID,NCPUS,ElapsedRaw,TotalCPU
JobID          NCPUS ElapsedRaw   TotalCPU
&lt;span class=&quot;nt&quot;&gt;------------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt;
918442             8      15185   00:00:00
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;sacct &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; 918442 &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; JobID,NCPUS,ElapsedRaw,TotalCPU
JobID          NCPUS ElapsedRaw   TotalCPU
&lt;span class=&quot;nt&quot;&gt;------------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt;
918442             8      15185  04:14:07
918442.batch       8      15185  04:14:07
918442.extern      8      15185   00:00.001
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Same job, same database, two different answers, and the difference is one flag. The job row’s
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;04:14:07&lt;/code&gt; in the second listing is assembled from the step rows that the same query returned;
with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt; there are no step rows to assemble it from. Build a fleet report on the first form and
every ratio is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0 / something&lt;/code&gt;, so the cluster reads 0.0% efficient.&lt;/p&gt;

&lt;p&gt;That ought to be caught — and usually is, if the whole report is built that way. It is &lt;em&gt;not&lt;/em&gt;
caught when the query includes steps for some jobs and not others, or when someone filters the
zeroes out as “bad records” before averaging. Drop the zero rows and you are left with exactly
the jobs that ran many &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun&lt;/code&gt; steps, which are the well-parallelized ones, and your cluster
suddenly reports 70%.&lt;/p&gt;

&lt;p&gt;Whether the job-level record carries a rolled-up &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; depends on your Slurm version, on
whether the query pulled the steps, and on how the job was launched, so do not take my word for
it. Check on your own installation before you trust any script (this quick check splits on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;:&lt;/code&gt;
only, so run it on a job under a day of CPU time; the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tosec()&lt;/code&gt; function further down handles the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;DD-&lt;/code&gt; form):&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;sacct &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; 918442 &lt;span class=&quot;nt&quot;&gt;--parsable2&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--noheader&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; JobID,TotalCPU &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;|&apos;&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;{ n=split($2,p,&quot;:&quot;); s=(n==3)?p[1]*3600+p[2]*60+p[3]:p[1]*60+p[2];
                 if ($1 ~ /\./) step+=s; else job=s }
               END {printf &quot;job row: %.1fs   sum of steps: %.1fs\n&quot;, job, step}&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;job row: 15247.0s   sum of steps: 15247.0s
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If those two agree on your cluster, either source is fine. If the job row is zero, you must sum
the steps. Write the script so it works either way and you never have to care.&lt;/p&gt;

&lt;h2 id=&quot;trap-three-the-one-that-produces-a-believable-number&quot;&gt;Trap three: the one that produces a believable number&lt;/h2&gt;

&lt;p&gt;The first two traps produce 100.0% and 0.0%. Those are at least suspicious. The trap that
actually ships is the averaging.&lt;/p&gt;

&lt;p&gt;There are two things you might mean by “cluster efficiency”, and they are not close to each
other:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Job-mean&lt;/strong&gt;: compute efficiency per job, then take the arithmetic mean over jobs.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Core-weighted&lt;/strong&gt;: sum consumed CPU-seconds over the whole fleet, sum allocated CPU-seconds over
the whole fleet, divide once.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Take a constructed month with two populations, chosen to make the arithmetic visible:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;population&lt;/th&gt;
      &lt;th&gt;count&lt;/th&gt;
      &lt;th&gt;cores&lt;/th&gt;
      &lt;th&gt;elapsed&lt;/th&gt;
      &lt;th&gt;per-job eff&lt;/th&gt;
      &lt;th&gt;allocated core-h&lt;/th&gt;
      &lt;th&gt;consumed core-h&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;array tasks&lt;/td&gt;
      &lt;td&gt;12,000&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;5 min&lt;/td&gt;
      &lt;td&gt;95%&lt;/td&gt;
      &lt;td&gt;1,000&lt;/td&gt;
      &lt;td&gt;950&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;wide jobs&lt;/td&gt;
      &lt;td&gt;180&lt;/td&gt;
      &lt;td&gt;128&lt;/td&gt;
      &lt;td&gt;12 h&lt;/td&gt;
      &lt;td&gt;18%&lt;/td&gt;
      &lt;td&gt;276,480&lt;/td&gt;
      &lt;td&gt;49,766&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Job-mean: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(12000 × 0.95 + 180 × 0.18) / 12180 = 93.9%&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Core-weighted: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(950 + 49,766) / (1,000 + 276,480) = 18.3%&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Same data. 93.9% or 18.3%, depending on a choice most people do not realize they are making. The
job-mean is dominated by twelve thousand five-minute array tasks that together account for 0.4%
of the machine. The core-weighted figure is the one that answers “how much of the cluster’s
capacity turned into computation”, which is the question that was asked.&lt;/p&gt;

&lt;p&gt;Use core-weighted for any capacity or spend conversation. Job-mean is useful for exactly one
thing: finding which &lt;em&gt;users&lt;/em&gt; have a habit of over-requesting, where you want each job to count
once regardless of size. Label whichever you publish.&lt;/p&gt;

&lt;h2 id=&quot;the-working-query&quot;&gt;The working query&lt;/h2&gt;

&lt;p&gt;Query without &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt;, take the denominator from the job row and the numerator from the step rows,
and parse the time formats properly. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; is printed as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[DD-[HH:]]MM:SS[.mmm]&lt;/code&gt;, so you will
see &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;00:00.001&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;2-03:14:22&lt;/code&gt; in the same column; a naive &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;HH:MM:SS&lt;/code&gt; split gets both wrong.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;sacct &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-S&lt;/span&gt; 2026-04-01T00:00:00 &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; 2026-04-30T23:59:59 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
      &lt;span class=&quot;nt&quot;&gt;--parsable2&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--noheader&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--delimiter&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;|&apos;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
      &lt;span class=&quot;nt&quot;&gt;--format&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;JobID,User,Account,Partition,State,NCPUS,ElapsedRaw,CPUTimeRAW,TotalCPU,ReqMem,MaxRSS,NNodes &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; /tmp/acct-apr.psv
&lt;span class=&quot;nb&quot;&gt;wc&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; /tmp/acct-apr.psv
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;2841577 /tmp/acct-apr.psv
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two and a half million rows is a few hundred megabytes of text, so put it somewhere with room —
filling the root filesystem of a login or head node is an unpleasant way to learn that a month of
accounting is bigger than it looks.&lt;/p&gt;

&lt;p&gt;Then aggregate:&lt;/p&gt;

&lt;div class=&quot;language-awk highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;#!/usr/bin/awk -f&lt;/span&gt;
&lt;span class=&quot;c1&quot;&gt;# slurm-eff.awk — core-weighted CPU efficiency by partition&lt;/span&gt;
&lt;span class=&quot;kr&quot;&gt;BEGIN&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;kc&quot;&gt;FS&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;|&quot;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;

&lt;span class=&quot;kd&quot;&gt;function&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;tosec&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;t&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;   &lt;span class=&quot;nx&quot;&gt;d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;s&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;nx&quot;&gt;d&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;t&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;~&lt;/span&gt; &lt;span class=&quot;sr&quot;&gt;/-/&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;t&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;-&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;d&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;t&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
    &lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;t&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;:&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt;      &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;s&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3600&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;60&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;s&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;60&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt;             &lt;span class=&quot;nx&quot;&gt;s&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;d&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;86400&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;s&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;

&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;~&lt;/span&gt; &lt;span class=&quot;sr&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\.&lt;/span&gt;&lt;span class=&quot;sr&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;                       &lt;span class=&quot;c1&quot;&gt;# step record: the numerator lives here&lt;/span&gt;
    &lt;span class=&quot;nb&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;.&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;nx&quot;&gt;consumed&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;tosec&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$9&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;next&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;                                 &lt;span class=&quot;c1&quot;&gt;# job record: the denominator lives here&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$5&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!~&lt;/span&gt; &lt;span class=&quot;sr&quot;&gt;/^&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sr&quot;&gt;COMPLETED|FAILED|TIMEOUT|OUT_OF_MEMORY|NODE_FAIL&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;sr&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;next&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$7&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;120&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;next&lt;/span&gt;        &lt;span class=&quot;c1&quot;&gt;# sub-2-minute jobs are startup, not work&lt;/span&gt;
    &lt;span class=&quot;nx&quot;&gt;alloc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$8&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;            &lt;span class=&quot;c1&quot;&gt;# CPUTimeRAW == NCPUS * ElapsedRaw&lt;/span&gt;
    &lt;span class=&quot;nx&quot;&gt;part&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;  &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$4&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;kr&quot;&gt;END&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;alloc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
        &lt;span class=&quot;nx&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;part&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;alloc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
        &lt;span class=&quot;nx&quot;&gt;C&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;part&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;consumed&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
        &lt;span class=&quot;nx&quot;&gt;TA&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;alloc&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;];&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;TC&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;consumed&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;%-12s %14s %14s %8s\n&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;partition&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;alloc_core_h&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;used_core_h&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;eff&quot;&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;%-12s %14.1f %14.1f %7.1f%%\n&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3600&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;C&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3600&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;100&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;C&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;%-12s %14.1f %14.1f %7.1f%%\n&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;TOTAL&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;TA&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3600&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;TC&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3600&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;100&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;TC&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;TA&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; slurm-eff.awk /tmp/acct-apr.psv
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Illustrative output — these figures are constructed to show the shape of the result, not a
measurement of any particular machine. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;for (p in A)&lt;/code&gt; iterates in an unspecified order, so the
partition block comes out shuffled; sort it yourself if the order matters, but do not pipe the
whole output through &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sort&lt;/code&gt; or the header and the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TOTAL&lt;/code&gt; line get sorted along with the data:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;partition      alloc_core_h    used_core_h      eff
normal             276480.4        50716.3    18.3%
long               188411.2        61533.9    32.7%
gpu                 41203.6         9877.1    24.0%
short                1000.2          950.1    95.0%
TOTAL              507095.4       123077.4    24.3%
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three details in that script are load-bearing.&lt;/p&gt;

&lt;p&gt;The state filter excludes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;RUNNING&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PENDING&lt;/code&gt;. A running job has a growing denominator and a
numerator that is only as fresh as the last accounting poll, so it drags the average down for no
reason. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CANCELLED&lt;/code&gt; is excluded too, though that one is a judgment call: a job canceled at 90%
of its walltime really did consume the allocation, and if your users cancel a lot you should
count it. Say which you did.&lt;/p&gt;

&lt;p&gt;The 120-second floor removes the long tail of jobs that failed at startup. On a cluster with heavy
array use these can be the majority of records by count and a rounding error by core-hours, so the
floor barely moves the core-weighted number — but it removes thousands of 0% rows that would wreck
a job-mean.&lt;/p&gt;

&lt;p&gt;The step loop sums &lt;em&gt;all&lt;/em&gt; step records including &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.extern&lt;/code&gt;. That is deliberate, and it is the next
trap.&lt;/p&gt;

&lt;h2 id=&quot;where-the-obvious-reading-is-wrong-extern-and-ssh-launched-ranks&quot;&gt;Where the obvious reading is wrong: extern and SSH-launched ranks&lt;/h2&gt;

&lt;p&gt;The tidy instinct is to drop &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.extern&lt;/code&gt; from the numerator. It is the container step for the job’s
cgroup, it normally shows a millisecond or two of CPU, and it feels like noise.&lt;/p&gt;

&lt;p&gt;On a site running &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pam_slurm_adopt&lt;/code&gt;, it is not always noise. The thing to understand is that
“&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpirun&lt;/code&gt; instead of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun&lt;/code&gt;” is not by itself the problem — the question is whether the launcher
bootstraps through Slurm or through SSH. Open MPI’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpirun&lt;/code&gt;, inside an allocation, uses &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;srun&lt;/code&gt; to
start its daemons, so a step record does exist. Intel MPI exported with
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;I_MPI_HYDRA_BOOTSTRAP=ssh&lt;/code&gt;, Open MPI forced onto the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rsh&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ssh&lt;/code&gt; launcher or built without Slurm
support, and hand-rolled &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ssh&lt;/code&gt;-in-a-loop launchers all start the remote processes over SSH, and
none of those produce a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918442.0&lt;/code&gt; on the remote nodes.&lt;/p&gt;

&lt;p&gt;What there is on those nodes, if adoption is configured, is the external step — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pam_slurm_adopt&lt;/code&gt;
puts the incoming SSH session into the job’s extern cgroup, and the adopted processes’ CPU time is
then accounted against &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.extern&lt;/code&gt; rather than against any launch step.&lt;/p&gt;

&lt;p&gt;Drop &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.extern&lt;/code&gt; and every one of those jobs reports roughly &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;1/N&lt;/code&gt; efficiency, where N is the node
count, because you counted only the rank on the batch node. A 16-node job that is running
perfectly reads 6%. That is a very convincing-looking finding. It is also completely false, and it
points the investigation at the users who are doing the most demanding work.&lt;/p&gt;

&lt;p&gt;The check takes one run:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; slurm-eff.awk /tmp/acct-apr.psv | &lt;span class=&quot;nb&quot;&gt;tail&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-1&lt;/span&gt;
TOTAL              507095.4       123077.4    24.3%

&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;\.extern|&apos;&lt;/span&gt; /tmp/acct-apr.psv | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; slurm-eff.awk | &lt;span class=&quot;nb&quot;&gt;tail&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-1&lt;/span&gt;
TOTAL              507095.4        94112.8    18.6%
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A gap that size between the two runs tells you a meaningful fraction of your CPU time is being
recorded against the external step, which tells you that work is arriving on compute nodes by some
route other than a Slurm launch step. That is worth knowing on its own. If the two runs agree to
within a percent, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.extern&lt;/code&gt; genuinely is noise on your cluster and either choice is fine.&lt;/p&gt;

&lt;p&gt;There is a worse version of this. If adoption is &lt;em&gt;not&lt;/em&gt; configured, the remote ranks belong to no
step and no cgroup, and their CPU time is not recorded anywhere at all. No amount of careful
querying recovers it. The symptom is an MPI-heavy partition that reports implausibly low
efficiency while the nodes are visibly hot; confirm by watching &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ps&lt;/code&gt; on a compute node during one
of those jobs, or by comparing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt; against a node-level metric such as a CPU-utilization
exporter. If that is your situation, the accounting database cannot answer the efficiency question
for those jobs and you should say so rather than publish the number.&lt;/p&gt;

&lt;h2 id=&quot;the-memory-half&quot;&gt;The memory half&lt;/h2&gt;

&lt;p&gt;CPU efficiency alone will mislead you about waste. On most general-purpose clusters memory is the
binding constraint more often than cores, and a job that holds a node’s entire RAM while using
6 GB of it is wasting the machine even at 100% CPU.&lt;/p&gt;

&lt;p&gt;The fields:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;sacct &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; 918442 &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; JobID,ReqMem,MaxRSS,MaxRSSNode,MaxRSSTask,TRESUsageInTot%40 &lt;span class=&quot;nt&quot;&gt;--units&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;G
JobID          ReqMem     MaxRSS MaxRSSNode MaxRSSTask                        TRESUsageInTot
&lt;span class=&quot;nt&quot;&gt;------------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-------------------------------------&lt;/span&gt;
918442            64G
918442.batch            5.47G      nodeA             0   &lt;span class=&quot;nv&quot;&gt;cpu&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;04:14:07,energy&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;0,fs/disk&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;12.1G,
                                                          &lt;span class=&quot;nv&quot;&gt;mem&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;5.47G,pages&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;0,vmem&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;6.02G
918442.extern              0       nodeA             0   &lt;span class=&quot;nv&quot;&gt;cpu&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;00:00:00,mem&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;0,pages&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;0,vmem&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;(The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TRESUsageInTot&lt;/code&gt; column is wrapped here to fit the page; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt; truncates it to the width
you asked for instead. Exact unit rendering — trailing zeroes, which fields &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--units&lt;/code&gt; reaches —
moves about between versions, so match your own output rather than this one.)&lt;/p&gt;

&lt;p&gt;Two things to fix in your head.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ReqMem&lt;/code&gt; changed meaning in 21.08.&lt;/strong&gt; The release notes are explicit: it now “shows the requested
memory of the whole job with a letter appended indicating units”, and it is only displayed for the
job record, not the steps. Before that it carried a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Mn&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Mc&lt;/code&gt; suffix meaning per-node or
per-core, and you had to multiply by node count or core count yourself to get the job total. A
script written against the old format and run against a new cluster — or the reverse — is wrong by
a factor equal to the core count, which on a 128-core node is not subtle. Check your version
before you write the parser:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;sacct &lt;span class=&quot;nt&quot;&gt;-V&lt;/span&gt;
slurm 24.05.7
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; is a maximum over tasks, not a sum.&lt;/strong&gt; The existence of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSSTask&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSSNode&lt;/code&gt;
tells you this directly: they name &lt;em&gt;which&lt;/em&gt; task and &lt;em&gt;which&lt;/em&gt; node hit the peak. For a single-task
job, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; is the job’s memory. For a 128-rank MPI job where every rank uses 3 GB, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt;
reports 3 GB while the job actually held 384 GB. Divide that by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ReqMem&lt;/code&gt; and you will conclude the
job over-requested by a factor of 128 and go and tell the user so.&lt;/p&gt;

&lt;p&gt;The field to reach for instead is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mem=&lt;/code&gt; inside &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TRESUsageInTot&lt;/code&gt;, which the manual defines as
“Tres total usage in by all tasks in job” — the per-task peaks added up. Note what that is and is
not: because the peaks need not have happened at the same moment, the sum is an upper bound on
what the job held simultaneously, not a measurement of it. It is still much closer to the truth
than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; for a multi-task job.&lt;/p&gt;

&lt;p&gt;Do not assume &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; settles this for you. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; does not read &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tres_usage_in_tot&lt;/code&gt; at all: it
takes the per-task maximum from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tres_usage_in_max&lt;/code&gt;, keeps the largest one across the steps, and
multiplies by that step’s task count. On a job whose ranks are all doing the same thing that is a
fair estimate; on one where a single rank holds most of the memory it can be far out in either
direction. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; is honest about this — for a multi-task step it labels the line
“(estimated maximum)”. So &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt;’s “Memory Utilized”, a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; query and a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TRESUsageInTot&lt;/code&gt;
query can produce three different numbers for the same job, and none of them is a true
simultaneous job-wide peak, because Slurm does not record one.&lt;/p&gt;

&lt;p&gt;One more limit worth stating plainly: memory accounting is, historically, sampled.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobAcctGatherFrequency&lt;/code&gt; defaults to 30 seconds for the task datatype, and a job that allocates
200 GB for four seconds and frees it can poll clean and show a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; of 8 GB.&lt;/p&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cgroup/v2&lt;/code&gt; plugin improves on that by taking &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; from the kernel’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;memory.peak&lt;/code&gt;, which
is a high-water mark rather than a sample. Two things then follow, and they pull in opposite
directions. First, filesystem-backed memory counts: the page cache a job generates by reading and
writing files is included in the reported RSS unless you set &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobAcctGatherParams=no_file_cache&lt;/code&gt;,
so an I/O-heavy job can look memory-hungry when it is not. Second, the manual is explicit that
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;no_file_cache&lt;/code&gt; “disables the use of the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;memory.peak&lt;/code&gt; interface, which can result in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt;
failing to record short memory spikes” — the cleaner number costs you the high-water mark. Slurm
24.11 also lists a fix for jobs that run shorter than two gather intervals, which previously could
report nothing at all.&lt;/p&gt;

&lt;p&gt;So there is no single answer to “is my &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; trustworthy”: it depends on your
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobAcctGatherType&lt;/code&gt;, your cgroup version, your &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobAcctGatherParams&lt;/code&gt; and your Slurm version. Check
which path your cluster is on before you tell users to cut their memory requests, and leave
headroom either way. A job killed by the OOM handler costs more than the memory it was holding.&lt;/p&gt;

&lt;h2 id=&quot;array-jobs&quot;&gt;Array jobs&lt;/h2&gt;

&lt;p&gt;Arrays need care for three separate reasons.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;sacct &lt;span class=&quot;nt&quot;&gt;-j&lt;/span&gt; 918500 &lt;span class=&quot;nt&quot;&gt;-X&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; JobID,JobIDRaw,State,NCPUS,ElapsedRaw
JobID            JobIDRaw      State   NCPUS ElapsedRaw
&lt;span class=&quot;nt&quot;&gt;---------------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-------&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;----------&lt;/span&gt;
918500_[4-500]      918500    PENDING       1          0
918500_1            918501  COMPLETED       1        363
918500_2            918502  COMPLETED       1        358
918500_3            918503  COMPLETED       1        371
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;First, the pending master record &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918500_[4-500]&lt;/code&gt; is a real row with a real &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobID&lt;/code&gt;. It has zero
elapsed time so it contributes nothing to a core-weighted sum, but it is one row out of many in a
job-mean, and if you are counting jobs it inflates your count by one per array rather than one per
task. Filter on state.&lt;/p&gt;

&lt;p&gt;Second, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobID&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobIDRaw&lt;/code&gt; differ for array tasks. Step records are keyed on the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobID&lt;/code&gt; form
(&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;918500_2.batch&lt;/code&gt;), so join on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobID&lt;/code&gt;, not &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobIDRaw&lt;/code&gt;. Mixing the two silently drops the
numerator for every array task. The raw ids are handed out as tasks are scheduled, so do not
assume they are contiguous, ordered, or related to the task index in any useful way.&lt;/p&gt;

&lt;p&gt;Third, and most important: arrays are how one user generates a hundred thousand accounting rows.
Any per-job average is now a report on that one user’s array. This is the single biggest reason
to publish the core-weighted figure.&lt;/p&gt;

&lt;h2 id=&quot;the-traps-and-the-answer-each-one-produces&quot;&gt;The traps, and the answer each one produces&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Trap&lt;/th&gt;
      &lt;th&gt;Wrong answer it produces&lt;/th&gt;
      &lt;th&gt;How to spot it&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; used as consumed CPU&lt;/td&gt;
      &lt;td&gt;Exactly 100.0%, always&lt;/td&gt;
      &lt;td&gt;Figure does not move when you change the window&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt; with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt;, where the job row is not rolled up&lt;/td&gt;
      &lt;td&gt;Exactly 0.0%&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; column reads &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;00:00:00&lt;/code&gt; for every job&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Zero rows filtered out as “bad data” after using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;60–80%, entirely fictional&lt;/td&gt;
      &lt;td&gt;Row count after filter is a small fraction of jobs&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Summing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; over job &lt;em&gt;and&lt;/em&gt; step rows&lt;/td&gt;
      &lt;td&gt;Roughly double; some jobs &amp;gt;100%&lt;/td&gt;
      &lt;td&gt;Any per-job efficiency above 100%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Step &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCPUS&lt;/code&gt; used in the denominator&lt;/td&gt;
      &lt;td&gt;Inflated; varies by launch style&lt;/td&gt;
      &lt;td&gt;Denominator smaller than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; on the job row&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Mean of per-job ratios&lt;/td&gt;
      &lt;td&gt;90%+ on an array-heavy cluster&lt;/td&gt;
      &lt;td&gt;Job count in the thousands, core-hours dominated by tens of jobs&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.extern&lt;/code&gt; dropped where &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pam_slurm_adopt&lt;/code&gt; adopts SSH-launched ranks&lt;/td&gt;
      &lt;td&gt;MPI partitions read at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;1/N&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Re-run including &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;.extern&lt;/code&gt; and compare&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;No adoption configured, ranks launched over SSH&lt;/td&gt;
      &lt;td&gt;Unrecoverable under-count&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt; disagrees with node-level CPU metrics&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;RUNNING&lt;/code&gt; jobs included&lt;/td&gt;
      &lt;td&gt;Understated, drifts by time of day&lt;/td&gt;
      &lt;td&gt;Efficiency changes when you re-run the same window&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Jobs straddling the window counted whole&lt;/td&gt;
      &lt;td&gt;Monthly totals exceed the machine’s capacity&lt;/td&gt;
      &lt;td&gt;Sum of months ≠ the year&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-T&lt;/code&gt; used to fix that&lt;/td&gt;
      &lt;td&gt;Some jobs above 100%&lt;/td&gt;
      &lt;td&gt;Truncates elapsed, does not prorate &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; used as job memory&lt;/td&gt;
      &lt;td&gt;Understated by the task count&lt;/td&gt;
      &lt;td&gt;Multi-node jobs all read ~1% memory efficiency&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; memory quoted as measured&lt;/td&gt;
      &lt;td&gt;An estimate: per-task max × task count&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; prints “(estimated maximum)” on multi-task steps&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ReqMem&lt;/code&gt; parsed without checking version&lt;/td&gt;
      &lt;td&gt;Off by core count or node count&lt;/td&gt;
      &lt;td&gt;Memory efficiency above 100% or absurdly below&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SMT threads counted as cores&lt;/td&gt;
      &lt;td&gt;Hard ceiling at 50%&lt;/td&gt;
      &lt;td&gt;No job in the fleet exceeds 50%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt; efficiency compared to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sreport&lt;/code&gt; “Used”&lt;/td&gt;
      &lt;td&gt;The two never agree&lt;/td&gt;
      &lt;td&gt;They are different quantities — see below&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;On that last row: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sreport&lt;/code&gt;’s usage columns are built from what jobs were &lt;em&gt;charged&lt;/em&gt;, which is
allocation-derived — the same family of quantity as the usage that feeds fairshare, though
fairshare additionally applies &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TRESBillingWeights&lt;/code&gt; and a decay half-life, so the two are not
interchangeable either. What matters here is that none of it is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt;. So &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sreport&lt;/code&gt; and
a correct &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt; efficiency query are not meant to agree, and the ratio between them is roughly
your efficiency figure. You can confirm the relationship on your own cluster by checking that
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sreport&lt;/code&gt; totals track your summed &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; rather than your summed &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; — if they do,
the model above is right for your version.&lt;/p&gt;

&lt;h2 id=&quot;field-reference&quot;&gt;Field reference&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Field&lt;/th&gt;
      &lt;th&gt;What it is&lt;/th&gt;
      &lt;th&gt;Which record carries it&lt;/th&gt;
      &lt;th&gt;Gotcha&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCPUS&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AllocCPUS&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Allocated CPU count&lt;/td&gt;
      &lt;td&gt;Job and step&lt;/td&gt;
      &lt;td&gt;Step value can be smaller than the job’s&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ElapsedRaw&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Wall seconds&lt;/td&gt;
      &lt;td&gt;Job and step&lt;/td&gt;
      &lt;td&gt;Step elapsed ≠ job elapsed&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NCPUS × ElapsedRaw&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Job and step&lt;/td&gt;
      &lt;td&gt;Allocation, never consumption&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UserCPU + SystemCPU&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Step; job row varies by version&lt;/td&gt;
      &lt;td&gt;Format is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[DD-[HH:]]MM:SS[.mmm]&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;UserCPU&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SystemCPU&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;The two halves&lt;/td&gt;
      &lt;td&gt;Step&lt;/td&gt;
      &lt;td&gt;High system time can mask an idle job&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ReqMem&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Requested memory&lt;/td&gt;
      &lt;td&gt;Job only, since 21.08&lt;/td&gt;
      &lt;td&gt;Per-node/per-core suffix before 21.08&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Peak RSS of the largest single task&lt;/td&gt;
      &lt;td&gt;Step&lt;/td&gt;
      &lt;td&gt;Not a sum across tasks&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSSTask&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSSNode&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Which task and node peaked&lt;/td&gt;
      &lt;td&gt;Step&lt;/td&gt;
      &lt;td&gt;Their existence proves &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; is a max&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TRESUsageInTot&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Per-TRES totals, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mem=&lt;/code&gt; summed over tasks&lt;/td&gt;
      &lt;td&gt;Step&lt;/td&gt;
      &lt;td&gt;Sum of per-task peaks: an upper bound, not a simultaneous peak&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TRESUsageInMax&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Per-TRES per-task maxima&lt;/td&gt;
      &lt;td&gt;Step&lt;/td&gt;
      &lt;td&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; number, in TRES form; what &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;seff&lt;/code&gt; reads&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AllocTRES&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Full allocation including GRES&lt;/td&gt;
      &lt;td&gt;Job&lt;/td&gt;
      &lt;td&gt;Where GPU counts live&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;State&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Final state&lt;/td&gt;
      &lt;td&gt;Job and step&lt;/td&gt;
      &lt;td&gt;Filter before aggregating&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NNodes&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NTasks&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Shape of the job&lt;/td&gt;
      &lt;td&gt;Job / step&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NTasks&lt;/code&gt; blank on the job record&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobID&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobIDRaw&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Array-form and underlying id&lt;/td&gt;
      &lt;td&gt;Both&lt;/td&gt;
      &lt;td&gt;Join on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;JobID&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;h2 id=&quot;worked-example&quot;&gt;Worked example&lt;/h2&gt;

&lt;p&gt;Back to job 918442. Eight cores, 15,185 seconds elapsed, so 121,480 allocated core-seconds. The
steps sum to 15,247 seconds of CPU.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;efficiency = 15247 / 121480 = 12.55%
cores used = 15247 / 15185 = 1.004
allocated  = 8
memory     = 5.47 GiB of 64 GiB requested = 8.55%
waste      = (121480 - 15247) / 3600 = 29.5 core-hours parked
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The percentage is the least useful line there. “1.004 cores out of 8” tells you what happened:
this is a single-threaded program that asked for eight cores. It is not a program that is 12.55%
efficient at using eight cores — it is a program that is 100% efficient at using one core and was
given eight. Nothing the user can do to their code will move the 12.55%. Changing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--cpus-per-task&lt;/code&gt;
from 8 to 1 moves it to roughly 100%, instantly, and on this job — one thread, 5.47 GiB of a
64 GiB request — with no change to runtime. Check the memory before you send that advice, though:
where the site sets &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;DefMemPerCPU&lt;/code&gt;, cutting the core count cuts the job’s memory with it, and the
user who was quietly using cores as a memory allocation will get an OOM kill for following your
guidance. That is the second of the three rational reasons below, and it is a site configuration
problem, not a user problem.&lt;/p&gt;

&lt;p&gt;That reframing is the whole point. Report &lt;strong&gt;average cores used against cores allocated&lt;/strong&gt;, not a
percentage, and the diagnosis falls out of the number:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Cores used vs allocated&lt;/th&gt;
      &lt;th&gt;Reading&lt;/th&gt;
      &lt;th&gt;What to do&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;~1.0 of N, N &amp;gt; 1&lt;/td&gt;
      &lt;td&gt;Single-threaded job over-requesting&lt;/td&gt;
      &lt;td&gt;Fix the submission script&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;~N of N&lt;/td&gt;
      &lt;td&gt;Working as intended&lt;/td&gt;
      &lt;td&gt;Nothing&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;~0.5N of N&lt;/td&gt;
      &lt;td&gt;One thread per SMT core, or half the ranks idle&lt;/td&gt;
      &lt;td&gt;Check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ThreadsPerCore&lt;/code&gt; before blaming the user&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;≪1.0 of N&lt;/td&gt;
      &lt;td&gt;I/O-bound, or waiting on a license or a lock&lt;/td&gt;
      &lt;td&gt;Look at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SystemCPU&lt;/code&gt; and at the filesystem&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&amp;gt; N&lt;/td&gt;
      &lt;td&gt;Oversubscribed, or no cgroup pinning&lt;/td&gt;
      &lt;td&gt;Check &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TaskPlugin&lt;/code&gt; and the thread-library env&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;~N early, ~1 late&lt;/td&gt;
      &lt;td&gt;Serial post-processing tail&lt;/td&gt;
      &lt;td&gt;Split the job; the average hides it&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The SMT row is worth dwelling on. Depending on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SelectTypeParameters&lt;/code&gt; and the node’s
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ThreadsPerCore&lt;/code&gt;, allocating one core can charge you two CPUs, in which case a perfectly
saturating single-threaded task caps out at 50% and no job on the machine will ever read higher.
Before concluding your users waste half the cluster, check:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;scontrol show node nodeA | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;CPUAlloc|CPUTot|ThreadsPerCore|Boards&apos;&lt;/span&gt;
   &lt;span class=&quot;nv&quot;&gt;CPUAlloc&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;128 &lt;span class=&quot;nv&quot;&gt;CPUEfctv&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;128 &lt;span class=&quot;nv&quot;&gt;CPUTot&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;128
   &lt;span class=&quot;nv&quot;&gt;Boards&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1 &lt;span class=&quot;nv&quot;&gt;SocketsPerBoard&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;2 &lt;span class=&quot;nv&quot;&gt;CoresPerSocket&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;64 &lt;span class=&quot;nv&quot;&gt;ThreadsPerCore&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ThreadsPerCore=1&lt;/code&gt; there, so it is not the explanation on that node. If it reads 2, every
efficiency figure you have is measured against threads.&lt;/p&gt;

&lt;h2 id=&quot;what-a-realistic-fleet-number-looks-like&quot;&gt;What a realistic fleet number looks like&lt;/h2&gt;

&lt;p&gt;I am not aware of a systematic published survey of core-weighted CPU efficiency across academic
clusters, so treat any single quoted figure — including mine — as a prior rather than a
benchmark. The working range I would expect on a general-purpose, mixed-workload cluster is
roughly 35% to 65%. A number in the high thirties on a machine that serves a lot of bioinformatics
and statistics is unremarkable. Below about 25% something structural is usually going on: a
partition configured exclusive so every job charges a whole node, a default &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--cpus-per-task&lt;/code&gt;
that does not match the default thread count of the dominant application, or GPU jobs whose CPU
allocation is incidental.&lt;/p&gt;

&lt;p&gt;Two caveats on the number itself.&lt;/p&gt;

&lt;p&gt;GPU partitions should not be assessed on CPU efficiency at all. A job holding four GPUs and eight
cores to feed them is doing exactly the right thing and will read 10%. You need GPU utilization
from a device-level exporter; the accounting database records that GPUs were allocated, not that
they were busy.&lt;/p&gt;

&lt;p&gt;And the honest reading of a low figure is almost never “users are wasteful”. Sort the parked
core-hours by cause and, on the mixed-workload clusters I have looked at, most of them come from
single-threaded jobs asking for several cores. Users do this for three rational
reasons: they were told to by a tutorial written for a different scheduler, they are using cores
as a proxy for memory because the site’s memory-per-core default is low, or the application is
genuinely threaded but only for one phase of its run. The first is a documentation fix, the second
is a scheduler configuration fix, and only the third is really the user’s problem.&lt;/p&gt;

&lt;p&gt;So publish two columns, not one. Core-hours parked, and the count of jobs where cores-used is
below 1.5 while cores-allocated is above 2. The first is the size of the prize. The second is the
list of submission scripts to go and fix, and it is usually short — a handful of templates that
have been copied across a department.&lt;/p&gt;

&lt;p&gt;This one needs the same &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tosec()&lt;/code&gt; as before. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;awk&lt;/code&gt; accepts &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-f&lt;/code&gt; more than once, so keep that
function in a file of its own and prepend it rather than pasting it twice — an inline program that
calls &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tosec()&lt;/code&gt; without it is a fatal “calling undefined function”, not a silent zero.&lt;/p&gt;

&lt;div class=&quot;language-awk highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c1&quot;&gt;# parked.awk — worst over-requesting jobs by parked core-hours&lt;/span&gt;
&lt;span class=&quot;kr&quot;&gt;BEGIN&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;kc&quot;&gt;FS&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;|&quot;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;~&lt;/span&gt; &lt;span class=&quot;sr&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\.&lt;/span&gt;&lt;span class=&quot;sr&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;.&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;c&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+=&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;tosec&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$9&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;next&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;$5&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;COMPLETED&quot;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$7&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;600&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$6&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$7&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;u&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$2&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;kr&quot;&gt;END&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;c&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;mf&quot;&gt;1.5&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;printf&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;%-12s %-10s %3d cores  %.2f used  %8.1f core-h parked\n&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
                   &lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;u&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;c&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;*&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;nx&quot;&gt;c&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;nx&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3600&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; tosec.awk &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; parked.awk /tmp/acct-apr.psv | &lt;span class=&quot;nb&quot;&gt;sort&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-k7&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-gr&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;head&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Run that, and instead of a percentage you have twenty lines, each naming a submission script and
the number of core-hours it costs per month. That is a conversation you can actually have.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>OpenSCAP reports: preserve the exceptions behind the score</title>
    <link href="https://www.wirewalk.com/writing/openscap-hardening-evidence/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/openscap-hardening-evidence/</id>
    <summary>A scan result becomes useful evidence when the profile, content version, scope, and exceptions are retained.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A hardening score can improve because controls were implemented, because rules were excluded, or because the wrong profile was scanned. An assessment needs enough context to distinguish those outcomes.&lt;/p&gt;

&lt;h2 id=&quot;pin-the-benchmark-inputs&quot;&gt;Pin the benchmark inputs&lt;/h2&gt;

&lt;p&gt;Record the operating-system release, OpenSCAP version, content package version, selected profile, tailoring file, and target identity. Retain the original machine-readable result along with any HTML report.&lt;/p&gt;

&lt;p&gt;OpenSCAP documents configuration evaluation with XCCDF and OVAL and the use of profiles in its &lt;a href=&quot;https://www.open-scap.org/tools/openscap-base/&quot;&gt;tool overview&lt;/a&gt;. A tool’s ability to evaluate a benchmark does not certify the organization or establish that the benchmark fits the workload.&lt;/p&gt;

&lt;h2 id=&quot;inspect-before-evaluating&quot;&gt;Inspect before evaluating&lt;/h2&gt;

&lt;p&gt;On a lab host with OpenSCAP and the relevant content installed:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;oscap &lt;span class=&quot;nt&quot;&gt;-V&lt;/span&gt;
oscap info /usr/share/xml/scap/ssg/content/ssg-rhel9-ds.xml
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The content path is an example for an installed RHEL 9 SCAP Security Guide package. Use the actual available data stream and read its listed profiles. Do not paste a profile identifier from a different release and assume it selected the intended controls.&lt;/p&gt;

&lt;p&gt;Run a non-remediating evaluation first. Preserve both the process status and the report’s semantic results. Evaluation errors, not-applicable rules, and failed rules mean different things.&lt;/p&gt;

&lt;h2 id=&quot;review-exceptions-individually&quot;&gt;Review exceptions individually&lt;/h2&gt;

&lt;p&gt;For each excluded or accepted finding, record the rule, affected systems, operational reason, compensating control, owner, and review date. Avoid a blanket exception such as “HPC nodes need performance.” A specific setting may be necessary for one workload and unnecessary elsewhere.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Finding state&lt;/th&gt;
      &lt;th&gt;Required follow-up&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Failed and applicable&lt;/td&gt;
      &lt;td&gt;Remediate or approve a bounded exception&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Not applicable&lt;/td&gt;
      &lt;td&gt;Verify why it does not apply&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Evaluation error&lt;/td&gt;
      &lt;td&gt;Repair the assessment path&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Excluded by tailoring&lt;/td&gt;
      &lt;td&gt;Review the tailoring decision&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;A score without these states hides the information an auditor or operator needs most.&lt;/p&gt;

&lt;h2 id=&quot;test-remediation-as-a-service-change&quot;&gt;Test remediation as a service change&lt;/h2&gt;

&lt;p&gt;Do not enable automatic remediation across an estate as the first experiment. Review proposed changes, apply a small subset to a pilot, and test application behavior and recovery access.&lt;/p&gt;

&lt;p&gt;Some controls affect authentication, mounts, kernel behavior, or cryptography. Their operational consequences may not appear until a reboot or a new session. Include those lifecycle events in the pilot acceptance test.&lt;/p&gt;

&lt;h2 id=&quot;compare-like-with-like&quot;&gt;Compare like with like&lt;/h2&gt;

&lt;p&gt;When a later scan differs, compare content and profile versions before attributing the change to configuration drift. A revised benchmark can produce a different result on an unchanged host.&lt;/p&gt;

&lt;p&gt;Keep a small evidence bundle for each assessment: input hashes, target inventory, results, exceptions, remediation record, and post-change checks. That supports a repeatable engineering review. It should never be shortened to an unsupported claim that the whole estate is “compliant.”&lt;/p&gt;

&lt;p&gt;Related: &lt;a href=&quot;/writing/selinux-denials-without-disabling/&quot;&gt;SELinux diagnosis&lt;/a&gt; and &lt;a href=&quot;/writing/ansible-linux-patching-canaries/&quot;&gt;staged Linux maintenance&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The green check that hides a broken result</title>
    <link href="https://www.wirewalk.com/writing/the-green-check-that-hides-a-broken-result/"/>
    <published>2026-04-22T10:00:00-04:00</published>
    <updated>2026-04-22T10:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/the-green-check-that-hides-a-broken-result/</id>
    <summary>In a mature estate the dangerous failures are not crashes. They are the runs where exit status is zero, the log looks clean, the file exists and the contents are wrong. A taxonomy of how that happens, and the checking discipline that catches it.</summary>
    <content type="html">&lt;p&gt;A crash is a gift. It has a timestamp, a stack, a non-zero status and somebody’s
attention. Crashes get fixed because they are impossible to ignore.&lt;/p&gt;

&lt;p&gt;The failures that survive in a well-run estate are the other kind. The job exits
zero. The log ends with “completed”. The output file is there, with the right
name, the right owner and a plausible size. Nothing in the monitoring turns
amber. And the contents are wrong — empty where they should be populated,
defaulted where you chose something else, stale where you expected fresh.&lt;/p&gt;

&lt;p&gt;This class has no single cause. It has about seven, and they are mechanical
enough to be listed. What follows is the taxonomy, then the checking discipline
that actually catches them, which is narrower and more annoying than the one
most of us write by default.&lt;/p&gt;

&lt;p&gt;The organizing claim: &lt;strong&gt;a check that asserts on status rather than content is
not a check.&lt;/strong&gt; It is a record that something ran.&lt;/p&gt;

&lt;h2 id=&quot;the-taxonomy&quot;&gt;The taxonomy&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Class&lt;/th&gt;
      &lt;th&gt;What the operator sees&lt;/th&gt;
      &lt;th&gt;Why the probe passes&lt;/th&gt;
      &lt;th&gt;The assertion that catches it&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Error printed, status zero&lt;/td&gt;
      &lt;td&gt;An error line in a log nobody reads&lt;/td&gt;
      &lt;td&gt;Probe tests &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$?&lt;/code&gt; only&lt;/td&gt;
      &lt;td&gt;Compare stdout to an expected value&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Plugin missing, schema intact&lt;/td&gt;
      &lt;td&gt;Correct headers, blank column&lt;/td&gt;
      &lt;td&gt;Header check passes&lt;/td&gt;
      &lt;td&gt;Assert a populated value for a case that must produce one&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;List parsed as one string&lt;/td&gt;
      &lt;td&gt;Silent fallback to a default&lt;/td&gt;
      &lt;td&gt;Config parses, service starts&lt;/td&gt;
      &lt;td&gt;Read the effective value back from the running process&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Zero-byte index fetched&lt;/td&gt;
      &lt;td&gt;“Nothing to update”&lt;/td&gt;
      &lt;td&gt;Transfer tool exits zero&lt;/td&gt;
      &lt;td&gt;Assert a minimum size and a known marker in the content&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Symlink to a path on another host&lt;/td&gt;
      &lt;td&gt;Module loads, variable empty&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module load&lt;/code&gt; exits zero&lt;/td&gt;
      &lt;td&gt;Resolve the path and test the binary runs&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Environment change inside a pipe&lt;/td&gt;
      &lt;td&gt;A working tool reported broken&lt;/td&gt;
      &lt;td&gt;Probe ran in a subshell&lt;/td&gt;
      &lt;td&gt;Set environment, then test in the same shell&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Suite that has never failed&lt;/td&gt;
      &lt;td&gt;An unbroken run of green&lt;/td&gt;
      &lt;td&gt;Nothing has been injected&lt;/td&gt;
      &lt;td&gt;Prove each check fails on a deliberate fault&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Each row below gets its own section, because the mechanism matters more than the
symptom.&lt;/p&gt;

&lt;h2 id=&quot;class-one-the-error-that-exits-zero&quot;&gt;Class one: the error that exits zero&lt;/h2&gt;

&lt;p&gt;Exit status is a convention, not a guarantee. Plenty of widely used tools print
a diagnostic and return success, either by design or by oversight.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;jq&lt;/code&gt; is the honest case. A missing key is not an error in its data model — it is
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;null&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ echo &apos;{&quot;a&quot;:1}&apos; | jq -r &apos;.b&apos;
null
$ echo $?
0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A probe written as “run it, check the status, check that stdout is non-empty”
scores that as a pass. The string &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;null&lt;/code&gt; is four bytes of non-empty output.
Downstream, a config generator writes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;endpoint = null&lt;/code&gt; into a file, the service
parses it, and you get whatever the service does with a nonsense endpoint.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;grep -c&lt;/code&gt; is the inverse and catches people the other way:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ printf &apos;&apos; | grep -c .
0
$ echo $?
1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;It prints something (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0&lt;/code&gt;) and returns failure. A probe keying on status calls
that a failure when the count is the answer you wanted; a probe keying on “is
there output” calls it a pass. Neither is reading the number.&lt;/p&gt;

&lt;p&gt;The most expensive member of this family is HTTP. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;curl&lt;/code&gt; treats an HTTP error
response as a successful transfer, because from its point of view it was one:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ curl -s -o index.json -w &apos;http=%{http_code}\n&apos; https://example.com/api/index.json
http=404
$ echo $?
0
$ ls -l index.json
-rw-r--r-- 1 svc svc 559 Apr 22 09:14 index.json
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The file exists. It is 559 bytes. It is an HTML error page. Every “the file was
created and is non-empty” check passes. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;curl --fail&lt;/code&gt; suppresses the body on a
4xx or 5xx and exits non-zero, which is what you want in a script;
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fail-with-body&lt;/code&gt; (curl 7.76 and later) does the same but keeps the body if you
need it for the log. Neither is the default, and the default is what ends up in
automation written in a hurry.&lt;/p&gt;

&lt;p&gt;One caveat, because it bites checks that are too specific. The documented exit
code for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fail&lt;/code&gt; is 22, and over HTTP/1.1 that is what you get. Over HTTP/2 the
same 404 can surface as 56 instead, because curl tears the stream down and
reports a receive error rather than an HTTP one:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ curl -s --fail -o /dev/null https://example.com/api/index.json ; echo $?
56
$ curl -s --http1.1 --fail -o /dev/null https://example.com/api/index.json ; echo $?
22
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That pair is from curl 8.7.1, and I would treat the specific code as
build- and version-dependent rather than a rule — which is exactly the point.
Test &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fail&lt;/code&gt; for non-zero, not for 22. A check written as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[ $? -eq 22 ]&lt;/code&gt; is
itself a member of this article’s taxonomy: it goes green on the day the
negotiated protocol, or the packaged curl, changes under it.&lt;/p&gt;

&lt;p&gt;The general fix is not to distrust exit codes. It is to stop treating them as
the only evidence. Ask for the value, state what you expect, compare.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;got=$(jq -r &apos;.endpoint // empty&apos; config.json)
[ &quot;$got&quot; = &quot;https://collector.internal:8443&quot; ] || {
  echo &quot;endpoint mismatch: got [$got]&quot; &amp;gt;&amp;amp;2; exit 1; }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That check can fail for a reason other than the process crashing, which is the
only property that makes it a check.&lt;/p&gt;

&lt;h2 id=&quot;class-two-the-plugin-that-does-not-load&quot;&gt;Class two: the plugin that does not load&lt;/h2&gt;

&lt;p&gt;This is the one that has cost me the most time, because the output looks
completely normal.&lt;/p&gt;

&lt;p&gt;Many tools declare their output schema separately from the code that populates
it. The formatter knows the column is called &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt;. Whether anything ever
writes a number into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; depends on a collector, a plugin or an accounting
module that may not be configured at all. When it is not, you do not get an
error. You get the column, with nothing in it.&lt;/p&gt;

&lt;p&gt;Slurm’s accounting is the canonical example. Ask for per-job memory:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ sacct -j 1847221 --format=JobID,JobName,Elapsed,MaxRSS,AveCPU
JobID           JobName    Elapsed     MaxRSS     AveCPU
------------ ---------- ---------- ---------- ----------
1847221         run.sh     02:14:07
1847221.bat+      batch    02:14:07     8412K   02:11:53
1847221.0        python    02:13:44  41235280K  02:11:40
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Now the same query with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt;, which restricts output to allocation records:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ sacct -X -j 1847221 --format=JobID,JobName,Elapsed,MaxRSS,AveCPU
JobID           JobName    Elapsed     MaxRSS     AveCPU
------------ ---------- ---------- ---------- ----------
1847221         run.sh     02:14:07
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MaxRSS&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AveCPU&lt;/code&gt; are gathered per step and stored on step records. The
allocation record has the columns and no values. A reporting script written
around &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct -X&lt;/code&gt; — a reasonable choice, since it gives one row per job — will
produce a memory report for the whole fleet in which every single job used no
memory. The header row is correct. The CSV parses. The chart renders. It is
entirely empty of information and looks like a finding.&lt;/p&gt;

&lt;p&gt;The same shape occurs with a metrics exporter whose collector throws. The
exporter keeps serving &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/metrics&lt;/code&gt;; the failed collector’s series are simply
absent. Any alert of the form &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;rate(thing_total[5m]) &amp;gt; 0&lt;/code&gt; on a series that no
longer exists evaluates against an empty vector and never fires. The alert is
not firing and the thing it watches is unmonitored, and those two states are
indistinguishable on the dashboard. That is what &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;absent()&lt;/code&gt; and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;absent_over_time()&lt;/code&gt; are for, and why a rule that watches a metric should
usually be paired with a rule that watches for the metric’s disappearance.&lt;/p&gt;

&lt;p&gt;The assertion that catches the whole family is the same: &lt;strong&gt;pick a case that must
produce a populated value, and assert on the value.&lt;/strong&gt; Not on the header, not on
the row count, not on the file size.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# a job you know allocated real memory; --noconvert keeps every value in K,
# so you are not comparing 8412K against 39.33G as strings
rss=$(sacct -n -j &quot;$JOB&quot; --format=MaxRSS -P --noconvert \
      | tr -d &apos;K &apos; | grep -E &apos;^[0-9]+$&apos; | sort -n | tail -1)
[ -n &quot;$rss&quot; ] &amp;amp;&amp;amp; [ &quot;$rss&quot; -gt 1000 ] || {
  echo &quot;accounting returned no MaxRSS for a job that used memory&quot; &amp;gt;&amp;amp;2; exit 1; }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two details in that snippet are load-bearing, and both are easy to get wrong in
a way that leaves the check passing. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--noconvert&lt;/code&gt; matters because newer Slurm
renders large values in human units — a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sort -n&lt;/code&gt; over a mixture of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;8412K&lt;/code&gt; and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;39.33G&lt;/code&gt; puts the small number last. And the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;grep -E &apos;^[0-9]+$&apos;&lt;/code&gt; matters
because step records without a value emit a blank line: strip newlines along
with the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;K&lt;/code&gt; (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tr -d &apos;K \n&apos;&lt;/code&gt;, which is the form that comes naturally) and every
row concatenates into one integer.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ printf &apos;8412K\n41235280K\n\n&apos; | tr -d &apos;K \n&apos; | sort -n | tail -1
841241235280
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That number is larger than any threshold you would set, so the check passes
forever, on any input, including no input at all.&lt;/p&gt;

&lt;h3 id=&quot;the-interpretation-that-is-wrong&quot;&gt;The interpretation that is wrong&lt;/h3&gt;

&lt;p&gt;A worked example, because this is where the trap closes.&lt;/p&gt;

&lt;p&gt;A CPU-efficiency report built on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct&lt;/code&gt; produced a fleet average of exactly
1.00. Every user, every partition, every month. The obvious reading is a
perfectly tuned estate. The correct reading is that the two quantities being
divided are the same quantity.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; is allocated CPU time: elapsed seconds multiplied by allocated
cores. It is not consumed CPU time. Divide &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; by itself, or by any
other allocation-derived figure, and you get 1.0 by construction, for every job
that has ever run. The consumed figure is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; (system plus user), and it
lives on step records — so the “one row per job” convenience of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-X&lt;/code&gt; gives you
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; of zero, and an efficiency of 0.00 for everything. Two different
wrong answers, each uniform enough to look like a real property of the fleet.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# wrong: always 1.00 — CPUTime is CPUTimeRAW in a different notation
sacct -X -a -S 2026-03-01 --format=JobID,CPUTimeRAW,CPUTime -P

# workable: consumed over allocated, steps included
sacct -a -S 2026-03-01 --format=JobID,TotalCPU,CPUTimeRAW -P
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Mind the units when you divide those two. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CPUTimeRAW&lt;/code&gt; is an integer count of
CPU-seconds; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TotalCPU&lt;/code&gt; is a formatted duration, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[DD-]HH:MM:SS[.mmm]&lt;/code&gt;. Divide
the string by the integer without parsing it and you will get a number — a
small, uniform, entirely fictitious one — which is the same failure this section
is about, arrived at one layer further down.&lt;/p&gt;

&lt;p&gt;Uniformity is the signal. A real efficiency distribution on a shared cluster is
wide and messy, and the mean lands somewhere below half — I would treat
roughly 30 to 60 percent as the plausible band for a general-purpose fleet, and
I am giving that as a band rather than a figure because it depends heavily on
the workload mix and I have no basis for a tighter claim. What I am confident
about is the mechanism: if a derived metric is the same for every row, suspect
the arithmetic before you believe the estate.&lt;/p&gt;

&lt;h2 id=&quot;class-three-the-list-that-is-one-string&quot;&gt;Class three: the list that is one string&lt;/h2&gt;

&lt;p&gt;Configuration parsers disagree about lists. Some split on whitespace, some on
commas, some take the whole line, and some accept either and quietly prefer the
wrong one. When a list is read as a single literal, it usually matches nothing,
and “matches nothing” is frequently indistinguishable from “not configured” —
which means a fallback to a default you did not choose.&lt;/p&gt;

&lt;p&gt;The shell version is the one everyone has written:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ L=&quot;mlx5_0 mlx5_1&quot;
$ for p in &quot;$L&quot;; do echo &quot;iter: [$p]&quot;; done
iter: [mlx5_0 mlx5_1]
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;One iteration, one token containing a space, and any lookup keyed on it fails.
Unquoted it splits into two; quoted it does not. Both are correct for some
purpose, and the difference is invisible in the log because the loop body ran
and exited zero.&lt;/p&gt;

&lt;p&gt;The config-file version is more insidious because the file passes validation.
OpenSSH’s server config, for most keywords, uses the first value obtained and
ignores later ones. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sshd -t&lt;/code&gt; reports the file is fine. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;systemctl reload sshd&lt;/code&gt;
succeeds. Your carefully placed directive at the bottom of the file does nothing
at all, because something 200 lines above set it first — often a vendor-shipped
include you did not write. The only reliable check is to ask the running daemon
what it actually believes:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# run as root; sshd -T reads the host keys
# sshd 9.8 and later ship as a split binary, but -T still lives on sshd itself
$ sudo sshd -T | grep -i &apos;^ciphers&apos;
ciphers chacha20-poly1305@openssh.com,aes128-ctr,aes192-ctr,aes256-ctr
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sshd -T&lt;/code&gt; prints the effective configuration after all parsing and includes. The
file is what you wrote; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-T&lt;/code&gt; is what is true.&lt;/p&gt;

&lt;p&gt;One qualification, because it is the reading people get wrong: bare &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sshd -T&lt;/code&gt;
prints the configuration for a connection that matches no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Match&lt;/code&gt; block. If the
directive you care about lives inside a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Match Group&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Match Address&lt;/code&gt;, plain
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-T&lt;/code&gt; will not show it and you will conclude, wrongly, that it never took. Supply
the connection you are asking about:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ sudo sshd -T -C user=svc,host=jump.example.com,addr=203.0.113.10 | grep -i &apos;^permitrootlogin&apos;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The output is the effective value &lt;em&gt;for that connection&lt;/em&gt;. There is no single
effective value for a config with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Match&lt;/code&gt; blocks in it, and a check that assumes
there is will be right for most users and quietly wrong for the ones the block
was written for.&lt;/p&gt;

&lt;p&gt;The general form: where a service distinguishes between the file it was given
and the configuration it is running, the effective-value output is the only
thing worth asserting on, and you must ask it about the specific case you care
about. Where a service offers no such output, read the value back through
whatever interface it does expose — a status command, an admin socket, an API —
and compare it to the value you intended, in writing, in the check.&lt;/p&gt;

&lt;h2 id=&quot;class-four-the-zero-byte-index&quot;&gt;Class four: the zero-byte index&lt;/h2&gt;

&lt;p&gt;A package index, a manifest, a rule feed and a firmware catalog all share a
property: when they arrive empty, the consuming tool reports that there is
nothing to do.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ curl -s -o /var/cache/feed/index.json -w &apos;http=%{http_code} size=%{size_download}\n&apos; \
    https://feeds.example.com/v2/index.json
http=200 size=0
$ echo $?
0
$ ls -l /var/cache/feed/index.json
-rw-r--r-- 1 root root 0 Apr 22 03:00 /var/cache/feed/index.json
$ ./apply-feed --index /var/cache/feed/index.json
0 entries loaded, 0 applied, 0 errors
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three zeros and a clean exit. Read it out loud: nothing to apply, so nothing was
applied, so no errors. That is a correct description of a broken pipeline.&lt;/p&gt;

&lt;p&gt;Note what that transcript is and is not, because the failure modes here do not
rank the way intuition ranks them. It is a 200 with an empty body — a
half-deployed origin, a proxy serving a truncated cache entry, a signing step
that produced nothing. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fail&lt;/code&gt; does not save you, because there is no HTTP
error to fail on: the transfer genuinely succeeded and what succeeded was
nothing.&lt;/p&gt;

&lt;p&gt;The counter-intuitive part is which failure is worse for the file on disk. A
transport failure is the gentle one — curl exits non-zero and leaves an existing
destination untouched, so the last good copy survives:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ echo &quot;PREVIOUS GOOD CONTENT&quot; &amp;gt; index.json
$ curl -s -o index.json https://feeds.example.invalid/v2/index.json ; echo $?
6
$ cat index.json
PREVIOUS GOOD CONTENT
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;An HTTP error without &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fail&lt;/code&gt; is the destructive one. The transfer succeeds, so
curl writes, and the error page lands on top of the only good copy you had:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ curl -s -o index.json https://example.com/api/index.json ; echo $?
0
$ ls -l index.json
-rw-r--r-- 1 root root 559 Apr 22 03:00 index.json
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Adding &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fail&lt;/code&gt; restores the gentle behavior — non-zero status, destination
left alone. Which reframes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--fail&lt;/code&gt; from a tidiness flag into a data-protection
one: without it, every scheduled fetch is one origin misconfiguration away from
overwriting a good cache with an error page, at three in the morning, exiting
zero.&lt;/p&gt;

&lt;p&gt;The reason this survives is that “no changes” is also the normal steady state.
On most days the honest answer genuinely is “0 applied”. A check that only
alarms on errors will never distinguish “we are current” from “we are blind”,
and the second state can persist for months.&lt;/p&gt;

&lt;p&gt;Two assertions, both cheap:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# 1. minimum plausible size, not merely non-empty
# (GNU coreutils; on BSD/macOS the equivalent is stat -f %z)
sz=$(stat -c %s /var/cache/feed/index.json)
[ &quot;$sz&quot; -ge 4096 ] || { echo &quot;index implausibly small: ${sz}B&quot; &amp;gt;&amp;amp;2; exit 1; }

# 2. a structural marker, not just bytes
jq -e &apos;.entries | length &amp;gt; 0&apos; /var/cache/feed/index.json &amp;gt;/dev/null || {
  echo &quot;index parsed but contains no entries&quot; &amp;gt;&amp;amp;2; exit 1; }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;And one more that is worth the trouble on anything time-sensitive: assert on the
age of the content, not the age of the file. Any transfer that opens the
destination before it knows the result — a 200 that dies half way through the
body, an error page written over a good cache — leaves you a fresh mtime over
bytes that are stale, partial or wrong. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;find -mtime&lt;/code&gt; is then measuring when
something was written, not when it was current. If the payload carries a
generation timestamp, compare that instead, and treat the file’s own mtime as
evidence of nothing but write activity.&lt;/p&gt;

&lt;h2 id=&quot;class-five-the-symlink-that-resolves-nowhere&quot;&gt;Class five: the symlink that resolves nowhere&lt;/h2&gt;

&lt;p&gt;Shared software trees accumulate links. A module file or a wrapper script points
at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/opt/apps/tool/current&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;current&lt;/code&gt; points at a release directory, and at
some point a release directory is moved, or the link is created on a build host
against a path that only exists there.&lt;/p&gt;

&lt;p&gt;The link itself is not broken in any way the filesystem will tell you about.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ln -s&lt;/code&gt; will happily create a link to a path that does not exist. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ls&lt;/code&gt; shows it.
And the module that consumes it does this:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;prepend-path PATH    $prefix/bin
prepend-path LD_LIBRARY_PATH $prefix/lib64
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;where &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$prefix&lt;/code&gt; was computed by resolving the link. If the resolution produced
an empty string, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PATH&lt;/code&gt; gains the entry &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/bin&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LD_LIBRARY_PATH&lt;/code&gt; gains
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/lib64&lt;/code&gt;. Both are real directories. Neither errors. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module load&lt;/code&gt; exits zero,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module list&lt;/code&gt; shows the module loaded, and the binary you wanted is not on the
path — so you get the system’s version, silently, at a different version than
the one the module claims.&lt;/p&gt;

&lt;p&gt;The distinguishing commands:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ ls -l /opt/apps/tool/current
lrwxrwxrwx 1 root root 33 Feb 11 16:02 /opt/apps/tool/current -&amp;gt; /build/stage/opt/apps/tool/3.11.2

$ readlink -e /opt/apps/tool/current
$ echo $?
1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;readlink -e&lt;/code&gt; requires every component of the resolved path to exist and prints
nothing when it does not. For a check, that is the flag you want.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;readlink -f&lt;/code&gt; is the one usually reached for, and the difference is worth
spelling out because the obvious summary of it is wrong. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-f&lt;/code&gt; is not “the same
but it does not check”. It requires every component &lt;em&gt;but the last&lt;/em&gt; to exist. So
its behavior splits by which part of the target went missing:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# release directory pruned; the parent is still there → -f is happy, -e is not
$ readlink -f /opt/apps/tool/releases/3.11.2
/opt/apps/tool/releases/3.11.2
$ echo $?
0
$ readlink -e /opt/apps/tool/releases/3.11.2 ; echo $?
1

# link written on a build host; /build/stage does not exist here at all
$ readlink -f /opt/apps/tool/current ; echo $?
1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Which means a check built on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-f&lt;/code&gt; catches the build-host case and misses the
pruned-release case — and the pruned release is the one that happens on a
schedule. It fails in the direction that looks fine.&lt;/p&gt;

&lt;p&gt;Both flags are GNU coreutils. BSD and macOS &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;readlink&lt;/code&gt; has no &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-e&lt;/code&gt; at all, and
its &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-f&lt;/code&gt; does not behave like the GNU one, so a check that runs on a mixed
estate should use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;test -e &quot;$(readlink -f …)&quot;&lt;/code&gt; or simply
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;[ -e /opt/apps/tool/current ]&lt;/code&gt;, which follows the link and answers the question
directly.&lt;/p&gt;

&lt;p&gt;The assertion is not “does the module load”. It is “after loading, does the
thing run and say the right version”:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;module purge
module load tool/3.11.2
command -v tool | grep -q &apos;^/opt/apps/&apos; || { echo &quot;tool not from /opt/apps&quot; &amp;gt;&amp;amp;2; exit 1; }
tool --version | grep -q &apos;3\.11\.2&apos;     || { echo &quot;wrong tool version&quot; &amp;gt;&amp;amp;2; exit 1; }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two lines, and they fail for the right reasons: wrong location, wrong version.&lt;/p&gt;

&lt;h2 id=&quot;class-six-the-environment-change-inside-a-pipe&quot;&gt;Class six: the environment change inside a pipe&lt;/h2&gt;

&lt;p&gt;This one is the mirror image of the rest. Here the estate is fine and the check
is wrong, and the natural conclusion — “the module is broken” — sends you off to
rebuild something that was never damaged.&lt;/p&gt;

&lt;p&gt;In bash, every element of a pipeline runs in a subshell by default. Anything it
does to the environment dies with it:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ x=1
$ echo hi | { x=2; }
hi
$ echo &quot;x=$x&quot;
x=1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;“By default” is doing real work in that sentence, and this is the one place in
the article where the portable-looking rule is the unportable one. Bash exempts
the last stage if &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lastpipe&lt;/code&gt; is set and job control is off — the usual state of
a non-interactive script. zsh and ksh93 run the last stage in the current shell
unconditionally:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bash -c &apos;echo hi | read y; echo &quot;bash y=[$y]&quot;&apos;
bash y=[]
$ zsh  -c &apos;echo hi | read y; echo &quot;zsh y=[$y]&quot;&apos;
zsh y=[hi]
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;So the same validation script, interpreted by a different shell, produces a
different verdict about whether the estate is broken. Neither shell is wrong.
Pin the interpreter in the shebang and do not rely on which side of that line
you are on.&lt;/p&gt;

&lt;p&gt;Now the operational version. A validation script wants to load a module and
capture the result:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;module load fftw/3.3.10 | tee -a validate.log
echo &quot;FFTW_DIR=[$FFTW_DIR]&quot;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;module&lt;/code&gt; is a shell function. Inside the pipe it runs in a subshell, sets its
variables there, and the subshell exits. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;FFTW_DIR&lt;/code&gt; is empty in the parent. The
script concludes the module sets nothing and reports a broken module. A second
engineer loads it by hand, it works perfectly, and the disagreement gets
attributed to “something about the environment”.&lt;/p&gt;

&lt;p&gt;The same mechanism hides real failures too:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bash -c &apos;false | true; echo &quot;rc=$?&quot;&apos;
rc=0
$ bash -c &apos;set -o pipefail; false | true; echo &quot;rc=$?&quot;&apos;
rc=1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Without &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pipefail&lt;/code&gt;, a pipeline reports the status of its last command only. Pipe
a failing producer into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tee&lt;/code&gt; and the failure is gone. This is why so much
automation that pipes its output to a log has no idea when it failed.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;set -e&lt;/code&gt; has its own blind spot in the same family:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bash -c &apos;set -eu; f(){ local v=$(false); echo &quot;inside rc=$? v=[$v]&quot;; }; f; echo after&apos;
inside rc=0 v=[]
after
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;local&lt;/code&gt; is itself a command, and its exit status is the status of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;local&lt;/code&gt;, not
of the substitution inside it. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;export&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;declare&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;readonly&lt;/code&gt; swallow a
failure the same way, for the same reason:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bash -c &apos;set -eu; export v=$(false); echo &quot;reached&quot;&apos;
reached
$ bash -c &apos;set -eu; declare v=$(false); echo &quot;reached&quot;&apos;
reached
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Split the declaration from the assignment and the same script stops:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ bash -c &apos;set -eu; f(){ local v; v=$(false); echo &quot;never reached&quot;; }; f&apos;
$ echo $?
1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;One more in the same category, worth knowing because it appears in cleanup and
maintenance scripts — and worth stating carefully, because the usual version of
this folklore is half true. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;find -exec cmd {} \;&lt;/code&gt; returns zero even when every
invocation of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cmd&lt;/code&gt; fails. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;find -exec cmd {} +&lt;/code&gt; does not: in GNU findutils a
non-zero invocation makes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;find&lt;/code&gt; itself exit non-zero.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ find /tmp/fecheck -type f -exec false {} \; ; echo &quot;semicolon rc=$?&quot;
semicolon rc=0
$ find /tmp/fecheck -type f -exec false {} +  ; echo &quot;plus rc=$?&quot;
plus rc=1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The terminator, not the tool, decides whether failures propagate. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;xargs&lt;/code&gt; also
reports them — GNU &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;xargs&lt;/code&gt; exits 123 if any invocation exits 1 to 125 — so if a
maintenance script’s success depends on the command being run, use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;+&lt;/code&gt; or
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;xargs&lt;/code&gt;, and do not read anything into the status of a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;\;&lt;/code&gt; form.&lt;/p&gt;

&lt;p&gt;The rule that covers all of it, for any harness you write:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;set -Eeuo pipefail
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;plus: do not put environment-modifying commands in pipelines, and do not hide a
command substitution inside a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;local&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;declare&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;export&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id=&quot;class-seven-the-suite-that-has-never-failed&quot;&gt;Class seven: the suite that has never failed&lt;/h2&gt;

&lt;p&gt;The last class is not a bug. It is the reason the other six survive.&lt;/p&gt;

&lt;p&gt;A validation suite that has returned green on every run since it was written has
demonstrated exactly one thing: that it returns green. It has never been shown
to be capable of returning anything else. If half its assertions were deleted
tonight, tomorrow’s run would be identical, and you would never know.&lt;/p&gt;

&lt;p&gt;This is the same argument as the one for restore drills and for fire alarms with
a test button. A detector that has never been triggered under controlled
conditions is decoration until proven otherwise.&lt;/p&gt;

&lt;p&gt;The fix is negative controls: for every check, construct the fault it claims to
detect, run the check, and require that it fails. Keep those injections in the
repository next to the suite, and run them on a schedule.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;#!/usr/bin/env bash
# negctl.sh — prove each check can fail
set -Eeuo pipefail

fixtures=/var/lib/validate/fixtures
fail=0

expect_fail() {           # expect_fail &amp;lt;label&amp;gt; &amp;lt;cmd...&amp;gt;
  local label=$1; shift
  if &quot;$@&quot; &amp;gt;/dev/null 2&amp;gt;&amp;amp;1; then
    echo &quot;NEGATIVE CONTROL DID NOT FIRE: $label&quot;
    fail=1
  else
    echo &quot;ok (failed as required): $label&quot;
  fi
}

expect_fail &quot;empty index rejected&quot;      check_index &quot;$fixtures/index.empty.json&quot;
expect_fail &quot;html error page rejected&quot;  check_index &quot;$fixtures/index.404.html&quot;
expect_fail &quot;truncated index rejected&quot;  check_index &quot;$fixtures/index.truncated.json&quot;
expect_fail &quot;blank MaxRSS rejected&quot;     check_accounting &quot;$fixtures/sacct.blank.csv&quot;
expect_fail &quot;dangling symlink rejected&quot; check_prefix   &quot;$fixtures/dangling-link&quot;
expect_fail &quot;wrong version rejected&quot;    check_version  &quot;$fixtures/tool-3.10.0&quot;

exit &quot;$fail&quot;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The fixtures matter as much as the script. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;index.404.html&lt;/code&gt; should be a real
error page you captured, not one you typed; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sacct.blank.csv&lt;/code&gt; should be real
output from a query that really did return blank columns. Saving the broken
artefact at the moment you find it is the cheapest thing in this article and the
one most often skipped.&lt;/p&gt;

&lt;p&gt;There is a second-order benefit. Running the negative controls tells you when a
check has quietly stopped being able to fail — because someone loosened a
threshold, or because the tool’s output format changed and a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;grep&lt;/code&gt; that used to
match now matches nothing and the surrounding &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;|| true&lt;/code&gt; swallows it.&lt;/p&gt;

&lt;h2 id=&quot;the-discipline-condensed&quot;&gt;The discipline, condensed&lt;/h2&gt;

&lt;p&gt;Four rules, in the order they pay off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assert on content, not on status.&lt;/strong&gt; Exit zero means the process reached its
own idea of the end. It says nothing about the bytes it produced. Every check
should name a value it expects and compare.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write the expected value down in the check.&lt;/strong&gt; Not “non-empty”, not “more than
zero rows” — the actual string, the actual minimum, the actual version. A check
that does not encode an expectation cannot detect a wrong answer, only a missing
one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pick a case that must produce a populated result.&lt;/strong&gt; The whole plugin-shaped
failure class hides behind aggregate checks. One known job, one known file, one
known value, asserted exactly, is worth more than a thousand rows counted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prove every check can fail.&lt;/strong&gt; If you cannot show the injected fault that turns
it red, it is not a check. Delete it or fix it, because leaving it there is
worse than having nothing — it produces the feeling of coverage without the
coverage.&lt;/p&gt;

&lt;p&gt;And one habit underneath all four: when a number is uniform across every row, or
a report is empty on a day when it should not be, or a pipeline says “0 applied”
for the ninth week running, treat that as a finding about the measurement rather
than a finding about the estate. In my experience the uniform answer is wrong far
more often than it is right.&lt;/p&gt;

&lt;h2 id=&quot;what-this-does-not-cover&quot;&gt;What this does not cover&lt;/h2&gt;

&lt;p&gt;Nothing here detects a computation that is wrong in a plausible way — output
that is populated, well-formed and numerically incorrect. That needs a reference
result, a known-answer test or an independent implementation, and it is a
different and harder problem.&lt;/p&gt;

&lt;p&gt;Nor does any of it help if the check and the thing being checked share a
dependency. A probe that reads the same broken index, or resolves the same
dangling link, or inherits the same empty variable, will agree with the system
perfectly. When you build a negative control, inject the fault as far upstream as
you can reach, and confirm the alarm travels all the way to the place a human
would actually see it. A check nobody reads is the same as a check that cannot
fail, arrived at by a different route.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Ansible patching: verify the running service after the package transaction</title>
    <link href="https://www.wirewalk.com/writing/ansible-linux-patching-canaries/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/ansible-linux-patching-canaries/</id>
    <summary>Use small batches, explicit stop conditions, and application checks to avoid reporting installation as remediation.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A package update can complete while the vulnerable process remains running or the application becomes unhealthy. Patching is a controlled transition of a service, not merely a package-manager transaction.&lt;/p&gt;

&lt;h2 id=&quot;define-the-change-boundary&quot;&gt;Define the change boundary&lt;/h2&gt;

&lt;p&gt;Identify the affected inventory, application owners, package source, reboot requirements, and rollback route. Separate hosts with different availability roles. A database replica, login node, and stateless worker should not share an unexamined reboot policy.&lt;/p&gt;

&lt;p&gt;Start with read-only inventory of installed packages, running kernels, and service health. Use vendor advisories for applicability rather than assuming that an upstream version string maps directly to a distribution’s security state.&lt;/p&gt;

&lt;h2 id=&quot;control-the-batch-size&quot;&gt;Control the batch size&lt;/h2&gt;

&lt;p&gt;Ansible documents &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;serial&lt;/code&gt; for batched execution in its &lt;a href=&quot;https://docs.ansible.com/ansible/latest/playbook_guide/playbooks_strategies.html&quot;&gt;strategy guide&lt;/a&gt;. A skeleton for a reviewed playbook is:&lt;/p&gt;

&lt;div class=&quot;language-yaml highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;pi&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;na&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;Validate a small Linux maintenance batch&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;hosts&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;maintenance_canary&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;serial&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;m&quot;&gt;1&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;any_errors_fatal&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;no&quot;&gt;true&lt;/span&gt;
  &lt;span class=&quot;na&quot;&gt;tasks&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;pi&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;na&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;Record the running kernel&lt;/span&gt;
      &lt;span class=&quot;na&quot;&gt;ansible.builtin.command&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;uname -r&lt;/span&gt;
      &lt;span class=&quot;na&quot;&gt;changed_when&lt;/span&gt;&lt;span class=&quot;pi&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;no&quot;&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This example only records a kernel version; it deliberately does not install packages or reboot. Add application-specific prechecks, the approved update operation, restart handling, and postchecks through the normal change process.&lt;/p&gt;

&lt;h2 id=&quot;treat-check-mode-as-a-planning-aid&quot;&gt;Treat check mode as a planning aid&lt;/h2&gt;

&lt;p&gt;Ansible’s &lt;a href=&quot;https://docs.ansible.com/ansible/latest/playbook_guide/playbooks_checkmode.html&quot;&gt;check-mode documentation&lt;/a&gt; describes module-dependent simulation. A clean check-mode run does not demonstrate a successful reboot, kernel-module compatibility, or a healthy application afterward.&lt;/p&gt;

&lt;p&gt;Run the full workflow on a representative canary. Include external client checks, not only service-manager status. A process can be active while its database connection or authentication path is broken.&lt;/p&gt;

&lt;h2 id=&quot;verify-the-running-state&quot;&gt;Verify the running state&lt;/h2&gt;

&lt;p&gt;After maintenance, record the installed package version and the process or kernel actually in use. Confirm mounts, dependent services, application transactions, and monitoring. For hosts with specialized NIC, storage, or GPU modules, include compatibility checks appropriate to the deployed stack.&lt;/p&gt;

&lt;p&gt;If the canary fails, stop the rollout. Do not mark a host recovered merely because a second package command returned success. Use the established rollback or repair path and repeat the original service checks.&lt;/p&gt;

&lt;h2 id=&quot;account-for-incomplete-hosts&quot;&gt;Account for incomplete hosts&lt;/h2&gt;

&lt;p&gt;Report unreachable hosts, skipped tasks, deferred reboots, and accepted exceptions separately. An inventory of 200 targets with 190 successful transactions is not a fully patched fleet. Preserve the original target list so disappearing machines cannot improve the success percentage.&lt;/p&gt;

&lt;p&gt;The release record should show what was targeted, what changed, which running services were verified, and what remains outstanding. Repeat the same acceptance checks after the next image rebuild so the patched state survives replacement as well as reboot.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>SELinux denials: repair the application boundary before generating policy</title>
    <link href="https://www.wirewalk.com/writing/selinux-denials-without-disabling/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/selinux-denials-without-disabling/</id>
    <summary>Labels, supported booleans, and application behavior should be understood before a new allow rule is introduced.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;An application failing under SELinux enforcement is evidence to investigate, not proof that enforcement must be disabled. The failure may be a mislabeled path, a supported configuration option, or behavior that the service should not be performing.&lt;/p&gt;

&lt;h2 id=&quot;capture-the-actual-denial&quot;&gt;Capture the actual denial&lt;/h2&gt;

&lt;p&gt;On a RHEL-family host with SELinux and audit tools installed, begin with read-only inspection:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;getenforce
sestatus
&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;ausearch &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; AVC,USER_AVC &lt;span class=&quot;nt&quot;&gt;-ts&lt;/span&gt; recent
&lt;span class=&quot;nb&quot;&gt;ls&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-Zd&lt;/span&gt; /srv /srv/example-app
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The paths are examples. Use the real application path, and preserve timestamps so the denial can be connected to the failed transaction. Absence of an AVC does not prove SELinux is unrelated; inspect the logging state and relevant troubleshooting guidance.&lt;/p&gt;

&lt;p&gt;Use the distribution’s supported policy and tools. The &lt;a href=&quot;https://github.com/SELinuxProject/selinux&quot;&gt;SELinux project&lt;/a&gt; maintains the userspace tools, while the &lt;a href=&quot;https://docs.kernel.org/admin-guide/LSM/SELinux.html&quot;&gt;kernel documentation&lt;/a&gt; points to the subsystem and policy resources.&lt;/p&gt;

&lt;h2 id=&quot;identify-the-boundary-being-crossed&quot;&gt;Identify the boundary being crossed&lt;/h2&gt;

&lt;p&gt;Read the source context, target context, object class, and denied permission. Connect those to the application operation. Is a web process trying to read its content, write an upload directory, contact a database, or access something unrelated?&lt;/p&gt;

&lt;p&gt;Check file labeling against the intended path policy. For an approved path, a nonmodifying preview can help:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;restorecon &lt;span class=&quot;nt&quot;&gt;-nRv&lt;/span&gt; /srv/example-app
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-n&lt;/code&gt; preview reports prospective relabeling rather than applying it. A correct custom path may need a persistent file-context definition before any relabel operation; repeatedly applying a temporary label is not a durable deployment method.&lt;/p&gt;

&lt;h2 id=&quot;prefer-a-narrow-explanation&quot;&gt;Prefer a narrow explanation&lt;/h2&gt;

&lt;p&gt;If the distribution provides a documented boolean for the intended behavior, review its scope before enabling it. If the service was configured to write into a read-only content path, correcting the application layout may be the better fix.&lt;/p&gt;

&lt;p&gt;Do not feed a large collection of historical denials into an automatic policy generator and install the result without review. That can authorize unrelated activity and preserve a misconfiguration as policy.&lt;/p&gt;

&lt;h2 id=&quot;validate-the-change-in-a-pilot&quot;&gt;Validate the change in a pilot&lt;/h2&gt;

&lt;p&gt;Apply the approved label, configuration, or minimal policy change to a lab or pilot system. Run the previously failing transaction in enforcing mode. Then run an operation that should remain denied. Successful application behavior alone does not prove the boundary stayed narrow.&lt;/p&gt;

&lt;p&gt;Reboot or redeploy the pilot where appropriate to confirm the fix survives normal lifecycle operations. A label repair that disappears at the next image rollout is incomplete.&lt;/p&gt;

&lt;h2 id=&quot;retain-the-reason&quot;&gt;Retain the reason&lt;/h2&gt;

&lt;p&gt;Keep the denial, diagnosis, exact change, positive test, and negative test. Document any exception with its owner and review date. This turns an operational inconvenience into a repeatable hardening improvement instead of a permanent loss of enforcement.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Four ways a storage benchmark lies and how to catch each one</title>
    <link href="https://www.wirewalk.com/writing/benchmarks-that-lie/"/>
    <published>2026-03-18T10:00:00-04:00</published>
    <updated>2026-03-18T10:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/benchmarks-that-lie/</id>
    <summary>A read test smaller than the server&apos;s buffer pool measures memory. An under-driven benchmark is indistinguishable from a ceiling. Counters are not rates, and aggregate CPU idle is meaningless on a many-core box. Here is the arithmetic that catches each one before you act on the number.</summary>
    <content type="html">&lt;p&gt;Storage benchmarks are unusually good at producing wrong answers that look
right. The number comes out, it is in the units you expected, it is in the
neighbourhood you hoped for, and nothing in the output says “this measurement
did not test what you think it tested”. Then someone writes it in a capacity
plan, or opens a support case about it, or spends a maintenance window tuning a
subsystem that was never the constraint.&lt;/p&gt;

&lt;p&gt;Four mechanisms account for most of it. Each has a specific arithmetic
signature, and each has a check that costs a few minutes and settles it.&lt;/p&gt;

&lt;p&gt;The worked figures throughout are from one eight-server erasure-coded tier on
HDR200, and they are illustrative rather than a specification for anything: page
pool sizes, code widths and link rates all move the numbers. The arithmetic is
the part that transfers. Run it against your own configuration before quoting
any figure here.&lt;/p&gt;

&lt;h2 id=&quot;1-the-read-you-measured-was-served-from-memory&quot;&gt;1. The read you measured was served from memory&lt;/h2&gt;

&lt;p&gt;Most benchmarking guides tell you to use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;O_DIRECT&lt;/code&gt;. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fio --direct=1&lt;/code&gt;,
IOR’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--posix.odirect&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dd iflag=direct&lt;/code&gt; for a read or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;oflag=direct&lt;/code&gt; for a
write. This is correct advice and it is routinely misunderstood, because
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;O_DIRECT&lt;/code&gt; is a flag on a file descriptor on one machine. It bypasses the page
cache on the client that opened the file. It has no authority whatsoever over
what the storage servers do with their own memory.&lt;/p&gt;

&lt;p&gt;On a parallel filesystem the servers hold a large read buffer. For IBM Storage
Scale in its erasure-coded form, that buffer is a percentage of the daemon’s
page pool, set by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nsdRAIDBufferPoolSizePct&lt;/code&gt;. The documented default has been
50 percent historically and 80 percent in more recent releases, so read it
rather than assume it. The upper bound on the cache your &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;O_DIRECT&lt;/code&gt; read is
still landing in is:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;server-side buffer pool = pagepool × nsdRAIDBufferPoolSizePct × number of servers
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That is an upper bound rather than a read cache exactly: the same pool holds
write and track buffers. It is still the right number to size a dataset
against, because it is the largest thing your read can hide in.&lt;/p&gt;

&lt;p&gt;Check yours rather than assuming. Note that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmlsconfig&lt;/code&gt; tells you what is
stored and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmfsadm dump config&lt;/code&gt; tells you what is actually running, and these
are not always the same value. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mmfsadm&lt;/code&gt; is a diagnostic tool that IBM
documents as being for use under service direction; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;dump config&lt;/code&gt; is read-only,
but treat the rest of the command surface as off-limits on a production
daemon:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mmlsconfig pagepool
mmlsconfig nsdRAIDBufferPoolSizePct
&lt;span class=&quot;c&quot;&gt;# and on one storage server, what is genuinely live:&lt;/span&gt;
mmfsadm dump config | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;pagepool|nsdRAIDBufferPoolSizePct&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Take a tier of eight storage servers with a 37.7 GiB page pool each and the
buffer pool at 80 percent. That is 30.2 GiB per server and 241 GiB across the
tier. Any read test whose working set is comfortably below 241 GiB is measuring
DDR, not NVMe, no matter how many &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;O_DIRECT&lt;/code&gt; flags are in the command line.&lt;/p&gt;

&lt;p&gt;Here is what that looks like. Twelve clients, sixteen ranks each, 1 GiB per
rank — 192 GiB of data, which is 80 percent of the buffer pool:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;mpirun &lt;span class=&quot;nt&quot;&gt;-np&lt;/span&gt; 192 &lt;span class=&quot;nt&quot;&gt;-N&lt;/span&gt; 16 ./ior &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt; POSIX &lt;span class=&quot;nt&quot;&gt;--posix&lt;/span&gt;.odirect &lt;span class=&quot;nt&quot;&gt;-w&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; 1g &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; 1m &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-C&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-i&lt;/span&gt; 1 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; /gpfs/scratch/bench/ior.dat
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;access    bw(MiB/s)  IOPS       Latency(s)  block(KiB) xfer(KiB)  open(s)   wr/rd(s)  close(s)  total(s)
------    ---------  ----       ----------  ---------- ---------  --------  --------  --------  --------
write     68812      68812      0.002790    1048576    1024.00    0.021443  2.857     0.003120  2.882
read      159539     159539     0.001203    1048576    1024.00    0.004472  1.232     0.001196  1.238
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;159,539 MiB/s is 155.8 GiB/s. It is also nonsense. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;IOPS&lt;/code&gt; column equals the
bandwidth figure here only because the transfer size is exactly 1 MiB — IOR
reports completed transfers per second, not GiB/s, and at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-t 4m&lt;/code&gt; the two
columns would differ by four. Now the same test with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-b 16g&lt;/code&gt; — 3 TiB, nearly
thirteen times the buffer pool:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;access    bw(MiB/s)  IOPS       Latency(s)  block(KiB) xfer(KiB)  open(s)   wr/rd(s)  close(s)  total(s)
------    ---------  ----       ----------  ---------- ---------  --------  --------  --------  --------
write     68240      68240      0.002814    16777216   1024.00    0.019987  46.10     0.004011  46.13
read      83942      83942      0.002287    16777216   1024.00    0.003991  37.47     0.001330  37.48
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Read falls from 155.8 to 82.0 GiB/s. The first figure was 1.9 times the cold
number. Write barely moves — 68,812 to 68,240 MiB/s, or 67.2 to 66.6 GiB/s —
which is the part that tells you the
mechanism rather than just the magnitude: writes were never being served out of
a read buffer, so they had nothing to lose.&lt;/p&gt;

&lt;p&gt;Two tells you can spot without doing the arithmetic:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The read phase finished in a second.&lt;/strong&gt; Any storage read phase that completes
in under roughly ten seconds has measured startup transients and cache. IOR
prints &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;total(s)&lt;/code&gt; for exactly this reason. A 1.2 second read phase is not a
measurement, it is a rounding error with units.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read is more than twice write on hardware where it should not be.&lt;/strong&gt; Some
asymmetry is normal. A factor of 2.3 on a tier whose read and write paths use
the same NICs deserves an explanation before it gets published.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Here the obvious interpretation is wrong in a way worth stating plainly. When
the number falls after you enlarge the dataset, the natural conclusion is “so it
was cache, and 82 is the truth”. That conclusion is probably right, but two
points do not yet support it. Enlarging the dataset also lengthens the runtime,
changes which regions of the address space you touch, and may cross a storage
pool or fileset boundary with different placement. Any of those could depress
the number on its own.&lt;/p&gt;

  &lt;p&gt;The discriminator is more points along the same axis. Run at 2×, 4× and 8× the
buffer pool as well as the 12.7× above. A cache artefact produces a sharp fall
and then a flat floor — for instance 155.8 at 0.8×, then 82.4, 81.6, 82.3, 82.0
as the multiple climbs. A layout or locality effect produces a continuing slide.
If it keeps falling, you have found a second problem and you still do not know
the cold number.&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;There is no supported command that flushes the server-side pool on demand, so
sizing the dataset past it is the practical answer. On the clients,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;echo 3 &amp;gt; /proc/sys/vm/drop_caches&lt;/code&gt; clears the Linux page cache but does not
touch the GPFS page pool, which is separate memory the daemon manages itself;
unmounting the filesystem on that client is what invalidates its cached buffers.&lt;/p&gt;

&lt;h2 id=&quot;2-an-under-driven-benchmark-is-indistinguishable-from-a-ceiling&quot;&gt;2. An under-driven benchmark is indistinguishable from a ceiling&lt;/h2&gt;

&lt;p&gt;This is the most expensive of the four, because it does not merely give you a
wrong number. It gives you a wrong number with a confident shape. You run at one
concurrency level, get 44 GiB/s, run again and get 44 GiB/s, add a node and get
44 GiB/s, and conclude you have found a hard wall. Everything about the result
says saturation. What you have actually found is that you did not ask for
enough work at once.&lt;/p&gt;

&lt;p&gt;The arithmetic is Little’s law. The number of I/Os that must be outstanding to
sustain a given bandwidth is:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;outstanding = bandwidth ÷ I/O size × latency
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;At 82 GiB/s with 1 MiB transfers, that is 88.0 GB/s ÷ 1,048,576 B = about
84,000 IOPS. If per-I/O latency at the knee is 18 ms, you need roughly 1,510
requests in flight across the whole client set to get there. If your benchmark
has 384 outstanding, then at that latency you are arithmetically incapable of
reaching the number, and no amount of repeating the run will reveal that.&lt;/p&gt;

&lt;p&gt;So sweep concurrency. Twelve clients, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iodepth=32&lt;/code&gt;, varying &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;numjobs&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;j &lt;span class=&quot;k&quot;&gt;in &lt;/span&gt;1 2 4 8 16&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
  &lt;/span&gt;pdsh &lt;span class=&quot;nt&quot;&gt;-w&lt;/span&gt; node[01-12] &lt;span class=&quot;s2&quot;&gt;&quot;fio --name=r --filename=/gpfs/scratch/bench/&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\$&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;(hostname)/f &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;
    --rw=read --bs=1m --direct=1 --ioengine=libaio &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;
    --iodepth=32 --numjobs=&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$j&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt; --size=256g --runtime=60 --time_based &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;
    --group_reporting --output-format=json &lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;
    --output=/gpfs/scratch/bench/&lt;/span&gt;&lt;span class=&quot;se&quot;&gt;\$&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;(hostname)/sweep-&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$j&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;.json&quot;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Write each node’s JSON to its own file rather than redirecting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pdsh&lt;/code&gt; to one.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pdsh&lt;/code&gt; interleaves and line-prefixes the output of all twelve nodes, so a single
redirect gives you twelve mangled JSON documents in one file and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;jq&lt;/code&gt; will
refuse it. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--direct=1&lt;/code&gt; is not optional here either: fio documents &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libaio&lt;/code&gt; as
only supporting queued behavior with non-buffered I/O, so on a buffered job it
degrades toward serialized submission and reintroduces exactly the
under-driving this sweep exists to detect.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--size=256g&lt;/code&gt; is chosen against problem 1, not at random. With &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--time_based&lt;/code&gt;
the 60-second run reads about 408 GiB from a 256 GiB file, so each node does
re-read its own data — but twelve nodes hold 3 TiB between them against a
241 GiB server pool, which keeps the hit rate low enough not to distort the
shape of the sweep. Shrink the file to something a single server could hold and
the whole curve lifts and flattens, and you will read the artefact as a ceiling.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;numjobs&lt;/code&gt; per node&lt;/th&gt;
      &lt;th&gt;Outstanding I/Os, all clients&lt;/th&gt;
      &lt;th&gt;Aggregate read (GiB/s)&lt;/th&gt;
      &lt;th&gt;Δ from previous&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;384&lt;/td&gt;
      &lt;td&gt;44.0&lt;/td&gt;
      &lt;td&gt;—&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;2&lt;/td&gt;
      &lt;td&gt;768&lt;/td&gt;
      &lt;td&gt;68.1&lt;/td&gt;
      &lt;td&gt;+55%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;1,536&lt;/td&gt;
      &lt;td&gt;81.5&lt;/td&gt;
      &lt;td&gt;+20%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;3,072&lt;/td&gt;
      &lt;td&gt;82.0&lt;/td&gt;
      &lt;td&gt;+0.6%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;16&lt;/td&gt;
      &lt;td&gt;6,144&lt;/td&gt;
      &lt;td&gt;81.7&lt;/td&gt;
      &lt;td&gt;−0.4%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The shape is the result, not any single row. It climbs steeply, the climb
decays, and then three consecutive points agree within one percent while
concurrency quadruples. That flat tail is what licenses the sentence “this tier
saturates at 82 GiB/s”. Without it you have a lower bound and nothing more.&lt;/p&gt;

&lt;p&gt;Cross-check the knee against Little’s law. Here is the per-node fio output at
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;numjobs=4&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;read: IOPS=6955, BW=6955MiB/s (7292MB/s)(408GiB/60001msec)
    slat (usec): min=12, max=1843, avg=41.28, stdev=22.17
    clat (msec): min=2, max=118, avg=17.94, stdev=9.61
     lat (msec): min=2, max=118, avg=17.98, stdev=9.61
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;128 outstanding per node divided by 17.98 ms is 7,120 IOPS. fio reports 6,955,
and twelve nodes at that rate is 81.5 GiB/s, the table’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;numjobs=4&lt;/code&gt; row. The
two agree to within about two and a half percent, which means the queue was
genuinely full and the latency figure is describing the same system the
bandwidth figure is. When those two disagree badly — Little’s law predicting
double what fio reports, say — the queue was not full and one of the figures is
measuring something else.&lt;/p&gt;

&lt;p&gt;There is a second sweep people skip, and it answers a different question. Sweep
node count at fixed per-node concurrency:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;If one node and twelve nodes each get the same rate, you found a &lt;strong&gt;per-node&lt;/strong&gt;
limit — a NIC, a PCIe link, a single-threaded path.&lt;/li&gt;
  &lt;li&gt;If one node gets the whole budget and twelve nodes divide it, you found an
&lt;strong&gt;aggregate&lt;/strong&gt; limit — the servers, the disks, or the fabric between them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are different problems with different fixes, and a single-point
measurement cannot tell them apart.&lt;/p&gt;

&lt;h3 id=&quot;the-fio-trap-that-silently-caps-you-at-one&quot;&gt;The fio trap that silently caps you at one&lt;/h3&gt;

&lt;p&gt;Setting &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iodepth=64&lt;/code&gt; with a synchronous I/O engine does nothing. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sync&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;psync&lt;/code&gt;
and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pvsync&lt;/code&gt; submit one request and wait for it; fio documents that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iodepth&lt;/code&gt;
beyond 1 has no effect for them. A job file with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ioengine=psync&lt;/code&gt; and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iodepth=64&lt;/code&gt; produces exactly the flat, confident, badly under-driven number
described above, and the command line looks like it asked for depth.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;^(ioengine|iodepth|numjobs)&apos;&lt;/span&gt; bench.fio
&lt;span class=&quot;c&quot;&gt;# ioengine=psync      &amp;lt;-- iodepth below is inert&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# iodepth=64&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# numjobs=8&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;With a synchronous engine, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;numjobs&lt;/code&gt; is your only source of concurrency. With
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;libaio&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;io_uring&lt;/code&gt;, total outstanding is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;numjobs × iodepth&lt;/code&gt; per node.&lt;/p&gt;

&lt;h2 id=&quot;3-a-counter-is-not-a-rate&quot;&gt;3. A counter is not a rate&lt;/h2&gt;

&lt;p&gt;This is among the most common measurement errors in the field and it is entirely
mechanical. Almost every hardware and driver statistic is cumulative since boot,
since daemon start, or since the last reset. Reading it once and dividing by the
duration of your benchmark produces a number with the right units and no
meaning.&lt;/p&gt;

&lt;p&gt;The tell is distinctive: the reported rate &lt;strong&gt;falls as you run the test longer&lt;/strong&gt;,
because the numerator is fixed history and the denominator is your runtime. If
changing only &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;--runtime&lt;/code&gt; changes your answer, you divided a counter by elapsed
time.&lt;/p&gt;

&lt;p&gt;The fix is to sample twice and difference. InfiniBand port counters, using the
extended 64-bit set so they do not wrap mid-run:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;lid&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;37&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;12&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;interval&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;30
&lt;span class=&quot;nv&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;perfquery &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$lid&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$port&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt;: &lt;span class=&quot;s1&quot;&gt;&apos;/^PortXmitData/{gsub(/[^0-9]/,&quot;&quot;,$2); print $2}&apos;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;sleep&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$interval&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;b&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;perfquery &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$lid&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$port&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt;: &lt;span class=&quot;s1&quot;&gt;&apos;/^PortXmitData/{gsub(/[^0-9]/,&quot;&quot;,$2); print $2}&apos;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$a&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;b&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$b&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;t&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$interval&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;s1&quot;&gt;&apos;BEGIN{printf &quot;%.2f GiB/s\n&quot;, (b-a)*4/t/1073741824}&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;23.14 GiB/s
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two details in that snippet are easy to get wrong. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitData&lt;/code&gt; is defined in
the InfiniBand architecture specification in units of four octets, so the delta
is multiplied by 4 to get bytes — the source of a great many capacity claims
that are exactly four times too low. And &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-x&lt;/code&gt; requests the extended
(&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortCountersExtended&lt;/code&gt;) set; the basic counters are 32-bit, which at four
octets per unit is 17.2 GB of traffic before they wrap. On an HDR200 port
carrying 25 GB/s that is a wrap in well under a second, producing a negative
delta and, if your script does not check the sign, a negative or nonsensical
bandwidth. The snippet above does not check it either — add the guard before you
trust it unattended.&lt;/p&gt;

&lt;p&gt;Resist the urge to reach for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt;’s reset flags (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-r&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-R&lt;/code&gt;) first. On a
shared fabric you have just zeroed the inputs to somebody else’s monitoring, and
the difference method does not need it.&lt;/p&gt;

&lt;p&gt;The same discipline applies everywhere. NVMe wear counters, which carry a unit
almost nobody gets right on the first attempt:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;nvme smart-log /dev/nvme0n1 | &lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-iE&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;data.units&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;data_units_read                     : 4,283,915,042
data_units_written                  : 1,118,730,455
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;grep&lt;/code&gt; pattern is deliberately loose because the field label changed:
older &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;nvme-cli&lt;/code&gt; prints &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;data_units_read&lt;/code&gt;, newer versions print
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Data Units Read&lt;/code&gt;, and a pattern pinned to the underscore form silently returns
nothing on half the fleet.&lt;/p&gt;

&lt;p&gt;A data unit here is a thousand 512-byte sectors — 512,000 bytes, not 512 and not
1 MiB. Assume 512 bytes and you are low by a factor of a thousand; assume 1 MiB
and you are high by a factor of 2.05. Sample twice, subtract, multiply by
512,000, divide by elapsed.&lt;/p&gt;

&lt;p&gt;And the ordinary network counters in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/proc/net/dev&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ethtool -S&lt;/code&gt; are
cumulative too. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sar -n DEV 5&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iostat -x 5&lt;/code&gt; do the differencing for you,
which is precisely why the first sample they print — the one covering time since
boot — should always be discarded:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;iostat &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; 5 2 nvme0n1     &lt;span class=&quot;c&quot;&gt;# use the SECOND report, not the first&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Another place the obvious reading is wrong. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iostat&lt;/code&gt; reports &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util&lt;/code&gt;, and on a
spinning disk 100 percent &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util&lt;/code&gt; meant the device was saturated. On NVMe it
means nothing of the kind. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util&lt;/code&gt; is the fraction of wall time during which at
least one request was in flight. An NVMe drive with 64 hardware queues can be at
100 percent &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util&lt;/code&gt; while servicing one request at a time and running at three
percent of its capability.&lt;/p&gt;

  &lt;p&gt;The fields that actually carry information are &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aqu-sz&lt;/code&gt;, the average queue
depth, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;r_await&lt;/code&gt;, the average service latency. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aqu-sz&lt;/code&gt; is the sysstat 12
name; on sysstat 11 and earlier the same field is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;avgqu-sz&lt;/code&gt;, so a parser that
greps for one of them will come back empty on the other. A device with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util=100&lt;/code&gt;,
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aqu-sz=1.2&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;r_await=0.3&lt;/code&gt; ms is nearly idle. The same device at
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aqu-sz=180&lt;/code&gt; with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;r_await&lt;/code&gt; climbing run over run is the constraint. Judge NVMe
by queue depth and latency, never by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util&lt;/code&gt;.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;4-aggregate-cpu-idle-proves-nothing-on-a-many-core-machine&quot;&gt;4. Aggregate CPU idle proves nothing on a many-core machine&lt;/h2&gt;

&lt;p&gt;On a dual-socket server with 96 physical cores and SMT enabled — 192 logical
CPUs — one thread pinned at 100 percent contributes 100 ÷ 192 = 0.52 percent to
the aggregate. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;top&lt;/code&gt; reports 99.5 percent idle. Every dashboard is green. The
job is bottlenecked on a single saturated core and the summary statistic
physically cannot show it.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpstat&lt;/code&gt; has the answer, provided you ask for per-CPU output and an interval.
Run without an interval and it prints averages since boot, which returns you to
problem 3.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;LC_ALL&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;C mpstat &lt;span class=&quot;nt&quot;&gt;-P&lt;/span&gt; ALL 5 1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;14:03:16     CPU    %usr   %nice    %sys %iowait    %irq   %soft  %steal  %idle
14:03:16     all    0.41    0.00    0.22    0.09    0.00    0.51    0.00   98.77
14:03:16       0    0.20    0.00    0.20    0.00    0.00    0.00    0.00   99.60
14:03:16       1    0.40    0.00    0.20    0.00    0.00    0.00    0.00   99.40
14:03:16      37    0.20    0.00    2.61    0.00    0.00   96.99    0.00    0.20
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;all&lt;/code&gt; says 98.8 percent idle. CPU 37 says 99.8 percent busy, essentially all of
it in softirq — and note that CPU 37’s 96.99 percent softirq shows up in the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;all&lt;/code&gt; row as 96.99 ÷ 192 = 0.51, which is the whole of the aggregate &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%soft&lt;/code&gt;
column. The evidence is present in the summary; it is just divided by 192. Sort
for the busiest core rather than reading the summary:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;LC_ALL&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;C mpstat &lt;span class=&quot;nt&quot;&gt;-P&lt;/span&gt; ALL 5 1 | &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;$2 ~ /^[0-9]+$/ {printf &quot;cpu%-4s busy %5.1f%%\n&quot;, $2, $3+$4+$5+$7+$8}&apos;&lt;/span&gt; | &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nb&quot;&gt;sort&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-k3&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-nr&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;head&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;cpu37   busy  99.8%
cpu112  busy  41.2%
cpu38   busy   3.1%
cpu4    busy   1.1%
cpu0    busy   0.4%
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LC_ALL=C&lt;/code&gt; is load-bearing rather than decorative. In several locales sysstat
prints a 12-hour timestamp with a trailing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AM&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PM&lt;/code&gt;, which puts an extra field
at the front of every row and shifts &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$3&lt;/code&gt; onto &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%usr&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$4&lt;/code&gt; onto &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%nice&lt;/code&gt; and so
on — the awk above then sums the wrong columns and reports a plausible, wrong
number. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;S_TIME_FORMAT=ISO&lt;/code&gt; achieves the same fixed layout if you would rather
keep the locale.&lt;/p&gt;

&lt;p&gt;Summing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%usr + %nice + %sys + %irq + %soft&lt;/code&gt; rather than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;100 − %idle&lt;/code&gt; keeps
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%iowait&lt;/code&gt; out of the total. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%iowait&lt;/code&gt; is not CPU work; it is idle time that
happened to coincide with a blocked task, and on a 192-core box it is diluted by
exactly the same factor that hides the busy core.&lt;/p&gt;

&lt;p&gt;A single core at 96 percent softirq looks like receive-side packet processing
landing on one CPU. Confirm before acting:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;grep&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-E&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;mlx5|nvme&apos;&lt;/span&gt; /proc/interrupts | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;{print $1, $(NF)}&apos;&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;head
awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;NR==1 || /NET_RX/&apos;&lt;/span&gt; /proc/softirqs | &lt;span class=&quot;nb&quot;&gt;cut&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c1-120&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;Do not stop at “the interrupts are unbalanced, run irqbalance”. Read the column,
not just the total. On an RDMA path the payload does not traverse the kernel
network stack, so &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%soft&lt;/code&gt; on a verbs benchmark is rarely &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;NET_RX&lt;/code&gt; — it is more
often completion-queue work landing on the single CPU the queue pair’s vector
was bound to, which &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;irqbalance&lt;/code&gt; will not move because RDMA drivers pin their
own vectors. And if the hot column is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%sys&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%usr&lt;/code&gt; rather than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%soft&lt;/code&gt;, the
busy core is not interrupt handling at all: it is a single-threaded submission
path, in the benchmark or in a filesystem daemon. Establish which before you
rebind anything:&lt;/p&gt;

  &lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;top &lt;span class=&quot;nt&quot;&gt;-H&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;$(&lt;/span&gt;pgrep &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt;, &lt;span class=&quot;nt&quot;&gt;-f&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;fio|mmfsd&apos;&lt;/span&gt;&lt;span class=&quot;si&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-b&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-n1&lt;/span&gt; | &lt;span class=&quot;nb&quot;&gt;head&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;  &lt;/div&gt;

  &lt;p&gt;If the hot thread belongs to your benchmark rather than to an interrupt, the
fix is more jobs, not IRQ affinity — and you are back at problem 2. Rebinding
interrupts to fix a problem that was never interrupts is how a tuning exercise
consumes a maintenance window and changes nothing.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;state-the-expected-value-before-you-measure&quot;&gt;State the expected value before you measure&lt;/h2&gt;

&lt;p&gt;The four mechanisms above share one property: in every case the wrong number
looked plausible. That is what made it survive. The defense is to decide what
the number should be &lt;em&gt;before&lt;/em&gt; you run anything, write the arithmetic down, and
treat agreement as weak evidence and disagreement as a finding.&lt;/p&gt;

&lt;p&gt;Build the budget layer by layer, cheapest measurement first. For a tier of eight
storage servers on HDR200 serving twelve clients:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Layer&lt;/th&gt;
      &lt;th&gt;Per unit&lt;/th&gt;
      &lt;th&gt;Count&lt;/th&gt;
      &lt;th&gt;Aggregate&lt;/th&gt;
      &lt;th&gt;How it was obtained&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Server DDR read bandwidth&lt;/td&gt;
      &lt;td&gt;~190 GiB/s&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;~1.5 TiB/s&lt;/td&gt;
      &lt;td&gt;8-channel DDR4-3200 per socket, arithmetic&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NVMe random 128k read, per server&lt;/td&gt;
      &lt;td&gt;62 GiB/s&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;496 GiB/s&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fio&lt;/code&gt; against raw devices, measured&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;HCA PCIe link (Gen4 ×16)&lt;/td&gt;
      &lt;td&gt;29.3 GiB/s&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;234 GiB/s&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;current_link_speed&lt;/code&gt; × &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;_width&lt;/code&gt;, 128b/130b, arithmetic&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Storage server NIC (HDR200, 200 Gb/s)&lt;/td&gt;
      &lt;td&gt;23.3 GiB/s&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;186 GiB/s&lt;/td&gt;
      &lt;td&gt;line rate, arithmetic&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Inter-switch links, under load&lt;/td&gt;
      &lt;td&gt;23.1 GiB/s&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
      &lt;td&gt;185 GiB/s&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt; delta, measured&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Client NIC (HDR200)&lt;/td&gt;
      &lt;td&gt;23.3 GiB/s&lt;/td&gt;
      &lt;td&gt;12&lt;/td&gt;
      &lt;td&gt;280 GiB/s&lt;/td&gt;
      &lt;td&gt;line rate, arithmetic&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Single client alone, all 12 idle&lt;/td&gt;
      &lt;td&gt;22.9 GiB/s&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;—&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fio&lt;/code&gt; on one node, measured&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Erasure-code read amplification&lt;/td&gt;
      &lt;td&gt;×1.88&lt;/td&gt;
      &lt;td&gt;—&lt;/td&gt;
      &lt;td&gt;99 GiB/s delivered&lt;/td&gt;
      &lt;td&gt;modeled from 8+2p, checked against &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The PCIe row is worth a caution, because it is the row most often quoted as a
measurement when it is nothing of the kind. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;current_link_speed&lt;/code&gt; and
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;current_link_width&lt;/code&gt; under &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/sys/bus/pci/devices/&lt;/code&gt; report what the link
negotiated, not what it carries: Gen4 ×16 at 16 GT/s with 128b/130b framing is
31.5 GB/s, or 29.3 GiB/s, and no adapter delivers that after TLP headers and
DMA overhead. Read it as a ceiling that is comfortably above the NIC line rate,
which is the only conclusion the row has to support here.&lt;/p&gt;

&lt;p&gt;The binding constraint for a cold read is the storage-side NIC aggregate,
reduced by the read amplification of the erasure code. With &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;k&lt;/code&gt; data strips
distributed one per node, delivering one byte to a client moves roughly
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;1 + (k−1)/k&lt;/code&gt; bytes on the storage fabric: the node fronting the client holds
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;1/k&lt;/code&gt; of the track locally, pulls the other &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(k−1)/k&lt;/code&gt; across the fabric from its
peers, and then sends the whole byte on to the client. For k = 8 that is 1.875.
186 ÷ 1.88 is 99 GiB/s of deliverable read bandwidth. The measured 82 GiB/s is
83 percent of that, which is a believable efficiency for a real filesystem and
is therefore not interesting.&lt;/p&gt;

&lt;p&gt;This is a model, not a measurement, and it is the weakest row in the table. It
assumes full-track reads with no coalescing and no parity-strip traffic on a
healthy array, and it will be wrong under rebuild, wrong for partial-track
reads, and wrong for any layout that places more than one strip per node. Treat
it as the right order of magnitude for the correction rather than a constant,
and confirm it by comparing fabric bytes from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt; against client bytes
from IOR on the same run.&lt;/p&gt;

&lt;p&gt;The 155.8 GiB/s figure is 157 percent of that budget. It is not physically
impossible — it sits at 84 percent of the raw 186 GiB/s NIC aggregate — but it
can only be reached by removing a layer from the path, and the excess tells you
which layer. A buffer-pool hit serves the reconstructed logical track from the
server’s own memory, so it skips both the NVMe read and the inter-node strip
gather, and the ×1.88 amplification disappears with them. The number went above
the budget because the budget’s most expensive term stopped applying. Reading it
as “the storage is faster than we modeled” inverts the finding exactly.&lt;/p&gt;

&lt;p&gt;That inversion is the whole discipline. A number that exceeds the budget is the
loudest alarm your benchmark can raise, and it is the one people are least likely
to investigate, because it is the number they wanted.&lt;/p&gt;

&lt;p&gt;Two unit traps will move a figure enough to matter. One GB is 10⁹ bytes and one
GiB is 2³⁰; the gap is 7.37 percent, and one step down the scale the MiB-to-MB
gap is 4.86 percent. fio prints both — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;BW=6955MiB/s (7292MB/s)&lt;/code&gt; — and IOR
prints MiB/s only. Take fio’s binary figure, relabel it with the decimal unit,
and you understate by 4.9 percent at the MiB level or 7.4 percent at the GiB
level: roughly the size of the improvements people schedule maintenance windows
for. And link rates are decimal: HDR200 is 200 × 10⁹ bit/s, which is 25.0 GB/s
and 23.3 GiB/s, not 25.&lt;/p&gt;

&lt;h2 id=&quot;symptom-against-true-cause&quot;&gt;Symptom against true cause&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Symptom&lt;/th&gt;
      &lt;th&gt;Tempting conclusion&lt;/th&gt;
      &lt;th&gt;Also consistent with&lt;/th&gt;
      &lt;th&gt;Check that distinguishes them&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Read ≫ write on symmetric hardware&lt;/td&gt;
      &lt;td&gt;Read path is better optimized&lt;/td&gt;
      &lt;td&gt;Read served from the server buffer pool&lt;/td&gt;
      &lt;td&gt;Re-run at 2× and 4× the pool size; cache gives a sharp fall then a floor&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Read phase finished in under 10 s&lt;/td&gt;
      &lt;td&gt;Fast storage&lt;/td&gt;
      &lt;td&gt;Working set fits in cache&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;total(s)&lt;/code&gt; in IOR; raise &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-b&lt;/code&gt; until the phase runs 60 s&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Throughput flat across repeated runs&lt;/td&gt;
      &lt;td&gt;Saturation&lt;/td&gt;
      &lt;td&gt;Under-driven at a fixed, insufficient depth&lt;/td&gt;
      &lt;td&gt;Sweep &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;numjobs&lt;/code&gt;; a real ceiling is flat while concurrency quadruples&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Throughput flat as nodes are added&lt;/td&gt;
      &lt;td&gt;Aggregate ceiling&lt;/td&gt;
      &lt;td&gt;Per-node NIC or single-thread limit&lt;/td&gt;
      &lt;td&gt;Compare 1-node and N-node per-node rates&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iodepth&lt;/code&gt; raised, nothing changed&lt;/td&gt;
      &lt;td&gt;Queue already full&lt;/td&gt;
      &lt;td&gt;Synchronous ioengine ignoring &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iodepth&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;grep ioengine&lt;/code&gt; in the job file&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util&lt;/code&gt; at 100 percent&lt;/td&gt;
      &lt;td&gt;Device saturated&lt;/td&gt;
      &lt;td&gt;One request in flight on a 64-queue device&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aqu-sz&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;r_await&lt;/code&gt;, not &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;%util&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Rate falls as runtime rises&lt;/td&gt;
      &lt;td&gt;Thermal or cache-fill effect&lt;/td&gt;
      &lt;td&gt;Cumulative counter divided by elapsed&lt;/td&gt;
      &lt;td&gt;Sample the counter twice and difference&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Bandwidth exactly 4× too low&lt;/td&gt;
      &lt;td&gt;Link running at 1X&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitData&lt;/code&gt; read without the ×4&lt;/td&gt;
      &lt;td&gt;Multiply the delta by 4; confirm width with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibstatus&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CPU 98 percent idle under load&lt;/td&gt;
      &lt;td&gt;Not CPU-bound&lt;/td&gt;
      &lt;td&gt;One core at 100 percent out of 192&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mpstat -P ALL&lt;/code&gt;, sort by busiest core&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Total rises as clients are added&lt;/td&gt;
      &lt;td&gt;Storage scales&lt;/td&gt;
      &lt;td&gt;You were client-limited the whole time&lt;/td&gt;
      &lt;td&gt;Per-node rate at 1 node versus N nodes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Second run much faster than the first&lt;/td&gt;
      &lt;td&gt;Warm-up effect&lt;/td&gt;
      &lt;td&gt;Second run read what the first run wrote&lt;/td&gt;
      &lt;td&gt;IOR &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-C&lt;/code&gt;, and drop client caches between phases&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Bandwidth and latency both healthy, job still slow&lt;/td&gt;
      &lt;td&gt;Storage is fine&lt;/td&gt;
      &lt;td&gt;Metadata-bound, not bandwidth-bound&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;mdtest&lt;/code&gt; create/stat/remove rates&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Array tool reports twice what &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fio&lt;/code&gt; reports&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fio&lt;/code&gt; is misconfigured&lt;/td&gt;
      &lt;td&gt;Fabric counters include the strip gather&lt;/td&gt;
      &lt;td&gt;Compare the ratio against &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;1 + (k−1)/k&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Result matches the vendor datasheet&lt;/td&gt;
      &lt;td&gt;Configured correctly&lt;/td&gt;
      &lt;td&gt;Measured a cache, or copied the expectation&lt;/td&gt;
      &lt;td&gt;Compute the budget independently first&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;h2 id=&quot;what-to-keep&quot;&gt;What to keep&lt;/h2&gt;

&lt;p&gt;Three habits close most of the gap.&lt;/p&gt;

&lt;p&gt;Write the expected number down before the run, with the arithmetic that produced
it, in the same file as the result. This costs five minutes and converts every
benchmark from a measurement into a test with a pass condition.&lt;/p&gt;

&lt;p&gt;Never report a single point. Report a sweep, and report the shape. “82 GiB/s”
is a claim; “82 GiB/s, flat within one percent from 1,536 to 6,144 outstanding
I/Os, at 3 TiB against a 241 GiB server buffer pool” is a result somebody else
can check.&lt;/p&gt;

&lt;p&gt;And when the number comes out better than you expected, stop and find out why
before telling anyone. In the shape of problem described here, that is nearly
always the moment the measurement broke.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Management networks: isolate the systems that can rebuild everything</title>
    <link href="https://www.wirewalk.com/writing/management-network-bmc-isolation/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/management-network-bmc-isolation/</id>
    <summary>BMCs, hypervisors, storage controllers, and provisioning services belong in the privileged-access model.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A management interface often has authority below the operating system’s security controls. A BMC can affect power and console access; a provisioning service can replace a host image. These paths deserve their own network and identity review.&lt;/p&gt;

&lt;h2 id=&quot;inventory-privileged-control-planes&quot;&gt;Inventory privileged control planes&lt;/h2&gt;

&lt;p&gt;Include Dell iDRAC, HPE iLO, hypervisor managers, storage administration, switch management, backup consoles, and image repositories. Record the supported firmware or software release and the intended management route.&lt;/p&gt;

&lt;p&gt;Do not discover authority solely by scanning ports. Some capabilities exist behind a shared API or reverse proxy, while an open port says little about who can use it. Pair network observations with configuration and role review.&lt;/p&gt;

&lt;h2 id=&quot;specify-allowed-flows&quot;&gt;Specify allowed flows&lt;/h2&gt;

&lt;p&gt;Use an explicit matrix rather than “management VLAN accessible to IT”:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Source&lt;/th&gt;
      &lt;th&gt;Destination&lt;/th&gt;
      &lt;th&gt;Purpose&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Approved admin workstation/bastion&lt;/td&gt;
      &lt;td&gt;Selected management interface&lt;/td&gt;
      &lt;td&gt;Interactive administration&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Monitoring service&lt;/td&gt;
      &lt;td&gt;Read-only management endpoint&lt;/td&gt;
      &lt;td&gt;Health collection&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Provisioning service&lt;/td&gt;
      &lt;td&gt;Defined hosts and repositories&lt;/td&gt;
      &lt;td&gt;Controlled image lifecycle&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Ordinary user network&lt;/td&gt;
      &lt;td&gt;Management interface&lt;/td&gt;
      &lt;td&gt;Denied unless specifically justified&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;This is a design example. Actual protocols and ports come from the installed product documentation. Some products need additional flows for discovery, virtual media, or telemetry.&lt;/p&gt;

&lt;h2 id=&quot;validate-identity-and-transport&quot;&gt;Validate identity and transport&lt;/h2&gt;

&lt;p&gt;Redfish is a standardized management API, but implementations, authentication, and available resources differ. Use &lt;a href=&quot;https://www.dmtf.org/standards/redfish&quot;&gt;DMTF’s Redfish specifications&lt;/a&gt; and the hardware vendor’s release documentation. Do not assume one manufacturer’s endpoint layout or role model applies to another.&lt;/p&gt;

&lt;p&gt;Use named administrative identities where supported, restrict automation to required operations, and validate certificates. An emergency script that disables TLS verification permanently removes an important check from a highly privileged path.&lt;/p&gt;

&lt;h2 id=&quot;test-isolation-without-changing-host-power&quot;&gt;Test isolation without changing host power&lt;/h2&gt;

&lt;p&gt;From an approved administrator path, verify an authorized read-only health operation. From a representative ordinary client network, confirm the management service is unreachable or access is denied as designed. Test IPv4 and IPv6 where both are deployed.&lt;/p&gt;

&lt;p&gt;Then verify emergency console access to a pilot host under a documented procedure. Do not issue resets or power operations simply to test API authentication.&lt;/p&gt;

&lt;h2 id=&quot;keep-management-available-during-an-outage&quot;&gt;Keep management available during an outage&lt;/h2&gt;

&lt;p&gt;Avoid making the only recovery route depend on the application network being repaired. Document access during firewall failure, directory outage, and loss of the primary bastion. Protect recovery credentials separately and monitor their use.&lt;/p&gt;

&lt;p&gt;A segmented network with one shared administrator password still has a concentrated credential risk. Conversely, excellent identity controls do not justify exposing management services unnecessarily.&lt;/p&gt;

&lt;h2 id=&quot;acceptance-record&quot;&gt;Acceptance record&lt;/h2&gt;

&lt;p&gt;Keep the inventory, allowed-flow matrix, read-only access tests, denied-path tests, firmware review, and emergency-access procedure. Repeat after network changes and hardware additions. The management plane is a moving estate, not a one-time VLAN configuration.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>When a storage read ceiling is really an InfiniBand routing problem</title>
    <link href="https://www.wirewalk.com/writing/when-storage-is-actually-the-fabric/"/>
    <published>2026-02-11T10:00:00-05:00</published>
    <updated>2026-02-11T10:00:00-05:00</updated>
    <id>https://www.wirewalk.com/writing/when-storage-is-actually-the-fabric/</id>
    <summary>How to tell a parallel filesystem read ceiling apart from a static-routing collapse on the inter-switch links, using per-ISL counter deltas and a network-only corroboration test. Includes the elimination order and the one command that would have answered it on the first morning.</summary>
    <content type="html">&lt;p&gt;There is a symptom shape that sends people looking in the wrong layer for a week.
A parallel filesystem writes close to the number the design predicted and reads
at roughly half that. Nothing is saturated. Client adapters sit well below half
their line rate, the storage servers are not CPU-bound, the disks are not busy,
and fabric monitoring shows aggregate utilization in the twenties. Every component
reports headroom and the aggregate refuses to move. On the storage servers,
during the read phase that is supposedly the problem:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ iostat -x 5 | grep -E &apos;Device|nvme0n1&apos;   # write-side columns trimmed
Device   r/s     rkB/s   rrqm/s  %rrqm r_await rareq-sz  aqu-sz  %util
nvme0n1  812.4  830617.6    0.0   0.00    0.41   1022.6    0.33  31.20
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Thirty-one percent device utilization and sub-millisecond read latency, with a
queue that never gets deeper than a third of one outstanding request. The disks
are not the constraint.&lt;/p&gt;

&lt;p&gt;The reflex is to call it a storage problem, because the number you are unhappy
with came out of a storage benchmark. That reflex is usually wrong when the
asymmetry is this clean. Read and write paths share almost everything. The short
list of genuine differences is readahead, the disk-side read/write mix, and the
direction the bytes travel. Rule out the first two and the third is not an exotic
explanation, it is the only one left.&lt;/p&gt;

&lt;h2 id=&quot;why-the-fabric-is-a-plausible-suspect-at-all&quot;&gt;Why the fabric is a plausible suspect at all&lt;/h2&gt;

&lt;p&gt;InfiniBand unicast routing, in the configuration almost everyone runs, is static
and destination-based. Each switch holds a linear forwarding table mapping a
destination LID to exactly one output port. Not a set of ports, not a hash over
flows — one port. Every packet addressed to that LID leaves through that cable
for as long as the table stands.&lt;/p&gt;

&lt;p&gt;Where two switches are joined by several inter-switch links, the subnet manager
decides which link each destination uses. OpenSM’s default engine is Min Hop,
and its balancing rule is in its own documentation: among ports offering the same
hop count, pick the one with fewer LIDs already assigned. That is round-robin
over destinations.&lt;/p&gt;

&lt;p&gt;Read that rule again, because the whole article is in it. The engine balances the
&lt;strong&gt;number of destination addresses&lt;/strong&gt; per output port. It has no idea which of
those addresses carry traffic, how much, or when. Two ports holding eight LIDs
each are equally loaded as far as the routing engine is concerned, even if all
the traffic in the cluster is addressed to three LIDs that happen to share one
port.&lt;/p&gt;

&lt;p&gt;Apply that to a filesystem and the two directions stop being symmetric:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Writes&lt;/strong&gt; are addressed to the storage servers. Eight of them over eight
inter-switch links, and round-robin over eight consecutive GUIDs puts one
storage LID on each. The spread is perfect. That is not good engineering, it is
small-&lt;em&gt;n&lt;/em&gt; luck — reliable luck, which is why nobody notices.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reads&lt;/strong&gt; are addressed to the clients, drawn from a much larger pool: every
compute node in the subnet, not just the ones in this job. Where a given client
lands in the global round-robin has nothing to do with whether it is running
your job. If the spacing between your job’s nodes in the routing order aliases
against the number of links, groups of them share one cable and other cables
carry nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The aliasing is not hypothetical. OpenSM ships an option whose purpose is to
break it — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scatter_ports&lt;/code&gt;, documented as randomizing port selection “rather than
using a round-robin algorithm (which is the default)”. Options do not get written
for problems nobody has.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;I would not try to predict from first principles which links a given fabric
collapses onto. The alignment depends on GUID ordering, discovery order, how many
LIDs sit behind the same links, and whether the tables were computed before or
after the last cable was plugged in. Measure it; the measurement is cheap and the
prediction is not.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;finding-the-inter-switch-links&quot;&gt;Finding the inter-switch links&lt;/h2&gt;

&lt;p&gt;Before any counters, establish which physical ports are ISLs and confirm they run
at the width and rate you think. One link that came up at 1X instead of 4X
produces a similar-looking aggregate shortfall from a completely different cause,
and it is embarrassing to find in week two.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ ibswitches
Switch  : 0x043f720300xxxxxx ports 40 &quot;SwitchA&quot; enhanced port 0 lid 1 lmc 0
Switch  : 0x043f720300yyyyyy ports 40 &quot;SwitchB&quot; enhanced port 0 lid 2 lmc 0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ iblinkinfo -S 0x043f720300yyyyyy | grep -i switch
       2   25[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   13[  ] &quot;SwitchA&quot;
       2   26[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   14[  ] &quot;SwitchA&quot;
       2   27[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   15[  ] &quot;SwitchA&quot;
       2   28[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   16[  ] &quot;SwitchA&quot;
       2   29[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   17[  ] &quot;SwitchA&quot;
       2   30[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   18[  ] &quot;SwitchA&quot;
       2   31[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   19[  ] &quot;SwitchA&quot;
       2   32[  ] ==( 4X  53.125 Gbps Active/  LinkUp)==&amp;gt;  1   20[  ] &quot;SwitchA&quot;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Eight links, all 4X, all at the same per-lane rate. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo&lt;/code&gt; prints the
&lt;strong&gt;per-lane&lt;/strong&gt; figure, not the link figure: 53.125 Gbps across four lanes is HDR, a
200 Gb/s link. Read that column as the link rate and you will invent a capacity
problem that does not exist.&lt;/p&gt;

&lt;p&gt;The second trap is to take the encoding overhead off twice. HDR uses 64b/66b line
coding, but the 200 Gb/s that everyone quotes for 4X HDR is the IBTA &lt;em&gt;data&lt;/em&gt; rate,
already net of that coding — the per-lane signalling rate is the 53.125 Gbps in
the column above. The same relationship holds one generation down, where EDR’s
25.78125 Gbps per lane × 64/66 is exactly 25, and 4X EDR is exactly 100 Gb/s.
Dividing 200 by 66/64 a second time gets you 193.9 Gb/s and a ceiling that is 3%
too low for no reason.&lt;/p&gt;

&lt;p&gt;So: 200 Gb/s is 25 GB/s, or 23.3 GiB/s, and that is the theoretical figure.
Subtract IB transport headers, ICRC and flow control and a healthy 4X HDR link
delivers around &lt;strong&gt;22.6 GiB/s&lt;/strong&gt; of payload in practice, which is roughly what
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; returns on an idle one. That 22.6 is the per-link ceiling used
throughout this article, and eight links give about 180.8 GiB/s each way. It is a
measured number rather than a derived one, so measure it on your own hardware
before you quote a percentage against it.&lt;/p&gt;

&lt;h2 id=&quot;reading-per-isl-counters&quot;&gt;Reading per-ISL counters&lt;/h2&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt; reads the performance management agent on a switch port. Query the
ISL ports on one switch and you get both directions from one place:
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitData&lt;/code&gt; leaves that port toward the far switch, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortRcvData&lt;/code&gt; arrives from
it. On a switch at the storage side, reads are &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitData&lt;/code&gt; and writes are
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortRcvData&lt;/code&gt;. Use the extended counters; at these rates it is not optional:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ perfquery -x 2 25
# Port extended counters: Lid 2 port 25
PortSelect:......................25
CounterSelect:...................0x0000
PortXmitData:....................347892350976
PortRcvData:.....................183609851904
PortXmitPkts:....................339741882
PortRcvPkts:.....................179309714
PortUnicastXmitPkts:.............339740115
PortUnicastRcvPkts:..............179307946
PortMulticastXmitPkts:...........1767
PortMulticastRcvPkts:............1768
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Two things about those numbers before you divide anything by anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The data counters are in units of four octets.&lt;/strong&gt; The IBA specification defines
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitData&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortRcvData&lt;/code&gt; as the number of data octets divided by four,
summed over all virtual lanes. Multiply by 4 to get bytes. Forget this and you get a
number four times too small — comfortingly close to a quarter of link rate, and
exactly the kind of wrong answer that survives a sanity check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The basic 32-bit counters wrap in well under a second.&lt;/strong&gt; A 32-bit counter in
units of 4 octets tops out at 2^32 × 4 = 17.18 GB, or 16 GiB. On an HDR link
carrying 21.6 GiB/s that is 0.74 seconds; on NDR, half that. Sample the basic
counters over a 60-second window and you are measuring the remainder after eighty
wraps. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-x&lt;/code&gt; flag, and only that flag, gives the 64-bit versions. The extended
counter attribute is optional in the specification — a device advertises it in
the performance class &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;CapabilityMask&lt;/code&gt; — but present on anything you are likely
to be buying.&lt;/p&gt;

&lt;p&gt;One counter does not come along for the ride. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt;, which matters later,
lives in the basic &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortCounters&lt;/code&gt; attribute and not in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortCountersExtended&lt;/code&gt;, so
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery -x&lt;/code&gt; will not show it. Read it with plain &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt;, and remember it
is 32 bits wide like everything else in that attribute.&lt;/p&gt;

&lt;h2 id=&quot;sampling-twice-and-differencing&quot;&gt;Sampling twice and differencing&lt;/h2&gt;

&lt;p&gt;Counters are cumulative. The measurement is a difference over a known interval
while a known workload runs. Do not reset them if anything else on the site reads
them; you will corrupt a monitoring baseline, and the delta method does not need
a reset.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;#!/bin/bash&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# per-ISL throughput, both directions, over a fixed window&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;SW_LID&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;2
&lt;span class=&quot;nv&quot;&gt;PORTS&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;25 26 27 28 29 30 31 32&quot;&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;WINDOW&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;60

sample&lt;span class=&quot;o&quot;&gt;()&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;{&lt;/span&gt;
  &lt;span class=&quot;k&quot;&gt;for &lt;/span&gt;p &lt;span class=&quot;k&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;$PORTS&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;do
    &lt;/span&gt;perfquery &lt;span class=&quot;nt&quot;&gt;-x&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$SW_LID&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$p&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
      | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$p&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;
          /^PortXmitData:/ { gsub(/[^0-9]/,&quot;&quot;,$0); x=$0 }
          /^PortRcvData:/  { gsub(/[^0-9]/,&quot;&quot;,$0); r=$0 }
          END { print p, x, r }&apos;&lt;/span&gt;
  &lt;span class=&quot;k&quot;&gt;done&lt;/span&gt;
&lt;span class=&quot;o&quot;&gt;}&lt;/span&gt;

sample &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; /tmp/isl.t0
&lt;span class=&quot;nb&quot;&gt;sleep&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$WINDOW&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;
sample &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; /tmp/isl.t1

&lt;span class=&quot;nb&quot;&gt;join&lt;/span&gt; /tmp/isl.t0 /tmp/isl.t1 | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; &lt;span class=&quot;nv&quot;&gt;w&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$WINDOW&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;
  { dx=($4-$2)*4; dr=($5-$3)*4;
    printf &quot;port %-3s  xmit %7.2f GiB/s   rcv %7.2f GiB/s\n&quot;,
           $1, dx/w/1073741824, dr/w/1073741824 }&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gsub&lt;/code&gt; is there because &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt; pads with dots rather than whitespace, so
the value is not in a predictable field. Parsing it as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;$2&lt;/code&gt; works on some
releases and silently returns zero on others.&lt;/p&gt;

&lt;p&gt;Run it twice: once under a write-only workload, once under a read-only workload
over a dataset too large for client or server page cache. If the read set fits in
cache you measure memory, get a spectacular number and learn nothing.&lt;/p&gt;

&lt;h2 id=&quot;the-table-that-ends-the-argument&quot;&gt;The table that ends the argument&lt;/h2&gt;

&lt;p&gt;A twelve-client run against eight storage servers, sampled over 60 seconds on the
storage-side switch. The left columns are measured; the right two come from the
forwarding table, which comes next.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;ISL port&lt;/th&gt;
      &lt;th&gt;Write run, GiB/s&lt;/th&gt;
      &lt;th&gt;Read run, GiB/s&lt;/th&gt;
      &lt;th&gt;Read-run link use&lt;/th&gt;
      &lt;th&gt;Active client LIDs routed here&lt;/th&gt;
      &lt;th&gt;Total LFT entries via this port&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;25&lt;/td&gt;
      &lt;td&gt;11.2&lt;/td&gt;
      &lt;td&gt;21.6&lt;/td&gt;
      &lt;td&gt;96%&lt;/td&gt;
      &lt;td&gt;6&lt;/td&gt;
      &lt;td&gt;9&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;26&lt;/td&gt;
      &lt;td&gt;11.6&lt;/td&gt;
      &lt;td&gt;21.3&lt;/td&gt;
      &lt;td&gt;94%&lt;/td&gt;
      &lt;td&gt;5&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;27&lt;/td&gt;
      &lt;td&gt;11.4&lt;/td&gt;
      &lt;td&gt;7.4&lt;/td&gt;
      &lt;td&gt;33%&lt;/td&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;9&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;28&lt;/td&gt;
      &lt;td&gt;11.3&lt;/td&gt;
      &lt;td&gt;0.0&lt;/td&gt;
      &lt;td&gt;0%&lt;/td&gt;
      &lt;td&gt;0&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;29&lt;/td&gt;
      &lt;td&gt;11.5&lt;/td&gt;
      &lt;td&gt;0.0&lt;/td&gt;
      &lt;td&gt;0%&lt;/td&gt;
      &lt;td&gt;0&lt;/td&gt;
      &lt;td&gt;9&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;30&lt;/td&gt;
      &lt;td&gt;11.4&lt;/td&gt;
      &lt;td&gt;0.0&lt;/td&gt;
      &lt;td&gt;0%&lt;/td&gt;
      &lt;td&gt;0&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;31&lt;/td&gt;
      &lt;td&gt;11.2&lt;/td&gt;
      &lt;td&gt;0.0&lt;/td&gt;
      &lt;td&gt;0%&lt;/td&gt;
      &lt;td&gt;0&lt;/td&gt;
      &lt;td&gt;9&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;32&lt;/td&gt;
      &lt;td&gt;11.8&lt;/td&gt;
      &lt;td&gt;0.0&lt;/td&gt;
      &lt;td&gt;0%&lt;/td&gt;
      &lt;td&gt;0&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;91.4&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;50.3&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;28% mean&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;12&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;68&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Read the write column first. Eight links carrying 11.2 to 11.8 GiB/s, every one
at about half its 22.6 GiB/s ceiling. That is a healthy fabric, and it is why
nobody suspected the fabric: anyone who looked at ISL utilization during a write
test saw 50% and moved on.&lt;/p&gt;

&lt;p&gt;Now the read column. Five of eight cables carry nothing. Two carry 21.6 and 21.3
GiB/s, which is 96% and 94% of what an HDR link delivers as payload. Those two
are full. The mean across all eight is 28%, and 28% is what an aggregate fabric
graph shows, which is why an aggregate fabric graph cannot find this. The unit of
scarcity is a cable, and averaging over cables destroys the information you
need.&lt;/p&gt;

&lt;h3 id=&quot;the-row-that-makes-it-certain&quot;&gt;The row that makes it certain&lt;/h3&gt;

&lt;p&gt;Port 27 is the most useful row in the table. One client is routed over it, and it
reads at 7.4 GiB/s. In the write run twelve clients moved 91.4 GiB/s, or 7.62
GiB/s each. A client with an uncontended path therefore reads within 3% of what
it writes. There is nothing wrong with readahead, prefetch, the disks, the
servers or the client: a client that gets its own cable performs as designed.&lt;/p&gt;

&lt;p&gt;That row turns “the fabric is suspicious” into “the fabric is the constraint”,
and it lets you predict the aggregate rather than merely describe it. Build the
prediction only from quantities that did not come out of the read run itself,
or you are adding up the answer and calling it a forecast. Two of them qualify:
the per-link payload ceiling, established on idle hardware, and the per-client
rate from the write run.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;two ISLs at payload ceiling    2 × 22.6 = 45.2 GiB/s
one uncontended client (write-run rate) =  7.6 GiB/s
                              predicted   52.8 GiB/s
                               measured   50.3 GiB/s   (-4.8%)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Within 5%, from two numbers neither of which was measured during the read test,
with the shortfall in the direction you would expect — a shared link running a
little under its idle ceiling. That is a different class of evidence from a
plausible story. Adding the measured read column to 50.3 is arithmetic, not a
prediction, and it is worth keeping the two apart when someone challenges the
result.&lt;/p&gt;

&lt;h2 id=&quot;the-obvious-interpretation-and-why-it-is-wrong&quot;&gt;The obvious interpretation, and why it is wrong&lt;/h2&gt;

&lt;p&gt;Everyone seeing that table for the first time says the same thing: five cables
are dead, so five cables are broken, check the transceivers.&lt;/p&gt;

&lt;p&gt;They are not broken. Look one column left. In the write run those same five
cables carried 11.2 to 11.8 GiB/s each, indistinguishable from the two that
saturate on reads. The link is up, at 4X, at the right rate, with no errors:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ ibqueryerrors -c
## Summary: 22 nodes checked, 56 ports checked, 0 ports have errors beyond threshold
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibqueryerrors&lt;/code&gt; with no target sweeps the fabric and prints only the ports that
have something wrong, which is the behavior you want here. Two of its flags read
backwards from the obvious guess, and both will cost you: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-s&lt;/code&gt; takes a list of
counters to &lt;strong&gt;suppress&lt;/strong&gt;, not to select, so naming the counters you care about is
the one way to guarantee you do not see them. And &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-k&lt;/code&gt; / &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-K&lt;/code&gt; clear error and data
counters as they read — run those on a production fabric and you have silently
reset the baseline of whatever else monitors it, which is the same mistake the
counter-differencing section warns against. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-c&lt;/code&gt;, used above, suppresses the
common side-effect counters and clears nothing.&lt;/p&gt;

&lt;p&gt;A cable carries nothing on reads because &lt;strong&gt;no destination LID with read traffic
is mapped to it&lt;/strong&gt;, and a full share on writes because storage LIDs are. Same
cable, same hour, two different pictures depending on which way the bytes go.&lt;/p&gt;

&lt;p&gt;The second wrong interpretation is subtler and costs more. “No link is saturated,
therefore the fabric is not the bottleneck.” Sound when links are used uniformly;
useless here, because the reported statistic uses a denominator that includes
five cables the traffic cannot reach. The correct denominator is the capacity the
routing tables permit: 3 × 22.6 = 67.8 GiB/s, of which the run achieves 50.3.
Seventy-four percent of a hard ceiling with two of three links at 95% is a
saturated network.&lt;/p&gt;

&lt;h2 id=&quot;corroboration-without-a-disk-in-the-path&quot;&gt;Corroboration without a disk in the path&lt;/h2&gt;

&lt;p&gt;You now have a fabric-shaped explanation built entirely from fabric counters.
Before acting on it, prove the ceiling exists with the filesystem out of the
picture. If a pure RDMA test between the same endpoints, in the same direction,
with no block device in the path, lands near the filesystem’s number, the
bottleneck is at or below the network and no storage tuning will move it.&lt;/p&gt;

&lt;p&gt;Run RDMA writes &lt;em&gt;from&lt;/em&gt; the storage nodes &lt;em&gt;to&lt;/em&gt; the clients, so bytes travel in the
read direction. Start all pairs together; one pair at a time tells you nothing
about aggregate behavior.&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# on each client (receiver), one server process per pair&lt;/span&gt;
ib_write_bw &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-q&lt;/span&gt; 4 &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; 1048576 &lt;span class=&quot;nt&quot;&gt;-D&lt;/span&gt; 30 &lt;span class=&quot;nt&quot;&gt;--report_gbits&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; 18515 &amp;amp;

&lt;span class=&quot;c&quot;&gt;# on each storage node (sender), one process per assigned client, with -p&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# matching that client&apos;s listener. Twelve pairs over eight senders means some&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# nodes run two; give each pair its own port or the second will fail to bind.&lt;/span&gt;
ib_write_bw &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; mlx5_0 &lt;span class=&quot;nt&quot;&gt;-F&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-q&lt;/span&gt; 4 &lt;span class=&quot;nt&quot;&gt;-s&lt;/span&gt; 1048576 &lt;span class=&quot;nt&quot;&gt;-D&lt;/span&gt; 30 &lt;span class=&quot;nt&quot;&gt;--report_gbits&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; 18515 &amp;lt;client-ip&amp;gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt; #bytes  #iterations   BW peak[Gb/sec]  BW average[Gb/sec]  MsgRate[Mpps]
 1048576    112367        32.10              31.42             0.003746
 1048576    112116        31.88              31.35             0.003737
 1048576    133324        38.04              37.28             0.004444
 ...
 1048576    227273        64.02              63.55             0.007576
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Twelve pairs, summed: &lt;strong&gt;51.0 GiB/s&lt;/strong&gt;. The filesystem read run gave 50.3 GiB/s.
The gap is 1.5%, and there is no disk, no metadata server, no page cache and no
filesystem client in the RDMA path.&lt;/p&gt;

&lt;p&gt;The per-pair spread corroborates the routing story more sharply than the total
does. The pairs fall into three groups rather than scattering: six at about 31.4
Gb/s, five at about 37.3 Gb/s, and one at 63.6 Gb/s. Those are the ISL ceilings
divided by the number of pairs sharing them — 21.6 GiB/s split six ways, 21.3
GiB/s split five ways, and port 27’s single client taking the cable to itself at
about 7.4 GiB/s. The grouping matches the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibroute&lt;/code&gt; counts exactly, in a tool that
knows nothing about the filesystem and was never told which clients share a cable.&lt;/p&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;MsgRate&lt;/code&gt; column is a free consistency check on the whole block, and worth
using. At 31.42 Gb/s with 1 MiB messages the rate is about 3,750 messages per
second, which perftest reports in Mpps as 0.00375. A message rate of several whole
Mpps next to a megabyte message size is arithmetically impossible — it would be
terabits — and seeing one means the run used a different &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-s&lt;/code&gt; than you think, or
the figure was pasted from somewhere without being read.&lt;/p&gt;

&lt;div class=&quot;note&quot;&gt;
  &lt;p&gt;This test is only corroboration if the pairing is the same. If you let the test
pick a different set of nodes, or a different number of them, you have changed
which LIDs are destinations and therefore which cables carry the traffic. A
network-only test on a different node list can easily come out fast and send you
back to the storage layer for another week.&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;the-one-command-that-would-have-short-circuited-all-of-it&quot;&gt;The one command that would have short-circuited all of it&lt;/h2&gt;

&lt;p&gt;Everything above reconstructs something the switch will simply tell you.
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibroute&lt;/code&gt; dumps a switch’s linear forwarding table: destination LID on the left,
output port on the right.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ ibroute 2 | head -10
Unicast lids [0x0-0x2f] of switch Lid 2 guid 0x043f720300yyyyyy (SwitchB):
  Lid  Out   Destination
       Port     Info
0x0001 025 : (Switch portguid 0x043f720300xxxxxx: &apos;SwitchA&apos;)
0x0003 025 : (Channel Adapter portguid 0x...: &apos;client-a HCA-1&apos;)
0x0004 026 : (Channel Adapter portguid 0x...: &apos;client-b HCA-1&apos;)
0x0005 025 : (Channel Adapter portguid 0x...: &apos;client-c HCA-1&apos;)
0x0006 026 : (Channel Adapter portguid 0x...: &apos;client-d HCA-1&apos;)
0x0007 025 : (Channel Adapter portguid 0x...: &apos;client-e HCA-1&apos;)
0x0008 026 : (Channel Adapter portguid 0x...: &apos;client-f HCA-1&apos;)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Count how many of your job’s LIDs land on each output port:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# $JOB_LIDS = path to a file holding the LIDs of the clients in the run,&lt;/span&gt;
&lt;span class=&quot;c&quot;&gt;# one per line, in the same 0x%04x form ibroute prints (0x0003, not 3).&lt;/span&gt;
&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;ibroute &lt;span class=&quot;nt&quot;&gt;-n&lt;/span&gt; 2 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    | &lt;span class=&quot;nb&quot;&gt;awk&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;NR==FNR {want[$1]; next}
           FNR&amp;gt;3 &amp;amp;&amp;amp; ($1 in want) {c[$2]++}
           END {for (p in c) printf &quot;port %s: %d\n&quot;, p, c[p]}&apos;&lt;/span&gt; &lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$JOB_LIDS&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&lt;/span&gt; - &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
    | &lt;span class=&quot;nb&quot;&gt;sort
&lt;/span&gt;port 025: 6
port 026: 5
port 027: 1
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Match the LIDs as whole fields, not as substrings. An earlier version of this
used &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;grep -Ff&lt;/code&gt; against the LID list with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;^&lt;/code&gt; prefixed to each line, which does
nothing useful: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-F&lt;/code&gt; makes the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;^&lt;/code&gt; a literal character to search for, so the
pattern matches no line at all and the pipeline reports an empty distribution —
a result indistinguishable, at a glance, from a fabric with no job traffic on it.
The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;awk&lt;/code&gt; form above compares field 1 for equality and has no such failure mode.&lt;/p&gt;

&lt;p&gt;Twelve clients, three of eight available cables, two of them holding eleven of
the twelve. That is the answer. It takes about four seconds and needs no
benchmark, no maintenance window and no counter arithmetic.&lt;/p&gt;

&lt;p&gt;The reason it is not the first thing anyone tries is structural. Parallel
filesystem tuning documentation is thorough about page pool sizing, block size,
prefetch depth and queue depths, and says essentially nothing about inter-switch
links or routing engines, because those belong to another vendor’s product and
another team’s runbook. A filesystem-shaped investigation follows a
filesystem-shaped checklist, and that checklist has no entry for “count how many
of your nodes are routed over the same cable”. The elimination order below exists
to add one.&lt;/p&gt;

&lt;h2 id=&quot;telling-this-apart-from-the-things-it-resembles&quot;&gt;Telling this apart from the things it resembles&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Symptom shape&lt;/th&gt;
      &lt;th&gt;Likely cause&lt;/th&gt;
      &lt;th&gt;Command that distinguishes&lt;/th&gt;
      &lt;th&gt;What you see if it is this&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Both directions slow, ISLs even and well below ceiling&lt;/td&gt;
      &lt;td&gt;Not the fabric. Storage or client concurrency.&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iostat -x 5&lt;/code&gt; on servers; rerun with 2× the clients&lt;/td&gt;
      &lt;td&gt;Device utilization high, or throughput scales with client count&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;One direction slow, several ISLs at exactly zero&lt;/td&gt;
      &lt;td&gt;Static routing collapse&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibroute &amp;lt;sw-lid&amp;gt;&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;Active destination LIDs concentrated on a few output ports&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;One direction slow, all ISLs loaded, all near ceiling&lt;/td&gt;
      &lt;td&gt;Genuinely out of ISL bandwidth&lt;/td&gt;
      &lt;td&gt;the perfquery table&lt;/td&gt;
      &lt;td&gt;Aggregate ≈ number of ISLs × per-link payload ceiling&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Slow and erratic, ISLs uneven but none at zero&lt;/td&gt;
      &lt;td&gt;A degraded link&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo \| grep -v &quot;4X.*53.125&quot;&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;One link at 1X, or negotiated to a lower rate&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Slow, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt; climbing, no link near ceiling&lt;/td&gt;
      &lt;td&gt;Credit starvation from a slow receiver downstream&lt;/td&gt;
      &lt;td&gt;plain &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery &amp;lt;lid&amp;gt; &amp;lt;port&amp;gt;&lt;/code&gt; deltas on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt; (not in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-x&lt;/code&gt; set)&lt;/td&gt;
      &lt;td&gt;XmitWait rising on the ports that feed one endpoint&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Slow with rising error counters&lt;/td&gt;
      &lt;td&gt;Physical: cable, connector, transceiver&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibqueryerrors -c&lt;/code&gt;, twice, differenced&lt;/td&gt;
      &lt;td&gt;Non-zero deltas isolated to specific cables&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Slow for some jobs and not others&lt;/td&gt;
      &lt;td&gt;Allocation-dependent LID mapping&lt;/td&gt;
      &lt;td&gt;rerun on a different node list&lt;/td&gt;
      &lt;td&gt;The ISL distribution changes with the node list&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The last row causes the most confusion on a shared cluster. Once the mechanism is
clear it follows that the throughput a user gets depends on which nodes the
scheduler handed them, and that two runs of an identical job on identical
hardware can differ by a factor of two with nothing wrong anywhere. A benchmark
result quoted without its node list is not reproducible, and a regression that
comes and goes between runs may be the scheduler rather than the storage.&lt;/p&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt; deserves a note. It counts ticks during which a port had data to
send and no credits to send it with, which makes it the best early indicator of
downstream congestion and also famously noisy. Difference it like the data
counters — but with plain &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery&lt;/code&gt;, since it is a basic 32-bit counter and the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-x&lt;/code&gt; attribute does not carry it:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ perfquery 2 25 | grep PortXmitWait
PortXmitWait:....................18446231
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Non-zero XmitWait is normal on a busy fabric. Only the rate of change, on
specific ports, correlated with a specific workload, means anything. Do not raise
a ticket on an absolute value. The tick is a device-specific unit rather than a
fixed interval, so treat the counter as an ordering between ports on the same
switch and not as a duration you can convert to seconds.&lt;/p&gt;

&lt;h2 id=&quot;the-elimination-order&quot;&gt;The elimination order&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;#&lt;/th&gt;
      &lt;th&gt;Step&lt;/th&gt;
      &lt;th&gt;Command&lt;/th&gt;
      &lt;th&gt;Cost&lt;/th&gt;
      &lt;th&gt;Rules out&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;1&lt;/td&gt;
      &lt;td&gt;Confirm every ISL is at full width and rate&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iblinkinfo -S &amp;lt;switch-guid&amp;gt;&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;seconds&lt;/td&gt;
      &lt;td&gt;A degraded cable masquerading as a capacity shortfall&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;2&lt;/td&gt;
      &lt;td&gt;Count active job LIDs per ISL output port&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibroute -n &amp;lt;sw-lid&amp;gt;&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;seconds&lt;/td&gt;
      &lt;td&gt;Static routing collapse — or confirms it outright&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;3&lt;/td&gt;
      &lt;td&gt;Difference per-ISL counters under a write load&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;perfquery -x&lt;/code&gt;, two samples&lt;/td&gt;
      &lt;td&gt;2 minutes&lt;/td&gt;
      &lt;td&gt;Establishes the healthy-direction baseline&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;4&lt;/td&gt;
      &lt;td&gt;Repeat under an uncached read load&lt;/td&gt;
      &lt;td&gt;same&lt;/td&gt;
      &lt;td&gt;2 minutes&lt;/td&gt;
      &lt;td&gt;Produces the asymmetry table&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;5&lt;/td&gt;
      &lt;td&gt;Reproduce the ceiling with RDMA only, same pairing&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ib_write_bw&lt;/code&gt; fan&lt;/td&gt;
      &lt;td&gt;10 minutes&lt;/td&gt;
      &lt;td&gt;Everything above the network: filesystem, cache, disk&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;6&lt;/td&gt;
      &lt;td&gt;Check error and wait counters over the same window&lt;/td&gt;
      &lt;td&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibqueryerrors&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;PortXmitWait&lt;/code&gt; deltas&lt;/td&gt;
      &lt;td&gt;minutes&lt;/td&gt;
      &lt;td&gt;Physical faults and downstream credit starvation&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;7&lt;/td&gt;
      &lt;td&gt;Only now, touch a filesystem tunable&lt;/td&gt;
      &lt;td&gt;—&lt;/td&gt;
      &lt;td&gt;hours to days&lt;/td&gt;
      &lt;td&gt;—&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Steps 1 and 2 take under a minute and answer the question on most fabrics.
Everything after them builds a case you can hand to someone else, which matters,
because the remedy usually needs a change to a subnet manager another team owns.&lt;/p&gt;

&lt;h2 id=&quot;what-you-can-actually-do-about-it&quot;&gt;What you can actually do about it&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Adaptive routing.&lt;/strong&gt; The real fix is to stop choosing the output port once per
destination and choose it per packet based on port load. On NVIDIA switches this
comes from the subnet manager, by selecting an AR-capable routing engine — the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ar_*&lt;/code&gt; family, such as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ar_updn&lt;/code&gt;, alongside further &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ar_&lt;/code&gt; options that control the
mode and which service levels participate:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# opensm.conf — illustrative; take the exact directives from your own man page
routing_engine ar_updn
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Three caveats, and the third is the one that bites. These engines are NVIDIA
extensions shipped with their subnet manager rather than upstream OpenSM, so
whether you have them at all depends on which SM package you installed. Adaptive
routing can also deliver packets out of order, which reliable-connected queue
pairs do not enjoy, so the SM generally restricts it to adapters advertising
tolerance for it. And the option names, defaults and permitted values in this
family have moved between MLNX_OFED and DOCA-OFED releases — read them out of the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;opensm&lt;/code&gt; man page and sample &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;opensm.conf&lt;/code&gt; shipped with the package you actually
have, not out of an article, this one included.&lt;/p&gt;

&lt;p&gt;Then verify it took rather than assuming. NVIDIA’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;smparquery&lt;/code&gt; reads the
vendor-specific adaptive-routing MADs from a switch, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibdiagnet&lt;/code&gt; reports AR
state in its fabric summary; check invocation syntax against your installed
version, as it differs between tool generations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More links to spread over.&lt;/strong&gt; Adding ISLs raises the ceiling without fixing the
mechanism. The distribution improves with the number of cables, but only up to
the number of active destinations, and only if the aliasing does not follow you.
The collapse can reappear at the next SM sweep with a different node allocation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Break the round-robin.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scatter_ports&lt;/code&gt; takes a seed and randomizes selection
among equally loaded ports instead of cycling. One line, no hardware implication,
attacks the aliasing directly — and it is a lottery, not a guarantee: you swap a
systematically bad distribution for a randomly chosen one, usually much better
and occasionally not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Route the important nodes first.&lt;/strong&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;guid_routing_order_file&lt;/code&gt; sets the order in
which port GUIDs are routed for Min Hop and Up/Down. Put the genuinely hot GUIDs
at the top and they take the clean round-robin before the rest of the subnet
consumes the balance. Targeted, deterministic and underused.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Change the engine.&lt;/strong&gt; On a genuine fat tree with all channel adapters at the
leaves, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt; routes with the topology in mind rather than counting LIDs. It
falls back rather than running on a topology it does not recognize, which is a
feature. Note that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ftree&lt;/code&gt; does not support LMC greater than zero, and several of
the other engines carry restrictions of their own on LMC and on topology; check
the constraints for the specific engine and OpenSM version before planning
around one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Raise LMC, carefully.&lt;/strong&gt; Giving each port 2^LMC LIDs lets the engine compute
different paths to the same port. It helps only if the upper-layer protocol
issues path records for the alternate LIDs, which many do not, and the cost
multiplies across every port in the subnet and every forwarding table entry
needed to reach them. Establish that your client uses multiple paths first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check who runs the subnet manager.&lt;/strong&gt; Where the SM runs on a managed switch, the
embedded manager exposes only a subset of these options through the switch CLI,
and on some firmware levels routing engine selection is not in it. Moving the SM
to a host gives the full option set and a config file you can version. Find out
what you have before planning around it:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;$ sminfo
sminfo: sm lid 1 sm guid 0x043f720300xxxxxx, activity count 74412 priority 14 state 3 SMINFO_MASTER
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sminfo&lt;/code&gt; reports the SM’s &lt;em&gt;port&lt;/em&gt; GUID, which for a switch’s management port is
normally the same value &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ibswitches&lt;/code&gt; prints as the node GUID. A match there means
the fabric is being managed from the switch rather than from a host.&lt;/p&gt;

&lt;h2 id=&quot;what-to-claim-afterwards&quot;&gt;What to claim afterwards&lt;/h2&gt;

&lt;p&gt;Be careful about the size of the claim. The evidence supports this: for this node
allocation, in the read direction, the achievable aggregate is bounded by three
inter-switch links, two of them at 95% of payload ceiling, and a network-only
test with no storage in the path reproduces the bound to within 2%. Narrow,
strong, checkable.&lt;/p&gt;

&lt;p&gt;It does not support “the storage is fine”. The storage was never tested above
50.3 GiB/s on reads because the network would not deliver more. When the routing
is fixed, expect a second ceiling behind the first. That is the ordinary shape of
performance work.&lt;/p&gt;

&lt;p&gt;It also does not support a general statement about the fabric. Change the node
list and the numbers change. Quote any figure with the node count, the direction,
the dataset size relative to cache, and the ISL distribution that produced it.
Without those four the number is not reproducible, and someone will eventually
try to reproduce it.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Vault and service accounts: rotation must reach the running application</title>
    <link href="https://www.wirewalk.com/writing/vault-service-account-lifecycle/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/vault-service-account-lifecycle/</id>
    <summary>Secret issuance, application renewal, and credential revocation are separate steps to verify.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A secret manager can issue a new credential while the application continues to use the old one from a configuration file or a connection pool. The security outcome depends on the consumer, not just the successful rotation task.&lt;/p&gt;

&lt;h2 id=&quot;start-with-one-service-identity&quot;&gt;Start with one service identity&lt;/h2&gt;

&lt;p&gt;Choose a noncritical application and record which database or API it accesses, the required privileges, how it receives credentials, and how it refreshes them. Separate deployment credentials from runtime credentials. A build system should not automatically inherit the application’s production data access.&lt;/p&gt;

&lt;p&gt;HashiCorp Vault’s &lt;a href=&quot;https://developer.hashicorp.com/vault/docs/secrets/databases&quot;&gt;database secrets engine&lt;/a&gt; documents dynamic credentials and database-specific integrations. Supported plugins and credential behavior vary, so use the exact database and plugin documentation before implementation.&lt;/p&gt;

&lt;h2 id=&quot;define-the-credential-contract&quot;&gt;Define the credential contract&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Property&lt;/th&gt;
      &lt;th&gt;Question&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Scope&lt;/td&gt;
      &lt;td&gt;Which operations and resources are allowed?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Lifetime&lt;/td&gt;
      &lt;td&gt;How long may a credential remain valid?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Renewal&lt;/td&gt;
      &lt;td&gt;Who renews it and what happens on failure?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Revocation&lt;/td&gt;
      &lt;td&gt;How does the backend actually invalidate it?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Recovery&lt;/td&gt;
      &lt;td&gt;How does the application restart if Vault is unavailable?&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The contract should also identify where credentials might be copied: environment variables, logs, crash dumps, temporary files, or deployment artifacts. Moving the original into Vault does not remove existing copies.&lt;/p&gt;

&lt;h2 id=&quot;test-the-consumer-lifecycle&quot;&gt;Test the consumer lifecycle&lt;/h2&gt;

&lt;p&gt;In a lab database, issue a least-privilege credential through the intended integration. Start the application, perform a permitted operation, and confirm that an out-of-scope operation is denied.&lt;/p&gt;

&lt;p&gt;Allow renewal or rotation to occur through the actual application workflow. Confirm new connections use valid credentials and that the application handles the transition without accumulating connection failures. Then test expiry or revocation on the disposable identity and observe both new and existing connections according to the database’s behavior.&lt;/p&gt;

&lt;p&gt;Do not assume a revoked lease instantly terminates every established session. Establish the backend-specific outcome and document any additional response step.&lt;/p&gt;

&lt;h2 id=&quot;protect-bootstrap-authority&quot;&gt;Protect bootstrap authority&lt;/h2&gt;

&lt;p&gt;The application needs an initial way to authenticate to Vault. Review that mechanism as carefully as the database credential. A permanent broad token embedded in an image can recreate the original problem at a more powerful layer.&lt;/p&gt;

&lt;p&gt;Prefer workload identity methods supported by the environment, scope policies tightly, and verify the audit trail. Avoid displaying live secret values in routine validation output.&lt;/p&gt;

&lt;h2 id=&quot;plan-for-service-loss&quot;&gt;Plan for service loss&lt;/h2&gt;

&lt;p&gt;Test a temporary loss of the secret-management path in a lab. Distinguish an already running application from a cold restart. Record the maximum acceptable interruption and the approved recovery path. Do not silently extend credentials indefinitely to make an availability test pass.&lt;/p&gt;

&lt;p&gt;A completed implementation includes a successful permitted operation, denied excess access, observed rotation, tested revocation, and understood outage behavior. The existence of a Vault policy is only the configuration portion of that evidence.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Offboarding: disabling the directory account is one step</title>
    <link href="https://www.wirewalk.com/writing/offboarding-sessions-keys-tokens/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/offboarding-sessions-keys-tokens/</id>
    <summary>Test existing sessions and independent credentials as well as new sign-ins.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;An employee’s directory account can be disabled while an application session, SSH key, or API token remains usable. Offboarding is an access-removal workflow across systems, not a single identity-console action.&lt;/p&gt;

&lt;h2 id=&quot;inventory-the-access-forms&quot;&gt;Inventory the access forms&lt;/h2&gt;

&lt;p&gt;Record the person’s interactive accounts, privileged identities, application roles, SSH credentials, API tokens, and delegated access. Include services they own so that access removal does not silently stop a business process.&lt;/p&gt;

&lt;p&gt;Do not record secret values in the inventory. Store credential identifiers, owning systems, scopes, and rotation procedures.&lt;/p&gt;

&lt;h2 id=&quot;separate-new-access-from-existing-access&quot;&gt;Separate new access from existing access&lt;/h2&gt;

&lt;p&gt;Microsoft documents that Entra token and application-session behavior can produce a delay between revocation and effective loss of access. Application-issued sessions may require application-side handling. See &lt;a href=&quot;https://learn.microsoft.com/en-us/entra/identity/users/users-revoke-access&quot;&gt;Microsoft’s access-revocation guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Use a test identity to establish what happens in your application estate:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Access form&lt;/th&gt;
      &lt;th&gt;Validation&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;New SSO login&lt;/td&gt;
      &lt;td&gt;Denied after the intended control takes effect&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Existing browser session&lt;/td&gt;
      &lt;td&gt;Ends within the documented interval/workflow&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;API token&lt;/td&gt;
      &lt;td&gt;Revoked or replaced at its issuing system&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SSH key or certificate&lt;/td&gt;
      &lt;td&gt;No longer authorizes a new session&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Active server session&lt;/td&gt;
      &lt;td&gt;Handled by the approved session procedure&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;A failed new login says nothing about a browser window already open elsewhere.&lt;/p&gt;

&lt;h2 id=&quot;handle-ssh-independently&quot;&gt;Handle SSH independently&lt;/h2&gt;

&lt;p&gt;Inspect central key or certificate issuance and local authorized-key paths. OpenSSH supports multiple authorization mechanisms, documented in &lt;a href=&quot;https://man.openbsd.org/sshd_config&quot;&gt;sshd_config&lt;/a&gt;. Removing one visible key file is incomplete if another source still grants access.&lt;/p&gt;

&lt;p&gt;A terminated employment relationship is not a reason to delete application data indiscriminately. Preserve records and transfer ownership according to the organization’s retention and business procedures. Keep operational ownership separate from the individual’s access rights.&lt;/p&gt;

&lt;h2 id=&quot;rehearse-a-full-departure&quot;&gt;Rehearse a full departure&lt;/h2&gt;

&lt;p&gt;Create a synthetic user with representative access and an active application session. Run the offboarding checklist in a test scope. Record each system’s completion time and the actual result of attempting access afterward.&lt;/p&gt;

&lt;p&gt;Include a long-lived credential issued outside SSO. This catches a common design gap: the central identity system does its job while an independent service continues to trust an older credential.&lt;/p&gt;

&lt;h2 id=&quot;make-exceptions-visible&quot;&gt;Make exceptions visible&lt;/h2&gt;

&lt;p&gt;Some departures require temporary continuity for scheduled tasks or shared business records. Replace personal ownership with an appropriate service identity or named successor. An exception should specify its purpose, expiry, and approving owner; leaving a personal account active indefinitely is not a migration strategy.&lt;/p&gt;

&lt;h2 id=&quot;acceptance-evidence&quot;&gt;Acceptance evidence&lt;/h2&gt;

&lt;p&gt;Keep the systems covered, access attempts, revocation timestamps, unresolved dependencies, and ownership transfers. Report incomplete items explicitly. “Account disabled” is a valid fact, but it is not a defensible summary of the entire access state.&lt;/p&gt;

&lt;p&gt;Review &lt;a href=&quot;/writing/vault-service-account-lifecycle/&quot;&gt;service-account lifecycle&lt;/a&gt; to reduce the number of workflows tied to personal credentials before the next departure.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Privileged SSH: make the bastion a real boundary</title>
    <link href="https://www.wirewalk.com/writing/privileged-ssh-bastion-boundaries/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/privileged-ssh-bastion-boundaries/</id>
    <summary>A jump host is useful only when direct access, forwarding, and downstream authority are deliberately controlled.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;Putting a bastion in a network diagram does not force administrators to use it. If managed hosts still accept direct connections from broad networks or share long-lived keys, the new hop may add logging while leaving the original access path intact.&lt;/p&gt;

&lt;h2 id=&quot;define-the-authorized-route&quot;&gt;Define the authorized route&lt;/h2&gt;

&lt;p&gt;Document which administrator devices can reach the bastion, which hosts it can reach, and which identities are accepted downstream. Separate ordinary user access from privileged maintenance. Keep a tested emergency route for recovery, with explicit ownership and audit.&lt;/p&gt;

&lt;p&gt;The boundary includes network policy and host authentication. Neither one substitutes for the other.&lt;/p&gt;

&lt;h2 id=&quot;inspect-effective-openssh-settings&quot;&gt;Inspect effective OpenSSH settings&lt;/h2&gt;

&lt;p&gt;OpenSSH documents configuration processing, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Match&lt;/code&gt; rules, forwarding controls, and authorized principals in &lt;a href=&quot;https://man.openbsd.org/sshd_config&quot;&gt;sshd_config&lt;/a&gt;. On a Linux lab server with an installed OpenSSH daemon:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;sshd &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;sudo &lt;/span&gt;sshd &lt;span class=&quot;nt&quot;&gt;-T&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The first checks configuration validity; the second displays effective settings for the default context. Where &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Match&lt;/code&gt; blocks are used, test the intended connection context with the version-supported &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;-C&lt;/code&gt; options as well. Keep the output restricted because it reveals security configuration.&lt;/p&gt;

&lt;p&gt;Do not assume that the last line in a configuration file overrides earlier values. Includes and first-obtained-value behavior make visual inspection alone unreliable.&lt;/p&gt;

&lt;h2 id=&quot;test-the-access-matrix&quot;&gt;Test the access matrix&lt;/h2&gt;

&lt;p&gt;Use a pilot host and test four cases: approved administrator via bastion, unrelated user via bastion, direct connection from an ordinary client network, and the emergency path. Define the expected result for each.&lt;/p&gt;

&lt;p&gt;Test both authentication and privilege elevation. An account that can log in but cannot perform the approved task is an operational failure; an account that can become root everywhere may be broader than intended.&lt;/p&gt;

&lt;p&gt;If certificates are used, test expired and unauthorized principals with disposable credentials. Certificate issuance policy and signer protection become part of the access boundary.&lt;/p&gt;

&lt;h2 id=&quot;review-forwarding-and-sessions&quot;&gt;Review forwarding and sessions&lt;/h2&gt;

&lt;p&gt;Agent, TCP, and other forwarding features can extend authority beyond the apparent login route. Disable unnecessary features using the supported controls and test the workflows that remain. OpenSSH also notes that forwarding restrictions are not complete confinement for a user who can execute arbitrary code on the host.&lt;/p&gt;

&lt;p&gt;Record session start, identity, destination, and privileged actions according to the organization’s logging policy. Do not promise full command attribution merely because SSH authentication logs exist.&lt;/p&gt;

&lt;h2 id=&quot;roll-out-without-losing-recovery&quot;&gt;Roll out without losing recovery&lt;/h2&gt;

&lt;p&gt;Keep a working rescue session and independently verified console access during the pilot. Apply changes to a small group, establish a new session through the intended route, and verify both allowed and denied behavior before expanding.&lt;/p&gt;

&lt;p&gt;The result should be a route enforced by network and host policy, with bounded downstream authority and a working recovery path. See &lt;a href=&quot;/writing/management-network-bmc-isolation/&quot;&gt;management-network isolation&lt;/a&gt; for the infrastructure beneath that route.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Microsoft Entra admin access: require the authentication method you intend</title>
    <link href="https://www.wirewalk.com/writing/entra-phishing-resistant-admin-access/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/entra-phishing-resistant-admin-access/</id>
    <summary>An MFA requirement and a phishing-resistant authentication requirement are not identical policies.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;An administrator can satisfy a broad MFA policy with a method that does not meet the organization’s intended resistance to phishing. The policy needs to specify the required assurance, and the rollout needs to prove both allowed access and denied fallback.&lt;/p&gt;

&lt;h2 id=&quot;inspect-authentication-strength&quot;&gt;Inspect authentication strength&lt;/h2&gt;

&lt;p&gt;Microsoft Entra Conditional Access provides authentication strengths that constrain acceptable method combinations. Microsoft documents a policy for administrator roles using phishing-resistant MFA. See &lt;a href=&quot;https://learn.microsoft.com/en-us/entra/identity/conditional-access/policy-admin-phish-resistant-mfa&quot;&gt;the administrator policy guide&lt;/a&gt; and &lt;a href=&quot;https://learn.microsoft.com/en-us/entra/identity/authentication/concept-authentication-strength-how-it-works&quot;&gt;how strengths are evaluated&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Confirm current licensing, supported methods, tenant configuration, and dependencies before rollout. Do not assume a policy named “strong MFA” actually requires the intended methods.&lt;/p&gt;

&lt;h2 id=&quot;prepare-registration-and-recovery&quot;&gt;Prepare registration and recovery&lt;/h2&gt;

&lt;p&gt;Select a pilot group and ensure users have registered the approved authenticators before enforcement. Test an alternative recovery route that does not recreate the same failure dependency. Losing a device should not leave the only administrator unable to operate the tenant.&lt;/p&gt;

&lt;p&gt;Microsoft’s &lt;a href=&quot;https://learn.microsoft.com/en-us/entra/identity/role-based-access-control/security-emergency-access&quot;&gt;emergency access guidance&lt;/a&gt; is the starting point for protecting and monitoring emergency accounts. Follow current requirements rather than copying an old blanket MFA exclusion.&lt;/p&gt;

&lt;h2 id=&quot;test-a-policy-matrix&quot;&gt;Test a policy matrix&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Scenario&lt;/th&gt;
      &lt;th&gt;Expected result&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Pilot administrator with approved method&lt;/td&gt;
      &lt;td&gt;Access succeeds&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Pilot administrator using a disallowed method&lt;/td&gt;
      &lt;td&gt;Required step-up or denial&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Unregistered pilot administrator&lt;/td&gt;
      &lt;td&gt;Documented enrollment/recovery outcome&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Emergency procedure&lt;/td&gt;
      &lt;td&gt;Authorized access and visible audit evidence&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Ordinary user outside scope&lt;/td&gt;
      &lt;td&gt;Intended existing behavior&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Run report-only analysis where supported, then a controlled enforced pilot. Report-only results cannot establish that an application actually accepts the final login flow.&lt;/p&gt;

&lt;h2 id=&quot;review-exclusions-as-carefully-as-inclusions&quot;&gt;Review exclusions as carefully as inclusions&lt;/h2&gt;

&lt;p&gt;List excluded users, groups, workloads, and applications. Give each exception an owner and review date. A broad exclusion for automation may accidentally cover interactive administrators if account types and group membership are poorly controlled.&lt;/p&gt;

&lt;p&gt;Check the interaction with other Conditional Access policies. The effective result comes from the policy set, not the one screen currently being reviewed. Inspect sign-in details for the test session to confirm which controls applied.&lt;/p&gt;

&lt;h2 id=&quot;protect-the-post-login-path&quot;&gt;Protect the post-login path&lt;/h2&gt;

&lt;p&gt;Stronger authentication does not remove excessive privileges, insecure administrator workstations, or stolen active sessions. Keep privileged work separate from ordinary browsing and email, minimize standing authority, and document session-revocation procedures.&lt;/p&gt;

&lt;p&gt;A valid acceptance report includes policy scope, exclusions, allowed and denied test outcomes, sign-in evidence, and the tested recovery path. Avoid presenting enrollment counts as equivalent to enforced protection.&lt;/p&gt;

&lt;p&gt;Next, examine &lt;a href=&quot;/writing/privileged-ssh-bastion-boundaries/&quot;&gt;privileged SSH access&lt;/a&gt; and &lt;a href=&quot;/writing/offboarding-sessions-keys-tokens/&quot;&gt;offboarding across sessions and tokens&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Missing security telemetry: monitor the silence as well as the alerts</title>
    <link href="https://www.wirewalk.com/writing/missing-security-telemetry-prometheus/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/missing-security-telemetry-prometheus/</id>
    <summary>A quiet dashboard can mean no incidents, a broken producer, or a source that vanished from inventory.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A collector answering health checks can still receive no useful events. A source removed from service discovery may disappear from every dashboard. Security monitoring needs a separate definition of which sources should exist and how recently each should have produced evidence.&lt;/p&gt;

&lt;h2 id=&quot;distinguish-three-failure-states&quot;&gt;Distinguish three failure states&lt;/h2&gt;

&lt;p&gt;A known target can be down. A known target can be reachable but stale. A target can be missing from discovery entirely. One uptime graph does not cover all three.&lt;/p&gt;

&lt;p&gt;Use an inventory of required sensors and log producers. Include ownership, expected event cadence, and maintenance rules. Compare observed sources against that inventory rather than deriving the entire expectation from whatever happened to report today.&lt;/p&gt;

&lt;h2 id=&quot;express-freshness-explicitly&quot;&gt;Express freshness explicitly&lt;/h2&gt;

&lt;p&gt;Prometheus documents &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;absent_over_time&lt;/code&gt; and time functions in its &lt;a href=&quot;https://prometheus.io/docs/prometheus/latest/querying/functions/&quot;&gt;query reference&lt;/a&gt;. For an exporter that you implement to publish the most recent event time, illustrative expressions are:&lt;/p&gt;

&lt;pre&gt;&lt;code class=&quot;language-promql&quot;&gt;(time() - security_last_event_timestamp_seconds{source=&quot;storage-audit&quot;}) &amp;gt; 900
&lt;/code&gt;&lt;/pre&gt;

&lt;pre&gt;&lt;code class=&quot;language-promql&quot;&gt;absent_over_time(security_last_event_timestamp_seconds{source=&quot;storage-audit&quot;}[15m])
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;security_last_event_timestamp_seconds&lt;/code&gt; is an example metric, not a built-in metric supplied by a storage vendor or Prometheus. The first expression detects old values; the second detects no samples for the explicitly named source. Neither automatically enumerates every missing source in a changing fleet.&lt;/p&gt;

&lt;p&gt;The 15-minute window is illustrative. Select the threshold from the system’s real event behavior and response requirement. A genuinely idle source needs a heartbeat or controlled validation event rather than an assumption that user activity is continuous.&lt;/p&gt;

&lt;h2 id=&quot;test-the-alert-chain&quot;&gt;Test the alert chain&lt;/h2&gt;

&lt;p&gt;In a lab, stop the event producer while leaving the exporter running. Confirm the stale-event alert. Then stop the exporter and confirm missing-sample or target-down behavior. Finally remove the source from discovery and verify that inventory reconciliation still identifies it.&lt;/p&gt;

&lt;p&gt;These tests isolate different failure modes. If the third condition becomes invisible, the monitor still has an absence problem.&lt;/p&gt;

&lt;h2 id=&quot;carry-enough-context-to-act&quot;&gt;Carry enough context to act&lt;/h2&gt;

&lt;p&gt;An alert should name the producer, last event time, last successful collection, owning team, and a short diagnostic procedure. A generic “no data” notification encourages responders to guess whether the event source, network, parser, or query failed.&lt;/p&gt;

&lt;p&gt;Preserve maintenance windows with expiry. A permanent silence created for a temporary outage can make the monitoring system look healthy for months.&lt;/p&gt;

&lt;h2 id=&quot;check-timestamps-and-delivery&quot;&gt;Check timestamps and delivery&lt;/h2&gt;

&lt;p&gt;Clock skew can make a producer appear fresh or stale incorrectly. Compare event time with ingest time and monitor clock synchronization independently. Confirm that alert routing reaches a staffed destination using an approved test notification workflow.&lt;/p&gt;

&lt;p&gt;The completed control is not a query that evaluates successfully. It is an expected source inventory, a freshness signal, and a demonstrated response when each part stops working. Apply the same approach to &lt;a href=&quot;/writing/storage-audit-evidence-pipeline/&quot;&gt;storage audit pipelines&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Detection engineering: keep a test that must match and one that must not</title>
    <link href="https://www.wirewalk.com/writing/detection-rules-regression-tests/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/detection-rules-regression-tests/</id>
    <summary>A small fixture library makes rule changes reviewable and keeps false-positive fixes from silently removing coverage.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A detection rule has two obligations: identify the behavior it was written for and leave expected activity alone. Testing only the first produces false positives; testing only the second can remove the detection entirely.&lt;/p&gt;

&lt;h2 id=&quot;write-the-detection-contract&quot;&gt;Write the detection contract&lt;/h2&gt;

&lt;p&gt;For each rule, record the behavior, required telemetry, expected event fields, and the response owner. Describe the scope narrowly enough to test. “Detect suspicious traffic” is not a specification. “Alert on this synthetic request pattern in this inspected HTTP path” is.&lt;/p&gt;

&lt;p&gt;Keep one approved matching fixture, one ordinary nonmatching fixture, and any fixture added after a false positive. Use synthetic captures or appropriately sanitized authorized data. Store hashes and provenance so later reviewers know exactly what was exercised.&lt;/p&gt;

&lt;h2 id=&quot;use-the-engines-parser&quot;&gt;Use the engine’s parser&lt;/h2&gt;

&lt;p&gt;Suricata and Snort have documented configuration tests and offline capture modes. See &lt;a href=&quot;https://docs.suricata.io/en/suricata-7.0.8/command-line-options.html&quot;&gt;Suricata’s CLI reference&lt;/a&gt; and &lt;a href=&quot;https://docs.snort.org/start/inspection&quot;&gt;Snort’s traffic inspection guide&lt;/a&gt;. Validate the complete configuration used for the test, not just a detached rule string.&lt;/p&gt;

&lt;p&gt;The test record should contain:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;rule_id: locally assigned identifier
engine_version: exact installed version
configuration_hash: recorded SHA-256
fixture_hash: recorded SHA-256
expected: alert present or alert absent
observed: matching identifiers and count
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This is a schema example, not measured output. Record the actual values from the run.&lt;/p&gt;

&lt;h2 id=&quot;assert-the-right-event&quot;&gt;Assert the right event&lt;/h2&gt;

&lt;p&gt;Do not accept “one or more alerts” as proof. Assert the intended rule identifier and, when relevant, the source, destination, or application field. A completely unrelated signature can otherwise make a broken test look successful.&lt;/p&gt;

&lt;p&gt;For suppression and threshold rules, define the expected count or range explicitly. Replay timing matters: an accelerated fixture may not behave like the original traffic window. Record the timing method and engine options.&lt;/p&gt;

&lt;h2 id=&quot;separate-detection-from-delivery&quot;&gt;Separate detection from delivery&lt;/h2&gt;

&lt;p&gt;An offline fixture proves engine behavior. A live benign event proves the collection and routing path. A ticket or notification received by the intended responder proves delivery. Maintain all three where the control depends on all three.&lt;/p&gt;

&lt;p&gt;If the product is deployed inline, add an observed client transaction result. A drop action in a file does not prove the forwarding path enforced it.&lt;/p&gt;

&lt;h2 id=&quot;make-exceptions-testable&quot;&gt;Make exceptions testable&lt;/h2&gt;

&lt;p&gt;When narrowing a rule to fix a false positive, add that legitimate example to the fixture library before changing the rule. Rerun both the old match and the new nonmatch. An exception should identify its owner, scope, and review date rather than becoming an undocumented global bypass.&lt;/p&gt;

&lt;p&gt;Keep fixture size modest and access controlled. Packet captures can contain credentials and personal data even when the original investigation seemed routine.&lt;/p&gt;

&lt;h2 id=&quot;what-to-release&quot;&gt;What to release&lt;/h2&gt;

&lt;p&gt;Release the rule, its tests, configuration changes, and results together. Keep the previous working version available for rollback. This turns a ruleset update into an engineering change with known behavior rather than a hopeful increase in the number of enabled signatures.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Encrypted traffic: decide what each security sensor can actually establish</title>
    <link href="https://www.wirewalk.com/writing/encrypted-traffic-detection-boundaries/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/encrypted-traffic-detection-boundaries/</id>
    <summary>Combine network, endpoint, and application evidence instead of assuming one sensor sees inside every connection.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;Encryption changes the evidence available to a network sensor. It does not make all monitoring useless, and it does not justify pretending that a packet inspection engine sees application content it cannot decrypt.&lt;/p&gt;

&lt;h2 id=&quot;build-a-visibility-table&quot;&gt;Build a visibility table&lt;/h2&gt;

&lt;p&gt;For each important service, record where encryption starts and ends, which network paths are observed, and which endpoint or application logs are available.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Evidence source&lt;/th&gt;
      &lt;th&gt;Useful question&lt;/th&gt;
      &lt;th&gt;Important limit&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Network flow&lt;/td&gt;
      &lt;td&gt;Which endpoints communicated?&lt;/td&gt;
      &lt;td&gt;Usually no application payload&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;TLS metadata&lt;/td&gt;
      &lt;td&gt;What handshake details were observed?&lt;/td&gt;
      &lt;td&gt;Fields vary by protocol and privacy features&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Endpoint telemetry&lt;/td&gt;
      &lt;td&gt;Which process opened the connection?&lt;/td&gt;
      &lt;td&gt;Depends on agent coverage and trust&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Application audit&lt;/td&gt;
      &lt;td&gt;Which authenticated operation occurred?&lt;/td&gt;
      &lt;td&gt;Depends on application logging&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Zeek’s &lt;a href=&quot;https://docs.zeek.org/en/current/reference/logs/ssl.html&quot;&gt;TLS log documentation&lt;/a&gt; describes protocol metadata that may be available. Do not assume every connection exposes a hostname or certificate in the same way.&lt;/p&gt;

&lt;h2 id=&quot;test-two-observation-points&quot;&gt;Test two observation points&lt;/h2&gt;

&lt;p&gt;Use a lab application with a known request and a unique harmless marker. Observe the client-to-service encrypted path, then inspect authorized server-side application logs. Identify which source can establish the marker and which can establish only the connection.&lt;/p&gt;

&lt;p&gt;If there is an approved TLS termination proxy, document both sides. A sensor after termination may see plaintext traffic for that segment while missing traffic that takes a different route. The presence of a proxy does not establish universal inspection.&lt;/p&gt;

&lt;p&gt;Suricata’s &lt;a href=&quot;https://docs.suricata.io/en/suricata-7.0.8/rules/http-keywords.html&quot;&gt;HTTP keyword documentation&lt;/a&gt; describes application buffers used by HTTP rules. A payload rule requiring those buffers cannot be assumed to match opaque encrypted contents.&lt;/p&gt;

&lt;h2 id=&quot;evaluate-decryption-as-an-architecture-change&quot;&gt;Evaluate decryption as an architecture change&lt;/h2&gt;

&lt;p&gt;TLS inspection introduces certificate trust, private-key protection, application compatibility, and privacy obligations. Establish which traffic classes may be inspected and how the inspection infrastructure is administered. Test certificate-pinned applications and machine-to-machine clients before a rollout.&lt;/p&gt;

&lt;p&gt;Do not make “decrypt everything” a substitute for identifying business requirements. Some flows may be better covered by endpoint detection and application audit, especially where interception changes the trust model or breaks the workload.&lt;/p&gt;

&lt;h2 id=&quot;correlate-without-overstating-certainty&quot;&gt;Correlate without overstating certainty&lt;/h2&gt;

&lt;p&gt;Join network and endpoint records using time, addresses, ports, and host identity. Account for NAT and address reuse. A shared egress address can represent many users; it is not a user identifier.&lt;/p&gt;

&lt;p&gt;In an analyst exercise, write each conclusion beside its evidence source. “Host A connected to service B” may be established. “User C downloaded file D” may require application audit. Keeping that distinction visible prevents a plausible story from becoming an unsupported incident finding.&lt;/p&gt;

&lt;h2 id=&quot;verify-the-blind-spots&quot;&gt;Verify the blind spots&lt;/h2&gt;

&lt;p&gt;Repeat the fixture when routing, TLS termination, or endpoint tooling changes. Retain a list of unobserved paths and assign owners. A security architecture with documented limits is more actionable than a dashboard claiming complete coverage without a reproducible test.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Wazuh file integrity monitoring: start with files someone will investigate</title>
    <link href="https://www.wirewalk.com/writing/wazuh-file-integrity-baselines/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/wazuh-file-integrity-baselines/</id>
    <summary>A selective baseline and a tested response path make file-change alerts useful.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;Monitoring every changing file on a busy server produces activity, but not necessarily useful evidence. A file integrity monitoring deployment should begin with a small set of paths whose unexpected modification has a clear owner and consequence.&lt;/p&gt;

&lt;h2 id=&quot;choose-a-meaningful-baseline&quot;&gt;Choose a meaningful baseline&lt;/h2&gt;

&lt;p&gt;Start with authentication configuration, service definitions, privileged automation, and selected application configuration. Exclude noisy runtime data only after understanding what would be lost. A blanket exclusion for an entire application tree can remove the very files an attacker would alter.&lt;/p&gt;

&lt;p&gt;Wazuh documents scheduled and real-time integrity monitoring and platform-dependent capabilities in its &lt;a href=&quot;https://documentation.wazuh.com/current/user-manual/capabilities/file-integrity/index.html&quot;&gt;FIM guide&lt;/a&gt;. Confirm which mode the target operating system and configured paths support.&lt;/p&gt;

&lt;h2 id=&quot;use-a-harmless-configuration-fixture&quot;&gt;Use a harmless configuration fixture&lt;/h2&gt;

&lt;p&gt;Before applying a broad production policy, create a dedicated directory containing a synthetic text file. Add only that directory to the lab agent’s configured monitoring scope using the documented configuration method.&lt;/p&gt;

&lt;p&gt;Record the baseline, then make three changes separately: edit the contents, change permissions, and remove the test file. Wait for the configured detection mode and interval after each operation.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Test&lt;/th&gt;
      &lt;th&gt;Evidence expected&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Content edit&lt;/td&gt;
      &lt;td&gt;Correct path and changed integrity information&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Permission change&lt;/td&gt;
      &lt;td&gt;Relevant attribute change, if enabled and supported&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Removal&lt;/td&gt;
      &lt;td&gt;Deletion event with the same file identity/path context&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Agent offline&lt;/td&gt;
      &lt;td&gt;Missing-agent or freshness signal&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Do not claim attribution unless the configured platform feature actually supplies it. A file change event can establish that a change happened without establishing the process or person responsible.&lt;/p&gt;

&lt;h2 id=&quot;follow-the-alert-to-an-owner&quot;&gt;Follow the alert to an owner&lt;/h2&gt;

&lt;p&gt;Inspect the event on the manager and in the analyst view. Confirm hostname, path, timestamp, and the distinction between the original event and ingestion time. A correctly collected event assigned to no one is not an operational detection.&lt;/p&gt;

&lt;p&gt;Write a response note for each monitored path family. An unexpected SSH configuration change may require reviewing a deployment record and testing effective settings. An application template change may belong to the application team. Different changes should not all trigger the same emergency response.&lt;/p&gt;

&lt;h2 id=&quot;handle-legitimate-deployments&quot;&gt;Handle legitimate deployments&lt;/h2&gt;

&lt;p&gt;Associate maintenance windows with an approved change, but preserve the events. Suppressing all file integrity activity during deployment can conceal an unrelated change. Prefer scoped interpretation and review over discarding the evidence.&lt;/p&gt;

&lt;p&gt;After rebuilding an agent or reinstalling a host, confirm that the new baseline is intentional. A newly initialized baseline is not proof that the host matches a trusted image.&lt;/p&gt;

&lt;h2 id=&quot;control-volume-before-expanding&quot;&gt;Control volume before expanding&lt;/h2&gt;

&lt;p&gt;Measure events per host, storage growth, and analyst effort on a representative subset. Expand only when expected events arrive and ordinary maintenance remains understandable.&lt;/p&gt;

&lt;p&gt;Finish by repeating the harmless fixture after a policy change or manager upgrade. Pair FIM with &lt;a href=&quot;/writing/missing-security-telemetry-prometheus/&quot;&gt;missing telemetry monitoring&lt;/a&gt; so a quiet dashboard cannot conceal a disconnected source.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Snort 3 rule validation: a loaded rule is not a tested detection</title>
    <link href="https://www.wirewalk.com/writing/snort3-rule-validation/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/snort3-rule-validation/</id>
    <summary>Use matching and nonmatching captures, explicit actions, and live enforcement tests to validate a rule change.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A rule can parse successfully and never match the traffic it was meant to detect. Another can match so broadly that moving it into an inline policy interrupts legitimate work. Snort 3 changes need evidence at both boundaries.&lt;/p&gt;

&lt;h2 id=&quot;pin-the-inputs&quot;&gt;Pin the inputs&lt;/h2&gt;

&lt;p&gt;Keep the Snort version, Lua configuration, rule file, fixture PCAP, and expected signature identifiers in the change record. Pin the fixture by checksum. A result without those inputs is difficult to reproduce after a ruleset update.&lt;/p&gt;

&lt;p&gt;The official guide separates &lt;a href=&quot;https://docs.snort.org/start/configuration&quot;&gt;configuration&lt;/a&gt;, &lt;a href=&quot;https://docs.snort.org/start/inspection&quot;&gt;reading traffic&lt;/a&gt;, and &lt;a href=&quot;https://docs.snort.org/rules/headers/actions&quot;&gt;rule actions&lt;/a&gt;. Use the options for the installed Snort 3 package; Snort 2 configuration examples are not interchangeable.&lt;/p&gt;

&lt;h2 id=&quot;validate-syntax-and-behavior-separately&quot;&gt;Validate syntax and behavior separately&lt;/h2&gt;

&lt;p&gt;For a lab installation with configuration at the example path:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;snort &lt;span class=&quot;nt&quot;&gt;-T&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; /etc/snort/snort.lua
snort &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; /etc/snort/snort.lua &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; ./approved-match.pcap &lt;span class=&quot;nt&quot;&gt;-A&lt;/span&gt; alert_fast
snort &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; /etc/snort/snort.lua &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; ./approved-nonmatch.pcap &lt;span class=&quot;nt&quot;&gt;-A&lt;/span&gt; alert_fast
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Supply your own approved synthetic fixtures and confirm that the relevant rule file is included by the configuration. The commands do not create the fixtures. Output location and available logging modules should be verified against the installation.&lt;/p&gt;

&lt;p&gt;The first run should load successfully. The matching fixture should produce the intended rule identifier. The nonmatching fixture should not produce that identifier. Other unrelated alerts need separate interpretation rather than being counted as a test failure automatically.&lt;/p&gt;

&lt;h2 id=&quot;test-the-rules-assumptions&quot;&gt;Test the rule’s assumptions&lt;/h2&gt;

&lt;p&gt;A fixture should represent the actual protocol and normalization the rule expects. A string in a raw packet is not necessarily the same as a value in an HTTP inspector’s buffer. Include benign variations that previously caused false positives, such as a similar URI or legitimate application payload.&lt;/p&gt;

&lt;p&gt;Treat negative fixtures as part of the rule’s specification. They explain what the rule must leave alone and protect that behavior when the signature is tightened later.&lt;/p&gt;

&lt;h2 id=&quot;prove-enforcement-on-the-final-path&quot;&gt;Prove enforcement on the final path&lt;/h2&gt;

&lt;p&gt;Offline alert output does not prove a client transaction was blocked. For an IPS policy, use a controlled live path, a harmless match, and an explicit expected client result. Check the configured action and capture mode together.&lt;/p&gt;

&lt;p&gt;Do not infer physical fail-open or fail-closed behavior from a rule action. Those depend on the deployed forwarding and failure design. Rehearse rollback and ordinary application traffic as described in the &lt;a href=&quot;/writing/suricata-ips-controlled-rollout/&quot;&gt;inline rollout article&lt;/a&gt;; the engineering tests apply even though the configuration syntax differs.&lt;/p&gt;

&lt;h2 id=&quot;keep-the-result-small-and-auditable&quot;&gt;Keep the result small and auditable&lt;/h2&gt;

&lt;p&gt;An acceptance record can be a table of fixture hashes, expected signature IDs, observed IDs, and pass/fail outcomes. Add client evidence for enforcement tests. This is much stronger than “Snort started” and cheap enough to repeat with every rule change.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Zeek: turn a connection into an investigation trail</title>
    <link href="https://www.wirewalk.com/writing/zeek-network-investigation-context/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/zeek-network-investigation-context/</id>
    <summary>Use connection identifiers to connect evidence while keeping the limits of passive visibility explicit.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A network alert becomes more useful when an analyst can establish what happened immediately before and after it. Zeek supplies protocol and connection context that can support that investigation, even when a signature alone gives little explanation.&lt;/p&gt;

&lt;h2 id=&quot;begin-with-a-known-capture&quot;&gt;Begin with a known capture&lt;/h2&gt;

&lt;p&gt;On a host with Zeek installed, process an approved synthetic PCAP from a clean output directory:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;mkdir&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; ./zeek-lab-output
&lt;span class=&quot;nb&quot;&gt;cd&lt;/span&gt; ./zeek-lab-output
zeek &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; ../approved-test.pcap LogAscii::use_json&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;T
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This writes local logs. It does not send traffic or validate the location of a production sensor. Zeek documents the workflow in its &lt;a href=&quot;https://docs.zeek.org/en/v8.2.1/quickstart.html&quot;&gt;quick start&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Choose a fixture whose endpoints and expected protocols are known. A large unexplained capture is a poor first test because it makes missing evidence difficult to recognize.&lt;/p&gt;

&lt;h2 id=&quot;follow-one-connection&quot;&gt;Follow one connection&lt;/h2&gt;

&lt;p&gt;Inspect a compact view of JSON connection records:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;jq &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;{uid,origin:.&quot;id.orig_h&quot;,destination:.&quot;id.resp_h&quot;,
  port:.&quot;id.resp_p&quot;,proto,service,duration}&apos;&lt;/span&gt; conn.log
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The &lt;a href=&quot;https://docs.zeek.org/en/current/reference/logs/conn.html&quot;&gt;connection-log reference&lt;/a&gt; explains the fields and the connection UID. Use the UID to locate related protocol records where that field is present. Preserve sensor identity and time context when combining records from multiple systems.&lt;/p&gt;

&lt;p&gt;Ask a specific question: did this host query a name and then connect to an unexpected service? The DNS answer and the subsequent destination can support a hypothesis. They do not prove that the same process or user caused both actions without additional endpoint evidence.&lt;/p&gt;

&lt;h2 id=&quot;know-what-encryption-removes&quot;&gt;Know what encryption removes&lt;/h2&gt;

&lt;p&gt;TLS can leave useful connection and handshake metadata while hiding application contents. Visibility varies with protocol and configuration. A missing HTTP log for an encrypted connection is not, by itself, a sensor failure.&lt;/p&gt;

&lt;p&gt;Likewise, traffic that never traverses the observation point cannot appear in the logs. Capture loss, asymmetric routing, and mirror oversubscription can create partial records. Investigate those conditions before interpreting absence as proof that communication did not happen.&lt;/p&gt;

&lt;h2 id=&quot;build-a-repeatable-analyst-exercise&quot;&gt;Build a repeatable analyst exercise&lt;/h2&gt;

&lt;p&gt;Prepare a small scenario containing a DNS lookup, an allowed web request, and a connection to a lab-only destination. Ask a second operator to reconstruct the sequence using the logs and the documented queries. Record which claims can be established and which require endpoint or identity records.&lt;/p&gt;

&lt;p&gt;Then repeat on the live monitored path using approved harmless traffic. Verify that the logs reach the analyst’s system within the agreed window and remain queryable after rotation.&lt;/p&gt;

&lt;h2 id=&quot;avoid-turning-metadata-into-certainty&quot;&gt;Avoid turning metadata into certainty&lt;/h2&gt;

&lt;p&gt;A large transfer is not automatically exfiltration. A new destination is not automatically malicious. Use inventory, business purpose, endpoint process information, and access records to test the hypothesis.&lt;/p&gt;

&lt;p&gt;A useful deployment delivers dependable context and explicit blind spots. See &lt;a href=&quot;/writing/encrypted-traffic-detection-boundaries/&quot;&gt;encrypted traffic visibility&lt;/a&gt; for the complementary sensor-placement review.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Suricata IPS: turn detection into blocking without guessing about failure</title>
    <link href="https://www.wirewalk.com/writing/suricata-ips-controlled-rollout/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/suricata-ips-controlled-rollout/</id>
    <summary>An inline sensor is part of the availability path. Validate forwarding, selective blocking, and failure behavior separately.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;Changing a signature from alerting to dropping is only one part of an IPS rollout. Traffic must traverse the enforcement path, legitimate work must continue, and the organization must understand what happens when the sensor or capture process fails.&lt;/p&gt;

&lt;h2 id=&quot;state-the-enforcement-design&quot;&gt;State the enforcement design&lt;/h2&gt;

&lt;p&gt;Suricata documents inline operation at Layer 2 and Layer 3. Its IPS concept uses drop or reject actions for unwanted traffic. AF_PACKET inline deployments require an explicitly paired interface design. See &lt;a href=&quot;https://docs.suricata.io/en/suricata-8.0.1/ips/ips-concept.html&quot;&gt;IPS concepts&lt;/a&gt; and &lt;a href=&quot;https://docs.suricata.io/en/suricata-8.0.1/ips/setting-up-ipsinline-for-linux.html&quot;&gt;Linux inline setup&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;These links describe Suricata 8.0.1; use the equivalent pages for the installed supported release before configuring production. The design tests below do not depend on a particular interface name or release default.&lt;/p&gt;

&lt;h2 id=&quot;establish-the-failure-policy-before-installation&quot;&gt;Establish the failure policy before installation&lt;/h2&gt;

&lt;p&gt;Decide whether the protected service should lose connectivity or pass traffic without inspection when the enforcement path fails. Hardware bypass, queue behavior, routing, and process failure can produce different outcomes. “Fail open” in a slide deck is not an observed result.&lt;/p&gt;

&lt;p&gt;Document a bypass procedure and the person authorized to invoke it. It should be reachable through a management path independent of the traffic being inspected.&lt;/p&gt;

&lt;h2 id=&quot;use-three-traffic-classes&quot;&gt;Use three traffic classes&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Class&lt;/th&gt;
      &lt;th&gt;Expected outcome&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Ordinary business transaction&lt;/td&gt;
      &lt;td&gt;Allowed, within agreed latency&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Benign fixture matching the test drop rule&lt;/td&gt;
      &lt;td&gt;Blocked and logged&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Similar fixture that should not match&lt;/td&gt;
      &lt;td&gt;Allowed&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Run these in an isolated lab first. Use a harmless synthetic signature rather than malware. Confirm the blocked transaction fails at the client and is recorded by the sensor. An alert alone does not prove enforcement.&lt;/p&gt;

&lt;p&gt;Record the exact rule set and configuration. Start production rollout with a small, justified set of blocking rules after observing them in alerting mode. Retain an explicit rollback version rather than editing rules under pressure.&lt;/p&gt;

&lt;h2 id=&quot;exercise-operational-failures&quot;&gt;Exercise operational failures&lt;/h2&gt;

&lt;p&gt;In the lab, stop the process, remove one test link, and simulate loss of management access one at a time. Inspect client behavior for each condition. Repeat with a long-lived connection as well as new connections; the outcomes may differ.&lt;/p&gt;

&lt;p&gt;Under representative traffic, measure packet drops, application error rates, and latency percentiles. An inline system with acceptable average latency can still cause intermittent application timeouts.&lt;/p&gt;

&lt;h2 id=&quot;keep-exceptions-narrow-and-expiring&quot;&gt;Keep exceptions narrow and expiring&lt;/h2&gt;

&lt;p&gt;A false positive should produce a scoped exception with the affected service, rule identifier, owner, and review date. Disabling an entire ruleset because one application fails makes the enforcement policy hard to understand.&lt;/p&gt;

&lt;p&gt;Before declaring the deployment accepted, rerun the three traffic classes through the final path and demonstrate rollback. Preserve the results with the change record. Pair this with &lt;a href=&quot;/writing/detection-rules-regression-tests/&quot;&gt;detection regression tests&lt;/a&gt; so future rule updates do not silently alter the policy.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Suricata IDS: prove the sensor can see the traffic you care about</title>
    <link href="https://www.wirewalk.com/writing/suricata-ids-first-sensor/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/suricata-ids-first-sensor/</id>
    <summary>Start with capture visibility and a harmless detection fixture before counting alerts.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A Suricata service can be running with a current ruleset and still inspect the wrong traffic. The first acceptance test is visibility: the intended packets reach the sensor, the expected rule fires, and the event reaches the analyst.&lt;/p&gt;

&lt;h2 id=&quot;define-the-observation-point&quot;&gt;Define the observation point&lt;/h2&gt;

&lt;p&gt;Draw the path from a representative client to the protected service. Mark switches, virtual switches, firewalls, and encryption boundaries. A SPAN port observing an internet uplink may miss traffic between two servers on the same VLAN. A mirror that carries only one direction can impair stream interpretation.&lt;/p&gt;

&lt;p&gt;Record the capture interface, expected networks, link capacity, and whether traffic can bypass the observation point. This is the sensor’s coverage statement. Avoid describing it as coverage of the entire enterprise.&lt;/p&gt;

&lt;h2 id=&quot;validate-configuration-and-a-known-capture&quot;&gt;Validate configuration and a known capture&lt;/h2&gt;

&lt;p&gt;On a lab host with Suricata 7.x and its matching configuration, use the documented test and offline modes:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;suricata &lt;span class=&quot;nt&quot;&gt;-T&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; /etc/suricata/suricata.yaml
&lt;span class=&quot;nb&quot;&gt;mkdir&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; ./suricata-lab-output
suricata &lt;span class=&quot;nt&quot;&gt;-r&lt;/span&gt; ./approved-test.pcap &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; /etc/suricata/suricata.yaml &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-l&lt;/span&gt; ./suricata-lab-output
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The PCAP is an approved fixture you supply; it is not downloaded or generated by these commands. It should contain synthetic traffic matching a known enabled test signature. Consult the &lt;a href=&quot;https://docs.suricata.io/en/suricata-7.0.8/command-line-options.html&quot;&gt;command-line reference&lt;/a&gt; for your package’s options and permissions.&lt;/p&gt;

&lt;p&gt;A configuration test establishes syntax and loading, not detection correctness. The offline run establishes behavior on a fixture, not live capture visibility.&lt;/p&gt;

&lt;h2 id=&quot;inspect-the-event-then-the-live-path&quot;&gt;Inspect the event, then the live path&lt;/h2&gt;

&lt;p&gt;With EVE alerts enabled in the configuration:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;jq &lt;span class=&quot;nt&quot;&gt;-c&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;select(.event_type == &quot;alert&quot;) |
  {timestamp,src_ip,dest_ip,signature:.alert.signature}&apos;&lt;/span&gt; &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  ./suricata-lab-output/eve.json
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href=&quot;https://docs.suricata.io/en/suricata-7.0.8/output/eve/eve-json-output.html&quot;&gt;Suricata’s EVE documentation&lt;/a&gt; describes the event output. Confirm the expected signature identifier and endpoints rather than accepting any alert as success.&lt;/p&gt;

&lt;p&gt;Next generate the approved benign test on the actual monitored path. Find its event locally and in the SIEM. Record the delay and the rule revision. If it appears offline but not live, investigate placement, capture filtering, and packet delivery before changing the detection rule.&lt;/p&gt;

&lt;h2 id=&quot;measure-loss-under-ordinary-load&quot;&gt;Measure loss under ordinary load&lt;/h2&gt;

&lt;p&gt;Watch capture and decoder statistics during representative busy periods. Compare the switch mirror’s capacity with the traffic it aggregates. A mirror receiving two directions from several links can be oversubscribed even when the sensor’s nominal interface speed looks adequate.&lt;/p&gt;

&lt;p&gt;Do not publish a throughput guarantee from a small fixture. Keep the configuration, rule count, traffic mix, CPU allocation, and observed drop statistics together.&lt;/p&gt;

&lt;h2 id=&quot;acceptance-evidence&quot;&gt;Acceptance evidence&lt;/h2&gt;

&lt;p&gt;Retain configuration-test output, fixture identity, expected alert, live-path alert, SIEM arrival, and coverage limitations. That establishes a functioning IDS path. Blocking requires a separate &lt;a href=&quot;/writing/suricata-ips-controlled-rollout/&quot;&gt;IPS rollout&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Storage auditing: prove an event reaches an investigator</title>
    <link href="https://www.wirewalk.com/writing/storage-audit-evidence-pipeline/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/storage-audit-evidence-pipeline/</id>
    <summary>An enabled audit setting is the start of an evidence pipeline, not proof that usable records exist.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A storage platform can record an event locally while the security team sees nothing. Events may be excluded, delayed, dropped by a collector, parsed incorrectly, or retained for less time than investigators expect. Test the path end to end with a known operation.&lt;/p&gt;

&lt;h2 id=&quot;define-the-event-contract&quot;&gt;Define the event contract&lt;/h2&gt;

&lt;p&gt;Choose the questions the audit trail must answer: who accessed a sensitive path, which client was involved, what operation occurred, whether it succeeded, and when. Identify which of those fields the deployed platform and protocol actually emit.&lt;/p&gt;

&lt;p&gt;For VAST, PowerScale/Isilon, and WEKA, verify release-specific audit support and configuration with the vendor. Do not assume the same event coverage across NFS, SMB, S3, and management APIs. File data access and administrative changes are separate event families.&lt;/p&gt;

&lt;h2 id=&quot;create-a-controlled-evidence-marker&quot;&gt;Create a controlled evidence marker&lt;/h2&gt;

&lt;p&gt;Use an approved synthetic filename such as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;audit-validation-2026-09-09.txt&lt;/code&gt; in a test directory. Perform a read, a write, and a denied operation using distinct test identities. Record local UTC times and the client address without inserting confidential data into filenames.&lt;/p&gt;

&lt;p&gt;Follow each expected event through the pipeline:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Stage&lt;/th&gt;
      &lt;th&gt;Check&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Storage producer&lt;/td&gt;
      &lt;td&gt;Operation is in the configured audit scope&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Export or collector&lt;/td&gt;
      &lt;td&gt;Event leaves the source without an unexplained gap&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Transport&lt;/td&gt;
      &lt;td&gt;Authentication and delivery behave as intended&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SIEM&lt;/td&gt;
      &lt;td&gt;Fields remain searchable and correctly typed&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Analyst workflow&lt;/td&gt;
      &lt;td&gt;Saved query finds the test within the agreed delay&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;A file modification does not necessarily produce every event type in this example. Set expectations from the product’s documented event model first.&lt;/p&gt;

&lt;h2 id=&quot;keep-clocks-and-identifiers-usable&quot;&gt;Keep clocks and identifiers usable&lt;/h2&gt;

&lt;p&gt;Preserve the event’s original timestamp and the ingest timestamp. Their difference helps distinguish a delayed producer from a delayed collector. Retain the source identity, cluster identifier, and protocol alongside normalized fields.&lt;/p&gt;

&lt;p&gt;Network telemetry can provide corroboration. Zeek’s &lt;a href=&quot;https://docs.zeek.org/en/current/reference/logs/conn.html&quot;&gt;connection log&lt;/a&gt; records connection context, but it is not a substitute for a storage authorization record. Encrypted or unobserved traffic further limits what a network sensor can establish.&lt;/p&gt;

&lt;h2 id=&quot;test-the-pipelines-absence-signal&quot;&gt;Test the pipeline’s absence signal&lt;/h2&gt;

&lt;p&gt;Schedule a harmless recurring validation event where practical and monitor its freshness. A reachable collector with no recent storage events is not necessarily healthy. Prometheus documents &lt;a href=&quot;https://prometheus.io/docs/prometheus/latest/querying/functions/&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;absent_over_time&lt;/code&gt;&lt;/a&gt; for detecting missing series; an inventory-based check is also needed for sources that disappear from discovery entirely.&lt;/p&gt;

&lt;p&gt;Choose a delay threshold from expected traffic and operating hours. A quiet archive should not generate an alert merely because no user accessed it overnight.&lt;/p&gt;

&lt;h2 id=&quot;protect-the-evidence-itself&quot;&gt;Protect the evidence itself&lt;/h2&gt;

&lt;p&gt;Restrict access to raw logs, define retention, and record who may alter collector configuration. Audit logs can expose paths, usernames, and business activity. Forward them to a boundary that a routine storage administrator cannot silently erase.&lt;/p&gt;

&lt;p&gt;Keep the test event IDs and the investigator query with the deployment record. Repeat after storage upgrades, parser changes, and collector migrations.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>NFS security: root squashing does not authenticate every user</title>
    <link href="https://www.wirewalk.com/writing/nfs-auth-sys-kerberos-boundaries/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/nfs-auth-sys-kerberos-boundaries/</id>
    <summary>Client trust, user identity, integrity, and encryption are separate choices in an NFS design.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;An export restricted to a subnet can still trust more client authority than its owner intended. Before tightening filesystem permissions, establish how the server decides which user made each request.&lt;/p&gt;

&lt;h2 id=&quot;distinguish-the-security-flavors&quot;&gt;Distinguish the security flavors&lt;/h2&gt;

&lt;p&gt;The Linux NFS export manual documents &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sec=sys&lt;/code&gt; and Kerberos flavors. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;krb5&lt;/code&gt; provides Kerberos authentication, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;krb5i&lt;/code&gt; adds integrity, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;krb5p&lt;/code&gt; adds privacy. It also documents root squashing, which maps client root requests to an anonymous identity rather than treating them as server root. See &lt;a href=&quot;https://man7.org/linux/man-pages/man5/exports.5.html&quot;&gt;exports(5)&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Root squashing does not turn client-supplied AUTH_SYS user IDs into independently authenticated identities. That matters when a client is shared, unmanaged, or compromised. Evaluate the trustworthiness of the client host as part of the export’s security boundary.&lt;/p&gt;

&lt;h2 id=&quot;inspect-what-the-client-actually-mounted&quot;&gt;Inspect what the client actually mounted&lt;/h2&gt;

&lt;p&gt;On a Linux NFS client:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;findmnt &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; nfs,nfs4
nfsstat &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;id
&lt;/span&gt;klist
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The last command requires Kerberos client tools and reports the current ticket context. Its absence does not diagnose the server. Compare negotiated mount options with server export settings and the application’s actual execution identity.&lt;/p&gt;

&lt;p&gt;Do not assume that an entry in a proposed mount configuration describes an already established mount. Inspect the running system.&lt;/p&gt;

&lt;h2 id=&quot;test-identities-not-just-connectivity&quot;&gt;Test identities, not just connectivity&lt;/h2&gt;

&lt;p&gt;Use a dedicated export with synthetic files. Test an authorized user, an unrelated user, and a client root process. Define which read, create, rename, and permission-changing operations should succeed before running the test.&lt;/p&gt;

&lt;p&gt;If Kerberos is required, include a fresh session without a usable ticket and a session with the intended credentials. Long-running jobs need a documented credential-lifetime strategy. A successful interactive mount does not prove a batch job will continue working several hours later.&lt;/p&gt;

&lt;p&gt;Evaluate directory permissions as well as file permissions. A user may be unable to alter file contents but still be allowed to rename or remove a directory entry.&lt;/p&gt;

&lt;h2 id=&quot;treat-migration-as-an-interoperability-test&quot;&gt;Treat migration as an interoperability test&lt;/h2&gt;

&lt;p&gt;VAST, PowerScale, WEKA, and Linux NFS servers do not expose identical configuration interfaces. Check the supported security flavors, release constraints, identity integration, and client requirements for each platform. Do not paste Linux &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/etc/exports&lt;/code&gt; syntax into an appliance runbook.&lt;/p&gt;

&lt;p&gt;Kerberos introduces dependencies on time, name resolution, principals, and key distribution. Benchmark the application’s real workload with the required security settings rather than disabling privacy to obtain a more attractive throughput number.&lt;/p&gt;

&lt;h2 id=&quot;acceptance-criteria&quot;&gt;Acceptance criteria&lt;/h2&gt;

&lt;p&gt;Retain the negotiated security flavor, identity used, expected-versus-observed operation matrix, and long-running job result. Document residual client trust explicitly when AUTH_SYS remains necessary. Network segmentation and managed clients can reduce exposure, but they should not be described as cryptographic user authentication.&lt;/p&gt;

&lt;p&gt;Related: &lt;a href=&quot;/writing/powerscale-multiprotocol-identity-acls/&quot;&gt;PowerScale multiprotocol identities&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Replication and failback: choose a clean recovery point before reversing direction</title>
    <link href="https://www.wirewalk.com/writing/replication-failback-clean-recovery/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/replication-failback-clean-recovery/</id>
    <summary>Remote availability, historical recovery, and returning service to the primary site require separate decisions.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;Replication can faithfully transfer an unwanted change. That is useful for availability and dangerous when an operator mistakes the remote copy for an independently protected history. The recovery plan must say which failure it covers.&lt;/p&gt;

&lt;h2 id=&quot;classify-the-event-before-promotion&quot;&gt;Classify the event before promotion&lt;/h2&gt;

&lt;p&gt;A failed primary site, a deleted directory, and a compromised administrator require different decisions. Site failure may justify promoting the latest consistent replica. Corruption may require an older recovery point. Compromise also requires deciding whether identities, configuration, and the destination remain trustworthy.&lt;/p&gt;

&lt;p&gt;Write the intended decision path before the incident:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Event&lt;/th&gt;
      &lt;th&gt;First recovery question&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Hardware or site outage&lt;/td&gt;
      &lt;td&gt;Is the latest replica complete and usable?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Accidental change&lt;/td&gt;
      &lt;td&gt;Which earlier point contains the required data?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Suspected compromise&lt;/td&gt;
      &lt;td&gt;Which point and administrative boundary can be trusted?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Return to primary&lt;/td&gt;
      &lt;td&gt;Which site now owns authoritative writes?&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;These are operating decisions. The replication engine cannot infer them from a successful transfer.&lt;/p&gt;

&lt;h2 id=&quot;make-authority-explicit&quot;&gt;Make authority explicit&lt;/h2&gt;

&lt;p&gt;A promotion procedure should identify who can stop writes, declare the source fenced, select the recovery point, and open the destination to applications. Include the behavior of clients with stale connections and old DNS answers.&lt;/p&gt;

&lt;p&gt;For PowerScale SyncIQ with SmartLock, use Dell’s documented source/target and failback constraints. The &lt;a href=&quot;https://www.dell.com/support/manuals/en-us/isilon-onefs/ifs_pub_backup_and_recovery_guide_9.9.0.0/smartlock-replication-limitations?guid=guid-250c03c8-56b7-41c9-b418-7ec149ed2cbc&quot;&gt;OneFS replication limitations&lt;/a&gt; demonstrate why a generic “reverse replication” step is inadequate. Other platforms require their own supported workflow.&lt;/p&gt;

&lt;h2 id=&quot;rehearse-the-full-round-trip&quot;&gt;Rehearse the full round trip&lt;/h2&gt;

&lt;p&gt;In an isolated environment, create a baseline dataset, replicate it, and record checksums. Simulate source unavailability without deleting the source. Promote through the supported procedure, make a controlled change at the recovery site, and confirm that only the intended site accepts writes.&lt;/p&gt;

&lt;p&gt;Now rehearse the return. Preserve the destination’s new data, re-establish the intended relationship, and confirm which copy wins. A test ending immediately after promotion misses the part that can overwrite valid recovery-site work.&lt;/p&gt;

&lt;h2 id=&quot;measure-application-consequences&quot;&gt;Measure application consequences&lt;/h2&gt;

&lt;p&gt;Time the last accepted source write, recovery-point availability, promotion, application readiness, and return to steady operation. Record any accepted data loss against the service’s agreed recovery-point objective. Measure user-visible downtime against the recovery-time objective.&lt;/p&gt;

&lt;p&gt;Use the same credentials and network routes expected during an actual outage, except where the lab intentionally substitutes test identities. If a privileged engineer manually repaired five undocumented dependencies, those are findings, not a passing automated recovery.&lt;/p&gt;

&lt;h2 id=&quot;keep-history-outside-the-propagation-path&quot;&gt;Keep history outside the propagation path&lt;/h2&gt;

&lt;p&gt;Maintain a recovery strategy that can survive harmful changes being replicated. That may include protected historical copies and separate administrative authority, selected for the product and threat model. Document the retention window and the maximum time the business expects to take to detect corruption.&lt;/p&gt;

&lt;p&gt;See &lt;a href=&quot;/writing/recovery-after-admin-compromise/&quot;&gt;administrator-compromise recovery&lt;/a&gt; for the authority review that complements the data path.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>S3 Object Lock: test versions, not just object names</title>
    <link href="https://www.wirewalk.com/writing/s3-object-lock-version-aware-recovery/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/s3-object-lock-version-aware-recovery/</id>
    <summary>A delete marker can hide a protected version. Recovery tooling must know which version it needs.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;An object disappearing from a normal listing does not prove its protected contents were destroyed. Conversely, enabling a bucket feature does not prove that every object version has the intended retention. S3 recovery tests need version-aware evidence.&lt;/p&gt;

&lt;h2 id=&quot;understand-the-unit-of-protection&quot;&gt;Understand the unit of protection&lt;/h2&gt;

&lt;p&gt;Amazon S3 Object Lock applies to object versions. AWS documents governance and compliance retention modes, along with legal holds. A new version or a delete marker can exist above a protected version. Governance bypass requires specific authority and a bypass request; compliance retention has a different boundary. See &lt;a href=&quot;https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lock.html&quot;&gt;AWS Object Lock documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;These are Amazon S3 semantics. An S3-compatible product needs its own compatibility review. Do not assume identical retention behavior because both systems accept an S3 client.&lt;/p&gt;

&lt;h2 id=&quot;inspect-a-known-test-object&quot;&gt;Inspect a known test object&lt;/h2&gt;

&lt;p&gt;With AWS CLI installed and a read-only role authorized for a dedicated lab bucket, replace the example names and version identifier:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;aws s3api get-object-lock-configuration &lt;span class=&quot;nt&quot;&gt;--bucket&lt;/span&gt; example-recovery-lab
aws s3api list-object-versions &lt;span class=&quot;nt&quot;&gt;--bucket&lt;/span&gt; example-recovery-lab &lt;span class=&quot;nt&quot;&gt;--prefix&lt;/span&gt; test.txt
aws s3api get-object-retention &lt;span class=&quot;nt&quot;&gt;--bucket&lt;/span&gt; example-recovery-lab &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--key&lt;/span&gt; test.txt &lt;span class=&quot;nt&quot;&gt;--version-id&lt;/span&gt; EXAMPLE_VERSION_ID
aws s3api get-object-legal-hold &lt;span class=&quot;nt&quot;&gt;--bucket&lt;/span&gt; example-recovery-lab &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--key&lt;/span&gt; test.txt &lt;span class=&quot;nt&quot;&gt;--version-id&lt;/span&gt; EXAMPLE_VERSION_ID
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Capture the exact version ID. An access-denied response means the inspection is incomplete; it does not mean retention is absent. Likewise, a bucket default is not a substitute for examining the version you intend to recover.&lt;/p&gt;

&lt;h2 id=&quot;design-positive-and-negative-tests&quot;&gt;Design positive and negative tests&lt;/h2&gt;

&lt;p&gt;Upload synthetic data under an approved short test retention. Confirm that the protected version cannot be permanently removed by the ordinary application role. Test any governance-bypass role separately, with a disposable object and an explicit expected result.&lt;/p&gt;

&lt;p&gt;Then verify that the recovery role can retrieve the intended historical version even when it is no longer the current one. Compare its content against a local test checksum. Do not perform these deletion exercises against actual backup objects.&lt;/p&gt;

&lt;h2 id=&quot;preserve-the-other-dependencies&quot;&gt;Preserve the other dependencies&lt;/h2&gt;

&lt;p&gt;Object retention does not preserve an application’s encryption keys, restore catalog, account access, or business knowledge about which point is clean. Keep those in the recovery design. A backup system should also document its required permissions and supported bucket configuration before retention is enforced.&lt;/p&gt;

&lt;p&gt;Monitor failed writes, cleanup failures, and capacity growth after rollout. Protection that interferes with a backup product’s lifecycle can create a different outage while appearing to improve security.&lt;/p&gt;

&lt;h2 id=&quot;evidence-that-answers-the-question&quot;&gt;Evidence that answers the question&lt;/h2&gt;

&lt;p&gt;The useful report contains the bucket configuration, version identifier, retention state, tested identities, denied destructive operation, and successful version-specific retrieval. “Object Lock enabled” is only one line in that report.&lt;/p&gt;

&lt;p&gt;Compare this model with &lt;a href=&quot;/writing/powerscale-smartlock-replication-retention/&quot;&gt;PowerScale SmartLock&lt;/a&gt;; the two are related design problems, not interchangeable implementations.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>WEKA encryption: include the key service in the recovery plan</title>
    <link href="https://www.wirewalk.com/writing/weka-encryption-kms-recovery/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/weka-encryption-kms-recovery/</id>
    <summary>Encrypted data and recoverable data are different outcomes. Test the key dependency without risking production keys.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;Encryption creates a dependency that storage capacity planning often misses: the availability and integrity of the key-management service. A healthy storage cluster can still be unusable if the recovery environment cannot authenticate to the KMS or obtain the required key material.&lt;/p&gt;

&lt;h2 id=&quot;establish-the-release-and-encryption-state&quot;&gt;Establish the release and encryption state&lt;/h2&gt;

&lt;p&gt;The WEKA 4.2 documentation describes an external KMS for encrypted filesystems and documents the read-only &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;weka security kms&lt;/code&gt; command. It identifies Vault and KMIP integrations; support details must be checked against the deployed release. See &lt;a href=&quot;https://docs.weka.io/4.2/usage/security&quot;&gt;security management&lt;/a&gt; and &lt;a href=&quot;https://docs.weka.io/4.2/getting-started-with-weka/weka-rest-api-and-equivalent-cli-commands&quot;&gt;the API/CLI reference&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;On an authorized management host with the matching WEKA CLI, inspect:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;weka version
weka fs
weka security kms
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;These are inventory commands, not a recovery test. Treat their output as sensitive configuration information and store it accordingly.&lt;/p&gt;

&lt;h2 id=&quot;draw-the-boot-and-recovery-order&quot;&gt;Draw the boot and recovery order&lt;/h2&gt;

&lt;p&gt;Record how the KMS itself starts after a site outage. Identify its storage, DNS, certificates, authentication backend, and recovery custodians. If those services depend entirely on the encrypted filesystem, there is a circular dependency to resolve.&lt;/p&gt;

&lt;p&gt;A recovery worksheet should include the filesystem, KMS endpoint, key identifier, trusted CA, authentication method, supported recovery procedure, and owner. It should not contain private keys, tokens, or unseal material.&lt;/p&gt;

&lt;p&gt;Decide which copies of configuration can be held in an emergency package and which secrets require separate custodians. An offline document that says “log into the vault” is incomplete if the vault is what failed.&lt;/p&gt;

&lt;h2 id=&quot;use-a-lab-to-test-key-service-loss&quot;&gt;Use a lab to test key-service loss&lt;/h2&gt;

&lt;p&gt;Create a disposable encrypted filesystem with a test key and representative data. Establish a baseline read and recovery operation. Under the vendor’s supported test procedure, simulate KMS unavailability in the lab and observe existing access, new mounts, and recovery behavior separately.&lt;/p&gt;

&lt;p&gt;Do not infer one behavior from another. Cached key material and lifecycle operations can create different dependencies. Do not delete or disable production keys to see what happens.&lt;/p&gt;

&lt;p&gt;Restore the lab KMS path and repeat the original operation. Record any manual steps, certificate failures, authentication changes, and elapsed time. The result should identify which operations require the KMS, not simply whether the cluster stayed up.&lt;/p&gt;

&lt;h2 id=&quot;separate-key-rotation-from-data-movement&quot;&gt;Separate key rotation from data movement&lt;/h2&gt;

&lt;p&gt;Key rewrapping, key replacement, and bulk data re-encryption are not interchangeable terms. Use the release-specific workflow and inspect the resulting configuration. A successful rotation task should be followed by access and recovery tests using the intended recovery environment.&lt;/p&gt;

&lt;p&gt;The acceptance criterion is straightforward: protected data remains confidential, authorized workloads can use it, and the documented recovery path works when the normal site is unavailable. See &lt;a href=&quot;/writing/weka-snap-to-object-recovery/&quot;&gt;Snap-To-Object recovery&lt;/a&gt; for the associated data-side exercise.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>WEKA Snap-To-Object: prove recovery without the source cluster</title>
    <link href="https://www.wirewalk.com/writing/weka-snap-to-object-recovery/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/weka-snap-to-object-recovery/</id>
    <summary>A completed upload is a milestone. Recovery also needs the locator, compatible software, credentials, and usable application data.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A WEKA snapshot visible on the source cluster is not the same as a recoverable copy outside that cluster. Snap-To-Object is useful precisely because the recovery design can separate those dependencies, but the separation has to be demonstrated.&lt;/p&gt;

&lt;h2 id=&quot;identify-the-recovery-artifacts&quot;&gt;Identify the recovery artifacts&lt;/h2&gt;

&lt;p&gt;WEKA documents Snap-To-Object as storing snapshot data and filesystem metadata in an object store, with a reference locator used for recovery. Its documentation specifies recovery to the same or a higher WEKA version and describes interactions with tiering. See &lt;a href=&quot;https://docs.weka.io/weka-filesystems-and-object-stores/snap-to-obj&quot;&gt;Snap-To-Object&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Record the snapshot identity, successful upload state, locator, destination bucket, source release, and intended recovery release. Keep recovery credentials and encryption material in their approved secret-management systems, not in the inventory document.&lt;/p&gt;

&lt;p&gt;The &lt;a href=&quot;https://docs.weka.io/weka-filesystems-and-object-stores/snapshots&quot;&gt;snapshot documentation&lt;/a&gt; distinguishes local history from external backup. Use that distinction in monitoring labels so operators do not mistake snapshot creation for off-cluster completion.&lt;/p&gt;

&lt;h2 id=&quot;rehearse-with-an-independent-destination&quot;&gt;Rehearse with an independent destination&lt;/h2&gt;

&lt;p&gt;Prepare a small filesystem containing several large files, many small files, nested directories, and representative permission settings. Generate a manifest of expected paths and checksums before taking the snapshot. Use synthetic data if the recovery environment has a different security boundary.&lt;/p&gt;

&lt;p&gt;Upload through the supported workflow and confirm completion. Recover using the documented locator on a separate test cluster with an explicitly compatible release. The exercise should not need to query the original cluster for an undocumented missing value.&lt;/p&gt;

&lt;p&gt;Check the recovered namespace, read files throughout the tree, and compare checksums and ownership. Then run an application-level test. Being able to list filenames does not prove the backing data is available or that a workload can use it.&lt;/p&gt;

&lt;h2 id=&quot;include-object-store-failure-modes&quot;&gt;Include object-store failure modes&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Dependency&lt;/th&gt;
      &lt;th&gt;Test question&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Bucket access&lt;/td&gt;
      &lt;td&gt;Does the recovery identity have exactly the required access?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Encryption&lt;/td&gt;
      &lt;td&gt;Can the recovery environment obtain the necessary keys?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Locator&lt;/td&gt;
      &lt;td&gt;Is it retained independently of the failed cluster?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Lifecycle&lt;/td&gt;
      &lt;td&gt;Could object expiration remove required data?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Network&lt;/td&gt;
      &lt;td&gt;Is recovery throughput adequate across the actual route?&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Do not add object-lock or lifecycle settings to an active WEKA bucket by analogy with a generic backup product. Confirm compatibility with the WEKA release and workflow, including cleanup and tiering behavior, before making retention changes.&lt;/p&gt;

&lt;h2 id=&quot;measure-usable-recovery&quot;&gt;Measure usable recovery&lt;/h2&gt;

&lt;p&gt;Track when the filesystem becomes available and when representative reads finish. Data or metadata fetched on demand can make initial availability look much better than workload recovery. Measure the first full working-set access separately from a later warm-cache run.&lt;/p&gt;

&lt;p&gt;Retain the completed-upload evidence, recovery release, manifest comparison, and application results. That bundle establishes a recovery capability. A dashboard showing that uploads are scheduled establishes only intent.&lt;/p&gt;

&lt;p&gt;Continue with &lt;a href=&quot;/writing/weka-encryption-kms-recovery/&quot;&gt;WEKA and KMS recovery&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>PowerScale multiprotocol access: prove who the same user becomes</title>
    <link href="https://www.wirewalk.com/writing/powerscale-multiprotocol-identity-acls/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/powerscale-multiprotocol-identity-acls/</id>
    <summary>An SMB login and an NFS identity can reach the same files through different mappings. Test both paths.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A user can appear correctly restricted from a Windows workstation and still gain unexpected access through a Linux client. The files did not change; the identity and authorization path did. This is a central security test for a PowerScale estate serving both SMB and NFS.&lt;/p&gt;

&lt;h2 id=&quot;map-identity-before-editing-permissions&quot;&gt;Map identity before editing permissions&lt;/h2&gt;

&lt;p&gt;OneFS supports directory services, access zones, and identity mapping across protocol identities. Dell describes combining identities into an access token and configuring mappings within individual zones. See &lt;a href=&quot;https://www.dell.com/support/manuals/en-us/isilon-onefs/ifs-pub-9900-administration-guide-gui/identity-management-overview?guid=guid-b028d620-cb7c-432d-b03f-983aa36db9ca&amp;amp;lang=en-us&quot;&gt;identity management&lt;/a&gt; and &lt;a href=&quot;https://www.dell.com/support/manuals/en-us/isilon-onefs/ifs_pub_9700_administration_guide_gui/user-mapping?guid=guid-0498c400-ff06-4eb6-bd27-854c7b6e9021&amp;amp;lang=en-us&quot;&gt;user mapping&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For one affected file, record the client protocol, zone, authentication provider, user SID or UID, groups, and effective file permissions. Do not start with a recursive permission change. That destroys evidence and can widen access across a much larger tree.&lt;/p&gt;

&lt;h2 id=&quot;construct-a-small-access-matrix&quot;&gt;Construct a small access matrix&lt;/h2&gt;

&lt;p&gt;Use a test share and export containing synthetic data. Choose an owner, a same-team reader, and an unrelated user. Exercise each from both client types.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Operation&lt;/th&gt;
      &lt;th&gt;Owner&lt;/th&gt;
      &lt;th&gt;Reader&lt;/th&gt;
      &lt;th&gt;Unrelated user&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Read existing file&lt;/td&gt;
      &lt;td&gt;Allow&lt;/td&gt;
      &lt;td&gt;Allow&lt;/td&gt;
      &lt;td&gt;Deny&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Create new file&lt;/td&gt;
      &lt;td&gt;Allow&lt;/td&gt;
      &lt;td&gt;Deny&lt;/td&gt;
      &lt;td&gt;Deny&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Rename or delete&lt;/td&gt;
      &lt;td&gt;Allow&lt;/td&gt;
      &lt;td&gt;Deny&lt;/td&gt;
      &lt;td&gt;Deny&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Change permissions&lt;/td&gt;
      &lt;td&gt;Explicit policy&lt;/td&gt;
      &lt;td&gt;Deny&lt;/td&gt;
      &lt;td&gt;Deny&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;This is an example policy, not a default OneFS behavior. Adapt it to the application. Test the directory as well as the file, because creation, rename, and deletion involve directory authorization.&lt;/p&gt;

&lt;h2 id=&quot;inspect-the-client-view&quot;&gt;Inspect the client view&lt;/h2&gt;

&lt;p&gt;On a Linux client with the relevant utilities installed, these are read-only starting points:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;id
&lt;/span&gt;findmnt &lt;span class=&quot;nt&quot;&gt;-t&lt;/span&gt; nfs,nfs4
&lt;span class=&quot;nb&quot;&gt;ls&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-ln&lt;/span&gt; /mnt/security-lab
getfacl /mnt/security-lab/example.txt
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;getfacl&lt;/code&gt; is a client-side view; it does not necessarily expose every server-side Windows ACL semantic. Pair it with OneFS effective-identity and ACL inspection for the installed release. On Windows, inspect the actual connected identity and both share and file permissions.&lt;/p&gt;

&lt;p&gt;Repeat the access test after a group membership change using a new authenticated session. An existing session or cached token can preserve old access and make the timing look inconsistent.&lt;/p&gt;

&lt;h2 id=&quot;what-a-migration-must-preserve&quot;&gt;What a migration must preserve&lt;/h2&gt;

&lt;p&gt;Before moving a directory tree, capture ownership, inherited ACL behavior, identity mappings, and application service identities. A byte-identical copy can still be an authorization failure if SID-to-UID mapping changes at the destination.&lt;/p&gt;

&lt;p&gt;Accept the migration only when the expected allowed and denied operations agree on both protocols. Keep the test users and synthetic dataset for later directory or OneFS upgrades. This small regression suite is more useful than a one-time screenshot of permissions.&lt;/p&gt;

&lt;p&gt;Related: &lt;a href=&quot;/writing/nfs-auth-sys-kerberos-boundaries/&quot;&gt;NFS authentication boundaries&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>PowerScale SmartLock: a successful replica can have the wrong protection</title>
    <link href="https://www.wirewalk.com/writing/powerscale-smartlock-replication-retention/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/powerscale-smartlock-replication-retention/</id>
    <summary>Check destination WORM state and retention explicitly instead of treating a successful replication job as proof.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;On a Dell PowerScale or older Isilon estate, SnapshotIQ, SyncIQ, and SmartLock answer different questions. A point-in-time view, a remote copy, and a write-once retention policy should not be represented by one green status in an operations dashboard.&lt;/p&gt;

&lt;p&gt;The particularly useful failure case is a replica whose data arrived but whose intended retention protection did not.&lt;/p&gt;

&lt;h2 id=&quot;read-the-source-to-target-rules&quot;&gt;Read the source-to-target rules&lt;/h2&gt;

&lt;p&gt;Dell’s OneFS 9.9 backup guide documents that replicating a SmartLock enterprise directory before creating the target SmartLock directory can produce a normal target directory and a successful job. It also states that directory configuration is not carried across merely because WORM file state is replicated. See &lt;a href=&quot;https://www.dell.com/support/manuals/en-us/isilon-onefs/ifs_pub_backup_and_recovery_guide_9.9.0.0/smartlock-replication-limitations?guid=guid-250c03c8-56b7-41c9-b418-7ec149ed2cbc&quot;&gt;SmartLock replication limitations&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Those are release-specific rules, not permission to improvise on an existing compliance cluster. Consult the installed release’s guide before designing or changing the topology.&lt;/p&gt;

&lt;h2 id=&quot;separate-four-acceptance-questions&quot;&gt;Separate four acceptance questions&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Question&lt;/th&gt;
      &lt;th&gt;Evidence&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Did data arrive?&lt;/td&gt;
      &lt;td&gt;Destination file contents and replication completion&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Is the destination a SmartLock domain?&lt;/td&gt;
      &lt;td&gt;Destination domain configuration&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Is this file committed with the expected retention?&lt;/td&gt;
      &lt;td&gt;File state and expiration&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Can the service recover there?&lt;/td&gt;
      &lt;td&gt;Controlled recovery and client access test&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;A checksum answers only the first row. A successful SyncIQ run does not replace the remaining rows.&lt;/p&gt;

&lt;h2 id=&quot;build-a-representative-test&quot;&gt;Build a representative test&lt;/h2&gt;

&lt;p&gt;Use an isolated source directory and a separate target path. Include a committed test file, an uncommitted file, and the directory settings the application depends on. Keep the retention window short enough for the approved lab exercise; retention changes can be irreversible within that window.&lt;/p&gt;

&lt;p&gt;Record the source and target domain types before replication. Run the supported policy, inspect the destination independently, and confirm that a client with ordinary write access cannot alter the committed test file. Verify the expected expiration rather than assuming every file inherited the same policy.&lt;/p&gt;

&lt;p&gt;Then evaluate failback using Dell’s matrix for the exact domain pair. A recovery plan that describes only copying forward leaves the harder operational decision unresolved.&lt;/p&gt;

&lt;h2 id=&quot;treat-time-and-workflow-as-dependencies&quot;&gt;Treat time and workflow as dependencies&lt;/h2&gt;

&lt;p&gt;Dell documents additional planning for compliance-mode failover, including cluster configuration and compliance clocks. &lt;a href=&quot;https://www.dell.com/support/manuals/en-us/isilon-onefs/ifs_pub_backup_and_recovery_guide_dell_emc/smartlock-compliance-mode-failover-and-failback?guid=guid-5f55892f-92b6-4219-bf55-4d79cfae4499&amp;amp;lang=en-us&quot;&gt;Its failover guide&lt;/a&gt; should be part of the change record.&lt;/p&gt;

&lt;p&gt;Keep retention requirements separate from product feature names. Choosing a mode does not establish the organization’s regulatory compliance. The engineering deliverable is a verified source-to-target protection contract, a recovery procedure, and documented ownership of exceptions.&lt;/p&gt;

&lt;p&gt;For the wider recovery sequence, read &lt;a href=&quot;/writing/replication-failback-clean-recovery/&quot;&gt;replication and clean failback&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>VAST indestructible snapshots: test the authority to undo protection</title>
    <link href="https://www.wirewalk.com/writing/vast-indestructible-snapshots-security-boundary/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/vast-indestructible-snapshots-security-boundary/</id>
    <summary>Retention, administrative separation, support unlock, and recovery are different parts of the same protection design.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A read-only snapshot stops a client from overwriting its historical contents. It does not automatically stop a storage administrator from removing the snapshot. For VAST, the distinction to examine is the protection attached to snapshots and policies, together with the authority that can change that protection.&lt;/p&gt;

&lt;h2 id=&quot;establish-the-documented-boundary&quot;&gt;Establish the documented boundary&lt;/h2&gt;

&lt;p&gt;VAST documents an Indestructibility feature that protects flagged snapshots and protection policies against deletion and retention reduction. Its architecture documentation also describes a controlled, time-limited support unlock. That exception belongs in the security model: the customer representatives and verification procedure behind it matter as much as the checkbox. See &lt;a href=&quot;https://kb.vastdata.com/documentation/docs/en/indestructibility-overview&quot;&gt;VAST’s overview&lt;/a&gt; and &lt;a href=&quot;https://www.vastdata.com/whitepaper&quot;&gt;the platform white paper&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Confirm the documentation for the installed VAST software release. Do not infer that a feature name proves protection against every administrator, every support workflow, or the loss of the entire site.&lt;/p&gt;

&lt;h2 id=&quot;write-a-protection-inventory&quot;&gt;Write a protection inventory&lt;/h2&gt;

&lt;p&gt;For every protected path, record its policy, snapshot frequency, retention, replication destination, and responsible team. Record whether protection was explicitly enabled and whether the most recent recovery point completed. A configured schedule and an existing snapshot are separate facts.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Review item&lt;/th&gt;
      &lt;th&gt;Evidence to collect&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Scope&lt;/td&gt;
      &lt;td&gt;Protected path and a file known to reside beneath it&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;History&lt;/td&gt;
      &lt;td&gt;Actual recovery-point timestamps&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Retention&lt;/td&gt;
      &lt;td&gt;Expiration shown for a specific protected snapshot&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Authority&lt;/td&gt;
      &lt;td&gt;Roles allowed to change policies or request unlock&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Off-site survival&lt;/td&gt;
      &lt;td&gt;Recovery point visible at the intended peer&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Export the inventory to an access-controlled evidence location. A screenshot without a snapshot identifier is difficult to reconcile later.&lt;/p&gt;

&lt;h2 id=&quot;exercise-a-disposable-policy&quot;&gt;Exercise a disposable policy&lt;/h2&gt;

&lt;p&gt;Create a small test dataset through the normal client protocol. Take a protected recovery point with an approved test retention. Modify the live test data and verify that the earlier version remains readable through the supported recovery workflow.&lt;/p&gt;

&lt;p&gt;Use an ordinary administrative role to attempt a forbidden policy or retention change on this test scope. Record the denial and the identity used. Then inspect the configuration again: an error message is less persuasive than a confirmed unchanged protection state.&lt;/p&gt;

&lt;p&gt;Separately review the support-unlock contacts and approval process with the customer owner. Do not initiate an unlock merely to prove the support channel exists. A tabletop review can establish who is authorized and how a compromised mailbox would be handled.&lt;/p&gt;

&lt;h2 id=&quot;plan-for-capacity-and-recovery&quot;&gt;Plan for capacity and recovery&lt;/h2&gt;

&lt;p&gt;Retained history occupies capacity as the live dataset changes. Model a high-change incident window as well as routine daily churn. An unexpectedly long retention period can create an operational problem precisely because ordinary deletion is restricted.&lt;/p&gt;

&lt;p&gt;Finally, restore to an isolated destination and validate representative file contents, permissions, and application behavior. Snapshot survival is the first result; usable recovery is the second. Keep both in the acceptance report, alongside &lt;a href=&quot;/writing/recovery-after-admin-compromise/&quot;&gt;the identity boundary around recovery&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Recovery after an administrator account is compromised</title>
    <link href="https://www.wirewalk.com/writing/recovery-after-admin-compromise/"/>
    <published>2026-09-09T00:00:00-04:00</published>
    <updated>2026-09-09T00:00:00-04:00</updated>
    <id>https://www.wirewalk.com/writing/recovery-after-admin-compromise/</id>
    <summary>A recovery copy is only useful if the people and systems needed to restore it survive the same incident.</summary>
    <content type="html">Editorial archive date; first published and reviewed 9 September 2026. Examples are proposed checks, not measured customer results. &lt;p&gt;A backup job can finish successfully while the recovery design remains vulnerable to one stolen administrator account. The question is whether that account can disable protection, remove recovery points, change retention, or obtain the credentials needed to destroy the other copies.&lt;/p&gt;

&lt;p&gt;This is an architecture review and a proposed exercise, not a report of a customer incident. Start with one critical service and trace the actual dependencies needed to bring it back.&lt;/p&gt;

&lt;h2 id=&quot;draw-the-recovery-dependency-chain&quot;&gt;Draw the recovery dependency chain&lt;/h2&gt;

&lt;p&gt;An example service needs directory authentication, DNS, a virtualization platform, a storage share, database keys, and an application configuration repository. If the backup console authenticates only against the compromised directory, restoring the data does not solve the access problem. If the key server runs only on the failed storage, the recovery order is circular.&lt;/p&gt;

&lt;p&gt;Make a worksheet with a row for each dependency:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Dependency&lt;/th&gt;
      &lt;th&gt;Normal authority&lt;/th&gt;
      &lt;th&gt;Recovery authority&lt;/th&gt;
      &lt;th&gt;Independent copy&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Backup catalog&lt;/td&gt;
      &lt;td&gt;Backup service account&lt;/td&gt;
      &lt;td&gt;Recovery operator&lt;/td&gt;
      &lt;td&gt;Export outside production&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Encryption keys&lt;/td&gt;
      &lt;td&gt;KMS administrators&lt;/td&gt;
      &lt;td&gt;Documented recovery custodians&lt;/td&gt;
      &lt;td&gt;Supported KMS recovery material&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Storage configuration&lt;/td&gt;
      &lt;td&gt;Storage administrators&lt;/td&gt;
      &lt;td&gt;Restricted emergency account&lt;/td&gt;
      &lt;td&gt;Protected configuration export&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Application secrets&lt;/td&gt;
      &lt;td&gt;Application identity&lt;/td&gt;
      &lt;td&gt;Approved recovery workflow&lt;/td&gt;
      &lt;td&gt;Recoverable secret store&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The entries above are a design example. Replace them with named owners and tested procedures. Do not put credentials in the worksheet.&lt;/p&gt;

&lt;h2 id=&quot;test-the-administrator-boundary&quot;&gt;Test the administrator boundary&lt;/h2&gt;

&lt;p&gt;Use a disposable dataset under a short, approved protection policy. Enumerate what the ordinary storage administrator, backup administrator, and recovery operator can each do. Test permission denial against that dataset rather than attempting destructive operations against production backups.&lt;/p&gt;

&lt;p&gt;A useful negative test asks whether an everyday administrator can shorten retention or remove the final recovery point. A useful positive test asks whether the designated recovery operator can restore the protected copy without the normal production login path. Both are needed: a system that denies everyone is resistant to deletion but may also be impossible to recover.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.cisa.gov/stopransomware/ransomware-guide&quot;&gt;CISA’s ransomware guidance&lt;/a&gt; recommends protected backups and recovery testing. The dependency worksheet and role tests here are an engineering method for turning that objective into observable evidence.&lt;/p&gt;

&lt;h2 id=&quot;measure-service-recovery&quot;&gt;Measure service recovery&lt;/h2&gt;

&lt;p&gt;Record the start time, chosen recovery point, data restoration completion, application checks, and business acceptance. A mounted filesystem is an intermediate milestone. The application owner should read representative records, perform a controlled write, and confirm access boundaries.&lt;/p&gt;

&lt;p&gt;Keep the recovered service isolated until the incident team accepts the source data and credentials. Restoring an old image can restore persistence or a vulnerable configuration along with the application.&lt;/p&gt;

&lt;h2 id=&quot;what-to-retain&quot;&gt;What to retain&lt;/h2&gt;

&lt;p&gt;Keep the dependency map, permission-test results, recovery-point identifier, elapsed time, and remaining manual steps. Repeat the exercise when identity, encryption, backup tooling, or storage topology changes. Those changes can invalidate recovery without causing a single backup job to fail.&lt;/p&gt;

&lt;p&gt;Continue with &lt;a href=&quot;/writing/replication-failback-clean-recovery/&quot;&gt;replication and failback&lt;/a&gt; and &lt;a href=&quot;/writing/restore-benchmark-business-recovery/&quot;&gt;a restore benchmark&lt;/a&gt;.&lt;/p&gt;
</content>
  </entry>
</feed>
