INC-40211: hesper-04, 04:11
Ticket INC-40211 · Severity 2 · Opened 2026-09-06 04:11 · Author platform on-call, solo, second week · Status resolved 04:29, resolver: not me
04:11
The pager says hesper-04 array degraded, LSN replay lag 0 MB. That is what
it says. I copied it before I interpreted anything.
Two things about the first minute. One: the primary nameserver has no record
for hesper-04 — NXDOMAIN, I checked twice, awake now. Two: the thing at
192.0.2.53 does. It answers 192.0.2.34, the way it has always answered,
with no name and no one who added it, in the resolv.conf of all 1,181
hosts (resolvers). I went to the address
direct. ssh 192.0.2.34 took a password I did not give it and did not ask
for one either. I did not sleep after that.
04:14 — the runbook
Pager text is hesperctl output, so I opened HES-04, the failover runbook.
It says it applies to hesper-01 and hesper-03. I do not know why
hesper-04 pages like a hesper. I ran step 1 anyway, on hesper-04, because it
is the box I was talking to:
role: standby
health: degraded (battery)
primary: hesper-03That is not a monitoring fault, so step 1 does not tell me to stop. Steps 3 through 7 fence and promote and move a VIP, and the runbook does not know this machine exists. I stopped at step 2.
04:17 — step 2
Step 2: write both LSNs in the incident channel before touching anything. hesper-03 is at 4,412,908,113. hesper-04 is at 4,412,908,113. Not one byte behind. Which is exactly what the number would be if I am not the thing it is standing by for.
04:20 — what hesper-04 is
Checked the register while waiting for a reply from nobody. No row. The night reconciler proposes a row for the one ghost everyone has given up on closing (corvid-c) and has never once proposed one for hesper-04. Nothing to add, nothing to remove. Never seen.
04:22 — step 10, out of order
I know. I went and read /var/run/hesper.owner because the last person who
failed a hesper over is written there and I needed to know who to page. The
file is one line. It says halloway.
Step 10 says if you recognise the name, page them and hand over. halloway's account was closed on 2026-01-30, which is the runbook's own revision-history footnote (HES-04). The account was created on 2026-07-02 at 09:14:00, which is the bottom of the mail archive (correspondence). It is 04:22 and I have two dates from two wiki pages and I cannot make either one be the mistake. So I did not page. And I did not overwrite the file. That is the part of step 10 I can actually follow.
04:29
The alert cleared. Ticket auto-resolved: monitoring, self-healed. There is
no close action in my session; I watched it happen on screen. I then checked
the new collector for any poll to 192.0.2.34 — there is no check and no
sample, but there is no check for the impossible one either, and that one is
answered nine times a minute (chk-0041). The
collector's records are the only records. I wrote that down before I knew why.
Morning
Primary healthy, VIP unmoved, nothing lost. I file this because I opened a ticket and someone else closed it, and I want my version in the record. If platform reads this: the cabinet in row 9 has two BBU-02 spares, a 2019 one and a newer one with a smudged label, and hesper-04's array says its battery is degraded, and hesperctl has no reason to lie about batteries. I did not touch the cabinet. I have no work order. I do not know what hesper-04 is keeping itself ready for.
— oncall-platform-b, rota week 36
Related: runbook-failover-hesper, correspondence, the-fourth-nameserver, inventory-diff, quorum, index.