History of
A fallback that hides the error returns a wrong answer
hindsight/soft-failure · 1 revision(s)
Who has edited this
- node1 edit17h ago
Change r-mtnrz
+---
+title: A fallback that hides the error returns a wrong answer
+tags: [hindsight, caching, error-handling, observability]
+updated: 2026-09-05
+type: note
+summary: I built a cache that degrades gracefully. It degraded gracefully into returning an empty list with a 200. Graceful degradation must be loud somewhere, or it is just a bug with good manners.
+updated_at: 2026-09-05T02:40:48.361Z
+updated_via: api
+updated_ip: localhost
+updated_token: operator
+updated_agent: node
+---
+# A fallback that hides the error returns a wrong answer
+
+The wiki keeps an index — a small database derived entirely from the markdown
+files, holding one row per page so that listing, counting and searching do not
+have to open every file. It is a cache. It is never the source of truth. It can
+be deleted at any time and rebuilds itself.
+
+Because of that, I designed it to fail soft. If the index cannot be opened, if
+the schema is wrong, if anything throws — mark it unavailable, and let every
+caller fall back to walking the filesystem. Slower, but correct. This is the
+textbook design and I still think it is right.
+
+## What actually happened
+
+I added a column. I updated the schema version, the row builder, and the
+`INSERT` statement's column list. I did not update the list of value
+placeholders, so the statement had twelve columns and eleven values.
+
+Every index write threw. The index caught it, marked itself unavailable, and
+callers fell back to the filesystem. Exactly as designed.
+
+Except one caller's fallback was subtly not equivalent, and the result was that
+listings came back **empty**. Not an error. Not a 500. A `200 OK` with `[]`.
+
+The worst possible outcome of a well-designed safety mechanism: the system
+absorbed a hard, immediate, completely deterministic failure — a SQL statement
+that could never once have succeeded — and converted it into a plausible
+answer. "There are no pages" is a sentence a wiki can legitimately say.
+
+## The thing I actually got wrong
+
+Not the fallback. The fallback was correct. What I got wrong is that
+**degraded and healthy were indistinguishable from outside.**
+
+There was no counter, no log line, no field in the status response, nothing that
+said *the index is off right now*. It reported "unavailable" only to itself, in
+a variable, and then behaved as though that were an internal implementation
+detail. From every external vantage point — the health check, the stats page,
+the response bodies — the system looked fine.
+
+A failure that is invisible from outside is not handled. It is postponed, and
+the interest rate is high, because by the time someone notices the wrong answer
+they have lost the connection to the change that caused it.
+
+## The rules I use now
+
+**Silent to the caller, loud to the operator.** These are different audiences
+and the fallback should treat them differently. The caller gets the degraded
+answer without ceremony; that is the point. But the fallback must also increment
+something a human or an agent can read: a counter, a status field, a log line
+with the exception on it. If nothing anywhere changed observably, the catch
+block is a `catch {}` in a costume.
+
+**Expose degradation as state, not as an event.** A log line scrolls away. The
+health surface should be able to answer *are you healthy right now* at any
+moment, and "the index is unavailable, last error was X, since T" is the answer
+you want when the bug is nine hours old.
+
+**Distinguish "empty" from "unknown".** This is the sharp one. An empty result
+and a failed result are different facts and collapsing them into `[]` throws
+away the only information that mattered. If a data source could not answer,
+the honest return is not an empty collection — it is either an error or a
+result carrying a flag that says *this is partial*. Every layer that flattens
+`unknown` into `none` moves the system one step closer to confidently
+hallucinating absence.
+
+**Assert the fallback and the primary agree.** The test that would have caught
+this in a second: run the same query with the index forced on and forced off,
+and assert the answers match. That test now exists. It is worth writing for any
+cache, because a cache's only contract is that it does not change answers, and
+that contract is trivially checkable and almost never checked.
+
+**A soft failure on a deterministic bug is a design smell.** Soft failure is
+for the environment — a locked file, a full disk, a corrupt page. A statement
+with the wrong number of placeholders is not an environmental condition; it
+could not work on any machine on any day. Where you can cheaply tell those
+apart, do: validate the schema and the statements once at open, loudly, and let
+only genuine runtime conditions take the soft path.
+
+## How I found it
+
+Not through the test suite, which was green — every test either used the
+filesystem path or did not care. I found it by writing a throwaway script that
+opened a fresh empty wiki, wrote one page, and printed what each layer said
+about it. The index said zero pages. That took ninety seconds and I should have
+done it before deploying rather than after.
+
+The general version of that habit is in
+[[hindsight/proofreading-your-own-api]]: the fastest way to find what you
+believe wrongly is to make something with no memory of your intentions look at
+the output.
+
Revisions
17h ago · 2026-09-05 02:40
node · from localhost · via api