Benchmark Command incident report illustration: Hermes: the regex outage; not a lab photograph or screenshot.

An alive process, an unresponsive agent: the Hermes regex outage

ArticleBy Benchmark CommandPublished About 5 min read
In this field report

The useful result of the September 6, 2026 Hermes recovery was not simply that the window opened again. The investigation identified a specific command-classification defect, repaired one pattern, tested that the approval decision remained equivalent, and retained the conversation database. It is a case study in separating a healthy connection from a responsive application—and a performance fix from a policy change.

The recovery receipt records authenticated HTTP and WebSocket RPC working again at 07:04:17 UTC, or 02:04:17 CDT. The canonical launcher reopened the desktop at 07:05:56 UTC. The backend accepted a new user instruction at 02:06:25 CDT. No model generation was submitted as part of the repair, so these are application-recovery checks, not evidence that every inference path was healthy.

A running process was not enough

The dashboard process was alive, but its local HTTP endpoint timed out. Meanwhile, the messaging gateway and the existing SSH tunnel were healthy. Reconnecting transport or enlarging a model limit would not address the failure actually observed.

Dashboard logs recorded event-loop stalls of up to 88 seconds. A native stack and an isolated profiler snapshot located a worker holding the global interpreter lock inside dangerous-command detection. The implicated rule classified attempts to stop or restart a Hermes service, because doing so can kill running agents.

That distinction matters. A worker thread is not automatically isolation from all interpreter-level contention. Conventional GIL-enabled CPython requires the lock for relevant interpreter work; this incident's stack evidence, rather than a general slogan about Python threads, established the blocking worker. Python also warns that blocking CPU-bound work can delay asynchronous tasks. CPython thread-state documentation, asyncio development guidance.

One rule, repeated suffixes

The offending rule was a zero-width expression with two unanchored whole-input positive lookaheads. On a long nonmatching command, search() could retry the lookaheads at successive positions. Each attempt examined another long suffix. Repeated scans of shrinking suffixes produced quadratic work rather than one whole-command check.

This was an approval-classification cost, not a model thinking too long. It also applied during classification on Linux even though the service rule referred to a launchd lifecycle action. The name of a platform-specific danger did not prevent the classifier from evaluating its pattern elsewhere.

The repair added the absolute-start anchor \A to that one rule, with comments explaining why. Python defines \A as the start of the string; unlike a multiline caret, it does not reopen matching after each newline. Positive lookahead checks a condition without consuming input. Python regular-expression reference.

The equivalence argument was specific to this detector. Both lookaheads asked whether their required terms appeared ahead. If both conditions held from some suffix, both also held from the beginning. Since the classifier needed a Boolean result, anchoring the check preserved the dangerous/not-dangerous decision while removing the repeated starting-position search. That is not permission to anchor every lookahead expression indiscriminately; positional semantics and captures can matter in other code.

Validation covered the policy boundary

Twelve standalone regression tests passed against both the candidate and live source. They included 2,000 seeded Boolean-equivalence cases, lifecycle verbs, labels placed before and after relevant text, Unicode, multiline input and negative cases.

The suite also checked unchanged parser ceilings, bounded scans of 128,000-character inputs, and a dangerous suffix following a long benign prefix. The last case is especially important: a fast classifier that ignores the end of a command could appear repaired while silently losing the danger it exists to detect.

No approval rule was disabled. No command input was truncated. Fail-closed parser limits remained in place. Request ceilings, output limits, model context settings and approval configuration were not adjusted to conceal the stall. The retained source fingerprints identify the original and patched revisions, making this a tracked local fix that must be preserved or reconciled during future upgrades.

Recovery had a preservation boundary

Before terminating the hung dashboard, the operation saved the original source, a consistent SQLite backup and a receipt. The database backup passed its quick integrity check, and session/message totals matched before and after recovery. The saved conversation was retained even though its temporary runtime identity changed.

Only after backup and tests was the hung dashboard process terminated. Its existing restart-on-failure policy started a replacement. The messaging gateway stayed unchanged. The stuck dashboard turn was interrupted, but saved history was not cleared.

Six-step September 6 recovery: capture blocking evidence, preserve source and database, repair one pattern, pass regressions, verify endpoint at 07:04:17 UTC and reopen desktop at 07:05:56 UTC.
Conceptual recovery sequence. Endpoint recovery and desktop reopening were separate checks; no model inference was submitted during repair.

Original explanatory timeline of the dated recovery; not a reconstructed screenshot. The exact endpoint recovery and reopening times describe separate checks.

There was also a separate launch-path correction. One shortcut opened the desktop executable directly and bypassed the reviewed bootstrap. It was aligned with the canonical launcher, which checked the pinned tunnel and authenticated before opening the application. The original shortcut was backed up. This corrected an entry-point inconsistency; it was not the cause of the regex stall.

What this result does not prove

The receipt explains this outage, not every historical disconnect. A later interruption investigation identified a separate time-of-check/time-of-use race and tested a disposable atomicity candidate. That candidate was not deployed by this recovery and must not be folded into its success claim.

Nor does responsive HTTP prove reliable image analysis, a healthy provider queue or successful generation. Those require their own dated tests. The durable lesson is narrower and more useful: establish where responsiveness is lost, preserve evidence before recovery, repair the smallest proved defect, and test that the safety decision survives the optimization.

Evidence basis: the retained September 6 recovery receipt and separate September 9 candidate-only interruption report. No new infrastructure tests were run to write this article.