Innovation is moving at unprecedented speed. Over the last few days, the signal I found worth paying attention to came from what happened when people could not reliably reach the intelligence they had started to depend on. A better model can expand what a team attempts. An interruption reveals which parts of that team's operating model still work without it.
My thesis is simple: every important AI workflow needs a degraded mode with a clear service promise. That promise should describe the useful work the organization can still complete when its preferred model, interface, or supporting service fails.
The recent evidence warrants precision. OpenAI's status history records elevated errors across ChatGPT and Codex on September 3 and regional problems on September 4. Anthropic's status page records September 3 errors affecting several models and subsequent recovery. These records establish interruptions in named services. They do not establish a common cause, a universal API outage, or the failure of every customer's workflow. The primary records appear in the references below.
The engineering implication extends beyond those incidents. As teams delegate more work, they also move knowledge about how that work happens into model sessions, connectors, and automated sequences. A service interruption can therefore remove both a worker and the instructions that another worker needs to continue.
Consider a hypothetical customer operations team that uses an agent to investigate delivery complaints. The agent reads order history, checks carrier events, drafts a response, and proposes a remedy. Customers care about whether the team receives their complaint, preserves the evidence, and resolves the issue. They rarely care which model helps the team do it.
That distinction changes the fallback design. During a model interruption, the team might preserve intake and case identifiers, show the latest verified carrier event, and queue judgment-heavy remedies for review. It can tell customers what it knows and when a person will respond. It should avoid inventing a resolution simply to maintain the appearance of automation.
Figure 1.1 separates the service promise from the intelligence that usually supports it. The middle layer gives the team a smaller, honest product during an interruption.
Figure 1.1: Preserve the service promise.
I would treat that smaller product as part of the product itself. A fallback that requires an engineer to improvise a database query under pressure offers little assurance. The organization needs a supported path that an on-call operator can actually use.
Model diversity can help when a particular provider or model fails. It cannot rescue a workflow whose shared identity service, document store, network route, or tool gateway has stopped working. Two model endpoints may still depend on the same broken path to the customer's records.
Figure 2.1 shows why counting providers can overstate resilience. The shared gateway in this illustrative architecture can block both routes even when both models remain healthy.
Figure 2.1: Find the shared dependency.
The substitution problem also reaches beyond connectivity. A replacement model may interpret tools differently, omit a field, or require a different context budget. A successful health check tells you that the endpoint responds. It does not tell you that the workflow still honors its promises.
I would reserve automatic failover for paths that the team has already tested with representative cases and approved for the relevant data. For other paths, a queue with a truthful status can deliver a better result than an untested substitute. Resilience requires a choice about what to preserve, what to simplify, and what to suspend.
An interruption rarely arrives between perfectly separated tasks. A tool may accept an update just before the agent loses its connection. The application then faces an ambiguous result: the absence of a response does not prove that the action failed.
Imagine that the delivery agent submitted a replacement order and timed out before receiving its identifier. Restarting the entire conversation could create a second order. The correct next step depends on the order system's durable record, not the model's recollection of its last attempt.
Figure 3.1 makes the ambiguity explicit. A recovery path should check the destination before it repeats an action whose outcome it does not know.
Figure 3.1: Reconcile before retrying.
This requirement changes what teams should store outside the conversation. A durable task record should capture the intended action, the operation identifier, the latest confirmed result, and the unresolved question. A replacement worker then needs to interpret a bounded record rather than reconstruct an entire exchange.
That record does not create exactly-once execution by itself. The destination needs to support deduplication or an authoritative lookup, and the team needs a manual reconciliation path when neither exists. The value lies in making uncertainty visible before another action compounds it.
For paid readers, the next section turns this architecture into a one-week operating exercise: define the minimum service, test the interruption, and clear the recovery backlog without duplicating work.