Fail-Open vs Fail-Closed: Security Checks That Won't Break Your Site
When a security check errors, should the request go through or be blocked? A framework for choosing fail-open or fail-closed, with patterns that make either safe.
Every security check you put in front of your application can fail. The database it reads from times out, a dependency throws, a config file is malformed. At that moment the check has to do something with the request in hand. It can let it through (fail open) or reject it (fail closed).
Both are legitimate choices. Picking the wrong one, or not choosing at all and letting an unhandled exception decide, is how a security feature turns into an outage or a hole.
The two failure modes
Fail open: on error, allow the request.
- ✅ The site stays up when the security layer breaks.
- ❌ While it’s broken, the protection is off.
Fail closed: on error, deny the request.
- ✅ Nothing gets through unchecked.
- ❌ A bug or outage in the security layer takes down everything behind it.
And a third, accidental mode: fail undefined. No error handling at all, so an exception crashes the handler and the platform returns a generic 500. In practice that’s fail-closed with a worse error message and no logging. If you haven’t chosen, this is probably what you have.
How to choose
Ask two questions:
1. What happens if an attacker gets through during the outage?
- If one bad request causes irreversible harm (reading someone else’s data, moving money, changing permissions), fail closed.
- If the harm is bounded and recoverable (a scraper copies some pages, some junk traffic reaches the origin for a few minutes), failing open is usually acceptable.
2. What happens to legitimate users if you fail closed?
- If the check is on a narrow path (an admin panel, a payment endpoint), failing closed inconveniences few people.
- If it’s in front of all traffic, failing closed turns any bug in the check into a full outage.
That gives a simple matrix:
| Harm if bypassed: low / recoverable | Harm if bypassed: high / irreversible | |
|---|---|---|
| Scope: all traffic | Fail open | Fail closed, and invest heavily in the check’s reliability |
| Scope: narrow path | Either; fail open is simpler | Fail closed |
Some examples:
- Authentication and authorization: fail closed. Always.
- Payment fraud checks: usually fail closed, or queue for manual review.
- Rate limiting and other abuse throttles in front of a whole site: usually fail open.
- Input validation: fail closed. Reject what you can’t validate.
Example: a rate limiter in front of everything
Picture a rate limiter that checks a shared counter store on every request to your site. Its job is to smooth out abuse and spikes, not to be the only thing protecting sensitive data. Your application still has authentication, authorization and its own validation.
If the limiter fails closed and its counter store becomes unreachable, every visitor gets an error. You’ve built a single point of failure whose failure is worse than the problem it solves. If it fails open, traffic goes unthrottled for the duration of the incident, and your origin handles the load it was handling before you had a limiter.
So the check defaults to “allow” when it can’t decide:
let allowed = true;
try {
allowed = await limiter.check(clientKey);
} catch (err) {
console.error('rate limiter unavailable, allowing request', err);
// allowed stays true
}
Making fail-open safe
Failing open is only responsible if you know when it’s happening. A silent fail-open is a protection that can be off for weeks without anyone noticing. Put these in place:
- Count it. Increment a metric or log a distinct marker (e.g.
reason: "check_error") every time the check fails open. - Alert on the rate, not individual events. A few transient errors are noise. A sustained rate means the protection is effectively off.
- Set timeouts. A dependency that hangs is worse than one that errors, because it holds the request. Bound every external call and treat a timeout as an error.
- Degrade in layers. If the primary lookup fails, a smaller in-memory list of the worst offenders can still apply. Partial protection beats none.
Making fail-closed safe
If a check must fail closed, its reliability is your site’s reliability. Treat it that way:
- Keep dependencies minimal. Every network call in the check is a way for it to fail.
- Cache last-known-good state. If config or policy can’t be refreshed, keep enforcing the previous version instead of erroring.
- Return a clear error. A specific status and message (“verification temporarily unavailable”) helps users and your support team far more than a generic 500.
- Test the failure path. Deliberately break the dependency in staging and watch what happens.
Circuit breakers
For checks that call remote services, a circuit breaker helps in both modes. After N consecutive failures, stop calling the dependency for a cool-down period and apply the failure policy immediately. This protects a struggling dependency from load, and it protects your latency from waiting on timeouts again and again.
let failures = 0;
let openUntil = 0;
async function guarded(fn, onOpen) {
if (Date.now() < openUntil) return onOpen();
try {
const result = await fn();
failures = 0;
return result;
} catch (err) {
if (++failures >= 5) openUntil = Date.now() + 30_000;
return onOpen();
}
}
Key takeaways
- Every security check needs an explicit failure policy. “Unhandled exception” is a policy, just a bad one.
- Fail closed where one bypass causes irreversible harm (auth, payments). Fail open where harm is bounded and the check covers all traffic (bot filtering).
- Fail-open without monitoring is protection that can silently vanish. Count and alert on it.
- Fail-closed makes the check’s reliability your site’s reliability. Minimize its dependencies and cache last-known-good state.