Rate-Limit Failed Logins Without Locking Out Victims

Any endpoint that checks a secret — a login, a password-protected resource, a one-time code — invites guessing. So you add a lockout: too many failures and further attempts get rejected. Reasonable. But the obvious version of this defense has a hole that turns it into a weapon, and the fix is a single design decision made before you write a line of code.

TL;DR: Never key a lockout on the protected resource alone. "5 failures and this account is locked" lets any attacker lock out every legitimate user by deliberately failing on their behalf — a denial of service you built and shipped yourself. Key the counter on (resource, client) instead. Then the attacker only locks themselves out, and honest users are untouched. Below is the pattern, done atomically in Redis, with a fallback so an outage doesn't disable the control.


The Problem

Here's the version almost everyone writes first:

// DON'T: keyed on the resource only
if (failures(resourceId) >= 5) reject();   // attacker fails 5x → locked for EVERYONE

Count failures per resource, lock the resource. It stops brute force — and it hands every griefer a button that says "make this account unusable." Fail five times against a victim's login and they can't get in. Fail against a shared or high-value resource and you've knocked it offline for its whole audience. You've converted a security control into an availability attack. This is a well-known anti-pattern (OWASP calls it out explicitly), and it's easy to ship because it "works" in every happy-path test.

The Fix: Key on (Resource, Client)

Scope the counter to the pair of the thing being protected and the party attempting it — typically the client IP:

private String key(String resourceId, String clientIp) {
    return "lockout:" + resourceId + ":" + (clientIp != null ? clientIp : "unknown");
}

public boolean isLockedOut(String resourceId, String clientIp) {
    String val = redis.opsForValue().get(key(resourceId, clientIp));
    int count = val != null ? Integer.parseInt(val) : 0;
    return count >= MAX_ATTEMPTS;
}

Now a single attacker IP gets locked out of that one resource. Everyone else, coming from other addresses, sails through. Combine this with a plain per-IP request throttle and you bound the distributed case too: every source is both rate-limited and attempt-capped, so spreading the attack across many IPs runs into the throttle instead.

Pick the "client" dimension for your context. Behind a CDN or proxy, make sure you read the real client address (the forwarded client IP your infra sets), not the edge's IP — otherwise every request looks like it comes from one place and you're back to a shared counter.

Increment and Expire Atomically

The counter needs to set its expiry exactly once — on the first failure — so the window is a true sliding window that starts when the attack starts. Do it in one atomic round-trip with a tiny Lua script, so there's no gap between "create the key" and "give it a TTL":

String lua =
    "local current = redis.call('INCR', KEYS[1]) " +
    "if current == 1 then " +
    "    redis.call('EXPIRE', KEYS[1], ARGV[1]) " +
    "end " +
    "return current";
redis.execute(new DefaultRedisScript<>(lua, Long.class),
    Collections.singletonList(key), String.valueOf(windowSeconds));

INCR creates the key at 1 and only then sets the expiry. No separate SET/EXPIRE race, no orphaned keys stuck without a TTL. A successful attempt clears the counter — guess right and your slate is clean.

Don't Fail Open on an Outage

If your datastore is the only thing between an attacker and unlimited guesses, a blip shouldn't silently disable the lockout. Keep a per-instance in-memory fallback that mirrors the same logic, and evict expired entries opportunistically so the fallback map can't grow without bound during a sustained outage:

memoryStore.compute(key, (k, existing) -> {
    if (existing == null || existing.expiresAtMillis() < now) {
        return new MemoryEntry(new AtomicInteger(1), now + windowMillis);
    }
    existing.count().incrementAndGet();
    return existing;
});

It's weaker than the shared counter — per-instance rather than cluster-wide — but it's a floor, not a hole. The security posture degrades gracefully instead of disappearing.

Choosing the Numbers

Two knobs — the attempt cap and the window length — are a UX-vs-security trade you should tune to your threat model and secret strength, not copy from a blog. A short window with a generous cap is friendlier to fat-fingered users; a long window with a tight cap is harsher on attackers and on the occasional confused human. Whatever you choose, make them configurable so you can adjust under real traffic without a redeploy, and log lockouts so you can see whether the threshold is catching attacks or annoying users.

Lessons Learned

  • The key is the security boundary. (resource, attacker) is almost always right; (resource) alone turns your lockout into a griefing tool.
  • Set the TTL atomically with the first increment. A Lua INCR + conditional EXPIRE avoids the classic no-TTL orphan and the race around it.
  • Fall back, don't fall open. A control that vanishes the moment your datastore blinks isn't a control. Degrade to a local floor instead.
  • Make the thresholds config, not constants. You'll want to tune them against real traffic, and you don't want a deploy to do it.

Ever shipped a "lockout" that became a DoS vector? Tell me about it below.

Building jo4.io - a URL shortener with analytics for developers who care about the details.