Pacing Bulk Async Jobs Under an LLM Rate Limit

An admin clicks a "reprocess everything" button. The handler does the obvious thing:

ids.forEach(this::processAsync);   // fan the whole batch onto the async pool

Hundreds of tasks land on a small, shared async pool at once. Most are gated by an upstream step, but the ones that pass finish in clusters — and each one fires a call to an external LLM API. Those calls go out together, and together they blow past the provider's rate limit. 429s everywhere, and the whole batch crawls as half the calls sit in backoff.

TL;DR: For most LLM APIs the binding limit is tokens-per-minute, not requests-per-minute — check your provider's docs and your own key's quota. Retry/backoff can absorb the occasional 429, but if your workload naturally bursts you'll spend the whole job fighting rate limits. The fix is to pace the dispatch — space each task a fixed interval apart — so the burst never forms. Retry becomes the safety net, not the strategy.


The Problem

forEach(async) is a burst generator. It hands the entire batch to the executor in a tight loop; the executor runs up to its pool size concurrently and queues the rest, which then run the instant a thread frees up. From the API's perspective that's a spike: N calls arriving inside a few seconds, N × (prompt + completion) tokens per minute, straight over the ceiling.

You can crank up retry/backoff, but that treats the symptom. You're still generating the burst; you're just apologizing for it afterward, slowly, one 429 at a time. And backoff has a nasty interaction with a shared pool — retrying tasks hold threads while they wait, starving unrelated real-time work that needs the same pool.

The Fix: Space the Dispatch

Put a small, fixed gap between each task's dispatch. A single-threaded scheduler is all it takes — the actual work is already async, so the scheduler thread only triggers the hand-off and returns immediately:

@Service
public class PacedDispatcher {

    static final long PACING_MS = 1500;

    private final ScheduledExecutorService scheduler =
        Executors.newSingleThreadScheduledExecutor(r -> {
            Thread t = new Thread(r, "paced-dispatch");
            t.setDaemon(true);
            return t;
        });

    public void dispatchPaced(List<Long> ids, LongConsumer work) {
        for (int i = 0; i < ids.size(); i++) {
            long id = ids.get(i);
            scheduler.schedule(() -> {
                try {
                    work.accept(id);
                } catch (Exception e) {
                    log.warn("Paced work failed for id {}: {}", id, e.getMessage());
                }
            }, (long) i * PACING_MS, TimeUnit.MILLISECONDS);
        }
    }
}

Item i fires at i × PACING_MS. Choose the interval from your token budget: estimate the average tokens per call, divide the per-minute quota by that, and leave headroom for whatever else is calling the API concurrently. The scheduler thread does almost no work — it just kicks off the already-async task at the right moment — so it never starves the shared pool.

The Deliberate Trade-offs

  • Fire-and-forget. dispatchPaced returns immediately and the caller reports the count as "dispatched." The admin gets instant feedback; the work drains in the background.
  • Drop on shutdown. If the app restarts, not-yet-fired items are simply dropped. For a best-effort reprocess that's fine — someone can re-trigger it. If your job must survive a restart, back it with a durable queue instead; don't bolt persistence onto a scheduler.
  • Pacing complements retry, doesn't replace it. Backoff still catches the odd 429 from concurrent traffic. Pacing just makes sure your own bulk action isn't the thing causing them.

Why Pacing Beats Shrinking the Pool

You could throttle by shrinking the async pool, but that pool is usually shared — starving it slows unrelated real-time requests to serve a background job. Pacing throttles only the bulk work, at the dispatch layer, without touching the pool's concurrency for everyone else. Right knob, right place. It also keeps the throttle explicit and local: the pacing constant lives next to the job it governs, so the next person reading the code sees exactly why calls are spaced, instead of inferring it from a pool size configured three layers away.

Know Which Limit Binds

The reason "just add more retries" fails is that people optimize against the wrong ceiling. If you assume the limit is requests-per-minute but it's actually tokens-per-minute, you'll size your concurrency by request count and still overshoot on tokens — a handful of long-prompt calls can exceed the budget that a hundred short ones wouldn't. Measure your real token-per-call distribution, find the binding limit, and pace against that.

Lessons Learned

  • Know which limit binds. For LLM APIs it's usually tokens-per-minute, not requests. Optimize against the real ceiling, not the one you assumed.
  • forEach(async) is a burst. Fanning a batch onto an executor in a loop produces a spike. If a downstream has a rate limit, space the dispatch.
  • Prevent the burst; keep retry as the net. Backoff should handle noise, not be your rate-limit strategy. Pace the source and the 429s mostly disappear.

How do you pace bulk work against a rate-limited API? Share your approach.

Building jo4.io - a URL shortener with analytics for developers.