All articles
AI Infrastructure 6 min readJun 2026

Keeping LLM costs sane at 100K queries a day

How queue-first design, caching, and per-provider throttling turned a runaway API bill into viable unit economics.

On a platform that tracks brand visibility across ChatGPT, Claude, Gemini, and Perplexity, the workload is brutal by design: the same questions, asked across four providers, tens of thousands of times a day. Done naively, the API bill grows faster than the product. Here's how we kept 100K+ daily queries economically viable.

Put a queue between intent and API call

The most important architectural decision wasn't a cost trick — it was making every LLM call go through a job queue instead of a synchronous request. A queue gives you back-pressure, retries, rate-limit handling, and a single place to enforce budget. Without it, a traffic spike becomes a billing spike and a pile of half-finished work.

ts
// Every provider call is a job, not an inline await
await queue.add(
  "llm-query",
  { provider, prompt, cacheKey },
  {
    attempts: 4,
    backoff: { type: "exponential", delay: 2000 },
    removeOnComplete: 1000,
  }
);

Cache aggressively — the same question is asked 10,000 times

When your workload has repetition, a cache is the cheapest optimization there is. Hash the normalized (provider, prompt, model) tuple and store the result with a freshness-appropriate TTL. For visibility tracking, answers only need to be a few hours fresh — so most of the day's queries never hit a provider at all.

  • Normalize prompts before hashing so trivial differences still hit the cache.
  • Set TTL by how fast the answer actually changes, not by habit.
  • Track cache hit-rate as a first-class cost metric.

Throttle per provider, not globally

Each provider has different rate limits, latencies, and prices. A single global limiter either underuses the fast providers or trips the strict ones. Give each provider its own token-bucket limiter tuned to its quota, and the queue naturally spreads load without 429 storms.

Queue-first design and aggressive caching were the difference between viable unit economics and a runaway bill.

Route to the cheapest model that clears the bar

Not every task needs a frontier model. Classification, extraction, and short structured outputs often run fine on a smaller, cheaper model. Reserve the expensive models for genuinely hard synthesis. A simple router that picks the model by task type quietly removes a large slice of spend.

Make cost observable

You can't control what you can't see. We logged tokens and estimated cost per job to Datadog and alerted on cost-per-thousand-queries, not just error rates. The moment a change made things more expensive, we knew — before the invoice did.

The takeaway

LLM cost control isn't one clever trick; it's an architecture. Queue the work, cache the repeats, throttle per provider, route by difficulty, and measure everything. Do that and per-query cost becomes a number you steer — not a surprise you absorb.

Building something with AI?

I'm Muhammad Usman Haider, a full stack ai engineer. If you want a system like the ones in this article, let's talk.

Start a project