Keeping LLM costs sane at 100K queries a day
How queue-first design, caching, and per-provider throttling turned a runaway API bill into viable unit economics.
On a platform that tracks brand visibility across ChatGPT, Claude, Gemini, and Perplexity, the workload is brutal by design: the same questions, asked across four providers, tens of thousands of times a day. Done naively, the API bill grows faster than the product. Here's how we kept 100K+ daily queries economically viable.
Put a queue between intent and API call
The most important architectural decision wasn't a cost trick — it was making every LLM call go through a job queue instead of a synchronous request. A queue gives you back-pressure, retries, rate-limit handling, and a single place to enforce budget. Without it, a traffic spike becomes a billing spike and a pile of half-finished work.
// Every provider call is a job, not an inline await
await queue.add(
"llm-query",
{ provider, prompt, cacheKey },
{
attempts: 4,
backoff: { type: "exponential", delay: 2000 },
removeOnComplete: 1000,
}
);Cache aggressively — the same question is asked 10,000 times
When your workload has repetition, a cache is the cheapest optimization there is. Hash the normalized (provider, prompt, model) tuple and store the result with a freshness-appropriate TTL. For visibility tracking, answers only need to be a few hours fresh — so most of the day's queries never hit a provider at all.
- Normalize prompts before hashing so trivial differences still hit the cache.
- Set TTL by how fast the answer actually changes, not by habit.
- Track cache hit-rate as a first-class cost metric.
Throttle per provider, not globally
Each provider has different rate limits, latencies, and prices. A single global limiter either underuses the fast providers or trips the strict ones. Give each provider its own token-bucket limiter tuned to its quota, and the queue naturally spreads load without 429 storms.
Queue-first design and aggressive caching were the difference between viable unit economics and a runaway bill.
Route to the cheapest model that clears the bar
Not every task needs a frontier model. Classification, extraction, and short structured outputs often run fine on a smaller, cheaper model. Reserve the expensive models for genuinely hard synthesis. A simple router that picks the model by task type quietly removes a large slice of spend.
Make cost observable
You can't control what you can't see. We logged tokens and estimated cost per job to Datadog and alerted on cost-per-thousand-queries, not just error rates. The moment a change made things more expensive, we knew — before the invoice did.
The takeaway
LLM cost control isn't one clever trick; it's an architecture. Queue the work, cache the repeats, throttle per provider, route by difficulty, and measure everything. Do that and per-query cost becomes a number you steer — not a surprise you absorb.
Building something with AI?
I'm Muhammad Usman Haider, a full stack ai engineer. If you want a system like the ones in this article, let's talk.
Start a project