# When to add usage metering to an AI feature, and how to do it in an afternoon

> A public AI route is a public spend endpoint. The three signals that say you need metering now, and the smallest implementation that actually protects you.

The story is always the same. Someone ships an AI feature behind a form with no
account required, posts it somewhere, and wakes up to a provider bill that is
four figures larger than expected. The traffic was not organic. Someone found
an endpoint that runs a language model, pointed a script at it, and your API
key paid for their weekend.

There is no per-request cap on a model API. There is no "you have spent enough
today" that fires by default. If a route can be called, it can be called ten
thousand times an hour, and every one of those is billed to you.

## Three signals that say "now"

Do not build metering into your prototype. Do build it before any of these is
true:

1. **The route is reachable without authentication.** This includes routes
   behind a session that anyone can create by signing up with a throwaway
   email. Rule of thumb: if getting a thousand requests costs an attacker
   nothing, the route is public.
2. **Any single request can cost more than a fraction of a cent.** Long
   context, tool loops, or a large `maxOutputTokens` all multiply. A tool-using
   turn on a strong model with a big document in context is meaningfully
   expensive; a hundred of them is a bill you will notice.
3. **You are about to price the product.** The moment a plan says "500 messages
   a month", you need a counter that is the same one enforcement reads.
   Retrofitting the counter after selling the plan means reconciling two
   numbers that never agreed.

If none apply (an internal tool, five colleagues, a key with a $20 monthly cap
on it), skip this. The provider dashboard's spend limit is enough.

## Do the provider-side limit first

Before you write any code: set a monthly spend limit in the Anthropic or OpenAI
console, and use a separate key per environment. This is a hard ceiling that
works even when your application logic is wrong, and it takes two minutes. It
is the only protection that cannot be bypassed by a bug in your own counting.

It is a ceiling, not a control: hitting it takes the feature down for everyone.
Treat it as the last line, not the plan.

## The wrong way

```ts
// Counts after the fact, in memory, per instance.
let requestsToday = 0;

export async function POST(request: Request) {
  requestsToday += 1;
  if (requestsToday > 1000) return new Response("Too many", { status: 429 });
  // ...
}
```

Three problems. It is per-instance, so ten serverless instances give an
attacker ten times the budget. It resets on every deploy and every cold start.
And it counts *requests*, which is not the thing that costs money: one request
with a 200k-token context costs more than a hundred short ones.

## The right way, in three layers

### Layer 1: a per-identity rate limit, before the model call

The cheapest effective control. It runs before you spend anything.

```ts
import { after } from "next/server";

export async function POST(request: Request) {
  const identity = await identify(request);   // user id, or hashed IP for anonymous
  const gate = await rateLimit.check(identity, { limit: 20, windowSeconds: 3600 });

  if (!gate.allowed) {
    return Response.json(
      { error: "Rate limit reached. Try again in a few minutes." },
      { status: 429, headers: { "retry-after": String(gate.retryAfterSeconds) } },
    );
  }
  // ...
}
```

Use a shared store: Redis, or your Postgres with a small `rate_limit` table
and an atomic upsert. Anything per-instance is decoration. Set the limit an
order of magnitude above what a real user does, not at it: the goal is to stop
scripts, not to annoy people.

For anonymous traffic, hash the IP with a server-side secret rather than
storing it raw. That is a per-identity counter and not a personal data store.

### Layer 2: record what each call actually cost

Every AI SDK result carries usage. Record it after the response is sent, so
accounting never adds latency to the user's request.

```ts
const result = streamText({
  model: languageModel(),
  messages,
  abortSignal: request.signal,
  // AI SDK 7: `onEnd` (it was `onFinish`), and `usage` covers every step.
  onEnd({ usage, finishReason, finalStep }) {
    after(
      recordUsage({
        userId: identity,
        feature: "chat",
        model: finalStep.response.modelId,
        inputTokens: usage.inputTokens ?? 0,
        outputTokens: usage.outputTokens ?? 0,
        finishReason,
      }),
    );
  },
  // An aborted stream skips `onEnd`. You were still billed for what ran.
  onAbort({ steps }) {
    after(
      recordUsage({
        userId: identity,
        feature: "chat",
        model: steps.at(-1)?.response.modelId ?? "unknown",
        inputTokens: steps.reduce((sum, step) => sum + (step.usage.inputTokens ?? 0), 0),
        outputTokens: steps.reduce((sum, step) => sum + (step.usage.outputTokens ?? 0), 0),
        finishReason: "aborted",
      }),
    );
  },
});
```

One row per call, with the model id the provider reports on it. Cost is
derived at read time from a price table you keep in code. Do not store a
computed cost: prices change and you will want to re-price history.

Record the aborted turn too. A step cut off mid-stream may not report its
tokens at all, so a Stop click undercounts slightly; the provider dashboard
is the source of truth for the invoice, your table is the source of truth
for who spent it.

### Layer 3: a quota check that reads the same numbers

Once layer 2 has data, the quota is a query:

```ts
const used = await tokensThisPeriod(userId);
const allowed = planLimits[plan].tokensPerMonth;
if (used >= allowed) {
  return Response.json({ error: "Monthly limit reached.", upgradeUrl: "/pricing" }, { status: 402 });
}
```

Two details that decide whether people trust the number. Show it before they
hit it: a counter in the UI at 80% turns a hard stop into an expected one. And
decide deliberately whether a failed generation consumes quota; the honest
answer is that you were billed, so it does, but say so in the plan description.

## What to measure

Per feature, not just in total. "AI spend is up 40%" is not actionable; "the
document-summary route is up 40% and its average input tokens tripled" is:
somebody raised a truncation limit.

Keep these four:

- tokens in and out, per feature per day
- calls per user, so you can find the outlier
- `finishReason` distribution: a rise in `"length"` means answers are being
  cut off and users are re-asking, paying twice
- p95 tokens per call, which catches context growing quietly

## The cheapest wins, before you optimise anything

- **Cap `maxOutputTokens` per feature.** Left unset, the SDK asks for the
  model's maximum (128,000 tokens on current Claude models). Leave room for
  reasoning, which counts against the cap, but a classification still does not
  need 4,000.
- **Use a smaller model where the task is easy.** Routing, classification and
  short extraction run fine on the cheap tier in `MODEL_CATALOGUE` (Claude
  Haiku 4.5, GPT-6 Luna) at a fraction of the default model's cost.
- **Truncate context.** Most "expensive" features are expensive because they
  send an entire document when the relevant paragraph would do.
- **Prompt caching** for a long, stable system prompt reused across requests:
  a provider feature that cuts repeat-context cost sharply.
- **Batch anything not interactive.** Both major providers offer a cheaper
  asynchronous tier for work nobody is waiting on.

---

Agentic Boilerplate: A Next.js repo your agent already knows. Free during launch, then $99 once.

- Site map for agents: https://agenticboilerplate.com/llms.txt
- Public API: https://agenticboilerplate.com/openapi.json
- Contact: agenticstudio@gmail.com
