# When AI work has to move to a background job, and what breaks if you wait

> Serverless functions have a hard timeout that no amount of streaming avoids. The four signals that mean the work no longer belongs in the request, and the smallest queue that fixes it.

Everything in a generated AI bundle runs inside the HTTP request. That is the
right default: it is simple, it streams, and there is nothing to operate. It
also has a hard ceiling, and the way you find the ceiling is a bug report that
says "it works for small files".

## The wall

A serverless function is killed at `maxDuration`. On Vercel with Fluid compute
(the default for new projects) that is 300 seconds by default, 300 at most on
Hobby and 800 on Pro. The number moves; there is always a number, and when you
hit it the function is terminated mid-work. The user sees a truncated stream or a network error. Nothing is
retried. Nothing is saved. The tokens you generated are billed anyway.

Streaming does not help. Streaming changes when the *first* byte arrives, not
how long the function runs. A twelve-minute batch job that streams progress is
still a twelve-minute function.

## Four signals it is time

1. **The work can exceed the timeout for realistic input.** Not average input:
   the 95th percentile. If summarising a 3-page PDF takes 8 seconds, a 200-page
   one takes eight minutes and will be uploaded on day one.
2. **Nobody is watching.** Nightly enrichment, re-embedding a corpus, scoring
   every new signup. If no human is waiting for the tokens, holding a request
   open is pure cost.
3. **It has to survive failure.** A provider 529, a deploy mid-run, an instance
   recycled. Inside a request, all three lose the work silently. A job can be
   retried from where it stopped.
4. **It fans out.** "Process these 400 rows" is 400 model calls. In a request
   they run in one function against one rate limit; as jobs they run at a
   concurrency you control, with per-item retries.

If none of these hold (a chat turn, a single classification, an extraction
from one pasted email), keep it in the request. A queue you did not need is
infrastructure you now maintain.

## The wrong way

```ts
export async function POST(request: Request) {
  const { documentIds } = await request.json();

  const summaries = [];
  for (const id of documentIds) {
    const doc = await loadDocument(id);
    summaries.push(await summarise(doc));   // 6s each, 40 documents
  }

  return Response.json({ summaries });
}
```

This dies at document seven. The first six were generated and billed, and
nothing recorded them. The user retries, and you pay for those six again. Add
`Promise.all` and it dies faster, now with provider 429s in the log.

Two variants of the same mistake are worth naming:

- **Fire-and-forget without a runtime that waits.** Calling an async function
  and returning immediately means the platform freezes the instance as soon as
  the response is sent. The work stops mid-flight. Next.js has `after()` for
  short post-response work; it is bounded by the same function lifetime and is
  not a job queue.
- **A cron route that does everything.** A single `/api/cron/process-all` that
  loops over the backlog hits the same wall, just at 3am where nobody sees it.

## The right way

### Step 1: make the unit of work one item

The first change is not the queue, it is the shape. One job = one document, one
row, one model call, with a status you can read:

```prisma
model SummaryJob {
  id         String   @id @default(cuid())
  documentId String
  status     JobStatus @default(QUEUED)   // QUEUED | RUNNING | DONE | FAILED
  attempts   Int      @default(0)
  result     String?
  error      String?
  createdAt  DateTime @default(now())
  updatedAt  DateTime @updatedAt

  @@unique([documentId])
  @@index([status, createdAt])
}
```

The unique constraint on `documentId` is the idempotency key: enqueueing twice
does not process twice. `attempts` is what stops a poisoned input retrying
forever.

### Step 2: accept fast, process later

```ts
export async function POST(request: Request) {
  const { documentIds } = await request.json();

  await db.summaryJob.createMany({
    data: documentIds.map((documentId) => ({ documentId })),
    skipDuplicates: true,
  });

  return Response.json({ queued: documentIds.length }, { status: 202 });
}
```

202 with a job id, not 200 with a result. The client polls a status route, or
you push over websockets. This single change removes the timeout from the
user-facing path entirely.

### Step 3: pick the smallest runner that fits

In rough order of how much you take on:

- **A hosted job runner** (Inngest, Trigger.dev, QStash). Fastest path on
  serverless: they call your endpoint, handle retries with backoff, give you a
  dashboard and step-level durability. You write a function, not a worker.
- **Your platform's cron** hitting a route that claims and processes a small
  batch (say ten items) and returns. Runs every minute. Crude, free, and
  perfectly adequate for a backlog that is not latency-sensitive. Claim rows
  with an atomic update so two overlapping runs cannot take the same item.
- **A real worker process** on a container host, pulling from a queue. Right
  when volume is steady and high; wrong as a first step, because it is a second
  deployable to operate.

### Step 4: make each job safe to run twice

Every job runner retries. Assume at-least-once delivery:

- Claim the row atomically (`update ... where status = 'QUEUED'` returning the
  row) so two runners cannot both take it.
- Write results with the job id as the key, so a duplicate write is a no-op.
- Cap `attempts`. After three, mark `FAILED` with the error and stop. A
  malformed PDF should not be retried a thousand times at model prices.
- Record token usage per job. This is where cost visibility comes from once the
  work is no longer in a request you can watch.

## The middle ground

Not everything needs a queue. Two patterns sit between:

- **`after()` for short trailing work.** Logging usage, writing an audit row,
  invalidating a cache. Runs after the response, still inside the function's
  lifetime. Do not put a model call in it that could exceed the remaining time.
- **Chunking in the client.** For a list of twenty documents, have the browser
  send twenty requests with a small concurrency limit. Each is short, progress
  is natural, and failures are per-item. It is not durable (a closed tab stops
  it), but it turns a timeout into a progress bar in an afternoon.

## How to know it worked

- Kill a deploy mid-batch. The queued items should still be processed
  afterwards.
- Enqueue the same document twice; assert one result row and one model call.
- Feed it something that always fails and confirm it stops after `attempts`
  reaches the cap, with the error recorded and visible.

---

Agentic Boilerplate: A Next.js repo your agent already knows. Free during launch, then $99 once.

- Site map for agents: https://agenticboilerplate.com/llms.txt
- Public API: https://agenticboilerplate.com/openapi.json
- Contact: agenticstudio@gmail.com
