# Tool retries and partial failures: what the AI SDK retries and what it does not

> The SDK retries the request to the model, never your tool's side effects. A tool that half-succeeded and then got called again is where duplicate charges come from.

You wire up a tool, ship it, and a week later a customer has two refunds for
one order. The logs show one conversation, one user message, and two calls to
`issueRefund` with identical arguments about four seconds apart.

Nobody wrote a retry loop. The retry came from somewhere else, and the tool was
not built to survive being called twice.

## What actually retries

Three different mechanisms can produce a second call, and they have nothing to
do with each other.

**1. The AI SDK's `maxRetries`.** This retries the HTTP request to the model
provider: a 429, a 500, a dropped connection. It defaults to 2, so a call can
be attempted three times. Crucially, it retries *the request*, and the request
is the whole conversation so far. If the previous step already ran your tool,
the tool is not re-run by this retry; the SDK is only re-asking the model.

**2. The model itself.** With `stopWhen: isStepCount(n)`, the model sees the
tool result and decides what to do next. A tool that returns something the
model reads as failure (`{ error: "timeout" }`, an empty object, a thrown
error surfaced as text) is an invitation to try again. The model does not know
your tool has side effects. It knows it did not get an answer.

**3. Your own infrastructure.** A function that times out and is retried by the
platform, a user who hits send twice, a client that resubmits.

The one that surprises people is (2). It looks like a retry loop nobody wrote,
because the loop is a conversation.

## The wrong way

```ts
export const issueRefund = tool({
  description: "Refund an order.",
  inputSchema: z.object({ orderRef: z.string(), amount: z.number() }),
  execute: async ({ orderRef, amount }) => {
    const charge = await payments.refund({ orderRef, amount });
    await db.order.update({ where: { ref: orderRef }, data: { refundedAt: new Date() } });
    await email.send({ template: "refund", orderRef });
    return { ok: true, refundId: charge.id };
  },
});
```

Three failure modes are baked in:

- If `email.send` throws, the refund has already happened but the tool reports
  failure. The model apologises and tries again. The customer gets two refunds
  and one email.
- If the whole thing succeeds but the response stream drops before the model
  reads the result, the next attempt starts from a conversation where the tool
  has no result, and calls it again.
- If the model is asked for "a refund for the last two orders" it may emit two
  parallel tool calls with the same `orderRef` because it misread the history.

## The right way

**Make the tool idempotent on a key the caller supplies.** The AI SDK gives you
`toolCallId` in the execute options. It is stable for one tool call, including
across the SDK's own request retries.

```ts
export const issueRefund = tool({
  description:
    "Issue a refund for one order. Only call this after the customer has confirmed the amount.",
  inputSchema: z.object({
    orderRef: z.string().regex(/^ORD-[0-9]{6}$/),
    amountCents: z.number().int().positive().max(50_000),
  }),
  // AI SDK 7: the user id arrives through `toolsContext`, checked against this
  // schema, never as an argument the model fills in.
  contextSchema: z.object({ userId: z.string() }),
  execute: async ({ orderRef, amountCents }, { toolCallId, context: { userId } }) => {
    const order = await db.order.findFirst({ where: { ref: orderRef, ownerId: userId } });
    if (!order) return { ok: false, reason: "no such order on this account" };
    if (order.refundedAt) return { ok: true, alreadyRefunded: true, refundId: order.refundId };

    // The payment provider dedupes on this key: a second call with the same
    // key returns the first result instead of moving money again.
    const charge = await payments.refund({ orderRef, amountCents, idempotencyKey: toolCallId });

    await db.order.update({ where: { id: order.id }, data: { refundedAt: new Date(), refundId: charge.id } });

    // Side effects that are not part of the money movement happen after, and
    // their failure does not fail the tool.
    void email.send({ template: "refund", orderRef }).catch((error) => {
      console.error("[refund] receipt email failed", { orderRef, message: String(error) });
    });

    return { ok: true, refundId: charge.id };
  },
});
```

Four things changed, and each maps to one of the failure modes:

- **An idempotency key.** Every payment API supports one. So does any `INSERT`
  with a unique constraint on that key. Without it, "did this already happen?"
  is unanswerable.
- **A pre-check that returns success.** `alreadyRefunded: true` tells the model
  the goal is achieved. Returning an error here is what causes the second
  attempt.
- **Ordering by reversibility.** The irreversible step happens once and first;
  the recoverable steps (email, analytics, cache invalidation) happen after and
  cannot fail the tool.
- **Failures are data.** `{ ok: false, reason }` ends the loop cleanly. A thrown
  error, by contrast, propagates as a tool error the model reads as "try
  differently".

## Deciding what a failure should look like

| Situation | Return | Why |
|---|---|---|
| Not found, no match, nothing to do | `{ ok: false, reason }` | The model tells the user; no retry |
| Already done | `{ ok: true, already: true }` | Ends the loop, no second side effect |
| Bad input that passed the schema | `{ ok: false, reason }` | The model can ask the user for a correction |
| Upstream 5xx, likely transient | throw | The SDK's retry is the right handler |
| Upstream 4xx, will never succeed | `{ ok: false, reason }` | Retrying wastes a step and a billed request |

The rule of thumb: **throw only when a retry could plausibly help.** Everything
else is a result.

## Bound the loop

```ts
streamText({
  model: languageModel(),
  tools: { issueRefund },
  toolsContext: { issueRefund: { userId: session.user.id } },
  stopWhen: isStepCount(6),
  maxRetries: 2,
  messages,
});
```

`isStepCount` (called `stepCountIs` before AI SDK 7) is the wall at the end of the conversation loop; `maxRetries` is
the transport budget for each individual request. Both are needed and they do
different jobs. Six steps is a reasonable default for search-and-answer; raise
it per surface for genuinely multi-step work, never globally.

## Partial failure in a multi-tool step

A model can emit several tool calls in one step, and they run concurrently. If
three succeed and one throws, the successful three have already happened, and
there is no rollback: you are not in a transaction.

The practical mitigations, in order of how much they help:

1. Make each tool individually idempotent (above). Then a re-run is harmless.
2. Never expose two tools that must succeed together. If two things must be
   atomic, that is one tool wrapping one transaction.
3. Return enough in each result for the model to describe what did happen. A
   user told "I refunded the order but could not cancel the subscription" is in
   a much better position than one told "something went wrong".

## Checking your work

- Call the tool twice with the same `toolCallId` in a test and assert the side
  effect happened once.
- Force the second, non-critical side effect to throw and assert the tool still
  reports success.
- Add an eval case that asks for the same action twice in one conversation and
  assert the second answer says it was already done.

---

Agentic Boilerplate: A Next.js repo your agent already knows. Free during launch, then $99 once.

- Site map for agents: https://agenticboilerplate.com/llms.txt
- Public API: https://agenticboilerplate.com/openapi.json
- Contact: agenticstudio@gmail.com
