> ## Documentation Index
> Fetch the complete documentation index at: https://docs.semicola.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Limits

> Design v3 clients around bounded tool calls, connection lifetime, payload size and seller latency.

v3 shares the platform's transport protections and adds a few bounds of its own. Build for them
from the start; they are what keep one slow seller or one large report from stalling everyone else.

## MCP bounds

| Bound                       | Limit                                                                                                          |
| --------------------------- | -------------------------------------------------------------------------------------------------------------- |
| One MCP connection          | Up to 30 minutes. Deploys may end it sooner.                                                                   |
| MCP session                 | Dropped after 60 minutes idle, or when the server restarts; re-`initialize` on `404 Session not found`.        |
| One tool call               | Up to 30 minutes, then `SERVICE_UNAVAILABLE`: the call may still have happened, so read state before retrying. |
| Structured response size    | Responses over 200 KB are truncated as a safety net. Page with `limit` and `cursor`.                           |
| Proposal executions         | One running `request_proposals` execution per buyer account.                                                   |
| Each seller in an execution | Bounded to 30 seconds. A seller that does not answer in time is reported as failed or pending.                 |
| Delivery ranges             | `get_delivery` accepts up to 90 days, or `{ "lifetime": true }`.                                               |
| Idempotency keys            | 16 to 255 characters from `A–Z a–z 0–9 _ . : -`. Replays are honored for 24 hours.                             |
| Confirmations               | A `pending_confirmation` expires after 15 minutes.                                                             |

## Rate limits

Semicola counts requests in fixed windows:

| Scope                                                              | Window     | Limit        |
| ------------------------------------------------------------------ | ---------- | ------------ |
| Each API key or OAuth token (REST and MCP)                         | 1 minute   | 100 requests |
| Each client IP, for signed-in app sessions and anonymous calls     | 1 minute   | 100 requests |
| Sign-in and other `/auth/*` routes, per IP                         | 15 minutes | 40 requests  |
| Signup, per IP                                                     | 1 hour     | 20 requests  |
| Password reset (request, verify and reset), per IP                 | 1 hour     | 12 requests  |
| OAuth client registration (`/auth/register`), per IP               | 1 hour     | 20 requests  |
| OAuth token and revocation (`/auth/token`, `/auth/revoke`), per IP | 15 minutes | 300 requests |
| Credentials that don't verify, per IP                              | 1 minute   | 30 requests  |

A bearer credential counts against itself, not the IP, once it verifies, so many agents behind one
host (Claude, ChatGPT) don't share one budget. A credential that doesn't verify (unknown, revoked or
malformed) counts against the IP instead, and after 30 of those in a minute the IP's API keys are
refused (`429`) until the window ends, except keys that verified in the last minute; an invalid
access token keeps getting the `401` challenge. A valid key's parallel requests are limited only by
its own 100 per minute, and a revoked key is refused on its next request. An expired access token
of ours counts against the IP but not as a failure: refresh it. CORS preflight (`OPTIONS`) requests
count against the IP too. An IPv6 client is counted per `/64` network; an IPv4 address written in
IPv6 form counts as that IPv4 address.
`/health`, `/.well-known/*` and public `/assets/*` aren't counted. The realtime streams aren't counted
either; each user may hold 10 open at a time.

Each account may hold 100 active API keys and register 50 browser origins; revoke or remove one to add
another.

Every counted response carries `RateLimit-Limit`, `RateLimit-Remaining` and `RateLimit-Reset` (seconds
until the window resets). Over the limit you get:

```http theme={null}
HTTP/1.1 429 Too Many Requests
Retry-After: 42
RateLimit-Limit: 100
RateLimit-Remaining: 0
RateLimit-Reset: 42

{ "data": null, "error": { "code": "RATE_LIMITED", "message": "Too many requests. Try again later." } }
```

Read `RateLimit-Remaining` on every response and slow down before it reaches zero.

### Backoff

1. **Honor `Retry-After`.** It says when to retry.
2. **Add jitter**, about ±25%, so clients don't all retry at the same moment.
3. **Cap retries** at 3 to 5, then show the error to the person.
4. **Be careful with writes.** Retry a create only with the same `Idempotency-Key` (REST) or
   `idempotencyKey` (MCP), so a retry can't create a duplicate.

### Polling cadence

| What you're waiting on                              | Poll at most                                             |
| --------------------------------------------------- | -------------------------------------------------------- |
| `GET /api/v2/buyer/campaigns/{id}/media-buy-status` | every 30 seconds                                         |
| A campaign after launch                             | every 15 to 30 seconds; prefer [webhooks](/buy/webhooks) |
| A proposal request (`get`, `kind: "proposal"`)      | every 5 to 10 seconds while sellers answer               |

If you poll faster and see `RATE_LIMITED`, slow down.

## Designing around them

* **Do not hold a call open.** `request_proposals` returns straight away with `running`; follow
  progress with `get` or in the Proposals widget instead of waiting on one call.
* **Keep pages small.** Use cursors and bounded `limit` values, and preserve partial results between
  pages.
* **Expect seller latency.** Sellers and their inventory systems often dominate response time even
  when Semicola itself is healthy.
* **Reuse keys only for true retries.** An idempotency key identifies one logical attempt.

<Note>
  Local development builds shorten some bounds so demos stay snappy, for example a two-minute tool
  call and a 20-second seller window.
</Note>
