MCP bounds
Rate limits
Semicola counts requests in fixed windows:
A bearer credential counts against itself, not the IP, once it verifies, so many agents behind one
host (Claude, ChatGPT) don’t share one budget. A credential that doesn’t verify (unknown, revoked or
malformed) counts against the IP instead, and after 30 of those in a minute the IP’s API keys are
refused (
429) until the window ends, except keys that verified in the last minute; an invalid
access token keeps getting the 401 challenge. A valid key’s parallel requests are limited only by
its own 100 per minute, and a revoked key is refused on its next request. An expired access token
of ours counts against the IP but not as a failure: refresh it. CORS preflight (OPTIONS) requests
count against the IP too. An IPv6 client is counted per /64 network; an IPv4 address written in
IPv6 form counts as that IPv4 address.
/health, /.well-known/* and public /assets/* aren’t counted. The realtime streams aren’t counted
either; each user may hold 10 open at a time.
Each account may hold 100 active API keys and register 50 browser origins; revoke or remove one to add
another.
Every counted response carries RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset (seconds
until the window resets). Over the limit you get:
RateLimit-Remaining on every response and slow down before it reaches zero.
Backoff
- Honor
Retry-After. It says when to retry. - Add jitter, about ±25%, so clients don’t all retry at the same moment.
- Cap retries at 3 to 5, then show the error to the person.
- Be careful with writes. Retry a create only with the same
Idempotency-Key(REST) oridempotencyKey(MCP), so a retry can’t create a duplicate.
Polling cadence
If you poll faster and see
RATE_LIMITED, slow down.
Designing around them
- Do not hold a call open.
request_proposalsreturns straight away withrunning; follow progress withgetor in the Proposals widget instead of waiting on one call. - Keep pages small. Use cursors and bounded
limitvalues, and preserve partial results between pages. - Expect seller latency. Sellers and their inventory systems often dominate response time even when Semicola itself is healthy.
- Reuse keys only for true retries. An idempotency key identifies one logical attempt.
Local development builds shorten some bounds so demos stay snappy, for example a two-minute tool
call and a 20-second seller window.