Rate Limits

Boltch enforces per-key rate limiting, a daily free-model quota per account, and propagates upstream provider 429s when every configured upstream key is throttled.

Free model daily quota

Every authenticated account may make 50 free-model requests per UTC day. The allowance is counted per account, so people sharing a network or household each receive their own.

Free tier fair use

The free allowance is intended for evaluating the API. Accounts that appear to exist mainly to farm it may receive a reduced free allowance, assessed from several signals together rather than any single one.

HTTP 429 — quota exhausted
HTTP/1.1 429 Too Many Requests
Retry-After: 34821
X-Free-Reset-At: 2026-07-30T00:00:00Z
Content-Type: application/json

{
  "type": "error",
  "error": {
    "type": "free_quota_exceeded",
    "message": "You have used all 50 free-model requests for today. Resets at midnight UTC (2026-07-30T00:00:00Z)."
  }
}

Free model health status

The status dot and Performance tab for each free model reflect real user traffic, not only scheduled probes. A probe that marks a model down is overridden when enough recent user requests succeed:

Per-key limits

The Worker maintains a sliding window per Boltch key. When a request exceeds the configured limit, the proxy returns HTTP 429 with error.type = "rate_limit_error" and a Retry-After header indicating when to retry.

Upstream propagation

If every upstream key for your account returns 429, Boltch surfaces the same rate_limit_error response and preserves the upstream Retry-After header value when one was provided.

Recommended client behavior

Response shape

HTTP 429
HTTP/1.1 429 Too Many Requests
Retry-After: 10
X-Request-ID: 9c8f6f2e-...
Content-Type: application/json

{
  "type": "error",
  "error": {
    "type": "rate_limit_error",
    "message": "Rate limit exceeded. Retry after 10 seconds."
  }
}