Rate Limits
Boltch enforces per-key rate limiting, a daily free-model quota per account, and propagates upstream provider 429s when every configured upstream key is throttled.
Free model daily quota
Every authenticated account may make 50 free-model requests per UTC day. The allowance is counted per account, so people sharing a network or household each receive their own.
- The counter resets at midnight UTC every day.
- The allowance is enforced in the database with an atomic conditional update, so issuing requests in parallel cannot exceed it.
- Each successful free-model response includes an
X-Free-Requests-Remainingheader showing how many requests remain for the current UTC day. - Once the quota is exhausted the proxy returns HTTP 429 with
error.type = "free_quota_exceeded"and anX-Free-Reset-Atheader containing the ISO 8601 reset timestamp. - Your current usage and reset time are shown on the dashboard Usage page.
- Paid models are billed from your balance and are not affected by this allowance.
Free tier fair use
The free allowance is intended for evaluating the API. Accounts that appear to exist mainly to farm it may receive a reduced free allowance, assessed from several signals together rather than any single one.
- Sharing a network or IP address with other accounts does not reduce your allowance. Households, offices and mobile carriers routinely share one address.
- Using a VPN does not reduce your allowance.
- Any reduction affects the free allowance only. Account access and paid usage are never restricted by it.
- If you believe your account was assessed incorrectly, contact support: every automated decision is recorded with its reasons and can be reversed.
HTTP/1.1 429 Too Many Requests
Retry-After: 34821
X-Free-Reset-At: 2026-07-30T00:00:00Z
Content-Type: application/json
{
"type": "error",
"error": {
"type": "free_quota_exceeded",
"message": "You have used all 50 free-model requests for today. Resets at midnight UTC (2026-07-30T00:00:00Z)."
}
}
Free model health status
The status dot and Performance tab for each free model reflect real user traffic, not only scheduled probes. A probe that marks a model down is overridden when enough recent user requests succeed:
- At least 5 recent requests are required before traffic can change the displayed status.
- The success threshold scales with volume: 80% for 5–19 requests, 90% for 20–99, 95% for 100+.
- The dashboard refreshes health dots every 30 seconds while the page is open — no reload required.
- The Performance drawer for an open free model also refreshes automatically within the same 30-second window.
Per-key limits
The Worker maintains a sliding window per Boltch key. When a request exceeds the configured limit, the proxy returns HTTP 429 with error.type = "rate_limit_error" and a Retry-After header indicating when to retry.
Upstream propagation
If every upstream key for your account returns 429, Boltch surfaces the same rate_limit_error response and preserves the upstream Retry-After header value when one was provided.
Recommended client behavior
- Honor the
Retry-Afterheader before retrying. - Apply exponential backoff with jitter when the header is absent.
- Spread bursty traffic across more time so fewer requests pile up on the same window.
- For free models, check
X-Free-Requests-Remainingand pause requests once it reaches 0 untilX-Free-Reset-At. - Ask the operator to add additional upstream keys to your account if 429s persist.
Response shape
HTTP/1.1 429 Too Many Requests
Retry-After: 10
X-Request-ID: 9c8f6f2e-...
Content-Type: application/json
{
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "Rate limit exceeded. Retry after 10 seconds."
}
}