The quick answer
exceeded retry limit, last status: 429 Too Many Requests does not by itself prove that your Codex subscription allowance is exhausted. First identify the account and provider, then read the underlying error. Temporary throttling can recover after waiting; an API balance or spending restriction needs the corresponding account action. A new chat or a local countdown cannot fix every 429.
What does “exceeded retry limit” mean?
Read the message as two clues: the client stopped retrying, and the last HTTP response was 429. This is a diagnostic interpretation of the supplied message, not an official guarantee about every client's retry count. The App Server protocol separates failed-attempt errors from upstream HTTP status information.
Copy the preceding error details if available. The HTTP status alone does not distinguish request throttling from API quota or billing restrictions. Repeatedly restarting the same task can add more requests without resolving the cause.
1. Identify which service rejected the request
Check the account and workspace in the client; in the CLI, use codex login status. Codex authentication distinguishes ChatGPT sign-in from API-key access.
- ChatGPT sign-in: inspect the same account's Codex usage dashboard. If a usage window is exhausted, follow Codex limit reached. An API billing balance is not your subscription allowance.
- OpenAI API key: inspect the returned error and the relevant API organization/project's usage and billing. Do not use a ChatGPT reset time to diagnose an API limit.
- Custom provider or gateway: check which provider the failing request actually reached. Its rate limits and billing may differ; OpenAI API error codes apply only when that service returns them. Check its own documentation and support channel.
2. Separate temporary throttling from exhausted quota
When available, inspect error.code, error.type and the message together. Do not infer a specific cause from a missing field. The API error reference distinguishes these cases:
- Request throttling: the response indicates a rate limit, possibly
slow_down. Reduce simultaneous tasks and burst traffic before trying again. - Credit balance:
credit_balance_exhaustedpoints to API credits. Check billing with the account owner; waiting for a Codex subscription reset will not replenish that balance. - Spending or usage restriction:
organization_spend_limit_exceeded,project_spend_limit_exceededororganization_usage_limit_exceededidentifies a different boundary. Check its applicable period and account settings with an administrator before authorizing any extra spending. - Broad quota type:
insufficient_quotamay be less specific thanerror.code. Read the detailed message instead of treating it as a temporary request-rate problem.
These are OpenAI API examples, not a promise that the Codex interface exposes each field. If the interface shows only 429, keep the cause unresolved until account information or a fuller error identifies it.
3. Retry only when waiting can help
For temporary API throttling, follow Retry-After when present. Without it, OpenAI recommends exponential backoff with jitter: progressively increase the delay and add randomness. Any retry loop you control should have both an attempt limit and a total time limit. If the required wait exceeds that budget, defer the task instead of shortening the server's delay.
For an interactive Codex task, pause repeated submissions, reduce concurrent work and try one small continuation after the indicated wait. Do not layer an aggressive manual loop on top of client retries. A credit or spending error is not repaired by backoff; there is no universal “wait five minutes” fix.
Before continuing an interrupted task, inspect files, command results and external actions already completed. Resume from a known step so a retry does not duplicate an operation.
401, 403, 500, 503 and interrupted streams
Use the actual message to choose the next check; these failures do not all mean a usage reset is due. See the error reference and Codex troubleshooting.
- 401: check authentication, the selected account and credential validity. Never paste a key into a support post.
- 403: inspect the stated access, policy or region restriction. Repeated requests do not grant permission.
- 500 or 503: a server error or temporary overload can warrant a delayed, bounded retry. If it persists, check provider service information and report the failure.
- Stream disconnected or timeout: the connection ended before the client received a complete result. This wording alone does not establish a quota problem or prove that no work completed. Inspect the underlying status, client logs and network path, then verify task progress before resuming.
What to record if the error persists
Keep the timestamp and timezone, client/version, model, sign-in method, provider, sanitized message, HTTP status and request ID if supplied. Note whether one small request fails too, whether other clients work, and which account checks you performed. Remove API keys, authorization headers, cookies and private task content before sharing logs through the provider's support channel.
If the evidence instead confirms an exhausted subscription window, use the Codex reset-time guide. The personal timer can remind you of a displayed recovery time; it neither removes API throttling nor confirms that access has returned.
