QuotchiGet Quotchi

Reference · Claude Code & Codex

The usage limits glossary

Session windows, cache writes, banked resets. Every term you’ll meet around AI coding limits, in plain words, with sources.

Last checked October 9, 2026. Limits change; the sources linked below are the final word.

See how Quotchi works
Quotchi reading a little notebook

About this glossary

39 terms you’ll meet when you hit, check or plan around Claude Code and Codex usage limits, each in plain words and taken from Anthropic’s and OpenAI’s own docs. Every term has its own link, so you can share /glossary#session-limit directly.

Limits and windows

Session limit (five-hour window)#

Claude’s short usage window on paid plans. Anthropic’s help center says it resets every five hours, and Claude Code’s error docs call it a rolling usage allowance. It applies across all models, so switching models doesn’t restore access. The message reads You've hit your session limit · resets 3:45pm. Anthropic doesn’t document exactly what starts the clock, so rely on the reset time it shows.

See when Claude Code resets.

Weekly limit#

The long window on Claude Pro and Max. For Max, Anthropic says it applies across all models and resets at a fixed time each week assigned to your account, regardless of when you subscribed. Every request counts against both the session and weekly allowance, so the weekly one can run out while session windows keep resetting. Message: You've hit your weekly limit · resets Mon 12:00am.

See Claude Code usage limits.

Model-specific limit#

A limit on one model family, shown as You've hit your Opus limit or You've hit your Sonnet limit. Unlike session and weekly limits, you can keep working by switching to a model outside that family with /model.

Fable weekly limit#

A separate weekly limit for Fable models on Max plans. When it’s reached, continuing on Fable uses usage credits, and Claude Code asks you to confirm first. If the prompt goes unanswered, nothing is sent.

Codex five-hour window#

A short usage window on some ChatGPT plans that include Codex. OpenAI’s help center says a new window starts with your first message after the previous one ends. OpenAI’s pricing page says Pro plans currently have no five-hour limit.

See Codex usage limits.

Codex weekly limit#

OpenAI says weekly limits may also apply on top of the five-hour window. The weekly period follows your previous weekly reset date, and using a banked or purchased reset can move it.

Seat allowance#

How Claude Team and Enterprise members are metered: each member draws on a per-seat allowance that resets on a rolling five-hour window and a weekly window, shared with Claude chat and Cowork. Its size depends on the seat tier (Standard or Premium).

Rate limit#

Often used loosely for any usage limit, but in the API sense it means requests or tokens per minute. Claude Code’s Request rejected (429) and Server is temporarily limiting requests errors are this kind, and the docs say the second is not your usage limit. Too many parallel subagents can trigger it on an API key.

Overloaded (529)#

A capacity error: Repeated 529 Overloaded errors. Claude Code’s docs say it isn’t your usage limit and doesn’t count against your quota. Retry, check status.claude.com, or switch models, since capacity is tracked per model.

Keeping going past a limit

Usage credits (extra usage)#

Anthropic’s pay-as-you-go top-up for Pro, Max 5x and Max 20x. Once enabled in Settings → Usage, you keep working after reaching your included limits, billed at standard API rates and separately from your subscription. Run /usage-credits in Claude Code to open the settings. Some older help-center links still use the name “extra usage”.

See limit reached: what to do.

Monthly spend limit#

The cap you set on usage-credit spending each month (or “unlimited”). Reaching it shows You've hit your monthly spend limit. Team and Enterprise admins can set spend limits for the organization, groups or individual members.

Limit reset (Claude)#

A one-time reset Anthropic occasionally gives eligible plans. You use it from Settings → Usage on the web or in Claude Desktop with Reset for free; it isn’t available in Claude Code in a terminal or IDE. It restores either your session or your weekly limit, and your weekly limit still resets on its usual day and time.

Banked reset (Codex)#

A one-time Codex usage-limit reset saved to your account until you use it or it expires. Using a full reset refreshes the five-hour and weekly windows and changes your weekly reset date. Redeem it from settings → usage, or from the /usage menu in the CLI.

See how to check Codex usage.

Codex credits#

The unit OpenAI uses to pay for eligible usage on credit-based plans. Plus and Pro users who reach their limit can buy credits to keep working without upgrading; Business, Edu and Enterprise plans with flexible pricing buy workspace credits.

Auto-continue#

Claude Code’s option to wait out a usage limit in the open session and resume the task after the reset. On by default from v2.1.234 in interactive sessions signed in with a claude.ai subscription. Controlled by Continue automatically at usage limit in /config; /rate-limit-options starts or cancels a wait.

Measuring usage

/usage (Claude Code)#

Shows plan usage bars and reset times, plus a breakdown of what used your allowance on paid plans. /cost and /stats open the same screen. The breakdown is estimated from local session history, so claude.ai and other devices aren’t in it. If the usage request is rate limited, it shows bars from the last 60 minutes labeled Showing last-known usage.

/status#

In Claude Code: version, model, account and connectivity, handy for spotting a stray API key. In Codex: session configuration and token usage, and, per OpenAI’s pricing page, your remaining limits in an active CLI session.

Status line#

A customizable line at the bottom of the terminal. Claude Code runs a script you choose and pipes it session data, including rate_limits.five_hour and rate_limits.seven_day for Pro and Max subscribers. Codex’s /statusline lets you add a rate-limits item to its footer.

See limits in the status line.

Codex usage dashboard#

The page at chatgpt.com/codex/settings/usage that shows your current Codex limits and reset times. It covers local and cloud use.

Token#

The unit models read and write text in. Claude Code charges by token consumption, and thinking tokens are billed as output tokens. Plan limits aren’t published in tokens or messages, but more tokens means more usage.

Context and caching

Context window#

The working memory of a session: conversation history, file contents, command output, CLAUDE.md, loaded skills and system instructions. Run /context in Claude Code to see what fills it. A full context isn’t a usage limit, but long context makes every request bigger.

Compaction#

Summarizing the conversation when the context window nears its limit, automatically or with /compact. The summarization is itself a request that reads the whole conversation, so it isn’t free. /clear starts fresh and costs nothing.

Prompt caching#

How the API avoids reprocessing the unchanged start (prefix) of each request. Claude Code re-sends your full context every turn; with caching, the repeated part is billed at the cheaper cached rate and only what changed is processed fully.

Cache read#

Tokens served from the prompt cache on a turn (cache_read_input_tokens), billed at the model’s cached token rate, below the standard input rate. A high read share means caching is working.

Cache write#

Tokens written to the cache on a turn (cache_creation_input_tokens), billed at the cache write rate. One-hour cache writes cost more than five-minute ones. If writes stay high turn after turn, something keeps changing your prefix.

Cache lifetime (TTL)#

How long a cached prefix survives inactivity. On a Claude subscription within plan usage, Claude Code requests one hour for the main conversation and five minutes for subagents and other requests. It drops to five minutes once you’re drawing on usage credits.

Cache miss#

A request that reprocesses content the cache already held, for example after a break longer than the TTL, a model switch, connecting an MCP server that loads tools upfront, or upgrading Claude Code. The first turn after one is slower and uses more.

Agents and extensions

Subagent#

A specialized assistant Claude Code spawns with its own context window, system prompt, tools and permissions. It returns a summary to the main conversation. Its requests count toward the same usage limits as your main conversation, so many subagents drain a window faster.

Agent teams#

Multiple Claude Code sessions coordinated by a lead, with a shared task list and direct messaging. Experimental and off by default (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1). Anthropic says teams use significantly more tokens, about 7x a standard session when teammates run in plan mode.

Hooks#

User-defined handlers (shell commands, HTTP endpoints, MCP tools, prompts or subagents) that run automatically at points in Claude Code’s lifecycle, such as PreToolUse, PostToolUse, UserPromptSubmit or Stop. A hook that filters test output before Claude reads it is a classic way to save tokens.

Skill#

A SKILL.md file of instructions or a workflow that Claude loads when relevant or when you type /skill-name. Skills load on demand, so moving specialized instructions out of CLAUDE.md into skills keeps every request smaller.

MCP server#

A program that gives Claude extra tools over the Model Context Protocol. Tool definitions are deferred by default, but every tool result still adds to context. /usage attributes usage to individual MCP servers.

Plan mode#

A permission mode where Claude researches and proposes a plan before editing anything. Anthropic recommends it for complex tasks to avoid expensive rework.

Effort level#

Controls how much the model thinks on each step. Higher effort means more thinking tokens; lower is faster and cheaper. Change it with /effort or in /model.

Fast mode#

A faster speed setting. In Claude Code, turning it on mid-session causes one cache miss, and it’s billed at fast mode rates. In Codex, Fast and Ultrafast modes use included limits and credits at different rates.

Plans and accounts

Pro#

Anthropic’s individual paid plan that includes Claude Code, listed at $20 a month or $17 a month billed yearly. It has a five-hour session limit and a weekly limit, shared with claude.ai.

See Max vs Pro.

Max 5x and Max 20x#

Anthropic’s higher individual plans at $100 and $200 a month (web subscriptions), with 5 and 20 times Pro’s per-session allowance. Weekly limits still apply, plus a separate Fable weekly limit.

API key override#

If ANTHROPIC_API_KEY is set, Claude Code uses it instead of your subscription and you pay API rates; environment variables take precedence over /login. Check the API key row in /status. In Codex, an API key is a deliberate way to run extra local chats at API rates.

Local vs cloud tasks (Codex)#

Codex runs locally (CLI, IDE extension, app) or in the cloud. OpenAI says both share your plan’s allowance and that cloud tasks may use more of it than local messages.

Where to go next

Quotchi turns the limits on this page into one glance: session and weekly windows, reset times and which tool has room, in the Mac menu bar. Free and local-first.

See how Quotchi works

Sources

Official pages we read on October 9, 2026. If anything here disagrees with them, they win.

FAQ

What’s the difference between a session limit and a weekly limit?

The session limit resets every five hours; the weekly limit resets once a week at a fixed time. Every Claude request counts against both.

Is a rate limit the same as a usage limit?

Not quite. Usage limits are your plan allowance. Rate limits in the API sense cap requests or tokens per minute, and some 429 errors in Claude Code are explicitly not your usage limit.

What are usage credits?

Anthropic’s pay-as-you-go top-up for Pro and Max, billed at standard API rates, that lets you keep working past your included limits.

What’s the difference between a cache read and a cache write?

A cache write stores the start of your request for reuse and is billed at the cache write rate. A cache read reuses it on a later turn at the cheaper cached rate.

Do subagents use my usage limit?

Yes. Claude Code’s docs say a subagent sends its own requests, which count toward the same usage limits as your main conversation.

What is a banked Codex reset?

A one-time Codex usage-limit reset saved to your account until you use it or it expires. It refreshes both windows and changes your weekly reset date.