This is the full developer documentation for Limitry # Introduction > Limitry tells your code whether a user, team, API key or agent may do something that costs you money, and keeps the count. Limitry answers one question for your code: may this subject do this action, at this cost? Ask it before anything that costs you money or could be abused, such as a model call, an image, a paid API request or an export. Limitry keeps the counts, windows, balances and holds behind the answer, so you don’t have to. A few words you’ll see throughout these docs: * A **subject** is whoever acts in your product: a user, a team, an API key, an AI agent. You choose the kinds and the ids. * A **limit** is a rule, such as “10 images per user per day”. * A **check** is the question your code asks. The answer is allowed or not, with what’s left and when it resets. ## Where to start [Section titled “Where to start”](#where-to-start) * [Quickstart](/docs/quickstart/): a first limit and a first check, in about five minutes. * [Add Limitry with your agent](/docs/agents-onboarding/): let Claude Code or Codex find the costly calls in your code and propose limits. * [How Limitry works](/docs/how-it-works/): how Limitry decides, what happens under load, and what happens if Limitry is down. Everything you can do in the app, you can also do through the [API](/docs/authentication/), the [command line](/docs/cli/) and the [MCP server](/docs/mcp/). # Use with AI agents > Let an AI agent work in your workspace through the MCP server, the command line or the API, and teach it how with the agent skill. AI agents can do in your workspace what you could do yourself, within the access you give them. They act as **you**, and the audit log records what they did. There are three ways to connect one, from least setup to most: | Your agent | Use | | --------------------------------------- | ------------------------------------------------------ | | Claude, ChatGPT or another app with MCP | The [MCP server](/docs/mcp/). Nothing to install. | | A coding agent with a terminal | The [command line](/docs/cli/), with `--json`. | | Your own software | The [API](/docs/quickstart/) or the [SDK](/docs/sdk/). | ## MCP server [Section titled “MCP server”](#mcp-server) Add `https://mcp.limitry.com/mcp` as a remote MCP server in your agent. You sign in once in the browser and choose **one workspace** and what the agent may do there, and that’s all it can do. See [MCP server](/docs/mcp/) for the details and safeguards. ## Command line [Section titled “Command line”](#command-line) Coding agents such as Claude Code, Cursor and Codex work well with the [command line](/docs/cli/): ```bash npm install -g @limitry/cli limitry login # you approve it in the browser limitry workspace list --json limitry workspace webhook-endpoints list --workspace acme --json ``` With `--json`, the agent gets the API’s exact JSON, errors as JSON on stderr, and an exit code of 0 or 1. In CI, or a sandbox without a browser, set `LIMITRY_TOKEN` to a [personal access token](/docs/authentication/) limited to what the agent needs. ## Teach your agent: the skill [Section titled “Teach your agent: the skill”](#teach-your-agent-the-skill) The agent skill is one file that teaches an agent everything on this page: how to sign in, choose the workspace, every command and operation, the error codes, and what to leave to a person. It’s at [`/docs/skill/SKILL.md`](https://limitry.com/docs/skill/SKILL.md), and it’s generated from the API, so it always matches. To install it for every project in Claude Code: ```bash mkdir -p ~/.claude/skills/limitry curl -fsSL https://limitry.com/docs/skill/SKILL.md \ -o ~/.claude/skills/limitry/SKILL.md ``` Or put it in a project’s `.claude/skills/limitry/`. Other agents that support skills take the same file in their own skills folder. ## Docs for agents [Section titled “Docs for agents”](#docs-for-agents) The documentation is also available as plain text: [`/docs/llms.txt`](/docs/llms.txt) indexes every page, and [`/docs/llms-full.txt`](/docs/llms-full.txt) has everything in one file. The site’s [`/llms.txt`](https://limitry.com/llms.txt) points to these, the API, the MCP server and the skill. ## What agents can’t do [Section titled “What agents can’t do”](#what-agents-cant-do) Some operations stay with people, even when an agent has full access: anything that sends your workspace’s data somewhere new or gives out access, such as creating API keys or adding webhook endpoints. The agent asks instead, and you approve or deny (see [Approvals](/docs/approvals/)). A secret created this way never reaches the agent. # Add Limitry with your agent > Let Claude Code or Codex find the costly calls in your code, propose limits, and wire them in, with you deciding. Your coding agent can add Limitry for you. It reads your code, proposes limits, asks you before creating any, and wraps the costly calls. ## 1. Connect the agent [Section titled “1. Connect the agent”](#1-connect-the-agent) Give it Limitry’s tools, through the [MCP server](/docs/mcp/) or the [command line](/docs/cli/), and the skill that explains how Limitry works ([Use with AI agents](/docs/agents/)). ## 2. Paste this prompt [Section titled “2. Paste this prompt”](#2-paste-this-prompt) ```text Add Limitry to this codebase. Find the calls that cost us money or that a user could abuse (LLM calls, paid APIs, emails, exports). Propose limits in a table: subject, action, amount, window, fail mode. Wait for my answer before creating any. Then create the ones I agree to, wrap the code with the Limitry SDK's guard or hold helpers, handle a refusal where we answer the user, and add a test for it. ``` ## 3. What it writes [Section titled “3. What it writes”](#3-what-it-writes) For a call whose cost you know up front, `guard` checks first and runs the work only if it’s allowed: ```ts import { createClient, createLimits, LimitExceeded } from "@limitry/sdk"; const limits = createLimits( createClient({ token: process.env.LIMITRY_API_KEY! }), ); const image = await limits.guard( { subject: { kind: "user", id: user.id }, action: "generate-image" }, () => generateImage(prompt), ); ``` When the cost is only known afterwards, as with tokens, or for “at most N at once”, `hold` reserves an estimate, runs the work, and commits what it actually used. If the work fails, it releases the hold: ```ts const answer = await limits.hold( { subject: { kind: "user", id: user.id }, action: "chat", estimate: 4000 }, async (held) => { const res = await llm(prompt); held.cost(res.usage.totalTokens); return res; }, ); ``` A refusal throws `LimitExceeded`, with `limit`, `reason` and `retryAfterSeconds`, for your code to turn into a 429 or a message. ## If Limitry is unreachable [Section titled “If Limitry is unreachable”](#if-limitry-is-unreachable) Each limit has a fail mode. The SDK remembers the fail mode from the last answer, and if Limitry doesn’t answer within 1.5 seconds (`timeoutMs`), it either lets the work run (`open`, the default) or throws `LimitExceeded` with the reason `unavailable` (`closed`). Your product never hangs waiting for Limitry. # Approvals > When an agent needs something it may not do alone, such as an API key or a webhook, it asks and a person decides. Some actions are never an agent’s to take alone, because they give out access or send your data somewhere new: creating an API key, or adding, removing or re-enabling a webhook endpoint. An agent that needs one **asks**, and a person who could do it themselves decides. ## How it works [Section titled “How it works”](#how-it-works) 1. **The agent asks,** with the action, its details and a reason, such as “I’m connecting acme-web and need an API key that can read the workspace”. Everyone who could approve it gets a notification. 2. **A person decides,** from the notification or the link the agent shows. You see exactly what will happen and why, and choose **Approve** or **Deny**. 3. **It happens as you,** with your permissions. The audit log records it under your name, along with the agent that asked. A request expires after 24 hours, and a person can decide it only once. ## Secrets stay out of the agent’s hands [Section titled “Secrets stay out of the agent’s hands”](#secrets-stay-out-of-the-agents-hands) When the action creates a secret, such as an API key or a webhook signing secret, the agent never sees it: * **Asked from the command line:** after you approve, the command line writes the secret straight into a file in your project. It’s never printed, so an agent working in the terminal can’t read it. ```bash limitry workspace approval-requests create --action workspace.apiKey.create \ --input '{"name":"acme-web","scopes":["workspace:read"]}' \ --reason "Connect acme-web" --json limitry workspace approval-requests redeem --wait --write-env .env # Done: … Wrote LIMITRY_API_KEY to .env (the value is not shown). ``` * **Asked by an agent over MCP:** the action runs when you approve, and **you** see the secret, once, on the approval page. ## For developers [Section titled “For developers”](#for-developers) `GET /v1/workspace/approval-actions` lists what an agent can ask for, each with its input schema and who approves it. `POST /v1/workspace/approval-requests` asks, and `GET /v1/workspace/approval-requests/{id}` reports the decision. Asking needs the `workspace.approvals:write` scope and a person’s credential: a personal access token, or an app a person connected, but not a workspace API key. # Authentication > Authenticate API requests with a workspace API key, a personal access token, or OAuth. The public API accepts a **workspace API key** in the standard Bearer header: ```http Authorization: Bearer YOUR_API_KEY ``` ```bash curl \ -H "Authorization: Bearer YOUR_API_KEY" \ https://api.limitry.com/v1/workspace ``` The response names the workspace the key belongs to: ```json { "id": "…", "name": "Example Workspace" } ``` ## API keys belong to a workspace [Section titled “API keys belong to a workspace”](#api-keys-belong-to-a-workspace) A key acts as a workspace, not a person. It keeps working if the person who created it leaves the workspace. It never carries a user session, so nobody can use it to sign in. ## Creating a key [Section titled “Creating a key”](#creating-a-key) Workspace owners create keys in the app, under **Settings → API Keys**: 1. Give it a name that says where it’s used, such as `production-backend`. 2. Choose what it may do (see [Scopes](#scopes)). It’s read-only unless you choose otherwise. 3. Copy the secret. We show it once and you can’t view it again, so put it in your secret manager straight away. The list of keys shows only what is safe to show: the name, the first few characters, the access, and when it was created and last used. If a secret is lost, create a new key and revoke the old one. ## Revoking and rotating [Section titled “Revoking and rotating”](#revoking-and-rotating) Owners can revoke a key at any time under **Settings → API Keys**. It stops working on the next request, and you can’t restore it. To rotate a key without downtime: 1. Create a new key. 2. Switch your integration to it. 3. Revoke the old key. ## Scopes [Section titled “Scopes”](#scopes) You give every key and token **scopes** when you create it, and it can do only what they allow. A scope names a group of endpoints and whether it may read or change them: `workspace:read` reads the workspace profile, `workspace.members:read` reads members and pending invitations, and `workspace.audit:read` reads the audit trail. The [API reference](https://api.limitry.com/v1/docs) shows the scope each endpoint needs. When you create a credential, you choose: | Access | What it gets | | --------------- | ------------------------------------------------------------------------------------------- | | **Read only** | The default. Every read scope that exists at creation. Scopes added later are not included. | | **Full access** | Everything, including scopes added to the API later. | | **Custom** | Exactly the scopes you tick. | Give an integration only what it needs. Calling an endpoint the credential wasn’t given returns `403 forbidden`: ```json { "error": { "code": "forbidden", "message": "This credential is not granted the workspace.audit:read scope", "requestId": "…" } } ``` You can’t change a credential’s scopes after you create it. To give an integration different access, create a new credential and revoke the old one. ## Personal access tokens [Section titled “Personal access tokens”](#personal-access-tokens) A personal access token (`pat_…`) acts as **you**, not as a workspace. Use one for your own scripts, CI jobs and the command line. Create it in the app under **Account → API tokens**, or let the command line create one. Its `login` opens your browser, and once you approve, the command line creates the token for that computer and stores it for you. A personal access token: * works in the workspaces you choose when you create it, all of yours or specific ones, with the [scopes](#scopes) you give it; * is also limited by your role in each workspace. For example, `GET /v1/workspace/audit-events` is for owners only, as the audit page is in the app. (A workspace API key is the workspace itself and has no role.) * expires when you choose: after 30 days, 90 days, a year, or never; * identifies you at `/v1/user/me`: ```bash curl \ -H "Authorization: Bearer YOUR_PERSONAL_ACCESS_TOKEN" \ https://api.limitry.com/v1/user/me ``` ```json { "id": "…", "name": "Casey Customer", "email": "casey@example.com", "workspaces": [ { "id": "…", "name": "Example", "slug": "example", "role": "owner" } ] } ``` Everywhere else, a personal access token has to say which workspace it’s acting in, with the `X-Workspace` header set to the workspace’s slug. The call then runs with your membership in that workspace and the token’s scopes: ```bash curl \ -H "Authorization: Bearer YOUR_PERSONAL_ACCESS_TOKEN" \ -H "X-Workspace: example" \ https://api.limitry.com/v1/workspace ``` Without `X-Workspace` the request fails with `400`. A workspace you’re not a member of, one the token wasn’t created for, or a scope the token doesn’t have, fails with `403`. Workspace API keys ignore the header, because a key already is its workspace, and they’re never accepted on `/v1/user/*`. Like API keys, we show tokens once, and you can revoke them from the same page. They stop working at once if they expire or your account is suspended. ## Apps and AI agents (OAuth) [Section titled “Apps and AI agents (OAuth)”](#apps-and-ai-agents-oauth) Apps, including AI agents connecting their tools, can act for a person without ever holding a key, because the product is an OAuth 2.1 authorization server. The app sends the person to sign in. They choose a workspace and approve what the app may do, and the app gets a short-lived access token for the API. * **Discovery:** `https://api.limitry.com/.well-known/oauth-authorization-server/auth` lists every endpoint. Apps can register themselves (dynamic client registration); registering grants nothing until a person approves. * **Scopes:** an app asks for the same [scopes](#scopes), such as `workspace.audit:read`, plus `offline_access` for a refresh token. The person can give it fewer. * **One person, one workspace:** a token acts for one person in one workspace, with that person’s role, so it needs no `X-Workspace` header. Request it for the resource `https://api.limitry.com/v1`. * **Disconnecting:** the person can disconnect the app at any time under **Account → Connected apps**. Its tokens stop working on the next request. ## When authentication fails [Section titled “When authentication fails”](#when-authentication-fails) A missing, malformed, invalid or revoked credential gets a `401` with the standard [error envelope](/docs/errors/): ```json { "error": { "code": "unauthorized", "message": "Unauthorized", "requestId": "…" } } ``` The response is the same whatever the cause, so it gives nothing away to someone guessing keys. Browser session cookies are never accepted on the public API. # Checks > Ask whether a subject may do an action at a cost, and what to do with the answer. A check is the question your code asks before doing something that costs: may this subject do this action, at this cost? Limitry looks at every enabled [limit](/docs/limits/) that matches, and answers. ```bash curl -X POST https://api.limitry.com/v1/checks \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: 7f9c…" \ -d '{"subject":{"kind":"user","id":"usr_123"},"action":"generate-image","cost":1}' ``` ```json { "id": "chk_9d2e…", "allowed": false, "reason": { "code": "limit_exceeded", "limit": "lim_3f0c…" }, "remaining": [ { "limit": "lim_3f0c…", "name": "Daily images", "unit": "images", "window": "day", "remaining": 0, "resetsAt": "2026-10-05T00:00:00.000Z" } ], "retryAfterSeconds": 41200 } ``` `cost` defaults to 1. ## Reading the answer [Section titled “Reading the answer”](#reading-the-answer) Every decision is HTTP 200, allowed or not, so branch on `allowed`. An error status means the request itself failed: a bad key, an invalid body, or calling too fast. `remaining` lists every limit that applied, so you can show “3 left today” without another call. When the answer is no, `reason` says which limit refused and `retryAfterSeconds` says when to try again. It’s `null` when waiting won’t help, such as an empty balance. ## Several limits and shared pools [Section titled “Several limits and shared pools”](#several-limits-and-shared-pools) When several limits match, the check is allowed only if the cost fits all of them, and then it counts against all of them. To draw from a shared pool as well, name up to 4 more subjects: ```json { "subject": { "kind": "user", "id": "usr_7" }, "shared": [{ "kind": "team", "id": "t_42" }], "action": "ai-call" } ``` Every limit on every subject must fit. Limitry charges all of them, or none. See [Pools, balances and budgets](/docs/pools-and-budgets/). ## Modes [Section titled “Modes”](#modes) | `mode` | What it does | | --------- | ---------------------------------------------------------------------------------------- | | `consume` | The default. Decides and counts. | | `preview` | Decides without counting: “would this be allowed?” | | `reserve` | Decides and holds the cost until you settle it. See [Reservations](/docs/reservations/). | ## Good to know [Section titled “Good to know”](#good-to-know) * Retries are safe with an `Idempotency-Key` header: for 24 hours the same key gets the same answer and counts nothing again. * If Limitry can’t be reached, follow each limit’s `failMode`. See [How Limitry works](/docs/how-it-works/). * The key needs the `checks:write` scope. * Limit changes reach checks within a minute. * Allowed checks count toward your [plan](/docs/plans/). Denied ones don’t. The same operation is `limitry checks create` on the [command line](/docs/cli/), and a tool on the [MCP server](/docs/mcp/). # Command line > Install the CLI, sign in, and call every API operation from a terminal or script. The command line uses the same public API as your code. It acts as **you**, with a personal access token, in the workspaces you allow. ## Install [Section titled “Install”](#install) ```bash npm install -g @limitry/cli ``` ## Sign in [Section titled “Sign in”](#sign-in) ```bash limitry login ``` Your browser opens. Sign in, choose which workspaces the command line may use and what it may do, and click **Allow**. The command line creates a personal access token for this computer, valid for 90 days and listed under **Account → API tokens**, and stores it in `~/.config/limitry/`. On a server or in CI there’s no browser, so pass a token instead, or set it in the environment. Then it writes nothing to disk: ```bash limitry login --token "$TOKEN" # or: echo "$TOKEN" | limitry login LIMITRY_TOKEN=pat_… limitry whoami ``` `limitry logout` revokes the token and forgets it. ## Choose a workspace [Section titled “Choose a workspace”](#choose-a-workspace) ```bash limitry workspace list limitry workspace use acme # the default for later commands limitry workspace audit-events list --workspace other-co # one call elsewhere ``` ## Every API operation is a command [Section titled “Every API operation is a command”](#every-api-operation-is-a-command) Each operation in the [API reference](https://api.limitry.com/v1/docs) is a command, named after its resource and action: ```bash limitry workspace webhook-endpoints list limitry workspace webhook-endpoints create --url https://example.com/hooks limitry workspace webhook-deliveries list we_123 --limit 10 limitry workspace webhook-endpoints delete we_123 --yes ``` * Ids in the path are arguments, and other inputs are flags. `limitry --help` lists them with descriptions. * `--data '{…}'` sends a whole JSON body. * Lists print one page. The last line gives the `--cursor` for the next. * Writes send an [idempotency key](/docs/idempotency/) for you. * An operation that can’t be undone asks for `--yes`. * `limitry api ` calls any endpoint directly. ## Scripts and agents [Section titled “Scripts and agents”](#scripts-and-agents) Add `--json` to any command to get the API’s exact JSON on stdout. Errors then go to stderr as JSON too, in the API’s [error envelope](/docs/errors/) with the HTTP status: ```bash limitry workspace webhook-endpoints list --json | jq '.items[].url' ``` ```json { "error": { "code": "forbidden", "message": "…", "status": 403, "requestId": "…" } } ``` The exit code is `0` on success and `1` on any failure, including a refused API call or a missing flag. # Errors & API rate limits > The error envelope, the error codes you can rely on, and how fast each credential may call the API. Every response outside the 2xx range uses the same envelope: ```json { "error": { "code": "invalid_request", "message": "Invalid request", "requestId": "9f0c1c2e-…", "docs": "https://limitry.com/docs/errors/#invalid_request", "details": [{ "path": "name", "message": "Required" }] } } ``` | Field | What it is | | ----------- | -------------------------------------------------------------------------------- | | `code` | One of a small, stable set of codes. Branch on this. | | `message` | A sentence for people. It may change, so don’t parse it. | | `requestId` | Matches the `X-Request-Id` response header. Include it when you contact support. | | `details` | Only on validation errors: one entry per field, with its `path` and `message`. | | `docs` | A link to the code’s entry on this page. | ## Error codes [Section titled “Error codes”](#error-codes) We may add new codes over time. Treat a code you don’t recognise as `internal_error`. ### `unauthorized` [Section titled “unauthorized”](#unauthorized) `401`. The credential is missing, malformed, invalid, revoked or expired. Send `Authorization: Bearer ` with a working credential, and create a new one if it was revoked or has expired. ### `forbidden` [Section titled “forbidden”](#forbidden) `403`. The credential is valid but isn’t allowed to do this. Either it doesn’t have the endpoint’s [scope](/docs/authentication/#scopes), or your role in the workspace doesn’t allow it, or a personal access token can’t act in the workspace named by `X-Workspace`. Use a credential with the access it needs; the `message` names the missing scope. ### `not_found` [Section titled “not\_found”](#not_found) `404`. The resource doesn’t exist, or isn’t yours to see. Check the path and the id; the API reports an id from another workspace as not found. ### `method_not_allowed` [Section titled “method\_not\_allowed”](#method_not_allowed) `405`. The path exists, but not with this method. Use one of the methods in the `Allow` response header. ### `invalid_request` [Section titled “invalid\_request”](#invalid_request) `400`. The request didn’t pass validation. `details` lists each invalid field with its `path` and `message`. ### `rate_limited` [Section titled “rate\_limited”](#rate_limited) `429`. The credential has used its quota (see [API rate limits](#api-rate-limits)). Wait the number of seconds in `Retry-After`, then retry with backoff. ### `plan_limit_reached` [Section titled “plan\_limit\_reached”](#plan_limit_reached) `402`. The workspace has used its plan’s monthly allowance for what this request needs. The message says which allowance and when it resets, and the workspace’s owners got an email and a notification at 80% and 100%. Upgrade the plan under **Settings → Billing**, or wait for the month to reset; retrying sooner gets the same answer. ### `idempotency_key_reused` [Section titled “idempotency\_key\_reused”](#idempotency_key_reused) `409`. This `Idempotency-Key` was already used for a different request. Use a new key for a new operation, and reuse a key only to retry the same request (see [Idempotency](/docs/idempotency/)). ### `idempotency_in_progress` [Section titled “idempotency\_in\_progress”](#idempotency_in_progress) `409`. A request with the same `Idempotency-Key` is still running. Wait the number of seconds in `Retry-After`, then retry with the same key. ### `internal_error` [Section titled “internal\_error”](#internal_error) `500`. Something failed on our side. Retry with backoff, and if it keeps happening, contact support with the `requestId`. ## API rate limits [Section titled “API rate limits”](#api-rate-limits) Every authenticated request counts against a quota for its credential: per API key on the workspace endpoints, per token on `/v1/user/*`. When the quota runs out, the API answers: ```http HTTP/1.1 429 Too Many Requests Retry-After: 60 ``` with `code: "rate_limited"`. Wait for `Retry-After`, then retry with backoff. Because limits are per credential, one busy integration doesn’t slow down another key’s requests. # Give AI agents a budget > Cap what each agent in your product can spend on model calls and tools, stop it cleanly at the cap, and let a person approve more. Say your product runs AI agents on your customers’ behalf, and each agent calls models and paid tools in a loop. A loop that goes wrong can spend a lot before anyone notices. Give every agent a budget in money, and have the agent check it before every step. ## 1. Create the budget [Section titled “1. Create the budget”](#1-create-the-budget) Count in cents. $5 a day for every agent: ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Agent daily budget","subjectKind":"agent","action":"*", "amount":500,"unit":"cents","window":"day","failMode":"closed"}' ``` `action` is `*`, so model calls and tool calls all draw from the same budget. The fail mode is `closed`: if Limitry can’t be reached, the agent stops rather than spending without a count. For a budget that never resets, such as $20 for one long task, use the window `total` instead of `day`. See [Pools, balances and budgets](/docs/pools-and-budgets/). ## 2. Check before every step [Section titled “2. Check before every step”](#2-check-before-every-step) Hold an estimate before each model or tool call, and settle what it actually cost: ```ts import { createClient, createLimits, LimitExceeded } from "@limitry/sdk"; const limits = createLimits( createClient({ token: process.env.LIMITRY_API_KEY! }), ); async function step(agentId: string, action: string, run: () => Promise) { return limits.hold( { subject: { kind: "agent", id: agentId }, action, estimate: 10 }, async (held) => { const result = await run(); held.cost(result.costCents); return result; }, ); } export async function runAgent(agentId: string, task: Task) { try { while (!task.done) { await step(agentId, "llm", () => nextModelCall(task)); } } catch (err) { if (err instanceof LimitExceeded) { return pause(task, "This agent reached its budget for today."); } throw err; } } ``` When the budget is used up, Limitry refuses the next step before it runs, so the agent stops between steps rather than in the middle of one. ## 3. Let a person approve more [Section titled “3. Let a person approve more”](#3-let-a-person-approve-more) With a `total` budget, the agent can ask for more credits through Limitry’s MCP server, and a workspace owner or admin approves or declines the request in the app. See [Approvals](/docs/approvals/). ## 4. Know when an agent stops [Section titled “4. Know when an agent stops”](#4-know-when-an-agent-stops) The first time an agent hits its budget in a window, your webhook endpoints receive [`limit.exceeded`](/docs/usage/#webhook-limitexceeded), with the agent’s id, so you can tell the person who started it. Related: [Reservations](/docs/reservations/), [Limit AI spend per user](/docs/guides/limit-ai-spend/). # Rate-limit your API's customers > Give each of your customers' API keys a fair share, with a burst limit and an hourly limit, and answer 429 with Retry-After. Say your product has its own public API, and each customer calls it with their own API key. You want each key to make at most 60 requests in any minute and 1,000 in any hour, so one customer’s script can’t slow everyone else down. ## 1. Create two limits [Section titled “1. Create two limits”](#1-create-two-limits) One over any minute for bursts, and one over any hour, both on the subject kind `apikey`: ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"API burst","subjectKind":"apikey","action":"api-request", "amount":60,"window":"minute","rolling":true}' curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"API per hour","subjectKind":"apikey","action":"api-request", "amount":1000,"window":"hour","rolling":true}' ``` Rolling windows count the last 60 seconds or 60 minutes at the moment of each request, so there’s no burst allowed at the top of the hour. ## 2. Check each request [Section titled “2. Check each request”](#2-check-each-request) Check before handling the request, using your customer’s key id as the subject. Both limits apply, and the request is allowed only if it fits both: ```ts import { createClient, createLimits, LimitExceeded } from "@limitry/sdk"; const limits = createLimits( createClient({ token: process.env.LIMITRY_API_KEY! }), ); export async function handle(request: Request, customerKeyId: string) { try { return await limits.guard( { subject: { kind: "apikey", id: customerKeyId }, action: "api-request" }, () => route(request), ); } catch (err) { if (err instanceof LimitExceeded) { return new Response("Too many requests", { status: 429, headers: { "Retry-After": String(err.retryAfterSeconds ?? 60) }, }); } throw err; } } ``` Both limits keep the default fail mode, `open`: if Limitry can’t be reached, requests go through rather than your API going down with it. ## 3. A bigger allowance for some customers [Section titled “3. A bigger allowance for some customers”](#3-a-bigger-allowance-for-some-customers) Every limit that matches a request applies, so a limit for one key (with `subjectId` set) can lower that key’s allowance but not raise it. To give customers on a bigger plan more, use a subject kind per plan, such as `apikey-pro` with its own limits, and send that kind for their keys. ## 4. See who is hitting it [Section titled “4. See who is hitting it”](#4-see-who-is-hitting-it) The [Usage](/docs/usage/) page shows, for each limit, the keys using it right now and how many of their requests were turned away. Related: [Rolling windows and concurrency](/docs/rolling-and-concurrency/), [Checks](/docs/checks/). # Cap jobs running at once > Let each workspace run at most a few heavy jobs at the same time, and free the slot even when a job crashes. Say exports are heavy, and you want each customer workspace to run at most 3 at once. The 4th should wait or be refused, and a crashed export mustn’t keep its slot forever. ## 1. Create the limit [Section titled “1. Create the limit”](#1-create-the-limit) ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Exports at once","subjectKind":"workspace","action":"export", "amount":3,"window":"concurrent"}' ``` ## 2. Take a slot for the job [Section titled “2. Take a slot for the job”](#2-take-a-slot-for-the-job) `hold` takes a slot before the job starts and gives it back when the job ends, whether it succeeds or throws: ```ts import { LimitExceeded } from "@limitry/sdk"; try { await limits.hold( { subject: { kind: "workspace", id: workspaceId }, action: "export", ttlSeconds: 600, }, async (held) => { for (const chunk of chunks) { await exportChunk(chunk); await held.extend(); // still running: keep the slot another 600 s } }, ); } catch (err) { if (err instanceof LimitExceeded) { // All 3 slots are taken. Try again in err.retryAfterSeconds. return requeue(job, err.retryAfterSeconds); } throw err; } ``` ## 3. If the worker dies [Section titled “3. If the worker dies”](#3-if-the-worker-dies) A slot is a reservation, and it lasts `ttlSeconds`. A job that’s still running extends it as it goes. A worker that crashes stops extending, and the slot frees itself within `ttlSeconds`. Pick a `ttlSeconds` comfortably longer than the gap between your `extend` calls. ## Combine with a daily limit [Section titled “Combine with a daily limit”](#combine-with-a-daily-limit) A concurrency limit doesn’t count use. To also cap exports per day, add a second limit on the same action with the window `day`. The `hold` above then has to fit both. Related: [Rolling windows and concurrency](/docs/rolling-and-concurrency/), [Reservations](/docs/reservations/). # Limit AI spend per user > Give each user a monthly model budget in dollars, charge what each call actually cost, and stop at the budget. Say each user on your free plan gets $2 of model calls a month. You don’t know what a call costs until it’s done, because it depends on the tokens used. So you hold an estimate before the call and settle the real cost after it. ## 1. Create the budget [Section titled “1. Create the budget”](#1-create-the-budget) Count in cents. $2 is 200: ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Free model budget","subjectKind":"user","action":"llm", "amount":200,"unit":"cents","window":"month","failMode":"closed"}' ``` The fail mode is `closed`: if Limitry can’t be reached, refuse the call rather than spend money nobody is counting. Choose `open` if you’d rather keep answering users. ## 2. Hold, call, settle [Section titled “2. Hold, call, settle”](#2-hold-call-settle) Estimate on the high side. Whatever the call doesn’t use, Limitry gives back straight away. ```ts import { createClient, createLimits, LimitExceeded } from "@limitry/sdk"; const limits = createLimits( createClient({ token: process.env.LIMITRY_API_KEY! }), ); // Your prices, in cents per million tokens. const priceCents = (usage: { inputTokens: number; outputTokens: number }) => Math.ceil((usage.inputTokens * 300 + usage.outputTokens * 1500) / 1_000_000); export async function answer(userId: string, prompt: string) { try { return await limits.hold( { subject: { kind: "user", id: userId }, action: "llm", estimate: 5 }, async (held) => { const res = await llm(prompt); held.cost(priceCents(res.usage)); return res.text; }, ); } catch (err) { if (err instanceof LimitExceeded) { return "You've used this month's AI allowance. It resets on the 1st."; } throw err; } } ``` If the model call throws, `hold` releases the estimate, so a failed call costs the user nothing. ## 3. Different budgets per plan [Section titled “3. Different budgets per plan”](#3-different-budgets-per-plan) Give paying users more by putting a second limit on a different subject kind, and sending that kind for them. For example, a `pro-user` limit of 2,000 cents a month, and `{ kind: "pro-user", id }` as the subject for Pro users. ## 4. Know when someone runs out [Section titled “4. Know when someone runs out”](#4-know-when-someone-runs-out) Add a webhook endpoint under **Settings → Webhooks**. The first time a user hits the budget in a month, you receive [`limit.exceeded`](/docs/usage/#webhook-limitexceeded), which you can use to show an upgrade offer. Related: [Reservations](/docs/reservations/), [Pools, balances and budgets](/docs/pools-and-budgets/). # Sell prepaid credits > Let customers buy credits, spend them across their team, and top up when they run out. Say customers buy credits in packs, and every member of a customer’s team spends from the same balance. That’s a balance limit on the team, with grants when they buy. ## 1. Create the balance [Section titled “1. Create the balance”](#1-create-the-balance) The window `total` never resets. The amount is what every team starts with, here 100 free credits: ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Credits","subjectKind":"team","action":"*", "amount":100,"unit":"credits","window":"total"}' ``` ## 2. Add credits when a customer pays [Section titled “2. Add credits when a customer pays”](#2-add-credits-when-a-customer-pays) In your payment webhook, grant what they bought. Use the payment’s id as the idempotency key, so a redelivered webhook can’t add credits twice: ```bash curl -X POST https://api.limitry.com/v1/limits/$LIMIT_ID/grants \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: payment_8231" \ -d '{ "subject": { "kind": "team", "id": "t_42" }, "amount": 5000, "note": "5,000-credit pack" }' ``` ## 3. Spend from the team’s balance [Section titled “3. Spend from the team’s balance”](#3-spend-from-the-teams-balance) Check as the user, and name the team as the shared subject. Any limits you set on the user still apply as well: ```ts const report = await limits.guard( { subject: { kind: "user", id: user.id }, shared: [{ kind: "team", id: user.teamId }], action: "report", cost: 25, }, () => buildReport(params), ); ``` However many teammates spend at the same moment, the balance never goes below zero. ## 4. When the balance runs out [Section titled “4. When the balance runs out”](#4-when-the-balance-runs-out) The check says no, with `retryAfterSeconds: null`, because waiting won’t help. Use that to offer a top-up rather than “try again later”. To show the balance in your product, read `remaining` from any check’s answer, or make a `preview` check, which counts nothing. Related: [Pools, balances and budgets](/docs/pools-and-budgets/), [Idempotency](/docs/idempotency/). # How Limitry works > Where Limitry decides, why counts stay exact under load, what happens if Limitry is down, and what data it keeps. ## Where Limitry decides a check [Section titled “Where Limitry decides a check”](#where-limitry-decides-a-check) Your check reaches Limitry at the location nearest the server that sent it. Each subject’s counts live in one place, so Limitry decides every check for that subject against the same numbers. When a check also draws from a shared pool (a team’s, say), Limitry first holds the cost on every subject, and charges them all only if every limit fits. Two things follow from this: * **Counts are exact.** Two requests racing for the last unit can’t both get it, and a pool shared by many users is never overspent. * **Changes take up to a minute.** A new or edited limit reaches checks within a minute. A revoked API key can keep working for checks for up to 60 seconds. ## When Limitry can’t be reached [Section titled “When Limitry can’t be reached”](#when-limitry-cant-be-reached) Every limit has a fail mode, which you choose per limit: * `open`, the default: your code goes ahead as if allowed. * `closed`: your code refuses, as if denied. The SDK applies it for you. If Limitry doesn’t answer within 1.5 seconds, `guard` and `hold` follow the fail mode of the limits they last saw. Choose `closed` where overspending would hurt more than refusing, such as an agent’s spending budget. ## Retries [Section titled “Retries”](#retries) Send an `Idempotency-Key` header with a check. For 24 hours, the same key gets the same answer and counts nothing again. Settling a reservation is safe to retry too. ## Work that crashes [Section titled “Work that crashes”](#work-that-crashes) A [reservation](/docs/reservations/) that nobody settles gives itself back when it expires, after 5 minutes by default. A worker that crashes mid-job never leaves a user blocked or a concurrency slot taken. ## What Limitry stores [Section titled “What Limitry stores”](#what-limitry-stores) * Your limits, and every change to them in the workspace’s audit log. * Counts per limit, per subject, per hour, kept as long as your plan’s usage history: 7 days on Free, 90 days on Pro, 13 months on Growth. * The subject kinds and ids you send. Limitry attaches no meaning to them, so send your own opaque ids (`usr_123`), not names or email addresses. Limitry never sees the work itself: no prompts, model output or files. A check carries a subject, an action name and a number. ## Access [Section titled “Access”](#access) API keys belong to a workspace and carry scopes: `checks:write` to check, `limits:read` and `limits:write` to read and manage limits. Give each service the narrowest key it needs. Personal access tokens act as the person who created them. See [Authentication](/docs/authentication/). # Idempotency > Retry writes safely with the Idempotency-Key header. A request can time out after the API has already acted on it. To retry a write without doing it twice, send an `Idempotency-Key` header: any unique string of up to 255 printable ASCII characters, such as a UUID you generate for each operation. ```bash curl -X POST \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Idempotency-Key: 6f1c2e9a-8b3d-4c47-9a51-0d2f7e4b8c13" \ -H "Content-Type: application/json" \ -d '{ … }' \ https://api.limitry.com/v1/… ``` The first request with a key runs normally. A retry with the same key and the same request gets the first response again, with the same status and body, and doesn’t run the operation a second time. A replayed response carries the header `Idempotency-Replayed: true`. ## The rules [Section titled “The rules”](#the-rules) * **Writes only.** The header applies to `POST`, `PUT`, `PATCH` and `DELETE`. Reads are safe to repeat anyway, and ignore it. * **One key per request.** Using a key again for a different request, with a different endpoint or body, fails with a `409` and the code `idempotency_key_reused`. Generate a new key for each operation, and reuse it only to retry that operation. * **A retry during the first request waits.** If the first request is still running, the retry gets `409 idempotency_in_progress` and a `Retry-After` header. Retry after that many seconds. * **You can retry server errors.** A `5xx` response isn’t stored, so a retry with the same key runs the request again. Any other response, a success or a `4xx`, is what every retry gets back. * **Per credential, for 24 hours.** A key belongs to the API key or token that sent it, so two integrations never collide, and it’s kept for 24 hours. Sending a key is optional. We recommend it on every write. # Limits > The rules checks enforce, how much of an action each subject may use per window. A limit says how much of an action a subject may use per window. For example: every `user` may `generate-image` 10 times per day. Subjects are whoever acts in your product, such as your users, teams, API keys or AI agents. You choose the labels, and Limitry only matches them. | Field | Meaning | | ------------- | ------------------------------------------------------------------------------------------------------- | | `name` | Your label for the limit. Unique in the workspace. | | `subjectKind` | Which subjects it applies to: `user`, `apikey`, `agent`, or any kind you use. | | `subjectId` | One subject, or empty for every subject of that kind. | | `action` | The action it limits, or `*` for every action. | | `amount` | How many units per window. | | `unit` | What the amount counts (`images`, `credits`, `cents`). Used for display. | | `window` | `minute`, `hour`, `day`, `month`, `total` or `concurrent`. See below. | | `rolling` | With `minute`, `hour` or `day`: count the last 60 seconds, minutes or 24 hours. Default `false`. | | `failMode` | What your code should do when Limitry can’t be reached: `open` (allow, the default) or `closed` (deny). | ## Windows [Section titled “Windows”](#windows) * `minute`, `hour`, `day` and `month` reset at UTC boundaries, unless the limit is [rolling](/docs/rolling-and-concurrency/). * `total` is a balance that never resets, for prepaid credits. See [Pools, balances and budgets](/docs/pools-and-budgets/). * `concurrent` allows at most `amount` at once. See [Rolling windows and concurrency](/docs/rolling-and-concurrency/). When several limits match a subject and action, all of them apply. ## Managing limits [Section titled “Managing limits”](#managing-limits) In the app, limits are under **Settings → Limits**. Owners and admins can change them, and every member can see them. Through the API, use a key or token with the `limits:read` or `limits:write` scope: ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Daily images","subjectKind":"user","action":"generate-image","amount":10,"unit":"images","window":"day"}' ``` The same operations are on the [command line](/docs/cli/) (`limitry limits create`) and the [MCP server](/docs/mcp/). The workspace’s audit log records every change, and a change reaches [checks](/docs/checks/) within a minute. # MCP server > Connect Claude or any MCP client to your workspace, what it can do there, and how you stay in control. AI agents that support the Model Context Protocol (MCP), such as Claude and many other assistants and coding agents, can work in your workspace through our MCP server: ```text https://mcp.limitry.com/mcp ``` ## Connect [Section titled “Connect”](#connect) Add the address above as a remote MCP server. In Claude, that’s **Settings → Connectors → Add custom connector**. The first time, your browser opens: 1. Sign in, if you aren’t already. 2. Choose the **one workspace** the agent will work in. 3. Choose what it may do. Everything that can change data starts unticked, so tick only what the agent needs. 4. Click **Allow**. There’s nothing to copy and no key to store. The agent gets its own sign-in, valid only for that workspace and that access. ## What the agent can do [Section titled “What the agent can do”](#what-the-agent-can-do) The agent’s tools are the [API](https://api.limitry.com/v1/docs) operations you allowed, such as reading the workspace or listing webhook deliveries. It always acts as **you**, with your role in the workspace: if you couldn’t do something in the app, the agent can’t either. Everything it changes appears in the workspace’s audit log under your name. Some actions are never available to agents, even with full access: anything that sends your workspace’s data somewhere new or gives out access, such as creating webhook endpoints or credentials. The agent can **ask** for them instead. You approve or deny in the app, and if the action creates a secret, the app shows it to you, never to the agent. See [Approvals](/docs/approvals/). ## Staying in control [Section titled “Staying in control”](#staying-in-control) * **Account → Connected apps** lists every agent you’ve connected. Disconnect one and it stops working at once. * The workspace’s **audit log** records what each agent did. * If you’re removed from the workspace, your agents lose access too. ## For developers [Section titled “For developers”](#for-developers) The server follows the MCP authorization specification. A request without a token gets a `401` pointing to `/.well-known/oauth-protected-resource/mcp`, which names the authorization server: `https://api.limitry.com/auth`, with dynamic client registration and PKCE. The authorization server issues tokens for the MCP server only. For your own scripts, the [command line](/docs/cli/) or the [SDK](/docs/sdk/) are simpler. See [Use with AI agents](/docs/agents/) for the agent skill. # Pagination > How list endpoints return results a page at a time. List endpoints return results newest first, a page at a time. Each page comes with a cursor that points to the next one. ```bash curl \ -H "Authorization: Bearer YOUR_API_KEY" \ "https://api.limitry.com/v1/workspace/audit-events?limit=25" ``` ```json { "items": [{ "id": "…", "action": "member.invited", "createdAt": "…" }], "nextCursor": "eyJ0IjoiMjAyNi0wOC0uLi4ifQ" } ``` To get the next page, send the cursor back as it is: ```bash curl \ -H "Authorization: Bearer YOUR_API_KEY" \ "https://api.limitry.com/v1/workspace/audit-events?limit=25&cursor=eyJ0IjoiMjAyNi0wOC0uLi4ifQ" ``` Keep going until `nextCursor` is `null`. That’s the last page. ## Good to know [Section titled “Good to know”](#good-to-know) * `limit` is from 1 to 100, and 25 by default. * Don’t build or change a cursor yourself. A cursor that isn’t one we gave you returns `invalid_request`. * Rows added while you’re paging don’t cause skipped or repeated items. * Items may gain new fields over time. We won’t rename existing fields. # Plans and limits > What each plan includes, what counts toward it, and why your product keeps working when you reach it. Every plan has every feature: pools, credits, budgets, reservations, rolling and concurrent limits, webhooks, the SDK and the MCP server. Plans differ in how much you can use. | | Free | Pro | Growth | Enterprise | | ----------------------- | ------- | ----------- | ------------ | ---------- | | Price | $0 | $39 / month | $299 / month | Talk to us | | Allowed checks / month | 100,000 | 1 million | 25 million | Custom | | Active subjects / month | 1,000 | 10,000 | 100,000 | Custom | | Usage history | 7 days | 90 days | 13 months | Custom | | Members | 3 | 10 | 25 | Custom | ## What counts [Section titled “What counts”](#what-counts) * **Allowed checks.** Denied checks never count, so a burst of abuse against your limits doesn’t use up your plan. A check that draws from a user and a team pool counts once. * **Active subjects:** the different users, keys, agents or teams with at least one allowed check this month, each counted once. You can see your usage in **Settings → Billing**, or with `GET /v1/workspace/usage`. It’s about a minute behind live checks. ## When you reach your plan’s limit [Section titled “When you reach your plan’s limit”](#when-you-reach-your-plans-limit) Your product keeps working. * At 80% and 100% of a monthly allowance, your workspace’s owners get an email and a notification, and the app shows a notice. * You get 10% grace. Past that, checks answer `402 plan_limit_reached` instead of a decision. The SDK treats that the same as Limitry being unreachable, so each limit’s fail mode applies. With `open`, the default, your product keeps working without metering. With `closed`, it refuses. * Upgrading lifts it within a minute. Otherwise it resets on the 1st of the month (UTC). If you call the API directly, handle a 402 the way you’d handle a timeout: apply your fail mode. The SDK’s `guard` and `hold` do this for you. # Pools, balances and budgets > Share an allowance across a team, sell prepaid credits, and cap what an agent can spend. Pools, balances and budgets aren’t separate features. Each one is an ordinary limit, used in a particular way. ## Pools: an allowance a team shares [Section titled “Pools: an allowance a team shares”](#pools-an-allowance-a-team-shares) Put the limit on the shared subject, for example 10,000 AI calls a month for each `team`. Then name the team on each user’s check: ```json { "subject": { "kind": "user", "id": "u_7" }, "shared": [{ "kind": "team", "id": "t_42" }], "action": "ai-call" } ``` The check must fit every limit on the user and every limit on the team. It charges all of them, or none. However many of the team’s users check at the same moment, the pool is never overspent. You can name up to 4 shared subjects, such as a team, its organization and a project. ## Balances: prepaid credits [Section titled “Balances: prepaid credits”](#balances-prepaid-credits) Set the window to `total`. The amount is the starting balance, and it never resets. When a customer buys credits, add them with a grant: ```bash curl -X POST https://api.limitry.com/v1/limits/$LIMIT_ID/grants \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "subject": { "kind": "team", "id": "t_42" }, "amount": 5000, "note": "October top-up" }' ``` In the app, open the balance under **Settings → Limits** and choose **Credits**. Limitry keeps every grant (`GET /v1/limits/{id}/grants`). When a balance runs out, the check’s `retryAfterSeconds` is `null`, because waiting won’t help and only more credits will. Owners and admins can add credits. An agent can ask to, and a person approves the request (`limit.grant`). ## Budgets: a spending cap for an agent [Section titled “Budgets: a spending cap for an agent”](#budgets-a-spending-cap-for-an-agent) A budget is a limit counted in money. For example, $50 a day for every agent: | Field | Value | | ------- | ----------------------------- | | Subject | `agent` (every agent, or one) | | Action | `*` | | Unit | `cents` | | Amount | `5000` | | Window | `day`, or `total` for one-off | Each check passes the call’s price in cents as `cost`. When you don’t know the price up front, [reserve](/docs/reservations/) an estimate and commit what the call actually cost. Once the budget is used up, every check says no until the window resets, or for a balance, until you add more. The [Limit AI spend](/docs/guides/limit-ai-spend/) guide puts this together end to end. # Quickstart > Create a limit and make your first check, in about five minutes. By the end of this page you’ll have a limit of 10 images per user per day, and a check that counts against it. ## 1. Create a workspace [Section titled “1. Create a workspace”](#1-create-a-workspace) [Sign up](https://app.limitry.com/sign-up), verify your email and create a workspace. Limits, keys and usage all belong to a workspace. ## 2. Create an API key [Section titled “2. Create an API key”](#2-create-an-api-key) In the app, go to **Settings → API Keys** and create a key with the `limits:write` and `checks:write` scopes. Copy the secret now, because it’s shown only once. Keep it in your secret manager or environment, not in code. ```bash export LIMITRY_API_KEY="…" ``` ## 3. Create a limit [Section titled “3. Create a limit”](#3-create-a-limit) ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Daily images","subjectKind":"user","action":"generate-image", "amount":10,"unit":"images","window":"day"}' ``` This applies to every subject of kind `user`. You can also create it in the app, under **Settings → Limits**. ## 4. Check before the work [Section titled “4. Check before the work”](#4-check-before-the-work) Ask whether user `usr_123` may generate one image: ```bash curl -X POST https://api.limitry.com/v1/checks \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"subject":{"kind":"user","id":"usr_123"},"action":"generate-image"}' ``` ```json { "id": "chk_…", "allowed": true, "reason": null, "remaining": [ { "limit": "lim_…", "name": "Daily images", "unit": "images", "window": "day", "remaining": 9, "resetsAt": "2026-10-06T00:00:00.000Z" } ], "retryAfterSeconds": null } ``` Run it ten more times and the eleventh answer is `"allowed": false`, with `retryAfterSeconds` telling you when the user can try again. A “no” is a normal answer (HTTP 200), not an error. ## 5. Use it from your code [Section titled “5. Use it from your code”](#5-use-it-from-your-code) In TypeScript, the [SDK](/docs/sdk/) wraps the check and the work together: ```ts import { createClient, createLimits } from "@limitry/sdk"; const limits = createLimits( createClient({ token: process.env.LIMITRY_API_KEY! }), ); const image = await limits.guard( { subject: { kind: "user", id: user.id }, action: "generate-image" }, () => generateImage(prompt), ); ``` If the user is over the limit, `guard` throws `LimitExceeded` and the work never runs. Next, read [Checks](/docs/checks/) for everything a check can do, or pick a [guide](/docs/guides/limit-ai-spend/) close to what you’re building. # Reservations > Hold an estimate now, and settle the actual amount when the work is done. Some work only knows its cost afterwards, like a model call billed by tokens or a job billed by minutes. A reservation holds an estimate while the work runs, then settles at the actual amount. In TypeScript, the SDK’s `hold` does all of this for you. See [Add Limitry with your agent](/docs/agents-onboarding/). The steps below are what it does through the API. ## 1. Reserve [Section titled “1. Reserve”](#1-reserve) Make a check with `mode: "reserve"` and your estimate as the `cost`: ```bash curl -X POST https://api.limitry.com/v1/checks \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "subject": { "kind": "user", "id": "u_42" }, "action": "ai-call", "cost": 2000, "mode": "reserve" }' ``` It decides like any other check. If it’s allowed, Limitry holds the estimate: other checks for that subject see it as used. The answer includes the reservation: ```json { "allowed": true, "remaining": [{ "limit": "lim_…", "remaining": 8000, "…": "…" }], "reservation": { "id": "rsv_…", "expiresAt": "2026-10-04T12:05:00Z" } } ``` If the answer is no, Limitry holds nothing and there’s no `reservation`. ## 2. Settle [Section titled “2. Settle”](#2-settle) When the work is done, commit what it actually used: ```bash curl -X POST https://api.limitry.com/v1/reservations/rsv_…/commit \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" -d '{ "cost": 1310 }' ``` If it used less than the estimate, Limitry gives the rest back at once. If it used more, Limitry charges the full amount, even past the limit, because the work did happen. The next check then sees the true total. If the work didn’t happen, release the hold instead: `POST /v1/reservations/{id}/release` gives all of it back. Settling is safe to retry: the same commit or release gets the same answer. Releasing a committed reservation, or committing a released one, is refused with `400`. ## Expiry [Section titled “Expiry”](#expiry) A hold lasts `ttlSeconds`, which you set on the check: 300 by default, 3,600 at most. A hold nobody settles gives itself back when it expires, so a crashed worker never keeps a subject blocked. If the work did finish after all, a late commit within 24 hours still charges what it used. For longer work, extend the hold as a heartbeat: `POST /v1/reservations/{id}/extend` with `{ "ttlSeconds": 600 }` makes it expire that long from now. A reservation can also hold a slot of a [concurrent limit](/docs/rolling-and-concurrency/). ## Good to know [Section titled “Good to know”](#good-to-know) * A subject can have at most 1,000 open reservations. Past that, a reserve is a normal “no” with `reason.code: "too_many_reservations"`. * A reserve counts as one check on your plan. Settling is free. * Reservations need the same `checks:write` scope as checks. # Rolling windows and concurrency > Limits like "100 in any 24 hours" and "at most 5 at once", which fixed windows can't express. ## Rolling windows [Section titled “Rolling windows”](#rolling-windows) By default, a `minute`, `hour` or `day` limit resets at the UTC boundary. A user could make 100 calls at 23:59 and 100 more at 00:01, and both bursts would be allowed. Set `rolling: true` and the limit counts the last 60 seconds, 60 minutes or 24 hours at the moment of each check instead. ```bash curl -X POST https://api.limitry.com/v1/limits \ -H "Authorization: Bearer $LIMITRY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name":"Calls in any 24 hours","subjectKind":"user","action":"call", "amount":100,"window":"day","rolling":true}' ``` Limitry counts use in 60 slices of the window: a second for a minute, a minute for an hour, 24 minutes for a day. The oldest slice counts in full, so a rolling limit never allows more than its amount, and may refuse up to one slice early. In a check’s answer, `resetsAt` (and `retryAfterSeconds` when refused) is when enough use has left the window. `month`, `total` and `concurrent` limits can’t be rolling. In the app, choose **In any 60 seconds**, **In any 60 minutes** or **In any 24 hours** as the window. ## Concurrency [Section titled “Concurrency”](#concurrency) A limit with the window `concurrent` caps work in progress: `amount` is how many at once. Each [reservation](/docs/reservations/) takes one slot, whatever its `cost`, and gives it back when it’s committed, released or expires. The `cost` still counts against the subject’s other limits. ```bash # At most 3 exports at once per user curl -X POST https://api.limitry.com/v1/limits ... \ -d '{"name":"Exports at once","subjectKind":"user","action":"export", "amount":3,"window":"concurrent"}' # Start one: the reservation holds a slot curl -X POST https://api.limitry.com/v1/checks ... \ -d '{"subject":{"kind":"user","id":"u_42"},"action":"export","mode":"reserve"}' # Finished, or failed: give the slot back curl -X POST https://api.limitry.com/v1/reservations/$ID/release ... ``` * A `consume` or `preview` check against a concurrent limit is allowed while a slot is free, but takes none. Instant work holds nothing. * When every slot is taken, `retryAfterSeconds` is when the earliest hold expires. A release can free one sooner. * Settle in a `finally` block. A hold you never settle keeps its slot until it expires. ### Long-running work [Section titled “Long-running work”](#long-running-work) A hold lasts `ttlSeconds`: 300 by default, 3,600 at most. For work that runs longer, extend the hold as a heartbeat: ```bash curl -X POST https://api.limitry.com/v1/reservations/$ID/extend ... \ -d '{"ttlSeconds":600}' ``` If your worker dies, it stops extending, and the slot frees itself. A limit can’t change to or from `concurrent`, because one counts use and the other counts open holds. Create a new limit instead. ## In reports [Section titled “In reports”](#in-reports) Rolling limits report like any other. On the Usage page, **Right now** shows the last window, to the hour. For a concurrent limit, the chart shows how much work the limit let start each day, and **Right now** shows the slots each subject holds, among subjects checked in the last hour. For both, Limitry sends `limit.exceeded` at most once an hour per subject. # TypeScript SDK > Call the API from TypeScript, with types generated from the API itself. The SDK is a small typed client. Every path, parameter and response takes its type from the same definitions that validate requests, so it always matches the [API reference](https://api.limitry.com/v1/docs). ## Install [Section titled “Install”](#install) ```bash npm install @limitry/sdk ``` ## Make a call [Section titled “Make a call”](#make-a-call) ```ts import { createClient, unwrap } from "@limitry/sdk"; const api = createClient({ token: process.env.API_KEY! }); const endpoints = unwrap(await api.GET("/workspace/webhook-endpoints")); for (const endpoint of endpoints.items) console.log(endpoint.url); ``` * With an **API key**, every call acts in the key’s workspace. * With a **personal access token**, name the workspace: `createClient({ token, workspace: "acme" })`. * `unwrap` returns the data, or throws an `ApiError` with the [error envelope](/docs/errors/)’s `code`, `message`, `status` and `requestId`. * Writes can send an [idempotency key](/docs/idempotency/): `api.POST("/workspace/webhook-endpoints", { body, headers: { "idempotency-key": id } })`. ## Limits in one wrapper [Section titled “Limits in one wrapper”](#limits-in-one-wrapper) `createLimits(client)` gives you `guard`, which checks and then runs the work, and `hold`, which reserves an estimate, runs the work and commits what it actually used. See [Add Limitry with your agent](/docs/agents-onboarding/) for both, and for how fail modes apply when Limitry can’t be reached. # Usage and reports > See how each limit is used over time and who is closest to it now, and get told when a limit is hit. Limitry counts every check per limit, per subject, per hour. Reports in the app and the API are at most about a minute behind. ## In the app [Section titled “In the app”](#in-the-app) The **Usage** page shows, for the limit you pick: * **Last 30 days:** units used per day (UTC), across all subjects. * **Right now:** the subjects in the current window, most used first, with what each has used and has left. Anyone who was turned away shows as **Over**, with how many checks were denied. ## Through the API [Section titled “Through the API”](#through-the-api) With a key or token that has `limits:read`: ```bash # A limit's use over time, by day (the default) or by hour curl "https://api.limitry.com/v1/usage?limitId=$LIMIT_ID&granularity=day" \ -H "Authorization: Bearer $LIMITRY_API_KEY" # Who is using it in the current window curl "https://api.limitry.com/v1/usage/subjects?limitId=$LIMIT_ID" \ -H "Authorization: Bearer $LIMITRY_API_KEY" ``` `/v1/usage` returns every bucket in the range, zeros included, each with `used`, `allowed` and `denied`. Add `subjectKind` and `subjectId` for one subject, and `from` and `to` (ISO times) for a different range. The default range is the last 30 days. `used` is net: Limitry gives back the unused part of a reservation, and all of a released one. For a balance (window `total`), what’s left includes the credits you’ve added. How far back you can look depends on your plan: 7 days on Free, 90 days on Pro, 13 months on Growth. ## Webhook: `limit.exceeded` [Section titled “Webhook: limit.exceeded”](#webhook-limitexceeded) The first time a limit turns a subject away in a window, your webhook endpoints receive `limit.exceeded`: ```json { "type": "limit.exceeded", "data": { "limit": "lim_…", "subject": { "kind": "user", "id": "u_42" }, "windowStart": "2026-10-05T00:00:00.000Z", "resetsAt": "2026-10-06T00:00:00.000Z" } } ``` It’s sent once per subject, limit and window, not on every denied check. `resetsAt` is `null` for a balance. A budget is a limit too, so an agent running out of budget sends this event as well. Add an endpoint under **Settings → Webhooks**. # Versioning > What can change in the API, what never changes, and how we announce deprecations. The API’s version is in its path: `/v1`. Within a version, we only make changes that don’t break a correctly written integration. The same promise covers the command line and the tools agents use, because their names come from the API’s operations. ## Changes we may make at any time [Section titled “Changes we may make at any time”](#changes-we-may-make-at-any-time) * New endpoints. * New optional request parameters and fields. * New fields in responses. * New error codes, webhook event types and scopes. * New commands, and new tools for agents. Write your integration so these don’t break it: * Ignore response fields you don’t recognise. * Treat an error `code` you don’t recognise as `internal_error`, and branch on `code`, never on `message`. * Treat cursors and ids as opaque strings. ## Changes we never make within `/v1` [Section titled “Changes we never make within /v1”](#changes-we-never-make-within-v1) * Removing or renaming an endpoint, a field or an operation, and so a command or an agent’s tool. * Changing the type or meaning of an existing field. * Making an optional parameter required. * Changing how requests authenticate, or which error a given failure returns. A change like these only ships in a new version, `/v2`, which runs alongside `/v1` while integrations move over. ## Deprecations [Section titled “Deprecations”](#deprecations) Before we retire anything in `/v1`, we announce it in the changelog and mark it **deprecated** in the [API reference](https://api.limitry.com/v1/docs), with what to use instead. Deprecated endpoints keep working for as long as `/v1` exists. # Webhooks > Receive workspace events at your endpoint, verify signatures, and handle retries. Webhooks send workspace events to an HTTPS endpoint you host, as they happen. Workspace owners manage endpoints in the app under **Settings → Webhooks**: add a URL, copy the signing secret (it’s shown once), and events start arriving. ## Managing endpoints through the API [Section titled “Managing endpoints through the API”](#managing-endpoints-through-the-api) Software can manage endpoints too, with a credential that has the `workspace.webhooks:write` scope. Reading them needs `workspace.webhooks:read`, and a personal access token must also belong to a workspace owner. ```bash curl -X POST \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Idempotency-Key: $(uuidgen)" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com/hooks" }' \ https://api.limitry.com/v1/workspace/webhook-endpoints ``` The response includes the endpoint’s signing secret, once. The API can also list endpoints, re-enable one that we switched off after repeated failures, send a test event, and list and redeliver deliveries. See the [API reference](https://api.limitry.com/v1/docs). AI agents connected to the workspace can read endpoints and delivery history, send test events and redeliver. They can’t add, remove or re-enable endpoints, because that changes where your workspace’s data goes, so a person does it. ## A delivery [Section titled “A delivery”](#a-delivery) Each delivery is an HTTP `POST` with a JSON body: ```json { "id": "evt_5f0c…", "type": "workspace.member.joined", "createdAt": "2026-08-20T12:00:00.000Z", "data": { "userId": "…", "email": "casey@example.com" } } ``` and three signature headers: ```http webhook-id: del_9a1b… webhook-timestamp: 1755691200 webhook-signature: v1,MEQCIB… ``` `webhook-id` identifies the delivery, and stays the same when it’s retried, so use it to ignore duplicates. The `id` in the body identifies the event. If you have several endpoints, each gets its own delivery of the same event; to process an event once across endpoints, deduplicate by its `id`. ## Verifying signatures [Section titled “Verifying signatures”](#verifying-signatures) We sign deliveries with your endpoint’s secret (`whsec_…`) in the same way as [Svix](https://docs.svix.com/receiving/verifying-payloads/how), so any standard Svix library can verify them: ```ts import { Webhook } from "svix"; const wh = new Webhook(process.env.WEBHOOK_SECRET); // Express-style handler; `payload` must be the RAW request body string. app.post("/webhooks", (req, res) => { let event; try { event = wh.verify(req.body, req.headers); } catch { return res.status(400).send("bad signature"); } // handle event… res.status(200).send("ok"); }); ``` To verify by hand: the signature is `v1,` followed by a Base64 HMAC-SHA256 of `` `${webhookId}.${timestamp}.${body}` ``, keyed with the secret after its `whsec_` prefix, Base64-decoded. Always verify the **raw** body, because re-serialising the JSON changes it and breaks the signature. Reject timestamps more than a few minutes old, so nobody can replay an old delivery. ## Respond quickly, process later [Section titled “Respond quickly, process later”](#respond-quickly-process-later) Answer with a `2xx` within 10 seconds. If processing takes longer, acknowledge first and do the work afterwards, because a timeout counts as a failed delivery. ## Retries and failures [Section titled “Retries and failures”](#retries-and-failures) We retry a failed delivery (anything but a `2xx`, or a timeout) automatically, with backoff, up to 6 attempts, with the same `webhook-id`. If an endpoint fails 20 deliveries in a row, we switch it off. Once your endpoint is working again, delete it and add it back, which gives it a new secret. Under **Settings → Webhooks → Deliveries** you can see each delivery’s status and attempts, send a test event (`type: "ping"`), and redeliver any recorded delivery. A redelivery has a new `webhook-id` but the same event `id`, so deduplicating by event still works. ## Events [Section titled “Events”](#events) | Type | Sent when | `data` | | ------------------------- | ------------------------------ | ----------------- | | `workspace.member.joined` | Someone accepts an invitation. | `userId`, `email` | | `ping` | You send a test from Settings. | a test message | Event payloads only ever gain fields, so write parsers that ignore fields they don’t know. `GET /v1/workspace/event-types` lists every event type and what it means. ## Choosing events [Section titled “Choosing events”](#choosing-events) An endpoint receives every event unless you choose otherwise. In **Settings → Webhooks**, pick “Only these” when you add it, or pass `eventTypes` when you create it through the API: ```bash curl -X POST https://api.limitry.com/v1/workspace/webhook-endpoints \ -H "Authorization: Bearer YOUR_API_KEY" -H "content-type: application/json" \ -d '{"url":"https://example.com/hooks","eventTypes":["workspace.member.joined"]}' ``` A test event (`ping`) always reaches the endpoint you’re testing. # Workspace usage > How much of your plan the workspace used this month, and its limits. We count some things per calendar month (UTC). Each count, called a meter, comes with your plan’s limit for it. See them under **Settings → Billing**, or through the API: ```bash curl -H "Authorization: Bearer YOUR_API_KEY" \ https://api.limitry.com/v1/workspace/usage ``` ```json { "periodStart": "2026-10-01T00:00:00.000Z", "periodEnd": "2026-11-01T00:00:00.000Z", "meters": [ { "id": "webhook_deliveries", "description": "Webhook deliveries this month (each event, each endpoint)", "used": 1520, "limit": null } ] } ``` `limit` is `null` when the plan has no limit for that meter. Counts start again at the beginning of each month. The key needs the `workspace:read` scope.