Skip to content

Add Limitry with your agent

Your coding agent can add Limitry for you. It reads your code, proposes limits, asks you before creating any, and wraps the costly calls.

Give it Limitry’s tools, through the MCP server or the command line, and the skill that explains how Limitry works (Use with AI agents).

Add Limitry to this codebase. Find the calls that cost us money or that a
user could abuse (LLM calls, paid APIs, emails, exports). Propose limits
in a table: subject, action, amount, window, fail mode. Wait for my answer
before creating any. Then create the ones I agree to, wrap the code with
the Limitry SDK's guard or hold helpers, handle a refusal where we answer
the user, and add a test for it.

For a call whose cost you know up front, guard checks first and runs the work only if it’s allowed:

import { createClient, createLimits, LimitExceeded } from "@limitry/sdk";
const limits = createLimits(
createClient({ token: process.env.LIMITRY_API_KEY! }),
);
const image = await limits.guard(
{ subject: { kind: "user", id: user.id }, action: "generate-image" },
() => generateImage(prompt),
);

When the cost is only known afterwards, as with tokens, or for “at most N at once”, hold reserves an estimate, runs the work, and commits what it actually used. If the work fails, it releases the hold:

const answer = await limits.hold(
{ subject: { kind: "user", id: user.id }, action: "chat", estimate: 4000 },
async (held) => {
const res = await llm(prompt);
held.cost(res.usage.totalTokens);
return res;
},
);

A refusal throws LimitExceeded, with limit, reason and retryAfterSeconds, for your code to turn into a 429 or a message.

Each limit has a fail mode. The SDK remembers the fail mode from the last answer, and if Limitry doesn’t answer within 1.5 seconds (timeoutMs), it either lets the work run (open, the default) or throws LimitExceeded with the reason unavailable (closed). Your product never hangs waiting for Limitry.