Use cases
The same check covers each of these. What changes is the limit you set and the subject you check.
Keep AI costs per user under control
Your product calls a model on each user's behalf, and you pay by the token. A few heavy users, or one script, can cost more than they pay you.
The rule. Each free user gets $2 of model calls a month. Hold an estimate before each call, and charge what it actually cost.
await limits.hold(
{
subject: { kind: "user", id: user.id },
action: "llm",
estimate: 5,
},
async (held) => {
const res = await llm(prompt);
held.cost(priceInCents(res.usage));
return res;
},
);Sell prepaid credits
Customers buy credits in packs, and everyone on their team spends from the same balance. When it runs out, they should be asked to top up, not told to wait.
The rule. Each team has a balance that never resets. Your payment webhook adds credits; every member's request draws from it.
await limits.guard(
{
subject: { kind: "user", id: user.id },
shared: [{ kind: "team", id: user.teamId }],
action: "report",
cost: 25,
},
() => buildReport(params),
);Rate-limit your own API's customers
Your customers call your API with their own keys. One customer's runaway loop shouldn't slow down everyone else.
The rule. Each key gets 60 requests in any minute and 1,000 in any hour. Over either, answer 429 with Retry-After.
await limits.guard(
{
subject: { kind: "apikey", id: keyId },
action: "api-request",
},
() => route(request),
);Give AI agents a budget
Agents in your product call models and paid tools in a loop. A loop that goes wrong can spend a lot before anyone looks.
The rule. Each agent gets $5 a day across every model and tool call. At the cap, it stops between steps, and a person can approve more.
await limits.hold(
{
subject: { kind: "agent", id: agentId },
action: "tool",
estimate: 10,
},
async (held) => {
const result = await callTool(input);
held.cost(result.costCents);
return result;
},
);Cap heavy jobs running at once
Exports, imports and renders are heavy. You want each customer to run only a few at a time, and a crashed job mustn't hold its place forever.
The rule. Each workspace runs at most 3 exports at once. A job takes a slot when it starts and gives it back when it ends, or when its worker stops responding.
await limits.hold(
{
subject: { kind: "workspace", id: workspaceId },
action: "export",
},
() => runExport(job),
);Something else?
If you can say it as "this many, per this, per this long", Limitry can probably enforce it. Tell us what you need, or read how limits work.