Skip to content

How Limitry works

Your check reaches Limitry at the location nearest the server that sent it. Each subject’s counts live in one place, so Limitry decides every check for that subject against the same numbers. When a check also draws from a shared pool (a team’s, say), Limitry first holds the cost on every subject, and charges them all only if every limit fits.

Two things follow from this:

  • Counts are exact. Two requests racing for the last unit can’t both get it, and a pool shared by many users is never overspent.
  • Changes take up to a minute. A new or edited limit reaches checks within a minute. A revoked API key can keep working for checks for up to 60 seconds.

Every limit has a fail mode, which you choose per limit:

  • open, the default: your code goes ahead as if allowed.
  • closed: your code refuses, as if denied.

The SDK applies it for you. If Limitry doesn’t answer within 1.5 seconds, guard and hold follow the fail mode of the limits they last saw. Choose closed where overspending would hurt more than refusing, such as an agent’s spending budget.

Send an Idempotency-Key header with a check. For 24 hours, the same key gets the same answer and counts nothing again. Settling a reservation is safe to retry too.

A reservation that nobody settles gives itself back when it expires, after 5 minutes by default. A worker that crashes mid-job never leaves a user blocked or a concurrency slot taken.

  • Your limits, and every change to them in the workspace’s audit log.
  • Counts per limit, per subject, per hour, kept as long as your plan’s usage history: 7 days on Free, 90 days on Pro, 13 months on Growth.
  • The subject kinds and ids you send. Limitry attaches no meaning to them, so send your own opaque ids (usr_123), not names or email addresses.

Limitry never sees the work itself: no prompts, model output or files. A check carries a subject, an action name and a number.

API keys belong to a workspace and carry scopes: checks:write to check, limits:read and limits:write to read and manage limits. Give each service the narrowest key it needs. Personal access tokens act as the person who created them. See Authentication.