Skip to content
Glossary

What is rate limiting?

The cheapest defence against abuse, runaway scripts and a surprise bill from your AI provider.

Plutonapps Engineering1 min read

In short

Rate limiting is a control that caps how many requests a user, API key or IP address can make in a given period, such as 100 per minute, and refuses or delays the rest. It protects an app from abuse, brute-force login attempts, runaway scripts and unexpected bills, and keeps one heavy user from slowing the app for everyone.

Also called: Throttling, Request limits

Why it matters when your prototype goes to production

A prototype trusts its users, because its only users are the team. A public app meets scripts that try thousands of passwords, bots that fill forms, and, if it calls an AI model, anyone who works out that each request costs you money. OWASP lists unrestricted resource consumption among the top ten API security risks.

What to limit

  • Sign-in, sign-up and password reset, by account and by IP address.
  • Anything that sends email or SMS.
  • Every call to a paid API, especially AI models, by user and in total.

Bell shows the AI version. Every model call passes one router that checks an off switch, a budget and a circuit breaker first. Each person has a daily limit in dollars and in calls: at 80% Bell drops to cheaper triage, at 100% it stops, and a global limit sits above both.

Common questions

What is the difference between rate limiting and throttling?

Rate limiting refuses requests over the limit; throttling slows them down. Many systems do both.

More on this: Production architecture & security · All glossary terms

Built something in Lovable you want people to rely on?

We are the engineers who take it the rest of the way — secured, tested, released and supported.