Rate Limiting
Rate limiting caps how quickly a user, tenant, or integration may make requests.
What is Rate Limiting?
Rate limiting controls the number or frequency of requests accepted within a defined interval. SaaS providers use it to protect service reliability, distribute capacity fairly, and enforce published API policies. Limits can apply to a user, tenant, endpoint, or concurrent workload rather than only to total requests.
Example in SaaS
A reporting API allows a customer a hypothetical 100 requests per minute and ten concurrent export jobs. An integration that exceeds either limit receives a response explaining that it should slow down. Its client queues work, backs off, and retries appropriately rather than launching an uncontrolled burst.
What to watch for
A rate limit differs from a monthly usage allowance. One controls speed or concurrency; the other can control total consumption or billing. Stripe’s documentation illustrates multiple limiter types. Explain response codes and retry behavior, and consider fairness at the tenant level so one account cannot exhaust shared capacity. Unlimited monthly requests can still be subject to a reasonable documented rate limit.
Source and further reading
Related terms: Usage-based pricing, API.