Rate limiting is an architectural pattern that restricts the rate of requests to services or APIs to prevent overload and abuse. It specifies allowed request frequencies, throttling behavior, and response signals. Common deployments include API gateways, CDNs, and in-service controls to ensure predictable availability and fair resource sharing.
Use this profile to understand the building block briefly, place it in the model, and switch to the 360° assessment when needed.
Theoretical construct: explains a term, principle, or mental model.
What organizes, connects, or makes decisions possible.
Rate limiting restricts the request rate to services or APIs to prevent overload and to handle abusive traffic more fairly.
In computer networks, rate limiting was used to control the rate of requests sent or received at network interfaces or protocol components; it was employed to help prevent DoS attacks and limit web scraping. In the web stack, the concept became more interoperable through standardized HTTP signaling: RFC 6585 (April 2012, IETF; M. Nottingham, R. Fielding) defines HTTP 429 “Too Many Requests” for the case where a client sends too many requests within a given time window. The RFC also specifies optional response details and the Retry-After header and states that responses with 429 must not be stored by caches.
Rate limiting acts like a barrier before the real service logic: first you define a counting scope (who/what is counted), a time window, and an upper bound for the allowed request frequency. For each incoming request, the limiter derives the counting key from that scope, updates the count over the time window, and compares it to the configured limit. If the request stays within the limit, it is forwarded to the service. If it exceeds the limit, it is rejected—typically with 429, optionally including Retry-After to guide when the client may try again. In distributed environments, the overall effect depends on how precisely the scope and counting strategy are implemented.
An allowed request frequency is expressed as an upper bound over a defined time window; exceeding it triggers rejection.
Enforcement is performed using a key such as an auth/cookie context or a resource; RFC 6585 explicitly does not mandate how users are identified or how requests are counted.
RFC 6585 defines 429 “Too Many Requests” as the standard signal for “too many requests”; additional details are optional.
A Retry-After header can indicate how long to wait before making another request.
Responses with 429 must not be stored by caches.
More precise rate limiting can require more resources for the rate limiters, creating a trade-off between accuracy and resource footprint.
Rate limiting is useful when endpoint capacity is threatened by high or uneven request rates—commonly at API gateways, fronting components, or via service-internal controls. It supports a predictable client backoff path via 429 and optionally Retry-After. Key limits and trade-offs matter: (1) returning 429 can consume resources under attack or very high volumes, so RFC 6585 notes servers are not required to use 429 in every case; (2) identification and counting behavior are deployment-specific because RFC 6585 leaves them open; (3) higher enforcement precision can increase operational cost, while coarse rules can unnecessarily impact legitimate traffic.
Where this building block is located in the topic model.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.