Autoscaling automatically adjusts an application's instance count and resource allocation to match real load. It improves availability and cost efficiency by dynamically scaling capacity — in cloud platforms, container orchestrators, or hybrid environments. Policies, metrics and thresholds determine scaling behaviour and safety limits.
Use this profile to understand the building block briefly, place it in the model, and switch to the 360° assessment when needed.
Theoretical construct: explains a term, principle, or mental model.
What organizes, connects, or makes decisions possible.
Autoscaling is the automatic scaling up and down of instances or resources so an application can absorb spikes and avoid excess capacity.
The approach emerged from the operational problem of serving fluctuating demand without permanently overprovisioning capacity. In cloud environments, capacity began to be adjusted automatically from load signals such as CPU, memory, or request rate instead of by manual intervention. Early systems used instance groups; Kubernetes later carried the pattern into pods and other workloads.
Think of autoscaling as a control loop with a safety envelope. Monitoring observes load and health. A policy compares those signals with targets and thresholds. The system then raises or lowers desired capacity: more instances, fewer instances, or larger or smaller resources. Min/max limits, delays, and buffers keep the system from oscillating during brief spikes.
Measurements such as CPU utilization, memory demand, request rate, or queue length provide the signal for scaling decisions.
The target number or size of instances is the value the system tries to reach.
A rule defines which change is triggered by which measurements.
Minimums, maximums, cooldowns, and hysteresis keep automatic adjustment within safe and stable boundaries.
Horizontal scaling adds or removes instances; vertical scaling changes the resources assigned to each instance.
Autoscaling is useful when traffic is bursty, daily patterns vary, or platforms must balance batch and real-time workloads. It can improve availability and cost-efficiency, but it still needs usable metrics, well-chosen thresholds, and enough lead time for provisioning. For stateful services, slow startup times, or very short spikes, the reaction may be too late or too costly.
Where this building block is located in the topic model.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.