iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
For an AWS workload, an SLI measures service behavior, an SLO sets the target for that measure, and an SLA states the commitment and consequences if service falls short. An error budget turns the gap between the SLO and perfect performance into a shared way to manage reliability and release risk.
What SLI, SLO, and SLA mean
| Term | Meaning | Example |
|---|---|---|
| SLI (service-level indicator) | A carefully defined quantitative measure of service behavior. | The proportion of eligible requests completed within a stated latency threshold. |
| SLO (service-level objective) | A target or range for an SLI, evaluated over a stated period. | At least 99.9% of eligible requests succeed during the evaluation window. |
| SLA (service-level agreement) | An agreement that describes expected service and what happens if the provider fails to deliver it. | A customer-facing commitment with specified remedies if its terms are not met. |
An internal SLO is not automatically a contractual SLA. Teams sometimes use “SLA” loosely, so define which meaning applies whenever the term appears in an operational policy or customer agreement.
Choose an SLI that represents the user experience
Start with what users need to accomplish, then select measurements that approximate that outcome. Request latency, error rate, and throughput are common indicators, but an easy-to-collect infrastructure metric is not necessarily a useful measure of service quality. Google SRE guidance recommends defining the measured population, threshold, aggregation or percentile, and evaluation window so another person can reproduce the result.
For example, “99% availability” leaves important questions unanswered: which requests count, what qualifies as success, and over what window? A more operationally useful objective might specify the percentage of eligible requests returning a successful result within a defined latency threshold during a rolling period. The exact threshold and population must reflect the service’s user-facing meaning of success; they are not universal defaults.
#1 Best Overall
- Population: State which requests or operations are included and how retries, health checks, and internal traffic are treated.
- Success condition: Define success in application terms, not just by whether a component is running.
- Aggregation: Specify whether the SLI uses a ratio, percentile, or another calculation.
- Window: State the calendar or rolling period over which attainment is evaluated.
Set an SLO that matches business needs
A target is a product and business decision as well as an engineering one. Consider the impact of failure, user expectations, available alternatives, workload criticality, and the cost and complexity of improving reliability. Google SRE cautions against setting objectives solely from current performance; a realistic initial target can be tightened as evidence about users and operations improves. AWS Well-Architected likewise advises aligning availability goals with business needs and the criticality of workload components.
More nines are not automatically better. A stricter target can require additional architecture, operational effort, and spending, while leaving less room for routine changes. Compare candidate objectives by user impact, measurement fidelity, dependency assumptions, cost and complexity, alert quality, and the effect on release speed. The right target is the one the organization can justify and operate, not the highest number it can publish.
Rank #2
Use AWS availability figures as design inputs, not promises
AWS Well-Architected defines availability as the percentage of time a workload is available for use and emphasizes that the result depends on the measurement period and the definition of “available.” Its 2024 Reliability Pillar revision gives these illustrative availability goals and yearly interruption allowances:
| Illustrative availability goal | Yearly interruption allowance |
|---|---|
| 99% | 3 days 15 hours |
| 99.9% | 8 hours 45 minutes |
| 99.95% | 4 hours 22 minutes |
| 99.99% | 52 minutes |
| 99.999% | 5 minutes |
These are AWS design examples, not recommendations for every workload or contractual guarantees. A workload’s own SLO needs a defined success condition and measurement window; an annual allowance should not be substituted for that specification.
Rank #3
Account for dependencies
End-to-end availability can be lower than the availability of any single component. AWS illustrates a workload and two hard, independent dependencies, each at 99.99% availability. Multiplying the three values gives about 99.97% theoretical end-to-end availability. This calculation assumes independence and that all three components must be available; shared failure modes or different architecture can invalidate the assumptions. Redundancy can improve theoretical availability, but only when its components and failure domains behave as intended.
Map the dependencies that can prevent users from completing the critical operation, then verify that the architecture and operational processes support the workload-specific objective. Include cost, performance, scaling, and operational complexity in the decision.
Rank #4
Calculate and interpret an error budget
An error budget is the amount of SLO miss tolerated over a specified evaluation window. For a success-percentage objective, calculate the allowed failure fraction as 1 − SLO target. At a 99.99% success target, the allowed failure fraction is 0.01%, as in Google SRE’s example. To express that as failed requests, multiply the fraction by the number of eligible requests in the window: with 1,000,000 eligible requests, 0.01% permits 100 failures. That request count is an illustration of the formula, not a Google or AWS workload target.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For an availability objective measured as healthy time, the budget can instead be expressed as unhealthy time within the window. The result depends on the chosen SLI and window, so specify both when reporting budget remaining. Google describes monthly budgets as common in its practice and quarterly resets as an option for mature services with very high objectives; neither calendar choice is mandatory.
Best Value
Google SRE chapter author Marc Alvidrez writes, “The error budget provides a clear, objective metric that determines how unreliable the service is allowed to be within a single quarter.” The point is not that every team must use a quarter, but that a defined budget makes acceptable risk explicit. Google SRE also describes “Hope is not a strategy” as its unofficial motto.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Turn budget consumption into an operating policy
Use the SLI and SLO as inputs to an agreed decision process, rather than treating the budget as a dashboard decoration. Google’s guidance describes pausing most changes after budget exhaustion, with exceptions for urgent security fixes and changes that address the errors driving the budget loss. Its workbook and example policy also illustrate postmortem and escalation thresholds. Those are Google examples, not universal rules.
- Measure: Calculate the SLI against the defined population and window.
- Assess the rate: Determine whether the remaining budget is being consumed quickly enough to threaten the objective.
- Choose a response: Continue, slow, or pause changes according to the severity and cause of the consumption.
- Apply exceptions deliberately: Name who can approve urgent work and what evidence qualifies for an exception.
- Resume against criteria: State what needs to be fixed or reviewed before normal release activity resumes.
Make the policy explicit before a high-pressure incident. It should name the budget window, decision owner, exception process, escalation and postmortem thresholds, and resumption criteria. Google’s example policy attributes roughly 70% of its outages to changes; that figure belongs to Google’s example and should not be treated as a universal industry rate.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallUse CloudWatch Application Signals with metric semantics in mind
Amazon CloudWatch Application Signals supports SLO tracking for services and critical operations. It can use standard latency and availability metrics, or other CloudWatch metrics and expressions; SLOs can use calendar or rolling intervals, and the service displays attainment and remaining error budget.
Validate the standard Availability metric against the application’s definition of success before using it as an SLI: it counts 5xx responses as faults and 4xx responses as successful responses. That classification can misrepresent a user-facing operation—for example, if a particular client error means the requested task did not succeed. Choosing a metric is not enough; its request population and success semantics must match the objective.
Quick Recap
Practical setup sequence
- Identify the service or critical operation whose user outcome matters.
- Choose a standard Application Signals metric or a CloudWatch metric or expression that represents that outcome.
- Define the target, eligible population, success condition, and calendar or rolling evaluation interval.
- Check that the displayed attainment and remaining budget agree with the organization’s own SLI calculation and operating policy.
- Decide how budget consumption affects alerts, release decisions, and escalation before relying on the dashboard operationally.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

