Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Code that works when every dependency is healthy is only part of a reliable system. Engineering also means deciding what happens when a dependency stalls, a fault spreads, or a feature becomes unavailable—and making sure the system’s response fits the consequences.

What does it mean to decide how software fails?

It means treating adverse conditions as part of the design, not as surprises to address only after an outage. A system should have an intended response when a component or dependency misbehaves: detect the problem, limit its effects, recover where appropriate, or stop in a safer state.

This is a useful way to think about senior engineering, not a measured dividing line between senior and junior developers. The available evidence does not compare career levels. The point is the judgment involved: understanding system boundaries, weighing risk, and choosing behavior for more than the expected path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Carnegie Mellon’s Software Engineering Institute (SEI) recommends anticipating ways a system could fail and writing requirements that can be analyzed, rather than specifying only normal operation. Its guidance focuses on dependable systems; it also cautions that practices have limitations and should be adapted to the mission and organization. SEI’s SPRUCE Project guidance is a useful starting point.

How can one failure spread through a system?

Consider a payment provider that stops responding promptly. A service waiting on it may hold requests open. If callers retry aggressively, those attempts can add work instead of restoring service. Connections may be exhausted, leaving other requests unable to proceed—even requests for features that do not need the payment provider.

This is an illustrative cascade, not a report of a particular incident. The title-matching DEV article uses this kind of dependency failure to make the point: a local problem can affect components beyond the one that first failed.

Analysis should distinguish a fault from the system-level failure it may cause. A fault might be a defect, a failed dependency, or an operational problem; it becomes consequential when activated and its effects propagate through interacting components. NASA’s safety analysis considers failure modes, their effects, and their likelihood, and describes techniques such as fault tree analysis and failure modes and effects analysis (FMEA). Those methods address safety-relevant systems and should not automatically be imposed on every application. NASA’s system-safety memorandum explains that context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should engineers decide before a fault occurs?

Start with the consequences and the functions that matter most. For each plausible failure scenario, decide how the system should respond and what evidence would show that the response works.

  • What could fail? Identify relevant dependencies, components, and operational conditions—not just code defects.
  • How will the system detect it? Define signals that distinguish a slow or unavailable dependency from ordinary variation.
  • Can the effects spread? Check whether waiting requests, retries, or shared resources could affect unrelated work.
  • What must remain available? Separate essential functions from optional features that can be suspended.
  • What is the right response? Depending on the system, return a cached result, queue work, reject requests quickly, continue in a degraded mode, or enter a safe state.
  • How will the team verify the choice? Specify observations and tests that demonstrate the intended behavior under failure.

SEI’s guidance calls for operational systems to detect an impending or active fault, signal it, and fail in an appropriate way. Redundancy and transition to a safe state are possible responses, not universal requirements. The right choice depends on the system’s purpose and the consequences of continuing incorrectly.

Which resilience patterns can contain a failure?

Patterns are tools for particular failure modes, not guarantees of reliability. The Microsoft Azure Well-Architected Framework’s resilience guidance distinguishes resilience from performance and scalability and recommends deliberate testing of failure behavior.

Timeouts

A timeout limits how long a caller waits for an operation. Without one, stalled dependencies can tie up resources indefinitely. A timeout does not make the dependency healthy; it bounds the wait and makes a defined fallback or error path possible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Circuit breakers

A circuit breaker stops sending requests to a dependency after failures cross a defined threshold, allowing the caller to fail quickly rather than repeatedly adding load. It must be configured and monitored carefully; an inappropriate threshold or recovery policy can cause needless failures or hide a dependency’s return to service.

Bulkheads

Bulkheads isolate resources or workloads so one troubled dependency is less likely to consume capacity needed by other functions. The goal is containment: a payment integration should not automatically be able to exhaust every connection or worker used by the application.

Redundancy

Redundant components or paths can preserve service when one component fails, but they add complexity and do not help if the alternatives share the same underlying fault. NASA’s safety memorandum discusses redundancy alongside independence, detection, isolation, and recovery for safety-relevant architecture.

Graceful degradation

When full service is impossible, a system may preserve core behavior while disabling optional capabilities. For example, a service might continue to show previously available information while a live update is unavailable. Whether stale data is acceptable depends on its use; in a safety-sensitive decision, returning old information could be worse than stopping.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These choices are tradeoffs. A customer-facing service may be more useful in a degraded state, while a safety-critical system may need to stop or transition to a safe state. Severity, likelihood, propagation, recovery needs, and cost all matter.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How can a team tell whether its failure design works?

A design diagram or the presence of a named pattern is not proof that failures are contained. Teams need operational signals and tests that exercise the behavior they intend to rely on.

  • Monitor for impending or active faults and make the signal visible to the people responsible for the service.
  • Test dependency delays and failures, including whether retries increase load or shared resources become unavailable.
  • Check that fallback behavior is acceptable for the data and user task involved.
  • Verify recovery as well as failure handling: a system should resume normal behavior safely when the dependency recovers.
  • Use analysis proportional to the consequences. Formal safety methods may be appropriate for safety-critical systems, but they are not a default burden for every low-risk service.

Resilience testing can expose mismatches between the intended design and actual behavior, but no single test or pattern proves that a system will withstand every failure. Evidence accumulates through analysis, monitoring, and tests that reflect the risks the system actually faces.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.