An AWS pilot light keeps essential recovery infrastructure and replicated data in a second Region, but leaves some application compute undeployed until recovery. That can reduce the recovery environment’s running footprint; it does not make failover automatic or prove the workload can recover on time. The hard work is making data, application dependencies, configuration, capacity, and traffic come together in a tested recovery process.
What does pilot light mean on AWS?
A pilot light is a disaster-recovery pattern in which core infrastructure and data-replication or backup resources are maintained in a recovery Region, while some application resources remain absent or inactive. If the primary Region fails, the recovery process creates or activates the missing resources, scales the environment, and redirects users.
The name is useful as a mental model, not as a precise architecture specification. The recovery Region must contain enough working infrastructure to support the data path and the steps required to start the application. What stays deployed depends on the workload and its services.
AWS’s Well-Architected Framework says, “The difference between pilot light and warm standby can sometimes be difficult to understand.” The practical distinction is how much of the application is already running and able to serve requests before a disaster.
#1 Best Overall
How does pilot light compare with warm standby?
Warm standby keeps a reduced but functional copy of the workload running in the recovery Region. It can accept traffic and is primarily scaled up during recovery. A pilot light leaves more application capacity or components to be deployed or activated first.
| Strategy | Published RPO | Published RTO | What AWS describes |
|---|---|---|---|
| Pilot light | Minutes in the AWS Well-Architected Framework; tens of minutes in AWS Prescriptive Guidance’s comparison for a full application-and-database stack | Tens of minutes in both sources | Core recovery infrastructure and data capabilities are maintained; additional application resources are brought up during recovery. |
| Warm standby | Seconds in the AWS Well-Architected Framework | Minutes in the AWS Well-Architected Framework | A reduced, functional environment is already running and can take traffic before it is scaled up. |
These are AWS strategy-level comparisons, not service guarantees or measured results for a particular application. AWS’s two cited pilot-light comparisons also differ on RPO: the Well-Architected Framework says minutes, while Prescriptive Guidance says tens of minutes for a full stack. Actual recovery and data-loss outcomes depend on workload design, service behavior, the recovery procedure, and test results.
Pilot light is not automatically cheaper in total operational terms. Less capacity may be running, but the design still requires replicated data, maintained deployment paths, recovery automation or operator procedures, and regular testing. AWS describes recovery patterns as trade-offs in cost, complexity, and recovery objectives.
Rank #2
What should be ready in the recovery Region?
Plan the recovery Region around the complete workload, not just the compute instances that appear on an architecture diagram. For each component, decide whether it must already be running, can be provisioned during recovery, or must be promoted or redirected.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Data: Identify what is replicated continuously, what is backed up, and how the recovery copy is restored or made writable.
- Application infrastructure: Define which resources are absent or reduced, and how they will be created and brought to production capacity.
- Dependencies: Account for services the application needs to start and serve traffic, including configuration, secrets, encryption keys, and any external or shared services.
- Traffic management: Choose how clients will be directed to the recovery Region and who or what is allowed to initiate the change.
- Operations: Assign ownership for declaring recovery, running the procedure, validating service health, and communicating status.
Set recovery objectives before choosing the pattern
Recovery time objective (RTO) is the acceptable time to restore service. Recovery point objective (RPO) is the acceptable data-loss interval. Establish both from business needs before selecting a strategy: a lower-cost design may not meet the required recovery objectives. Apply the objectives to the complete workload, including data availability, application startup, dependency checks, and traffic redirection—not just to a single service’s replication metric.
Keep the recovery environment deployable
Infrastructure as code can make it easier to create missing resources consistently, but code alone does not prevent drift or prove that the recovery Region is ready. Keep application and infrastructure changes synchronized across Regions, and verify that the recovery deployment can reach the required capacity. A deployment that succeeds is only one input to a recovery test.
How should data replication and backups work together?
Replication helps make recent changes available in another Region; it is not a substitute for backups. If data is corrupted, deleted, or otherwise damaged in the source, replication may carry that damage to the recovery copy. AWS recommends backing up data in its Region and copying those backups to the recovery Region.
Design and test both paths: the path to the latest replicated data and the path to a known-good backup or recovery point. Confirm who can access the recovery data, how it becomes usable by the application, and how the restored application is checked before traffic is sent to it.
Plan database promotion explicitly
Some database topologies require promoting a cross-Region read replica before the application can use it for writes. The application must then connect to the promoted database, not continue using the old endpoint or read-only replica. The exact promotion, endpoint, and validation steps depend on the database service and topology; document the actual service-specific procedure rather than assuming all replicas fail over the same way.
Rank #4
How does a multi-Region failover work?
A useful runbook treats failover as a sequence with decision points and validation gates. The precise commands and console paths are service- and architecture-specific, so record the tested actions for the resources actually in use.
- Declare the event. Define the conditions for starting recovery, who can authorize it, and how to avoid competing recovery attempts.
- Assess the data state. Determine whether the replicated data is suitable for recovery or whether the procedure should use a known-good backup.
- Promote or restore data services. Perform the database promotion, restore, or other data action required by the topology, then validate that the application can use the resulting data endpoint.
- Provision and scale the application. Create or activate missing resources, apply current configuration, and scale to the capacity needed for recovery.
- Validate dependencies and health. Check that the application can reach its data, keys, secrets, and other required services, and verify its behavior before sending production traffic.
- Redirect traffic. Use the chosen routing mechanism and confirm that requests are reaching the recovery Region. Include any DNS caching or client behavior relevant to the architecture in the tested procedure.
- Monitor and stabilize. Check service health, error rates, and data behavior after the change, and define how the team will handle a prolonged recovery or a later return to the primary Region.
AWS lists Route 53, Application Recovery Controller, Global Accelerator, and CloudFront as possible multi-Region routing options. They are not interchangeable: choose based on the architecture, the recovery objectives, and whether routing control should be manual or automated. Document the actual control and validation steps for the selected option.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What can make a recovery runbook fail?
Configuration drift between Regions
Changes that reach the primary Region but not the recovery Region can leave the application unable to start or behave correctly after failover. Manage infrastructure and application changes so both Regions remain aligned, and include drift checks in the recovery-readiness process.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchBest Value
Control-plane assumptions
Recovery procedures can depend on control-plane actions to create resources or change routing. AWS guidance recommends considering that reliance and points to Application Recovery Controller readiness checks and routing controls. Whether a particular design remains operable during a specific control-plane incident is architecture-dependent; validate the exact recovery path rather than assuming it.
Encryption and access gaps
Cross-Region encryption design varies by service. AWS discusses Region-scoped KMS keys and multi-Region KMS keys as design choices, but the correct setup depends on the services and data paths involved. Verify that the recovery workload can access the required keys and secrets, and test those access paths from the recovery Region.
Unverified traffic changes
A routing change is not proof that the application is ready. Make traffic redirection contingent on checks that the recovery application is healthy and connected to usable data. The selected routing mechanism and client behavior determine which additional checks matter.
How should a pilot-light design be tested?
A diagram documents intended components; a deployment proves that resources can be created. Neither alone establishes that the workload can recover. AWS recommends testing disaster recovery, automating recovery where appropriate, and managing configuration drift.
- Exercise the runbook, including authorization, data promotion or restoration, application provisioning, scaling, validation, and traffic changes.
- Measure the workload’s actual RTO and RPO during the exercise; do not substitute AWS’s strategy-level ranges for those measurements.
- Record failed steps, manual work, stale configuration, missing access, and dependencies that were not apparent from the design.
- Retest after significant infrastructure, application, database, security, or routing changes.
- Test the backup recovery path as well as replication, so the team knows how it will respond to corrupted or unusable replicated data.
The outcome to record is not merely whether the drill completed. It is how long the full service took to recover, what data point was recovered, which actions required people, and what prevented the runbook from meeting its objectives.
Where does AWS Elastic Disaster Recovery fit?
AWS presents Elastic Disaster Recovery (AWS DRS) as an option for teams considering pilot light or warm standby. AWS describes continual data protection with RPO measured in seconds and RTO measured in minutes, while only replication resources remain deployed until failover or a drill. Those are AWS’s service descriptions, not a guarantee for every workload or a result from a particular implementation. Evaluate the service against the application’s recovery procedure and test outcomes rather than treating its stated figures as the whole system’s recovery time.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

