Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A Python server monitor stays running reliably when its failure is diagnosed and the right supervisor owns its lifecycle: systemd for a Linux service, Docker for a container, or Kubernetes for a Pod. First determine whether the process exited, was deliberately stopped, its container restarted, or it is still running but no longer doing useful work. A restart rule can recover from some failures, but it cannot fix the underlying bug or detect every hung process.

Find out what stopped—or whether it stopped at all

The title alone does not identify the cause. An exception, resource limit, signal, deployment action, or other event might explain a particular incident, but none can be assumed without logs and runtime details. Start by collecting evidence before changing restart settings.

  1. Check the application logs. Preserve the last lines before the incident, including startup, shutdown, exceptions, and health-state changes. Use Python’s logging facility, and send logs to a destination that survives process or container restarts.
  2. Check the supervisor’s view. Record the systemd service status and journal, Docker container state and logs, or Kubernetes Pod state, events, and logs, depending on where the monitor runs.
  3. Identify the exit outcome. Note the process exit status or termination signal, whether the container stopped or restarted, and whether an operator or deployment deliberately stopped it.
  4. Check for a live but unhealthy process. If the process remains present, determine whether it is responsive and making useful progress. A process-lifecycle restart policy alone may not detect a deadlock or stalled worker.

These checks distinguish four materially different cases: the Python process exited; its container stopped or restarted; a person or deployment intentionally stopped it; or the process is alive but unresponsive. The evidence determines which recovery mechanism makes sense.

Choose one supervisor for the deployment

Use the lifecycle manager that owns the service at its deployment layer. Avoid adding overlapping restart mechanisms without understanding which one will act; Docker specifically warns, “Don’t combine Docker restart policies with host-level process managers, as this creates conflicts.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
  • Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM)
  • Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
  • CanaKit Turbine Black Case for the Raspberry Pi 5
  • CanaKit Low Noise Bearing System Fan
  • Mega Heat Sink - Black Anodized
Where it runs Lifecycle supervisor What its restart behavior covers Key caveat
Linux host service systemd Configured process exits, selected signal terminations, timeouts, and—when configured and supported—watchdog expiry. A deliberate systemctl stop is not automatically undone. Choose behavior according to whether a clean exit should mean the work is complete. [systemd.service]
Docker container Docker restart policy Container exit and restart behavior according to the selected policy. Policy semantics differ for clean exits, manual stops, and daemon restarts. Docker activates a policy only after it has observed the container running successfully for at least 10 seconds. [Docker documentation]
Kubernetes Pod Pod restart policy and kubelet probes Container lifecycle; liveness can trigger a restart, while readiness controls whether the Pod receives traffic. Startup probes allow initialization time. Probe actions are distinct, and an overly aggressive liveness probe can create cascading failures. [Kubernetes documentation]

Linux service managed by systemd

In a systemd unit, Restart=on-failure covers nonzero exits, certain signal terminations, timeouts, and watchdog expiry. Restart=always also restarts after a clean exit, which is undesirable when a clean exit means the monitor has finished its work. Neither setting means that a deliberate systemctl stop will be undone. Configure the unit’s start-rate behavior as well, so repeated failures do not turn into uncontrolled rapid restarts; inspect the journal when restarts recur. Watchdog recovery requires the service to send keep-alive notifications within the configured deadline. See the current systemd.service documentation for unit directives and their precise behavior.

Docker container

Docker provides no, on-failure, always, and unless-stopped policies. They differ in what happens after clean exits, manual stops, and Docker daemon restarts, so select one based on the intended lifecycle rather than treating them as interchangeable. A policy becomes active only after Docker has observed at least 10 seconds of successful container runtime. Also, the terminal command attached to a container can exit while Docker continues restarting that container; check Docker’s container state rather than assuming the CLI session owns its lifetime. Review Docker’s automatic-start documentation for the current policy semantics.

Rank #2
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
  • Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM)
  • Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
  • CanaKit Premium High-Gloss Raspberry Pi 4 Case with Integrated Fan Mount, CanaKit Low Noise Bearing System Fan
  • CanaKit 3.5A USB-C Raspberry Pi 4 Power Supply (US Plug) with Noise Filter, Set of Heat Sinks, Display Cable - 6 foot (Supports up to 4K60p)
  • CanaKit USB-C PiSwitch (On/Off Power Switch for Raspberry Pi 4)

Kubernetes Pod

Set an appropriate Pod restart policy, then decide whether the workload needs startup, liveness, and readiness probes. A failed liveness probe eventually causes the kubelet to kill and restart the container under the Pod’s restart policy. A failed readiness probe marks the Pod not ready and removes it from service traffic; it does not restart the process. A startup probe gives a slow-initializing application time to start before liveness checks apply. The Kubernetes probe documentation explains the actions and configuration fields.

Use health checks for the failure a restart can fix

A supervisor generally reacts to lifecycle events; a health check can identify an application that is still alive but not healthy. In Kubernetes, choose each probe according to what its result should do:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
UCTRONICS 19” 1U Rack Mount for Raspberry Pi with SSD Mounting Brackets, Thumbscrews Front Removable Bracket Supports Up to 4 Raspberry Pi 5, 3B/3B+, 4B and 4 SSDs, Option SD Card Adapter
  • Design for Raspberry Pi: Supports installation of 4 Raspberry Pis and 4 ssds, compatible with any 2.5” Solid State Drive (7mm/9mm) and Rpi 4B/3B+, and other B/B+ models.
  • The SSD mounting bracket also has two holes reserved for the SD card extension adapter ASIN: B09CKRDFTH, which allows you to access the SD card from the front of the rack.
  • Easy to Setup: Just use two included thumbscrews to mount the rackmount, which adopts a screw-in design, which helps you install and replace quickly and easily, no tools needed!
  • Applications: This is a hardware solution to get ingenious use of the Raspberry Pi, with this kit and open source software OpenMediaVault, you can use the Pi as a NAS Server, Surveillance station, or even a Web server.
  • Optional accessories: Single mounting bracket: B09GFQLPTY; Micro SD card extension adapter ASIN: B09CKRDFTH. I/O Panel: B09FXRQPFM
  • Startup: allow legitimate initialization to finish before liveness checking begins.
  • Liveness: detect an internal failure from which restarting the container is likely to help.
  • Readiness: indicate whether the instance should receive service traffic. A failed readiness check does not, by itself, restart the process.

Keep each check low-cost and test only a condition relevant to its action. For example, if a shared dependency is unavailable, an instance may need to stop receiving traffic without being restarted; restarting every instance will not restore the dependency. Kubernetes warns that incorrect liveness probes can cause cascading failures, and that poor readiness-probe implementations can contribute to resource starvation. Use a startup probe where a long initialization period is normal rather than weakening liveness until it stops being useful.

Kubernetes documentation lists defaults of periodSeconds: 10, timeoutSeconds: 1, failureThreshold: 3, and successThreshold: 1. These are configuration defaults, not universal recommendations. Set thresholds and timeouts to match observed startup and recovery behavior; in particular, a one-second timeout may be too short for a check that can legitimately take longer.

Rank #4
Pironman 5-MAX Raspberry Pi 5 Case Dual NVMe M.2 SSD PCIe, Mini PC NAS RAID 0/1 Hailo-8L AI Accelerator PWM Tower Cooler+Dual RGB Fans, OLED Module, Safe Shutdown, Standard HDMI (RPI5 Not Included)
  • [ULTIMATE RASPBERRY PI 5 CASE & MINI PC] - Unlock the full potential of your Raspberry Pi 5 with the Pironman 5-MAX — the most advanced Raspberry Pi 5 Case for power users. This high-performance Raspberry Pi 5 Cooling Case features dual NVMe M.2 slots with RAID 0/1 support, AI accelerator compatibility ( e.g. Hailo-8l M.2 AI), a PCIe Gen2 switch, a PWM tower cooler + dual RGB fans and a smart OLED display. With its dual transparent panels and optimized cable management (including full-size HDMI), it’s the ideal Raspberry Pi 5 Enclosure for building a high-speed NAS, AI edge computing device, or Home Assistant hub. (Raspberry Pi NOT Included)
  • [DUAL NVMe M.2 SLITS & NAS RAID SUPPORT] - Supercharge your storage with the best Raspberry Pi 5 NVMe Case solution. Featuring two expandable NVMe M.2 slots (2230-2280) powered by a built-in PCIe Gen2 switch, this Raspberry Pi 5 NAS Case supports RAID 0/1 for ultra-fast data setups. Whether you're using a high-speed NVMe SSD or a Hailo-8L AI accelerator, Pironman 5-MAX delivers the ultimate performance boost for advanced Raspberry Pi 5 AI applications and edge computing
  • [ADVANCED COOLING SYSTEM] - Engineered for high-performance builds, Pironman 5-MAX features a powerful tower cooler, one PWM fan, and dual RGB fans for enhanced airflow. The dual transparent panel design improves ventilation while showcasing vibrant RGB lighting. Ideal for cooling both the Raspberry Pi 5 and dual NVMe SSDs or AI accelerators like Hailo-8L, it ensures stable operation under heavy workloads with low noise and long-term durability
  • [SMART OLED DISPLAY WITH VIBRATION WAKE-UP] - Pironman 5-MAX features a 0.96" OLED screen that delivers real-time system insights including CPU usage, memory, temperature, IP address, and disk status. With customizable display options and auto sleep mode, the screen can be instantly reactivated by a light tap thanks to the built-in vibration sensor—offering a smarter and more interactive experience
  • [ENHANCED FUNCTIONALITY] - Pironman 5-MAX empowers your Raspberry Pi 5 with advanced features like safe shutdown via a metal power button, customizable RGB lighting, dual full-size HDMI ports, vibration-triggered OLED wake-up, and an external GPIO extender. It also includes RTC battery support for timekeeping and seamless Home Assistant integration. With detailed guides, online tutorials, and full technical support from SunFounder, setup and use are effortless and worry-free
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Make Python shutdown and crash evidence useful

A restart loop without retained logs can erase the clues needed to identify the original fault. Record when the monitor starts and stops, exceptions, health transitions, and context relevant to a restart. Ensure the log destination persists through the kind of restart in use.

Handle shutdown signals with care. Python’s documentation states, “Python signal handlers are always executed in the main Python thread of the main interpreter, even if the signal was received in another thread.” Keep handlers minimal: do not perform blocking cleanup or acquire locks there. Use a thread-safe mechanism such as a synchronization primitive to tell worker threads to stop, and make cleanup bounded so shutdown cannot hang indefinitely. Let systemd, Docker, or Kubernetes own restart decisions. See Python’s signal documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If evidence points to a segmentation fault or a failure in native code, Python’s faulthandler can emit Python and C stack traces to help investigate. It is a diagnostic aid for faults such as native crashes, not a remedy for ordinary Python exceptions, a missing supervisor, or a process that is merely unresponsive.

Quick Recap

Bestseller No. 1
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM); CanaKit Turbine Black Case for the Raspberry Pi 5
$259.95
Bestseller No. 2
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM); Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
$159.99

Apply the fix without hiding the cause

  1. Keep the incident evidence. Save logs, exit status or signal, service or container state, and relevant deployment events before changing settings.
  2. Classify the failure. Separate an unexpected exit from an intentional stop, a container lifecycle event, and a live but unhealthy application.
  3. Configure the owner of the deployment. Use systemd for a host service, Docker restart policy for a container, or Kubernetes restart policy and probes for a Pod.
  4. Validate the recovery path. Confirm that the selected policy handles the failure you observed and that a deliberate stop still behaves as intended. For Kubernetes, verify separately what startup, liveness, and readiness failures do.
  5. Watch the next failure. Retained logs and supervisor state should make it possible to tell whether the change recovered the service or merely restarted it again.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.