Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Production software is not finished when its feature works. It becomes a continuing service: people depend on it, business expectations change, components fail, and someone must notice, respond, and improve the system. The hardest lessons are often about setting realistic reliability goals, seeing problems clearly, learning from incidents, and making operational work part of the roadmap.

What does production teach you that tutorials usually do not?

A tutorial can show how to build a feature under known conditions. Operating a business application adds uncertainty: real users take unexpected paths, dependencies behave differently over time, and an outage can interrupt a business workflow. The job therefore includes more than implementation. Teams need to know whether the workflow is succeeding, how failures affect users, who owns the service, and how to restore it safely.

The following lessons are grounded in published accounts from Google Cloud Customer Reliability Engineering, GitHub, Atlassian, Meta, and Carnegie Mellon University’s Software Engineering Institute. They are company examples and engineering guidance, not a claim that every team or author has had the same experience.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should a team decide how reliable a service must be?

Start with user impact, then choose a measurable service-level objective (SLO): a target for the level of service users should receive over a defined period. Google Cloud Customer Reliability Engineering describes an SLO as a reliability level below which users will be unhappy, and advises balancing that target against user expectations and engineering expense. Its guidance puts the principle plainly: “Your SLO sets a minimum reliability requirement, something strictly less than 100%.”

That does not mean a team should tolerate avoidable failures. It means reliability targets are product and business decisions as well as engineering decisions. A target that is more demanding than users need can absorb time and money that might have improved other parts of the product, while a target that is too lax can undermine trust or disrupt important work.

Google Cloud’s 2019 article uses 90% and 99.95% SLOs as illustrative examples of objectives that can call for different rollout practices; they are not universal recommendations. It also describes a service that is 10 times more reliable as “100 times more expensive to run.” That is an illustrative cost comparison in the article, not a measured rule that applies to every service. The useful lesson is to make the reliability-versus-cost tradeoff explicit rather than assume that pursuing the highest possible availability is always the right choice.

What should teams measure beyond whether the application is up?

Availability alone does not tell a team whether users can complete the work they came to do. Define indicators tied to the service’s actual behavior, make them visible to the people responsible for the service, and ensure alerts give responders actionable information.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Look for slow experiences hidden by averages

An average can look acceptable while a meaningful group of users experiences severe delays. In its account of cloud reliability work, Atlassian says it had focused on metric averages without enough attention to important 90th- and 99th-percentile values. Those percentiles help reveal the slower end of the experience. The case is a practical warning: choose measurements that can expose problems affecting a portion of users, not just a summary that smooths them away.

Rank #2
Income and Expense Log Book - Bookkeeping Record Book/Tracker
  • Income And Expense Log Book: This Income and Expense Record Book(8.5" x 10.5") is a necessary item for any small business owner or entrepreneur. It is an essential part of any business - helping you understand your overall earnings to determine if you are profitable.
  • Daily Tracking and Weekly Overview: let our log tell you if you are profitable today! There are two pages per week to help you you track your income and expenses. At the end of each day or week, you can note whether you made a profit or a loss for the day.
  • Clear P&L Statement For Your Business: This income and expense book makes it easy to see your expenses and how they fluctuate from time to time. This makes it easy for you to decide where you can cut back on expenses and assess your total annual net profit.
  • Main Features: Expense Review + Income Review + Weekly Pages + Summary of The Year + Twin-Wire Binding + Waterproof Cover + Rounded corner design + Thicker paper
  • Effective Organization: This budget book has a twin-wire binding and you can easily lay it flat at 180°. This effective design can help you work better and bring you great convenience in the process of using.

Make ownership and service expectations discoverable

When a service behaves unexpectedly, responders should be able to find its owner, contacts, service tier, and relevant quality expectations without guesswork. GitHub’s Engineering Fundamentals program used scorecards for availability, security, and accessibility, and recorded service metadata such as tier, owner, sponsor, and contact information. GitHub also describes connecting unmet requirements to action items associated with the service repository. That approach makes operational expectations part of the engineering workflow instead of leaving them in an informal checklist.

Meta’s internal SLICK system standardized and surfaced service-level indicator (SLI) and SLO information, integrating reliability data into workflows and incident response. Meta reported per-minute metric granularity and up to two years of retention in its December 2021 account. Those are specifications of Meta’s internal system at that time, not a retention requirement for every team.

How can an incident become useful learning?

An incident is useful only if the organization can understand what happened and complete improvements that reduce the chance or impact of a recurrence. Google Cloud CRE recommends written postmortems after significant SLO hits and near misses, with concrete follow-up work. It quotes an SRE motto: “Hope is not a strategy.” The point is to replace assumptions about recovery with a record of events, contributing conditions, and actions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Blameless analysis is not the same as avoiding accountability or ignoring harmful decisions. It asks teams to examine the conditions that shaped people’s choices: alert quality, training, workload, procedures, and available tools. Google Cloud CRE writes, “A blameless culture recognizes that people will do what makes sense to them at the time.” The same article says: “Rather we should seek to make improvements in the system to positively influence the person’s actions during the next emergency.”

Atlassian describes tracking whether incidents recur and how long post-incident actions take to complete. These measures help distinguish a postmortem that merely documents a problem from one that changes the system. An action should have an owner and a clear definition of completion; otherwise, important fixes can remain indefinitely on a list.

What operational work belongs on the roadmap?

Feature requests are visible, but reliability, observability, security, and technical-debt work compete for the same engineering capacity. If teams do not explicitly prioritize that work, the system can become harder to change and incidents harder to diagnose even as features continue to ship.

GitHub says its Engineering Fundamentals program was created to address technical debt, reliability, and observability as enterprise needs and platform innovation grew. Atlassian’s retrospective describes technical complexity, observability gaps, and root-cause work accumulating during a large migration and feature drought; later feature demand made it difficult to reserve roadmap time for that debt. These are specific company accounts, not industry-wide measurements. Carnegie Mellon University’s Software Engineering Institute maintains a collection of technical-debt resources, including research reviews and organizational recommendations, reflecting that the subject has both research and practice dimensions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A workable roadmap makes operational improvements visible alongside feature work. For each significant risk, identify the user or business consequence, the evidence that would show improvement, an owner, and a place in planning. That makes tradeoffs explicit and gives the team a way to revisit deferred work rather than letting it disappear.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does moving to distributed services make a system easier to operate?

Not automatically. A distributed architecture can support independent services and different scaling or development needs, but it also introduces interactions, failure modes, and operational responsibilities across those services. Atlassian reports that moving from a smaller number of monolithic codebases to more distributed services brought unintended complexity and lower confidence in adding capabilities. Its response included changes to hiring, training, tools, and fail-safe processes.

This account does not prove that monoliths are always preferable or that distributed systems inevitably fail. It shows why a migration should be evaluated as a change in operating model, not just a change in code organization. Before splitting a system, ask which specific constraint the new boundaries solve and whether the team can support the resulting monitoring, ownership, deployment, and recovery work.

What to put in place before maintaining a business application

  • Define success in user terms: identify the workflow that matters and set an SLO that reflects its expected service level.
  • Make problems observable: choose indicators that show user impact, including slow-tail behavior where it matters, and ensure alerts lead to useful investigation.
  • Assign ownership: keep service contacts, expectations, and operational requirements easy to find.
  • Plan for recovery: document how to respond, restore service, and roll back risky changes; Atlassian’s operational reviews include data integrity and recovery, monitoring, alerting, logging, on-call plans, security, deployments, and rollbacks.
  • Close the learning loop: record significant incidents and near misses, assign specific follow-up actions, and track recurrence and completion.
  • Reserve roadmap capacity: prioritize reliability, observability, and debt reduction explicitly alongside feature delivery.
  • Choose architecture for a reason: weigh the flexibility a change offers against the additional complexity and operating work it creates.

Further reading

For a deeper treatment of SLOs, incident response, and operating services, see Google’s book Site Reliability Engineering: How Google Runs Production Systems, linked from Google Cloud Customer Reliability Engineering’s article on reducing the impact of production incidents.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.