Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A useful managed IT services service-level agreement (SLA) turns broad promises into measurable commitments: what the provider covers, who owns each security and recovery task, how performance is measured, and what happens when service falls short. Use this checklist to review a new agreement or renewal. Set targets around your business impact and purchased services; the official guidance cited here does not establish universal response-time, backup, recovery, or incident-notification numbers.

What should a managed IT services SLA define?

An SLA should describe more than a provider’s headline response time. NIST’s Computer Security Resource Center glossary defines an SLA in terms of responsibilities, service type, expected performance—including response times—reporting, resolution, and termination. NIST Special Publication 800-35 also discusses roles, performance measurement, costs, remedies, service periods, and handling sensitive data.

Use the agreement and its attachments together. Make sure service descriptions, security schedules, backup terms, escalation procedures, and exit provisions do not contradict one another. NIST SP 800-35 is October 2003 guidance, not a current contract template or legal advice; requirements also depend on your jurisdiction, industry, and data obligations.

Checklist: define the service boundary and owners

List the work the provider is actually responsible for, and distinguish routine IT operations from separately scoped security services. CISA’s MSP guidance advises customers to understand the provider’s access and the security work covered by contract. UK NCSC guidance likewise calls for clear responsibilities in an MSP agreement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Covered environment: Identify users, endpoints, servers, locations, cloud services, and configurations included. State business hours, after-hours coverage, and any geographic limits.
  • Included and excluded work: Spell out whether help desk support, patching, system hardening, monitoring, threat detection, incident response, backup administration, and restoration are included. Explain how out-of-scope incidents are handled and priced or escalated.
  • Dependencies: Name customer prerequisites such as supported systems, licensing, connectivity, access, approvals, and timely contact updates. State what the provider may do if a prerequisite is missing.
  • Named roles: Assign provider and customer owners for access approvals, changes, incident decisions, communications, recovery, and service reviews. Define who can authorize disruptive actions such as isolating a device or disabling an account.
  • Subcontractors and data: Disclose whether subcontractors may access systems or data, what controls and accountability apply to them, and how the customer is notified of material changes. Specify sensitive-data handling, staff controls, and any required security qualifications.

Checklist: make response-time targets measurable

A fast acknowledgement is not the same as fixing an issue. For each priority, define both the clock mechanics and what service outcome the target promises. NIST and NCSC call for clear response expectations, but the cited guidance does not set a universal number of minutes or hours. Negotiate targets that match business impact, coverage purchased, and the provider’s actual scope.

  • Severity trigger: Tie each priority to concrete impact—for example, number or type of users affected, critical business function unavailable, or suspected security compromise. Avoid letting the provider assign severity without an agreed definition and review route.
  • Coverage and contact: State service hours, holiday coverage, supported contact channels, and how an urgent ticket is distinguished from a routine request.
  • Clock rules: Specify when time begins, what information is needed to open a ticket, permitted pause conditions, and whether the clock runs outside business hours.
  • Separate milestones: Define acknowledgement, active response, workaround, restoration, and final resolution separately. If a restoration or resolution target is not committed, do not infer one from the response target.
  • Escalation: Name escalation levels, when each is triggered, who is contacted, and how unresolved or worsening incidents reach a decision-maker.
  • Measurement and reporting: Identify the ticketing or monitoring system of record, how missed targets are calculated, report frequency, exclusions, and how either party can challenge an inaccurate record. If availability is included, define the calculation period and exclusions.

Checklist: make backups and recovery verifiable

A promise to “back up” systems is incomplete unless the agreement says what is protected and demonstrates that restoration works. CISA recommends isolated backups and regular testing; its MSP customer guidance stresses recovery exercises. NIST NCCoE’s April 2020 guide is specifically about MSPs conducting, maintaining, and testing backup files.

  • Coverage: Enumerate protected data, applications, systems, and configurations. Explicitly list exclusions and who is responsible for anything outside the backup service.
  • Recovery objectives: Agree on a recovery point objective (RPO), the tolerable amount of recent data loss, and a recovery time objective (RTO), the target time to restore service. Set backup frequency in relation to the RPO. Treat an objective as a target, not a guarantee, unless the contract explicitly commits to that result.
  • Retention and storage: State retention periods, storage locations, separation from production systems, and how copies are protected from unauthorized access or deletion. Define encryption and key ownership, privileged access, and the customer’s ability to obtain backup copies.
  • Job monitoring and restoration: Assign who monitors failed backup jobs, investigates failures, performs restores, approves recovery priorities, and provides status updates. Specify restore assistance and any limits or dependencies.
  • Test evidence: Set a restore-test cadence and scope, success criteria, evidence delivered to the customer, and remediation deadlines when a test fails. An exercise should prove that selected data or systems can be recovered, not merely that a backup job completed.
  • Offline media: External media can be one isolated-storage option, but it is not a complete backup program. If used, specify device capacity, encryption, custody, handling, and rotation practices appropriate to the environment.

Checklist: assign security and incident-response duties

Security responsibilities cross the provider/customer boundary. CISA’s May 11, 2022 joint advisory on MSPs and their customers says customers should understand provider access and contractual scope, including which party owns hardening, detection, and incident response. CISA’s customer guidance also calls for clear separation between IT operations and security services.

  • Preventive controls: Assign responsibility for system hardening, updates, privileged and remote access, and multifactor authentication (MFA) where applicable. State who approves exceptions and how they are tracked.
  • Monitoring and detection: Say whether alert and log monitoring is included, what systems and hours it covers, who triages alerts, and when the customer is contacted. Do not treat ordinary IT support as an implied security monitoring service.
  • Notification and coordination: Define what constitutes a customer-notifiable event, the notification deadline, contact path, and initial information to provide. Set the deadline to fit the contract, applicable law, sector rules, and incident severity; the cited sources do not supply one universal SLA deadline.
  • Investigation records: Define log and record retention, customer access, secure transfer, and preservation during an investigation. Specify how the provider will support the customer’s investigation and any required external reporting.
  • Containment and remediation: Assign who can isolate systems, revoke credentials, or take other disruptive steps; define approval and emergency exceptions. Set remediation acceptance criteria, escalation, and how unresolved risk is reported.
  • Plans and exercises: Coordinate the customer’s incident-response and recovery plans with the provider’s process. Name decision-makers, communications routes, and exercise expectations.

Checklist: measure performance, remedies, and exit

NIST SP 800-35 recommends defining how compliance will be assessed and how often monitoring occurs, alongside service costs, remedies, and the agreement period. Put the practical mechanics in writing rather than relying on informal account reviews.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Governance: Set service-report cadence, review meetings, metric owners, a dispute process, and how changes to users, systems, or business requirements trigger a scope review.
  • Remedies: If negotiated, define service credits or other remedies, how to claim them, exclusions, and caps in the agreement. Do not assume a credit is the customer’s exclusive remedy; that depends on the contract and applicable law.
  • Continuity: State how critical support, communications, and recovery coordination continue during a provider outage or serious disruption.
  • Termination and handoff: Define notice and termination conditions, transition assistance, data export format and timing, secure deletion and confirmation, credential revocation, and cooperation with a replacement provider.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to compare MSP proposals on equal terms

Normalize the scope before comparing prices or headline targets. A proposal with a short response commitment may cover fewer assets, narrower hours, or acknowledgement only, while another may include broader restoration or security work.

Quick Recap

Best Value
Disaster Recovery
  • Used Book in Good Condition
  1. Match the estate: Compare the same users, endpoints, sites, cloud services, and service hours; identify exclusions and customer prerequisites.
  2. Normalize service levels: Compare severity definitions, clock start and pause rules, acknowledgement versus restoration commitments, escalation, and measurement method.
  3. Compare security ownership: Mark which offer includes hardening, monitoring, incident response, notification, investigation support, and remediation—and which leaves each task to the customer.
  4. Compare recovery capability: Check protected workloads, RPO/RTO commitments, isolation, retention, restore assistance, testing evidence, and failure remediation.
  5. Review accountability and exit: Compare subcontractor controls, reporting, remedies, continuity provisions, data return/deletion, and transition assistance.
  6. Record gaps before selection: For every item marked unclear or excluded, get a written answer, price, owner, and contractual commitment before signing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.