Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You should worry about a growing replication queue when the standby falls further behind than your application can tolerate, or when the WAL kept for that standby starts consuming the disk headroom you need. A lag number that rises for a few seconds is rarely the signal. A backlog that keeps growing while the standby is connected, and that is not shrinking, is.

This guide uses PostgreSQL physical streaming replication as the concrete example, because the PostgreSQL documentation is the authoritative reference for the metrics and behaviour described here. The view names, column names and thresholds below do not transfer unchanged to MySQL, Kafka, or managed database migration services, so check their own documentation before applying the same logic.

Start with the freshness your application actually needs

There is no universal number of seconds or bytes at which every PostgreSQL system should page someone. A reporting replica that refreshes once an hour can tolerate minutes of delay. A read replica serving logged-in users, or a standby that is the failover target for a payment system, may need delay measured in seconds. Before you look at any metric, write down three things:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The maximum replay delay that read traffic on the standby can tolerate.
  • The maximum time a failover can take before the standby must be promotable.
  • The disk budget the primary can spend on WAL retained for replication.

Those three numbers are your alert thresholds. Everything else in this article is about measuring them correctly.

#1 Best Overall
Sale
UGREEN NAS DH2300 2-Bay for Beginners & Personal Users, Phone Backup
  • Entry-level NAS Personal Storage:UGREEN NAS DH2300 is your first and best NAS made easy. It is designed for beginners who want a simple, private way to store videos, photos and personal files, which is intuitive for users moving from cloud storage or external drives and move away from scattered date across devices. This entry-level NAS 2-bay perfect for personal entertainment, photo storage, and easy data backup (doesn't support Docker or virtual machines).
  • Set Your Devices Free, Expand Your Digital World: This unified storage hub supports massive capacity up to 64TB.*Storage drives not included. Stop Deleting, Start Storing. You can store 22 million 3MB images, or 2 million 30MB songs, or 43K 1.5GB movies or 67 million 1MB documents! UGREEN NAS is a better way to free up storage across all your devices such as phones, computers, tablets and also does automatic backups across devices regardless of the operating system—Window, iOS, Android or macOS.
  • The Smarter Long-term Way to Store: Unlike cloud storage with recurring monthly fees, a UGREEN NAS enclosure requires only a one-time purchase for long-term use. For example, you only need to pay $459.98 for a NAS, while for cloud storage, you need to pay $719.88 per year, $2,159.64 for 3 years, $3,599.40 for 5 years. You will save $6,738.82 over 10 years with UGREEN NAS! *NAS cost based on DH2300 + 12TB HDD; cloud cost based on 12TB plan (e.g. $59.99/month).
  • Blazing Speed, Minimal Power: Equipped with a high-performance processor, 1GbE port, and 4GB RAM on Board, this NAS handles multiple tasks with ease. File transfers reach up to 125MB/s—a 1GB file takes only 8 seconds. Don't let slow clouds hold you back; they often need over 100 seconds for the same task. The difference is clear.
  • Let AI Better Organize Your Memories: UGREEN NAS uses AI to tag faces, locations, texts, and objects—so you can effortlessly find any photo by searching for who or what's in it in seconds. It also automatically finds and deletes similar or duplicate photo, backs up live photos and allows you to share them with your friends or family with just one tap. Everything stays effortlessly organized, powered by intelligent tagging and recognition.

What the lag columns in pg_stat_replication actually measure

The view pg_stat_replication on the primary has one row for each standby directly connected to that server. It exposes WAL positions (sent_lsn, write_lsn, flush_lsn, replay_lsn) and three interval columns, write_lag, flush_lag and replay_lag. Each interval describes how long recent WAL took to reach a particular stage on the standby.

replay_lag is the one most people mean

For an asynchronous standby, the PostgreSQL documentation states that replay_lag approximates the delay before recent transactions become visible to queries. That is the number that matters for read freshness, because it describes what a user querying the standby would see.

write and flush lag tell you where the pipeline is slow

If write_lag is large, the standby is slow to receive or write WAL, which often points to network throughput or the standby’s I/O. If write_lag is small but flush_lag or replay_lag is large, WAL is arriving but not being flushed or applied quickly enough. That usually means the standby’s replay process, a long-running query on the standby that blocks replay, or disk pressure on the standby is the bottleneck.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What lag values do not tell you

This is the most common misreading. The PostgreSQL documentation is explicit that the reported lag times are not predictions of how long the standby will take to catch up at the current replay rate. A standby showing 90 seconds of replay_lag does not mean it will be current in 90 seconds. Lag is a measurement of recent WAL progress, not an estimate of remaining work.

Rank #2
Sale
UGREEN NAS DXP2800 2-Bay for Advanced Home Users, Remote Workers & Creators
  • 【Advanced Home Data & Media Hub】For advanced home users who need phone backup, file storage, and centralized data management. Centralize family photos, 4K videos, movies, computer backups, and personal files in one place while running multiple apps for home entertainment and everyday data management. Suitable for households with growing digital libraries and multiple NAS use cases.
  • 【Built for Creators, Media Servers & Advanced Apps】Powered by the Intel N100 Quad-Core CPU, 8GB DDR5 RAM, 2.5GbE networking, and dual M.2 NVMe slots, DXP2800 handles large files and heavier workloads with ease. Run Docker, virtual machines, and media server applications compatible with Plex—ideal for content creators, tech enthusiasts, and advanced home users managing 4K videos, RAW photos, personal media libraries, and multiple NAS apps.
  • 【Up to 80TB for Growing Digital Libraries】 Supports up to 80TB of storage using two HDD bays and two M.2 NVMe SSD slots for family photos, movies, RAW photos, 4K videos, work files, and device backups. AI photo management supports recognition of people, objects, scenes, and locations, album organization, and duplicate photo detection. HDDs and SSDs are not included.
  • 【AI-powered Home Surveillance】Turn DXP2800 into a centralized home surveillance hub by connecting compatible network cameras and storing recordings locally on your NAS. AI-powered features include Face Recognition, People Detection, and Pet Detection, helping advanced home users review important events more efficiently while managing home surveillance and personal data in one place.
  • 【One data Center Across Your Devices】Keep files from desktops, laptops, phones, tablets, and other devices together instead of scattered across cloud accounts and external drives. Access, back up, organize, and share data across Windows, macOS, Android, iOS, web browsers, and compatible smart TVs—ideal for creators and advanced home users working across multiple devices.

Idle systems also produce confusing values. When the standby has caught up and the primary is generating little WAL, the lag columns can eventually become NULL rather than zero. A NULL in an idle period is not a failure. A NULL during a period when the primary is busy is worth investigating, because it means the standby has stopped reporting progress.

Use the byte gap to answer the catch-up question

To estimate whether the standby is gaining or losing ground, compare WAL positions over time rather than reading a single lag value. Run this on the primary:

SELECT application_name,
       client_addr,
       state,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn)) AS replay_gap,
       replay_lag
FROM pg_stat_replication;

Sample this every minute or so and record the replay_gap value. The question to ask is whether the gap is shrinking, flat, or growing. A flat or growing gap while the standby is connected means replay is not keeping pace with newly generated WAL. A shrinking gap means the standby is catching up, even if replay_lag still looks high.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a growing queue is a real problem

A growing queue is worth acting on when one of these conditions holds:

Rank #3
BUFFALO LinkStation 210 2TB 1-Bay NAS Network Attached Storage with HDD Hard Drives Included NAS Storage that Works as Home Cloud or Network Storage Device for Home
  • Value NAS with RAID for centralized storage and backup for all your devices. Check out the LS 700 for enhanced features, cloud capabilities, macOS 26, and up to 7x faster performance than the LS 200.
  • Connect the LinkStation to your router and enjoy shared network storage for your devices. The NAS is compatible with Windows and macOS*, and Buffalo's US-based support is on-hand 24/7 for installation walkthroughs. *Only for macOS 15 (Sequoia) and earlier. For macOS 26, check out our LS 700 series.
  • Subscription-Free Personal Cloud – Store, back up, and manage all your videos, music, and photos and access them anytime without paying any monthly fees.
  • Storage Purpose-Built for Data Security – A NAS designed to keep your data safe, the LS200 features a closed system to reduce vulnerabilities from 3rd party apps and SSL encryption for secure file transfers.
  • Back Up Multiple Computers & Devices – NAS Navigator management utility and PC backup software included. NAS Navigator 2 for macOS 15 and earlier. You can set up automated backups of data on your computers.
  • The replay gap has grown across several consecutive samples and is not returning to a baseline after normal load drops.
  • The observed replay delay exceeds the freshness objective you wrote down earlier.
  • replay_lsn has stopped advancing while sent_lsn keeps moving.
  • The standby has disconnected, so the view no longer shows it and the primary is retaining WAL for it.

A single spike during a bulk load or a large vacuum on the primary is not the same as a sustained trend. Give the standby time to recover from known events before escalating.

Diagnose the direction: data not arriving, or replay falling behind

The next decision depends on which position is moving.

Observation Likely meaning First check
sent_lsn advancing, write_lsn lagging WAL is sent but not written on the standby Network throughput, standby disk I/O
write_lsn advancing, replay_lsn flat WAL is received but replay is blocked or slow Long-running standby queries, replay CPU, standby disk
sent_lsn flat, row missing from the view Standby is disconnected Standby log, network path, authentication, replication slot state
Lag NULL with primary idle Standby is caught up and nothing new is generated Confirm WAL generation rate before treating it as a fault

When WAL is arriving but replay is blocked, a standby that serves queries is often the cause. Long-running read transactions on the standby can delay replay conflicts, depending on your settings. Check the standby’s activity and its replay settings before changing anything on the primary.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Replication slots and disk risk

A replication slot tells the primary to keep WAL until the consumer has read it. That protects continuity: a standby that briefly disconnects can resume from where it stopped instead of needing a full base backup. The cost is that a disconnected or stalled consumer can cause WAL to accumulate on the primary.

Rank #4
BUFFALO LinkStation 210 4TB 1-Bay NAS Network Attached Storage with HDD Hard Drives Included NAS Storage that Works as Home Cloud or Network Storage Device for Home
  • Value NAS with RAID for centralized storage and backup for all your devices. Check out the LS 700 for enhanced features, cloud capabilities, macOS 26, and up to 7x faster performance than the LS 200.
  • Connect the LinkStation to your router and enjoy shared network storage for your devices. The NAS is compatible with Windows and macOS*, and Buffalo's US-based support is on-hand 24/7 for installation walkthroughs. *Only for macOS 15 (Sequoia) and earlier. For macOS 26, check out our LS 700 series.
  • Subscription-Free Personal Cloud – Store, back up, and manage all your videos, music, and photos and access them anytime without paying any monthly fees.
  • Storage Purpose-Built for Data Security – A NAS designed to keep your data safe, the LS200 features a closed system to reduce vulnerabilities from 3rd party apps and SSL encryption for secure file transfers.
  • Back Up Multiple Computers & Devices – NAS Navigator management utility and PC backup software included. NAS Navigator 2 for macOS 15 and earlier. You can set up automated backups of data on your computers.

The PostgreSQL documentation warns that slots can retain enough WAL to fill the primary’s pg_wal space. If that happens, the primary can stop accepting writes, which turns a standby problem into an outage on the primary. This is the main reason a growing replication queue can become urgent even when the standby’s freshness looks acceptable.

Check slot retention and free space

On PostgreSQL 13 and later, pg_replication_slots exposes active, restart_lsn, wal_status and safe_wal_size. Check it together with free space on the volume holding pg_wal:

SELECT slot_name,
       active,
       wal_status,
       pg_size_pretty(safe_wal_size) AS safe_wal_remaining,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal
FROM pg_replication_slots;

An inactive slot with a growing retained_wal value is the most urgent pattern you can see on the primary. Act on it before the standby’s lag becomes a user-facing problem.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The max_slot_wal_keep_size tradeoff

Setting max_slot_wal_keep_size limits how much WAL a slot can retain, and PostgreSQL enforces that cap at checkpoint time. This protects the primary’s disk. The tradeoff is that if WAL required by a slot is removed because the slot fell too far behind, its standby may no longer be able to continue replication from that slot. Recovering usually means rebuilding the standby from a new base backup.

Best Value
Synology DS225+ Private Cloud Media Server - Stream, Back Up Photos & Share Files, Intel CPU for Hardware Transcoding (2-Bay Diskless NAS)
  • Your Personal Streaming Server - Build your own Netflix-style media library and stream 4K movies, shows and photos to any device without monthly fees
  • Create Your Own Cloud - Store your entire photo, video and music collection; access from anywhere with fast 282 MB/s transfer speeds
  • Creator-Grade Backup Solution - Protect your irreplaceable content with automated backups to cloud services, external drives and remote NAS
  • Multi-Layered Data Protection - Combine RAID redundancy, automated backups and snapshot technology to prevent data loss from any cause
  • Smart Home Surveillance - Support up to 30 IP cameras with AI detection, instant alerts and secure remote monitoring

Treat the cap as a recovery decision, not a cleanup switch. Before you set it, decide how long a standby can be disconnected before you would rather rebuild it than keep disk reserved, and make sure you monitor wal_status so you know when a slot is approaching that limit.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Setting your own alert threshold

Build the threshold from the inputs you already have. A practical starting structure is:

  • Freshness alert: warn when replay_lag or the replay gap exceeds a fraction of your freshness objective, and page when it stays above the full objective for longer than your normal recovery window.
  • Trend alert: warn when the replay gap grows across several consecutive samples while the standby is connected.
  • Disk alert: warn when the retained WAL for any slot, or the free space on pg_wal‘s volume, crosses a level that leaves enough time to respond at the observed WAL generation rate.
  • Slot state alert: alert on any inactive slot, and on any slot whose wal_status moves toward a state where its required WAL is no longer guaranteed.

The exact numbers depend on your workload. A primary that generates a gigabyte of WAL an hour has very different headroom from one that generates a gigabyte a minute, so calculate time-to-full from your measured WAL rate rather than copying a fixed byte value from another system.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A short decision checklist

  • Is the standby serving reads, acting as a failover target, feeding analytics, or consuming changes? Each use has a different tolerance.
  • Is the replay gap shrinking, flat, or growing over the last several samples?
  • Is the standby connected, and are sent_lsn and replay_lsn still moving?
  • Is any replication slot inactive, and how much WAL is it retaining compared with free space?
  • If the cap were reached tomorrow, do you have a documented path to rebuild the standby?

If the answer to the first three questions shows a sustained trend that breaks your freshness objective, or the last two show disk or slot risk, the queue deserves immediate attention. If the lag is high only during known load and recovers on its own, it is usually a capacity question for later.

Quick Recap

Bestseller No. 3
BUFFALO LinkStation 210 2TB 1-Bay NAS Network Attached Storage with HDD Hard Drives Included NAS Storage that Works as Home Cloud or Network Storage Device for Home
BUFFALO LinkStation 210 2TB 1-Bay NAS Network Attached Storage with HDD Hard Drives Included NAS Storage that Works as Home Cloud or Network Storage Device for Home
2TB capacity – 1 Drive bay, HDD included.; Made in Japan – Quality Devices.; 24/7 US-based support, with 2-year warranty, including hard drives.
$153.99
Bestseller No. 4
BUFFALO LinkStation 210 4TB 1-Bay NAS Network Attached Storage with HDD Hard Drives Included NAS Storage that Works as Home Cloud or Network Storage Device for Home
BUFFALO LinkStation 210 4TB 1-Bay NAS Network Attached Storage with HDD Hard Drives Included NAS Storage that Works as Home Cloud or Network Storage Device for Home
4TB capacity – 1 Drive bay, HDD included.; Made in Japan – Quality Devices.; 24/7 US-based support, with 2-year warranty, including hard drives.
$192.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.