What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A p95 latency value tells you where the slowest 5% of measured requests begin—not how slow those requests get, how many requests were measured, or whether users are meeting a service objective. Read it alongside its population, time window, request volume, distribution, and measurement method. A single p95 line can reveal a degrading tail that an average hides, but it cannot describe latency on its own.

What p95 latency means

p95 is the latency value at or below which 95% of observations fall for a defined population and measurement interval. If p95 is 200 ms, 95% of the measured observations took 200 ms or less; roughly 5% took longer. The value is a rank in a distribution, not an average or a promise about every request. The population might be requests to one endpoint, one region, one server, or an entire fleet, so name it along with the time window. Google Cloud’s latency SLO documentation describes percentile groups in this way for measurements such as one-minute intervals.

Why p95 adds information—and what it leaves out

Request latency is a distribution. A mean can remain steady while a subset of requests gets much slower, making p95 useful for spotting tail degradation that an average may conceal. Google’s SRE book illustrates this with typical latency of about 50 ms and 5% of requests taking 20 times longer. That is an explanatory example, not a general benchmark or a measured result that applies to every service. Google’s SRE guidance on service-level objectives discusses latency as a distribution.

But p95 shows only one cut through that distribution. It does not reveal whether the slowest 5% are just above the percentile or many times slower, nor whether a separate group of requests is especially affected. Pair it with request volume and other useful views of the distribution, and drill into relevant dimensions such as endpoint, region, or instance when investigating a change. Google’s monitoring guidance describes using percentiles, logging or sampling, dashboards, and drill-downs to understand system behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can you average p95 across servers?

No. A quantile is not composable by averaging quantile values. Averaging each server’s p95 does not produce the fleet’s p95; it gives an average of server-level percentile values, which is a different statistic. To estimate fleet p95, combine compatible histogram observations and calculate the percentile from the combined data. Prometheus states: “Using histograms, the aggregation is perfectly possible with the histogram_quantile() function.” Prometheus Authors’ histogram guide explains the distinction between histogram and summary aggregation.

Classic Prometheus histograms

For a classic histogram, aggregate the bucket rates while retaining the le bucket-boundary label, then calculate the percentile:

Rank #2
D DICHEN 4K DMA Video Fuser, HDMI & DisplayPort Dual PC Display Combiner, Supports 4K60Hz/2K144Hz, Low Latency AV Processor (DC 6th Gen HDMI Fuser - Pink)
  • Professional 4K Video Fuser Display Combiner The D DICHEN 4K video fuser display combiner is designed for dual-PC visual routing, AV signal workflows, monitor testing, streaming setups, and professional desktop integration. It helps combine and route display signals in a clean, stable, and efficient hardware workflow.
  • Supports 4K60Hz & 2K144Hz Visual Output Built for high-resolution visual performance, this display fuser supports up to 4K at 60Hz and 2K at 144Hz, delivering smooth image output, clear picture quality, and reliable signal handling for demanding video, presentation, and workstation environments.
  • HDMI & DisplayPort Connection Design Equipped with HDMI and DisplayPort connectivity, this video fuser is suitable for dual-computer setups, display routing, signal testing, and multi-device desktop workflows. The clear interface layout helps simplify installation and daily operation.
  • Low-Latency Hardware Workflow Engineered for stable and responsive visual signal processing, this device supports low-latency display routing for streaming desks, AV testing, video production workflows, lab environments, and professional hardware validation tasks.
  • Includes USB Tutorial Drive, Cables & Power Adapter The package includes the video fuser unit, USB tutorial drive, connection cables, and power adapter to support a smoother setup experience. Suitable for authorized research, display testing, AV workflow setup, system validation, and professional electronics projects.
histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))

Native Prometheus histograms

For a native histogram, aggregate the histogram rates and calculate the percentile from the combined histogram:

histogram_quantile(0.95, sum(rate(http_request_duration_seconds[5m])))

These examples use a five-minute rate window; that interval is illustrative, not a universal recommendation. Choose a window suited to the service and the purpose of the query or alert. The exact expression also depends on metric type and labels. Prometheus documents these patterns and the function’s estimation assumptions in its query functions reference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Vogzone for MCX556A-EDAT ConnectX-5 Ex 100GbE VPI 2xQSFP28 PCIe4.0x16 NIC
  • 【Controller】: 100GbE PCI-E NIC with Mellanox connectX-5 VPI controller,which provide high performance and flexible solutions with up to two ports of 100GbE connectivity, 750ns latency, up to 200 million messages per second (Mpps). and a record setting 197Mpps when running an open source Data Path Development Kit (DPDK) PCIe (Gen 4.0).
  • 【Data Rate】:Dual QSFP28 Ports(10GbE/25GbE/40GbE/50GbE/100GbE) and EDR let you connect to network cable for meeting the demands of data center environments.PCIe v4.0 (16.0GT/s) x16(Compatible with 2.0/1.1/3.0); X16 Lane.
  • 【Technical Support】:iPXE, DPDK, iSCSI, UEFI, TCP/IP, UDP/IP, Jumbo Frames, RDMA(RoCE v1, RoCE V2),ASAP², VMDq, SR-IOV, RSS, IPsec, IB, IEEE1588.
  • 【Supported Operating Systems】:Windows; Windows Server; Linux Stable Kernel version; Ubuntu; Vmware ESXi; Citrix XenServer; Deepin; RHEL/CENTOS; Freebsd; OFED AND WINOF-2; Mikrotik; Debian; BCLINUX; ALIOS; Euler; KYLIN; etc.
  • 【I/O virtualization, multi-VM support】:SR-IOV technology enables efficient management of I/O resources of virtual machines by sharing physical resources. And Infiniband technology fully meets the needs of high bandwidth and low latency in big data, its aggregation on virtual I/O and flat network architecture provide a huge pipeline that can be dynamically distributed on demand to improve availability and load balancing.

Why a p95 can be misleading

Histogram buckets make the percentile an estimate

A histogram stores observations in buckets, not as every original value. A percentile calculated from it is therefore estimated. Bucket boundaries, bucket resolution, and assumptions about where observations lie inside a bucket affect the result. If p95 is near a service threshold, coarse bucket placement can shift the estimate enough to change how the result appears relative to that threshold. Finer native-histogram resolution can narrow the uncertainty, but does not turn a histogram estimate into an exact observation. Prometheus explains the interpolation assumptions and resolution trade-offs in its histogram guide and query functions reference.

Low request volume weakens tail readings

Sample count matters. A percentile from a sparse interval has less information behind it, and histogram bucket placement may make p95 and p99 resolve to the same bucket. Google Cloud gives an example in which fewer than 20 samples put those two percentiles in the same bucket; that is an example of the limitation, not a universal minimum sample threshold. Google Cloud’s distribution-metrics documentation explains how bucket count and width, the measured distribution, and sample count affect percentile estimates. Display request count or another volume indicator next to the percentile so a quiet interval is not mistaken for a precise reading.

How to compare p95 values fairly

Two p95 numbers are not meaningfully comparable unless they describe compatible measurements. Check these dimensions before concluding that one service, region, or release is faster:

  • Request population: compare the same endpoints or request classes, and make clear whether each value is per-instance, per-region, or fleet-wide.
  • Measurement boundary: distinguish client-side latency from server-side latency. Client-side collection can capture user-affecting delays that server-side measurements miss, while load-balancer metrics have their own measurement boundaries. See Google’s SRE guidance and Google Cloud’s load-balancing metric documentation.
  • Time interval and aggregation: use comparable windows and the same aggregation approach; a short spike and a long-window percentile answer different questions.
  • Traffic volume: note request counts, since estimates from very different sample sizes can have different reliability.
  • Estimation method: check histogram resolution and percentile calculation method before comparing values near a threshold.
  • Relevant threshold: relate the percentile to a threshold meaningful for the service’s users, rather than treating a lower number as automatically better in every context.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use p95 carefully in an SLO

A p95 line should not stand in for both ordinary performance and severe tail behavior. Google Cloud distinguishes request-based latency objectives, which count the share of requests meeting a threshold, from window-based objectives, which count intervals meeting a condition. Its guidance says percentile-group data is a case for a window-based SLO. A separate tail-focused objective can complement a typical-performance objective; set thresholds and compliance periods according to user needs rather than treating vendor examples as universal targets. See Google Cloud’s latency SLO documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Zunate USB Mouse Receiver 2.4G Wireless Adapter for Mouse
  • HIGH COMPATIBILITY: Designed specifically fit for Spectrum wireless mouse, ensuring excellent functionality and seamless integration.
  • STABLE WIRELESS TECHNOLOGY: Equipped with advanced 2.4GHz wireless technology, this adapter ensures a stable and reliable signal transfer for uninterrupted performance.
  • EASY REPLACEMENT SOLUTION: This 2.4GHz mouse receiver provides a simple solution to replace lost original receivers, allowing you to maintain productivity without any disruptions to your workflow.
  • COMPACT AND PORTABLE DESIGN: The receiver's compact design makes it easy to store and carry, ensuring you can take your wireless setup wherever you go without the hassle of bulky accessories.
  • RELIABLE DESIGN: Each mouse receiver undergoes comprehensive factory testing to meet strict quality standards, ensuring a reliable and professional user experience.

A practical way to read a p95 dashboard

  1. Identify what is measured. Confirm the request population, latency boundary, and whether the value is for an instance, endpoint, region, or fleet.
  2. Check the window and volume. Read the time range and request count together; treat a sparse interval with caution.
  3. Confirm the calculation. If you need fleet p95, aggregate compatible histogram observations before calculating the quantile—do not average server p95 values.
  4. Account for estimation. Inspect bucket resolution when the value is near a threshold or changes in a way that seems inconsistent with the underlying behavior.
  5. Investigate the tail. Compare useful dimensions and consult additional distribution views or request-level evidence to learn how severe the slow requests are.
  6. Judge against the right objective. Use an SLO that matches the user-facing question, and do not make one percentile carry every measure of service quality.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.