Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HyperLogLog (HLL) estimates how many distinct values appear in a set or stream without keeping a complete list of those values. It replaces exactness with compact state, making it useful for large-scale counts such as daily unique page visits—provided an estimate is sufficient for the job.

What cardinality means—and what HyperLogLog returns

Cardinality is the number of distinct elements in a collection or data stream. If a page receives many visits, its daily cardinality is the number of different visitors, not the total number of visits. HLL produces an estimate of that distinct count; it does not preserve the member list or provide an exact count. Google Research describes cardinality estimation as determining the number of distinct elements in a data stream (Google Research).

That distinction determines whether HLL fits: it can answer “how many?” approximately, but it cannot tell you which users were counted or verify membership. Redis gives examples including daily unique visits, unique users who played a song, and unique viewers of a video (Redis HyperLogLog documentation).

How the estimate works

At a high level, an implementation hashes each input value, assigns the hash to one of many registers, and records information about rare patterns—such as unusually long runs of leading zeros—in each register. A new distinct value contributes evidence through its hash pattern. Across many registers, the observed frequency of rare events helps estimate how many distinct values were seen.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is an intuition, not the full estimator. Real implementations may add corrections and change how they represent small or large sketches. Redis, for example, documents sparse and dense representations in its implementation (Redis implementation details). Accordingly, HLL implementations should not be assumed to have identical memory use or error behavior.

How accurate is HyperLogLog?

Accuracy figures depend on the implementation and its configuration. A standard error describes estimator behavior across outcomes; it is not a promise that every individual result will fall within that percentage of the true count.

Rank #2
Sale
Introduction to Algorithms, fourth edition
  • color: White
  • INTRODUCTION TO ALGORITHMS, FOURTH EDITION
Implementation and configuration Documented figure How to interpret it
Redis HyperLogLog 0.81% standard error Redis’s documented figure for its implementation, not a universal guarantee for HLL (Redis).
Apache DataSketches HLL at LgK=14 0.0065 relative standard error (0.65%), calculated as 0.8326 / √(214) A configured figure for DataSketches, not Redis or every HLL implementation. DataSketches also describes confidence contours and cautions that error behavior is not necessarily Gaussian (Apache DataSketches).

These values are useful for understanding particular implementations, not for treating an estimate as an individual-result guarantee. If a count must be exact—for an audit, for example—or you need the underlying members, use a method that retains the necessary exact data rather than relying on an HLL estimate alone.

Why sketches can be combined—and what that does not enable

HLL’s compact summaries support union-style aggregation: sketches representing separate groups of observations can be combined to estimate the distinct count across their union. Redis provides PFMERGE and multi-key PFCOUNT; Apache DataSketches provides an HLL union operator (Apache DataSketches HLL documentation).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Data Structures and Algorithms in Python
  • Used Book in Good Condition

A union is not the same as an intersection or a difference. DataSketches says its HLL sketches do not intrinsically provide those operations because the resulting error would be poor. Do not infer accurate “users in both groups” or “users in group A but not B” results merely from a library’s ability to merge sketches. Alternative estimators have been proposed in research, but that does not make those operations standard, universally implemented HLL features (Research on set-operation estimators).

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Using HyperLogLog in Redis

Redis exposes its HLL functionality through commands. The following example estimates the number of distinct visitor identifiers recorded for a day:

  1. PFADD visitors:2026-10-03 user-17 user-42 user-17 adds identifiers to the sketch for that day. Repeated values do not turn the result into a visit count.
  2. PFCOUNT visitors:2026-10-03 returns Redis’s estimate of the sketch’s cardinality.
  3. PFMERGE visitors:week visitors:2026-10-01 visitors:2026-10-02 visitors:2026-10-03 combines the sketches into a sketch for the union. Then use PFCOUNT visitors:week to estimate the union’s cardinality.

Use a consistent identifier and a clearly defined time window for every sketch you intend to combine. Redis documents a maximum footprint of up to 12 KB per HLL and a 0.81% standard error for its implementation; these are Redis-specific properties, not generic HLL sizing rules (Redis PFCOUNT command reference). Redis describes single-key PFCOUNT as O(1) with a small average constant and multi-key calls as O(N) in the number of keys. Those complexity statements apply to Redis’s command behavior, not to every HLL library (Redis PFCOUNT command reference).

Quick Recap

SaleBestseller No. 2
Introduction to Algorithms, fourth edition
Introduction to Algorithms, fourth edition
color: White; INTRODUCTION TO ALGORITHMS, FOURTH EDITION
$91.50
SaleBestseller No. 3
Data Structures and Algorithms in Python
Data Structures and Algorithms in Python
Used Book in Good Condition
$118.92
SaleBestseller No. 5
Data Structures and Algorithms Made Easy: Data Structures and Algorithmic Puzzles
Data Structures and Algorithms Made Easy: Data Structures and Algorithmic Puzzles
Binding: paperback; Language: english; It ensures you get the best usage for a longer period
$29.41
Best Value
Sale
Data Structures and Algorithms Made Easy: Data Structures and Algorithmic Puzzles
  • Binding: paperback
  • Language: english
  • It ensures you get the best usage for a longer period

When HLL is a good fit

  • Use it when you need approximate distinct counts over large streams and compact summaries that can be combined, such as unique visits or viewers.
  • Check the implementation when memory limits or error tolerance matter. Configuration, estimator behavior, and representation affect the trade-off; DataSketches documents configurable HLL sizes, while Redis documents its own sparse and dense representations.
  • Choose another approach when the task requires an exact audit count, a membership list, or accurate intersection and difference results that your selected tool does not support.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.