What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To find missing data in a time series, check both for null values in existing rows and for timestamps that should exist but do not. The second check requires a known sampling schedule; an irregular event stream has no fixed set of expected timestamps unless you define a business rule for expected events.

What counts as missing data in a time series?

Missingness has two distinct forms, and a dataset can have either or both:

  • Explicit missing values: A timestamp row exists, but one or more measurements are null, such as NaN, NaT, or None.
  • Implicit missing observations: A measurement row is absent altogether, so an expected timestamp is missing from the dataset.

A null scan will not find a timestamp that has no row. Conversely, comparing timestamps with an expected schedule will not reveal a null value in a row that is present. Run both checks.

Prepare timestamps before checking for gaps

Establish what the data represents

Identify the timestamp column, measurement columns and units, timezone, entity or sensor key, and documented collection schedule. Keep an unchanged copy of the raw data so you can audit parsing or cleaning decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

Normalize and validate time

Parse timestamps, convert them to one explicit timezone, review parse failures, and sort by entity and time. Check for duplicate timestamps within each entity: duplicates can distort counts and obscure comparisons with the expected schedule. If the source schedule is defined in UTC, analyze timestamps in UTC. Daylight-saving-time transitions can repeat or skip local clock times, which can otherwise look like duplicate or missing observations.

Check for null measurements in rows that exist

In pandas, use isna() or notna() to identify missing values. These methods recognize missing-value sentinels according to the data type, including NaN, NaT, and None. Do not test missing values with equality comparisons: pandas notes that these sentinels do not compare equal to themselves. See the pandas missing-data documentation.

For each measurement column, report the null count and its share of observed rows. Calculate the share as null rows divided by the total number of observed rows, and state the denominator; it describes values in rows present, not absent timestamps. Keep counts by entity as well as overall when the dataset contains multiple sensors or other series.

Find timestamps that should exist but do not

Use a declared cadence, not a guess

Choose the expected interval from the source specification—for example, every five minutes, hourly, or daily. Inferring a frequency from an incomplete file can mistake the gaps you are trying to find for the normal schedule. If the data is event-based or intentionally irregular, define which events should have occurred using the relevant business rules instead of generating a fixed-frequency index.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compare expected and observed timestamps

For each entity, build the expected timestamp sequence over the period being checked, then compare it with that entity’s observed timestamps. In pandas, DatetimeIndex and date_range can represent the expected schedule; comparing the expected and observed indexes exposes absent timestamps. reindex or asfreq can align observations to a declared frequency and make gaps visible. Review the pandas time-series documentation for these frequency and alignment tools.

Define the start and end of the interval you expect to cover before counting gaps. An absent timestamp inside that interval is different from a file that simply begins late or ends early. Record leading and trailing truncation separately from internal gaps.

Classify and investigate the gaps

Turn absent timestamps into auditable gap records rather than treating them as an undifferentiated missing count. Group consecutive absent timestamps into runs and record:

  • the run’s first and last expected timestamps and its duration;
  • the affected entity or sensor and measurement columns;
  • whether it is an isolated point, a contiguous outage, or a recurring calendar pattern;
  • whether it is internal to the expected range or reflects leading or trailing truncation;
  • a status such as expected, unknown, or suspected failure, with the reason for that label.

Compare each run with maintenance logs, holidays, operating hours, sensor state, ingestion jobs, and timezone changes. A gap may be valid by design; do not label every absent timestamp as a system failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate the missingness pattern

Combine numerical summaries with plots. Summarize null counts and absent timestamps by day, week, month, and entity; plot the series with missingness markers; and inspect distributions before and after gaps. NIST recommends graphical and numerical checks when looking for data-quality problems, including unusual values and missing data. Its guidance includes scatter plots, histograms, and numerical summaries in its exploratory data analysis discussion.

A lag plot can help assess serial correlation, randomness, and outliers, but it does not replace the timestamp comparison. NIST’s univariate time-series guidance assumes equally spaced observations and says irregularly spaced analysis is outside that section’s scope. For that guidance, see NIST’s time-series analysis section.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Decide whether to leave gaps or fill them

Finding missing data does not mean it should automatically be imputed. Preserve a missingness flag and document the reason for any treatment. The appropriate choice depends on the cadence, gap length and pattern, domain limits, confidence about the cause, and how the result will be used.

  • Leave values missing: Often preferable when the cause is unknown or filling would imply measurements that were never made.
  • Delete rows or intervals: Consider only when removing them will not bias the analysis or discard important time periods.
  • Forward-fill or backward-fill: Can suit values that are expected to remain valid until updated, but may create artificial flat periods or use future information.
  • Interpolate: May be reasonable for short gaps in a smoothly changing signal; it can conceal abrupt changes and is not appropriate for every variable.
  • Use model-based imputation: Can account for relationships among observations or variables, but depends on assumptions that need validation.

Scikit-learn defines imputation as inferring missing values from known data; its imputation guide describes available methods. Evaluate a proposed method by hiding some observed values, imputing them, and comparing estimates with the known values—or use a domain-specific validation rule. Treat interpolation, filling, and imputation as downstream analytical choices, not as part of detecting gaps.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to report for a dataset

There is no universal percentage that describes how much time-series data is missing. Report figures calculated for the dataset being analyzed, with its organization and extraction year identified where relevant. A useful summary includes:

  • observed row count and explicit null count by measurement column;
  • the expected schedule, expected timestamp count, and time range used;
  • the number of absent timestamps and grouped gap runs, by entity;
  • null and gap rates with their denominators clearly stated;
  • how leading and trailing truncation, expected closures, and unknown gaps were classified.

Pandas APIs and behavior are versioned; consult the documentation for the version used in your environment.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.