Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Removing 40% of a dataset’s rows does not prove that those rows were bad data. In my own account, I can substantiate that 27% of the removed rows were wrong; the rest are not proven wrong simply because I deleted them. Without the dataset, denominator, and supporting evidence, those percentages remain a personal claim—not a general error rate.
The practical lesson is to treat deletion as a decision that needs evidence, a stated purpose, and a record others can check. A suspicious row may be an error, a valid but unusual observation, or a case that cannot now be resolved.
What do the 40% and 27% figures actually tell us?
The figures describe one person’s cleaning experience. They do not establish how often data cleaning removes valid rows in other datasets. The year, dataset, denominator, deletion criteria, and evidence behind the claim are not specified here, so the percentages cannot be independently evaluated or generalized.
There is also an important distinction between “not proven wrong” and “proven correct.” If evidence for a deleted row is missing, its status may be unresolved. That uncertainty should be reported as uncertainty—not silently counted as either a confirmed error or a valid record.
#1 Best Overall
How do I know if a data row is wrong?
Start with the purpose of the analysis. The U.S. Government Accountability Office’s guide to assessing data reliability frames reliability in relation to accuracy, completeness, and applicability for the intended purpose. A value can be unusual without being inaccurate, and a record can be accurate but unsuitable for a particular analysis.
Use checks to flag records for investigation, not as automatic proof that every flagged row should be deleted. Depending on the data, checks might look for invalid ranges, duplicate records, schema violations, failed linkage criteria, or conflicts with known domain constraints. The right rules depend on what a valid record means for the task.
Classify what the evidence supports
- Confirmed error: Evidence shows the record is incorrect under a defined rule or against a reliable source.
- Suspected anomaly: A check has raised a concern, but the available evidence does not establish that the row is wrong.
- Unresolved case: There is not enough evidence to determine whether the record is valid.
- Valid but unusual: Review supports keeping the record even though it differs from typical observations.
This is a practical working distinction, not a universal formal standard. Keeping categories separate prevents a flag from being mistaken for a finding.
Should I delete outliers from my dataset?
Not solely because they are outliers. First ask whether the value is possible and relevant for the analysis, then check whether it reflects a data-entry problem, a measurement issue, or a real but uncommon event. Removing genuine extremes can change the population represented by the data and alter results; retaining errors can do so as well.
Choose the action that matches the evidence and the consequences. Depending on the case, that may mean correcting a value from a reliable source, retaining it, excluding it under a documented rule, or marking it unresolved and assessing how it affects the analysis. If a decision depends on a threshold, make that threshold explicit and consider whether conclusions change under reasonable alternatives.
Why deletion has two kinds of risk
Deleting an invalid record can prevent it from distorting an analysis. Deleting a valid one can exclude relevant evidence. The balance between those risks depends on the use case: an error that matters greatly in one decision may have little effect in another.
Rank #3
Data linkage provides a useful analogy, but its terminology applies directly to matching records, not automatically to every cleaning task. In linkage, a false match joins records that should not be joined, while a missed match fails to join records that belong together. The ONS guidance on precision and recall in data linkage explains why quality cannot be reduced to one number. For linkage, precision concerns the proportion of proposed matches that are correct; recall concerns how many true matches were found. In other cleaning tasks, choose measures that reflect the specific errors at stake.
Automated rules can apply consistent checks at scale, while targeted human review can help resolve ambiguous cases. Statistics Netherlands describes automatic detection and correction alongside selective manual editing. The ONS notes that clerical review is laborious as a primary linkage method, but can help with ambiguous cases and checking proposed matches. These are examples of approaches, not a universal threshold for deciding which rows to review.
How do I document data cleaning?
Keep enough information for another analyst to understand what changed and why. The record should distinguish the evidence available at the time from later interpretations, and should make it possible to trace a decision back to its inputs and rule.
Rank #4
- Python Data Science Handbook
- Preserve a read-only original or a recoverable version before transformations.
- Record the dataset source, version, and date received.
- Describe each rule, its purpose, and the code or procedure used to apply it.
- Log the action for each affected row or group: retained, corrected, excluded, or left unresolved.
- Document the rationale, supporting evidence, reviewer decision, and disagreements where relevant.
- Report how many rows were flagged, reviewed, corrected, excluded, or left uncertain, using clearly defined denominators.
The older GAO-03-273G guidance—now superseded—states that “Reliability does not mean that computer-processed data are error-free.” It also emphasizes documenting assessment work and warns that deleting original files can leave reliability undetermined. Those points are useful context, while GAO-20-283G is the current guide cited above.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What if I deleted valid data?
Recover the original or a versioned copy first, then identify which rows were removed and which rule removed them. Reassess the cases using the original evidence and the analysis purpose. If records were incorrectly excluded, restore them and rerun affected analyses; record the correction and its impact rather than overwriting the history of the earlier decision.
If the original records or evidence no longer exist, do not claim certainty that the deleted rows were invalid. Report that their status cannot be verified, explain what files or corroboration are missing, and describe how that limitation affects the conclusions. The superseded GAO guidance specifically notes that loss of original files can prevent a reliable assessment.
Best Value
How can I tell whether data cleaning introduced bias?
Check whether exclusions are concentrated in particular groups, time periods, sources, or value ranges. Compare retained and removed records on relevant characteristics where possible, and assess whether the decision rule affects groups differently. If the evidence is available, rerun the analysis with plausible alternative treatments of uncertain cases—such as retaining them, excluding them, or applying a justified correction—and compare the conclusions.
For linkage, examine both false matches and missed matches rather than relying on one accuracy figure. For other cleaning tasks, define the error types first and choose measures that capture both incorrect retention and incorrect exclusion. A clinical-research review found substantial variation across data-processing methods and stressed reporting the measured error rate, uncertainty, and method. Its clinical setting and error definitions do not provide a benchmark for general row deletion.
The ICO’s guidance on accuracy and statistical accuracy concerns AI and data protection, rather than general-purpose cleaning rules, and the page says its guidance is under review. It may be relevant to those specific contexts, but it does not establish a universal standard for deciding when a row should be removed.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

