Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A semantic similarity score of 0.87 does not prove that two records describe the same person, organization, or other entity. It is a threshold applied to a score whose meaning depends on the model, comparison method, data, and task. A pair can look semantically close while conflicting on identity-defining details; a real match can also score lower because its identifiers are incomplete or outdated.

What a 0.87 threshold actually tells you

A threshold is a decision boundary: a system uses it to decide which record pairs to treat as links, candidates, or non-matches. The score is not automatically a probability, confidence level, or measure of identity. Without knowing how a particular model produces its score and how that score was evaluated on the target data, 0.87 has no universal interpretation.

The UK Government’s data-linkage quality guidance puts the general point plainly: “In all linkage methods, some choice must generally be made about an evidentiary threshold for classifying record pairs as links or not.” The threshold is part of a decision process, not proof that the decision is correct.

Why similar records can still be different entities

Semantic matching identifies records with similar meaning or attributes. Identity resolution asks a stricter question: does the available evidence distinguish the entity the task is about? Common descriptions, shared locations, or similar names may make two records resemble one another without making them the same entity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, two organization records might describe similar services but contain conflicting legal identifiers or locations. That hypothetical illustrates why a high similarity score should be checked against fields that matter to identity; it is not a reported case from the cited sources. Conversely, genuine matches can score lower when identifying details are missing, misspelled, changed over time, or stale. The Government guidance notes that linkage errors can occur regardless of method and depend on the quality and completeness of identifying data.

False links and missed links have different costs

A false link joins records that do not refer to the same entity. A missed link leaves apart records that do. The Government guidance notes that false links can result when identifiers are shared or fail to distinguish entities; missed links can result from errors, changes, or missing or weak identifiers.

Two measures help describe these different errors:

  • Precision asks what proportion of assigned links are true. Low precision means more false links among the pairs the system accepted.
  • Recall asks what proportion of true matches were identified. Low recall means more genuine matches were missed.

Raising a threshold may reduce false positives while excluding valid matches, but the actual trade-off depends on the score system and task. A broad candidate-discovery workflow may tolerate lower precision if people review the candidates later. A sensitive process that creates merged records may place greater weight on avoiding false links. There is no context-free best threshold.

What published threshold results do—and do not—show

A 2026 study in Frontiers in Artificial Intelligence, “Detecting reconciliation discrepancies in tabular data using transformers,” reports experiments involving 185,909 tables for its proposed semantic tabular-reconciliation method. Its reported metrics vary by task, so they are evidence about that method and evaluation—not validation of 0.87 for an unspecified system or dataset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Data Quality Assessment
  • Used Book in Good Condition
Study evaluation Reported threshold result What it applies to
Large-scale relationship-identification experiments At τ=0.9, precision 0.958; F1 scores 0.77–0.87 The study’s relationship-identification experiments
Representative discrepancy-detection case At τ=0.7, precision 0.91, recall 0.91, and F1 0.912 The study’s representative case
Representative discrepancy-detection case At τ=0.8, recall 0.79 and F1 0.857 The same representative case
Representative discrepancy-detection case At τ=0.9, precision 0.958 and recall 0.676 The same representative case; stricter threshold, higher precision and lower recall than at τ=0.7

The two appearances of 0.958 precision refer to different evaluations: one is from relationship-identification experiments, and the other is from the representative discrepancy-detection case. They should not be combined into a single result. The study does not establish that a score of 0.87 is appropriate for a different model, data population, or downstream decision.

How to evaluate a linkage threshold for your data

  1. Define the downstream decision. Decide what a link will cause—such as a candidate being queued for review or records being merged—and whether false links or missed links are more harmful in that use.
  2. Test representative labeled pairs. Use pairs from the population the system will actually process. Report precision and recall, then inspect the types of errors instead of relying on the cutoff alone. The Government guidance recommends making linkage decisions in light of the intended analysis.
  3. Separate candidate discovery from acceptance when needed. A similarity score can identify pairs worth checking without being sufficient evidence for a final link. Add exact or otherwise discriminative evidence where the task calls for it.
  4. Keep uncertainty usable. The Government guidance recommends retaining less-than-certain links and providing link-level measures. That lets end users adjust decisions and conduct sensitivity analysis instead of forcing every pair into a definitive yes-or-no result.
  5. Check clusters, not only pairs. A series of plausible pairwise links can create a transitive group whose endpoints may not be sufficiently supported. Review the resulting groups against the same identity evidence used for individual links.
  6. Re-evaluate after meaningful changes. New data, a different score construction, or a changed downstream use can alter what the same numeric cutoff means in practice.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What matching tools can make explicit

AWS’s documentation for creating a rule-based matching workflow with the Advanced rule type describes combining exact and fuzzy conditions and documents transitive matching as a capability available through the API. This is an implementation example, not evidence that the feature guarantees correct identity resolution. When evaluating any approach, compare the evidence it uses, how it handles uncertainty, whether it forms transitive clusters, and whether its reported evaluation matches your task and data.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.