Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Hard negative mining means choosing training examples that a model finds hard to tell apart from the correct match, while those examples are still wrong under the task’s own relevance rule. Near-miss examples like these show a retriever or reranker where the boundary between “relevant” and “almost relevant” lies. In the sources covered here, the trained model is a neural retriever, reranker, or embedding model, not a general-purpose LLM. LLMs appear mainly as tools that generate or label negatives. None of the work reviewed shows that the technique improves an LLM’s general reasoning.

The sequence below follows what these papers imply: define what counts as relevant, mine candidates that look right, remove the ones that are secretly correct, check that generated examples do not hand the model an easy shortcut, and then measure the result against a baseline on held-out retrieval or ranking tests.

What “almost right” means in hard negative mining

The general definition comes from Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka’s 2020 arXiv paper Contrastive Learning with Hard Negative Samples. Its abstract argues that “as with metric learning, contrastive learning of representations benefits from hard negative samples (i.e., points that are difficult to distinguish from an anchor point).” The key phrase is difficult to distinguish from an anchor point. A hard negative sits close to the anchor in the model’s view, yet it is labeled as a non-match.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In retrieval and reranking, the working definition is narrower and more practical: a hard negative is a page that ranks highly for a query but is irrelevant to it. Wasserman et al. frame the problem this way in DocReRank, presented at EMNLP 2025, which targets training multimodal retrieval-augmented generation (RAG) rerankers (ACL Anthology entry). The two definitions are related but not identical. The first describes how representations are trained; the second describes a candidate judged against a relevance rule. When you write a training specification, the second is the one you must verify.

An illustration, not an experimental result: suppose a query asks how reciprocal rank fusion combines ranked lists. A passage explaining a weighted-score fusion method shares vocabulary such as “ranking,” “merge,” and “score,” so a retriever may place it near the top. If that passage does not explain how reciprocal rank fusion works, it is a valid hard negative. If another passage in the corpus does explain it, that passage is a false negative and must not be used as a negative.

What the model is actually trained to do

The phrase “teaching an LLM” in this topic covers a real but narrower setup. The training objective in these papers is a contrastive or ranking loss that rewards placing the positive above the negatives for a query. The model being trained is a retriever or a reranker. When an LLM enters the pipeline, it has one of three jobs:

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
  • Generating a query for a page, so the page becomes a negative for a new query that is similar in form but not answerable from it (DocReRank).
  • Producing a passage-grounded answer and judging answerability, so ambiguous candidates can be relabeled or excluded. The ARHN paper uses open-source LLMs for this step.
  • Synthesizing negatives from a positive, with controls intended to violate a specific query requirement (Zhang et al., 2026).

The papers do not support using this technique as a route to general reasoning gains in a chat model. The “almost right” in this topic refers to passages and documents judged against a query.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where hard negatives come from

Three candidate sources appear in the literature. They trade cost, control, and risk differently.

In-batch and random negatives

Contrastive training commonly uses other examples in the same training batch, or random examples from the corpus, as negatives. They need no mining step and are cheap to produce. Because they are usually unrelated to the query, the model can reject them easily. They set a floor for difficulty rather than a boundary.

Passive mining from a retriever’s top results

Run the current retriever over the corpus, take the high-ranking results for each query, remove the labeled positives, and keep the rest as candidates. This is the most direct form of “almost right,” because the retriever itself surfaced these passages. DocReRank calls this passive mining, and its limits are covered below.

Generated negatives

Start from a positive page or passage and generate a negative. DocReRank generates a query that the page cannot answer. Zhang et al. generate passages that deliberately violate a stated requirement. Generation gives control over hardness and diversity, but it adds a verification burden.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Option Where candidates come from Control over hardness and diversity Ability to catch false negatives Main risk Added cost
In-batch or random negatives Other examples in the training batch or corpus Low; examples are mostly unrelated and easy to reject Low chance of a hidden match, but not checked Teaches little about the relevance boundary Lowest extra work; no retriever pass or generation step
Passive mining (DocReRank’s description) Top-ranked results from the current retriever over the corpus Limited to what the corpus contains; the paper describes limited diversity and insufficient hardness Requires a relevance check on each candidate; the paper describes frequent false negatives The retriever’s blind spots become the training set A full pass of the current retriever over the corpus; cost not stated in the cited paper
Generated from a positive (DocReRank, 2025) A query generated for a page the query cannot answer High; the paper reports diverse, targeted negatives The paper includes false-negative verification Generated text may drift in form or context Generation model calls plus verification; cost not stated in the cited paper
LLM-generated negatives with requirement violations (Zhang et al., 2026) Generated passages that violate a specific query requirement High in principle; the paper proposes counterfactual perturbations Depends on the checks applied; the paper warns about degraded retrieval when checks are weak Source-identity shortcuts and a generative-discriminative gap Generation plus training-time controls; cost not stated in the cited paper

How to mine hard negatives: a working sequence

One practitioner asked on a Reddit machine-learning learning thread how to mine hard negatives before feeding data to a model, and whether there are constraints on doing so (the discussion thread). That is a single community example, not a measure of how common the question is. The answer is a sequence of checks rather than one algorithm.

  1. Write the task and the relevance rule. State the retrieval or ranking task, the model being trained, and what a passage must contain to count as relevant. Without this rule you cannot tell a hard negative from a false one.
  2. Choose a candidate source. Use in-batch or random negatives as a floor, mine the current retriever’s top results for hard cases, or generate targeted negatives from positives when the corpus lacks boundary cases. The papers do not establish a single best mix.
  3. Remove or relabel likely false negatives. Read or score high-ranking candidates for alternate relevance, partial answers, and annotation gaps. ARHN’s answerability workflow, described below, is one automated way to do this.
  4. Check generated examples for a named violation. A generated negative should fail one specific requirement of the query. If it only changes the topic or rewords the positive, discard it. Also look for style or source markers the model could use as a shortcut.
  5. Train and record the setup. Keep the training data, sampler settings, and number of negatives per query so the run can be reproduced.
  6. Evaluate on held-out retrieval or ranking data against a baseline. Compare with a model trained without the mined negatives, or with your previous recipe, on queries and documents that were not used for mining. Report the metric, the test set, and the baseline together.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Failure modes and the proposed fixes

False negatives: hard does not mean good

A mined candidate can be relevant, only partly answer the query, or contain the answer even though the training labels call it irrelevant. Training on such a passage gives contradictory supervision: similar content is rewarded in one place and penalized in another. The ARHN paper by Choi et al., an arXiv preprint dated April 13, 2026 (ARHN: Answer-Centric Relabeling of Hard Negatives with Open-Source LLMs for Dense Retrieval), describes these ambiguous negatives as a source of noisy, inconsistent supervision.

Its proposed workflow generates a passage-grounded answer signal, ranks candidates by answerability, relabels any passage ranked above the original positive, and excludes answer-bearing passages from the negative set. The paper presents this as its approach for dense retrieval. It is not a guaranteed fix for every corpus, so check the relabeled set yourself before training on it.

Passive mining limits

DocReRank describes passive mining as restricted to what the retriever can find in the available corpus. It names four limits: limited diversity, insufficient hardness, low controllability, and frequent false negatives. Its alternative starts from a page and a positive query, then generates a query that is similar in form and context but not answerable from that page. The paper reports that this supports diverse, targeted negatives and false-negative verification. Those results are the paper’s own and were measured on multimodal RAG rerankers, so do not assume they transfer directly to a text-only dense retriever.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generated negatives and shortcut learning

Zhang et al.’s arXiv preprint, dated May 31, 2026 (When Hard Negatives Hurt: Bridging the Generative-Discriminative Gap in Hard Negative Synthesis for Retrieval), warns that LLM-generated negatives can degrade retrieval in two situations. The first is when generation is generic or drifts in topic. The second is when training can tell examples apart by their source rather than their relevance. That second case is a shortcut: the model learns “generated versus real” instead of “relevant versus not.” The paper proposes counterfactual perturbations that explicitly violate a query requirement, along with query-view entropy maximization to reduce source-identity shortcuts. Because the paper is recent, treat this as an emerging method rather than an established practice.

Constraints on mining hard negatives

The direct answer to the practitioner’s question about constraints is that there are few fixed numeric rules, but several conditions must hold:

  • A written relevance rule for each query type. Without it, candidates cannot be classified.
  • Hardness is relative to the current model, the data, and the relevance definition. If the candidate pool is built from model output, re-mine after major retraining, because a passage that was hard for one checkpoint may be easy for the next.
  • No single hardness threshold or negative count has been established as universally best. Set them by evaluation on your target task.
  • False-negative checks before training, not after. Exclude or relabel any candidate that answers the query.
  • Keep held-out evaluation queries and documents out of the mining pool, so the test measures generalization rather than memorization.
  • Any performance claim needs the dataset, metric, and comparison baseline. This article does not quote benchmark numbers from these papers; cite results only with that context attached.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.