Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To reduce popularity bias and repetitive recommendations, first identify whose experience is harmed and how; then intervene at the data, model, preference-elicitation, ranking, or session level. Measure relevance alongside discovery and repeated exposure, and test the changes with users. Popular items are not automatically a problem: the issue is when popularity-driven exposure limits a system’s value or harms a stakeholder.

What popularity bias is—and when it matters

High interaction counts can reflect genuine quality, broad appeal, price, promotion, or simply that an item has been shown more often. Recommendation logs may therefore capture both user preference and prior exposure. If recommendations generate the interactions later used to train the system, an early exposure imbalance can reinforce itself: popular items receive more recommendations, accumulate more interactions, and become still more likely to appear.

A 2024 survey on popularity bias in recommender systems defines the problem in terms of impact: popularity becomes bias when recommendations focus on popular items so much that they limit system value or cause harm to stakeholders. That means the right question is not simply “Are popular items appearing?” but “What useful options or fair opportunities are being crowded out, for whom, and with what evidence?”

Define the harm before choosing a fix

Depending on the application, the problem might be less discovery of relevant niche items, repeated exposure to near-identical choices, reduced catalog coverage, or limited exposure for providers. State the system’s purpose, the stakeholders who could be affected, and the observable change that would count as improvement. There is no evidence-backed universal popularity cutoff, diversity quota, or repeat limit that suits every recommender.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why recommendations keep repeating

Skewed data and exposure

Interaction histories are not necessarily a neutral record of what people would choose if they had seen the full catalog. Logging coverage, candidate generation, where an item appeared, and how often it was exposed can all shape which items collect feedback. Popularity may be a real signal, but counts alone do not distinguish that signal from accumulated opportunity.

Feedback from recommendations

When a system’s output affects what users can interact with—and those interactions return as training data—the model can amplify its own previous choices. Repetition can also happen within a single list or between sessions, even when each individual recommendation appears relevant. Auditing only one ranked list will miss patterns that emerge over time.

Where to intervene

Mitigation can happen before training, during model learning, while eliciting preferences, after candidate scoring, or across a session. The right point depends on whether the diagnosed issue is data representation, the learned objective, limited knowledge of user interests, list composition, or repeat exposure.

Intervention stage What to change Key trade-off or caution
Before training Audit representation and logs; consider reweighting or otherwise addressing data skew. Do not remove popular items indiscriminately: their popularity may reflect real quality or preference. (2024 survey: Springer Nature)
During learning Incorporate popularity-aware regularization, constraints, or joint objectives. Tune the intervention against relevance and the specific harm being addressed. In-process approaches are the most common mitigation type in the surveyed literature, not a guarantee of the best result for every system. (2024 survey: Springer Nature)
Preference elicitation Explore a broader range of interests instead of asking only about likely hits; multi-armed bandits are one studied method. Exploration changes what the system learns from the elicitation process. Google Research authors reported in 2021 that popularity bias in preference elicitation contributes to popularity bias in recommendation. (Google Research)
After candidate scoring Rerank with list-level or feature-level diversification, novelty, or exposure objectives. Keep an application-appropriate relevance floor. Diversity can trade off against accuracy, and greater diversity does not necessarily mean greater serendipity. (2024 survey: Springer Nature; Kotkov, Veijalainen, and Wang: SOG article)
Across a session Track repeat exposure and similarity between consecutive lists; consider adjusting diversity over time. Validate beyond simulation before relying on reported engagement gains. (Li et al., 2026: ACM Transactions on the Web)

A practical sequence for reducing bias and repetition

  1. Specify the application-specific problem

    Choose the outcome to improve—for example, more relevant long-tail discovery, less repeated exposure, or a fairer distribution of item exposure. Identify the affected users or providers and define what evidence would demonstrate improvement.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  2. Audit popularity, exposure, and feedback

    Measure the popularity distribution in both training data and recommendations. Check logging and catalog representation, candidate generation, item position or exposure effects, and whether system outputs feed into future training interactions. Separate popularity that plausibly reflects user preference or item quality from popularity amplified by earlier exposure.

  3. Match the intervention to the cause

    Use the intervention-stage table to choose a starting point. A reranker cannot repair missing or poorly logged feedback by itself; data changes alone may not stop a feedback loop; and preference exploration addresses what the system learns about interests rather than directly controlling every final list.

  4. Evaluate several outcomes together

    Compare relevance or accuracy with intra-list diversity, novelty, serendipity, catalog coverage or exposure across popularity groups, and repetition across time. Segment results by user group, item popularity, or stakeholder when those distinctions matter. Compare alternatives on the same data and interaction horizon; there is no universal metric or threshold for success.

  5. Test with people, not only offline metrics

    Use offline evaluation to screen approaches and make results reproducible, then use human evaluation, experiments, or field studies to establish whether people find the recommendations useful and less repetitive. The 2024 survey reviewed 123 papers and found that computational experiments dominate, while human-in-the-loop and field evaluations are comparatively rare. Offline improvement alone does not establish user benefit.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to judge diversification methods

List-level diversification

Reranking can reduce near-duplicates or balance item features within a recommendation list while preserving a relevance floor. A serendipity-oriented greedy algorithm (SOG) is one published example: Kotkov, Veijalainen, and Wang’s article, first published in 2018 and appearing in a 2020 journal volume, reports better diversity and serendipity than other algorithms in its evaluation, alongside a trade-off against accuracy-oriented algorithms. Treat that as a result for the studied setting, not a guarantee for another catalog or audience.

Dynamic control over time

A 2026 ACM paper on dynamic, fine-grained homogeneity control reports a 4.35% increase in average session length and a 25.65% increase in long-term engagement over state-of-the-art baselines in the KuaiRand simulated environment. The authors’ findings are simulation results, not live-deployment outcomes. The study suggests that moderate homogeneity may help early in an interaction and become harmful later, but a production system should validate any time-varying policy with users and real-world evidence before treating those reported gains as expected performance.

Common mistakes to avoid

  • Equating popularity with bias: popular recommendations may be appropriate when they match user interests and serve the application’s purpose.
  • Removing popular items without diagnosis: this can suppress valid preference or quality signals rather than correct an exposure problem.
  • Optimizing diversity alone: additional variety can reduce relevance, and variety is not automatically useful discovery or serendipity.
  • Claiming success from one metric: report relevance alongside discovery, exposure, and repeat behavior over the relevant time horizon.
  • Treating simulations or offline metrics as user proof: use user studies or field evaluation to establish real-world value.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.