Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

News API retrieves headlines; it does not extract keywords for you. A dependable Python workflow is to collect a batch of title values, clean them carefully, rank unigrams and bigrams with TF-IDF, then enrich the result with named entities or noun phrases when your use case needs people, companies, places, or readable concepts.

This approach treats headlines as a corpus instead of pretending that one short headline contains enough evidence for statistical ranking.

Define the output before choosing an algorithm

“Keyword” can mean several different outputs:

  • Keywords: individual terms such as inflation or wildfires.
  • Keyphrases: multi-word concepts such as interest rate or machine learning.
  • Named entities: people, organizations, locations, products, laws, events, dates, and monetary values.
  • Topics: broader themes inferred from many headlines.
  • Tags or search terms: labels chosen for a taxonomy or retrieval system, which are not always the most statistically salient words.

For a dashboard or topic monitor, start with corpus-level TF-IDF and bigrams. For one headline, use entities or noun phrases instead; TF-IDF needs document-to-document comparison.

Choose the News API endpoint

/v2/top-headlines for current batches

Use top-headlines for country- or category-based dashboards, breaking-news displays, and small recent batches. It supports country, category, sources, q, pagination, and a documented maximum pageSize of 100. News API requires an API key. Country and category cannot be combined with sources.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

/v2/everything for analysis corpora

For historical or search-based analysis, everything is usually a better corpus source. It can search fields with searchIn=title, constrain dates and languages, filter domains or sources, and sort by relevance, popularity, or publication time.

params = {
    "q": "artificial intelligence OR machine learning",
    "searchIn": "title",
    "language": "en",
    "from": "2026-08-01",
    "to": "2026-08-18",
    "sortBy": "publishedAt",
    "pageSize": 100,
    "apiKey": NEWS_API_KEY,
}

The dates above are an example, not a permanent query; generate them dynamically in an application.

Fetch and validate headlines

Keep the key outside source control, request a bounded page, and fail clearly on transport or API errors.

import os
import requests

NEWS_API_KEY = os.environ["NEWS_API_KEY"]

response = requests.get(
    "https://newsapi.org/v2/top-headlines",
    params={
        "country": "us",
        "category": "technology",
        "pageSize": 100,
        "apiKey": NEWS_API_KEY,
    },
    timeout=30,
)
response.raise_for_status()
payload = response.json()

if payload.get("status") != "ok":
    raise RuntimeError(payload.get("message", "News API request failed"))

headlines = [
    article["title"]
    for article in payload.get("articles", [])
    if article.get("title")
]

News API article objects include fields such as title, description, url, publishedAt, and content. For this task, select article["title"] explicitly. The documented content value may be truncated to 200 characters, so it is not a substitute for full article text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Clean titles without destroying meaning

Cleaning is a scoring aid, not permission to erase entities. Keep the original title for display and create a normalized copy for matching.

import re

NEWS_STOPWORDS = {
    "says", "say", "said", "report", "reports", "reported",
    "new", "latest", "live", "update", "updates", "breaking",
    "amid", "after", "before", "over", "could", "would", "may",
    "watch", "video",
}

def clean_headline(text: str) -> str:
    text = re.sub(r"[[^]]*]", " ", text)       # [Updated], [Video]
    text = re.sub(r"([^)]*)", " ", text)        # optional labels
    text = re.sub(r"https?://S+", " ", text)
    text = re.sub(r"[^ws'-]", " ", text)
    text = re.sub(r"s+", " ", text).strip().lower()

    return " ".join(
        token for token in text.split()
        if token not in NEWS_STOPWORDS
        and not token.isdigit()
        and len(token) > 2
    )

documents = [clean_headline(title) for title in headlines]
documents = [doc for doc in documents if doc]

Do not blindly remove punctuation: C++, COVID-19, U.S., AI-powered, and S&P 500 can lose their identity. A production cleaner should test representative titles and decide how acronyms, hyphens, dates, and publisher suffixes are handled. General English stop words remove terms such as “the” and “of”; a custom news list is needed for boilerplate such as “says” and “amid”. Words like “war”, “state”, or “trade” may be meaningful in a particular corpus, so review before adding them.

Rank corpus keywords with TF-IDF

TF-IDF raises terms that are frequent in one document but less common across the supplied collection. It ranks distinctiveness within your corpus, not objective importance or newsworthiness.

import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer

vectorizer = TfidfVectorizer(
    stop_words="english",
    ngram_range=(1, 2),
    min_df=2,
    max_df=0.85,
    sublinear_tf=True,
)

matrix = vectorizer.fit_transform(documents)
terms = vectorizer.get_feature_names_out()
scores = matrix.sum(axis=0).A1
ranking = sorted(zip(terms, scores), key=lambda item: item[1], reverse=True)

for term, score in ranking[:20]:
    print(f"{term}: {score:.3f}")

Bigrams preserve meaning that unigrams lose: interest rate, climate change, and stock market are more useful than isolated fragments. Trigrams can be tested with ngram_range=(1, 3), but become sparse in small collections.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Tune document-frequency thresholds

  • min_df=2 requires a term to occur in at least two headline documents. It reduces noise but excludes one-off breaking-news names.
  • Use min_df=1 for a small batch when rare entities matter.
  • max_df=0.85 removes terms appearing in most documents. Adjust it when a common term is actually your target.

With one or two documents, TF-IDF has little statistical value. Collect dozens or hundreds of comparable headlines, or switch to noun phrases, named entities, a fixed vocabulary, or embeddings.

Get keywords for each headline

def keywords_for_document(row_index, top_n=8):
    row = matrix[row_index].toarray().ravel()
    indices = np.argsort(row)[::-1]
    return [
        (terms[i], float(row[i]))
        for i in indices
        if row[i] > 0
    ][:top_n]

for index, title in enumerate(headlines[:5]):
    print(title)
    print(keywords_for_document(index))

These scores show what is distinctive in each title relative to the downloaded collection; they do not measure importance outside that collection.

Add entities and noun phrases

TF-IDF can rank generic verbs above a company or person. A local spaCy model supplies a complementary linguistic signal.

import spacy

nlp = spacy.load("en_core_web_sm")

ENTITY_LABELS = {"PERSON", "ORG", "GPE", "LOC", "PRODUCT", "EVENT", "LAW"}

def extract_entities(text):
    doc = nlp(text)
    return [
        (ent.text, ent.label_)
        for ent in doc.ents
        if ent.label_ in ENTITY_LABELS
    ]

def extract_noun_phrases(text):
    doc = nlp(text)
    return [
        chunk.text.lower()
        for chunk in doc.noun_chunks
        if len(chunk.text.split()) <= 5
    ]

A practical hybrid pipeline extracts TF-IDF terms, entities, and noun phrases; removes overlaps; then boosts entities when the product monitors people, organizations, places, or products. Preserve the original spelling for display while using normalized forms for deduplication. Entity accuracy depends on the model, language, capitalization, spelling, and headline context; short or ambiguous titles can be misclassified.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Deduplicate before scoring

Syndicated stories can make one event appear artificially important. Start with exact and normalized-title deduplication:

unique_titles = list(dict.fromkeys(headlines))

For stronger controls, deduplicate by canonical URL, compare normalized-title similarity, group by source and publication time, or cluster embeddings. Different publishers may describe the same event with different wording, so identical text is not required for a duplicate. If one publisher supplies most rows, calculate source-specific rankings or balance the corpus.

Choose a method for the job

Method Best use Main trade-off
Frequency counts Fast prototype Rewards repeated boilerplate
TF-IDF Explainable corpus ranking Needs multiple documents and is corpus-dependent
RAKE Simple phrase extraction Sensitive to stop words and punctuation
Noun phrases Readable keyphrases Depends on parser quality
Named entities People, companies, places Misses general concepts
TextRank Unsupervised phrases More complex and unstable on short text
Embeddings or KeyBERT-style methods Synonyms and paraphrases More compute and model choices
LLM extraction Structured labels and explanations Cost, latency, privacy, consistency, and evaluation concerns
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Only generic words appear

Expand the custom news stop list, inspect the top 100 terms, raise min_df, and remove duplicated or single-source headlines. Do not permanently remove a word merely because it is common in one category.

Results are one-word fragments

Use ngram_range=(1, 2) or noun chunks. Compare the original phrase with normalized output so readers see “interest rates” rather than an artificial stem.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

The result is empty

Check the API status, key, rate limit, response size, missing titles, and whether cleaning removed every token. Return an empty or low-confidence result for titles such as “Markets react” rather than inventing keywords.

API requests fail

if response.status_code == 429:
    # Back off and retry according to your account and response headers.
    pass

Also handle timeouts, invalid keys, non-ok payloads, malformed JSON, and empty result sets. Retry policy should follow the response and account plan, not an arbitrary fixed assumption.

Dates and languages are inconsistent

News API documents publishedAt timestamps in UTC. Keep UTC for filtering and deduplication, converting only for presentation. English stop words and spaCy’s English model do not generalize to other languages; use language-specific tokenization, models, and stop lists.

Evaluate usefulness instead of trusting plausible output

  1. Collect 50–100 representative headlines.
  2. Have a reviewer mark useful keywords and phrases.
  3. Run the pipeline and measure precision at 5 or 10 results.
  4. Inspect false positives such as attribution verbs and false negatives such as missed entities.
  5. Tune stop words, frequency thresholds, n-gram range, deduplication, and entity weighting.

Repeat this by category, source mix, language, and time period. Vocabulary changes with the news cycle, and a list that looks reasonable can still be unstable or dominated by one publisher.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Production considerations

  • Pagination and caching: store the retrieval time, query, source, and API response so rankings can be reproduced without unnecessary requests.
  • Monitoring: watch for vocabulary drift, empty batches, sudden source concentration, and changes in entity quality.
  • Privacy: keep headline retention and third-party processing aligned with your policy; local spaCy and scikit-learn avoid sending text to a managed NLP API.
  • Rights and access: the article URL may require separate processing subject to publisher terms, robots rules, copyright, access controls, and your News API plan. Do not assume the content field is full text.
  • Commercial use: News API’s pricing page, checked August 18, 2026, lists a $0 Developer plan, $449/month Business plan, and $1,749/month Advanced plan, with different limits, delays, support, and SLA terms. It states that the Developer plan is for development and testing rather than staging, production, or internal production-like use. Recheck current terms and prices before deployment.

For managed enrichment, Google Cloud Natural Language, Amazon Comprehend, and Azure AI Language can provide entities or key phrases, but compare current regional pricing, quotas, language coverage, latency, data handling, and lock-in with local processing before choosing one.

A complete minimal pipeline

import os, re, requests, numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer

API_KEY = os.environ["NEWS_API_KEY"]

def clean(text):
    text = re.sub(r"[[^]]*]|https?://S+", " ", text)
    text = re.sub(r"[^ws'-]", " ", text.lower())
    text = re.sub(r"s+", " ", text).strip()
    stop = {"says", "said", "reports", "reported", "new", "latest", "live", "breaking", "update", "amid", "after", "before"}
    return " ".join(t for t in text.split() if len(t) > 2 and t not in stop and not t.isdigit())

r = requests.get("https://newsapi.org/v2/top-headlines", params={
    "country": "us", "category": "technology", "pageSize": 100, "apiKey": API_KEY
}, timeout=30)
r.raise_for_status()
data = r.json()
if data.get("status") != "ok":
    raise RuntimeError(data.get("message", "News API error"))

headlines = list(dict.fromkeys(a["title"] for a in data.get("articles", []) if a.get("title")))
documents = [clean(h) for h in headlines]
documents = [d for d in documents if d]
if not documents:
    raise RuntimeError("No usable headlines were returned")

v = TfidfVectorizer(stop_words="english", ngram_range=(1, 2), min_df=1, max_df=.90, sublinear_tf=True)
matrix = v.fit_transform(documents)
scores = np.asarray(matrix.sum(axis=0)).ravel()
for i in np.argsort(scores)[::-1][:20]:
    print(f"{v.get_feature_names_out()[i]}: {scores[i]:.3f}")

The Bottom Line

Use /v2/top-headlines for a current batch, extract only non-empty title values, clean conservatively, and rank unigrams plus bigrams with TF-IDF across many deduplicated headlines. Add named entities or noun phrases when names and readable concepts matter; use embeddings when synonym and paraphrase matching is central.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.