News API retrieves headlines; it does not extract keywords for you. A dependable Python workflow is to collect a batch of title values, clean them carefully, rank unigrams and bigrams with TF-IDF, then enrich the result with named entities or noun phrases when your use case needs people, companies, places, or readable concepts.
This approach treats headlines as a corpus instead of pretending that one short headline contains enough evidence for statistical ranking.
Define the output before choosing an algorithm
“Keyword” can mean several different outputs:
- Keywords: individual terms such as inflation or wildfires.
- Keyphrases: multi-word concepts such as interest rate or machine learning.
- Named entities: people, organizations, locations, products, laws, events, dates, and monetary values.
- Topics: broader themes inferred from many headlines.
- Tags or search terms: labels chosen for a taxonomy or retrieval system, which are not always the most statistically salient words.
For a dashboard or topic monitor, start with corpus-level TF-IDF and bigrams. For one headline, use entities or noun phrases instead; TF-IDF needs document-to-document comparison.
Choose the News API endpoint
/v2/top-headlines for current batches
Use top-headlines for country- or category-based dashboards, breaking-news displays, and small recent batches. It supports country, category, sources, q, pagination, and a documented maximum pageSize of 100. News API requires an API key. Country and category cannot be combined with sources.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
/v2/everything for analysis corpora
For historical or search-based analysis, everything is usually a better corpus source. It can search fields with searchIn=title, constrain dates and languages, filter domains or sources, and sort by relevance, popularity, or publication time.
params = {
"q": "artificial intelligence OR machine learning",
"searchIn": "title",
"language": "en",
"from": "2026-08-01",
"to": "2026-08-18",
"sortBy": "publishedAt",
"pageSize": 100,
"apiKey": NEWS_API_KEY,
}
The dates above are an example, not a permanent query; generate them dynamically in an application.
Fetch and validate headlines
Keep the key outside source control, request a bounded page, and fail clearly on transport or API errors.
import os
import requests
NEWS_API_KEY = os.environ["NEWS_API_KEY"]
response = requests.get(
"https://newsapi.org/v2/top-headlines",
params={
"country": "us",
"category": "technology",
"pageSize": 100,
"apiKey": NEWS_API_KEY,
},
timeout=30,
)
response.raise_for_status()
payload = response.json()
if payload.get("status") != "ok":
raise RuntimeError(payload.get("message", "News API request failed"))
headlines = [
article["title"]
for article in payload.get("articles", [])
if article.get("title")
]
News API article objects include fields such as title, description, url, publishedAt, and content. For this task, select article["title"] explicitly. The documented content value may be truncated to 200 characters, so it is not a substitute for full article text.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Clean titles without destroying meaning
Cleaning is a scoring aid, not permission to erase entities. Keep the original title for display and create a normalized copy for matching.
import re
NEWS_STOPWORDS = {
"says", "say", "said", "report", "reports", "reported",
"new", "latest", "live", "update", "updates", "breaking",
"amid", "after", "before", "over", "could", "would", "may",
"watch", "video",
}
def clean_headline(text: str) -> str:
text = re.sub(r"[[^]]*]", " ", text) # [Updated], [Video]
text = re.sub(r"([^)]*)", " ", text) # optional labels
text = re.sub(r"https?://S+", " ", text)
text = re.sub(r"[^ws'-]", " ", text)
text = re.sub(r"s+", " ", text).strip().lower()
return " ".join(
token for token in text.split()
if token not in NEWS_STOPWORDS
and not token.isdigit()
and len(token) > 2
)
documents = [clean_headline(title) for title in headlines]
documents = [doc for doc in documents if doc]
Do not blindly remove punctuation: C++, COVID-19, U.S., AI-powered, and S&P 500 can lose their identity. A production cleaner should test representative titles and decide how acronyms, hyphens, dates, and publisher suffixes are handled. General English stop words remove terms such as “the” and “of”; a custom news list is needed for boilerplate such as “says” and “amid”. Words like “war”, “state”, or “trade” may be meaningful in a particular corpus, so review before adding them.
Rank corpus keywords with TF-IDF
TF-IDF raises terms that are frequent in one document but less common across the supplied collection. It ranks distinctiveness within your corpus, not objective importance or newsworthiness.
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
vectorizer = TfidfVectorizer(
stop_words="english",
ngram_range=(1, 2),
min_df=2,
max_df=0.85,
sublinear_tf=True,
)
matrix = vectorizer.fit_transform(documents)
terms = vectorizer.get_feature_names_out()
scores = matrix.sum(axis=0).A1
ranking = sorted(zip(terms, scores), key=lambda item: item[1], reverse=True)
for term, score in ranking[:20]:
print(f"{term}: {score:.3f}")
Bigrams preserve meaning that unigrams lose: interest rate, climate change, and stock market are more useful than isolated fragments. Trigrams can be tested with ngram_range=(1, 3), but become sparse in small collections.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Tune document-frequency thresholds
min_df=2requires a term to occur in at least two headline documents. It reduces noise but excludes one-off breaking-news names.- Use
min_df=1for a small batch when rare entities matter. max_df=0.85removes terms appearing in most documents. Adjust it when a common term is actually your target.
With one or two documents, TF-IDF has little statistical value. Collect dozens or hundreds of comparable headlines, or switch to noun phrases, named entities, a fixed vocabulary, or embeddings.
Get keywords for each headline
def keywords_for_document(row_index, top_n=8):
row = matrix[row_index].toarray().ravel()
indices = np.argsort(row)[::-1]
return [
(terms[i], float(row[i]))
for i in indices
if row[i] > 0
][:top_n]
for index, title in enumerate(headlines[:5]):
print(title)
print(keywords_for_document(index))
These scores show what is distinctive in each title relative to the downloaded collection; they do not measure importance outside that collection.
Add entities and noun phrases
TF-IDF can rank generic verbs above a company or person. A local spaCy model supplies a complementary linguistic signal.
import spacy
nlp = spacy.load("en_core_web_sm")
ENTITY_LABELS = {"PERSON", "ORG", "GPE", "LOC", "PRODUCT", "EVENT", "LAW"}
def extract_entities(text):
doc = nlp(text)
return [
(ent.text, ent.label_)
for ent in doc.ents
if ent.label_ in ENTITY_LABELS
]
def extract_noun_phrases(text):
doc = nlp(text)
return [
chunk.text.lower()
for chunk in doc.noun_chunks
if len(chunk.text.split()) <= 5
]
A practical hybrid pipeline extracts TF-IDF terms, entities, and noun phrases; removes overlaps; then boosts entities when the product monitors people, organizations, places, or products. Preserve the original spelling for display while using normalized forms for deduplication. Entity accuracy depends on the model, language, capitalization, spelling, and headline context; short or ambiguous titles can be misclassified.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Deduplicate before scoring
Syndicated stories can make one event appear artificially important. Start with exact and normalized-title deduplication:
unique_titles = list(dict.fromkeys(headlines))
For stronger controls, deduplicate by canonical URL, compare normalized-title similarity, group by source and publication time, or cluster embeddings. Different publishers may describe the same event with different wording, so identical text is not required for a duplicate. If one publisher supplies most rows, calculate source-specific rankings or balance the corpus.
Choose a method for the job
| Method | Best use | Main trade-off |
|---|---|---|
| Frequency counts | Fast prototype | Rewards repeated boilerplate |
| TF-IDF | Explainable corpus ranking | Needs multiple documents and is corpus-dependent |
| RAKE | Simple phrase extraction | Sensitive to stop words and punctuation |
| Noun phrases | Readable keyphrases | Depends on parser quality |
| Named entities | People, companies, places | Misses general concepts |
| TextRank | Unsupervised phrases | More complex and unstable on short text |
| Embeddings or KeyBERT-style methods | Synonyms and paraphrases | More compute and model choices |
| LLM extraction | Structured labels and explanations | Cost, latency, privacy, consistency, and evaluation concerns |
Troubleshoot common failures
Only generic words appear
Expand the custom news stop list, inspect the top 100 terms, raise min_df, and remove duplicated or single-source headlines. Do not permanently remove a word merely because it is common in one category.
Results are one-word fragments
Use ngram_range=(1, 2) or noun chunks. Compare the original phrase with normalized output so readers see “interest rates” rather than an artificial stem.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
The result is empty
Check the API status, key, rate limit, response size, missing titles, and whether cleaning removed every token. Return an empty or low-confidence result for titles such as “Markets react” rather than inventing keywords.
API requests fail
if response.status_code == 429:
# Back off and retry according to your account and response headers.
pass
Also handle timeouts, invalid keys, non-ok payloads, malformed JSON, and empty result sets. Retry policy should follow the response and account plan, not an arbitrary fixed assumption.
Dates and languages are inconsistent
News API documents publishedAt timestamps in UTC. Keep UTC for filtering and deduplication, converting only for presentation. English stop words and spaCy’s English model do not generalize to other languages; use language-specific tokenization, models, and stop lists.
Evaluate usefulness instead of trusting plausible output
- Collect 50–100 representative headlines.
- Have a reviewer mark useful keywords and phrases.
- Run the pipeline and measure precision at 5 or 10 results.
- Inspect false positives such as attribution verbs and false negatives such as missed entities.
- Tune stop words, frequency thresholds, n-gram range, deduplication, and entity weighting.
Repeat this by category, source mix, language, and time period. Vocabulary changes with the news cycle, and a list that looks reasonable can still be unstable or dominated by one publisher.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Production considerations
- Pagination and caching: store the retrieval time, query, source, and API response so rankings can be reproduced without unnecessary requests.
- Monitoring: watch for vocabulary drift, empty batches, sudden source concentration, and changes in entity quality.
- Privacy: keep headline retention and third-party processing aligned with your policy; local spaCy and scikit-learn avoid sending text to a managed NLP API.
- Rights and access: the article URL may require separate processing subject to publisher terms, robots rules, copyright, access controls, and your News API plan. Do not assume the
contentfield is full text. - Commercial use: News API’s pricing page, checked August 18, 2026, lists a $0 Developer plan, $449/month Business plan, and $1,749/month Advanced plan, with different limits, delays, support, and SLA terms. It states that the Developer plan is for development and testing rather than staging, production, or internal production-like use. Recheck current terms and prices before deployment.
For managed enrichment, Google Cloud Natural Language, Amazon Comprehend, and Azure AI Language can provide entities or key phrases, but compare current regional pricing, quotas, language coverage, latency, data handling, and lock-in with local processing before choosing one.
A complete minimal pipeline
import os, re, requests, numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
API_KEY = os.environ["NEWS_API_KEY"]
def clean(text):
text = re.sub(r"[[^]]*]|https?://S+", " ", text)
text = re.sub(r"[^ws'-]", " ", text.lower())
text = re.sub(r"s+", " ", text).strip()
stop = {"says", "said", "reports", "reported", "new", "latest", "live", "breaking", "update", "amid", "after", "before"}
return " ".join(t for t in text.split() if len(t) > 2 and t not in stop and not t.isdigit())
r = requests.get("https://newsapi.org/v2/top-headlines", params={
"country": "us", "category": "technology", "pageSize": 100, "apiKey": API_KEY
}, timeout=30)
r.raise_for_status()
data = r.json()
if data.get("status") != "ok":
raise RuntimeError(data.get("message", "News API error"))
headlines = list(dict.fromkeys(a["title"] for a in data.get("articles", []) if a.get("title")))
documents = [clean(h) for h in headlines]
documents = [d for d in documents if d]
if not documents:
raise RuntimeError("No usable headlines were returned")
v = TfidfVectorizer(stop_words="english", ngram_range=(1, 2), min_df=1, max_df=.90, sublinear_tf=True)
matrix = v.fit_transform(documents)
scores = np.asarray(matrix.sum(axis=0)).ravel()
for i in np.argsort(scores)[::-1][:20]:
print(f"{v.get_feature_names_out()[i]}: {scores[i]:.3f}")
The Bottom Line
Use /v2/top-headlines for a current batch, extract only non-empty title values, clean conservatively, and rank unigrams plus bigrams with TF-IDF across many deduplicated headlines. Add named entities or noun phrases when names and readable concepts matter; use embeddings when synonym and paraphrase matching is central.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

