The defensible way to collect YouTube comments is through the YouTube Data API, not by scraping YouTube pages. For a video, start with commentThreads.list to retrieve top-level comments and any replies included in each thread. When you need every reply to a particular top-level comment, call comments.list with that comment’s parentId. Paginate until your documented stopping point, track quota use, and treat every conclusion as an estimate from the comments you collected—not as the opinion of all viewers.
Why page scraping is the wrong starting point
YouTube’s API Services Developer Policies state: “You and your API Clients must not, and must not encourage, enable, or require others to, directly or indirectly, scrape YouTube Applications or Google Applications, or obtain scraped YouTube data or content.” A browser script that parses YouTube HTML or automates the site can therefore create a policy problem as well as a brittle engineering project. API responses give you documented fields, pagination tokens and quota accounting instead.
The API route still has limits. Comments can be disabled, videos can disappear, moderation can remove items, and a response is not a census of every viewer. State the collection date, target selection, exclusions and reply policy in your analysis so another person can understand exactly what your dataset represents.
What you need before collecting
- A Google Cloud project with the YouTube Data API enabled and an API key (or another permitted credential appropriate to your application).
- The target video’s ID, or a channel ID when your question concerns channel-related threads.
- A written sampling plan: which videos, date window, languages, reply depth, pagination cutoff and exclusions you will use.
- Storage for raw responses and a separate, cleaned table. Keep the original response so you can audit transformations.
Choose the correct API method
Video threads with commentThreads.list
Pass videoId to retrieve comment threads for one video. Request part=snippet for top-level comment data. Add replies when you want replies that YouTube includes inline with a thread. An inline thread response may not contain every reply, so do not label it “complete” merely because a replies object is present.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Complete replies with comments.list
For a specific top-level comment, pass its ID as parentId to comments.list. This endpoint returns replies in pages of up to 100 and supplies nextPageToken when another page exists. Follow that token until it is absent or until your predeclared boundary is reached.
Channel-related retrieval
The thread method documents both channelId and allThreadsRelatedToChannelId for channel-related retrieval. Confirm which interpretation matches your question; a channel-wide project needs an explicit video-selection rule rather than an assumption that one request represents the entire channel.
A reproducible collection workflow
- Define the question. “What questions do viewers ask?” requires different sampling than “How did viewers react to this launch?” Write the unit of analysis (comment, thread, video or channel) first.
- Resolve IDs. Record each video or channel ID, title, publication date and the reason it was selected. Do not silently replace unavailable videos.
- Request only needed parts. Use
snippetfor top-level comments, andsnippet,repliesonly when inline replies are useful. Smaller responses are easier to store and inspect. - Paginate. Save every request’s parameters and
nextPageToken. Stop at a stated date, page count, comment count or complete-token condition. - Expand replies selectively. Use
comments.listbyparentIdfor threads where reply completeness matters. This costs additional calls, so prioritize threads according to your research question. - Log exclusions. Record disabled comments, unavailable videos, deleted text, language filters, duplicate IDs and moderation-related gaps.
- Freeze the raw data. Keep immutable JSON responses with collection timestamps, then create a derived table for cleaning and analysis.
Runnable Python collector
The script below collects top-level threads for one video, follows thread pages, and writes raw JSON. Replace the key and ID with your values. It requests inline replies but does not claim those replies are complete.
import json
import time
import requests
API_KEY = "YOUR_API_KEY"
VIDEO_ID = "VIDEO_ID"
BASE = "https://www.googleapis.com/youtube/v3/commentThreads"
rows = []
page_token = None
while True:
params = {
"key": API_KEY,
"part": "snippet,replies",
"videoId": VIDEO_ID,
"maxResults": 100,
}
if page_token:
params["pageToken"] = page_token
response = requests.get(BASE, params=params, timeout=30)
response.raise_for_status()
payload = response.json()
rows.extend(payload.get("items", []))
page_token = payload.get("nextPageToken")
if not page_token:
break
time.sleep(0.1)
with open("youtube_threads.json", "w", encoding="utf-8") as f:
json.dump(rows, f, ensure_ascii=False, indent=2)
print(f"Saved {len(rows)} thread objects")
For complete replies to a selected top-level comment, make a second request for its ID:
import requests
params = {
"key": "YOUR_API_KEY",
"part": "snippet",
"parentId": "TOP_LEVEL_COMMENT_ID",
"maxResults": 100,
}
replies = []
while True:
data = requests.get(
"https://www.googleapis.com/youtube/v3/comments",
params=params,
timeout=30,
).json()
replies.extend(data.get("items", []))
token = data.get("nextPageToken")
if not token:
break
params["pageToken"] = token
print(f"Collected {len(replies)} replies")
Equivalent cURL and Node.js requests
cURL
curl -G "https://www.googleapis.com/youtube/v3/commentThreads"
--data-urlencode "key=YOUR_API_KEY"
--data-urlencode "part=snippet,replies"
--data-urlencode "videoId=VIDEO_ID"
--data-urlencode "maxResults=100"
Node.js
const params = new URLSearchParams({
key: 'YOUR_API_KEY',
part: 'snippet,replies',
videoId: 'VIDEO_ID',
maxResults: '100'
});
const res = await fetch(`https://www.googleapis.com/youtube/v3/commentThreads?${params}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const page = await res.json();
console.log(page.items?.length ?? 0, page.nextPageToken ?? 'done');
Quota, pagination and scale planning
The comments.list method costs one quota unit per call and accepts 1–100 results per page. Invalid requests can still consume at least one point. Google’s API overview lists a default allocation of 10,000 units per day for most endpoints at the time of its documentation, but says defaults can change and describes a quota-extension request. Treat 10,000 as a planning reference, not a guaranteed allowance for every project.
Estimate calls before a large run: one call per thread page, plus one call per reply page you expand. Cache completed pages, avoid retrying permanent 4xx errors, and use exponential backoff for transient failures. Store the response status and request parameters so a partial run can resume without duplicating records.
Turn comments into defensible insights
Clean without erasing meaning
Preserve the original text and comment ID. Create analysis fields for normalized case, tokenization, language, spam flags and redacted personal information. Keep emojis, punctuation and slang in a separate representation because they can carry sentiment or emphasis.
Code recurring questions and themes
Start with a small hand-coded sample. Define labels such as setup questions, feature requests, factual corrections, troubleshooting reports and off-topic posts. Apply the codebook to a validation sample, revise ambiguous rules, then scale with an automated classifier only after measuring its errors.
Recommended Free Tools
Use sentiment as an aggregate, not a verdict
YouTube’s derived-metrics policy permits aggregate viewer-sentiment analysis from comment analysis subject to its conditions. It does not authorize inferring or estimating sensitive protected attributes. Do not infer age, race, health status, political affiliation or other protected characteristics from usernames, language or sentiment. Report counts and proportions for your collected sample, and explain uncertainty caused by spam, irony, multilingual text and unequal commenting rates.
Account for thread structure
Separate top-level comments from replies. Report the number of videos, threads, commenters and comments, and say whether one person could contribute multiple comments. A highly active thread can dominate a naive frequency count, so consider per-video rates or a cap per commenter when that matches your question.
What published studies can—and cannot—tell you
Shajari, Agarwal and Alassad’s 2023 study analyzed 20 channels, 7,782 videos, 294,199 commenters and 596,982 comments while studying suspicious coordinated commenter behavior. Those are that study’s dataset counts, not a measurement of all YouTube activity.
Heydari, Zhang, Appel, Wu and Ranade’s 2019 “YouTube Chatter” paper compared comment rates, reply rates, thread lengths, comment lengths, profanity rates and simple classifiers in specific political and apolitical channel groups. Use such findings as study-specific evidence, never as a platform-wide baseline.
Troubleshooting common failures
HTTP 403 or quota errors
Check that the API is enabled, the key is valid and the project has remaining quota. Reduce unnecessary reply expansion, wait for quota reset, or request additional quota through Google’s documented process. An invalid parameter can also consume a unit.
An empty result when comments should exist
Verify the video ID rather than the full watch URL, confirm comments are enabled, and check that your credential belongs to the intended project. A removed, private or restricted video can produce no usable comments.
Replies appear incomplete
This is expected when relying only on inline replies. Take the top-level comment ID and paginate comments.list with parentId.
Duplicates after resuming a job
Deduplicate by stable comment or thread IDs, not by text. Persist the last successful page token and request parameters, and keep raw pages for audit.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
Classifier quality is poor
Review false positives by language, sarcasm, profanity and topic vocabulary. Expand the hand-labeled validation set, split results by video, and publish the coding rules alongside aggregate counts.
Or skip the browser setup
If your goal is to capture a clean visual record of a YouTube page or another site—not to retrieve comment text—ScreenshotNeo provides a single-request screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
Use the ScreenshotNeo API documentation for options such as full-page capture, CSS selectors, custom JavaScript, waits, headers, cookies, device presets, PDFs, signed links and bulk jobs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The Free plan includes 1,000 shots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
Privacy, retention and reporting checklist
- Remove API keys from source control and logs.
- Document collection date, timezone, video/channel selection and pagination boundary.
- Explain how disabled comments, deleted content, duplicates, spam and languages were handled.
- Use aggregate reporting where possible and avoid publishing unnecessary usernames or personal data.
- Label automated classifications as estimates and include a human-reviewed error check.
- State clearly that commenters are self-selected and your sample does not represent silent viewers.
Frequently Asked Questions
Can I collect every comment on YouTube?
No. Availability depends on enabled comments, video access, moderation and API responses. Describe the exact sample and exclusions instead of claiming complete platform coverage.
Should I retrieve replies for every thread?
Only when your question requires them. Inline replies may be incomplete; expanding every thread increases calls and quota use, so prioritize and document the rule.
How should I report sentiment results?
Give the coding method, validation approach, sample size and uncertainty. Present aggregate results for collected comments and do not infer sensitive protected attributes.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →

