Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A 2017 research project showed how human judgments and machine learning could turn years of Wikipedia discussions into measurable evidence about personal attacks. Researchers associated with Jigsaw and the Wikimedia Foundation assembled more than 100,000 human-labeled comments, trained a classifier, and used it to examine 63 million English Wikipedia discussion comments from 2004 to 2015. The often-repeated figure of 13,500 “nastygrams” refers to the personal-attack subset highlighted in contemporaneous coverage—not to every kind of trolling, harassment, hate speech, or abuse.

What the 13,500 comments actually were

The project studied personal attacks as defined in the context of Wikipedia’s community policies. That is narrower than the broad everyday meaning of “troll.” A comment could be hostile or disruptive without meeting the study’s attack definition, while other harmful behaviors—such as hate speech, threats, stalking, or coordinated harassment—were outside the specific label being modeled.

MIT Technology Review described the resulting historical analysis as more than 13,500 personal attacks accompanied by more than 100,000 less abusive posts. The headline number is a useful shorthand for the article’s story, but it should not be silently treated as an official count of every abusive comment on Wikipedia or as a universal measure of online toxicity.

How researchers built the dataset

A large historical corpus

The authors processed a public dump of English Wikipedia’s full history and assembled 63 million discussion comments posted between 2004 and 2015. This gave them a long time span and enough material to study patterns that would be impractical to inspect manually.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Human judgments came first

They crowdsourced labels for 115,737 comments. Each comment received ten judgments, and 11.7% were labeled attacks by majority vote. The labeled set contained 37,611 randomly sampled comments and 78,126 comments selected near block events. In those two samples, the reported attack shares were 0.9% and 16.9%, respectively.

The different rates are important. Comments near blocks are deliberately enriched for difficult or potentially abusive cases; they are useful for finding examples and training a model, but they are not a representative estimate of attack prevalence across all Wikipedia discussions.

A classifier extended the analysis

The human-labeled comments supplied examples for a text classifier. The paper reports that its best classifier performed comparably, under the study’s evaluation procedure and metrics, to aggregating judgments from three crowd workers. That result means the model could approximate a particular aggregate of crowd labels on this dataset. It does not show that the system understood intent like a trained moderator, made perfect decisions, or would perform equally well on another site, language, or period.

Why machine labeling mattered

Human review is valuable but expensive and slow at the scale of tens of millions of comments. Once trained, a classifier could apply a consistent research label across the much larger historical corpus. That made it possible to ask questions about when attacks appeared, how they related to discussion outcomes, and how often existing moderation actions followed them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MIT Technology Review reported that roughly one in ten attacks resulted in moderator action. This was the project’s historical, model-based estimate from Wikipedia’s records, not a current Wikipedia rate and not a general rule for online communities. It illustrates the kind of system-level question that becomes possible when a large archive can be screened consistently.

What the study can—and cannot—say about trolls

It can make a narrow behavior measurable

The work demonstrates a repeatable method: define a policy-relevant behavior, collect multiple human judgments, train a model on those examples, and use the model to analyze a much larger archive. The paper’s authors describe the contribution as combining crowdsourcing and machine learning to analyze personal attacks at scale.

It cannot solve online harassment by itself

Personal attacks are only one part of harmful online behavior. Definitions vary between communities, and people can disagree about sarcasm, criticism, quoted language, cultural references, or context. The original report also raised concerns about language nuance, whether an algorithm matches real moderators, and whether users might change their wording to evade detection.

The evidence comes from English Wikipedia talk-page discussions written from 2004 through 2015. Applying the approach elsewhere would require new validation for the target platform’s rules, language, users, time period, and consequences of errors. A model that is acceptable for retrospective research may be unsuitable for automatically deleting posts or banning people.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Human annotation versus automated classification

Approach Strength Limitation Role in this project
Multiple human judgments Captures community-based interpretation and disagreement Costly and slow at very large scale Created the labeled examples and majority-vote reference
Text classifier Can process millions of archived comments consistently Reflects its training data and can miss context or new evasive language Extended labels from the sample to the 63-million-comment corpus
Moderator decision Can incorporate policy, history, intent, and situational context Requires time and may vary between reviewers Used as a historical outcome for comparison, not as the classifier’s ground truth
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why sampling near blocks helps—and misleads

Random sampling provides a way to estimate what ordinary discussion looks like. The block-adjacent sample supplies many more likely attacks, which helps a learning algorithm encounter rare examples. Combining both improves research coverage, but the two samples answer different questions. A 16.9% attack share in the enriched sample must not be reported as the attack rate for Wikipedia as a whole; the random sample’s reported share was 0.9%.

Could the method work on another platform?

Only after testing it on that platform’s own data and rules. A responsible deployment would need:

  • Local labels: multiple reviewers trained on the community’s definition of a personal attack.
  • Coverage checks: evaluation across dialects, languages, topics, newcomer and veteran users, and changes over time.
  • Error-cost analysis: separate thresholds for research, warnings, ranking, and enforcement, because a false accusation can be more damaging than a missed insult—or vice versa.
  • Human review: an appeal path and moderator oversight for consequential actions.
  • Drift monitoring: periodic relabeling to detect new slang, coded language, and changes in community norms.

Without those checks, a model can reproduce the biases and blind spots of its labels while giving them an appearance of objectivity.

The broader significance for online discussion

The project’s value was methodological as much as operational. It showed that a community archive could support quantitative study of interpersonal harm without pretending that a classifier had solved the underlying social problem. Researchers could compare patterns over time, examine relationships between attacks and moderation, and identify where additional human investigation was warranted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Lucas Dixon, identified in the contemporaneous report as Jigsaw’s chief research scientist, expressed the goal this way: “Our goal is to see how can we help people discuss the most controversial and important topics in a productive way all across the Internet.” The study offers one tool toward that goal: making a narrowly defined behavior visible at scale so communities can ask better questions about prevention and response.

What readers should take away

  • The 13,500 figure refers to a personal-attack subset in a historical Wikipedia analysis, not all online abuse.
  • Human judgments—ten per comment—formed the basis for the automated labels.
  • The classifier enabled retrospective analysis of 63 million English discussion comments from 2004–2015.
  • Its reported performance approximated an aggregate of crowd judgments under the paper’s evaluation method, not professional moderation in every setting.
  • Using the approach on another platform or language requires fresh labels, validation, and safeguards.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.