Word2Vec does not assign semantic roles such as Agent, Patient, or Recipient to words in a sentence. It learns a static vector for each word from patterns of use. Those vectors can encode useful lexical and selectional regularities, and they can serve as features in a semantic role labeling (SRL) system, but identifying who did what to whom requires a predicate, its sentence context, and a role-classification model.
Two different meanings of “semantic roles”
The phrase can point to two related but distinct ideas:
- Relations in a word-vector space: words used in similar contexts tend to have nearby vectors, and some recurring relationships appear as directions or offsets between vectors.
- Semantic role labeling: an SRL system analyzes a particular sentence, finds a predicate, identifies its arguments, and labels what each argument does in relation to that predicate.
Word2Vec addresses the first problem directly. It can contribute evidence to the second, but a standalone Word2Vec vector is not a sentence-level role label.
What Word2Vec represents
Word2Vec learns distributed representations from neighboring words in a training corpus. The original Skip-gram work describes these representations as capturing “a large number of precise syntactic and semantic word relationships.” A word is represented by a numerical vector whose useful properties come from distributional patterns: words appearing in comparable contexts often occupy comparable regions of the vector space.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems#1 Best Overall
- NLP: The Essential Guide to Neuro-Linguistic Programming
That representation is static. In the usual Word2Vec model, one spelling has one learned vector regardless of the sentence. The vector for bank does not change when the sentence means a financial institution rather than a river edge. Word2Vec also does not encode a sentence’s full word order. As the original authors put it, “An inherent limitation of word representations is their indifference to word order and their inability to represent idiomatic phrases.”
How vector offsets express lexical relationships
Some relationships can be approximated by subtracting and adding vectors. A well-known illustration is:
vector(“King”) − vector(“Man”) + vector(“Woman”) ≈ vector(“Queen”)
The intuition is that the difference between king and man captures part of a gender-and-status direction; adding that direction to woman lands near queen. This is a regularity among lexical representations, not an analysis of an event involving a king, a man, and a woman.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #2
Mikolov, Yih, and Zweig reported almost 40% accuracy on the syntactic analogy questions in their 2013 evaluation. That figure belongs specifically to that paper’s analogy task. The same study separately evaluated semantic regularities with SemEval-2012 Task 2; its result should not be merged with the analogy percentage or interpreted as SRL accuracy.
What semantic role labeling actually asks
SRL is predicate-centered. The system first considers a predicate, usually a verb, and then determines which phrases are its arguments and what roles they play. The task has been defined as “analyzing clause predicates in text by identifying arguments and tagging them with semantic labels indicating the role they play with respect to the predicate.”
For the sentence “Mr. Smith sent the report to me this morning,” a role analysis can identify:
| Phrase | Role relative to “sent” | What the label means |
|---|---|---|
| Mr. Smith | Agent | The entity performing the sending |
| the report | Object | The thing sent |
| me | Recipient | The entity receiving it |
| this morning | Temporal | When the event occurred |
The same word can receive different roles with different predicates or sentence structures. A vector learned for report cannot by itself determine whether a particular occurrence is an object, a topic, or part of a different construction. The predicate, candidate phrase, syntax, and surrounding context are required.
Can Word2Vec tell who did what to whom?
Not on its own. Word2Vec can provide clues about which arguments are plausible for a predicate. For example, distributional statistics may indicate that a verb commonly occurs with people as agents and documents as objects. Those are selectional preferences: expectations about the kinds of entities that tend to fill a role.
Selectional preferences are useful because syntax is not always reliable or sufficient. A prepositional phrase may attach ambiguously, parsing may be wrong, or the same syntactic pattern may support several roles. A classifier can use word-vector similarity to supplement those signals. However, plausibility is not identification. A verb’s typical preferences cannot settle every sentence, and a static vector cannot resolve context-dependent meaning or long-distance dependencies by itself.
What the selectional-preference evidence shows
Zapirain, Agirre, Màrquez, and Surdeanu (2013) evaluated WordNet-based and distributional selectional-preference models for semantic role classification on CoNLL-2005 data annotated under PropBank. In those experiments:
- Selectional-preference models outperformed a lexical-matching baseline.
- Distributional approaches performed better than the WordNet alternatives that were tested.
- Second-order distributional similarity was the strongest of the evaluated preference variants.
- Modeling preferences around prepositions as well as verbs improved prepositional-phrase classification compared with verb-only preferences.
The authors reported 20 F1 points of in-domain improvement and almost 40 F1 points of out-of-domain improvement over their lexical baseline when measuring the preference models in isolation. When the features were added to a state-of-the-art SRC system, they reported 17% in-domain and 13% out-of-domain error reduction. In end-to-end SRL, the change produced small but statistically significant improvements and affected approximately 4% of argument candidates.
Rank #4
These numbers describe that study’s models, baselines, domains, annotation scheme, and evaluation procedure. They are not a general Word2Vec score, a universal SRL guarantee, or a comparison with every later system. The paper’s error analysis also found that preference features can introduce mistakes because they model syntactic structure imperfectly. They complement syntax and contextual evidence rather than replacing them.
How embeddings fit into a real SRL system
Modern SRL architectures normally combine several information sources:
- Sentence context: the words around the predicate and candidate argument.
- Predicate representation: which predicate is being analyzed and where it occurs.
- Candidate-argument representation: the span or token that may receive a role.
- Syntax or structural cues: dependency, constituency, or learned structural patterns.
- Selectional preferences: distributional evidence about plausible argument types.
- Subword information: character-level representations that help with morphology and rare words.
A 2019 TACL system described an SRL setup using randomly initialized word embeddings, pretrained embeddings, character embeddings, sentence encoding, and explicit predicate–argument information. This is the careful way to state Word2Vec’s practical role: an embedding can contribute features to a model that assigns roles. The embedding itself does not assign those roles.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why analogy accuracy and SRL F1 are not comparable
| Aspect | Word-vector analogy evaluation | Semantic role classification evaluation |
|---|---|---|
| Primary task | Complete lexical or syntactic relationships between word vectors | Label predicate arguments in sentences |
| Typical representation | Static word vectors and vector arithmetic | Sentence, predicate, argument, syntactic, and distributional features |
| Example evidence | “King − Man + Woman” near “Queen” | Agent, Object, Recipient, and Temporal labels for “sent” |
| Reported metric | Analogy accuracy; almost 40% in the cited 2013 syntactic-analogy evaluation | F1 differences, error reduction, and end-to-end changes on CoNLL-2005/PropBank-style data |
| What the result supports | Some regularities are geometrically encoded in the lexicon | Specific features can improve role classification under the tested conditions |
An analogy result therefore cannot prove that a model understands event roles. Conversely, an SRL improvement from distributional features does not mean that vector arithmetic alone performs semantic parsing.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchBest Value
Practical answer: when should you use Word2Vec for roles?
Use it as a supporting feature
Word2Vec is useful when you need broad lexical similarity or selectional-preference evidence and have a model that can combine it with sentence context and predicate–argument structure. It can be especially helpful when a parser or surface pattern leaves a role ambiguous.
Do not use it as the role decision
Do not infer “who did what to whom” by applying an analogy formula to the words in a sentence. That approach ignores argument boundaries, predicate sense, word order, negation, tense, and the possibility that a phrase is idiomatic.
Choose evaluation that matches the claim
If your goal is lexical relationships, report analogy or semantic-relatedness results on the relevant benchmark. If your goal is SRL, report role-labeling metrics on an annotated SRL dataset and identify the predicate inventory, argument scheme, domain, and baseline. Never present one task’s score as evidence for the other.
Further reading
For a broader treatment, Daniel Jurafsky and James H. Martin’s freely available third-edition draft of Speech and Language Processing includes separate coverage of embeddings and semantic role labeling and argument structure. It is a general NLP textbook, not a manual claiming that Word2Vec alone performs SRL.
The Bottom Line
Word2Vec captures distributional relationships among words and can supply selectional-preference features, but semantic roles are assigned only when a model analyzes a predicate, its arguments, and their sentence context. Vector analogies reveal lexical geometry; they are not a substitute for semantic role labeling.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

