Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
FlashText is a Python library for finding or replacing many keywords in text without running a separate search for every term. It scans each document in one pass, supports whole-word matching, and selects the longest matching phrase when entries overlap. It can be useful for corpus-scale keyword extraction and alias normalization, but its published speed figures are historical results from the library author—not a guarantee for your workload.
What FlashText does
FlashText’s KeywordProcessor lets you build a keyword dictionary and then extract matching terms or replace them with canonical labels. For example, you can map the alias “Big Apple” to “New York,” then use the same processor to identify or normalize that phrase in documents.
The original paper describes searching or replacing keywords in one pass over a document, with scan complexity O(N), where N is the number of characters. In the paper’s comparison, regex-based searching is described as O(M × N), where M is the number of terms. This distinction is most relevant when the keyword list is large; actual runtime still depends on the text, dictionary, implementation, and workload.
FlashText uses trie-based processing related to Aho-Corasick. Rather than independently scanning the document once for each keyword, it traverses the text while consulting the keyword structure. The paper explains the algorithm and its complexity in Vikash Singh’s 2017 paper.
#1 Best Overall
How matching handles word boundaries and overlaps
Whole-word matches
FlashText is designed to match complete words or phrases rather than arbitrary substrings. A keyword “Apple” does not match inside “Pineapple.” That makes it useful when a substring hit would be misleading, though the definition of a boundary matters for your data: punctuation, Unicode text, and domain-specific tokenization can affect what counts as a word boundary.
Longest match wins
If your dictionary contains “Machine,” “Learning,” and “Machine learning,” and the full phrase appears, FlashText selects “Machine learning” rather than returning the shorter overlapping entries. This longest-match behavior helps avoid splitting a known phrase into smaller labels. Review the result against your own terminology, especially where phrase boundaries are ambiguous.
Rank #2
Case and spans
The repository documents case-sensitive mode and the ability to return span information for matches. Those options are useful if capitalization distinguishes terms in your corpus or if a later processing step needs the original character positions. See the official repository documentation for the API details.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteInstall and try the Python API
The package is installed with pip and exposes the KeywordProcessor class. This small example defines canonical names with aliases, extracts matches, and replaces an alias in text:
from flashtext import KeywordProcessor
processor = KeywordProcessor()
processor.add_keyword("Big Apple", "New York")
processor.add_keyword("New Delhi", "NCR region")
text = "The Big Apple and New Delhi are mentioned in this report."
print(processor.extract_keywords(text))
# ['New York', 'NCR region']
print(processor.replace_keywords(text))
# The New York and NCR region are mentioned in this report.
The second argument to add_keyword is the canonical value returned for that keyword. You can add entries individually, or load aliases from a dictionary or list as shown in the repository. For a large corpus, build the processor once and reuse it across documents rather than rebuilding its keyword structure for each one.
- Install: Run
pip install flashtextin the Python environment that will run your processing job. - Build the dictionary: Create a
KeywordProcessorand add each keyword, optionally associating an alias with its canonical label. - Choose the operation: Call
extract_keywords(text)to retrieve matches, orreplace_keywords(text)to produce normalized text. - Validate representative documents: Check punctuation, capitalization, overlapping terms, and boundary behavior before applying the mapping to the full corpus.
The package’s installation instructions, examples, and MIT license are available in the FlashText GitHub repository. PyPI lists version 2.7, uploaded on 16 February 2018, and names Vikash Singh as maintainer; that release date alone does not establish the package’s current support status. Check the repository and package listing for current activity before adopting it for a new production system: FlashText on PyPI.
FlashText versus regex: which should you use?
| Need | FlashText | Regex |
|---|---|---|
| Search many dictionary terms in one document | Trie-based, one-pass approach; the paper gives O(N) scan complexity. | The paper characterizes its regex comparison as O(M × N), with M terms and N document characters; performance depends on the pattern and implementation. |
| Require whole-word matching | Designed for word-boundary matching; “Apple” does not match inside “Pineapple.” | Possible with boundary assertions, but those rules must be represented in the pattern and may need tailoring. |
| Resolve overlapping phrases | Prefers the longest matching phrase. | Outcome depends on pattern construction and alternation order; it may need explicit handling. |
| Extract or normalize aliases | Provides extraction and replacement through the same keyword processor. | Can do both, but the patterns and replacement logic must be assembled for the task. |
| Specialized syntax or custom match rules | Focused on dictionary keyword matching and its documented boundary behavior. | More suitable when the search requires general pattern syntax rather than a fixed keyword dictionary. |
The table reflects the design and comparison in the original paper and the official API documentation, not a benchmark of every regex engine or dataset. If you need patterns such as optional characters, numeric ranges, or complex context-sensitive rules, regex may be the better fit. If you have a fixed list of terms and want dictionary-based extraction or replacement, FlashText offers a direct API and explicit longest-match behavior.
Free tools Windows power users keep installed
One-click scans. No signup required.
What the published speed figures mean
Vikash Singh reported that a regex process took 24 hours on one million documents with 2,000 keywords. For an expanded workload of millions of documents and more than 10,000 keywords, he reported that the earlier approach would take more than 10 days, while his custom FlashText implementation completed the keyword-extraction workload in 15 minutes. These are the author’s 2017 implementation figures, not an independent, controlled comparison or a promise of the same result on other hardware, text, or software versions. The figures are described in the 2017 paper.
Quick Recap
Best Value
When FlashText is a good fit—and what to check
- Consider it when you have a large, mostly fixed dictionary and need to extract terms or replace aliases throughout many documents.
- Check boundaries with realistic samples if your text contains unusual punctuation, mixed scripts, or tokens whose boundaries do not follow ordinary word rules.
- Inspect overlaps when one term is contained in a longer phrase, since the longest match takes precedence.
- Choose regex instead when your matching problem depends on general pattern syntax or complex conditions that go beyond a keyword dictionary.
- Verify maintenance fit before using it as a production dependency. PyPI’s listed 2.7 release was uploaded in 2018; the release date is a reason to check current repository and package activity, not proof by itself that the project is unsupported.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

