Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

FlashText is a Python library for finding or replacing many keywords in text without running a separate search for every term. It scans each document in one pass, supports whole-word matching, and selects the longest matching phrase when entries overlap. It can be useful for corpus-scale keyword extraction and alias normalization, but its published speed figures are historical results from the library author—not a guarantee for your workload.

What FlashText does

FlashText’s KeywordProcessor lets you build a keyword dictionary and then extract matching terms or replace them with canonical labels. For example, you can map the alias “Big Apple” to “New York,” then use the same processor to identify or normalize that phrase in documents.

The original paper describes searching or replacing keywords in one pass over a document, with scan complexity O(N), where N is the number of characters. In the paper’s comparison, regex-based searching is described as O(M × N), where M is the number of terms. This distinction is most relevant when the keyword list is large; actual runtime still depends on the text, dictionary, implementation, and workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FlashText uses trie-based processing related to Aho-Corasick. Rather than independently scanning the document once for each keyword, it traverses the text while consulting the keyword structure. The paper explains the algorithm and its complexity in Vikash Singh’s 2017 paper.

How matching handles word boundaries and overlaps

Whole-word matches

FlashText is designed to match complete words or phrases rather than arbitrary substrings. A keyword “Apple” does not match inside “Pineapple.” That makes it useful when a substring hit would be misleading, though the definition of a boundary matters for your data: punctuation, Unicode text, and domain-specific tokenization can affect what counts as a word boundary.

Longest match wins

If your dictionary contains “Machine,” “Learning,” and “Machine learning,” and the full phrase appears, FlashText selects “Machine learning” rather than returning the shorter overlapping entries. This longest-match behavior helps avoid splitting a known phrase into smaller labels. Review the result against your own terminology, especially where phrase boundaries are ambiguous.

Case and spans

The repository documents case-sensitive mode and the ability to return span information for matches. Those options are useful if capitalization distinguishes terms in your corpus or if a later processing step needs the original character positions. See the official repository documentation for the API details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install and try the Python API

The package is installed with pip and exposes the KeywordProcessor class. This small example defines canonical names with aliases, extracts matches, and replaces an alias in text:

from flashtext import KeywordProcessor

processor = KeywordProcessor()
processor.add_keyword("Big Apple", "New York")
processor.add_keyword("New Delhi", "NCR region")

text = "The Big Apple and New Delhi are mentioned in this report."

print(processor.extract_keywords(text))
# ['New York', 'NCR region']

print(processor.replace_keywords(text))
# The New York and NCR region are mentioned in this report.

The second argument to add_keyword is the canonical value returned for that keyword. You can add entries individually, or load aliases from a dictionary or list as shown in the repository. For a large corpus, build the processor once and reuse it across documents rather than rebuilding its keyword structure for each one.

  1. Install: Run pip install flashtext in the Python environment that will run your processing job.
  2. Build the dictionary: Create a KeywordProcessor and add each keyword, optionally associating an alias with its canonical label.
  3. Choose the operation: Call extract_keywords(text) to retrieve matches, or replace_keywords(text) to produce normalized text.
  4. Validate representative documents: Check punctuation, capitalization, overlapping terms, and boundary behavior before applying the mapping to the full corpus.

The package’s installation instructions, examples, and MIT license are available in the FlashText GitHub repository. PyPI lists version 2.7, uploaded on 16 February 2018, and names Vikash Singh as maintainer; that release date alone does not establish the package’s current support status. Check the repository and package listing for current activity before adopting it for a new production system: FlashText on PyPI.

FlashText versus regex: which should you use?

Need FlashText Regex
Search many dictionary terms in one document Trie-based, one-pass approach; the paper gives O(N) scan complexity. The paper characterizes its regex comparison as O(M × N), with M terms and N document characters; performance depends on the pattern and implementation.
Require whole-word matching Designed for word-boundary matching; “Apple” does not match inside “Pineapple.” Possible with boundary assertions, but those rules must be represented in the pattern and may need tailoring.
Resolve overlapping phrases Prefers the longest matching phrase. Outcome depends on pattern construction and alternation order; it may need explicit handling.
Extract or normalize aliases Provides extraction and replacement through the same keyword processor. Can do both, but the patterns and replacement logic must be assembled for the task.
Specialized syntax or custom match rules Focused on dictionary keyword matching and its documented boundary behavior. More suitable when the search requires general pattern syntax rather than a fixed keyword dictionary.

The table reflects the design and comparison in the original paper and the official API documentation, not a benchmark of every regex engine or dataset. If you need patterns such as optional characters, numeric ranges, or complex context-sensitive rules, regex may be the better fit. If you have a fixed list of terms and want dictionary-based extraction or replacement, FlashText offers a direct API and explicit longest-match behavior.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the published speed figures mean

Vikash Singh reported that a regex process took 24 hours on one million documents with 2,000 keywords. For an expanded workload of millions of documents and more than 10,000 keywords, he reported that the earlier approach would take more than 10 days, while his custom FlashText implementation completed the keyword-extraction workload in 15 minutes. These are the author’s 2017 implementation figures, not an independent, controlled comparison or a promise of the same result on other hardware, text, or software versions. The figures are described in the 2017 paper.

When FlashText is a good fit—and what to check

  • Consider it when you have a large, mostly fixed dictionary and need to extract terms or replace aliases throughout many documents.
  • Check boundaries with realistic samples if your text contains unusual punctuation, mixed scripts, or tokens whose boundaries do not follow ordinary word rules.
  • Inspect overlaps when one term is contained in a longer phrase, since the longest match takes precedence.
  • Choose regex instead when your matching problem depends on general pattern syntax or complex conditions that go beyond a keyword dictionary.
  • Verify maintenance fit before using it as a production dependency. PyPI’s listed 2.7 release was uploaded in 2018; the release date is a reason to check current repository and package activity, not proof by itself that the project is unsupported.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.