Sumy is a free, open-source Python package and command-line tool for extractive text summarization. It ranks and selects sentences from plain text, local files, or HTML rather than generating rewritten prose. The current PyPI release observed on August 18, 2026 is Sumy 0.12.0, which requires Python 3.8 or newer. Install it locally, choose among eight classical algorithms, and keep source-level traceability without an API key or cloud account.
What automated summarization means
Extractive summarization selects sentences (or sentence fragments) already present in a document. Abstractive summarization generates new wording, usually with a language model. Sumy is primarily an extractive, single-document toolkit: it is designed to condense one document, not synthesize evidence across a collection or rewrite text in a requested style.
That distinction explains both its appeal and its limits. Selected sentences are easy to compare with the source and can run offline, but Sumy will not reliably resolve pronouns, reconcile contradictions, or explain facts spread across several sentences.
What the Sumy library provides
- Python API and a command-line interface.
- Plain-text, local-file, and HTML parsing.
- Sentence-count and percentage-based summary lengths.
- Eight classical summarizers: Luhn, Edmundson, LSA, LexRank, TextRank, SumBasic, KL-Sum, and Reduction.
- A
sumy_evalutility for comparing output with a reference summary. - Apache License 2.0 metadata and local execution without a mandatory subscription or API key.
See the PyPI package page and the official repository for release-specific details. The package classifiers list numerous languages, but tokenizer, stemming, stop-word coverage, and practical quality vary by language and version.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors#1 Best Overall
Install Sumy
Use Python 3.8 or newer for the current package metadata:
python --version
python -m pip install sumy
The project also documents uv and a Git installation:
uv pip install sumy
uv pip install git+https://github.com/miso-belica/sumy.git
Verify the command-line entry point:
sumy --help
If the command is missing, activate the virtual environment that received the installation or call the executable from that environment’s binary directory. Do not name your application sumy.py or create a local directory named sumy; either can shadow the installed package.
Your first Sumy summarizer
from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer
from sumy.nlp.stemmers import Stemmer
from sumy.utils import get_stop_words
LANGUAGE = "english"
SENTENCES_COUNT = 3
text = """
Python is a widely used programming language. It is popular for automation,
web development, data analysis, and machine learning. Its large ecosystem
contains libraries for many different tasks. Developers often choose Python
because its syntax is relatively easy to read and its community is large.
"""
parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))
stemmer = Stemmer(LANGUAGE)
summarizer = LsaSummarizer(stemmer)
summarizer.stop_words = get_stop_words(LANGUAGE)
for sentence in summarizer(parser.document, SENTENCES_COUNT):
print(sentence)
Tokenizer determines sentence and word boundaries. The optional Stemmer and stop-word list help LSA compare related terms. The summarizer returns sentence objects; iterating over them prints the selected source sentences.
Recommended Free Tools
Summarize a local text file
from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lex_rank import LexRankSummarizer
parser = PlaintextParser.from_file("article.txt", Tokenizer("english"))
summarizer = LexRankSummarizer()
for sentence in summarizer(parser.document, 5):
print(sentence)
- Normalize files to UTF-8 when you preprocess them.
- Reject empty or near-empty input before calling the algorithm.
- Keep original sentence positions if you need audit trails or later reordering.
- Escape or sanitize generated output before inserting it into HTML.
Summarize an HTML page
from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lex_rank import LexRankSummarizer
parser = HtmlParser.from_url(
"https://example.com/article",
Tokenizer("english")
)
summarizer = LexRankSummarizer()
for sentence in summarizer(parser.document, 5):
print(sentence)
URL parsing is not a full article-extraction system. Navigation, cookie notices, comments, advertisements, malformed markup, client-side rendering, login walls, rate limits, robots restrictions, or a network outage can all produce incomplete or irrelevant text. For production, fetch the page with a controlled HTTP client, check status and timeouts, extract the article body, and pass cleaned text to PlaintextParser.
Rank #2
Use Sumy from the command line
sumy lex-rank --length=10
--url=https://en.wikipedia.org/wiki/Automatic_summarization
sumy lex-rank --language=uk --length=30
--url=https://uk.wikipedia.org/wiki/Україна
sumy luhn --language=czech
--url=https://www.zdrojak.cz/clanky/automaticke-zabezpeceni/
sumy edmundson --language=czech --length=3%
--url=https://cs.wikipedia.org/wiki/Bitva_u_Lipan
--length accepts a sentence count or a percentage in the documented examples. Run sumy --help against the installed release because options and defaults can change.
How Sumy’s algorithms differ
LSA
Latent Semantic Analysis represents term relationships and favors sentences associated with important concepts. It is a useful first baseline for documents with several themes, but short inputs provide little statistical signal and sentence order may need restoration.
LexRank
LexRank builds a similarity graph in which sentences are nodes; graph centrality, inspired by PageRank, identifies sentences representative of the document. It is a strong general baseline for informational or news-like text, not a universally best method. The original method is described at arXiv.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
TextRank
TextRank also ranks a sentence-similarity graph. It belongs to the same broad family as LexRank, but the implementations and scoring details are not identical.
Luhn
This heuristic emphasizes clusters of significant terms. It can suit keyword-heavy technical writing, while overvaluing repeated terminology and underrepresenting context.
Edmundson
Edmundson can combine cue words, title relevance, and sentence position. It is most useful when your application can define domain-specific signals.
SumBasic
SumBasic favors frequent words and is a simple research baseline. Frequency can capture a topic, but repeated terms can make the result redundant.
KL-Sum
KL-Sum greedily chooses sentences that make the summary’s word distribution resemble the source. It can improve vocabulary coverage, although greedy choices do not guarantee global coherence.
Reduction
Reduction scores sentences through their relationships with other sentences and is related to TextRank-style similarity methods. Test it empirically on your corpus rather than assuming a ranking advantage.
Choose an algorithm with a small evaluation matrix
| Use case | First algorithms to try | Reason |
|---|---|---|
| General article | LexRank, TextRank, LSA | Useful classical baselines |
| Keyword-heavy technical material | Luhn, LexRank | Rewards salient terms and central sentences |
| Several themes | LSA, LexRank | Concept and centrality signals |
| Frequency baseline | SumBasic | Simple comparison point |
| Domain cue words | Edmundson | Explicit heuristic features |
| Vocabulary coverage | KL-Sum | Distribution-oriented selection |
| Research comparison | Test several | Results are corpus- and task-dependent |
Hold the corpus, language, summary length, metrics, and post-processing constant. Compare at least three algorithms on representative documents before choosing one.
Language, tokenization, and ordering
LANGUAGE = "english"
parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))
Use the exact language identifier accepted by the installed release. Declared language metadata does not guarantee equal segmentation or stemming quality for every script. Test a short sample first, preserve Unicode punctuation and accents, and avoid silently stripping non-Latin characters.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Importance ranking is not narrative ordering. If the output feels disjointed, retain each selected sentence’s original index and sort by source position, unless your application specifically needs rank order.
Evaluate summary quality
With a reference summary, run the documented evaluator:
sumy_eval lex-rank reference_summary.txt
--url=https://en.wikipedia.org/wiki/Automatic_summarization
sumy_eval lsa reference_summary.txt
--language=czech
--url=https://www.zdrojak.cz/clanky/automaticke-zabezpeceni/
Lexical-overlap metrics can help detect regressions, but they are not a verdict. Review coverage of essential points, redundancy, factual consistency with the source, sentence order, readability, and usefulness for the actual task. A summary can overlap strongly with a reference yet omit a crucial qualification.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot common failures
Import or command errors
- Install into the interpreter you run:
python -m pip install --upgrade sumy. - Check the environment with
python -c "import sumy; print(sumy)". - Activate the correct virtual environment.
- Rename local
sumy.pyfiles orsumydirectories.
Tokenizer or language errors
Verify the language name supported by your installed version and test tokenization on a short string before processing a large document.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Best Value
Empty or poor summaries
- Print the parsed document to confirm that useful sentences were extracted.
- Remove HTML boilerplate and repeated headings.
- Lower the requested sentence count when the source is short.
- Try LexRank, LSA, and TextRank.
- Deduplicate or reorder selected sentences in application code.
Network and encoding problems
Handle HTTP status codes, timeouts, and inaccessible pages outside Sumy. Normalize input to UTF-8 and preserve Unicode characters.
Short or sensitive documents
A two-sentence input cannot yield a meaningful five-sentence summary. Local execution reduces transmission to a hosted service, but privacy compliance still requires appropriate access controls, logging, retention, and storage practices.
Sumy compared with modern summarization options
| Option | Advantages | Trade-offs |
|---|---|---|
| Sumy | Local, lightweight, extractive, source-traceable | Limited fluency, no rewriting or deep synthesis |
| Custom NLTK, spaCy, or Gensim pipeline | Custom linguistic features and scoring | More engineering; not a drop-in Sumy CLI replacement |
| Transformer or local LLM | Abstractive wording and broader synthesis | Model size, hardware, latency, validation, and licensing |
| Cloud model API | Managed scale, long context, instructions, multimodal workflows | Usage cost, vendor dependence, data governance, changing behavior |
For hosted choices, compare current terms rather than treating prices as permanent: Hugging Face Inference Providers, its pricing, Gemini API pricing, Amazon Bedrock, and Claude pricing. A focused alternative implementation is lexrank; it is not automatically more accurate.
When Sumy is the right choice
- Local or privacy-sensitive processing where extractive traceability is valuable.
- Lightweight scripts, teaching, prototypes, and reproducible baselines.
- Applications that can validate input extraction and review sentence coherence.
Choose a transformer or hosted model when you need fluent paraphrasing, style instructions, synthesis across many documents, or managed scaling. Sumy is useful because it is simple and transparent—not because it offers the capabilities of a modern generative model.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

