Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sumy is a free, open-source Python package and command-line tool for extractive text summarization. It ranks and selects sentences from plain text, local files, or HTML rather than generating rewritten prose. The current PyPI release observed on August 18, 2026 is Sumy 0.12.0, which requires Python 3.8 or newer. Install it locally, choose among eight classical algorithms, and keep source-level traceability without an API key or cloud account.

What automated summarization means

Extractive summarization selects sentences (or sentence fragments) already present in a document. Abstractive summarization generates new wording, usually with a language model. Sumy is primarily an extractive, single-document toolkit: it is designed to condense one document, not synthesize evidence across a collection or rewrite text in a requested style.

That distinction explains both its appeal and its limits. Selected sentences are easy to compare with the source and can run offline, but Sumy will not reliably resolve pronouns, reconcile contradictions, or explain facts spread across several sentences.

What the Sumy library provides

  • Python API and a command-line interface.
  • Plain-text, local-file, and HTML parsing.
  • Sentence-count and percentage-based summary lengths.
  • Eight classical summarizers: Luhn, Edmundson, LSA, LexRank, TextRank, SumBasic, KL-Sum, and Reduction.
  • A sumy_eval utility for comparing output with a reference summary.
  • Apache License 2.0 metadata and local execution without a mandatory subscription or API key.

See the PyPI package page and the official repository for release-specific details. The package classifiers list numerous languages, but tokenizer, stemming, stop-word coverage, and practical quality vary by language and version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Sumy

Use Python 3.8 or newer for the current package metadata:

python --version
python -m pip install sumy

The project also documents uv and a Git installation:

uv pip install sumy
uv pip install git+https://github.com/miso-belica/sumy.git

Verify the command-line entry point:

sumy --help

If the command is missing, activate the virtual environment that received the installation or call the executable from that environment’s binary directory. Do not name your application sumy.py or create a local directory named sumy; either can shadow the installed package.

Your first Sumy summarizer

from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer
from sumy.nlp.stemmers import Stemmer
from sumy.utils import get_stop_words

LANGUAGE = "english"
SENTENCES_COUNT = 3

text = """
Python is a widely used programming language. It is popular for automation,
web development, data analysis, and machine learning. Its large ecosystem
contains libraries for many different tasks. Developers often choose Python
because its syntax is relatively easy to read and its community is large.
"""

parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))
stemmer = Stemmer(LANGUAGE)
summarizer = LsaSummarizer(stemmer)
summarizer.stop_words = get_stop_words(LANGUAGE)

for sentence in summarizer(parser.document, SENTENCES_COUNT):
    print(sentence)

Tokenizer determines sentence and word boundaries. The optional Stemmer and stop-word list help LSA compare related terms. The summarizer returns sentence objects; iterating over them prints the selected source sentences.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Summarize a local text file

from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lex_rank import LexRankSummarizer

parser = PlaintextParser.from_file("article.txt", Tokenizer("english"))
summarizer = LexRankSummarizer()

for sentence in summarizer(parser.document, 5):
    print(sentence)
  • Normalize files to UTF-8 when you preprocess them.
  • Reject empty or near-empty input before calling the algorithm.
  • Keep original sentence positions if you need audit trails or later reordering.
  • Escape or sanitize generated output before inserting it into HTML.

Summarize an HTML page

from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lex_rank import LexRankSummarizer

parser = HtmlParser.from_url(
    "https://example.com/article",
    Tokenizer("english")
)
summarizer = LexRankSummarizer()

for sentence in summarizer(parser.document, 5):
    print(sentence)

URL parsing is not a full article-extraction system. Navigation, cookie notices, comments, advertisements, malformed markup, client-side rendering, login walls, rate limits, robots restrictions, or a network outage can all produce incomplete or irrelevant text. For production, fetch the page with a controlled HTTP client, check status and timeouts, extract the article body, and pass cleaned text to PlaintextParser.

Use Sumy from the command line

sumy lex-rank --length=10 
  --url=https://en.wikipedia.org/wiki/Automatic_summarization

sumy lex-rank --language=uk --length=30 
  --url=https://uk.wikipedia.org/wiki/Україна

sumy luhn --language=czech 
  --url=https://www.zdrojak.cz/clanky/automaticke-zabezpeceni/

sumy edmundson --language=czech --length=3% 
  --url=https://cs.wikipedia.org/wiki/Bitva_u_Lipan

--length accepts a sentence count or a percentage in the documented examples. Run sumy --help against the installed release because options and defaults can change.

How Sumy’s algorithms differ

LSA

Latent Semantic Analysis represents term relationships and favors sentences associated with important concepts. It is a useful first baseline for documents with several themes, but short inputs provide little statistical signal and sentence order may need restoration.

LexRank

LexRank builds a similarity graph in which sentences are nodes; graph centrality, inspired by PageRank, identifies sentences representative of the document. It is a strong general baseline for informational or news-like text, not a universally best method. The original method is described at arXiv.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TextRank

TextRank also ranks a sentence-similarity graph. It belongs to the same broad family as LexRank, but the implementations and scoring details are not identical.

Luhn

This heuristic emphasizes clusters of significant terms. It can suit keyword-heavy technical writing, while overvaluing repeated terminology and underrepresenting context.

Edmundson

Edmundson can combine cue words, title relevance, and sentence position. It is most useful when your application can define domain-specific signals.

SumBasic

SumBasic favors frequent words and is a simple research baseline. Frequency can capture a topic, but repeated terms can make the result redundant.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

KL-Sum

KL-Sum greedily chooses sentences that make the summary’s word distribution resemble the source. It can improve vocabulary coverage, although greedy choices do not guarantee global coherence.

Reduction

Reduction scores sentences through their relationships with other sentences and is related to TextRank-style similarity methods. Test it empirically on your corpus rather than assuming a ranking advantage.

Choose an algorithm with a small evaluation matrix

Use case First algorithms to try Reason
General article LexRank, TextRank, LSA Useful classical baselines
Keyword-heavy technical material Luhn, LexRank Rewards salient terms and central sentences
Several themes LSA, LexRank Concept and centrality signals
Frequency baseline SumBasic Simple comparison point
Domain cue words Edmundson Explicit heuristic features
Vocabulary coverage KL-Sum Distribution-oriented selection
Research comparison Test several Results are corpus- and task-dependent

Hold the corpus, language, summary length, metrics, and post-processing constant. Compare at least three algorithms on representative documents before choosing one.

Language, tokenization, and ordering

LANGUAGE = "english"
parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))

Use the exact language identifier accepted by the installed release. Declared language metadata does not guarantee equal segmentation or stemming quality for every script. Test a short sample first, preserve Unicode punctuation and accents, and avoid silently stripping non-Latin characters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Importance ranking is not narrative ordering. If the output feels disjointed, retain each selected sentence’s original index and sort by source position, unless your application specifically needs rank order.

Evaluate summary quality

With a reference summary, run the documented evaluator:

sumy_eval lex-rank reference_summary.txt 
  --url=https://en.wikipedia.org/wiki/Automatic_summarization

sumy_eval lsa reference_summary.txt 
  --language=czech 
  --url=https://www.zdrojak.cz/clanky/automaticke-zabezpeceni/

Lexical-overlap metrics can help detect regressions, but they are not a verdict. Review coverage of essential points, redundancy, factual consistency with the source, sentence order, readability, and usefulness for the actual task. A summary can overlap strongly with a reference yet omit a crucial qualification.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Import or command errors

  • Install into the interpreter you run: python -m pip install --upgrade sumy.
  • Check the environment with python -c "import sumy; print(sumy)".
  • Activate the correct virtual environment.
  • Rename local sumy.py files or sumy directories.

Tokenizer or language errors

Verify the language name supported by your installed version and test tokenization on a short string before processing a large document.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Empty or poor summaries

  1. Print the parsed document to confirm that useful sentences were extracted.
  2. Remove HTML boilerplate and repeated headings.
  3. Lower the requested sentence count when the source is short.
  4. Try LexRank, LSA, and TextRank.
  5. Deduplicate or reorder selected sentences in application code.

Network and encoding problems

Handle HTTP status codes, timeouts, and inaccessible pages outside Sumy. Normalize input to UTF-8 and preserve Unicode characters.

Short or sensitive documents

A two-sentence input cannot yield a meaningful five-sentence summary. Local execution reduces transmission to a hosted service, but privacy compliance still requires appropriate access controls, logging, retention, and storage practices.

Sumy compared with modern summarization options

Option Advantages Trade-offs
Sumy Local, lightweight, extractive, source-traceable Limited fluency, no rewriting or deep synthesis
Custom NLTK, spaCy, or Gensim pipeline Custom linguistic features and scoring More engineering; not a drop-in Sumy CLI replacement
Transformer or local LLM Abstractive wording and broader synthesis Model size, hardware, latency, validation, and licensing
Cloud model API Managed scale, long context, instructions, multimodal workflows Usage cost, vendor dependence, data governance, changing behavior

For hosted choices, compare current terms rather than treating prices as permanent: Hugging Face Inference Providers, its pricing, Gemini API pricing, Amazon Bedrock, and Claude pricing. A focused alternative implementation is lexrank; it is not automatically more accurate.

When Sumy is the right choice

  • Local or privacy-sensitive processing where extractive traceability is valuable.
  • Lightweight scripts, teaching, prototypes, and reproducible baselines.
  • Applications that can validate input extraction and review sentence coherence.

Choose a transformer or hosted model when you need fluent paraphrasing, style instructions, synthesis across many documents, or managed scaling. Sumy is useful because it is simple and transparent—not because it offers the capabilities of a modern generative model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.