Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To highlight matches in PDF, Word, or other Tika-extracted text, make sure Solr indexes that text in the field you search, that the field is stored, and that your request enables highlighting for that field. Start with Solr’s Unified Highlighter: it is the default and generally the best first choice. For long documents, choose an offset strategy deliberately because faster highlighting can require more index storage.

How Solr highlighting works with Tika-extracted files

Solr Cell uses Apache Tika to extract text and metadata from binary documents such as PDFs and Office files. Solr can highlight that text only after it has been mapped into an indexed field and the request asks Solr to highlight that field. Tika does not independently produce search-result highlights.

  1. Extract: Solr Cell’s ExtractingRequestHandler passes the file to Tika for parsing. The extraction module must be enabled.
  2. Map: Configure the extracted text to go into the Solr field you intend to query and display. For example, Solr Cell can map Tika’s content output to a field such as _text_ with fmap.content.
  3. Index: Ensure the destination field is indexed and stored. The stored value gives standard hl.fl highlighting access to the text it needs to return in snippets.
  4. Query and highlight: Search the intended field and request snippets from that same stored field.

Keep the three field roles straight: the field receiving extracted text, the field targeted by the query, and the field named for highlighting. They may be the same field, but if they are not, configure and test the relationship explicitly. A correctly parsed file can still yield no highlight if its extracted text was mapped elsewhere.

Request highlight snippets

Use hl=true to enable highlighting and hl.fl to name the field whose text should appear in snippets. For example, if your query searches content:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Solr in Action
  • Used Book in Good Condition
q=search terms
hl=true
hl.method=unified
hl.fl=content
hl.snippets=2
hl.fragsize=180
hl.tag.pre=<mark>
hl.tag.post=</mark>
hl.encoder=html

This is a parameter example, not a complete request URL: add these parameters to the request mechanism and endpoint used by your Solr deployment. Solr returns highlights in a separate highlighting section of the response, keyed by document ID and field, rather than replacing the document’s ordinary field values.

  • hl.snippets sets the maximum number of snippets per field.
  • hl.fragsize sets an approximate fragment size; it is not a promise of an exact character count.
  • hl.tag.pre and hl.tag.post wrap matched text in the chosen markup.
  • hl.encoder=html escapes the stored text for HTML output while leaving the configured highlight tags unescaped. Retain that escaping when rendering snippets so document text is not treated as markup.

Use a stored text field for this standard highlighting workflow. Also align the field analyzer with the query field’s analysis: if the terms produced at index or query time do not line up with the text being highlighted, the expected words may not be marked.

Choose a highlighter

Unified Highlighter: the starting point

hl.method=unified selects the Unified Highlighter, which is Solr’s default and recommended starting point for most workloads. It follows Lucene query semantics more accurately than the Original Highlighter and supports multiple offset sources. That makes it a sensible general choice for query types and documents that are not all simple, short text fields.

When to evaluate another method

A default is not a substitute for validating your query patterns and latency requirements. If your application depends on unusual query types, very long fields, or a strict response-time target, test representative documents and queries with the deployed Solr version before settling on a method and offset configuration. The official guidance provides configuration options, not a universal performance benchmark that predicts which setup will be fastest for your corpus.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an offset strategy for long fields

The highlighter needs to locate matches in the text. Solr can obtain offsets from analysis at query time or from information recorded in the index. The trade-off is typically less index overhead versus less highlighting work on long fields.

Strategy Field configuration Trade-off and suitable use
Analysis offsets No additional offset storage option specified Lowest index overhead among these choices, but Solr does more work analyzing stored text during highlighting. This work grows with the amount and complexity of text.
Postings offsets storeOffsetsWithPositions=true Adds index data, but can substantially speed highlighting on long fields by making offsets available from postings.
Light term vectors termVectors=true without the other term-vector options Adds index data. Consider this when wildcard highlighting on large fields matters; it can avoid falling back to analysis for wildcard queries.
Full term vectors Term vectors, positions, and offsets enabled Adds substantial index weight. It is mainly justified when another application use already requires full term vectors.

These are schema-level trade-offs, not request parameters. Compare them using the field sizes, query mix, index capacity, and latency targets of your own deployment. Avoid enabling full term vectors solely because highlighting is needed without first checking whether a lighter offset source meets the requirement.

Configure Tika extraction and field mapping

In Solr Cell, fmap.content can map Tika’s extracted content to a Solr field such as _text_. Make hl.fl name the actual stored destination field, not a Tika metadata name or an unrelated field. The field also needs to be the one your query can match, with compatible analysis.

The capture parameter can copy selected XHTML elements, such as paragraphs, into supplementary fields while preserving the main extracted content. This is useful when an application needs a separate field for a particular content structure; it does not remove the need to highlight the field containing the text that matched the query.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For Solr 10, the extraction backend is Tika Server. The Solr 10 configuration option tikaserver.recursive=true enables recursive extraction of embedded documents, such as email attachments or files inside archives. Confirm the behavior and parameter names against the documentation for the exact Solr release you deploy, since extraction backends and defaults vary by release.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Diagnose empty or unexpected snippets

  • No highlight field in the response: Confirm hl=true, check that hl.fl names a valid field, and inspect the response’s highlighting section under the matching document ID.
  • The field is returned but has no snippet: Check that the field is stored and that extracted text actually landed in it. Then verify that the query matched that field and that field analysis is compatible with query analysis.
  • Highlights disappear with a multi-field query: Check hl.requireFieldMatch. When enabled, it can exclude a highlight field that does not match the query field.
  • Phrase or wildcard matches are absent: Review hl.usePhraseHighlighter and hl.highlightMultiTerm. Both default to true in the cited Solr guide, but behavior and defaults should be checked for your deployed release.
  • Long text is only partly considered: Review hl.maxAnalyzedChars. The cited guide gives a default of 51,200 characters; verify the applicable release setting and select an offset source suited to the field size.
  • Files with attachments or embedded content are incomplete: Check whether recursive extraction is enabled and supported for your Solr version and extraction setup. Test representative embedded documents rather than assuming every parser or file structure behaves identically.

Protect Solr when parsing complex or untrusted files

The Solr 9.10 guide warns that local, in-process extraction can allow parser failures to affect the Solr JVM. Running extraction through an external Tika Server provides process isolation and allows it to scale independently. This is an operational choice as well as an extraction configuration choice: assess file trust, parser complexity, and how the service will be operated in your environment.

Validate the complete workflow

Before relying on highlights in production, exercise the path from uploaded file to displayed snippet with documents that resemble your actual inputs. Include PDFs, Office files, encrypted files, and documents containing embedded attachments where those formats are in scope. For each, verify the extracted text, destination field, query match, response snippet, and rendered markup. This distinguishes extraction or mapping failures from highlighting and display problems.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.