To highlight matches in PDF, Word, or other Tika-extracted text, make sure Solr indexes that text in the field you search, that the field is stored, and that your request enables highlighting for that field. Start with Solr’s Unified Highlighter: it is the default and generally the best first choice. For long documents, choose an offset strategy deliberately because faster highlighting can require more index storage.
How Solr highlighting works with Tika-extracted files
Solr Cell uses Apache Tika to extract text and metadata from binary documents such as PDFs and Office files. Solr can highlight that text only after it has been mapped into an indexed field and the request asks Solr to highlight that field. Tika does not independently produce search-result highlights.
- Extract: Solr Cell’s
ExtractingRequestHandlerpasses the file to Tika for parsing. The extraction module must be enabled. - Map: Configure the extracted text to go into the Solr field you intend to query and display. For example, Solr Cell can map Tika’s
contentoutput to a field such as_text_withfmap.content. - Index: Ensure the destination field is indexed and stored. The stored value gives standard
hl.flhighlighting access to the text it needs to return in snippets. - Query and highlight: Search the intended field and request snippets from that same stored field.
Keep the three field roles straight: the field receiving extracted text, the field targeted by the query, and the field named for highlighting. They may be the same field, but if they are not, configure and test the relationship explicitly. A correctly parsed file can still yield no highlight if its extracted text was mapped elsewhere.
Request highlight snippets
Use hl=true to enable highlighting and hl.fl to name the field whose text should appear in snippets. For example, if your query searches content:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
q=search terms
hl=true
hl.method=unified
hl.fl=content
hl.snippets=2
hl.fragsize=180
hl.tag.pre=<mark>
hl.tag.post=</mark>
hl.encoder=html
This is a parameter example, not a complete request URL: add these parameters to the request mechanism and endpoint used by your Solr deployment. Solr returns highlights in a separate highlighting section of the response, keyed by document ID and field, rather than replacing the document’s ordinary field values.
hl.snippetssets the maximum number of snippets per field.hl.fragsizesets an approximate fragment size; it is not a promise of an exact character count.hl.tag.preandhl.tag.postwrap matched text in the chosen markup.hl.encoder=htmlescapes the stored text for HTML output while leaving the configured highlight tags unescaped. Retain that escaping when rendering snippets so document text is not treated as markup.
Use a stored text field for this standard highlighting workflow. Also align the field analyzer with the query field’s analysis: if the terms produced at index or query time do not line up with the text being highlighted, the expected words may not be marked.
Choose a highlighter
Unified Highlighter: the starting point
hl.method=unified selects the Unified Highlighter, which is Solr’s default and recommended starting point for most workloads. It follows Lucene query semantics more accurately than the Original Highlighter and supports multiple offset sources. That makes it a sensible general choice for query types and documents that are not all simple, short text fields.
When to evaluate another method
A default is not a substitute for validating your query patterns and latency requirements. If your application depends on unusual query types, very long fields, or a strict response-time target, test representative documents and queries with the deployed Solr version before settling on a method and offset configuration. The official guidance provides configuration options, not a universal performance benchmark that predicts which setup will be fastest for your corpus.
Free tools Windows power users keep installed
One-click scans. No signup required.
Choose an offset strategy for long fields
The highlighter needs to locate matches in the text. Solr can obtain offsets from analysis at query time or from information recorded in the index. The trade-off is typically less index overhead versus less highlighting work on long fields.
| Strategy | Field configuration | Trade-off and suitable use |
|---|---|---|
| Analysis offsets | No additional offset storage option specified | Lowest index overhead among these choices, but Solr does more work analyzing stored text during highlighting. This work grows with the amount and complexity of text. |
| Postings offsets | storeOffsetsWithPositions=true |
Adds index data, but can substantially speed highlighting on long fields by making offsets available from postings. |
| Light term vectors | termVectors=true without the other term-vector options |
Adds index data. Consider this when wildcard highlighting on large fields matters; it can avoid falling back to analysis for wildcard queries. |
| Full term vectors | Term vectors, positions, and offsets enabled | Adds substantial index weight. It is mainly justified when another application use already requires full term vectors. |
These are schema-level trade-offs, not request parameters. Compare them using the field sizes, query mix, index capacity, and latency targets of your own deployment. Avoid enabling full term vectors solely because highlighting is needed without first checking whether a lighter offset source meets the requirement.
Rank #4
Configure Tika extraction and field mapping
In Solr Cell, fmap.content can map Tika’s extracted content to a Solr field such as _text_. Make hl.fl name the actual stored destination field, not a Tika metadata name or an unrelated field. The field also needs to be the one your query can match, with compatible analysis.
The capture parameter can copy selected XHTML elements, such as paragraphs, into supplementary fields while preserving the main extracted content. This is useful when an application needs a separate field for a particular content structure; it does not remove the need to highlight the field containing the text that matched the query.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →For Solr 10, the extraction backend is Tika Server. The Solr 10 configuration option tikaserver.recursive=true enables recursive extraction of embedded documents, such as email attachments or files inside archives. Confirm the behavior and parameter names against the documentation for the exact Solr release you deploy, since extraction backends and defaults vary by release.
Diagnose empty or unexpected snippets
- No highlight field in the response: Confirm
hl=true, check thathl.flnames a valid field, and inspect the response’shighlightingsection under the matching document ID. - The field is returned but has no snippet: Check that the field is stored and that extracted text actually landed in it. Then verify that the query matched that field and that field analysis is compatible with query analysis.
- Highlights disappear with a multi-field query: Check
hl.requireFieldMatch. When enabled, it can exclude a highlight field that does not match the query field. - Phrase or wildcard matches are absent: Review
hl.usePhraseHighlighterandhl.highlightMultiTerm. Both default totruein the cited Solr guide, but behavior and defaults should be checked for your deployed release. - Long text is only partly considered: Review
hl.maxAnalyzedChars. The cited guide gives a default of 51,200 characters; verify the applicable release setting and select an offset source suited to the field size. - Files with attachments or embedded content are incomplete: Check whether recursive extraction is enabled and supported for your Solr version and extraction setup. Test representative embedded documents rather than assuming every parser or file structure behaves identically.
Protect Solr when parsing complex or untrusted files
The Solr 9.10 guide warns that local, in-process extraction can allow parser failures to affect the Solr JVM. Running extraction through an external Tika Server provides process isolation and allows it to scale independently. This is an operational choice as well as an extraction configuration choice: assess file trust, parser complexity, and how the service will be operated in your environment.
Validate the complete workflow
Before relying on highlights in production, exercise the path from uploaded file to displayed snippet with documents that resemble your actual inputs. Include PDFs, Office files, encrypted files, and documents containing embedded attachments where those formats are in scope. For each, verify the extracted text, destination field, query match, response snippet, and rendered markup. This distinguishes extraction or mapping failures from highlighting and display problems.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

