Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To add syntax colors to a code diff without disturbing the added and deleted styling, highlight the complete old and new versions of each file separately, convert the highlighter’s output into character ranges, apply those ranges to the original source text, and fall back to plain text whenever the result cannot be trusted. The green and red change backgrounds stay exactly where they were, and keywords, strings, and comments gain a second visual cue inside each line.
Why a plain diff is hard to read
Reviewers of a pull request often see the same symptom: the code is technically visible, but almost all of it appears in one color. In a write-up on this approach, the author describes a pull-request reading guide where the diff already marked additions with green backgrounds and deletions with red backgrounds, and already showed line numbers and definition links. Those cues answer the question “what changed?” They do not answer “what kind of token is this?” A keyword, a string literal, and a comment look alike, so the reader has to work out the structure of each line by hand.
Syntax colors address that second question. The goal is not to replace the change signal. It is to add a layer of token information on top of it, so that a line can be both marked as added and readable as code.
Why the diff cannot be highlighted line by line
The obvious shortcut is to run a highlighter over the lines shown in the diff. It fails for a structural reason: a diff view interleaves lines from two different file versions. Deleted lines belong to the old file, added lines belong to the new file, and unchanged context lines belong to both. Syntax is not local to a single line. A string that opens on one line and closes several lines later, or a block comment that begins above the visible hunk, changes how every later token is classified.
#1 Best Overall
If the highlighter reads only the visible fragment, it can guess the wrong state at the top of the fragment. The colors then look plausible but are wrong, which is worse than no colors at all.
The five parts of the approach
1. Parse each file version independently
Build two complete snapshots for each changed file: the old version and the new version. Parse both in full, including the context that the diff collapses. Each snapshot gets its own tokenization, so a multiline string that is open in the old file cannot influence the parse of the new file, and the reverse is also true. A line that the diff shows as deleted is highlighted using the old snapshot’s state; a line shown as added uses the new snapshot’s state.
The cost is that the renderer needs the full file contents, not just the hunk text. Collapsed regions must still be fetched or reconstructed, even though they are not drawn.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsRank #2
2. Treat the highlighter output as token ranges, not markup
The reported implementation uses highlight.js to produce token ranges. It does not insert the highlighter’s generated HTML into the page. Instead, it checks that the decoded text from the highlighter matches the source text, then applies each token’s start and end positions to the original string. The page’s own text remains the authoritative copy, and the highlighter contributes only the classification of spans within it.
This separation matters for safety and for correctness. Generated HTML can carry its own escaping rules, and any mismatch between that HTML and the source can alter what the reader sees. Working with ranges keeps the displayed characters identical to the file’s characters.
3. Verify equivalence before applying anything
Before any color is applied, the renderer compares the highlighter’s decoded text with the source text. If they differ, the file is rendered as plain code. This check is the gatekeeper for the whole approach: ranges computed against one string must never be painted onto a slightly different one.
Rank #3
4. Keep character positions aligned
Other features in the same interface rely on positions in the source. In the reported case, definition links use character offsets into the same source text. If the colored output adds, removes, or normalizes any character, those offsets drift, and links land on the wrong symbol. Whitespace and Unicode need particular care, because a single visible glyph can occupy more than one code unit. The token ranges therefore have to be computed on the same string the links use, with no normalization step in between.
Free tools Windows power users keep installed
One-click scans. No signup required.
5. Layer colors on top of change backgrounds
The added and deleted backgrounds stay as they were. Token colors are applied as an additional layer inside each line, so a green added line still reads as added, and a keyword inside it is also visually distinct. The article’s own design principle is stated plainly: “A color feature should not prevent the diff from rendering.”
What happens when highlighting fails
The fallback is what makes the feature safe to ship. Any of the following conditions causes the affected file to render as plain code, with its change backgrounds and line numbers intact:
- The file type is unknown, or no lexer is registered for its extension.
- The highlighter throws an error while tokenizing.
- The decoded highlighter text does not match the source text.
The failure is contained to the one file. The rest of the pull request continues to render with colors where they succeed.
Edge cases the reported tests cover
The article describes regression tests for four categories of input. Each one targets a place where token ranges commonly go wrong:
- Multiline strings, where the token state carries across line boundaries.
- Collapsed context, where the visible hunk is a subset of the file being parsed.
- Renamed files, where the old and new paths differ and the wrong language or snapshot could be selected.
- Emoji and HTML-like characters, where multi-unit characters and literal angle brackets test both offset handling and escaping.
Checklist for building the same behavior in your own renderer
- Fetch the complete old and new contents of each changed file, not only the hunk text.
- Choose the language from the file’s path, and decide in advance what happens for renamed files (use the new path for the new snapshot and the old path for the old snapshot).
- Tokenize each snapshot separately and capture token start and end offsets.
- Decode the highlighter’s text and compare it with the source. On mismatch, render plain code.
- Apply ranges to the original string as wrapper elements, escaping all text as you go.
- Confirm that any link or selection feature still uses offsets computed on the unmodified source.
- Add regression tests for multiline strings, collapsed context, renames, emoji, and angle brackets before enabling the feature.
What the source does and does not establish
The description above comes from a September 18, 2026 article by AIWithGhost, “Adding syntax colors without changing the diff.” The source was available to us as a search-result excerpt; the full page could not be opened directly at the time of review, so the details above are limited to what that excerpt reports. The article is a description of one implementation, not an independent audit of its code.
Best Value
- Used Book in Good Condition
The source also does not establish that highlight.js is the only suitable choice. Its claims are about a specific design. Comparing it with other approaches would require evaluating source-text preservation, correct handling of old and new snapshots, offset stability, behavior on unknown or malformed input, and rendering cost, and the source does not report measurements for the last of these.
The screenshots in the source show a presentation change. They do not measure whether reviews become faster or more accurate, and nothing here should be read as evidence of that effect.
Original article: https://aiwithghost.com/news/news-adding-syntax-colors-without-changing-the-diff/
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

