The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →To render Unicode correctly with wkhtmltoimage, make the entire pipeline UTF-8 and provide fonts that contain every script you use. Save the HTML as UTF-8, put <meta charset="utf-8"> early in the document head, pass --encoding UTF-8 on the command line, and install a suitable fallback font in the same environment and user context that runs the renderer. Encoding fixes misread bytes; fonts supply glyphs; neither step alone guarantees correct Arabic shaping, Indic joining, combining marks, or emoji.
The reliable fix, in one command and one HTML file
Start with a UTF-8 fixture and force the renderer’s input encoding:
wkhtmltoimage --encoding UTF-8 unicode-test.html unicode-test.png
The corresponding HTML should declare its charset before any content that depends on decoding:
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<style>
body { font-family: "Noto Sans", "DejaVu Sans", sans-serif; }
</style>
</head>
<body>
<p>English — Ελληνικά — Русский — 中文 — العربية — हिन्दी — 日本語 — 😀</p>
</body>
</html>
A wkhtmltopdf project issue records a Unicode problem that was fixed by adding --encoding UTF-8 (a 2018 user report). The libwkhtmltox documentation likewise says that settings supplied to PDF and image C bindings must be UTF-8 encoded strings. Treat those as separate checks: the command-line flag controls decoding of input, while CSS and installed fonts determine whether characters can be drawn.
#1 Best Overall
- Used Book in Good Condition
Understand the three failure layers
1. Bytes are decoded with the wrong character set
If UTF-8 bytes are interpreted as a legacy code page, accented text can become question marks, replacement characters, or unrelated symbols. This happens before font selection. Ensure the file is actually saved as UTF-8, declare the charset, and pass the explicit renderer option.
2. The selected fonts lack glyphs
A correct Unicode code point still renders as an empty square when no installed font contains its glyph. A desktop may have broad coverage while a minimal container or service account has only a few Latin fonts. Install a font covering the required script and make it visible to the same runtime user. Keep a CSS fallback stack such as "Noto Sans", "DejaVu Sans", sans-serif.
3. The browser engine cannot shape the script
Arabic joining, Indic shaping, combining marks, and emoji involve more than individual glyph lookup. After bytes and fonts are verified, the bundled legacy Qt WebKit engine may still produce incorrect shaping or missing emoji. In that case, changing the flag or adding another Latin font will not solve the rendering limitation; compare with a modern renderer or change the capture approach.
A repeatable diagnostic procedure
- Verify the source bytes. Open the HTML with a hex or text inspection tool. Confirm that the non-ASCII text is UTF-8 rather than a legacy code page. If an application supplied the text, decode incoming bytes explicitly as UTF-8 before constructing the HTML.
- Declare the charset early. Put
<meta charset="utf-8">in the<head>, before content that depends on decoding. Do not rely on a browser guessing from a late declaration. - Force the renderer setting. Run
wkhtmltoimage --encoding UTF-8 input.html output.pngand record the exact wkhtmltoimage version used. Reproducing the same binary matters when comparing machines. - Reduce the case to one line. Use the fixture above with a Latin accent, Greek or Cyrillic, one CJK character, an Arabic word, a Hindi word, Japanese text, and an emoji. A minimal fixture tells you whether the failure is decoding, coverage, or shaping.
- Check glyph coverage. Install a font containing the missing script and retain a CSS fallback list. Confirm that the runtime user, container, or server account can discover that font; installing it only for your desktop login does not make it available to a service.
- Check shaping behavior. If Latin, CJK, and isolated characters work but Arabic joining, Indic reordering, combining marks, or emoji remain wrong, test the same fixture with another rendering engine. The legacy WebKit component may be the constraint.
- Compare environments. Render with the same container image or server installation used in production, including font packages, locale, binary version, and user permissions. A result from a developer workstation is not proof that the deployment image has the same resources.
Application and Qt integration
The command-line option cannot repair text that your application has already decoded incorrectly. Keep the conversion explicit at every boundary. Qt’s encoding guidance notes that in Qt 4, constructing a QString from a narrow const char * can use Latin-1; use an explicit UTF-8 conversion instead.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #2
- Used Book in Good Condition
// Prefer an explicit UTF-8 conversion in Qt 4
QString title = QString::fromUtf8(utf8Bytes);
The same principle applies to language wrappers around libwkhtmltox: pass a Unicode string or a known UTF-8 byte sequence, not a locale-dependent narrow string. The bindings’ settings are expected to be UTF-8 encoded. Keep the HTML declaration and the renderer option even when the wrapper claims to use Unicode, because they protect separate stages of the pipeline.
Font selection and deployment
Use a deliberate fallback stack
Specify fonts in CSS rather than depending on a machine’s default choice. Put the font with the broadest expected coverage first, then add practical fallbacks and a generic family. Test every script you publish; a stack that covers Cyrillic may still lack Devanagari, CJK extensions, or emoji.
Install fonts where the renderer runs
For a container, bake the required font packages into the image. For a service, install them for the account that launches wkhtmltoimage and verify that the process can read them. Re-run the minimal fixture after rebuilding the image. If boxes appear only in production, compare the image and user context before changing HTML.
Do not confuse a missing glyph with a decoding error
Question marks and replacement characters often indicate that bytes were decoded incorrectly. Empty squares with otherwise correct punctuation usually indicate missing glyph coverage. Capture the same string in a font known to support the script to separate the two cases.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallRank #3
Commands and fixtures you can keep in a regression test
Store the UTF-8 fixture in source control and render it in the same job that builds your image or PDF assets:
wkhtmltoimage --encoding UTF-8 unicode-test.html artifacts/unicode-test.png
Include strings that exercise your actual product: accented names, right-to-left text, combining marks, CJK, Devanagari, and the emoji you permit. Compare the output after changes to the binary, base image, font packages, or CSS. The goal is not a universal benchmark; it is an early warning that a deployment no longer has the bytes, glyphs, or shaping behavior your pages require.
Troubleshooting by symptom
| Symptom | Likely cause | Fix |
|---|---|---|
| Every non-ASCII character becomes a question mark or replacement symbol | Input was decoded with a legacy encoding, or the renderer was not told to use UTF-8 | Verify the file bytes, decode application input explicitly as UTF-8, add the early meta declaration, and run with --encoding UTF-8. |
| Latin text works but Chinese, Arabic, Hindi, or Japanese is a square | No installed font contains the required glyph | Install a font covering that script, add a CSS fallback stack, and verify visibility to the production user or container. |
| Only one deployment environment fails | Different fonts, image, wkhtmltoimage version, locale, or permissions | Reproduce with the exact production image and account; compare font packages and binary versions rather than changing the source text first. |
| Arabic letters do not join or Indic marks are misplaced | Legacy Qt WebKit shaping limitation after bytes and glyphs are correct | Confirm with the minimal fixture, then evaluate a renderer with the shaping support your document needs. |
| Emoji is absent while surrounding text is correct | Missing emoji glyphs or an engine limitation | Check emoji-capable font coverage and test the same file with another engine. The wkhtmltoimage engine may not support the required emoji behavior. |
| Adding the meta tag changes nothing | The source bytes are already wrong, or the character is not in any available font | Inspect bytes first, then inspect font coverage. Charset declaration and font fallback solve different problems. |
Reliability, performance, and maintenance considerations
Unicode correctness is primarily a reproducibility problem. Keep the renderer version, HTML fixture, CSS font stack, font packages, container image, and runtime user under the same deployment controls. Rendering a small diagnostic page before a large batch can fail fast when a base image loses a font.
Adding more fallback fonts can increase installation size and font discovery work, but it is safer than depending on an undocumented host default. Conversely, no amount of extra font coverage fixes a renderer that cannot shape a script or emoji sequence correctly. Decide whether your output needs simple multilingual glyphs or full modern-browser shaping before committing to wkhtmltoimage for a long-lived pipeline.
Rank #4
When you need to preserve wkhtmltoimage for existing templates, isolate the Unicode-sensitive pages and test them explicitly. When shaping limitations are unacceptable, compare a modern browser-based capture path using the same fixture and deployment constraints. The comparison should cover encoding control, font fallback, script shaping and emoji, container reproducibility, and the maintenance status of the underlying browser engine.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your goal is a clean screenshot rather than maintaining a wkhtmltoimage runtime, ScreenshotNeo accepts one request and returns a PNG, JPEG, WebP, or PDF. It removes cookie-consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools to Claude, Cursor, and other MCP clients.
For a direct call, see the ScreenshotNeo API documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
You can still control capture details when needed: full-page loading, CSS-selector element capture, dark mode, device or custom viewport, retina scale, PDF paper and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, ad and tracker blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk calls for up to 100 URLs, and usage reporting. Every feature is included on every plan.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; Growth is $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing provides two months free. Create a free ScreenshotNeo account to start without a card.
Best Value
Frequently Asked Questions
Does the --encoding UTF-8 option install fonts?
No. It controls how wkhtmltoimage decodes input. Missing squares still require a font that contains the glyph and is visible to the renderer’s runtime user.
Why can isolated Arabic or Hindi characters look correct while words do not?
Individual glyph coverage can be present even when the bundled Qt WebKit engine cannot perform the joining, reordering, or combining behavior required by the script.
Do image and PDF bindings need different Unicode handling?
The libwkhtmltox documentation describes UTF-8 encoded strings for settings passed to both PDF and image C bindings, so keep the same explicit UTF-8 handling for either output.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

