Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

JavaScript’s String.length counts UTF-16 code units—not necessarily Unicode code points, visible characters, or words. Use it when you need the units JavaScript uses for string indexing. For other jobs, choose a count that matches the requirement: code points, grapheme clusters, or language-aware word segments.

What does String.length count?

A JavaScript string is represented as UTF-16 code units, and text.length returns the number of those units. A character outside the Basic Multilingual Plane is represented by a surrogate pair, so that one Unicode code point contributes two to length. MDN explains that the result can therefore differ from the number of Unicode characters: String: length – JavaScript.

"A".length       // 1
"😀".length      // 2

This is not a bug: UTF-16 code units are the unit used by JavaScript’s string indexing model. It is simply the wrong measure if a product requirement means “how many characters will a person perceive?”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which unit should you count?

Need Count JavaScript approach What it does not tell you
JavaScript string indexing UTF-16 code units text.length Not a count of code points or perceived characters.
Unicode code points Code points [...text].length Combining marks and multi-code-point emoji may still count separately.
User-facing character limit Grapheme clusters Intl.Segmenter with granularity: "grapheme" Not a byte count or rendered-width measurement.
Words in text Word-like segments Intl.Segmenter with granularity: "word"; count segments whose isWordLike is true Results follow segmentation rules and the selected locale.

How do you count Unicode code points?

For a code-point count, spread the string into an array or iterate over it. JavaScript string iteration treats a valid surrogate pair as one code point, unlike indexing or length.

const codePointCount = (text) => [...text].length;

codePointCount("😀"); // 1

Code points are not the same as characters users perceive. A letter followed by a combining mark can be two code points, and an emoji can be built from multiple code points. MDN’s string guide describes these distinctions and emoji sequences: String – JavaScript.

How do you count emojis and other perceived characters?

Use grapheme segmentation when the requirement is a user-facing character limit. A grapheme cluster is an approximation of a user-perceived character; it can keep together a base letter with combining marks or an emoji sequence that consists of several code points. Unicode’s default segmentation rules define boundaries for grapheme clusters, words, and sentences: Unicode Text Segmentation, UAX #29, Version 49.

const graphemeSegmenter = new Intl.Segmenter("en", {
  granularity: "grapheme",
});

const graphemeCount = (text) =>
  [...graphemeSegmenter.segment(text)].length;

Choose a locale appropriate to the content or application, and test the target runtimes and locale support if the count controls validation or stored data. MDN documents Intl.Segmenter for grapheme-level segmentation: Internationalization – JavaScript.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you count words, including text without spaces?

Splitting on whitespace is a quick approximation, but it does not reliably identify words around punctuation or in languages that do not separate words with spaces. Use word segmentation and count only the segments marked isWordLike.

const wordSegmenter = new Intl.Segmenter("en", {
  granularity: "word",
});

const wordCount = (text) =>
  [...wordSegmenter.segment(text)]
    .filter((part) => part.isWordLike).length;

Set the locale to suit the text being analyzed rather than assuming one segmentation choice fits every language. Intl.Segmenter provides locale-sensitive word segmentation, including cases where whitespace splitting is insufficient, as described in MDN’s internationalization guide.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Are character counts the same as bytes or display width?

No. A grapheme-cluster count is useful for a user-facing character limit, but it does not measure storage or transport size. If a system imposes a byte limit, measure bytes in the required encoding separately. Nor does grapheme count determine visual width: rendered width depends on how text is displayed, so use an appropriate layout or measurement method for that requirement.

Practical implementation checklist

  • Use text.length for UTF-16 code units and JavaScript indexing.
  • Use [...text].length for Unicode code points, not for perceived characters.
  • Use grapheme segmentation for a user-facing character count.
  • Use word segmentation and isWordLike for word counts.
  • Select a suitable locale and verify runtime support when the count affects validation or persistence.
  • Measure bytes or rendered width separately when those are the actual constraints.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.