Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a hash to track each token and Ruby’s scan method to find tokens in the file. For a file that fits comfortably in memory, File.read is concise; for a large file, File.foreach reads it line by line. Your regular expression determines what counts as a word, so choose case and punctuation rules deliberately.

Count words in a file with Ruby

This whole-file version follows the example in the official Ruby FAQ:

freq = Hash.new(0)
File.read("example").scan(/w+/) { |word| freq[word] += 1 }
freq.keys.sort.each { |word| puts "#{word}: #{freq[word]}" }

Replace example with your file’s path. Hash.new(0) gives an unseen token a starting count of zero, so each match can be counted with freq[word] += 1. The call to scan(/w+/) finds matches and the final loop prints them in alphabetical order.

For the FAQ’s sample input, the output is:

and: 1
is: 3
line: 3
one: 1
this: 3
three: 1
two: 1

Process a large file line by line

File.read loads the file contents as a whole. If that is not suitable for your input size, use File.foreach, which calls its block with each successive line, and scan each line as it arrives. Ruby documents this behavior in its IO reference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
freq = Hash.new(0)

File.foreach(path) do |line|
  line.scan(/w+/) { |word| freq[word] += 1 }
end

freq.sort_by { |word, count| [-count, word] }.each do |word, count|
  puts "#{word}: #{count}"
end

This still keeps the frequency hash in memory, and that hash grows with the number of distinct tokens. The final sort ranks words by count from highest to lowest; when counts tie, the word is used as an alphabetical tie-breaker. To print alphabetical order instead, use freq.keys.sort as in the first example.

Choose what counts as a word

The FAQ’s /w+/ pattern is a practical baseline, not a universal definition of a word. It finds runs of word characters; punctuation separates matches. As a result, a form such as don't is split at the apostrophe, and well-being is split at the hyphen. Numbers may also be included in matches.

For other token rules—such as preserving apostrophes, treating hyphenated forms as one token, or counting multilingual text—adjust the regular expression or use a tokenizer designed for the text. The rule you choose changes the counts.

Decide whether capitalization matters

The FAQ example is case-sensitive: Ruby and ruby become separate hash keys. To combine them, normalize each match before incrementing its count:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
freq = Hash.new(0)

File.foreach(path) do |line|
  line.scan(/w+/) do |word|
    word = word.downcase
    freq[word] += 1
  end
end

Choose this only if case distinctions should not matter for your task; otherwise, leave the original token unchanged.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Account for the file’s encoding

Ruby’s File documentation describes text-mode defaults, including UTF-8 as the default external encoding and BOM detection for UTF-8 and UTF-16 variants. For multilingual or irregularly encoded files, confirm the input encoding and decide how invalid byte sequences should be handled. The token pattern also needs to match the characters you intend to count.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.