Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteAn inverted index maps each search term to the documents that contain it. In Elixir, you can build one with maps and Enum, keeping a term’s document postings and its frequency in each document. Reuse the same tokenizer for queries, then score matching documents with a clearly specified TF-IDF formula.
What the index stores
Use stable document IDs so each posting can refer back to its source text. For a small in-memory corpus, a map from IDs to text is enough:
documents = %{
"doc_1" => "Elixir builds an index. Elixir is useful.",
"doc_2" => "Search uses an index.",
"doc_3" => "Elixir supports search."
}
The inverted index reverses that relationship: its outer map is a dictionary of unique terms, and each inner map is a posting list keyed by document ID. A posting value can be the number of times that term occurs in that document.
%{
"elixir" => %{"doc_1" => 2, "doc_3" => 1},
"index" => %{"doc_1" => 1, "doc_2" => 1},
"search" => %{"doc_2" => 1, "doc_3" => 1}
}
Document IDs alone are sufficient for boolean retrieval. Frequencies support ranking; token positions can support phrase or proximity matching, and character offsets can help locate text for highlighting.
#1 Best Overall
Choose one tokenizer for documents and queries
Tokenization is part of search behavior: it determines where terms begin and end and how they are normalized. This example treats runs of letters and digits as terms, lowercases them, and drops one-character tokens. It does not remove stop words or stem words.
defmodule SimpleSearch do
def tokenize(text) do
text
|> String.downcase()
|> then(&Regex.scan(~r/[p{L}p{N}]+/u, &1))
|> List.flatten()
|> Enum.filter(&(String.length(&1) > 1))
end
end
The Unicode character classes make this less limited than splitting only on ASCII punctuation, but it remains a deliberately simple tokenizer, not a complete language-aware analyzer. Hyphenated words become separate terms; accents are not removed, and stemming and stop-word handling are absent. Change these decisions only deliberately, and use the same rules at index time and query time. A mismatch can make a query term differ from the term stored in the index.
This tokenizer returns strings only. If the search must support phrase matching, retain each token’s position. For highlighting, preserve or derive character offsets in the original text as well.
Count terms and build postings
A frequency map counts repetitions within a document. A map is appropriate here because repeated terms must increase a count; a MapSet is for unique membership, such as a vocabulary or a set of distinct document IDs.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minutedefmodule SimpleSearch do
def tokenize(text) do
text
|> String.downcase()
|> then(&Regex.scan(~r/[p{L}p{N}]+/u, &1))
|> List.flatten()
|> Enum.filter(&(String.length(&1) > 1))
end
def term_counts(text) do
text
|> tokenize()
|> Enum.frequencies()
end
def build_index(documents) do
Enum.reduce(documents, %{}, fn {doc_id, text}, index ->
text
|> term_counts()
|> Enum.reduce(index, fn {term, frequency}, acc ->
update_in(acc[term], fn
nil -> %{doc_id => frequency}
postings -> Map.put(postings, doc_id, frequency)
end)
end)
end)
end
end
index = SimpleSearch.build_index(documents)
Enum traverses enumerables such as lists and maps; it can also work with streams. That does not make a pipeline constant-time: each traversal still processes its input. This map-based version keeps the entire index in memory and is intended for a small, inspectable corpus.
Calculate TF-IDF with an explicit convention
Term frequency (TF) measures how often a term appears in a document. Document frequency (DF) counts how many distinct documents contain the term. It is not the total number of occurrences across the corpus. In the example, elixir has DF 2, because it appears in two documents even though it appears twice in doc_1.
Rank #3
There is no single universal TF-IDF formula. The implementation below uses raw term frequency and smoothed IDF:
tfidf(term, document) = count(term, document) × (ln((N + 1) / (df(term) + 1)) + 1)
Recommended Free Tools
Here, N is the number of indexed documents. Adding one to numerator and denominator avoids a zero denominator and keeps the IDF weight defined for any term in the index. Adding one after the logarithm keeps the weight positive, including for terms found in every document. This convention does not normalize for document length, and the example does not apply cosine normalization.
defmodule SimpleSearch do
# Keep tokenize/1, term_counts/1, and build_index/1 from above.
def idf(index, term, document_count) do
document_frequency =
index
|> Map.get(term, %{})
|> map_size()
:math.log((document_count + 1) / (document_frequency + 1)) + 1
end
def tfidf(index, term, doc_id, document_count) do
frequency =
index
|> Map.get(term, %{})
|> Map.get(doc_id, 0)
frequency * idf(index, term, document_count)
end
end
For elixir, the index has N = 3 and df = 2, so its IDF is ln(4/3) + 1, approximately 1.288. Its TF-IDF for doc_1 is 2 × 1.288, approximately 2.575; for doc_3, it is approximately 1.288. These values follow the formula above and are not length-normalized.
Other systems use different TF transforms, IDF smoothing, or document-length normalization. Elastic documents one example using square-root TF, a smoothed logarithmic DF expression, and inverse-square-root document-length normalization; that is one system’s formula, not a universal requirement. Elastic also describes BM25 as its default similarity and as a variation of TF-IDF, with term frequency, document frequency, and document length among its ranking inputs. BM25 is related to TF-IDF, but the names do not mean the scoring formulas are interchangeable. See Elastic’s similarity documentation.
Score a query and rank matching documents
Tokenize the query with the same function used for documents. This scorer adds the TF-IDF weight for each query term present in a document. Repeated query terms therefore contribute repeatedly; unknown terms match no postings and add nothing. Ties are resolved by ascending document ID for deterministic output.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
defmodule SimpleSearch do
# Keep the earlier functions in this module.
def search(index, query, document_count) do
query
|> tokenize()
|> Enum.reduce(%{}, fn term, scores ->
postings = Map.get(index, term, %{})
Enum.reduce(postings, scores, fn {doc_id, _frequency}, acc ->
score = tfidf(index, term, doc_id, document_count)
Map.update(acc, doc_id, score, &(&1 + score))
end)
end)
|> Enum.sort_by(fn {doc_id, score} -> {-score, doc_id} end)
end
end
SimpleSearch.search(index, "ELIXIR index", map_size(documents))
For this corpus, doc_1 receives the weight for two occurrences of elixir plus one of index; doc_3 receives one elixir weight; and doc_2 receives one index weight. The output is a sorted list of document IDs and scores. Because this is an additive query-term scorer rather than a cosine-similarity implementation, it does not build or normalize query and document vectors.
Check behavior before extending the index
These small checks cover case and punctuation normalization, empty text, repeated terms, unknown terms, and deterministic tie ordering:
assert SimpleSearch.tokenize("Elixir, INDEX!") == ["elixir", "index"]
assert SimpleSearch.tokenize("") == []
assert SimpleSearch.term_counts("Elixir elixir") == %{"elixir" => 2}
assert SimpleSearch.search(index, "no_such_term", map_size(documents)) == []
# Two documents with the same single-term weight sort by document ID.
tie_index = %{"search" => %{"doc_2" => 1, "doc_1" => 1}}
assert SimpleSearch.search(tie_index, "search", 2) ==
Enum.sort_by(SimpleSearch.search(tie_index, "search", 2), fn {id, score} -> {-score, id} end)
For a production search system, decide explicitly whether the tokenizer needs language-specific boundaries, stemming, stop words, positions, or offsets; whether postings need counts or more payload; and whether a different scoring model or length normalization is appropriate. In-memory maps are easy to understand, but persistence, larger collections, advanced analysis, and production ranking may call for a dedicated search engine. Elixir’s documentation lists stable v1.20.4 and supported Erlang/OTP versions 27, 28, and 29; check the official Elixir documentation for the version relevant to your project.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →

