Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Vincent Granville built XLLM to find trustworthy, relevant sources for advanced research in statistics, machine learning, and computer science. Despite the word “LLM” in its name, XLLM is not a neural language model trained from scratch: it is a domain-focused search and retrieval system built from selected web content, taxonomies, token dictionaries, association tables, and rules.

Why build a specialized system instead of using a general chatbot?

In an article published by DataScienceCentral on January 13, 2024, Granville described a practical problem: his questions were specialized, and the tools he tried did not reliably surface the references and links he wanted. He said OpenAI did not return links for his queries, while Google, Bing, and individual site-search boxes produced inconsistent results.

His response was to automate discovery across carefully selected sources and organize results around expert subject areas. The goal was not to beat general chatbots at everyday conversation. Granville put the distinction plainly: “It does not replace OpenAI / GPT for the general public: that was not the goal.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What “from scratch” means in this project

For XLLM, “from scratch” means engineering a custom retrieval application rather than training a large transformer from random initialization. Granville said the system uses no neural networks and has no actual training. Its output is driven by crawled material, a curated taxonomy, dictionaries, association tables, and query-processing rules.

#1 Best Overall

He frames the broader aim in terms of retrieval, augmentation, and generation (RAG), but the documented core is retrieval-oriented. This distinction matters if you are asking whether you can make an LLM without an API: you can build a useful, independent search-and-retrieval system without calling a general-purpose model API, but that is not the same as creating a generative foundation model.

How XLLM’s data pipeline works

1. Select sources and organize them by subject

The project begins with selected repositories that provide useful content and taxonomies. Wolfram was the initial source. Granville described subsets of Wikipedia and his own books as possible additions, not as confirmed parts of the initial crawl. Content is grouped by category so that a user can choose the areas relevant to a query instead of searching an undifferentiated web-sized corpus.

Granville reported that the Wolfram crawl contained about 15,000 webpages and roughly 1 GB before compression. He also characterized that collection as about 1% of human knowledge; that is his framing, not an independently established measure of coverage. For the math domain, he described a taxonomy with about 5,000 categories.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Extract information and build dictionaries

The system gathers categories, tokens, links, tags, metadata, related items, and other navigation information. It builds a dictionary of successive tokens found in sentences, titles, or category entries, preserving relationships among terms and the content in which they appear.

3. Calculate associations and store summaries

XLLM calculates associations between tokens and multi-token expressions, including pointwise mutual information (PMI). It stores related content and category counts in nested hash tables. These tables provide the structured index used later to connect a query with relevant material.

4. Match query phrases and retrieve related content

When a user submits a query, the system looks for matching n-gram subsets in a sorted dictionary and retrieves their associated information. The text-processing rules must account for accents, stop words, autocorrection, stemming, singularization, capitalization, punctuation, and multi-token names. Granville uses “Saint-Petersburg” to illustrate how generic token handling can damage a name’s meaning if its parts are processed carelessly.

Two versions serve users and developers

Version Purpose What it processes
XLLM-short End-user version Loads the final summary tables.
XLLM Developer version Processes the full crawled data and generates the tables.

Granville said both versions should return the same results when XLLM-short is using current tables. The split separates the heavier data-processing work from the version intended for everyday queries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does building your own system require a GPU?

Not necessarily. XLLM avoids conventional neural-network training, so it does not follow the hardware pattern associated with training a large model. Its requirements instead depend on the scale of the crawl, the indexing and table-building work, and how the application is deployed; Granville’s account does not specify a hardware configuration that readers can treat as a general requirement.

By contrast, training a large neural model from scratch generally requires GPU infrastructure. An I-TEK guide published in a 2023 context repeats a rule of thumb of 20 training tokens per parameter and gives roughly $25,000 as an illustrative estimate for training a 7-billion-parameter model. Those are secondary-source estimates, not universal costs: actual hardware, methods, prices, and location affect the result.

When can a small, curated system be better?

A specialized index can be a better fit when the task depends on domain relevance, source quality, useful links, and controllable retrieval rather than broad conversational ability. XLLM’s design keeps the collection and taxonomy focused, allowing expert users to select relevant categories. Granville described its strengths as speed, efficiency, scalability, flexibility, replicability, and a simple architecture.

That approach also sets limits. Results depend on which sources were selected and how thoroughly they cover the question; a curated crawl cannot stand in for all web knowledge. It is designed around Granville’s expert-research needs and similar users, not around the broad needs of a general public chatbot.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to judge it against search engines or chatbots

Granville presented a “random walks” search example and invited comparisons with OpenAI, Bing, Google, Bard, and Wolfram’s own search box. There is no single universal metric for deciding which tool is best: the answer depends on the user and task. A practical comparison should examine:

  • Source trustworthiness: Are the returned materials appropriate and credible for the topic?
  • Links and citations: Can the user follow results back to useful source material?
  • Domain specificity: Does the system understand the specialist vocabulary and categories involved?
  • Latency and coverage: How quickly does it respond, and what sources or topics are missing?
  • Ranking control: Can a researcher adjust parameters or select categories to influence results?
  • Audience fit: Does it serve an expert researcher’s task, a lay user’s question, or both?

Plans described in 2024 are not proof of current availability

Granville discussed possible advertiser keywords, a paid version with larger live tables and more parameter tuning, and a future Web API through GenAItechLab.com. These were presented as ideas or plans in the January 2024 article; they do not establish that any of those offerings are currently available.

A learning resource for conventional LLM construction

For readers who want to study neural language-model implementation rather than a table-driven retrieval system, Sebastian Raschka’s Build a Large Language Model (From Scratch) is a relevant step-by-step book with code. It is an optional learning companion, not a source Granville identifies as part of XLLM’s development.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.