What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
José Henrique Oliveira de Carvalho’s portfolio assistant uses local embedding generation and PostgreSQL with pgvector to retrieve information from Markdown files, then sends the selected context to an LLM through Groq. The pipeline is local only in part: response generation uses a hosted service. Its value is as a concrete example of how source preparation, chunking, retrieval, and relevance filtering work together—not as a benchmark or a universal recipe.
How the portfolio assistant’s RAG pipeline works
The assistant answers questions about Carvalho’s background, experience, projects, and technical decisions using information he maintains in versioned Markdown files. The flow is:
- Maintain source material: Store profile, experience, and project information in Markdown with structured frontmatter.
- Parse and split documents: Use LangChain’s
RecursiveCharacterTextSplitterin Markdown mode to create smaller passages. - Enrich the text: Add probable visitor questions to the content before embedding it.
- Generate embeddings locally: Run
Xenova/multilingual-e5-smallthrough@huggingface/transformerson CPU. - Store and retrieve: Save source content and vectors in PostgreSQL with pgvector, then search for passages close to a visitor’s question.
- Filter and answer: Pass sufficiently relevant retrieved context to the LLM; when no result qualifies, avoid adding arbitrary context.
The project’s broader stack includes Bun, Elysia, TypeScript, Drizzle ORM, and openai/gpt-oss-120b. Carvalho reports that generation happens through Groq, so the system’s embeddings are local but its complete answer-generation path is not.
Recommended Free Tools
How the documents are chunked and enriched
Carvalho reports setting the Markdown splitter to chunkSize: 800 and chunkOverlap: 50. These are settings from his implementation, not proven optimal values for other content or models. Splitting lets retrieval return focused passages rather than entire documents; overlap preserves some continuity where a fact crosses a chunk boundary.
#1 Best Overall
He also appends likely user questions to source text before embedding it. The rationale is that a visitor may phrase a query differently from the way a biography or project description is written. Including question-like wording can give matching a closer textual signal without changing the LLM. It is an implementation choice, not evidence that this enrichment will improve every corpus.
How local embeddings are generated
The project uses Xenova/multilingual-e5-small with Transformers.js on CPU. Carvalho reports mean pooling, normalization, and 384-dimensional output. He prefixes stored passages with passage: and incoming questions with query:, details tied to this model and implementation.
Rank #2
- TypeScript implements a superset of syntax for strictly typed development, facilitating deep static analysis and enhanced development environment integration. The compiler translates source into standard script formats, ensuring parity across any runtime.
- TypeScript is ideal for front-end developers, full-stack engineers, and software architects who build large-scale web applications. It serves those looking to improve code excellence, reduce bugs through static checking, and maintain complex projects more.
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
Those prefixes and processing steps should not be copied blindly to another embedding model. The example demonstrates a local embedding stage; it does not establish a general embedding configuration or a measured quality advantage.
How PostgreSQL and pgvector retrieve passages
The project stores both the original text and its embedding in PostgreSQL. Its query uses pgvector’s <=> cosine-distance operator, orders results by ascending distance, and asks for five results. The pgvector documentation confirms support for cosine distance.
Carvalho then applies a project-specific cosine-distance cutoff of < 0.35. A distance score is a retrieval signal, not a universal definition of relevance: the cutoff depends on the model, data, and query behavior. Passing this filter also cannot guarantee that a generative model will never produce an unsupported answer.
What happens when retrieval finds nothing useful?
When no retrieved result passes the cutoff, Carvalho says he does not inject arbitrary context and instead gives the LLM a basic instruction not to invent information. This is a sensible fallback for a source-grounded assistant: irrelevant context can be worse than no context. Still, an instruction is not a guarantee against hallucination, so the answer should be treated as potentially uncertain when the source material does not support it.
Do you need a dedicated vector database?
For this personal portfolio, Carvalho says PostgreSQL with pgvector was sufficient. If an application already relies on PostgreSQL, storing ordinary records and vectors together can avoid introducing a separate database for a straightforward workload. That does not mean PostgreSQL is always the right choice; workload size, query complexity, operational needs, and latency targets matter.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →pgvector performs exact nearest-neighbor search by default. Its documentation also describes HNSW and IVFFlat indexes for approximate search. Approximate indexing can trade recall for speed, so adopting it is a workload decision rather than an automatic upgrade. The project article does not report using either index or provide a scale threshold or benchmark comparing database options.
Best Value
What this implementation illustrates—and what it does not
The main lesson Carvalho draws is that “the LLM is not the whole system.” Source representation, chunking, embeddings, retrieval, and relevance filtering all shape what context reaches generation. He describes the design as suitable for his portfolio, not as a universal architecture, and notes that a dedicated vector database may make sense for larger or more complex workloads.
The implementation details and author’s rationale are described in Carvalho’s project article. They should be read as a report of one system, not independent quality testing: no hardware tests, answer-quality evaluation, cost comparison, or scale benchmark is established there.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

