iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
You can build semantic search without an LLM generating answers: turn documents and queries into embeddings, then retrieve the closest matching passages from a vector database. The example behind this title uses a FastAPI backend, Hugging Face embeddings, and Qdrant. It is a retrieval prototype—not proof of zero operating cost, guaranteed relevance, or production durability.
What “semantic search without an LLM” means
Semantic search matches a query to text with similar meaning, rather than relying only on exact keyword overlap. In this design, an embedding model converts both stored documents and the search query into vectors; Qdrant compares the query vector with stored vectors and returns matching source text. The response is retrieval, not a newly generated answer.
“Without an LLM” here means without an answer-generation model in the search response. The pipeline still relies on a machine-learning embedding model, hosted externally in the tutorial’s example. Returning source passages avoids a generated-answer step, but does not guarantee that results are relevant, complete, or correct.
Free tools Windows power users keep installed
One-click scans. No signup required.
How the example pipeline works
Umer Abdullah’s September 23, 2026 DEV Community tutorial describes a separate frontend and Python API. The frontend is hosted on GitHub Pages; its FastAPI backend accepts JSON document uploads at /upload, obtains embeddings from Hugging Face, and stores vectors with payloads in a Qdrant client initialized with location=":memory:". Its /search endpoint embeds the query, can filter by category, retrieves up to three vector matches, and returns scores with source text. The example creates a 384-dimensional cosine collection and includes a small JSON sample dataset. See the tutorial.
#1 Best Overall
The tutorial lists FastAPI, Uvicorn, python-multipart, qdrant-client, and requests among its requirements. For a related overview of Qdrant’s collection, data-loading, and search workflow, consult the official Qdrant quickstart. That quickstart does not validate the tutorial’s deployment choices or persistence behavior.
How to build a similar prototype
- Prepare the backend. Create a Python API with FastAPI and run it with Uvicorn. Include the tutorial’s listed dependencies:
fastapi,uvicorn,python-multipart,qdrant-client, andrequests. - Choose an embedding service. The example requests embeddings from Hugging Face. Configure the service and credentials using its current documentation; the tutorial’s architecture depends on an external inference provider.
- Create the vector collection. The example uses a 384-dimensional vector configuration with cosine similarity. The embedding model’s output dimension and the collection configuration must match.
- Load documents. Send JSON documents to
/upload, generate an embedding for each document, and store each vector with its associated payload, such as category and source text. - Search. Send a query to
/search, embed it with the same model, optionally supply a category filter, and return up to three matches with their scores and source text. - Connect the frontend. Host the frontend separately (the tutorial uses GitHub Pages) and configure it to call the backend’s deployed API endpoint.
These steps describe the tutorial’s design, not a tested deployment recipe. Review its request validation, error handling, embedding endpoint details, and CORS configuration for your own application; the example’s broad CORS setting should not be assumed appropriate for production.
What the “$0” and speed claims do—and do not—show
The tutorial presents the architecture as costing $0 per month and reports vector matching in 2–10 milliseconds. It does not provide a reproducible benchmark setup, corpus size, hardware, traffic pattern, or independent measurements. Treat both figures as the author’s claims about the example, not as a typical result or a guarantee. End-to-end response time also involves more than the vector comparison, including embedding inference and network requests.
The cost claim is conditional on service allowances and usage. Hugging Face’s current Inference Providers pricing documentation says free users receive $0.10 in monthly credits, subject to change; usage beyond the credits is pay-as-you-go. As Hugging Face puts it, “Past the free-tier credits, you get charged for every inference request based on the compute time x price of the underlying hardware.” A small experiment may fit within available credits, but the figure does not establish that production use will remain free.
Rank #3
Storage, uptime, and production readiness
In-memory vectors are a prototype choice
The tutorial initializes Qdrant in memory. Before relying on this design, decide where durable vectors and payloads will live, how data will be restored after a restart or redeployment, and how document updates will be reflected in the index. The tutorial and Qdrant quickstart do not establish that this in-memory deployment provides persistence across lifecycle events.
Plan for hosting behavior
The tutorial warns that a free backend may sleep after inactivity and reports cold starts of 30–60 seconds. Those are the author’s reported figures, not independently verified timings or a statement of current Render behavior. Check your chosen host’s current plan details and policies before relying on any uptime, cold-start, or periodic-ping approach.
Rank #4
- Used Book in Good Condition
Review the application before deployment
- Validate uploaded JSON and set appropriate request-size limits.
- Check authentication and access controls for upload and search endpoints.
- Restrict CORS to the frontend origins that should be allowed to call the API.
- Define useful error handling for embedding-provider failures and unavailable storage.
- Decide how credentials, stored content, backups, and index refreshes will be managed.
These are implementation checks prompted by the example’s design, not findings from a security audit or executed test. The tutorial should not be treated as production-ready without that review.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsChoices to make before adapting the design
| Decision | Example in the tutorial | What to assess |
|---|---|---|
| Embedding generation | Hosted Hugging Face inference | Hosted inference brings provider availability, network latency, credentials, and usage charges into the design. Local embedding generation is an alternative to assess, but the tutorial supplies no benchmark comparing the two. |
| Vector storage | Qdrant client in memory | In-memory storage is used by the example; choose and verify a persistent storage and recovery approach for durable data. |
| Backend hosting | Free hosting is discussed; the article reports possible sleeping and cold starts | Check the current host’s policy and whether its uptime and response behavior fit your use case. A paid or always-on option changes the operating-cost assumptions. |
| Search response | Up to three retrieved passages with scores and source text | Retrieval-only results preserve source material without generating an answer. A generated answer would be a separate capability and is outside this example’s no-answer-generation design. |
The tutorial demonstrates one arrangement, not a fair comparison across these choices. Select based on persistence, latency, traffic, operating budget, and the consequences of an unavailable or incorrect result.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

