Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A reliable way to populate a knowledge graph from unstructured text with an LLM is a fixed pipeline, not a single prompt. Parse each document, split it into chunks, fix a target schema of entity and relationship types, have the model extract typed entities and relations (triples such as Acme Corp, ACQUIRED, Widget Labs), write them to a graph with a link back to the supporting passage, and then validate and prune the result before anyone queries it. Neo4j and Microsoft both document concrete versions of this pattern. The choices that most affect quality are the schema, the provenance links, and how the output is checked.

This guide is for developers and data engineers who need to turn a corpus of documents into a queryable graph. A question common in public discussions is how to build a knowledge graph from thousands of unstructured documents. At that scale the repeatable stages matter more than any one prompt.

Why extraction is a pipeline, not a single prompt

A single prompt that reads a document and returns a graph breaks down quickly. Documents exceed the model’s context window, facts are spread across paragraphs, model output has to be checked against a schema, and the same real-world entity appears under many names. Neo4j’s knowledge graph builder treats these concerns as separate stages: loading and parsing, chunking, schema definition, entity and relation extraction, graph writing, and cleanup. Some of those components are optional, as described in the first step below.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build the pipeline step by step

1. Load documents and split them into units

Extract plain text from each source file and give every document and every chunk a stable identifier, so that any fact can later be traced to a specific passage. Split the text into units that fit comfortably inside the model’s context window. Chunk size is a trade-off. Small units keep provenance tight but can separate a subject from the sentence that states its relation. Large units give the model more context but make the output harder to check. Test both on your own corpus.

Neo4j’s builder can also write a lexical layer of nodes for documents and chunks alongside the entity layer, and it can compute embeddings for chunks as an optional component. Neo4j’s guide reports its best results on long-form English text. The builder is less suited to tabular data such as spreadsheets, and to images, diagrams and slides. If your corpus is mostly in those formats, plan a separate path for them rather than expecting the same extraction quality.

2. Fix the target schema before extraction

Decide which entity types and relation types the application needs, and for each relation type state which entity types may appear at each end. An explicit schema is the strongest constraint you can give the model. Neo4j’s pipeline can also generate a schema automatically, but an automatic schema reflects the text rather than your domain. Review it before use, and do not assume it will cover the types your queries require.

An illustrative schema for a company-news corpus might look like this. It is an example of the shape, not a recommended standard.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Entity types: Organization, Person, Product, Location
  • Relation types: ACQUIRED (Organization to Organization), FOUNDED (Person to Organization), LAUNCHED (Organization to Product), HEADQUARTERED_IN (Organization to Location)

The schema does two jobs. It steers the extraction, and it gives the pruning step an objective standard to check against.

3. Extract typed entities and relations

For each unit, ask the model for three things: entities, each with a name, a type from your schema, and a short description or attributes; relations, each with a source entity, a relation type and a target entity; and, for each relation, the supporting text. Microsoft’s standard GraphRAG method works along these lines. It prompts an LLM to extract named entities and descriptions from each text unit and to describe relationships between entity pairs, and it produces summaries of those descriptions that later steps can aggregate.

Where your model provider supports it, request structured output that validates against a schema. Neo4j’s guide recommends this for type safety and reliability. Two cautions apply. Supported integrations and API behavior change over time, so confirm that your provider and library version support the feature before relying on it. The guide also labels the knowledge graph builder feature experimental, so pin library versions and retest your extraction after every upgrade.

An illustrative record for one sentence, ‘In March, Acme Corp agreed to acquire Widget Labs’, looks like this:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Entity: Acme Corp (id: acme-corp, type: Organization)
Entity: Widget Labs (id: widget-labs, type: Organization)
Relation: acme-corp, ACQUIRED, widget-labs
Evidence: 'In March, Acme Corp agreed to acquire Widget Labs'
Chunk: news-0412-chunk-07

The sentence reports an agreement, not a completed acquisition. If your schema has an ACQUIRED type, the extraction prompt has to say whether pending deals count, or the schema has to carry a status property. Otherwise the graph will record plans as completed facts.

4. Keep provenance on every node and relation

Store the document and chunk identifiers that support each entity and relation, and keep a short supporting quote with each relation. Microsoft’s output documentation records text-unit references for entities and for relationship identifiers found in text units, which gives reviewers a path back to the source. When a fact appears in several chunks, keep the full list of supporting units rather than only the first. That list shows how widely a claim is supported.

Provenance is also your cheapest error-correction tool. A reviewer who sees a suspicious edge can open the passage and decide whether the edge follows from it. A fix to the prompt or the chunking can then be traced to every edge it affected.

5. Resolve mentions and aggregate evidence

Mentions are not identities. ‘Acme’, ‘Acme Corp’ and ‘Acme Corporation’ may be one company or three, and two people may share a name. Merge mentions only on keys you trust, such as an internal customer ID, a company registry number or a product SKU. Send ambiguous cases to a review queue rather than letting the model decide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Microsoft’s method summarizes entity and relation descriptions across all the occurrences it finds. That is aggregation: it combines what each passage says about an entity. It is not a general solution to identity resolution, and it should not be treated as one.

6. Validate, prune and write

Before writing, check every record against the schema. The entity type must be on the list, the relation type must be allowed, and each endpoint must have the type the schema requires. Reject or repair records that fail, and prune disallowed types. Neo4j’s builder can apply a configured schema and cleanup operations, which covers this stage. Write only validated output, and use deterministic identifiers so that rerunning the pipeline updates existing nodes instead of duplicating them.

Measure quality on a manually reviewed sample drawn from your own corpus. Include the relation types you care about most, and score precision and recall per relation type rather than one overall number. The official sources describe schema checks, pruning and output structures. They do not establish a standard evaluation benchmark or a single validation recipe that is best for every corpus, so the measurement has to be yours.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choosing an extraction approach

Microsoft’s documentation describes two extraction styles that are useful to compare. The standard method uses LLM prompts for entities, relations and summaries. Its FastGraphRAG description is a cheaper, co-occurrence-oriented alternative that produces a noisier graph, which the documentation says is less directly relevant outside GraphRAG. Neo4j’s builder is a different kind of choice: a configurable pipeline whose main decisions are the schema and whether to use structured output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Approach Documented tradeoff What to measure before committing
Standard LLM extraction and summarization (Microsoft GraphRAG) Prompts an LLM for entities, relations and summaries, and aggregates descriptions across text units. Cost figures: not stated in Microsoft’s documentation. Relation precision, schema adherence, cost per document, and coherence across chunks.
FastGraphRAG, co-occurrence-oriented construction (Microsoft) Described as cheaper, with a noisier graph that is less directly relevant outside GraphRAG. Cost figures: not stated in Microsoft’s documentation. Graph noise on your sample, whether your retrieval tasks still work, and actual compute cost.
Schema-driven structured output (Neo4j knowledge graph builder) Schema validation and structured output for supported provider integrations; maturity caveats are covered in step 3. Provider support, malformed-output rate, schema fit, and stability across version upgrades.

Choose the standard LLM route when the graph will be queried directly by applications or people and relation precision matters. Choose the cheaper co-occurrence route only when cost dominates and the graph’s job is the GraphRAG retrieval it was designed for, since the documentation frames its output as noisy outside that setting. In either case, the comparison that counts is the one you run on your own corpus.

Failure modes to check

  • Duplicate entities: the same organization or person appears under several nodes. Sort nodes by name similarity, review the closest pairs, and tighten your merge keys.
  • Unsupported edges: a relation is present but the cited passage does not state it. Sample edges and open their evidence. A high rate points to the extraction prompt or the chunk size.
  • Type drift: the model returns synonyms such as ‘bought’ instead of ACQUIRED. Enumerate relation types and reject anything outside the list.
  • Status errors: pending, planned or rumored events are recorded as completed facts, as in the ‘agreed to acquire’ example. Add a status property, or state in the prompt whether pending events count.
  • Direction errors: a directional relation is reversed, for example with acquirer and target swapped. Check endpoint order for every directional relation type on the reviewed sample.
  • Malformed or truncated output: a response fails to parse or stops mid-record. Log the failing unit, retry with validation, and confirm that the unit plus the prompt fits the context window.
  • Duplicates on rerun: rerunning the pipeline creates a second copy of the graph. Use deterministic identifiers for documents, chunks and entities.

Where to build hands-on skills

Neo4j GraphAcademy lists a course on constructing knowledge graphs with Neo4j GraphRAG for Python. The listing covers schema definition, chunking strategies, extraction prompts and pipeline parameters. Course availability and any terms can change, so check the listing directly before planning around it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.