Skip to content
leventaricanPublic

About

Minimal, readable GraphRAG in Node.js: text to knowledge graph to answers, every step visible in the terminal

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

GraphRAG

A minimal, readable GraphRAG in Node.js — about 400 lines, one dependency, no database.

Feed it plain text, it builds a knowledge graph. Ask it a question, it answers from that graph — and shows you every step in the terminal.

$ node query.js "Which policies did Bob Turner sell?"

┌──────────────────────────────────────────────┐
│ ENTITY MATCHING                              │
│ ▸ bob turner                                 │
└──────────────────────────────────────────────┘
             │
             ▼
┌──────────────────────────────────────────────┐
│ 1-HOP TRAVERSAL                              │
│ 8 relevant triples                           │
└──────────────────────────────────────────────┘

  Context triples:
  + Bob Turner ──[SOLD]──▶ Policy P-100
  + Bob Turner ──[SOLD]──▶ Policy P-200
  + Alice Chen ──[HOLDS]──▶ Policy P-100
  ...

What it does

  • Builds a knowledge graph from text. An LLM turns sentences into (subject, relation, object) triples.
  • Answers questions from the graph. It finds the entities in your question, collects their neighbours, and lets the LLM answer only from those facts.
  • Shows its work. Every pipeline step is printed as a box: matched entities, retrieved triples, model, answer. When an answer is wrong, you can see where it went wrong.
  • Keeps the graph in a plain file. graph.jsonl has one triple per line — readable, greppable, editable by hand.
  • Grows incrementally. Load more texts any time; duplicate triples are skipped.
  • Runs on Claude or on a local model. Use the Anthropic API, or any OpenAI-compatible server such as llama.cpp — no API key needed.

What it deliberately does not do: embeddings, vector search, a database, a web UI. It is meant for learning how GraphRAG works, not for large data.

Quick start

Requires Node.js ≥ 20.6.

npm install

Then pick a model backend:

Option A — Claude (best results)

cp .env.example .env        # put your ANTHROPIC_API_KEY in .env

npm run extract -- example.txt
npm run query -- "Which policies did Bob Turner sell?"

The API is billed separately from a Claude Pro subscription — add credits at console.anthropic.com.

Option B — local model (free, offline)

# terminal 1: start any OpenAI-compatible server, e.g. llama.cpp
llama-server -m gemma-3-1b-it-Q4_K_M.gguf --port 8080

# terminal 2
export LLM_BASE_URL=http://localhost:8080/v1
export LLM_MODEL=gemma-3-1b          # optional, only shown in the output

node extract.js example.txt
node query.js "Which policies did Bob Turner sell?"

Small models (~1B) are fine for watching the pipeline, but they extract fewer and sloppier triples — expect wrong answers on multi-step questions.

The example: a fictional insurer

example.txt describes Harbor Mutual, a made-up insurance company: customers, policies, claims, a broker, a claims handler, an appraiser, an underwriter and a reinsurer.

It is fictional on purpose: the model cannot know any of it from training, so a correct answer proves the graph did the work.

Questions to try, from easy to hard:

Question What it tests
Which policies did Bob Turner sell? direct lookup
Which claims does Carol Davis handle? all edges of one entity
Who handles the claim of Alice Chen? multi-hop: customer → policy → claim → handler
Who is the broker of the customer whose claim Dave Wilson appraised? long chain — likely fails with MAX_HOPS = 1, try 2

Usage

Load text

npm run extract -- example.txt                    # from a file
echo "Alice Chen is a customer of Harbor Mutual." \
  | npm run extract -- -                          # from stdin
npm run extract -- example.txt my-graph.jsonl     # into another graph file

Ask questions

npm run query -- "Who handles the claim of Alice Chen?"
npm run query -- "Who handles the claim of Alice Chen?" my-graph.jsonl

Start over

rm graph.jsonl

How it works

extract.js:  text ──▶ LLM ──▶ triples ──▶ append to graph.jsonl (skip duplicates)

query.js:    question ──▶ entity matching ──▶ graph traversal ──▶ context ──▶ LLM ──▶ answer
  1. Extract — the LLM reads the text and returns triples such as:
    {"subject":"Alice Chen","relation":"HOLDS","object":"Policy P-100"}
    {"subject":"Claim C-1042","relation":"FILED_UNDER","object":"Policy P-100"}
    {"subject":"Carol Davis","relation":"HANDLES","object":"Claim C-1042"}
  2. Match — words in the question are compared with entity names; the best 5 matches are the starting points.
  3. Traverse — all triples touching those entities are collected, then their neighbours (MAX_HOPS), up to MAX_TRIPLES.
  4. Answer — the triples go to the LLM as context, with the instruction to answer only from these facts.

Tuning knobs are at the top of query.js:

Constant Default Effect
MAX_HOPS 1 how far to walk from the matched entities
MAX_TRIPLES 30 maximum facts sent to the LLM

Limitations

  • Entity matching is word overlap. "Alice" matches "Alice Chen", but synonyms or typos do not.
  • Graph quality depends on the model. If extraction misses a fact, no question can find it.
  • Everything is loaded into memory on each run — fine for thousands of triples, not for millions.

Files

File Purpose
extract.js text → triples → graph.jsonl
query.js question → graph lookup → answer
llm.js model call: Anthropic API or local OpenAI-compatible server
ui.js terminal boxes and arrows
example.txt fictional insurance example
graph.jsonl the knowledge graph (created on first extract, git-ignored)
graphrag.md background: what GraphRAG is and where it gets hard

License

MIT

About

Minimal, readable GraphRAG in Node.js: text to knowledge graph to answers, every step visible in the terminal

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages