A final project for the Graph Database course (Semester 6) that integrates Neo4j, Large Language Models (LLM), Graph Analytics, and Retrieval-Augmented Generation (RAG) to build an intelligent medical knowledge graph system.
MediGraph constructs a medical knowledge graph from 6 structured datasets covering 100 diseases, their symptoms, medications, precautions, workout recommendations, and diet recommendations. On top of this graph, four AI-powered tiers are implemented:
| Tier | Component | Description |
|---|---|---|
| 1 | Text-to-Cypher | LLM translates natural language questions into Cypher queries and executes them against Neo4j |
| 1 | Graph Analytics | PageRank, Community Detection (Louvain), Shortest Path, Degree Centrality |
| 2 | ML on Graph | FastRP Node Embeddings + KNN Similarity Search |
| 3 | LLM Graph Builder | Entity and relation extraction from raw text to populate Neo4j |
| 4 | Graph-Augmented RAG | Retrieve context from the graph, augment the LLM prompt, generate a final answer |
| Label | Properties | Count |
|---|---|---|
| Disease | name, description | 101 |
| Symptom | name | 230 |
| Medication | name | 402 |
| Precaution | name | 336 |
| Workout | name | 234 |
| Diet | name | 309 |
| Total | 1,612 |
| Relationship | From | To | Count |
|---|---|---|---|
| HAS_SYMPTOM | Disease | Symptom | 1,139 |
| TREATED_WITH | Disease | Medication | 500 |
| RECOMMENDED_DIET | Disease | Diet | 496 |
| HAS_PRECAUTION | Disease | Precaution | 400 |
| RECOMMENDED_WORKOUT | Disease | Workout | 400 |
| Total | 2,935 |
(Disease) -[:HAS_SYMPTOM]-> (Symptom)
(Disease) -[:TREATED_WITH]-> (Medication)
(Disease) -[:HAS_PRECAUTION]-> (Precaution)
(Disease) -[:RECOMMENDED_WORKOUT]-> (Workout)
(Disease) -[:RECOMMENDED_DIET]-> (Diet)
Six CSV files are required to run this project:
| File | Description |
|---|---|
Diseases_and_Symptoms_dataset.csv |
96,088 rows — binary symptom indicators per disease (230 symptoms, 100 diseases) |
description.csv |
Text description for each of the 100 diseases |
medications.csv |
List of medications per disease |
precautions.csv |
Precautionary measures per disease |
workout.csv |
Workout recommendations per disease |
diets.csv |
Diet recommendations per disease |
The LLM (Llama 3.3 70B via OpenRouter) receives the graph schema and a natural language question, then generates and executes a read-only Cypher query against Neo4j.
Example flow:
User: "Penyakit apa yang memiliki gejala fever dan headache?"
-> LLM generates Cypher
-> Neo4j executes query
-> Returns: common cold, strep throat, acute bronchitis, ...
Four graph algorithms are applied to the medical knowledge graph:
- PageRank — identifies the most central/influential nodes in the medical network
- Community Detection (Louvain) — groups diseases that share similar symptoms, medications, or treatments
- Shortest Path — finds the shortest connection path between any two medical entities
- Degree Centrality — ranks nodes by the number of direct connections
- FastRP (Fast Random Projection) node embeddings are generated for all nodes
- KNN (K-Nearest Neighbors) similarity search finds diseases, symptoms, or medications that are most similar based on their graph neighborhood
Raw medical text is fed to the LLM, which extracts entities and relationships and uses them to populate new nodes and edges into Neo4j automatically.
- User submits a medical question
- Relevant context is retrieved from the Neo4j knowledge graph (symptoms, medications, precautions, diet, workout)
- Retrieved context is injected into the LLM prompt
- LLM generates a comprehensive, graph-grounded answer
- Google Colab (recommended, GPU T4 used in this project)
- Neo4j (installed inside Colab VM via apt)
- OpenRouter API Key (for LLM access via Llama 3.3 70B)
All dependencies are installed inside the notebook:
pip install neo4j openai pandasNeo4j is installed and started directly inside the Colab VM:
sudo apt-get install neo4j -y
sudo service neo4j startThe Neo4j Graph Data Science (GDS) plugin is downloaded and configured for graph algorithm support.
- Open
FP GRAF MediGraph (1).ipynbin Google Colab - In Colab Secrets, add your
OPENROUTER_API_KEYfrom openrouter.ai - Run all cells in order from top to bottom
- When prompted, upload the 6 CSV dataset files
| Cell | Description |
|---|---|
| Cell 1 | Install dependencies |
| Cell 2 | Configure API key and Neo4j credentials |
| Cell 3 | Upload 6 CSV files |
| Cell 4 | Load and preview datasets |
| Cell 5 | Create Neo4j uniqueness constraints |
| Cell 6 | Ingest all data into Neo4j (nodes + relationships) |
| Cell 7 | Verify graph — count nodes and relationships |
| Cell 8 | Text-to-Cypher demo (Tier 1) |
| Cell 9 | PageRank analytics (Tier 1, GDS) |
| Cell 10 | Community Detection / Louvain (Tier 1, GDS) |
| Cell 11+ | Shortest Path, Degree Centrality, FastRP, KNN, LLM Graph Builder, Graph-Augmented RAG |
| Component | Technology |
|---|---|
| Graph Database | Neo4j (installed in Colab VM) |
| Graph Algorithms | Neo4j Graph Data Science (GDS) 2.6.8 |
| LLM | Llama 3.3 70B via OpenRouter API |
| LLM Client | OpenAI Python SDK (pointed to OpenRouter) |
| Data Processing | pandas, numpy |
| Language | Python 3.x |
| Environment | Google Colab (GPU T4) |
- Graph contains 1,612 nodes and 2,935 relationships across 6 entity types
- Text-to-Cypher successfully translates Indonesian and English natural language questions into valid Cypher queries
- PageRank identifies "Hydration" (Diet node) as the most central node in the entire medical network
- Louvain detects disease communities — e.g., musculoskeletal diseases (arthritis, bursitis, carpal tunnel, etc.) clustered together
Abyansyah Dewanto Undergraduate Student — Information Systems GitHub: @abyansyah052
This repository is for educational purposes as part of a Graph Database course final project.

