# When the customer's data is a network: Neo4j, Neptune and GraphRAG

> If your SQL needs JOIN after JOIN just to trace a fraud ring, the problem may be the data model, not the query.

Bản gốc: https://fdetimes.net/en/guides/graph-databases-neo4j-neptune-graphrag/

You are in your second week at a fintech company. The risk team hands you three tables: accounts, payment cards and login IP addresses. Their question sounds simple: "Which accounts share a card, email or IP with an account confirmed as fraudulent, including indirectly through two or three layers?"

You open Postgres and write the first query, then a second one twice as long. By the third layer the query is a pile of JOINs, it runs slowly, and nobody in the room is confident it is correct.

This is the moment an FDE needs to recognise early: the customer's data is not a set of tables but a network. Spot it early and you pick the right tool in week one, instead of rewriting the whole pipeline in month two.

## When is data really a network?

AWS defines a graph database as a database that stores data as a network of entities and the relationships between them. In the same documentation, AWS states plainly that purpose-built graph databases deliver the most value for highly connected datasets, while simple tabular data is better suited to a relational database.

The second sentence matters as much as the first. Order lists, payroll, individual transaction logs: leave them in Postgres. The real signal lies in the customer's question, when it follows a chain of "who connects to whom, through what, over how many hops".

The use cases AWS lists all share that shape: social networks, product recommendations, fraud detection through shared emails, cards and IPs, route optimisation, knowledge management. When a customer describes a problem, listen for whether they are describing a path that runs through several entities.

**Điểm mấu chốt:** Choose a graph database because the questions follow relationships, not because the data is big.

## One fraud ring, written two ways

Picture a relational schema: an `accounts` table, plus join tables `account_cards(account_id, card_id)` and `account_ips(account_id, ip)`. Finding accounts that directly share a card with a flagged account, for a single attribute type alone, already takes this:

```sql
SELECT DISTINCT ac2.account_id
FROM accounts f
JOIN account_cards ac1 ON ac1.account_id = f.id
JOIN account_cards ac2 ON ac2.card_id = ac1.card_id
WHERE f.flagged = true AND ac2.account_id <> f.id;
```

To add IPs, you UNION another similar block. To go one layer further (A shares a card with B, B shares an IP with C), the number of JOINs multiplies with every combination of attribute types. Each new layer is another rewrite.

The same question in a graph model: accounts, cards, emails and IPs are all nodes; a `USES` relationship links an account to the things it uses. The Cypher query on Neo4j:

```cypher
MATCH (f:Account {flagged: true})-[:USES*2..6]-(a:Account)
WHERE a <> f
RETURN DISTINCT a.id
```

Each "shared" step is two edges (account to card, card back to another account), so `*2..6` means going at most three layers out. Changing the depth means changing a number. Adding a new attribute type, such as phone numbers, needs no change to the query either, only new nodes and `USES` edges.

That is the advantage to show the risk team with their own eyes. AWS also advertises that performance stays stable as graph data grows; treat that as a vendor claim and measure it on the customer's data before promising anything.

## Neo4j or Neptune: ask about language and infrastructure first

Graph query languages remain fragmented: Neo4j's Cypher, Apache TinkerPop's Gremlin, the W3C's SPARQL for RDF, and GQL as the ISO standard. On the modelling side there are two main schools: the labelled property graph (nodes and edges carry properties, as in the example above) and RDF.

| Aspect | Neo4j | Amazon Neptune |
|---|---|---|
| Data model | Property graph | Both property graph and RDF |
| Language | Cypher; from version 2025.06, new features go only into Cypher 25, and Cypher 5 is frozen | Gremlin, openCypher, SPARQL |
| GraphRAG | Positions itself as the "knowledge layer" for AI, promotes Agentic GraphRAG | Managed GraphRAG via Amazon Bedrock Knowledge Bases |

Architecturally, a Neptune cluster has one primary writer instance and up to 15 read replicas, all sharing a single cluster volume spread across multiple AZs. Alongside it sits Neptune Analytics, a separate in-memory analytics engine that complements the Neptune database when large volumes of graph data need analysing.

On site, the table boils down to a few very practical questions. If the customer already runs everything on AWS and the operations team does not want another system, Neptune reduces friction. If the customer's data is already RDF or an industry ontology, Neptune with SPARQL is the natural choice.

If the customer's own team will write and maintain queries over the long term, Cypher is usually easier to read for people who know SQL. Because Neptune supports openCypher, you can prototype in Cypher and then check each feature before committing to a platform. Record the Cypher version in the handover documentation, since Cypher 5 no longer receives new features.

## Which questions does GraphRAG answer?

Graphs have a second use: as the foundation for AI question-answering systems. Microsoft's GraphRAG research argues that conventional RAG fails on global questions aimed at an entire text corpus. That is the motivation for a graph-based approach.

Imagine a customer with several thousand incident reports. "What caused the incident on the 12th?" is a local question: vector search finding the right few passages is enough.

"Which cause has recurred most often over the past three years?" is different. No passage contains the answer ready-made, and retrieving the top-k most similar passages gives only a skewed view.

A graph helps here because it gathers entities (equipment, suppliers, fault types) and the relationships between them across the whole collection, so the system answers from structure rather than from a few disconnected passages. But if 90% of the customer's questions are local, GraphRAG only adds cost and complexity.

## Step by step on site

Start from the questions, not the tools. Collect 15 to 20 real questions the customer needs answered, and mark which ones run through several relationship hops and which span the whole document collection.

Next, sketch the node-and-edge model on a whiteboard with the customer's domain experts. For the fintech example, four node types and one edge type are enough for a demo. Then load a slice of real data, rewrite the two or three hardest questions in Cypher or Gremlin, and set them beside the SQL versions so the customer can compare for themselves.

Only then choose the platform, based on the customer's existing infrastructure, the language their team will maintain, and whether heavy analysis should move to an in-memory engine such as Neptune Analytics.

## Common mistakes

The most common mistake is using a graph database for data that is inherently tabular, just because "GraphRAG is hot". The second is modelling everything as a node, including attributes that should simply be properties, which bloats the graph and slows queries.

The third is letting variable-length queries with no upper bound, such as `[:USES*]`, run on production data; in a dense network they can traverse almost the entire graph. The fourth is trusting marketing figures, such as the 100k+ queries per second Neptune advertises, instead of benchmarking on the customer's actual workload.

For developers moving into FDE roles, this skill is easy to demonstrate. A small repo with the same question written in SQL and in Cypher, plus notes on why you chose that model, says more than "knows Neo4j" on a CV.

When reading a job description, if the problem falls into graph's familiar use cases, such as fraud detection, recommendations or knowledge management, put that repo at the top of your application.

The next time a SQL query starts joining a table to itself, stop and ask the customer how many relationship hops their real question goes through.

**Thử ngay tuần này:**

- Build a small graph of about 20 accounts, 10 cards and 10 IPs on a local Neo4j, then write a query that finds accounts at most 2 steps from a flagged account.
- Rewrite that exact query in SQL on Postgres and note the number of JOINs, plus how long it took you to write and to read back.
- Take 20 real questions the customer often asks their internal chatbot and classify which are local questions and which span the whole collection.

## Nguồn

- [What is a Graph database? (AWS)](https://aws.amazon.com/nosql/graph/)

- [Graph database - Wikipedia](https://en.wikipedia.org/wiki/Graph_database)

- [What Is Amazon Neptune? (AWS Docs)](https://docs.aws.amazon.com/neptune/latest/userguide/intro.html)

- [Amazon Neptune (AWS)](https://aws.amazon.com/neptune/)

- [Neo4j](https://neo4j.com)

- [Cypher Manual: Introduction (Neo4j Docs)](https://neo4j.com/docs/cypher-manual/current/introduction/)

- [From Local to Global: A Graph RAG Approach to Query-Focused Summarization](https://arxiv.org/abs/2404.16130)
