TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Build a Dependency Graph With Neo4j and Python to Trace Vulnerable Packages
Learning Hub

How to Build a Dependency Graph With Neo4j and Python to Trace Vulnerable Packages

A hands-on introduction to graph databases: install Neo4j locally without Docker or a cloud account, model a package dependency graph in Python, and write the multi-hop Cypher query that finds every...

August 21, 2026 16 Min Read
41

Most application data is really about relationships. A package depends on another package. A maintainer owns a package. A service calls another service. You can store all of that in tables, and for a long time that works fine. Then someone asks a question like this one: which maintainers are exposed if this one small, deeply buried dependency turns out to be vulnerable? In a relational database that question means walking a dependency chain of unknown depth: a self-join that might need to run two, three, or five times depending on how deep the chain goes, with no way to know the depth in advance.

Table Of Content

  • What You Will Build
  • Prerequisites
  • What a Graph Database Actually Stores
  • Why Traversal Beats Joins as the Question Gets Deeper
  • Step 1: Install a JDK for Neo4j
  • Step 2: Download and Start Neo4j Community Server
  • Step 3: Connect From Python
  • Step 4: Model the Graph Before You Load Any Data
  • Mistake to Avoid: Treating “Name” as the Whole Identity
  • Mistake to Avoid: CREATE When You Mean MERGE
  • Step 5: Create Constraints Before You Load Data
  • Step 6: Load the Graph With Parameters, MERGE, and UNWIND
  • Step 7: Write Your First Cypher Queries
  • Step 8: The Multi-Hop Query That Justifies the Whole Thing
  • Step 9: Confirm the Index Is Actually Being Used
  • How to Verify Everything Works End to End
  • Shutting Down and Cleaning Up
  • Next Steps

A graph database is built for exactly that question. Instead of computing relationships at query time by matching foreign keys across tables, it stores the relationships themselves as data, so following a connection is a direct lookup rather than a search. This tutorial builds one from scratch: you will install Neo4j Community Server locally with no Docker and no cloud account, model a small package dependency graph in Python, and write the query that answers the “who’s exposed” question no matter how many hops deep the vulnerable package is buried.

Every command in this tutorial was run for real against a local Neo4j 5.26.14 instance while writing it. The node counts, query results, and query plans you will see below are the actual captured output, not invented numbers. You do not need any prior graph database or Cypher experience. If you can write a Python dictionary and read a SQL join, you already have what you need.

What You Will Build

  • A local Neo4j Community Server, installed and started without Docker, without a cloud account, and without administrator rights
  • A small Python script that loads a demo package dependency graph: packages, the maintainers who own them, and the DEPENDS_ON edges between them
  • Uniqueness constraints that make package lookups fast and stop you from accidentally creating duplicate nodes
  • A multi-hop Cypher query that finds every maintainer exposed to a vulnerable dependency, however many packages deep it is buried
  • Two deliberately broken versions of the loading logic, so you can see exactly what goes wrong when you use CREATE instead of MERGE, and when you forget that a package’s identity needs more than just its name

Prerequisites

  • Windows, macOS, or Linux. Commands below are shown for Windows first, with the macOS/Linux equivalent noted where it differs.
  • Python 3.10 or newer and pip (this tutorial used Python 3.13.14)
  • Comfort with basic Python (dictionaries, functions) and running commands in a terminal. No prior graph database, Cypher, or Java experience is assumed.
  • About 500 MB of free disk space for a JDK and Neo4j Community Server
  • No Docker, no cloud account, and no credit card. This tutorial deliberately avoids all three so you can follow along on a bare machine.

What a Graph Database Actually Stores

A few terms are worth defining before they show up in code. A node is one thing in your data: a package, a maintainer, an order. It is the rough equivalent of a row. A relationship is a stored, directed connection between exactly two nodes, such as one package depending on another. Unlike a foreign key, a relationship in a graph database is a real record on disk, not a value that has to be matched against another table at query time. A property is a key and value stored on a node or relationship, like version: "3.2.0". A label tags a node with a type, such as Package or Maintainer, the rough equivalent of a table name. Cypher is Neo4j’s query language: instead of describing joins, you draw the shape of the pattern you are looking for, like (a)-[:DEPENDS_ON]->(b).

Why Traversal Beats Joins as the Question Gets Deeper

This is the architectural idea usually called index-free adjacency, and it is the reason a graph database exists as a separate category of tool rather than just being a table with extra syntax. When a node is written to disk, its record already stores the disk address of its first relationship. Following that relationship to the next node, and the next, is a pointer lookup: roughly constant time, regardless of how large the overall database gets. A relational join, by contrast, has to look up matching rows through an index each time, which is typically logarithmic in the size of the table, and a chain of unknown depth means you do not know in advance how many joins to write. That difference does not matter much for a one-hop question like “who maintains this package.” It matters enormously for a five-hop question like “which maintainers are exposed to this vulnerable dependency, no matter how deep it is buried,” which is exactly the query this tutorial builds toward.

This tutorial is inspired by freeCodeCamp’s Neo4j and Python handbook, which covers general knowledge graph modeling in more depth. The dataset, the security angle, and every script here are original and were built and tested independently for this post.

Step 1: Install a JDK for Neo4j

Neo4j Community Server runs on the Java virtual machine, so it needs a JDK before it can start. If you already have Java 17 or Java 21 installed, run java -version to confirm and skip ahead to Step 2. If not, here is a way to get one without administrator rights and without touching your system’s default Java installation, using the install-jdk Python package as a downloader:

pip install install-jdk
python -c "import jdk; print(jdk.install('17'))"

That downloads an OpenJDK 17 build (Eclipse Temurin) into your own user profile and prints where it landed. On the machine used for this tutorial, that looked like this:

C:\Users\Administrator\.jdk\jdk-17.0.20+8

Yours will show your own username instead. If you would rather install Java the ordinary way, download an installer directly from Eclipse Temurin’s release page and pick 17 or 21; either works identically for everything below.

Either way, point your terminal at that JDK and confirm it runs. On Windows (PowerShell):

$env:JAVA_HOME = "C:\Users\<you>\.jdk\jdk-17.0.20+8"
$env:PATH = "$env:JAVA_HOME\bin;$env:PATH"
java -version

On macOS/Linux (bash/zsh):

export JAVA_HOME="$HOME/.jdk/jdk-17.0.20+8"
export PATH="$JAVA_HOME/bin:$PATH"
java -version

You should see something like this (the exact build number will differ slightly by platform and date):

openjdk version "17.0.20" 2026-07-21
OpenJDK Runtime Environment Temurin-17.0.20+8 (build 17.0.20+8)
OpenJDK 64-Bit Server VM Temurin-17.0.20+8 (build 17.0.20+8, mixed mode, sharing)

Neo4j’s current long-term support release, 5.26, supports both Java 17 and Java 21, so either Java version works for everything in this tutorial.

Step 2: Download and Start Neo4j Community Server

Download the Community Server archive for your platform. This tutorial uses 5.26.14, the current LTS release.

On Windows:

curl -L -o neo4j.zip https://dist.neo4j.org/neo4j-community-5.26.14-windows.zip
tar -xf neo4j.zip
cd neo4j-community-5.26.14

On macOS/Linux:

curl -L -o neo4j.tar.gz https://dist.neo4j.org/neo4j-community-5.26.14-unix.tar.gz
tar -xf neo4j.tar.gz
cd neo4j-community-5.26.14

(curl and tar both ship built in on current Windows, macOS, and Linux, so no extra tools are needed.)

Before starting Neo4j for the first time, set its initial password. This only works before the first start; changing it later requires a different command.

Windows:

.\bin\neo4j-admin.bat dbms set-initial-password "choose-a-real-password"

macOS/Linux:

bin/neo4j-admin dbms set-initial-password "choose-a-real-password"

You should see:

Changed password for user 'neo4j'. IMPORTANT: this change will only take effect if performed before the database is started for the first time.

Now start the server. Windows:

.\bin\neo4j.bat console

macOS/Linux:

bin/neo4j console

console mode runs Neo4j in the foreground and prints its logs directly to your terminal. That is deliberate for this tutorial: it needs no administrator rights to register a Windows service or a systemd unit, and you can stop it at any time with Ctrl+C. Leave this terminal window open and running; open a second terminal for the Python work in the next steps. Here is the real startup log from this tutorial’s test run:

Starting Neo4j.
INFO  Logging config in use: File 'neo4j-community-5.26.14\conf\user-logs.xml'
INFO  Starting...
INFO  This instance is ServerId{c0ac7a34} (c0ac7a34-3903-4c88-b652-bedbba59c0c9)
INFO  ======== Neo4j 5.26.14 ========
INFO  Anonymous Usage Data is being sent to Neo4j, see https://neo4j.com/docs/usage-data/
INFO  Bolt enabled on localhost:7687.
INFO  HTTP enabled on localhost:7474.
INFO  Remote interface available at http://localhost:7474/
INFO  Started.

Two ports matter here: 7687 speaks Bolt, the binary protocol Neo4j’s drivers use, and 7474 serves a web interface and a plain HTTP API. Verify the server is actually listening by opening http://localhost:7474/ in a browser, or by curling it from your second terminal:

curl http://localhost:7474/
{"bolt_routing":"neo4j://localhost:7687","query":"http://localhost:7474/db/{databaseName}/query/v2","transaction":"http://localhost:7474/db/{databaseName}/tx","bolt_direct":"bolt://localhost:7687","neo4j_version":"5.26.14","neo4j_edition":"community"}

Step 3: Connect From Python

Install the official driver in your second terminal:

pip install neo4j

This tutorial used driver version 6.2.0, which requires Python 3.10 or newer. Create connect_test.py:

from neo4j import GraphDatabase

URI = "bolt://localhost:7687"
AUTH = ("neo4j", "choose-a-real-password")

driver = GraphDatabase.driver(URI, auth=AUTH)
driver.verify_connectivity()
print("Connected.")

records, summary, keys = driver.execute_query("RETURN 1 AS x")
print("Query result:", records[0]["x"])

driver.close()

Run it:

python connect_test.py
Connected.
Query result: 1

Two things worth calling out for later. First, the driver object is expensive to create and cheap to reuse: create one when your program starts and keep it, rather than opening a new one per request. Second, notice the query used a plain string with no user data spliced in. Every query below that includes data uses a $parameter placeholder instead of Python string formatting. That is not just a security habit; Neo4j caches a query’s execution plan keyed by its text, so parameterized queries reuse a cached plan while string-formatted ones force a fresh plan to be built every time.

Step 4: Model the Graph Before You Load Any Data

This tutorial’s graph has two node types and two relationship types:

  • Package nodes, with name, ecosystem, version, and vulnerable properties
  • Maintainer nodes, with a name property
  • (:Maintainer)-[:MAINTAINS]->(:Package) relationships
  • (:Package)-[:DEPENDS_ON]->(:Package) relationships

Before writing any loading code, it is worth walking through two mistakes that are easy to make here, because both produce a graph that looks fine until you query it.

Mistake to Avoid: Treating “Name” as the Whole Identity

A package name is not globally unique. An npm package called acme-crypto-utils and a pip package with the exact same name are two completely different pieces of software that happen to share a string. If you model a package’s identity as just its name, your loading code will silently treat those two packages as the same node.

Here is what that looks like when it actually happens. This loads an npm package and a pip package, both named acme-crypto-utils, using a MERGE keyed only on name:

MERGE (p:BuggyPackage {name: $name})
SET p.ecosystem = $ecosystem, p.vulnerable_dep = $vulnerable_dep

After loading both packages, one npm (flagged as depending on something vulnerable) and one pip (not), here is the real result:

BuggyPackage nodes after loading BOTH the npm and pip package: [{'name': 'acme-crypto-utils', 'ecosystem': 'pip', 'vulnerable_dep': False}]

Only one node exists. The pip package’s write landed on the exact same node as the npm package and silently overwrote its fields, because as far as MERGE was concerned, they were the same thing. There was no error. Nothing crashed. The data is just quietly wrong, and it would stay wrong until someone noticed a security report that should have applied to the npm package instead pointed at the pip one.

The fix is to key on everything that actually determines identity. This tutorial uses a composite key of (name, ecosystem):

MERGE (p:Package {name: $name, ecosystem: $ecosystem})
SET p.version = $version, p.vulnerable = $vulnerable

With the full key, the same two packages correctly land as two separate nodes. You will set this up for real as a database constraint in Step 5, so the mistake becomes structurally impossible rather than something you just have to remember.

Mistake to Avoid: CREATE When You Mean MERGE

Cypher has two ways to add a node: CREATE always makes a new one, no matter what already exists. MERGE checks whether a matching node already exists first, and only creates one if it does not. If you use CREATE for data you expect to load more than once, such as re-running an import script after fixing a bug, you get duplicates. Here is that happening for real: the same CREATE statement run twice with identical data.

run 1: nodes_created = 1
run 2: nodes_created = 1
total acme-widget nodes after running the same CREATE twice: 2

Two nodes, same name, same properties, no error. This is why every loading query in this tutorial uses MERGE: it makes re-running your import script safe, which matters the moment you need to fix a typo in your dataset or add one more package without starting over.

Step 5: Create Constraints Before You Load Data

A uniqueness constraint does two things at once: it enforces that the property or properties you choose can never be duplicated, and it automatically creates a backing index so lookups on that key are fast. Creating constraints before loading data, not after, means bad data gets rejected on the way in rather than needing to be found and cleaned up later. Create create_constraints.py:

from neo4j import GraphDatabase

URI = "bolt://localhost:7687"
AUTH = ("neo4j", "choose-a-real-password")
driver = GraphDatabase.driver(URI, auth=AUTH)

with driver.session(database="neo4j") as session:
    session.run(
        "CREATE CONSTRAINT package_key IF NOT EXISTS "
        "FOR (p:Package) REQUIRE (p.name, p.ecosystem) IS UNIQUE"
    ).consume()

    session.run(
        "CREATE CONSTRAINT maintainer_name IF NOT EXISTS "
        "FOR (m:Maintainer) REQUIRE m.name IS UNIQUE"
    ).consume()

    for row in session.run("SHOW CONSTRAINTS").data():
        print(row["name"], "->", row["labelsOrTypes"], row["properties"])

driver.close()
python create_constraints.py
maintainer_name -> ['Maintainer'] ['name']
package_key -> ['Package'] ['name', 'ecosystem']

The package_key constraint is exactly the fix from the previous section, now enforced by the database itself instead of relying on every script that touches this data to remember it.

Step 6: Load the Graph With Parameters, MERGE, and UNWIND

This tutorial’s demo dataset is a small, entirely fictional package ecosystem: acme-web-app depends on acme-auth-lib and acme-ui-kit, both of which depend on acme-crypto-utils, which depends on acme-legacy-hash, a package flagged as vulnerable. A separate pip package also happens to be named acme-crypto-utils, to exercise the composite key from Step 5. None of these are real packages; the vulnerability label DEMO-2026-001 is a made-up placeholder, not a real advisory ID.

Loading many rows one at a time means one network round trip per row. UNWIND turns a list of rows into one round trip for the whole batch: it takes a parameter list and expands it into a stream of rows inside a single Cypher query. Combined with MERGE, that gives you an idempotent bulk load. Create load_graph.py:

from neo4j import GraphDatabase

URI = "bolt://localhost:7687"
AUTH = ("neo4j", "choose-a-real-password")

packages = [
    {"name": "acme-web-app", "ecosystem": "npm", "version": "3.2.0", "vulnerable": False, "advisory": None},
    {"name": "acme-auth-lib", "ecosystem": "npm", "version": "1.9.4", "vulnerable": False, "advisory": None},
    {"name": "acme-ui-kit", "ecosystem": "npm", "version": "2.4.1", "vulnerable": False, "advisory": None},
    {"name": "acme-crypto-utils", "ecosystem": "npm", "version": "0.8.2", "vulnerable": False, "advisory": None},
    {"name": "acme-legacy-hash", "ecosystem": "npm", "version": "0.1.3", "vulnerable": True, "advisory": "DEMO-2026-001"},
    {"name": "acme-data-pipeline", "ecosystem": "pip", "version": "4.0.0", "vulnerable": False, "advisory": None},
    {"name": "acme-crypto-utils", "ecosystem": "pip", "version": "2.1.0", "vulnerable": False, "advisory": None},
]

# Each row says "the package on the left depends on the package on the right."
depends_on = [
    {"from_name": "acme-web-app", "from_eco": "npm", "to_name": "acme-auth-lib", "to_eco": "npm"},
    {"from_name": "acme-web-app", "from_eco": "npm", "to_name": "acme-ui-kit", "to_eco": "npm"},
    {"from_name": "acme-auth-lib", "from_eco": "npm", "to_name": "acme-crypto-utils", "to_eco": "npm"},
    {"from_name": "acme-ui-kit", "from_eco": "npm", "to_name": "acme-crypto-utils", "to_eco": "npm"},
    {"from_name": "acme-crypto-utils", "from_eco": "npm", "to_name": "acme-legacy-hash", "to_eco": "npm"},
    {"from_name": "acme-data-pipeline", "from_eco": "pip", "to_name": "acme-crypto-utils", "to_eco": "pip"},
]

maintains = [
    {"maintainer": "Priya", "package": "acme-web-app", "eco": "npm"},
    {"maintainer": "Priya", "package": "acme-auth-lib", "eco": "npm"},
    {"maintainer": "Jordan", "package": "acme-auth-lib", "eco": "npm"},
    {"maintainer": "Jordan", "package": "acme-ui-kit", "eco": "npm"},
    {"maintainer": "Sam", "package": "acme-crypto-utils", "eco": "npm"},
    {"maintainer": "Sam", "package": "acme-legacy-hash", "eco": "npm"},
    {"maintainer": "Riley", "package": "acme-data-pipeline", "eco": "pip"},
    {"maintainer": "Alex", "package": "acme-crypto-utils", "eco": "pip"},
]

driver = GraphDatabase.driver(URI, auth=AUTH)

with driver.session(database="neo4j") as session:
    summary = session.run(
        """
        UNWIND $rows AS row
        MERGE (p:Package {name: row.name, ecosystem: row.ecosystem})
        SET p.version = row.version,
            p.vulnerable = row.vulnerable,
            p.advisory = row.advisory
        """,
        rows=packages,
    ).consume()
    print("Package nodes created:", summary.counters.nodes_created)

    summary = session.run(
        """
        UNWIND $rows AS row
        MATCH (a:Package {name: row.from_name, ecosystem: row.from_eco})
        MATCH (b:Package {name: row.to_name, ecosystem: row.to_eco})
        MERGE (a)-[:DEPENDS_ON]->(b)
        """,
        rows=depends_on,
    ).consume()
    print("DEPENDS_ON relationships created:", summary.counters.relationships_created)

    summary = session.run(
        """
        UNWIND $rows AS row
        MERGE (m:Maintainer {name: row.maintainer})
        WITH m, row
        MATCH (p:Package {name: row.package, ecosystem: row.eco})
        MERGE (m)-[:MAINTAINS]->(p)
        """,
        rows=maintains,
    ).consume()
    print("Maintainer nodes created:", summary.counters.nodes_created)
    print("MAINTAINS relationships created:", summary.counters.relationships_created)

    counts = session.run(
        "MATCH (n) RETURN labels(n)[0] AS label, count(*) AS total ORDER BY label"
    ).data()
    print("Node counts:", counts)

driver.close()
python load_graph.py
Package nodes created: 7
DEPENDS_ON relationships created: 6
Maintainer nodes created: 5
MAINTAINS relationships created: 8
Node counts: [{'label': 'Maintainer', 'total': 5}, {'label': 'Package', 'total': 7}]

Run the exact same script again without changing anything, to confirm MERGE actually does what Step 4 claimed:

Package nodes created: 0
DEPENDS_ON relationships created: 0
Maintainer nodes created: 0
MAINTAINS relationships created: 0
Node counts: [{'label': 'Maintainer', 'total': 5}, {'label': 'Package', 'total': 7}]

Every creation counter drops to zero on the second run. Nothing was duplicated, because every node and relationship already existed and MERGE matched them instead of creating new ones.

Step 7: Write Your First Cypher Queries

With data loaded, start with simple, single-hop questions. Every query in this and the next step follows the same shape: open a session, call session.run(cypher).data() to get results back as a list of plain Python dictionaries, and print them. Create queries.py with this pattern:

from neo4j import GraphDatabase

URI = "bolt://localhost:7687"
AUTH = ("neo4j", "choose-a-real-password")
driver = GraphDatabase.driver(URI, auth=AUTH)

with driver.session(database="neo4j") as session:
    rows = session.run(
        """
        MATCH (m:Maintainer)-[:MAINTAINS]->(p:Package {name: 'acme-crypto-utils', ecosystem: 'npm'})
        RETURN m.name AS maintainer
        ORDER BY m.name
        """
    ).data()
    print(rows)

driver.close()

Run it, asking who maintains the npm acme-crypto-utils package:

python queries.py
[{'maintainer': 'Sam'}]

From here on, each query below is shown on its own for readability. Drop each one into the same session.run(cypher).data() shape above, inside the same with driver.session(...) as session: block, to run it. What does acme-web-app depend on directly?

MATCH (:Package {name: 'acme-web-app', ecosystem: 'npm'})-[:DEPENDS_ON]->(dep:Package)
RETURN dep.name AS dependency
ORDER BY dependency
[{'dependency': 'acme-auth-lib'}, {'dependency': 'acme-ui-kit'}]

Both of these are one-hop patterns. A relational query for either would look almost identical to the Cypher above once written as a join. The difference shows up in the next step.

Step 8: The Multi-Hop Query That Justifies the Whole Thing

Cypher supports variable-length paths: [:DEPENDS_ON*1..5] means “follow one to five DEPENDS_ON relationships in a row,” without you having to know or write out how many hops that actually takes for any given package. First, the full transitive dependency chain of acme-web-app:

MATCH (:Package {name: 'acme-web-app', ecosystem: 'npm'})-[:DEPENDS_ON*1..5]->(dep:Package)
RETURN DISTINCT dep.name AS dependency
ORDER BY dependency
[{'dependency': 'acme-auth-lib'}, {'dependency': 'acme-crypto-utils'}, {'dependency': 'acme-legacy-hash'}, {'dependency': 'acme-ui-kit'}]

That correctly reaches three hops deep (acme-web-app to acme-auth-lib to acme-crypto-utils to acme-legacy-hash) in a single query, with no recursive CTE and no application-side loop walking the graph one level at a time. Now the actual question this tutorial set out to answer: which maintainers own a package that depends, directly or transitively, on something flagged as vulnerable?

MATCH (m:Maintainer)-[:MAINTAINS]->(p:Package)-[:DEPENDS_ON*1..5]->(vuln:Package {vulnerable: true})
RETURN DISTINCT m.name AS maintainer, p.name AS owns, vuln.advisory AS advisory
ORDER BY maintainer, owns
[{'maintainer': 'Jordan', 'owns': 'acme-auth-lib', 'advisory': 'DEMO-2026-001'}, {'maintainer': 'Jordan', 'owns': 'acme-ui-kit', 'advisory': 'DEMO-2026-001'}, {'maintainer': 'Priya', 'owns': 'acme-auth-lib', 'advisory': 'DEMO-2026-001'}, {'maintainer': 'Priya', 'owns': 'acme-web-app', 'advisory': 'DEMO-2026-001'}, {'maintainer': 'Sam', 'owns': 'acme-crypto-utils', 'advisory': 'DEMO-2026-001'}]

Five results, and they make sense once you trace them: Priya is exposed twice, through acme-web-app and through acme-auth-lib. Jordan is exposed twice, through acme-auth-lib and acme-ui-kit. Sam is exposed once, directly, since Sam maintains both acme-crypto-utils and the vulnerable acme-legacy-hash itself. As a control, confirm the pip packages are correctly untouched, since acme-crypto-utils in pip is a different node from the npm one flagged as depending on something vulnerable:

MATCH (p:Package {ecosystem: 'pip'})
OPTIONAL MATCH (p)-[:DEPENDS_ON*1..5]->(vuln:Package {vulnerable: true})
RETURN p.name AS pip_package, vuln.name AS vulnerable_dependency
ORDER BY pip_package
[{'pip_package': 'acme-crypto-utils', 'vulnerable_dependency': None}, {'pip_package': 'acme-data-pipeline', 'vulnerable_dependency': None}]

Both pip packages come back with no vulnerable dependency, exactly as expected: they are genuinely separate nodes from their npm namesake, not the same node with mixed-up data, because Step 5’s composite constraint made that mistake impossible.

Step 9: Confirm the Index Is Actually Being Used

A constraint only speeds up a lookup if your query actually provides the full key it was built on. EXPLAIN shows you the query plan Neo4j chose without running the query, so you can check this directly instead of assuming it. Create explain_plan.py:

from neo4j import GraphDatabase

URI = "bolt://localhost:7687"
AUTH = ("neo4j", "choose-a-real-password")
driver = GraphDatabase.driver(URI, auth=AUTH)


def print_plan(plan, depth=0):
    print("  " * depth + plan["operatorType"])
    for child in plan.get("children", []):
        print_plan(child, depth + 1)


with driver.session(database="neo4j") as session:
    print("--- matching on the full (name, ecosystem) constraint key ---")
    plan = session.run(
        "EXPLAIN MATCH (p:Package {name: 'acme-web-app', ecosystem: 'npm'}) RETURN p"
    ).consume().plan
    print_plan(plan)

    print()
    print("--- matching on name only, not the full constraint key ---")
    plan = session.run(
        "EXPLAIN MATCH (p:Package {name: 'acme-web-app'}) RETURN p"
    ).consume().plan
    print_plan(plan)

driver.close()
python explain_plan.py
--- matching on the full (name, ecosystem) constraint key ---
ProduceResults@neo4j
  NodeUniqueIndexSeek@neo4j

--- matching on name only, not the full constraint key ---
ProduceResults@neo4j
  Filter@neo4j
    NodeByLabelScan@neo4j

The first query goes straight to NodeUniqueIndexSeek: it uses the constraint’s backing index to jump directly to the matching node. The second query, missing the ecosystem half of the key, falls back to NodeByLabelScan plus Filter: it reads every single Package node in the database and checks each one’s name in memory. On seven packages that difference is invisible. On a real dataset with a few hundred thousand packages, it is the difference between a lookup that returns instantly and one that scans the whole label every time it runs. The lesson: a constraint only helps the queries that actually use its full key.

How to Verify Everything Works End to End

Before moving on, confirm each of these, in order:

  • http://localhost:7474/ responds with a JSON banner showing "neo4j_edition":"community"
  • connect_test.py prints Connected. and Query result: 1
  • SHOW CONSTRAINTS lists both package_key and maintainer_name
  • load_graph.py reports 7 Package nodes and 5 Maintainer nodes on the first run, and all-zero creation counters on a second run
  • The multi-hop exposure query returns exactly Jordan, Priya (twice each), and Sam, and never Riley or Alex
  • EXPLAIN on a full-key lookup shows NodeUniqueIndexSeek, not NodeByLabelScan

If the exposure query in Step 8 comes back empty, the most likely cause is that load_graph.py was only run partially, for example if it was interrupted between creating the Package nodes and creating the DEPENDS_ON relationships. Re-running the whole script is safe (that’s the point of MERGE) and will fill in whatever is missing.

Shutting Down and Cleaning Up

Switch to the terminal running neo4j console and press Ctrl+C. You should see Neo4j log a clean shutdown before the process exits. Because this was started in console mode rather than as a service, there is nothing left running in the background afterward, and nothing to uninstall: deleting the neo4j-community-5.26.14 folder removes the server and every bit of data you loaded into it.

Next Steps

This tutorial deliberately used the smallest setup that still teaches the real ideas: a local server, a tiny fictional dataset, and no cloud dependency. From here:

  • Neo4j Aura’s free tier removes the local install entirely once you are ready to keep a graph running persistently, at the cost of needing a hosted account
  • The APOC library adds hundreds of procedures for things this tutorial’s plain Cypher does not cover, like loading directly from CSV or JSON files
  • Neo4j’s Graph Data Science library adds graph algorithms (PageRank, community detection, shortest path) that go beyond pattern matching, useful if “which maintainer is exposed” grows into “which maintainer sits at the center of the most dependency chains”
  • If this pattern is useful for real software, not a demo dataset, look at Software Bill of Materials tools like Syft alongside the OSV vulnerability database instead of hand-typing package data; the graph modeling ideas here transfer directly
  • If you are running a service backed by more than one dependency source, our guide on scaling enterprise knowledge graphs without building a monolith covers what changes once a single local instance like this one isn’t enough
  • For the other side of this problem, catching malicious packages before they ever reach a dependency graph like this one, see our tutorials on npm and pip dependency cooldowns and detecting malicious npm preinstall scripts

Tags:

cypherGraph Databasesneo4jPythonSupply Chain Security

Share

A self-serve cafeteria salad bar with individual bins, tongs, and stacked trays, illustrating selective access instead of an all-or-nothing choice
Previous Post

Cloudflare’s Task-Based OAuth Consent Turns All-or-Nothing Agent Permissions Into a Choice

Six antique padlocks of different shapes and sizes with six mismatched keys laid out on a white background
Next Post

A Critical WordPress SSO Login Bypass Stayed Hidden in Six Untracked Plugin Editions

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026