src.dackar.knowledge_graph.py2neo

Created on March, 2025

@author: wangc, mandd

Attributes

logger

_SAFE_TOKEN_RE

Classes

Py2Neo

Functions

_safe_token(value, kind)

Validate that value is a safe Neo4j identifier (label or relationship type).

Module Contents

src.dackar.knowledge_graph.py2neo.logger[source]
src.dackar.knowledge_graph.py2neo._SAFE_TOKEN_RE[source]
src.dackar.knowledge_graph.py2neo._safe_token(value, kind)[source]

Validate that value is a safe Neo4j identifier (label or relationship type).

Used by the schema-governed batch-ingestion helpers (upsert_nodes_batch / upsert_edges_batch) to guard against Cypher injection through interpolated labels / relationship types.

Parameters:
  • value (str) – identifier string to validate.

  • kind (str) – human-readable descriptor used in the error message (e.g. “label”).

Returns:

the original value unchanged if it passes validation.

Return type:

str

Raises:

ValueError – if value is not a string or does not match [A-Za-z_][A-Za-z0-9_]*.

class src.dackar.knowledge_graph.py2neo.Py2Neo(uri, user, pwd)[source]
__uri[source]
__user[source]
__pwd[source]
__driver = None[source]
close()[source]

Close the python neo4j connection

restart()[source]

Restart the python neo4j connection

create_node(label, properties)[source]

Create a new graph node

Parameters:
  • label (str) – node label will be used by neo4j

  • properties (dict) – node attributes

static _create_node(tx, label, properties)[source]

Create a new graph node

Parameters:
  • tx (obj) – python neo4j active session that can be used to execute queries

  • label (str) – node label will be used by neo4j

  • properties (dict) – node attributes

create_relation(l1, p1, l2, p2, lr, pr=None)[source]

create graph relation

Parameters:
  • l1 (str) – first node label

  • p1 (dict) – first node attributes

  • l2 (str) – second node label

  • p2 (dict) – second node attributes

  • lr (str) – relationship label

  • pr (dict, optional) – attributes for relationship. Defaults to None.

static _create_relation(tx, l1, p1, l2, p2, lr, pr)[source]

create graph relation

Parameters:
  • tx (obj) – python neo4j active session that can be used to execute queries

  • l1 (str) – first node label

  • p1 (dict) – first node attributes

  • l2 (str) – second node label

  • p2 (dict) – second node attributes

  • lr (str) – relationship label

  • pr (dict, optional) – attributes for relationship. Defaults to None.

find_nodes(label, properties=None)[source]

Find the node in neo4j graph database

Parameters:
  • label (str) – node label

  • properties (dict, optional) – node attributes. Defaults to None.

Returns:

list of nodes

Return type:

list

static _find_nodes(tx, label, properties)[source]

Find the node in neo4j graph database

Parameters:
  • tx (obj) – python neo4j active session that can be used to execute queries

  • label (str) – node label

  • properties (dict, optional) – node attributes. Defaults to None.

Returns:

list of nodes

Return type:

list

load_csv_for_nodes(file_path, label, attribute)[source]

Load CSV file to create nodes

Parameters:
  • file_path (str) – file path for CSV file, location is relative to ‘dbms.directories.import’ or ‘server.directories.import’ in neo4j.conf file

  • label (str) – node label

  • attribute (dict) – node attribute from the CSV column names

static _load_csv_nodes(tx, file_path, label, attribute)[source]
load_csv_for_relations(file_path, l1, p1, l2, p2, lr, pr=None)[source]

Load CSV file to create node relations

Parameters:
  • file_path (str) – file path for CSV file, location is relative to ‘dbms.directories.import’ or ‘server.directories.import’ in neo4j.conf file

  • l1 (str) – first node label

  • p1 (dict) – first node attribute from the CSV column names

  • l2 (str) – second node label

  • p2 (dict) – second node attribute from the CSV column names

  • lr (str) – relationship label

  • pr (dict, optional) – of attributes for relation. Defaults to None.

static _load_csv_for_relations(tx, file_path, l1, p1, l2, p2, lr, pr)[source]
query(query, parameters=None, db=None)[source]

User provided Cypher query statements for python neo4j driver to use to query database

Parameters:
  • query (str) – user provided Cypher query statements

  • parameters (dict, optional) – dictionary that provide key/value pairs for query statement to use. Defaults to None.

  • db (str, optional) – name for database. Defaults to None.

Returns:

returned list of queried results.

Return type:

list

write(query, parameters=None, db=None)[source]

Execute a write Cypher query inside a managed (auto-committing) transaction.

Parameters:
  • query (str) – Cypher write query string.

  • parameters (dict, optional) – parameter map bound into the query. Defaults to None.

  • db (str, optional) – target database name; uses the driver default when None. Defaults to None.

reset(db=None)[source]

Reset the database, delete all records, use it with care

Parameters:

db (str, optional) – target database name; uses the driver default when None. Defaults to None.

static _reset(tx)[source]
get_all(db=None)[source]

Get all records from database

Parameters:

db (str, optional) – target database name; uses the driver default when None. Defaults to None.

Returns:

list of all records

Return type:

list

static _get_all(tx)[source]
upsert_nodes_batch(nodes, db=None)[source]

Batch-upsert a collection of nodes, grouped by label.

Nodes are merged on their id attribute so that repeated calls are idempotent. Each dict in nodes must have "label" and "attrs" keys; attrs must contain "id". Used by the schema-governed KG batch-ingestion workflows.

Parameters:
  • nodes (Sequence[dict]) – sequence of {"label": str, "attrs": dict} dicts.

  • db (str, optional) – target database name; uses the driver default when None. Defaults to None.

upsert_edges_batch(edges, db=None)[source]

Batch-upsert a collection of relationships, grouped by endpoint labels and type.

Each dict in edges must contain "from_label", "to_label", "type", "from" (source node id), "to" (target node id), and an optional "attrs" dict for relationship properties. Endpoints are matched on their id attribute. Used by the schema-governed KG batch-ingestion workflows.

Parameters:
  • edges (Sequence[dict]) – sequence of edge descriptor dicts.

  • db (str, optional) – target database name; uses the driver default when None. Defaults to None.

load_dataframe_for_nodes(df, labels, properties)[source]

Load pandas dataframe to create nodes

Parameters:
  • df (pandas.DataFrame) – DataFrame for loading

  • labels (str) – node label

  • properties (list) – node properties from the dataframe column names

load_dataframe_for_relations(df, l1='sourceLabel', p1='sourceNodeId', l2='targetLabel', p2='targetNodeId', lr='relationshipType', pr=None)[source]

Load dataframe to create node relations

Parameters:
  • df (pandas.DataFrame) – DataFrame for relationships

  • l1 (str) – first node label

  • p1 (str) – first node ID

  • l2 (str) – second node label

  • P2 (str) – second node ID

  • lr (str) – relationship label

  • pr (list, optional) – of attributes for relation. Defaults to None.