src.dackar.knowledge_graph.py2neo¶
Created on March, 2025
@author: wangc, mandd
Attributes¶
Classes¶
Functions¶
|
Validate that value is a safe Neo4j identifier (label or relationship type). |
Module Contents¶
- src.dackar.knowledge_graph.py2neo._safe_token(value, kind)[source]¶
Validate that value is a safe Neo4j identifier (label or relationship type).
Used by the schema-governed batch-ingestion helpers (
upsert_nodes_batch/upsert_edges_batch) to guard against Cypher injection through interpolated labels / relationship types.- Parameters:
value (str) – identifier string to validate.
kind (str) – human-readable descriptor used in the error message (e.g. “label”).
- Returns:
the original value unchanged if it passes validation.
- Return type:
str
- Raises:
ValueError – if value is not a string or does not match
[A-Za-z_][A-Za-z0-9_]*.
- class src.dackar.knowledge_graph.py2neo.Py2Neo(uri, user, pwd)[source]¶
-
- create_node(label, properties)[source]¶
Create a new graph node
- Parameters:
label (str) – node label will be used by neo4j
properties (dict) – node attributes
- static _create_node(tx, label, properties)[source]¶
Create a new graph node
- Parameters:
tx (obj) – python neo4j active session that can be used to execute queries
label (str) – node label will be used by neo4j
properties (dict) – node attributes
- create_relation(l1, p1, l2, p2, lr, pr=None)[source]¶
create graph relation
- Parameters:
l1 (str) – first node label
p1 (dict) – first node attributes
l2 (str) – second node label
p2 (dict) – second node attributes
lr (str) – relationship label
pr (dict, optional) – attributes for relationship. Defaults to None.
- static _create_relation(tx, l1, p1, l2, p2, lr, pr)[source]¶
create graph relation
- Parameters:
tx (obj) – python neo4j active session that can be used to execute queries
l1 (str) – first node label
p1 (dict) – first node attributes
l2 (str) – second node label
p2 (dict) – second node attributes
lr (str) – relationship label
pr (dict, optional) – attributes for relationship. Defaults to None.
- find_nodes(label, properties=None)[source]¶
Find the node in neo4j graph database
- Parameters:
label (str) – node label
properties (dict, optional) – node attributes. Defaults to None.
- Returns:
list of nodes
- Return type:
list
- static _find_nodes(tx, label, properties)[source]¶
Find the node in neo4j graph database
- Parameters:
tx (obj) – python neo4j active session that can be used to execute queries
label (str) – node label
properties (dict, optional) – node attributes. Defaults to None.
- Returns:
list of nodes
- Return type:
list
- load_csv_for_nodes(file_path, label, attribute)[source]¶
Load CSV file to create nodes
- Parameters:
file_path (str) – file path for CSV file, location is relative to ‘dbms.directories.import’ or ‘server.directories.import’ in neo4j.conf file
label (str) – node label
attribute (dict) – node attribute from the CSV column names
- load_csv_for_relations(file_path, l1, p1, l2, p2, lr, pr=None)[source]¶
Load CSV file to create node relations
- Parameters:
file_path (str) – file path for CSV file, location is relative to ‘dbms.directories.import’ or ‘server.directories.import’ in neo4j.conf file
l1 (str) – first node label
p1 (dict) – first node attribute from the CSV column names
l2 (str) – second node label
p2 (dict) – second node attribute from the CSV column names
lr (str) – relationship label
pr (dict, optional) – of attributes for relation. Defaults to None.
- query(query, parameters=None, db=None)[source]¶
User provided Cypher query statements for python neo4j driver to use to query database
- Parameters:
query (str) – user provided Cypher query statements
parameters (dict, optional) – dictionary that provide key/value pairs for query statement to use. Defaults to None.
db (str, optional) – name for database. Defaults to None.
- Returns:
returned list of queried results.
- Return type:
list
- write(query, parameters=None, db=None)[source]¶
Execute a write Cypher query inside a managed (auto-committing) transaction.
- Parameters:
query (str) – Cypher write query string.
parameters (dict, optional) – parameter map bound into the query. Defaults to None.
db (str, optional) – target database name; uses the driver default when None. Defaults to None.
- reset(db=None)[source]¶
Reset the database, delete all records, use it with care
- Parameters:
db (str, optional) – target database name; uses the driver default when None. Defaults to None.
- get_all(db=None)[source]¶
Get all records from database
- Parameters:
db (str, optional) – target database name; uses the driver default when None. Defaults to None.
- Returns:
list of all records
- Return type:
list
- upsert_nodes_batch(nodes, db=None)[source]¶
Batch-upsert a collection of nodes, grouped by label.
Nodes are merged on their
idattribute so that repeated calls are idempotent. Each dict in nodes must have"label"and"attrs"keys;attrsmust contain"id". Used by the schema-governed KG batch-ingestion workflows.- Parameters:
nodes (Sequence[dict]) – sequence of
{"label": str, "attrs": dict}dicts.db (str, optional) – target database name; uses the driver default when None. Defaults to None.
- upsert_edges_batch(edges, db=None)[source]¶
Batch-upsert a collection of relationships, grouped by endpoint labels and type.
Each dict in edges must contain
"from_label","to_label","type","from"(source node id),"to"(target node id), and an optional"attrs"dict for relationship properties. Endpoints are matched on theiridattribute. Used by the schema-governed KG batch-ingestion workflows.- Parameters:
edges (Sequence[dict]) – sequence of edge descriptor dicts.
db (str, optional) – target database name; uses the driver default when None. Defaults to None.
- load_dataframe_for_nodes(df, labels, properties)[source]¶
Load pandas dataframe to create nodes
- Parameters:
df (pandas.DataFrame) – DataFrame for loading
labels (str) – node label
properties (list) – node properties from the dataframe column names
- load_dataframe_for_relations(df, l1='sourceLabel', p1='sourceNodeId', l2='targetLabel', p2='targetNodeId', lr='relationshipType', pr=None)[source]¶
Load dataframe to create node relations
- Parameters:
df (pandas.DataFrame) – DataFrame for relationships
l1 (str) – first node label
p1 (str) – first node ID
l2 (str) – second node label
P2 (str) – second node ID
lr (str) – relationship label
pr (list, optional) – of attributes for relation. Defaults to None.