src.dackar.RCA.equipment_similarity.equipment_spec_store

equipment_spec_store — EquipmentSpecStore.

Thin wrapper around ChromaRecordStore for the equipment_specs doc_type. One document per plant component; embedding_text is the natural-language spec string produced by EquipmentSpecBuilder.

The full Stage 1-6 enrichment pipeline (NLP, NER, summarization) is not used here — equipment specs are structured, not unstructured prose, so a minimal record with just the spec text and identity metadata is sufficient.

Attributes

logger

JsonDict

EQUIPMENT_SPECS_DOC_TYPE

Classes

EquipmentSpecStore

Manages the equipment_specs Chroma collection.

Module Contents

src.dackar.RCA.equipment_similarity.equipment_spec_store.logger[source]
src.dackar.RCA.equipment_similarity.equipment_spec_store.JsonDict[source]
src.dackar.RCA.equipment_similarity.equipment_spec_store.EQUIPMENT_SPECS_DOC_TYPE = 'EQUIPMENT_SPECS'[source]
class src.dackar.RCA.equipment_similarity.equipment_spec_store.EquipmentSpecStore(chroma_store)[source]

Manages the equipment_specs Chroma collection.

Parameters:

chroma_store (Any) – A ChromaRecordStore instance (from storage/chroma_store.py). The store is used as-is; no modifications are made to its configuration.

chroma_store[source]
doc_type = 'EQUIPMENT_SPECS'[source]
upsert_component(component_id, spec_text, metadata=None)[source]

Upsert one component spec into the Chroma collection.

Parameters:
  • component_id (str) – KG element_usage node ID. Used as doc_id and for stable deduplication (same component_id → same Chroma record_id).

  • spec_text (str) – Natural-language spec string from EquipmentSpecBuilder.

  • metadata (Optional[JsonDict]) – Optional extra metadata (component_label, component_type, etc.) stored alongside the embedding for retrieval context.

Return type:

None

upsert_batch(components)[source]

Upsert a batch of component spec dicts.

Each dict must have keys: component_id, spec_text, and optionally metadata.

Returns the number of records upserted.

Parameters:

components (List[JsonDict])

Return type:

int

find_similar(query_text, top_k=10, exclude_ids=None)[source]

Find components with specs most similar to query_text.

Parameters:
  • query_text (str) – Natural-language description of the target equipment (built from kg_context by EquipmentSimilarityResolver._build_query_text()).

  • top_k (int) – Maximum number of results to return (after exclude_ids filtering).

  • exclude_ids (Optional[List[str]]) – component_ids to exclude from results (the target component itself).

Returns:

Each Document has page_content (spec text) and metadata including component_id, the fused RRF _score, the raw dense distance _vector_score, and other stored properties. Returns [] if the collection has not been populated yet.

Return type:

List[langchain_core.documents.Document]

component_count()[source]

Return the number of component specs currently in the collection. Returns 0 if the collection has not been populated.

Return type:

int