src.dackar.RCA.equipment_similarity.equipment_spec_store¶
equipment_spec_store — EquipmentSpecStore.
Thin wrapper around ChromaRecordStore for the equipment_specs doc_type.
One document per plant component; embedding_text is the natural-language
spec string produced by EquipmentSpecBuilder.
The full Stage 1-6 enrichment pipeline (NLP, NER, summarization) is not used here — equipment specs are structured, not unstructured prose, so a minimal record with just the spec text and identity metadata is sufficient.
Attributes¶
Classes¶
Manages the |
Module Contents¶
- src.dackar.RCA.equipment_similarity.equipment_spec_store.EQUIPMENT_SPECS_DOC_TYPE = 'EQUIPMENT_SPECS'[source]¶
- class src.dackar.RCA.equipment_similarity.equipment_spec_store.EquipmentSpecStore(chroma_store)[source]¶
Manages the
equipment_specsChroma collection.- Parameters:
chroma_store (Any) – A
ChromaRecordStoreinstance (fromstorage/chroma_store.py). The store is used as-is; no modifications are made to its configuration.
- upsert_component(component_id, spec_text, metadata=None)[source]¶
Upsert one component spec into the Chroma collection.
- Parameters:
component_id (str) – KG element_usage node ID. Used as
doc_idand for stable deduplication (same component_id → same Chroma record_id).spec_text (str) – Natural-language spec string from
EquipmentSpecBuilder.metadata (Optional[JsonDict]) – Optional extra metadata (component_label, component_type, etc.) stored alongside the embedding for retrieval context.
- Return type:
None
- upsert_batch(components)[source]¶
Upsert a batch of component spec dicts.
Each dict must have keys:
component_id,spec_text, and optionallymetadata.Returns the number of records upserted.
- Parameters:
components (List[JsonDict])
- Return type:
int
- find_similar(query_text, top_k=10, exclude_ids=None)[source]¶
Find components with specs most similar to
query_text.- Parameters:
query_text (str) – Natural-language description of the target equipment (built from kg_context by
EquipmentSimilarityResolver._build_query_text()).top_k (int) – Maximum number of results to return (after exclude_ids filtering).
exclude_ids (Optional[List[str]]) – component_ids to exclude from results (the target component itself).
- Returns:
Each Document has
page_content(spec text) andmetadataincludingcomponent_id, the fused RRF_score, the raw dense distance_vector_score, and other stored properties. Returns[]if the collection has not been populated yet.- Return type:
List[langchain_core.documents.Document]