src.dackar.RCA.summarizers.reliability_summarizer¶
Doc-type aware summarization utilities for nuclear plant reliability documents.
Designed to integrate with your existing pipeline: - pdfParser.py creates document_index.json (text_md_path, tables_paths, figures, provenance) :contentReference[oaicite:3]{index=3} - mdParser.py parses Markdown into sections and emits chunks.jsonl :contentReference[oaicite:4]{index=4}
This module provides: 1) detect_doc_type(): Identify SOP/CR/WO/ECA/OTHER from doc_name/source_path and early text 2) section_role_from_title(): Tag sections as purpose/steps/evidence/etc. based on doc type 3) build_prompt(): Exact JSON-output prompts per doc type and view type 4) ollama_generate_json(): Call Ollama and parse JSON robustly 5) quality gates: validate_retrieval_summary_json(), validate_rca_frame_json()
All outputs are strict JSON dicts, suitable for: - storing alongside chunk records in chunks.jsonl - feeding to Chroma embedding pipeline (e.g., embedding flattened JSON)
No external dependencies beyond ‘requests’ and the standard library.
Attributes¶
Classes¶
Context for a chunk being summarized. |
|
Seed signals derived from your NER layer and tag regex extraction. |
Functions¶
|
Detect document type using filename/path hints + early extracted text. |
|
Infer a section role from its title based on doc type. |
|
Build the citations block from trusted ChunkContext provenance. |
|
Overwrite provenance fields on a model result with trusted values. |
Return the empty retrieval_summary contract for a chunk. |
|
|
Return the empty rca_frame contract for a chunk. |
|
Build an Ollama prompt that produces STRICT JSON for the required view. |
|
Call Ollama and return a parsed JSON dict. |
Return the first standalone JSON object embedded in |
|
|
Parse JSON from the model output. |
|
Flag a missing or mis-typed |
Validate retrieval_summary shape and return flags (empty => pass). |
|
Validate rca_frame shape and return flags (empty => pass). |
|
|
Robust summarize call with retries. |
Convert retrieval_summary JSON into a dense string for embedding. |
|
Convert rca_frame JSON into an embedding-friendly string. |
Module Contents¶
- class src.dackar.RCA.summarizers.reliability_summarizer.ChunkContext[source]¶
Context for a chunk being summarized.
Input format¶
doc_id: str
doc_type: DocType
chunk_id: str
section_path: str (e.g., “7 Procedure > 7.2 Stroke Time Verification”)
page_start: int
page_end: int
authority_level: str in {“mandatory”,”guidance”,”informational”,”unknown”}
section_role: str (derived label such as “steps”, “evidence”, “analysis”…)
Output format¶
Used to populate the “citations” block of JSON outputs.
- class src.dackar.RCA.summarizers.reliability_summarizer.NERSeed[source]¶
Seed signals derived from your NER layer and tag regex extraction.
Input format¶
- {
“systems”: […], # from NER label “syst” :contentReference[oaicite:5]{index=5} “equipment_ids”: […], # regex-derived (recommend you add) “components”: […], # from ast_*, comp_* labels :contentReference[oaicite:6]{index=6} “mechanisms”: […], # from “deg_mech” :contentReference[oaicite:7]{index=7} “outcomes”: […], # from “fail_type_n” + “event” :contentReference[oaicite:8]{index=8} “surveillance_actions”: […], # from surv_ops_v / surv_ops_n :contentReference[oaicite:9]{index=9} “maintenance_actions”: […], # from mnt_ops :contentReference[oaicite:10]{index=10} “properties”: […], # from prop :contentReference[oaicite:11]{index=11} “tools”: […], # from surv_tool + mnt_tool :contentReference[oaicite:12]{index=12} “fm_ids”: […], # failure-mode identifiers (regex-derived) “measurements”: [{…}], # list of measurement dicts “doc_refs”: […], # referenced document identifiers “alarm_ids”: […], # alarm identifiers “temporal_refs”: […], # absolute/relative time references “temporal_relations”: [{…}], # list of temporal-relation dicts “temporal_qualifiers”: […], # qualifiers (e.g. “intermittent”) “locations”: [{…}], # list of location dicts “conjectures”: […] # hedged/uncertain statements
}
Output format¶
A JSON-serializable dict used inside prompts. The model is instructed to ONLY include items if supported by CHUNK_TEXT.
- to_json()[source]¶
Serialize the seed to the NER_SEED dict embedded in prompts.
- Returns:
One JSON-serializable key per public field of this dataclass, each defaulting to an empty list when unset.
fm_ids(failure-mode identifiers) is included so it reaches the model throughbuild_prompt(); the key set mirrors the class field list.- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer._DOC_TYPE_HINTS: List[Tuple[DocType, List[str]]] = [('SOP', ['standard operating procedure', 'procedure', 'operating procedure', 'sop']), ('CR',...[source]¶
- src.dackar.RCA.summarizers.reliability_summarizer.detect_doc_type(doc_name, source_path, early_text)[source]¶
Detect document type using filename/path hints + early extracted text.
Inputs¶
doc_name: optional filename from document_index[“doc_name”] :contentReference[oaicite:13]{index=13} source_path: optional path from document_index[“source_path”] :contentReference[oaicite:14]{index=14} early_text: first 2-5k chars of cleaned text from the first parsed section (string)
Output¶
DocType: “SOP”|”CR”|”WO”|”ECA”|”OTHER”
- Parameters:
doc_name (Optional[str])
source_path (Optional[str])
early_text (str)
- Return type:
- src.dackar.RCA.summarizers.reliability_summarizer._ROLE_PATTERNS: Dict[DocType, List[Tuple[str, str]]][source]¶
- src.dackar.RCA.summarizers.reliability_summarizer.section_role_from_title(doc_type, section_title)[source]¶
Infer a section role from its title based on doc type.
Inputs¶
doc_type: SOP|CR|WO|ECA|OTHER section_title: the markdown heading title
Output¶
role: string (e.g., “steps”, “evidence”, “analysis”, …) or “unknown”
- Parameters:
doc_type (DocType)
section_title (str)
- Return type:
str
- src.dackar.RCA.summarizers.reliability_summarizer._ENTITY_LIST_FIELDS: Tuple[str, ...] = ('systems', 'equipment_ids', 'components')[source]¶
- src.dackar.RCA.summarizers.reliability_summarizer._RETRIEVAL_SUMMARY_LIST_FIELDS: Tuple[str, ...] = ('symptoms_outcomes', 'mechanisms', 'diagnostics', 'corrective_actions', 'numbers_limits',...[source]¶
- src.dackar.RCA.summarizers.reliability_summarizer._RCA_FRAME_LIST_FIELDS: Tuple[str, ...] = ('observed', 'hypotheses', 'tests_to_confirm', 'candidate_actions', 'constraints')[source]¶
- src.dackar.RCA.summarizers.reliability_summarizer._trusted_citations(ctx)[source]¶
Build the citations block from trusted ChunkContext provenance.
- Parameters:
ctx (ChunkContext)
- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer._apply_trusted_provenance(obj, ctx, view_type)[source]¶
Overwrite provenance fields on a model result with trusted values.
The model is never trusted to report its own
chunk_id,doc_type,view_type, orcitations; these are replaced in place from the ChunkContext so a fabricated value can never be stored as authoritative provenance. Content fields are left untouched for the validator to check.- Parameters:
obj (Dict[str, Any])
ctx (ChunkContext)
view_type (ViewType)
- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer.empty_retrieval_summary(ctx)[source]¶
Return the empty retrieval_summary contract for a chunk.
Used both as the prompt output skeleton (forcing keys/types) and as the canonical shape checked by validate_retrieval_summary_json().
- Parameters:
ctx (ChunkContext) – Trusted provenance; populates the
citationsblock.- Returns:
Every contract key with an empty value of the correct type.
- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer.empty_rca_frame(ctx)[source]¶
Return the empty rca_frame contract for a chunk.
Used both as the prompt output skeleton and as the canonical shape checked by validate_rca_frame_json().
- Parameters:
ctx (ChunkContext) – Trusted provenance; populates the
citationsblock.- Returns:
Every contract key with an empty value of the correct type.
- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer._COMMON_SYSTEM_CONTRACT = Multiline-String[source]¶
Show Value
"""You are an information extraction engine for nuclear plant reliability documents. Return ONLY valid JSON. Do not include markdown, comments, or extra text. Rules: - Use ONLY facts explicitly present in the provided CHUNK_TEXT. - If a field cannot be filled from CHUNK_TEXT, use an empty array [] and add a short note to unknowns (retrieval_summary only). - Preserve numbers, limits, units, and step numbers exactly as written. - Prefer exact phrases from CHUNK_TEXT for technical terms. - Do NOT invent causes, steps, thresholds, or conclusions. - Output must match the requested JSON structure exactly (keys and types). """
- src.dackar.RCA.summarizers.reliability_summarizer.build_prompt(doc_type, view_type, ctx, ner_seed, chunk_text)[source]¶
Build an Ollama prompt that produces STRICT JSON for the required view.
Inputs¶
doc_type: SOP|CR|WO|ECA|OTHER view_type: retrieval_summary|rca_frame ctx: ChunkContext (doc_id/doc_type/chunk_id/section_path/pages/authority/role) ner_seed: NERSeed (JSON-serializable grouped entity hints) chunk_text: string (cleaned text; include table text if relevant)
Output¶
prompt: string for Ollama (/api/chat or /api/generate)
- Parameters:
doc_type (DocType)
view_type (ViewType)
ctx (ChunkContext)
ner_seed (NERSeed)
chunk_text (str)
- Return type:
str
- src.dackar.RCA.summarizers.reliability_summarizer.ollama_generate_json(prompt, model=None, timeout=90)[source]¶
Call Ollama and return a parsed JSON dict.
Inputs¶
prompt: str (must instruct “JSON only”) model: optional model name; defaults to env OLLAMA_MODEL or “mistral:latest” timeout: request timeout seconds
Output¶
dict parsed from model output
Behavior¶
Prefers /api/chat (non-stream for easier JSON parsing)
Falls back to /api/generate
If output contains extra text, attempts to extract the first JSON object via regex
- Parameters:
prompt (str)
model (Optional[str])
timeout (int)
- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer._extract_first_json_object(text)[source]¶
Return the first standalone JSON object embedded in
text.Scans each
{and attemptsraw_decodefrom that position, so a single object is recovered even with extra text on either side (e.g.prefix {"a": 1} suffix {"b": 2}yields{"a": 1}). A greedy first-brace-to-last-brace match would instead span both objects and fail with “Extra data”. ReturnsNonewhen no position decodes to an object.- Parameters:
text (str)
- Return type:
Optional[Dict[str, Any]]
- src.dackar.RCA.summarizers.reliability_summarizer._parse_json_strict(text)[source]¶
Parse JSON from the model output. If the output includes extra text, extract the first {…} block.
Input: text (string) Output: dict Raises: ValueError on failure
- Parameters:
text (str)
- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer._citation_flags(obj)[source]¶
Flag a missing or mis-typed
citationsblock against the contract.- Parameters:
obj (Dict[str, Any])
- Return type:
List[str]
- src.dackar.RCA.summarizers.reliability_summarizer.validate_retrieval_summary_json(obj)[source]¶
Validate retrieval_summary shape and return flags (empty => pass).
Checks every field of the empty_retrieval_summary() contract for presence and type, including the nested entities and citations blocks, using the shared field tables so the gate cannot drift from the skeleton.
Input: dict (parsed JSON) Output: List[str] flags
- Parameters:
obj (Dict[str, Any])
- Return type:
List[str]
- src.dackar.RCA.summarizers.reliability_summarizer.validate_rca_frame_json(obj)[source]¶
Validate rca_frame shape and return flags (empty => pass).
Checks every field of the empty_rca_frame() contract for presence and type, including the nested citations block, using the shared field tables.
Input: dict (parsed JSON) Output: List[str] flags
- Parameters:
obj (Dict[str, Any])
- Return type:
List[str]
- src.dackar.RCA.summarizers.reliability_summarizer.summarize_with_retry(doc_type, view_type, ctx, ner_seed, chunk_text, model=None, timeout=90, max_tries=3, sleep_sec=0.25)[source]¶
Robust summarize call with retries.
Inputs¶
doc_type/view_type/ctx/ner_seed/chunk_text: see build_prompt() model: optional Ollama model name max_tries: integer retry count
Output¶
Parsed JSON dict (retrieval_summary or rca_frame)
Notes
Retry ladder: 1) normal prompt 2) add “Your previous response was invalid JSON. Return valid JSON only.” 3) truncate chunk text to reduce failure risk
- Parameters:
doc_type (DocType)
view_type (ViewType)
ctx (ChunkContext)
ner_seed (NERSeed)
chunk_text (str)
model (Optional[str])
timeout (int)
max_tries (int)
sleep_sec (float)
- Return type:
Dict[str, Any]
- src.dackar.RCA.summarizers.reliability_summarizer.flatten_retrieval_summary_for_embedding(summary)[source]¶
Convert retrieval_summary JSON into a dense string for embedding.
Only fields with content are emitted; empty sections are dropped so they do not dilute the embedding text.
Input: retrieval_summary dict Output: string
- Parameters:
summary (Dict[str, Any])
- Return type:
str
- src.dackar.RCA.summarizers.reliability_summarizer.flatten_rca_frame_for_embedding(rca)[source]¶
Convert rca_frame JSON into an embedding-friendly string.
Only fields with content are emitted; empty sections are dropped so they do not dilute the embedding text.
Input: rca_frame dict Output: string
- Parameters:
rca (Dict[str, Any])
- Return type:
str