src.dackar.RCA.summarizers.reliability_summarizer

Doc-type aware summarization utilities for nuclear plant reliability documents.

Designed to integrate with your existing pipeline: - pdfParser.py creates document_index.json (text_md_path, tables_paths, figures, provenance) :contentReference[oaicite:3]{index=3} - mdParser.py parses Markdown into sections and emits chunks.jsonl :contentReference[oaicite:4]{index=4}

This module provides: 1) detect_doc_type(): Identify SOP/CR/WO/ECA/OTHER from doc_name/source_path and early text 2) section_role_from_title(): Tag sections as purpose/steps/evidence/etc. based on doc type 3) build_prompt(): Exact JSON-output prompts per doc type and view type 4) ollama_generate_json(): Call Ollama and parse JSON robustly 5) quality gates: validate_retrieval_summary_json(), validate_rca_frame_json()

All outputs are strict JSON dicts, suitable for: - storing alongside chunk records in chunks.jsonl - feeding to Chroma embedding pipeline (e.g., embedding flattened JSON)

No external dependencies beyond ‘requests’ and the standard library.

Attributes

DocType

ViewType

_DOC_TYPE_HINTS

_ROLE_PATTERNS

_ENTITY_LIST_FIELDS

_RETRIEVAL_SUMMARY_LIST_FIELDS

_RCA_FRAME_LIST_FIELDS

_CITATION_FIELD_TYPES

_COMMON_SYSTEM_CONTRACT

Classes

ChunkContext

Context for a chunk being summarized.

NERSeed

Seed signals derived from your NER layer and tag regex extraction.

Functions

detect_doc_type(doc_name, source_path, early_text)

Detect document type using filename/path hints + early extracted text.

section_role_from_title(doc_type, section_title)

Infer a section role from its title based on doc type.

_trusted_citations(ctx)

Build the citations block from trusted ChunkContext provenance.

_apply_trusted_provenance(obj, ctx, view_type)

Overwrite provenance fields on a model result with trusted values.

empty_retrieval_summary(ctx)

Return the empty retrieval_summary contract for a chunk.

empty_rca_frame(ctx)

Return the empty rca_frame contract for a chunk.

build_prompt(doc_type, view_type, ctx, ner_seed, ...)

Build an Ollama prompt that produces STRICT JSON for the required view.

ollama_generate_json(prompt[, model, timeout])

Call Ollama and return a parsed JSON dict.

_extract_first_json_object(text)

Return the first standalone JSON object embedded in text.

_parse_json_strict(text)

Parse JSON from the model output.

_citation_flags(obj)

Flag a missing or mis-typed citations block against the contract.

validate_retrieval_summary_json(obj)

Validate retrieval_summary shape and return flags (empty => pass).

validate_rca_frame_json(obj)

Validate rca_frame shape and return flags (empty => pass).

summarize_with_retry(doc_type, view_type, ctx, ...[, ...])

Robust summarize call with retries.

flatten_retrieval_summary_for_embedding(summary)

Convert retrieval_summary JSON into a dense string for embedding.

flatten_rca_frame_for_embedding(rca)

Convert rca_frame JSON into an embedding-friendly string.

Module Contents

src.dackar.RCA.summarizers.reliability_summarizer.DocType[source]
src.dackar.RCA.summarizers.reliability_summarizer.ViewType[source]
class src.dackar.RCA.summarizers.reliability_summarizer.ChunkContext[source]

Context for a chunk being summarized.

Input format

  • doc_id: str

  • doc_type: DocType

  • chunk_id: str

  • section_path: str (e.g., “7 Procedure > 7.2 Stroke Time Verification”)

  • page_start: int

  • page_end: int

  • authority_level: str in {“mandatory”,”guidance”,”informational”,”unknown”}

  • section_role: str (derived label such as “steps”, “evidence”, “analysis”…)

Output format

Used to populate the “citations” block of JSON outputs.

doc_id: str[source]
doc_type: DocType[source]
chunk_id: str[source]
section_path: str[source]
page_start: int[source]
page_end: int[source]
authority_level: str = 'unknown'[source]
section_role: str = 'unknown'[source]
class src.dackar.RCA.summarizers.reliability_summarizer.NERSeed[source]

Seed signals derived from your NER layer and tag regex extraction.

Input format

{

“systems”: […], # from NER label “syst” :contentReference[oaicite:5]{index=5} “equipment_ids”: […], # regex-derived (recommend you add) “components”: […], # from ast_*, comp_* labels :contentReference[oaicite:6]{index=6} “mechanisms”: […], # from “deg_mech” :contentReference[oaicite:7]{index=7} “outcomes”: […], # from “fail_type_n” + “event” :contentReference[oaicite:8]{index=8} “surveillance_actions”: […], # from surv_ops_v / surv_ops_n :contentReference[oaicite:9]{index=9} “maintenance_actions”: […], # from mnt_ops :contentReference[oaicite:10]{index=10} “properties”: […], # from prop :contentReference[oaicite:11]{index=11} “tools”: […], # from surv_tool + mnt_tool :contentReference[oaicite:12]{index=12} “fm_ids”: […], # failure-mode identifiers (regex-derived) “measurements”: [{…}], # list of measurement dicts “doc_refs”: […], # referenced document identifiers “alarm_ids”: […], # alarm identifiers “temporal_refs”: […], # absolute/relative time references “temporal_relations”: [{…}], # list of temporal-relation dicts “temporal_qualifiers”: […], # qualifiers (e.g. “intermittent”) “locations”: [{…}], # list of location dicts “conjectures”: […] # hedged/uncertain statements

}

Output format

A JSON-serializable dict used inside prompts. The model is instructed to ONLY include items if supported by CHUNK_TEXT.

systems: List[str][source]
equipment_ids: List[str][source]
components: List[str][source]
mechanisms: List[str][source]
outcomes: List[str][source]
surveillance_actions: List[str][source]
maintenance_actions: List[str][source]
properties: List[str][source]
tools: List[str][source]
fm_ids: List[str] = [][source]
measurements: List[Dict[str, Any]] = [][source]
doc_refs: List[str] = [][source]
alarm_ids: List[str] = [][source]
temporal_refs: List[str] = [][source]
temporal_relations: List[Dict[str, str]] = [][source]
temporal_qualifiers: List[str] = [][source]
locations: List[Dict[str, str]] = [][source]
conjectures: List[str] = [][source]
to_json()[source]

Serialize the seed to the NER_SEED dict embedded in prompts.

Returns:

One JSON-serializable key per public field of this dataclass, each defaulting to an empty list when unset. fm_ids (failure-mode identifiers) is included so it reaches the model through build_prompt(); the key set mirrors the class field list.

Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer._DOC_TYPE_HINTS: List[Tuple[DocType, List[str]]] = [('SOP', ['standard operating procedure', 'procedure', 'operating procedure', 'sop']), ('CR',...[source]
src.dackar.RCA.summarizers.reliability_summarizer.detect_doc_type(doc_name, source_path, early_text)[source]

Detect document type using filename/path hints + early extracted text.

Inputs

doc_name: optional filename from document_index[“doc_name”] :contentReference[oaicite:13]{index=13} source_path: optional path from document_index[“source_path”] :contentReference[oaicite:14]{index=14} early_text: first 2-5k chars of cleaned text from the first parsed section (string)

Output

DocType: “SOP”|”CR”|”WO”|”ECA”|”OTHER”

Parameters:
  • doc_name (Optional[str])

  • source_path (Optional[str])

  • early_text (str)

Return type:

DocType

src.dackar.RCA.summarizers.reliability_summarizer._ROLE_PATTERNS: Dict[DocType, List[Tuple[str, str]]][source]
src.dackar.RCA.summarizers.reliability_summarizer.section_role_from_title(doc_type, section_title)[source]

Infer a section role from its title based on doc type.

Inputs

doc_type: SOP|CR|WO|ECA|OTHER section_title: the markdown heading title

Output

role: string (e.g., “steps”, “evidence”, “analysis”, …) or “unknown”

Parameters:
  • doc_type (DocType)

  • section_title (str)

Return type:

str

src.dackar.RCA.summarizers.reliability_summarizer._ENTITY_LIST_FIELDS: Tuple[str, ...] = ('systems', 'equipment_ids', 'components')[source]
src.dackar.RCA.summarizers.reliability_summarizer._RETRIEVAL_SUMMARY_LIST_FIELDS: Tuple[str, ...] = ('symptoms_outcomes', 'mechanisms', 'diagnostics', 'corrective_actions', 'numbers_limits',...[source]
src.dackar.RCA.summarizers.reliability_summarizer._RCA_FRAME_LIST_FIELDS: Tuple[str, ...] = ('observed', 'hypotheses', 'tests_to_confirm', 'candidate_actions', 'constraints')[source]
src.dackar.RCA.summarizers.reliability_summarizer._CITATION_FIELD_TYPES: Dict[str, type][source]
src.dackar.RCA.summarizers.reliability_summarizer._trusted_citations(ctx)[source]

Build the citations block from trusted ChunkContext provenance.

Parameters:

ctx (ChunkContext)

Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer._apply_trusted_provenance(obj, ctx, view_type)[source]

Overwrite provenance fields on a model result with trusted values.

The model is never trusted to report its own chunk_id, doc_type, view_type, or citations; these are replaced in place from the ChunkContext so a fabricated value can never be stored as authoritative provenance. Content fields are left untouched for the validator to check.

Parameters:
Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer.empty_retrieval_summary(ctx)[source]

Return the empty retrieval_summary contract for a chunk.

Used both as the prompt output skeleton (forcing keys/types) and as the canonical shape checked by validate_retrieval_summary_json().

Parameters:

ctx (ChunkContext) – Trusted provenance; populates the citations block.

Returns:

Every contract key with an empty value of the correct type.

Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer.empty_rca_frame(ctx)[source]

Return the empty rca_frame contract for a chunk.

Used both as the prompt output skeleton and as the canonical shape checked by validate_rca_frame_json().

Parameters:

ctx (ChunkContext) – Trusted provenance; populates the citations block.

Returns:

Every contract key with an empty value of the correct type.

Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer._COMMON_SYSTEM_CONTRACT = Multiline-String[source]
Show Value
"""You are an information extraction engine for nuclear plant reliability documents.
Return ONLY valid JSON. Do not include markdown, comments, or extra text.

Rules:
- Use ONLY facts explicitly present in the provided CHUNK_TEXT.
- If a field cannot be filled from CHUNK_TEXT, use an empty array [] and add a short note to unknowns (retrieval_summary only).
- Preserve numbers, limits, units, and step numbers exactly as written.
- Prefer exact phrases from CHUNK_TEXT for technical terms.
- Do NOT invent causes, steps, thresholds, or conclusions.
- Output must match the requested JSON structure exactly (keys and types).
"""
src.dackar.RCA.summarizers.reliability_summarizer.build_prompt(doc_type, view_type, ctx, ner_seed, chunk_text)[source]

Build an Ollama prompt that produces STRICT JSON for the required view.

Inputs

doc_type: SOP|CR|WO|ECA|OTHER view_type: retrieval_summary|rca_frame ctx: ChunkContext (doc_id/doc_type/chunk_id/section_path/pages/authority/role) ner_seed: NERSeed (JSON-serializable grouped entity hints) chunk_text: string (cleaned text; include table text if relevant)

Output

prompt: string for Ollama (/api/chat or /api/generate)

Parameters:
Return type:

str

src.dackar.RCA.summarizers.reliability_summarizer.ollama_generate_json(prompt, model=None, timeout=90)[source]

Call Ollama and return a parsed JSON dict.

Inputs

prompt: str (must instruct “JSON only”) model: optional model name; defaults to env OLLAMA_MODEL or “mistral:latest” timeout: request timeout seconds

Output

dict parsed from model output

Behavior

  • Prefers /api/chat (non-stream for easier JSON parsing)

  • Falls back to /api/generate

  • If output contains extra text, attempts to extract the first JSON object via regex

Parameters:
  • prompt (str)

  • model (Optional[str])

  • timeout (int)

Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer._extract_first_json_object(text)[source]

Return the first standalone JSON object embedded in text.

Scans each { and attempts raw_decode from that position, so a single object is recovered even with extra text on either side (e.g. prefix {"a": 1} suffix {"b": 2} yields {"a": 1}). A greedy first-brace-to-last-brace match would instead span both objects and fail with “Extra data”. Returns None when no position decodes to an object.

Parameters:

text (str)

Return type:

Optional[Dict[str, Any]]

src.dackar.RCA.summarizers.reliability_summarizer._parse_json_strict(text)[source]

Parse JSON from the model output. If the output includes extra text, extract the first {…} block.

Input: text (string) Output: dict Raises: ValueError on failure

Parameters:

text (str)

Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer._citation_flags(obj)[source]

Flag a missing or mis-typed citations block against the contract.

Parameters:

obj (Dict[str, Any])

Return type:

List[str]

src.dackar.RCA.summarizers.reliability_summarizer.validate_retrieval_summary_json(obj)[source]

Validate retrieval_summary shape and return flags (empty => pass).

Checks every field of the empty_retrieval_summary() contract for presence and type, including the nested entities and citations blocks, using the shared field tables so the gate cannot drift from the skeleton.

Input: dict (parsed JSON) Output: List[str] flags

Parameters:

obj (Dict[str, Any])

Return type:

List[str]

src.dackar.RCA.summarizers.reliability_summarizer.validate_rca_frame_json(obj)[source]

Validate rca_frame shape and return flags (empty => pass).

Checks every field of the empty_rca_frame() contract for presence and type, including the nested citations block, using the shared field tables.

Input: dict (parsed JSON) Output: List[str] flags

Parameters:

obj (Dict[str, Any])

Return type:

List[str]

src.dackar.RCA.summarizers.reliability_summarizer.summarize_with_retry(doc_type, view_type, ctx, ner_seed, chunk_text, model=None, timeout=90, max_tries=3, sleep_sec=0.25)[source]

Robust summarize call with retries.

Inputs

doc_type/view_type/ctx/ner_seed/chunk_text: see build_prompt() model: optional Ollama model name max_tries: integer retry count

Output

Parsed JSON dict (retrieval_summary or rca_frame)

Notes

Retry ladder: 1) normal prompt 2) add “Your previous response was invalid JSON. Return valid JSON only.” 3) truncate chunk text to reduce failure risk

Parameters:
Return type:

Dict[str, Any]

src.dackar.RCA.summarizers.reliability_summarizer.flatten_retrieval_summary_for_embedding(summary)[source]

Convert retrieval_summary JSON into a dense string for embedding.

Only fields with content are emitted; empty sections are dropped so they do not dilute the embedding text.

Input: retrieval_summary dict Output: string

Parameters:

summary (Dict[str, Any])

Return type:

str

src.dackar.RCA.summarizers.reliability_summarizer.flatten_rca_frame_for_embedding(rca)[source]

Convert rca_frame JSON into an embedding-friendly string.

Only fields with content are emitted; empty sections are dropped so they do not dilute the embedding text.

Input: rca_frame dict Output: string

Parameters:

rca (Dict[str, Any])

Return type:

str