src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer

IncidentIndex — historical episode database.

Stores pre-computed IncidentFingerprints and provides fast candidate retrieval via an inverted index over event types.

Lifecycle:
  1. build_from_history() — run detector + extractor on raw events_df

  2. add() / add_batch() — incremental updates

  3. get_candidates() — O(|query_event_set|) pre-filter lookup

  4. save() / load() — parquet + JSON persistence

Attributes

_log

_EPISODES_PARQUET

_INVERTED_INDEX_JSON

_EMD_META_JSON

_INDEX_META_JSON

_COMPLEX_COLS

Classes

IncidentIndex

Stores and manages pre-computed IncidentFingerprints for the historical

Functions

_df_to_events(events_df)

Converts a raw events_df to a list of UnifiedEvents.

_fingerprint_to_row(fp)

Converts an IncidentFingerprint to a flat dict for DataFrame storage.

_coerce_ts(value)

Coerces a pandas Timestamp or datetime to a plain datetime.

_serialise_df(df)

Returns a copy of df with frozenset/list/dict columns JSON-encoded as strings.

_deserialise_df(df)

Returns a copy of df with JSON-string columns decoded back to Python objects.

Module Contents

src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._log[source]
src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._EPISODES_PARQUET = 'episodes.parquet'[source]
src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._INVERTED_INDEX_JSON = 'inverted_index.json'[source]
src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._EMD_META_JSON = 'emd_meta.json'[source]
src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._INDEX_META_JSON = 'index_meta.json'[source]
src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._COMPLEX_COLS = ('event_set', 'event_seq', 'freq_vec', 'source_types')[source]
class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer.IncidentIndex(config)[source]

Stores and manages pre-computed IncidentFingerprints for the historical episode database.

Internal storage:

episodes_df — pd.DataFrame, one row per episode. _inverted_index — dict[str, set[str]]: event_type → set of episode_ids.

Used for O(|query_event_set|) Jaccard pre-filtering.

Parameters:

config (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config.SearchConfig)

config[source]
episodes_df: pandas.DataFrame[source]
_inverted_index: dict[str, set[str]][source]
emd_normalization_factor: float | None = None[source]
build_timestamp: datetime.datetime | None = None[source]
asset_scope: list[str] = [][source]
build_from_history(events_df, rho_query, query_duration)[source]

Populates the index from a raw historical events_df.

Steps:
  1. Convert events_df rows to UnifiedEvent list.

  2. Run EpisodeDetector.detect() to find episode boundaries.

  3. Group events by episode boundary.

  4. Derive IncidentFingerprint for each episode via IncidentExtractor._derive_fingerprint().

  5. Store fingerprints via add_batch().

Parameters:
  • events_df (pandas.DataFrame) – Raw historical event log. Expected columns: raw_id, asset_id, source, event_type, timestamp_start, timestamp_end. No episode_id column required.

  • rho_query (float) – Reference density from the query incident (events per second). Passed to EpisodeDetector.

  • query_duration (float) – D_query in seconds. Used for KDE bandwidth and minimum episode duration filter.

Return type:

None

Notes

events_df is not modified in place. episode_id is derived from the episode window_start (EP_{asset}_{window_start:%Y%m%dT%H%M%S}), so ids are stable across rebuilds. The add path upserts by episode_id: re-building over the same data replaces episodes in place (idempotent), while genuinely new episodes append — call reset() first for a clean rebuild.

Stage 1 is intended to be built once against a representative query: rho_query/query_duration come from that query and calibrate episode boundaries (delta * rho_query; bandwidth query_duration/4), then the index is persisted via save() (save/load is not a per-query cache). Results are sensitive to these values — see rca_pattern_matching.md for guidance on choosing representative rho_query/query_duration.

reset()[source]

Clears all stored episodes and the inverted index.

Return type:

None

add(fingerprint)[source]

Adds a single fingerprint to the index (upsert by episode_id).

If an episode with the same episode_id already exists it is replaced, and the inverted index is rebuilt once to drop the old episode’s stale postings. The common insert-new path stays incremental. Less efficient than add_batch() for many insertions because the inverted index is updated per fingerprint rather than rebuilt once.

Parameters:

fingerprint (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint)

Return type:

None

add_batch(fingerprints)[source]

Adds multiple fingerprints in a single operation (upsert by episode_id).

Existing episodes whose episode_id appears in the incoming batch are replaced, and duplicate ids within the batch collapse to the last occurrence, so the append path never produces duplicate episode_ids. Rebuilds the inverted index once after all insertions, which is more efficient than repeated add() calls.

Parameters:

fingerprints (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint])

Return type:

None

get_candidates(query_event_set)[source]

Returns episode_ids of historical episodes that share at least one event type with the query.

Uses the inverted index for O(|query_event_set|) lookup. Episodes with no event-type overlap cannot have Jaccard > 0 and are excluded.

Parameters:

query_event_set (frozenset[str]) – event_set from the query IncidentFingerprint.

Returns:

De-duplicated list of episode_ids sharing >= 1 event type with the query (candidates are accumulated in a set, so no id repeats).

Return type:

list[str]

compute_emd_normalization_factor(max_pairs=1000)[source]

Computes the empirical maximum raw L1 distance across historical episode pairs.

Used to normalise EMD scores when emd_normalization_mode=”empirical_max”. Should be called once after build_from_history() and before search().

Algorithm:
  • Extract all freq_vec dicts from episodes_df.

  • If N*(N-1)/2 <= max_pairs: evaluate all pairs.

  • Else: draw max_pairs random pairs without replacement (seeded).

  • For each pair (a, b): raw_l1 = Σ_t |a.get(t,0) - b.get(t,0)|

  • Store the maximum observed L1 distance.

Parameters:

max_pairs (int) – Maximum number of pairs to evaluate. If the index has fewer pairs than this, all pairs are used.

Returns:

The empirical maximum raw L1 distance (float >= 0). Returns 1.0 if index is empty or contains only one episode (to avoid division by zero and provide a sensible fallback).

Return type:

float

Side effect:

Sets self.emd_normalization_factor to the computed value.

save(path)[source]

Persists the index to disk.

episodes_df is saved as parquet with complex columns (frozenset, list, dict) JSON-serialised to strings. The inverted index is saved as JSON (sets → sorted lists). EMD metadata (normalization factor) is saved as JSON. All files are written atomically via a .tmp rename.

Parameters:

path (str) – Directory path. Created if it does not exist.

Return type:

None

classmethod load(path, config)[source]

Loads a persisted index from disk.

Reconstructs episodes_df (deserialising complex columns from JSON strings), the inverted index, and EMD metadata if available.

Parameters:
Returns:

Populated IncidentIndex.

Raises:

FileNotFoundError if core files (parquet, inverted index) are missing. –

Return type:

IncidentIndex

_rebuild_inverted_index()[source]

Rebuilds the inverted index from scratch from episodes_df.

Return type:

None

__len__()[source]
Return type:

int

__repr__()[source]
Return type:

str

src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._df_to_events(events_df)[source]

Converts a raw events_df to a list of UnifiedEvents.

Parameters:

events_df (pandas.DataFrame)

Return type:

list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]

src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._fingerprint_to_row(fp)[source]

Converts an IncidentFingerprint to a flat dict for DataFrame storage.

Parameters:

fp (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint)

Return type:

dict

src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._coerce_ts(value)[source]

Coerces a pandas Timestamp or datetime to a plain datetime.

pandas NaT subclasses datetime, so the NaT check must come first. pandas Timestamp also subclasses datetime, so to_pydatetime() must be tried before the plain isinstance guard.

Return type:

datetime.datetime

src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._serialise_df(df)[source]

Returns a copy of df with frozenset/list/dict columns JSON-encoded as strings.

Parameters:

df (pandas.DataFrame)

Return type:

pandas.DataFrame

src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._deserialise_df(df)[source]

Returns a copy of df with JSON-string columns decoded back to Python objects.

Parameters:

df (pandas.DataFrame)

Return type:

pandas.DataFrame