src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer¶
IncidentIndex — historical episode database.
Stores pre-computed IncidentFingerprints and provides fast candidate retrieval via an inverted index over event types.
- Lifecycle:
build_from_history() — run detector + extractor on raw events_df
add() / add_batch() — incremental updates
get_candidates() — O(|query_event_set|) pre-filter lookup
save() / load() — parquet + JSON persistence
Attributes¶
Classes¶
Stores and manages pre-computed IncidentFingerprints for the historical |
Functions¶
|
Converts a raw events_df to a list of UnifiedEvents. |
Converts an IncidentFingerprint to a flat dict for DataFrame storage. |
|
|
Coerces a pandas Timestamp or datetime to a plain datetime. |
|
Returns a copy of df with frozenset/list/dict columns JSON-encoded as strings. |
|
Returns a copy of df with JSON-string columns decoded back to Python objects. |
Module Contents¶
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._EPISODES_PARQUET = 'episodes.parquet'[source]¶
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._INVERTED_INDEX_JSON = 'inverted_index.json'[source]¶
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._EMD_META_JSON = 'emd_meta.json'[source]¶
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._INDEX_META_JSON = 'index_meta.json'[source]¶
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._COMPLEX_COLS = ('event_set', 'event_seq', 'freq_vec', 'source_types')[source]¶
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer.IncidentIndex(config)[source]¶
Stores and manages pre-computed IncidentFingerprints for the historical episode database.
- Internal storage:
episodes_df — pd.DataFrame, one row per episode. _inverted_index — dict[str, set[str]]: event_type → set of episode_ids.
Used for O(|query_event_set|) Jaccard pre-filtering.
- Parameters:
config (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config.SearchConfig)
- build_from_history(events_df, rho_query, query_duration)[source]¶
Populates the index from a raw historical events_df.
- Steps:
Convert events_df rows to UnifiedEvent list.
Run EpisodeDetector.detect() to find episode boundaries.
Group events by episode boundary.
Derive IncidentFingerprint for each episode via IncidentExtractor._derive_fingerprint().
Store fingerprints via add_batch().
- Parameters:
events_df (pandas.DataFrame) – Raw historical event log. Expected columns: raw_id, asset_id, source, event_type, timestamp_start, timestamp_end. No episode_id column required.
rho_query (float) – Reference density from the query incident (events per second). Passed to EpisodeDetector.
query_duration (float) – D_query in seconds. Used for KDE bandwidth and minimum episode duration filter.
- Return type:
None
Notes
events_df is not modified in place. episode_id is derived from the episode window_start (EP_{asset}_{window_start:%Y%m%dT%H%M%S}), so ids are stable across rebuilds. The add path upserts by episode_id: re-building over the same data replaces episodes in place (idempotent), while genuinely new episodes append — call reset() first for a clean rebuild.
Stage 1 is intended to be built once against a representative query: rho_query/query_duration come from that query and calibrate episode boundaries (delta * rho_query; bandwidth query_duration/4), then the index is persisted via save() (save/load is not a per-query cache). Results are sensitive to these values — see rca_pattern_matching.md for guidance on choosing representative rho_query/query_duration.
- add(fingerprint)[source]¶
Adds a single fingerprint to the index (upsert by episode_id).
If an episode with the same episode_id already exists it is replaced, and the inverted index is rebuilt once to drop the old episode’s stale postings. The common insert-new path stays incremental. Less efficient than add_batch() for many insertions because the inverted index is updated per fingerprint rather than rebuilt once.
- Parameters:
fingerprint (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint)
- Return type:
None
- add_batch(fingerprints)[source]¶
Adds multiple fingerprints in a single operation (upsert by episode_id).
Existing episodes whose episode_id appears in the incoming batch are replaced, and duplicate ids within the batch collapse to the last occurrence, so the append path never produces duplicate episode_ids. Rebuilds the inverted index once after all insertions, which is more efficient than repeated add() calls.
- Parameters:
fingerprints (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint])
- Return type:
None
- get_candidates(query_event_set)[source]¶
Returns episode_ids of historical episodes that share at least one event type with the query.
Uses the inverted index for O(|query_event_set|) lookup. Episodes with no event-type overlap cannot have Jaccard > 0 and are excluded.
- Parameters:
query_event_set (frozenset[str]) – event_set from the query IncidentFingerprint.
- Returns:
De-duplicated list of episode_ids sharing >= 1 event type with the query (candidates are accumulated in a set, so no id repeats).
- Return type:
list[str]
- compute_emd_normalization_factor(max_pairs=1000)[source]¶
Computes the empirical maximum raw L1 distance across historical episode pairs.
Used to normalise EMD scores when emd_normalization_mode=”empirical_max”. Should be called once after build_from_history() and before search().
- Algorithm:
Extract all freq_vec dicts from episodes_df.
If N*(N-1)/2 <= max_pairs: evaluate all pairs.
Else: draw max_pairs random pairs without replacement (seeded).
For each pair (a, b): raw_l1 = Σ_t |a.get(t,0) - b.get(t,0)|
Store the maximum observed L1 distance.
- Parameters:
max_pairs (int) – Maximum number of pairs to evaluate. If the index has fewer pairs than this, all pairs are used.
- Returns:
The empirical maximum raw L1 distance (float >= 0). Returns 1.0 if index is empty or contains only one episode (to avoid division by zero and provide a sensible fallback).
- Return type:
float
- Side effect:
Sets self.emd_normalization_factor to the computed value.
- save(path)[source]¶
Persists the index to disk.
episodes_df is saved as parquet with complex columns (frozenset, list, dict) JSON-serialised to strings. The inverted index is saved as JSON (sets → sorted lists). EMD metadata (normalization factor) is saved as JSON. All files are written atomically via a .tmp rename.
- Parameters:
path (str) – Directory path. Created if it does not exist.
- Return type:
None
- classmethod load(path, config)[source]¶
Loads a persisted index from disk.
Reconstructs episodes_df (deserialising complex columns from JSON strings), the inverted index, and EMD metadata if available.
- Parameters:
path (str) – Directory path written by save().
config (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config.SearchConfig) – SearchConfig to attach to the loaded index.
- Returns:
Populated IncidentIndex.
- Raises:
FileNotFoundError if core files (parquet, inverted index) are missing. –
- Return type:
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._df_to_events(events_df)[source]¶
Converts a raw events_df to a list of UnifiedEvents.
- Parameters:
events_df (pandas.DataFrame)
- Return type:
list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._fingerprint_to_row(fp)[source]¶
Converts an IncidentFingerprint to a flat dict for DataFrame storage.
- Parameters:
fp (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint)
- Return type:
dict
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer._coerce_ts(value)[source]¶
Coerces a pandas Timestamp or datetime to a plain datetime.
pandas NaT subclasses datetime, so the NaT check must come first. pandas Timestamp also subclasses datetime, so to_pydatetime() must be tried before the plain isinstance guard.
- Return type:
datetime.datetime