src.dackar.RCA.log_pattern_recognition.rca_pattern_search¶
rca_pattern_search — RCA Temporal Pattern Matching
- Two-stage pipeline:
Stage 1 (offline): KDE-based episode detection from historical event logs Stage 2 (online): Coarse-to-fine similarity retrieval
Typical usage:
from dackar.RCA.log_pattern_recognition.rca_pattern_search import (
SearchConfig, IncidentIndex, PatternSearcher, IncidentExtractor,
)
cfg = SearchConfig()
# --- Stage 1: build historical index (run offline) ---
index = IncidentIndex(cfg)
index.build_from_history(events_df, rho_query=rho, query_duration=dur)
index.save("/path/to/index")
# --- Stage 2: query at runtime ---
index = IncidentIndex.load("/path/to/index", cfg)
extractor = IncidentExtractor(cfg)
query_fp = extractor.extract(alarm_log, soe_log, telemetry, "INC_001", t0, t1)
searcher = PatternSearcher(index, cfg)
results = searcher.search(query_fp)
Submodules¶
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.density
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.extractor
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.indexer
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.metrics
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.searcher
Classes¶
Operational configuration for the pattern search subsystem. |
|
Central configuration for the RCA pattern search pipeline. |
|
Detects episode boundaries in a continuous historical event log using |
|
Converts raw source records to UnifiedEvents and derives IncidentFingerprints. |
|
Stores and manages pre-computed IncidentFingerprints for the historical |
|
Public output type for PatternSearcher.search(). |
|
Pre-computed similarity representations for a single incident or detected |
|
Canonical representation of a single event from any source. |
|
Retrieves the top-k most similar historical episodes for a query incident. |
Functions¶
|
Weighted combination of the three metric scores. |
|
Frequency-based similarity between two event count vectors. |
|
Set-based similarity between two event sets. |
|
Normalised Longest Common Subsequence similarity. |
Package Contents¶
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.PatternSearchConfig[source]¶
Operational configuration for the pattern search subsystem.
Kept separate from SearchConfig (which tunes the similarity algorithm) and from CrossPatternConfig (Phase 2).
- enable_signal_episode_search: master switch; when False the subsystem
is completely bypassed and historical_signal_episodes.json is not written.
- index_staleness_window_days: episode index older than this is flagged “stale”;
links built against stale results are capped at confidence 0.70 (§4.11).
search_config: algorithm tuning (thresholds, weights, window expansion, etc.).
- enable_signal_episode_search: bool = False¶
- index_staleness_window_days: int = 30¶
- search_config: SearchConfig¶
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.SearchConfig[source]¶
Central configuration for the RCA pattern search pipeline.
All parameters are tunable and should be validated empirically against historical data before production use.
- beta: float = 0.2¶
- delta: float = 0.5¶
- kde_bandwidth: float | str = 'auto'¶
- freq_threshold: int = 5¶
- min_jaccard: float = 0.3¶
- top_k: int = 5¶
- alpha: float = 0.3333333333333333¶
- beta_w: float = 0.3333333333333333¶
- gamma: float = 0.3333333333333333¶
- weight_profile: str = 'equal'¶
- emd_normalization_mode: str = 'tv'¶
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.EpisodeDetector(config)[source]¶
Detects episode boundaries in a continuous historical event log using kernel density estimation (KDE) over event timestamps.
The detection threshold is relative to the query incident density, making the method self-calibrating: a busier query requires historically busier windows to qualify as matching episodes.
- Pipeline (per detect() call):
Convert event timestamps to float seconds since earliest event.
Evaluate Gaussian KDE on a fine time grid.
Threshold: mask = KDE(t) >= delta * rho_query.
Extract contiguous masked regions as raw episode boundaries.
Apply beta buffer expansion to each boundary.
Merge overlapping expanded boundaries.
Discard episodes shorter than query_duration / 10.
- Parameters:
config (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config.SearchConfig)
- config¶
- compute_reference_density(query_events, window_start, window_end)[source]¶
Computes the reference event density for the query incident.
rho_query = N_query / D_query
N_query: events with timestamp_start in [window_start, window_end]. D_query: (window_end - window_start).total_seconds()
All sources contribute equally. Returns 0.0 if duration <= 0.
- Parameters:
query_events (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent])
window_start (datetime.datetime)
window_end (datetime.datetime)
- Return type:
float
- detect(historical_events, rho_query, query_duration)[source]¶
Detects episode boundaries in the historical event log.
- Parameters:
historical_events (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]) – Full flat event list, all sources merged. No episode_id expected at this stage.
rho_query (float) – Reference density from compute_reference_density(). Units: events per second.
query_duration (float) – D_query in seconds. Used for KDE bandwidth and minimum episode duration filter.
- Returns:
List of (episode_start, episode_end) tuples, sorted ascending. These are expanded boundaries (beta already applied) ready for fingerprinting. Empty list if no qualifying episodes found.
- Return type:
list[tuple[datetime.datetime, datetime.datetime]]
- bandwidth_scan(historical_events, rho_query, query_duration, bandwidths=None)[source]¶
Multi-scale diagnostic: counts detected episodes at different bandwidths.
Helps operators validate episode segmentation by showing how many episodes are detected when the smoothing scale varies. Useful when query timescale may not match historical episode timescales (e.g., fast transient query for slow degradation history, or vice versa).
- Parameters:
historical_events (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]) – Full flat event list, all sources merged.
rho_query (float) – Reference density from compute_reference_density(). Units: events per second.
query_duration (float) – D_query in seconds.
bandwidths (Optional[list[float]]) – Explicit bandwidth list in seconds. If None, defaults to [D/32, D/16, D/8, D/4, D/2, D, 2D, 4D] for broad coverage.
- Returns:
dict mapping bandwidth (float, seconds) to episode count (int), sorted by bandwidth ascending.
- Return type:
dict[float, int]
- _run_detection(t_seconds, t_epoch, rho_query, query_duration, bw)[source]¶
Core detection pipeline: KDE → threshold → extract → expand → merge → filter.
Called by detect() and bandwidth_scan() to avoid code duplication.
- Parameters:
t_seconds (numpy.ndarray) – Event timestamps as float seconds since t_epoch.
t_epoch (datetime.datetime) – Reference time (datetime).
rho_query (float) – Reference density (events/second).
query_duration (float) – Duration of query window in seconds.
bw (float) – Bandwidth in seconds (already resolved, > 0).
- Returns:
Expanded, merged, filtered episode boundaries.
- Return type:
list[tuple[datetime.datetime, datetime.datetime]]
- assign_episode_ids(historical_events, episode_boundaries)[source]¶
Assigns episode_id to each historical event based on detected boundaries.
An event is assigned to the episode whose expanded boundary contains its timestamp_start. Events outside all boundaries retain episode_id = None (background noise).
If an event falls within multiple boundaries (should not occur after merging, handled defensively), it is assigned to the first match and a WARNING is logged.
Episode ids: “EP_{asset_id}_{index:05d}” where asset_id is the dominant asset among events in that episode.
- Parameters:
historical_events (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]) – Full flat event list.
episode_boundaries (list[tuple[datetime.datetime, datetime.datetime]]) – Output of detect(), sorted (start, end) tuples.
- Returns:
New list of UnifiedEvents with episode_id populated. Original list is not mutated.
- Return type:
list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.IncidentExtractor(config)[source]¶
Converts raw source records to UnifiedEvents and derives IncidentFingerprints.
- Two public entry points:
to_unified_events() — normalise all three sources into a flat list extract() — full pipeline: expand window → filter → fingerprint
_derive_fingerprint() is a staticmethod so that indexer.py can call it directly on pre-filtered episode event lists without instantiating an extractor.
- Parameters:
config (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config.SearchConfig)
- config¶
- to_unified_events(alarm_log, soe_log, telemetry_summaries, *, anomaly_severity_threshold=_DEFAULT_SEVERITY_THRESHOLD)[source]¶
Converts raw schema records to a flat UnifiedEvent list.
- Applies filtering defaults:
Alarms with state == “suppressed” are excluded.
Anomalies with promoted_to_kg_event == False are excluded. If the field is absent, severity_score >= anomaly_severity_threshold is used as the inclusion gate.
All SOE records are included.
- Parameters:
alarm_log (dict) – Dict with key “alarms”: list of alarm dicts.
soe_log (dict) – Dict with key “records”: list of SOE record dicts.
telemetry_summaries (list[dict]) – List of telemetry summary dicts, each with key “anomalies”: list of anomaly dicts.
anomaly_severity_threshold (float) – Fallback gate when promoted_to_kg_event absent.
- Returns:
Flat list of UnifiedEvent, unsorted.
- Return type:
list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]
- extract(alarm_log, soe_log, telemetry_summaries, incident_id, window_start, window_end, metadata=None)[source]¶
Full extraction pipeline for a single query incident.
- Steps:
Apply beta buffer to compute expanded window.
Call to_unified_events() for all sources.
Filter events to those with timestamp_start in expanded window.
Compute density over the expanded window (consistent with historical episode density).
Derive event_set, event_seq, freq_vec via _derive_fingerprint().
- Parameters:
alarm_log (dict) – Raw alarm log dict.
soe_log (dict) – Raw SOE log dict.
telemetry_summaries (list[dict]) – List of telemetry summary dicts.
incident_id (str) – Identifier for this query incident.
window_start (datetime.datetime) – Incident window start (before buffer expansion).
window_end (datetime.datetime) – Incident window end (before buffer expansion).
metadata (Optional[dict]) – Optional dict; “known_rca” and “asset_id” are read if present.
- Returns:
IncidentFingerprint with expanded window stored as window_start/end.
- Return type:
src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint
- static _derive_fingerprint(events, freq_threshold)[source]¶
Derives the three similarity representations from a list of events.
High-frequency event types (count > freq_threshold) are excluded from event_set and event_seq but retained in freq_vec.
- Parameters:
events (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]) – Events belonging to a single episode or incident.
freq_threshold (int) – Count above which a type is considered high-frequency.
- Returns:
(event_set, event_seq, freq_vec)
- Return type:
tuple[frozenset[str], list[str], dict[str, int]]
- _parse_alarm_log(alarm_log)[source]¶
- Parameters:
alarm_log (dict)
- Return type:
list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]
- _parse_soe_log(soe_log)[source]¶
- Parameters:
soe_log (dict)
- Return type:
list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]
- _parse_telemetry_summary(summary, severity_threshold)[source]¶
- Parameters:
summary (dict)
severity_threshold (float)
- Return type:
list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.UnifiedEvent]
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.IncidentIndex(config)[source]¶
Stores and manages pre-computed IncidentFingerprints for the historical episode database.
- Internal storage:
episodes_df — pd.DataFrame, one row per episode. _inverted_index — dict[str, set[str]]: event_type → set of episode_ids.
Used for O(|query_event_set|) Jaccard pre-filtering.
- Parameters:
config (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config.SearchConfig)
- config¶
- episodes_df: pandas.DataFrame¶
- _inverted_index: dict[str, set[str]]¶
- emd_normalization_factor: float | None = None¶
- build_timestamp: datetime.datetime | None = None¶
- asset_scope: list[str] = []¶
- build_from_history(events_df, rho_query, query_duration)[source]¶
Populates the index from a raw historical events_df.
- Steps:
Convert events_df rows to UnifiedEvent list.
Run EpisodeDetector.detect() to find episode boundaries.
Group events by episode boundary.
Derive IncidentFingerprint for each episode via IncidentExtractor._derive_fingerprint().
Store fingerprints via add_batch().
- Parameters:
events_df (pandas.DataFrame) – Raw historical event log. Expected columns: raw_id, asset_id, source, event_type, timestamp_start, timestamp_end. No episode_id column required.
rho_query (float) – Reference density from the query incident (events per second). Passed to EpisodeDetector.
query_duration (float) – D_query in seconds. Used for KDE bandwidth and minimum episode duration filter.
- Return type:
None
Notes
events_df is not modified in place. episode_id is derived from the episode window_start (EP_{asset}_{window_start:%Y%m%dT%H%M%S}), so ids are stable across rebuilds. The add path upserts by episode_id: re-building over the same data replaces episodes in place (idempotent), while genuinely new episodes append — call reset() first for a clean rebuild.
Stage 1 is intended to be built once against a representative query: rho_query/query_duration come from that query and calibrate episode boundaries (delta * rho_query; bandwidth query_duration/4), then the index is persisted via save() (save/load is not a per-query cache). Results are sensitive to these values — see rca_pattern_matching.md for guidance on choosing representative rho_query/query_duration.
- add(fingerprint)[source]¶
Adds a single fingerprint to the index (upsert by episode_id).
If an episode with the same episode_id already exists it is replaced, and the inverted index is rebuilt once to drop the old episode’s stale postings. The common insert-new path stays incremental. Less efficient than add_batch() for many insertions because the inverted index is updated per fingerprint rather than rebuilt once.
- Parameters:
fingerprint (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint)
- Return type:
None
- add_batch(fingerprints)[source]¶
Adds multiple fingerprints in a single operation (upsert by episode_id).
Existing episodes whose episode_id appears in the incoming batch are replaced, and duplicate ids within the batch collapse to the last occurrence, so the append path never produces duplicate episode_ids. Rebuilds the inverted index once after all insertions, which is more efficient than repeated add() calls.
- Parameters:
fingerprints (list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint])
- Return type:
None
- get_candidates(query_event_set)[source]¶
Returns episode_ids of historical episodes that share at least one event type with the query.
Uses the inverted index for O(|query_event_set|) lookup. Episodes with no event-type overlap cannot have Jaccard > 0 and are excluded.
- Parameters:
query_event_set (frozenset[str]) – event_set from the query IncidentFingerprint.
- Returns:
De-duplicated list of episode_ids sharing >= 1 event type with the query (candidates are accumulated in a set, so no id repeats).
- Return type:
list[str]
- compute_emd_normalization_factor(max_pairs=1000)[source]¶
Computes the empirical maximum raw L1 distance across historical episode pairs.
Used to normalise EMD scores when emd_normalization_mode=”empirical_max”. Should be called once after build_from_history() and before search().
- Algorithm:
Extract all freq_vec dicts from episodes_df.
If N*(N-1)/2 <= max_pairs: evaluate all pairs.
Else: draw max_pairs random pairs without replacement (seeded).
For each pair (a, b): raw_l1 = Σ_t |a.get(t,0) - b.get(t,0)|
Store the maximum observed L1 distance.
- Parameters:
max_pairs (int) – Maximum number of pairs to evaluate. If the index has fewer pairs than this, all pairs are used.
- Returns:
The empirical maximum raw L1 distance (float >= 0). Returns 1.0 if index is empty or contains only one episode (to avoid division by zero and provide a sensible fallback).
- Return type:
float
- Side effect:
Sets self.emd_normalization_factor to the computed value.
- save(path)[source]¶
Persists the index to disk.
episodes_df is saved as parquet with complex columns (frozenset, list, dict) JSON-serialised to strings. The inverted index is saved as JSON (sets → sorted lists). EMD metadata (normalization factor) is saved as JSON. All files are written atomically via a .tmp rename.
- Parameters:
path (str) – Directory path. Created if it does not exist.
- Return type:
None
- classmethod load(path, config)[source]¶
Loads a persisted index from disk.
Reconstructs episodes_df (deserialising complex columns from JSON strings), the inverted index, and EMD metadata if available.
- Parameters:
path (str) – Directory path written by save().
config (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.config.SearchConfig) – SearchConfig to attach to the loaded index.
- Returns:
Populated IncidentIndex.
- Raises:
FileNotFoundError if core files (parquet, inverted index) are missing. –
- Return type:
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.combined_score(j, n, e, alpha, beta_w, gamma)[source]¶
Weighted combination of the three metric scores.
Score = alpha · J + beta_w · NLCS + gamma · EMD
Inputs are assumed to be valid (alpha + beta_w + gamma ≈ 1). No re-normalisation is applied here; the caller (PatternSearcher) is responsible for passing a coherent weight triple via SearchConfig.
- Parameters:
j (float)
n (float)
e (float)
alpha (float)
beta_w (float)
gamma (float)
- Return type:
float
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.emd_similarity(a, b, normalization_factor=None)[source]¶
Frequency-based similarity between two event count vectors.
Captures the repetition signal that Jaccard and NLCS deliberately discard.
Implementation uses the Total Variation (TV) distance between the two probability distributions derived by normalising the count vectors. For categorical distributions with unit ground distance between any two distinct types, TV distance equals the (unit-ground) Earth Mover’s Distance:
TV(P, Q) = 0.5 · Σ_t |P(t) − Q(t)| where P(t) = a[t] / Σa
This is always in [0, 1], so:
emd_similarity = 1 − TV(P, Q)
- Alternative (raw-count) normalisation:
If normalization_factor is provided the raw L1 distance between the unnormalised count vectors is used instead:
raw_emd = Σ_t |a.get(t, 0) − b.get(t, 0)| emd_similarity = max(0.0, 1 − raw_emd / normalization_factor)
Suitable when an empirically derived or vocabulary-size-based upper bound is available (see spec open point on normalisation).
- Edge cases:
Both empty → 1.0 (identical — neither has any events) One empty → 0.0 (maximally dissimilar)
- Parameters:
a (dict[str, int])
b (dict[str, int])
normalization_factor (float | None)
- Return type:
float
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.jaccard(a, b)[source]¶
Set-based similarity between two event sets.
Returns 0.0 when both sets are empty (undefined ratio treated as no similarity rather than perfect similarity, which is the safer default for retrieval purposes).
- Parameters:
a (frozenset[str])
b (frozenset[str])
- Return type:
float
- src.dackar.RCA.log_pattern_recognition.rca_pattern_search.nlcs(a, b)[source]¶
Normalised Longest Common Subsequence similarity.
NLCS(A, B) = |LCS(A, B)| / max(|A|, |B|)
Operates on deduplicated ordered sequences so high-frequency events do not dominate the ordering signal.
Returns 0.0 when both sequences are empty.
- Parameters:
a (list[str])
b (list[str])
- Return type:
float
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.HistoricalSignalEpisode[source]¶
Public output type for PatternSearcher.search().
Represents a single historical signal episode retrieved for a query incident. Carries all three metric scores individually (§5 of the integration plan) and an index_status field that governs cross-pattern linkage eligibility (§4.11).
Sentinel (no_episodes_indexed) instances have episode_id == “” and similarity_to_current == 0.0. Callers must check index_status before attempting linkage.
- episode_id: str¶
- asset_id: str¶
- window_start: datetime.datetime | None¶
- window_end: datetime.datetime | None¶
- source_types: list[str]¶
- event_set: frozenset[str]¶
- event_seq: list[str]¶
- freq_vec: dict[str, int]¶
- similarity_to_current: float¶
- jaccard_score: float¶
- nlcs_score: float¶
- emd_score: float¶
- weight_profile: str¶
- matched_events: set[str]¶
- query_only_events: set[str]¶
- episode_only_events: set[str]¶
- episode_density: float¶
- known_rca: str | None¶
- linked_doc_ids: list[str]¶
- index_status: str¶
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.IncidentFingerprint[source]¶
Pre-computed similarity representations for a single incident or detected historical episode. This is the unit of comparison in the retrieval pipeline.
Derived from a list of UnifiedEvents by IncidentExtractor.extract() or EpisodeDetector after episode boundary assignment.
- The three representations serve distinct metrics:
event_set → Jaccard (what types occurred, ignoring order/repetition) event_seq → NLCS (what types occurred and in what order) freq_vec → EMD (how many times each type occurred)
High-frequency event types (count > freq_threshold) are excluded from event_set and event_seq but retained in freq_vec.
- episode_id: str¶
- asset_id: str¶
- window_start: datetime.datetime¶
- window_end: datetime.datetime¶
- density: float¶
- event_set: frozenset[str]¶
- event_seq: list[str]¶
- freq_vec: dict[str, int]¶
- known_rca: str | None = None¶
- source_types: list[str] = []¶
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.UnifiedEvent[source]¶
Canonical representation of a single event from any source.
All three input sources (alarm, SOE, anomaly) are normalised into this structure before any further processing.
- Lifecycle:
Created by IncidentExtractor.to_unified_events()
episode_id is None until EpisodeDetector assigns membership
timestamp_end is nullable and not used in current similarity metrics but carried for traceability and future use
- raw_id: str¶
- asset_id: str¶
- source: str¶
- event_type: str¶
- timestamp_start: datetime.datetime¶
- timestamp_end: datetime.datetime | None¶
- episode_id: str | None = None¶
- class src.dackar.RCA.log_pattern_recognition.rca_pattern_search.PatternSearcher(index, config)[source]¶
Retrieves the top-k most similar historical episodes for a query incident.
- Pipeline per search() call:
Inverted-index lookup: episode_ids sharing ≥ 1 event type with query.
Jaccard pre-filter: discard candidates below config.min_jaccard.
NLCS computation on survivors.
EMD computation on survivors.
Combined score: weighted sum using the resolved weight profile.
Rank descending by combined_score; return the top-k HistoricalSignalEpisodes.
The coarse-to-fine design avoids computing NLCS and EMD on clearly dissimilar episodes (those failing the Jaccard gate).
- Parameters:
- index¶
- config¶
- search(query, weight_profile=None, staleness_window_days=None)[source]¶
Retrieves top-k most similar historical episodes for a query fingerprint.
Returns list[HistoricalSignalEpisode] with index_status populated on every result (§4.11):
“indexed” — normal result from a current, populated index
“no_episodes_indexed” — index is empty; returns a single sentinel episode
- “stale” — index is older than staleness_window_days; results
returned but flagged; link_confidence capped downstream
- Parameters:
query (src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.IncidentFingerprint) – IncidentFingerprint for the query incident.
weight_profile (Optional[str]) – Weight profile override; None → config.weight_profile.
staleness_window_days (Optional[int]) – If set and index.build_timestamp is known, episodes from an index older than this are marked “stale”. None disables the staleness check. By design the window lives on PatternSearchConfig.index_staleness_window_days; the orchestrator reads it there and passes it in here (PatternSearcher itself only carries a SearchConfig).
- Return type:
list[src.dackar.RCA.log_pattern_recognition.rca_pattern_search.models.HistoricalSignalEpisode]
Notes
All three metric scores (jaccard, nlcs, emd) are individually visible on every returned HistoricalSignalEpisode (§5).
matched_events, query_only_events, episode_only_events are derived from event_set comparison.
- _resolve_weights(weight_profile)[source]¶
Returns (alpha, beta_w, gamma) for the given weight profile name.
Delegates to SearchConfig.resolve_weights() which owns the profile registry and “custom” fallback logic.
- Raises:
ValueError for unrecognised profile names. –
- Parameters:
weight_profile (str)
- Return type:
tuple[float, float, float]