src.dackar.RCA.pm_compliance.vocabulary_loader

PMVocabularyLoader — load health-status vocabulary from the DACKAR data directory.

Drives analyze_degradation and compute_pm_found_defect_rate in effectiveness_analyzer. Falls back to hardcoded stems when data_dir is None or the CSV files are absent, so callers without the data path are unaffected.

Attributes

_COLS

_NEG_FILE

_POS_FILE

_NEU_FILE

_FALLBACK

Classes

PMVocabularyLoader

Load and cache health-status vocabulary keyed by data_dir path.

Functions

_read_terms(path, cols)

Return lowercased, stripped terms from cols of a keyword CSV.

_term_matches(term, blob)

Return True if term appears in blob at a leading word boundary.

matches_any(blob, terms)

Return True if any term in terms matches within blob.

Module Contents

src.dackar.RCA.pm_compliance.vocabulary_loader._COLS: Sequence[str] = ('Nouns', 'Verbs', 'Adjectives')[source]
src.dackar.RCA.pm_compliance.vocabulary_loader._NEG_FILE = 'health_status_keywords_negative.csv'[source]
src.dackar.RCA.pm_compliance.vocabulary_loader._POS_FILE = 'health_status_keywords_positive.csv'[source]
src.dackar.RCA.pm_compliance.vocabulary_loader._NEU_FILE = 'health_status_keywords_neutral.csv'[source]
src.dackar.RCA.pm_compliance.vocabulary_loader._FALLBACK: Dict[str, FrozenSet[str]][source]
src.dackar.RCA.pm_compliance.vocabulary_loader._read_terms(path, cols)[source]

Return lowercased, stripped terms from cols of a keyword CSV.

Parameters:
  • path (pathlib.Path)

  • cols (Sequence[str])

Return type:

FrozenSet[str]

src.dackar.RCA.pm_compliance.vocabulary_loader._term_matches(term, blob)[source]

Return True if term appears in blob at a leading word boundary.

Multi-word phrases (containing spaces) are matched as substrings. Single-word terms use a left-side word boundary so that the vocabulary term is matched as a prefix of a word — “crack” matches “crack”, “cracks”, “cracked”, “cracking” and “leak” matches “leakage”, but neither matches in the middle of an unrelated word (e.g., “crack” does not match “firecracker” because ‘c’ in “cracker” has no leading word boundary).

Parameters:
  • term (str)

  • blob (str)

Return type:

bool

src.dackar.RCA.pm_compliance.vocabulary_loader.matches_any(blob, terms)[source]

Return True if any term in terms matches within blob.

Parameters:
  • blob (str)

  • terms (FrozenSet[str])

Return type:

bool

class src.dackar.RCA.pm_compliance.vocabulary_loader.PMVocabularyLoader[source]

Load and cache health-status vocabulary keyed by data_dir path.

Usage:

vocab = PMVocabularyLoader.load(cfg.data_dir)
is_degrading = matches_any(blob, vocab["degrading"])
_cache: Dict[str, Dict[str, FrozenSet[str]]][source]
classmethod load(data_dir)[source]

Return {"degrading": ..., "improving": ...} vocabulary sets.

When data_dir is None, returns the hardcoded fallback stems so existing behaviour is fully preserved.

Parameters:

data_dir (Optional[pathlib.Path])

Return type:

Dict[str, FrozenSet[str]]

classmethod clear_cache()[source]

Invalidate the vocabulary cache (useful in tests).

Return type:

None