Orthographically-informed ASR evaluation


source

Alignment

def Alignment(
    reference:tuple[str, ...], hypothesis:tuple[str, ...], operations:tuple[str, ...]
)->None:

An edit-distance alignment and its operation counts.


source

normalize_for_evaluation

def normalize_for_evaluation(
    text:str
)->str:

Apply the NFC and whitespace normalization recommended for ASR evaluation.


source

llm_wer

def llm_wer(
    reference:str, hypothesis:str, judge:Callable[[str, str], bool]
)->float:

Return Sarvam-style LLM-WER using an application-supplied conservative judge.

The judge receives each non-exact aligned span and must return True only when it is certainly semantically and phonetically equivalent. The final numerator is capped at the reference length, preventing repeated hallucinations from dominating aggregate results.


source

oiwer

def oiwer(
    hypothesis:str, reference_variations:Sequence[Sequence[str]], reference:Optional[str]=None
)->float:

Return Orthographically-Informed WER as a fraction.

reference fixes the denominator to the original transcript, as in an OI benchmark. When omitted, the first alternative in every segment is treated as that original transcript.


source

oiwer_alignment

def oiwer_alignment(
    hypothesis:str, reference_variations:Sequence[Sequence[str]]
)->Alignment:

Align a hypothesis with ordered OI reference segments.

Each inner sequence contains valid alternatives for one reference span. An alternative may contain several words, which supports compound splitting, merging, acronyms, and inverse-text-normalization variants.


source

cer

def cer(
    reference:str, hypothesis:str
)->float:

Return Character Error Rate as a fraction, after Unicode NFC normalization.


source

wer

def wer(
    reference:str, hypothesis:str
)->float:

Return conventional WER as a fraction, after Unicode NFC normalization.

# CER basic tests
assert cer("", "") == 0.0
assert cer("hello", "hello") == 0.0
# one substitution in a 3-char string
assert abs(cer("abc", "axc") - 1/3) < 1e-9
# full deletion
assert cer("hi", "") == 1.0
# CER is more lenient than WER for single-char typos
assert cer("hello world", "helo world") < wer("hello world", "helo world")
# NFC normalisation applies
assert cer("caf\u00e9", "cafe\u0301") == 0.0
print("CER tests passed")
assert wer('वह डॉक्टर के पास गया', 'वह doctor के पास गया') == 0.2
variations = [['वह'], ['डॉक्टर', 'doctor'], ['के'], ['पास'], ['गया']]
assert oiwer('वह doctor के पास गया', variations) == 0.0
assert oiwer('मुझे 56849 चाहिए', [['मुझे'], ['five six eight four nine', '56849'], ['चाहिए']], 'मुझे five six eight four nine चाहिए') == 0.0
assert oiwer('one two three four', [['one two', '12'], ['three four', '34']]) == 0.0
assert llm_wer('नहीं', 'नहीं नहीं नहीं नहीं', lambda _ref, _hyp: False) == 1.0
assert llm_wer('वह डॉक्टर के पास गया', 'वह doctor के पास गया', lambda ref, hyp: {ref, hyp} == {'डॉक्टर', 'doctor'}) == 0.0