Victoria Akano, Olufade F. W. Onifade, Nancy Fugate Woods · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23169768
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automated English-to-American Sign Language (ASL) translation has transformative potential to overcome systemic communication and instructional barriers faced by Deaf and Hard-of-Hearing (DHH) learners in Science, Technology, Engineering, and Mathematics (STEM) education. However, evaluating Sign Language Translation (SLT) systems faces an acute methodological crisis: standard machine translation metrics (such as BLEU, ROUGE, and chrF) rely strictly on n-gram co-occurrence against human-annotated reference translations. In specialized technical domains, bilingual parallel corpora are virtually non-existent, and the absence of standardized written sign orthography causes n-gram metrics to penalize linguistically authentic sign glosses while rewarding unnatural word-for-word Signed Exact English (SEE). To resolve this fundamental bottleneck, this paper presents ASL-GEMS (ASL Gloss Evaluation Metrics System), an open-source, reference-free, linguistically grounded evaluation framework specifically designed for English-to-ASL machine translation pipelines. ASL-GEMS mathematically formalizes canonical ASL morphosyntactic regularities across three decoupled dimensions: (1) Structural Completeness (C), measuring semantic fidelity by verifying the retention of meaning-bearing content lemmas; (2) Grammar Consistency (G), penalizing English syntactic intrusions, including copulas, articles, non-fronted temporal indicators, and misplaced Wh-question words; and (3) Fluency Score (F), which models natural visual signing pacing through an empirically motivated target compression ratio (R_comp in [0.40, 0.90]). The framework also integrates automated Out-of-Vocabulary (OOV) tracking to assess domain-boundary fallbacks (fingerspelling) and incorporates an edit-distance Word Error Rate (WER) module for speech-driven multimodal translation pipelines. We evaluate ASL-GEMS on a curated 35-sentence benchmark corpus categorized into five structural typologies: simple declarative, interrogative (Wh-), temporal, out-of-vocabulary, and complex multi-clause sentences. Empirical benchmarking reveals that while a word-for-word SEE baseline suffers severe grammatical degradation (Grammar Consistency: 0.8043, Fluency: 0.8521, with persistent article retention penalties averaging 0.1329), the proposed ASL syntactic restructuring engine achieves perfect Structural Completeness (1.0000), flawless Grammar Consistency (1.0000), and a superior Fluency score of 0.9916, operating at a natural compression ratio of 0.8249. Paired non-parametric statistical significance testing confirms that these improvements are highly significant (Wilcoxon signed-rank W = 0.0, p < 0.0001 for grammar; W = 6.0, p < 0.0001 for fluency). Crucially, ASL-GEMS identified an empirical OOV rate of 51.4%, accurately directing unrecognized technical jargon to fingerspelling. We release the ASL-GEMS evaluation suite and benchmark corpus as an open-source standard to support reproducible, reference-free assessment in low-resource sign language processing. Keywords: Sign Language Machine Translation, Reference-Free Evaluation, ASL-GEMS, Topic-Comment Grammar, STEM Accessibility, Out-of-Vocabulary Fallback.
No comments yet — start the discussion below.