Yidong Sun, Dongxu Liu, Jiale Zhang, Youcheng Wang · Computers 2026 · 2026
DOI: 10.3390/computers15090576
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Tibetan-to-English machine translation (MT) models frequently falter under extreme domain data scarcity, often producing translations that violate the distinctive agglutinative rules of Tibetan and suffer from domain-specific stylistic mismatches. To overcome these limitations, we propose Semantic-Syntax Prealignment (SSPA), an innovative corpus generation framework. SSPA constructs high-quality pseudo-parallel pairs by explicitly minimizing the deviation between the syntactic-semantic profiles of generated samples and professional reference texts. Specifically, source-target structural representations are standardized through length-unified truncation and terminology normalization, followed by a dual-domain alignment process that maximizes syntactic cosine similarity under rigorous structural constraints. We further augment these aligned frames via a cross-length dynamic filling mechanism, which is integrated with an Expectation-over-Transformation (EOT)-based style regularization mechanism specifically adapted for stylistic perturbations, to simulate authentic linguistic variations. Extensive evaluations on our newly constructed Tibetan Medicine-Tibetan English (TM-TE) dataset demonstrate that SSPA significantly outperforms existing competitive baselines. Notably, SSPA achieves a BLEU-4 score of 36.2 and improves long-sentence BLEU-4 by 16.8 points, with a parser-verified grammatical compliance rate of 96.2%. The framework exhibits remarkable cross-domain adaptability and stylistic consistency, offering a robust, versatile solution for low-resource Tibetan professional domain MT.
No comments yet — start the discussion below.