Agus Siswanto, Bambang Tutuko, Firdaus, Jasmir · International journal of intelligent engineering and systems 2026 · 2026
DOI: 10.22266/ijies2026.1031.01
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Biomedical Named Entity Recognition (BioNER) is essential for extracting clinical entities from biomedical text, yet its performance in few-shot settings is constrained by limited annotated data and severe token-level class imbalance.Existing studies typically address these challenges separately through either data augmentation or sampling strategies.This paper proposes IBUA-BioNER, a unified framework that jointly tackles both issues by integrating improved Balanced Undersampling (iBUS) and Adaptive Mention-Replacement (ADAMER).iBUS reduces redundant non-entity tokens while preserving informative contextual information surrounding entities.ADAMER adaptively generates augmented samples according to the imbalance level of each sentence, producing more samples for highly imbalanced instances and fewer for relatively balanced ones.Experiments using BioBERT on NCBI-Disease, BC5CDR-Chemical, and BC2GM under 2%, 5%, and 10% few-shot settings showed that combining iBUS with ADAMER improved the F1-score over the no-augmentation baseline in eight of the nine corpus and few-shot combinations, with the largest gain of 3.40 percentage points on NCBI-Disease at the 5% setting.Aggregated across the nine settings, the proposed framework outperformed all three COSINER-based configurations in every setting (Wilcoxon signed-rank p = 0.0039), whereas its advantage over the strongest competing configuration within an individual setting was generally not statistically significant.
No comments yet — start the discussion below.