Plamen Milev, B. Bahov · Big Data and Cognitive Computing 2026 · 2026
DOI: 10.3390/bdcc10090314
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Proactive novel topic detection in text streams lacks standardized evaluation: systems are tested on different data, under different definitions of novelty, with no shared ground truth. The goal of this study is a standardized, reproducible way to evaluate proactive topic detection under identical conditions. We introduce the Synthetic Fire Drill, a benchmark methodology that injects articles about fictitious topics, generated by large language models, into real news streams under controlled semantic distance, volume, and timing. Since the topics are fictitious, every detection outcome is classifiable against complete ground truth, and a validation battery separates topical detection from generation artifacts. We demonstrate the benchmark on a broad general news corpus and a focused science-technology corpus, with two generators, four injection schedules, and 16 scenarios spanning a calibrated distance gradient, evaluating four detection approaches (the micro-cluster stream clusterers DBSTREAM and DenStream, the windowed topic-modeling detector BERTrend, and a cosine novelty detector) under two sentence encoders. The experiment reveals that every stream-clustering detection, across all runs, seeds, and encoders, occurs through absorption into pre-existing clusters, never through the formation of new ones, with the absorbing cluster’s topical relatedness depending on the corpus, not the detector; only the windowed paradigm produces an emergence signal. The approaches prove complementary along mechanism lines, with each covering difficulty regions the others miss. Finally, the benchmark’s ground truth demonstrates that the detection rate alone is a misleading metric, requiring a computable quality dimension like enrichment to evaluate true system performance. The findings’ scope is bounded by two embedding-model families, two corpora from different eras, and 16 scenarios. The synthetic articles, embeddings, ground truth, evaluation protocol, and refresh recipe are publicly released.
No comments yet — start the discussion below.