Łukasz Brzozowski, Marek Ga̧golewski, Grzegorz Siudem · Information Sciences 2026 · 2026
DOI: 10.1016/j.ins.2026.124196
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Generating realistic synthetic citation, patent, or component dependency networks is essential for benchmarking community detection, graph visualisation, and network data mining algorithms. We present a systematic comparison of generators of directed graphs that are nearly acyclic and have a ground-truth community structure. We evaluate 12 methods across 7 real citation networks and 26 metrics. We propose the practice of reversing directions of edges in static generators to break cycles and induce a citation-like flow, which markedly improves the performance of a degree-corrected Stochastic Block Model. Our methodological approach to evaluating community detection benchmarks distinguishes between endogenous and exogenous mesoscopic similarities, with the latter proving more important. This distinction reveals that high-parameter models gain much of their advantage by memorising planted community statistics, which limits their fidelity when the evaluation is blind to those labels. Finally, we introduce the Citation Seeder (CS) algorithm, an iterative generator grounded in the Price-Pareto model of citation networks, with interpretable parameters and generation. CS achieves competitive results against the best-performing baselines while using up to four orders of magnitude fewer parameters, providing an interpretable account of a network’s structure and, on a temporal hold-out, extrapolating its future in-degree distribution more faithfully than descriptive baselines.
No comments yet — start the discussion below.