Huy Quoc To, Q Tran, H Zhang, Guangyan Huang, Ming Liu · Deakin Research Online (Deakin University) 2026 · 2026
DOI: 10.26187/deakin.34007652
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Long-form scientific writing remains a significant challenge for large language models (LLMs), requiring coherent, well-structured, and citation-aware outputs over thousands of words. Existing benchmarks mainly focus on short-to-long tasks such as summarization, leaving long-to-long scientific generation underexplored. In this paper, we introduce LongSciGen, a benchmark for long-form scientific text generation that combines human-written survey papers with high-quality synthetic section-level data. This dual design enables both realistic document-level evaluation and scalable section-wise training. We further propose LongSciEval, a unified evaluation framework that assesses generation quality at both the section and document levels in terms of structure, coverage, and citation quality. We evaluate 13 state-of-the-art LLMs across different output lengths and training settings. Results reveal clear performance differences across model scales and highlight the impact of output length on generation quality. Our work provides a practical benchmark, a comprehensive evaluation framework, and actionable insights for designing long-form scientific text generation systems.
No comments yet — start the discussion below.