
Wangfan Li, Sofia Abilene Campos Hernandez, Carlos Toxtli · Proceedings of the Human Factors and Ergonomics Society Annual Meeting 2026 · 2026
DOI: 10.1177/10711813261475162
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Researchers increasingly use large language models (LLMs) to revise manuscripts in response to peer reviews, yet this adoption is largely unprincipled and risks the ironies of automation. Prompt-Level Supervisory Alignment (PLSA) applies Supervisory Control Theory as a structured prompting strategy for LLM-assisted manuscript revision. We extend PLSA’s planning function to multi-round peer review, where the planner must anticipate concerns that later reviewers will raise. We construct RevPlan-Bench, a corpus of multi-round manuscripts whose ground truth is the issues expert reviewers raised across three or more review cycles, and score 23,256 revision plans spanning five information conditions, three LLM backbones, and four prompting variants. First-round reviews substantially improve coverage of future concerns; multi-agent debate degrades performance; and explicit anticipation prompting, our central pre-registered hypothesis, adds no practical value beyond the reviews themselves, an informative null. Information design is the dominant lever; the researcher remains the final arbiter of scholarly claims.
No comments yet — start the discussion below.