Rohith Reddy Bellibatlu, Pratyush Jena, Morla Triveni, Edward Raff, Wenbin Zhang · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.19584052
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Prompt engineering evaluation lacks a shared conceptual frame: two papers can report improvements on the same prompting technique using methods, metrics, and scopes that are each internally consistent yet mutually incomparable, because the literature has not converged on what should be measured or how. This survey provides a measurement-theoretic taxonomy of prompt-engineering evaluation, tracing the incomparability to two long-standing tensions in measurement theory: construct validity vs. measurement reliability, and reference-anchored vs. judgment-anchored evaluation. We propose a seven-dimensional taxonomy of the evaluation choices a study makes: how outputs are scored, on what task, with which metric, how automated the pipeline is, over how many models and datasets, whether rephrasing the prompt was tested, and whether a third party could reproduce the result. We apply it to a corpus of 151 papers assembled by a PRISMA 2020 systematic review. The coding falls into three recurring configurations, which we name Benchmark-Automation, Judge-Mediated, and Expert-Anchored, and which are largely a compact summary of how a study scores its outputs. We report two negative results about them: the dimensions the archetype rule does not read carry no independent trace of the partition, and the archetypes are not associated with held-out robustness or reproducibility practice once a structural dependency is removed. The same coding reveals three recurring failure modes, each a configuration that does not fully support the claim the paper makes with it; 45.0% of the corpus exhibits at least one. We ask readers to use the failure-mode catalog and the companion seven-question design checklist, which are the parts that survive this scrutiny intact. Benchmark-Automation papers (66.2%) live in the high-reliability, narrow-validity region: fine for ground-truth tasks, under-specified for open-ended generation. Judge-Mediated work (17.2%) addresses the construct gap, but the corpus has not converged on a shared judge-calibration protocol. Robustness and contamination-resistant evaluation remain minority practices in the corpus. Reporting a result's sensitivity to rephrasing, and the exact version of any LLM judge, should become default requirements rather than optional extras.
No comments yet — start the discussion below.