Hanif Sajid · · 2026
DOI: 10.31222/osf.io/ja7bs_v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In this paper, I leverage three automatic prompt optimization (APO) algorithms to search for optimal prompts that enable large language models to classify policy documents accurately, without requiring task-specific training data. I evaluate the performance of these optimized prompts using OpenAI’s GPT-4o-mini model, comparing them against prompts created by domain experts as a baseline. Experimental results show that all three algorithms significantly outperform the baseline. Among the algorithms, APE, which induces simple prompts from input-output example pairs, significantly outperforms GRIPS and performs comparably to ProTeGi. Additionally, ProTeGi demonstrates similar performance to GRIPS. However, neither the optimized prompts nor the human-designed ones achieve a research-grade performance (defined as an F1-score of 70%) with the GPT-4o-mini model. To reach research-grade performance in policy document classification using APO methods, future work could explore the use of more powerful models, hybrid APO systems, cross-model evaluations, and inference-time enhancements.
No comments yet — start the discussion below.