Victoria Shi, Ari Holtzman, Peter West · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.2190.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Common intuitions suggest that the largest available model is always best, but what about when one LLM predicts or stands in for another? We present scale affinity: the observation that models of a similar size are often best at matching another model’s behavior, meaning bigger is not always better. We first establish this distributionally: testing the ability to predict LLM generations across up to 23 models in 5 data domains, we find the best predictor of a given model’s text is most often a similarly-sized model within the same family, and the trend against pure scale holds across family boundaries. The effect is not an artifact of likelihood-based measurement: similarly-sized models also reproduce each other’s errors on multiple-choice benchmarks beyond what matched accuracy explains, and smaller in-family models are better adversarial predictors of another model’s moves. Scale can eventually help with behavior matching, but typically only at much larger sizes and with more observed context than similarly-sized models require, suggesting that scale affinity and scaling laws (bigger is better) are opposing factors that trade off to determine how effectively one LLM can predict another. Scale affinity has significant implications when using models to stand in for other systems, suggesting that matching capability to the target can be a better choice than simply using the strongest LLM available.
No comments yet — start the discussion below.