Pranay M. Mahendrakar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23049431
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. In 2023 a large language model was reported to solve text-based analogy problems zero-shot at or above the level of college students. Two critiques followed. They showed that performance on letter-string analogies collapses when the alphabet is permuted or replaced by symbols, while human performance does not. The original authors replied that the failures come from an auxiliary difficulty with counting, and that GPT-4 solves the permuted problems at a human level once it can write and execute code. This paper reads all three rounds in full, including the preprint and published versions of the reply, and sorts their claims into three groups. Settled: the replications agree, and the unaided model fails the counterfactual variants under answer-only prompts. Moved: the evidence. The third round's decisive result is for a model-plus-interpreter system, while its title claim remains about language models. Unsettled: which operations count as auxiliary, and whether the models represent the new alphabet at all. Published checks from both sides, run in different studies with different alphabets, suggest splitting "counting" into two operations: GPT-4 names the one-step successor of a letter in a permuted alphabet almost perfectly, but identifies the interval (of up to two steps) between two given letters about one time in ten. A study of four newer models, however, finds every model worse at naming items two steps away, and its authors conclude that the models do not build representations of novel alphabets on the fly. If the difficulty lies in recognising the relation in the source pair, it lies in what componential theories of analogy call inference; if it lies in representing the order, it lies in encoding, which the same theories also count as part of analogy. On that decomposition, either way, the reply's "auxiliary" label is not yet earned. The paper states a procedure for reading capacity claims from counterfactual tasks, in which the decomposition of the task is fixed before data are seen, and proposes five experiments on existing materials that could decide between the readings. Every number in it is taken, or summed, from the published sources.
No comments yet — start the discussion below.