Prakhar Gautam, Jitenda Singh Thakur, Ashish Mishra · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23076774
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Subject–verb agreement (SVA) errors are frequent in learner and machine-generated English, and minimal-pair benchmarksshow that language models (LMs) still confuse the agreement controller with an intervening noun. We test whether anexplicit, parse-grounded checker is a stronger detector than LM scoring. The proposed Parse-grounded AgreementConsistency Checker (PACC) locates the subject of every finite verb in a dependency parse, reads the verb’s number from itssurface form through a morphological lexicon, and flags conflicts. Because off-the-shelf parsers are trained on well-formedtext and often misread the very verb that carries the error, PACC adds agreement-neutralised re-parsing (ANR): when themain clause has no analysable finite verb, candidate verbs are replaced one at a time by their number-neutral past-tense form,the sentence is re-parsed, and the original form is checked against the recovered subject. We evaluate on the six English SVAparadigms of BLiMP (6,000 minimal pairs), split once into 1,200 development pairs used for rule design and 4,800 held-outtest pairs scored in a single run. On the test pairs, PACC with the medium spaCy parser selected on development data reaches96.5% pairwise accuracy (95% bootstrap CI 96.0–96.9), against 86.6% (85.7–87.6) for the released GPT-2 scores of theBLiMP authors and 52.0% for a Kneser–Ney 5-gram trained here on 6.0 million tokens. The largest gain is on attractors insiderelative clauses (89.4% vs 65.9% for GPT-2). A post-hoc run with a transformer-based parser raises accuracy to 98.1% andrelative-clause accuracy to 98.0%, at 51 rather than 670 sentences per second. Used as a stand-alone detector on 9,600 singlesentences, PACC reaches F1 = 96.3% at a 2.96% false-positive rate, while per-word LM scores barely separate acceptablefrom unacceptable sentences (ROC-AUC 0.578 for GPT-2). Ablations show that reading number from the surface form andANR contribute most. Manual inspection of the remaining errors attributes them to attractor attachment, gaps in irregularplural morphology, and at least 17 test pairs whose “acceptable” member itself contains an agreement error.
No comments yet — start the discussion below.