Jiajun Ma · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22853597
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Public comment submitted in a personal capacity to NIST (TEVV-Athlon@nist.gov) on 19 September 2026 on NIST AI 200-2 ipd, an initial public draft of the TEVV-Athlon framework for evaluating AI systems. The comment makes one primary recommendation, with replacement text, for Section 2.4 — a minimum reporting summary for readers who did not conduct the evaluation, stating which Blocks were not measured, what human judges were shown before judging, and who conducted the assessment relative to the system's developer — and one secondary recommendation for Section 2.1, that an evaluation declare the adoption or reliance stage it is designed for, on a published scale, with design proportionate to it. The corpus at doi:10.5281/zenodo.22247885 is offered as an existing four-stage scale and as a worked example of instrumented human verification. Use of AI assistants in preparing the comment is disclosed in the document.
No comments yet — start the discussion below.