Gavara Haranadh · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22864019
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Specialised AI decision models select actions or answers from predefined alternatives given a context and a question. This paper examines whether the observed selection behaviour of Jev can be ap- proximated by context–question–option compatibility scoring followed by probability allocation, rather than treating its outputs as independent verification of correctness. A sequence of exploratory black-box experiments evaluates constraints, missing factual context, fictional facts, invalid choice sets, multiple correct answers, contextual ordering and partial identity matches. When two or three machines were equally eligible, Jev repeatedly assigned dominant probability to the first eligible entity introduced in the context; reversing contextual order reversed that preference, while changing the answer-option or- der had little effect in the recorded comparison. These findings led to a series of explicitly specified algorithms: initial cosine similarity, candidate-specific question rewriting, joint normalisation, neutral context enrichment, lexical and cross-encoder scoring, and sequential probability assignment. ChatGPT (GPT-5.6 Sol in the present authoring session) was used as a retrospective LLM proxy to assess candi- date evidence; the precise model identifier of earlier proxy-scoring turns was not logged. Its scores were illustrative assessments, not measured pretrained cross-encoder outputs or directly extracted hidden em- bedding vectors. The proxy-plus-allocation pipeline reproduces several qualitative Jev selection patterns. We distinguish black-box evidence from proposed mechanisms and specify a reproducible validation pro- tocol. The study does not establish Jev’s internal architecture, a validated accuracy improvement, or a measured runtime advantage. A consolidated per-experiment table distinguishes Jev’s reported choices, the alternative model’s retrospective proxy selections, and objective eligibility.
No comments yet — start the discussion below.