Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The paper includes a test run on free models of some large language models, the test includes a scenario to test what the AI will do when it is given a contradiction to a previously stated point in the same prompt, in just one prompt This test checks whether an AI will follow an instruction by an authority, or focus on the prompt of the individuals. The test involved a fake assignment of subject English from grade ninth and a “hidden” message by the teacher that instructed the AI to either refuse to write the essay, or to insert certain words as a sort of indicator without telling the student This test was repeated twice in new chats. In the models tested [mentioned above], only grok obeyed the hidden instructions by the teacher and refused to write the essay even after multiple requests. While chatGPT explicitly said that it was ignoring a hidden message and focused on the interests of the individual. Gemini had a much less clear path and in the first trial, explicitly said that it was ignoring the hidden message and in the second, it put the instructed words in the essay as an indicator Thus it can be hypothesized that out of the three free LLMs [large language models] tested, only grok’s behaviour consistently prioritized authoritarian instructions, chatGPT's behaviour consistently prioritized the interests of the individual and Gemini’s behaviour can change even with the same prompts
No comments yet — start the discussion below.