TrinityLabo · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22767805
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Avatar-Bounded AI Evaluation proposes a containment-oriented framework for evaluating advanced AI systems without granting the underlying models unrestricted access to external tools, networks, filesystems, or execution environments. The central idea is to separate model intelligence from external agency. Each model is coupled to a constrained virtual embodiment, or avatar, whose sensory and action capabilities are explicitly bounded. All interactions with the evaluation environment are mediated through this avatar and an independent policy gate. The architecture therefore aims to preserve a fixed external action envelope even as the cognitive capability of the underlying model increases. The framework is first defined for single-agent evaluation and then extended to multi-agent examination arenas, where multiple AI systems can be placed inside the same closed environment. Communication between agents is restricted to observable channels, enabling controlled study of cooperation, collusion, rule violations, deceptive coordination, whistleblowing, supervision sensitivity, and emergent group behavior. Formally, the work introduces an avatar-mediated causal structure in which no direct model-to-environment action edge is permitted. The proposed design emphasizes complete mediation, bounded action spaces, controlled communication, and explicit separation between intelligence growth and externally available authority. The manuscript also outlines experimental protocols for testing whether advanced AI systems attempt unauthorized coordination, exploit evaluation rules, behave differently under observation, or cooperate with adversarial agents. This work is a theoretical research manuscript. It does not report deployed-agent experiments, human-subject studies, or empirical claims of demonstrated safety. Its contribution is the specification of a containment and evaluation architecture, formal properties of the proposed boundary, testable hypotheses, and a structured research program for single- and multi-agent AI evaluation. Keywords: AI safety; AI containment; multi-agent systems; AI evaluation; embodied AI; agentic AI; sandboxing; AI control; collusion; evaluation integrity; frontier AI; constrained agency. Author: TrinityLaboVersion: 1.0Status: Theoretical research manuscript; not peer reviewed.
No comments yet — start the discussion below.