Guorui Chen, Yifan Xia, Xinliang Ma, Z. Merrick Li · Symmetry 2026 · 2026
DOI: 10.3390/sym18101644
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Hidden-state probes are widely used to study safety-related behavior in language and multimodal models, yet their results can change with the data sources, behavior labels, extracted layer, and probe family. We introduce TriSAIL, an evaluation protocol that records these choices and tests how they affect the resulting claim. TriSAIL uses three response-derived states—benign, non-refusal jailbreak, and borderline/refusal—and assigns train, validation, and test data by source. We extract hidden states from the final input position in the first generation forward pass, before a generated token is returned. The operating layer is selected using validation data only. We evaluate the protocol on five text-only LLMs and six MLLMs with logistic, kNN, and SVM probes. The LLM results vary substantially across held-out attack families, probe families, and source assignments. The MLLM probes achieve high label separability, while source-ID decoding, prediction–source association, and residualization tests show strong source alignment in the same representations. Because the main MLLM matrix couples sources and labels, these results describe source-sensitive separability. Source-transfer conclusions use source-held-out columns; mixed accuracy is a descriptive overall summary. TriSAIL reports these results with probe sensitivity and source diagnostics so that each score is interpreted under the conditions that produced it.
No comments yet — start the discussion below.