Junyi Tao · Stanford Digital Repository 2026 · 2026
DOI: 10.25740/jk225ym4523
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Explaining the behavior of a complex system--in particular, a trained neural network--has been a central interest of interpretability research, and more broadly of cognitive science and philosophy of science. It is also the topic of this thesis, which collects two papers I co-authored during my master's program. What makes an explanation a good one? One can always describe a model's behavior in a thin way, by listing out the full input-output mapping. Such a description, though perfectly faithful, does not seem to be a good explanation: it does not tell us how the behavior is causally generated; it is hardly cognitively tractable for us; moreover, there is no hope that it can help us predict how the model will behave beyond the already observed cases. A good explanation, by contrast, should do better on each of these aspects: it should tell a causal story; it should do so compactly enough for us to grasp; ideally it should let us anticipate behavior we have not yet seen. I pursue explanations with these virtues throughout this thesis. It is worth distinguishing three increasingly demanding explanatory goals, which the two projects of this thesis can be situated in: what algorithm a model implements, how generalizable that algorithm is, and why it was learned at all.
No comments yet — start the discussion below.