Shubham Waghmare · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22886297
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Despite the widespread adoption of supervised fine-tuning (SFT) and direct preference optimization (DPO), little is known about how internal representations evolve across the complete language model training lifecycle. We conduct a controlled longitudinal analysis of internal representations across 25 checkpoints spanning pretraining, SFT, and DPO within the same language model. We pretrain a 334M-parameter decoder-only transformer from scratch on 12.4B FineWeb-Edu tokens, then analyze checkpoints spanning subsequent SFT on Tulu 3 and DPO using UltraFeedback and HH-RLHF. We quantify representation similarity using Centered Kernel Alignment (CKA) and Representational Similarity Analysis (RSA) across all 25 layers, observing substantial representational change during SFT concentrated in middle layers (CKA dropping to 0.53), whereas DPO induces minimal representational drift (CKA > 0.998) despite successful preference optimization. These trends are consistent across both preference datasets and both similarity metrics. Our findings provide empirical evidence that SFT and DPO influence pretrained representations in fundamentally different ways, offering new insight into how alignment reshapes language models after pretraining.
No comments yet — start the discussion below.