Gavara Haranadh · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22979838
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Structured neural network pruning physically deletes entire channels and intermediate fea- ture projections, delivering immediate memory and computational savings without requiring specialized sparse hardware accelerators. However, dominant channel selection heuristics— predominantly magnitude-based (weight norms) or local activation statistics (mean acti- vation, variance)—treat channels as isolated scalar components, ignoring high-dimensional inter-channel covariance structures. Consequently, they either prune channels carrying criti- cal independent task information or retain clusters of mutually redundant features. Further- more, conventional post-pruning recovery relies either on computationally expensive iterative fine-tuning or rigid linear weight merging that ignores nonlinear activation dynamics. This paper presents CGAS-NC (Covariance-Guided Activation Selection with Nonlinear Compensation), an end-to-end framework for structural channel pruning and representation recovery. CGAS-NC introduces: (1) a mathematically unified Contribution-Aware Covari- ance Metric that identifies redundant donor channels by penalizing mutual correlation while explicitly weighting downstream task sensitivity; (2) dynamic layer-budgeting mechanisms based on the participation-ratio Effective Rank (¯ reff) and layer redundancy mass; and (3) an Adaptive Recovery Hierarchy spanning zero-overhead linear weight folding, orthogonal Chebyshev polynomial compensation, and lightweight basis surrogates. Crucially, CGAS-NC executes true structural pruning by physically contracting weight matrices and eliminating intermediate dimensions, rather than relying on simulated zero- masking. Furthermore, while legacy structural pruning methods (such as Network Slim- ming) depend fundamentally on Batch Normalization scaling parameters or assume simple piecewise-linear activations (like ReLU), they are intrinsically inapplicable to modern Trans- formers and Large Language Models (LLMs) that employ LayerNorm or RMSNorm and smooth or gated activations (GELU, SwiGLU, SiLU). In contrast, because CGAS-NC op- erates strictly on empirical activation covariances and downstream projection sensitivity, it is fundamentally activation-agnostic and normalization-independent. Consequently, CGAS- NC provides a deployable structural pruning mechanism that is also theoretically and prac- tically suitable for integration into LLM training and fine-tuning pipelines. Empirically, this work provides an exhaustive evaluation spanning over 60 distinct con- figurations across multiple benchmark datasets and architectural paradigms. Across multi- seed training-time structural pruning experiments on feed-forward networks (MNIST and Fashion-MNIST across 5 random seeds), CGAS-NC achieves up to 50% channel reduction with < 0.3% accuracy degradation. On post-training structural pruning of pre-trained Large Language Models (GPT-2 124M on WikiText-2), dynamic Effective-Rank allocation prunes 56% more transformer feed-forward channels than uniform pruning at stable token perplex- ity, while closed-form ridge projection compensation recovers perplexity from 61.08 to 36.96 (+24.12 PPL). Furthermore, while this paper presents completed empirical benchmarks on multi-layer networks and GPT-2, active investigations are currently underway extending 1CGAS-NC to deep residual convolutional networks (ResNet-18, ResNet-50) and compact modern foundation language models (such as LLaMA-3.2 and Gemma-2)
No comments yet — start the discussion below.