Kartikey Singh · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23049559
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
AS-ViT (Adaptive Subspace Vision Transformer) introduces an architectural bridge between Scientific Machine Learning (SciML) parameter decoupling and dense multi-task visual perception. Multi-task learning backbones routinely suffer from negative transfer caused by conflicting task gradient trajectories across shared transformer weights. Rather than relying on computationally heavy post-hoc gradient surgery (such as PCGrad or CAGrad), AS-ViT integrates continuous Partition of Unity (PoU) gating and autonomous parameter cleavage—inspired by Adaptive Mesh Refinement (AMR) in PDEs—directly into Vision Transformer feed-forward layers. Key Contributions:• Autonomous Cleavage: Vectorized Jacobian autodiff profiles gradient conflict online. When chronic negative transfer is detected, shared feed-forward layers dynamically cleave into task-dedicated expert subspaces without human re-architecting.• Partition of Unity (PoU) Routing: Implements compact-support weighting matrices satisfying a strict unity sum, ensuring zero router collapse and provably bounded backpropagation variance.• NYUv2 Benchmark Superiority: Delivers an overall multi-task gain of ΔM = +5.84% over monolithic ViT-B/16 across simultaneous 13-class semantic segmentation (mIoU: 44.82%), monocular metric depth (RMSE: 0.512 m), and 3D surface normals (MAE: 18.24°).• Inference Efficiency: Achieves a 2.22× throughput speedup (92.4 FPS vs. 41.6 FPS on an NVIDIA Tesla T4) compared to static 4-expert MoE architectures via dynamic sparse expert gating.• In-the-Wild Generalization: Demonstrates robust zero-shot multi-task inference on uncalibrated smartphone imagery across rural terrain and dynamic outdoor lighting. Codebase & Master Notebooks:https://github.com/KartikeyaGangwar/as-vit-multitask Preprint Lineage (SciML Foundations):This work represents the computer vision extension of the adaptive subspace framework established in:1. Exact Boundary-Preserving & Null-Space Constrained PINNs (DOI: 10.5281/zenodo.22132799)2. AS-PINN: Adaptive Subspace Physics-Informed Neural Networks (DOI: 10.5281/zenodo.22822521)
No comments yet — start the discussion below.