Ashish Prasad, Saurav Kumar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22917683
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper presents a two-phase training framework that combines Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) to enhance reasoning and IoT control capabilities in small language models. The proposed approach enables a reasoning-enhanced model to interpret real-time sensor data and control actuators on resource-constrained edge devices. The trained model is deployed on a Raspberry Pi 4 using GGUF quantization and Ollama, demonstrating a fully local, cloud-independent sense-reason-act loop for Physical AI and intelligent IoT applications.
No comments yet — start the discussion below.