Mahran Jazi, Ilai Bistritz, Nicholas Bambos, Irad Ben-Gal · Information Sciences 2026 · 2026
DOI: 10.1016/j.ins.2026.124093
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Collaboration between edge devices has the potential to scale up machine learning (ML) by enabling access to unprecedented amounts of data. Federated learning (FL) is a collaborative algorithm in which clients learn from each other without sharing private data. However, edge devices tend to have different data distributions because they are naturally exposed to different data sources. This heterogeneity, also known as non-independent and identically distributed (non-IID) data, has been shown to decrease the accuracy of FL. We study how limited data sharing among users can alleviate this performance degradation. Such sharing can occur naturally on a social graph or be incentivized by the platform. We evaluated the performance gains of data sharing on MNIST, CIFAR-10, and CIFAR-100 across topologies including complete graphs, clusters, and stochastic block models. We empirically demonstrate that modest data sharing between neighbors on a social graph enhances learning performance in the non-IID case. Interestingly, we found that data sharing could also improve performance in the IID case. By normalizing the dataset sizes, we verified that this boost is significant even if sharing does not increase the number of data points per client. Therefore, data sharing is an efficient technique for improving social federated learning (SFL).
No comments yet — start the discussion below.