Large Language Models (LLMs) are increasingly adopted in edge applications, with notable advantages compared to cloud-based approaches, such as enhanced privacy, reduced latency, and improved energy efficiency. While deploying pre-trained models at the edge is common, training or fine-tuning them locally poses significant challenges due to data remaining on-device in distributed, heterogeneous environments and the ones related to computational constraints and communication overhead. In this context, a well-established solution is Federated Learning (FL), where clients train models locally, without sharing their data, and a global model is obtained by aggregating their parameters, weighted by the number of examples. This method may be inadequate for LLMs, as it overlooks the varying information content of training examples. In fact, longer sequences often contain more informative structures, offering richer learning signals. To investigate this issue, we evaluate the training performance of LLMs in a federated learning setting using two aggregation methods: standard FedAvg and a token-based variant that weights updates based on the number of tokens processed locally. We conduct experiments using lightweight LLMs, specifically SmolLM2, comparing performance using different open-source datasets from the healthcare field. Experimental results demonstrate that token-based FedAvg reaches the performance of standard FedAvg and, in some cases, slightly surpasses it.
Investigating Edge Fine-Tuning of Large Language Models in a Federated Environment
Lorenzo Colombi
;Michela Vespa;Francesco Resca;Edoardo Di Caro;Elena Bellodi;Mauro Tortonesi;Cesare Stefanelli
2025
Abstract
Large Language Models (LLMs) are increasingly adopted in edge applications, with notable advantages compared to cloud-based approaches, such as enhanced privacy, reduced latency, and improved energy efficiency. While deploying pre-trained models at the edge is common, training or fine-tuning them locally poses significant challenges due to data remaining on-device in distributed, heterogeneous environments and the ones related to computational constraints and communication overhead. In this context, a well-established solution is Federated Learning (FL), where clients train models locally, without sharing their data, and a global model is obtained by aggregating their parameters, weighted by the number of examples. This method may be inadequate for LLMs, as it overlooks the varying information content of training examples. In fact, longer sequences often contain more informative structures, offering richer learning signals. To investigate this issue, we evaluate the training performance of LLMs in a federated learning setting using two aggregation methods: standard FedAvg and a token-based variant that weights updates based on the number of tokens processed locally. We conduct experiments using lightweight LLMs, specifically SmolLM2, comparing performance using different open-source datasets from the healthcare field. Experimental results demonstrate that token-based FedAvg reaches the performance of standard FedAvg and, in some cases, slightly surpasses it.I documenti in SFERA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


