Performance of LLMs on Computing Systems for Deployment in IoT Devices
摘要
In this study, the authors explore the performance of different Large Language Models such as BART-Base, GPT Neo and DistilGPT-2 on hardware devices. These models are fine-tuned on a general dataset and tested on systems with various computing capabilities, from high-end servers and cloud infrastructures to more resource-constrained embedded devices. The main objective is to determine how fast a model can handle the input when given, the precision of text summarisation and the similarity between the machine-generated translation and the reference translations. The novelty of this research lies in finding the compromise between the speed of processing and the precision in generating the output. This approach aims to determine which model and system performs the best for future deployment in Internet of Things (IoT) devices.