Comparative Study of Multimodel Architecture in Generative AI
摘要
In this paper, various Kolmogorov-Arnold Network (KAN), namely Smooth KAN (S-KAN), Wavelength KAN (WAV-KAN), Basic KAN (B-KAN), and Temporal KAN (T-KAN) models, is compared. The comparison is done based upon how they are processing the multimodal data. These models are the variation of KAN model either by incorporating the functional or structural changes to enhance their scalability and efficiency in accessing complex multimodal data. Wavelength KAN used the strength of wavelet-based concepts, which helped to access the low- and high-frequency data components to enhance the model. To reduce the overfitting, Smooth KAN utilizes the functional transformations. To address temporal dependencies and sequential flow of data, KAN is modified for time-sensitive information as T-KAN. The datasets used to compare the performance of these models are image, audio, and time-series data. The performance is compared in terms of accuracy, interpretability, computational efficiency, and generalizability and provide the strengths and limitations of each KAN model. The analysis made in this paper can be utilized in the future to develop KAN-based models.