Clustering in genomic data is critical for understanding genetic variation, population structure, and disease associations. The high-dimensional and complex nature of genomic data frequently results in suboptimal accuracy from traditional clustering algorithms, whether or not dimensionality reduction techniques are used. The potential of a genome-specific LLM and the embeddings it produces for clustering genomic data are examined in this work. Using three benchmark genome datasets, we test the effectiveness of this method to determine whether LLM-based embeddings can improve clustering accuracy. Our results indicate that while the fine-tuned embeddings considerably increase accuracy in just two datasets, the pre-trained embeddings perform similarly to conventional clustering techniques. These findings demonstrate the possibility of optimizing genome-specific LLMs to improve clustering results, but they also underline the necessity of more research to fully utilize LLM embeddings for the analysis of genomic data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Use of Large Language Models to Cluster Genomic Data

  • Reem Al-Saidi,
  • Ziad Kobti,
  • Thorsten Strufe

摘要

Clustering in genomic data is critical for understanding genetic variation, population structure, and disease associations. The high-dimensional and complex nature of genomic data frequently results in suboptimal accuracy from traditional clustering algorithms, whether or not dimensionality reduction techniques are used. The potential of a genome-specific LLM and the embeddings it produces for clustering genomic data are examined in this work. Using three benchmark genome datasets, we test the effectiveness of this method to determine whether LLM-based embeddings can improve clustering accuracy. Our results indicate that while the fine-tuned embeddings considerably increase accuracy in just two datasets, the pre-trained embeddings perform similarly to conventional clustering techniques. These findings demonstrate the possibility of optimizing genome-specific LLMs to improve clustering results, but they also underline the necessity of more research to fully utilize LLM embeddings for the analysis of genomic data.