Pretrained models may fail to capture immunological sequences
摘要
Pretrained models, originally developed for vision and textual data, are not a panacea and may fail to fully represent the complexity of sequences in immunological tasks. In studying pretrained immunological sequence modules of a renowned immunogenicity prediction model, pMTnet, we observe that our carefully designed model removing (ablating) the pretrained T cell receptor (TCR) autoencoder in pMTnet can even improve the prediction accuracy. Furthermore, we note the TCR pretraining data, used by pretrained modules within pMTnet, dramatically deviates from a broader and more representative TCR repertoire. Such findings underscore the impacts of the blend of heterogeneous representations and distribution discrepancy in immunological sequences, which necessitate appropriate coordination of different pretrained models and representative databases.