Research on clinical trials requires substantial background and technical knowledge. Large language models (LLMs) have already made a significant impact in various fields. We attempted to use general-purpose LLMs to assist newcomers in clinical trials, enabling them to quickly begin their work. In our work, we demonstrated that in the domain of clinical trials, the current cutting-edge LLMs can provide excellent recommendations for feature selection. By utilizing the features suggested by LLMs, we achieved a 2.5% improvement in AUC compared to complex neural network models when using simpler algorithms. We have also demonstrated that by adjusting the prompts, LLM can play a significant role in the feature extraction process. By adjusting the prompts for certain features suggested by LLM, LLM-assisted feature extraction achieved 100% accuracy in a random sample covering approximately 10% of the entire dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Facilitating Feature Selection and Extraction in Clinical Trials with Large Language Models

  • Jiaji Guo,
  • Wen Sun,
  • Shiting Wen,
  • Di Wu,
  • Yipeng Zhou

摘要

Research on clinical trials requires substantial background and technical knowledge. Large language models (LLMs) have already made a significant impact in various fields. We attempted to use general-purpose LLMs to assist newcomers in clinical trials, enabling them to quickly begin their work. In our work, we demonstrated that in the domain of clinical trials, the current cutting-edge LLMs can provide excellent recommendations for feature selection. By utilizing the features suggested by LLMs, we achieved a 2.5% improvement in AUC compared to complex neural network models when using simpler algorithms. We have also demonstrated that by adjusting the prompts, LLM can play a significant role in the feature extraction process. By adjusting the prompts for certain features suggested by LLM, LLM-assisted feature extraction achieved 100% accuracy in a random sample covering approximately 10% of the entire dataset.