Large language models (LLMs) adeptly navigate the intricate dependencies between sentences and texts, accurately predicting subsequent words or phrases. Such capabilities underscore LLMs’ supre-macy in sequential information processing and their profound understanding of complex sequence patterns. Consistently, we posit that LLMs can more effectively discern fine-grained temporal features and interpret long-term behavioral patterns in scenarios, which requires sophisticated temporal modeling for accurate long-term action anticipation (LTA). In this study, we introduce a novel methodology, termed Adapting Large Language Model for Long-Term Action Anticipation (LLMAction), which incorporates a distinctive action adapter module. The module enables the training of the LLM with minimal parameters, yet fully harnesses its robust sequence modeling prowess for LTA. To substantiate our hypothesis, comprehensive experiments are conducted on two benchmark datasets: Breakfast and 50 Salads. Our method achieves superior results to state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLMAction: Adapting Large Language Model for Long-Term Action Anticipation

  • Binglu Wang,
  • Yao Tian,
  • Changhe Wang,
  • Le Yang

摘要

Large language models (LLMs) adeptly navigate the intricate dependencies between sentences and texts, accurately predicting subsequent words or phrases. Such capabilities underscore LLMs’ supre-macy in sequential information processing and their profound understanding of complex sequence patterns. Consistently, we posit that LLMs can more effectively discern fine-grained temporal features and interpret long-term behavioral patterns in scenarios, which requires sophisticated temporal modeling for accurate long-term action anticipation (LTA). In this study, we introduce a novel methodology, termed Adapting Large Language Model for Long-Term Action Anticipation (LLMAction), which incorporates a distinctive action adapter module. The module enables the training of the LLM with minimal parameters, yet fully harnesses its robust sequence modeling prowess for LTA. To substantiate our hypothesis, comprehensive experiments are conducted on two benchmark datasets: Breakfast and 50 Salads. Our method achieves superior results to state-of-the-art approaches.