<p>Semantic role identification is essential for understanding human activities by explicitly modelling the functional relationships between entities and actions. However, existing approaches predominantly rely on shallow syntactic cues or rule-based parsing, which limits their ability to capture rich semantic interactions and to generalize across diverse sentence structures. To overcome these limitations, this paper proposes RINet (Role Identification Network), a novel attention-based transformer framework specifically designed for semantic role identification in mutual human activity sentences. RINet leverages a lightweight, domain-adapted Bidirectional Encoder Representations from Transformers (BERT) architecture to generate contextualized token embeddings and directly map sentence components to semantically grounded roles such as agent, patient, giver, and receiver. To facilitate systematic evaluation of this newly defined task, we introduce RIT_3V, the first benchmark dataset for activity-centric semantic role identification, annotated across three syntactic variants and verb forms to capture linguistic diversity and semantic consistency. The new (RIT 3V) dataset is meticulously designed for role identification from text and comprises three versions that represent different sentence categories. Experimental results demonstrate that RINet consistently outperforms recurrent neural networks and standard transformer baselines, achieving Precision scores of 0.87, 0.82, and 0.78; Recall scores of 0.91, 0.85, and 0.81; F1-scores of 0.89, 0.83, and 0.79; and classification accuracies of 90.01%, 85.38%, and 80.20% across the V1, V2, and V3 dataset versions, respectively. By bridging the gap between syntactic pattern recognition and semantic understanding, this work establishes the first transformer-based benchmark for role identification in human activity descriptions and provides a scalable foundation for advanced activity reasoning and knowledge extraction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RINet: an attention-based transformer with the RIT-3V benchmark for semantic role identification in human activities

  • Vivek Tiwari,
  • Anam Arshad,
  • Mayank Lovanshi,
  • Rahul Shrivastava

摘要

Semantic role identification is essential for understanding human activities by explicitly modelling the functional relationships between entities and actions. However, existing approaches predominantly rely on shallow syntactic cues or rule-based parsing, which limits their ability to capture rich semantic interactions and to generalize across diverse sentence structures. To overcome these limitations, this paper proposes RINet (Role Identification Network), a novel attention-based transformer framework specifically designed for semantic role identification in mutual human activity sentences. RINet leverages a lightweight, domain-adapted Bidirectional Encoder Representations from Transformers (BERT) architecture to generate contextualized token embeddings and directly map sentence components to semantically grounded roles such as agent, patient, giver, and receiver. To facilitate systematic evaluation of this newly defined task, we introduce RIT_3V, the first benchmark dataset for activity-centric semantic role identification, annotated across three syntactic variants and verb forms to capture linguistic diversity and semantic consistency. The new (RIT 3V) dataset is meticulously designed for role identification from text and comprises three versions that represent different sentence categories. Experimental results demonstrate that RINet consistently outperforms recurrent neural networks and standard transformer baselines, achieving Precision scores of 0.87, 0.82, and 0.78; Recall scores of 0.91, 0.85, and 0.81; F1-scores of 0.89, 0.83, and 0.79; and classification accuracies of 90.01%, 85.38%, and 80.20% across the V1, V2, and V3 dataset versions, respectively. By bridging the gap between syntactic pattern recognition and semantic understanding, this work establishes the first transformer-based benchmark for role identification in human activity descriptions and provides a scalable foundation for advanced activity reasoning and knowledge extraction.