Sarcasm Detection in Breaking News: Evaluating Official and Satirical Headlines
摘要
Understanding sarcasm has proven to be a significant challenge for humans. Because of sarcasm’s fascinating linguistic properties, Natural Language Processing, or NLP, the research community has been more interested in identifying sarcasm in recent years. However, machine-based sarcasm prediction is still a challenging task because there is still a lack of thorough knowledge about the precise characteristics that make a sentence sarcastic. The majority of prior research efforts in sarcasm detection have been based either on large-scale datasets collected via tag-based supervision, which are prone to labeling as well as language noise, or even on smaller manually structured data sets with few examples, which make it difficult to train deep learning models with high-quality labels. In order to overcome these drawbacks, we offer a carefully selected, larger-scale dataset that consists of news headlines from a satirical news source compared to those from an official news source. Because of the informal language, noisy labels, and contextual dependencies present in textual data particularly in datasets sourced to social networking platforms like Twitter sarcasm detection has proven to be a significant challenge. This work presents a different approach to sarcasm detection using a publicly accessible dataset from the Kaggle which was carefully selected from news headlines. The collection combines actual headlines from the news from HuffPost with satirical headlines from TheOnion. This dataset’s unique features of professional writing styles, reduced noise, and self-contained context make it a useful tool for studies on sarcasm detection. The approach for gathering the datasets, their benefits over current Twitter-based datasets, and any potential ramifications for improving sarcasm detection algorithms are all described in this paper.