Over the past decade, we can see a notable increase in the use of multimodal posts (text \(+\) images) on social media platforms. In such multimodal posts, images themselves possess the property of multimodality, which means they exhibit different kinds of content (textual/non-textual/both). In the past, limited research has been devoted to extracting the sentiment of users from multimodal images. The main objective of this paper is to categorize multimodal social media political images into distinct classes based on the content they are displaying and subsequently extract the user’s political views from the images. This research proposes a hybrid approach named textual and non-textual image sentiment analysis (TaNTISA). TaNTISA (TaNT \(+\) ISA) is a combination of two tasks: (1) textual and non-textual image classification (TaNT) and (2) image sentiment analysis (ISA). TaNTISA contains three modules that perform the sentiment analysis of multimodal images. The first module is the transfer-learning-based textual and non-textual image classifier named TaNT. TaNT is the main module of this research work, which classifies an image into textual, non-textual, or combined image classes with an accuracy of 94%. The second and third modules of TaNTISA extract the content (text and face) from different types of images and analyze their sentiment, respectively.