The Impact of Alternative Data on Default Probability: Analyzing the Italian E-commerce Sector with NLP and Network Structures
摘要
E-commerce is a key sector in the Italian economy, with online companies becoming some of the largest and most profitable businesses. However, this growth comes with increased risk exposure. This study aims to investigate the relationship between alternative data (contextual factors, Text-Driven Data Enrichment) and the probability of default for Italian e-commerce companies. To date, no studies have examined how these alternative data affect the default probability within this sector. To address this gap, the ongoing research analyzes a dataset of Italian companies, focusing on sector-specific indicators. In the Italian e-commerce market, companies are identified by a unique code that indicates the sales channel but not the types of products marketed. To overcome this limitation, a natural language processing (NLP) model is applied to the es of these companies, allowing us to identify the types of goods or services sold by each company. These enriched data provide a more comprehensive understanding and serve as the foundation for a classification model that uses standard interpretable algorithms for default prediction. Additionally, network structures identified by companies’ similarities are used to extract new insights, which can support decision-making processes for stakeholders at different levels of the supply chain. The model’s performance is evaluated in various scenarios, comparing results with and without the inclusion of alternative data. Key performance metrics are analyzed to demonstrate how integrating alternative data enhances default prediction models.