Text-based measurement of regional economic information
摘要
This paper proposes a methodology to extract information about regional economic conditions from newspaper text in real time. The approach relies on large-scale collections of news articles that are summarized using unsupervised machine learning to generate topics capturing recurring themes in economic reporting. Because the method uses the full corpus of regional news and avoids restrictive keyword selection, it minimizes human judgment and allows the data to reveal economically relevant patterns in news coverage. I apply the methodology to Canada, a large and economically diverse country, and show that the resulting topics contain information about fluctuations in economic indicators such as manufacturing activity and unemployment at both the national and provincial levels. The results are robust to alternative choices of the number of topics. A composite index constructed from the topic measures provides a summary indicator of economic information contained in the news. While some topics display similar associations with economic outcomes across provinces, others capture region-specific developments, highlighting the ability of the approach to uncover geographically heterogeneous economic signals in news data.