Analyzing Data Science Labor Market Trends in Poland Using NLP Techniques
摘要
This paper responds to the need for in-depth analyzis of trends in data science and the requirements of data science professionals in the long term. The primary goal of the research is to analyze trends within the Polish data science job market. The study covers job openings in Poland from 2015 to 2023, containing more than one million job advertisements. This extensive dataset, which spans nine years, allows us to present changes and trends in the demand for data science specialists in the Polish labor market. This study utilized a variety of data analyzis tools and techniques at various stages of the data workflow: (1) data acquisition—Beautiful Soup and Apache Spark were used to extract data from the website; (2) data cleaning and preparation—Pandas, NLTK, and Spacy were used to clean and prepare the data for analyzis; (3) exploratory analyzis—LLM Mistral, Pandas and Matplotlib were used to conduct exploratory analyzis of the data. Understanding this evolution is critical not only for data science professionals, but also for educators who develop relevant curricula, policymakers who design skills training programs and HR departments that deal with talent acquisition in the dynamic labor market.