This study focused on exploring different patterns, testing different hypotheses, and establishing relationships between salaries, skill sets, experience, and other dimensions across diverse data science job roles. The dataset used in this study is referred from the Kaggle data source. The dataset’s attributes include expertise area, employment type, expertise level, company location across the continents, company size, and the dependent variable salary class. Exploratory data analysis and data visualization were used to understand the different patterns, different hypotheses were established, and statistical tests were incorporated to identify the statistical significance of the different types of relationships. A machine learning technique called Catboost regressor was used to boost the gradient over decision trees. The focus was to leverage tree-based ensemble techniques (Catboost) to ensure the great ability of this algorithm to handle outliers and optimize trees with the right pruning. One of the great advantages of this approach is to ensure that it captures nonlinear relationships (if any), which was not leveraged in previous studies or other research, and the emphasis was on traditional equation-based regression methods capturing linear relationships. The current study determines which variable has the maximum influence on the dependent variable (salary grades) for the different data science-related job roles, thus helping to create the optimum strategy in capacity building for the diverse organizations in the areas of data science-related skill sets and job roles. The insights from this study can also be leveraged by job aspirants to create their own optimal strategy for their growth plan in data science. The study can also use to estimate the probable salary slab for the different types of data science-related job roles along with the other factors (experience, expertise area, employment type, expertise level, company location across the continents, and company size), thus helping in optimal human resource planning and budgeting.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Demystifying Patterns and Trends Across Different Data Science Skills and Job Roles: An Analysis of Key Determinants Using Machine Learning

  • Avisek Kundu,
  • Nitesh Dhar Badgayan,
  • Seeboli Ghosh Kundu,
  • Geetha V. Sharma

摘要

This study focused on exploring different patterns, testing different hypotheses, and establishing relationships between salaries, skill sets, experience, and other dimensions across diverse data science job roles. The dataset used in this study is referred from the Kaggle data source. The dataset’s attributes include expertise area, employment type, expertise level, company location across the continents, company size, and the dependent variable salary class. Exploratory data analysis and data visualization were used to understand the different patterns, different hypotheses were established, and statistical tests were incorporated to identify the statistical significance of the different types of relationships. A machine learning technique called Catboost regressor was used to boost the gradient over decision trees. The focus was to leverage tree-based ensemble techniques (Catboost) to ensure the great ability of this algorithm to handle outliers and optimize trees with the right pruning. One of the great advantages of this approach is to ensure that it captures nonlinear relationships (if any), which was not leveraged in previous studies or other research, and the emphasis was on traditional equation-based regression methods capturing linear relationships. The current study determines which variable has the maximum influence on the dependent variable (salary grades) for the different data science-related job roles, thus helping to create the optimum strategy in capacity building for the diverse organizations in the areas of data science-related skill sets and job roles. The insights from this study can also be leveraged by job aspirants to create their own optimal strategy for their growth plan in data science. The study can also use to estimate the probable salary slab for the different types of data science-related job roles along with the other factors (experience, expertise area, employment type, expertise level, company location across the continents, and company size), thus helping in optimal human resource planning and budgeting.