Innovating Supervised and Semi Supervised Learning: Reducing Label Dependency, Enhancing Data Efficiency and Developing Scalable, Adaptive Models for Real Time Applications
摘要
This research work aims to discuss the difficulties and the progress made in extracting data using machine learning algorithms, more specifically those of the supervised and semi-supervised types. Although, these methods play a crucial role in converting unstructured data into useful information, they come across certain challenges such as over-dependency on a large number of labelled datasets, which are time-consuming and expensive to create. The quality of data holds the noise and imbalance factors that obstruct the model from working properly and lead to biases. Besides, these models lack generalization capabilities and are unfit when applied to different datasets thus requiring constant reviewing and tuning. Semi-supervised learning seems to be a viable solution for these two issues through the use of labelled and unlabeled data, but it complicates the model accuracy and handles irrelevant data. To establish these problems, this research employs a convergent parallel mixed research design that employs quantitative, as well as qualitative, data collection techniques. Outcomes involve revising criteria of data labelling, improving feature selection, and hyperparameter tuning. The research also examines approaches to deal with big data and design algorithms that are robust and able to learn in real time in an environment that is constantly changing. In order to fill these gaps and support the on-going development of machine learning methodologies, this research seeks to add to the body of knowledge in this fields, and thus the continuing improvement of data extraction systems within different contexts. It is hoped that the results will be useful to practitioners and researchers in presenting better working practices for employing ML methods for the best extraction of data, safely and efficiently.