Estimate Initial and Final Data Bounding for Big Data Progressive Sampling in Parallel Computing Environment
摘要
Big Data Progressive Sampling requires initial and final data bounding values to generate optimal number of samples in order to train any learning algorithm. Any learning algorithm can be trained for minimal hypothesis space. Approximating data bounding to Rademacher Averages of a given hypothesis space and learning algorithm ensure final data bounds with ε-Approximation. This final Data bounding value serves as stopping criterion for Progressive Sampling. Statistically optimal sample size estimation technique ensure minimal initial sample size. In this work, we propose a methodology to approximate initial and final Data bounding to Progressive Sampling using local Rademacher in big data Hadoop environment with experimental values along with computational parameters. Also the performance of learning algorithm to various highly complex big datasets for their corresponding initial minimal value and final Data bound to Rademacher Averages are experimentally shown in Hadoop Distributed environment.