Web Mining and Scientific Rigor
摘要
In this final chapter, we discuss how reproducible and replicable research is facilitated by automated data collection, as well as the pitfalls—like selection bias and representativity issues—that one must consider when using scraped web data. Finally, we outline best practices for organizing code and data for transparency and integrity, before concluding with an overview of sampling concerns and external validity.