A proposal is introduced for National Statistical Offices (NSOs) to produce representative information, on multiple topics, more frequently, jointly using household survey data and social network (SN) posts. SN posts by respondents are tagged using survey responses to train machine learning (ML) algorithms. Trained algorithms tag large volumes of recent posts. Those by the same author are tallied to tag users and form a non-random large sample for statistical production. Further monitoring is carried out by tagging posts dated between survey rounds from new “large samples.” To deal with selection bias in user populations, sociodemographic (SD) variables collected during the survey are used to develop SD-tagged large sample databases of recent authors. They will be SD post-stratified during thematic studies to mitigate selection bias. Suggestions to link user-respondent survey responses and network posts are explored. An author can be tagged according to many surveys, leading to a multi-topic large-sample, to explore interactions rarely studied.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Little Bird Told Me: Blueprint Toward Frequent and Representative Official Statistical Information, Through Joint Use of Social Network Posts and Survey Data

  • Alfredo Bustos,
  • Silvia Fraustro,
  • Noemí López,
  • Ricardo Olvera

摘要

A proposal is introduced for National Statistical Offices (NSOs) to produce representative information, on multiple topics, more frequently, jointly using household survey data and social network (SN) posts. SN posts by respondents are tagged using survey responses to train machine learning (ML) algorithms. Trained algorithms tag large volumes of recent posts. Those by the same author are tallied to tag users and form a non-random large sample for statistical production. Further monitoring is carried out by tagging posts dated between survey rounds from new “large samples.” To deal with selection bias in user populations, sociodemographic (SD) variables collected during the survey are used to develop SD-tagged large sample databases of recent authors. They will be SD post-stratified during thematic studies to mitigate selection bias. Suggestions to link user-respondent survey responses and network posts are explored. An author can be tagged according to many surveys, leading to a multi-topic large-sample, to explore interactions rarely studied.