<p>The construction of control groups of scientists is often a daunting effort. This paper presents <Emphasis FontCategory="NonProportional">sosia</Emphasis>, an open-source Python-based software designed to&#xa0;efficiently query the Scopus database via RESTful API. <Emphasis FontCategory="NonProportional">sosia</Emphasis> searches for researchers with publication profiles similar to a given researcher up to a given year based on all main standard bibliometric indicators. The user can choose flexibly a set of parameters to restrict the search to more or less narrow boundaries upfront and obtain additional similarity indicators to select a subset of authors after the search. Advanced settings also allow narrowing the search to a list of affiliations and to minimize the possible errors arising from ambiguous author profiles. One basic search can be set up in a few command lines and the average time of computation goes between 60 and 300&#xa0;minutes. We discuss the functioning, characteristics, limitations and possible extension of the software.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Finding Doppelgängers in Scopus: how to build scientists control groups using sosia

  • Michael E. Rose,
  • Stefano H. Baruffaldi

摘要

The construction of control groups of scientists is often a daunting effort. This paper presents sosia, an open-source Python-based software designed to efficiently query the Scopus database via RESTful API. sosia searches for researchers with publication profiles similar to a given researcher up to a given year based on all main standard bibliometric indicators. The user can choose flexibly a set of parameters to restrict the search to more or less narrow boundaries upfront and obtain additional similarity indicators to select a subset of authors after the search. Advanced settings also allow narrowing the search to a list of affiliations and to minimize the possible errors arising from ambiguous author profiles. One basic search can be set up in a few command lines and the average time of computation goes between 60 and 300 minutes. We discuss the functioning, characteristics, limitations and possible extension of the software.