Using public data to evaluate global dynamics in biology publishing and the human microbiome A DISSERTATION SUBMITTED TO THE FACULTY OF THE UNIVERSITY OF MINNESOTA BY Richard John Abdill III IN PARTIAL FULFILLMENT OF THE REQUIREMENTS FOR THE DEGREE OF DOCTOR OF PHILOSOPHY Advisor: Ran Blekhman February 2022 Copyright 2022 Richard John Abdill III i Acknowledgements Any consideration of research done at the University of Minnesota is incomplete without consideration of the unique circumstances of land-grant institutions, and I can’t ignore that the resources provided to me here owe much to the dispossession of the indigenous peoples who lived and continue to live in the place now known as Minneapolis. The campus where I studied is situated on Dakhóta land, personally stolen by Zebulon Pike in an 1805 treaty, the first in a century full of broken promises (Čhaŋtémaza and McKay 2020). The university continues to benefit from an endowment seeded by tens of thousands of acres of land seized throughout the 19th century (R. Lee et al. 2020), some of which is maintained by the state even today (“Treasury & Endowments” n.d.). I offer this land acknowledgement in recognition of my time spent in “the land where the waters reflect the clouds,” the home of Native peoples from time immemorial, and I encourage you to learn more about the original stewards of this place and how we can work toward redressing these injustices. As for my work specifically, I did not complete it alone, and I am indebted to many people who supported me over the last four years. I am especially grateful to my advisor, Dr. Ran Blekhman, whose kindness, exacting standards and unflagging patience has made me the researcher I am today. I would also like to express my gratitude to my committee members, Drs. Chad Myers, Frank Albert and R. Stephanie Huang, whose generous feedback has guided me through the final years of my project. ii I also thank Dr. Beth Adamowicz and Samantha Graham, with whom I collaborated on several of the chapters presented here. I would also like to thank Dr. Sambhawa Priya, Dr. Casey Greene, Dr. Dan Knights, and Lena Morrill Gavarró for discussions that were foundational in developing my research plan. I am also grateful to Dr. Laura Grieneisen and the other members of the Blekhman lab, who played a huge part in my development as a scientist and became my close friends. I also thank Dr. Serghei Mangul, Dr. Humberto Debat, Clarissa F.D. Carniero and Dr. Olavo B. Amaral, from whom I learned a great deal in our collaborations. This work also would not have happened without Frank Hanson, who sat me in front of a Linux terminal in 2013 and told me to get to work. Frank’s guidance and trust made me a programmer—and forever changed the course of my life. I’m grateful to my parents, Richard and Nancy Abdill, who always supported my interests, even when it meant dealing with a 10-year-old whose favorite hobby was spreadsheets. I must also thank Kevin LaCherra, Rob Gindes, John Banusiewicz, Nick Werner and Dr. Kevin Hemer—I count myself very lucky to call them friends, but please, no one tell them I said so. I also thank my son, Abram, who was born in between chapters two and three of this dissertation and provided all the joy I could need to get me through the late nights of work. Most importantly, I thank my wife, Lauren, whose companionship has been the lodestar of my journey to, and through, graduate school. Her inspiration, kindness and ceaseless support is the reason I’m here at all. iii This work is dedicated to Lauren Abdill. She deserves far more, but this is the best I’ve got. iv Abstract The last decade has seen a rapid growth of research interest in the human microbiome, the bacterial communities that live on and in human bodies. Complex dynamics in the world of academic publishing have changed over almost the same period, during which time it has become increasingly common for researchers in the life sciences to share their manuscripts online as “preprints,” making them freely available prior to conventional peer review. Both are analyzed here by using publicly available data to quantify trends, patterns, and risks developing in both systems. First, I evaluate the growth of preprints posted to bioRxiv.org in its first five years. I describe the process for compiling the first comprehensive characterization of biology preprints and evaluate patterns in readership and outcomes for more than 37,000 manuscripts. Secondly, I examine changing trends in country-level preprint authorship and find patterns of international collaboration that suggest the supposedly egalitarian practice of preprinting may be replicating some of the same entrenched power dynamics found in conventional publishing. Third, I apply a similar approach to metadata for publicly available human microbiome samples: I characterize the origins of more than 440,000 samples and use country-level research activity to determine which countries are overrepresented in the characterization of the human microbiome relative to their population. Finally, I describe how we used this metadata in the construction of the largest human gut microbiome dataset in the world, consisting of more than 170,000 samples in one unified dataset that can be used to characterize broad patterns in microbial ecological dynamics. v Table of Contents Acknowledgements ........................................................................................................................................... i Abstract ........................................................................................................................................................... iv Table of Contents ............................................................................................................................................. v List of Tables ................................................................................................................................................. vii List of Figures .............................................................................................................................................. viii Introduction ...................................................................................................................................................... 1 Preprints and publishing in biology ............................................................................................................ 2 The human microbiome ............................................................................................................................... 3 Chapter One: Tracking the popularity and outcomes of all bioRxiv preprints ................................................ 7 Summary ...................................................................................................................................................... 8 Background .................................................................................................................................................. 8 Results ........................................................................................................................................................ 10 Preprint submissions ............................................................................................................................. 11 Preprint downloads ............................................................................................................................... 13 Preprint authors ..................................................................................................................................... 15 Publication outcomes ............................................................................................................................ 17 Discussion .................................................................................................................................................. 26 Methods ..................................................................................................................................................... 31 Data availability ........................................................................................................................................ 44 Chapter Two: International authorship and collaboration across bioRxiv preprints ...................................... 45 Summary .................................................................................................................................................... 46 Background ................................................................................................................................................ 46 Results ........................................................................................................................................................ 48 Country-level bioRxiv participation over time ..................................................................................... 48 Preprint adoption relative to overall scientific output .......................................................................... 51 Patterns and imbalances in international collaboration ........................................................................ 53 Differences in preprint downloads and publication rates ..................................................................... 58 Preprint publication patterns between countries and journals .............................................................. 60 Discussion .................................................................................................................................................. 62 Methods ..................................................................................................................................................... 67 Data availability ........................................................................................................................................ 77 vi Chapter Three: Public human microbiome data are dominated by highly developed countries .................... 78 Summary .................................................................................................................................................... 79 Background ................................................................................................................................................ 79 Results ........................................................................................................................................................ 81 Samples per country ............................................................................................................................. 82 Samples per country relative to population .......................................................................................... 87 Discussion .................................................................................................................................................. 89 Methods ..................................................................................................................................................... 96 Data availability ...................................................................................................................................... 100 Chapter Four: Compendium of 170,000 uniformly processed 16S rRNA human gut microbiome samples from public repositories ................................................................................................................................ 101 Summary .................................................................................................................................................. 102 Background .............................................................................................................................................. 103 Results ...................................................................................................................................................... 105 Sample composition ............................................................................................................................ 106 Patterns of variation across the compendium ..................................................................................... 110 Methods ................................................................................................................................................... 112 Discussion ................................................................................................................................................ 117 Data availability ...................................................................................................................................... 118 Chapter Five: Discussion .............................................................................................................................. 119 Summary of results .................................................................................................................................. 119 Future work: bibliometrics ...................................................................................................................... 121 Future work: microbiome compendium .................................................................................................. 122 Bibliography ................................................................................................................................................. 130 Appendix A: Chapter One Supplementary Material .................................................................................... 145 Appendix B: Chapter Two Supplementary Material .................................................................................... 152 Appendix C: Chapter Three Supplementary Material .................................................................................. 157 vii List of Tables Chapter One Table 1-1. Unique preprint authors per year. .............................................................................................. 16 Table 1-2. Median downloads by publication status. .................................................................................. 23 Chapter Two Table 2-1. Preprints per country. ................................................................................................................. 50 Chapter Three Table 3-1. Samples per country. .................................................................................................................. 84 Table 3-2. Samples by body site. ................................................................................................................ 86 Table 3-3. Samples and population by region. ............................................................................................ 88 Appendix A Supplementary Table A-1. Top 15 institutions by author count. .............................................................. 150 Supplementary Table A-2. Total preprints published per journal. ............................................................ 151 viii List of Figures Chapter One Figure 1-1. Preprints over time. ...................................................................................................................... 12 Figure 1-2. Distribution of all recorded downloads of bioRxiv preprints. ..................................................... 14 Figure 1-3. Characteristics of the bioRxiv preprints published in journals. ................................................... 18 Figure 1-4. Journals that have published the most preprints. ......................................................................... 20 Figure 1-5. Preprint downloads by publishing journal. .................................................................................. 22 Figure 1-6. Time to publication by journal. ................................................................................................... 25 Chapter Two Figure 2-1. Preprints per country. .................................................................................................................. 49 Figure 2-2. BioRxiv adoption per country. .................................................................................................... 52 Figure 2-3. Contributor countries. .................................................................................................................. 55 Figure 2-4. Preprint outcomes. ....................................................................................................................... 59 Figure 2-5. Overrepresentation of US preprints. ............................................................................................ 61 Chapter Three Figure 3-1. Global microbiome representation. ............................................................................................. 83 Chapter Four Figure 4-1. Sample pipeline progress. .......................................................................................................... 106 Figure 4-2. Most prevalent bacteria at three taxonomic levels. ................................................................... 108 Figure 4-3. Sample composition. ................................................................................................................. 109 Figure 4-4. Principal coordinates analysis. .................................................................................................. 111 Appendix A Supplementary Figure A-1. Downloads per preprint by months available. ................................................. 145 Supplementary Figure A-2. Proportion of downloads per preprint by months available. ........................... 146 Supplementary Figure A-3. Annual publication rates and estimates. .......................................................... 147 Supplementary Figure A-4. Multiple perspectives on per-preprint download counts. ................................ 148 Supplementary Figure A-5. Total downloads per preprint from each year. ................................................. 149 Appendix B Supplementary Figure B-1. Preprint collaboration. ..................................................................................... 154 Supplementary Figure B-2. Correlation between three measurements of international collaboration. ....... 155 Supplementary Figure B-3. International collaboration correlations. .......................................................... 156 Appendix C Supplementary Figure C-1. Samples per year. ............................................................................................. 158 1 Introduction Uranus had a wobble. Its path didn’t make sense. The planet had been discovered in 1781, and by the mid-1800s it had moved far enough through its orbit around the Sun that astronomers were starting to notice that the numbers didn’t add up: Given what they knew about the planet and its position in the solar system, it frequently popped up in places it shouldn’t have been. Sometimes it was too far along in its orbit, then later it wasn’t quite far enough. Two compelling theories emerged: Either Isaac Newton’s theories were wrong, or there was another planet out there, with gravity strong enough to alter its neighbors’ movement through space. Years passed without any developments, but French astronomer Urbain Le Verrier had access to plenty of meticulously documented observations of Uranus over time. He worked on the problem for well over a year, trying to figure out what combination of mass, position, and orbit could account for this orbital wobble. In September 1846, he sent a letter to a colleague at an observatory in Germany: Look in this exact spot. Something should be there. It will look like a star, but it’s not. Days later, he held the electrifying response in his hands: “The planet whose position you predicted really exists.” Using decades of data and a lifetime of mathematics, Le Verrier had discovered Neptune without ever looking through a telescope. Le Verrier’s approach is not limited to planet-hunting: In scenarios for which direct observation is impossibly convoluted, we can gain valuable insights into the dynamics of complex systems by characterizing how they alter the state of other, more practically 2 observed entities. We consider two such systems here: First, the careers of biologists and the publishing decisions they make to share results and maintain their livelihoods. Second, the study of the human microbiome and how we’ve come to know the things we do about the microbial communities with which we are constantly interacting. Whether it’s biology preprints or microbiome sequencing, we observe a similar problem in both scenarios: As a consequential change in practice moves from niche interest into broad popularity, the communities involved frequently look very different from the communities affected, and as growth explodes outward, things move too quickly for anyone to notice the patterns emerging. In both cases, this work focuses on the characterization of this growth, important patterns in participation, and the systemic consequences that arise from tens of thousands of seemingly individual decisions. Preprints and publishing in biology The first two chapters deal with global dynamics in scholarly biology publishing, particularly the growing popularity of preprints. The term “preprint” itself is something of an anachronism in the internet age, but the appeal is straightforward: Peer review can take months or years so sharing your manuscript online means you can short-circuit this process and gain most of the benefits of publication on your own timeline: You can tell the world about your findings, yes, but preprints can also be used to establish priority for an exciting discovery, solicit feedback from the field, find new collaborators, and showcase your work to funders and employers. For readers, it also gives a valuable preview of what work is 3 making its way through conventional channels—if exciting papers are being shared as preprints, reading only published papers means you’re permanently six months behind. Researchers in fields such as physics and computer science have been posting preprints for decades, but the practice didn’t gain a foothold in the life sciences until the early 2010s. Chapter 1 describes a quantification of this growth: Prior to its publication, there was very little data even about the volume of preprints in biology, which were clearly growing in popularity, but in directions that were unclear. We quantify how many preprints have been posted over time, the topics of those preprints, and importantly, how many of them are later published, and by whom. I build on this in Chapter 2, which augments the initial dataset with information about the institutional affiliation of authors. I use this data to evaluate country-level trends in authorship and international collaboration. While chapter one explains why editors, publishers and institutions were motivated to paying special attention to preprints, chapter two examines who is being excluded by those programs and how, without more careful consideration, the “democratizing” influence of preprints could end up reproducing the same inequities found in conventional publishing. The human microbiome The final two chapters take a similar approach to the evaluation of human microbiome research around the world: Chapter 3 describes a database I built indexing sample-level metadata for all publicly available human microbiome samples available from the largest genomic repositories in the world. I then examine which countries hold a disproportionate 4 amount of influence in our understanding of the human microbiome and examine the consequences of the field’s exclusion of broad segments of the world population. Chapter 4 describes a new dataset I generated by processing many of those samples: To quantify large-scale variation in the human gut microbiome, we built a compendium of taxonomic information on the makeup of more than 170,000 amplicon sequencing samples, more than three times as large as any comparable dataset available today. Understanding the behavior and ecological dynamics of microbes are important, in short, because they are everywhere. They inhabit the skin and hair of primates, they coat the roots of plants and the surface of our eyeballs (Lu and Liu 2016). The most recent estimate found the average human male has 30 trillion human cells but is still greatly outnumbered by the 38 trillion bacterial cells making up his personal microbiome (Sender, Fuchs, and Milo 2016). Studies have found these microbes, located mostly in the gut, are deeply involved with health and disease. In humans, some bacteria have been found to improve the effectiveness of cancer immunotherapy treatments (Matson et al. 2018), and microbiome diversity was also found to be a highly accurate predictor of whether lymphoma patients would develop potentially lethal bloodstream infections during treatment (Montassier et al. 2016). The microbiome is suspected to play a role in many other conditions as well: Links to obesity are the subject of ongoing debate and analysis (Sze and Schloss 2016), as are links to autism (Sharon et al. 2019; Saurman, Margolis, and Luna 2020) and Alzheimer’s disease (Vogt et al. 2017). 5 In some of these relationships, we understand the possible mechanisms through which these effects arise. In colorectal cancer, for example, metabolites generated by Fusobacterium nucleatum were found to directly influence β-catenin signaling and substantially increase tumor proliferation (Rubinstein et al. 2013), and some opportunistic pathogens can induce gut inflammation that drives out commensal competitors (Zeng, Inohara, and Nuñez 2017). In most cases, however, the reasons for observed correlations are unknown: Are hosts responding to the presence of particular microbes, or are human cells cultivating beneficial bacterial communities around themselves? A recurring frustration in these studies is the statistical challenge of dealing with noisy, high-dimensional datasets generated by DNA-based surveys of the microbiome. Fecal samples may contain hundreds or thousands of detected microbial taxa, making it challenging to find the relevant signal in even the largest studies. A better understanding of broad patterns of covariance in the human microbiome would enable researchers to summarize their data in fewer dimensions, giving microbiome studies a similar boost as is offered by gene pathway analysis: Though there may be no relevant links to an individual taxon, patterns may be more apparent looking at larger groups of taxa. It has been my goal in this PhD work to quantify trends and patterns that are frequently discussed but rarely evaluated: We give a lot of attention to the “human microbiome,” but which humans? Editors and publishers are developing new programs to incorporate preprints into the notoriously inequitable academic publishing system, but who is writing them? Which preprints are being published, and by whom? On such a large scale, these 6 questions can only be properly addressed using thorough and detailed collection of metadata—this document describes my attempt at collecting it. I am proud of the work I present here and am humbled by the opportunity to share it with you. Thanks for reading. —Rich 7 Chapter One: Tracking the popularity and outcomes of all bioRxiv preprints The content in this chapter is based on work previously published in eLife. Copyright retained by the authors and available via Creative Commons Attribution License v4. Abdill RJ and Blekhman R, 2019. Meta-Research: Tracking the popularity and outcomes of all bioRxiv preprints. eLife, 8:e45133. DOI: 10.7554/eLife.45133. Please refer to Appendix A for supplementary materials. 8 Summary The growth of preprints in the life sciences has been reported widely and is driving policy changes for journals and funders, but little quantitative information has been published about preprint usage. Here, we report how we collected and analyzed data on all 37,648 preprints uploaded to bioRxiv.org, the largest biology-focused preprint server, in its first five years. The rate of preprint uploads to bioRxiv continues to grow (exceeding 2,100 in October 2018), as does the number of downloads (1.1 million in October 2018). We also find that two-thirds of preprints posted before 2017 were later published in peer-reviewed journals, and we find a relationship between the number of downloads a preprint has received and the impact factor of the journal in which it is published. We also describe Rxivist.org, a web application that provides multiple ways to interact with preprint metadata. Background In the 30 days of September 2018, four leading biology journals—The Journal of Biochemistry, PLOS Biology, Genetics and Cell—published 85 full-length research articles. The preprint server bioRxiv (pronounced “Bio Archive”) had posted this number of preprints by the end of September. Preprints allow researchers to make their results available as quickly and widely as possible, short-circuiting the delays and requests for extra experiments often associated with peer review (Berg et al. 2016; Powell 2016; Raff, Johnson, and Walter 2008; Snyder 2013; Hartgerink 2015; Vale 2015; Royle 2014). 9 Physicists have been sharing preprints using the service now called arXiv.org since 1991 (Verma 2017), but early efforts to facilitate preprints in the life sciences failed to gain traction (Cobb 2017; Desjardins-Proulx et al. 2013). An early proposal to host preprints on PubMed Central (Varmus 1999; Smaglik 1999) was scuttled by the National Academy of Sciences, which successfully negotiated to exclude work that had not been peer-reviewed (Marshall 1999; Kling, Spector, and Fortuna 2004). Further attempts to circulate biology preprints, such as NetPrints (Delamothe et al. 1999), Nature Precedings (Kaiser 2017), and The Lancet Electronic Research Archive (McConnell and Horton 1999), popped up (and then folded) over time (“ERA Home” 2005). The preprint server that would catch on, bioRxiv, was not founded until 2013 (Callaway 2013). Now, biology publishers are actively trawling preprint servers for submissions (Barsh et al. 2016; Vence 2017), and more than 100 journals accept submissions directly from the bioRxiv website (“Submit a Manuscript” n.d.). The National Institutes of Health now allows researchers to cite preprints in grant proposals (“Reporting Preprints and Other Interim Research Products” n.d.), and grants from the Chan Zuckerberg Initiative require researchers to post their manuscripts to preprint servers (“Science Funding” 2019; Champieux 2018). Preprints are influencing publishing conventions in the life sciences, but many details about the preprint ecosystem remain unclear. We know bioRxiv is the largest of the biology preprint servers (Anaya 2018), and sporadic updates from bioRxiv leaders show steadily increasing submission numbers (Sever 2018). Analyses have examined metrics such as total downloads (Serghiou and Ioannidis 2018) and publication rate (Schloss 2017), but 10 long-term questions remain open. Which fields have posted the most preprints, and which collections are growing most quickly? How many times have preprints been downloaded, and which categories are most popular with readers? How many preprints are eventually published elsewhere, and in what journals? Is there a relationship between a preprint’s popularity and the journal in which it later appears? Do these conclusions change over time? Here, we aim to answer these questions by collecting metadata about all 37,648 preprints posted to bioRxiv from its launch through November 2018. As part of this effort, we have developed Rxivist (pronounced “Archivist”): a website, API and database (available at https://rxivist.org and gopher://origin.rxivist.org) that provide a fully featured system for interacting programmatically with the periodically indexed metadata of all preprints posted to bioRxiv. Results We developed a Python-based web crawler to visit every content page on the bioRxiv website and download basic data about each preprint across the site’s 27 subject-specific categories: title, authors, download statistics, submission date, category, DOI, and abstract. The bioRxiv website also provides the email address and institutional affiliation of each author, plus, if the preprint has been published, its new DOI and the journal in which it appeared. For those preprints, we also used information from Crossref to determine the date of publication. We have stored these data in a PostgreSQL database; snapshots of the 11 database are available for download, and users can access data for individual preprints and authors on the Rxivist website and API. Additionally, a repository is available online at https://doi.org/10.5281/zenodo.2465689 that includes the database snapshot used for this manuscript, plus the data files used to create all figures. Code to regenerate all the figures in this paper is included there and on GitHub (https://github.com/blekhmanlab/rxivist/blob/master/paper/figures.md). See Methods and Supplementary Information for a complete description. Preprint submissions The most apparent trend that can be pulled from the bioRxiv data is that the website is becoming an increasingly popular venue for authors to share their work, at a rate that increases almost monthly. There were 37,648 preprints available on bioRxiv at the end of November 2018, and more preprints were posted in the first 11 months of 2018 (18,825) than in all four previous years combined (Figure 1-1a). The number of bioRxiv preprints doubled in less than a year, and new submissions have been trending upward for five years (Figure 1-1b). The largest driver of site-wide growth has been the neuroscience collection, which has had more submissions than any bioRxiv category in every month since September 2016 (Figure 1-1b). In October 2018, it became the first of bioRxiv’s collections to contain 6,000 preprints (Figure 1-1a). The second-largest category is bioinformatics (4,249 preprints), followed by evolutionary biology (2,934). October 2018 was also the first month in which bioRxiv posted more than 2,000 preprints, increasing its total preprint count by 6.3% (2,119) in 31 days. 12 Figure 1-1. Preprints over time. (a) The number of preprints (y-axis) at each month (x- axis), with each category depicted as a line in a different color. Inset: The overall number of preprints on bioRxiv in each month. (b) The number of preprints posted (y-axis) in each month (x-axis) by category. The category color key is provided below the figure. 13 Preprint downloads Using preprint downloads as a metric for readership, we find that bioRxiv’s usage among readers is also increasing rapidly (Figure 1-2). The total download count in October 2018 (1,140,296) was an 82% increase over October 2017, which itself was a 115% increase over October 2016 (Figure 1-2a). BioRxiv preprints were downloaded almost 9.3 million times in the first 11 months of 2018, and in October and November 2018, bioRxiv recorded more downloads (2,248,652) than in the website’s first two and a half years (Figure 1-2b). The overall median downloads per paper is 279 (Figure 1-2b, inset), and the genomics category has the highest median downloads per paper, with 496 (Figure 1-2c). The neuroscience category has the most downloads overall—it overtook bioinformatics in that metric in October 2018, after bioinformatics spent nearly four and a half years as the most downloaded category (Figure 1-2d). In total, bioRxiv preprints were downloaded 19,699,115 times from November 2013 through November 2018, and the neuroscience category’s 3,184,456 total downloads accounts for 16.2% of these (Figure 1-2d). However, this is driven mostly by that category’s high volume of preprints: the median downloads per paper in the neuroscience category is 269.5, while the median of preprints in all other categories is 281 (Figure 1-2c; Mann–Whitney U test p=0.0003). 14 Figure 1-2. Distribution of all recorded downloads of bioRxiv preprints. (a) The downloads recorded in each month, with each line representing a different year. The lines reflect the same totals as the height of the bars in Figure 1-2b. (b) A stacked bar plot of the downloads in each month. The height of each bar indicates the total downloads in that month. Each stacked bar shows the number of downloads in that month attributable to each category; the colors of the bars are described in the legend in Figure 1-1. Inset: A histogram showing the site-wide distribution of downloads per preprint. The median download count for a single preprint is 279, marked by the yellow dashed line. (c) The distribution of downloads per preprint, broken down by category. Each box illustrates that category’s first quartile, median, and third quartile. The vertical dashed yellow line indicates the overall median downloads for all preprints. (d) Cumulative downloads over time of all preprints in each category. The top seven categories at the end of the plot (November 2018) are labeled using the same category color-coding as above. 15 We also examined traffic numbers for individual preprints relative to the date that they were posted to bioRxiv, which helped create a picture of the change in a preprint’s downloads by month (Supp. Figure A-1). We can see that preprints typically have the most downloads in their first month, and the download count per month decays most quickly over a preprint’s first year on the site. The most downloads recorded in a preprint’s first month is 96,047, but the median number of downloads a preprint receives in its debut month on bioRxiv is 73. The median downloads in a preprint’s second month falls to 46, and the third month median falls again, to 27. Even so, the average preprint at the end of its first year online is still being downloaded about 12 times per month, and some papers don’t have a “big” month until relatively late, receiving the majority of their downloads in their sixth month or later (Supp. Figure A-2). Preprint authors While data about the authors of individual preprints is easy to organize, associating authors between preprints is difficult due to a lack of consistent unique identifiers (see Methods). We chose to define an author as a unique name in the author list, including middle initials but disregarding letter case and punctuation. Keeping this in mind, we find that there are 170,287 individual authors with content on bioRxiv. Of these, 106,231 (62.4%) posted a preprint in 2018, including 84,339 who posted a preprint for the first time (Table 1-1), indicating that total authors increased by more than 98% in 2018. 16 Year New authors Total authors 2013 608 608 2014 3,873 4,012 2015 7,584 8,411 2016 21,832 24,699 2017 52,051 61,239 2018 84,339 106,231 Table 1-1. Unique preprint authors per year. “New authors” counts authors posting preprints in that year that had never posted before; “Total authors” includes researchers who may have already been counted in a previous year but are also listed as an author on a preprint posted in that year. Data for table pulled directly from database. An SQL query to generate these numbers is provided in the Methods section. Even though 129,419 authors (76.0%) are associated with only a single preprint, the mean preprints per author is 1.52 because of a skewed rate of contributions also found in conventional publishing (Rørstad and Aksnes, 2015): 10% of authors account for 72.8% of all preprints, and the most prolific researcher on bioRxiv, George Davey Smith, is listed on 97 preprints across seven categories. 1,473 authors list their most recent affiliation as Stanford University, the most represented institution on bioRxiv (Supp. Table A-1). Though the majority of the top 100 universities (by author count) are based in the United States, five of the top 11 are from Great Britain. These results rely on data provided by authors, however, and is confounded by varying levels of specificity: while 530 authors report their affiliation as “Harvard University,” for example, there are 528 different 17 institutions that include the phrase “Harvard,” and the four preprints from the “Wyss Institute for Biologically Inspired Engineering at Harvard University” don’t count toward the “Harvard University” total. Publication outcomes In addition to monthly download statistics, bioRxiv also records whether a preprint has been published elsewhere, and in what journal. In total, 15,797 bioRxiv preprints have been published, or 42.0% of all preprints on the site (Figure 1-3a), according to bioRxiv’s records linking preprints to their external publications. Proportionally, evolutionary biology preprints have the highest publication rate of the bioRxiv categories: 51.5% of all bioRxiv evolutionary biology preprints have been published in a journal (Figure 1-3b). Examining the raw number of publications per category, neuroscience again comes out on top, with 2,608 preprints in that category published elsewhere (Figure 1-3c). When comparing the publication rates of preprints posted in each month, we see that more recent preprints are published at a rate close to zero, followed by an increase in the rate of publication every month for about 12–18 months (Figure 1-3a). A similar dynamic was observed in a study of preprints posted to arXiv; after recording lower rates in the most recent time periods, Larivière et al. found publication rates of arXiv preprints leveled out at about 73% (Larivière et al. 2014). Of bioRxiv preprints posted between 2013 and the end of 2016, 67.0% have been published; if 2017 papers are included, that number falls to 64.0%. Of preprints posted in 2018, only 20.0% have been printed elsewhere (Figure 1- 3a). 18 Figure 1-3. Characteristics of the bioRxiv preprints published in journals. (a) The proportion of preprints that have been published (y-axis), broken down by the month in which the preprint was first posted (x-axis). (b) The proportion of preprints in each category that have been published elsewhere. The dashed line marks the overall proportion of bioRxiv preprints that have been published and is at the same position as the dashed line in panel 3a. (c) The number of preprints in each category that have been published in a journal. These publication statistics are based on data produced by bioRxiv’s internal system that links publications to their preprint versions, a difficult challenge that appears to rely heavily 19 on title-based matching. To better understand the reliability of the linking between preprints and their published versions, we selected a sample of 120 preprints that were not indicated as being published, and manually validated their publication status using Google and Google Scholar (see Methods). Overall, 37.5% of these “unpublished” preprints had actually appeared in a journal. We found earlier years to have a much higher false-negative rate: 53% of the evaluated “unpublished” preprints from 2015 had actually been published, though that number dropped to less than 17% in 2017 (Supp. Figure A-3). While a more robust study would be required to draw more detailed conclusions about the “true” publication rate, this preliminary examination suggests the data from bioRxiv may be an underestimation of the number of preprints that have been published. Overall, 15,797 bioRxiv preprints have appeared in 1,531 different journals (Figure 1-4). Scientific Reports has published the most, with 828 papers, followed by eLife and PLOS ONE with 750 and 741 papers, respectively. However, considering the proportion of preprints of the total papers published in each journal can lead to a different interpretation. For example, Scientific Reports published 398 bioRxiv preprints in 2018, but this represents 2.36% of the 16,899 articles it published in that year, as indexed by Web of Science (Supp. Table A-2). In contrast, eLife published almost as many bioRxiv preprints (394), which means more than a third of their 1,172 articles from 2018 first appeared on bioRxiv. GigaScience had the highest proportion of articles from preprints in 2018 (49.4% of 89 articles), followed by Genome Biology (39.9% of 183 articles) and Genome Research 20 (36.7% of 169 articles). Incorporating all years in which bioRxiv preprints have been published (2014–2018), these are also the three top journals. Figure 1-4. Journals that have published the most preprints. The bars indicate the number of preprints published by each journal, broken down by the bioRxiv categories to which the preprints were originally posted. 21 Some journals have accepted a broad range of preprints, though none have hit all 27 of bioRxiv’s categories—PLOS ONE has published the most diverse category list, with 26. (It has yet to publish a preprint from the clinical trials collection, bioRxiv’s second smallest.) Other journals are much more specialized, though in expected ways. Of the 172 bioRxiv preprints published by the Journal of Neuroscience, 169 were in neuroscience, and three were from animal behavior and cognition. Similarly, NeuroImage has published 211 neuroscience papers, two in bioinformatics, and one in bioengineering. It should be noted that these counts are based on the publications detected by bioRxiv and linked to their preprint, so some journals—for example, those that more frequently rewrite the titles of articles—may be underrepresented here. When evaluating the downloads of preprints published in individual journals (Figure 1-5), there is a significant positive correlation between the median downloads per paper and journal impact factor (JIF): in general, journals with higher impact factors (Clarivate Analytics 2018) publish preprints that have more downloads. For example, Nature Methods (2017 JIF 26.919) has published 119 bioRxiv preprints; the median download count of these preprints is 2,266. By comparison, PLOS ONE (2017 JIF 2.766) has published 719 preprints with a median download count of 279 (Figure 1-5). In this analysis, each data point in the regression represented a journal, indicating its JIF and the median downloads per paper for the preprints it had published. We found a significant positive correlation between these two measurements (Kendall’s τb=0.5862, p=1.364×10- 6). We also found a similar, albeit weaker, correlation when we performed another analysis 22 in which each data point represented a single preprint (n=7,445; Kendall’s τb=0.2053, p=9.311×10-152; see Methods). Figure 1-5. Preprint downloads by publishing journal. Each box illustrates the journal’s first quartile, median, and third quartile, as in Figure 1-2c. Colors correspond to journal access policy as described in the legend. Inset: A scatterplot in which each point represents an academic journal, showing the relationship between median downloads of the bioRxiv preprints published in the journal (x-axis) against its 2017 journal impact factor (y-axis). The size of each point is scaled to reflect the total number of bioRxiv preprints published by that journal. The regression line in this plot was calculated using the “lm” function in the R ‘stats’ package, but all reported statistics use the Kendall rank correlation coefficient, which does not make as many assumptions about normality or homoscedasticity. 23 It is important to note that we did not evaluate when these downloads occurred, relative to a preprint’s publication. While it looks like accruing more downloads makes it more likely that a preprint will appear in a higher impact journal, it is also possible that appearances in particular journals drive bioRxiv downloads after publication. The Rxivist dataset has already been used to begin evaluating questions like this (Kramer 2019), and further study may be able to unravel the links, if any, between downloads and journals. If journals are driving post-publication downloads on bioRxiv, however, their efforts are curiously consistent: preprints that have been published elsewhere have almost twice as many downloads as preprints that have not (Table 1-2; Mann–Whitney U test, p<2.2×10- 16). Among papers that have not been published, the median number of downloads per preprint is 208. For preprints that have been published, the median download count is 394 (Mann–Whitney U test, p<2.2×10-16). When preprints published in 2018 are excluded from this calculation, the difference between published and unpublished preprints shrinks, but is still significant (Table 1-2; Mann–Whitney U test, p<2.2×10-16). Though preprints posted in 2018 received more downloads in 2018 than preprints posted in previous years did (Supp. Figure A-4), it appears they have not yet had time to accumulate as many downloads as papers from previous years (Supp. Figure A-5). Posted Published Unpublished 2017 and earlier 465 414 Through 2018 394 208 Table 1-2. Median downloads by publication status. See Methods section for description of tests used. 24 We also retrieved the publication date for all published preprints using the Crossref “Metadata Delivery” API (“Metadata Delivery REST API” 2018). This, combined with the bioRxiv data, gives us a comprehensive picture of the interval between the date a preprint is first posted to bioRxiv and the date it is published by a journal. These data show the median interval is 166 days, or about 5.5 months. 75% of preprints are published within 247 days of appearing on bioRxiv, and 90% are published within 346 days (Figure 1-6a). The median interval we found at the end of November 2018 (166 days) is a 23.9% increase over the 134-day median interval reported by bioRxiv in mid-2016 (Inglis and Sever 2016). We also used these data to further examine patterns in the properties of the preprints that appear in individual journals. The journal that publishes preprints with the highest median age is Nature Genetics, whose median interval between bioRxiv posting and publication is 272 days (Figure 1-6b), a significant difference from every journal except Genome Research (Kruskal–Wallis rank sum test, p<2.2×10-16; Dunn’s test q<0.05 comparing Nature Genetics to all other journals except Genome Research, after Benjamini–Hochberg correction). Among the 30 journals publishing the most bioRxiv preprints, the journal with the most rapid transition from bioRxiv to publication is G3, whose median, 119 days, is significantly different from all journals except Genetics, mBio, and The Biophysical Journal (Figure 1-5). 25 Figure 1-6. Time to publication by journal. (a) A histogram showing the distribution of publication intervals. The x-axis indicates the time between preprint posting and journal publication; the y-axis indicates how many preprints fall within the limits of each bin. The yellow line indicates the median; the same data is also visualized using a boxplot above the histogram. (b) The publication intervals of preprints, broken down by the journal in which each appeared. The journals in this list are the 30 journals that have published the most total bioRxiv preprints; the plot for each journal indicates the density distribution of the preprints published by that journal, excluding any papers that were posted to bioRxiv after publication. Portions of the distributions beyond 1,000 days are not displayed. 26 It is important to note that this metric does not directly evaluate the production processes at individual journals. Authors submit preprints to bioRxiv at different points in the publication process and may work with multiple journals before publication, so individual data points capture a variety of experiences. For example, 122 preprints were published within a week of being posted to bioRxiv, and the longest period between preprint and publication is 3 years, 7 months and 2 days, for a preprint that was posted in March 2015 and not published until October 2018 (Figure 1-6a). Discussion Biology preprints have a growing presence in scientific communication, and we now have ongoing, detailed data to quantify this process. The ability to better characterize the preprint ecosystem can inform decision-making at multiple levels. For authors, particularly those looking for feedback from the community, our results show bioRxiv preprints are being downloaded more than one million times per month, and that an average paper can receive hundreds of downloads in its first few months online (Supp. Figure A-1). Serghiou and Ioannidis (2018) evaluated download metrics for bioRxiv preprints through 2016 and found an almost identical median for downloads in a preprint’s first month; we have expanded this to include more detailed longitudinal traffic metrics for the entire bioRxiv collection (Figure 1-2b). For readers, we show that thousands of new preprints are being posted every month. This tracks closely with a widely referenced summary of submissions to preprint servers 27 (“Monthly Statistics for October 2018” 2018) generated monthly by PrePubMed (http://www.prepubmed.org) and expands on submission data collected by researchers using custom web scrapers of their own (Stuart 2016, 2017; Holdgraf 2016). There is also enough data to provide some evidence against the perception that research in preprint is less rigorous than papers appearing in journals (“Methods, Preprints and Papers” 2017; Vale 2015). In short, the majority of bioRxiv preprints do appear in journals eventually, and potentially with very few differences: an analysis of published preprints that had first been posted to arXiv.org found that “the vast majority of final published papers are largely indistinguishable from their pre-print versions” (Klein et al. 2016). A 2016 project measured which journals had published the most bioRxiv preprints (Schmid 2016); despite a six-fold increase in the number of published preprints since then, 23 of the top 30 journals found in their results are also in the top 30 journals we found (Figure 1-5). For authors, we also have a clearer picture of the fate of preprints after they are shared online. Among preprints that are eventually published, we found that 75% have appeared in a journal by the time they had spent 247 days (about eight months) on bioRxiv. This interval is similar to results from Larivière et al. showing preprints on arXiv were most frequently published within a year of being posted there (Larivière et al. 2014), and to a later study examining bioRxiv preprints that found “the probability of publication in the peer-reviewed literature was 48% within 12 months” (Serghiou and Ioannidis 2018). Another study published in spring 2017 found that 33.6% of preprints from 2015 and earlier had been published (Schloss 2017); our data through November 2018 show that 68.2% of 28 preprints from 2015 and earlier have been published. Multiple studies have examined the interval between submission and publication at individual journals (Himmelstein 2016a; Royle 2015; Powell 2016), but the incorporation of information about preprints is not as common. We also found a positive correlation between the impact factor of journals and the number of downloads received by the preprints they have published. This finding in particular should be interpreted with caution. Journal impact factor is broadly intended to be a measurement of how citable a journal’s “average” paper is (Garfield 2006), though it morphed long ago into an unfounded proxy for scientific quality in individual papers (The PLoS Medicine Editors 2006). It is referenced here only as an observation about a journal- level metric correlated with preprint downloads; there is no indication that either factor is influencing the other, nor that download numbers play a direct role in publication decisions. More broadly, our granular data provide a new level of detail for researchers looking to evaluate many remaining questions. What factors may impact the interval between when a preprint is posted to bioRxiv and when it is published elsewhere? Does a paper’s presence on bioRxiv have any relationship to its eventual citation count once it is published in a journal, as has been found with arXiv (e.g. Feldman, Lo, and Ammar 2018; Wang, Glänzel, and Chen 2018; Schwarz and Kennicutt 2004)? What can we learn from “altmetrics” as they relate to preprints, and is there value in measuring a preprint’s impact using methods rooted in online interactions rather than citation count (Haustein 2018)? One study, published before bioRxiv launched, found a significant association between Twitter 29 mentions of published papers and their citation counts (Thelwall et al. 2013)—have preprints changed this dynamic? Researchers have used existing resources and custom scripts to answer questions like these. Himmelstein found that only 17.8% of bioRxiv papers had an “open license” (Himmelstein 2016b), for example, and another study examined the relationship between Facebook “likes” of preprints and “traditional impact indicators” such as citation count but found no correlation for papers on bioRxiv (Ringelhan, Wollersheim, and Welpe 2015). Since most bioRxiv data is not programmatically accessible, many of these studies had to begin by scraping data from the bioRxiv website itself. The Rxivist API allows users to request the details of any preprint or author on bioRxiv, and the database snapshots enable bulk querying of preprints using SQL, C and several other languages (PostgreSQL Global Development Group 2018) at a level of complexity currently unavailable using the standard bioRxiv web interface. Using these resources, researchers can now perform detailed and robust bibliometric analysis of the website with the largest collection of preprints in biology, the one that, beginning in September 2018, held more biology preprints than all other major preprint servers combined (Anaya 2018). In addition to our analysis here that focuses on big-picture trends related to bioRxiv, the Rxivist website provides many additional features that may interest preprint readers and authors. Its primary feature is sorting and filtering preprints based by download count or mentions on Twitter, to help users find preprints in particular categories that are being discussed either in the short term (Twitter) or over the span of months (downloads). 30 Tracking these metrics could also help authors gauge public reaction to their work. While bioRxiv has compensated for a low rate of comments posted on the site itself (Inglis and Sever 2016) by highlighting external sources such as tweets and blogs, Rxivist provides additional context for how a preprint compares to others on similar topics. Several other sites have attempted to use social interaction data to “rank” preprints, though none incorporate bioRxiv download metrics. The “Assert” web application (https://assert.pub) ranks preprints from multiple repositories based on data from Twitter and GitHub. The “PromisingPreprints” Twitter bot (https://twitter.com/PromPreprint) accomplishes a similar goal, posting links to bioRxiv preprints that receive an exceptionally high social media attention score (Altmetric Support 2018) from Altmetric (https://www.altmetric.com) in their first week on bioRxiv (De Coster 2017). Arxiv Sanity Preserver (http://www.arxiv-sanity.com) provides rankings of arXiv.org preprints based on Twitter activity, though its implementation of this scoring (Karpathy 2018) is more opinionated than that of Rxivist. Other websites perform similar curation, but they rely on user interactions within the sites themselves: SciRate (https://scirate.com), Paperkast (https://paperkast.com) and upvote.pub allow users to vote on articles that should receive more attention (van der Silk et al. 2018), though upvote.pub is no longer online (“Upvote.pub Snapshot” 2018). By comparison, Rxivist doesn’t rely on user interaction— by pulling “popularity” metrics from Twitter and bioRxiv, we aim to decouple the quality of our data from the popularity of the website itself. 31 In summary, our approach provides multiple perspectives on trends in biology preprints: (1) the Rxivist.org website, where readers can prioritize preprints and generate reading lists tailored to specific topics; (2) a dataset that can provide a foundation for developers and bibliometric researchers to build new tools, websites and studies that can further improve the ways we interact with preprints and (3) an analysis that brings together a comprehensive summary of trends in bioRxiv preprints and an examination of the crossover points between preprints and conventional publishing. Methods The Rxivist website. We attempted to put the Rxivist data to good use in a relatively straightforward web application. Its main offering is a ranked list of all bioRxiv preprints that can be filtered by areas of interest. The rankings are based on two available metrics: either the count of PDF downloads, as reported by bioRxiv, or the number of Twitter messages linking to that preprint, as reported by Crossref (https://crossref.org). Users can also specify a timeframe for the search—for example, one could request the most downloaded preprints in microbiology over the last two months or view the preprints with the most Twitter activity since yesterday across all categories. Each preprint and each author is given a separate profile page, populated only by Rxivist data available from the API. These include rankings across multiple categories, plus a visualization of where the download totals for each preprint (and author) fall in the overall distribution across all 37,000 preprints and 170,000 authors. 32 The Rxivist API and dataset. The full data described in this paper is available through Rxivist.org, a website developed for this purpose. BioRxiv data is available from Rxivist in two formats: (1) SQL “database dumps” are currently pulled and published weekly on zenodo.org. (See Supplementary Information for a visualization and description of the schema.) These convert the entire Rxivist database into binary files that can be loaded by the free and open-source PostgreSQL database management system to provide a local copy of all collected data on every article and author on bioRxiv.org. (2) We also provide an API (application programming interface) from which users can request information in JSON format about individual preprints and authors, or search for preprints based on similar criteria available on the Rxivist website. Complete documentation is available at https://www.rxivist.org/docs. While the analysis presented here deals mostly with overall trends on bioRxiv, the primary entity of the Rxivist API is the individual research preprint, for which we have a straightforward collection of metadata: title, abstract, DOI (digital object identifier), the name of any journal that has also published the preprint (and its new DOI), and which collection the preprint was submitted to. We also collected monthly traffic information for each preprint, as reported by bioRxiv. We use the PDF download statistics to generate rankings for each preprint, both site-wide and for each collection, over multiple timeframes (all-time, year to date, etc.). In the API and its underlying database schema, “authors” exist separately from “preprints” because an author can be associated with multiple preprints. They are recorded with three main pieces of data: name, institutional affiliation and a 33 unique identifier issued by ORCID. Like preprints, authors are ranked based on the cumulative downloads of all their preprints, and separately based on downloads within individual bioRxiv collections. Emails are collected for each researcher but are not necessarily unique (see “Consolidation of author identities” below). Web crawler design. To collect information on all bioRxiv preprints, we developed an application that pulled preprint data directly from the bioRxiv website. The primary issue with managing this data is keeping it up to date: Rxivist aims to essentially maintain an accurate copy of a subset of bioRxiv’s production database, which means routinely running a web crawler against the website to find any new or updated content as it is posted. We have tried to find a balance between timely updates and observing courteous web crawler behavior; currently, each preprint is re-crawled once every two to three weeks to refresh its download metrics and publication status. The web crawler itself uses Python 3 and requires two primary modules for interacting with external services: Requests-HTML (Reitz 2018) is used for fetching individual web pages and pulling out the relevant data, and the psycopg2 module (Di Gregorio and Varrazzo 2018) is used to communicate with the PostgreSQL database that stores all the Rxivist data (PostgreSQL Global Development Group 2017). PostgreSQL was selected over other similar database management systems because of its native support for text search, which, in our implementation, enables users to search for preprints based on the contents of their titles, abstracts and author list. The API, spider and web application are all hosted within separate Docker containers (Docker Inc 2018), a decision we made to simplify the logistics required for others to deploy the 34 components on their own: Docker is the only dependency, so most workstations and servers should be able to run any of the components. New preprints are recorded by parsing the section of the bioRxiv website that lists all preprints in reverse-chronological order. At this point, a preprint’s title, URL and DOI are recorded. The bioRxiv webpage for each preprint is then crawled to obtain details only available there: the abstract, the date the preprint was first posted, and monthly download statistics are pulled from here, as well as information about the preprint’s authors—name, email address and institution. These authors are then compared against the list of those already indexed by Rxivist, and any unrecognized authors have profiles created in the database. Consolidation of author identities. Authors are most reliably identified across multiple papers using the bioRxiv feature that allows authors to specify an identifier provided by ORCID (https://orcid.org), a nonprofit that provides a voluntary system to create unique identification numbers for individuals. These ORCID (Open Researcher and Contributor ID) numbers are intended to serve approximately the same role for authors that DOIs do for papers (Haak 2012), providing a way to identify individuals whose other information may change over time. 29,559 bioRxiv authors, or 17.4%, have an associated ORCID. If an individual included in a preprint’s list of authors does not have an ORCID already recorded in the database, authors are consolidated if they have an identical name to an existing Rxivist author. 35 There are certainly authors who are duplicated within the Rxivist database, an issue arising mostly from the common complaint of unreliable source data. 68.4% of indexed authors have at least one email address associated with them, for example, including 7,085 (4.40%) authors with more than one. However, of the 118,490 email addresses in the Rxivist database, 6,517 (5.50%) are duplicates that are associated with more than one author. Some of these are because real-life authors occasionally appear under multiple names, but other duplicates are caused by uploaders to bioRxiv using the same email address for multiple authors on the same preprint, making it far more difficult to use email addresses as unique identifiers. There are also cases like one from 2017, in which 16 of the 17 authors of a preprint were listed with the email address “test@test.com.” Inconsistent naming patterns cause another chronic issue that is harder to detect and account for. For example, at one point thousands of duplicate authors were indexed in the Rxivist database with various versions of the same name—including a full middle name, or a middle initial, or a middle initial with a period, and so on—which would all have been recorded as separate people if they did not all share an ORCID, to say nothing of authors who occasionally skip specifying a middle initial altogether. Accommodations could be made to account for inconsistencies such as these (using institutional affiliation or email address as clues, for example), but these methods also have the potential to increase the opposite problem of incorrectly combining different authors with similar names who intentionally introduce slight modifications such as a middle initial to help differentiate themselves. One allowance was made to normalize author names: when the web crawler 36 searches for name matches in the database, periods are now ignored in string matches, so “John Q. Public” would be a match with “John Q Public.” The other naming problem we encountered was of the opposite variety: multiple authors with identical names (and no ORCID). For example, the Rxivist profile for author “Wei Wang” is associated with 40 preprints and 21 different email addresses but is certainly the conglomeration of multiple researchers. A study of more than 30,000 Norwegian researchers found that when using full names rather than initials, the rate of name collisions was 1.4% (Aksnes 2008). Retrieval of publication date information. Publication dates were pulled from the Crossref Metadata Delivery API (“Metadata Delivery REST API” 2018) using the publication DOI numbers provided by bioRxiv. Dates were found for all but 31 (0.2%) of the 15,797 published bioRxiv preprints. Because journals measure publication date in different ways, several metrics were used. If a “published—online” date was available from Crossref with a day, month and year, then that was recorded. If not, “published—print” was used, and the Crossref “created” date was the final option evaluated. Requests for which we received a 404 response were assigned a publication date of 1 Jan 1900, to prevent further attempts to fetch a date for those entries. It appears these articles were published, but with DOIs that were not registered correctly by the destination journal; for consistency, these results were filtered out of the analysis. There was no practical way to validate the nearly 16,000 values retrieved, but anecdotal evaluation reveals some inconsistencies. For example, the preprint with the longest interval before publication (1,371 days) has a publication date reported by Crossref of 1 Jul 2018, when it appeared in 37 IEEE/ACM Transactions on Computational Biology and Bioinformatics 15(4). However, the IEEE website lists a date of 15 Dec 2015, two and a half years earlier, as that paper’s “publication date,” which they define as “the very first instance of public dissemination of content.” Since every publisher is free to make their own unique distinctions, these data are difficult to compare at a granular level. Calculation of download rankings. The web crawler’s “ranking” step orders preprints and authors based on download count in two populations (overall and by bioRxiv category) and over several periods: all-time, year-to-date, and since the beginning of the previous month. The last metric was chosen over a “month-to-date” ranking to avoid ordering papers based on the very limited traffic data available in the first days of each month—in addition to a short lag in the time bioRxiv takes to report downloads, an individual preprint’s download metrics may only be updated in the Rxivist database once every two or three weeks, so metrics for a single month will be biased in favor of those that happen to have been crawled most recently. This effect is not eliminated in longer windows but is diminished. The step recording the rankings takes a more unusual approach to loading the data. Because each article ranking step could require more than 37,000 “insert” or “update” statements, and each author ranking requires more than 170,000 of the same, these modifications are instead written to a text file on the application server and loaded by running an instance of the Postgres command-line client “psql,” which can use the more efficient “copy” command, a change that reduced the duration of the ranking process from several hours to less than one minute. 38 Reporting of small p-values. In several locations, p-values are reported as “<2.2×10-16”. It is important to note that this is an inequality, and these p-values are not necessarily identical. The upper limit, 2.2×10−16, is not itself a particularly meaningful number and is an artifact of the limitations of the floating-point arithmetic used by R, the software used in the analysis. 2.2×10−16 is the “machine epsilon,” or the smallest number that can be added to 1.0 that would generate a result measurably different from 1.0. Though smaller numbers can be represented by the system, those smaller than the machine epsilon are not reported by default; we elected to do the same. Data preparation. Several steps were taken to organize the data that was used for this paper. First, the production data being used for the Rxivist API was copied to a separate “schema”—a PostgreSQL term for a named set of tables. This was identical to the full database but had a specifically circumscribed set of preprints. Once this was copied, the table containing the associations between authors and each of their papers “article_authors”) was pruned to remove references to any articles that were posted after 30 Nov 2018, and any articles that were not associated with a bioRxiv collection. For unknown reasons, 10 preprints (0.03%) could not be associated with a bioRxiv collection; because the bioRxiv profile page for some papers does not specify which collection it belongs to, these papers were ignored. Once these associations were removed, any articles meeting those criteria were removed from the “articles” table. References to these articles were also removed from the table containing monthly bioRxiv download metrics for each paper (“article_traffic”). We also removed all entries from the “article_traffic” table that 39 recorded downloads after November 2018. Next, the table containing author email addresses (“author_emails”) was pruned to remove emails associated with any author that had zero preprints in the new set of papers; those authors were then removed from the “authors” table. Before evaluating data from the table linking published preprints to journals and their post- publication DOI (“article_publications”), journal names were consolidated to avoid under- counting journals with spelling inconsistencies. First, capitalization was stripped from all journal titles, and inconsistent articles (“The Journal of…” vs. “Journal of…”; “and” vs. “&” and so on) were removed. Then, the list of journals was reviewed by hand to remove duplication more difficult to capture automatically: “PNAS” and “Proceedings of the National Academy of Sciences,” for example. Misspellings were rare, but one publication in “integrrative biology” did appear. See figures.md in the project’s GitHub repository (https://github.com/blekhmanlab/rxivist/blob/master/paper/figures.md) for a full list of corrections made to journal titles. We also evaluated preprints for publication in “predatory journals,” organizations that use irresponsibly low academic standards to bolster income from publication fees (Xia et al. 2015). A search for 1,345 journals based on the list compiled by Stop Predatory Journals (https://predatoryjournals.com) showed that bioRxiv publication data did not include any instances of papers appearing in those journals (“List of Predatory Journals” 2018). It is important to note that the absence of this information does not necessarily indicate that preprints have not appeared in these journals—we 40 performed this search to ensure our analysis of publication rates was not inflated with numbers from illegitimate publications. Reproduction of figures. Two files are needed to recreate the figures in this manuscript: a compressed database backup containing a snapshot of the data used in this analysis, and a file called figures.md storing the SQL queries and R code necessary to organize the data and draw the figures. The PostgreSQL documentation for restoring database dumps should provide the necessary steps to “inflate” the database snapshot, and each figure and table is listed in figures.md with the queries to generate comma-separated values files that provide the data underlying each figure. (Those who wish to skip the database reconstruction step will find CSVs for each figure provided along with these other files.) Once the data for each figure is pulled into files, executing the accompanying R code should create figures containing the exact data as displayed here. Tallying institutional authors and preprints. When reporting the counts of bioRxiv authors associated with individual universities, there are several important caveats. First, these counts only include the most recently observed institution for an author on bioRxiv: if someone submits 15 preprints at Stanford, then moves to the University of Iowa and posts another preprint afterward, that author will be associated with the University of Iowa, which will receive all 16 preprints in the inventory. Second, this count is also confounded by inconsistencies in the way authors report their affiliations: for example, “Northwestern University,” which has 396 preprints, is counted separately from “Northwestern University 41 Feinberg School of Medicine,” which has 76. Overlaps such as these were not filtered, though commas in institution names were omitted when grouping preprints together. Evaluation of publication rates. Data referenced in this manuscript is limited to preprints posted through the end of November 2018. However, determining which preprints had been published in journals by the end of November required refreshing the entries for all 37,000 preprints after the month ended. Consequently, it is possible that papers published after the end of November (but not after the first weeks of December) are included in the publication statistics. Estimation of ranges for true publication rates. To evaluate the sensitivity of the system bioRxiv uses to detect published versions of preprints, we pulled a random sample of 120 preprints that had not been marked as published on bioRxiv.org—30 preprints from each year between 2014 and 2017. We then performed a manual online literature search for each paper to determine whether they had been published. The primary search method was searching on Google.com for the preprint’s title and the senior author’s last name. If this did not return any results that looked like publications, other author names were added to the search to replace the senior author’s name. If this did not return any positive results, we also checked Google Scholar (https://scholar.google.com) for papers with similar titles. If any of the preprint’s authors, particularly the first and last authors, had Google Scholar profiles, they were reviewed for publications on subject matter similar to the preprint. If a publication looked similar to the preprint, a visual comparison between the preprint and published paper’s abstract and introduction was used to determine if they were simply 42 different versions of the same paper. The paper was marked as a true negative if none of these returned positive results, or if the suspected published paper described a study that was different enough that the preprint effectively described a different research project. Once all 120 preprints had been evaluated, the results were used to approximate a false- negative rate to each year—the proportion of preprints that had been incorrectly excluded from the list of published papers. The sample size for each year (30) was used to calculate the margin of error using a 95% confidence interval (17.89 percentage points). This margin was then used to generate the minimum and maximum false-negative rates for each year, which were then used to calculate the minimum and maximum number of incorrectly classified preprints from each year. These numbers yielded a range for each year’s actual publication rate; for 2015, for example, bioRxiv identified 1,218 preprints (out of 1,774) that had been published. The false-negative rate and margin of error suggest between 197 and 396 additional preprints have been published but not detected, yielding a final range of 1,415–1,614 preprints published in that year. To evaluate the specificity of the publication detection system, we pulled 40 samples (10 from each of the years listed above) that bioRxiv had listed as published and found that all 40 had been accurately classified. Though this helps establish that bioRxiv is not consistently finding all preprint publications, it should be noted that the determination of a more precise estimation for publication rates would require deeper analysis and sampling. 43 Calculation of publication intervals. There are 15,797 distinct preprints with an associated date of publication in a journal, a corpus too large to allow detailed manual validation across hundreds of journal websites. Consequently, these dates are only as accurate as the data collected by Crossref from the publishers. We attempted to use the earliest publication date, but researchers have found that some publishers may be intentionally manipulating dates associated with publication timelines (Royle 2015), particularly the gap between online and print publication, which can inflate journal impact factor (Tort, Targino, and Amaral 2012). Intentional or not, these gaps may be inflating the time to press measurements of some preprints and journals in our analysis. In addition, there are 66 preprints (0.42%) that have a publication date that falls before the date it was posted to bioRxiv; these were excluded from analyses of publication interval. Counting authors with middle initials. To obtain the comparatively large counts of authors using one or two middle initials, results from a SQL query were used without any curation. For the counts of authors with three or four middle initials, the results of the database call were reviewed by hand to remove “author” names that look like initials but are actually the name of consortia (“International IBD Genetics Consortium”) or authors who provided non-initialized names using all capital letters. 44 Data availability There are multiple web links to resources related to this project: ● The Rxivist application is available on the web at https://rxivist.org and via Gopher at gopher://origin.rxivist.org. ● The source for the web crawler and API is available at https://github.com/blekhmanlab/rxivist (copy archived at https://github.com/elifesciences-publications/rxivist). ● The source for the Rxivist website is available at https://github.com/blekhmanlab/rxivist_web (copy archived at https://github.com/elifesciences-publications/ rxivist_web). ● Data files used to generate the figures in this manuscript are available on Zenodo at https://doi.org/10.5281/zenodo.2465689, as is a snapshot of the database used to create the files. 45 Chapter Two: International authorship and collaboration across bioRxiv preprints The content in this chapter is based on work previously published in eLife. Copyright retained by the authors and available via Creative Commons Attribution License v4. Abdill RJ, Adamowicz EM and Blekhman R., 2020. Meta-Research: International authorship and collaboration across bioRxiv preprints. eLife, 9:e58496. DOI: 10.7554/eLife.58496. Please refer to Appendix B for supplementary materials. 46 Summary Preprints are becoming well established in the life sciences, but relatively little is known about the demographics of the researchers who post preprints and those who do not, or about the collaborations between preprint authors. Here, based on an analysis of 67,885 preprints posted on bioRxiv, we find that some countries, notably the United States and the United Kingdom, are overrepresented on bioRxiv relative to their overall scientific output, while other countries (including China, Russia, and Turkey) show lower levels of bioRxiv adoption. We also describe a set of “contributor countries” (including Uganda, Croatia and Thailand): researchers from these countries appear almost exclusively as non-senior authors on international collaborations. Lastly, we find multiple journals that publish a disproportionate number of preprints from some countries, a dynamic that almost always benefits manuscripts from the US. Background Preprints are being shared at an unprecedented rate in the life sciences (Narock and Goldstein 2019; Abdill and Blekhman 2019b): since 2013, more than 90,000 preprints have been posted to bioRxiv.org, the largest preprint server in the field, including a total of 29,178 in 2019 alone (Abdill and Blekhman 2019a). In addition to allowing researchers to share their work independently of publication at a traditional journal, there is evidence that published papers receive more citations if they first appeared as preprints (Fu and Hughey 2019; Fraser et al. 2020). Some journals also search preprint servers to solicit submissions 47 (Barsh et al. 2016; Vence 2017), and there are various initiatives to encourage and facilitate the peer review of preprints, such as In Review (https://www.researchsquare.com/publishers/in-review), Review Commons (https://www.reviewcommons.org), and Preprint Review (eLife 2020). However, relatively little is known about who is benefiting from the growth of preprints or how this new approach to publishing is affecting different populations of researchers (Penfold and Polka 2020). Academic publishing has grappled for decades with hard-to-quantify concerns about factors of success that are not directly linked to research quality. Studies have found bias in favor of wealthy, English-speaking countries in citation count (Akre et al. 2011) and editorial decisions (Nuñez et al. 2019; Saposnik et al. 2014; Okike et al. 2008; Ross et al. 2006), and there have long been concerns regarding how peer review is influenced by factors such as institutional prestige (C. J. Lee et al. 2013). Preprints have been praised as a democratizing influence on scientific communication (Berg et al. 2016), but a critical question remains: where do they come from? More specifically, which countries are participating in the preprint ecosystem, how are they working with each other, and what happens when they do? Here, we aim to answer these questions by analyzing a dataset of all preprints posted to bioRxiv between its launch in 2013 and the end of 2019. After collecting author-level metadata for each preprint, we used each author’s institutional affiliation to summarize country-level participation and outcomes. 48 Results Country-level bioRxiv participation over time We retrieved author data for 67,885 preprints for which the most recent version was posted before 2020. First, we attributed each preprint to a single country, using the affiliation of the last individual in the author list, considered by convention in the life sciences to be the “senior author” who supervised the work (see Methods). North America, Europe and Australia dominate the top spots (Figure 2-1a): 26,598 manuscripts (39.2%) have a last author from the United States (US), followed by 7151 manuscripts (10.5%) from the United Kingdom (UK; Figure 2-1b), though China (4.1%), Japan (1.9%) and India (1.8%) are the sources of more than 1200 preprints each (Table 2-1). Brazil, with 704 manuscripts, has the 15th-most preprints and is the first South American country on the list, followed by Argentina (163 preprints) in 32nd place. South Africa (182 preprints) is the first African country on the list, in 29th place, followed by Ethiopia (57 preprints) in 42nd place (Supp. Table B-1). It is noticeable that South Africa and Ethiopia both have high opt-in rates for a program operated by PLOS journals that enabled submissions to be sent directly to bioRxiv (PLOS 2019). We found similar results when we looked at which countries were most highly represented based on authorship at any position (Table 2-1). Overall, US authors appear on the most bioRxiv preprints—34,676 manuscripts (51.1%) include at least one US author (Figure 2-1c). 49 Figure 2-1. Preprints per country. (a) A heat map indicating the number of preprints per country, based on the institutional affiliation of the senior author. The color coding uses a log scale. (b) The total preprints attributed to the seven most prolific countries. The x-axis indicates total preprints listing a senior author from a country; the y-axis indicates the country. The ‘Other’ category includes preprints from all countries not listed in the plot. (c) Similar to panel b, but showing the total preprints listing at least one author from the country in any position, not just the senior position. (d) Proportion of total senior-author preprints from each country (y-axis) over time (x-axis), starting in November 2013 and continuing through December 2019. Each colored segment indicates the proportion of total preprints attributed to a single country (using same color scheme as panels (b and c), as of the end of the month indicated on the x-axis. 50 Over time, the country-level proportions on bioRxiv have remained remarkably stable (Figure 2-1d), even as the number of preprints grew exponentially. For example, at the end of 2015 Germany accounted for 4.7% of bioRxiv’s 2460 manuscripts, and at the end of 2019 it was responsible for 5.4% of 67,885 preprints. However, the proportion of preprints from countries outside the top seven contributing countries is growing slowly (Figure 2-1d): from 19.4% at the end of 2015 to 23.1% at the end of 2019, by which time bioRxiv was hosting preprints from senior authors affiliated with 136 countries. Country Preprints, senior author Preprints, any author United States 26,598 (39.2%) 34,676 (51.1%) United Kingdom 7151 (10.5%) 11,578 (17.1%) (Unknown) 4985 (7.3%) 17,635 (26.0%) Germany 3668 (7.3%) 7157 (10.5%) France 2863 (4.2%) 5218 (7.7%) China 2778 (4.1%) 4609 (6.8%) Canada 2380 (3.5%) 4409 (6.5%) Australia 1755 (2.6%) 3260 (4.8%) Switzerland 1364 (2.0%) 2779 (4.1%) Netherlands 1291 (1.9%) 2764 (4.1%) Japan 1263 (1.9%) 2287 (3.4%) India 1212 (1.8%) 1769 (2.6%) Table 2-1. Preprints per country. All 11 countries with more than 1000 preprints attributed to a senior author affiliated with that country. The percentages in the ‘Preprints, any author’ column sum to more than 100% because preprints may be counted for more than one country. A full list of countries is provided in Supplementary Table B-1. 51 Preprint adoption relative to overall scientific output We noted that some patterns may be obscured by countries that had hundreds or thousands of times as many preprints as other countries, so we re-evaluated these ranks after adjusting for overall scientific output (Figure 2-2a). The corrected measurement, which we call “bioRxiv adoption,” is the proportion of preprints from each country divided by the proportion of worldwide research outputs from that country (see Methods). The US posted 26,598 preprints and published about 3.5 million citable documents, for a bioRxiv adoption score of 2.31 (Figure 2-2b). Nine of the 12 countries with adoption scores above 1.0 were from North America and Europe, but Israel has the third-highest score (1.67) based on its 640 preprints. Ethiopia has the 10th-highest bioRxiv adoption (1.08): though only 57 preprints list a senior author with an affiliation in Ethiopia, the country had a total of 15,820 citable documents published between 2014 and 2019 (Supp. Table B-2). 52 Figure 2-2. BioRxiv adoption per country. (a) Correlation between two scientific output metrics. Each point is a country; the x-axis (log scale) indicates the total citable documents attributed to that country from 2014 to 2019, and the y-axis (also log scale) indicates total senior-author preprints attributed to that country overall. The red line demarcates a ‘bioRxiv adoption’ score of 1.0, which indicates that a country’s share of bioRxiv preprints is identical to its share of general scholarly outputs. Countries to the left of this line have a bioRxiv adoption score greater than 1.0. A score of 2.0 would indicate that its share of preprints is twice as high as its share of other scholarly outputs (See Discussion for more about this measurement.) (b) The countries with the 10 highest and 10 lowest bioRxiv adoption scores. The x-axis indicates each country’s adoption score, and the y-axis lists each country in order. All panels include only countries with at least 50 preprints. 53 By comparison, some countries are present on bioRxiv at much lower frequencies than would be expected, given their overall participation in scientific publishing (Figure 2-2b): Turkey published 249,086 citable documents from 2014 through 2019 but was the senior author on only 80 preprints, for a bioRxiv adoption score of 0.10. Russia (283 preprints), Iran (123 preprints) and Malaysia (78 preprints) all have adoption scores below 0.18. The largest country with a low adoption score is China (3,176,571 citable documents; 2778 preprints; bioRxiv adoption = 0.26), which published more than 15% of the world’s citable documents (according to SCImago) but was the source of only 4.1% of preprints (Figure 2-2a). Patterns and imbalances in international collaboration After analyzing preprints using senior authorship, we also evaluated interactions within manuscripts to better understand collaborative patterns found on bioRxiv. We found the number of authors per paper increased from 3.08 in 2014 to 4.56 in 2019 (Supp. Figure B-1). The monthly average authors per preprint has increased linearly with time (Pearson’s r = 0.949, p=8.73×10-38), a pattern that has also been observed, at a less dramatic rate, in published literature (Adams et al. 2005; Wuchty, Jones, and Uzzi 2007; Bordons, Aparicio, and Costas 2013). Examining the number of countries represented in each preprint (Supp. Figure B-1), we found that 24,927 preprints (36.7%) included authors from two or more countries; 3041 preprints (4.5%) were from four or more countries, and one preprint, “Fine- mapping of 150 breast cancer risk regions identifies 178 high confidence target genes,” listed 343 authors from 38 countries, the most countries listed on any single preprint. The 54 mean number of countries represented per preprint is 1.612, which has remained stable since 2014 despite steadily growing author lists overall. We then looked at countries appearing on at least 50 international preprints to examine basic patterns in international collaboration. We found that several countries with comparatively low output contributed almost exclusively to international collaborations: for example, of the 76 preprints that had an author with an affiliation in Vietnam, 73 (96%) had an author from another country. Upon closer examination, we found these countries were part of a larger group, which we call “contributor countries,” that (1) appear mostly on preprints with authors from other countries, but (2) seldom as the senior author. For this analysis, we defined a contributor country as one that has contributed to at least 50 international preprints but appears in the senior author position of less than 20% of them. (We excluded countries with less than 50 preprints to minimize the effect of dynamics that could be explained by countries with just one or two labs that frequently worked with international collaborators.) 17 countries met these criteria (Supp. Figure A-3): for example, of the 84 international preprints that had an author with an affiliation in Uganda, only 5 (6%) had an author from Uganda in the senior author position. This figure was also less than 12% for Vietnam, Tanzania, Slovakia and Indonesia: by comparison, the figure for the US was 48.7% (Figure 2-3a). 55 Figure 2-3. Contributor countries. (a) Bar plot indicating the international senior author rate (y-axis) by country (x-axis) – that is, of all international preprints with a contributor from that country, the percentage of them that include a senior author from that country. All 17 contributor countries are listed in red, with the five countries with the highest senior- author rates (in grey) for comparison. (b) A bar plot with the same y-axis as panel (a). The x-axis indicates the international collaboration rate, or the proportion of preprints with a contributor from that country that also include at least one author from another country. (c) is a bar plot indicating the total international preprints featuring at least one author from that country (the median value per country is 19). (d) On the left are the 17 contributor countries. On the right are the countries that appear in the senior author position of preprints that were co-authored with contributor countries. (Supervising countries with 25 or fewer preprints with contributor countries were excluded from the figure.) The width of the ribbons connecting contributor countries to senior-author countries indicates the number of preprints supervised by the senior-author country that included at least one author from the contributor country. Statistically significant links were found between four 56 combinations of supervising countries and contributors: Australia and Bangladesh (Fisher’s exact test, q = 1.01 × 10−11); the UK and Thailand (q = 9.54 × 10−4); the UK and Greece (q = 6.85 × 10−3); and Australia and Vietnam (q = 0.049). All p-values reflect multiple-test correction using the Benjamini–Hochberg procedure. In addition to a high percentage of international collaborations and a low percentage of senior-author preprints, another characteristic of contributor countries is a comparatively low number of preprints overall. To define this subset of countries more clearly, we examined whether there was a relationship between any of the three factors we identified across all countries with at least 50 international preprints. We found consistent patterns for all three (see Methods). First, countries with fewer international collaborations also tend to appear as senior author on a smaller proportion of those preprints (Supp. Figure B-3a). Second, we also observed a negative correlation between total international collaborations and international collaboration rate—that is, the proportion of preprints a country contributes to that include at least one contributor from another country (Supp. Figure B-3b). This indicates that countries with mostly international preprints (Figure 2- 3b) also tended to have fewer international collaborations (Figure 2-3c) than other countries. Third, we found a negative correlation between international collaboration rate and the proportion of international preprints for which a country appears as senior author (Supp. Figure B-3c), demonstrating that countries that appear mostly on international preprints (Figure 2-3b) are less likely to appear as senior author of those preprints (Figure 2-3). Similar patterns have been observed in previous studies: González-Alcaide et al., 2017 found countries ranked lower on the Human Development Index participated more 57 frequently in international collaborations, and a review of oncology papers found that researchers from low- and middle-income countries collaborated on randomized control trials, but rarely as senior author (Wong et al. 2014). After generating a list of preprints with authors from contributor countries, we examined which countries appeared most frequently in the senior author position of those preprints (Figure 2-3d). Among the 2133 preprints with an author from a contributor country, 494 (23.2%) had a senior author listing an affiliation in the US (Supp. Table B-3). The UK was listed as senior author on the next-most preprints with contributor countries, at 318 (14.9%), followed by Germany (4.2%) and France (3.1%). Given the large differences in preprint authorship between countries, we tested which of these senior-author relationships was disproportionately large. Using Fisher’s exact test (see Methods), we found four links between contributor countries and senior-author countries that were significant (Supp. Table B-4). The strongest link is between Bangladesh and Australia: of the 82 international preprints with an author from Bangladesh, 22 list a senior author with an affiliation in Australia. Authors in Vietnam appear with disproportionate frequency on preprints with a senior author in Australia as well (9 of 67 preprints). The other two links are to senior authors in the UK, with contributing authors from Thailand (50 of 187 preprints) and Greece (41 of 155 preprints). 58 Differences in preprint downloads and publication rates After quantifying which countries were posting preprints, we also examined whether there were differences in preprint outcomes between countries. We obtained monthly download counts for all preprints, as well as publication status, the publishing journal, and date of publication for all preprints flagged as “published” on bioRxiv (see Methods). We then evaluated country-level patterns for the 36 countries with at least 100 senior-author preprints. When evaluating downloads per preprint, we used only download numbers from each preprint’s first six months online, which would capture most downloads for most preprints (Abdill and Blekhman 2019b) while minimizing the effect of the long tail of downloads that would be longer for countries that were earlier adopters. Using this measurement, the median number of PDF downloads per preprint is 210 (Figure 2-4a). Among countries with at least 100 preprints, Austria has the highest median downloads per preprint, with 261.5, followed by Germany (235.0), Switzerland (233.0) and the US (233.0). Argentina has the lowest median, at 138.5 downloads; next-fewest were Taiwan (142), Brazil (145) and Russia (145). To examine whether these results were influenced by changes in downloads per preprint over time, we re-analyzed the data after dividing each preprint’s download count by the median download count of all preprints posted in the same month. The country-level medians of the adjusted downloads per paper are highly correlated (Spearman’s rho = 0.989, p=3.33×10-109) with the unadjusted median downloads per paper, indicating there is no influence from countries posting more preprints during times in which 59 many preprints were downloaded in general. Across all countries with at least 100 preprints, there was a weak correlation between total preprints attributed to a country and the median downloads per preprint (Figure 2-4b), and another correlation between median downloads per preprint and each country’s publication rate (Figure 2-4c). Figure 2-4. Preprint outcomes. All panels include countries with at least 100 senior- author preprints. (a) A box plot indicating the number of downloads per preprint for each country. The dark line in the middle of the box indicates the median, and the ends of each box indicate the first and third quartiles, respectively. “Whiskers” and outliers were omitted from this plot for clarity. The red line indicates the overall median. (b) A plot showing the relationship (Spearman’s ρ = 0.485, p=0.00274) between total preprints and downloads. Each point represents a single country. The x-axis indicates the total number of senior- author preprints attributed to the country. The y-axis indicates the median number of 60 downloads for those preprints. (c) A plot showing the relationship (Spearman’s ρ = 0.777, p=2.442 × 10−8) between downloads and publication rate. Each point represents a single country. The x-axis indicates the median number of downloads for all preprints listing a senior author affiliated with that country. The y-axis indicates the proportion of preprints posted before 2019 that have been published. (d) A bar plot indicating the proportion of preprints posted before 2019 that are now flagged as ‘published’ on the bioRxiv website. The x-axis (and color scale) indicates the proportion, and the y-axis lists each country. The red line indicates the overall publication rate. Next, we examined country-level publication rates by assigning preprints posted prior to 2019 to countries using the affiliation of the senior author, then measuring the proportion of those preprints flagged as “published” on the bioRxiv website. Overall, 62.6% of pre- 2019 preprints were published (Supp. Table B-5). Ireland had the highest publication rate (49/67 = 73.1%; Figure 2-4d), followed by New Zealand (100/142; 70.4%) and Switzerland (505/724; 69.8%). China (588/1355, 43.4%) had the lowest publication rate, with Iran and Taiwan also in the bottom three. Preprint publication patterns between countries and journals After evaluating the country-level publication rates, we examined which journals were publishing these preprints and whether there were any meaningful country-level patterns (Figure 2-5). We quantified how many senior-author preprints from each country were published in each journal and used the χ² test (with Yates’s correction for continuity) to examine whether a journal published a disproportionate number of preprints from a given country, based on how many preprints from that country were published overall. To 61 minimize the effect of journals with differing review times, we limited the analysis to preprints posted before 2019, resulting in a total of 23,102 published preprints. Figure 2-5. Overrepresentation of US preprints. (a) A heat map indicating all disproportionately strong (q < 0.05) links between countries and journals, for journals that have published at least 15 preprints from that country. Columns each represent a single country, and rows each represent a single journal. Colors indicate the raw number of preprints published, and the size of each square indicates the statistical significance of that link—larger squares represent smaller q-values. See Supplementary Table B-6 for the results of each statistical test. (b) A bar plot indicating the degree to which US preprints are over- or under-represented in a journal’s published bioRxiv preprints. The y-axis lists all the journals that published at least 15 preprints with a US senior author. The x-axis indicates the overrepresentation of US preprints compared to the expected number: for 62 example, a value of ‘0%’ would indicate the journal published the same proportion of US preprints as all journals combined. A value of ‘100%’ would indicate the journal published twice as many U. preprints as expected, based on the overall representation of the US among published preprints. Journals for which the difference in representation was less than 15% in either direction are not displayed. The red bars indicate which of these relationships were significant using the Benjamini–Hochberg-adjusted results from χ² tests shown in panel A. After controlling the false-discovery rate using the Benjamini–Hochberg procedure, we found 63 significant links between journals and countries, of journal–country links with at least 15 preprints (Figure 2-5a). 11 countries had links to journals that published a disproportionate number of their preprints, but the US had more links than any other country. 33 of the 63 significant links were between a journal and the US: the US is listed as the senior author on 41.7% of published preprints, but accounts for 74.5% of all bioRxiv preprints published in Cell, 72.7% of preprints published in Science, and 61.0% of those published in Proceedings of the National Academy of Sciences (PNAS) (Figure 2-5b). Discussion Our study represents the first comprehensive, country-level analysis of bioRxiv preprint publication and outcomes. While previous studies have split up papers into “USA” and “everyone else” categories in biology (Fraser et al. 2020) and astrophysics (Schwarz and Kennicutt 2004), our results provide a broad picture of worldwide participation in the 63 largest preprint server in biology. We show that the US is by far the most highly represented country by number of preprints, followed distantly by the UK and Germany. By adjusting preprint counts by each country’s overall scientific output, we were able to develop a “bioRxiv adoption” score (Figure 2-2), which showed the US and the UK are overrepresented while countries such as Turkey, Iran and Malaysia are underrepresented even after accounting for their comparatively low scientific output. Open science advocates have argued that there should not be a “one size fits all” approach to preprints and open access (Debat and Babini 2019; ALLEA 2018; Becerril-García 2019; Mukunth 2019), and further research is required to determine what drives certain countries to preprint servers— what incentives are present for biologists in Finland but not Greece, for example. There is also more to be done regarding the trade-offs of using a more distributed set of repositories that are specific to disciplines or countries (e.g. INA-Rxiv in Indonesia or PaleorXiv for paleontology), which could also influence the observed levels of bioRxiv adoption. Our results make it clear that those reading bioRxiv (or soliciting submissions from the platform) are reviewing a biased sample of worldwide scholarship. There are two findings that may be particularly informative about the state of open science in biology. First, we present evidence of contributor countries—countries from which authors appear almost exclusively in non-senior roles on preprints led by authors from more prolific countries (Figure 2-3). While there are many reasons these dynamics could arise, it is worth noting that the current corpus of bioRxiv preprints contains the same familiar disparities observed in published literature (Mammides et al. 2016; Burgman, Jarrad, and 64 Main 2015; Wong et al. 2014; González-Alcaide et al. 2017). Critically, we found the three characteristics of contributor countries (low international collaboration count, high international collaboration rate, low international senior author rate) are strongly correlated with each other (Figure 2-3). When looking at international collaboration using pairwise combinations of these three measurements, countries fall along tidy gradients (Supp. Figure B-2)—which means not only that they can be used to delineate properties of contributor countries, but that if a country fits even one of these criteria, they are more likely to fit the other two as well. Second, we found numerous country-level differences in preprint outcomes, including a positive correlation at the country level between downloads per preprint and publication rate (Figure 2-4c). This raises an important consideration that when evaluating the role of preprints, some benefits may be realized by authors in some countries more consistently than others. If one of the goals of preprinting one’s work is to solicit feedback from the community (Sarabipour et al. 2019; Sever et al. 2019), what are the implications of the average Brazilian preprint receiving 37% fewer downloads than the average Dutch preprint? Do preprint authors from the most-downloaded countries (mostly in western Europe) have broader social-media reach than authors in low-download countries such as Argentina and Taiwan? What role does language play in outcomes, and why do countries that get more downloads also tend to have higher publication rates? We also found some journals had particularly strong affinities for preprints from some countries over others: even when accounting for differing publication rates across countries, we found dozens of 65 journal–country links that disproportionately favored the US and UK. While it’s possible this finding is coincidental, it demonstrates that journals can embrace preprints while still perpetuating some of the imbalances that preprints could be theoretically alleviating. Our study has several limitations. First, bioRxiv is not the only preprint server hosting biology preprints. For example, arXiv’s Quantitative Biology category held 18,024 preprints at the end of 2019 (https://arxiv.org/help/stats/2019_by_area/index), and repositories such as Indonesia’s INA-Rxiv (https://osf.io/preprints/inarxiv/) hold multidisciplinary collections of country-specific preprints. We chose to focus on bioRxiv for several reasons: primarily, bioRxiv is the preprint server most broadly integrated into the traditional publishing system (Barsh et al. 2016; Vence 2017; eLife 2020). In addition, bioRxiv currently holds the largest collection of biology preprints, with metadata available in a format we were already equipped to ingest (Abdill and Blekhman 2019c). Analyzing data from only a single repository also avoids the issue of different websites holding metadata that is mismatched or collected in different ways. Comparing publication rates between repositories would also be difficult, particularly because bioRxiv is one of the few with an automated method for detecting when a preprint has been published. Second, this “worldwide” analysis of preprints is explicitly biased toward English-language publishing. BioRxiv accepts submissions only in English, and the primary motivation for this work was the attention being paid to bioRxiv by organizations based mostly in the US and western Europe. In addition, bibliometrics databases such as Scopus and Web of Science have well-documented biases in favor of English-language publications (Mongeon and 66 Paul-Hus 2016; Archambault et al. 2006; de Moya-Anegón et al. 2007), which could influence observed publication rates and the bioRxiv adoption scores that depend on scientific output derived from Scopus. There were also 4985 preprints (7.3%) for which we were not able to confidently assign a country of origin. An evaluation of these (see Methods) showed that the most prolific countries were also underrepresented in the “unknown” category, compared to the 148 other countries with at least one author. While it is impractical to draw country-specific conclusions from this, it suggests that the preprint counts for countries with comparatively low participation may be slightly higher than reported, an issue that may be exacerbated in more granular analyses, such as at the institutional level. Country-level differences in metrics such as downloads and publication rate may also be confounded with field-level differences: on average, genomics preprints are downloaded twice as many times as microbiology preprints (Abdill and Blekhman 2019b), for example, so countries with a disproportionate number of preprints in a particular field could receive more downloads due to choice of topic, rather than country of origin. Further study is required to determine whether these two factors are related and in which direction. In summary, we find country-level participation on bioRxiv differs significantly from existing patterns in scientific publishing. Preprint outcomes reflect particularly large differences between countries: comparatively wealthy countries in Europe and North America post more preprints, which are downloaded more frequently, published more consistently, and favored by the largest and most well-known journals in biology. While 67 there are many potential explanations for these dynamics, the quantification of these patterns may help stakeholders make more informed decisions about how they read, write and publish preprints in the future. Methods Ethical statement. This study was submitted to the University of Minnesota Institutional Review Board (study #00008793), which determined the work did not qualify as human subjects research and did not require IRB oversight. Preprint metadata. We used existing data from the Rxivist web crawler (Abdill and Blekhman 2019c) to build a list of URLs for every preprint on bioRxiv.org. We then used this list as the input for a new tool that collects author data: we recorded a separate entry for each author of each preprint, and stored name, email address, affiliation, ORCID identifier, and the date of the most recent version of the preprint that has been indexed in the Rxivist database. While the original web crawler performs author consolidation during the paper index process (i.e. “Does this new paper have any authors we already recognize?”), this new tool creates a new entry for each preprint; we make no connections for authors across preprints in this analysis and infer author country separately for every author of every paper. It is also important to note that for longitudinal analyses of preprint trends, each preprint is associated with the date on its most recent version, which means a paper first posted in 2015, but then revised in 2017, would be listed in 2017. The final version of the preprint metadata was collected in the final weeks of January 2020—because 68 preprints were filtered using the most recent known date, those posted before 2020, but revised in the first month of 2020, were not included in the analysis. In addition, 95 preprints were excluded because the bioRxiv website repeatedly returned errors when we tried to collect the metadata, leaving a total of 67,885 preprints in the analysis. Of these, there were 2409 manuscripts (3.6%) for which we were unable to scrape affiliation data for at least one author, including 137 preprints with no affiliation information for any author. bioRxiv maintains an application programmatic interface (API) that provides machine- readable data about their holdings. However, the information it exposes about authors and their affiliations is not as complete as the information available from the website itself, and only the corresponding author’s institutional affiliation is included (https://api.biorxiv.org/). Therefore, we used the more complete data in the Rxivist database (Abdill and Blekhman 2019b), which includes affiliations for all authors. All data on published preprints was pulled directly from bioRxiv. However, it is also possible, if not likely, that the publication of many preprints goes undetected by its system. Fraser et al., 2020 developed a method of searching for published preprints in Scopus and Crossref databases and found most had already been picked up by bioRxiv’s detection process, though bioRxiv states that preprints published with new titles or authors can go undetected (https://www.biorxiv.org/about-biorxiv), and preliminary data suggests this may affect thousands of preprints (Abdill and Blekhman 2019b). How these effects differ by country of origin remains unclear—perhaps authors from some countries are more likely 69 to have their titles changed by journal editors, for example—but bias at the country level may also be more pronounced for other reasons. The assignment of Digital Object Identifiers (DOIs) to papers provides a useful proxy for participation in the “western” publishing system. Each published bioRxiv preprint is listed with the DOI of its published version, but DOI assignment is not yet universally adopted. Boudry and Chartron (2017) examined papers from 2015 indexed by PubMed and found DOI assignment varied widely based on the country of the publisher. 96% of publications in Germany had a DOI, for example, plus 98% of UK publications and more than 99% of Brazilian publications. However, only 31% of papers published in China had DOIs, and just 2% (33 out of 1582) of papers published in Russia. There are 45 countries that overlap between our analysis and that of Boudry and Chartron (2017). Of these, we found a modest correlation (Spearman’s rho = 0.295, p=0.0489) between a country’s preprint publication rate and the rate at which publishers in that country assigned DOIs (Supp. Table B-8). This indicates that countries with higher rates of DOI issuance (for publications dating back to 1955) also tend to have higher observed rates of preprint publication. Attribution of preprints. Throughout the analysis, we define the “senior author” for each preprint as the author appearing last in the author list, a longstanding practice in biomedical literature (Riesenberg and Lundberg 1990; Buehring, Buehring, and Gerard 2007) corroborated by a 2003 study, which found that 91% of publications indicated a corresponding author that was in the first- or last-author position (Mattsson, Sundberg, and Laget 2011). Among the 59,562 preprints for which the country was known for the first 70 and last author, 7965 (13.4%) preprints included a first author associated with a different country than the senior author. When examining international collaboration, we also considered whether more nuanced methods of distributing credit would be more informative. Our primary approach— assigning each preprint to the one country appearing in the senior author position—is considered straight counting (Gauffriau et al. 2008). We repeated the process using complete-normalized counting (Supp. Table B-7), which splits a single credit among all authors of a preprint. So, for a preprint with 10 authors, if six authors are affiliated with an institution in the UK, the UK would receive 0.6 “credits” for that preprint. We found the complete-normalized preprint counts to be almost identical to the counts distributed based on straight counting (Pearson’s r = 0.9971, p=4.48×10-197). While there are numerous proposals for proportioning differing levels of recognition to authors at different positions in the author list (e.g. Hagen 2013; Kim and Diesner 2015), the close link between the complete-normalized count and the count based on senior authorship indicates that senior authors are at least an accurate proxy for the overall number of individual authors, at the country level. When computing the average authors per paper, the harmonic mean is used to capture the average “contribution” of an author, as in Glänzel and Schubert (2005)—in short, this shows that authors were responsible for about one-third of a preprint in 2014, but less than one-fourth of a preprint as of 2019. 71 Data collection and management. All bioRxiv metadata was collected in a relational PostgreSQL database (https://www.postgresql.org). The main table, “article_authors,” recorded one entry for each author of each preprint, with the author-level metadata described above. Another table associated each unique affiliation string with an inferred institution (see “Institutional affiliation assignment” below), with other tables linking institutions to countries and preprints to publications. (See the repository storing the database snapshot for a full description of the database schema.) Analysis was performed by querying the database for different combinations of data and outputting them into CSV files for analysis in R (R Core Team 2017). For example, data on “authors per preprint” was collected by associating all the unique preprints in the “article_authors” table with a count of the number of entries in the table for that preprint. Similar consolidation was done at many other levels as well—for example, since each author is associated with an affiliation string, and each affiliation string is associated with an institution, and each institution is associated with a country, we can build queries to evaluate properties of preprints grouped by country. Contributor countries. The analysis described in the “Collaboration” section measured correlations between three country-level descriptors, calculated for all countries that contributed to more than 50 international preprints: i. International collaborations: The total number of international preprints including at least one author from that country. 72 ii. International collaboration rate: Of all preprints listing an author from that country, the proportion of them that includes at least one author from another country. iii. International senior-author rate: Of all the international collaborations associated with a country, the proportion of them for which that country was listed as the senior author. We examined disproportionate links between contributor countries and senior-author countries by performing one-tailed Fisher’s exact tests between each contributor country and each senior-author country, to test the null hypothesis that there is no association between the classifications “preprints with an author from the contributor country” and preprints with a senior author from the senior-author country.” To minimize the effect of partnerships between individual researchers affecting country-level analysis, the senior- author country list included only countries with at least 25 senior-author preprints that include a contributor country, and we only evaluated links between contributor countries and senior-author countries that included at least five preprints. We determined significance by adjusting p-values using the Benjamini–Hochberg procedure. BioRxiv adoption. When evaluating bioRxiv participation, we corrected for overall research output, as documented by SCImago Journal and Country Rank portal, which counts articles, conference papers, and reviews in Scopus-indexed journals (https://www.scimagojr.com; https://www.scimagojr.com/help.php). We added the totals of these “citable documents” from 2014 through 2019 for each country with at least 3000 citable documents and 50 preprints. We used these totals to generate a productivity- 73 adjusted score, termed “bioRxiv adoption,” by taking the proportion of preprints with a senior author from that country and dividing it by that country’s proportion of citable documents from 2014 to 2019. While SCImago is not specific to life sciences research, it was chosen over alternatives because it had consistent data for all countries in our dataset. A shortcoming of combining data SCImago and the Research Organization Registry (see below) is that they use different criteria for the inclusion of separate states. In most cases, SCImago provides more specific distinctions than ROR: for example, Puerto Rico is listed separately from the US in the SCImago dataset, but not in the ROR dataset. We did not alter these distinctions—as a result, nations with disputed or complex borders may have slightly inflated bioRxiv adoption scores. For example, preprints attributed to institutions in Hong Kong are counted in the total for China, but the 108,197 citable documents from Hong Kong in the SCImago dataset are not included in the China total. Visualization. All figures were made with R and the ggplot2 package (Wickham 2009), with colors from the RcolorBrewer package (Neuwirth 2014; Woodruff and Brewer 2017). World maps were generated using the Equal Earth projection (Šavrič, Patterson, and Jenny 2019) and the rnaturalearth R package (South 2017), following the procedure described in Le et al. (2020). Code to reproduce all figures is available on GitHub (Abdill, 2020; https://github.com/blekhmanlab/biorxiv_countries; copy archived at https://github.com/elifesciences-publications/biorxiv_countries). Institutional affiliation assignment. We used the Research Organization Registry (ROR) API to translate bioRxiv affiliation strings into canonical institution identities (Research 74 Organization Registry 2019). We launched a local copy of the database using their included Docker configuration and linked it to our web crawler’s container, to allow the two applications to communicate. We then pulled a list of every unique affiliation string observed on bioRxiv and submitted them to the ROR API. We used the response’s “chosen” field, indicating the ROR application’s confidence in the assignment, to dictate whether the assignment was recorded. Any affiliation strings that did not have an assigned result were put into a separate “unknown” category. As with any study of this kind, we are limited by the quality of available metadata. Though we can efficiently scrape data from bioRxiv, data provided by authors can be unreliable or ambiguous. There are 465 preprints, for example, in which multiple or all authors on a paper are listed with the same ORCID, ostensibly a unique personal identifier, and there are hundreds of preprints for which authors do not specify any affiliation information at all, including in the PDF manuscript itself. We are also limited by the content of the ROR system (https://ror.org/about/): Though there are tens of thousands of institutions in the dataset and its basis, the Global Research Identifier Database (https://www.grid.ac/stats), has extensive coverage around the world, the translation of affiliation strings is likely more effective for regions that have more extensive coverage. Country-level accuracy of ROR assignments. Across 67,885 total preprints, we indexed 488,660 total author entries, one for each author of each preprint. These entries each included one of 136,456 distinct affiliation strings, which we processed using the ROR API before making manual corrections. 75 We first focused on assigning countries to preprints that were in the “unknown” category. We started by manually adding institutional assignments to “unknown” affiliation strings that were associated with 10 or more authors. We then used sub-strings within affiliation strings to find matches to existing institutions, and finally generated a list of individual words that appeared most frequently in “unknown” affiliation strings. We searched this list for words indicating an affiliation that was at least as specific as a country (e.g. “Italian,” “Boston,” “Guangdong”) and associated any affiliation strings that included that word with an institution in the corresponding country. Finally, we evaluated any authors still in the “unknown” category by searching for the presence of a country-specific top-level domain in their email addresses—for example, uncategorized authors with an email address ending in “.nl” were assigned to the Netherlands. Generic domains such as “.com” were not categorized, except for “.edu,” which was assigned to the US. While these corrections would have negatively impacted the institution-level accuracy, it was a more practical approach to generate country-level observations. There were also corrections made that placed more affiliations into the “unknown” category—there is an ROR institution called “Computer Science Department,” for example, that contained spurious assignments. Prior to correction, 23,158 (17%) distinct affiliation strings were categorized as “unknown,” associated with 71,947 authors. After manual corrections, there were 20,099 unknown affiliation strings associated with 51,855 authors. 76 There were also corrections made to existing institutional assignments, which are used to make the country-level inferences about author location. It appears the ROR API struggles with institutions that are commonly expressed as acronyms—affiliation strings including “MIT,” for example, was sometimes incorrectly coded not as “Massachusetts Institute of Technology” in the US, but as “Manukau Institute of Technology” in New Zealand, even when other clues within the affiliation string indicated it was the former. Other affiliation strings were more broadly opaque— “Centre for Research in Agricultural Genomics (CRAG) CSIC-IRTA-UAB-UB,” for example. A full list of manual edits is included in the “manual_edits.sql” file. In total, 12,487 institutional assignments were corrected, affecting 52,037 author entries (10.6%). Prior to the corrections, an evaluation of the ROR assignments in a random sample (n = 488) found the country-level accuracy was 92.2 ± 2.4%, at a 95% confidence interval. After an initial round of corrections, the country-level accuracy improved to 96.5 ± 1.6%. (These samples were sized to evaluate errors in the assignment of institutions rather than countries, which, once institution-level analysis was removed from the study, became irrelevant.) After another round of corrections that assigned countries to 14,690 authors in the “unknown” category, we pulled another random sample of corrected affiliations— using a 95% confidence interval, the sample size required to detect 96.5% assignment accuracy with a 2% margin of error was calculated to be 325 (Naing, Winn, and Rusli 2006). Manually evaluating the country assignments of this sample showed the country- level accuracy of the corrected affiliations was 95.7 ± 2.2%. 77 Though preprints that were assigned a country could be categorized with high accuracy, we also sought to characterize the preprints that remained in the “unknown” category after corrections, to evaluate whether there was a bias in which preprints were categorized at all. Among the successfully classified preprints, the distribution across countries is heavily skewed—the 27 most prolific countries (15%) account for 95.3% of categorized preprints. Accordingly, characterizing the prevalence of individual countries would require an impractically large sample made up of a large portion of all uncategorized preprints. Instead, we split the countries into two groups: the first contained the 27 most prolific countries. The second group contained the remaining 148 countries, which account for the remaining 2960 preprints (4.7%). We used this as the prevalence in our sample size calculation. Using a 95% confidence interval and a precision of 0.00235 (half the prevalence), the sample size (with correction for a finite population of 4,985) was calculated to be 307 (Naing, Winn, and Rusli 2006). Within this sample, we found that preprints with a senior author in the bottom 148 countries were present at a prevalence of 12.6 ± 3.9%. Data availability All data has been deposited in a versioned repository at Zenodo.org. Source data files have been provided for all figures, along with the code used to generate each plot. Code used to collect and analyze data has been deposited at https://github.com/blekhmanlab/biorxiv_countries. 78 Chapter Three: Public human microbiome data are dominated by highly developed countries The content in this chapter is based on work currently in press at PLOS Biology. Copyright retained by the authors and available via Creative Commons Attribution License v4. Abdill RJ, Adamowicz EM and Blekhman R, 2022. Public human microbiome data are dominated by highly developed countries. PLOS Biology. DOI: 10.1371/journal.pbio.3001536. Please refer to Appendix C for supplementary materials. 79 Summary The importance of sampling from globally representative populations has been well established in human genomics. In human microbiome research, however, we lack a full understanding of the global distribution of sampling in research studies. This information is crucial to better understand global patterns of microbiome-associated diseases and to extend the health benefits of this research to all populations. Here, we analyze the country of origin of all 444,829 human microbiome samples that are available from the world’s three largest genomic data repositories, including the Sequence Read Archive (SRA). The samples are from 2,592 studies of 19 body sites, including 220,017 samples of the gut microbiome. We show that more than 71% of samples with a known origin come from Europe, the United States, and Canada, including 46.8% from the United States alone, despite the country representing only 4.3% of the global population. We also find that central and southern Asia is the most underrepresented region: Countries such as India, Pakistan, and Bangladesh account for more than a quarter of the world population but make up only 1.8 percent of human microbiome samples. These results demonstrate a critical need to ensure more global representation of participants in microbiome studies. Background A growing body of research shows the human microbiome has broad relevance to human health and disease. However, identifying the specific connections between the microbiome and human health requires a broad survey of both human populations and their most 80 common health conditions. Even among healthy individuals, human microbiome composition varies between populations in ways that are still being uncovered: Geography and geographic relocation has been found to have an influence on microbiome composition (Yatsunenko et al. 2012; Vangay et al. 2018; Kaplan et al. 2019), as have host genetic variation and ethnicity (Goodrich et al. 2014; Blekhman et al. 2015; Brooks et al. 2018). Diet (Johnson et al. 2019), lifestyle (Clemente et al. 2015) and patterns in antibiotic use (Forslund et al. 2013) have all been linked to microbiome composition, with other studies considering the influence of locational factors such as pollution (Mutlu et al. 2018). Even within countries, interacting factors such as income, race, and education have critical impacts on health outcomes that could be mediated by the human microbiome (Amato et al. 2021). Some microbiome studies have specifically collected and compared data from global sites (Fragiadakis et al. 2019; Groussin et al. 2021), but large gaps and disparities still exist in which microbiomes are being studied on a global scale. The human microbiome has been linked to a growing number of social, medical and economic factors not directly related to host genetics, which reinforces the urgent need to evaluate the microbiomes of many populations (Amato et al. 2021; Ishaq et al. 2021). Other genomics fields have developed similar gaps, in which disproportionate attention is paid to the majority populations of wealthy countries: Genome-wide association studies (GWAS), for example, have been primarily conducted in populations with European ancestry (Medina-Gomez et al. 2015; Gurdasani et al. 2019). As a result, polygenic risk scores (PRSs) from these studies have poorer accuracy when applied to non-European 81 groups, limiting the possible benefits of this research—including personalized medicine, early disease screening, and risk prediction—to European-descended populations (De La Vega and Bustamante 2018; Peterson et al. 2019; Cai et al. 2021). There has been a concerted effort in genomics to include non-European individuals in GWAS studies, concurrent with calls to build research infrastructure and capacity globally (Gurdasani et al. 2019). It is likewise critical to identify underrepresented populations and locations in both genomics and microbiome research; otherwise, the benefits of host–microbiome research may only extend to a subset of the global population. To investigate the geographic distribution of microbiome studies, we used metadata on all human microbiome datasets in the BioSample database, which includes metadata describing samples in the Sequence Read Archive (SRA), DNA Data Bank of Japan, and European Nucleotide Archive (Nakamura et al. 2013). Our data includes the country of origin and time of release for more than 444,000 samples, including both 16S amplicon sequencing and shotgun metagenomic sequencing, released over the last 11 years. These samples from the three largest genomic databases represent a large majority of all human microbiome samples that have been published. Results We downloaded metadata for 444,829 human microbiome samples across 19 body sites and 2,592 studies. This data is available from the BioSample database maintained by the National Center for Biotechnology Information (NCBI), which includes metadata 82 describing raw sequencing data deposited in multiple international repositories, including SRA (Barrett et al. 2012). While sample-level genomic sequencing data is uploaded to SRA, information such as geographic origin is saved separately to an entry in the BioSample database. Samples per country As expected, we found the number of human microbiome samples with publicly available data has been increasing over time, from three microbiome samples in 2010 to 123,302 in 2020, the first year in which more than 100,000 human microbiome samples were released (Supp. Figure C-1). Though there were microbiome studies conducted prior to 2010, that was the first year of the BioSample database, which all depositors must now use if they submit sequencing data to the Sequence Read Archive. The most commonly used attribute in this subset of samples is the geographic origin of the sample, which is available for 99.5% of samples (Supp. Table C-1). Using this attribute, we were able to determine the country of origin for 382,711 (86%) human microbiome samples (Figure 3-1a), which originated in 115 different countries. We found that 178,960 samples (40.2%) were from the United States, almost five times more than any other country (Table 3-1). China has the next-most samples, with 36,162 (8.1%), followed by the United Kingdom, Denmark, Australia, and the Netherlands. China is the only Asian country in the top 14; the first South American country is Chile, in 16th place with 3616 samples (0.8%). Malawi is the first African country, in 19th place with 3052 samples (0.7%). 83 Figure 3-1. Global microbiome representation. (a) Total samples by country. The color of each country indicates the total number of samples originating in that country. (b) Relative representation by country. The color of each country indicates its representation in human microbiome datasets, relative to its share of world population. Red colors mark countries that are overrepresented relative to their population, and blue colors mark countries that are underrepresented. Countries with zero samples in the dataset are marked with dark blue. (c) Cumulative microbiome samples by world region. The x-axis indicates the year; the y-axis indicates the cumulative microbiome samples available at the end of that year. Colors indicate the cumulative microbiome samples from each of the world regions specified in the legend. The colored bar to the right of the plot indicates the share of the world population living in each of the regions using the same colors. (d) Proportion of annual samples. The x-axis indicates the year, and the y-axis indicates the proportion of samples from each world region published in that year. Colors correspond to the world regions shown in panel C. 10 100 1000 10000 178960 Total samples 0 Representation a b d c Overrepresented Underrepresented 0 100,000 200,000 300,000 400,000 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 Sa m ple s ( cu m ula tiv e) Region Oceania Northern Africa and Western Asia Central and Southern Asia Australia/New Zealand Latin America and the Caribbean Sub−Saharan Africa Eastern and South−Eastern Asia Europe and Northern America 0.00 0.25 0.50 0.75 1.00 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 Sa m ple p ro po rti on (y ea r) 0 100,000 200,000 300,000 400,000 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 Sa m ple s ( cu m ula tiv e) Region Oceania Northern Africa and Western Asia Central and Southern Asia Australia/New Zealand Latin America and the Caribbean Sub−Saharan Africa Eastern and South−Eastern Asia Europe and Northern America 0.00 0.25 0.50 0.75 1.00 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 Sa m ple p ro po rti on (y ea r) W orld population 84 Position Country Samples Share 1 United States 178,960 40.2% unknown 62,118 14.0% 2 China 36,162 8.1% 3 United Kingdom 16,076 3.6% 4 Denmark 11,497 2.6% 5 Australia 9,266 2.1% 6 Netherlands 9,173 2.1% 7 Canada 8,829 2.0% 8 Finland 7,855 1.8% 9 Italy 6,265 1.4% 10 Germany 5,531 1.2% 11 Spain 5,517 1.2% 12 Sweden 5,248 1.2% 13 Israel 4,831 1.1% 14 New Zealand 4,354 1.0% 15 Japan 4,298 1.0% 16 Chile 3,616 0.8% 17 Bangladesh 3,502 0.8% 18 France 3,402 0.8% 19 Malawi 3,052 0.7% 20 India 2,997 0.7% Rest of world 52,280 11.8% Table 3-1. Samples per country. 85 We also evaluated patterns specific to body sites. The number of countries represented in each body site is roughly proportional to the number of overall samples, with the most frequently sampled body site, the human gut, also holding data from the most countries, 96 (Table 3-2). This number drops quickly, however: For example, there are 44 countries represented in the skin microbiome category, and only 22 in the nasopharyngeal microbiome. Even if we consider only the 115 countries that appear in this dataset, it appears most body sites exclude most countries. When we consider body sites per country, rather than countries per body site, we can also evaluate the best-characterized country- level microbiomes: China has samples in 17 of the 19 body sites, the most of any country (Supp. Table C-2), followed by the United States with 16. The first South American country on the list, Brazil, has only 9, and South Africa, the first African country, appears in 8 body sites. Next, we used these data to assess country-level patterns at the five most prevalent body sites: the gut, mouth, skin, vagina, and lung (Supp. Table C-3). The United States has the most samples in all five. The skin microbiome category differs notably from the overall top 10: Though the U.S. and China again appear at the top, the remainder of the top 10 includes Chile, Bangladesh, Papua New Guinea, Hong Kong, India, Puerto Rico, Australia and Peru. However, this is also the body site with the most lopsided difference between the United States and the rest of the world: The U.S. total (19,706 samples) is 12.6 times that of the number two country, China (1562 samples), and more than 50 times that of the 10th country, Peru (391 samples). 86 Body site Samples Countries gut 220,017 96 human metagenome* 69,697 58 oral 47,798 63 skin 36,593 44 vaginal 17,784 31 lung 17,307 30 nasopharyngeal 15,646 22 feces 6,858 13 reproductive system 3,180 6 blood 2,707 9 saliva 2,503 15 milk 2,060 9 urinary tract 1,187 4 tracheal 520 2 sputum 364 3 eye 359 8 semen 203 3 bile 45 2 skeleton 1 1 * Samples under the “human metagenome” label refer to an NCBI category that does not specify a particular body site. Table 3-2. Samples by body site. 87 Samples per country relative to population To examine patterns of under- and overrepresentation of countries, we compared human microbiome sample counts to each country’s population, according to United Nations estimates for 2020 (“World Population Prospects 2019, Online Edition. Rev. 1” 2019). The United States is dramatically overrepresented relative to its population: Though the country has about 4.3% of the global population, 40.2% of human microbiome samples originate there. Proportionally, Denmark is the most overrepresented country, with 11,497 samples from a country of about 5.8 million people (Figure 3-1b). Of the 235 countries and territories included in the United Nations population estimates, 120 have zero human microbiome samples available in these public databases. To gain a better understanding of global representation in microbiome research, we grouped countries using the eight United Nations Sustainable Development Goals regions (“SDG Indicators” n.d.). We found that 71.2 percent of samples with a known location come from Europe and Northern America, a region that holds only 14.3 percent of the world’s population (Table 3-3). Proportionally, Australia/New Zealand has the most lopsided presence in the database: The region’s 30.3 million people is 0.4 percent of the population, but account for 3.1 percent of samples (Figure 3-1c). Central/Southern Asia is the most underrepresented region: It holds 25.8 percent of the population but makes up only 1.8 percent of microbiome samples. Northern Africa and Western Asia is the next- most underrepresented region, followed by Sub-Saharan Africa, which is home to 14.0% of the world’s population but is the source of 4.2% of human microbiome samples. 88 Region Samples 2020 population (estimated, in thousands) % of samples % of samples (known location) % of population Representation proportion* Europe and Northern America 272,544 1,116,506 61.3% 71.2% 14.3% 4.97 Eastern and South- Eastern Asia 49,007 2,346,709 11.0% 12.8% 30.1% 0.43 Sub-Saharan Africa 18,651 1,094,366 4.2% 4.9% 14.0% 0.35 Latin America and the Caribbean 15,264 653,962 3.4% 4.0% 8.4% 0.49 Australia/New Zealand 13,620 30,322 3.1% 3.6% 0.4% 9.14 Central and Southern Asia 6,685 2,014,709 1.5% 1.7% 25.8% 0.07 Northern Africa and Western Asia 5,621 525,869 1.3% 1.5% 6.7% 0.22 Oceania 1,178 12,356 0.3% 0.3% 0.2% 1.94 Unknown 62,259 14.0% Least developed countries 15,254 1,057,438 3.4% 4.0% 13.6% 0.29 Rest of world 367,457 6,737,361 82.6% 96.0% 86.4% 1.11 Unknown 62,118 14.0% *Representation proportion calculated by dividing a regions percentage of known samples by its percentage of population. Table 3-3. Samples and population by region. These proportions indicate a person in Europe or Northern America is roughly 14 times more likely to be studied in a microbiome project than someone from Sub-Saharan Africa. The 47 countries on the United Nations list of “least developed countries” account for about 89 14 percent of the world’s population (“About LDCs” 2013), but 3.4 percent of microbiome samples; 29 of those countries have no samples at all (Supplementary Table 4). We also found that although samples from Europe and Northern America are over-represented, in recent years there is more representation for samples from other regions, most prominently eastern and south-eastern Asia (Figure 3-1d). Discussion Our results show that the global distribution of human microbiome sampling is heavily skewed towards North American and European populations, both in total samples (Figure 3-1a) and in samples adjusted for population (Figure 3-1b). The United States is by far the greatest contributor to the database (Table 3-1), though this is slowly beginning to change as other countries’ contributions grow (Figure 3-1d). This neglect of most of the world’s population represents a disparity in microbiome research that could limit the health benefits of microbiome research to those countries and populations whose microbiomes have been extensively sampled and studied. Since only a subset of the world’s populations are currently being studied, the associations between the microbiome and disease may not hold in undersampled populations (Gupta, Paul, and Dutta 2017; He et al. 2018). For example, Gupta et al. identified several differences in the microbiome of healthy individuals from various geographic locations and lifestyles across the globe; without a consistent “healthy” microbiome across global populations, identifying microbiome-disease associations is nearly impossible (Gupta, Paul, and Dutta 2017). He et al. also found that microbiome- 90 based models for predicting metabolic disease failed when applied to populations outside of the geographical location in which they were developed (He et al. 2018). Additionally, by only sampling a subset of the global population, the diseases studied in the context of the microbiome are limited to diseases which impact that subset. Helminth parasite infections, for example, are common in tropical and subtropical regions of the world, but rare in North American and European populations. Undersampling of the microbiota from populations where these infections are common has led to a lack of clear understanding of the role of the microbiome in helminth colonization and resistance (S. C. Lee et al. 2014). To ensure greater global equity in the benefits of microbiome research, many stakeholders—funders, researchers and journals, to name a few—should consider how to ethically prioritize and incentivize improved global representation of microbiome samples, as they have begun to do in genomics with efforts such as the H3Africa initiative (H3Africa Consortium et al. 2014). Others have also highlighted opportunities for growth in the microbiome field, such as developing infrastructure and processes in low-resource settings (Soo et al. 2017; Mulder et al. 2017), building more comprehensive microbial reference databases and pursuing more flexible and affordable sequencing technologies (Brewster et al. 2019). Importantly, this approach should be grounded in benefitting the populations and communities sampled, rather than simply using these microbiomes as a tool to improve health in North American and European countries (Benezra 2020; Delgado and Baedke 2021). Ongoing discussion of “helicopter research” (e.g. Haelewaters, Hofmann, and Romero-Olivares 2021) sheds light on ethical objections to “solving” research disparities 91 with what essentially becomes charity, rather than collaboration: Researchers from wealthy countries obtain funding to do research in developing countries, “helicopter in” to collect data, then leave to publish their papers (Rochmyaningsih 2018). The result is more data from that country, but as part of a project that may not address the problems and priorities of the country under study. Local researchers, if they are consulted at all, may be excluded from authorship on the papers that are then hidden behind paywalls, written in a language they may not speak—part of much broader issues in scientific communication (Amano, González-Varo, and Sutherland 2016; Ramírez-Castañeda 2020). Researchers from the so- called “Global North” (as we are) would benefit from deferring to experienced scientists in these countries to find out how to avoid common extractive tropes in imbalanced collaborations (e.g. Armenteras 2021; Haelewaters, Hofmann, and Romero-Olivares 2021). Research and discussion in other fields may also help scientists trying to build more inclusive research projects: Though there are no easy answers, essays in applied ecology (Nuñez et al. 2019; Baker, Eichhorn, and Griffiths 2019; Pettorelli et al. 2021), ocean science (Belhabib 2021), botany (Antonelli 2020), geography (Noxolo 2017; Eichhorn, Baker, and Griffiths 2020), and conservation (Hazlett et al. 2020), among many others, deal with the hallmarks and dangers of colonial science (de Vos 2020), and how researchers can change their approach to knowledge production. The reasons for, and solutions to, global disparities in scientific research go far beyond the scope of this paper, and indeed of the microbiome field. There are broader issues of global representation in science that we and others have discussed, for example, in terms of 92 authorship (Abdill, Adamowicz, and Blekhman 2020), language (Amano, González-Varo, and Sutherland 2016) and the makeup of editorial boards (Nuñez et al. 2019). The complex history and current conditions driving these disparities requires a comprehensive assessment of global socio-political factors that we, as biologists based in North America, are not able to fully address. However, the necessity of such an assessment to solve these problems illustrates an important possible reason that these problems continue to perpetuate. Most microbiome researchers are not trained in social or political science and lack the appropriate tools to assess and address these problems. The more intentional inclusion of social scientists in microbiome projects may help address not only country- level imbalances, but also remediate harmful conventions used to deal with other issues like race (De Wolfe et al. 2021). Despite ongoing challenges, there have been several recent success stories of microbiome initiatives set in, driven by, and focused on, countries and populations who have been historically left out of microbiome research. One such example is the recently convened Microbiome Task Force from the H3Africa Consortium; their goals are to harmonize and perform meta-analyses of microbiome data from H3Africa, build capacity and knowledge- sharing among members, and provide data analysis support to researchers (“Microbiome Working Group” 2020). The Pan-African Bioinformatics Network (H3ABioNet), which has worked extensively in genomics research capacity building in Africa, also recently hosted a hackathon wherein they began work on a data portal for African microbiome samples (Fadlelmola et al. 2021). In South America, the Brazilian Microbiome project and 93 the recently proposed Ecuadorian Microbiome project both seek to advance microbiome research capacity in their respective countries and create local infrastructure to support these goals (Pylro et al. 2016; Díaz et al. 2021). Initiatives such as H3Africa’s African Collaborative Center for Microbiome and Genomics Research (ACCME) (C. Adebamowo et al., n.d.) may be ideally positioned to make progress in these trends, though as research activity grows in these underrepresented countries, using public metadata may become a less viable measure of these disparities: ACCME’s two existing microbiome publications, for example, do not have information about data availability (Dareng et al. 2016; S. N. Adebamowo et al. 2017), and ongoing discussions about issues such as data sovereignty (Gewin 2021) raise important questions about whether making data publicly available is a just and sustainable approach to biomedical research in countries or populations with comparatively little power in the global research ecosystem (Fox 2020; Tsosie, Fox, and Yracheta 2021; “CARE Principles of Indigenous Data Governance — Global Indigenous Data Alliance” n.d.). There are several limitations to our study. Metadata quality is the primary hurdle in characterizing samples (Gonçalves and Musen 2019): For example, our results suggest data for some microbiome samples are misclassified as “Homo sapiens” data rather than “human metagenome” data, which makes them much more difficult to locate. As a result, some of the countries listed here with zero samples do have microbiome studies that were either submitted to databases that are challenging to access in bulk (e.g. Zenodo) or mislabeled in the SRA. However, the number of these misclassified samples is likely to be 94 minor, and given the magnitude of differences observed in our study, this is unlikely to affect our main results (see Methods). It is also possible that not all samples identified as human in this study are indeed from humans and could, for example, include studies using human gut microbiota transferred into mice. We also did not evaluate differences in host phenotypic information: Most samples are missing even basic information such as sex (77 percent missing) and age (79 percent missing), and the most prevalent tag indicating host health status, “host_disease,” is only available for 7.8 percent of samples (Supp. Table C- 1). Consequently, we do not have sufficient information to draw conclusions about differences in geographic distribution between “healthy” and “disease” samples. Though disease-specific analysis is beyond the scope of our dataset, it would be interesting to investigate differences in the types of microbiome studies, and the questions they ask, on a global scale: If the human microbiome is generally under-studied in a given country, it’s likely that diseases prevalent in that country may also be lacking information about microbiome associations. We have also limited our database search to three databases (Sequence Read Archive, DNA Data Bank of Japan, and European Nucleotide Archive); it is possible that different patterns of global representation are present in other databases, such as MG-RAST (Wilke et al. 2016) and gcMeta (Shi et al. 2019), though they are orders of magnitude smaller than the NCBI holdings. In addition, as it has been estimated that 20 percent of microbiome papers do not have publicly available data (Eckert et al. 2020), our study only examines the subset of microbiome studies that also shared their data in the largest international repositories. 95 Samples collected from the same host could occur in longitudinal studies or datasets in which biological replicates were submitted as separate BioSamples, a pattern that is difficult to evaluate across multiple studies that may identify subjects differently, if at all. If longitudinal studies happen more frequently in some regions than others, it’s possible that the reported proportions of samples between countries could differ from the proportions of human subjects. However, given the differences in sample numbers between countries, this is unlikely to change the main results from our study. Moreover, since we are using sample collection as a proxy for investment in microbiome research in a given country, the identity of the subject may not be as relevant—indeed, it is likely more costly to perform a longitudinal study with subject follow-up than it is to recruit more subjects for a single sample each. Still, if longitudinal sampling is more common in studies in North America and Europe (which seems likely, given the extensive infrastructure and funding required for following patients long-term), it is possible that the gap between the “Global North” and the rest of the world in terms of microbiome sampling is smaller than our results suggest, if we were to count subjects rather than samples. However, given the magnitude of the difference between countries in our study, we do not believe repeated sampling from the same individuals in the Global North alone can account for such drastic disparities in sample numbers. To conclude, we analyzed the geographic origins of almost a half-million samples from the largest genomic repositories in the world. We find evidence that the human microbiome field may be encountering some of the same flaws that arose in human genomics (Need 96 and Goldstein 2009; Popejoy and Fullerton 2016), in which much of the world is excluded and progress is focused on the priorities of the wealthy. The field would benefit from a more global perspective on investigating the human microbiome’s relationship to health and disease. Methods A list of samples was exported from the NCBI BioSample database (https://www.ncbi.nlm.nih.gov/biosample) using the search string “txid408170[Organism] AND biosample sra[filter] AND “public”[filter]”, which requests all samples classified under the “human gut metagenome” category in the NCBI Taxonomy (https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi). The resulting sample IDs and all associated tags were loaded into a PostgreSQL database. We repeated this for all categories described as human metagenomes (Table 3-2). We note that the term “human gut metagenome” does not describe the sequencing technique used to generate the microbiome data, including shotgun metagenomics and amplicon sequencing -- specifically, 301,700 samples (72.0%) are associated with sequencing runs that list the library strategy as “AMPLICON.” We then looked in other NCBI categories nested beneath the “organismal metagenomes” category that were not explicitly labeled “human” but were likely to contain some human samples (“Organismal Metagenomes” n.d.). We downloaded the metadata for samples classified under any NCBI category that was the “generic” version of a human one we’d 97 already collected—the “blood metagenome” category is the generic version of the “human blood metagenome” category, for example (Supplementary Table 5). We downloaded all sample data for any generic categories that had at least 1,000 samples, then evaluated the metadata to find which samples indicated they were taken from a human host. To do this, we used the value of the “host_taxid” field or, if that was blank, the value of “host,” to create a putative “host” value, and manually flagged any that explicitly indicated the sample was from a human—references to “human” or “homo sapiens,” for example, or if the host included words such as “patient” or “crew member” and did not indicate another species. We evaluated 4,395 unique “host” values for 173,038 samples and found 501 values assigned to 29,934 samples (17.3%) that indicated the host was a human. These were also included in the analysis. The sample data was collected between April and June 2021; to minimize the effect of collecting some body sites after others, only samples dated prior to 2021 were included here. We then used the NCBI eUtils API to find “runs” associated with each sample, so we could ensure all the BioSamples were associated with actual sequencing data. In the NCBI system, “runs” are the entities associated with sequencing data. We also used this API to obtain information on publication date, library strategy and the dates on which samples became publicly available. This resulted in a collection of 444,829 samples across 19 body sites (Table 3-2) after removing several hundred samples that were missing dates or sequencing data. 98 Representation proportions. To determine which countries were over- or underrepresented relative to their populations, we obtained the 2020 population estimates for all countries as estimated by the United Nations (“World Population Prospects 2019, Online Edition. Rev. 1” 2019). We used this to calculate two percentages for each country, one for the country’s share of the global population, and another for the country’s share of human microbiome samples. We then calculated a representation index: For countries with a higher sample percentage than population percentage, we divided the former by the latter to obtain a number indicating how many times more samples are present than expected. For countries with a lower sample percentage than population percentage, we took the negative reciprocal of this number, indicating (in negative numbers) the number one would have to multiply the sample count by to get the number that would be proportionally representative. The interim result leaves overrepresented countries with positive scores, and underrepresented countries with negative scores. After removing the scores for countries with 50 or fewer samples, we scaled the positive scores to fall between 0 and 100, and separately scaled the negative scores to fall between 0 and -100. We then plotted these on the map using the “pseudo_log” transformation to add more variation in the color-coding for the countries with middling scores. For the regional calculations (Figure 3-1c–d), we used top-level classifications from the same United Nations document. Antarctica is not included in a region, so those samples were added to the “Unknown” category for region- level calculations. 99 To better understand gaps in what data may be available outside of these large, centralized repositories evaluated here, we selected several countries with zero attributed samples and did a literature search to determine whether human microbiome studies had been performed there and, if so, where the data is stored. For example, we could not confirm any samples available from Kazakhstan (population 18.7 million), in central Asia, but a human gut microbiome study from there was published in 2020 (Yegorov et al. 2020); its raw sequencing data (but no phenotypic information) is available on Zenodo, a scientific data repository with many submissions but no way of searching for samples or projects. Another Kazakhstan microbiome study (Kushugulova et al. 2018) is linked to data in the Sequence Read Archive (BioProject PRJEB17632), but with incorrect metadata: samples are classified as human sequencing data, rather than metagenomic, an issue addressed directly in the SRA submission instructions (“BioSample: Organism Information” n.d.). In addition, geolocation metadata was submitted, but listed the country of origin as Germany, the location of the senior author (and presumably the sequencing center), rather than Kazakhstan, the geographical source of the sample, as requested by NCBI (NCBI n.d.). A study in Honduras (population 9.9 million) includes SRA data with accurate geolocation information (BioProject PRJEB31759), but the samples were again classified under “Homo sapiens” rather than “human metagenome” (Walters et al. 2020). All figures were made using R and the ggplot2 package (Wickham 2009). Maps use the Equal Earth projection (Šavrič, Patterson, and Jenny 2019) and the rnaturalearth R package (South 2017). 100 Data availability All data has been deposited at Zenodo.org and is available at https://doi.org/10.5281/zenodo.5351179. This repository also includes the code used for data collection, along with the code used to generate each plot. 101 Chapter Four: Compendium of 170,000 uniformly processed 16S rRNA human gut microbiome samples from public repositories The content in this chapter is based on work by Richard J. Abdill, Samantha Graham and Ran Blekhman. 102 Summary While other bioinformatic specialties such as RNA-seq can take advantage of large compendia for meta-analyses or as training data for machine learning tasks, there has been no comparable resource for the 16S rRNA sequencing technology commonly used to quantify microbiome composition. Such resources would be valuable for hypothesis generation and developing ecological models of the human microbiome, but the computational burden of building such a resource is a significant hurdle to teams that may have analytical expertise but lack bioinformatics experience. To help close this gap, we have processed, profiled and compiled 5.7 terabases of human gut microbiome sequencing data to help identify large-scale patterns in microbiome composition across more than 80 percent of all publicly available 16S rRNA amplicon sequencing samples. Here, we share a novel dataset of 170,000 human gut microbiome samples processed with a single pipeline and combined into a unified dataset larger than any in the world. We find Firmicutes are almost universally present in human gut samples, particularly Bacilli and Clostridia, and identify multiple microbial classes that play a central role in differentiating samples across many phenotypes and locations. We are optimistic this dataset will help teach us more about the microbial ecology of the human gut and serve as a springboard for future research. 103 Background The human microbiome, particularly that of the digestive system, is of increasing interest as an important factor in understanding the etiology of disease. Studies show gut microbes have an intricate and bidirectional relationship with their hosts: Some microbes help digest materials that the host is unable to degrade on its own (Flint et al. 2008); others secrete vitamins that would otherwise be missing from the host’s diet (Smith, McCoy, and Macpherson 2007). Certain taxa have been linked, with varying levels of confidence, to conditions such as cystic fibrosis (Dayama et al. 2020), obesity (Sze and Schloss 2016) and Alzheimer’s disease (Vogt et al. 2017). A recurring frustration in these studies is the statistical challenge of dealing with sparse, high-dimensional datasets generated by DNA- based surveys of the microbiome. Fecal samples may contain hundreds or thousands of detected microbial taxa, making it challenging to find the relevant signal in even the largest studies. A better understanding of broad patterns of covariance in the human microbiome would enable researchers to summarize their data in fewer dimensions, giving microbiome studies a similar boost as is offered by gene pathway analysis: Though there may be no relevant links to an individual taxon, patterns may be more apparent looking at larger groups of taxa. In an environment as noisy and complex as the human gut microbiome, important patterns of variation may only become apparent after collecting thousands or tens of thousands of samples: With 100 samples, it’s easy to dismiss a small fluctuation, even a “significant” one, as noise or coincidence. With 70,000 samples, however, even weak relationships can 104 be characterized with more certainty. Though large compendia such as recount2 (Collado- Torres et al. 2017) have been developed for transcriptomic analysis, the human microbiome field does not have a comparable database of samples processed using a uniform bioinformatics pipeline. Quantifying these patterns requires sample sizes that are practically impossible to achieve in a single study but have already been collected by the field at large. The National Institutes of Health have directly invested more than $1 billion in human microbiome research (NIH Human Microbiome Portfolio Analysis Team 2019), and raw data for tens of thousands of microbiome samples are uploaded to the National Center for Biotechnology Information’s Sequence Read Archive (SRA) every year (Abdill, Adamowicz, and Blekhman 2021), plus many more from other samples from collaborators in the International Nucleotide Sequence Database Collaboration (INSDC), which includes organizations in Japan and the European Union. Although these are world-class repositories for a huge variety of genomic data, this raw data is difficult to manage at scale: The primary utility of compendia such as recount2 is that these sequencing reads have already been processed, curated and combined into a unified dataset. To mitigate this gap in the bioinformatic capabilities of the human microbiome field, we present here a collection of more than 170,000 human gut microbiome samples processed using state-of- the-art tools and combined into a single dataset that we are optimistic will be useful not only as a target for meta-analyses, but as a training resource in machine learning projects that could push the limits of microbial ecology into new territory. 105 Results We started with a list of 245,627 entries categorized as “human gut metagenome” samples and used metadata and the results of genomic analysis to filter out non-applicable projects and samples: Some projects evaluating fungi used 18S amplicon sequencing rather than 16S, for example, and many older projects used pyrosequencing technologies that must be processed separately. We then applied quality filters and straightforward read-merging processes (see Methods) to generate project-level results, which were then combined into a single large dataset. The final compendium contains 170,964 samples across 483 projects, which together contained 4048 unique taxa. We estimate this represents 80.8 percent of all 16S human gut microbiome samples available from INSDC repositories, excluding samples using older pyrosequencing technology, which must be processed differently (Figure 4-1). For analysis, we filtered this dataset to exclude rare taxa and poorly characterized reads (see Methods), leaving 164,898 samples containing 501 taxa for analysis. Of the 482 projects, the largest contained 4667 samples, and the median was 157.5 samples. 106 Figure 4-1. Sample pipeline progress. This figure visualizes how many samples were present at each step of the dataset compilation process and the reasons samples were excluded. Sample composition Taxonomic assignment returned 20 microbial phyla, 32 classes, 77 orders, 125 families, and 380 genera, not counting “NA” entries in which reads are categorized if confident classification cannot be made at that taxonomic level. The most prevalent phylum was Firmicutes, which was present in 99.8 percent of samples (Figure 4-2) and had a mean abundance of 50.4 percent per sample. This appears to be driven in large part by the Lactobacillales order, which appears in 90.1 percent of samples. Samples had a median 245,627 “Human gut metagenome” samples 10,752 in projects with less than 50 samples 234,875 Samples entering pipeline 31,509 unresolved processing errors 31,887 non-applicable data (fungi, non-human host, ITS sequencing, etc.) 171,479 Samples completed pipeline 807 removed during quality filtering and trimming 170,672 Samples in final compendium 107 Shannon diversity index of 1.745, which we found to follow a left-skewed distribution favoring diversity indices nearer to 1 (Figure 4-3A). After merging and filtering, samples contained a median of 36,968 reads. We found this distribution was essentially log-normal (Figure 4-3B), following a heavily right-skewed and covering samples that contained as many as 9.75 million reads. There are broad differences across microbiome samples—the average sample contains 55.8 ± 36 genus categorizations (of 501 observed genera) and 27.1 ± 14 families (of 163 observed). However, at the phylum level, there are several patterns that quickly emerge (Figure 4-3C): As many others have observed (e.g. Magne et al. 2020), samples can be immediately categorized by Firmicutes and Bacteroidota, the two most abundant phyla in the compendium. The next-most variance is observed among members of the Proteobacteria phylum, though not notorious members such as Salmonella and Campylobacter. A small subset of samples (the far-left side of Figure 4-3C) is composed almost entirely of other phyla; some of these may be unusual human samples (i.e. from very ill subjects, or meconium from newborns), and others may be misclassified samples from non-human hosts that could be filtered out by further analysis. 108 Figure 4-2. Most prevalent bacteria at three taxonomic levels. The vertical axis indicates the taxon, and the horizontal axis indicates the total number of samples in which that taxon appeared. Colors indicate the bacterial phylum to which the taxon belongs; colors are defined in the top panel and apply to the five most prevalent taxa in the dataset. Campilobacterota Fusobacteriota Cyanobacteria Bacteria (Unassigned) Verrucomicrobiota Desulfobacterota Bacteroidota Actinobacteriota Proteobacteria Firmicutes 0 50,000 100,000 150,000 164,898 Ph ylu m Proteobacteria Alphaproteobacteria Verrucomicrobiota Verrucomicrobiae Desulfobacterota Desulfovibrionia Actinobacteriota Coriobacteriia Actinobacteriota Actinobacteria Firmicutes Negativicutes Bacteroidota Bacteroidia Proteobacteria Gammaproteobacteria Firmicutes Clostridia Firmicutes Bacilli 0 50,000 100,000 150,000 164,898 Cl as s Actinobacteriota Coriobacteriia Coriobacteriales Actinobacteriota Actinobacteria Bifidobacteriales Firmicutes Negativicutes Veillonellales Firmicutes Bacilli Erysipelotrichales Firmicutes Clostridia Peptostreptococcales Proteobacteria Gammaproteobacteria Enterobacterales Firmicutes Clostridia Oscillospirales Bacteroidota Bacteroidia Bacteroidales Firmicutes Clostridia Lachnospirales Firmicutes Bacilli Lactobacillales 0 50,000 100,000 150,000 164,898 Samples Or de r 109 Figure 4-3. Sample composition. A) A histogram of the Shannon diversity per sample. The x-axis indicates the Shannon diversity index of a single sample (calculated using relative abundances consolidated at the family level), and the y-axis indicates how many samples fall into each bin. The vertical red line indicates the median. B) A histogram of the read counts per sample, reflecting the final read counts after trimming, filtering and merging of paired-end reads. The x-axis indicates library size (using a log scale), and the y-axis indicates how many samples fall into each bin. The vertical red line indicates the median. C) A stacked bar plot indicating the phylum-level composition of samples. The x- axis indicates sample, and the y-axis indicates the relative abundance of each phylum. The colored bars each indicate a separate phylum; the five phyla with the highest variance are displayed individually, and all others are combined into an “other” category. (Because technical and print limitations prevent us from visualizing all samples here, this figure shows a random set of 5,000 samples.) median: 1.745 0 2,500 5,000 7,500 0 1 2 3 4 Shannon diversity Sa m ple s A median: 36,968 0 5,000 10,000 15,000 1,000 10,000 100,000 1,000,000 10,000,000 Merged reads (log) Sa m ple s B 0% 25% 50% 75% 100% Sample Re lat ive a bu nd an ce C Phylum FirmicutesBacteroidota Proteobacteria Actinobacteria Verrucomicrobiota other 110 Patterns of variation across the compendium We then used principal coordinates analysis to evaluate broad patterns in microbiome composition in more dimensions (Figure 4-4A). We were able to visualize all 164,898 samples across two axes that captured about 20.7% and 13.2% of overall variation, respectively. After plotting each sample along those axes, we then evaluated the sample- level abundance of bacterial classes to find which were most strongly correlated to those axes, revealing which taxa were most responsible for differentiating samples in the plot. The result is one of the most comprehensive visualizations of worldwide variation in the human gut microbiome. Classes such as Negativicutes and Clostridia (both of the phylum Firmicutes) pull samples in one direction, while taxa such as Fusobacteriia and Cyanobacteriia, both less commonly observed in the human microbiome at high abundances, pull samples in the opposite direction. Evaluating sample density helps explain this dynamic (Figure 4-4B): We find samples are heavily concentrated to the left of the origin, indicating a location that is negatively correlated with the less common taxa. Taxa such as Bacilli, Actinobacteria and Desulfovibrionia pull in directions orthogonal to this axis and appear to account for variation mostly independent of the others. 111 Figure 4-4. Principal coordinates analysis. A) A PCoA plot of 164,898 human gut microbiome samples. Each point represents a single sample; the color indicates the most abundant microbial phylum in the sample. B) A density plot indicating the number of samples in each region of the plot. Each hexagon represents a range of values along both axes, and the color indicates the samples within those bounds. In both plots, the lines indicate correlation of bacterial phyla with the two displayed components—that is, a long, horizontal line to the right indicates a strong positive correlation between that phylum and the “MDS1” axis, and a long, vertical line downward indicates a strong negative correlation with the “MDS2” axis. Diagonal lines indicate correlation with both axes. See Methods for the process used to generate this figure. 112 Methods We retrieved metadata for all BioSamples categorized in the NCBI Taxonomy (Schoch et al. 2020) under “human gut metagenome” on 9 October 2021. After removing samples that could not be associated with a BioProject or sequencing run, selected only those for which the library source was “genomic” or “metagenomic,” excluding the values “metatranscriptomic,” “transcriptomic,” “viral RNA,” “synthetic” and “other.” Of these, we then limited the dataset to samples with a “library strategy” value of “amplicon” (and not values such as “WGS” and “RNA-Seq”). This left 245,627 samples across 1,437 projects. We then excluded projects that contained less than 50 samples meeting our criteria, leaving us with a list of 234,875 samples in 811 projects. Because pyrosequencing technologies developed by companies such as 454 Life Sciences and Ion Torrent require different processing, we then sought to remove projects containing pyrosequencing data. We used the SRA Toolkit APIs to retrieve sequencing instrument information for each sample and evaluated projects that reported using 454 or Ion Torrent instruments. We found 10 such instruments (“454 GS FLX Titanium,” “Ion Torrent PGM,” “454 GS FLX,” etc.), but manual review revealed one instrument was consistently mislabeled: In projects such as PRJNA685914 and PRJNA605031, the sequencing instrument was reported as “454 GS” even though the authors report elsewhere in the project that they used sequencers such as Illumina’s MiSeq platform, which we were not attempting to remove. More careful examination revealed these projects performed their analysis using Mothur (Schloss et al. 2009), a popular microbiome analysis tool, and “454 113 GS” is Mothur’s default entry for the “instrument” field when uploading to SRA (Westcott and Schloss 2021). To avoid removing applicable samples, we removed projects reporting using any of the pyrosequencing instruments except for “454 GS.” Sample retrieval. Samples were exported from the database one project at a time; each project had a file listing all accessions of runs associated with samples meeting the criteria described above. This file was used as the input for the “fasterq-dump” tool (SRA Tools Wiki n.d.) from the SRA Toolkit maintained by NCBI. This tool downloads the data from the Sequence Read Archive and splits the information into FASTQ files for downstream processing. We used samples from all available INSDC members—of the completed samples, 126,452 were from the Sequence Read Archive (74 percent), 38,971 were from European Nucleotide Archive, and 5249 were from the DNA Data Bank of Japan. Amplicon processing. If the number of files for forward reads matched the number of files for reverse reads, we processed the project as paired-end sequencing. If there was a mismatch, or there were no reverse reads, we processed the project as single-ended data. In both cases, we used DADA2 to process the data (B. J. Callahan et al. 2016). We used very general settings that we believed would be effective across as many studies as possible: We did not trim a set number of bases from either end, nor did we limit the maximum length of a read. We removed reads shorter than 20 nucleotides, reads with any ambiguous (“N”) base calls, and any reads that aligned to the phiX genome (almost certainly present as a control in Illumina sequencing runs). We also disabled quality-based truncation of reads. Paired-end reads were merged with a minimum overlap of 20 bases. In 114 some cases, the process of merging reads failed, and close to zero forward reads were merged with their paired reverse read. While this likely indicates a sequencing strategy that involves non-overlapping reads (or reads with very minimal overlap), we discarded the reverse reads rather than concatenating them, to avoid situations in which merging failed because of low-quality calls or mismatched forward and reverse read files. In those cases, the reverse reads were removed, and these projects were re-processed as single-ended data. If the number of forward reads did not match the number of reverse reads in a sample, we attempted to use DADA2 to detect the sequence identifier field in the FASTQ file to match the samples that could be salvaged. If this was unsuccessful, we removed the reverse reads and reprocessed these as single-ended data as well. Pipeline success. Due to resource constraints, we did not attempt to process projects with fewer than 50 samples, which accounted for 10,752 out of 245,627 samples (Figure 1). Of the 234,875 samples we processed, we found 31,887 samples (13.6 percent) contained non- applicable data—projects that targeted fungi or archaea, for example, and projects that used shotgun or nanopore sequencing instead of 16S amplicon sequencing. Another 31,509 samples (13.4 percent) were excluded because acceptable results they could not be processed using the automated pipeline: DADA2 identified excessive chimeric reads in some projects, for example, and we excluded any project in which, in a set of 10 samples, five or more samples in which more than 25 percent of their reads were identified as chimeras. Several projects were also excluded because they contained samples associated with multiple sequencing runs or were associated with samples that could not be 115 downloaded. 807 samples were removed from projects because all their reads were filtered out. Fraction of all samples in compendium. Of the 234,875 samples entering the pipeline, we found 13.6 percent were not applicable, and of the 171,479 that were successfully processed, 807 (0.57 percent) were removed by quality filters (Figure 1). Using these two factors, we can extrapolate that of the 245,627 samples categorized as “human gut metagenome,” 212,280.3 samples are applicable, and 211,281.3 would pass quality filtering. This means that our final compendium of 170,672 samples represents about 80.8 percent of all available 16S human gut microbiome samples. Analysis. We combined project-level taxonomy tables into one large matrix containing 170,672 samples for rows (from 483 projects) and 4048 unique taxonomic identifiers for columns, making up the compendium. However, not all samples available in the compendium were used in the analysis presented here: We removed 4732 samples with fewer than 1000 reads total, and 3547 taxa that appeared in fewer than 1000 samples (0.6 percent). After taxa were removed, we again evaluated the sample read counts and removed another 32 samples that then had less than 1000 reads. We then looked at taxonomic assignments in each sample and removed 1010 samples for which more than 10% of all reads were lacking a phylum-level assignments. The result was a dataset with 164,898 samples from 482 projects, containing 501 taxa. 116 Principal coordinates analysis. We began by building a distance matrix between all points. We used the Aitchison distance by using the centered log-ratio transformation on read counts then calculating the Euclidean distance between each sample. We added a pseudo-count of 1 to each zero value in the taxonomy table to accommodate the logarithms in the CLR transformation; computational limitations prevented us from using a more sophisticated method of imputation because of the size of the taxonomy table. We then performed multidimensional scaling (Figure 4-4) using the “divide and conquer” approach described by Delicado and Pachon-Garcia (2020). We extracted 8 principal coordinates and used 16 points as the overlap between partitions. Limitations. The use of a “one size fits all” pipeline has several drawbacks related not to the automated nature of the approach, but to the lack of information about the sequencing strategy (or strategies) employed by each project. For example, the merging process in paired-end datasets would be much more effective with more knowledge of study design and primer choices, particularly in cases where the amplicon length was greater than the read length and the reads did not overlap. In addition, DADA2 recommends building separate error models for each sequencing run (B. Callahan n.d.), but only project could be reliably inferred. There may also be samples in the compendium that are not from humans, but closely resemble human microbiome samples in their makeup. Though we removed obviously non-applicable samples (18S sequencing datasets of the mycobiome, for example), we did not pursue more stringent filtering to minimize the number of legitimate samples removed from the compendium. Our intent was to give researchers as much 117 information as possible to allow future studies to filter the full dataset as stringently as required. Discussion This dataset presents several opportunities to advance our understanding of the human gut microbiome. Most straightforwardly, it provides an unprecedented amount of information about microbiome composition in humans around the world, which can be evaluated to understand global variation in microbial communities and provide more context for the identification of aberrant patterns and conditions. Some of this analysis may be hindered by the inconsistent quality of sample-level metadata, a long-standing weakness of data sourced in bulk from public repositories (e.g. Gonçalves et al. 2017). However, new techniques are being developed to use natural language processing to augment this information (Hawkins et al. 2021), and imputation of missing fields may be possible—if only 1 percent of samples are tagged with a property, for example, that would still mean the compendium holds about 1,700 samples for training a classifier. In addition, we believe this dataset holds potential as a training set for the next generation of machine learning analyses of the human microbiome, enabling the creation of models that can be applied to future studies that will not have the sample-size advantage of this compendium. The field is only now beginning to explore techniques such as neural networks that others have been using for several years, but application of these approaches has so far been limited. 118 This database will bridge a key gap in meta-analysis of microbiome data: raw data requires computationally intensive processing prior to analysis, and the results published with papers (generally relative abundance tables, if available) are challenging to combine because of differences across studies in how data is processed (Prodan et al. 2020). By processing all samples using a uniform pipeline and developing a more standardized set of metadata fields, we will greatly increase the findability and interoperability of available sequencing data. By unifying so many isolated datasets, this compendium will enable analysis of patterns and associations within the human microbiome at an unprecedented scale. (See Chapter Five for further discussion of opportunities for analysis of this dataset.) Data availability The compendium, including sequencing data and metadata database, will be publicly shared when further analysis is complete and has shared online as a preprint. 119 Chapter Five: Discussion Summary of results Here, I presented work using publicly available data to shed light on the expansion of biology preprints and human microbiome research. In Chapter 1, we investigated the growth of bioRxiv.org, the largest preprint repository in the life sciences. We described the exponential growth of monthly preprint submissions, led by neuroscience, bioinformatics, and genomics. We also found that annual preprint publication rates may be as high as 90 percent, led by fields such as evolutionary biology and genomics, with fields such as scientific communication falling behind the average. Scientific Reports and eLife published the most preprints, and journals such as GigaScience, Genome Biology had the highest proportion of their publications first appear as preprints. We found median downloads per preprint were highest for fields such as genomics, synthetic biology and bioinformatics, and lowest in smaller fields such as physiology, pathology and immunology, though it should be noted that this research was done prior to the launch of medRxiv, the sister site to bioRxiv for biomedical research, and prior to the COVID-19 pandemic, which dramatically changed the way preprints in those fields were used and who was using them (Fraser et al. 2021). Chapter 2 expanded on this by annotating each preprint with information about the affiliations of its authors, which we used to evaluate patterns in international authorship and collaboration among 67,885 preprints. We found that a number of countries with 120 comparatively low output contributed almost exclusively to international collaborations: Authors in countries such as Vietnam almost never appear on preprints without international collaborators, part of a larger pattern of “contributor countries,” which appear mostly in international collaborations but almost never as senior author—for example, of the 84 international preprints with an author from Uganda, only 5 had a senior author there. We also found many journals with an exceptionally strong preference toward preprints with a senior author in the United States: Almost all preprints published by Cell, Genetic Epidemiology and Plant Direct, for example, were from the United States, while authors from other countries appeared more frequently in journals such as Biology Open, Frontiers in Immunology and Royal Society Open Science. Chapter 3 examined a similar dynamic at play in human microbiome research: We found 71% of publicly available human microbiome samples with a known origin come from Europe, the United States, and Canada, including 46.8% from the United States alone, despite the country representing only 4.3% of the global population. More critically, we found countries in central Asia such as India, Pakistan, and Bangladesh account for more than a quarter of the world population but make up only 1.8 percent of human microbiome samples. Western Asia is also underrepresented, as is much of Africa, a critical shortcoming as our understanding of the human microbiome moves closer to the development of clinical interventions that may be developed in a context that excludes much of the world. 121 Chapter 4 described a compendium of 170,964 human gut microbiome samples processed with a uniform pipeline and brought together into a single dataset. We used this dataset to characterize broad patterns in microbiome composition: We found the Firmicutes phylum was almost universally present, though at levels ranging almost linearly between 1 percent and 100 percent. The Bacilli and Clostridia classes were the most prevalent classes. We also used multidimensional scaling to evaluate the major axes of variation across all samples and found 20 percent of global microbiome differences could be attributed mostly to eight classes, including Bacteroidia, Desulfovibrionia and Negativicutes in one direction and Fusobacteriia, Cyanobacteriia and Campylobacteria in the other. Another 13 percent of variation appears to be attributable mostly to levels of Actinobacteria and Bacilli, though this only scratches the surface of the analyses we hope will be enabled by the compilation of this dataset. Future work: bibliometrics The continued analysis of preprints and their place in broader patterns in publishing is critical to building a more equitable system of sharing and evaluating research, and our work has already been used to inform policy discussions and proposals. The research in Chapter 1 was highlighted in the “Plan U” policy proposal calling for universal access to research (Sever, Eisen, and Inglis 2019). It was also cited by a journal announcing an official policy to accept preprints as submissions (Poremski et al. 2019) and by the Council of Science Editors, which used it to inform revisions to their long-standing “White Paper 122 on Publication Ethics” to include guidance on preprint servers (Cox 2019). Encouraging similar analyses going forward would provide scientists, publishers, policymakers, funders, advocates, and research institutions with actionable information about who is posting preprints, how they are benefiting from the practice, and, importantly, how these answers are changing over time. Data-driven investigations of policy changes by journals and funders is critical not only to evaluate progress toward their stated goals, but also to determine what side-effects appear over time. The ideals of open science are appealing in theory, but there is much to be done to determine their impact on the dissemination of research and how emerging tools such as preprints, open-source software projects, and open data initiatives are altering the research landscape and career trajectories of their practitioners. Future work: microbiome compendium My collaborators and I believe this dataset holds the potential for years of projects and investigations: Data mining. Rather than using the dataset to definitively answer questions about microbiome dynamics, researchers looking for patterns in the large, heterogeneous dataset can use their findings to determine “first-pass” inferences and generate hypotheses that can later be evaluated with a more targeted approach (Bhandary et al. 2018). The quantification (and correction) of batch effects is one particularly appealing application of this dataset: Even factors as mundane as sample storage (Blekhman et al. 2016) and choice of DNA 123 extraction kit (Olomu et al. 2020) have unique effects on observed microbiome composition; if annotations can be improved for an adequate number of samples, the imputation of information across the full compendium may yield useful insights not previously possible in studies of far fewer samples. Data accessibility. The compendium, particularly its database of sample metadata, provides a foundation for future work improving the adherence of biomedical datasets to the FAIR principles of data management (Wilkinson et al. 2016), which will significantly increase the utility of microbiome data in the SRA and make its data findable, accessible, interoperable, and reusable to a broader set of researchers. For example, one of the defined components of the principle of findability is that “metadata clearly and explicitly include the identifier of the data they describe,” a critical requirement for which the SRA currently falls short. SRA metadata is spread across a half-dozen entities, each with different accession numbers and a system of linking that covers a wide range of potential applications but can be difficult to navigate (Sequence Read Archive Submissions Staff 2011). The physical sample is described using an SRA “sample” object, usually grouped within a “study” object with other samples (and its own metadata). The library prep and sequencing platform is described by a separate “experiment” object, and none of these contain the actual sequencing data, which is stored in an SRA “run” object, which also has its own metadata. We have reconciled all of these into two entities: studies and samples. This will also improve adherence to another findability component, “(meta)data are registered or indexed in a searchable resource.” Without the nesting-doll 124 metadata structure, it is now much more straightforward for users to find the samples and data they’re looking for, and further improvements can now be built on top of the infrastructure we have already established. The other relevant FAIR component is accessibility: “(Meta)data are retrievable by their identifier using a standardised communications protocol.” While the focus is generally placed on the “communications protocol” aspect of this requirement, a pressing issue with SRA data is that it is not retrievable by its identifier, for the same reason samples are not particularly findable: Locating a relevant BioProject requires obtaining a list of BioSamples, determining the SRA sample ID for each BioSample, sending queries for each sample’s SRA run accessions, then sending requests to download each run separately. Users can now retrieve all samples in a project with a single command or download all genomic data from a sample using its SRA sample accession. Contextualization of novel datasets. The compendium could also be useful to researchers entering the analysis phase of their own microbiome studies to contextualize their datasets using information drawn from other studies, even without specialized machine-learning models. For example, in a project with 30 fecal microbiome samples taken from infants, a principal coordinates analysis (PCoA) of those 30 samples is likely to be heavily influenced by random variation and irrelevant correlations—in a study that may include hundreds of bacterial genera, it’s difficult, if not impossible, to determine which are most relevant in all but the most pronounced effects. However, researchers could query the Human Microbiome Compendium and export results for all infant fecal microbiome samples, 125 combined and formatted identically to the output from their own samples, and plot all these together, as Shin et al. (2016) did in their study of eye microbiota and contact lens wearing. Currently, the approach requires data managers to sift through thousands of samples to locate the studies with applicable data, process each study separately to generate relative abundance tables, then combine them with their new data before they can begin to perform their analysis. This compendium obviates much of this process and makes it easier to determine which samples to include. Dimensionality reduction. The original intent for this dataset was as the training set for dimensionality reduction algorithms, which is still an appealing avenue of investigation: When the microbiome composition of a sample is characterized using 16S amplicon sequencing, taxonomic specificity is generally limited to the genus level (Jovel et al. 2016). In studies seeking to differentiate between groups, it is common to summarize this data by zooming out taxonomically—essentially, trying to reduce noise by consolidating genera into families, or families into orders (e.g. Brenner et al. 2018). However, combining variables, rather than discarding them, may give us a more holistic view of the ecological dynamics at play. Little has so far been published on using microbiome composition as the output of a model, or of inferring ecological characteristics of the microbiome using “multi- omics” datasets: Wilmanski et al. (2019) found blood metabolites could be used to predict alpha (within-sample) diversity in fecal microbiome samples, and Morton et al. (2017) predicted soil microbiome composition using the pH of soil samples. Morgan et al. (2015) 126 examined associations between host gene expression and the microbiome but did not perform any prediction. Though these studies took very different approaches, they share a crucial methodological approach: dimensionality reduction. This technique summarizes high-dimensional datasets by finding combinations of factors that account for a large proportion of observed variance between subjects. Models in all these studies were trained not on relative abundance profiles, but on computationally reduced representations of those profiles. Wilmanski et al. (2019) used Shannon diversity, a single number reflecting the number of taxa present in a sample and how even their abundances are (Shannon 1948; DeJong 1975). Morton et al. (2017) used the isometric log transform, which divides taxa into “balance trees” that measure the proportional differences between subpopulations within the sample (Egozcue et al. 2003). Morgan et al. (2015) used principal component analysis to summarize both microbiome composition and host gene expression, generating nine “gene principal components” and nine “clade principal components” that were then compared. Dimensionality reduction has been used in microbiome studies for a decade, and in other machine-learning contexts for far longer. These reduced dimensions can then be used to build models with fewer predictors that are less noisy, less sparse, and, perhaps most importantly, less intercorrelated—a common issue with machine learning models that rely on techniques such as variable selection (Wolf and Bileschi 2005). Microbiome researchers have only now begun to investigate the utility of deep learning techniques to their work: Oh and Zhang (2019) used autoencoders to build latent representations of shotgun 127 metagenomic datasets and found the approach improved models for discriminating between case and control subjects across several diseases. García-Jiménez et al. (2020) used an autoencoder trained on samples from a single study (n=4,724) to predict microbiome composition of the maize rhizosphere using environmental variables such as plant age, rainfall and temperature. There are few other examples, but we are optimistic that this dataset, in performing the bioinformatic analysis of so many samples at once, could help remove one of the primary computational barriers to their development. Another chronic problem in microbiome studies is small sample sizes, which can make dimensionality reduction difficult and unreliable, as described by years of literature on linear discriminant analysis (Chen et al. 2000; Rui Huang et al. 2002). Similar issues have been encountered by researchers trying to model gene expression: High-dimensional datasets generated by next-generation sequencing can enable us to look at tens of thousands of factors for each sample; whether those factors are bacterial species or genes is mostly irrelevant. It’s much easier to expand the number of genes in an analysis than it is to expand the number of subjects—huge, “wide” datasets are more likely to include the relevant signal, but finding it becomes increasingly difficult without more samples. Recent research in RNA-seq studies of rare diseases has attempted to solve this problem using transfer learning, with encouraging results (Taroni et al. 2019; Banerjee et al. 2020). A new approach to defining gene expression pathways powered by “big data” has demonstrated that important variables can be extracted from large public compendia to be used in studies for which large sample sizes are impractical or impossible (Way and Greene 2019). To my 128 knowledge, this approach has never been applied to the microbiome. Just as transfer learning could be used to reduce tens of thousands of transcripts to less than 1,000 latent variables, I propose it could be used to reduce thousands of microbial amplicon sequence variants down to a much more manageable (and less noisy) size. A practical, reusable method of dimensionality reduction would make it easier to extract meaningful insights from limited sample sizes. The output of these models, the microbiome equivalent of gene “pathway” analysis, can then be used to train more specific models that examine relationships between microbes and any number of factors, from clinical outcomes to the levels of metabolites or gene transcripts. Gene pathway analysis has revealed critical insights (e.g. Walsh et al. 2008; Kauffmann et al. 2008; Deaglio et al. 2010) that may have remained hidden forever if analysis was limited to either individual genes or the entire genome. There is currently no equivalent approach in microbiome research—determining broad patterns of covariance in gut microbial communities will enable us and other researchers to interrogate their data from a new and powerful perspective. Transfer learning can be used for dimensionality reduction of microbiome data, to break down wide datasets into more manageable pieces. These latent variables can then be used to examine complex relationships between gene expression and microbial communities, enabling us to use existing datasets to find gene–microbe interactions that can be examined with more specificity and get us closer to understanding how gut microbes influence their hosts. The investigation of causal mechanisms in host–microbe interactions is hampered 129 by a dearth of precise, testable hypotheses. The use of latent variables to evaluate community-level patterns can help us move closer to identifying how microbes and humans coexist. Those latent variables can now be much more confidently generated using our compendium. Simulated microbiomes. Another intriguing area is the potential for the generation of “simulated microbiomes” informed by models trained on existing samples. Evidence in transcriptomics, for example, shows some genes tend to appear as “differentially expressed” regardless of the biological question (Crow et al. 2019). It would be interesting to ask similar questions of patterns in the human microbiome: Many gut microbiota have been implicated in specific conditions, but how many of these results are because some microbes are always going to be implicated? These questions require many more samples than have been collected by a single study—a problem resolved by our compendium. In summary, the compendium described in Chapter 4 presents myriad opportunities for further analysis, both as an object of study itself and as a computational component of new methods that could be applied to future microbiome studies. The bioinformatics legwork of compilation is complete—now we can get to the fun part. 130 Bibliography Abdill, Richard J., Elizabeth M. Adamowicz, and Ran Blekhman. 2020. “International Authorship and Collaboration across bioRxiv Preprints.” eLife 9 (July): e58496. ———. 2021. “Public Human Microbiome Data Dominated by Highly Developed Countries.” bioRxiv. https://doi.org/10.1101/2021.09.02.458641. Abdill, Richard J., and Ran Blekhman. 2019a. “Complete Rxivist Dataset of Scraped bioRxiv Data.” Zenodo. https://doi.org/10.5281/ZENODO.2529922. ———. 2019b. “Tracking the Popularity and Outcomes of All bioRxiv Preprints.” eLife 8 (April). https://doi.org/10.7554/eLife.45133. ———. 2019c. “Rxivist.org: Sorting Biology Preprints Using Social Media and Readership Metrics.” PLoS Biology 17 (5): e3000269. “About LDCs.” 2013. United Nations Office of the High Representative for the Least Developed Countries, Landlocked Developing Countries and the Small Island Developing States (UN-OHRLLS). September 19, 2013. http://unohrlls.org/about-ldcs/. Adams, James D., Grant C. Black, J. Roger Clemmons, and Paula E. Stephan. 2005. “Scientific Teams and Institutional Collaborations: Evidence from U.S. Universities, 1981–1999.” Research Policy 34 (3): 259–85. Adebamowo, Clement, Non-African Collaborators, Sally Adebamowo, and Charles Rotimi. n.d. “African Collaborative Center for Microbiome and Genomics Research (ACCME).” https://h3africa.org/index.php/accme/. Adebamowo, Sally N., Bing Ma, Davide Zella, Ayotunde Famooto, Jacques Ravel, Clement Adebamowo, and ACCME Research Group. 2017. “Mycoplasma Hominis and Mycoplasma Genitalium in the Vaginal Microbiota and Persistent High-Risk Human Papillomavirus Infection.” Frontiers in Public Health 5 (June): 140. Akre, Olof, Francesco Barone-Adesi, Andreas Pettersson, Neil Pearce, Franco Merletti, and Lorenzo Richiardi. 2011. “Differences in Citation Rates by Country of Origin for Papers Published in Top- Ranked Medical Journals: Do They Reflect Inequalities in Access to Publication?” Journal of Epidemiology and Community Health 65 (2): 119–23. Aksnes, Dag W. 2008. “When Different Persons Have an Identical Author Name. How Frequent Are Homonyms?” Journal of the American Society for Information Science and Technology 59 (5): 838–41. ALLEA. 2018. “Systemic Reforms and Further Consultation Needed to Make Plan S a Success.” European Federation of Academies of Sciences and Humanities. December 12, 2018. https://allea.org/systemic-reforms-and-further-consultation-needed-to-make-plan-s-a-success/. Altmetric Support. 2018. “How Is the Altmetric Attention Score Calculated?” 2018. https://help.altmetric.com/support/solutions/articles/6000060969-how-is-the-altmetric-attention- score-calculated. Amano, Tatsuya, Juan P. González-Varo, and William J. Sutherland. 2016. “Languages Are Still a Major Barrier to Global Science.” PLoS Biology 14 (12): e2000933. Amato, Katherine R., Marie-Claire Arrieta, Meghan B. Azad, Michael T. Bailey, Josiane L. Broussard, Carlijn E. Bruggeling, Erika C. Claud, et al. 2021. “The Human Gut Microbiome and Health 131 Inequities.” Proceedings of the National Academy of Sciences of the United States of America 118 (25). https://doi.org/10.1073/pnas.2017947118. Anaya, Jordan. 2018. PrePubMed: Analyses (version 674d5aa). Github. https://github.com/OmnesRes/prepub. Antonelli, Alexandre. 2020. “Director of Science at Kew: It’s Time to Decolonise Botanical Collections.” The Conversation, June 19, 2020. http://theconversation.com/director-of-science-at-kew-its-time- to-decolonise-botanical-collections-141070. Archambault, Éric, Étienne Vignola-Gagné, Grégoire Côté, Vincent Larivière, and Yves Gingrasb. 2006. “Benchmarking Scientific Output in the Social Sciences and Humanities: The Limits of Existing Databases.” Scientometrics 68 (3): 329–42. Armenteras, Dolors. 2021. “Guidelines for Healthy Global Scientific Collaborations.” Nature Ecology & Evolution 5 (9): 1193–94. Baker, Kate, Markus P. Eichhorn, and Mark Griffiths. 2019. “Decolonizing Field Ecology.” Biotropica 51 (3): 288–92. Banerjee, Jineta, Robert J. Allaway, Jaclyn N. Taroni, Aaron Baker, Xiaochun Zhang, Chang In Moon, Jaishri O. Blakeley, et al. 2020. “Integrative Analysis Identifies Candidate Tumor Microenvironment and Intracellular Signaling Pathways That Define Tumor Heterogeneity in NF1.” bioRxiv. https://doi.org/10.1101/2020.01.13.904771. Barrett, Tanya, Karen Clark, Robert Gevorgyan, Vyacheslav Gorelenkov, Eugene Gribov, Ilene Karsch- Mizrachi, Michael Kimelman, et al. 2012. “BioProject and BioSample Databases at NCBI: Facilitating Capture and Organization of Metadata.” Nucleic Acids Research 40 (D1): D57–63. Barsh, Gregory S., Casey M. Bergman, Christopher D. Brown, Nadia D. Singh, and Gregory P. Copenhaver. 2016. “Bringing PLOS Genetics Editors to Preprint Servers.” PLoS Genetics 12 (12): e1006448. Becerril-García, Arianna. 2019. “AmeliCA vs Plan S: Same Target, Two Different Strategies to Achieve Open Access.” AmeliCA (blog). http://amelica.org/index.php/en/2019/02/10/amelica-vs-plan-s- same-target-two-different-strategies-to-achieve-open-access/. Belhabib, Dyhia. 2021. “Ocean Science and Advocacy Work Better When Decolonized.” Nature Ecology & Evolution 5 (6): 709–10. Benezra, Amber. 2020. “Race in the Microbiome.” Science, Technology & Human Values 45 (5): 877–902. Berg, Jeremy M., Needhi Bhalla, Philip E. Bourne, Martin Chalfie, David G. Drubin, James S. Fraser, Carol W. Greider, et al. 2016. “Preprints for the Life Sciences.” Science 352 (6288): 899–901. Bhandary, Priyanka, Arun S. Seetharam, Zebulun W. Arendsee, Manhoi Hur, and Eve Syrkin Wurtele. 2018. “Raising Orphans from a Metadata Morass: A Researcher’s Guide to Re-Use of Public ‘Omics Data.” Plant Science: An International Journal of Experimental Plant Biology 267 (February): 32– 47. “BioSample: Organism Information.” n.d. NCBI. Accessed November 27, 2021. https://www.ncbi.nlm.nih.gov/biosample/docs/organism/. Blekhman, Ran, Julia K. Goodrich, Katherine Huang, Qi Sun, Robert Bukowski, Jordana T. Bell, Timothy D. Spector, et al. 2015. “Host Genetic Variation Impacts Microbiome Composition across Human Body Sites.” Genome Biology 16 (September): 191. Blekhman, Ran, Karen Tang, Elizabeth A. Archie, Luis B. Barreiro, Zachary P. Johnson, Mark E. Wilson, Jordan Kohn, et al. 2016. “Common Methods for Fecal Sample Storage in Field Studies Yield 132 Consistent Signatures of Individual Identity in Microbiome Sequencing Data.” Scientific Reports 6 (August): 31519. Bordons, María, Javier Aparicio, and Rodrigo Costas. 2013. “Heterogeneity of Collaboration and Its Relationship with Research Impact in a Biomedical Field.” Scientometrics 96 (2): 443–66. Boudry, Christophe, and Ghislaine Chartron. 2017. “Availability of Digital Object Identifiers in Publications Archived by PubMed.” Scientometrics 110 (3): 1453–69. Brenner, David, Andreas Hiergeist, Carolin Adis, Benjamin Mayer, André Gessner, Albert C. Ludolph, and Jochen H. Weishaupt. 2018. “The Fecal Microbiome of ALS Patients.” Neurobiology of Aging 61 (January): 132–37. Brewster, Ryan, Fiona B. Tamburini, Edgar Asiimwe, Ovokeraye Oduaran, Scott Hazelhurst, and Ami S. Bhatt. 2019. “Surveying Gut Microbiome Research in Africans: Toward Improved Diversity and Representation.” Trends in Microbiology 27 (10): 824–35. Brooks, Andrew W., Sambhawa Priya, Ran Blekhman, and Seth R. Bordenstein. 2018. “Gut Microbiota Diversity across Ethnicities in the United States.” PLoS Biology 16 (12): e2006842. Buehring, Gertrude Case, Jessica E. Buehring, and Patrick D. Gerard. 2007. “Lost in Citation: Vanishing Visibility of Senior Authors.” Scientometrics 72 (3): 459–68. Burgman, Mark, Frith Jarrad, and Ellen Main. 2015. “Decreasing Geographic Bias in Conservation Biology.” Conservation Biology: The Journal of the Society for Conservation Biology 29 (5): 1255–56. Cai, Mingxuan, Jiashun Xiao, Shunkang Zhang, Xiang Wan, Hongyu Zhao, Gang Chen, and Can Yang. 2021. “A Unified Framework for Cross-Population Trait Prediction by Leveraging the Genetic Correlation of Polygenic Traits.” American Journal of Human Genetics 108 (4): 632–55. Callahan, Benjamin. n.d. “A DADA2 Workflow for Big Data (1.4 or Later).” Accessed November 28, 2021. https://benjjneb.github.io/dada2/bigdata.html. Callahan, Benjamin J., Paul J. McMurdie, Michael J. Rosen, Andrew W. Han, Amy Jo A. Johnson, and Susan P. Holmes. 2016. “DADA2: High-Resolution Sample Inference from Illumina Amplicon Data.” Nature Methods 13 (7): 581–83. Callaway, Ewen. 2013. “Preprints Come to Life.” Nature 503 (7475): 180. “CARE Principles of Indigenous Data Governance — Global Indigenous Data Alliance.” n.d. Accessed December 9, 2021. https://www.gida-global.org/care. Champieux, Robin. 2018. “Gathering Steam: Preprints, Librarian Outreach, and Actions for Change.” The Official PLOS Blog. October 15, 2018. https://blogs.plos.org/plos/2018/10/gathering-steam- preprints-librarian-outreach-and-actions-for-change/. Čhaŋtémaza, and Monica Siems McKay. 2020. “Where We Stand: The University of Minnesota and Dakhóta Treaty Lands.” Open Rivers: Rethinking Water, Place & Community, no. 17. https://doi.org/10.24926/2471190X.7588. Chen, Li-Fen, Hong-Yuan Mark Liao, Ming-Tat Ko, Ja-Chen Lin, and Gwo-Jong Yu. 2000. “A New LDA- Based Face Recognition System Which Can Solve the Small Sample Size Problem.” Pattern Recognition 33 (10): 1713–26. Clarivate Analytics. 2018. “Journal Citation Reports Science Edition.” 2018. Clemente, Jose C., Erica C. Pehrsson, Martin J. Blaser, Kuldip Sandhu, Zhan Gao, Bin Wang, Magda Magris, et al. 2015. “The Microbiome of Uncontacted Amerindians.” Science Advances 1 (3). https://doi.org/10.1126/sciadv.1500183. 133 Cobb, Matthew. 2017. “The Prehistory of Biology Preprints: A Forgotten Experiment from the 1960s.” PLoS Biology 15 (11): e2003995. Collado-Torres, Leonardo, Abhinav Nellore, Kai Kammers, Shannon E. Ellis, Margaret A. Taub, Kasper D. Hansen, Andrew E. Jaffe, Ben Langmead, and Jeffrey T. Leek. 2017. “Reproducible RNA-Seq Analysis Using recount2.” Nature Biotechnology 35 (4): 319–21. Cox, Jennifer. 2019. “Preprint Servers: CSE Editorial Policy Committee White Paper Update.” Science 42 (1): 24. Crow, Megan, Nathaniel Lim, Sara Ballouz, Paul Pavlidis, and Jesse Gillis. 2019. “Predictability of Human Differential Gene Expression.” Proceedings of the National Academy of Sciences of the United States of America 116 (13): 6491–6500. Dareng, E. O., B. Ma, A. O. Famooto, S. N. Adebamowo, R. A. Offiong, O. Olaniyan, P. S. Dakum, et al. 2016. “Prevalent High-Risk HPV Infection and Vaginal Microbiota in Nigerian Women.” Epidemiology and Infection 144 (1): 123–37. Dayama, Gargi, Sambhawa Priya, David E. Niccum, Alexander Khoruts, and Ran Blekhman. 2020. “Interactions between the Gut Microbiome and Host Gene Regulation in Cystic Fibrosis.” Genome Medicine 12 (1): 12. Deaglio, Silvia, Semra Aydin, Maurizia Mello Grand, Tiziana Vaisitti, Luciana Bergui, Giovanni D’Arena, Giovanna Chiorino, and Fabio Malavasi. 2010. “CD38/CD31 Interactions Activate Genetic Pathways Leading to Proliferation and Migration in Chronic Lymphocytic Leukemia Cells.” Molecular Medicine 16 (3-4): 87–91. Debat, Humberto, and Dominique Babini. 2019. “Plan S in Latin America: A Precautionary Note.” PeerJ. https://doi.org/10.7287/peerj.preprints.27834v1. De Coster, W. 2017. “A Twitter Bot to Find the Most Interesting bioRxiv Preprints.” Gigabase or Gigabyte. August 8, 2017. https://gigabaseorgigabyte.wordpress.com/2017/08/08/a-twitter-bot-to-find-the- most-interesting-biorxiv-preprints/. DeJong, T. M. 1975. “A Comparison of Three Diversity Indices Based on Their Components of Richness and Evenness.” Oikos 26 (2): 222–27. Delamothe, T., R. Smith, M. A. Keller, J. Sack, and B. Witscher. 1999. “Netprints: The next Phase in the Evolution of Biomedical Publishing.” BMJ 319 (7224): 1515–16. De La Vega, Francisco M., and Carlos D. Bustamante. 2018. “Polygenic Risk Scores: A Biased Prediction?” Genome Medicine 10 (1): 100. Delgado, Abigail Nieves, and Jan Baedke. 2021. “Does the Human Microbiome Tell Us Something about Race?” Humanities and Social Sciences Communications 8 (1): 1–12. Delicado, Pedro, and Cristian Pachon-Garcia. 2020. “Multidimensional Scaling for Big Data.” arXiv [stat.CO]. arXiv. http://arxiv.org/abs/2007.11919. Desjardins-Proulx, Philippe, Ethan P. White, Joel J. Adamson, Karthik Ram, Timothée Poisot, and Dominique Gravel. 2013. “The Case for Open Preprints in Biology.” PLoS Biology 11 (5): e1001563. De Wolfe, Travis J., Mohammed Rafi Arefin, Amber Benezra, and María Rebolleda Gómez. 2021. “Chasing Ghosts: Race, Racism, and the Future of Microbiome Research.” mSystems 6 (5): e0060421. Díaz, Magdalena, Pablo Jarrín-V, Raquel Simarro, Pablo Castillejo, Gabriela N. Tenea, and C. Alfonso Molina. 2021. “The Ecuadorian Microbiome Project: A Plea to Strengthen Microbial Genomic Research.” Neotropical Biodiversity 7 (1): 223–37. 134 Di Gregorio, F., and D. Varrazzo. 2018. psycopg2 (version 2.7.5). Github. https://github.com/psycopg/psycopg2. Docker Inc. 2018. Docker (version 18.06.1-ce). https://www.docker.com. Eckert, Ester M., Andrea Di Cesare, Diego Fontaneto, Thomas U. Berendonk, Helmut Bürgmann, Eddie Cytryn, Despo Fatta-Kassinos, et al. 2020. “Every Fifth Published Metagenome Is Not Available to Science.” PLoS Biology 18 (4): e3000698. Egozcue, J. J., V. Pawlowsky-Glahn, G. Mateu-Figueras, and C. Barceló-Vidal. 2003. “Isometric Logratio Transformations for Compositional Data Analysis.” Mathematical Geology 35 (3): 279–300. Eichhorn, Markus P., Kate Baker, and Mark Griffiths. 2020. “Steps towards Decolonising Biogeography.” Frontiers of Biogeography 12 (1). https://doi.org/10.21425/F5FBG44795. eLife. 2020. “New from eLife: Invitation to Submit to Preprint Review.” May 13, 2020. https://elifesciences.org/inside-elife/d0c5d114/new-from-elife-invitation-to-submit-to-preprint- review. “ERA Home.” 2005. The Lancet Electronic Research Archive. 2005. https://web.archive.org/web/20050422224839/http://www.thelancet.com/era. Fadlelmola, Faisal M., Kais Ghedira, Yosr Hamdi, Mariem Hanachi, Fouzia Radouani, Imane Allali, Anmol Kiran, et al. 2021. “H3ABioNet Genomic Medicine and Microbiome Data Portals Hackathon Proceedings.” Database: The Journal of Biological Databases and Curation 2021 (April). https://doi.org/10.1093/database/baab016. Feldman, Sergey, Kyle Lo, and Waleed Ammar. 2018. “Citation Count Analysis for Papers with Preprints.” arXiv [cs.DL]. arXiv. http://arxiv.org/abs/1805.05238. Flint, Harry J., Edward A. Bayer, Marco T. Rincon, Raphael Lamed, and Bryan A. White. 2008. “Polysaccharide Utilization by Gut Bacteria: Potential for New Insights from Genomic Analysis.” Nature Reviews. Microbiology 6 (2): 121–31. Forslund, Kristoffer, Shinichi Sunagawa, Jens Roat Kultima, Daniel R. Mende, Manimozhiyan Arumugam, Athanasios Typas, and Peer Bork. 2013. “Country-Specific Antibiotic Use Practices Impact the Human Gut Resistome.” Genome Research 23 (7): 1163–69. Fox, Keolu. 2020. “The Illusion of Inclusion - The ‘All of Us’ Research Program and Indigenous Peoples’ DNA.” The New England Journal of Medicine 383 (5): 411–13. Fragiadakis, Gabriela K., Samuel A. Smits, Erica D. Sonnenburg, William Van Treuren, Gregor Reid, Rob Knight, Alphaxard Manjurano, et al. 2019. “Links between Environment, Diet, and the Hunter- Gatherer Microbiome.” Gut Microbes 10 (2): 216–27. Fraser, Nicholas, Liam Brierley, Gautam Dey, Jessica K. Polka, Máté Pálfy, Federico Nanni, and Jonathon Alexis Coates. 2021. “The Evolving Role of Preprints in the Dissemination of COVID-19 Research and Their Impact on the Science Communication Landscape.” PLoS Biology 19 (4): e3000959. Fraser, Nicholas, Fakhri Momeni, Philipp Mayr, and Isabella Peters. 2020. “The Relationship between bioRxiv Preprints, Citations and Altmetrics.” Quantitative Science Studies, April, 1–39. Fu, Darwin Y., and Jacob J. Hughey. 2019. “Releasing a Preprint Is Associated with More Attention and Citations for the Peer-Reviewed Article.” eLife 8 (December). https://doi.org/10.7554/eLife.52646. García-Jiménez, Beatriz, Jorge Muñoz, Sara Cabello, Joaquín Medina, and Mark D. Wilkinson. 2020. “Predicting Microbiomes through a Deep Latent Space.” bioRxiv. 135 Garfield, Eugene. 2006. “The History and Meaning of the Journal Impact Factor.” JAMA: The Journal of the American Medical Association 295 (1): 90–93. Gauffriau, Marianne, Peder Olesen Larsen, Isabelle Maye, Anne Roulin-Perriard, and Markus von Ins. 2008. “Comparisons of Results of Publication Counting Using Different Methods.” Scientometrics 77 (1): 147–76. Gewin, Virginia. 2021. “How to Include Indigenous Researchers and Their Knowledge.” Nature 589 (7841): 315–17. Glänzel, Wolfgang, and András Schubert. 2005. “Analysing Scientific Networks Through Co-Authorship.” In Handbook of Quantitative Science and Technology Research: The Use of Publication and Patent Statistics in Studies of S&T Systems, edited by Henk F. Moed, Wolfgang Glänzel, and Ulrich Schmoch, 257–76. Dordrecht: Springer Netherlands. Gonçalves, Rafael S., and Mark A. Musen. 2019. “The Variable Quality of Metadata about Biological Samples Used in Biomedical Experiments.” Scientific Data 6 (February): 190021. Gonçalves, Rafael S., Martin J. O’Connor, Marcos Martínez-Romero, John Graybeal, and Mark A. Musen. 2017. “Metadata in the BioSample Online Repository Are Impaired by Numerous Anomalies.” arXiv [cs.DB]. arXiv. http://arxiv.org/abs/1708.01286. González-Alcaide, Gregorio, Jinseo Park, Charles Huamaní, and José M. Ramos. 2017. “Dominance and Leadership in Research Activities: Collaboration between Countries of Differing Human Development Is Reflected through Authorship Order and Designation as Corresponding Authors in Scientific Publications.” PloS One 12 (8): e0182513. Goodrich, Julia K., Jillian L. Waters, Angela C. Poole, Jessica L. Sutter, Omry Koren, Ran Blekhman, Michelle Beaumont, et al. 2014. “Human Genetics Shape the Gut Microbiome.” Cell 159 (4): 789– 99. Groussin, Mathieu, Mathilde Poyet, Ainara Sistiaga, Sean M. Kearney, Katya Moniz, Mary Noel, Jeff Hooker, et al. 2021. “Elevated Rates of Horizontal Gene Transfer in the Industrialized Human Microbiome.” Cell 184 (8): 2053–67.e18. Gupta, Vinod K., Sandip Paul, and Chitra Dutta. 2017. “Geography, Ethnicity or Subsistence-Specific Variations in Human Microbiome Composition and Diversity.” Frontiers in Microbiology 8 (June): 1162. Gurdasani, Deepti, Inês Barroso, Eleftheria Zeggini, and Manjinder S. Sandhu. 2019. “Genomics of Disease Risk in Globally Diverse Populations.” Nature Reviews. Genetics 20 (9): 520–35. H3Africa Consortium, Charles Rotimi, Akin Abayomi, Alash ‘le Abimiku, Victoria May Adabayeri, Clement Adebamowo, Ezekiel Adebiyi, et al. 2014. “Research Capacity. Enabling the Genomic Revolution in Africa.” Science 344 (6190): 1346–48. Haak, Laure. 2012. “The O in ORCID.” ORCiD. December 5, 2012. https://orcid.org/blog/2012/12/06/o- orcid. Haelewaters, Danny, Tina A. Hofmann, and Adriana L. Romero-Olivares. 2021. “Ten Simple Rules for Global North Researchers to Stop Perpetuating Helicopter Research in the Global South.” PLoS Computational Biology 17 (8): e1009277. Hagen, Nils T. 2013. “Harmonic Coauthor Credit: A Parsimonious Quantification of the Byline Hierarchy.” Journal of Informetrics 7 (4): 784–91. 136 Hartgerink, C. H. J. 2015. “Publication Cycle: A Study of the Public Library of Science (PLOS).” 2015. https://www.authorea.com/users/2013/articles/36067-publication-cycle-a-study-of-the-public- library-of-science-plos/_show_article. Haustein, Stefanie. 2018. “Scholarly Twitter Metrics.” arXiv [cs.SI]. arXiv. http://arxiv.org/abs/1806.02201. Hawkins, Nathaniel T., Marc Maldaver, Anna Yannakopoulos, Lindsay A. Guare, and Arjun Krishnan. 2021. “Systematic Tissue Annotations of--Omics Samples by Modeling Unstructured Metadata.” bioRxiv. https://www.biorxiv.org/content/10.1101/2021.05.10.443525v2.abstract. Hazlett, Megan A., Kate M. Henderson, Ilana F. Zeitzer, and Joshua A. Drew. 2020. “The Geography of Publishing in the Anthropocene.” Conservation Science and Practice 2 (10). https://doi.org/10.1111/csp2.270. He, Yan, Wei Wu, Hui-Min Zheng, Pan Li, Daniel McDonald, Hua-Fang Sheng, Mu-Xuan Chen, et al. 2018. “Regional Variation Limits Applications of Healthy Gut Microbiome Reference Ranges and Disease Models.” Nature Medicine 24 (10): 1532–35. Himmelstein, Daniel. 2016a. “The History of Publishing Delays.” February 10, 2016. https://blog.dhimmel.com/history-of-delays/. ———. 2016b. “The Licensing of bioRxiv Preprints.” Satoshi Village. December 5, 2016. https://blog.dhimmel.com/biorxiv-licenses/. Holdgraf, C. R. 2016. “Question 1: How Has the Published Articles Rate Changed?” 2016. https://predictablynoisy.com/scrape-biorxiv. Huang, Rui, Qingshan Liu, Hanqing Lu, and Songde Ma. 2002. “Solving the Small Sample Size Problem of LDA.” In Object Recognition Supported by User Interaction for Service Robots, 3:29–32 vol.3. Inglis, John R., and Richard Sever. 2016. “bioRxiv: A Progress Report.” ASAPbio Blog. 2016. http://asapbio.org/biorxiv. Ishaq, Suzanne L., Francisco J. Parada, Patricia G. Wolf, Carla Y. Bonilla, Megan A. Carney, Amber Benezra, Emily Wissel, et al. 2021. “Introducing the Microbes and Social Equity Working Group: Considering the Microbial Components of Social, Environmental, and Health Justice.” mSystems, July, e0047121. Johnson, Abigail J., Pajau Vangay, Gabriel A. Al-Ghalith, Benjamin M. Hillmann, Tonya L. Ward, Robin R. Shields-Cutler, Austin D. Kim, et al. 2019. “Daily Sampling Reveals Personalized Diet-Microbiome Associations in Humans.” Cell Host & Microbe 25 (6): 789–802.e5. Jovel, Juan, Jordan Patterson, Weiwei Wang, Naomi Hotte, Sandra O’Keefe, Troy Mitchel, Troy Perry, et al. 2016. “Characterization of the Gut Microbiome Using 16S or Shotgun Metagenomics.” Frontiers in Microbiology 7 (April): 459. Kaiser, Jocelyn. 2017. “The Preprint Dilemma.” Science 357 (6358): 1344–49. Kaplan, Robert C., Zheng Wang, Mykhaylo Usyk, Daniela Sotres-Alvarez, Martha L. Daviglus, Neil Schneiderman, Gregory A. Talavera, et al. 2019. “Gut Microbiome Composition in the Hispanic Community Health Study/Study of Latinos Is Shaped by Geographic Relocation, Environmental Factors, and Obesity.” Genome Biology 20 (1): 219. Karpathy, Andrej. 2018. Arxiv Sanity Preserver (version 8e52b8b). Github. https://github.com/karpathy/arxiv-sanity-preserver. Kauffmann, A., F. Rosselli, V. Lazar, V. Winnepenninckx, A. Mansuet-Lupo, P. Dessen, J. J. van den Oord, A. Spatz, and A. Sarasin. 2008. “High Expression of DNA Repair Pathways Is Associated with Metastasis in Melanoma Patients.” Oncogene 27 (5): 565–73. 137 Kim, Jinseok, and Jana Diesner. 2015. “Coauthorship Networks: A Directed Network Approach Considering the Order and Number of Coauthors.” Journal of the Association for Information Science and Technology 66 (12): 2685–96. Klein, Martin, Peter Broadwell, Sharon E. Farb, and Todd Grappone. 2016. “Comparing Published Scientific Journal Articles to Their Pre-Print Versions.” In Proceedings of the 16th ACM/IEEE-CS on Joint Conference on Digital Libraries, 153–62. JCDL ‘16. New York, NY, USA: Association for Computing Machinery. Kling, Rob, Lisa B. Spector, and Joanna Fortuna. 2004. “The Real Stakes of Virtual Publishing: The Transformation of E-Biomed into PubMed Central.” Journal of the American Society for Information Science and Technology 55 (2): 127–48. Kramer, Bianca. 2019. “Rxivist Analysis.” Google Docs. https://docs.google.com/spreadsheets/d/18- zIlfgrQaGo6e4SmyfzMTY7AN1dUYiI5l6PyX5pWtg/edit. Kushugulova, Almagul, Sofia K. Forslund, Paul Igor Costea, Samat Kozhakhmetov, Zhanagul Khassenbekova, Maira Urazova, Talgat Nurgozhin, et al. 2018. “Metagenomic Analysis of Gut Microbial Communities from a Central Asian Population.” BMJ Open 8 (7): e021682. Larivière, Vincent, Cassidy R. Sugimoto, Benoit Macaluso, Staša Milojević, Blaise Cronin, and Mike Thelwall. 2014. “arXiv E-Prints and the Journal of Record: An Analysis of Roles and Relationships.” Journal of the Association for Information Science and Technology 65 (6): 1157– 69. Lee, Carole J., Cassidy R. Sugimoto, Guo Zhang, and Blaise Cronin. 2013. “Bias in Peer Review.” Journal of the American Society for Information Science and Technology 64 (1): 2–17. Lee, Robert, Tristan Ahtone, Margaret Pearce, Kalen Goodluck, Geoff McGhee, Cody Leff, and Marty Two Bulls Jr. 2020. “Land-Grab Universities: University of Minnesota.” High Country News, 2020. https://www.landgrabu.org/universities/university-of-minnesota. Lee, Soo Ching, Mei San Tang, Yvonne A. L. Lim, Seow Huey Choy, Zachary D. Kurtz, Laura M. Cox, Uma Mahesh Gundra, et al. 2014. “Helminth Colonization Is Associated with Increased Diversity of the Gut Microbiota.” PLoS Neglected Tropical Diseases 8 (5): e2880. Le, Trang T., Daniel S. Himmelstein, Ariel A. Hippen Anderson, Matthew R. Gazzara, and Casey S. Greene. 2020. “Analysis of ISCB Honorees and Keynotes Reveals Disparities.” bioRxiv. https://doi.org/10.1101/2020.04.14.927251. “List of Predatory Journals.” 2018. Stop Predatory Journals. 2018. https://predatoryjournals.com/journals/. Lu, Louise J., and Ji Liu. 2016. “Human Microbiota and Ophthalmic Disease.” The Yale Journal of Biology and Medicine 89 (3): 325–30. Magne, Fabien, Martin Gotteland, Lea Gauthier, Alejandra Zazueta, Susana Pesoa, Paola Navarrete, and Ramadass Balamurugan. 2020. “The Firmicutes/Bacteroidetes Ratio: A Relevant Marker of Gut Dysbiosis in Obese Patients?” Nutrients 12 (5). https://doi.org/10.3390/nu12051474. Mammides, Christos, Uromi M. Goodale, Richard T. Corlett, Jin Chen, Kamaljit S. Bawa, Hetal Hariya, Frith Jarrad, et al. 2016. “Increasing Geographic Diversity in the International Conservation Literature: A Stalled Process?” Biological Conservation 198 (June): 78–83. Marshall, E. 1999. “PNAS to Joint PubMed Central--on Condition.” Science 286 (5440): 655–56. Matson, Vyara, Jessica Fessler, Riyue Bao, Tara Chongsuwat, Yuanyuan Zha, Maria-Luisa Alegre, Jason J. Luke, and Thomas F. Gajewski. 2018. “The Commensal Microbiome Is Associated with Anti-PD- 1 Efficacy in Metastatic Melanoma Patients.” Science 359 (6371): 104–8. 138 Mattsson, Pauline, Carl Johan Sundberg, and Patrice Laget. 2011. “Is Correspondence Reflected in the Author Position? A Bibliometric Study of the Relation between Corresponding Author and Byline Position.” Scientometrics 87 (1): 99–105. McConnell, J., and R. Horton. 1999. “Lancet Electronic Research Archive in International Health and Eprint Server.” The Lancet 354 (9172): 2–3. Medina-Gomez, Carolina, Janine Frédérique Felix, Karol Estrada, Marjoline Josephine Peters, Lizbeth Herrera, Claudia Jeanette Kruithof, Liesbeth Duijts, et al. 2015. “Challenges in Conducting Genome-Wide Association Studies in Highly Admixed Multi-Ethnic Populations: The Generation R Study.” European Journal of Epidemiology 30 (4): 317–30. “Metadata Delivery REST API.” 2018. Crossref. 2018. https://www.crossref.org/services/metadata- delivery/rest-api/. “Methods, Preprints and Papers.” 2017. Nature Biotechnology 35 (12): 1113. “Microbiome Working Group.” 2020. August 18, 2020. https://h3africa.org/index.php/microbiome-working- group/. Mongeon, Philippe, and Adèle Paul-Hus. 2016. “The Journal Coverage of Web of Science and Scopus: A Comparative Analysis.” Scientometrics 106 (1): 213–28. Montassier, Emmanuel, Gabriel A. Al-Ghalith, Tonya Ward, Stephane Corvec, Thomas Gastinne, Gilles Potel, Phillipe Moreau, Marie France de la Cochetiere, Eric Batard, and Dan Knights. 2016. “Pretreatment Gut Microbiome Predicts Chemotherapy-Related Bloodstream Infection.” Genome Medicine 8 (1): 49. “Monthly Statistics for October 2018.” 2018. PrePubMed. 2018. http://www.prepubmed.org/monthly_stats/. Morgan, Xochitl C., Boyko Kabakchiev, Levi Waldron, Andrea D. Tyler, Timothy L. Tickle, Raquel Milgrom, Joanne M. Stempak, et al. 2015. “Associations between Host Gene Expression, the Mucosal Microbiome, and Clinical Outcome in the Pelvic Pouch of Patients with Inflammatory Bowel Disease.” Genome Biology 16 (April): 67. Morton, James T., Jon Sanders, Robert A. Quinn, Daniel McDonald, Antonio Gonzalez, Yoshiki Vázquez- Baeza, Jose A. Navas-Molina, et al. 2017. “Balance Trees Reveal Microbial Niche Differentiation.” mSystems 2 (1). https://doi.org/10.1128/mSystems.00162-16. Moya-Anegón, Félix de, Zaida Chinchilla-Rodríguez, Benjamín Vargas-Quesada, Elena Corera-Álvarez, Francisco José Muñoz-Fernández, Antonio González-Molina, and Victor Herrero-Solana. 2007. “Coverage Analysis of Scopus: A Journal Metric Approach.” Scientometrics 73 (1): 53–78. Mukunth, Vasudevan. 2019. “India Will Skip Plan S, Focus on National Efforts in Science Publishing.” The Wire: Science. October 26, 2019. https://science.thewire.in/the-sciences/plan-s-open-access- scientific-publishing-article-processing-charge-insa-k-vijayraghavan/. Mulder, Nicola J., Ezekiel Adebiyi, Marion Adebiyi, Seun Adeyemi, Azza Ahmed, Rehab Ahmed, Bola Akanle, et al. 2017. “Development of Bioinformatics Infrastructure for Genomics Research.” Global Heart 12 (2): 91–98. Mutlu, Ece A., Işın Y. Comba, Takugo Cho, Phillip A. Engen, Cemal Yazıcı, Saul Soberanes, Robert B. Hamanaka, et al. 2018. “Inhalational Exposure to Particulate Matter Air Pollution Alters the Composition of the Gut Microbiome.” Environmental Pollution 240 (September): 817–30. Naing, L., T. Winn, and B. N. Rusli. 2006. “Practical Issues in Calculating the Sample Size for Prevalence Studies.” Archives of Orofacial Sciences, no. 1: 9–14. 139 Nakamura, Yasukazu, Guy Cochrane, Ilene Karsch-Mizrachi, and International Nucleotide Sequence Database Collaboration. 2013. “The International Nucleotide Sequence Database Collaboration.” Nucleic Acids Research 41 (D1): D21–24. Narock, Tom, and Evan B. Goldstein. 2019. “Quantifying the Growth of Preprint Services Hosted by the Center for Open Science.” Publications 7 (2): 44. NCBI. n.d. “Biosample Attributes.” BioSample. Accessed May 11, 2021. https://www.ncbi.nlm.nih.gov/biosample/docs/attributes/. Need, Anna C., and David B. Goldstein. 2009. “Next Generation Disparities in Human Genomics: Concerns and Remedies.” Trends in Genetics: TIG 25 (11): 489–94. Neuwirth, Erich. 2014. “RColorBrewer: ColorBrewer Palettes. R Package Version 1.1-2.” The R Foundation. https://CRAN.R-project.org/package=RColorBrewer. NIH Human Microbiome Portfolio Analysis Team. 2019. “A Review of 10 Years of Human Microbiome Research Activities at the US National Institutes of Health, Fiscal Years 2007-2016.” Microbiome 7 (1): 31. Noxolo, Patricia. 2017. “Introduction: Decolonising Geographical Knowledge in a Colonised and Re- Colonising Postcolonial World.” Area 49 (3): 317–19. Nuñez, Martin A., Jos Barlow, Marc Cadotte, Kirsty Lucas, Erika Newton, Nathalie Pettorelli, and Philip A. Stephens. 2019. “Assessing the Uneven Global Distribution of Readership, Submissions and Publications in Applied Ecology: Obvious Problems without Obvious Solutions.” The Journal of Applied Ecology 56 (1): 4–9. Oh, Min, and Liqing Zhang. 2019. “DeepMicro: Deep Representation Learning for Disease Prediction Based on Microbiome Data.” bioRxiv. https://doi.org/10.1101/785626. Okike, Kanu, Mininder S. Kocher, Charles T. Mehlman, James D. Heckman, and Mohit Bhandari. 2008. “Nonscientific Factors Associated with Acceptance for Publication in The Journal of Bone and Joint Surgery (American Volume).” The Journal of Bone and Joint Surgery. American Volume 90 (11): 2432–37. Olomu, Isoken Nicholas, Luis Carlos Pena-Cortes, Robert A. Long, Arpita Vyas, Olha Krichevskiy, Ryan Luellwitz, Pallavi Singh, and Martha H. Mulks. 2020. “Elimination of ‘Kitome’ and ‘Splashome’ Contamination Results in Lack of Detection of a Unique Placental Microbiome.” BMC Microbiology 20 (1): 157. “Organismal Metagenomes.” n.d. NCBI Taxonomy Browser. Accessed June 18, 2021. https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?mode=Undef&id=410656&lvl=3 &keep=1&srchmode=1&unlock. Penfold, Naomi C., and Jessica K. Polka. 2020. “Technical and Social Issues Influencing the Adoption of Preprints in the Life Sciences.” PLoS Genetics 16 (4): e1008565. Peterson, Roseann E., Karoline Kuchenbaecker, Raymond K. Walters, Chia-Yen Chen, Alice B. Popejoy, Sathish Periyasamy, Max Lam, et al. 2019. “Genome-Wide Association Studies in Ancestrally Diverse Populations: Opportunities, Methods, Pitfalls, and Recommendations.” Cell 179 (3): 589– 603. Pettorelli, Nathalie, Jos Barlow, Martin A. Nuñez, Romina Rader, Philip A. Stephens, Thomas Pinfield, and Erika Newton. 2021. “How International Journals Can Support Ecology from the Global South.” The Journal of Applied Ecology 58 (1): 4–8. 140 PLOS. 2019. “Trends in Preprints.” October 8, 2019. https://plos.org/blog/announcement/trends-in- preprints/. Popejoy, Alice B., and Stephanie M. Fullerton. 2016. “Genomics Is Failing on Diversity.” Nature. Poremski, Daniel, Bruno Falissard, Jörg Fegert, Andreas Witt, Anna E. Ordóñez, Andrés Martin, and Daniel Shuen Sheng Fung. 2019. “Moving from ‘Personal Communication’ to ‘Available Online at’: Preprint Servers Enhance the Timeliness of Scientific Exchange.” Child and Adolescent Psychiatry and Mental Health 13 (October): 42. PostgreSQL Global Development Group. 2017. PostgreSQL (version 9.6.6). https://www.postgresql.org. ———. 2018. “Procedural Languages.” PostgreSQL Documentation Version 9.4.20. 2018. https://www.postgresql.org/docs/9.4/xplang.html. Powell, Kendall. 2016. “Does It Take Too Long to Publish Research?” Nature 530 (7589): 148–51. Prodan, Andrei, Valentina Tremaroli, Harald Brolin, Aeilko H. Zwinderman, Max Nieuwdorp, and Evgeni Levin. 2020. “Comparing Bioinformatic Pipelines for Microbial 16S rRNA Amplicon Sequencing.” PloS One 15 (1): e0227434. Pylro, Victor S., Tsai S. Mui, Jorge L. M. Rodrigues, Fernando D. Andreote, Luiz F. W. Roesch, and Working Group Supporting the INCT Microbiome. 2016. “A Step Forward to Empower Global Microbiome Research Through Local Leadership.” Trends in Microbiology 24 (10): 767–71. Raff, Martin, Alexander Johnson, and Peter Walter. 2008. “Painful Publishing.” Science 321 (5885): 36. Ramírez-Castañeda, Valeria. 2020. “Disadvantages in Preparing and Publishing Scientific Papers Caused by the Dominance of the English Language in Science: The Case of Colombian Researchers in Biological Sciences.” PloS One 15 (9): e0238372. R Core Team. 2017. “R: A Language and Environment for Statistical Computing.” Vienna, Austria: R Foundation for Statistical Computing. https://www.R-project.org/. Reitz, Kenneth. 2018. Requests-HTML (version 0.9.0). Github. https://github.com/kennethreitz/requests- html. “Reporting Preprints and Other Interim Research Products.” n.d. National Institutes of Health. Accessed January 7, 2019. https://grants.nih.gov/grants/guide/notice-files/NOT-OD-17-050.html. Research Organization Registry. 2019. ROR API (version a3b153c). Github. https://github.com/ror- community/ror-api. Riesenberg, D., and G. D. Lundberg. 1990. “The Order of Authorship: Who’s on First?” JAMA: The Journal of the American Medical Association 264 (14): 1857. Ringelhan, Stefanie, Jutta Wollersheim, and Isabell M. Welpe. 2015. “I Like, I Cite? Do Facebook Likes Predict the Impact of Scientific Work?” PloS One 10 (8): e0134389. Rochmyaningsih, Dyna. 2018. “Did a Study of Indonesian People Who Spend Most of Their Days under Water Violate Ethical Rules?” Science, July. https://doi.org/10.1126/science.aau8972. Ross, Joseph S., Cary P. Gross, Mayur M. Desai, Yuling Hong, Augustus O. Grant, Stephen R. Daniels, Vladimir C. Hachinski, Raymond J. Gibbons, Timothy J. Gardner, and Harlan M. Krumholz. 2006. “Effect of Blinded Peer Review on Abstract Acceptance.” JAMA: The Journal of the American Medical Association 295 (14): 1675–80. Royle, Stephen. 2014. “What the World Is Waiting for.” Quantixed (blog). October 17, 2014. https://quantixed.org/2014/10/17/what-the-world-is-waiting-for/. 141 ———. 2015. “Waiting to Happen II: Publication Lag Times.” Quantixed (blog). March 16, 2015. https://quantixed.org/2015/03/16/waiting-to-happen-ii-publication-lag-times/. Rubinstein, Mara Roxana, Xiaowei Wang, Wendy Liu, Yujun Hao, Guifang Cai, and Yiping W. Han. 2013. “Fusobacterium Nucleatum Promotes Colorectal Carcinogenesis by Modulating E-Cadherin/β- Catenin Signaling via Its FadA Adhesin.” Cell Host & Microbe 14 (2): 195–206. Saposnik, Gustavo, Bruce Ovbiagele, Stavroula Raptis, Marc Fisher, and S. Claiborne Johnston. 2014. “Effect of English Proficiency and Research Funding on Acceptance of Submitted Articles to Stroke Journal.” Stroke; a Journal of Cerebral Circulation 45 (6): 1862–68. Sarabipour, Sarvenaz, Humberto J. Debat, Edward Emmott, Steven J. Burgess, Benjamin Schwessinger, and Zach Hensel. 2019. “On the Value of Preprints: An Early Career Researcher Perspective.” PLoS Biology 17 (2): e3000151. Saurman, Virginia, Kara G. Margolis, and Ruth Ann Luna. 2020. “Autism Spectrum Disorder as a Brain- Gut-Microbiome Axis Disorder.” Digestive Diseases and Sciences, February. https://doi.org/10.1007/s10620-020-06133-5. Šavrič, Bojan, Tom Patterson, and Bernhard Jenny. 2019. “The Equal Earth Map Projection.” International Journal of Geographical Information Science: IJGIS 33 (3): 454–65. Schloss, Patrick D. 2017. “Preprinting Microbiology.” mBio 8 (3). https://doi.org/10.1128/mBio.00438-17. Schloss, Patrick D., Sarah L. Westcott, Thomas Ryabin, Justine R. Hall, Martin Hartmann, Emily B. Hollister, Ryan A. Lesniewski, et al. 2009. “Introducing Mothur: Open-Source, Platform- Independent, Community-Supported Software for Describing and Comparing Microbial Communities.” Applied and Environmental Microbiology 75 (23): 7537–41. Schmid, M. W. 2016. crawlBiorxiv (version e2af128). Github. https://github.com/MWSchmid/crawlBiorxiv. Schoch, Conrad L., Stacy Ciufo, Mikhail Domrachev, Carol L. Hotton, Sivakumar Kannan, Rogneda Khovanskaya, Detlef Leipe, et al. 2020. “NCBI Taxonomy: A Comprehensive Update on Curation, Resources and Tools.” Database: The Journal of Biological Databases and Curation 2020 (January). https://doi.org/10.1093/database/baaa062. Schwarz, Greg J., and Robert C. Kennicutt Jr. 2004. “Demographic and Citation Trends in Astrophysical Journal Papers and Preprints.” arXiv [astro-Ph]. arXiv. http://arxiv.org/abs/astro-ph/0411275. “Science Funding.” 2019. Chan Zuckerberg Initiative. March 20, 2019. https://chanzuckerberg.com/science/science-funding/. “SDG Indicators.” n.d. United Nations Sustainable Development Goals. Accessed May 18, 2021. https://unstats.un.org/sdgs/indicators/regional-groups. Sender, Ron, Shai Fuchs, and Ron Milo. 2016. “Revised Estimates for the Number of Human and Bacteria Cells in the Body.” PLoS Biology 14 (8): e1002533. Sequence Read Archive Submissions Staff. 2011. Understanding SRA Search Results. National Center for Biotechnology Information (US). Serghiou, Stylianos, and John P. A. Ioannidis. 2018. “Altmetric Scores, Citations, and Publication of Studies Posted as Preprints.” JAMA: The Journal of the American Medical Association 319 (4): 402–4. Sever, Richard. 2018. “Public Twitter Post.” Twitter. November 1, 2018. https://twitter.com/cshperspectives/status/1058002994413924352. Sever, Richard, Michael Eisen, and John Inglis. 2019. “Plan U: Universal Access to Scientific and Medical Research via Funder Preprint Mandates.” PLoS Biology 17 (6): e3000273. 142 Sever, Richard, Ted Roeder, Samantha Hindle, Linda Sussman, Kevin-John Black, Janet Argentine, Wayne Manos, and John R. Inglis. 2019. “bioRxiv: The Preprint Server for Biology.” bioRxiv, November. https://doi.org/10.1101/833400. Shannon, C. E. 1948. “A Mathematical Theory of Communication.” Bell System Technical Journal. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x. Sharon, Gil, Nikki Jamie Cruz, Dae-Wook Kang, Michael J. Gandal, Bo Wang, Young-Mo Kim, Erika M. Zink, et al. 2019. “Human Gut Microbiota from Autism Spectrum Disorder Promote Behavioral Symptoms in Mice.” Cell 177 (6): 1600–1618.e17. Shin, Hakdong, Kenneth Price, Luong Albert, Jack Dodick, Lisa Park, and Maria Gloria Dominguez-Bello. 2016. “Changes in the Eye Microbiota Associated with Contact Lens Wearing.” mBio 7 (2): e00198. Shi, Wenyu, Heyuan Qi, Qinglan Sun, Guomei Fan, Shuangjiang Liu, Jun Wang, Baoli Zhu, et al. 2019. “gcMeta: A Global Catalogue of Metagenomics Platform to Support the Archiving, Standardization and Analysis of Microbiome Data.” Nucleic Acids Research 47 (D1): D637–48. Silk, Noon van der, Aram Harrow, Jaiden Mispy, Dave Bacon, Jackson Bates, Steven Flammia, Jonathan Oppenheim, et al. 2018. “About.” 2018. Smaglik, Paul. 1999. “E-Biomed Becomes PubMed Central.” The Scientist Magazine. September 26, 1999. https://www.the-scientist.com/news/e-biomed-becomes-pubmed-central-56359. Smith, Karen, Kathy D. McCoy, and Andrew J. Macpherson. 2007. “Use of Axenic Animals in Studying the Adaptation of Mammals to Their Commensal Intestinal Microbiota.” Seminars in Immunology 19 (2): 59–69. Snyder, Solomon H. 2013. “Science Interminable: Blame Ben?” Proceedings of the National Academy of Sciences of the United States of America 110 (7): 2428–29. Soo, Cassandra Claire, Freedom Mukomana, Scott Hazelhurst, and Michele Ramsay. 2017. “Establishing an Academic Biobank in a Resource-Challenged Environment.” South African Medical Journal = Suid-Afrikaanse Tydskrif Vir Geneeskunde 107 (6): 486–92. South, Andy. 2017. “World Map Data from Natural Earth [R Package Rnaturalearth Version 0.1.0],” March. https://cran.r-project.org/package=rnaturalearth. SRA Tools Wiki. n.d. Github. Accessed December 26, 2021. https://github.com/ncbi/sra-tools. Stuart, Tim. 2016. “bioRxiv.” 2016. http://timoast.github.io/blog/2016-03-01-biorxiv/. ———. 2017. “bioRxiv 2017 Update.” 2017. http://timoast.github.io/blog/biorxiv-2017-update/. “Submit a Manuscript.” n.d. bioRxiv. Accessed November 30, 2018. https://www.biorxiv.org/submit-a- manuscript. Sze, Marc A., and Patrick D. Schloss. 2016. “Looking for a Signal in the Noise: Revisiting Obesity and the Microbiome.” mBio 7 (4). https://doi.org/10.1128/mBio.01018-16. Taroni, Jaclyn N., Peter C. Grayson, Qiwen Hu, Sean Eddy, Matthias Kretzler, Peter A. Merkel, and Casey S. Greene. 2019. “MultiPLIER: A Transfer Learning Framework for Transcriptomics Reveals Systemic Features of Rare Disease.” Cell Systems 8 (5): 380–94.e4. Thelwall, Mike, Stefanie Haustein, Vincent Larivière, and Cassidy R. Sugimoto. 2013. “Do Altmetrics Work? Twitter and Ten Other Social Web Services.” PloS One 8 (5): e64841. The PLoS Medicine Editors. 2006. “The Impact Factor Game. It Is Time to Find a Better Way to Assess the Scientific Literature.” PLoS Medicine 3 (6): e291. 143 Tort, Adriano B. L., Zé H. Targino, and Olavo B. Amaral. 2012. “Rising Publication Delays Inflate Journal Impact Factors.” PloS One 7 (12): e53374. “Treasury & Endowments.” n.d. University of Minnesota Comptroller’s Office. Accessed January 29, 2022. https://controller.umn.edu/treasury-endowments/index.html. Tsosie, Krystal S., Keolu Fox, and Joseph M. Yracheta. 2021. “Genomics Data: The Broken Promise Is to Indigenous People.” Nature. “Upvote.pub Snapshot.” 2018. Internet Archive. 2018. https://web.archive.org/web/20180430180959/https://upvote.pub/. Vale, Ronald D. 2015. “Accelerating Scientific Publication in Biology.” Proceedings of the National Academy of Sciences of the United States of America 112 (44): 13439–46. Vangay, Pajau, Abigail J. Johnson, Tonya L. Ward, Gabriel A. Al-Ghalith, Robin R. Shields-Cutler, Benjamin M. Hillmann, Sarah K. Lucas, et al. 2018. “US Immigration Westernizes the Human Gut Microbiome.” Cell 175 (4): 962–72.e10. Varmus, Harold. 1999. “E-BIOMED: A Proposal for Electronic Publications in the Biomedical Sciences.” National Institutes of Health. Archive.org Snapshot, 18 Oct 2015. 1999. https://web.archive.org/web/20151018182443/https://www.nih.gov/about/director/pubmedcentral/ ebiomedarch.htm. Vence, T. 2017. “Journals Seek Out Preprints.” The Scientist, 2017. https://www.the-scientist.com/news- opinion/journals-seek-out-preprints-32183. Verma, Inder M. 2017. “Preprint Servers Facilitate Scientific Discourse.” Proceedings of the National Academy of Sciences of the United States of America 114 (48): 12630. Vogt, Nicholas M., Robert L. Kerby, Kimberly A. Dill-McFarland, Sandra J. Harding, Andrew P. Merluzzi, Sterling C. Johnson, Cynthia M. Carlsson, et al. 2017. “Gut Microbiome Alterations in Alzheimer’s Disease.” Scientific Reports 7 (1): 13537. Vos, Asha de. 2020. “The Problem of ‘Colonial Science.’” Scientific American, July 1, 2020. https://www.scientificamerican.com/article/the-problem-of-colonial-science/. Walsh, Tom, Jon M. McClellan, Shane E. McCarthy, Anjené M. Addington, Sarah B. Pierce, Greg M. Cooper, Alex S. Nord, et al. 2008. “Rare Structural Variants Disrupt Multiple Genes in Neurodevelopmental Pathways in Schizophrenia.” Science 320 (5875): 539–43. Walters, William A., Faviola Reyes, Giselle M. Soto, Nathanael D. Reynolds, Jamie A. Fraser, Ricardo Aviles, David R. Tribble, et al. 2020. “Epidemiology and Associated Microbiota Changes in Deployed Military Personnel at High Risk of Traveler’s Diarrhea.” PloS One 15 (8): e0236703. Wang, Z., W. Glänzel, and Y. Chen. 2018. “How Self-Archiving Influences the Citation Impact of a Paper: A Bibliometric Analysis of arXiv Papers and Non-arXiv Papers in the Field of Information and Library Science.” STI 2018 Conference Proceedings, September, 323–30. Way, Gregory P., and Casey S. Greene. 2018. “Extracting a Biologically Relevant Latent Space from Cancer Transcriptomes with Variational Autoencoders.” In Biocomputing 2018, 80–91. WORLD SCIENTIFIC. ———. 2019. “Discovering Pathway and Cell Type Signatures in Transcriptomic Compendia with Machine Learning.” Annual Review of Biomedical Data Science 2 (1): 1–17. Westcott, Sarah, and Patrick D. Schloss. 2021. Mothur “Sracommand.cpp” (version 1e6af22). GitHub. https://github.com/mothur/mothur/blob/ba42f8ddea4f30e8cf261633dd0ea5133f1f559a/source/com mands/sracommand.cpp#L80. 144 Wickham, Hadley. 2009. ggplot2: Elegant Graphics for Data Analysis. Springer Science & Business Media. Wilke, Andreas, Jared Bischof, Wolfgang Gerlach, Elizabeth Glass, Travis Harrison, Kevin P. Keegan, Tobias Paczian, et al. 2016. “The MG-RAST Metagenomics Database and Portal in 2015.” Nucleic Acids Research 44 (D1): D590–94. Wilkinson, Mark D., Michel Dumontier, I. Jsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, et al. 2016. “The FAIR Guiding Principles for Scientific Data Management and Stewardship.” Scientific Data 3 (March): 160018. Wilmanski, Tomasz, Noa Rappaport, John C. Earls, Andrew T. Magis, Ohad Manor, Jennifer Lovejoy, Gilbert S. Omenn, Leroy Hood, Sean M. Gibbons, and Nathan D. Price. 2019. “Blood Metabolome Predicts Gut Microbiome α-Diversity in Humans.” Nature Biotechnology 37 (10): 1217–28. Wolf, L., and S. Bileschi. 2005. “Combining Variable Selection with Dimensionality Reduction.” In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), 2:801–6 vol. 2. Wong, Janice C., Kimberly A. Fernandes, Shubarna Amin, Zarnie Lwin, and Monika K. Krzyzanowska. 2014. “Involvement of Low- and Middle-Income Countries in Randomized Controlled Trial Publications in Oncology.” Globalization and Health 10 (December): 83. Woodruff, Andy, and Cynthia Brewer. 2017. Colorbrewer. Github. https://github.com/axismaps/colorbrewer. “World Population Prospects 2019, Online Edition. Rev. 1.” 2019. United Nations, Department of Economic and Social Affairs, Population Division. 2019. https://population.un.org/wpp/Download/Standard/Population/. Wuchty, Stefan, Benjamin F. Jones, and Brian Uzzi. 2007. “The Increasing Dominance of Teams in Production of Knowledge.” Science 316 (5827): 1036–39. Xia, Jingfeng, Jennifer L. Harmon, Kevin G. Connolly, Ryan M. Donnelly, Mary R. Anderson, and Heather A. Howard. 2015. “Who Publishes in ‘predatory’ Journals?” Journal of the Association for Information Science and Technology 66 (7): 1406–17. Yatsunenko, Tanya, Federico E. Rey, Mark J. Manary, Indi Trehan, Maria Gloria Dominguez-Bello, Monica Contreras, Magda Magris, et al. 2012. “Human Gut Microbiome Viewed across Age and Geography.” Nature 486 (7402): 222–27. Yegorov, Sergey, Dmitriy Babenko, Samat Kozhakhmetov, Lyudmila Akhmaltdinova, Irina Kadyrova, Ayaulym Nurgozhina, Madiyar Nurgaziyev, et al. 2020. “Psoriasis Is Associated With Elevated Gut IL-1α and Intestinal Microbiome Alterations.” Frontiers in Immunology 11 (October): 571319. Zeng, M. Y., N. Inohara, and G. Nuñez. 2017. “Mechanisms of Inflammation-Driven Bacterial Dysbiosis in the Gut.” Mucosal Immunology 10 (1): 18–26. 145 Appendix A: Chapter One Supplementary Material Supplementary Figure A-1. Downloads per preprint by months available. The x-axis indicates the number of months a given preprint has been available online; the y-axis indicates the number of downloads the preprint received in that month. Information is summarized by box plots indicating the first quartile, median and third quartile for each month designation. 146 Supplementary Figure A-2. Proportion of downloads per preprint by months available. The x-axis indicates the number of months a given preprint has been available online; the y-axis indicates the proportion of downloads received in a preprint’s first year that were received in that month. For example, if a preprint received 100 downloads in its first 12 months, and received 22 downloads in its first month, its value here would be 0.22. Information is summarized by box plots indicating the first quartile, median and third quartile for each month designation. 147 Supplementary Figure A-3. Annual publication rates and estimates. The x-axis indicates year, and the y-axis indicates the proportion of preprints posted in that year that were later published. The point in each year indicates the observed rate, and the error bars indicate the 95 percent confidence interval for an estimate of the proportion that was published, when accounting for publications not detected by the bioRxiv system (see Methods). 148 Supplementary Figure A-4. Multiple perspectives on per-preprint download counts. (a) A box plot indicating the number of downloads per preprint, limited to a preprint’s first month online. The x-axis indicates the year in which the preprint was posted, and the y- axis indicates downloads. (b) A box plot indicating the number of downloads received in each preprint’s best month of downloads. The x-axis indicates the year in which the preprint was first posted, and the y-axis indicates the maximum monthly download count for each preprint. (c) A box plot indicating the number of downloads per preprint received in 2018. The x-axis indicates the year a preprint was first posted, and the y-axis indicates the number of downloads received by each preprint in 2018. 149 Supplementary Figure A-5. Total downloads per preprint from each year. A series of density plots indicating the total downloads for preprints posted in each displayed year. The grey bars on the right side indicate the year being visualized. The x-axis indicates the total downloads received for a single preprint (log scale), and the y-axis indicates the likelihood a given preprint would be observed with that number of downloads, a smoothed approximation of a histogram. 150 Supplementary Table A-1. Top 15 institutions by author count. Each institution’s count of total preprints is based on the number of papers posted by authors currently listed with those affiliations, but preprints attributed to authors from multiple institutions count toward the total for all institutions mentioned. A paper with multiple authors from the same institution is counted only once for that institution. Institution Authors Preprints Stanford University Ž,‘ Ž,’“ University of Oxford Ž,Ž”• ”’• University of Cambridge Ž,Ž’” –• University of Washington ”• —’” University College London –’Ž — University of Pennsylvania — “ University of Michigan —‘ – University of California, San Francisco “’ “ŽŽ University of California, San Diego •“ ”“ Imperial College London ’‘ • University of Edinburgh —— – University of California, Berkeley —•’ “•– Yale University “““ ‘”• Duke University ““ ‘•‘ Harvard University “‘• ““ 151 Supplementary Table A-2. Total preprints published per journal. This table compares the total number of articles published by the top 20 journals that have published the most bioRxiv preprints, compared to how many of those articles appeared on bioRxiv. The list is ordered by the proportion of 2018 publications first appearing on bioRxiv. Total article counts are from the number of works in the “article” category as indexed by Web of Science (Clarivate Analytics). Journals marked with an asterisk have a large number of published works categorized on Web of Science as “meeting abstracts”; for consistency, those are not included in the counts here, though it is possible some of the preprints published by these journals fall into that category. PQRS PQRT through PQRS Journal Total Preprints Proportion Total Preprints Proportion GigaScience •– —— —–.——% ˜™š –› œ—.››% Genome Biology •˜ ™˜ ˜–.•–% ,—š • š.•% Genome Research ž– žœ ˜ž.ž–% ,›œ› ™— ™.›ž% eLife ™œ ˜–— ˜˜.žœ% ž,• ™š› œ.œž% Nature Methods ˜™ —ž ˜˜.š•% •—™ – —.›š% PLOS Computational Bio. —–› š˜ ˜.œœ% ˜,˜› ˜˜™ ›.•% Nature Genetics •• šš œ–.œž% ,ž– œš ›.ž–% G˜ ˜—— –š œ™.žœ% ,–žœ œ—ž œ.š—% Genetics œ™ž ™ œš.™œ% ,•–– œž– —.™% Bioinformatics •œ– œ›– œš.œ% —,ž–— —™œ ›.›ž% PLOS Genetics šš ™ œœ.™œ% ˜,••š œ–œ ™.šœ% Molecular Bio. & Evolution œœ• —™ œ›.ž% ,š—— š™ ›.™% PLOS Biology ˜˜˜ ž™ œ›.œ% ,˜™– ›• ™.•˜% Genome Bio. and Evolution –– —› œ›.›% ,š™ ›– ž.–—% Molecular Bio. of the Cell* œš ˜• ™.ž™% œ,›œ •• —.˜™% mBio ˜™˜ ž— ™.ž% œ,˜•˜ ›• —.š˜% NeuroImage •žœ ™ ˜.š™% š,•ž œ— —.˜% Biophysical Journal* —•œ žš ˜.—–% ˜,›˜ ›– ˜.š% BMC Bioinformatics —šž ž ˜.˜•% ˜,™ž ˜™ —.˜% Journal of Neuroscience ™™™ –• œ.ž% ž,šž– ™œ œ.žœ% 152 Appendix B: Chapter Two Supplementary Material The supplementary tables from this chapter are too large to practically reproduce in print. Their legends have been reproduced below, but the content has been archived in CSV format at Zenodo and is available for download here: https://doi.org/10.5281/zenodo.3909824 Supplementary Table B-1. Preprints per country. Each row represents a single country, sorted in descending order by the “senior_author” and “any_author” columns. The “alpha2” column indicates the two-letter country code defined in ISO 3166-1. “country” indicates the country name as recorded in the ROR dataset. “senior_author” lists the number of bioRxiv preprints for which the final author in the author list specified an affiliation to an institution in that country. “any_author” lists the number of bioRxiv preprints for which at least one author (in any position) specified an affiliation to an institution in that country. Supplementary Table B-2. Country productivity and bioRxiv adoption. Each row represents a single country, sorted in descending order by the “citable_total” and “senior_author_preprints” columns. The “alpha2” column indicates the two-letter country code defined in ISO 3166-1. “country” indicates the country name as recorded in the SCImago dataset. The “y2014” through “y2018” columns list the total number of citable documents attributed to that country in the SCImago dataset for the year specified. “citable_total” indicates the sum of all citable documents from that country from 2014 through 2018. “senior_author_preprints” lists the number of senior-author preprints attributed to that country from 2013 through 2019. Supplementary Table B-3. Combinations of senior authors with collaborator countries. Each row represents a combination of two countries, sorted alphabetically by the “contributor” and “senior” columns. The “contributor” column indicates the name of the contributor country. “senior” indicates the name of the country that appears as a senior author. “count” lists the number of preprints that include at least one author listing an affiliation from the country in the “contributor” column and a senior author listing an affiliation from the country in the “senior” column. Supplementary Table B-4. Links between contributor countries and the senior- author countries they write with. Each row represents a combination of two countries. The “contributor” column indicates the name of the contributor country. “senior” indicates the name of the country that appears as a senior author. “p” lists the p-value of a Fisher’s exact test, as described in the “Methods” section. “with” lists the number of preprints that 153 include an author from the “contributor” country and a senior author from the “senior” country. “without” lists the number of preprints that include an author from the “contributor” country but do not list a senior author from the “senior” country. “seniortotal” lists the total number of senior-author preprints attributed to the country in the “senior” column. “padj” lists the p-value from the “p” column, adjusted to control the false- discovery rate using the Benjamini–Hochberg procedure. Supplementary Table B-5. Published pre-2019 preprints by country. Each row represents a country, sorted in descending order by the “published” and “total” columns. The “country” column indicates the country name as recorded in the ROR dataset. “total” lists the number of preprints last updated prior to 2019 that list a senior author who declared an affiliation in the specified country. “published” lists, of the preprints counted in the “total” column, the number that are listed as published on the bioRxiv website. Supplementary Table B-6. Journal–country links. Each row represents a combination of country and journal, sorted in ascending order using the “padj” column, then descending order using the “preprints” column. “country” indicates the name of a country as recorded in the ROR dataset. “journal” indicates the name of a journal that has published preprints from the specified country. “preprints” indicates the number of preprints last updated prior to 2019 that were published by the specified journal that list a senior author affiliated with the specified country. “expected” indicates the number of preprints we would expect the specified journal to have published from the specified country, if the country and journal both published the same number of papers, but the journal’s publications mirrored the country-level proportions observed in published bioRxiv preprints overall. “p” indicates the p-value of a chi-squared test as described in the “Methods” section. “padj” lists the p- value from the “p” column, adjusted to control the false-discovery rate using the Benjamini–Hochberg procedure. “journaltotal” lists the total preprints published by the specified journal that were last updated on bioRxiv prior to 2019. “countrytotal” lists the total preprints posted to bioRxiv prior to 2019 that list a senior author affiliated with the specified country. Supplementary Table B-7. Preprint counting methods at the country level. Each row represents a country, sorted in descending order using the “cn_total” and “straight_count” columns. The “country” column is the country name as recorded in the ROR dataset. “cn_total” lists the number of preprints attributed to that country using the completenormalizing counting technique. “straight_count” lists the number of preprints attributed to that country using the straight-counting technique. 154 Supplementary Table B-8. Publication rates and DOI usage. Each row represents a country, sorted alphabetically. The “doi_rate” field lists the percentage of published papers from that country issued a Digital Object Identifier (DOI), according to Boudry and Chartron (2017). The “pub_rate” field lists the proportion of preprints from that country posted before 2019 that have been published. Supplementary Figure B-1. Preprint collaboration. (a) shows the average number of authors per paper over time. The x-axis indicates the year; the y-axis indicates the harmonic mean authors per preprint. Each point indicates the average of papers posted in a single month; the blue line indicates the six-month moving average. (b) illustrates the number of countries per preprint, over time. The x-axis indicates time; the y-axis indicates the arithmetic mean countries per preprint. Each point indicates the average unique countries found in all preprints posted in a single month. The blue line indicates the six-month moving average. 155 Supplementary Figure B-2. Correlation between three measurements of international collaboration. This figure is an alternative presentation of the same data as the three panels in Supplementary Figure B-3. Each point represents a country, and the size of the point indicates the total international preprints associated with that country. The x-axis indicates the proportion of preprints with a contributor from that country that also include at least one contributor from another country. The y-axis indicates the proportion of those preprints for which that country appears in the senior author position. 156 Supplementary Figure B-3. International collaboration correlations. Each point represents a country; the red points indicate those in the ‘contributor country’ category. Blue lines indicate lines of best fit for each plot, though they are unrelated to the Spearman correlations reported for these relationships. (a) A scatter plot showing the relationship (Spearman’s ρ=0.781, p=1.09×10−14) between a country’s total international preprints (x- axis; log scale) and the proportion of those preprints for which they are the senior author (y-axis). (b) A scatter plot showing the relationship (Spearman’s ρ=−0.578, p=3.68×10−7) between a country’s total international preprints (x-axis; log scale) and the proportion of preprints with a contributor from that country that also include at least one contributor from another country (y-axis). (c) A scatter plot showing the relationship (Spearman’s ρ=−0.572, p=5.32×10−7) between the proportion of preprints with a contributor from that country that also include at least one contributor from another country (x-axis) and the proportion of those preprints for which that country appears in the senior author position (y-axis). 157 Appendix C: Chapter Three Supplementary Material Supplementary data files and documentation are available at https://doi.org/10.5281/zenodo.5351180. The supplementary tables from this chapter are too large to practically reproduce in print. Their legends have been reproduced below, but the content has been archived in CSV format at Zenodo and is available for download here: https://doi.org/zenodo.5351180 Supplementary Table C-1. Samples per tag. Each row represents a single metadata field available for BioSample entries. The “samples” column indicates how many samples have a value for that field. Supplementary Table C-2. Samples per body site per country. This contains similar data to Supplementary Table C-2, except no countries or body sites are omitted. Each column is a single NCBI Taxonomy entry. Each row is a country, and each cell represents the number of samples from that country that appeared in that body site. Supplementary Table C-3. Top 10 countries by body site. Each column holds a list of the 10 countries with the most samples in a single body site. The “unknown” category is omitted here. Supplementary Table C-4. Country-level data. Each row represents a single country or territory as defined by the United Nations. There are 10 columns; see the supplementary documentation for a description of them. Supplementary Table C-5. NCBI Taxonomy IDs. Each row represents a single body site. The “human” column indicates the ID used to identify samples explicitly labeled as human (e.g. “human gut metagenome”); the “generic” column indicates the ID used to identify samples not labeled as human (e.g. “gut metagenome”). 158 Supplementary Figure C-1. Samples per year. The x-axis indicates the year; the y-axis indicates the number of microbiome samples released in that year. Colors indicate the region of origin for each sample and match the colors used in Figure 3-1c and 3-1d.