1,034 research outputs found

    Evolution of Privacy Loss in Wikipedia

    Full text link
    The cumulative effect of collective online participation has an important and adverse impact on individual privacy. As an online system evolves over time, new digital traces of individual behavior may uncover previously hidden statistical links between an individual's past actions and her private traits. To quantify this effect, we analyze the evolution of individual privacy loss by studying the edit history of Wikipedia over 13 years, including more than 117,523 different users performing 188,805,088 edits. We trace each Wikipedia's contributor using apparently harmless features, such as the number of edits performed on predefined broad categories in a given time period (e.g. Mathematics, Culture or Nature). We show that even at this unspecific level of behavior description, it is possible to use off-the-shelf machine learning algorithms to uncover usually undisclosed personal traits, such as gender, religion or education. We provide empirical evidence that the prediction accuracy for almost all private traits consistently improves over time. Surprisingly, the prediction performance for users who stopped editing after a given time still improves. The activities performed by new users seem to have contributed more to this effect than additional activities from existing (but still active) users. Insights from this work should help users, system designers, and policy makers understand and make long-term design choices in online content creation systems

    Measuring internet activity: a (selective) review of methods and metrics

    Get PDF
    Two Decades after the birth of the World Wide Web, more than two billion people around the world are Internet users. The digital landscape is littered with hints that the affordances of digital communications are being leveraged to transform life in profound and important ways. The reach and influence of digitally mediated activity grow by the day and touch upon all aspects of life, from health, education, and commerce to religion and governance. This trend demands that we seek answers to the biggest questions about how digitally mediated communication changes society and the role of different policies in helping or hindering the beneficial aspects of these changes. Yet despite the profusion of data the digital age has brought upon us—we now have access to a flood of information about the movements, relationships, purchasing decisions, interests, and intimate thoughts of people around the world—the distance between the great questions of the digital age and our understanding of the impact of digital communications on society remains large. A number of ongoing policy questions have emerged that beg for better empirical data and analyses upon which to base wider and more insightful perspectives on the mechanics of social, economic, and political life online. This paper seeks to describe the conceptual and practical impediments to measuring and understanding digital activity and highlights a sample of the many efforts to fill the gap between our incomplete understanding of digital life and the formidable policy questions related to developing a vibrant and healthy Internet that serves the public interest and contributes to human wellbeing. Our primary focus is on efforts to measure Internet activity, as we believe obtaining robust, accurate data is a necessary and valuable first step that will lead us closer to answering the vitally important questions of the digital realm. Even this step is challenging: the Internet is difficult to measure and monitor, and there is no simple aggregate measure of Internet activity—no GDP, no HDI. In the following section we present a framework for assessing efforts to document digital activity. The next three sections offer a summary and description of many of the ongoing projects that document digital activity, with two final sections devoted to discussion and conclusions

    Social Structure and Mechanisms of Collective Production: Evidence from Wikipedia

    Get PDF
    In my dissertation I propose three counterintuitive social mechanisms to alleviate the risk that collective production will fail to maintain participant involvement and respond to demand. My first study, based on a panel dataset of edits and views of articles in the English Wikipedia, shows that, although collective production lacks a price-like mechanism to estimate demand for the goods it produces, consumers’ contributions act as such a signal to expert producers. In the second paper I examine the theory that collective production participation is greatest when social norms of collaboration are obeyed. Using a large panel dataset of production networks and normrelated behavior in Wikipedia, I show that social norm infringement is not completely detrimental to participation because norm enforcement increases the likelihood that the beneficiary producer continues participating. In my third paper, I rely on interviews with experienced Wikipedia producers to examine whether producers’ ties to non-participants in collective production increase the likelihood of turnover, and whether producers’ embeddedness in collective production reduces turnover risk. Surprisingly, I find that producers with networks rich in ties to non-producers and with a task-oriented approach to collective production are those least likely to stop participating
    • …
    corecore