<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD with MathML3 v1.4 20241031//EN" "JATS-journalpublishing1-4-mathml3.dtd">
<article article-type="research-article" xml:lang="EN" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">LIBER</journal-id>
<journal-title-group>
<journal-title>LIBER QUARTERLY</journal-title>
</journal-title-group>
<issn pub-type="epub">2213-056X</issn>
<publisher>
<publisher-name>openjournals.nl</publisher-name>
<publisher-loc>The Hague, The Netherlands</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">lq.23531</article-id>
<article-id pub-id-type="doi">10.53377/lq.23531</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Relevance and Diversity &#x2013; Towards Selection Criteria for Archiving the German Web</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-9718-0698</contrib-id>
<name>
<surname>Woldering</surname>
<given-names>Britta</given-names>
</name>
<email>b.woldering@dnb.de</email>
<xref ref-type="aff" rid="aff1"/>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-2576-4640</contrib-id>
<name>
<surname>Gottschalk</surname>
<given-names>Simon</given-names>
</name>
<email>gottschalk@L3S.de</email>
<xref ref-type="aff" rid="aff2"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Heinrichs</surname>
<given-names>Randi</given-names>
</name>
<email>randi.heinrichs@leuphana.de</email>
<xref ref-type="aff" rid="aff3"/>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-3374-2193</contrib-id>
<name>
<surname>Nejdl</surname>
<given-names>Wolfgang</given-names>
</name>
<email>nejdl@L3S.de</email>
<xref ref-type="aff" rid="aff2"/>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-0892-610X</contrib-id>
<name>
<surname>Neuberger</surname>
<given-names>Christoph</given-names>
</name>
<email>christoph.neuberger@weizenbaum-institut.de</email>
<xref ref-type="aff" rid="aff4"/>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-3799-1146</contrib-id>
<name>
<surname>Weller</surname>
<given-names>Katrin</given-names>
</name>
<email>katrin.weller@gesis.org</email>
<xref ref-type="aff" rid="aff5"/>
</contrib>
<aff id="aff1">Domain Access and Engagement &#x2013; Science, German National Library, Frankfurt a.M., Germany</aff>
<aff id="aff2">L3S Research Center, Leibniz Universit&#x00E4;t Hannover, Hannover, Germany</aff>
<aff id="aff3">Center for Digital Cultures, Leuphana University Luneburg, Luneburg, Germany</aff>
<aff id="aff4">Weizenbaum Institute, Freie Universit&#x00E4;t Berlin, Berlin, Germany</aff>
<aff id="aff5">Data Services for the Social Sciences, GESIS &#x2013; Leibniz Institute for the Social Sciences, Cologne, Germany</aff>
</contrib-group>
<pub-date pub-type="epub">
<month>07</month>
<year>2026</year>
</pub-date>
<volume>36</volume>
<fpage>1</fpage>
<lpage>30</lpage>
<permissions>
<copyright-statement>Copyright 2026, The copyright of this article remains with the author</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://www.liberquarterly.eu/article/10.53377/lq.23531"/>
<abstract>
<p>The motivation for this paper is the need of the German National Library to massively increase its web archiving activities as the 12,000 snapshots collected per year are considered not to be enough to reflect the German web appropriately. As collecting everything which could be defined as the German web is impossible, the concept of selecting an &#x201C;exemplary diversity&#x201D; is introduced and substantiated by the guiding principles of relevance and diversity. The long-term goal is to find methods to (semi-)automatise the selection process and the prerequisite for this are well-defined selection criteria and sources of relevant as well as diverse topics that can be exploited. First steps to operationalise the guiding principles are outlined and supplemented by technological challenges and requirements to support the selection process as well as in the light of the advancement of web technology. Finally, the paper proposes a mix of curational and automatised measures and sources of information that should be taken into consideration to cater for a good coverage of relevance and diversity, supplemented by technological aspects regarding alternative representations of web content in web archives to be discussed to bring the content of web archives closer to the original user experience.</p>
</abstract>
<kwd-group>
<kwd>web archiving</kwd>
<kwd>national libraries</kwd>
<kwd>digitalisation</kwd>
<kwd>national web</kwd>
<kwd>selection criteria</kwd>
<kwd>web technology</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1. Introduction</title>
<p>With this paper, we reflect on challenges posed to Germany&#x2019;s National Library (Deutsche Nationalbibliothek, DNB) in the context of the new media and publishing landscape that is shaped by digitalisation processes. The DNB&#x2019;s mandate is to collect media works related to Germany, i.e. published in Germany, in German language and about Germany. Since 2006, the legal deposit mandate of the DNB comprises media works in non-physical form, hence also the collection of websites. The DNB started web archiving in 2012 and currently the web archive of the DNB holds about 70,000 snapshots of about 9,000 websites, selected by librarians. This is in opposition to 17 million registered .de-domains and presumably even more websites related to Germany registered under generic domains. Therefore, there is a need for the DNB to massively increase its web archiving activities. If the collection activities are to be extended by factor 100 or even 1,000, it becomes clear that a manually curated selection of what should be archived is impossible. The long-term goal is to find methods to (semi-)automatise the selection process. A prerequisite for this are well-defined selection criteria and sources of relevant as well as diverse topics that can be exploited to create a collection that represents the German web. As an overview of the totality of the German web reflecting the collection mandate is not at hand, a representative selection is impossible. The guiding principle the DNB has chosen to escape this dilemma, is &#x201C;exemplary diversity&#x201D; &#x2013; a term created by Christoph Classen at a workshop on web archiving at the DNB in 2018 and describing the attempt to select a variety as broad and unbiased as possible without the aspiration to be complete or representative. A second principle is the mission of the DNB (2024) to preserve the &#x201C;cultural heritage.&#x201D; Our research question is therefore: How can the DNB implement and technically realise the principles of &#x201C;exemplary diversity&#x201D; and &#x201C;cultural heritage&#x201D; when selecting content on the German web?</p>
<sec id="s1a">
<title>1.1. Background: The Role of Digital Transformations</title>
<p>Digitalisation has fundamentally changed the public sphere since around the mid-1990s. Large parts of publishing and public communication have shifted to the Internet. Media use &#x2013; such as the reception of news (<xref ref-type="bibr" rid="r6">Behre et al., 2024</xref>; <xref ref-type="bibr" rid="r55">Newman et al., 2023</xref>) &#x2013; is also increasingly taking place online. In contrast, older media, especially printed media, are becoming less relevant. The digital transformation not only affects traditional mass media but goes far beyond this: the whole society is mirrored in the digital public sphere including cultural, political and commercial activities. As a result of digitalisation, the public sphere is also expanding into new areas, such as everyday communication, which has become particularly visible in social media. There are several characteristics which clearly distinguish the digital public sphere from the analogue public sphere and which must be considered when it comes to preserving the digital cultural heritage.</p>
<list list-type="bullet">
<list-item><p><italic>Loss of memory</italic>: Contrary to the assumption that the Internet does not forget, it is apparent that considerable parts of the offerings published in the digital public sphere are not permanently accessible but disappear (<xref ref-type="bibr" rid="r16">Chapekis et al., 2024</xref>). This also applies to scientific sources (<xref ref-type="bibr" rid="r32">Gomes et al., 2021</xref>) and journalistic offerings (<xref ref-type="bibr" rid="r27">Duffy &#x0026; Flynn, 2021</xref>).</p></list-item>
<list-item><p><italic>Increasing quantity</italic>: Due to lower technological barriers and thus the broad participation of the former &#x201C;audience&#x201D;, the amount of information increases considerably, especially on social media platforms.</p></list-item>
<list-item><p><italic>Decreasing quality</italic>: The ability to bypass media (disintermediation) means that there is no need for quality checks anymore before publication (<xref ref-type="bibr" rid="r12">Bruns, 2018</xref>). Journalism lost its role as &#x201C;gatekeeper&#x201D;. Platform companies have little incentive to check quality (<xref ref-type="bibr" rid="r72">Zuboff, 2019</xref>). As a result, there are major differences in quality between digital offerings.</p></list-item>
<list-item><p><italic>Collapse of contexts</italic>: Boundaries are dissolving between forms of communication, including between mass communication and individual communication, between different media (convergence), between the public and private sphere or journalism and advocacy (advertising, public relations). This &#x201C;collapse of contexts&#x201D; (<xref ref-type="bibr" rid="r19">Davis &#x0026; Jurgenson, 2014</xref>) makes it difficult to assign offerings to specific categories.</p></list-item>
<list-item><p><italic>Individualisation</italic>: The Internet expands the range of possible choice and thus promotes the individualisation of use (high-choice media environment; <xref ref-type="bibr" rid="r66">Van Aelst et al., 2017</xref>), both through the active choice of users and through algorithmic selection.</p></list-item>
<list-item><p><italic>Inequality</italic>: Despite the large number of offerings, there is a strong focus on a small number of offerings on the audience and thus a considerable inequality in the distribution of visibility and influence (<xref ref-type="bibr" rid="r39">Hindman, 2009</xref>).</p></list-item>
</list>
<p>As a consequence of the digitalisation of the publication market and the communication in cultural, political and commercial activities, the collection mandates of national libraries were extended in many countries to comprise digital media to preserve the cultural heritage in the digital age (<xref ref-type="bibr" rid="r4">Altenh&#x00F6;ner &#x0026; Schrimpf, 2014</xref>; <xref ref-type="bibr" rid="r62">Schmitt, 2011</xref>; <xref ref-type="bibr" rid="r68">Vlassenroot et al., 2019</xref>, <xref ref-type="bibr" rid="r69">2021</xref>). This also applies to the German National Library and the law it is based on. According to the German National Library Act (Gesetz &#x00FC;ber die Deutsche Nationalbibliothek; DNBG), it is the task of the DNB to collect all media works (&#x201E;Medienwerke&#x201C;). Media works are defined in &#x00A7; 3 (1) as follows: &#x201C;Media works are all representations in writing, image and sound that are distributed in physical form or made accessible to the public in non-physical form&#x201D; (translated by the authors). The task is not limited to distinct media. What is required is the publication, i.e. public accessibility of media works, and this includes websites and social media.</p>
<p>The collection mandate (<xref ref-type="bibr" rid="r24">DNB, 2023</xref>) applies without restriction: &#x201C;Our collection mandate aims for completeness, regardless of the form of the media works or the form of publication&#x201D; (translated by the authors). Completeness might be aimed at regarding distinct media like e-books or e-journals, but it cannot be achieved in the case of web archiving for several reasons. First, the totality of the &#x201C;German web&#x201D; is unclear as it comprises much more than the .de domain: all websites with relation to Germany, German language, culture or history as well as websites whose owner&#x2019;s place of business is in Germany. Besides the DENIC list,<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> it should be hard to get a complete overview of websites fulfilling these criteria, let alone to keep up with the constant change in new websites arising and websites which are taken down. Second, even after identifying a set of relevant websites, to crawl and archive all their versions over time is an impossible task since websites are constantly updated, content is added or deleted without notice and it is impossible to archive all versions (<xref ref-type="bibr" rid="r14">Br&#x00FC;gger, 2018</xref>).</p>
<p>As completeness of the collection is impossible to achieve in archiving the German web, the DNB is confronted for the first time in its history with the question of how to define selection criteria (<xref ref-type="bibr" rid="r22">DNB, 2016</xref>) and to justify them (<xref ref-type="bibr" rid="r4">Altenh&#x00F6;ner &#x0026; Schrimpf, 2014</xref>). A key aspect of the collection mandate of the DNB is to collect &#x201C;without judging content&#x201D; (<xref ref-type="bibr" rid="r23">Deutsche Nationalbibliothek [DNB], 2017</xref>, p. 11) which does not conform with the idea of a content-related selection process. However, it can be interpreted that the selection criteria should be as neutral and consensual as possible and well justified. Guiding principles are the terms &#x201C;exemplary diversity&#x201D; and &#x201C;cultural heritage.&#x201D; How these principles are to be interpreted must be reviewed permanently, also with the participation of different social groups. The meaning of these principles cannot be definitively fixed. In view of transnationally oriented media such as the web, the limitation of the collection mandate to an individual nation also requires a new, differentiated interpretation.</p>
</sec>
<sec id="s1b">
<title>1.2. Scope and Structure of the Paper</title>
<p>The motivation for this article is the DNB&#x2019;s need to massively increase its web archiving activities. Before proposing our guiding principles, we provide a closer reflection on the challenges of archiving the web regarding content selection and technology in section two. Section three then addresses social media as a special case of web archiving. In section four, we introduce our approach for operationalising the guiding principles of relevance and diversity in four dimensions. The section is completed by the technological requirements to address the challenges of the advanced web technology as well as to support a (semi-)automated selection process, complemented by aspects of transparency to be considered in the selection process. Finally, in section five, recommendations for further steps to operationalise the concept of relevance and diversity and how to cope with the abundance of the German web are given.</p>
<p>When full archiving is not possible, decisions about selection processes should be the starting point for building and expanding a web archive. Therefore, this paper focuses on the discussion about how to define processes for deciding what should be archived. Based on these measures could be identified and applied to enable the selection process. It is obvious that a broad expansion of a web archive bears technological challenges for the whole life cycle of web archiving as well as organisational and financial issues, but these must be addressed in a later step after first having defined content related selection criteria. The technological, organisational and financial limitations that will occur when implementing the collection concept will have an impact on the scale, but not on the basic content selection decisions. Hence, in this paper, technological issues are only discussed in the context of enabling the suggested selection criteria while the full operationalisation to address these issues must be discussed subsequently in a different context. The problems and developments regarding crawling, indexing, searching and use of web archives as well as long term preservation and legal issues (e.g., data privacy and copyright) are out of scope of this paper.</p>
</sec>
</sec>
<sec id="s2">
<title>2. Challenges in Archiving the Web</title>
<p>The capacity of digital technology to collect, connect and save information seems to suggest unlimited access to knowledge and all-encompassing archives. However, those promises of comprehensiveness distort the understanding of digital and furthermore new web archives and more attention needs to be paid to the politics of selection and practicability. Scholars in the emerging field of critical data studies have pointed out that a rhetoric has emerged around digitalisation and datafication processes that goes far beyond their technical capacities and forgets or even obscures the work required to secure, maintain and care for these systems (<xref ref-type="bibr" rid="r17">Christin, 2020</xref>; <xref ref-type="bibr" rid="r25">D&#x2019;Ignazio &#x0026; Klein, 2020</xref>; <xref ref-type="bibr" rid="r28">Elish &#x0026; Boyd, 2017</xref>; <xref ref-type="bibr" rid="r42">Irani, 2015</xref>). What at first sounds straightforward, namely digitising analogue books, storing digital publications and now also archiving websites, involves its own and very specific &#x201C;infra-political processes&#x201D;, as <xref ref-type="bibr" rid="r64">Thylstrup (2018)</xref> emphasises in Politics of Mass Digitization. <xref ref-type="bibr" rid="r65">Thylstrup et al. (2021)</xref> bring critical debates around big data into dialogue with archival theories, emphasising that it is not the supposed sense of certainty but the limitations and uncertainties that must be considered and examined to understand and think through the newly emerging regimes of cultural memory.</p>
<p>Building on the post-structuralist theories of the archive in the mid-20th century &#x2013; including the critical positions developed in response from feminist, queer and post-colonial perspectives (<xref ref-type="bibr" rid="r8">Best, 2018</xref>; <xref ref-type="bibr" rid="r15">Butler, 2009</xref>; <xref ref-type="bibr" rid="r36">Hartman, 1997</xref>) &#x2013; cultural studies has been discussing an archival turn, which is now receiving renewed scholarly attention in relation to digital transformations. In classical cultural studies theories, such as those of <xref ref-type="bibr" rid="r30">Foucault (1970)</xref> and <xref ref-type="bibr" rid="r21">Derrida and Prenowitz (1995)</xref>, archives have been described as the places in which the order of knowledge is formed since ancient times. Archives provide the foundations and infrastructures of what can be researched, remembered or known and thus also shape what is not considered important, not collected and potentially forgotten. This must be taken into account by the institutions and practitioners responsible for these new archives, both in terms of diversity and relevance of content and the technical conditions of selecting, crawling and preserving.</p>
<sec id="s2a">
<title>2.1. Challenges Regarding Content</title>
<p>The DNB&#x2019;s mission is to archive everything that was published in Germany, in German language and about Germany. With traditional publications these could be matched to selection criteria that would work in a setting of formal publishing cultures. However, on the web, publishing does not follow formal and standardised procedures. Therefore, a new selection strategy needs to be defined that takes into account content-related criteria. Regarding print media the DNB applies formal criteria like a minimum circulation figure or a minimum volume to decide whether a publication is ingested in the collection. In the case of web archiving the selection criteria cannot be determined purely formally as the characteristics of the web as media differs significantly from print media as is shown in the list below, hence criteria must be defined content-related. The selection must be based not only on the mission of archive libraries to preserve cultural heritage for the long term, but also on the current user demand. The question of expectations of future generations can only be answered hypothetically. These selection criteria must be justified, made transparent and be revisited regularly.</p>
<p>The digital public sphere has characteristics that make the systematic description and selection of web offers for archiving difficult. This results in a number of challenges for the DNB:</p>
<list list-type="bullet">
<list-item><p><italic>Web offers and their characteristics:</italic> The unknown totality of web offers makes a representative selection impossible. There is no continuous recording of all offerings as in the case of press and broadcasting, for which complete directories are available.</p></list-item>
<list-item><p><italic>Unit of selection:</italic> It is not easy to determine how to delimit the units to be collected. External content can be embedded on a website. Providers (organisations, individuals) can operate a large number of accounts on different digital platforms at the same time (multichannel communication). The discursive context is also of interest, i.e. the connection of a large number of contributions in networks on social media (e.g. through links or hashtags) or in discussion forums (threads). Topics are dealt with in many places on the web. They spread across the boundaries of platforms (spill-over). National borders are also frequently crossed.</p></list-item>
<list-item><p><italic>Categories</italic>: Formats of publication continue to develop or are only vaguely defined. Predefined formats or genres would help to systematise the digital offerings (e.g. <xref ref-type="bibr" rid="r44">Kampes, 2021</xref>).</p></list-item>
<list-item><p><italic>Continuous production</italic>: The web is not a completed, stable publication like a newspaper edition; continuous updates of different parts of the web that partly appear in frequent intervals need to be handled.</p></list-item>
<list-item><p><italic>Variability</italic>: Offers are not uniform for the whole audience as in mass media, but are personalised to individual users. Generative AI provides individual responses on demand based on external content, lacking transparency with regard to the origin of the data used and the production process. In addition, the presentation varies depending on the used technical device.</p></list-item>
<list-item><p><italic>Usage data</italic>: There is a lack of valid data on the different uses of digital offers. Platforms are only willing to share data to a limited extent.</p></list-item>
<list-item><p><italic>Multimedia</italic>: Elements such as text, photos, graphics, videos, audio, and animation must be archived together.</p></list-item>
</list>
<p>Overall, the digital public sphere is a comprehensive, less structured representation of society as a whole, containing various modes of access to the world (knowledge, entertainment, culture, advertising and other forms of persuasive communication, everyday communication).</p>
<p>A further basic part of the definition of selection criteria is the decision whether basically everything on the (national) web may be collected or if certain categories of content must be excluded. This question is mainly relevant for dealing with potentially illegal content. This could be, for example, violence, hate speech and online harassment, or sexual content &#x2013; depending on respective legal frameworks that are in place. It can be argued that everything on the web can be regarded as published, hence is part of the collection mandate. As legality of the content cannot be examined in detail before archiving, the DNB applies the common rule that it takes down illegal content and excludes it from being indexed, if appropriate, as soon as the DNB is notified of illegal content in its web archive. Generally, everything on the German web is seen as part of the historical tradition and worth to be archived.</p>
</sec>
<sec id="s2b">
<title>2.2. Technological Challenges</title>
<p>Defining selection criteria for a national web archive is first of all content-related, but technological aspects are closely interlinked to content-related decisions, e.g. the appraisal of the relevance of a website or how to deal with personalisation and the dynamics of websites or the use of focused crawlers.</p>
<p>The web has seen a dramatic alteration since the first initiatives for web archiving, specifically the foundation of the Internet Archive in 1996. Consequently, Laursen and M&#x00F8;ldrup-Dalum recognised in <xref ref-type="bibr" rid="r47">2017</xref> that &#x201C;the archive separates itself increasingly from the live web the archive tries to preserve&#x201D;. Essential reasons for the divergence between the live web and the archived web are, on the one hand, the development of the web structure and the role of links as well as the continuous dynamisation and personalisation of the web. Both characteristics immediately affect the performances and modes of operation of classic web crawlers, whose selection strategies are only to a limited extent applicable to today&#x2019;s web.</p>
<p>Hyperlinks were the central element of the Hypertext Markup Language (HTML) and interlinked all available information in the World Wide Web (<xref ref-type="bibr" rid="r7">Berners-Lee, 1999</xref>). Web crawlers typically follow this paradigm and traverse hyperlinks to find potentially relevant websites. However, the role of hyperlinks has seen several developments (<xref ref-type="bibr" rid="r37">Helmond, 2019</xref>). Initially only being utilised as navigational elements, their distribution was soon used as an indicator of a website&#x2019;s importance with the PageRank algorithm developed in 1996 as a famous example. Since then, new types of hyperlinks, such as deep links and app links (<xref ref-type="bibr" rid="r37">Helmond, 2019</xref>), dynamic links and asynchronous triggered links, have been established. Social media have introduced platform-specific link characteristics like the usage of URL shorteners as well as the focus on intra-website links (e.g., retweets) have led to further fragmentation of the web (<xref ref-type="bibr" rid="r37">Helmond, 2019</xref>; <xref ref-type="bibr" rid="r49">Lee et al., 2009</xref>). Consequently, the paradigm that each URL uniquely identifies a page on the web, initially only queried via the GET method,<xref ref-type="fn" rid="fn2"><sup>2</sup></xref> has been softened up.</p>
<p>On top of the link structure, new web technologies have made web archiving more complex. The advent of CSS, Javascript and client-side executed scripts (Ajax) had already been identified as a major challenge to web archiving in 2012 which called for new web page processing in indexing methodologies (<xref ref-type="bibr" rid="r33">Gyllstrom et al., 2012</xref>) and JavaScript had been responsible for 33.2&#x0025; more missing resources in 2012 than in 2005 (<xref ref-type="bibr" rid="r11">Brunelle et al., 2016</xref>). Nowadays, websites are typically built in a modular way, consisting of several files and relying on complex server-side processes and user interaction which require server-side preservation strategies for preservation of their contents (<xref ref-type="bibr" rid="r67">Bo&#x010D;yt&#x0117; &#x0026; de Vos, 2018</xref>).</p>
<p>Web crawlers have only partially been adapted to these changes of the web structure. Focused crawling techniques aim to identify web pages related to a topic of interest, but they require the identification of topic-relevant web pages, which has become increasingly inefficient and noisy as the complexity of the web has grown (<xref ref-type="bibr" rid="r1">Aggarwal, 2019</xref>). While Google&#x2019;s web crawler has allowed the use of POST requests since 2021,<xref ref-type="fn" rid="fn3"><sup>3</sup></xref> there remains an open question of which requests are suitable for web archiving. Also, the web archive should reflect the structure of the web since users want to conduct analyses of the granularity of single websites and the structure of the web in general (<xref ref-type="bibr" rid="r40">Holzmann &#x0026; Nejdl, 2021</xref>). Hence, new web crawlers need to drift away from the original concepts of the web and instead find new solutions that enable them to extract and store as complete information as possible from the contents of web pages.</p>
<p>Due to the personalisation of the web, two users can see different versions of the same web page at the same time.<xref ref-type="fn" rid="fn4"><sup>4</sup></xref> <xref ref-type="bibr" rid="r45">Kelly et al. (2013)</xref> have identified different cases potentially leading to such a situation, including different website layouts (e.g., mobile vs desktop mode) and local website content based on the user&#x2019;s IP address or GPS position (e.g., local news and weather forecasts). Exceptional cases of personalisation are advertisements and social media: recommendations for posts and products or displayed trends can be tailored to user groups, and different user groups might be shown different content via microtargeting for political influence (<xref ref-type="bibr" rid="r60">Ribeiro et al., 2019</xref>).</p>
<p>To deal with personalised content, web crawlers can either attempt to disable any personalised information, e.g., by declaring themselves as bots in the request header, or by taking the role of different personae to retrieve different views from the same website. In the former case, the challenge lies in the definition of user profiles and the identification of websites that actually offer a relevant diversity in information regarding personalisation. When viewing a snapshot in a web archive, users should get access to metadata specifying which personalisation is behind the respective stored snapshots (<xref ref-type="bibr" rid="r45">Kelly et al., 2013</xref>).</p>
</sec>
<sec id="s2c">
<title>2.3. Resource Limitations</title>
<p>Limited resources are one of the main reasons why the web cannot be archived in its entirety. <xref ref-type="bibr" rid="r49">Lee et al. (2009)</xref> define three dimensions of resource requirements when creating web archives:</p>
<list list-type="bullet">
<list-item><p>Scalability: the number of web pages that a crawler can proceed &#x2013; specifically, if the time and resource usage of the web crawler scales linearly with the number of pages to be crawled.</p></list-item>
<list-item><p>Performance: the speed at which the crawler discovers the web.</p></list-item>
<list-item><p>Resource usage: the CPU and RAM consumption during crawling. In addition, the disk space is a major requirement: the Wayback Machine of the Internet Archive consumed 57 PetaByte<xref ref-type="fn" rid="fn5"><sup>5</sup></xref> in December 2021.</p></list-item>
</list>
<p>While we do not yet have exact ways to measure the size of the German web, some data may help to illustrate the potential dimensions: <xref ref-type="fig" rid="fg001">Figure 1</xref> is based on the analysis of the content of .de snapshots in the Internet Archive between 1994 and 2013 performed by <xref ref-type="bibr" rid="r41">Holzmann et al. (2016)</xref>, and highlights the challenges regarding the size of the German web offers with more than 5 billion snapshots until 2013 including 400 million images. Both numbers undoubtedly have increased substantially in recent years.</p>
<fig id="fg001">
<label>Fig. 1:</label>
<caption><p>Number of snapshots of .de URLs in the Internet Archive per file type between November 1994 and September 2013. The application data type contains file formats such as PDF and JSON.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="figures/LIBER_2026_36_Woldering_fig1.jpg"/></fig>
</sec>
</sec>
<sec id="s3">
<title>3. Online Platforms and Social Media</title>
<p>Finally, for preserving the web and our digital heritage, archival institutions also need to consider activities that are happening within a variety of online platforms. In the context of the web being shaped through platformisation, platforms are viewed as &#x201C;(re-)programmable digital infrastructures that facilitate and shape personalised interactions among end-users and complementors, organised through the systematic collection, algorithmic processing, monetisation, and circulation of data&#x201D; (<xref ref-type="bibr" rid="r58">Poell et al., 2019</xref>, p. 3). Online platforms exist for various domains and purposes, including notably e-commerce or entertainment. Also, a broad category of online platforms is social media.</p>
<p>Over the years, social media platforms have grown into a viable part of the web. They are not only being used by individuals, but also by numerous official actors and institutions, such as media, politicians and government agencies, cultural institutions, NGOs, or celebrities. Such platforms used both for personal and professional communication settings are of particular interest, contribute significantly to the information that can be found on the web, and thus are also relevant from a web archiving perspective. Their relevance for today&#x2019;s web landscape is not just due to size but also comes from the unique quality of their content. With their multifaceted contributions from different user communities they represent important cultural and even historical artefacts: For example, they carry the native traces of several protest movements such as the Arab Spring (<xref ref-type="bibr" rid="r35">Harlow &#x0026; Johnson, 2011</xref>) or the Occupy Wall Street (<xref ref-type="bibr" rid="r20">DeLuca et al., 2012</xref>) movement, and have been used to raise awareness to societal challenges at large scale, for example with the #metoo hashtag pointing out sexual harassment.</p>
<p>Archiving from web platforms, however, comes with its own additional challenges for the web archiving community (<xref ref-type="bibr" rid="r69">Vlassenroot et al., 2021</xref>). The collaboration between the Library of Congress (LoC) in the U.S. and Twitter Inc. has been the only attempt so far to professionally preserve a social media platform in its entirety &#x2013; and it did not lead to the desired success. The decision by the LoC to shift to selective collection (<xref ref-type="bibr" rid="r50">Library of Congress, 2017</xref>) is linked to some of the more general challenges that web archivists have to face for preserving social media platform content, e.g., capturing multimedia content shared through social media and handling outgoing links to other websites or other platform content.</p>
<p>Furthermore, platforms are not stable and the entire platform context may change over time. In the case of Twitter, the takeover by Elon Musk has led to numerous changes of how the platform operates including changes of the platform brand from Twitter to X. This has affected the platform&#x2019;s user community and how users interact on and with the platform. It has also heavily reduced the options for accessing data from the platform. The availability of APIs for accessing online platforms is frequently changing, with the closure of the previous Facebook API in 2018<xref ref-type="fn" rid="fn6"><sup>6</sup></xref> being another impactful event for everyone aiming at collecting and preserving platform data. Partly closures of APIs as access points have inspired implementing alternative approaches such as web scraping, but these are also often technically restricted by the platforms. There is ongoing work to find new solutions for enabling data access (e.g., at the Social Media Archive SOMAR)<xref ref-type="fn" rid="fn7"><sup>7</sup></xref> and work on new models for making online platform data accessible for research while also respecting data protection regulations (<xref ref-type="bibr" rid="r29">European Digital Media Observatory (EDMO), 2022</xref>). Overall, the question of what can be accessed and thus archived from online platforms has to reflect different technical challenges, but also legal constraints (including the ones imposed by platforms through their terms of services) and ethical considerations (for example around collecting sensitive information from vulnerable groups).</p>
<p>Various efforts to collect and preserve social media content exist, at national libraries (<xref ref-type="bibr" rid="r69">Vlassenroot et al., 2021</xref>) as well as in academia ranging from individual researchers to academic institutions interested in social media content as research data (<xref ref-type="bibr" rid="r70">Weller &#x0026; Kinder-Kurlanda, 2016</xref>). Therefore, the selection processes are often influenced by specific current research interests, rather than by specific preservation strategies and long-term perspectives.</p>
<p>A selection strategy for social media archiving as part of a national web archive will have to take several factors into consideration, including:</p>
<list list-type="bullet">
<list-item><p><italic>Platform selection</italic>: As a first step it needs to be decided which platforms should be included in archival processes. This decision might have to be made based on practical considerations, e.g., technical feasibility. It could also be informed by the size of the user community, or the impact a platform has in a specific country.</p></list-item>
<list-item><p><italic>Content selection</italic>: Assuming that full coverage of an entire platform is likely not feasible, additional strategies have to be found to create subsets. This could for example be based on topics (requiring good strategies to translate a given topical interest into a search query on the platform), random data, person-centred or networked approaches.</p></list-item>
<list-item><p><italic>Format selection:</italic> Finally, strategies have to be defined on how to archive content and make it accessible. A fundamental question here could be to decide between purely text-based formats and multimedia content, potentially even preserving the &#x201C;look and feel&#x201D; of the platform.</p></list-item>
</list>
<p>Considering social media as part of a general web archiving strategy is desirable for preserving the cultural heritage of the social web. The sheer size of content produced on social media platforms makes selection strategies necessary, but this is challenging due to its very ephemeral nature. Furthermore, access limitations as well as the need to quickly react to new developments (new platforms, new platform features, content changes like deletions or edits) requires very specialised setups and dedicated resources.</p>
</sec>
<sec id="s4">
<title>4. Operationalisation</title>
<p>In order to fulfil its collecting mission of exemplary diversity, the DNB strives for a selection based on relevance and diversity as central criteria. There is an obvious tension between relevance and diversity as selection criteria: relevance demands strict selection and the focusing of attention on a few important elements in a dimension. In contrast, the demand for diversity requires the broad representation of as many topics, actors or opinions as possible, regardless of their importance and quality. Both criteria are important for liberal democracy as well as for an archive of the German web.</p>
<sec id="s4a">
<title>4.1. Relevance and Diversity</title>
<p>Relevance is measured in approaches of communication studies via operationalisations of visibility and influence in the public sphere (as a concept see <xref ref-type="bibr" rid="r61">Schatz &#x0026; Schulz, 1992</xref>). Indicators for visibility can be, for example, the usage of offers, number of citations and mentions of actors, and the number of contributions to a topic. Indicators for influence can be agenda-setting effects and opinion power ascribed to (&#x201C;leading&#x201D;) media and actors like politicians in high positions. A special approach to define the relevance of topics is the &#x201C;news value&#x201D; (<xref ref-type="bibr" rid="r31">Galtung &#x0026; Ruge, 1965</xref>).</p>
<p>Relevance in the public sphere can be supplemented by relevance indicators defined from the perspective of functional subsystems of society (such as politics, economy, science, art, or sports). These subsystems have their own measures of importance. For example, prominence of scientists in public often differs from their reputation in the science system itself. The system-specific relevance becomes obvious through the roles of the actors in the system&#x2019;s own hierarchy (e.g., the chancellor in the political system) or the results of internal recognition mechanisms (e.g., awards). Following <xref ref-type="bibr" rid="r9">Bourdieu (1985)</xref>, the definition of relevance at the heteronomous pole of a field is oriented toward outside influences and audience tastes, while at the autonomous pole relevance assessment follows internal standards of quality and self-defined societal goals of professions.</p>
<p>Other forms of social differentiation also have an impact on relevance assessments, e.g., regions, milieus, or classes. Therefore, relevance must always be defined in a distinct perspective. This can lead to a variety of different definitions of relevance. For the purpose at hand, it is important to follow the most consensual and stable relevance assessments possible.</p>
<p>Diversity as a norm means that all interests of population groups and all manifestations of society shall be represented in the public sphere, regardless of the quality, visibility or influence of these interests and manifestations. Diversity can refer to different dimensions of content, including events, topics, opinions, actors, spaces, and formats. The norm of diversity has been intensively discussed and empirically investigated in the context of media and public sphere (e.g., <xref ref-type="bibr" rid="r38">Hendrickx et al., 2022</xref>; <xref ref-type="bibr" rid="r43">Joris et al., 2020</xref>; <xref ref-type="bibr" rid="r51">Loecherbach et al., 2020</xref>; <xref ref-type="bibr" rid="r54">Napoli, 1999</xref>) to which the DNB can refer.</p>
<p>The operationalisation of diversity as well as of relevance depends on several normative decisions. In the first step, dimensions of diversity must be defined for both sides: for society&#x2019;s representation and preferences of audience groups. Then, in the second step, the different manifestations must be found in one dimension, for example, forms of everyday culture in different milieus or different opinions on a given issue. Ultimately, it is a question of capacity, determining how much diversity can be documented and archived, since the manually curated selection of diverse material is complex and time-consuming. Relevance and diversity can be operationalised meticulously and with high complexity which detracts the practical manageability and raises many questions of detail. A compromise between a documented rationale and practicability needs to be found.</p>
</sec>
<sec id="s4b">
<title>4.2. Dimensions of Relevance and Diversity</title>
<p>Regarding relevance and diversity the following four dimensions need to be considered.</p>
<sec id="s4b1">
<title>4.2.1. Dimension &#x201C;Topics&#x201D;</title>
<p>The first decision in the selection process for web archiving is about the topics. Different from print media or broadcast, there are no established thematic structures or genres for online media so far (<xref ref-type="bibr" rid="r44">Kampes, 2021</xref>). Kampes distinguishes online offers into information based, e-commerce, social media, search engines and games. She analysed the elements of the titles of information based online offers and condensed the following 23 top genres which are defined further in her thesis: digital, sports, lifestyle, guidance, finances, knowledge, entertainment, regional, culture, cars, health, travel, family, business, style, games, football, career, politics, forum, news, real estate, newsletter (translation by the authors).</p>
<p>The assumption is that information providers label their offers to be understood intuitively by their users and to be found easily by search engines. Hence, although there are no established structures or genres for online media, a de facto structure created by search engine optimisation can be defined. It is important to point out that the results of the analysis show a long tail distribution, and the above mentioned 23 genres are the top ones which are used very often while a lot more topics or genres are used much less.</p>
<p>Having such top topics as starting points, they have to be explored further regarding their subsystems, and finally with regard to relevance and diversity within these subsystems.</p>
</sec>
<sec id="s4b2">
<title>4.2.2. Dimension &#x201C;Actors&#x201D;</title>
<p>The dimension &#x201C;actors&#x201D; comprises individual persons and collective actors, such as organisations, social movements etc. Relevance can be defined by formal positions, roles which are connected with high visibility (attention, range, prominence), high esteem (positive rating, reputation, appreciation) or influence (power in different aspects). This can be determined by lists, rankings, statistics, experts etc.</p>
<p>Depending on the topic and the respective social subsystem separate taxonomies have to be used or created. For example, in politics roles and positions are highly formalised which facilitates the selection process: top institutions on state and federal state level as well as organised interest groups. In contrast, the topic lifestyle is barely formalised and characterised by high volatility (fashions, trends) and a strong differentiation by milieus. To approach this, several subsystems need to be distinguished, and their taxonomies defined. They might be defined sociologically, i.e. lifestyle distinguished by its aspiration for singularity (<xref ref-type="bibr" rid="r59">Reckwitz, 2020</xref>), distinction (<xref ref-type="bibr" rid="r10">Bourdieu, 2001</xref>) and experience (<xref ref-type="bibr" rid="r63">Schulze, 1992</xref>), or by different forms of culture, e.g. popular culture. Popular culture again might be differentiated into fields like fashion, leisure activities or lifestyle products.</p>
<p>This shows that relevance and diversity are closely linked and that even the relevance aspect needs diversity. To gain even more diversity besides relevant actors, public and amateur actors should be considered. For example, in politics, members of the public, honorary politicians, citizens&#x2019; movements or, in lifestyle, people interested in food, health, nature and environment, all kinds of leisure activities and amateur communities in this field. Further aspects of diversity can be included by considering socio-demographic characteristics, e.g., age, gender, religion, level of education, phase of life or milieu.</p>
</sec>
<sec id="s4b3">
<title>4.2.3. Dimension &#x201C;Opinion&#x201D;</title>
<p>For the dimension &#x201C;opinion&#x201D; also the aspects of relevance and diversity apply. Opinion leaders, trendsetters, influencers in a broader sense are indicators for relevance in the subsystems of the different topics and again have to be chosen in a large variety to show the whole &#x2013; or at least a large &#x2013; spectrum of opinions.</p>
</sec>
<sec id="s4b4">
<title>4.2.4. Dimension &#x201C;Space&#x201D;</title>
<p>For the dimension &#x201C;space&#x201D;, the different spatial levels should be taken into account, i.e. transnational, international, federal, states, regional, or local aspects of a topic.</p>
</sec>
</sec>
<sec id="s4c">
<title>4.3. Technological Requirement and Automated Assessment</title>
<p>Manually curating selection processes require high personal resources. Therefore, technologies that can support or conduct the selection process will be needed to fully implement the selection based on relevance and diversity.</p>
<sec id="s4c1">
<title>4.3.1. Technological Requirements</title>
<p>As a first step towards this goal, a set of technological requirements for the algorithms and tools to be created can be identified to address the challenges and the strategies for operationalisation discussed above:</p>
<list list-type="bullet">
<list-item><p>Technical operationalisation of selection criteria that are defined around the concepts of relevance and diversity and their interdependencies.</p></list-item>
<list-item><p>Web crawling strategies that cope with the structure of today&#x2019;s dynamic web.</p></list-item>
<list-item><p>Consideration of personalisation of content, either by targeting at the archival of non-personalised contents or by disclosing the identities adopted when archiving personalised content in the metadata of the crawl.</p></list-item>
<list-item><p>Re-visit policies that define the frequency of archiving websites, which should be adaptive with respect to the domain, topics and type of a website.</p></list-item>
<list-item><p>Rules to identify national contents that are in line with the specific characteristics of the targeted nation. To determine websites beyond the .de domain to be collected, subdomains, language and geo data can be factored in as shown in strategies of other national libraries (<xref ref-type="bibr" rid="r68">Vlassenroot et al., 2019</xref>).</p></list-item>
<list-item><p>Methodologies to ensure the quality of archived websites regarding appropriateness of the sources and the preservation of their original contents and look and feel.</p></list-item>
</list>
<p>Facing this number of technological requirements, the concepts and methodologies behind traditional web archiving need to be questioned as key aspects of today&#x2019;s web can no longer be reconciled with traditional methods of web archiving: web archive content is not presented in a personalised way, dynamic and embedded components of web pages are often not archived, and user interactions in the archived web are oftentimes not possible.</p>
</sec>
<sec id="s4c2">
<title>4.3.2. Automated Assessment of Relevance and Diversity</title>
<p>Sources of information that allow the automated identification of relevant topics include news articles (<xref ref-type="bibr" rid="r48">Leban et al., 2014</xref>), Wikipedia (<xref ref-type="bibr" rid="r53">Miz et al., 2020</xref>), social networks (<xref ref-type="bibr" rid="r56">Panagiotou et al., 2016</xref>) and scientific literature (<xref ref-type="bibr" rid="r18">Chen, 2006</xref>). However, automated topic identification based on these sources typically focuses on the detection of current trends and events (i.e., related to the criteria of news value as in Section 4.1), thus neglecting topics without strong temporal focus. Further, a critical analysis by <xref ref-type="bibr" rid="r34">Hanada et al. (2013)</xref> on the use of link analysis metrics such as link counts and article lengths for document selection has revealed that such metrics are more correlated to popularity than to quality and importance, hence questioning their use as unbiased selection strategies. In contrast, an approach for an unbiased and diverse selection of websites is the random selection of web contents, e.g., from the DENIC list.</p>
<p>With the rise of Large Language Models (LLMs) in recent years, a pressing question is whether LLMs can be used as tools to advance web archiving and address the challenges and issues raised. To date, while the use of LLMs for tasks such as web scraping has shown promising use cases (<xref ref-type="bibr" rid="r57">Padoan &#x0026; Vinciguerra, 2024</xref>), their use for web crawling has not been fully explored (<xref ref-type="bibr" rid="r3">Al Galib et al., 2023</xref>). Although LLMs may struggle to identify trends and new websites on their own, we foresee their use to navigate the web for crawling and relevance assessment purposes. LLMs have recently been used to simulate human behaviour as demonstrated in Turing Experiments in different settings (<xref ref-type="bibr" rid="r2">Aher et al., 2023</xref>). These capabilities of LLMs have been utilised to simulate web navigation behaviour of users, targeting at solving specific tasks such as searching for weather reports (<xref ref-type="bibr" rid="r46">Lai et al., 2024</xref>). Adapting these approaches to the needs of web crawlers is a promising direction, allowing crawlers to emulate human-like navigation. Such LLM-based web navigation could specifically be used in focused crawling, which extracts relevant links from web pages and automatically assesses page relevance. While doing so, the LLM could even be prompted or fine-tuned to include the relevance and diversity criteria in its decision.</p>
<p><xref ref-type="bibr" rid="r5">Arabzadeh et al. (2025)</xref> have demonstrated that LLMs can accurately assess the relevance of items like websites or documents, showing high alignment with human labels. However, first indications exist that LLMs may not be neutral in their assessment. For example, <xref ref-type="bibr" rid="r71">Ye et al. (2025)</xref> have recently demonstrated cases of LLMs&#x2019; biases involved in the judgement processes, such as &#x201C;the tendency to assign more credibility to statements made by authority figures, regardless of actual evidence&#x201D;. Better understanding these processes is critical when applying LLMs for judging the relevance and diversity criteria in the context of web archiving. Another challenge in using LLMs for relevance assessment lies in their inherent intransparency despite the need to make selection criteria and decisions transparent as discussed in the following.</p>
</sec>
</sec>
<sec id="s4d">
<title>4.4. Establishing High-quality and Transparent Selection and Archiving Procedures</title>
<p>Where archiving of the full (national) web is not possible and selection criteria need to be applied by an institution, transparency about the selection criteria that have led to the inclusion or exclusion of websites can directly impact the perceived quality of the archive. Selection criteria directly influence the composition of the collection and therefore become an integral part of the archive themselves. Only if selection criteria are documented and possible inappropriate content is defined, users will be able to comprehend why specific contents are included and what might be missing. Transparency is particularly important for sensitive topics, for example, when preparing a topical collection on LGBTQ&#x002B; websites: users of the archive could receive information on why certain websites were included. This can help to avoid &#x201C;blind spots&#x201D; (<xref ref-type="bibr" rid="r26">Donig et al., 2023</xref>) and to reflect on potential biases that may still be inherent to the collection.</p>
<p>Apart from transparency on selection criteria, quality-control should also respect aspects of transparency on technical dimensions. <xref ref-type="bibr" rid="r52">Maemura et al. (2018)</xref> define the following three elements of provenance or transparency criteria to document a specific crawl:</p>
<list list-type="bullet">
<list-item><p><italic>Scoping</italic>: e.g., motivation for the collection, resulting in the focus, the seed-list, the crawl timing and configuration, inclusions and exclusions</p></list-item>
<list-item><p><italic>Process</italic>: How are scheduled and unscheduled events or process anomalies handled?</p></list-item>
<list-item><p><italic>Context</italic>: legal context, institutional setting and mandate, policies and guidelines, available resources for web archiving</p></list-item>
</list>
<p>Selection criteria can be presented in general for the complete crawl, or individually for each crawled website, e.g., as part of their metadata.</p>
</sec>
</sec>
<sec id="s5">
<title>5. Conclusion and Recommendation</title>
<p>This paper was motivated by the German National Library&#x2019;s need to extend its web archiving activities massively and to define selection criteria which can be supported by semi-automated processes. As collecting everything which could be defined as the German web is impossible, the concept of selecting an &#x201C;exemplary diversity&#x201D; was introduced and substantiated by the guiding principles of relevance and diversity. Technological challenges and requirements to support the selection process as well as in the light of the advancement of web technology were described. Finally, the factors to be considered for establishing a selection strategy for social media archiving as part of a web archive were outlined.</p>
<sec id="s5a">
<title>5.1. Measures of Relevance and Diversity</title>
<p>A mix of curational and automatised measures and sources of information should be taken into consideration to cater for a good coverage of relevance and diversity. A selection of such measures is presented in <xref ref-type="table" rid="tb001">Table 1</xref>, also summarising the automatised measures as outlined in chapter 4.3.</p>
<table-wrap id="tb001">
<label>Table 1:</label>
<caption><p>Selection of curational and automatised measures.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top"/>
<th align="left" valign="top">Curational</th>
<th align="left" valign="top">Automatised</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Relevance</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Round tables with experts, e.g. scientific associations and organised groups of society</p></list-item>
<list-item><p>Representative population surveys regarding relevant topics in society</p></list-item>
<list-item><p>Thematic campaigns consulting citizens</p></list-item>
<list-item><p>Citizen/user suggestion scheme</p></list-item>
</list></td>
<td align="left" valign="top">Intelligent crawlers trained with large language models<break/>Exploit preselected types of offerings<break/>
<list list-type="bullet">
<list-item><p>News offers</p></list-item>
<list-item><p>News aggregators</p></list-item>
<list-item><p>Trend monitors</p></list-item>
<list-item><p>Ranking lists</p></list-item>
<list-item><p>Wikipedia</p></list-item>
<list-item><p>Hashtags on social media</p></list-item>
<list-item><p>Web tracking data/studies<sup><xref ref-type="fn" rid="fn11">1</xref></sup></p></list-item>
<list-item><p>Web resources in e-publications</p></list-item>
<list-item><p>Scientific topics from Google Scholar, <ext-link ext-link-type="uri" xlink:href="http://arxiv.org">arxiv.org</ext-link>, &#x2026;</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">Diversity</td>
<td align="left" valign="top"/>
<td align="left" valign="top">Random selections from<break/>
<list list-type="bullet">
<list-item><p>the DENIC list</p></list-item>
<list-item><p>Crawls of websites having a German address in their legal notice</p></list-item>
</list></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="fn11"><p><sup>1</sup>For example, the GESIS Panel Digital Behavioural Data Sample (<ext-link ext-link-type="uri" xlink:href="https://www.gesis.org/gesis-panel">https://www.gesis.org/gesis-panel</ext-link>).</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s5b">
<title>5.2. Representations of Web Content in an Archive</title>
<p>Apart from the content-related selection process guided by such measures, alternative representations of web content in web archives need to be discussed to bring the content of web archives closer to the original user experience. Hence, there is a need to question the following three aspects:</p>
<list list-type="bullet">
<list-item><p><italic>Web crawlers</italic>: There exists a variety of web crawlers (ranging from command line tools such as GNU Wget<xref ref-type="fn" rid="fn8"><sup>8</sup></xref> to Heritrix<xref ref-type="fn" rid="fn9"><sup>9</sup></xref> of the Internet Archive), tools and frameworks relevant to the creation of web archives that are explained in more detail by <xref ref-type="bibr" rid="r68">Vlassenroot et al. (2019)</xref>. The suitability of these tools for operationalising complex selection criteria elaborated in this article must be examined.</p></list-item>
<list-item><p><italic>Data structures</italic>: The majority of web crawlers create or process content in the WARC file format,<xref ref-type="fn" rid="fn10"><sup>10</sup></xref> recognised as the standard format for web archive content, particularly by national libraries. The expressivity of WARC archives in the context of the requirements, specifically with respect to the metadata and transparency requirements, needs to be assessed.</p></list-item>
<list-item><p><italic>Interaction</italic>: The use of web archives rarely reflects the experiences in the live web, due to incomplete, non-preserving website snapshots and incomplete crawls. To bring the content of web archives closer to the original user experience, alternative representations of web content in web archives need to be discussed.</p></list-item>
</list>
</sec>
<sec id="s5c">
<title>5.3. Next Steps</title>
<p>As the next steps to establish a semi-automated selection process, we envision the following activities:</p>
<list list-type="bullet">
<list-item><p>The concept of relevance and diversity has to be elaborated further, based on the genres approach developed by <xref ref-type="bibr" rid="r44">Kampes (2021)</xref>. These genres have to be investigated further regarding their relevant sub-topics, actors, regional differences and sources of information and how to evaluate these results on a regular basis. Social media should be part of this approach, both as a source of information and as an object to be archived.</p></list-item>
<list-item><p>To support the aspect of diversity, the web archiving institution should aim to open the selection process to various groups of persons to participate. In the table above possible groups and measures for participation are listed. These are complementing each other and should ensure that different views and aspects in the selection process are considered.</p></list-item>
<list-item><p>The use of Large Language Models (LLMs) for web crawling should be further explored, in particular focused crawling as an example application of LLMs to extract relevant links from a web page and to automatically assess the relevance.</p></list-item>
</list>
<p>Having outlined the necessary next steps towards selection criteria for archiving the German web, we aim to elaborate them further in collaborative projects. As future work, concrete operationalisations should be explored, implemented and evaluated, also considering scalability and resource-efficiency. Balancing out the aim for completeness against the current restrictions and limitations requires solutions that facilitate the necessary selection process in a transparent and understandable way. Our concept for using relevance and diversity as main selection criteria is the proposed way of addressing this tension.</p>
</sec>
</sec>
</body>
<back>
<fn-group>
<title>Notes</title>
<fn id="fn1"><p>DENIC provides the Domain Name System (DNS) for the .de domain.</p></fn>
<fn id="fn2"><p><ext-link ext-link-type="uri" xlink:href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Basics_of_HTTP/Evolution_of_HTTP">https://developer.mozilla.org/en-US/docs/Web/HTTP/Basics_of_HTTP/Evolution_of_HTTP</ext-link>.</p></fn>
<fn id="fn3"><p><ext-link ext-link-type="uri" xlink:href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Basics_of_HTTP/Evolution_of_HTTP">https://developer.mozilla.org/en-US/docs/Web/HTTP/Basics_of_HTTP/Evolution_of_HTTP</ext-link>.</p></fn>
<fn id="fn4"><p>This applies especially to content which requires an account to be accessible, e.g. content behind paywalls or in Facebook groups.</p></fn>
<fn id="fn5"><p><ext-link ext-link-type="uri" xlink:href="https://archive.org/web/petabox.php">https://archive.org/web/petabox.php</ext-link>.</p></fn>
<fn id="fn6"><p><ext-link ext-link-type="uri" xlink:href="https://techcrunch.com/2018/07/02/facebook-rolls-out-more-api-restrictions-and-shutdowns/">https://techcrunch.com/2018/07/02/facebook-rolls-out-more-api-restrictions-and-shutdowns/</ext-link>.</p></fn>
<fn id="fn7"><p><ext-link ext-link-type="uri" xlink:href="https://socialmediaarchive.org/">https://socialmediaarchive.org/</ext-link>.</p></fn>
<fn id="fn8"><p><ext-link ext-link-type="uri" xlink:href="https://www.gnu.org/software/wget/">https://www.gnu.org/software/wget/</ext-link>.</p></fn>
<fn id="fn9"><p><ext-link ext-link-type="uri" xlink:href="https://github.com/internetarchive/heritrix3/wiki">https://github.com/internetarchive/heritrix3/wiki</ext-link>.</p></fn>
<fn id="fn10"><p><ext-link ext-link-type="uri" xlink:href="https://github.com/internetarchive/heritrix3/wiki">https://github.com/internetarchive/heritrix3/wiki</ext-link>.</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="r1"><mixed-citation>Aggarwal, K. (2019). An efficient focused web crawling approach. In M. N. Hoda (Ed.), <italic>Software Engineering: Proceedings of CSI 2015</italic> (Vol. 731, pp. 131&#x2013;138). Springer. <ext-link ext-link-type="doi" xlink:href="10.1007/978-981-10-8848-3_13">https://doi.org/10.1007/978-981-10-8848-3_13</ext-link></mixed-citation></ref>
<ref id="r2"><mixed-citation>Aher, G., Arriaga, R. I., &#x0026; Kalai, A. T. (2023). Using large language models to simulate multiple humans and replicate human subject studies. In <italic>Proceedings of the 40th International Conference on Machine Learning (ICML &#x2019;23)</italic> (pp. 337&#x2013;371).</mixed-citation></ref>
<ref id="r3"><mixed-citation>Al Galib, A., Mehedi, M. H. K., &#x0026; Rasel, A. A. (2024). Large scale web crawling and distributed search engines: techniques, challenges, current trends, and future prospects. In N. H. Zakaria, N. S. Mansor, H. Husni, &#x0026; F. Mohammed (Eds.), <italic>Computing and Informatics: 9th International Conference (ICOCI &#x2019;23)</italic>, Kuala Lumpur, Malaysia, September 13&#x2013;14, 2023, Revised Selected Papers (pp. 17&#x2013;29). Springer Nature. <ext-link ext-link-type="doi" xlink:href="10.1007/978-981-99-9589-9_2">https://doi.org/10.1007/978-981-99-9589-9_2</ext-link></mixed-citation></ref>
<ref id="r4"><mixed-citation>Altenh&#x00F6;ner, R., &#x0026; Schrimpf, M. (2014). Lost in tradition?: Systematische und technische Aspekte der Erwerbung von Internetpublikationen in Archivbibliotheken. In A. Sch&#x00FC;ller-Zwierlein &#x0026; M. Hollmann (Eds.), <italic>Diachrone Zug&#x00E4;nglichkeit als Prozess: Kulturelle &#x00DC;berlieferung in systematischer Sicht</italic> (pp. 297&#x2013;328). De Gruyter Saur. <ext-link ext-link-type="doi" xlink:href="10.1515/9783110311846.297">https://doi.org/10.1515/9783110311846.297</ext-link></mixed-citation></ref>
<ref id="r5"><mixed-citation>Arabzadeh, N., &#x0026; Clarke, C. L. (2025). Benchmarking LLM-based relevance judgment methods. In <italic>Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval</italic> (pp. 3194&#x2013;3204). <ext-link ext-link-type="doi" xlink:href="10.1145/3726302.3730305">https://doi.org/10.1145/3726302.3730305</ext-link></mixed-citation></ref>
<ref id="r6"><mixed-citation>Behre, J., H&#x00F6;lig, S., &#x0026; M&#x00F6;ller, J. (2024). <italic>Reuters Institute Digital News Report 2024: Ergebnisse f&#x00FC;r Deutschland</italic>. Verlag Hans-Bredow-Institut. <ext-link ext-link-type="doi" xlink:href="10.21241/ssoar.94461">https://doi.org/10.21241/ssoar.94461</ext-link></mixed-citation></ref>
<ref id="r7"><mixed-citation>Berners-Lee, T. (1999). <italic>Weaving the Web: The original design and ultimate destiny of the World Wide Web by its inventor</italic>. Harper.</mixed-citation></ref>
<ref id="r8"><mixed-citation>Best, S. (2018). <italic>None like us: Blackness, belonging, aesthetic life: Theory Q</italic>. Duke University Press. <ext-link ext-link-type="doi" xlink:href="10.1215/9781478002581">https://doi.org/10.1215/9781478002581</ext-link></mixed-citation></ref>
<ref id="r9"><mixed-citation>Bourdieu, P. (1985). The market of symbolic goods. <italic>Poetics</italic>, <italic>14</italic>(1&#x2013;2), 13&#x2013;44. <ext-link ext-link-type="doi" xlink:href="10.1016/0304-422X(85)90003-8">https://doi.org/10.1016/0304-422X(85)90003-8</ext-link></mixed-citation></ref>
<ref id="r10"><mixed-citation>Bourdieu, P. (2001). <italic>Distinction: A social critique of the judgement of taste</italic>. Routledge.</mixed-citation></ref>
<ref id="r11"><mixed-citation>Brunelle, J. F., Kelly, M., Weigle, M. C., &#x0026; Nelson, M. L. (2016). The impact of JavaScript on archivability. <italic>International Journal on Digital Libraries</italic>, <italic>17</italic>(2), 95&#x2013;117. <ext-link ext-link-type="doi" xlink:href="10.1007/s00799-015-0140-8">https://doi.org/10.1007/s00799-015-0140-8</ext-link></mixed-citation></ref>
<ref id="r12"><mixed-citation>Bruns, A. (2018). <italic>Gatewatching and news curation: Journalism, social media, and the public sphere</italic>. Peter Lang.</mixed-citation></ref>
<ref id="r14"><mixed-citation>Br&#x00FC;gger, N. (2018). <italic>The archived web: Doing history in the digital age</italic>. MIT Press.</mixed-citation></ref>
<ref id="r15"><mixed-citation>Butler, B. (2009). &#x2018;Othering&#x2019; the archive - From exile to inclusion and heritage dignity: The case of Palestinian archival memory. <italic>Archive and Museum Informatics</italic>, <italic>9</italic>(1), 57&#x2013;69.</mixed-citation></ref>
<ref id="r16"><mixed-citation>Chapekis, A., Bestvater, S., Remy, E., &#x0026; Rivero, G. (2024). <italic>When online content disappears</italic>. PEW Research Center Report. <ext-link ext-link-type="uri" xlink:href="https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/">https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/</ext-link></mixed-citation></ref>
<ref id="r17"><mixed-citation>Christin, A. (2020). <italic>Metrics at work: Journalism and the contested meaning of algorithms</italic>. Princeton University Press.</mixed-citation></ref>
<ref id="r18"><mixed-citation>Chen, C. (2006). CiteSpace II: Detecting and visualizing emerging trends and transient patterns in scientific literature. <italic>Journal of the American Society for Information Science and Technology</italic>, <italic>57</italic>(3), 359&#x2013;377. <ext-link ext-link-type="doi" xlink:href="10.1002/asi.20317">https://doi.org/10.1002/asi.20317</ext-link></mixed-citation></ref>
<ref id="r19"><mixed-citation>Davis, J. L., &#x0026; Jurgenson, N. (2014). Context collapse: Theorizing context collusions and collisions. Information, <italic>Communication &#x0026; Society</italic>, <italic>17</italic>(4), 476&#x2013;485. <ext-link ext-link-type="doi" xlink:href="10.1080/1369118X.2014.888458">https://doi.org/10.1080/1369118X.2014.888458</ext-link></mixed-citation></ref>
<ref id="r20"><mixed-citation>DeLuca, K. M., Lawson, S., &#x0026; Sun, Y. (2012). Occupy wall street on the public screens of social media: The many framings of the birth of a protest movement. <italic>Communication, Culture and Critique</italic>, <italic>5</italic>(4), 483&#x2013;509. <ext-link ext-link-type="doi" xlink:href="10.1111/j.1753-9137.2012.01141.x">https://doi.org/10.1111/j.1753-9137.2012.01141.x</ext-link></mixed-citation></ref>
<ref id="r21"><mixed-citation>Derrida, J., &#x0026; Prenowitz, E. (1995). Archive fever: A freudian impression. <italic>Diacritics</italic>, <italic>25</italic>(2), 9&#x2013;63. <ext-link ext-link-type="doi" xlink:href="10.2307/465144">https://doi.org/10.2307/465144</ext-link></mixed-citation></ref>
<ref id="r22"><mixed-citation>Deutsche Nationalbibliothek. (2016). <italic>Strategischer Kompass</italic>. <ext-link ext-link-type="uri" xlink:href="https://d-nb.info/1112299254/34">https://d-nb.info/1112299254/34</ext-link></mixed-citation></ref>
<ref id="r23"><mixed-citation>Deutsche Nationalbibliothek. (2017). <italic>Zum Sammelauftrag der Deutschen Nationalbibliothek</italic>. <ext-link ext-link-type="uri" xlink:href="https://www.dnb.de/SharedDocs/Downloads/DE/Ueber-uns/zumSammelauftragDNB.pdf?__blob=publicationFile&amp;v=4">https://www.dnb.de/SharedDocs/Downloads/DE/Ueber-uns/zumSammelauftragDNB.pdf?__blob&#x003D;publicationFile&#x0026;v&#x003D;4</ext-link></mixed-citation></ref>
<ref id="r24"><mixed-citation>Deutsche Nationalbibliothek. (2023). Sammelauftrag. <ext-link ext-link-type="uri" xlink:href="https://www.dnb.de/DE/Professionell/Sammeln/sammeln_node.html">https://www.dnb.de/DE/Professionell/Sammeln/sammeln_node.html</ext-link></mixed-citation></ref>
<ref id="r25"><mixed-citation>D&#x2019;Ignazio, C., &#x0026; Klein, L. (2020). <italic>Data feminism</italic>. MIT Press.</mixed-citation></ref>
<ref id="r26"><mixed-citation>Donig, S., Eckl, M., Gassner, S., &#x0026; Rehbein, M. (2023). Web archive analytics: Blind spots and silences in distant readings of the archived web. <italic>Digital Scholarship in the Humanities</italic>, <italic>38</italic>(3), 1033&#x2013;1048. <ext-link ext-link-type="doi" xlink:href="10.1093/llc/fqad014">https://doi.org/10.1093/llc/fqad014</ext-link></mixed-citation></ref>
<ref id="r27"><mixed-citation>Duffy, C., &#x0026; Flynn, K. (2021, September 10). Some of the most iconic 9/11 news coverage is lost. Blame Adobe Flash. <italic>CNN Business</italic>. <ext-link ext-link-type="uri" xlink:href="https://edition.cnn.com/2021/09/10/tech/digital-news-coverage-9-11/index.html">https://edition.cnn.com/2021/09/10/tech/digital-news-coverage-9-11/index.html</ext-link></mixed-citation></ref>
<ref id="r28"><mixed-citation>Elish, M. C., &#x0026; Boyd, D. (2017). <italic>Situating methods in the magic of big data and artificial intelligence</italic>. SSRN. <ext-link ext-link-type="uri" xlink:href="https://ssrn.com/abstract=3040201">https://ssrn.com/abstract&#x003D;3040201</ext-link></mixed-citation></ref>
<ref id="r29"><mixed-citation>European Digital Media Observatory (EDMO). (2022). <italic>Report of the European Digital Media Obervatory&#x2019;s Working Group on platform-to-researcher data access</italic>. (EDMO WP-Content). European Digital Media Observatory. <ext-link ext-link-type="uri" xlink:href="https://edmo.eu/wp-content/uploads/2022/02/Report-of-the-European-Digital-Media-Observatorys-Working-Group-on-Platform-to-Researcher-Data-Access-2022.pdf">https://edmo.eu/wp-content/uploads/2022/02/Report-of-the-European-Digital-Media-Observatorys-Working-Group-on-Platform-to-Researcher-Data-Access-2022.pdf</ext-link></mixed-citation></ref>
<ref id="r30"><mixed-citation>Foucault, M. (1970). <italic>The order of things: An archaeology of the human sciences</italic>. Pantheon Books.</mixed-citation></ref>
<ref id="r31"><mixed-citation>Galtung, J., &#x0026; Ruge, M. H. (1965). The structure of foreign news: The presentation of the Congo, Cuba und Cyprus crises in four foreign newspapers. <italic>Journal of Peace Research</italic>, <italic>2</italic>(1), 64&#x2013;90. <ext-link ext-link-type="doi" xlink:href="10.1177/002234336500200104">https://doi.org/10.1177/002234336500200104</ext-link></mixed-citation></ref>
<ref id="r32"><mixed-citation>Gomes, D., Demidova, E., Winters, J., &#x0026; Risse, T. (Eds.). (2021). <italic>The past web: Exploring web archives</italic>. Springer. <ext-link ext-link-type="doi" xlink:href="10.1007/978-3-030-63291-5">https://doi.org/10.1007/978-3-030-63291-5</ext-link></mixed-citation></ref>
<ref id="r33"><mixed-citation>Gyllstrom, K. A., Eickhoff, C., de Vries, A. P., &#x0026; Moens, M.-F. (2012). The downside of markup: Examining the harmful effects of CSS and javascript on indexing today&#x2019;s web. In <italic>The Proceedings of the 21st ACM International Conference on Information and Knowledge Management (CIKM &#x2019;12)</italic> (pp. 1990&#x2013;1994). The Association for Computing Machinery. <ext-link ext-link-type="doi" xlink:href="10.1145/2396761.2398558">https://doi.org/10.1145/2396761.2398558</ext-link></mixed-citation></ref>
<ref id="r34"><mixed-citation>Hanada, R., Cristo, M., &#x0026; Pimentel, M. D. G. C. (2013). How do metrics of link analysis correlate to quality, relevance and popularity in Wikipedia? In <italic>Proceedings of the 19th Brazilian Symposium on Multimedia and the web (WebMedia &#x2019;13)</italic> (pp. 105&#x2013;112). <ext-link ext-link-type="doi" xlink:href="10.1145/2526188.2526198">https://doi.org/10.1145/2526188.2526198</ext-link></mixed-citation></ref>
<ref id="r35"><mixed-citation>Harlow, S., &#x0026; Johnson, T. J. (2011). The Arab Spring&#x007C; Overthrowing the Protest Paradigm?: How The New York Times, Global Voices and Twitter Covered the Egyptian Revolution. <italic>International Journal of Communication</italic>, <italic>5</italic>, 1358&#x2013;1374.</mixed-citation></ref>
<ref id="r36"><mixed-citation>Hartman, S. (1997). <italic>Scenes of subjection: Terror, slavery, and self-making in Nineteenth-Century America</italic>. Oxford University Press.</mixed-citation></ref>
<ref id="r37"><mixed-citation>Helmond, A. (2019). A historiography of the hyperlink: Periodizing the Web through the changing role of the hyperlink. In <italic>The SAGE Handbook of Web History</italic> (pp. 227&#x2013;241). Sage Publications Ltd.</mixed-citation></ref>
<ref id="r38"><mixed-citation>Hendrickx, J., Ballon, P., &#x0026; Ranaivoson, H. (2022). Dissecting news diversity: An integrated conceptual framework. <italic>Journalism</italic>, <italic>23</italic>(8), 1751&#x2013;1769. <ext-link ext-link-type="doi" xlink:href="10.1177/1464884920966881">https://doi.org/10.1177/1464884920966881</ext-link></mixed-citation></ref>
<ref id="r39"><mixed-citation>Hindman, M. S. (2009). <italic>The myth of digital democracy</italic>. Princeton University Press.</mixed-citation></ref>
<ref id="r40"><mixed-citation>Holzmann, H., &#x0026; Nejdl, W. (2021). A holistic view on web archives. In D. Gomes, E. Demidova, J. Winters, &#x0026; T. Risse (Eds.), <italic>The past web: Exploring web archives</italic> (pp. 85&#x2013;99). Springer. <ext-link ext-link-type="doi" xlink:href="10.1007/978-3-030-63291-5">https://doi.org/10.1007/978-3-030-63291-5</ext-link></mixed-citation></ref>
<ref id="r41"><mixed-citation>Holzmann, H., Nejdl, W., &#x0026; Avisheh, A. (2016). The Dawn of today&#x2019;s popular domains: A study of the archived German Web over 18 years. In <italic>Proceedings of the 16th ACM/IEEE-CS on Joint Conference on Digital Libraries</italic>. <ext-link ext-link-type="doi" xlink:href="10.1145/2910896.2910901">https://doi.org/10.1145/2910896.2910901</ext-link></mixed-citation></ref>
<ref id="r42"><mixed-citation>Irani, L. (2015). Difference and dependence among digital workers: The case of amazon mechanical turk. <italic>South Atlantic Quarterly</italic>, <italic>114</italic>(1), 225&#x2013;234. <ext-link ext-link-type="doi" xlink:href="10.1215/00382876-2831665">https://doi.org/10.1215/00382876-2831665</ext-link></mixed-citation></ref>
<ref id="r43"><mixed-citation>Joris, G., De Grove, F., Van Damme, K., &#x0026; De Marez, L. (2020). News diversity reconsidered: A systematic literature review unravelling the diversity in conceptualizations. <italic>Journalism Studies</italic>, <italic>21</italic>(13), 1893&#x2013;1912. <ext-link ext-link-type="doi" xlink:href="10.1080/1461670X.2020.1797527">https://doi.org/10.1080/1461670X.2020.1797527</ext-link></mixed-citation></ref>
<ref id="r44"><mixed-citation>Kampes, C. F. (2021). <italic>Angebotsfragmentierung online: Empirische Analysen struktureller Differenzierung von Medienangeboten und Medienanbietern im Online-Medienmarkt</italic> [Doctoral dissertation, Universit&#x00E4;t D&#x00FC;sseldorf]. Universit&#x00E4;t D&#x00FC;sseldorf Docserv. <ext-link ext-link-type="uri" xlink:href="https://docserv.uni-duesseldorf.de/servlets/DocumentServlet?id=58308">https://docserv.uni-duesseldorf.de/servlets/DocumentServlet?id&#x003D;58308</ext-link></mixed-citation></ref>
<ref id="r45"><mixed-citation>Kelly, M., Brunelle, J. F., Weigle, M. C., &#x0026; Nelson, M. L. (2013). A method for identifying personalized representations in web archives. <italic>D-Lib Magazine</italic>, <italic>18</italic>(11&#x2013;12), 1&#x2013;11. <ext-link ext-link-type="doi" xlink:href="10.1045/november2013-kelly">https://doi.org/10.1045/november2013-kelly</ext-link></mixed-citation></ref>
<ref id="r46"><mixed-citation>Lai, H., Liu, X., Iong, I. L., Yao, S., Chen, Y., Shen, P., Yu, H., Zhang, H., Zhang, X., Dong, Y., &#x0026; Tang, J. (2024). AutoWebGLM: A large language model-based web navigating agent. In <italic>Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD &#x2019;24)</italic> (pp. 5295&#x2013;5306). <ext-link ext-link-type="doi" xlink:href="10.1145/3637528.3671620">https://doi.org/10.1145/3637528.3671620</ext-link></mixed-citation></ref>
<ref id="r47"><mixed-citation>Laursen, D., &#x0026; M&#x00F8;ldrup-Dalum, P. (2017). Looking back, looking forward: 10 years of web development to collect, preserve and access the Danish web. In N. Br&#x00FC;gger (Ed.), <italic>Web 25: Histories from the first 25 years of the World Wide Web</italic> (pp. 207&#x2013;228). Peter Lang.</mixed-citation></ref>
<ref id="r48"><mixed-citation>Leban, G., Fortuna, B., Brank, J., &#x0026; Grobelnik, M. (2014). Event registry: learning about world events from news. In <italic>Companion: Proceedings of the 23rd International Conference on World Wide Web (WWW &#x2019;14)</italic> (pp. 107&#x2013;110). <ext-link ext-link-type="doi" xlink:href="10.1145/2567948.2577024">https://doi.org/10.1145/2567948.2577024</ext-link></mixed-citation></ref>
<ref id="r49"><mixed-citation>Lee, H.-T., Leonard, D., Wang, X., &#x0026; Loguinov, D. (2009). IRLbot: Scaling to 6 billion pages and beyond. <italic>ACM Transactions on the Web</italic>, <italic>3</italic>(3), 1&#x2013;34. <ext-link ext-link-type="doi" xlink:href="10.1145/1541822.1541823">https://doi.org/10.1145/1541822.1541823</ext-link></mixed-citation></ref>
<ref id="r50"><mixed-citation>Library of Congress. (2017). <italic>Update on the Twitter Archive at the Library of Congress</italic>. [White paper]. Library of Congress - Blogs. <ext-link ext-link-type="uri" xlink:href="https://blogs.loc.gov/loc/files/2017/12/2017dec_twitter_white-paper.pdf">https://blogs.loc.gov/loc/files/2017/12/2017dec_twitter_white-paper.pdf</ext-link></mixed-citation></ref>
<ref id="r51"><mixed-citation>Loecherbach, F., Moeller, J., Trilling, D., &#x0026; van Atteveldt, W. (2020). The unified framework of media diversity: A systematic literature review. <italic>Digital Journalism</italic>, <italic>8</italic>(5), 605&#x2013;642. <ext-link ext-link-type="doi" xlink:href="10.1080/21670811.2020.1764374">https://doi.org/10.1080/21670811.2020.1764374</ext-link></mixed-citation></ref>
<ref id="r52"><mixed-citation>Maemura, E., Worby, N., &#x0026; Milligan, I. (2018). If these crawls could talk: Studying and documenting web archives provenance. <italic>Journal of the Association for Information Science and Technology</italic>, <italic>69</italic>(10), 1223&#x2013;1233. <ext-link ext-link-type="doi" xlink:href="10.1002/asi.24048">https://doi.org/10.1002/asi.24048</ext-link></mixed-citation></ref>
<ref id="r53"><mixed-citation>Miz, V., Hanna, J., Aspert, N., Ricaud, B., &#x0026; Vandergheynst, P. (2020). What is trending on Wikipedia? Capturing trends and language biases across Wikipedia editions. In <italic>Companion Proceedings of the Web Conference 2020 (WWW &#x2019;20)</italic> (pp. 794&#x2013;801). <ext-link ext-link-type="doi" xlink:href="10.1145/3366424.3383567">https://doi.org/10.1145/3366424.3383567</ext-link></mixed-citation></ref>
<ref id="r54"><mixed-citation>Napoli, P. M. (1999). Deconstructing the diversity principle. <italic>Journal of Communication</italic>, <italic>49</italic>(4), 7&#x2013;34. <ext-link ext-link-type="doi" xlink:href="10.1111/j.1460-2466.1999.tb02815.x">https://doi.org/10.1111/j.1460-2466.1999.tb02815.x</ext-link></mixed-citation></ref>
<ref id="r55"><mixed-citation>Newman, N., Fletcher, R., Eddy, K., Robertson, C. T., &#x0026; Nielsen, R. K. (2023). <italic>Reuters Institute Digital News Report 2023</italic>. Reuters Institute and University of Oxford. <ext-link ext-link-type="uri" xlink:href="https://reutersinstitute.politics.ox.ac.uk/sites/default/files/2023-06/Digital_News_Report_2023.pdf">https://reutersinstitute.politics.ox.ac.uk/sites/default/files/2023-06/Digital_News_Report_2023.pdf</ext-link></mixed-citation></ref>
<ref id="r56"><mixed-citation>Panagiotou, N., Katakis, I., &#x0026; Gunopulos, D. (2016). Detecting events in online social networks: Definitions, trends and challenges. In S. Michaelis, N. Piatkowski, M. Stolpe (Eds.), <italic>Lecture Notes in Computer Science: Vol. 9580. Solving large scale learning tasks. Challenges and algorithms</italic> (pp. 42&#x2013;84). Springer, Cham. <ext-link ext-link-type="doi" xlink:href="10.1007/978-3-319-41706-6_2">https://doi.org/10.1007/978-3-319-41706-6_2</ext-link></mixed-citation></ref>
<ref id="r57"><mixed-citation>Padoan, L., &#x0026; Vinciguerra, M. (2024). ScrapeGraphAI [Online post]. <ext-link ext-link-type="uri" xlink:href="https://github.com/ScrapeGraphAI/Scrapegraph-ai">https://github.com/ScrapeGraphAI/Scrapegraph-ai</ext-link></mixed-citation></ref>
<ref id="r58"><mixed-citation>Poell, T., Nieborg, D., &#x0026; van Dijck, J. (2019). Platformisation. <italic>Internet Policy Review</italic>, <italic>8</italic>(4), 1&#x2013;13. <ext-link ext-link-type="doi" xlink:href="10.14763/2019.4.1425">https://doi.org/10.14763/2019.4.1425</ext-link></mixed-citation></ref>
<ref id="r59"><mixed-citation>Reckwitz, A. (2020). <italic>The society of singularities</italic>. Polity.</mixed-citation></ref>
<ref id="r60"><mixed-citation>Ribeiro, F. N., Koustuv, S., Babaei, M., Henrique, L., Messias, J., Benevenuto, F., Goga, O., Gummadi, K. P., &#x0026; Redmiles, E. M. (2019). On microtargeting socially divisive Ads: A case study of Russia-Linked Ad campaigns on Facebook. In Association for Computing Machinery (Ed.), <italic>Proceedings of the 2019 Conference on Fairness, Accountability, and Transparency (FAT* &#x2019;19)</italic>, January 29&#x2013;31, 2019, Atlanta, GA, USA (pp. 140&#x2013;149). ACM. <ext-link ext-link-type="doi" xlink:href="10.1145/3287560.3287580">https://doi.org/10.1145/3287560.3287580</ext-link></mixed-citation></ref>
<ref id="r61"><mixed-citation>Schatz, H., &#x0026; Schulz, W. (1992). Qualit&#x00E4;t von Fernsehprogrammen: Kriterien und Methoden zur Beurteilung von Programmqualit&#x00E4;t im dualen Fernsehen. <italic>Media Perspektiven</italic>, <italic>11</italic>, 690&#x2013;712.</mixed-citation></ref>
<ref id="r62"><mixed-citation>Schmitt, T. M. (2011). <italic>Cultural governance: Zur Kulturgeographie des UNESCO-Welterberegimes</italic>. Franz Steiner Verlag.</mixed-citation></ref>
<ref id="r63"><mixed-citation>Schulze, G. (1992). <italic>Die Erlebnisgesellschaft: Kultursoziologie der Gegenwart</italic>. Campus Verlag.</mixed-citation></ref>
<ref id="r64"><mixed-citation>Thylstrup, N. B. (2018). <italic>The politics of mass digitization</italic>. MIT Press.</mixed-citation></ref>
<ref id="r65"><mixed-citation>Thylstrup, N. B., Agostinho, D., Ring, A., D&#x2019;Ignazio, C., &#x0026; Veel, K. (Eds.). (2021). <italic>Uncertain Archives: Critical Keywords for Big Data</italic>. MIT Press.</mixed-citation></ref>
<ref id="r66"><mixed-citation>Van Aelst, P., Str&#x00F6;mb&#x00E4;ck, J., Aalberg, T., Esser, F., de Vreese, C., Matthes, J., &#x0026; Stanyer, J. (2017). Political communication in a high-choice media environment: A challenge for democracy? <italic>Annals of the International Communication Association</italic>, <italic>41</italic>(1), 3&#x2013;27. <ext-link ext-link-type="doi" xlink:href="10.1080/23808985.2017.1288551">https://doi.org/10.1080/23808985.2017.1288551</ext-link></mixed-citation></ref>
<ref id="r67"><mixed-citation>Bo&#x010D;yt&#x0117;, R., &#x0026; de Vos, J. (2018). <italic>Server-side Preservation of Dynamic Websites</italic> [White paper]. Universiteit van Amsterdam. <ext-link ext-link-type="uri" xlink:href="https://publications.beeldengeluid.nl/pub/633/Juli2018_Server-Side-Preservation-of-Dynamic-Websites_R_Bocyte.pdf">https://publications.beeldengeluid.nl/pub/633/Juli2018_Server-Side-Preservation-of-Dynamic-Websites_R_Bocyte.pdf</ext-link></mixed-citation></ref>
<ref id="r68"><mixed-citation>Vlassenroot, E., Chambers, S., Di Pretoro, E., Geeraert, F., Haesendonck, G., Michel, A., &#x0026; Mechant, P. (2019). Web archives as a data resource for digital scholars. <italic>International Journal of Digital Humanities</italic>, <italic>1</italic>, 185&#x2013;111. <ext-link ext-link-type="doi" xlink:href="10.1007/s42803-019-00007-7">https://doi.org/10.1007/s42803-019-00007-7</ext-link></mixed-citation></ref>
<ref id="r69"><mixed-citation>Vlassenroot, E., Chambers, S., Lieber, S., &#x0026; Michel, A. (2021). Web-archiving and social media: An exploratory analysis. <italic>International Journal of Digital Humanities</italic>, <italic>2</italic>, 107&#x2013;128. <ext-link ext-link-type="doi" xlink:href="10.1007/s42803-021-00036-1">https://doi.org/10.1007/s42803-021-00036-1</ext-link></mixed-citation></ref>
<ref id="r70"><mixed-citation>Weller, K., &#x0026; Kinder-Kurlanda, K. E. (2016). A manifesto for data sharing in social media research. In W. Nejdl, W. Hall, P. Parigi, &#x0026; S. Staab (Eds.), <italic>Proceedings of the 2016 ACM Web Science Conference (WebSci &#x2019;16)</italic>, Hannover, Germany, May 22&#x2013;25, 2016 (pp. 166&#x2013;172). ACM. <ext-link ext-link-type="doi" xlink:href="10.1145/2908131.2908172">https://doi.org/10.1145/2908131.2908172</ext-link></mixed-citation></ref>
<ref id="r71"><mixed-citation>Ye, J., Wang, Y., Huang, Y., Chen, D., Zhang, Q., Moniz, N., Gao, T., Geyer, W., Huang, C., Chen, P., Chawla, N. V. &#x0026; Zhang, X. (2025, April). <italic>Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge</italic>. [Published as a conference paper at the International Conference on Learning Representations (ICLR) 2025]. <ext-link ext-link-type="uri" xlink:href="http://ICLR-2025-justice-or-prejudice-quantifying-biases-in-llm-as-a-judge-Paper-Conference.pdf">ICLR-2025-justice-or-prejudice-quantifying-biases-in-llm-as-a-judge-Paper-Conference.pdf</ext-link>. <ext-link ext-link-type="doi" xlink:href="10.48550/arXiv.2410.02736">https://doi.org/10.48550/arXiv.2410.02736</ext-link></mixed-citation></ref>
<ref id="r72"><mixed-citation>Zuboff, S. (2019). <italic>The age of surveillance capitalism: The fight for a human future at the new frontier of power</italic>. PublicAffairs.</mixed-citation></ref>
</ref-list>
</back>
</article>
