Relevance and Diversity – Towards Selection Criteria for Archiving the German Web

Authors

DOI:

https://doi.org/10.53377/lq.23531

Abstract

The motivation for this paper is the need of the German National Library to massively increase its web archiving activities as the 12,000 snapshots collected per year are considered not to be enough to reflect the German web appropriately. As collecting everything which could be defined as the German web is impossible, the concept of selecting an “exemplary diversity” is introduced and substantiated by the guiding principles of relevance and diversity. The long-term goal is to find methods to (semi-)automatise the selection process and the prerequisite for this are well-defined selection criteria and sources of relevant as well as diverse topics that can be exploited. First steps to operationalise the guiding principles are outlined and supplemented by technological challenges and requirements to support the selection process as well as in the light of the advancement of web technology. Finally, the paper proposes a mix of curational and automatised measures and sources of information that should be taken into consideration to cater for a good coverage of relevance and diversity, supplemented by technological aspects regarding alternative representations of web content in web archives to be discussed to bring the content of web archives closer to the original user experience.

Downloads

Download data is not yet available.

References

Aggarwal, K. (2019). An efficient focused web crawling approach. In M. N. Hoda (Ed.), Software Engineering: Proceedings of CSI 2015 (Vol. 731, pp. 131–138). Springer. https://doi.org/10.1007/978-981-10-8848-3_13

Aher, G., Arriaga, R. I., & Kalai, A. T. (2023). Using large language models to simulate multiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning (ICML ’23) (pp. 337–371).

Al Galib, A., Mehedi, M. H. K., & Rasel, A. A. (2024). Large scale web crawling and distributed search engines: techniques, challenges, current trends, and future prospects. In N. H. Zakaria, N. S. Mansor, H. Husni, & F. Mohammed (Eds.), Computing and Informatics: 9th International Conference (ICOCI ’23), Kuala Lumpur, Malaysia, September 13–14, 2023, Revised Selected Papers (pp. 17–29). Springer Nature. https://doi.org/10.1007/978-981-99-9589-9_2

Altenhöner, R., & Schrimpf, M. (2014). Lost in tradition?: Systematische und technische Aspekte der Erwerbung von Internetpublikationen in Archivbibliotheken. In A. Schüller-Zwierlein & M. Hollmann (Eds.), Diachrone Zugänglichkeit als Prozess: Kulturelle Überlieferung in systematischer Sicht (pp. 297–328). De Gruyter Saur. https://doi.org/10.1515/9783110311846.297

Arabzadeh, N., & Clarke, C. L. (2025). Benchmarking LLM-based relevance judgment methods. In Proceedings of the 48th International ACM SIGIR Conferenceon Research and Development in Information Retrieval (pp. 3194–3204). https://doi.org/10.1145/3726302.3730305

Behre, J., Hölig, S., & Möller, J. (2024). Reuters Institute Digital News Report 2024: Ergebnisse für Deutschland. Verlag Hans-Bredow-Institut. https://doi.org/10.21241/ssoar.94461

Berners-Lee, T. (1999). Weaving the Web: The original design and ultimate destiny of the World Wide Web by its inventor. Harper. Best, S. (2018). None like us: Blackness, belonging, aesthetic life: Theory Q. Duke University Press. https://doi.org/10.1215/9781478002581

Bourdieu, P. (1985). The market of symbolic goods. Poetics, 14(1–2), 13–44. https://doi.org/10.1016/0304-422X(85)90003-8

Bourdieu, P. (2001). Distinction: A social critique of the judgement of taste. Routledge. Brunelle, J. F., Kelly, M., Weigle, M. C., & Nelson, M. L. (2016). The impact of JavaScript on archivability. International Journal on Digital Libraries, 17(2), 95–117. https://doi.org/10.1007/s00799-015-0140-8

Bruns, A. (2018). Gatewatching and news curation: Journalism, social media, and the public sphere. Peter Lang.

Brügger, N. (2018). The archived web: Doing history in the digital age. MIT Press.

Butler, B. (2009). ‘Othering’ the archive - From exile to inclusion and heritage dignity: The case of Palestinian archival memory. Archive and Museum Informatics, 9(1), 57–69.

Chapekis, A., Bestvater, S., Remy, E., & Rivero, G. (2024). When online content disappears. PEW Research Center Report. https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/

Christin, A. (2020). Metrics at work: Journalism and the contested meaning of algorithms. Princeton University Press.

Chen, C. (2006). CiteSpace II: Detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of the American Society for Information Science and Technology, 57(3), 359–377. https://doi.org/10.1002/asi.20317

Davis, J. L., & Jurgenson, N. (2014). Context collapse: Theorizing context collusions and collisions. Information, Communication & Society, 17(4), 476–485. https://doi.org/10.1080/1369118X.2014.888458

DeLuca, K. M., Lawson, S., & Sun, Y. (2012). Occupy wall street on the public screens of social media: The many framings of the birth of a protest movement. Communication, Culture and Critique, 5(4), 483–509. https://doi.org/10.1111/j.1753-9137.2012.01141.x

Derrida, J., & Prenowitz, E. (1995). Archive fever: A freudian impression. Diacritics, 25(2), 9–63. https://doi.org/10.2307/465144

Deutsche Nationalbibliothek. (2016). Strategischer Kompass. https://d-nb.info/1112299254/34

Deutsche Nationalbibliothek. (2017). Zum Sammelauftrag der Deutschen Nationalbibliothek. https://www.dnb.de/SharedDocs/Downloads/DE/Ueber-uns/zumSammelauftragDNB.pdf?__blob=publicationFile&v=4

Deutsche Nationalbibliothek. (2023). Sammelauftrag. https://www.dnb.de/DE/Professionell/Sammeln/sammeln_node.html

D’Ignazio, C., & Klein, L. (2020). Data feminism. MIT Press

Donig, S., Eckl, M., Gassner, S., & Rehbein, M. (2023). Web archive analytics: Blind spots and silences in distant readings of the archived web. Digital Scholarship in the Humanities, 38(3), 1033–1048. https://doi.org/10.1093/llc/fqad014

Duffy, C., & Flynn, K. (2021, September 10). Some of the most iconic 9/11 news coverage is lost. Blame Adobe Flash. CNN Business. https://edition.cnn.com/2021/09/10/tech/digital-news-coverage-9-11/index.html

Elish, M. C., & Boyd, D. (2017). Situating methods in the magic of big data and artificial intelligence. SSRN. https://ssrn.com/abstract=3040201

European Digital Media Observatory (EDMO). (2022). Report of the European Digital Media Obervatory’s Working Group on platform-to-researcher data access. (EDMO WP-Content). European Digital Media Observatory. https://edmo.eu/wp-content/uploads/2022/02/Report-of-the-European-Digital-Media-Observatorys-Working-Group-on-Platform-to-Researcher-Data-Access-2022.pdf

Foucault, M. (1970). The order of things: An archaeology of the human sciences. Pantheon Books.

Galtung, J., & Ruge, M. H. (1965). The structure of foreign news: The presentation of the Congo, Cuba und Cyprus crises in four foreign newspapers. Journal of Peace Research, 2(1), 64–90. https://doi.org/10.1177/002234336500200104

Gomes, D., Demidova, E., Winters, J., & Risse, T. (Eds.). (2021). The past web: Exploring web archives. Springer. https://doi.org/10.1007/978-3-030-63291-5

Gyllstrom, K. A., Eickhoff, C., de Vries, A. P., & Moens, M.-F. (2012). The downside of markup: Examining the harmful effects of CSS and javascript on indexing today’s web. In The Proceedings of the 21st ACM International Conference on Information and Knowledge Management (CIKM ’12) (pp. 1990–1994). The Association for Computing Machinery. https://doi.org/10.1145/2396761.2398558

Hanada, R., Cristo, M., & Pimentel, M. D. G. C. (2013). How do metrics of link analysis correlate to quality, relevance and popularity in Wikipedia? In Proceedings of the 19th Brazilian Symposium on Multimedia and the web (WebMedia ’13) (pp. 105–112). https://doi.org/10.1145/2526188.2526198

Harlow, S., & Johnson, T. J. (2011). The Arab Spring| Overthrowing the Protest Paradigm?: How The New York Times, Global Voices and Twitter Covered the Egyptian Revolution. International Journal of Communication, 5, 1358–1374.

Hartman, S. (1997). Scenes of subjection: Terror, slavery, and self-making in Nineteenth-Century America. Oxford University Press.

Helmond, A. (2019). A historiography of the hyperlink: Periodizing the Web through the changing role of the hyperlink. In The SAGE Handbook of Web History (pp. 227–241). Sage Publications Ltd.

Hendrickx, J., Ballon, P., & Ranaivoson, H. (2022). Dissecting news diversity: An integrated conceptual framework. Journalism, 23(8), 1751–1769. https://doi.org/10.1177/1464884920966881

Hindman, M. S. (2009). The myth of digital democracy. Princeton University Press.

Holzmann, H., & Nejdl, W. (2021). A holistic view on web archives. In D. Gomes, E. Demidova, J. Winters, & T. Risse (Eds.), The past web: Exploring web archives (pp. 85–99). Springer. https://doi.org/10.1007/978-3-030-63291-5

Holzmann, H., Nejdl, W., & Avisheh, A. (2016). The Dawn of today’s popular domains: A study of the archived German Web over 18 years. In Proceedings of the 16th ACM/IEEE-CS on Joint Conference on Digital Libraries. https://doi.org/10.1145/2910896.2910901

Irani, L. (2015). Difference and dependence among digital workers: The case of amazon mechanical turk. South Atlantic Quarterly, 114(1), 225–234. https://doi.org/10.1215/00382876-2831665

Joris, G., De Grove, F., Van Damme, K., & De Marez, L. (2020). News diversity reconsidered: A systematic literature review unravelling the diversity in conceptualizations. Journalism Studies, 21(13), 1893–1912. https://doi.org/10.1080/1461670X.2020.1797527

Kampes, C. F. (2021). Angebotsfragmentierung online: Empirische Analysen struktureller Differenzierung von Medienangeboten und Medienanbietern im Online-Medienmarkt [Doctoral dissertation, Universität Düsseldorf]. Universität Düsseldorf Docserv. https://docserv.uni-duesseldorf.de/servlets/DocumentServlet?id=58308

Kelly, M., Brunelle, J. F., Weigle, M. C., & Nelson, M. L. (2013). A method for identifying personalized representations in web archives. D-Lib Magazine, 18(11–12), 1–11. https://doi.org/10.1045/november2013-kelly

Lai, H., Liu, X., Iong, I. L., Yao, S., Chen, Y., Shen, P., Yu, H., Zhang, H., Zhang, X., Dong, Y., & Tang, J. (2024). AutoWebGLM: A large language model-based web navigating agent. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’24) (pp. 5295–5306). https://doi.org/10.1145/3637528.3671620

Laursen, D., & Møldrup-Dalum, P. (2017). Looking back, looking forward: 10 years of web development to collect, preserve and access the Danish web. In N. Brügger (Ed.), Web 25: Histories from the first 25 years of the World Wide Web (pp. 207–228). Peter Lang.

Leban, G., Fortuna, B., Brank, J., & Grobelnik, M. (2014). Event registry: learning about world events from news. In Companion: Proceedings of the 23rd International Conference on World Wide Web (WWW ’14) (pp. 107–110). https://doi.org/10.1145/2567948.2577024

Lee, H.-T., Leonard, D., Wang, X., & Loguinov, D. (2009). IRLbot: Scaling to 6 billion pages and beyond. ACM Transactions on the Web, 3(3), 1–34. https://doi.org/10.1145/1541822.1541823

Library of Congress. (2017). Update on the Twitter Archive at the Library of Congress. [White paper]. Library of Congress - Blogs. https://blogs.loc.gov/loc/files/2017/12/2017dec_twitter_white-paper.pdf

Loecherbach, F., Moeller, J., Trilling, D., & van Atteveldt, W. (2020). The unified framework of media diversity: A systematic literature review. Digital Journalism, 8(5), 605–642. https://doi.org/10.1080/21670811.2020.1764374

Maemura, E., Worby, N., & Milligan, I. (2018). If these crawls could talk: Studying and documenting web archives provenance. Journal of the Association for Information Science and Technology, 69(10), 1223–1233. https://doi.org/10.1002/asi.24048

Miz, V., Hanna, J., Aspert, N., Ricaud, B., & Vandergheynst, P. (2020). What is trending on Wikipedia? Capturing trends and language biases across Wikipedia editions. In Companion Proceedings of the Web Conference 2020 (WWW ’20) (pp. 794–801). https://doi.org/10.1145/3366424.3383567

Napoli, P. M. (1999). Deconstructing the diversity principle. Journal of Communication, 49(4), 7–34. https://doi.org/10.1111/j.1460-2466.1999.tb02815.x

Newman, N., Fletcher, R., Eddy, K., Robertson, C. T., & Nielsen, R. K. (2023). Reuters Institute Digital News Report 2023. Reuters Institute and University of Oxford. https://reutersinstitute.politics.ox.ac.uk/sites/default/files/2023-06/Digital_News_Report_2023.pdf

Panagiotou, N., Katakis, I., & Gunopulos, D. (2016). Detecting events in online social networks: Definitions, trends and challenges. In S. Michaelis, N. Piatkowski, M. Stolpe (Eds.), Lecture Notes in Computer Science: Vol. 9580. Solving large scale learning tasks. Challenges and algorithms (pp. 42–84). Springer, Cham. https://doi.org/10.1007/978-3-319-41706-6_2

Padoan, L., & Vinciguerra, M. (2024). ScrapeGraphAI [Online post]. https://github.com/ScrapeGraphAI/Scrapegraph-ai

Poell, T., Nieborg, D., & van Dijck, J. (2019). Platformisation. Internet Policy Review, 8(4), 1–13. https://doi.org/10.14763/2019.4.1425

Reckwitz, A. (2020). The society of singularities. Polity.

Ribeiro, F. N., Koustuv, S., Babaei, M., Henrique, L., Messias, J., Benevenuto, F., Goga, O., Gummadi, K. P., & Redmiles, E. M. (2019). On microtargeting socially divisive Ads: A case study of Russia-Linked Ad campaigns on Facebook. In Association for Computing Machinery (Ed.), Proceedings of the 2019 Conference on Fairness, Accountability, and Transparency (FAT* ’19), January 29–31, 2019, Atlanta, GA, USA (pp.140–149). ACM. https://doi.org/10.1145/3287560.3287580

Schatz, H., & Schulz, W. (1992). Qualität von Fernsehprogrammen: Kriterien und Methoden zur Beurteilung von Programmqualität im dualen Fernsehen. Media Perspektiven, 11, 690–712.

Schmitt, T. M. (2011). Cultural governance: Zur Kulturgeographie des UNESCOWelterberegimes. Franz Steiner Verlag.

Schulze, G. (1992). Die Erlebnisgesellschaft: Kultursoziologie der Gegenwart. Campus Verlag.

Thylstrup, N. B. (2018). The politics of mass digitization. MIT Press.

Thylstrup, N. B., Agostinho, D., Ring, A., D’Ignazio, C., & Veel, K. (Eds.). (2021). Uncertain Archives: Critical Keywords for Big Data. MIT Press.

Van Aelst, P., Strömbäck, J., Aalberg, T., Esser, F., de Vreese, C., Matthes, J., & Stanyer, J. (2017). Political communication in a high-choice media environment: A challenge for democracy? Annals of the International Communication Association, 41(1), 3–27. https://doi.org/10.1080/23808985.2017.1288551

Bočytė, R., & de Vos, J. (2018). Server-side Preservation of Dynamic Websites [White paper]. Universiteit van Amsterdam. https://publications.beeldengeluid.nl/pub/633/Juli2018_Server-Side-Preservation-of-Dynamic-Websites_R_Bocyte.pdf

Vlassenroot, E., Chambers, S., Di Pretoro, E., Geeraert, F., Haesendonck, G., Michel, A., & Mechant, P. (2019). Web archives as a data resource for digital scholars. International Journal of Digital Humanities, 1, 185–111. https://doi.org/10.1007/s42803-019-00007-7

Vlassenroot, E., Chambers, S., Lieber, S., & Michel, A. (2021). Web-archiving and social media: An exploratory analysis. International Journal of Digital Humanities, 2, 107–128. https://doi.org/10.1007/s42803-021-00036-1

Weller, K., & Kinder-Kurlanda, K. E. (2016). A manifesto for data sharing in social media research. In W. Nejdl, W. Hall, P. Parigi, & S. Staab (Eds.), Proceedings of the 2016 ACM Web Science Conference (WebSci ’16), Hannover, Germany, May 22–25, 2016 (pp. 166–172). ACM. https://doi.org/10.1145/2908131.2908172

Ye, J., Wang, Y., Huang, Y., Chen, D., Zhang, Q., Moniz, N., Gao, T., Geyer, W., Huang, C., Chen, P., Chawla, N. V. & Zhang, X. (2025, April). Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge. [Published as a conference paper at the International Conference on Learning Representations (ICLR) 2025]. ICLR-2025-justice-orprejudice-quantifying-biases-in-llm-as-a-judge-Paper-Conference.pdf. https://doi.org/10.48550/arXiv.2410.02736

Zuboff, S. (2019). The age of surveillance capitalism: The fight for a human future at the new frontier of power. PublicAffairs.

Downloads

Published

2026-07-17

Issue

Section

Case studies

How to Cite

Woldering, B., Gottschalk, S., Heinrichs, R., Nejdl, W., Neuberger, C., & Weller, K. (2026). Relevance and Diversity – Towards Selection Criteria for Archiving the German Web. LIBER Quarterly: The Journal of the Association of European Research Libraries, 36(1), 1-30. https://doi.org/10.53377/lq.23531