Revolutionizing Web Preservation: The Innovative Design of WebKurator.de

The digital landscape is vast and continuously evolving, making the documentation and preservation of its contents a daunting challenge. In a groundbreaking research paper, the authors introduce WebKurator.de, a state-of-the-art platform that tackles this issue by merging regional and topical web curation. Spearheaded by a team from German institutions, including the University of Passau and the Deutsche Nationalbibliothek, this platform aims to enhance how we archive and access culturally significant online resources.

The Challenge of Web Curation

As the primary source of documenting cultural, economic, and social dynamics, the web acts as a digital repository of our collective history. National libraries, striving to preserve this information, face significant hurdles due to the existing models of web directories which predominantly focus on a single dimension—typically either topical or geographical. Current systems like Curlie fall short, leading to incomplete curation that complicates searches for combined criteria, such as local businesses based on specific topics.

Introducing Two-Dimensional Curation

The beauty of the WebKurator.de platform lies in its two-dimensional curation model, which distinctly separates topics from geographic information. This innovative approach allows users to conduct precise searches, making it easier to locate relevant sites based both on subject matter and location. For example, one can search specifically for "delicatessen shops in Brunswick, Lower Saxony," and get relevant results integrated seamlessly.

Utilizing Automation and User Collaboration

One of the defining features of WebKurator.de is its integration of large language models (LLMs) to automate tedious tasks involved in web curation, such as categorizing topics and extracting addresses. This automation is complemented by a collaborative user-driven model, where users can suggest new websites and contribute to the curation process. However, it's essential to note that while LLMs filter the information efficiently, human moderators still oversee the process to ensure accuracy and quality, thereby creating a robust system that blends technology with human expertise.

Significant Statistics from the German Imprints Dataset

At the core of WebKurator.de is the German Imprints Dataset—a comprehensive collection amassing 5.54 million websites, with 3.14 million featuring imprint pages enriched with geo-coordinates and topic labels. Impressively, 85.17% of these websites are situated in Germany. This rich dataset serves as the foundation for the platform's curation activities and offers an extensive range of entries for users and researchers to explore.

A New Era of Web Curation

With the launch of WebKurator.de, web curation takes a giant leap forward. The platform not only preserves the rich tapestry of web content but also enhances accessibility and usability through its architectural features, structured categorization, and collaborative functionalities. As it expands its capabilities to encompass other German-speaking regions, the platform promises to be a pivotal resource in the digital archiving of web content. This initiative stands as a testament to the potential that modern technology holds in reshaping how we interact with our digital heritage.

As we move further into the digitized future, platforms like WebKurator.de will become integral to preserving the essence of our online culture for generations to come.

Authors: Michael Dinzinger, Natanael Arndt, Ben Böck, Jelena Mitrović, Michael Granitzer