DH20206 presentation: A never-ending story: supporting access to museum and library collections (in the UK)

I'm posting my DH2026 talk abstract here, as they can be slightly hard to find in the huge official book of abstracts. You can access and cite the Book of Abstracts at https://doi.org/10.5281/zenodo.21495909

Mia Ridge, British Library, United Kingdom; Museum Data Service, United Kingdom

Arran Rees, Museum Data Service, United Kingdom

Ross Parry, University of Leicester, United Kingdom

In 2016, the Culture White Paper set out the UK government's 'ambition and strategy for the cultural sectors'. It included a vision of users enjoying 'a seamless experience online', with the ability 'to access particular collections in depth as well as search across all collections (Department for Digital, Culture, Media & Sport 2016)'.

In 2026, users have many options for accessing UK cultural heritage collections online, but the experience is far from seamless. Infrastructure for digital cultural heritage is a patchwork of open access images on non-profit platforms, paywalls and commercial databases, research datasets and various forms of 'free to view' collections – some collections can be accessed in depth, others only as metadata or catalogue records, while others have no online presence (Gosling et al. 2022). Searching across all collections is impossible.

Focusing on access to UK museum and library collections, this paper discusses two significant components of the UK's digital cultural heritage access for researchers.

The first is an aggregator of museum collections data, while the second is a brief description of the many options for accessing digitised items from the British Library. It closes with a call to support the long-term preservation and sustainability of cultural heritage collections online.

Aggregating museum collection data: introducing the Museum Data Service

The museum sector in the UK, as in many regions, is highly diverse. Large, national institutions can publish collection records with images, rich metadata and more via specialist websites and APIs, while tiny local or specialist museums might not be able to easily access their own collections records. While some of the estimated 80 million catalogue records held in c. 1700 accredited museums are available online, they were difficult to find and lacked persistent identifiers (Gosling et al. 2022). Finding museum records required using hundreds of websites with dozens of different interfaces, and collating museum collections for research at scale was almost impossible.

To address this challenge, the Museum Data Service (MDS) was launched in September 2024 to connect and share all the object records across all UK museums, large and small.  The MDS offers UK museums, and institutions with museum-like collections, the ability to share their object records with other museums, researchers and the public, on their own terms.

The MDS platform responds to the sector's difficulties using earlier aggregation platforms (such as Culture Grid and Europeana), particularly the difficulties museums had in exporting collections data and formatting it for aggregation. Sharing records with Europeana requires organisations to prepare their data by mapping their records to the Europeana Data Model (Europeana PRO, n.d.). In contrast, the MDS works with each contributing museum's own levels of technical capacity – it acknowledges the reality of the inconsistent approaches to data standards across the sector and accepts data in any format. The MDS team maps incoming fields to align them with the ‘Spectrum Units of Information’ names, without reformatting or standardising the data.  Museums can choose an appropriate Creative Commons licence for their dataset, reflecting their own access policies and attitude to risk management; this is also designed to encourage wider uptake by museums.

The real benefit of the MDS for Digital Humanities researchers is the meaningful connections it enables between museum collections and researchers (Parry et al. 2025). The MDS lets anyone search for types of collections, or specific types of objects within hundreds of museum collections. A user-friendly search interface enables detailed searches and sets of search results can be exported to support ongoing research by individuals. A specialist API supports the creation of bespoke datasets to support data analysis and modelling in heritage studies, for example to study questions of bias in cataloguing.

The fact that MDS does not ask collections to use a standard schema reduces the barriers to participation for museums. This also means that researchers can have access to the unique raw records that make up museum data, rather than a 'cleaned' version (Rawson and Muñoz 2016). The MDS does not make assumptions about the types of data that researchers might want, and it does not attempt to meet every need in the heritage sector. Rather, it aims to be an intrinsic part of a wider ecosystem for museum and heritage information, including participation in academic research projects (Bailey-Ross et al. 2024; Bailey et al. 2024). However, the MDS acknowledges that some researchers would prefer a slightly more standardised version of the data and is exploring options for this.

We hope to inspire future research using the museum data available on the platform, and to understand how the MDS might be developed to support emerging computational methods and Digital Humanities research questions.

The paradox of access enabled by commercial digitisation

While many nations have funded the digitisation of their national collections, organisations in the UK have been required by successive governments to seek public-private partnerships to digitise their collections. Under these deals, commercial companies who pay for collections digitisation can charge for access to those collections for a certain number of years. While research or direct government funding has paid for a certain amount of digitisation in the past, it is dwarfed by commercial funding. As a blog post by a British Library curator puts it:

'It has long been the goal of the British Library to make some of its digitised newspapers freely available online, but we also want to see the BNA [British Newspaper Archive] succeed as it has been doing, without which we could not have reached such a huge collection overall of digitised newspapers, nor the rate at which they are being produced (currently around half a million pages are being added to the BNA every month).' (McKernan 2021)

While most material digitised by commercial companies is eventually made more freely available, paywalls are one important factor in the fragmentation of access to collections. This was extensively documented in a Digital Collections Audit commissioned for the AHRC-funded research programme Towards a National Collection (TaNC) (Gosling et al. 2022).

The variety of sources of access to digitised collections from the British Library as in September 2023 helps illustrate this point.  The British Library's websites and catalogues hosted hundreds of thousands digitised books, manuscripts, 3D objects and sound files. Open access images and metadata were also available on Flickr Commons, Europeana and Wikimedia Commons. Newspapers were available on the British Newspaper Archive. Other commercial databases held selected other items, while other items were available via union catalogues. The Library's Research Repository had more open access items, from digitised images to datasets created via research projects like Living with Machines.  A large backlog of items was being processed for publication online. Altogether, this made answering an apparently simple enquiry about the availability of a particular item or collection online surprisingly complex.

However, this patchwork of platforms had one positive effect. While library users lost access to online catalogues, digitised and born-digital collections and the many services hosted by the British Library in a ransomware attack in October 2023, collections hosted on commercial and non-profit third-party platforms (including the Research Repository) continued to be available. The saying 'Lots of Copies Keep Stuff Safe' has never been more pertinent.

At the time of writing, the UK government is again investing in digital infrastructure, including the TaNC project and its successor, N-RICH. N-RICH has commissioned a range of activities to  'examine the scope, costs, risks, impacts and benefits of a future digital research infrastructure for cultural heritage in the UK' (Towards a National Collection 2025).

Digital records are fragile: the need to fund long-term access

Enabling access to collections online is one thing. Sustaining that access in the longer-term is another. The ransomware attack that took out most of the British Library's services was just the most dramatic incident that reduced online access to collections. Websites and platforms gradually disappearing when funding ends has eroded earlier digitisation efforts (Dunning 2009). For example, the results of the first big digitisation effort in the UK, the New Opportunities Fund c. 2000 – 2004, have long since been lost. Culture Grid, the UK's aggregator for Europeana, itself built on an earlier no-longer-supported project (the Peoples Network Discover Service), failed to find a sustainable financial model in the 2010s (Gosling 2019; Poole 2015).

More recently, collections sites have struggled to stay online through surges of bots scraping their sites to feed AI models. For example, the British Library's Research Repository was being scraped so heavily that it effectively experienced a Denial of Service attack until the team were able to put preventative measures in place.

Resourcing digital preservation and sustainability is not easy. In order to make the MDS a relatively lean operation, an early decision was made to exclude images and other media from the data aggregated. In addition to reducing financial costs, this reduces the environmental overhead of the service. However, not only does this add additional retrieval steps for researchers wanting to access images, it also means that the MDS cannot act as a backup of last resort for museums who lose access to their image assets.

Looking back over the long history of digitisation in museums, libraries and archives in the UK, it is clear that new economic models are needed to ensure the long-term availability of digital cultural heritage collections. Coming up with these models will require creativity and persuasive powers worthy of the collections they would protect.

References

  • Bailey, Rebecca, Javier Pereda, Chris Michaels, and Tom Callahan. 2024. Unlocking the Potential of Digital Collections. A Call to Action. Arts and Humanities Research Council. https://doi.org/10.5281/zenodo.13838916.
  • Bailey-Ross, Claire, Emily Burgess, and Panagiotis Papageorgiou. 2024. User Research: UK Gallery, Library, Archive and Museum  (GLAM) Digital Collections Infrastructure. Towards a National Collection. https://doi.org/10.5281/zenodo.12751226.
  • British Library. 2024. Learning Lessons from the Cyber-Attack: British Library Cyber Incident Review. British Library. https://www.bl.uk/home/british-library-cyber-incident-review-8-march-2024.pdf/.
  • Department for Digital, Culture, Media & Sport. 2016. The Culture White Paper. Department for Digital, Culture, Media & Sport. https://www.gov.uk/government/publications/culture-white-paper.
  • Dunning, Alastair. 2009. ‘Digitising the Past: Next Steps for Public-Sector Digitisation’. Digital Information-Order Or Anarchy? http://eprints.rclis.org/handle/10760/18048.
  • Europeana PRO. n.d. ‘Europeana Data Model’. Accessed 7 May 2026. https://pro.europeana.eu/page/edm-documentation.
  • Gosling, Kevin. 2019. ‘Nothing New except What Has Been Forgotten’. Collections Trust, November 27. https://collectionstrust.org.uk/blog/nothing-new-except-what-has-been-forgotten/.
  • Gosling, Kevin, Gordon McKenna, and Adrian Cooper. 2022. Digital Collections Audit. Towards a National Collection. Collections Trust. https://doi.org/10.5281/zenodo.6379581.
  • McKernan, Luke. 2021. ‘Free to View Online Newspapers’. The Newsroom Blog, August 9. https://blogs-archive.bl.uk/thenewsroom/2021/08/free-to-view-online-newspapers.html.
  • Parry, Ross, Stef De Sabbata, Andrew Ellis, Helen Hardy, and Mia Ridge. 2025. ‘A Framework for Sustainable and Scalable Cultural Data Integration and Analysis: Using the Large Dataset of the Museum Data Service’. 2025 11th International Symposium on System Security, Safety, and Reliability (ISSSR), April, 417–24. https://doi.org/10.1109/ISSSR65654.2025.00062.
  • Poole, Nick. 2015. ‘Guest Blog: Aggregation & the Culture Grid’. Museums Computer Group, May 31. https://museumscomputergroup.org.uk/culture-grid/.
  • Rawson, Katie, and Trevor Muñoz. 2016. ‘Against Cleaning’. Curating Menus, July 6. http://www.curatingmenus.org/articles/against-cleaning/.
  • ‘Spectrum’. n.d. Collections Trust. Accessed 7 May 2026. https://collectionstrust.org.uk/spectrum/.
  • Towards a National Collection. 2025. ‘N-RICH Prototype’. July. https://www.nationalcollection.org.uk/n-rich-prototype.

New data paper and datasets from crowdsourcing on Living with Machines

After lots of hard work by me, Nilo Pedrazzini, Miguel V., Arianna Ciula and Barbara McGillivray, we have a data paper in the Journal of Open Humanities Data: Language of Mechanisation Crowdsourcing Datasets from the Living with Machines Project.

And huge thanks to the thousands of Zooniverse volunteers who annotated 19th century newspaper articles to create the datasets we've published alongside the data paper!

Abstract: We present the ‘Language of Mechanisation’ datasets with examples of re-use in visualisations and analysis. These reusable CSV files, published on the British Library’s Research Repository, contain automatically-transcribed text from 19th century British newspaper articles. Volunteers on the Zooniverse crowdsourcing platform took part in tasks that asked ‘How did the word x change over time and place?’ They annotated articles with pre-selected meanings (senses) for the words coach, car, trolley and bike.

The datasets can support scholarship on a range of historical and linguistic research areas, including research on crowdsourcing and online volunteering behaviours, data processing and data visualisations methodologies.

The two datasets described are at:

Keynote video 'Evolutionary Innovations: Collections as Data in the AI era' for Making Meaning 2024

Making Meaning 2024: Mia Ridge Keynote

My slides for #SLQMakingMeaning #CollectionsAsData, 'Evolutionary Innovations: Collections as Data in the AI era', are online at https://zenodo.org/records/10795641

‘Collections as data’ describes the movement to publish open data from museum, library and archive collections that began in the noughties. The benefits of machine learning for better discoverability and research with digitised/born digital collections are alluring. And the popularity of generative AI – and an increased awareness of the biases it reinscribes – has focused attention on responsible computational access to collections – but what does this mean in practical terms? Mia will share examples from the British Library and the Living with Machines data science project.

'Enriching lives: connecting communities and culture with the help of machines': my EuropeanaTech 2023 keynote

Panorama lit by natural light of a seaside town
The video for my opening keynote on 'Enriching lives: connecting communities and culture with the help of machines' for the EuropeanaTech 2023 conference is now online.

The EuropeanaTech 2023 conference was held in The Hague, the Netherlands and online from 10 – 12 October 2023. My slides are online.

My abstract: I’ll begin with an overview of current developments in AI and machine learning, then present work with crowdsourcing from the Living with Machines project to think about what AI means for online volunteers and communities around digital cultural heritage. I’ll share new thinking on ‘volunteer enrichment’ – participation in crowdsourcing that not only enriches and enhances collections records, but also enriches the lives of volunteers. How can we embed GLAM values when we apply AI and machine learning tools in our work?

In preparing my keynote I revisited my keynote for EuropeanaTech 2011, and reflected on work on crowdsourcing, data science and AI at the British Library, the Collective Wisdom project and Living with Machines since then.

2022: an overview(ish)

A work-in-progress post about what I got up to last year.

The biggest thing I did in 2022 was co-curate an exhibition at Leeds City Museum for the British Library and Living with Machines project.

My work on crowdsourcing for Living with Machines was a 'Research Highlight of the year' for the Alan Turing Institute.

November: I was invited to the Archives nationales de France conference 'Crowdsourcing et patrimoine culturel écrit', where I spoke on Crowdsourcing as connection: a constant star over a sea of change / Établir des connexions : un invariant des projets de crowdsourcing par Mia Ridge, British Library, Royaume-Uni

Also in November, I took part in a panel on 'International Infrastructures for the Digital Humanities' – video below – for the Building Infrastructures event. The panel was chaired by Paul Arthur and the other panellists were Toma Tasovac, Alexandra Pretrulevich, Langa Khumalo, Juan Steyn and Ruth Ahnert.

In December I gave an online keynote on 'Citizen Science as Public History?' for the conference 'When publics co-produce history in museums: skills, methodologies and impact of participation' at The Luxembourg Centre for Contemporary and Digital History (C²DH), University of Luxembourg.