DH20206 presentation: A never-ending story: supporting access to museum and library collections (in the UK)

I'm posting my DH2026 talk abstract here, as they can be slightly hard to find in the huge official book of abstracts. You can access and cite the Book of Abstracts at https://doi.org/10.5281/zenodo.21495909

Mia Ridge, British Library, United Kingdom; Museum Data Service, United Kingdom

Arran Rees, Museum Data Service, United Kingdom

Ross Parry, University of Leicester, United Kingdom

In 2016, the Culture White Paper set out the UK government's 'ambition and strategy for the cultural sectors'. It included a vision of users enjoying 'a seamless experience online', with the ability 'to access particular collections in depth as well as search across all collections (Department for Digital, Culture, Media & Sport 2016)'.

In 2026, users have many options for accessing UK cultural heritage collections online, but the experience is far from seamless. Infrastructure for digital cultural heritage is a patchwork of open access images on non-profit platforms, paywalls and commercial databases, research datasets and various forms of 'free to view' collections – some collections can be accessed in depth, others only as metadata or catalogue records, while others have no online presence (Gosling et al. 2022). Searching across all collections is impossible.

Focusing on access to UK museum and library collections, this paper discusses two significant components of the UK's digital cultural heritage access for researchers.

The first is an aggregator of museum collections data, while the second is a brief description of the many options for accessing digitised items from the British Library. It closes with a call to support the long-term preservation and sustainability of cultural heritage collections online.

Aggregating museum collection data: introducing the Museum Data Service

The museum sector in the UK, as in many regions, is highly diverse. Large, national institutions can publish collection records with images, rich metadata and more via specialist websites and APIs, while tiny local or specialist museums might not be able to easily access their own collections records. While some of the estimated 80 million catalogue records held in c. 1700 accredited museums are available online, they were difficult to find and lacked persistent identifiers (Gosling et al. 2022). Finding museum records required using hundreds of websites with dozens of different interfaces, and collating museum collections for research at scale was almost impossible.

To address this challenge, the Museum Data Service (MDS) was launched in September 2024 to connect and share all the object records across all UK museums, large and small.  The MDS offers UK museums, and institutions with museum-like collections, the ability to share their object records with other museums, researchers and the public, on their own terms.

The MDS platform responds to the sector's difficulties using earlier aggregation platforms (such as Culture Grid and Europeana), particularly the difficulties museums had in exporting collections data and formatting it for aggregation. Sharing records with Europeana requires organisations to prepare their data by mapping their records to the Europeana Data Model (Europeana PRO, n.d.). In contrast, the MDS works with each contributing museum's own levels of technical capacity – it acknowledges the reality of the inconsistent approaches to data standards across the sector and accepts data in any format. The MDS team maps incoming fields to align them with the ‘Spectrum Units of Information’ names, without reformatting or standardising the data.  Museums can choose an appropriate Creative Commons licence for their dataset, reflecting their own access policies and attitude to risk management; this is also designed to encourage wider uptake by museums.

The real benefit of the MDS for Digital Humanities researchers is the meaningful connections it enables between museum collections and researchers (Parry et al. 2025). The MDS lets anyone search for types of collections, or specific types of objects within hundreds of museum collections. A user-friendly search interface enables detailed searches and sets of search results can be exported to support ongoing research by individuals. A specialist API supports the creation of bespoke datasets to support data analysis and modelling in heritage studies, for example to study questions of bias in cataloguing.

The fact that MDS does not ask collections to use a standard schema reduces the barriers to participation for museums. This also means that researchers can have access to the unique raw records that make up museum data, rather than a 'cleaned' version (Rawson and Muñoz 2016). The MDS does not make assumptions about the types of data that researchers might want, and it does not attempt to meet every need in the heritage sector. Rather, it aims to be an intrinsic part of a wider ecosystem for museum and heritage information, including participation in academic research projects (Bailey-Ross et al. 2024; Bailey et al. 2024). However, the MDS acknowledges that some researchers would prefer a slightly more standardised version of the data and is exploring options for this.

We hope to inspire future research using the museum data available on the platform, and to understand how the MDS might be developed to support emerging computational methods and Digital Humanities research questions.

The paradox of access enabled by commercial digitisation

While many nations have funded the digitisation of their national collections, organisations in the UK have been required by successive governments to seek public-private partnerships to digitise their collections. Under these deals, commercial companies who pay for collections digitisation can charge for access to those collections for a certain number of years. While research or direct government funding has paid for a certain amount of digitisation in the past, it is dwarfed by commercial funding. As a blog post by a British Library curator puts it:

'It has long been the goal of the British Library to make some of its digitised newspapers freely available online, but we also want to see the BNA [British Newspaper Archive] succeed as it has been doing, without which we could not have reached such a huge collection overall of digitised newspapers, nor the rate at which they are being produced (currently around half a million pages are being added to the BNA every month).' (McKernan 2021)

While most material digitised by commercial companies is eventually made more freely available, paywalls are one important factor in the fragmentation of access to collections. This was extensively documented in a Digital Collections Audit commissioned for the AHRC-funded research programme Towards a National Collection (TaNC) (Gosling et al. 2022).

The variety of sources of access to digitised collections from the British Library as in September 2023 helps illustrate this point.  The British Library's websites and catalogues hosted hundreds of thousands digitised books, manuscripts, 3D objects and sound files. Open access images and metadata were also available on Flickr Commons, Europeana and Wikimedia Commons. Newspapers were available on the British Newspaper Archive. Other commercial databases held selected other items, while other items were available via union catalogues. The Library's Research Repository had more open access items, from digitised images to datasets created via research projects like Living with Machines.  A large backlog of items was being processed for publication online. Altogether, this made answering an apparently simple enquiry about the availability of a particular item or collection online surprisingly complex.

However, this patchwork of platforms had one positive effect. While library users lost access to online catalogues, digitised and born-digital collections and the many services hosted by the British Library in a ransomware attack in October 2023, collections hosted on commercial and non-profit third-party platforms (including the Research Repository) continued to be available. The saying 'Lots of Copies Keep Stuff Safe' has never been more pertinent.

At the time of writing, the UK government is again investing in digital infrastructure, including the TaNC project and its successor, N-RICH. N-RICH has commissioned a range of activities to  'examine the scope, costs, risks, impacts and benefits of a future digital research infrastructure for cultural heritage in the UK' (Towards a National Collection 2025).

Digital records are fragile: the need to fund long-term access

Enabling access to collections online is one thing. Sustaining that access in the longer-term is another. The ransomware attack that took out most of the British Library's services was just the most dramatic incident that reduced online access to collections. Websites and platforms gradually disappearing when funding ends has eroded earlier digitisation efforts (Dunning 2009). For example, the results of the first big digitisation effort in the UK, the New Opportunities Fund c. 2000 – 2004, have long since been lost. Culture Grid, the UK's aggregator for Europeana, itself built on an earlier no-longer-supported project (the Peoples Network Discover Service), failed to find a sustainable financial model in the 2010s (Gosling 2019; Poole 2015).

More recently, collections sites have struggled to stay online through surges of bots scraping their sites to feed AI models. For example, the British Library's Research Repository was being scraped so heavily that it effectively experienced a Denial of Service attack until the team were able to put preventative measures in place.

Resourcing digital preservation and sustainability is not easy. In order to make the MDS a relatively lean operation, an early decision was made to exclude images and other media from the data aggregated. In addition to reducing financial costs, this reduces the environmental overhead of the service. However, not only does this add additional retrieval steps for researchers wanting to access images, it also means that the MDS cannot act as a backup of last resort for museums who lose access to their image assets.

Looking back over the long history of digitisation in museums, libraries and archives in the UK, it is clear that new economic models are needed to ensure the long-term availability of digital cultural heritage collections. Coming up with these models will require creativity and persuasive powers worthy of the collections they would protect.

References

  • Bailey, Rebecca, Javier Pereda, Chris Michaels, and Tom Callahan. 2024. Unlocking the Potential of Digital Collections. A Call to Action. Arts and Humanities Research Council. https://doi.org/10.5281/zenodo.13838916.
  • Bailey-Ross, Claire, Emily Burgess, and Panagiotis Papageorgiou. 2024. User Research: UK Gallery, Library, Archive and Museum  (GLAM) Digital Collections Infrastructure. Towards a National Collection. https://doi.org/10.5281/zenodo.12751226.
  • British Library. 2024. Learning Lessons from the Cyber-Attack: British Library Cyber Incident Review. British Library. https://www.bl.uk/home/british-library-cyber-incident-review-8-march-2024.pdf/.
  • Department for Digital, Culture, Media & Sport. 2016. The Culture White Paper. Department for Digital, Culture, Media & Sport. https://www.gov.uk/government/publications/culture-white-paper.
  • Dunning, Alastair. 2009. ‘Digitising the Past: Next Steps for Public-Sector Digitisation’. Digital Information-Order Or Anarchy? http://eprints.rclis.org/handle/10760/18048.
  • Europeana PRO. n.d. ‘Europeana Data Model’. Accessed 7 May 2026. https://pro.europeana.eu/page/edm-documentation.
  • Gosling, Kevin. 2019. ‘Nothing New except What Has Been Forgotten’. Collections Trust, November 27. https://collectionstrust.org.uk/blog/nothing-new-except-what-has-been-forgotten/.
  • Gosling, Kevin, Gordon McKenna, and Adrian Cooper. 2022. Digital Collections Audit. Towards a National Collection. Collections Trust. https://doi.org/10.5281/zenodo.6379581.
  • McKernan, Luke. 2021. ‘Free to View Online Newspapers’. The Newsroom Blog, August 9. https://blogs-archive.bl.uk/thenewsroom/2021/08/free-to-view-online-newspapers.html.
  • Parry, Ross, Stef De Sabbata, Andrew Ellis, Helen Hardy, and Mia Ridge. 2025. ‘A Framework for Sustainable and Scalable Cultural Data Integration and Analysis: Using the Large Dataset of the Museum Data Service’. 2025 11th International Symposium on System Security, Safety, and Reliability (ISSSR), April, 417–24. https://doi.org/10.1109/ISSSR65654.2025.00062.
  • Poole, Nick. 2015. ‘Guest Blog: Aggregation & the Culture Grid’. Museums Computer Group, May 31. https://museumscomputergroup.org.uk/culture-grid/.
  • Rawson, Katie, and Trevor Muñoz. 2016. ‘Against Cleaning’. Curating Menus, July 6. http://www.curatingmenus.org/articles/against-cleaning/.
  • ‘Spectrum’. n.d. Collections Trust. Accessed 7 May 2026. https://collectionstrust.org.uk/spectrum/.
  • Towards a National Collection. 2025. ‘N-RICH Prototype’. July. https://www.nationalcollection.org.uk/n-rich-prototype.

Upcoming talks and travel

Poster for a talk at Trinity College Dublin with illustrations of a man working in a factory and a soldier in a trench with sandbags and barbed wire
Trinity lecture poster

Get in touch if you'd like to meet for a chat about crowdsourcing / digital participation, digital scholarship, digital humanities or AI / machine learning in libraries, archives and museums! Or, indeed, the Museum Data Service, though the MDS contact page might be a better bet.

I'll be at the Digital Humanities 2026 conference in Seoul, South Korea at the end of July. Say hi if you'll be there too!

I was at Digital Humanities 2025 in Lisbon July 14 – 19, presenting with the British Library's Universal Viewer team and on a panel.

In September 2025 I was part of a residency at the Lorenz Center in Leiden for 'Enriching Digital Heritage with LLMs and Linked Open Data', and presented online in a programme on 'Múza, nebo hrůza? AI v literatuře, knihovnách a kultuře vůbec' at the Knihovny současnosti 2025 (Contemporary Libraries 2025) conference.

I also gave a keynote at the first RIDLE:HE (Research in Digital Learning in Higher Education) conference in Newcastle.

Previous activities are listed on '2024, an overview', etc.

Recent books

I'm currently working on chapters for the final Living with Machines book.

Chapter 5: Analysing the language of mechanisation in nineteenth-century British newspapers by Barbara McGillivray, Nilo Pedrazzini, Arianna Ciula, Jon Lawrence, Tiffany Ong, Mia Ridge, Miguel Vieira is now online for 'early access'.

In January 2023, Collaborative Historical Research in the Age of Big Data: Lessons from an Interdisciplinary Project by Ruth Ahnert, Emma Griffin, me and Giorgia Tolfo was published by Cambridge University Press.

In 2021 I wrote another book with 15 or so brilliant co-authors: The Collective Wisdom Handbook: perspectives on crowdsourcing in cultural heritage

My edited volume on 'Crowdsourcing our Cultural Heritage' for Ashgate, featuring chapters from some of the most amazing people working in the field was published in October 2014 and reprinted a few times subsequently. You can read my introduction on the OU repository: Crowdsourcing Our Cultural Heritage: Introduction.

By day, I usually at work at home at the British Library, so drop me a line if you'd like to meet for coffee and a chat. My availability for events is limited, but you can drop me a line if you'd like to book me for an event.

Some recent papers

Some publications are listed or accessible at my ORCID page, my Open University repository page, Humanities Commons page, Zenodo, and my Zotero page.

This page is rarely up-to-date or complete, but here's a summary of talks, fellowships, writing, etc in 2023, 2022, 2021, 2020, 2019, 2018, 2017, 2016, 2015, 2014, 2013, 2012 and 2011. You can also follow me on twitter (@mia_out) mastodon @mia@hcommons.social / https://hcommons.social/@mia for updates. I'm also on bluesky @miaout.bsky.social.

Previous papers are generally listed at miaridge.com or on my blog, Open Objects.

'Community Engagement and Special Collections' talk

In April 2024 I was one of four presenters at the Association for Manuscripts and Archives in Research Collections (AMARC)'s Spring Meeting on 'Community Engagement and Special Collections', sharing our work on 'successful projects and strategies for engaging public audiences in meaningful ways through in-person events and digital outreach activities'

I presented on 'Living with Machines: Crowdsourcing transcriptions for digitised historical collections of the British industrial revolution'. The video from the seminar is below.

Keynote video 'Evolutionary Innovations: Collections as Data in the AI era' for Making Meaning 2024

Making Meaning 2024: Mia Ridge Keynote

My slides for #SLQMakingMeaning #CollectionsAsData, 'Evolutionary Innovations: Collections as Data in the AI era', are online at https://zenodo.org/records/10795641

‘Collections as data’ describes the movement to publish open data from museum, library and archive collections that began in the noughties. The benefits of machine learning for better discoverability and research with digitised/born digital collections are alluring. And the popularity of generative AI – and an increased awareness of the biases it reinscribes – has focused attention on responsible computational access to collections – but what does this mean in practical terms? Mia will share examples from the British Library and the Living with Machines data science project.

'Enriching lives: connecting communities and culture with the help of machines': my EuropeanaTech 2023 keynote

Panorama lit by natural light of a seaside town
The video for my opening keynote on 'Enriching lives: connecting communities and culture with the help of machines' for the EuropeanaTech 2023 conference is now online.

The EuropeanaTech 2023 conference was held in The Hague, the Netherlands and online from 10 – 12 October 2023. My slides are online.

My abstract: I’ll begin with an overview of current developments in AI and machine learning, then present work with crowdsourcing from the Living with Machines project to think about what AI means for online volunteers and communities around digital cultural heritage. I’ll share new thinking on ‘volunteer enrichment’ – participation in crowdsourcing that not only enriches and enhances collections records, but also enriches the lives of volunteers. How can we embed GLAM values when we apply AI and machine learning tools in our work?

In preparing my keynote I revisited my keynote for EuropeanaTech 2011, and reflected on work on crowdsourcing, data science and AI at the British Library, the Collective Wisdom project and Living with Machines since then.