Example benefits and risks for typical digital preservation business cases
This section provides example benefit and risk statements for four typical digital preservation scenarios. Use this to build a convincing and realistic proposal.
The four scenarios are:
-
Developing a digital preservation strategy/roadmap
-
Increasing staffing complement
-
Repository migration
-
Procuring a new digital preservation system
The statements provide suggestions that can be developed into fuller descriptions that take in your particular organizational context. Note that these examples are not considered to be comprehensive and should not be included in a business case without careful consideration. Many of the risks may need careful consideration that results in changes or mitigations in your Implementation Plan rather than explicit articulation as a significant risk to the project.
Developing a digital preservation strategy/roadmap
Potential benefits and risks associated with developing a new strategy or roadmap for digital preservation within a particular organization.
Benefits |
Risks |
|
|
Increasing Staffing Complement
Potential benefits and risks associated with increasing the digital preservation staffing levels with new roles.
Benefits |
Risks |
|
|
Repository migration to cloud
Potential benefits and risks associated with migrating from a home grown and on-premise digital preservation system to a cloud hosted system.
Benefits |
Risks |
|
|
Procuring a Digital Preservation System: Benefits and Risks
Potential benefits and risks associated with procuring and establishing a new digital preservation system.
Benefits |
Risks |
|
|
Step-by-step-guide to building a business case
An outline of the the main steps to follow in constructing your business case including research, drafting, validation and delivery. Use this to plan out how you will develop your proposal.
1) Establish a foundation for your business case
Having the right foundation in place from which to launch a business case can be critical, especially if your organization is relatively new to digital preservation. The first section of this toolkit, Understand your digital preservation readiness, provides a detailed discussion of actions that may be useful to undertake before developing a business case, including:
-
Modeling your digital preservation maturity levels
-
Gaining an awareness of useful context and previous digital preservation track record
-
Understanding your staffing and skill levels
-
Developing a digital preservation policy
-
Establishing a digital preservation steering group
2) Perform pre-business case research and planning
In order for your business case to be successful, it is important to gather some key information about your organization and its stakeholders before you start writing. Initially you could establish expectations for the business case itself by:
-
Finding someone who has submitted successful business cases at your organization and learning from their experience
-
Understanding how much research/preparation is appropriate for this business case
-
Understanding the organization’s expectations for a business case
-
Establishing whether the organization expects you to use a specific template for your business case?
-
Understanding the format (text, diagrams, tables, presentations) and extent (one-pager, 30 pager, elevator pitch, 10 minute presentations etc.) for different audiences.
-
Identifying organizational allies who can support the research and development of your business case. For example see “Digital Preservation and Enterprise Architecture Collaboration at the University of Melbourne” page 155.
3) Understand your organizational context
Understanding the context within which your business case will function is an essential part of preparing to develop the case itself. It is important to factor in the many different aspects of your organization that your business case will depend on. A PESTLE (Political, Economic, Social, Technological, Environmental, and Legal) analysis is one way to structure and guide an assessment of these contextual issues but some critical elements to consider are discussed in more detail below.
4) Align with your organization’s mission and strategy
Any business case should further the broader aims of the organization, which makes it essential that you align your proposal with current strategic objectives. Typically these will be described in an organizational strategy or policy, but also perhaps in lower level departmental or sectoral strategic plans. There may also be a wider organizational mission that should be considered. Your business plan stands a much better chance of success if you can clearly demonstrate how it will advance these organizational objectives either directly, or by providing enabling actions that will provide the foundation for objectives to be achieved. So ensure you understand your organization’s objectives and ideally gain a feel for the decision makers opinion on current priorities.
Some organizations will have established policy frameworks which can define constraints within which operations are conducted. Alignment with a digital preservation policy may prove beneficial (see Understand your digital preservation readiness).
Growing concerns about the impact of our actions on the environment and the climate make it especially important to consider how the activities outlined within a business case might have an ecological impact. Most organizations will have an environmental policy which should be considered carefully in relation to your business case.
5) Identify your audience
Before you start writing your business case, you need to determine your audience(s) so that you can adopt the right tone and language, as well as being prepared with the right background knowledge. This will ensure that your document is impactful. Consider the following: which organizations, departments or individuals will assess your business case? Which parts of the organization (like teams or committees) will be affected by it, and to what extent? By identifying the people to whom you are presenting your business case, you will gain a better understanding of your audience and their priorities. Think about who will be evaluating your business case or making the decision to implement your proposal, even if they are not a part of your organization. These external stakeholders might have different needs and expectations. Other practical tips include:
-
Network and find a champion for your business case amongst the decision-makers
-
Be mindful of staff turnover - make the most of time spent courting champions, and don’t rely too much on one person
-
Be aware of your organization’s appetite for risk, and don’t over-promise.
-
Identify “hooks” that may play well to key stakeholders, e.g. avoiding reputational risks or facilitating a favored project. And at the same time avoid topics or arguments that may not play well with key stakeholders.
6) Choose the right moment to develop and submit your business case
Choosing the right time to develop and submit a Business Case can be crucial in achieving success, aligning with organizational processes, avoiding distractions and maximizing the potential support for building your case.
The business case should not come as a surprise to anyone and stakeholders should ideally have been brought on-board with key elements of the proposal as much as possible before they see the detail.
There may be constraints on when you can submit a business case within your organization’s planning and funding cycles. If there is insufficient funding left in this cycle, then it may, for example, be more realistic to delay a business case until the next funding year. Some organizations have specific cycles and deadlines for submitting business cases on a quarterly or annual basis.
There may be optimum times for building a Business Case, such as when there is better availability of your business analyst support staff and other stakeholders who can make a useful contribution.
Avoiding periods of busy, or peak activity within your organization could be important. There may be periods of organizational downtime, for example outside term times at an academic institution. Staff turnover might leave critical posts empty or lead to the loss of critical champions for your business case.
7) Develop a realistic budget
Decision-makers will want to know how much money you will need to realize your plans. It is crucial you provide them with a clear overview of all costs involved, including staffing, training, procurement, maintenance, and development. Make sure you have an idea of available budgets before you start with your business case, for example by looking at previous examples for similar activities or projects within your organization. Consider the following questions:
-
Does funding work on a project basis or can a new and permanent funding stream be established?
-
Which components of the project will your organization provide funding for?
-
If your organization has any guidelines or constraints on levels of contingency etc., it is vital that you keep to these.
-
Consider breaking your proposal into separate phases, each with clearly defined objectives and success criteria if you think that will be more acceptable to your organization than making a significant one-off commitment to fund a single large proposal.
8) Draft your business case
The roles and responsibilities for writing the business case should be clear to everyone across the organization. This may be straightforward for a small organization, but may involve more people as the size and scale of the proposal increases. Prior to embarking on the full business case, options should be explored to establish whether there is scope for a pre-business case to be presented without detailed costings and other time-consuming content being included. This ‘concept document’ can be used to test the organization’s appetite for a proposal (focusing on a SWOT analysis, strategic fit, need and viability) to make sure the time invested in developing a business case is not wasted.
When it has been established that the full business case is required, the sponsor of the proposal (ideally someone as senior as possible in the organization) should authorize commencement. Work should then begin to populate the most appropriate template with well-written, clear and focused content.
By the time writing starts, the key facts, evidence, costs, benefits, value, impact and other detailed areas of the proposal should all have been established. But it may be appropriate for the designated lead author to ask for specialist input from other key people (e.g. IT, finance, HR, third-party providers, technical writers) to ensure that the content and language is accurate and effective. See the section “Template for building a business case” for more help with the detailed content of your business case.
The document will almost certainly benefit from being checked by at least one person who can objectively assess the proposal without confirmation bias, especially if they have prior experience of assessing other business cases. Authors should approach the creation of the proposal with a flexible mindset and be open to ideas and changes to content and structure as the proposal comes together. Independent advice and feedback on your business case might be especially valuable. If possible seek guidance from the Digital Preservation Coalition and/or your peers at other organizations.
In many instances, organizations may prefer to initially respond to a summary pitch (often a slide deck) that represents the fuller proposal. The quality, clarity and persuasiveness of this summary is a critical component in establishing the value of the idea. If the purpose of the pitch is to provide a decision-making group with a foundation for discussion on the merits of the idea, it is vital that they start that discussion with positive feelings about the effort that has gone into drafting the proposal.
Understand your digital preservation readiness
This section considers what context and preparations are needed before producing a business case. Use this to establish the best foundation for your business case.
Establish your level of digital preservation maturity
Your organization’s digital preservation maturity can be quickly identified using the DPC Rapid Assessment Model (DPC RAM) which provides a set of organizational and service level capabilities that are rated on a simple and consistent set of maturity levels. The model enables organizations to monitor their progress as they develop and improve their preservation capability and infrastructure and to set future maturity goals. Understanding your maturity level helps in identifying gaps in current capability and areas to be developed in the future.
Understand the history and track record of digital preservation at your organization
Researching the history of digital preservation work at your organization can provide context for your business case, avoid previously encountered pitfalls and ensure learning from previous work. In particular, you should seek to understand and acknowledge investment (e.g. in systems, staffing, training, professional organizational membership) successes, failures, and missed opportunities:
-
Have there been digital preservation business cases at your institution before? If they were successful, can you show that the money was well spent? If they were unsuccessful, can you address or at least acknowledge what went wrong in any new business case?
-
How has your organization’s digital preservation capacity changed over time? You may be able to establish this with DPC RAM scores (see above).
-
Make sure to acknowledge any existing investment in basic digital preservation services like digital storage, systems, staff, and training.
-
Check whether those responsible for IT provision are aligned with the need for a digital preservation solution and understand what value it brings in addition to storage and fixity.
Assess organizational policies that might impact significantly on future digital preservation activity, including the need for a dedicated Digital Preservation Policy:
-
Do you have existing policies that establish a requirement to preserve digital content e.g. collection policies, records management policy, open access policy.
-
Consider whether any other policies will be relevant e.g. closure periods or embargos, IPR, data protection, environmental or potential liability issues.
-
Do you have a Digital Preservation Policy? Establishing a policy can provide a helpful foundation from which to build further infrastructure and activity, as well as acting as a useful advocacy tool. See the DPC Digital Preservation Policy Toolkit for more information.
Understand key disasters, risks, and missed opportunities relating to digital content that has already been identified by colleagues, external audits, customers/users, and other stakeholders:
-
Has there been previous data loss, reputational damage or legal damage due to poor digital preservation practice at your institution?
-
Do stakeholders have live, documented digital preservation concerns that need addressing?
-
Has your institution missed out on a benefit that might be obtained by better investment in digital preservation? What benefits are other institutions able to realise that you cannot?
-
There may also be examples where good digital preservation practice has previously “saved the day”/minimised the impact of an event/protected your organization’s reputation.
Understand current staffing, roles, skills and competencies
Successful digital preservation requires skilled staff and a key component of a business case might be a request for new staff and/or additional training and development opportunities. It is, therefore, important to understand which members of staff currently work on digital preservation activities, where they sit within the organization, what skills they have, and where gaps might exist in terms of both roles and competencies. Which members of staff are within scope will be context dependent but may include: digital preservation practitioners, other information management professionals, IT staff, digital content creators, and more.
The DPC’s Competency Audit Toolkit (DPC CAT) can be used to assess current staffing capabilities within an organization, identifying where gaps exist in relation to current and target digital preservation capabilities as identified during a DPC RAM assessment.
Understand and utilise relevant governance structures
Do you have a steering group or advisory board of relevant stakeholders that can guide digital preservation practice at your organization? Establishing a steering group can provide cross-organizational support, advice and experience of real value to your business case.
Assess the target of your preservation activity: the digital content
Establishing the scope of the digital content to be preserved, which is of relevance to your business case, will be fundamental to understanding requirements and potential solutions.
-
It may be necessary to undertake an audit of current and expected collections within the organization to gather information on the size and nature of the digital content they contain. This information will clarify areas of need and help facilitate decision-making about what should be included within the business case.
-
To help structure this process and gather the information in a reusable format you may wish to capture the information in a Digital Asset Register, if you do not already maintain one. This will then provide a useful tool for both developing the business case and ongoing management of the digital content within your collections. To help gather the information needed, you may wish to use characterisation or content profiling tools to capture data on volumes and types of digital content within the collections. Information on what tools are available can be found on the COPTR tools registry.
-
For content which you do not currently collect, or which does not yet exist, consult with key members of staff (either within your own, or within comparable organizations if possible) to establish well-informed estimates of the size, complexity and challenges of such material. You must be able to justify and defend any figures derived in this way if they are crucial to your business case, so record your sources and your working even if this will not be included in the final business case.
-
The scale and scope of content may well increase over time. Estimates of how volume is likely to change may help to inform a future proof solution.
Identify systems and workflows currently in use
It will be useful to establish the systems and workflows used to create, manage and store the digital content that will be preserved, as well as the systems they interoperate with. This will help you assess current digital preservation capability, gaps in capacity and associated risks. This will be important for identifying potential improvements and solutions and understanding how these solutions fit into the existing system landscape at your organization. This is likely to require conversations with colleagues across several areas, including IT and administrative staff, any third-party solution providers, and potentially external creators and users of digital content.
Understand current practice at other organizations
Gain an understanding of digital preservation capacity and practice at other, similar organizations. This can aid with benchmarking your current capacity as well as helping to set the goals that will be achieved if your business case is funded. It may even be possible to learn lessons from others who have already taken a similar challenge to which your business case seeks to address.
-
There may be particular organizations that your organization typically compares itself with, possibly engaging in peer review or collaboration. These might make useful points of comparison for your business case, particularly if they have already implemented work similar to what you are proposing.
-
If you already have connections with colleagues at similar organizations, you may be able to gather this information through informal conversations. If you wish to make new connections this might be done through membership of networks such as the DPC, the NDSA, or Australasia Preserves or at conferences and events such as iPres, IDCC, or PASIG.
-
DPC members can consult information gathered through a yearly comparison of member RAM assessments - contact your DPC Champion for further information.
Introduction to the Business Case Toolkit
This section provides an overview and introduction to business cases and this toolkit. Use this to find out more about the toolkit and how to apply it.
What is a Business Case and what is it for?
Digital preservation within any organization requires investment in people, processes and technology. If an organization would like to change, adapt or expand its digital preservation activities, it is likely that you will need to prepare a business case for further investment in supporting infrastructure, staffing and/or training. A business case will outline the resources required and provide justification for undertaking that investment; evaluating the benefit, cost and risk of alternative options and providing a rationale for a proposed solution.
What is the Digital Preservation Business Case Toolkit?
This Toolkit provides a guide to writing a business case focused on digital preservation activities. While there is no perfect formula for writing a business case (it all depends on what the business case is for, who it is aimed at and how your organization wants you to present it) this toolkit provides a variety of ways to get you thinking about what should go into your business case.
The Toolkit starts with factors to consider when planning your business case, before offering templates and a step-by-step guide to drafting, constructing and delivering your finished case:
-
What makes a good digital preservation business case captures experience pooled from within the DPC on what makes a good policy. Use this to avoid obvious pitfalls and take advantage of approaches that have been successful for other DPC Members.
-
Understand your digital preservation readiness considers what context and preparations are needed before producing a business case. Use this to establish the best foundation for your business case.
-
Step by step guide to building a business case outlines the main steps to follow in constructing your business case including research, drafting, validation and delivery. Use this to plan out how you will develop your proposal.
-
Template for building your business case explores the detail that makes up an evidenced and convincing business case. Use this to guide the development of the content of your proposal.
-
Example benefits and risks for typical digital preservation business cases provides motivations for 4 common business case targets. Use this to develop a strong case for your project.
-
Business case hints and tips provides a summary of helpful ideas on writing business cases and selling your proposal to your organization. Use this to ensure you create a positive and convincing business case.
-
Further resources on business cases provides a list of additional guidance materials relating to business cases.
Who is this toolkit for?
This Toolkit is for anyone who would like to create a business case focused on digital preservation. It is targeted at practitioners (and their managers) who are working with digital resources and would like to obtain funds to expand their digital preservation activities. The Toolkit is primarily aimed at those seeking further funds from within their organization, but could also provide useful information for those writing a bid for project funds from an external funding body.
A great deal of internal advocacy and preparation is often required before you get to the stage of preparing and submitting a business case for a major digital preservation activity, such as establishing a new digital preservation system. For support with pre-business case advocacy activities see the Executive Guide on Digital Preservation and the DPC Rapid Assessment Model.
If your Digital Preservation Business Case has been approved, consult the DPC Procurement Toolkit for advice on the next steps!
Introduction
Providing access to digital content is a core activity for digital preservation practitioners, making up one of the six functional entities of the Reference Model for an Open Archival Information System (OAIS) or ISO 14721:2012. It is widely recognised within the digital preservation community that there is little point in preserving content for the long term if there is not an intention to facilitate access at some point (either now or at a future date), but even providing simpler forms of access to content can be challenge for digital preservation practitioners (as described in Developing an Access Strategy for Born Digital Archival Material), This is particularly the case where large volumes of content and more complex methodologies are employed.
Computational methods of providing access to digital content and metadata are generally considered to be more advanced techniques, certainly a step up from the more standard models of access, for example where a user can browse an online catalogue and view or download one file at a time.
The Levels of Born Digital Access from DLF is a helpful and practical resource which articulates three levels of access under a series of headings, moving from the most simple to the more advanced. Computational access techniques are included at the highest level of the model where the ‘Tools’ section of level 3 states that an organization should “Provide remote access and sophisticated tools for exploring, rendering, and interpretation of data; provide hardware and software to support access to legacy/obscure content, including emulation services.” Examples given within the supporting information include:
-
“Provide open and web-based remote access to materials, including via programming interfaces” and
-
“Provide software for exploring, rendering, and interpreting materials, such as text mining, data visualization, annotation, and natural language processing tools.”
Similarly, the DPC’s Rapid Assessment Model (DPC RAM) puts computational access techniques at the highest level of the model. Level 4 of 'Discovery and access' states that “Advanced resource discovery and access tools are provided, such as faceted searching, data visualization or custom access via APIs”.
There is typically no one-size-fits-all approach to digital preservation and this also extends to access strategies. Organizations are encouraged to weigh up their own priorities, resources and the needs of their users in informing their own approach. So whilst it is acknowledged that not all practitioners will strive for the highest levels of either of these models, many in the community are curious about understanding and exploring these more advanced approaches of access in order to inform their own decision making.
The access strategies of an organization should of course be aligned with the needs of their users. The growing desire for users to be able to carry out their own computational processing on archival metadata for example is mentioned in Born digital archive cataloguing and description. Whilst user needs are not covered in any great detail in our online resource, a helpful introduction can be found in Understanding user needs and it is acknowledged that engaging with users should be a key step in establishing appropriate access strategies.
What is the purpose of this guide?
Computational access is a term mentioned with increasing frequency by those in the digital preservation community. Many practitioners are aware it might be helpful to them (and indeed to their users), but do not have an understanding of what exactly it entails, how it is best applied and, perhaps most importantly, where to start. To add to the challenge, computational access raises professional and ethical concerns. These well-founded but sometimes partially formed concerns, in combination with a lack of practical experience and know-how, mean computational access has been relatively slow to develop within the digital preservation community despite its potential to help with our ambitions to facilitate greater access to digital archives.
The topic was highlighted as a priority by DPC members at the DPC unconference in June 2021. It was clear from discussions that digital preservation practitioners felt this was an area they would like to explore, but one of the key barriers was simply not knowing where to start. This guide has been created to provide an introduction to this topic, and to help the community move forward in applying computational access techniques.
Who is this guide for?
This guide is primarily aimed at digital preservation practitioners with no prior knowledge of computational access. It is a beginner’s guide, intended to provide an overview of key topics as well as tips on getting started and examples of a range of different implementations. It does not hold all the answers, but instead aims to move practitioners towards an understanding of computational access terms and approaches and give them the necessary information and resources to consider whether these techniques could be used to provide access to the digital archives that they hold.
It is not specifically aimed at researchers or users of collections who might want to use computational access techniques to analyze and understand digital collections. Other resources that will help those users are available – see, for example, the Programming Historian and GLAM Workbench. It is important, however, that digital preservation practitioners keep potential users and use cases in mind whilst reading and using this guide
Definitions
Computational access is often linked with terms such as text mining, machine learning and artificial intelligence, so much so that there is understandable confusion around what each of these concepts entails and where the areas of overlap occur. This section provides clear definitions of key terms and the relationships between them.
Computational Access
The term computational access relates to the ability to enable users to access collections within a digital preservation repository (in a machine-readable manner, e.g., via download or API) in order to analyze, interrogate, or extract new meaning from that material (through, e.g., data or text mining, machine learning) as a means of investigating a particular research question.
This term is first found in the literature in a report by HathiTrust and is closely linked to ‘Collections as Data’ which similarly urges organizations to make their collections available as data, therefore making it possible for users to compute over the material. The term is closely linked to a number of definitions, highlighted in bold above; these and other related terms are discussed below in alphabetical order.
Algorithms
An algorithm is a set of coded instructions normally followed to solve specific problems. A real-world example would be baking a cake; other examples are the process of doing laundry, or the method used to solve a logistic problem.
Artificial Intelligence (AI)
Artificial Intelligence (AI) has been defined by the UK Parliament as:
‘Technologies with the ability to perform tasks that would otherwise require human intelligence, such as visual perception, speech recognition, and language translation.’ They also add that ‘AI systems today usually have the capacity to learn or adapt to new experiences or stimuli.’ AI in the UK: ready, willing and able?
Artificial Intelligence typically takes one specific task, which would normally be done by a human, and provides a method of reliably automating it. An example of this is classifying traffic signs, or recognizing the handwriting of a particular scribe. It especially comes in handy when working with large amounts of material which are extremely time consuming to process manually. Recent developments relating to AI and archival thinking and practice are discussed in the article Archives and AI: An Overview of Current Debates and Future Perspectives.
AI can be split into Broad AI and Narrow AI, and these terms are further defined below.
Broad AI represents a system that is sophisticated and adaptive, able to perform any cognitive task based on its sensory perceptions, previous experience, and learned skills. Currently this type of AI is not achievable due to technical limitations. Read more about steps towards a broad AI.
Narrow AI is task focused. This type of AI is very good at doing one single task, for example, classifying documents into different topics. It is also referred to as Weak AI.
Sometimes Narrow AI becomes so good at a specific task that it can give the impression that it is able to think for itself, so falling under the Broad AI marker. An example of this would be the quick improvement of Voice Assistants, such as Alexa. However, current technology is only able to support Narrow AI.
More information on the differences between these two terms can be found here: Distinguishing between Narrow AI, General AI and Super AI.
Algorithms are a large component of AI, but differ slightly, as AI takes the use of these a step further. AI is basically a set of algorithms that can modify and create new algorithms in response to learned inputs and data, as opposed to relying solely on the inputs it was designed to recognize as triggers.
Computer Vision
Computer vision is a sub domain of AI that focuses on deriving meaningful information from digital images. In an archives context, computer vision could be used to generate metadata for a set of uncatalogued digital images to enable more effective processing, or search and retrieval. A good example of this can be found here: Libraries Use Computer Vision to Explore Photo Archives. As digital images can be of a complicated nature, machine learning is the methodology typically used to carry out this task. You can find out more about computer vision here: What is computer vision?
Data Mining
Data mining is the discipline of finding patterns, correlations, and anomalies in data. A broad range of techniques can be used in data mining, including AI. The data mining workflow can be roughly split into data gathering, data preparation, training, and data analysis. AI is most commonly used during the training stage; this is when the algorithm is trained in a specific task. However, as this discipline mainly focuses on using large amounts of data, AI and algorithms can also be used to aid other steps of the workflow. For example, an algorithm could be written to gather certain information from the web. Find out more about data mining here: Data Mining: What it is & why it matters
Machine Learning
Like computer vision and Natural Language Processing (NLP), machine learning is a sub domain of AI. It differs slightly from computer vision and NLP, as it focuses more on the infrastructure and models than on the techniques and material that are being inputted. Machine learning takes AI to the next level; not only are the algorithms adaptable, but they are also able to perform a task without being explicitly programmed to do so. Find out more about machine learning here: Machine learning. An example of using machine learning on archival collections can be seen on the Archives Hub blog: Machine Learning with Archive Collections.
When talking about AI and machine learning, the terms supervised and unsupervised are sometimes used. These refer to the different approaches that can be taken when applying machine learning. Supervised learning is where a labelled dataset will be used for the algorithm to learn from. A labelled dataset contains items that are tagged (mostly by humans) with an informative label; one example of this is a labelled dataset of images with names of the people who appear in them attached. Unsupervised learning uses a dataset that has not been labelled. This is the biggest difference between these two approaches, but a more nuanced explanation can be found here: Supervised vs. Unsupervised Learning: What’s the Difference?
A term that is also closely linked to machine learning is deep learning. This is a more complex form of machine learning where deep neural networks are used to resemble the complex structure of the human brain. Read more about the differences between deep learning and machine learning here: Deep Learning vs. Machine Learning – What’s The Difference?
Natural Language Processing (NLP)
NLP, just like computer vision and machine learning, is another sub domain of AI. This sub domain focuses on the ability of a computer program to understand human language as it is spoken and written. You can read more about NLP here: Natural Language Processing (NLP). Most of the time, due to the complexity of human language, machine learning will be used alongside NLP to produce better results. An example of this is text classification, where due to the ambiguous and unstructured nature of human language, this approach has only been able to evolve since the use of machine learning.
Text Mining
Text mining is very similar to data mining, the biggest difference being that instead of collecting data in general, text mining focuses solely on collecting text. Therefore, while text mining is data mining, data mining is not necessarily text mining. Text mining typically results in a large quantity of unstructured text which is difficult to analyze, so more advanced methods such as machine learning are often used in association with it to help make sense of the resulting dataset. Read more about text mining and its relationship to machine learning and NLP here: What is Text Mining, Text Analytics and Natural Language Processing?
As is apparent from the definitions as described above, there are close relationships between many of the terms used. The diagram below illustrates some of these areas of overlap.

Approaches to computational access
There is more than one way to implement computational access, and the range of approaches can be confusing to those who are new to this topic. This section looks at the pros and cons of the different approaches and provides real-life examples of how they are being applied by different organizations.
Computational access can be approached in several different ways. Four approaches have been identified:
The one an organization selects will depend on its resources and priorities, the needs of its users, and any legal and ethical concerns relating to the collections or material being made available. It is also possible for organizations to apply several approaches, depending on their collections and users. The approaches described in this resource appear in order of complexity, with simpler solutions first, followed by those that require greater commitment of time and resources.
This recording from Leontien Talboom at our launch event in July 2022 provides a helpful introduction to computational access and a clear overview of the four different approaches.
Leontien Talboom, University College London - An introduction to computational access
Ethics of computational access
No guide to computational access would be complete without discussion of some of the ethical issues which should be considered when applying these techniques. This section provides a summary of some of the key considerations and signposts further sources of information.
The digital preservation community has inherited and further developed a sophisticated understanding of the ethics surrounding the provision of access to heritage collections, an understanding which is having to become even more sophisticated in respect of the provision of digital access. Some of the challenges around the ethics of access, with a particular focus on approaches and tools that have been applied at the Australian Institute of Aboriginal and Torres Strait Islander Studies (AIATSIS) are summarized in Exploring ethical considerations for providing access to digital heritage collections.
The provision of computational access will require a similar evolution in understanding for several reasons. Firstly, because computational access, as defined in this guidance, generally involves the use of algorithms or other computational methods, some of which are not in themselves commonly understood or easily explainable. This leads to issues with maintaining transparency and accountability with respect to, and hence trust in, their use and the conclusions and outcomes that arise from that. Arguably, we should for this reason be circumspect in adopting such techniques for our own (digital preservation) purposes.
Secondly, a common factor in the use of algorithms and other computational methods is the desire to be able to work at scale, processing larger amounts of data at one time than was previously possible using more traditional methods, and being able to combine disparate datasets to create a clearer picture. The bigger the scale, the bigger the potential for harm and unexpected outcomes, particularly when the data being worked on are about people. This potential for harm is heightened still further by the fact that the regulatory environment for the use of algorithms and computational methods (in all contexts, not just for the provision of computational access to digital heritage collections) is itself in an unsophisticated state of development, with discussion ongoing and the absence of any form of widely held consensus on the topic. There is a good discussion of some of these questions here: Algorithmic accountability for the public sector.
In large part then, there is an element of ‘watch this space’. However, a first step could be to enhance our documentation of any ‘sets’ of material or data to which computational access is anticipated or offered. Within the AI and data science communities increased attention is being paid to dataset documentation; see, for example, Datasheets for Datasets – Microsoft Research or the work by Eun Seo Jo and Timnit Gebru from the machine learning community. While some of the questions asked as prompts to documentation may seem odd to those used to describing digital assets for non-computational use, this highlights the information needed in order to ensure that they can use it safely. Similarly, projects such as The Data Nutrition Project and Data Hazards flag up some of the already known dangers for those wishing to analyze data using computational means and a blog from the Archives Hub describes some of the challenges of understanding the effectiveness of, and bias within the tools.
When providing access to collections, improving the documentation we offer to support users to use it more safely would seem to be the least we, as practitioners, can do. It would provide users with a context for the material and help them to make an informed and ethically responsible choice as to how to use it. Although what users ultimately do with this material is their responsibility, there is perhaps still a debate to be had about whether, in light of the increased potential for harm and the nascent regulatory environment, we feel it is our responsibility to police the uses to which our collections are put even more stringently than in the past. A good way to get started is to provide terms of use, discussed in the approaches section of this guide. Should we prohibit certain types of use? And if so, how do we design systems and procedures to prevent these?
Benefits and drawbacks
Understanding the benefits of using computational access is key to helping practitioners decide whether it is something they should explore. Awareness of the hurdles that might be encountered along the way is equally important in aiding those decisions and in helping understand the potential pitfalls once implementation begins. This section summarizes some of the key points to consider.
Benefits
Computational access opens up potential benefits both for digital preservation practitioners and for users of their digital collections. Some of the key benefits are listed below:
-
Handling bulk requests will no longer need manual intervention by staff (though note that staff time may well be spent supporting users in other ways).
-
It is empowering for the digital preservation community, as the importance of documentation and contextualization can be emphasized. See Nothing About Us Without Us for a discussion on this theme.
-
Computational access techniques enable and encourage the possibility of collaborative working with other disciplines, such as digital humanities/cultural analytics/computational social science, especially when considering access through an API, and this may bring wider professional benefits.
-
It allows digital preservation practitioners to more effectively meet the needs of users, empowering them to explore collections in novel and innovative ways and opening up possibilities for new types of research. Platforms such as GLAM Workbench are helpful in demonstrating what is possible.
-
It provides opportunities for developing new digital skills, both for digital preservation practitioners and for users.
Drawbacks
Alongside the benefits described above, there are several potential drawbacks to using computational access techniques. These are listed below:
-
Unintended bias or privacy concerns in the collections may not be revealed until the material is made available at scale.
-
Loss of control. It can be unclear who is using the material, therefore there could be unanticipated and unpredicted outcomes. This could be partly controlled with the implementation of terms of use or a policy on the use of derivative datasets, such as that published by HathiTrust.
-
Providing computational access is a long-term commitment with associated costs. It needs to be planned and resourced correctly for it to be successful.
-
Users may become reliant on computational access services that are being trialled by an organization. Managing user expectations when prototyping new types of access is important to factor in.
-
Data comes with an erroneous aura of objectivity, which can lead to the idea that it is unbiased; this is far from the truth for a lot of organizations. For more detail see Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning.
Constraints
There are also several constraints or barriers that are stopping the community moving forward with computational access. Some ideas to help overcome these can be found in the ‘Practical Steps’ section of this guide.
-
Lack of technical skills and resources within an organization to implement this approach.
-
Sustainability of the implemented architecture. Digital infrastructures and tools need to be appropriately resourced in order to ensure they are maintained over time.
-
Depending on the type of collection you want to enable access to, licensing and other copyright restrictions could make computational access difficult. An example of this is UK copyright laws around data and text mining.
-
Lack of use cases or demand for computational access can make it difficult to know what approach to implement to meet user requirements.
-
The available tools that can enable and help with this type of access are not necessarily tailored towards the requirements of digital preservation practitioners, which can make it time consuming to implement and maintain.
Practical steps for moving forward with computational access
Taking the first steps to implement or work with a new technique or technology can be a challenge, particularly when you do not know how, or even where, to start. This section offers practical tips to help digital preservation practitioners move forward, as well as case studies demonstrating how others have tackled the challenge.
Circumstances will vary widely in terms of the resources and technical skills available to an organization as well as in the type of material they are hoping to provide access to. However, the following ideas should provide a useful starting point for further exploration, whatever your circumstances. These '23 things' focus on actions aimed at both individual practitioners and at organizations, though of course there will be some overlap and you are encouraged to consider all points.
Steps for an individual practitioner
-
Do something! It is easy to be overwhelmed by the possibilities available, but the best way to learn is to experiment. For example, if you can safely do so, consider making a small dataset available and seeing how people use it.
-
There are great free resources available to learn more about computational techniques – use these to build your capacity. Good places to start are the Programming Historian or Library Carpentry. There are links to more resources at the end of this guide.
-
Talk to people you do not normally talk to. A computational access project will benefit from a multi-disciplinary approach, and the more people with different perspectives you can talk to about it, the better. This could include people within your organization or in the wider community.
-
Scope out your project meaningfully. While it might be tempting to just say ‘I want all the data’, that would not work in an 'analogue' project and will not work in a digital one either.
-
Check the licences and terms of access that apply to the material you are working on, collate these in a document so you know, and can share, what users can legally do with your content.
-
Look for existing datasets you can use to enrich the data you will be working with. For example, you might work in an organization that has catalogue entries you could incorporate into your project as metadata.
-
Familiarize yourself with the various approaches to computational access discussed in this guide. They might all be useful to different types of user, but there are implications for the skills and resources you need for your project.
-
Computational approaches and their outputs can lend work an aura of ‘objectivity’ that should be treated with caution. Keep in mind that your outputs might look more authoritative than perhaps they should, so think about how you can explain and qualify them. Make your assumptions and processes transparent to the user if you can.
Steps for an organization
-
Consider how computational access may align with, or support, organizational strategy. Being able to demonstrate a link may help you to make the case to explore new methods of access and gain the support of colleagues.
-
Apply an active outreach approach. Speak to your users, set up user groups, conduct interviews. Try to help them articulate what they want (or might want if they knew about it) from computational access.
-
Consider your own colleagues as a user community for computational access. While you might aspire to attract a new group of external users, computational access can also be of great benefit internally, by, for example, increasing your own understanding of your collections.
-
A good way to generate ideas is to create a single dataset in a familiar format such as CSV, and host an internal hackathon, giving staff time to play with the dataset and see what they come up with.
-
Think carefully about where the expertise to guide service provision for computational access lies in your organization, and how you can engage the right people. Bear in mind that expertise and responsibility might not always go together.
-
Consider sustainability from the start. Remember that users may come to depend on the services you are planning to deliver, and you have a responsibility to think about how you will sustain them. Consider the resource commitment you will need and the environmental impact of your project. Though not specifically focusing on computational access, the article Towards Environmentally Sustainable Digital Preservation discusses the need for a paradigm shift in archival practice and touches on the environmental impact of mass digitization and instant access to digital content.
-
Make documentation an integral part of your project from the very beginning. Part of the appeal of computational access is that others can build on your work in new and unexpected ways. A well-documented service, ideally with worked examples, will help facilitate this.
-
Consider funding or empowering other users to approach your collections in new computational ways. For example, you could run a content hacking competition or workshop for postgraduate students and/or community practitioners to imagine potential uses of your content.
-
Find a public domain or CC0 ‘No Rights Reserved’ collection, dataset or set of metadata and publicly release it. This can serve as a trial run to work through your processes, and as a way of engaging with users and working out what to do next.
-
Search for existing standards you can use. Do not reinvent the wheel if other organizations have done already done similar work. For example, if you adopt similar data output forms to other organizations, it will make it easier for users who are already familiar with the standard.
-
Ask lots of awkward questions during the procurement of any system. This might feel uncomfortable but will save problems in the long run. Engage expert colleagues to help you with this if needed.
-
When testing your system, and when you are in production, be sure to question your results – are you able to identify and describe bias in the results? Do not take the computer's answer as right! The following article is an interesting reflection on selection and technological bias of the Living with Machines project: Living with Machines: Exploring bias in the British Newspaper Archive.
-
It is important that users can contact you with questions or to request further data from your collection, so provide a primary contact or a contact form as part of your project.
-
Make sure you have a defined plan to transition discussions of ethics to concrete actions in systems, processes, and collaborations.
-
Going forward, embed computational access in your day-to-day collection accessioning and processing. New methods of access may require adjustments to organizational policy, accessioning and appraisal workflows and conversations or agreements with donors and depositors. The Reconfiguration of the Archive as Data to Be Mined is a helpful article which discusses changing practices brought about by the move to online digital records.
Further practical tips for getting started can be found in '50 things' which has been published as part of the Collections as Data framework and inspired this section of the guide. Though 50 things is not aimed specifically at digital preservation practitioners many of the tips and ideas are helpful and their key message "start simple and engage others in the process" is great advice!
Subcategories
Template for building a Business Case
This section provides guidance on the content that will be useful to include in your business case, but it will likely need to be adapted to the structure used in your organization’s template.



















































































































































