Data Rescue Project
Lynda Kellam, Halle Burns, Mikala Narlock, Lena Bohman, Kathleen Burlingame, Sebastian Majstorovic, Tess Grynoch, and Amy Nurnberger
![]() |
![]() |
In January 2025, the United States federal data landscape changed overnight. Within days of the new presidential administration taking office, it became clear that datasets across dozens of federal agencies were at risk of removal or alteration — particularly those related to climate change, environmental science, and sexual orientation and gender identity. These efforts moved with alarming speed and breadth, affecting agencies from the Centers for Disease Control and Prevention to the Bureau of Labor Statistics, the Department of Agriculture to the Department of Education.
The scale of what was at stake is difficult to overstate. Federal data underpins everyday American life in ways that most people never consider: it enables phone applications to function, informs home-buyers about local schools, supports food supply chains, and provides the evidentiary foundation for public health decisions, journalism, and academic research. When that data is removed or altered without notice, the consequences ripple into communities, classrooms, and research institutions. The volume of data at risk, and the speed at which it was being removed, made it immediately clear that no single institution could respond alone.
In February 2025, representatives from three research data organizations — the Data Curation Network, the Research Data Access and Preservation Association (RDAP), and IASSIST — convened to coordinate a response. Initially the focus was on information sharing: identifying what data were at risk, what had already been removed, and where researchers could still find what they needed. Within days, however, it became clear that a coordinated, active effort to capture and preserve at-risk data was the only adequate response. The Data Rescue Project was formalized, and its Steering Committee took shape.
At the heart of the Data Rescue Project is a deliberately designed rescue workflow. Rather than capturing files indiscriminately, the DRP prioritized curation: rescuing not just datasets but the associated documentation that makes data usable over time. Crucially, the workflow ensured that essential metadata was captured alongside each dataset, making rescued materials discoverable and preventing duplication of effort across a distributed volunteer base. The workflow was designed to be accessible to beginners and experts alike, lowering the barrier to participation without sacrificing rigor.
In just fourteen months, the DRP facilitated the capture of nearly 3,000 datasets (2,944 as of April 2026) from 97 US federal offices. This is not a collection of raw files sitting on a hard drive; it is a curated, documented, discoverable archive, built to preservation standards, that ensures the data remains usable for researchers, journalists, and the public long into the future.
A conscious strategic decision shaped the scope of this work. Following the first Trump administration, several organizations had launched specifically to protect climate and environmental data. The DRP identified a significant gap: social sciences data, covering demographics, labor, health, education, and more, had no equivalent advocacy community. By focusing on this underserved area, the DRP ensured that datasets which might otherwise have been permanently lost were captured and preserved.
None of this was possible without people. Through coordinated recruitment, training events, hackathons, and train-the-trainer sessions, the DRP mobilized more than 500 volunteers to participate in the rescue process. Coordinating that many people while accounting for time zones and differing levels of technical expertise required thoughtful care and connection. The DRP operated during a period of significant stress and uncertainty for many of its participants, and the Steering Committee made a deliberate commitment to support volunteer wellbeing alongside the technical work. That human dimension is inseparable from the project's success.
The preserved datasets are not sitting in an inaccessible archive. Many have been deposited in DataLumos, a crowd-sourced repository for US federal data hosted by ICPSR at the University of Michigan, where they have been downloaded more than 15,000 times. Given the distributed nature of data rescue efforts more broadly, other datasets have been deposited across multiple repositories. To address the discoverability challenge this creates, the DRP developed a dedicated portal aggregating rescued data from across the ecosystem, amplifying not just their own work but the efforts of allied organizations.
The real-world impact of accessible, preserved data is already visible. For example, the Homeland Infrastructure Foundation-Level Data (HIFLD) was discontinued in September 2025. This resource provided essential data for city planners, analysts, architects, and more. Thanks to the concerted effort of the DRP, the data were captured. A few months later, the data was used to create a new portal to serve the same users.
The Data Rescue Project has never treated its work as a one-time emergency response. From the outset, the Steering Committee has been building toward extensibility and emphasizing sustainability. The DRP is developing a comprehensive data rescue playbook: a practical guide that will enable other institutions and communities facing similar threats to launch their own rapid-response workflows quickly and effectively. The need for such a resource is not hypothetical: war, political change, and climate disaster will require coordinated data rescue responses in many countries and contexts in the years ahead. The DRP has demonstrated that such a response is achievable, and is now ensuring that others can replicate it.
The Data Rescue Project is proof that protecting the digital legacy is as much a human endeavor as a technical one. It required solidarity and a firm commitment to ensure that public access to public data remains a public good. In fourteen months, a community of volunteers has preserved nearly 3,000 federal datasets, trained hundreds of participants, and built the infrastructure for others to follow. The data belongs to the American people. The DRP made sure they could keep it.























































































































































