← Writing

The UK's vision for using patient health data

My take-aways and follow-up thoughts from the Westminster Health Forum on "Next steps for the use of patient data in the UK"

Keir Starmer announced in April 2025 that the government, along with the Wellcome Trust, will invest up to £600M to organise access to NHS data and make it easier to use for research. The initiative is being driven through the Health Data Research Service (HDRS). This half day conference took stock of the current state of research datasets in the UK, considerations and challenges, and laid out an initial vision for HDRS.

Current state

I was surprised by how many health research data environments already exist in the UK. These are secure databases that store a curated set of healthcare data, with a process for researchers to get access, and some type of council that oversees its use. They are commonly referred to as ‘Trusted Research Environments’ (TREs), ‘Secure Data Environments’ (SDEs) or ‘Data Safe Havens’ (DSHs). There were presentations from folks representing OpenSAFELY, Our Future Health and MS Register, with a number of other SDEs mentioned.

There have been several efforts to review and catalogue what is out there, for example, Research Data Scotland was created to help use the existing health data assets in Scotland, and the Health Data Gateway was established in 2020 to be a centralised catalogue for UK health datasets. Data and Analytics Research Environments UK (DARE UK) published a 2025 landscape review on the UK’s sensitive-data research infrastructure (see below). Research Data Scotland commissioned an independent report in 2026 on the maturity of health-data infrastructure in the devolved nations, which found that they have significant experience in providing access to data that exceeds the publicly documented capabilities. HDRS has also surveyed the current state in preparation for setting its strategic goals.

What research data environments exist today?

Health Data Gateway has 69 data custodians holding just over 1200 datasets. That likely means about 69 separate databases that store data from 1200 different sources.

A map of the UK showing markers for each TRE and lines connecting related ones, like the Secure Data Havens in Scotland. The table below summarizes the same information. From: p18 DARE UK 2025 landscape review

UK Health data environments
Country Full country Regional Other (disease-specific, study-specific, or partial datasets) How to find a dataset
England NHS England Secure Data Environment 10 Regional SDEs covering England, listed in the Health Data Gateway. Health Data Gateway has 42 data custodians and over 600 datasets that include data from England. Search Health Data Gateway for England
Scotland Scottish National Safe Haven 4 regional safe havens: Grampian, Edinburgh and South-East Scotland, Fife and Tayside, Greater Glasgow and Clyde. See the Scottish safe haven network. These don't have matching datasets in them. Research Data Scotland has metadata for 143 datasets, the majority of which are related to healthcare. Research Data Scotland
or for Public Health Scotland search: Health Data Gateway
Wales SAIL Databank contains over 100 datasets, primarily Welsh data but with some data from the rest of the UK as well Already centralised Search Health Data Gateway for SAIL
Northern Ireland NITRE (Northern Ireland Trusted Research Environment) is the umbrella service for getting data. The Honest Broker Service can provide extracts of health data needed. Already centralised Search Health Data Gateway for Honest Broker Service

Some interesting findings from DARE’s 2025 landscape review?

DARE got responses from 63 organisations that support TREs in the UK holding multiple types of data. There is about 1/3 health data, 1/3 administrative data and 1/3 other. The most commons types of data are all health related: secondary care (77%), research cohorts (67%) and primary care (56%) From: p24 DARE UK 2025 landscape review

Every nation has TREs, although there is a proliferation in England, with perhaps more coordination in Scotland, Wales and Northern Ireland. These mostly cost less than £1M annually, with a few more expensive ones.

This is a U-shaped bar graph with most budgets under £1M and another chunk around £4.5M - £5+M. From: p52 DARE UK 2025 landscape review

Trust depends on transparent security and privacy practices and how the data is used. The majority of TREs follow the NHS Data Security and Protection Toolkit (71% - though this is required for NHS data anyway) and ISO/IEC 27001 accreditation (61%). About half have also self-evaluated against SATRE, a framework of best practices for building, operating and evaluating TREs developed by the UK TRE community.

Perhaps surprisingly, the majority of platforms have a manual step to check that data sets don’t contain sensitive information before being exported from the platform. If the goal is to consolidate and scale these environments that could be a bottleneck, though any automation would have to earn trust in order to be implemented.

Most TREs do have ways to engage the public and measure the effectiveness of that engagement, though there aren’t clear best practices in this area. There is no information in the report about how these environments are used or the impact they have beyond a count of projects and users.

46% have fewer than 25 projects annually. 13% have more than 500. 44% have 11-50 users monthly. 13% have more than 500 users monthly.

From: p21,22 DARE UK 2025 landscape review

A couple of key emerging focuses were mentioned.

  1. Linking between datasets and environments
    1. When the data has been collected for different purposes and in different ways linking can be a huge and sometimes intractable challenge. For example, if data for a study is collected and stored with an anonymising identifier it may not retain a way to link that individual to other datasets. Even when there is a unique identifier, like NHS number, it can be wrong or missing in some cases.
    2. Datasets in different environments can also contain duplicate information that can be hard to identify. Also, one dataset might be partially missing data and linking to another can skew the picture.
  2. Enabling more capabilities for training AI
    1. AI means a lot of different things. People can train simple models with the tools that are commonly available today (like R or Python). But people are getting more interested in larger and more complex models that often require more computing resources than it would make sense to maintain for one research environment. To do this you need a way to securely send data to the large computing resources available in the UK.
    2. AI brings new considerations that need a policy before using the data in this way. For example, is there a possibility that the model could leak sensitive data? Full report.

Common themes

Trust

Trust is earned in spoonfuls and lost in buckets. Trust was a big theme throughout, and it was generally acknowledged that for HDRS to succeed it needed to earn the trust of the public. There have been two previous attempts at something similar that failed because of this. In 2013, Care.data tried to add GP data to existing Hospital Episode Statistics, but lost patient and GP trust and was shelved. That happened again in 2021 with GPDPR which, at this point, seems abandoned rather than postponed. In both cases, transparency and communication were lacking about who would have access and how the data would be used.

In the meantime, OpenSAFELY has managed to make data accessible from all GPs in England, and has done so uncontroversially. The unique thing about OpenSAFELY, as I understand it, is that it doesn’t move the raw data to a central store; instead it has a method to securely run analyses against all the individual GP databases and return the aggregated results. This is possible because GP data is ‘harmonised’ and there are just two systems to integrate with, which means two branches of the code that need to be maintained. Hospital data is not harmonised, so bringing that together involves a lot more work, which makes a federated solution harder.

The organisation Understanding Patient Data published a report in July 2026 on the state of public understanding and trust of healthcare data uses, which was presented by Anna Steere. The report found that people feel comfortable about health data use when they see evidence that data delivers value, is used responsibly, is subject to robust oversight, and that public voices genuinely influence decisions. A large majority of the population are supportive of using the data for research to inform better care, but that doesn’t mean blanket approval. Most people trust their GPs with their data (69%), but only 34% trust the government and 35% support collaboration with private tech companies. The British Medical Association has also stated that GPs should remain in control of their data in the context of the single patient record.

Impact

Dr Elizabeth Ford made the point, echoed by others, that she would like to see the conversation around TREs move from “Is it private?” to “Is it being used for public good?”, because the value is in the outcomes and not in simply having the data stores. There was also some discussion about how much should the use case(s) be defined prior to data collection. This is a tough one. For sure, the best data quality comes when you know exactly what you need it for and collect it for that purpose. But I don’t see that as entirely compatible with the efficiencies of having a consolidated, linked, dataset that can be used for many projects. In that case, you will be using it for purposes that you didn’t think about during collection. Moreover, someone else made the point that much of the current data is collected for the purpose of caring for the patient, which may not be exactly what you need for the research. This is one that can’t be solved perfectly, I think the best situation will be a foundation of standardised datasets that will mostly cover many research questions, with ongoing needs for more specialised or ad-hoc data collections for particular studies.

In terms of impact, providers (the people who are generating the data) should also get some benefit in terms of meeting their responsibilities, like using what we learn to help improve turnaround times for diagnostic and treatment pathways. Measuring the impact of TREs in a standardised way would be a valuable addition. Currently, the landscape report only covers the number of projects and users. Health Data Gateway does link datasets to the publications that used them, though I’m not clear that it is a complete list, and it isn’t in a reportable format. That means that right now, we don’t know what benefits these environments are driving.

Access

Access was compared to a leaky pipe; there are many environments, many and varied datasets going in, but just drops being accessed by researchers on the other side. Getting access can take months or years, and that is a dealbreaker for a lot of projects since many studies in academia are only funded for a few years. Often a project will need data from multiple sources to build a longitudinal picture, but there are different processes to access each database and it can be hard to know what data is available in each and hard to impossible to be able to join them up. Services like Health Data Gateway try to help by giving people one place to search through what is available and how to apply for access, and that is also one of the first projects that HDRS is working on. The phrase “single front door” was used a lot to describe where we should get to.

Current regulation, like GDPR, can also make using and sharing data complicated, especially with private companies. Industry wants to invest in countries with good data sources, but they need to know what is available and not have too many barriers. This of course needs to be balanced against public trust.

Vision for HDRS

Dr Melanie Ivarsson, the CEO of HDRS, shared her vision for the programme, which was set up following the Sudlow review and recommendations on uniting the UK’s health data. The first year primarily involves planning the strategy, while continuing to maintain existing infrastructure. HDRS is primarily focused on four areas: discovery, clinical trials, innovation development, and evaluation and evidence generation.

While surveying what is out there and planning on how to execute the vision in the Sudlow report, they are funding the SDEs and running about 40 projects to address the current challenges. For example, by unifying the access process for the SDEs, deploying Scotland’s solution for images in England, and linking together pathology, genetic and health data. She also mentioned that they are looking at ways to connect the data without centralising it.

Concluding thoughts

It is clear to me that this is not a technical challenge, it is a political one. There is existing technology and experience that can support a variety of solutions. The crux is how to design it so that it gains the public’s trust, it is used for good, there is transparency, and also it brings the benefits of investment from private companies. That last one does seem to be a key motivating factor for the project.

Although all the language in the reports that have been commissioned is about helping patients and helping the NHS, the funding and project are part of the Life Sciences Sector Plan, a plan to support life sciences companies with the goal of attracting investment and innovation in the UK. Apart from the data initiative, the plan has goals to streamline regulation and remove overhead for clinical trials and bringing innovations to market. Leaving the specifics of how private companies can access the data ambiguous was a key factor in losing support for this type of effort the previous two times. The model for access is really going to matter here.

HDRS’s full strategy should be published by the end of this year.