← Writing

▶️ On the Record: Inaugural Lecture by Professor Seb Bacon (2026)

When I watched this talk I felt like Prof Seb Bacon had already figured out all the things I wanted to think about over the next 5-10 years.

This is a wonderful synthesis of ideas about transparency, defaults and systems of accountability when it comes to data and research.

Sometimes it feels like there is a tension between data privacy and healthcare research; that the benefits of one directly mean compromise on the other, and you have to choose what trade-offs you are willing to make. Perhaps it feels like that because it is like that in our default way of operating. Once your data is shared you don’t always have control or oversight over how it may be used in the future. Prof Bacon gave the example of the Havasupai Tribe in Arizona who donated DNA samples in the 90s for the purpose of investigating type 2 diabetes, which had high rates in the Tribe. In 2003, they found out that their DNA had been used for other purposes that they had not consented to, like research about inbreeding, alcoholism and migration patterns that conflicted with the Tribe’s own beliefs.

He mentions that privacy concerns aren’t necessarily about keeping secrets, many of us don’t feel that we have anything to hide, but rather about context and power. The problem comes when someone with power uses our data in a context that we haven’t consented to or takes the data out of context. This can lead to people refusing to share data for any purpose. That means it is essential to build trust and legitimacy through transparency and accountability.

He then describes how OpenSAFELY’s model of research is different. Researchers never see the raw data, just the output of their analysis, and every research decision is logged in a public audit trail. This creates a social contract for handling the data with care, and makes it easy for anyone to review how the data has been used, while the actual data is kept private.

Another interesting idea is that “hidden systems concentrate power and visible systems distribute power”. Prof Bacon makes the case that institutions never share more than is necessary. Openness invites scrutiny and criticism, so it isn’t in their interest. But without transparency, intentions can drift over time, and things that don’t serve them can be swept under the carpet. That is particularly relevant as private companies are building their own large healthcare datasets, often at larger scales than the public and non-profit offerings. For example, Epic (the company I used to work for), Oracle Health, TriNetX and Truveta all have databases with billions of data points from a few hundred million patients. They are promoted as a means to further scientific research, and I do not dispute the intention of the people involved. However, there is no public audit trail of how the data has been used, only the list of publications they have chosen to highlight. There are few public details about who has access and for what purpose, while some of these clearly sell the data to pharmaceutical companies. I know there is valuable research happening with these datasets, but good intentions and “just trust us” aren’t enough.

During my time at Epic I did research on the Cosmos dataset. The vision is to make insights available faster than traditional research, and all the people I worked with really were driven by the mission. In this case the data is explicitly not for sale. I do feel that the research it enables is overall a net positive, but do have some reservations about the inherent conflict of interest. We are only human, and however mission-driven a company is, there are always other motives at play as well: profit, winning, growth. I don’t have inside knowledge on how other companies operate, but the lack of public transparency is still evident. There isn’t a way for the public to know what analyses never got published. Methods and code for creating the dataset are not transparent, with intellectual property offered as the justification. Even with the best intentions this is a problem.

As an example, Epic has published several studies showing that patients who use MyChart (Epic’s patient portal) do better on some metric (fewer no-shows, more preventative screenings, shorter hospital stays). If the result had shown worse outcomes I’m almost certain it would not have been published with that narrative. Similarly, I doubt that any analysis that casts a negative light on one of Epic’s customers would be published. Those questions are likely not even asked. This gives companies a lot of power in shaping the narrative. Transparency distributes power because it allows us to assess the claims properly, and to sound the alarm before the damage is done.

Prof Bacon mentions that debates in Parliament were considered private and leaking them was illegal. But through pressure and legal mandates publishing them is now the default. Although I’d felt uncomfortable about the conflict of interest, this lecture explained the issue and possible solutions with much more clarity than I had reached so far. We changed the default in Parliament, and we can change the default when it comes to large private datasets.