Opinion

Health-data sharing across provinces is the cornerstone of the government’s new AI strategy, but will our privacy really be protected?

Health-data sharing across provinces is the cornerstone of the government’s new AI strategy, but will our privacy really be protected?

The federal government has made health-data sharing across provinces a cornerstone of Canada's new AI strategy, with promises it will be “privacy preserving.” The premise is that Canada can generate economic benefit from health-data sharing while still safeguarding Canadians' personal information.

The claim is that with health-data sharing, Canada can strengthen research, attract investment, usher in AI abundance, and improve patient care. The argument is compelling. Canada's health-care systems generate some of the world's richest health-data resources. Every laboratory result, prescription, scan, and hospital visit creates information that can improve patient care.

But before we rush to connect more data, we should ask: do the approaches described as "privacy preserving" actually preserve privacy?

Much of today's health-data sharing discussion focuses on fragmentation or data silos. Health records remain scattered across hospitals, provinces, and incompatible information systems. Remove those barriers and innovation will follow.

Fragmentation matters. But it is not our biggest problem.

Canada has never built the institutions or standards needed to govern AI-enabled health-data sharing. Connecting records is only part of the challenge. We must also ensure that privacy is genuinely protected.

Preserving privacy is harder than it appears.

More than 25 years ago, computer scientist Latanya Sweeney demonstrated that removing names from a dataset did not make it anonymous. By combining only a few pieces of publicly available information like city, sex, and date of birth, she showed that half of the individuals could be re-identified. Since then, commercial databases have exploded, data brokers have proliferated, and data breaches are all too common.

The recent breach of Alberta's provincial voter registry illustrates why Sweeney's work matters today. Nearly three million Albertans had their names, addresses, and contact information exposed. The breach contained no health information. But that is precisely the point.

Every major data breach creates more information that can be linked against datasets once considered anonymous. A health dataset that appeared to be safely de-identified and shared five years ago may no longer be anonymous today. This means that de-identified health-care data shared for economic purposes may later be re-identified.

Privacy is not a property of the dataset alone, however. It is a property of the surrounding information environment. As that environment changes, so too does privacy risk. AI is a significant part of the changing environment.

Canada's AI strategy touts an approach called “federated AI” or “federated analytics” as their way to protect privacy. These methods eliminate the need to move patient data between institutions, but they also shift the privacy problem from the dataset to the model.

Research has shown that AI model parameters can reveal information about their training data, including entire reconstructions of what data they were trained on. Anonymizing the dataset does not necessarily protect the model from re-identification attacks, particularly as publicly available information continues to expand through commercial and political data breaches.

Risk does not mean we don’t advance with health-data sharing. Connecting health data is an important and critical step toward improving the health of Canadians. 

Our research with cancer patients across Canada found strong support for sharing health data to improve care, provided there are protections around unintended data use and reidentification.

Patients already contribute data to clinical trials, public health surveillance, and health services research. These activities operate within systems of law, consent, ethics review and public accountability. AI-enabled health data sharing should meet the same standard.

This is why Canada's recent investment of more than $100-million in the VITAL platform deserves careful scrutiny. Although VITAL keeps health data within each province by sharing AI models rather than patient records, privacy risks remain.

Before provincial governments participate in platforms like VITAL, they should require independent testing against privacy attacks and ongoing reassessment as new methods of re-identification emerge.

Terms such as "secure," "sovereign," and "privacy-preserving" should describe demonstrated technical properties, not simply policy aspirations. Public trust should rest on evidence, not assumptions.

Privacy is not only an ethical or policy question. It is also a scientific and regulatory one.

Like in most areas of health care, new technologies should be evaluated on evidence, not assumptions. If Canada wants to become a global leader in health AI, it should also become a global leader in demonstrating that "privacy-preserving" AI actually preserves privacy.

Dean Regier is director of the faculty of medicine’s Academy of Translational Medicine and an associate professor at the School of Population and Public Health at the University of British Columbia. Bret Nestor is a research associate at UBC’s Centre for Health Services and Policy Research, and School of Population and Public Health.

The Hill Times