
Making Data Sharing Work in the EU: Why Data Provenance and Machine Unlearning Matter
Without data provenance and machine unlearning, a single unlawful data point can force the invalidation of an entire dataset or AI model, which is a disproportionate outcome that jeopardises the European Data Strategy.
As the EU's Data Act and the European Health Data Space Regulation (EHDS) come into effect, datasets will increasingly combine information from multiple sources and actors, raising urgent questions about privacy, trade secrets, and intellectual property. Existing legislation offers little guidance on what happens when a dataset is found to contain unlawfully processed data — leaving the invalidation of entire datasets and AI models as the default response.
This policy brief shows how the interdependent mechanisms of Data Provenance and Machine Unlearning can support compliance with EU data governance obligations by providing guidance on how to identify and remove unlawfully processed data without the invialitaion of while preserving the utility of legitimate data resources.
In this policy brief, members of Workpackage 2 show how two technical mechanisms — data provenance and machine unlearning — can close a critical governance gap: how to identify and remove unlawfully processed data without invalidating entire datasets or AI models.
Drawing on empirical testing with health data, the authors demonstrate that both exact and approximate unlearning methods can operate effectively on complex, mixed data.
Recommendations:
EU and national regulators should promote provenance-by-design, requiring shared datasets to carry standardised metadata on data origin, legal basis, usage restrictions, and transformations.
Regulators should formally recognise machine unlearning as a legitimate and proportionate mechanism for complying with data deletion, withdrawal, and revocation obligations.
The European Commission should prioritise standardisation of provenance frameworks and unlearning mechanisms through targeted funding and coordination with standard-setting bodies.
Legislators should clarify liability rules in complex, mixed-data environments to prevent disproportionate invalidation of datasets and models.
Regulatory authorities should be equipped with specialised technical expertise to evaluate provenance records and unlearning claims.
by
Pratiksha Ashok, Inge Graef, Pradeep Kumar, Patricia Prüfer, Berkay Serceoglu, Teun Siebers
/
Read more