The classic trap
Recital 68 does not impose a direct obligation, but it shapes the interpretation of Article 10 of the AI Act on training data governance. In practice, providers of high-risk AI systems get cornered: they must demonstrate that their datasets are relevant, representative and high-quality, yet they often lack access to European common data spaces (notably EHDS for health) which are still being deployed. The classic trap is to train a model on a scraped dataset, with no provenance traceability, and then be unable to prove the legality of the collection before the EU AI Office or the CNPD for the personal data dimension.
How to read this recital in practice
Recital 68 sends three operational signals to Luxembourg providers:
- Prioritise European data spaces (EHDS health, mobility space, finance space) over opaque non-EU datasets, since these spaces provide a solid legal basis.
- Document the provenance chain of every dataset (source, legal basis, consent, anonymisation) because Article 10(2) requires this traceability.
- Register now with European Digital Innovation Hubs (EDIH Luxembourg run by Luxinnovation) and Testing and Experimentation Facilities (TEF) to benefit from validated sectoral datasets.
- For healthcare, anticipate the rollout of EHDS which will become the preferred path to access medical data for AI training, in coordination with the CNPD and Agence eSante.
The quality test nobody runs properly
Most providers confuse volume and quality. A dataset of 10 million radiology images scraped from the internet has no legal value if you cannot prove: the legal basis for collection, demographic representativeness (bias), the absence of leaked identifying personal data, and adequacy to the use case. This is exactly what Recital 68 aims to fix by steering providers toward governed data spaces.
How Luxgap automates this risk
Our Luxgap Dataset Provenance Tracker turns your training datasets into evidence that holds up before the EU AI Office and the CNPD. The tool plugs into your MLOps pipelines (MLflow, Databricks, AWS SageMaker, Azure ML, Hugging Face Hub) and automatically tracks, for every dataset used, its provenance, legal basis, representativeness score and Article 10 AI Act compliance, without ever asking your data scientists to fill in a form.
- Automatically detects every new dataset injected into a training pipeline, from its first use, via native MLflow and SageMaker hooks.
- Classifies provenance (European data space, licensed public dataset, client data, scraping) and computes a legal risk score against Recital 68 and Article 10.
- Checks eligibility for EHDS, the finance space or the mobility space, and alerts when a governed alternative dataset is available for the same use case.
- Detects residual personal data through automated PII scanning and computes a demographic bias score on sensitive attributes.
- Generates the Article 10(2) data governance sheet ready to embed into the Annex IV technical documentation of the AI Act.
- Produces a time-stamped, cryptographically sealed PDF report, admissible before the EU AI Office and the CNPD during a compliance audit.
Available as a complement to a Luxgap DPO or CISO mandate or as a dedicated SaaS module depending on your scope. Request a tailored quote and our teams will run a demonstration on your real pipelines, with a free 48-hour blind audit to measure the exposure of your current datasets before any commitment.