The classic trap
Recital 89 creates a dangerous grey zone: many integrators believe that using an open-source building block (Hugging Face, LangChain, scikit-learn, Apache 2.0 fine-tuned models) exempts them from AI Act obligations. This is wrong. The exemption targets the third party publishing the tool, not the provider integrating it into an AI system placed on the market. The EU AI Office and national AI market surveillance authorities will examine the real value chain first: who integrated what, with which upstream documentation, and what contractual or technical guarantees the provider obtained from the open-source component.
The concrete test: exempt third party or responsible provider?
- You publish a model on GitHub under MIT licence without commercialising it: exemption likely, but documentation (model card, data sheet) strongly recommended by the recital.
- You integrate Llama, Mistral or a Hugging Face model into a SaaS billed to your clients: you are a provider under article 3, the exemption does not apply to you.
- You redistribute an open-source component with substantial modifications: the initial provider's responsibility fades, yours begins.
- You use a poorly documented open-source component (no model card, no data sheet, no training data description): you inherit the documentation gap and will need to fill it yourself to meet downstream obligations.
The documentation practices encouraged by the recital
The legislator explicitly cites model cards (Hugging Face / Google format) and data sheets for datasets (Gebru et al. 2018) as de facto standards. Adopting these formats upstream accelerates your downstream compliance: annex IV (technical documentation), article 13 (transparency), article 53 (GPAI obligations).
How Luxgap automates this risk
Our Luxgap Open Source AI Lineage automatically reconstructs your organisation's AI value chain and qualifies, for each detected component, your legal status under the AI Act: exempt third party, integrating provider, or responsible redistributor. The tool scans your GitHub/GitLab repositories, Python environments (requirements.txt, poetry.lock), MLflow registries, Hugging Face Hub and Docker containers to materialise the real dependency, without developer self-declaration.
- Automatically detects every open-source model, library or pipeline integrated into your production AI systems via native GitHub, GitLab, Hugging Face, MLflow and Azure ML connectors.
- Qualifies each component's licence (Apache 2.0, MIT, GPL, RAIL, Llama Community) and flags licences incompatible with commercial use or with the recital 89 exemption.
- Automatically retrieves available upstream model cards and data sheets and identifies documentation gaps you will need to fill to comply with annex IV.
- Generates missing model cards and data sheets in standard Hugging Face format, pre-filled from your pipelines' technical metadata.
- Alerts in real time when an upstream component changes licence or disappears (Llama 2 to Llama 3 Community Licence case) to anticipate legal requalification.
- Produces a timestamped PDF report enforceable before the AI surveillance authority, demonstrating the complete value chain mapping and your status qualification.
Available as a complement to a Luxgap DPO or CISO mandate or as a dedicated SaaS module depending on your perimeter. Request a tailored quote and our teams prepare a demonstration on your real AI stack, with a free 48-hour blind audit to map your open-source components and measure your exposure before any commitment.