The classic trap
Recital 74 sheds light on Article 15 covering accuracy, robustness and cybersecurity of high-risk AI systems. In practice, providers declare flattering performance metrics (95% accuracy, 0.92 F1-score) measured on idealised test sets, then the system silently degrades in production. The EU AI Office expects metrics declared in the instructions for use, measurable, and stable throughout the entire lifecycle. The CNPD, on the personal data side, already sanctions misleading statements about the performance of scoring or profiling algorithms.
The 'consistent performance throughout the lifecycle' test
The recital imposes continuous metrology logic. Concretely, you must be able to demonstrate:
- Declared accuracy metrics (accuracy, precision, recall, F1, AUC) with measurement methodology and reference dataset.
- Robustness thresholds against adversarial inputs, noisy data and distribution drift.
- The cybersecurity level achieved with regard to the generally acknowledged state of the art (NIST AI RMF, ISO/IEC 42001, ENISA AI Threat Landscape).
- A post-market monitoring system that alerts as soon as real metrics diverge from declared metrics.
- Instructions for use written for non-technical deployers, without misleading marketing language.
How Luxgap automates this risk
Our Luxgap AI Performance Sentinel turns the metrics declared in the instructions for use into an enforceable performance contract, continuously measured against your production AI. The tool connects an observability agent to your inference endpoints (MLflow, Azure ML, AWS SageMaker, Vertex AI, Databricks, self-hosted models on LuxConnect or eBRC) and recomputes accuracy, robustness and drift indicators in real time against the thresholds declared to the provider or deployer.
- Continuously recomputes accuracy metrics (accuracy, precision, recall, F1, AUC) on your real inference flows and compares them with values declared in the instructions for use.
- Detects input and output distribution drift through statistical tests (KS, PSI, Wasserstein) with Teams or Slack alerts as soon as a degradation threshold is crossed.
- Runs adversarial attack batteries (FGSM, PGD, prompt injection for LLMs) to measure robustness according to NIST AI RMF and ENISA benchmarks.
- Generates an instructions-for-use document compliant with Article 13 and Recital 74, written for non-technical deployers, with declared performance indicators free of misleading statements.
- Produces a timestamped, cryptographically sealed PDF report, enforceable before the EU AI Office during an inspection, demonstrating consistent system quality throughout its lifecycle.
Available as part of a Luxgap DPO or CISO mandate or as a dedicated SaaS module depending on your scope. Request a tailored quote and our teams will prepare a demonstration on your actual models, with a free 48h blank audit to measure the gap between your declared metrics and observed performance.