> For the complete documentation index, see [llms.txt](https://www.headlesslaw.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://www.headlesslaw.com/ai-act/articles/10.md).

# Article 10 — Data and data governance

*In force · Consolidated version of 27 July 2026 · Checked against EUR-Lex on 30 Sep 2026 ·* [*Official source*](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02024R1689-20260727)

1. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and [testing data](https://www.headlesslaw.com/definitions/ai-act/testing-data) sets that meet the quality criteria referred to in paragraphs 2, 3 and 4 of this Article and in [Article 4a(1)](/ai-act/articles/4a.md) whenever such data sets are used.
2. Training, validation and testing data sets shall be [subject](https://www.headlesslaw.com/definitions/ai-act/subject) to data governance and management practices appropriate for the [intended purpose](https://www.headlesslaw.com/definitions/ai-act/intended-purpose) of the high-risk [AI system](https://www.headlesslaw.com/definitions/ai-act/ai-system). Those practices shall concern in particular:

   **(a)** the relevant design choices;

   **(b)** data collection processes and the origin of data, and in the case of [personal data](https://www.headlesslaw.com/definitions/ai-act/personal-data), the original purpose of the data collection;

   **(c)** relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation;

   **(d)** the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent;

   **(e)** an assessment of the availability, quantity and suitability of the data sets that are needed;

   **(f)** examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations;

   **(g)** appropriate measures to detect, prevent and mitigate possible biases identified according to point (f);

   **(h)** the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed.
3. Training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used. Those characteristics of the data sets may be met at the level of individual data sets or at the level of a combination thereof.
4. Data sets shall take into account, to the extent required by the intended purpose, the characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting within which the high-risk AI system is intended to be used.
5. \[Deleted by [Regulation (EU) 2026/1744](https://eur-lex.europa.eu/eli/reg/2026/1744/oj) of 8 July 2026. See [Article 4a — Processing of special categories of personal data for bias detection and correction](/ai-act/articles/4a.md).]
6. For the development of high-risk AI systems not using techniques involving the training of AI models, paragraphs 2, 3 and 4 of this Article and [Article 4a(1)](/ai-act/articles/4a.md) shall apply only to the testing data sets.

***

[Chapter III — High-risk AI systems](/ai-act/chapters/iii.md) · [← Article 9](/ai-act/articles/9.md) · [Article 11 →](/ai-act/articles/11.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://www.headlesslaw.com/ai-act/articles/10.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
