What providers actually disclosed
The Commission's template asks a fixed set of yes/no questions. These are the answers, taken from the filed documents themselves — which box the provider ticked, not what anyone said about it afterwards.
7 filed summaries · read 16 August 2026 · every cell links to its document
What the answers show
- 7 of 7 report crawling the open web, and 7 report using publicly available datasets. On this evidence, crawling is universal.
- On whether rightsholders were paid, the answers split: 1 say yes, 1 say no, and 5 tick “Other” — the template's escape hatch. Every OpenAI filing takes the escape hatch; xAI answers the question.
- 3 of 7 say data from user interactions with the model was used in training. 6 say user data from the provider's other products was — a distinction the headline question alone would hide, and the reason both rows are shown separately.
A ticked box is what the provider chose to state on a regulated form. It is not an audit, and this register has no way to verify it against what a model was actually trained on.
| Question in the template | OpenAI ChatGPT Images 2.0 | OpenAI GPT-5.2 | OpenAI GPT-5.4 nano | OpenAI GPT-5.5 | OpenAI GPT-5.6 Luna | xAI Grok 4.5 | xAI Grok Voice Think Fast 2.0 |
|---|---|---|---|---|---|---|---|
| Publicly available datasets | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Commercially licensed data | Other | Other | Other | Other | Other | Yes | No |
| Crawled the open web | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Trained on user interactions with the model | Yes | No | No | No | No | Yes | Yes |
| Trained on user data from other products | Yes | Yes | Yes | Yes | Yes | Yes | No |
| Crawler named | GPTBot Purposes of the crawler | GPTBot Purposes of the crawler | GPTBot Purposes of the crawler | GPTBot Purposes of the crawler | GPTBot Purposes of the crawler | xAI Web Crawler Purposes of the crawler | xAI Web Crawler 3 Purposes of the crawle |
| Summary last updated | 30 July 2026 | 30 July 2026 | 31 July 2026 | 29 July 2026 | 23 July 2026 | 8 July 2026 | 29 July 2026 |
| The document |
The sentence behind each answer
A tick mark is easy to misread, so the text it was read from is printed here. If any of it is wrong, the document is one click away and a correction is welcome.
OpenAI — ChatGPT Images 2.0
- Publicly available datasets — Yes
- Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
- Commercially licensed data — Other
- Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
- Crawled the open web — Yes
- Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
- Trained on user interactions with the model — Yes
- Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☒ Yes ☐ No Was data collected from
- Trained on user data from other products — Yes
- Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide
OpenAI — GPT-5.2
- Publicly available datasets — Yes
- Have you used publicly available datasets to train the model? ☒ Yes ☐ No 2 If yes, specify the modality(ies) of the content covered by the d
- Commercially licensed data — Other
- Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
- Crawled the open web — Yes
- Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
- Trained on user interactions with the model — No
- Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
- Trained on user data from other products — Yes
- Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide
OpenAI — GPT-5.4 nano
- Publicly available datasets — Yes
- Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
- Commercially licensed data — Other
- Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
- Crawled the open web — Yes
- Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
- Trained on user interactions with the model — No
- Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
- Trained on user data from other products — Yes
- Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide
OpenAI — GPT-5.5
- Publicly available datasets — Yes
- Have you used publicly available datasets to train the model? ☒ Yes ☐ No 2 If yes, specify the modality(ies) of the content covered by the d
- Commercially licensed data — Other
- Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
- Crawled the open web — Yes
- Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
- Trained on user interactions with the model — No
- Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
- Trained on user data from other products — Yes
- Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide
OpenAI — GPT-5.6 Luna
- Publicly available datasets — Yes
- Have you used publicly available datasets to train the model? ☒ Yes ☐ No 2 If yes, specify the modality(ies) of the content covered by the d
- Commercially licensed data — Other
- Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
- Crawled the open web — Yes
- Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
- Trained on user interactions with the model — No
- Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
- Trained on user data from other products — Yes
- Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide
xAI — Grok 4.5
- Publicly available datasets — Yes
- Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
- Commercially licensed data — Yes
- Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☒ Yes ☐ No If yes, sp
- Crawled the open web — Yes
- Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): xAI Web Crawler Purposes of th
- Trained on user interactions with the model — Yes
- Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☒ Yes ☐ No Was data collected from
- Trained on user data from other products — Yes
- Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide
xAI — Grok Voice Think Fast 2.0
- Publicly available datasets — Yes
- Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
- Commercially licensed data — No
- Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☒ No If yes, sp
- Crawled the open web — Yes
- Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): xAI Web Crawler 3 Purposes of
- Trained on user interactions with the model — Yes
- Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☒ Yes ☐ No Was data collected from
- Trained on user data from other products — No
- Was data collected from user interactions with the provider’s other services or products used to train the model? ☐ Yes ☒ No If yes, provide