Trained On AI training-data disclosure

What providers actually disclosed

The Commission's template asks a fixed set of yes/no questions. These are the answers, taken from the filed documents themselves — which box the provider ticked, not what anyone said about it afterwards.

7 filed summaries · read 16 August 2026 · every cell links to its document

What the answers show

  • 7 of 7 report crawling the open web, and 7 report using publicly available datasets. On this evidence, crawling is universal.
  • On whether rightsholders were paid, the answers split: 1 say yes, 1 say no, and 5 tick “Other” — the template's escape hatch. Every OpenAI filing takes the escape hatch; xAI answers the question.
  • 3 of 7 say data from user interactions with the model was used in training. 6 say user data from the provider's other products was — a distinction the headline question alone would hide, and the reason both rows are shown separately.

A ticked box is what the provider chose to state on a regulated form. It is not an audit, and this register has no way to verify it against what a model was actually trained on.

Question in the template OpenAI ChatGPT Images 2.0 OpenAI GPT-5.2 OpenAI GPT-5.4 nano OpenAI GPT-5.5 OpenAI GPT-5.6 Luna xAI Grok 4.5 xAI Grok Voice Think Fast 2.0
Publicly available datasets Yes Yes Yes Yes Yes Yes Yes
Commercially licensed data Other Other Other Other Other Yes No
Crawled the open web Yes Yes Yes Yes Yes Yes Yes
Trained on user interactions with the model Yes No No No No Yes Yes
Trained on user data from other products Yes Yes Yes Yes Yes Yes No
Crawler named GPTBot Purposes of the crawlerGPTBot Purposes of the crawlerGPTBot Purposes of the crawlerGPTBot Purposes of the crawlerGPTBot Purposes of the crawlerxAI Web Crawler Purposes of the crawlerxAI Web Crawler 3 Purposes of the crawle
Summary last updated 30 July 202630 July 202631 July 202629 July 202623 July 20268 July 202629 July 2026
The document PDF PDF PDF PDF PDF PDF PDF

The sentence behind each answer

A tick mark is easy to misread, so the text it was read from is printed here. If any of it is wrong, the document is one click away and a correction is welcome.

OpenAI — ChatGPT Images 2.0
Publicly available datasets — Yes
Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
Commercially licensed data — Other
Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
Crawled the open web — Yes
Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
Trained on user interactions with the model — Yes
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☒ Yes ☐ No Was data collected from
Trained on user data from other products — Yes
Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide

Open the filed summary

OpenAI — GPT-5.2
Publicly available datasets — Yes
Have you used publicly available datasets to train the model? ☒ Yes ☐ No 2 If yes, specify the modality(ies) of the content covered by the d
Commercially licensed data — Other
Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
Crawled the open web — Yes
Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
Trained on user interactions with the model — No
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
Trained on user data from other products — Yes
Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide

Open the filed summary

OpenAI — GPT-5.4 nano
Publicly available datasets — Yes
Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
Commercially licensed data — Other
Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
Crawled the open web — Yes
Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
Trained on user interactions with the model — No
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
Trained on user data from other products — Yes
Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide

Open the filed summary

OpenAI — GPT-5.5
Publicly available datasets — Yes
Have you used publicly available datasets to train the model? ☒ Yes ☐ No 2 If yes, specify the modality(ies) of the content covered by the d
Commercially licensed data — Other
Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
Crawled the open web — Yes
Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
Trained on user interactions with the model — No
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
Trained on user data from other products — Yes
Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide

Open the filed summary

OpenAI — GPT-5.6 Luna
Publicly available datasets — Yes
Have you used publicly available datasets to train the model? ☒ Yes ☐ No 2 If yes, specify the modality(ies) of the content covered by the d
Commercially licensed data — Other
Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☐ No ☒ Other (s
Crawled the open web — Yes
Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): GPTBot Purposes of the crawler
Trained on user interactions with the model — No
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☐ Yes ☒ No Was data collected from
Trained on user data from other products — Yes
Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide

Open the filed summary

xAI — Grok 4.5
Publicly available datasets — Yes
Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
Commercially licensed data — Yes
Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☒ Yes ☐ No If yes, sp
Crawled the open web — Yes
Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): xAI Web Crawler Purposes of th
Trained on user interactions with the model — Yes
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☒ Yes ☐ No Was data collected from
Trained on user data from other products — Yes
Was data collected from user interactions with the provider’s other services or products used to train the model? ☒ Yes ☐ No If yes, provide

Open the filed summary

xAI — Grok Voice Think Fast 2.0
Publicly available datasets — Yes
Have you used publicly available datasets to train the model? ☒ Yes ☐ No If yes, specify the modality(ies) of the content covered by the dat
Commercially licensed data — No
Have you concluded transactional commercial licensing agreement(s) with rightsholder(s) or with their representatives? ☐ Yes ☒ No If yes, sp
Crawled the open web — Yes
Were crawlers used by the provider or on behalf of? ☒ Yes ☐ No If yes, specify crawler name(s)/identifier(s): xAI Web Crawler 3 Purposes of
Trained on user interactions with the model — Yes
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? ☒ Yes ☐ No Was data collected from
Trained on user data from other products — No
Was data collected from user interactions with the provider’s other services or products used to train the model? ☐ Yes ☒ No If yes, provide

Open the filed summary