HOME ABOUT NEWS RESEARCHEVENTS CONTACT
DATASET DIALOGUES
Interactive Website
Kate Stonehill
Dataset Dialogues is a dialogue between artists and generative AI. The project takes the artists associated with Sonic Screen and asks a question: how have they – or their work – been incorporated as training data for Generative AI? It then invites the artists to reflect on becoming data for generative AI imagery, bringing these reflections back into dialogue with the generative AI systems themselves.
Modern generative AI models require vast amounts of training data, and have been trained using significant amounts of copyrighted material. Despite calls for transparency, the training data used by AI developers remains largely hidden from public view. However, as Kate Crawford argues, the ongoing integration of more facets of our lives with AI systems brings renewed urgency to the question of understanding datasets: “The new internet-scale datasets require new investigative methods, new research questions. What political and cultural inflections are baked into training sets? Who and what is represented? What is rendered invisible and unintelligible? Who profits from all this data, and at whose expense?”
One particular dataset – LAION-5B – is a rare exception to the secrecy surrounding datasets used by the AI frontier labs. LAION-5B is an open-source dataset that was released in 2022 by a German non-profit, LAION. It contains approximately five billion image-text pairs, which in turn are used to help AI systems build a lexicon of the visual world, as they learn the association between a word and its representative image.
LAION-5B was initially created for research purposes; its authors, in fact, argued that it should not be used for ‘industrial products’, though this warning has been ignored. The goal of LAION-5B was “to conduct basic research into dataset curation. Specifically, its authors wanted to create an image training set with purely automatic methods – with no humans in the mix”. LAION-5B was built using another dataset – Common Crawl. Common Crawl is a scraping of web data that happens every month – essentially, a snapshot of the Internet, frozen in time, increasingly used to feed hungry AI models.
To ascertain how artists show up in the dataset, I conducted searches of LAION-5B for artist names. It is important to note that searching for specific artworks would have surfaced different results. However, given that many artworks have fairly generic names – and the sheer quantities of data (5 billion images) involved – this would have been a far more unwieldy task. In fact, even having a somewhat more generic name proved to be disadvantageous in unearthing genuine search results.
Art and artists become data, but what about the artists themselves? This project was motivated by a desire to close a generative AI loop, and re-insert humans into a machine-led process that has all but erased them. In interviews, the artists reflect on the images, offering stories and context for the images in question. Their reflections reveal things that the machines learning from them could never have known.
Finally, this project also includes an interactive chatbot. Curious visitors are invited to converse with a chatbot that personifies the dataset in question, LAION-5B. This element uses generative AI to build knowledge about the dataset itself.
One particular dataset – LAION-5B – is a rare exception to the secrecy surrounding datasets used by the AI frontier labs. LAION-5B is an open-source dataset that was released in 2022 by a German non-profit, LAION. It contains approximately five billion image-text pairs, which in turn are used to help AI systems build a lexicon of the visual world, as they learn the association between a word and its representative image.
LAION-5B was initially created for research purposes; its authors, in fact, argued that it should not be used for ‘industrial products’, though this warning has been ignored. The goal of LAION-5B was “to conduct basic research into dataset curation. Specifically, its authors wanted to create an image training set with purely automatic methods – with no humans in the mix”. LAION-5B was built using another dataset – Common Crawl. Common Crawl is a scraping of web data that happens every month – essentially, a snapshot of the Internet, frozen in time, increasingly used to feed hungry AI models.
To ascertain how artists show up in the dataset, I conducted searches of LAION-5B for artist names. It is important to note that searching for specific artworks would have surfaced different results. However, given that many artworks have fairly generic names – and the sheer quantities of data (5 billion images) involved – this would have been a far more unwieldy task. In fact, even having a somewhat more generic name proved to be disadvantageous in unearthing genuine search results.
Art and artists become data, but what about the artists themselves? This project was motivated by a desire to close a generative AI loop, and re-insert humans into a machine-led process that has all but erased them. In interviews, the artists reflect on the images, offering stories and context for the images in question. Their reflections reveal things that the machines learning from them could never have known.
Finally, this project also includes an interactive chatbot. Curious visitors are invited to converse with a chatbot that personifies the dataset in question, LAION-5B. This element uses generative AI to build knowledge about the dataset itself.