Showing posts with label scraping. Show all posts
Showing posts with label scraping. Show all posts

Monday, July 08, 2024

Have You Been Scraped? Uncovering AI's Training Data

In the age of generative AIart and large language models, the question of concern for any creative artist (I am loath to call them "content creators" but social media does): 

Has our work been used to train AI without our knowledge or consent? A new tool offers some answers and a way to take action.


The website haveibeentrained.com allows users to search through vast, publicly researched AI training datasets like LAION 5B using text prompts. Curious about my own digital footprint, I decided to give it a try.

https://haveibeentrained.com/


Searching my name yielded numerous images from other Norm Hansons, but among them was a familiar face—my own. A self-portrait rock painting I'd posted long ago as a profile picture on Artists at Large had made its way into the dataset. While not overly distressed by this single instance, it did give me pause.

More concerning was the discovery that my charcoal sketch of Sir John Monash, created for an exhibition in 2018, had been scraped from my website. This unauthorized use of my work felt like a violation of my artistic rights.


Fortunately, the website offers a small measure of control. For individual images, users can tick a box that adds the image to a "Do Not Train" register, signaling to participating groups that you don't want your work included in future neural network training sets. For broader protection, entire domains can be registered.

It's worth noting that these actions are somewhat akin to closing the stable door after the horse has bolted. The data has already been used in training existing models. However, it's currently our best option for protecting our work moving forward.

This situation highlights a critical need for transparency and ethical behavior from those creating large language models, whether for legitimate research, commercial interests, or other purposes. As AI continues to evolve, so too must our understanding of its implications for creative rights and data privacy.

Have you checked if your work has been used in AI training datasets? Share your experiences and thoughts in the comments below.

Saturday, July 06, 2024

Is Image Glazing Worth the Hassle?

Protecting our art is becoming crucial in today's AI-driven digital world. But is image glazing the answer? Here's my experience so far.

Cara, a popular alternate platform to Instagram at the moment, has yet to offer built-in glazing. They suggest using Glaze, a separate software. Sounds simple. Not quite.

Setting up Glaze is a bit of a headache:

- It's free, 👍 but requires downloading large zip files 👎

- You need a hefty amount of disk space 👎

- The process feels outdated and tedious 👎

took ~88 minutes to Glaze
The real kicker? It's slow.👎👎👎 We're talking almost an hour and a half per image for my 9x5 submissions on older hardware. Ouch.

tool ~83 minutes to Glaze

You can batch them 👍 but it's one at a time.👎

But here's the biggest issue: you can't tell if it worked 🤞. The image looks the same, and there's no way to verify if it's actually protecting your style.🤞

So, is it worth it? Excuse me if I'm a little sceptical. The process is time-consuming, and we're essentially trusting a black box. 

Can it really save our unique mark-making styles from AI theft?

Tuesday, July 02, 2024

Protecting Creative Work in the Age of AI Scraping

As creatives in the digital age, we're facing a new challenge: how to protect our work from indiscriminate scraping by AI companies. While tools like Creative Commons licensing have been a go-to solution, their effectiveness against AI data collection is questionable.

Creative Commons: A False Sense of Security?

I've long relied on Creative Commons to share my work while maintaining some control. My license specifies attribution, non-commercial use, and (previously) share-alike terms. However, I'm beginning to question whether this offers real protection against AI scraping.

The Reality of AI Data Collection

Many companies, often hiding behind research organizations, are scraping vast amounts of online data to train AI models. This process often ignores licensing terms and lacks proper attribution or curation.


Changing Tactics

In response, I've updated my blog's license from "share-alike" to "no derivatives," hoping to prevent AI from copying my style. However, the legal landscape around this issue remains unclear, especially in Europe.

New Technological Defences

A promising development is the creation of tools that embed changes in image files. These alterations are invisible to humans but can disrupt AI training, potentially "poisoning" the dataset. Glaze and Nightshade are two such tools, though they're still in development and can be resource-intensive to use.

The Path Forward

Despite these efforts, I'm still uncertain about how to confidently share my work with those who behave ethically while protecting it from misuse. As creatives, we need to stay informed about these issues and continue seeking effective solutions to protect our work in the AI era.

What are your thoughts on protecting creative work in the age of AI? Have you found any effective strategies?


I've prepared this blog post with some good advice and a little rewording from Claude.AI