Showing posts with label stablediffusion. Show all posts
Showing posts with label stablediffusion. Show all posts

Sunday, April 09, 2023

The “A Picture is worth a 1000 Words” Project

This might be a bit long but it documents an interesting self-referencing project I envisioned to have Deep Learning AI create and reuse something in a self-referencing strange loop. I consider this a personal Art Project BUT it was interesting, and may I suggest worth reading. 
Does AI really know what it is doing?

 1. First up I gave chatGPT the following task.

It actually only wrote 499 words but they do seem intelligent and coherent enough (in a roundabout way). They are reproduced unedited below.

2. I took this text and duplicated it and literally added “two words.” then reproduced it on a single A4 sheet (I had to make it a landscape orientation and make the font tiny)

3.  I photographed it (to make it an image).

4. Then I submitted the same 1000-word quote as a prompt and a photograph of the text to Stable Diffusion as a reference/starter image. Not sure what it is trying to portray (if anything). Yet I can see how it is related to my 1000 word project.


The text from ChatGPT …..

The phrase "a picture is worth a thousand words" is a timeless adage that has been passed down through generations. It speaks to the idea that visual imagery can convey a wealth of meaning and emotion in a way that words cannot always capture. This quote has become so ubiquitous that it is often used in everyday conversation, advertising, and even in academic circles. However, the idea behind it is much deeper and more complex than the simple phrase suggests.

To truly understand the significance of this quote, we must first explore the relationship between words and images. Words are powerful tools for communication, but they have their limitations. They rely on the reader's ability to interpret them correctly, and they can be subject to different interpretations depending on the reader's perspective. Images, on the other hand, are more universally understood. They can communicate ideas and emotions in a way that is less dependent on interpretation and more dependent on the visual experience.

One of the main advantages of visual imagery is that it can capture complex emotions and ideas that are difficult to express in words. For example, an image of a sunset over a calm ocean can evoke a sense of peace and tranquility that is hard to capture in words. Similarly, an image of a child's face can convey a range of emotions from joy to sadness, without the need for a single word. These emotions and ideas can be conveyed in an instant through an image, while it would take much longer to express them through words.

Another advantage of visual imagery is that it can convey information quickly and efficiently. This is particularly true in the age of social media and the internet, where people are bombarded with information every day. A well-crafted image can capture a person's attention and convey a message much more effectively than a block of text. This is why social media platforms like Instagram and Pinterest have become so popular, as they allow people to share images and visual content in a way that is easily digestible and engaging.

Despite these advantages, it is important to remember that images and words are not mutually exclusive. In fact, they often work together to convey meaning and emotion. For example, a news article may use words to describe an event, but it is often accompanied by an image that helps to bring the story to life. Similarly, a marketing campaign may rely on both words and images to convey its message and connect with its audience.

In conclusion, the phrase "a picture is worth a thousand words" captures the idea that visual imagery can convey a wealth of meaning and emotion in a way that words cannot always capture. While words are powerful tools for communication, images have their advantages, particularly in their ability to capture complex emotions and ideas quickly and efficiently. Ultimately, it is the combination of words and images that can create a truly powerful and engaging message.

…..Actually only 499 words

I also submitted the prompt to Dall.E (although I couldn’t submit a reference/starter image). These don’t make a lot of sense to me. Although they are photorealistic?

And also a couple of other systems, similar to MidJourney, again not making a lot of sense.

But my favourite is still Style Transfer. I can see a little creative self-reference here, I’m just not sure it was created intentionally


All images, other than those created in Dall.E 4, where created at @nightcafe studio

PS: If you want to consider "worth" to be represented by bandwidth or file space required. Then an image will cost at least between 5 and 10 times that of the equivalent raw text, or even 100 to 300 times if you need something of higher resolution. SO a picture costs a lot more than 1000 words to move around the internet. 


Monday, January 02, 2023

If a picture is worth 1000 words…

What value will a picture have if it's based on 10, or maybe 100 words?

I’m a chronic dyslexic and words don’t flow easily for me. However, I had decided it was about time to update my Instagram profile picture (it still had me behind a COVID Mask). So why not try and find what Stable Diffusion (a leading text-to-image AI system) had to offer.

Prompt: "Profile artwork for instagram, using watercolour markers"  7 words

Plus a starter image, a head and shoulders photo of me. Hmmm, I’m not a girl and my eyes are roughly the same size. So it didn’t really take much notice of my picture perhaps I need to tell it I have grey hair and a beard

Prompt: "Portrait artist, gray hair goatee beard Profile artwork for instagram, using watercolour markers head and shoulders" 16 words

I add a few negatives  with a -0.3 weight to discourage the process going there

"ugly, scary, poorly drawn face, out of frame, cut off" 10 more words, 26 in total

A mighty improvement and without a starter photo this time BUT how are these an artist and blue rinse in my hair? Seriously?

The last image set was choosen with the same starter but used the additional modifiers (artistic portrait preset image). It added a lot of words some for a positive match some negative, for a total of 88 words

Prompt: "Portrait artist, gray hair goatee beard Profile artwork for instagram, using watercolour markers head and shoulders portrait, 8k resolution concept art portrait by Greg Rutkowski, Artgerm, WLOP, Alphonse Mucha dynamic lighting hyperdetailed intricately detailed Splash art trending on Artstation triadic colors Unreal Engine 5 volumetric lighting” 56words

"ugly, tiling, poorly drawn hands, poorly drawn feet, poorly drawn face, out of frame, extra limbs, disfigured, deformed, body out of frame, blurry, bad anatomy, blurred, watermark, grainy, signature, cut off, draft" 32 negative words

Well it’s a lot more realistically rendered, but still poorly cropped (head chopped up and at best a single shoulder, etc… etc and... Why is this an artist? So do I need to craft a longer more descriptive prompt, or just run a lot more prompts?

The images above have been created on Nightcafe using Stable Diffusion interface

It actually took less time to have a play with my ecoline watercolour brush pens and a pitt pen than to assemble the above.


So my new Instagram profile picture is handdrawn, I trust you understand why.


PS I must humbly apologise to @Greg Rutkowski (I hadn't noticed he was mentioned in the final prompt, hmm the problems with presets) and I do understand why he is not happy.

Wednesday, December 21, 2022

Opt-in & Opt-out Talk about #AIart and Have I Been Trained website

 Nothing brings the finer points of a debate into the spotlight than personal involvement.

The big AIart news of the last few days is that Stable AI will allow artist to opt out of being included in the next dataset being used to train the neural network to be used for Stable Diffusion 3. Well at least you can opt-out for the next couple of weeks, so follow this up now. 

There is a site, that can check whether you have been included in the massive Laion-5B & Laion-400m neural networks used in Stable Diffusion & Google's Imogen, yes they were trained on 5.8 billion images. I haven't established its legitimacy as an ethical stance, but it does seem legit. It doesn't appear to be a sneaky way to get more images (with so much hacking and phising establishing trust is a big issue in the #AIart discussions)



However, when you see you are one of the artist whose images have been used to train the Lioan Neural Network and had no idea they are in there, some things suddenly confront you.


It took me very little time using the image search feature, to find two of my works. They were part of my Retracing Darwin exhibition in early 2010 and very early examples of my personal technique I call photoimpression. I actually don't mind if others study my method and even create examples of their own. I would like them to acknowledge me, which is becoming a hollow wish on today's web. I don't blame redbubble either, I am sure they didn't know and/or had not given permission either,

I really don't want my work, especially my own special techniques, style, mark-making colouring or composition used without my permission. This is the stuff that makes my work original. Firstly because I know I'll never be acknowledged, that others could profit from this work or contribution to this work to what is presented as their original, I also actually find a lot of the so-called art generated by these AI's a bit scary and I don't approve, and finally I find the whole process a bit morally questionable and not ethical.

So I've made up my mind, Now I do want to opt out.

Now I have to find out how

Damn! I have to do it image by image. Cest La Vie

Sunday, October 09, 2022

Unnatural Crops Explained

The developments in Text to Image AIart generated images are amazing. Things are changing, and largely improving in image quality almost weekly. However one thing I had noticed that was staying fairly constant was unnatural looking crops, the weird truncating of the subjects, particularly people. Surely this was not a new artistic trend I had no knowledge of, or perhaps the artist who's work is being used to train these systems had an aversion to conventional composition.

Example of a headless figure based on stable diffusion prompt

Am I an artist now? John Singer Sargent


Turns out there is a simpler explanation (see quote from a NovelAI blog post below). The unnatural crops are a result of the training set being converted to a square format (so the images are the same ratio) and just arbitrarily using the center of the image.

 Aspect Ratio Bucketing

One common issue of existing image generation models is that they are very prone to producing images with unnatural crops. This is due to the fact that these models are trained to produce square images. However, most photos and artworks are not square. However, the model can only work on images of the same size at the same time, and during training, it is common practice to operate on multiple training samples at once to optimize the efficiency of the GPUs used. As a compromise, square images are chosen, and during training, only the center of each image is cropped out and then shown to the image generation model as a training example.

Monday, October 03, 2022

Are we having fun yet?

It hasn't taken long for my AIart works to fall foul of the censors. Ok not the real censors, the perceived NFW (not for work, standards of the USA based morals on the internet) Whilst I can only see blur in one of the images generated, I expect that it was being interpreted as NUDITY. Which isn't surprising for anything Salvador Dali inspired

I also believe the warning is probably coming from Stable Diffusion itself. That might be a very good thing, indicating they are concerned with the ethics (aka complying with social norms) of the types of images they are being asked to generate. I do know that certain words are not allowed in the text submitted. Again, a bit restrictive but overall, a good thing.

Back to what I was playing with, just having fun. Which I learnt from a Catherine Price in a recent Ted Talk, involves Play, Connection & the Flow State. I like self-referential (meta) subject so I have submitted the phrase "Are we having fun yet?" Firstly the phrase just by itself and the results were totally underwhelming. Perhaps AI doesn't know what fun is yet?

Then by 3 well known artists  / cartoonists / photographers each as modifiers. The results are more encouraging. I've picked the most appropriate in each group of four thumbnails.

Moral: Be careful what you ask for!