259
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this post was submitted on 11 Jan 2024
259 points (99.6% liked)
Technology
37702 readers
577 users here now
A nice place to discuss rumors, happenings, innovations, and challenges in the technology sphere. We also welcome discussions on the intersections of technology and society. If it’s technological news or discussion of technology, it probably belongs here.
Remember the overriding ethos on Beehaw: Be(e) Nice. Each user you encounter here is a person, and should be treated with kindness (even if they’re wrong, or use a Linux distro you don’t like). Personal attacks will not be tolerated.
Subcommunities on Beehaw:
This community's icon was made by Aaron Schneider, under the CC-BY-NC-SA 4.0 license.
founded 2 years ago
MODERATORS
Interestingly, LLMs seem to show emerging semiotic organization. By analyzing the activation space of the neural network, related concepts seem to get trained into similar activation patterns, which is what allows LLMs to zero shot relationships when executed at a "temperature" (randomness level) in the right range.
Pairing an LLM with a stable diffusion model, allows the resulting AI to... well, judge by yourself: https://llm-grounded-diffusion.github.io/
I'm unconvinced that the fact they're getting better at following instructions, like putting objects where the prompter specifies, or changing the colour, or putting the right number of them, etc means the model actually understands what the objects mean beyond their appearance. It doesn't understand the cultural meanings attached to each object, and thus is unable to truly make a decision about why it should place an apple rather than an orange, or how the message within the picture changes when it's a red sports car rather than a beige people-carrier.
Well, that's part of the LLM step, so let's put it to the test:
Image prompt:
Create an ad for a car. The target audience are young adults, with high income, and thrill seeking. Come up with a brand name fitting such a car, and include it in the image. For the car, come up with a shape, and color, best fitting of the target audience. Come up with an image layout, art style, and camera angle, best fitting of the target audience. Include between zero and two additional items that will make the overall picture more attractive to the target audience.
ChatGPT:
[('an advertisement for a car', [45, 58, 422, 396]), ('a high-performance car', [123, 193, 266, 128]), ('an edgy and dynamic brand name logo', [188, 10, 136, 30]), ('a sleek, red sports car', [168, 254, 176, 102])] Background prompt: An adrenaline-pumping car advertisement targeting young adults with high income and a thrill-seeking spirit. The layout includes a bold brand name logo, a sleek red sports car, and a dynamic composition to captivate the audience. Negative prompt: additional items
How did it know to pick a "sleek red sports car"? Or the rest of the elements.
Because some of its training data included some of the many, many websites out there that describe marketing techniques. However, your example has actually proved my point - the red sports car is a car for insecure middle-aged men needing a mid-life crisis penis extension. The LLM has entirely missed that cultural association, and has basically suggested a red sports car for a young audience, when an alternate colour would actually be more appropriate - because it doesn't actually understand what a red sports car means.
It also hasn't actually picked any distinctive elements that couldn't be found on a website offering generic marketing advice. "A dynamic composition" is obvious, but it hasn't specified any details about what the composition should look like. It hasn't detailed any of the surrounding scenery. It says you should include a brand name logo, which was obvious because you prompted it to come up with a brand name, but it's failed to detail what those should actually look like. The entirety of the elements it's created here is "sleek red sports car", which has a cultural connotation inappropriate to the target audience, and the rest you could literally get from any search for "how do I create an advert for a car?"