71
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this post was submitted on 19 Nov 2024
71 points (100.0% liked)
Technology
37702 readers
406 users here now
A nice place to discuss rumors, happenings, innovations, and challenges in the technology sphere. We also welcome discussions on the intersections of technology and society. If it’s technological news or discussion of technology, it probably belongs here.
Remember the overriding ethos on Beehaw: Be(e) Nice. Each user you encounter here is a person, and should be treated with kindness (even if they’re wrong, or use a Linux distro you don’t like). Personal attacks will not be tolerated.
Subcommunities on Beehaw:
This community's icon was made by Aaron Schneider, under the CC-BY-NC-SA 4.0 license.
founded 2 years ago
MODERATORS
The article includes a link to the rest of the chat before that line: https://gemini.google.com/share/6d141b742a13
The message immediately preceeding:
To which it responded:
My only guesses for what happened:
Or maybe that Google engineer was right when he said that one of their AI chatbots is sentient
This is just a standard prompt hack. This will always exist with llms. They don't have any real understanding of language so safety protocols can't actually ban topics, only sets of words and phrases.
There was an extensive set of prompts working toward elder abuse before the result in question.
My guess is that the redditor who discovered it disguised it to look like homework and reproduced the hack, and added the "brother" to create more authentic rage bait.
This feels to me like the LLM misinterpreted it as some kind of fictional villain talk and started to autocomplete it.
Could also be the model simply breaking. There was a time when Sydney (Bing AI or whatever they call it now) had to be constrained to 10 messages per context and having some sort of supervisor on top of itself because it would occasionally throw a fit or start threatening the user for no reason.
This is probably just a regurgitated comment scraped from somewhere on reddit.
Twitter is another possibility. The LLM could have learned how to write like a bubbling barrel of radioactive toxic waste, and then just applied those lessons in longer format.
The preceding message is really quite an undefined input, as the user copy/pasted some questions from their assignment without phrasing it as a question or cleaning up the formatting.
I wonder what kind of outputs you would get from LLMs if you'd been talking sensibly on certain subjects then started to feed it garbage input. It feels like this might be what happened here.