Technology

75436 readers

1812 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

L4s@hackingne.ws

429

Google Researchers’ Attack Prompts ChatGPT to Reveal Its Training Data (www.404media.co)

submitted 2 years ago by stopthatgirl7@kbin.social to c/technology@lemmy.world

96 comments fedilink hide all child comments

ChatGPT is full of sensitive private information and spits out verbatim text from CNN, Goodreads, WordPress blogs, fandom wikis, Terms of Service agreements, Stack Overflow source code, Wikipedia pages, news blogs, random internet comments, and much more.

you are viewing a single comment's thread
view the rest of the comments

[–] fubo@lemmy.world 7 points 2 years ago

It doesn’t have to have a copy of all copyrighted works it trained from in order to violate copyright law, just a single one.

Sure, which would create liability to that one work's copyright owner; not to every author. Each violation has to be independently shown: it's not enough to say "well, it recited Harry Potter so therefore it knows Star Wars too;" it has to be separately shown to recite Star Wars.

It's not surprising that some works can be recited; just as it's not surprising for a person to remember the full text of some poem they read in school. However, it would be very surprising if all works from the training data can be recited this way, just as it's surprising if someone remembers every poem they ever read.