memes

13727 readers

2881 users here now

Community rules

1. Be civil

No trolling, bigotry or other insulting / annoying behaviour

2. No politics

This is non-politics community. For political memes please go to !politicalmemes@lemmy.world

3. No recent reposts

Check for reposts when posting a meme, you can only repost after 1 month

4. No bots

No bots without the express approval of the mods or the admins

5. No Spam/Ads

No advertisements or spam. This is an instance rule and the only way to live.

A collection of some classic Lemmy memes for your enjoyment

Sister communities

!tenforward@lemmy.world : Star Trek memes, chat and shitposts
!lemmyshitpost@lemmy.world : Lemmy Shitposts, anything and everything goes.
!linuxmemes@lemmy.world : Linux themed memes
!comicstrips@lemmy.world : for those who love comic stories.

founded 2 years ago

MODERATORS

Tenthrow@lemmy.world

The_Picard_Maneuver@lemmy.world

The_Picard_Maneuver@startrek.website

276

... (lemmynsfw.com)

submitted 2 months ago by samunder@lemmynsfw.com to c/memes@lemmy.world

13 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] brucethemoose@lemmy.world 22 points 2 months ago* (last edited 2 months ago) (4 children)

I know it's a meme, but the idea that transformers models 'remember' anything is a common misconception.

They have zero memory. When you submit a prompt, it feeds your entire chat history as one big prompt and... forgets it immediately, with no impact on the model itself. It's like its frozen in time, and copied, unfrozen, and thrown away every time it answers.

[–] UrPartnerInCrime@sh.itjust.works 6 points 2 months ago

Currently

[–] Zorque@lemmy.world 4 points 2 months ago

This has been a joke since before anything resembling the modern "AI" boom. Basically since murderous future AI was a think in popular media, at least since Terminator if not earlier. People would joke about treating their appliances kindly so that "Skynet" won't kill them in the future.

[–] kn33@lemmy.world 2 points 2 months ago (1 children)

Am I misunderstanding your comment or does it completely ignore context windows? Not that context windows are long-term, but it's not zero.

[–] brucethemoose@lemmy.world 5 points 2 months ago* (last edited 2 months ago)

The context window is indeed the LLM's memory.

...But its also muddy.

Many LLMs get 'dumber' and less attentive as their context windows grow, and OpenAI's models just happen to be one of these. It's awful close to the full 128K, even with the full GPT-4. Mistral models are also really bad at long context understanding while, conversely, I find that Google Gemini and Qwen 2.5 are really good close to their limits.

There are attempts to try and measure this performance objectively, like: https://github.com/NVIDIA/RULER

[–] samunder@lemmynsfw.com -3 points 2 months ago (1 children)

Yeah, yeah, let's see how Google will achieve more memory with their new Titan architecture

[–] brucethemoose@lemmy.world 5 points 2 months ago

It's still ephemeral, chats don't change the underlying language model, but yes it's interesting.