TechTakes

2096 readers

96 users here now

Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

founded 2 years ago

MODERATORS

dgerard@awful.systems

Facebook Pushes Its Llama 4 AI Model to the Right, Wants to Present “Both Sides” [404 Media] (www.404media.co)

submitted 3 months ago by BlueMonday1984@awful.systems to c/techtakes@awful.systems

9 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] corbin@awful.systems 5 points 3 months ago

It's well-known folklore that reinforcement learning with human feedback (RLHF), the standard post-training paradigm, reduces "alignment," the degree to which a pre-trained model has learned features of reality as it actually exists. Quoting from the abstract of the 2024 paper, Mitigating the Alignment Tax of RLHF (alternate link):

LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax.