this post was submitted on 20 Mar 2025
503 points (99.6% liked)

Technology

67050 readers
6380 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] melpomenesclevage@lemmy.dbzer0.com 28 points 1 day ago* (last edited 1 day ago) (1 children)

i hear there's a tool called (I think) 'nepenthe' that creates a loop for an LLM, if you use that in combination with a fairly tight blacklist of IP's you're certain are LLM crawlers, I bet you could do a lot of damage, and maybe make them slow their shit down, or do this in a more reasonable way.

[–] PrivacyDingus@lemmy.world 7 points 1 day ago (1 children)

nepenthe

It's a Markov-chain-based text generator which could be difficult for people to implement on repos depending upon how they're hosting them. Regardless, any sensibly-built crawler will have rate limits. This means that although Nepenthe is an interesting thought exercise, it's only going to do anything to things knocked together by people who haven't thought about it, not the Big Big companies with the real resources who are likely having the biggest impact.

might hit a few times, or maybe there's a version that can puff stuff up the data in the sense of space, and salt it in the sense of utility.