Can't figure out how to download an embedded PDF (lemmy.ml)

submitted 1 year ago* (last edited 1 year ago) by nooneescapesthelaw@lemmy.ml to c/piracy@lemmy.dbzer0.com

31 comments fedilink hide all child comments

Ok so I want to download this embedded PDF/document in order to physically print it. The website allows me to view it as much as I want, but is asking me to fork over 25 + tax USD so i can download the document.

Obviously, i don't want to do that, so I try to download the embedded document via inspect element. But, the weird thing is it not actually loading a pdf, but like really small pictures of each page:

So, my question is basically how can I download this document in order to print it?

PdF link: https://www.sbcaplanroom.com/jobs/2477/plans/goleta-sanitary-district-biosolids-and-energy-phase-1-project/?preview=200647

you are viewing a single comment's thread
view the rest of the comments

[-] princessnorah@lemmy.blahaj.zone 10 points 1 year ago

Okay so, PDF documents are actually already “a collection of images” basically. This website is clearly trying to make it an extra step harder by loading the images individually as you browse the document. You could manually save/download all the images and use a tool to turn it back into a pdf. I haven’t heard of a tool that does this automatically, but it should be possible for a web scraper to make the GET requests sequentially then stitch the pdf back together.

[-] Pantoffel@feddit.de 5 points 1 year ago

I would go this route as well. As a developer this sounds easy enough. It you don't get vertical sequences of images, but instead a grid of images, then I would apply traditional image stitching techniques. There are tons of libraries for that on github.

load more comments (1 replies)

this post was submitted on 04 Nov 2023

65 points (92.2% liked)