Save full webpage

后端未结

关注

 6  378

I\'ve bumped into a problem while working at a project. I want to \"crawl\" certain websites of interest and save them as \"full web page\" including styles and images in or

相关标签:

6条回答

失恋的感觉

2020-12-06 04:04

You actually need to parse the html and all css files that are referenced, which is NOT easy. However a fast way to do it is to use an external tool like wget. After installing wget you could run from the command line wget --no-parent --timestamping --convert-links --page-requisites --no-directories --no-host-directories -erobots=off http://example.com/mypage.html

This will download the mypage.html and all linked css files, images and those images linked inside css. After installing wget on your system you could use php's system() function to control programmatically wget.

NOTE: You need at least wget 1.12 to properly save images that are references through css files.

0 讨论(0)
发布评论:

提交评论
- 加载中...
生来不讨喜

2020-12-06 04:11

Whatever app is going to do the work (your code, or code that you find) is going to have to do exactly that: download a page, parse it for references to external resources and links to other pages, and then download all of that stuff. That's how the web works.

But rather than doing the heavy lifting yourself, why not check out curl and wget? They're standard on most Unix-like OSes, and do pretty much exactly what you want. For that matter, your browser probably does, too, at least on a single page basis (though it'd also be harder to schedule that).

0 讨论(0)
发布评论:

提交评论
- 加载中...
-上瘾入骨i

2020-12-06 04:18

Is there a way to do this without read and save each and every link on the page?

Short answer: No.

Longer answer: if you want to save every page in a website, you're going to have to read every page in a website with something on some level.

It's probably worth looking into the Linux app wget, which may do something like what you want.

One word of warning - sites often have links out to other sites, which have links to other sites and so on. Make sure you put some kind of stop if different domain condition in your spider!

0 讨论(0)
发布评论:

提交评论
- 加载中...
無奈伤痛

2020-12-06 04:22

I'm not sure if you need a programming solution to 'crawl websites' or personally need to save websites for offline viewing, but if its the latter, there's a great app for Windows — Teleport Pro and SiteCrawler for Mac.

0 讨论(0)
发布评论:

提交评论
- 加载中...
长发绾君心

2020-12-06 04:28
If you prefer an Objective-C solution, you could use the WebArchive class from Webkit.
It provides a public API that allows you to store whole web pages as .webarchive file. (Like Safari does when you save a webpage).

Some nice features of the webarchive format:
- completely self-contained (incl. css, scripts, images)
- QuickLook support
- Easy to decompose
0 讨论(0)
发布评论:

提交评论
- 加载中...
陌清茗

2020-12-06 04:29

You can use IDM (internet downloader management) for downloading full webpages, there's also HTTrack.

0 讨论(0)
发布评论:

提交评论
- 加载中...