urllib是python的一個(gè)獲取url(Uniform Resource Locators,統(tǒng)一資源定址符)了,可以利用它來抓取遠(yuǎn)程的數(shù)據(jù)進(jìn)行保存,本文整理了一些關(guān)于urllib使用中的一些關(guān)于header,代理,超時(shí),認(rèn)證,異常處理處理方法。
1.基本方法
urllib.request.urlopen(url, data=None, [timeout, ]*, cafile=None, capath=None, cadefault=False, context=None)
直接用urllib.request模塊的urlopen()獲取頁面,page的數(shù)據(jù)格式為bytes類型,需要decode()解碼,轉(zhuǎn)換成str類型。
from urllib import requestresponse = request.urlopen(r'http://python.org/') # <http.client.HTTPResponse object at 0x00000000048BC908> HTTPResponse類型page = response.read()page = page.decode('utf-8')urlopen返回對象提供方法:
1、簡單讀取網(wǎng)頁信息
import urllib.request response = urllib.request.urlopen('http://python.org/') html = response.read() 2、使用request
urllib.request.Request(url, data=None, headers={}, method=None)
使用request()來包裝請求,再通過urlopen()獲取頁面。
import urllib.request req = urllib.request.Request('http://python.org/') response = urllib.request.urlopen(req) the_page = response.read() 3、發(fā)送數(shù)據(jù),以登錄知乎為例
''''' Created on 2016年5月31日 @author: gionee ''' import gzip import re import urllib.request import urllib.parse import http.cookiejar def ungzip(data): try: print("嘗試解壓縮...") data = gzip.decompress(data) print("解壓完畢") except: print("未經(jīng)壓縮,無需解壓") return data def getXSRF(data): cer = re.compile('name=/"_xsrf/" value=/"(.*)/"',flags = 0) strlist = cer.findall(data) return strlist[0] def getOpener(head): # cookies 處理 cj = http.cookiejar.CookieJar() pro = urllib.request.HTTPCookieProcessor(cj) opener = urllib.request.build_opener(pro) header = [] for key,value in head.items(): elem = (key,value) header.append(elem) opener.addheaders = header return opener # header信息可以通過firebug獲得 header = { 'Connection': 'Keep-Alive', 'Accept': 'text/html, application/xhtml+xml, */*', 'Accept-Language': 'en-US,en;q=0.8,zh-Hans-CN;q=0.5,zh-Hans;q=0.3', 'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64; rv:46.0) Gecko/20100101 Firefox/46.0', 'Accept-Encoding': 'gzip, deflate', 'Host': 'www.zhihu.com', 'DNT': '1' } url = 'http://www.zhihu.com/' opener = getOpener(header) op = opener.open(url) data = op.read() data = ungzip(data) _xsrf = getXSRF(data.decode()) url += "login/email" email = "登錄賬號" password = "登錄密碼" postDict = { '_xsrf': _xsrf, 'email': email, 'password': password, 'rememberme': 'y' } postData = urllib.parse.urlencode(postDict).encode() op = opener.open(url,postData) data = op.read() data = ungzip(data) print(data.decode())
新聞熱點(diǎn)
疑難解答
圖片精選