PDF下载地址:爬虫PDF
一、爬虫简介
1.模块准备
requests
lxml
selenium
scrapy
命令:
pip install 模块名
2.爬虫的分类
-
全网爬虫
百度、搜狗、Google、bing
-
聚焦爬虫
针对单一网站的搜索工具
3.爬虫的注意点
-
变化非常迅速
今天代码没问题—>明天代码不行了
- 时间问题
- 身份问题
- 网站变化
- 网站更新、升级
- 网站崩溃、倒闭
-
爬虫爬取数据的来源
网址
- 网址的构成(数据来源)
- 协议
- 域名
- 资源目录
- 搜索关键字
- 网址的构成(数据来源)
-
爬虫的流程
- 分析网址
- 构建爬虫
- 数据清洗
- 数据保存
-
爬虫代码问题
一个代码只适用于一个网站
-
爬虫只能爬取可以获取的数据
-
学习爬虫的过程中,不要过多的发起请求
4.爬虫的作用
获取数据
- 文字
- 图片
- 视频
- 音频
为了批量的获取数据、可以进行自动化的操作、快速的获取
5.反爬虫
学习如何识别对应网站的反爬策略,从而找到可以通过反爬策略的方法
二、抓包分析
1.数据来源
数据包的形式进行数据的传输
客户端 —- 发起请求[请求数据包] —> 服务端
客户端 <— 返回数据[响应数据包] —- 服务端
2.什么是抓包?
在客户端与服务器交互的时候,对其相互传输的数据包进行解析从而获取需要的内容
补充:每个网站都是独立的存在,因为开发人员的不同,不同的网站所实行的方法是不一样的
3.如何抓包(开发者选项、调试控制台)
需要开启抓包后,进行请求操作,才会有数据包被抓到
请求方法
- GET
- POST
4.如何找到需要的资源包
方法一:过滤
Doc中的资源包
方法二:搜索(在空白处按Ctrl+ F )
寻找数据,构建爬虫的步骤
分析网站 —> 寻找目标数据包 —> 构造请求
注:搜索的范围只有:响应、响应头、请求头
5.如何构造请求
状态码
- 200-300:请求正常
- 300-400:重定向
- 400-500:服务器问题
- 大于或等于500:客户端问题
# 导入模块
import requests # 发送请求的模块
url = 'https://www.xxsy.net/'
headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'referer':'https://www.baidu.com/link?url=suO4Oz7K_kkQ_LRXty3QNX7Y-iCJFCOabR0R_V_z_T_&wd=&eqid=ca9a0fcb006058fc000000066953763c',
'cookie':'newstatisticUUID=1767077440_8196224563',
}
# 构建并发起get请求
res = requests.get(url=url, headers=headers)
# res.status_code # 查看状态码
print(res.status_code)
# res.text # 打印响应文本信息
print(res.text)
# res.content # 响应数据的二进制形式
print(res.content)
说明
res.status_code:查看状态码res.text:响应文本信息(字符串类型数据)res.content:响应数据的二进制形式(流媒体数据)
代码过期问题
- 身份问题:cookie是会过期
补充
不同的网站,以及同一网站不同资源路径,都可能存在不一样的校验方式
同一、相似的资源路径下通常校验规则一样
所以说一般的爬虫流程:分析每一个页面所需要构建的请求方式
作业
需求:
- 熟悉浏览器抓包流程,并可以找到目标资源
- 获取出【https://www.xxsy.net/】网站中,无限追书板块的10本书名,将书名打印出来(选做),**提示:使用正则**
# 导入模块
import requests # 发送请求的模块
# 导入正则模块
import re
url = 'https://www.xxsy.net/'
headers = {
'user -agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'referer':'https://www.baidu.com/link?url=suO4Oz7K_kkQ_LRXty3QNX7Y-iCJFCOabR0R_V_z_T_&wd=&eqid=ca9a0fcb006058fc000000066953763c',
'cookie':'newstatisticUUID=1767077440_8196224563',
}
# 构建并发起get请求
res = requests.get(url=url, headers=headers)
# 定义正则查找条件
module_re = r'<div class="text-t34 text-l-c-1 font-semibold truncate break-all wrap-word hover:text-l-brand-primary-0 block" data-v-540fce0b>(.*?)</div>'
book_html = re.findall(module_re, res.text)
for i in range(len(book_html)):
print(f'{i+1}.{book_html[i]}')
注:有些网站当中的属性值以及标签,是由于、框架进行设置的,存在不确定性,所以具体的正则表达式的写法,需根据代码当中的响应内容来进行
三、登录流程分析
1.权证流
权限(VIP)、凭证(是否登录)
2.登录流程分析
限流问题
单一时间段内,限制同一IP、用户的访问(请求)次数
测试网址
如何快捷的找到登录目标资源包
- 通过搜索关键字:login、submit
- 通过浏览器手动过滤出一些常见的资源包
- 第三方抓包工具
代码
import requests
# 请求地址
url = 'https://www.bqgam.com/user/action.html'
# 请求头内容
headers = {
'Host': 'www.bqgam.com',
'Content-Type': 'application/x-www-form-urlencoded; charset=UTF-8',
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
}
# 负载
data = 'action=login&username=19896223228&password=wzw200508'
# 发起请求
res = requests.post(url=url, headers=headers, data=data)
print(res.text)
3.复杂登录流程分析
测试网站
复杂登录流程分析
输入用户名密码 —> 校验 —> 登录
图片验证码原理
就是每个验证码都会有一个独自的ID值,这个id值跟当前验证码的结果组成一个键值对,提交请求时,会自动携带图片验证码的ID,从而校验提交的验证码是否与键值对当中的验证码一致
Cookie:Hm_lvt_60b7389344d5e30b600d3767cdf28d50=1767158046,1767158565,1767159422; Hm_lpvt_60b7389344d5e30b600d3767cdf28d50=1767159422; HMACCOUNT=306E98BBEBC08A58Cookie:Hm_lvt_60b7389344d5e30b600d3767cdf28d50=1767158046,1767158565,1767159422; HMACCOUNT=306E98BBEBC08A58; ASP.NET_SessionId=pvcyfaa0uywkwb2qh3gbv5hp; Hm_lpvt_60b7389344d5e30b600d3767cdf28d50=1767159429
实战
import requests
# http://www.fbook.net/Member/Captcha?t=0.10251672893148756
# 验证码图片地址
code_url = 'http://www.fbook.net/Member/Captcha'
code_res = requests.get(code_url)
# 将验证码图片保存到指定文件夹
with open('验证码图片/code_01.gif', 'wb') as f:
f.write(code_res.content)
# 登录
# 登录网址
login_url = 'http://www.fbook.net/Member/Login/'
# 登录请求头
login_headers = {
'Content-Type': 'application/x-www-form-urlencoded; charset=UTF-8',
'Cookie': f'ASP.NET_SessionId={code_res.cookies.get("ASP.NET_SessionId")}; Hm_lvt_60b7389344d5e30b600d3767cdf28d50=1767158046,1767158565,1767159422,1767173244; HMACCOUNT=306E98BBEBC08A58; max=AFA58E2DD36DBFADB7F80B766ADEA6A55691F5404576EF7CA6C692327A33F85DB04FDD8E2067594506B40884CA3D81C28EB30E02A47544B9A6DFEF889CFF8F371A5FA4AA1652CFF41A7FBD4E3FD4626647FB2815E376FF70D111A75692B1292A0F91CB6F6DB41F697FA1A67130D8327C6E4BAB22; Hm_lpvt_60b7389344d5e30b600d3767cdf28d50=1767173476',
}
# 登录负载
login_data = f'loginName=huihuia24&loginPass=wzw200508&captcha={input("请输入验证码:")}'
# 登录并发送请求
login_res = requests.post(url=login_url, headers=login_headers, data=login_data)
# 输出结果
print(login_res.text)
四、IP代理
1.代理网址
2.代理的原理
什么情况会使用代理
- 自己的主机被网站限流了
- 不想暴露自己的ip的时候
获取一个代理服务器,转发我们的请求以及回传的响应
原理
主机—-发起请求–>代理—-转发请求–>服务器
主机<–转发响应—-代理<–回传响应—-服务器
3.代理的使用
Chrome浏览器使用代理
注:在使用代理前需要关闭所有正在运行的Google浏览器窗口,否则代理无效
打开cmd通过命令:
chrome --proxy-server=代理ip:端口号
例:
Chrome --proxy-server=210.16.160.222:7890
4.python使用代理
示例
import requests
# 使用代理访问https://www.xxsy.net/
url = 'https://www.xxsy.net/'
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36",
}
# 配置代理
proxy = {
'http': 'http://210.16.160.222:7890',
'https': 'http://210.16.160.222:7890',
}
res = requests.get(url=url, headers=headers, proxies=proxy)
print(res.text)
IP池的使用:搭建代理池—-根据代理IP时长来进行设计使用流程
5.搭建ip代理池
代理池的校验方式
- 使用前校验
- 适合在代理ip提供商提供的代理ip质量不是很好的情况下使用,此方法来提高你的代理IP的可靠性
- 使用时校验
- 使用在代理IP提供商提供的代理IP质量足够高的时候使用,此方法来节省时间
- 需要注意的是,此方法会导致子啊数据爬取工程中导致数据丢失,因此需要对出问题的地方进行标识或记录,然后在代码运行后对标识或记录的丢失位置进行重新采集
IP池的使用
搭建代理池–>根据代理IP时长来进行设计使用
-
针对短效IP
使用
redis数据库来进行存储此类IP,并且设置对应的有效时长 -
长效IP
使用本地文件的形式进行存储,
不管以上两种IP都需要根据它的有效时长来进行合理的代理IP更新
五、批量爬取和翻页规则分析
案例网址:点击此处
1.翻页规则
楼盘网—-新房页面翻页规则
主资源路径(https://cs.loupan.com/xinfang/)+(pn),n为数字1-97来进行翻页更新数据
- 第一页网址:https://cs.loupan.com/xinfang/
- 第二页网址:https://cs.loupan.com/xinfang/p2/
- 第三页网址:https://cs.loupan.com/xinfang/p3/
楼盘网—-商业地产页面翻页规则
主资源路径(https://cs.loupan.com/business/)+(pn),n为数字1-22来进行翻页更新数据
示例
获取站长素材(链接)—-图片页面的前五页数据
# 导入模块
import requests
# 请求头内容
headers = {
'cookie':'_clck=l8ww91%5E2%5Eg2h%5E0%5E2196; _clsk=7urkiw%5E1767681834863%5E8%5E1%5Ev.clarity.ms%2Fcollect',
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'referer':'https://sc.chinaz.com/jianli/',
}
def zhanzhang_tupian():
# 设置爬取范围1-5
for i in range(1,6):
# 根据翻页规则定
if i == 1:
tupian_url = 'https://sc.chinaz.com/tupian/'
else:
tupian_url = f'https://sc.chinaz.com/tupian/index_{i}.html'
# 代理IP
proxy = {
'http': '120.92.212.16:8890',
'https': '120.92.212.16:8890',
}
# 发起请求
tupian_res = requests.get(url=tupian_url, headers=headers, proxies=proxy)
# 设置编码格式防止乱码
tupian_res.encoding = 'UTF-8'
# 短点处
tupian_res
if __name__ == '__main__':
zhanzhang_tupian()
补充
当我们去批量爬取数据的时候,需要非常注意网站的限流问题
- 使用代理—-减少单一IP的过量访问
- 添加等待时间—-防止单一IP在一个时间段内访问频率过高
2.无特定规则翻页分析
潇湘书院(链接)
第一本书
第一页:https://www.xxsy.net/chapter/25866531901067904/69470529228165448
第二页:https://www.xxsy.net/chapter/25866531901067904/69474434292962500第二本书
第一页:https://www.xxsy.net/chapter/26169087201414504/70256336176203430
第二页:https://www.xxsy.net/chapter/26169087201414504/70256433886714439
层级爬取
书籍章节目录:https://www.xxsy.net/chapterlist/ +书籍编号
书籍章节内容:https://www.xxsy.net/chapter/ +书籍编号+书籍章节编号
示例
# 导入模块
import requests # 发送请求的模块
# 导入正则模块
import re
url = 'https://www.xxsy.net/'
headers = {
'user -agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'referer':'https://www.xxsy.net/',
'cookie':'newstatisticUUID=1767077440_8196224563',
}
# 构建并发起get请求
res = requests.get(url=url, headers=headers)
# 获取榜单中书的ID
rank_book_list = re.findall(r'<a href="/book/(d{14,17})" target="_blank">.*?</a>',res.text)[-30:]
rank_chapter_dict = dict()
for i in rank_book_list[:5]:
chapter_url = f'https://www.xxsy.net/chapterlist/{i}'
chapter_res = requests.get(url=chapter_url, headers=headers)
# 获取每个章节的ID
chapter_list = re.findall(r'<a href="/chapter(.*?)" target="_blank" title=".*?" style="padding:8px 0;display:flex;">',chapter_res.text)[0:10]
chapter_list[0] = re.search(r'/chapter(.*)',chapter_list[0][-50:]).group(1)
rank_chapter_dict[i] = chapter_list
for key in rank_chapter_dict:
for chapter_id in rank_chapter_dict[key]:
chapter_url = f'https://www.xxsy.net/chapter{chapter_id}'
chapter_res_if = requests.get(url=chapter_url, headers=headers)
# 获取标题
title = re.findall(r'<h1 .*?>(.*?)</h1>', chapter_res_if.text)[0]
# 获取内容
book_text = tuple(map(lambda x:x+'n',re.findall(r'<p>(.*?)</p>',chapter_res_if.text)))
# 储存至文件
with open(f'book/{key}.txt','a',encoding='utf-8') as f:
f.write(title+'n')
f.writelines(book_text)
流程:获取书籍编号–>获取对应书籍的章节–>获取章节中的具体内容
拓展
map()函数
map函数是用来统一修改可迭代类型的函数
test_list = [1,2,3,4,5,6]
test_list_new = list(map(lambda x:x+1,test_list))
3.针对搜索规则批量爬取分析
经过分析,网站数据是根据搜索内容来进行翻页获取规则的
六、正则解析
1.高阶函数
map
统一操作函数
语法:map(处理函数,可迭代类型)
map返回的是一个可迭代的map对象
注:map对象是不可查看的,需要转化为序列类型才可见(列表、集合、元组)
map函数是用来统一修改可迭代类型的函数
test_list = [1,2,3,4,5,6]
test_list_new = list(map(lambda x:x+1,test_list))
zip
将两个参数合并为一个item对象(组合)
语法:zip(序列类型1(可以包含集合),序列类型2(可以包含集合))->合并函数
注:两个参数的元素个数必须一致
zip返回一个zip对象,需要转换为可见对象->list/tuple/dict/set
list1 = [1,2,3,4,5]
list2 = ['a','b','c','d','e']
zipped1 = tuple(zip(list1,list2))
zipped2 = list(zip(list1,list2))
zipped3 = dict(zip(list1,list2))
zipped4 = set(zip(list1,list2))
print(zipped1)
print(zipped2)
print(zipped3)
print(zipped4)
"""
((1, 'a'), (2, 'b'), (3, 'c'), (4, 'd'), (5, 'e'))
[(1, 'a'), (2, 'b'), (3, 'c'), (4, 'd'), (5, 'e')]
{1: 'a', 2: 'b', 3: 'c', 4: 'd', 5: 'e'}
{(1, 'a'), (3, 'c'), (5, 'e'), (2, 'b'), (4, 'd')}
"""
filter
过滤函数
语法:filter(过滤函数,序列类型(包含集合))
filter返回一个filter对象,不可见,需要转换为可见对象
list_test = [345,454,455,322,456,754]
new_list = filter(lambda x:x>400,list_test)
print(list(new_list))
"""
[454, 455, 456, 754]
"""
enumerate
枚举函数
语法:enumerate(序列类型(不包含集合))
将序列类型里面的数据与它所对应的下标进行组合,返回一个对应的item,
返回enumerate对象,不可见,需转换为序列类型
list_test = [345,454,455,322,456,754]
new_list1 = tuple(enumerate(list_test))
new_list2 = list(enumerate(list_test))
new_list3 = dict(enumerate(list_test))
new_list4 = set(enumerate(list_test))
print(new_list1)
print(new_list2)
print(new_list3)
print(new_list4)
"""
((0, 345), (1, 454), (2, 455), (3, 322), (4, 456), (5, 754))
[(0, 345), (1, 454), (2, 455), (3, 322), (4, 456), (5, 754)]
{0: 345, 1: 454, 2: 455, 3: 322, 4: 456, 5: 754}
{(0, 345), (5, 754), (3, 322), (4, 456), (1, 454), (2, 455)}
"""
2.正则解析
s:匹配换行符–>因为通配符.无法匹配换行
?:取消贪婪模式–>为了防止过多的获取内容
例:
# 内容
<h1>asjdhfkjash</h1>
# 贪婪模式
<.*> ---> <h1>asjdhfkjash</h1>
# 取消贪婪模式
<.*?> --> <h1>,</h1>
():分组–>体现在findall方法与search、macth方法之间的区别
3.正则解析
案例网址
正则解析的优点并不是在于可以很准确的获取数据,而是它可以获取出任何你想要获取的数据
案例一
获取网站中的电影名、电影评分,电影年份、电影时长
import requests
import re
url = 'https://www.imdb.com/chart/top/?ref_=fn_nv_menu'
headers = {
'cookie':'session-id=133-9525601-0942401; session-id-time=2082787201l; ad-oo=0; ci=eyJhZ2VTaWduYWwiOiJBRFVMVCIsImlzR2RwciI6ZmFsc2V9; ubid-main=130-2984191-2615442; session-token=YoLKgL4H8W+ckpct8UwBW3dolcE+Px4dE12RAKFPsa5B+utgA4wzSF6Ghxw3hhSK9zh6ohd/cb6MDEt+WYJKaP33GzQWKV8EcmnIcHkYnKxgZR/zlmKTUUfhfR3slUt0u54gpSBLO548XnP/rMJs46CQeS715Mgr9djGMSRQobrTWKdR5ts02pqU8o+FYgtgK8pRFvNrLtRq8WnEGo0ukFNgFzqhvu3YPsjpDeHT1vx+EkCv1Y5E4B4TsUcmL04M59aYMN0OWxlL+h0i+boAk0HlHw4ds+l2HbtK8U/kSL15TdCQLTDRO8kBX3OdMolUmHa2Qhea6jnb+S1w0xa5LFD8wTVadNLm; csm-hit=tb:3587K4XC6086NRMSH4V8+s-BM1GD1E6FSAYGFZWCA95|1767778920598&t:1767778920598&adb:adblk_yes',
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'referer':'https://www.imdb.com/find/?q=top%2050%20movies&ref_=chttp_nv_srb_sm'
}
# 时间转换
def time_conversion(second):
hour = second // 3600
minute = (second % 3600) // 60
return f'{hour}h{minute}m'
res = requests.get(url,headers=headers)
# 电影名
movie_title = re.findall(r'"url":".*?","name":"(.*?)","description"', res.text)
# 电影评分
movie_rating = re.findall(r'"worstRating":1,"ratingValue":(.*?),', res.text)
# 电影年份
movie_year = re.findall(r'"releaseYear":{"year":(d{4}),', res.text)
# 电影时长
movie_duration = re.findall(r'"seconds":(d{4,5}),', res.text)
# 转换为xh xm 格式
movie_duration_new = [time_conversion(int(i)) for i in movie_duration]
# 拼接输出
max_len = min(len(movie_title), len(movie_rating), len(movie_year), len(movie_duration_new))
for i in range(max_len):
title = movie_title[i]
rating = movie_rating[i]
year = movie_year[i]
duration = movie_duration_new[i]
print(f'电影名称:{title}')
print(f'评分:{rating}')
print(f'年份:{year}')
print(f'时长:{duration}')
print("*"*80)
案例二
爬取猫眼电影TOP100榜1-10页的数据
import requests
import re
headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'referer':'https://www.maoyan.com/',
'cookie':'__mta=47143448.1767771051803.1767790754242.1767791308864.15; _lxsdk_cuid=19b975d6e6fc8-00450a48ddbee6-26061a51-1fa400-19b975d6e6fc8; uuid_n_v=v1; uuid=CA56C320EB9A11F09C6983D1745B13D9E85113618BC240D390361010F21FD821; _csrf=f3bcb09860b1c551112859ec8c32f4a92c56ca73d60be976066f96037dc10791; _lx_utm=utm_source%3DBaidu%26utm_medium%3Dorganic; _lxsdk=CA56C320EB9A11F09C6983D1745B13D9E85113618BC240D390361010F21FD821; _ga=GA1.1.1487119828.1767771051; __mta=47143448.1767771051803.1767771051803.1767791295955.2; _ga_WN80P4PSY7=GS2.1.s1767789851$o2$g1$t1767791308$j46$l0$h0; _lxsdk_s=19b987c7124-7a5-7d3-24e%7C%7C31',
}
result = [i*10 for i in range(0,10)]
for i in result:
url = f'https://www.maoyan.com/board/4?timeStamp=1767789872419&offset={i}'
res = requests.get(url,headers=headers)
# 电影名
movie_title = re.findall(r'data-val="{movieId:d{3,7}}">(.*?)</a></p>',res.text)
# 主演
starring = re.findall(r'主演:(.*?)s*</p>',res.text)
# 上映时间
release_time = re.findall(r'<p class="releasetime">上映时间:(.*?)</p>',res.text)
# 评分
rating = re.findall(r'<p class="score"><i class="integer">(d.)</i><i class="fraction">(d)</i></p>',res.text)
rating_new = [float(i[0]+i[1]) for i in rating]
# 输出
max_len = min(len(movie_title),len(starring),len(release_time),len(rating_new))
for u in range(max_len):
title_date = movie_title[u]
starring_date = starring[u]
release_time_date = release_time[u]
rating_new_date = rating_new[u]
print(f'电影名:{title_date}')
print(f'主演:{starring_date}')
print(f'上映时间:{release_time_date}')
print(f'评分:{rating_new_date}')
print('*'*50)
六、XPATH解析
1.XPATH的使用
XPATH的写法
如果没有限定具体标签时,会找到同级的所有当前标签
:表示分级
htmlbody
[n]:下标限制,n为整数,注:XPATH中下标从1开始
html/body/div[7]
元素[@属性="属性值"]:属性限制
html/body/div[@class="pages"]
//:表示当前节点下任意层级的标签
//div[@class="pages"]
2.提取数据
text
text():提取标签当中的内容
html/body/div[7]/div[3]/div[4]/div[1]/ul/li[1]/div[1]/h2/a/text()
@属性
@属性:获取标签中的属性值
html/body/div[7]/div[3]/div[4]/div[1]/ul/li[1]/div[1]/h2/a/@href
实战一
使用XPATH获取猫眼TOP100榜1-10页的数据
import requests
from lxml import etree # 导入xpath方法
headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'referer':'https://www.maoyan.com/',
'cookie':'__mta=47143448.1767771051803.1767790754242.1767791308864.15; _lxsdk_cuid=19b975d6e6fc8-00450a48ddbee6-26061a51-1fa400-19b975d6e6fc8; uuid_n_v=v1; uuid=CA56C320EB9A11F09C6983D1745B13D9E85113618BC240D390361010F21FD821; _csrf=f3bcb09860b1c551112859ec8c32f4a92c56ca73d60be976066f96037dc10791; _lx_utm=utm_source%3DBaidu%26utm_medium%3Dorganic; _lxsdk=CA56C320EB9A11F09C6983D1745B13D9E85113618BC240D390361010F21FD821; _ga=GA1.1.1487119828.1767771051; __mta=47143448.1767771051803.1767771051803.1767791295955.2; _ga_WN80P4PSY7=GS2.1.s1767789851$o2$g1$t1767791308$j46$l0$h0; _lxsdk_s=19b987c7124-7a5-7d3-24e%7C%7C31',
}
# 评分拼接
def splice(a, b):
rating_list = []
max_len1 = min(len(a),len(b))
for i in range(max_len1):
a1 = a[i]
b1 = b[i]
rating_info = float(a1 + b1)
rating_list.append(rating_info)
return rating_list
result = [i*10 for i in range(0,10)]
for i in result:
url = f'https://www.maoyan.com/board/4?timeStamp=1767789872419&offset={i}'
res = requests.get(url,headers=headers)
# res
# 将text转换为html页面信息,因为获取的数据是字符串信息,字符串数据不具备网页结构,因此需要转换为html页面结构,才能使用XPATH
res_html = etree.HTML(res.text)
# 电影名
movie_title = res_html.xpath('//div[4]/div/div/div[1]/dl/dd/div/div/div[1]/p/a/text()')
# 主演
starring = tuple(map(lambda x:x.strip(), res_html.xpath('//div[4]/div/div/div[1]/dl/dd/div/div/div[1]/p[2]/text()')))
# 上映时间
release_time = res_html.xpath('//div[4]/div/div/div[1]/dl/dd/div/div/div[1]/p[3]/text()')
# 评分
num1 = res_html.xpath('//div[4]/div/div/div[1]/dl/dd/div/div/div[2]/p/i[1]/text()')
num2 = res_html.xpath('//div[4]/div/div/div[1]/dl/dd/div/div/div[2]/p/i[2]/text()')
rating = splice(num1,num2)
# 输出结果
max_len = min(len(movie_title),len(starring),len(release_time),len(rating))
for m in range(max_len):
movie_title_m = movie_title[m]
starring_m = starring[m]
release_time_m = release_time[m]
rating_m = rating[m]
print(f'电影名:{movie_title_m}')
print(starring_m)
print(release_time_m)
print(rating_m)
print('*'*80)
如果碰到标签数据分散的情况下,需要将数据综合时,应当先把他们所共有的部分获取出来,然后对其单独的获取数据
实战二
获取潇湘书院排行榜所有书名、ID及作者
import requests
from lxml import etree
url = 'https://www.xxsy.net/rank'
headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'cookie':'newstatisticUUID=1767077440_8196224563',
'referer':'https://www.xxsy.net/'
}
res = requests.get(url,headers=headers)
res_html = etree.HTML(res.text)
# 书名
books_name = res_html.xpath('//div[@id="app"]/div[4]/div[2]/div[2]/div/div[2]/ul/li/p/a/text()')
# 图书ID
books_id = res_html.xpath('//div[@id="app"]/div[4]/div[2]/div[2]/div/div[2]/ul/li/p/a/@href')
# 作者
books_author = res_html.xpath('//div[@id="app"]/div[4]/div[2]/div[2]/div/div[2]/ul/li/span[1]/text()')
# 拼接网址
books_https = tuple(map(lambda x:'https://www.xxsy.net'+x,books_id))
# 输出
books_list = [f'书名:{name}n作者:{author}n链接:{https}n{'*'*80}' for name,author,https in zip(books_name,books_author,books_https)]
for book in books_list:
print(book)
七、HTML异常处理
1.HTML异常
- 响应数据与浏览器当中的数据存在差异
- 1-100页,其中23出现问题,这个页面没有了
实战一
爬取房天下新房1-10页中的数据
import requests
from lxml import etree
# 清除函数
def clear_string(z, s_list):
s = f'{z}'.join(s_list).strip().replace(' ', '')
return s
# 装饰器,防止网页无法加载时报错结束进程
def requests_error(func):
def requests_error_except(url):
try:
res = requests.get(url=url, headers=headers)
except Exception as e:
print(f'{url}---->{e}')
else:
func(res)
return requests_error_except
@requests_error
def res_try(res):
res_html = etree.HTML(res.text)
elements_li = res_html.xpath('//div[@id="bx1"]/div/div[1]/div/div/div/ul/li')
# 新房名
house_name = tuple(map(lambda x: clear_string('', x.xpath('.//div[@class="nlcd_name"]/a/text()')), elements_li))
# 户型
house_type = tuple(map(lambda x: clear_string('', x.xpath('.//div[@class="house_type clearfix"]//text()')), elements_li))
# 地址
address = tuple(map(lambda x: clear_string('', x.xpath('.//div[@class="address"]//@title')), elements_li))
# 标签
fangyuan = tuple(map(lambda x: clear_string('/', x.xpath('.//div[@class="fangyuan"]//text()')), elements_li))
# 价格
house_price = tuple(map(lambda x: clear_string('/', x.xpath('./div[1]/div[2]/div[5]//text()')), elements_li))
# 电话
house_phone = tuple(map(lambda x: clear_string('', x.xpath('.//div[@class="tel"]/p//text()')), elements_li))
house_list = []
max_len = min(len(house_name),len(house_type),len(address),len(fangyuan),len(house_price),len(house_phone))
for u in range(max_len):
a = house_name[u]
b = house_type[u]
c = address[u]
d = fangyuan[u]
e = house_price[u]
f = house_phone[u]
house_list.append(f'小区名:{a},户型:{b},地址:{c},标签:{d},价格:{e},电话:{f}')
for li in house_list:
with open('book/houses.txt', 'a', encoding='utf-8') as f:
f.write(f'{li}n')
if __name__ == '__main__':
for i in range(1, 11):
url = f'https://newhouse.fang.com/house/s/b9{i}/'
headers = {
'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'cookie': 'global_cookie=zawiyz5btr1jlt6qzqud4vuzm10mk270vgn; city=www1; otherid=accb6736de6a81dacfbdaf6477400308; csrfToken=exWSPIZB1k-XXrJ7s1mDAWgF; unique_cookie=U_6ixbd5vz73mipuw3li39h244m2omk88nvgd*9',
}
if i == 4:
url = f'https://www.google.com/'
house = res_try(url)
实战二
爬取一房网-青岛新房的数据
import requests
from lxml import etree
from tqdm import trange # 进度条
# 清除函数
def clear_string(z, s_list):
s = f'{z}'.join(s_list).strip().replace(' ', '')
return s
# 装饰器,防止网页无法加载时报错结束进程
def requests_error(func):
def requests_error_except(url):
try:
res = requests.get(url=url, headers=headers)
except Exception as e:
print(f'{url}---->{e}')
else:
func(res)
return requests_error_except
@requests_error
def res_try(res):
res_html = etree.HTML(res.text)
elements_li = res_html.xpath('//div[@class="property-item-box"]')
# 新房名
house_name = tuple(map(lambda x: clear_string('', x.xpath('.//p[@class="property-item-title"]/span[1]/text()')), elements_li))
# 户型
house_type = tuple(map(lambda x: clear_string('', x.xpath('.//p[3]//text()')), elements_li))
# 地址
address = tuple(map(lambda x: clear_string('', x.xpath('.//div[@class="property-item-main"]/p[2]//text()')), elements_li))
# 房源
fangyuan = tuple(map(lambda x: clear_string('/', x.xpath('.//p[@class="property-item-title"]/span[@class="el-tag el-tag--mini el-tag--dark"]/text()')), elements_li))
# 价格
house_price = tuple(map(lambda x: clear_string('/', x.xpath('.//div[@class="price-info"]/span/text()')), elements_li))
# 标签
house_label = tuple(map(lambda x: clear_string('/', x.xpath('.//div[@class="property-item-bottom"]//text()')), elements_li))
house_list = []
max_len = min(len(house_name), len(house_type), len(address), len(fangyuan), len(house_price), len(house_label))
for u in range(max_len):
a = house_name[u]
b = house_type[u]
c = address[u]
d = fangyuan[u]
e = house_price[u]
f = house_label[u]
house_list.append(f'小区名:{a},户型:{b},地址:{c},房源:{d},价格:{e},标签:{f}')
for li in house_list:
with open('book/qingdao_houses.txt', 'a', encoding='utf-8') as f:
f.write(f'{li}n')
if __name__ == '__main__':
for i in trange(1, 11):
url = f'http://www.ifang.com/newHouse/list/&/pg{i}/'
headers = {
'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
}
res_try(url)
八、SQL数据库操作
1.mysql配置
打开配置文件
vim /etc/mysql/mysql.conf.d/mysqld.cnf
将配置文件中的**bind-address**设置为0.0.0.0
注:修改配置文件,需要暂停mysql服务
service mysql stop # 暂停mysql服务
service mysql start # 开启mysql服务
2.测试连接数据库
本地下载**pymsql**模块
pip install pymysql
测试连接
from pymysql import Connect
mysql = Connect(
host = "localhost",
# host = "127.0.0.1",
port = 3306,
user = "root",
password = "huihuia24",
charset = "utf8"
)
msyql # 断点
实战一
爬取马蜂窝-去旅行-正在热卖中的数据,并存入数据库
from pymysql import Connect
import requests
from lxml import etree
url = 'https://www.mafengwo.cn/localdeals//'
headers = {
'cookie':'mfw_uuid=69674779-4986-8e88-cbe4-c30ff333c37d; oad_n=a%3A3%3A%7Bs%3A3%3A%22oid%22%3Bi%3A1029%3Bs%3A2%3A%22dm%22%3Bs%3A15%3A%22www.mafengwo.cn%22%3Bs%3A2%3A%22ft%22%3Bs%3A19%3A%222026-01-14+15%3A36%3A25%22%3B%7D; __mfwc=direct; __mfwa=1768376185854.89689.1.1768376185854.1768376185854; __mfwlv=1768376185; __mfwvn=1; uva=s%3A92%3A%22a%3A3%3A%7Bs%3A2%3A%22lt%22%3Bi%3A1768376186%3Bs%3A10%3A%22last_refer%22%3Bs%3A24%3A%22https%3A%2F%2Fwww.mafengwo.cn%2F%22%3Bs%3A5%3A%22rhost%22%3BN%3B%7D%22%3B; __mfwurd=a%3A3%3A%7Bs%3A6%3A%22f_time%22%3Bi%3A1768376186%3Bs%3A9%3A%22f_rdomain%22%3Bs%3A15%3A%22www.mafengwo.cn%22%3Bs%3A6%3A%22f_host%22%3Bs%3A3%3A%22www%22%3B%7D; __mfwuuid=69674779-4986-8e88-cbe4-c30ff333c37d; bottom_ad_status=0; PHPSESSID=ohha4pifncjblkb81uq9ai9i73; __omc_chl=; __omc_r=; __mfwb=76ff3f541d7e.9.direct; __mfwlt=1768377937; w_tsfp=ltvuV0MF2utBvS0Q76Lqk0ynETsjdD84h0wpEaR0f5thQLErU5mD2YV9ucLwNHTY5sxnvd7DsZoyJTLYCJI3dwMSRMmQd9wX31uQxoIjjYdAAhMxEZ7dWlMbJL8kuDRHe3hCNxS00jA8eIUd379yilkMsyN1zap3TO14fstJ019E6KDQmI5uDW3HlFWQRzaLbjcMcuqPr6g18L5a5T/Z5Qiufw18BLlK1UObgXtLW3Er4BLqIuBZMh//I53+SqA=',
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
}
res = requests.get(url, headers=headers)
res_html = etree.HTML(res.text)
# 获取所有的li标签
all_li = res_html.xpath('//div[@class="bd hotSales"]/ul/li')
# 关闭连接
res.close()
# 类型
mark_tag = tuple(map(lambda x:(x.xpath('.//div[@class="mark-tag"]/text()') or ['暂无'])[0], all_li))
# 信息
res_information = tuple(map(lambda x:(x.xpath('.//div[@class="caption"]/h3/text()') or ['暂无'])[0], all_li))
# 出售情况
sell = tuple(map(lambda x:(x.xpath('.//div[@class="caption"]/span[1]/text()') or ['暂无'])[0].split('|'), all_li))
# 店铺
sold = tuple(map(lambda x:x.pop(),sell))
sell2 = tuple(map(lambda x:x[0].replace(' ','') if x else '暂无', sell))
# 价格
price = tuple(map(lambda x: '¥0起' if (s:=''.join(x.xpath('.//div[@class="caption"]/span[2]//text()')).replace(' ','')) == '¥起' else s, all_li))
# 将数据保存至数据库
# 连接数据库
mysql = Connect(
host = "localhost",
# host = "127.0.0.1",
port = 3306,
user = "root",
password = "huihuia24",
charset = "utf8mb4"
)
# 开启游标
cursor = mysql.cursor()
# 创建数据库(如果数据库存在就不会创建)
cursor.execute('create database if not exists mafengwo;')
# 进入数据库
cursor.execute('use mafengwo;')
# 建表
cursor.execute('create table if not exists localdeals('
'id int primary key auto_increment,'
'mark_tag varchar(255) not null,'
'res_information varchar(255) not null,'
'sell char(30) not null,'
'sold char(30) not null,'
'price char(20) not null);')
# 插入数据
insert_sql = 'insert into localdeals(mark_tag,res_information,sell,sold,price) values (%s,%s,%s,%s,%s);'
"""单条数据插入"""
# cursor.execute(insert_sql, (数据))
"""
多条插入
cursor.executemany(insert_sql, data)
多条插入时date的数据格式
data = (
(mark_tag1,information1,sell1,sold1,price1),
(mark_tag2,information2,sell2,sold2,price2),
(mark_tag3,information3,sell3,sold3,price3)
)
"""
data = tuple(map(lambda x:(x[1],res_information[x[0]],sell2[x[0]],sold[x[0]],price[x[0]]),enumerate(mark_tag)))
cursor.executemany(insert_sql, data)
# 提交事务
mysql.commit()
# 关闭游标
cursor.close()
# 关闭数据库连接
mysql.close()
实战二
获取gratisography网中的图片链接,并保存至数据库
import requests
from pymysql import Connect
from lxml import etree
from tqdm import trange
headers = {
'cookie': '_ga=GA1.1.828024986.1768446989; _hjSession_2556496=eyJpZCI6IjJiYTdkNDU4LTQwNTEtNGRhOC1hNDE0LWUyMWRlZjQ1MGNiMCIsImMiOjE3Njg0NDY5OTA5NjEsInMiOjAsInIiOjAsInNiIjowLCJzciI6MCwic2UiOjAsImZzIjoxLCJzcCI6MX0=; _hjSessionUser_2556496=eyJpZCI6ImQ1N2VlYjVmLWRhM2QtNTZmZC04MWIxLTA3MTM3OGVjODUyNyIsImNyZWF0ZWQiOjE3Njg0NDY5OTA5NjAsImV4aXN0aW5nIjp0cnVlfQ==; _ga_TNFKZTM4P0=GS2.1.s1768446989$o1$g1$t1768447380$j52$l0$h0',
'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
}
mysql = Connect(
host='localhost',
port=3306,
user='root',
password='huihuia24',
charset='utf8'
)
# 开启游标
cou = mysql.cursor()
# 创建数据库
cou.execute('create database if not exists gratisography;')
# 进入数据库
cou.execute('use gratisography')
# 创建表
cou.execute('create table if not exists img_url('
'id int primary key auto_increment,'
'name varchar(500) not null,'
'data longblob not null);')
for i in trange(1, 11):
url = f'https://gratisography.com/page/{i}/?s=girl'
res = requests.get(url, headers=headers)
res_html = etree.HTML(res.text)
# 图片地址
img_type = res_html.xpath('//div[@class="off-canvas-content"]/section[2]/div/div/article/div/a[1]/img/@src')
img_info = tuple(map(lambda x: x.replace('-800x525', ''), img_type))
res.close()
# 将图片地址进行请求
for a in range(len(img_info)):
img_url_b = requests.get(url= img_info[a], headers=headers)
img_s = img_url_b.content
img_url_b.close()
# with open(f'girl/{img_info[a]}', 'wb', ) as f:
# f.write(img_s)
# 插入数据
sql = 'insert into img_url(name,data) values(%s,%s);'
cou.execute(sql, (img_info[a], img_s))
# 提交
mysql.commit()
# 关闭
cou.close()
mysql.close()
实战三
获取飞猪旅行-旅游度假-丽江中1-3页数据,并储存至数据库中
import requests
from lxml import etree
from pymysql import Connect
# 图片URL处理
def img_url_info(u):
img_data_list = []
for l in u:
res_img = requests.get(l,headers=headers)
res_b = res_img.content
img_data_list.append(res_b)
res_img.close()
img_data_tuple = tuple(img_data_list)
return img_data_tuple
# 链接数据库
mysql = Connect(
host = 'localhost',
port = 3306,
user = 'root',
passwd = 'huihuia24',
charset = 'utf8'
)
# 开启游标
cur = mysql.cursor()
# 创建数据库
cur.execute('create database if not exists pachong;')
# 进入数据库
cur.execute('use pachong;')
# 建表
cur.execute('create table if not exists fliggy('
'id int primary key auto_increment,'
'tag varchar(50) not null,'
'img_info longblob not null,'
'title varchar(500) not null,'
'address varchar(500) not null,'
'hotel_star varchar(50) not null,'
'rating decimal(3,1) not null,'
'price varchar(50) not null);')
# 爬取数据
headers = {
'cookie':'lid=tb764317057; wk_cookie2=19e5a18fda9e5ad850c05da184c6a41e; wk_unb=UUpgRsNstd9bfuqdKg%3D%3D; cna=1OHjIeP5NmUCAdyoNOTl3rxr; dnk=tb764317057; tracknick=tb764317057; havana_lgc_exp=1799134684908; lgc=; cookie2=1c38e67d82ce54a5b965cd32b31e8eb9; sgcookie=E100d%2BrbIrsZOUIv433Te5daGExY56POchz1s%2FktWzPR9ovYrr%2BVeaP%2FkTv6%2FHhSMt2tb2kBy1Qx0vApJIO%2BTaoC8orxzKpy3fKwfQaarj6NtrZNs6hLto83RMTAI9wzlYoE; cancelledSubSites=empty; t=67bfa247584aa29997f00573aeb0c26c; csg=a83283aa; sn=; _tb_token_=77633e36e6bd6; mtop_partitioned_detect=1; _m_h5_tk=12dba58ee68d6291653b5a96272cd2b2_1768740632889; _m_h5_tk_enc=4eaa6e69520b33d1dbd2ab5fdd64c4c9; xlly_s=1; isg=BAIC-EEM37FXjcNtw0Ec6_2AUwhk0wbtvQOkFEwbLnUgn6IZNGNW_YjMSZvjz36F; _mw_us_time_=1768734760944; x5sec=7b22733b32223a2263323331323134363939333636323639222c22617365727665723b33223a22307c434b654473387347454c61322f706345476738794d6a45774e7a51784f544d324d7a67794f7a45777a4e4c7132762f2f2f2f2f2f41513d3d227d; tfstk=gA4ZU9T8mNQZMj6ua70V8YLdA2utlqW7srMjiSVm1ADihmg2m8wH5lstcxRqL8gb1Pa1oky0pcMgfSG2u5HQhEN_ukL4wSU16q1tW53xoT65PrPT6qF1HwPNdWfmMqlgyPzlp53xo9OBog_z6RwhZfnio6ungb-ioxcM-6loNqYimFxHKvhmoqmimeAnZjJMoEccTWDKiqD0nqfE-vhmox2mobPusSYENcfdICT2ngHjbYViLUPL8fmMX5DeoE4UYckkIv8Doylai5b9qUSjEuHsVYehRFu4tjyEVlWy7-P0womUudf_ES2T8V3Pc1nzxkZSbkWMuvasg00mYIY0TqqimrulrplLx5Z0WRR2ofU_Pmk-YsYxc24Squ2wM_FnSj2x2r6JWAV0wzESzNJ-_kVr8g-6HX0gLrEwnnoi9Xk5T6-SezG1aeiJ5nKxx4hEF1G6Dnni9Xk5T6-vDD0-TY1s1',
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'accept':'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
'referer':'https://www.fliggy.com/'
}
for i in range(1,4):
if i == 1:
url = 'https://travelsearch.fliggy.com/index.htm?spm=181.15077045.1398723350.13.7ab1620dOHgDhv&searchType=product&keyword=%E4%B8%BD%E6%B1%9F&category=MULTI_SEARCH'
else:
url = f'https://travelsearch.fliggy.com/index.htm?spm=181.15077045.1398723350.13.7ab1620dOHgDhv&searchType=product&keyword=%E4%B8%BD%E6%B1%9F&category=MULTI_SEARCH&pagenum={i}'
res_z = requests.get(url, headers=headers)
res_z_html = etree.HTML(res_z.text)
# 全部div
div_all = res_z_html.xpath('//div[@class="page-products-block-left clear-fix"]/div[1]/div/div')
# 标签
tag = tuple(map(lambda x: x.xpath('./div[1]/a/span[1]/text()')[0], div_all))
# 图片地址
img_url = tuple(map(lambda x: ('https:' + x.xpath('./div[1]/a/div/img/@data-src')[0].replace('_200x200xz','')).replace('https:https://','https://'), div_all))
# 图片二进制元组
img_info = img_url_info(img_url)
# 标题
title = tuple(map(lambda x:x.xpath('./div[2]/div[1]/a/h3/div/text()')[0], div_all))
# 地址
address = tuple(map(lambda x:(x.xpath('./div[2]/div[1]/h4/text()') or ['未知'])[0].replace(' | ',''), div_all))
# 酒店星级
hotel_star = tuple(map(lambda x:(x.xpath('./div[2]/p[1]//text()') or ['未知'])[0], div_all))
# 评分
rating = tuple(map(lambda x:(x.xpath('./div[2]/p[3]//text()') or ['0.0'])[0], div_all))
# 价格
price = tuple(map(lambda x:''.join(x.xpath('./div[3]/div/div/span//text()') or ['未知']), div_all))
# 关闭
res_z.close()
# 数据拼接
data = tuple(map(lambda x:(x[1],img_info[x[0]],title[x[0]],address[x[0]],hotel_star[x[0]],rating[x[0]],price[x[0]]),enumerate(tag)))
# 将数据传入数据库中
sql = 'insert into fliggy(tag, img_info, title, address, hotel_star, rating, price) values (%s,%s,%s,%s,%s,%s,%s)'
cur.executemany(sql,data)
# 提交
mysql.commit()
# 关闭
cur.close()
mysql.close()
3.拓展
网络编码转换
from urllib import parse
# 导入urllib
s = '香港'
print(parse.quote(s))
# %E9%A6%99%E6%B8%AF
九、多线程-异步爬虫
实现批量爬虫,多爬虫统一运行
1.多线程
多线程的本质就是多个“人”一起爬取内容
多线程多用于IO(读写)的操作(文件打开,数据存储,爬虫等待)
多线程的使用
不使用多线程
import time
start = time.time()
def sleep_time(x):
print(x)
time.sleep(2)
for i in range(1,11):
sleep_time(i)
print(time.time() - start)
# 运行时间20秒
使用多线程
import time
from threading import Thread # 线程模块
start = time.time()
def sleep_time(x):
print(x)
time.sleep(2)
tasks = [Thread(target=sleep_time,args=(i,))for i in range(1,11)]
for t in tasks:
# 启动线程任务
t.start()
for x in tasks:
# 守护线程
x.join()
print(time.time() - start)
# 运行时间2秒
实战
使用多线程对飞卢小说网中的推理灵异内的书籍内容进行爬取
import requests
import re
from lxml import etree
from threading import Thread
def book_id():
book_id_url = 'https://b.faloo.com/'
headers ={
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
}
book_id_res = requests.get(url=book_id_url, headers=headers)
book_id_html = etree.HTML(book_id_res.text)
books_id = tuple(map(lambda x: re.findall(r'//b.faloo.com/(d+).html?1', x)[0],book_id_html.xpath('//div[@class="TenLeft"]/div[2]/div[1]/div[2]/ul/li/div/a/@href')))
return books_id
def book_count(id_info):
f = open(f'book/{id_info}.txt', 'a', encoding='utf-8')
for i in range(1,41):
book_count_url = f'https://b.faloo.com/{id_info}_{i}.html'
headers ={
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'cookie':'host4chongzhi=b.faloo.com; readline=1; fontFamily=1; fontsize=16; vip_img_width=3; font_Color=666666; nc_rela=2; novelrelative=1458703; favorates28=1458703%2C2%7C1436865%2C4; autobuychapters28=1458703%2C2%7C1436865%2C4; bgcolor=%23FFFFFE; curr_url=https%3A//b.faloo.com/1436865.html%3F1',
}
book_count_res = requests.get(url=book_count_url, headers=headers)
book_count_html = etree.HTML(book_count_res.text)
text = tuple(map(lambda x:x+'n',book_count_html.xpath('//div[@class="noveContent"]/p/text()')))
f.writelines(text)
if __name__ == '__main__':
book_id_info = book_id()
tasks = [Thread(target=book_count,args=(i,)) for i in book_id_info]
for i in tasks:
i.start()
for i in tasks:
i.join()
2.异步-协程
异步的使用
只用于一个场景–>等待IO,遇到IO就切换到下一个任务
import asyncio # 异步模块
# 声明异步函数
async def count(x):
# await 标记需要等待的IO操作
await asyncio.sleep(1)
print(x)
# 异步容器 任务表
loop = asyncio.new_event_loop()
# 启动任务(单次)
# loop.run_until_complete(count(1))
# 启动任务(多次)
ap = [loop.create_task(count(i)) for i in range(10)]
loop.run_until_complete(asyncio.wait(ap))
实战
使用异步对赶集招聘中的数据进行爬取
需要导入异步请求模块
pip install aiohttp
import asyncio
import aiohttp
from lxml import etree
async def download(page:int):
url = f'https://cs.ganji.com/tech/pn{page}/'
headers ={
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/143.0.0.0 Safari/537.36',
'cookie':'id58=CkwAG2lvK0U6FAFwEvZYAg==; 58tj_uuid=dac3aab2-2276-469e-a743-f8b8d18f8420; ngj_city_id=2; ngj_city_name=%E4%B8%8A%E6%B5%B7; ngj_city_listname=sh; new_uv=2',
}
async with aiohttp.ClientSession() as session:
# 使用代理的写法
# async with session.get(url, headers=headers, proxy='http://代理ip:端口号') as res:
async with session.get(url, headers=headers) as res:
# text()字符串
# read()二进制数据
# json()的json类型的数据
text = await res.text()
return text
# 无法直接解析数据
# loop = asyncio.new_event_loop()
# tasks = [loop.create_task(download(i)) for i in range(1,11)]
# loop.run_until_complete(asyncio.wait(tasks))
# 解析函数
def func(fut):
result = fut.result()
html = etree.HTML(result)
a_all = html.xpath('//div[@class="position"]/div[4]/div/a[@class="ibox"]')
# 标题
title = tuple(map(lambda x:x.xpath('./ul[1]/li[1]/text()')[0],a_all))
# 薪水
salary = tuple(map(lambda x:''.join(x.xpath('./ul[1]/li[2]//text()')).strip('n '),a_all))
# 标签
label = tuple(map(lambda x:'/'.join(x.xpath('./ul[1]/div/span/span/text()')),a_all))
# 公司
company = tuple(map(lambda x: x.xpath('./ul[2]/li[1]/object/a/text()')[0].strip('n '), a_all))
# 地址
address = tuple(map(lambda x:''.join(x.xpath('./ul[2]/li[2]//text()')).strip('n ').replace(' | ','|'),a_all))
# 拼接数据
data = tuple(map(lambda x:(x[1],salary[x[0]],label[x[0]],company[x[0]],address[x[0]]), enumerate(title)))
pass
# 分配任务的异步函数
async def main():
tasks = list()
for page in range(1,11):
c = download(page) # 创建协程对象
task = asyncio.ensure_future(c) # 将任务对象进行异步封装
task.add_done_callback(func) # 指定回调函数
tasks.append(task)
await asyncio.gather(*tasks)
if __name__ == '__main__':
asyncio.run(main())
十、selenium
1.selenium工具准备
-
Chrome浏览器
-
Chrome驱动
-
驱动版本查看
Chrome浏览器地址栏输入:
chrome://version
2.selenium的使用
需要安装第三方模块
pip install selenium
当使用selenium这种自动化程序时,会有对应的“指纹”信息
在网页控制台终端输入:
navigator.webdriver
反检测模块
pip install undetected_chromedriver
示例
from selenium import webdriver
import undetected_chromedriver as webdriver_1
# 使用代理
options = webdriver.ChromeOptions()
options.add_argument('--proxy-server=socks4://39.104.16.201:3128')
chrome = webdriver_1.Chrome()
# 注:在运行前尽量关闭所有Google浏览器
chrome.get('https://www.youku.com/ku/webhome')
input()
十一、Scrapy框架基础
1.Scrapy框架
Scrapy爬虫框架是基于异步的爬虫框架
主要封装的的功能:多线程+异步操作(twisted网络异步通信架构)
多机协同工作(redis服务)
Scrapy框架是一个完整的爬虫
完整的爬虫包括:爬虫构造,数据获取,数据库初始化,数据保存
2.模块准备
pip install scrapy
3. 框架的使用
创建项目
scrapy startproject 项目名
项目设置
构建爬虫
数据存储
十二、js逆向-代码反编译
案例网址:oklink
1.js逆向
有些网站不仅仅通过一些常规字段来进行身份检验,如:cookie、sessionid,因为这些常见字段具有一定缺陷,那么此时就会由前端js代码来生成一个对应的校验字段,来进行身份校验,这种校验是算法的校验。
x-apikey比较:
LWIzMWUtNDU0Ny05Mjk5LWI2ZDA3Yjc2MzFhYmEyYzkwM2NjfDI4ODAxODAzNDA1MzAzMTA=
3Yjc2MzFhYmEyYzkwM2NjfDI4ODAxODA2MzE5NTY5MTE=
当看到某个加密字段每次都会变化,但会有一段始终不变,可以推断它是一个签名算法
签名算法:由一个或者多个变化的值,加上固定值[盐],通过加密算法得来
2.寻找目标位置
Ctrl+shift+f进行查找
这些数据是通过加密算法来的:看搜索出来的上下文当中是否包含[parse、encrypt、md5、rsa、aes]这类型的字符
得到的js代码
function getApiKey() {
var e = (new Date).getTime(), t = encryptApiKey();
return e = encryptTime(e),
comb(t, e)
}
API_KEY = "a2c903cc-b31e-4547-9299-b6d07b7631ab"
function encryptApiKey() {
var e = API_KEY, t = e.split(""), n = t.splice(0, 8);
return e = t.concat(n).join("")
}
function encryptTime(e) {
var t = (1 * e + s).toString().split("")
, n = parseInt(10 * o.o.mathRandom(), 10)
, r = parseInt(10 * o.o.mathRandom(), 10)
, i = parseInt(10 * o.o.mathRandom(), 10);
return t.concat([n, r, i]).join("")
}
function comb(e, t) {
var n = "".concat(e, "|").concat(t);
return a.A.btoa(n)
}
转换后
from time import time
from random import randint
from base64 import b64encode
import requests
def getApiKey():
e = int(time()*1000)
t = encryptApiKey()
e = encryptTime(e)
apikey = comb(t, e)
return apikey
API_KEY = "a2c903cc-b31e-4547-9299-b6d07b7631ab"
def encryptApiKey():
e = API_KEY
t = list(e)
n,t = t[:8],t[8:]
t.extend(n)
e = ''.join(t)
return e
def encryptTime(e):
t = list(str((1 * e + 1111111111111)))
n = str(randint(0,9))
r = str(randint(0,9))
i = str(randint(0,9))
t.extend([n,r,i])
return ''.join(t)
def comb(e, t):
n = f'{e}|{t}'
apikey = b64encode(n.encode()).decode()
return apikey
key = getApiKey()
url = f'https://www.oklink.com/api/explorer/v1/btc/blocks?offset=0&limit=100&t={int(time()*1000)}'
headers ={
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/144.0.0.0 Safari/537.36',
'cookie':'devId=db003635-abc4-45aa-94d4-3698e67bde1e; ok_site_info=9FjOikHdpRnblJCLiskTJx0SPJiOiUGZvNmIsIyUVJiOi42bpdWZyJye; locale=zh_CN; ok-exp-time=1769069055106; first_ref=https%3A%2F%2Fwww.qklw.com%2F; fingerprint_id=db003635-abc4-45aa-94d4-3698e67bde1e; fp_s=-1; okg.currentMedia=md; traceId=2020190762744410001; __cf_bm=fQb5tROqasRo6vByWpjb4JlO_2Xqq3HrpX0hnayvhkY-1769076274-1.0.1.1-HORhLlqOB6gzsJAlkGkDRb5KM4cQ1k.MzM8Ik0vN10Hdpgt3UEN7i44WjbxbpA78JfOq6gI0U2lZwiXklJHRn2seNJbwDLeYzS3uVD9Z8MQ; ok-ses-id=iVwVAiHXnes+wD/HMLUOA0QIZ2QKUX9KZ18WRIIiM/Tsdl6ALbbVAT/kyOyxWCINLAy9JBLxKLbEEOKuiCMi2C2Gk/vccjYy1+SQojTseDiF2IH2Yhdn4HfS5fVN/zbP; _monitor_extras={"deviceId":"K7C9deCj86-hyeCzzPkH5t","eventId":98,"sequenceNumber":98}',
'x-apikey':key,
'devid':'db003635-abc4-45aa-94d4-3698e67bde1e'
}
res = requests.get(url,headers=headers)
res
3.练习
vivo社区
使用js逆向对vivo社区进行爬取
md5的加密特征
- 不可逆的加密算法
- 明文与密文一对一
- 进行加密后的结果为32位16进制的字符组成==>0-9,a-f
import requests
import json
from hashlib import md5
from time import time
from random import randint
for page in range(1,11):
# js
# e. params.nonce = Object(je.md5)(n + "" + parseInt(1e7 * Math.random(), 10) + 1, 32)
times = int(time() * 1000)
nonce = md5(f'{times}{randint(1000000, 9999999) + 1}'.encode()).hexdigest() # hexdigest()获取32位加密结果
url = 'https://bbs.vivo.com.cn/api/community/index'
headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/144.0.0.0 Safari/537.36',
'cookie':'cookieId=2af71ac1-d2eb-1171-be0a-a4da24190f7d1767335490049; sessionId=9ad7e8ac-3f6b-abc2-5611-caf9e978768f',
'referer':'https://bbs.vivo.com.cn/newbbs/',
'content-type':'application/json;charset=UTF-8',
}
# 因为前后端数据传输需要的是json格式的数据,可能需要对data进行序列化操作
data = json.dumps({"lastId":"","pageNum":page,"pageSize":10,"imgSpecs":["t577x324","t577x4096"],"timestamp":times,"nonce":nonce})
res = requests.post(url,headers=headers,data=data)
res
4.HOOK
因为有些网站会干扰调试,那么此时就需要进行HOOK注入,替换对应的索引文件
HOOK注入
十三、js逆向-webpack
1.常见的加密算法
-
签名算法
不可逆的加密算法,通常明文和密文是一对一关系,在进行加密的时候,在明文前后加入一些固定的字符串(盐)进行加密==>md5加密,sha系列
- MD5的密文==>16进制的字符组成0-9,a-f
-
对称加密算法
可逆的算法,可以加密也可以解密,明文和密文之间是对称的关系,对称加密往往要传入3个参数
-
明文(data、messge)
-
钥匙(key)
-
偏移量(iv)
aes、des,加密后的内容==>
a-z,A-Z,0-9,+,/,=
-
-
非对称加密算法
明文和密文存在一对多的关系,可以解密,会存在私钥、公钥,公钥只能用来加密,私钥用来解密
2.python调用js代码
模块准备:pyexecjs
pip install pyexecjs
execjs模块使用示例
function a(x,y){
return x+y;
}
# 导入模块
import execjs
# 加载js代码
with open('text.js', 'r', encoding='utf-8') as f:
nodejs = f.read()
# 转换为可执行的代码
ctx = execjs.compile(nodejs)
# call方法:调用js中的函数,参数1:函数名,参数2:传入值
a = ctx.call('a',3,5)
print(a)
# 8
案例展示
1.oklink
API_KEY = "a2c903cc-b31e-4547-9299-b6d07b7631ab"
function encryptApiKey() {
var e = API_KEY
, t = e.split("")
, r = t.splice(0, 8);
return e = t.concat(r).join("")
}
p = 1111111111111
function encryptTime(e) {
var t = (1 * e + p).toString().split("")
, r = parseInt(10 * Math.random(), 10)
, n = parseInt(10 * Math.random(), 10)
, o = parseInt(10 * Math.random(), 10);
return t.concat([r, n, o]).join("")
}
function comb(e, t) {
var r = "".concat(e, "|").concat(t);
return Buffer(r).toString('base64')
}
function getApiKey() {
var e = (new Date).getTime()
, t = encryptApiKey();
return e = encryptTime(e),
comb(t, e)
}
import execjs
with open('oklink.js','r',encoding='utf-8') as f:
nodejs = f.read()
ctx = execjs.compile(nodejs)
a = ctx.call('getApiKey')
print(a)
2.vivo社区
var n = Date.now();
function md5(e, t) {
function r(e, t) {
return e << t | e >>> 32 - t
}
function n(e, t) {
var r, n, u, o, l;
return u = 2147483648 & e,
o = 2147483648 & t,
l = (1073741823 & e) + (1073741823 & t),
(r = 1073741824 & e) & (n = 1073741824 & t) ? 2147483648 ^ l ^ u ^ o : r | n ? 1073741824 & l ? 3221225472 ^ l ^ u ^ o : 1073741824 ^ l ^ u ^ o : l ^ u ^ o
}
function u(e, t, u, o, l, i, c) {
return e = n(e, n(n(function(e, t, r) {
return e & t | ~e & r
}(t, u, o), l), c)),
n(r(e, i), t)
}
function o(e, t, u, o, l, i, c) {
return e = n(e, n(n(function(e, t, r) {
return e & r | t & ~r
}(t, u, o), l), c)),
n(r(e, i), t)
}
function l(e, t, u, o, l, i, c) {
return e = n(e, n(n(function(e, t, r) {
return e ^ t ^ r
}(t, u, o), l), c)),
n(r(e, i), t)
}
function i(e, t, u, o, l, i, c) {
return e = n(e, n(n(function(e, t, r) {
return t ^ (e | ~r)
}(t, u, o), l), c)),
n(r(e, i), t)
}
function c(e) {
var t, r = "", n = "";
for (t = 0; t <= 3; t++)
r += (n = "0" + (e >>> 8 * t & 255).toString(16)).substr(n.length - 2, 2);
return r
}
var d, a, s, p, f, h, m, A, g, v = e, b = Array();
for (b = function(e) {
for (var t, r = e.length, n = r + 8, u = 16 * ((n - n % 64) / 64 + 1), o = Array(u - 1), l = 0, i = 0; i < r; )
l = i % 4 * 8,
o[t = (i - i % 4) / 4] = o[t] | e.charCodeAt(i) << l,
i++;
return l = i % 4 * 8,
o[t = (i - i % 4) / 4] = o[t] | 128 << l,
o[u - 2] = r << 3,
o[u - 1] = r >>> 29,
o
}(v),
h = 1732584193,
m = 4023233417,
A = 2562383102,
g = 271733878,
d = 0; d < b.length; d += 16)
a = h,
s = m,
p = A,
f = g,
m = i(m = i(m = i(m = i(m = l(m = l(m = l(m = l(m = o(m = o(m = o(m = o(m = u(m = u(m = u(m = u(m, A = u(A, g = u(g, h = u(h, m, A, g, b[d + 0], 7, 3614090360), m, A, b[d + 1], 12, 3905402710), h, m, b[d + 2], 17, 606105819), g, h, b[d + 3], 22, 3250441966), A = u(A, g = u(g, h = u(h, m, A, g, b[d + 4], 7, 4118548399), m, A, b[d + 5], 12, 1200080426), h, m, b[d + 6], 17, 2821735955), g, h, b[d + 7], 22, 4249261313), A = u(A, g = u(g, h = u(h, m, A, g, b[d + 8], 7, 1770035416), m, A, b[d + 9], 12, 2336552879), h, m, b[d + 10], 17, 4294925233), g, h, b[d + 11], 22, 2304563134), A = u(A, g = u(g, h = u(h, m, A, g, b[d + 12], 7, 1804603682), m, A, b[d + 13], 12, 4254626195), h, m, b[d + 14], 17, 2792965006), g, h, b[d + 15], 22, 1236535329), A = o(A, g = o(g, h = o(h, m, A, g, b[d + 1], 5, 4129170786), m, A, b[d + 6], 9, 3225465664), h, m, b[d + 11], 14, 643717713), g, h, b[d + 0], 20, 3921069994), A = o(A, g = o(g, h = o(h, m, A, g, b[d + 5], 5, 3593408605), m, A, b[d + 10], 9, 38016083), h, m, b[d + 15], 14, 3634488961), g, h, b[d + 4], 20, 3889429448), A = o(A, g = o(g, h = o(h, m, A, g, b[d + 9], 5, 568446438), m, A, b[d + 14], 9, 3275163606), h, m, b[d + 3], 14, 4107603335), g, h, b[d + 8], 20, 1163531501), A = o(A, g = o(g, h = o(h, m, A, g, b[d + 13], 5, 2850285829), m, A, b[d + 2], 9, 4243563512), h, m, b[d + 7], 14, 1735328473), g, h, b[d + 12], 20, 2368359562), A = l(A, g = l(g, h = l(h, m, A, g, b[d + 5], 4, 4294588738), m, A, b[d + 8], 11, 2272392833), h, m, b[d + 11], 16, 1839030562), g, h, b[d + 14], 23, 4259657740), A = l(A, g = l(g, h = l(h, m, A, g, b[d + 1], 4, 2763975236), m, A, b[d + 4], 11, 1272893353), h, m, b[d + 7], 16, 4139469664), g, h, b[d + 10], 23, 3200236656), A = l(A, g = l(g, h = l(h, m, A, g, b[d + 13], 4, 681279174), m, A, b[d + 0], 11, 3936430074), h, m, b[d + 3], 16, 3572445317), g, h, b[d + 6], 23, 76029189), A = l(A, g = l(g, h = l(h, m, A, g, b[d + 9], 4, 3654602809), m, A, b[d + 12], 11, 3873151461), h, m, b[d + 15], 16, 530742520), g, h, b[d + 2], 23, 3299628645), A = i(A, g = i(g, h = i(h, m, A, g, b[d + 0], 6, 4096336452), m, A, b[d + 7], 10, 1126891415), h, m, b[d + 14], 15, 2878612391), g, h, b[d + 5], 21, 4237533241), A = i(A, g = i(g, h = i(h, m, A, g, b[d + 12], 6, 1700485571), m, A, b[d + 3], 10, 2399980690), h, m, b[d + 10], 15, 4293915773), g, h, b[d + 1], 21, 2240044497), A = i(A, g = i(g, h = i(h, m, A, g, b[d + 8], 6, 1873313359), m, A, b[d + 15], 10, 4264355552), h, m, b[d + 6], 15, 2734768916), g, h, b[d + 13], 21, 1309151649), A = i(A, g = i(g, h = i(h, m, A, g, b[d + 4], 6, 4149444226), m, A, b[d + 11], 10, 3174756917), h, m, b[d + 2], 15, 718787259), g, h, b[d + 9], 21, 3951481745),
h = n(h, a),
m = n(m, s),
A = n(A, p),
g = n(g, f);
return 32 == t ? c(h) + c(m) + c(A) + c(g) : c(m) + c(A)
}
function res(){
data = md5(n + "" + parseInt(1e7 * Math.random(), 10) + 1, 32)
return data;
}
console.log(res())
import execjs
with open('vivo.js','r',encoding='utf-8') as f:
nodejs = f.read()
ctx = execjs.compile(nodejs)
a = ctx.call('res')
print(a)
3.webpack
代码容器—-因为一个实际的模块当中存在非常多的方法,并且有可能存在相互反复的调用,所以我们不能保证一个一个去找到,所有的配置函数,因此应该找到调用这个加密函数的代码容器来进行执行当中的代码
Promise(函数)-then(参数)
Promise(function).then(x)Promise里面的function函数的结果会传入到then方法的参数当中
十四、滑块验证码
验证码是通过什么来校验的
| ID | data |
|---|---|
| abc01 | 170px-175px |
1.案例
案例网址:慈善中国
清除网站cookie
应用–>cookie–>右键清除
2.ddddocr
下载ddddocr模块
pip install ddddocr
实例
import requests
from ddddocr import DdddOcr
from base64 import b64decode,b64encode
# 创建ocr实例
# ocr:检测文字
# det:用作目标检测,如扫码
ocr = DdddOcr(ocr=False,det=False)
get_img_url = 'https://cszg.mca.gov.cn/biz/ma/csmh/filter/getSlideCaptcha.html'
img_headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/144.0.0.0 Safari/537.36'
}
img_res = requests.get(get_img_url, headers=img_headers)
oriImage,cutImage = img_res.json().get('c').get('oriImage'),img_res.json().get('c').get('cutImage')
ori_img,cut_img = b64decode(oriImage),b64decode(cutImage)
with open(f'img/ori_img.png', 'wb') as f:
f.write(ori_img)
with open(f'img/cut_img.png', 'wb') as f:
f.write(cut_img)
position = ocr.slide_match(cut_img,ori_img,simple_target=True)['target']
print(position)
x1,y1,x2,y2 = position
slide_cap_url = f'https://cszg.mca.gov.cn/biz/ma/csmh/filter/slideCaptchaCheck.html?slidevalue={b64encode(str(x1).encode()).decode()}'
slide_cap_headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/144.0.0.0 Safari/537.36',
'cookie':f'https_waf_cookie=92fbb60c-23e4-4884d27f5352a8a01883cce7d8f068eed942; JSESSIONID={img_res.cookies.get('JSESSIONID')}',
}
slide_cap_res = requests.get(slide_cap_url, headers=slide_cap_headers)
print(slide_cap_res.json())
3.opencv-python
下载模块
pip install opencv-python
import cv2
# 读取背景图片和缺口图片
cut_img = cv2.imread('img/cut_img.png')
ori_img = cv2.imread('img/ori_img.png')
# 识别图片边缘
cut_edge = cv2.Canny(cut_img, 100, 200)
ori_edge = cv2.Canny(ori_img, 100, 200)
# 转换图片格式
cut_pic = cv2.cvtColor(cut_edge, cv2.COLOR_GRAY2BGRA)
ori_pic = cv2.cvtColor(ori_edge, cv2.COLOR_GRAY2BGRA)
# 缺口匹配
res = cv2.matchTemplate(ori_pic, cut_pic, cv2.TM_CCOEFF_NORMED)
# 寻找最佳匹配
position = cv2.minMaxLoc(res)
print(position)
十四、字体反爬
案例网址:懂车帝
1.字体反爬原理
实际数据是由一个网络编码来进行占位,后期通过js脚本去将对应的编码替换成对应的“字符”
字体查看工具:BEJSON
import requests
url = 'https://www.dongchedi.com/motor/pc/sh/sh_sku_list?aid=1839&app_name=auto_web_pc'
headers = {
'cookie':'ttwid=1%7Cl6DOgeb_8xSC-4cLqLudIpUdQ0bswFBs1wzOE0fOB9o%7C1769166483%7C20d2f5b5520bb67ac42cc9ab957ae1362bfcf71b73c2d4160f5bf3401b490443; tt_webid=7598512060565030462; tt_web_version=new; is_dev=false; is_boe=false; _ga=GA1.1.184307655.1769166485; x-web-secsdk-uid=a1d77778-418f-40c3-9a8b-bb34d7c0abfc; s_v_web_id=verify_mkqs1q83_vMr6oxVB_XxS7_4jEF_ApS1_9gubbliP4yXJ; city_name=%E9%95%BF%E6%B2%99; rit_city=%E9%95%BF%E6%B2%99; _ga_YB3EWSDTGF=GS2.1.s1769166485$o1$g1$t1769167384$j56$l0$h0',
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/144.0.0.0 Safari/537.36',
'referer':'https://www.dongchedi.com/usedcar/x-x-x-x-x-x-x-x-x-x-x-x-x-x-x-x-x-x-x-x-x-110000-x-x-x-x-x-x',
'content-type':'application/x-www-form-urlencoded'
}
data = '&sh_city_name=全国&page=1&limit=20'
res = requests.post(url, headers=headers, data=data)
font_dict = {
'uE439':'0',
'uE54C':'1',
'uE463':'2',
'uE49D':'3',
'uE41D':'4',
'uE411':'5',
'uE534':'6',
'uE3EB':'7',
'uE4E3':'8',
'uE45D':'9',
'uE40A':'万'
}
sh_price = tuple(map(lambda x:x.get('sh_price'),res.json().get('data').get('search_sh_sku_info_list')))
sh_price = tuple(map(lambda x:x.split('.'),sh_price))
sh_Price = []
for i,f in sh_price:
I = ''
F = ''
for i_test in i:
I += font_dict[i_test]
for f_test in f:
F += font_dict[f_test]
sh_Price.append(I+'.'+F)
official_price = tuple(map(lambda x:x.get('official_price'),res.json().get('data').get('search_sh_sku_info_list')))
official_price = tuple(map(lambda x:x.split('.'),official_price))
official_Price = []
for i,f in official_price:
I = ''
F = ''
for i_test in i:
I += font_dict[i_test]
for f_test in f:
F += font_dict[f_test]
official_Price.append(I+'.'+F)
2.fontTools
模块准备
pip install fontTools
案例网址:猫眼电影-国内票房榜
import requests
from fontTools.ttLib import TTFont # 加载、读取原始字体文件的方法
from lxml import etree
# 读取字体文件
Font_Dict = dict()
font1 = TTFont('猫眼font/20a70494.woff')
# 保存为XML文件
font1.saveXML('猫眼font/20a70494.xml')
index1 = tuple(map(lambda x:x.lower().replace('uni',r'u'),font1.getGlyphOrder()[2:]))
font_dict1 = dict(zip(index1,[7,3,6,1,2,8,0,4,9,5]))
font2 = TTFont('猫眼font/e3dfe524.woff')
# 保存为XML文件
font2.saveXML('猫眼font/e3dfe524.xml')
index2 = tuple(map(lambda x:x.lower().replace('uni',r'u'),font2.getGlyphOrder()[2:]))
font_dict2 = dict(zip(index2,[9,2,4,1,5,3,6,8,0,7]))
font3 = TTFont('猫眼font/75e5b39d.woff')
# 保存为XML文件
font3.saveXML('猫眼font/75e5b39d.xml')
index3 = tuple(map(lambda x:x.lower().replace('uni',r'u'),font3.getGlyphOrder()[2:]))
font_dict3 = dict(zip(index3,[0,3,8,2,6,7,9,4,1,5]))
font4 = TTFont('猫眼font/2a70c44b.woff')
# 保存为XML文件
font4.saveXML('猫眼font/2a70c44b.xml')
index4 = tuple(map(lambda x:x.lower().replace('uni',r'u'),font4.getGlyphOrder()[2:]))
font_dict4 = dict(zip(index4,[7,5,3,9,0,2,6,4,1,8]))
Font_Dict.update(font_dict1)
Font_Dict.update(font_dict2)
Font_Dict.update(font_dict3)
Font_Dict.update(font_dict4)
url = 'https://www.maoyan.com/board/1'
headers = {
'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/144.0.0.0 Safari/537.36',
'cookie':'__mta=47143448.1767771051803.1769171627927.1769171630945.34; _lxsdk_cuid=19b975d6e6fc8-00450a48ddbee6-26061a51-1fa400-19b975d6e6fc8; uuid_n_v=v1; uuid=CA56C320EB9A11F09C6983D1745B13D9E85113618BC240D390361010F21FD821; _lxsdk=CA56C320EB9A11F09C6983D1745B13D9E85113618BC240D390361010F21FD821; _ga=GA1.1.1487119828.1767771051; _csrf=41b68d1e6a40d78c0860c78aa6776dce2b5ef4c99bf4e27f8ca934ccd6bafc18; global-guide-isclose=true; hotMovieIds=1478868,1322627,1498191,1142033,78463,1525137,1504573,338414,1487265,1422798,1355532,356895,1528369,1529812,1399234,1405509,1504556,1425266,1603579,1432807,1618267,1323,1573941,1535383,1299948,1572282,1614889,1548343,1502290,1428854,1547279,346650,1559720,1393510,1526767,1501173,1485017,376806,4430,247161,583,1502849,1302195,7284,285540,1295,1340,233631,32124,1487834,1528899; old-moviepage-ci=70; __mta=47143448.1767771051803.1769171570301.1769171591436.8; _lx_utm=utm_source%3DBaidu%26utm_medium%3Dorganic; _ga_WN80P4PSY7=GS2.1.s1769171394$o6$g1$t1769171630$j57$l0$h0; _lxsdk_s=19bead51e7d-5c-fbd-cda%7C%7C15',
'referer':'https://www.maoyan.com/board/4?timeStamp=1767789872419&offset=0',
}
res = requests.get(url,headers=headers)
res_html = etree.HTML(res.text)
res
realtime = res_html.xpath('//dl[@class="board-wrapper"]/dd/div/div/div[2]/p[1]/span/span/text()')
total_boxoffice = res_html.xpath('//dl[@class="board-wrapper"]/dd/div/div/div[2]/p[2]/span/span/text()')
realtime = tuple(map(lambda x:x.split('.'), realtime))
total_boxoffice = tuple(map(lambda x:x.split('.'), total_boxoffice))
# 替换
def replace(l,s:str):
L = []
for i,f in l:
I = ''
F = ''
for i_test in i:
I += str(Font_Dict[i_test.encode('unicode-escape').decode()])
for f_test in f:
F += str(Font_Dict[f_test.encode('unicode-escape').decode()])
L.append(I+s+F)
return L
realtime = replace(realtime,'.')
total_boxoffice = replace(total_boxoffice,'.')
print(realtime)
print(total_boxoffice)






如有错误,请反馈给作者,谢谢🌹