加载中...
python爬虫表情包
发表于:2022-04-20 | 分类: python
字数统计: 243 | 阅读时长: 1分钟 | 阅读量:

在这里插入图片描述
引入模块

1
2
import requests //用于请求网页
import re //正则表达式,用于解析筛选网页中的信息

要爬的网址

1
2
3
4
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:98.0) Gecko/20100101 Firefox/98.0'
}
response = requests.get('https://qq.yh31.com/zjbq/',headers=headers) //请求网页

完整代码

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
import requests
import re
import os

image = '表情包'
if not os.path.exists(image):
os.mkdir(image)
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:98.0) Gecko/20100101 Firefox/98.0'
}
response = requests.get('https://qq.yh31.com/zjbq/',headers=headers)
response.encoding = 'GBK'
response.encoding = 'utf-8'
print(response.request.headers)
print(response.status_code)
t = '<img src="(.*?)" alt="(.*?)" width="160" height="120">'
result = re.findall(t, response.text)
for img in result:
print(img)
res = requests.get(img[0])
print(res.status_code)
s = img[0].split('.')[-1] //截取地址末尾,得到表情包格式,如.jpg .git
with open(image + '/' + img[1] + '.' + s, mode='wb') as file:
file.write(res.content)

原文章

下一篇:
Django JWT token 登录注册
本文目录
本文目录