Scrapy 爬虫入门

文章发布时间:

Scrapy 适合抓取直接包含目标数据的 HTML 页面。如果数据要等浏览器执行 JavaScript,或滚动、点击后才出现,先检查页面是否调用了数据接口;没有可直接请求的接口时,再考虑 Selenium 或 Playwright。

Scrapy 的请求流程

先看请求和数据怎样在各组件之间流转。下面这张流程图来自腾讯云开发者社区:

Scrapy 请求和数据流转示意图

基本流程是:

1
Spider 产生请求 -> Engine 调度 -> Downloader 下载页面 -> Spider 解析响应 -> Pipeline 处理数据

入门时主要写两块:

  • spider:页面从哪里进来,数据怎么取,下一页怎么跟;
  • pipeline:取出来的数据要不要清洗、去重、保存。

请求调度、下载、请求去重和并发控制由 Scrapy 负责;需要持续抓取、翻页和处理数据时,不必再围绕 requests.get() 自己拼这些环节。

创建项目

建议先用熟悉的工具创建并激活虚拟环境;我通常用 uv。然后安装 Scrapy:

1
pip install scrapy

创建项目:

1
2
scrapy startproject quote_spider
cd quote_spider

创建完成后,目录大概是这样:

Scrapy 项目目录示意图

常用文件先记这几个:

1
2
3
4
5
quote_spider/
items.py # 定义数据字段
pipelines.py # 清洗、保存数据
settings.py # 并发、下载延迟、管道开关等配置
spiders/ # 放具体爬虫

用 shell 先试页面

写 spider 前,先在 Scrapy shell 里试 CSS 或 XPath 选择器,确认它们能从实际响应中取到数据:

1
scrapy shell https://quotes.toscrape.com/

进去后先试:

1
response.css("div.quote").get()

能拿到一段 HTML,说明页面内容在服务端返回的 HTML 里,Scrapy 可以直接抓。

再取第一条 quote:

1
2
3
4
quote = response.css("div.quote")[0]
quote.css("span.text::text").get()
quote.css("small.author::text").get()
quote.css("div.tags a.tag::text").getall()

这里最常用的就三个写法:

1
2
3
::text        取文本
::attr(href) 取属性
.getall() 取全部结果

如果取不到数据,先检查响应内容和选择器,不要直接在 spider 里反复试。

第一个 spider

在 quote_spider/spiders/quotes.py 新建文件:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
import scrapy


class QuotesSpider(scrapy.Spider):
name = "quotes"
allowed_domains = ["quotes.toscrape.com"]

async def start(self):
yield scrapy.Request("https://quotes.toscrape.com/", self.parse)

def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}

运行:

1
scrapy crawl quotes

终端里能看到 quote 数据,就说明抓取和解析跑通了。

网上很多旧教程会写:

1
start_urls = ["https://quotes.toscrape.com/"]

start_urls 在简单项目里仍能用。这里采用新版官方教程中的 async def start(),以后要调整入口请求时可以直接修改这个方法。

保存成文件

先别急着写数据库。入门阶段最适合先导出 JSONL:

1
scrapy crawl quotes -O quotes.jsonl

-O 是覆盖输出,-o 是追加输出。开发调试时我更喜欢 -O,不容易把旧数据混进去。

JSONL 的好处是每行一条数据:

1
{"text": "The world as we have created it...", "author": "Albert Einstein", "tags": ["change", "deep-thoughts", "thinking", "world"]}

后面要导入 MySQL、MongoDB,或者丢给脚本二次处理,都方便。

把下一页也跟上

先在 shell 里看下一页链接:

1
response.css("li.next a::attr(href)").get()

返回的是:

1
/page/2/

改 spider:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
import scrapy


class QuotesSpider(scrapy.Spider):
name = "quotes"
allowed_domains = ["quotes.toscrape.com"]

async def start(self):
yield scrapy.Request("https://quotes.toscrape.com/", self.parse)

def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}

next_page = response.css("li.next a::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)

这里用 response.follow(),别自己拼域名。它会把 /page/2/ 转成完整 URL,并交给 Scrapy 调度。

再次运行:

1
scrapy crawl quotes -O quotes.jsonl

看到日志里出现 /page/2/、/page/3/,说明翻页已经跑起来了。

什么时候需要 Item

前面直接 yield dict 没问题。项目小的时候,这样最快。

但如果字段多了,建议在 items.py 里写清楚:

1
2
3
4
5
6
7
import scrapy


class QuoteItem(scrapy.Item):
text = scrapy.Field()
author = scrapy.Field()
tags = scrapy.Field()

spider 里改成:

1
2
3
4
5
6
7
8
9
10
from quote_spider.items import QuoteItem


def parse(self, response):
for quote in response.css("div.quote"):
yield QuoteItem(
text=quote.css("span.text::text").get(),
author=quote.css("small.author::text").get(),
tags=quote.css("div.tags a.tag::text").getall(),
)

Item 把字段集中定义在一处。后面加 Pipeline 或保存到数据库时,更容易检查字段名。

Pipeline 放清洗逻辑

Spider 里别塞太多处理逻辑。Spider 负责抓,Pipeline 负责处理。

比如去掉 quote 两边的中文引号,顺手清一下空标签:

1
2
3
4
5
6
class CleanQuotePipeline:
def process_item(self, item, spider):
item["text"] = item["text"].strip("“”")
item["author"] = item["author"].strip()
item["tags"] = [tag.strip() for tag in item["tags"] if tag.strip()]
return item

在 settings.py 打开:

1
2
3
ITEM_PIPELINES = {
"quote_spider.pipelines.CleanQuotePipeline": 300,
}

后面的数字是执行顺序,数字越小越早执行。比如你可以先清洗,再去重,最后写数据库。

抓取时常见的问题

数据是不是接口返回的

现在很多网站都是前后端分离。直接请求页面 HTML,可能只有空壳,真正数据来自接口。

先打开浏览器开发者工具,在 Network 里找 JSON 接口。如果接口可以直接请求,仍然可以用 Scrapy 抓取。

频率有没有太高

开发阶段建议先保守一点:

1
2
3
CONCURRENT_REQUESTS = 8
DOWNLOAD_DELAY = 1
ROBOTSTXT_OBEY = True

别一开始就把并发开很大。我曾因为抓取过快被封过两个账号,先确认目标站点的规则,再逐步调整频率。

日志是不是太吵

调试时可以看详细日志,跑稳定后再降级:

1
LOG_LEVEL = "INFO"

如果要写到文件:

1
LOG_FILE = "quotes.log"

数据会不会重复

Scrapy 默认会对请求 URL 去重,但不会帮你判断两条 Item 是否重复。如果一个详情页从多个入口都能进去,就要在 Pipeline 里用唯一键去重,比如详情页 URL、业务 ID 或内容 hash。