feat(monitor): 抖音接入博主监控
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s

上游爬虫本身不缺抖音能力(三模式、四项指标、二级评论都与小红书对等、指标还是同名同列),
缺的全在监控层的适配。这次把「平台之间不一样」的管子集中到一个新模块,再把散落的
xhs 硬编码接上去。

* 新增 api/monitor/adapters.py:产物目录名、jsonl 字段别名、目标链接形态与正则、
  通知链接模板。不放进 platforms.py 是因为那个模块被 describe_all() 整个序列化进
  /api/config/platforms 交给前端,塞进正则和目录名会让爬虫内部细节漏进 API 载荷。
  代价是两个注册表可能漂移,用一条测试钉住「声明接通就必须有适配器」。
* 两个必须知道的坑,都在这版里处理掉了:
  1) 抖音的平台 id 是 dy,而 store 把产物写在 douyin/ 下(store/douyin/_store_impl.py:47)。
     不改就是 ingest 一个文件都读不到 —— 不报错,只是 0 条,然后被冒充成「疑似登录失效」。
  2) 抖音的作品没有 note_id(叫 aweme_id)、评论也用 aweme_id 指作品。ingest 第一步是
     `if not note_id: continue`,不映射就逐条全丢。
  另外抖音顶层评论的 parent_comment_id 是字符串 "0",归一成空串,免得前端多出悬空的父节点。
* 顺带把「东西抓到了、只是没落在期望目录里」单独识别出来。这类故障的现象和登录失效
  一模一样,按登录失效报会把人指去查完全错误的方向。
* 修两个既有 bug(今天只有小红书所以无害,加抖音就踩响):
  - service.py update_task 换目标时漏传 task.platform,回落到默认小红书
  - scheduler.py 取 cookie 没传 platform,抖音任务会读着小红书那份 cookie 不动
* 行为变更(已与用户确认):cookie 闸门改成「没 cookie 且没开 CDP」才跳过。
  CDP 模式下登录态来自被接管的浏览器,粘不粘 cookie 由不得它决定;不放行的话,
  选了「接管已有 Chrome」却没粘 cookie 的用户会看到任务永远不触发,而且不报错。
  副作用是开启了 CDP 的小红书任务也不再被该闸门拦住 —— 语义上是对的。
* 目标输入框的示例链接与措辞改由能力矩阵提供(notes_label 抖音说「作品」、小红书说
  「笔记」;「建议只填纯 ID」是小红书专属劝告,抖音链接不带令牌,不再显示)。

测试 +22 条(858 通过),其中最关键的是「抖音作品/评论不被静默丢弃」与「产物目录名
不等于平台 id」两条 —— 都是把最难查的失败模式钉死在回归网里。

注意:抖音这条路的**端到端尚未验证**,需要一份可用的抖音登录态(CDP 那台 Chrome 里
登录,或导出一份 cookie)。单测覆盖的是解析与入库,真实抓取还没跑过。
This commit is contained in:
2026-10-10 14:55:28 +08:00
parent e348de48d3
commit 06718a1351
17 changed files with 863 additions and 105 deletions
+233
View File
@@ -0,0 +1,233 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/adapters.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""平台适配:两个平台之间**不一样**的那些管子。
监控层的大部分是平台中立的 —— 调度、入库、差分、报表、封面缓存、通知发送都与平台无关。
真正随平台变化的只有四样东西:
1. 爬虫把产物**落在哪个目录**(这里有个坑,见 ``artifact_dir``)
2. jsonl 里**字段叫什么**(抖音的作品没有 ``note_id``,叫 ``aweme_id``)
3. **目标链接**长什么样(怎么拼、怎么从链接里抠出 id)
4. 通知里的作品链接怎么拼
集中在这里,是为了让「加一个平台」变成在一处补一份数据,而不是去五个文件里找硬编码。
**为什么不放进 platforms.py**:那个模块被 ``describe_all()`` 整个序列化进
``GET /api/config/platforms`` 交给前端(连 ``**capability`` 一起),把正则、目录名、字段别名
塞进去会让爬虫的内部细节漏进 API 载荷,也会让「改适配」有动到接口形状的风险。
分工与既有的 schedule.py(算术)↔ scheduler.py(循环)一致。
"""
import re
from dataclasses import dataclass
from typing import Any, Dict, Mapping, Pattern, Tuple
from .platforms import PLATFORM_XHS
PLATFORM_DY = "dy"
def _first_cover(record: Dict[str, Any], fields: Tuple[str, ...]) -> str:
"""封面地址:取第一个非空字段,再取逗号分隔的第一段。
一条规则同时适配两边,所以不需要 per-platform 的函数:
小红书的 ``image_list`` 是 ``"url1,url2,..."``(要切第一段),
抖音的 ``cover_url`` 本身就是单个地址(切了等于没切)。
"""
for name in fields:
raw = record.get(name)
if raw:
return str(raw).split(",")[0].strip()
return ""
@dataclass(frozen=True)
class PlatformAdapter:
"""一个平台的全部「管子」。
字段别名的方向是**规范名 -> 该平台 jsonl 里的键**,读作「我们的列 ← 他们的键」。
"""
# 爬虫落盘用的目录名。**不等于平台 id**:抖音的平台 id 是 ``dy`` 而目录是 ``douyin``。
# 这不是笔误,是上游 store 里写死的(store/douyin/_store_impl.py:47)。改错这里的
# 后果是 ingest 一个文件都找不到 —— 它不会报错,只会落进「没抓到数据」分支,
# 然后被误报成「疑似登录失效」。
artifact_dir: str
web_base: str
creator_path: str
note_path: str
# 从链接里抠 id。是元组而不是单个正则,因为同一个平台可能有多种链接形态
# (抖音的作品链接还带 ?modal_id= 那种),按顺序试,第一个匹配的胜出。
# 每个正则必须恰好有一个捕获组。
creator_url_res: Tuple[Pattern, ...]
note_url_res: Tuple[Pattern, ...]
# 也允许直接粘贴裸 id —— 但两边的 id 形状不同,所以分开。
creator_bare_re: Pattern
note_bare_re: Pattern
# 短链(v.douyin.com 这种)无法在不发请求的情况下还原出 id,解析时单独报错,
# 好过存一个聚不出目标的值进去。
short_link_hosts: Tuple[str, ...]
note_fields: Mapping[str, str]
comment_fields: Mapping[str, str]
cover_fields: Tuple[str, ...]
def note_url(self, note_id: str) -> str:
"""作品的可点击链接。拼法与监控目标的链接是同一个形状 —— 通知里给的就是
人能直接点开看的那一个。"""
return f"{self.web_base}{self.note_path}/{note_id}"
def note_field(self, record: Dict[str, Any], name: str) -> Any:
"""按规范名读作品记录里的原始值(没有就是 None)。"""
return record.get(self.note_fields.get(name, name))
def comment_field(self, record: Dict[str, Any], name: str) -> Any:
return record.get(self.comment_fields.get(name, name))
def cover(self, record: Dict[str, Any]) -> str:
return _first_cover(record, self.cover_fields)
def parent_comment_id(self, record: Dict[str, Any]) -> str:
"""父评论 id,顶层评论一律归一成空串。
抖音顶层评论的 ``reply_id`` 是字符串 ``"0"``,小红书是 ``""`` —— 把 "0" 原样
存进去,前端就会多出一堆指向不存在的父评论的边。
"""
raw = self.comment_field(record, "parent_comment_id")
if raw is None:
return ""
raw = str(raw).strip()
return "" if raw in ("", "0") else raw
# 小红书 id 是 24 位 hex,允许稍宽一点,让格式变化退化成「仍然接受」而不是「拒绝」。
_XHS_BARE_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
XHS = PlatformAdapter(
artifact_dir="xhs",
web_base="https://www.xiaohongshu.com",
creator_path="/user/profile",
note_path="/explore",
creator_url_res=(re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)"),
),
creator_bare_re=_XHS_BARE_RE,
note_bare_re=_XHS_BARE_RE,
short_link_hosts=(),
note_fields={
"note_id": "note_id",
"title": "title",
"note_url": "note_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "type",
"published_at": "time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "note_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("image_list",),
)
# 抖音的 id 形状与小红书完全不同(见 media_platform/douyin/help.py:101-164):
# 作品 aweme_id 纯数字,如 7525082444551310602
# 博主 sec_user_id 形如 MS4wLjABAAAA...,含 - 和 _,**变长**(实测样本 55 字符,更长的也常见),
# 而小红书那条裸 id 规则封顶 64 —— 所以两条规则必须分开,否则长一点的 sec_uid
# 会被拒,表现为「粘贴了一个完全正确的链接却说无法识别」。
# 另外抖音**不需要 xsec_token**,裸链接就能用,比小红书简单。
DY = PlatformAdapter(
artifact_dir="douyin",
web_base="https://www.douyin.com",
creator_path="/user",
note_path="/video",
creator_url_res=(re.compile(r"douyin\.com/user/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"douyin\.com/video/(\d+)"),
# 带 modal_id 的链接:在别人主页或搜索结果里点开视频就是这个形态。
re.compile(r"[?&]modal_id=(\d+)"),
),
# 用长度而不是前缀来区分两者:sec_uid 是 20 字符以上的变长串,作品 id 是 19 位数字。
# 用前缀(MS4wLjABAAAA)更精确,但上游的 parse_creator_info_from_url 对裸 id 一律
# 照单全收,万一有别的前缀就会被我这里挡掉 —— 门槛设在长度上,两边都放得进,
# 又不会把 19 位的作品号误当成博主。
creator_bare_re=re.compile(r"^[A-Za-z0-9_-]{20,128}$"),
note_bare_re=re.compile(r"^\d{8,25}$"),
short_link_hosts=("v.douyin.com",),
note_fields={
"note_id": "aweme_id",
"title": "title",
"note_url": "aweme_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "aweme_type",
"published_at": "create_time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "aweme_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("cover_url",),
)
ADAPTERS: Dict[str, PlatformAdapter] = {
PLATFORM_XHS: XHS,
PLATFORM_DY: DY,
}
class UnknownPlatformError(ValueError):
"""平台还没有适配器。"""
def adapter(platform: str) -> PlatformAdapter:
try:
return ADAPTERS[platform]
except KeyError as exc:
raise UnknownPlatformError(f"平台 {platform} 还没有适配器") from exc
def has_adapter(platform: str) -> bool:
return platform in ADAPTERS
def artifact_dir(platform: str) -> str:
"""该平台的爬虫会把 jsonl 落在哪个子目录下。
``runner`` 用它判断产物是否真的出现过,``ingest`` 用它定位文件 —— 两处必须用
同一个值,否则会出现「文件在,但两边找的目录不是同一个」这种最难查的错。
"""
return adapter(platform).artifact_dir
+94 -25
View File
@@ -45,6 +45,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .platforms import PLATFORM_XHS
from .models import (
EVENT_AUTH_FAILURE,
@@ -171,8 +172,11 @@ def find_run_files(
Glob rather than reconstructing the name: both the crawler type and the date
are runtime-dependent. Returns lists because a crawl crossing midnight
produces one file per day.
``platform`` 是**监控层的平台 id**,而爬虫落盘的目录名未必同名(抖音的 id 是
``dy``、目录是 ``douyin``),所以这里经 adapters 解析 —— 调用方不必知道这个差异。
"""
jsonl_dir = out_dir / platform / "jsonl"
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
if not jsonl_dir.is_dir():
return [], []
@@ -182,6 +186,30 @@ def find_run_files(
)
def _misplaced_output_dirs(out_dir: Path, expected: str) -> List[str]:
"""在 out_dir 下找「有产物、但目录名不是期望的那个」的目录。
这是专门为**最难查的那类故障**准备的:产物目录名与平台对不上时,ingest 一个文件
都找不到,现象和「登录态失效」一模一样 —— 而实际上登录好好的、数据也抓到了,
只是没人去对的地方读。上游哪天改了 store 的目录名,这里能直接把实情说出来。
"""
found = []
try:
children = list(out_dir.iterdir())
except OSError:
return found
for child in children:
if child.name == expected or not child.is_dir():
continue
try:
if any(child.glob("jsonl/*_contents_*.jsonl")):
found.append(child.name)
except OSError:
continue
return sorted(found)
async def _emit(
session: AsyncSession,
run: MonitorRun,
@@ -271,13 +299,18 @@ async def _ingest_notes(
run: MonitorRun,
records: List[Dict[str, Any]],
is_baseline: bool,
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert notes, write metric snapshots, and emit new-note/delta events."""
"""Upsert notes, write metric snapshots, and emit new-note/delta events.
记录里的字段一律经 ``adapter`` 读。抖音的作品没有 ``note_id``(叫 ``aweme_id``),
按名字硬取的话每条记录都会在下面第一行被 continue 掉 —— 一条都不报错地全丢。
"""
now = get_current_timestamp()
new_count = 0
for record in records:
note_id = record.get("note_id")
note_id = adapter.note_field(record, "note_id")
if not note_id:
continue
@@ -288,21 +321,20 @@ async def _ingest_notes(
)
)
title = (record.get("title") or "")[:500]
raw_images = record.get("image_list") or ""
cover = raw_images.split(",")[0] if raw_images else ""
title = (adapter.note_field(record, "title") or "")[:500]
cover = adapter.cover(record)
if note is None:
note = MonitorNote(
task_id=run.task_id,
note_id=note_id,
title=title,
note_url=record.get("note_url") or "",
note_url=adapter.note_field(record, "note_url") or "",
cover=cover,
creator_hash=record.get("creator_hash") or "",
creator_name=record.get("nickname") or "",
source_kind=record.get("type") or "",
published_at=_as_int(record.get("time")),
creator_hash=adapter.note_field(record, "creator_hash") or "",
creator_name=adapter.note_field(record, "creator_name") or "",
source_kind=adapter.note_field(record, "source_kind") or "",
published_at=_as_int(adapter.note_field(record, "published_at")),
first_seen_run_id=run.id,
first_seen_at=now,
last_seen_run_id=run.id,
@@ -330,8 +362,9 @@ async def _ingest_notes(
if cover:
note.cover = cover
# 昵称也要刷:作者改昵称是常事,只在首次入库写一次会一直显示旧的。
if record.get("nickname"):
note.creator_name = record["nickname"]
creator_name = adapter.note_field(record, "creator_name")
if creator_name:
note.creator_name = creator_name
note.last_seen_run_id = run.id
note.last_seen_at = now
@@ -429,14 +462,18 @@ async def _ingest_comments(
records: List[Dict[str, Any]],
is_baseline: bool,
previous_run_started_at: Optional[int],
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert comments and emit events for ones never seen before."""
"""Upsert comments and emit events for ones never seen before.
与作品同理,评论记录也要经 ``adapter`` 读:抖音的评论用 ``aweme_id`` 指作品。
"""
now = get_current_timestamp()
new_count = 0
for record in records:
comment_id = record.get("comment_id")
note_id = record.get("note_id")
comment_id = adapter.comment_field(record, "comment_id")
note_id = adapter.comment_field(record, "note_id")
if not comment_id or not note_id:
continue
@@ -450,19 +487,22 @@ async def _ingest_comments(
if exists is not None:
continue
create_time = _as_int(record.get("create_time"))
create_time = _as_int(adapter.comment_field(record, "create_time"))
session.add(
MonitorComment(
task_id=run.task_id,
note_id=note_id,
comment_id=comment_id,
content=(record.get("content") or "")[:2000],
nickname=record.get("nickname") or "",
creator_hash=record.get("creator_hash") or "",
content=(adapter.comment_field(record, "content") or "")[:2000],
nickname=adapter.comment_field(record, "creator_name") or "",
creator_hash=adapter.comment_field(record, "creator_hash") or "",
create_time=create_time,
like_count=parse_count(record.get("like_count")),
sub_comment_count=_as_int(record.get("sub_comment_count")) or 0,
parent_comment_id=record.get("parent_comment_id") or "",
like_count=parse_count(adapter.comment_field(record, "like_count")),
sub_comment_count=_as_int(
adapter.comment_field(record, "sub_comment_count")
)
or 0,
parent_comment_id=adapter.parent_comment_id(record),
first_seen_run_id=run.id,
first_seen_at=now,
)
@@ -523,6 +563,9 @@ async def ingest_run(
)
return IngestResult(status=RUN_FAILED, error=run.error_message)
adapter = adapters.adapter(task.platform)
subdir = adapters.artifact_dir(task.platform)
contents_paths, comment_paths = find_run_files(out_dir, task.platform)
contents = [record for path in contents_paths for record in _read_jsonl(path)]
comments = [record for path in comment_paths for record in _read_jsonl(path)]
@@ -538,6 +581,32 @@ async def ingest_run(
if not contents:
run.status = RUN_PARTIAL
# 先排除「东西抓到了,只是没落在我们找的那个目录里」。这种故障的现象和登录失效
# 一模一样,但登录其实是好的 —— 按登录失效报会把人指到完全错的方向去查。
misplaced = _misplaced_output_dirs(out_dir, subdir)
if misplaced:
run.error_message = (
f"crawler wrote into {misplaced} but platform {task.platform} "
f"expects {subdir}"
)
await _emit(
session,
run,
EVENT_NO_DATA,
f"采集产物目录与平台不匹配(实际 {misplaced}、期望 {subdir}),本次未读到任何作品",
severity="error",
payload={
"out_dir": str(out_dir),
"expected": subdir,
"found": misplaced,
},
)
return IngestResult(
status=RUN_PARTIAL,
error=run.error_message,
comments_fetched=len(comments),
)
# Blaming the cookie is only honest if nothing else is authenticating.
# A sibling task that just succeeded proves the login works, so the
# fault is with this target (bad/expired per-creator token, an empty
@@ -587,10 +656,10 @@ async def ingest_run(
comments_fetched=len(comments),
is_baseline=is_baseline,
)
result.new_notes = await _ingest_notes(session, run, contents, is_baseline)
result.new_notes = await _ingest_notes(session, run, contents, is_baseline, adapter)
if task.enable_comments:
result.new_comments = await _ingest_comments(
session, run, comments, is_baseline, previous_started_at
session, run, comments, is_baseline, previous_started_at, adapter
)
run.new_notes = result.new_notes
+4 -1
View File
@@ -36,6 +36,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .models import (
EVENT_AUTH_FAILURE,
EVENT_NEW_NOTE,
@@ -164,7 +165,9 @@ async def build_run_message(
payload = _load_payload(event.payload_json)
title = payload.get("title") or event.target_id
note_id = payload.get("note_id") or event.target_id
url = f"https://www.xiaohongshu.com/explore/{note_id}"
# 链接形状按平台来。抖音的作品是 /video/{id},写死小红书域名的话,
# 群里点进去会是一个 404 —— 而这正是通知唯一要它干的事。
url = adapters.adapter(task.platform).note_url(note_id)
lines.append(f"> [{title}]({url})")
if len(new_notes) > 10:
lines.append(f"> …等共 {len(new_notes)} 篇")
+24 -5
View File
@@ -29,12 +29,13 @@ Two distinct things are recorded here, and conflating them would be misleading:
platform modules, not assumed -- all seven implement search/detail/creator;
the real differences are in which interaction metrics they capture.
* ``monitor_wired`` says whether the *monitoring layer* has been hooked up. It
currently covers only Xiaohongshu: ``runner.py`` pins the platform,
``ingest.py`` reads a fixed ``xhs/jsonl`` directory, and ``service.py`` only
parses Xiaohongshu target URLs.
covers Xiaohongshu and Douyin. The parts where those two differ -- which
directory the crawler writes into, what the jsonl fields are called, what a
target URL looks like -- live in ``adapters.py``; the rest of the layer is
platform-neutral.
A platform can therefore be fully crawlable by upstream and still not usable for
monitoring, which is exactly the state of the other six today.
monitoring, which is exactly the state of the other five today.
"""
from typing import Any, Dict, List, Optional
@@ -61,13 +62,31 @@ PLATFORM_CAPABILITIES: Dict[str, Dict[str, Any]] = {
"comment_levels": 2,
"media": True,
"monitor_wired": True,
"target_hints": {
"creator": "https://www.xiaohongshu.com/user/profile/5f58bd990000000001003753",
"note": "https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=...",
"creator_label": "博主主页",
"note_label": "笔记",
# 只有小红书的链接带会过期的 xsec_token。抖音的链接不带令牌,永久有效,
# 那句「建议只填纯 ID」的劝告对它没有意义。
"token_expires": True,
},
},
"dy": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": False,
"monitor_wired": True,
# 用户可见的示例链接(前端的目标输入框用它做 placeholder)。放这里是因为
# 它属于「这个平台长什么样」的能力描述;真正干活的管子(正则、目录名、
# 字段别名)在 adapters.py。
"target_hints": {
"creator": "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE",
"note": "https://www.douyin.com/video/7525082444551310602",
"creator_label": "博主主页",
"note_label": "作品",
},
},
"ks": {
"crawler_modes": ["search", "detail", "creator"],
+20 -16
View File
@@ -38,7 +38,7 @@ from ..schemas import (
SaveDataOptionEnum,
)
from ..services import crawler_manager
from . import app_settings, covers, notify
from . import adapters, app_settings, covers, notify
from .db import get_session
from .ingest import IngestResult, ingest_run
from .models import (
@@ -68,30 +68,33 @@ _PLATFORM_ENUM = {
"zhihu": PlatformEnum.ZHIHU,
}
_XHS_WEB_BASE = "https://www.xiaohongshu.com"
_CREATOR_PATH = "/user/profile"
_NOTE_PATH = "/explore"
# Timeout used when the caller does not care; tasks carry their own.
DEFAULT_RUN_TIMEOUT_SECONDS = 3600
def build_target_url(value: str, kind: str) -> str:
def build_target_url(value: str, kind: str, platform: str) -> str:
"""Turn a stored target into a URL the crawler's parser accepts.
Always emits a full URL rather than a bare id: the XHS parser accepts a bare
24-hex id only, so the URL form is the safer universal input. The
``xsec_token`` is appended when present but is deliberately optional -- it
expires, and the id alone is what keeps a long-running task alive.
Always emits a full URL rather than a bare id: both platforms' parsers accept
a bare id only in a narrower form, so the URL is the safer universal input.
The shape itself is platform-specific and comes from ``adapters``.
"""
path = _CREATOR_PATH if kind == MODE_CREATOR else _NOTE_PATH
return f"{_XHS_WEB_BASE}{path}/{value}"
spec = adapters.adapter(platform)
path = spec.creator_path if kind == MODE_CREATOR else spec.note_path
return f"{spec.web_base}{path}/{value}"
def build_target_urls(mode: str, targets: Iterable[MonitorTarget]) -> List[str]:
def build_target_urls(
mode: str, targets: Iterable[MonitorTarget], platform: str
) -> List[str]:
"""存储的目标 -> 爬虫接受的 URL。
``xsec_token`` 只有小红书有,而且是会过期的刷新令牌 —— 有就带上,没有就算了。
抖音恒为空,所以这一段对它是天然的 no-op,不需要平台分支。
"""
urls = []
for target in targets:
url = build_target_url(target.external_id, target.kind)
url = build_target_url(target.external_id, target.kind, platform)
if target.xsec_token:
url = f"{url}?xsec_token={target.xsec_token}"
if target.xsec_source:
@@ -159,7 +162,7 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
raise ValueError(f"Monitor task {task_id} has no enabled targets")
platform = task.platform
urls = build_target_urls(task.mode, targets)
urls = build_target_urls(task.mode, targets, platform)
cookie = await get_cookie(session, platform)
strategy = await _strategy_settings(session, platform)
# System-wide switch. On a headless server the crawler must attach to the
@@ -249,7 +252,8 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
if run is None or task is None:
raise ValueError(f"Run {run_id} or task {task_id} vanished during execution")
if exit_code == -1 and not (out_dir / "xhs").exists():
# 目录名按平台解析 —— 抖音的平台 id 是 dy 而产物目录是 douyin,写死就永远判不准。
if exit_code == -1 and not (out_dir / adapters.artifact_dir(task.platform)).exists():
# run_and_wait returns -1 when the process could not start or timed out.
run.status = RUN_TIMEOUT
run.finished_at = get_current_timestamp()
+22 -14
View File
@@ -67,8 +67,9 @@ class MonitorScheduler:
def __init__(self) -> None:
self._loop_task: Optional[asyncio.Task] = None
self._stopping = asyncio.Event()
# Avoids logging "no cookie" on every single tick.
self._warned_no_cookie = False
# Avoids logging "no cookie" on every single tick. Per platform, because
# warning once for Xiaohongshu must not silence the warning for Douyin.
self._warned_no_cookie: set = set()
async def start(self) -> None:
if self._loop_task is not None and not self._loop_task.done():
@@ -203,19 +204,26 @@ class MonitorScheduler:
if task is None:
return
# No cookie means every run would report an auth failure. Leave the
# task due rather than advancing: it starts working the moment the
# user pastes one.
cookie = await get_cookie(session)
# 没有 cookie 就跳过,是为了不让任务每轮白跑一趟出个认证失败。任务留在
# due 状态而不推进 —— 用户一粘上 cookie 它就能自己跑起来。
#
# **但开着 CDP 时必须放行**:那种模式下登录态来自被接管的那个浏览器,
# 粘不粘 cookie 根本轮不到它决定成败。不放行的话,选了「接管已有 Chrome」
# 却没粘 cookie 的用户会发现任务永远不被触发,而且什么错都不报。
cookie = await get_cookie(session, task.platform)
if not cookie:
if not self._warned_no_cookie:
print(
"[monitor.scheduler] no XHS cookie configured; "
"scheduled tasks will not run until one is set"
)
self._warned_no_cookie = True
return
self._warned_no_cookie = False
cdp_enabled = await app_settings.get_value(
session, "cdp_enabled", fallback=False
)
if not cdp_enabled:
if task.platform not in self._warned_no_cookie:
print(
f"[monitor.scheduler] no {task.platform} cookie configured; "
"scheduled tasks will not run until one is set or CDP is enabled"
)
self._warned_no_cookie.add(task.platform)
return
self._warned_no_cookie.discard(task.platform)
# Advance before running so a crash mid-run cannot cause an immediate
# re-fire, and so a long outage coalesces into a single run instead
+24 -19
View File
@@ -19,7 +19,6 @@
"""Task CRUD and dashboard queries for the monitoring layer."""
import asyncio
import re
from typing import Any, Dict, List, Optional
from urllib.parse import parse_qs, urlparse
@@ -28,7 +27,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import app_settings, covers, platforms, schedule
from . import adapters, app_settings, covers, platforms, schedule
from .db import get_session
from .platforms import PLATFORM_XHS
from .models import (
@@ -75,11 +74,8 @@ def _require_clock_fields(mode: str, hours: list, days: list) -> None:
if mode == schedule.MODE_WEEKLY and not days:
raise ValueError("按周调度至少要选一个星期")
_CREATOR_URL_RE = re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)")
_NOTE_URL_RE = re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)")
# XHS user ids and note ids are 24-char hex; allow a slightly wider range so a
# format change degrades into "still accepted" rather than "rejected".
_BARE_ID_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
# 各平台的链接形态、id 形状、短链域名都在 adapters.py —— 那里是「平台之间不一样」
# 的东西的唯一出处,所以这里不再留任何平台字面量。
class TargetParseError(ValueError):
@@ -95,27 +91,36 @@ def parse_target_input(
Storing the id separately from the token is what keeps a long-running task
alive: tokens expire, ids do not.
URL shapes are platform-specific. Only Xiaohongshu is wired, so anything else
is rejected here as well as at task creation -- parsing a Douyin link as if it
were a Xiaohongshu one would be worse than refusing it.
URL shapes are platform-specific and come from ``adapters``. A platform with
no adapter is rejected here as well as at task creation -- parsing a Douyin
link as if it were a Xiaohongshu one would be worse than refusing it.
"""
if platform != PLATFORM_XHS:
if not adapters.has_adapter(platform):
raise TargetParseError(f"暂不支持解析该平台({platform})的目标链接")
spec = adapters.adapter(platform)
raw = (value or "").strip()
if not raw:
raise TargetParseError("Empty target")
creator_mode = mode == MODE_CREATOR
expected = "博主主页" if creator_mode else "笔记"
external_id = ""
if raw.startswith("http") or "/" in raw:
# xhslink.com and other short links are not resolvable without a network
# round-trip, so only the direct profile/explore forms are supported.
match = _CREATOR_URL_RE.search(raw) if mode == MODE_CREATOR else _NOTE_URL_RE.search(raw)
if not match:
expected = "博主主页" if mode == MODE_CREATOR else "笔记"
if any(host in raw for host in spec.short_link_hosts):
# 短链要联网跳一次才知道指向谁,而这里没有网络可跳。明确拒绝好过存一个
# 解析不出 id 的值 —— 那会变成一个永远抓不到东西、还不报错的任务。
raise TargetParseError(f"{expected}短链无法解析,请粘贴完整链接:{raw}")
patterns = spec.creator_url_res if creator_mode else spec.note_url_res
for pattern in patterns:
match = pattern.search(raw)
if match:
external_id = match.group(1)
break
if not external_id:
raise TargetParseError(f"无法从链接中解析出{expected} ID:{raw}")
external_id = match.group(1)
elif _BARE_ID_RE.match(raw):
elif (spec.creator_bare_re if creator_mode else spec.note_bare_re).match(raw):
external_id = raw
else:
raise TargetParseError(f"无法识别的目标:{raw}")
@@ -260,7 +265,7 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
now = get_current_timestamp()
seen: set[str] = set()
for value in payload["targets"]:
parsed = parse_target_input(value, task.mode)
parsed = parse_target_input(value, task.mode, task.platform)
if parsed["external_id"] in seen:
continue
seen.add(parsed["external_id"])