feat(monitor): 抖音接入博主监控
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s

上游爬虫本身不缺抖音能力(三模式、四项指标、二级评论都与小红书对等、指标还是同名同列),
缺的全在监控层的适配。这次把「平台之间不一样」的管子集中到一个新模块,再把散落的
xhs 硬编码接上去。

* 新增 api/monitor/adapters.py:产物目录名、jsonl 字段别名、目标链接形态与正则、
  通知链接模板。不放进 platforms.py 是因为那个模块被 describe_all() 整个序列化进
  /api/config/platforms 交给前端,塞进正则和目录名会让爬虫内部细节漏进 API 载荷。
  代价是两个注册表可能漂移,用一条测试钉住「声明接通就必须有适配器」。
* 两个必须知道的坑,都在这版里处理掉了:
  1) 抖音的平台 id 是 dy,而 store 把产物写在 douyin/ 下(store/douyin/_store_impl.py:47)。
     不改就是 ingest 一个文件都读不到 —— 不报错,只是 0 条,然后被冒充成「疑似登录失效」。
  2) 抖音的作品没有 note_id(叫 aweme_id)、评论也用 aweme_id 指作品。ingest 第一步是
     `if not note_id: continue`,不映射就逐条全丢。
  另外抖音顶层评论的 parent_comment_id 是字符串 "0",归一成空串,免得前端多出悬空的父节点。
* 顺带把「东西抓到了、只是没落在期望目录里」单独识别出来。这类故障的现象和登录失效
  一模一样,按登录失效报会把人指去查完全错误的方向。
* 修两个既有 bug(今天只有小红书所以无害,加抖音就踩响):
  - service.py update_task 换目标时漏传 task.platform,回落到默认小红书
  - scheduler.py 取 cookie 没传 platform,抖音任务会读着小红书那份 cookie 不动
* 行为变更(已与用户确认):cookie 闸门改成「没 cookie 且没开 CDP」才跳过。
  CDP 模式下登录态来自被接管的浏览器,粘不粘 cookie 由不得它决定;不放行的话,
  选了「接管已有 Chrome」却没粘 cookie 的用户会看到任务永远不触发,而且不报错。
  副作用是开启了 CDP 的小红书任务也不再被该闸门拦住 —— 语义上是对的。
* 目标输入框的示例链接与措辞改由能力矩阵提供(notes_label 抖音说「作品」、小红书说
  「笔记」;「建议只填纯 ID」是小红书专属劝告,抖音链接不带令牌,不再显示)。

测试 +22 条(858 通过),其中最关键的是「抖音作品/评论不被静默丢弃」与「产物目录名
不等于平台 id」两条 —— 都是把最难查的失败模式钉死在回归网里。

注意:抖音这条路的**端到端尚未验证**,需要一份可用的抖音登录态(CDP 那台 Chrome 里
登录,或导出一份 cookie)。单测覆盖的是解析与入库,真实抓取还没跑过。
This commit is contained in:
2026-10-10 14:55:28 +08:00
parent e348de48d3
commit 06718a1351
17 changed files with 863 additions and 105 deletions
+233
View File
@@ -0,0 +1,233 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/adapters.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""平台适配:两个平台之间**不一样**的那些管子。
监控层的大部分是平台中立的 —— 调度、入库、差分、报表、封面缓存、通知发送都与平台无关。
真正随平台变化的只有四样东西:
1. 爬虫把产物**落在哪个目录**(这里有个坑,见 ``artifact_dir``)
2. jsonl 里**字段叫什么**(抖音的作品没有 ``note_id``,叫 ``aweme_id``)
3. **目标链接**长什么样(怎么拼、怎么从链接里抠出 id)
4. 通知里的作品链接怎么拼
集中在这里,是为了让「加一个平台」变成在一处补一份数据,而不是去五个文件里找硬编码。
**为什么不放进 platforms.py**:那个模块被 ``describe_all()`` 整个序列化进
``GET /api/config/platforms`` 交给前端(连 ``**capability`` 一起),把正则、目录名、字段别名
塞进去会让爬虫的内部细节漏进 API 载荷,也会让「改适配」有动到接口形状的风险。
分工与既有的 schedule.py(算术)↔ scheduler.py(循环)一致。
"""
import re
from dataclasses import dataclass
from typing import Any, Dict, Mapping, Pattern, Tuple
from .platforms import PLATFORM_XHS
PLATFORM_DY = "dy"
def _first_cover(record: Dict[str, Any], fields: Tuple[str, ...]) -> str:
"""封面地址:取第一个非空字段,再取逗号分隔的第一段。
一条规则同时适配两边,所以不需要 per-platform 的函数:
小红书的 ``image_list`` 是 ``"url1,url2,..."``(要切第一段),
抖音的 ``cover_url`` 本身就是单个地址(切了等于没切)。
"""
for name in fields:
raw = record.get(name)
if raw:
return str(raw).split(",")[0].strip()
return ""
@dataclass(frozen=True)
class PlatformAdapter:
"""一个平台的全部「管子」。
字段别名的方向是**规范名 -> 该平台 jsonl 里的键**,读作「我们的列 ← 他们的键」。
"""
# 爬虫落盘用的目录名。**不等于平台 id**:抖音的平台 id 是 ``dy`` 而目录是 ``douyin``。
# 这不是笔误,是上游 store 里写死的(store/douyin/_store_impl.py:47)。改错这里的
# 后果是 ingest 一个文件都找不到 —— 它不会报错,只会落进「没抓到数据」分支,
# 然后被误报成「疑似登录失效」。
artifact_dir: str
web_base: str
creator_path: str
note_path: str
# 从链接里抠 id。是元组而不是单个正则,因为同一个平台可能有多种链接形态
# (抖音的作品链接还带 ?modal_id= 那种),按顺序试,第一个匹配的胜出。
# 每个正则必须恰好有一个捕获组。
creator_url_res: Tuple[Pattern, ...]
note_url_res: Tuple[Pattern, ...]
# 也允许直接粘贴裸 id —— 但两边的 id 形状不同,所以分开。
creator_bare_re: Pattern
note_bare_re: Pattern
# 短链(v.douyin.com 这种)无法在不发请求的情况下还原出 id,解析时单独报错,
# 好过存一个聚不出目标的值进去。
short_link_hosts: Tuple[str, ...]
note_fields: Mapping[str, str]
comment_fields: Mapping[str, str]
cover_fields: Tuple[str, ...]
def note_url(self, note_id: str) -> str:
"""作品的可点击链接。拼法与监控目标的链接是同一个形状 —— 通知里给的就是
人能直接点开看的那一个。"""
return f"{self.web_base}{self.note_path}/{note_id}"
def note_field(self, record: Dict[str, Any], name: str) -> Any:
"""按规范名读作品记录里的原始值(没有就是 None)。"""
return record.get(self.note_fields.get(name, name))
def comment_field(self, record: Dict[str, Any], name: str) -> Any:
return record.get(self.comment_fields.get(name, name))
def cover(self, record: Dict[str, Any]) -> str:
return _first_cover(record, self.cover_fields)
def parent_comment_id(self, record: Dict[str, Any]) -> str:
"""父评论 id,顶层评论一律归一成空串。
抖音顶层评论的 ``reply_id`` 是字符串 ``"0"``,小红书是 ``""`` —— 把 "0" 原样
存进去,前端就会多出一堆指向不存在的父评论的边。
"""
raw = self.comment_field(record, "parent_comment_id")
if raw is None:
return ""
raw = str(raw).strip()
return "" if raw in ("", "0") else raw
# 小红书 id 是 24 位 hex,允许稍宽一点,让格式变化退化成「仍然接受」而不是「拒绝」。
_XHS_BARE_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
XHS = PlatformAdapter(
artifact_dir="xhs",
web_base="https://www.xiaohongshu.com",
creator_path="/user/profile",
note_path="/explore",
creator_url_res=(re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)"),
),
creator_bare_re=_XHS_BARE_RE,
note_bare_re=_XHS_BARE_RE,
short_link_hosts=(),
note_fields={
"note_id": "note_id",
"title": "title",
"note_url": "note_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "type",
"published_at": "time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "note_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("image_list",),
)
# 抖音的 id 形状与小红书完全不同(见 media_platform/douyin/help.py:101-164):
# 作品 aweme_id 纯数字,如 7525082444551310602
# 博主 sec_user_id 形如 MS4wLjABAAAA...,含 - 和 _,**变长**(实测样本 55 字符,更长的也常见),
# 而小红书那条裸 id 规则封顶 64 —— 所以两条规则必须分开,否则长一点的 sec_uid
# 会被拒,表现为「粘贴了一个完全正确的链接却说无法识别」。
# 另外抖音**不需要 xsec_token**,裸链接就能用,比小红书简单。
DY = PlatformAdapter(
artifact_dir="douyin",
web_base="https://www.douyin.com",
creator_path="/user",
note_path="/video",
creator_url_res=(re.compile(r"douyin\.com/user/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"douyin\.com/video/(\d+)"),
# 带 modal_id 的链接:在别人主页或搜索结果里点开视频就是这个形态。
re.compile(r"[?&]modal_id=(\d+)"),
),
# 用长度而不是前缀来区分两者:sec_uid 是 20 字符以上的变长串,作品 id 是 19 位数字。
# 用前缀(MS4wLjABAAAA)更精确,但上游的 parse_creator_info_from_url 对裸 id 一律
# 照单全收,万一有别的前缀就会被我这里挡掉 —— 门槛设在长度上,两边都放得进,
# 又不会把 19 位的作品号误当成博主。
creator_bare_re=re.compile(r"^[A-Za-z0-9_-]{20,128}$"),
note_bare_re=re.compile(r"^\d{8,25}$"),
short_link_hosts=("v.douyin.com",),
note_fields={
"note_id": "aweme_id",
"title": "title",
"note_url": "aweme_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "aweme_type",
"published_at": "create_time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "aweme_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("cover_url",),
)
ADAPTERS: Dict[str, PlatformAdapter] = {
PLATFORM_XHS: XHS,
PLATFORM_DY: DY,
}
class UnknownPlatformError(ValueError):
"""平台还没有适配器。"""
def adapter(platform: str) -> PlatformAdapter:
try:
return ADAPTERS[platform]
except KeyError as exc:
raise UnknownPlatformError(f"平台 {platform} 还没有适配器") from exc
def has_adapter(platform: str) -> bool:
return platform in ADAPTERS
def artifact_dir(platform: str) -> str:
"""该平台的爬虫会把 jsonl 落在哪个子目录下。
``runner`` 用它判断产物是否真的出现过,``ingest`` 用它定位文件 —— 两处必须用
同一个值,否则会出现「文件在,但两边找的目录不是同一个」这种最难查的错。
"""
return adapter(platform).artifact_dir
+94 -25
View File
@@ -45,6 +45,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .platforms import PLATFORM_XHS
from .models import (
EVENT_AUTH_FAILURE,
@@ -171,8 +172,11 @@ def find_run_files(
Glob rather than reconstructing the name: both the crawler type and the date
are runtime-dependent. Returns lists because a crawl crossing midnight
produces one file per day.
``platform`` 是**监控层的平台 id**,而爬虫落盘的目录名未必同名(抖音的 id 是
``dy``、目录是 ``douyin``),所以这里经 adapters 解析 —— 调用方不必知道这个差异。
"""
jsonl_dir = out_dir / platform / "jsonl"
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
if not jsonl_dir.is_dir():
return [], []
@@ -182,6 +186,30 @@ def find_run_files(
)
def _misplaced_output_dirs(out_dir: Path, expected: str) -> List[str]:
"""在 out_dir 下找「有产物、但目录名不是期望的那个」的目录。
这是专门为**最难查的那类故障**准备的:产物目录名与平台对不上时,ingest 一个文件
都找不到,现象和「登录态失效」一模一样 —— 而实际上登录好好的、数据也抓到了,
只是没人去对的地方读。上游哪天改了 store 的目录名,这里能直接把实情说出来。
"""
found = []
try:
children = list(out_dir.iterdir())
except OSError:
return found
for child in children:
if child.name == expected or not child.is_dir():
continue
try:
if any(child.glob("jsonl/*_contents_*.jsonl")):
found.append(child.name)
except OSError:
continue
return sorted(found)
async def _emit(
session: AsyncSession,
run: MonitorRun,
@@ -271,13 +299,18 @@ async def _ingest_notes(
run: MonitorRun,
records: List[Dict[str, Any]],
is_baseline: bool,
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert notes, write metric snapshots, and emit new-note/delta events."""
"""Upsert notes, write metric snapshots, and emit new-note/delta events.
记录里的字段一律经 ``adapter`` 读。抖音的作品没有 ``note_id``(叫 ``aweme_id``),
按名字硬取的话每条记录都会在下面第一行被 continue 掉 —— 一条都不报错地全丢。
"""
now = get_current_timestamp()
new_count = 0
for record in records:
note_id = record.get("note_id")
note_id = adapter.note_field(record, "note_id")
if not note_id:
continue
@@ -288,21 +321,20 @@ async def _ingest_notes(
)
)
title = (record.get("title") or "")[:500]
raw_images = record.get("image_list") or ""
cover = raw_images.split(",")[0] if raw_images else ""
title = (adapter.note_field(record, "title") or "")[:500]
cover = adapter.cover(record)
if note is None:
note = MonitorNote(
task_id=run.task_id,
note_id=note_id,
title=title,
note_url=record.get("note_url") or "",
note_url=adapter.note_field(record, "note_url") or "",
cover=cover,
creator_hash=record.get("creator_hash") or "",
creator_name=record.get("nickname") or "",
source_kind=record.get("type") or "",
published_at=_as_int(record.get("time")),
creator_hash=adapter.note_field(record, "creator_hash") or "",
creator_name=adapter.note_field(record, "creator_name") or "",
source_kind=adapter.note_field(record, "source_kind") or "",
published_at=_as_int(adapter.note_field(record, "published_at")),
first_seen_run_id=run.id,
first_seen_at=now,
last_seen_run_id=run.id,
@@ -330,8 +362,9 @@ async def _ingest_notes(
if cover:
note.cover = cover
# 昵称也要刷:作者改昵称是常事,只在首次入库写一次会一直显示旧的。
if record.get("nickname"):
note.creator_name = record["nickname"]
creator_name = adapter.note_field(record, "creator_name")
if creator_name:
note.creator_name = creator_name
note.last_seen_run_id = run.id
note.last_seen_at = now
@@ -429,14 +462,18 @@ async def _ingest_comments(
records: List[Dict[str, Any]],
is_baseline: bool,
previous_run_started_at: Optional[int],
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert comments and emit events for ones never seen before."""
"""Upsert comments and emit events for ones never seen before.
与作品同理,评论记录也要经 ``adapter`` 读:抖音的评论用 ``aweme_id`` 指作品。
"""
now = get_current_timestamp()
new_count = 0
for record in records:
comment_id = record.get("comment_id")
note_id = record.get("note_id")
comment_id = adapter.comment_field(record, "comment_id")
note_id = adapter.comment_field(record, "note_id")
if not comment_id or not note_id:
continue
@@ -450,19 +487,22 @@ async def _ingest_comments(
if exists is not None:
continue
create_time = _as_int(record.get("create_time"))
create_time = _as_int(adapter.comment_field(record, "create_time"))
session.add(
MonitorComment(
task_id=run.task_id,
note_id=note_id,
comment_id=comment_id,
content=(record.get("content") or "")[:2000],
nickname=record.get("nickname") or "",
creator_hash=record.get("creator_hash") or "",
content=(adapter.comment_field(record, "content") or "")[:2000],
nickname=adapter.comment_field(record, "creator_name") or "",
creator_hash=adapter.comment_field(record, "creator_hash") or "",
create_time=create_time,
like_count=parse_count(record.get("like_count")),
sub_comment_count=_as_int(record.get("sub_comment_count")) or 0,
parent_comment_id=record.get("parent_comment_id") or "",
like_count=parse_count(adapter.comment_field(record, "like_count")),
sub_comment_count=_as_int(
adapter.comment_field(record, "sub_comment_count")
)
or 0,
parent_comment_id=adapter.parent_comment_id(record),
first_seen_run_id=run.id,
first_seen_at=now,
)
@@ -523,6 +563,9 @@ async def ingest_run(
)
return IngestResult(status=RUN_FAILED, error=run.error_message)
adapter = adapters.adapter(task.platform)
subdir = adapters.artifact_dir(task.platform)
contents_paths, comment_paths = find_run_files(out_dir, task.platform)
contents = [record for path in contents_paths for record in _read_jsonl(path)]
comments = [record for path in comment_paths for record in _read_jsonl(path)]
@@ -538,6 +581,32 @@ async def ingest_run(
if not contents:
run.status = RUN_PARTIAL
# 先排除「东西抓到了,只是没落在我们找的那个目录里」。这种故障的现象和登录失效
# 一模一样,但登录其实是好的 —— 按登录失效报会把人指到完全错的方向去查。
misplaced = _misplaced_output_dirs(out_dir, subdir)
if misplaced:
run.error_message = (
f"crawler wrote into {misplaced} but platform {task.platform} "
f"expects {subdir}"
)
await _emit(
session,
run,
EVENT_NO_DATA,
f"采集产物目录与平台不匹配(实际 {misplaced}、期望 {subdir}),本次未读到任何作品",
severity="error",
payload={
"out_dir": str(out_dir),
"expected": subdir,
"found": misplaced,
},
)
return IngestResult(
status=RUN_PARTIAL,
error=run.error_message,
comments_fetched=len(comments),
)
# Blaming the cookie is only honest if nothing else is authenticating.
# A sibling task that just succeeded proves the login works, so the
# fault is with this target (bad/expired per-creator token, an empty
@@ -587,10 +656,10 @@ async def ingest_run(
comments_fetched=len(comments),
is_baseline=is_baseline,
)
result.new_notes = await _ingest_notes(session, run, contents, is_baseline)
result.new_notes = await _ingest_notes(session, run, contents, is_baseline, adapter)
if task.enable_comments:
result.new_comments = await _ingest_comments(
session, run, comments, is_baseline, previous_started_at
session, run, comments, is_baseline, previous_started_at, adapter
)
run.new_notes = result.new_notes
+4 -1
View File
@@ -36,6 +36,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .models import (
EVENT_AUTH_FAILURE,
EVENT_NEW_NOTE,
@@ -164,7 +165,9 @@ async def build_run_message(
payload = _load_payload(event.payload_json)
title = payload.get("title") or event.target_id
note_id = payload.get("note_id") or event.target_id
url = f"https://www.xiaohongshu.com/explore/{note_id}"
# 链接形状按平台来。抖音的作品是 /video/{id},写死小红书域名的话,
# 群里点进去会是一个 404 —— 而这正是通知唯一要它干的事。
url = adapters.adapter(task.platform).note_url(note_id)
lines.append(f"> [{title}]({url})")
if len(new_notes) > 10:
lines.append(f"> …等共 {len(new_notes)} 篇")
+24 -5
View File
@@ -29,12 +29,13 @@ Two distinct things are recorded here, and conflating them would be misleading:
platform modules, not assumed -- all seven implement search/detail/creator;
the real differences are in which interaction metrics they capture.
* ``monitor_wired`` says whether the *monitoring layer* has been hooked up. It
currently covers only Xiaohongshu: ``runner.py`` pins the platform,
``ingest.py`` reads a fixed ``xhs/jsonl`` directory, and ``service.py`` only
parses Xiaohongshu target URLs.
covers Xiaohongshu and Douyin. The parts where those two differ -- which
directory the crawler writes into, what the jsonl fields are called, what a
target URL looks like -- live in ``adapters.py``; the rest of the layer is
platform-neutral.
A platform can therefore be fully crawlable by upstream and still not usable for
monitoring, which is exactly the state of the other six today.
monitoring, which is exactly the state of the other five today.
"""
from typing import Any, Dict, List, Optional
@@ -61,13 +62,31 @@ PLATFORM_CAPABILITIES: Dict[str, Dict[str, Any]] = {
"comment_levels": 2,
"media": True,
"monitor_wired": True,
"target_hints": {
"creator": "https://www.xiaohongshu.com/user/profile/5f58bd990000000001003753",
"note": "https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=...",
"creator_label": "博主主页",
"note_label": "笔记",
# 只有小红书的链接带会过期的 xsec_token。抖音的链接不带令牌,永久有效,
# 那句「建议只填纯 ID」的劝告对它没有意义。
"token_expires": True,
},
},
"dy": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": False,
"monitor_wired": True,
# 用户可见的示例链接(前端的目标输入框用它做 placeholder)。放这里是因为
# 它属于「这个平台长什么样」的能力描述;真正干活的管子(正则、目录名、
# 字段别名)在 adapters.py。
"target_hints": {
"creator": "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE",
"note": "https://www.douyin.com/video/7525082444551310602",
"creator_label": "博主主页",
"note_label": "作品",
},
},
"ks": {
"crawler_modes": ["search", "detail", "creator"],
+20 -16
View File
@@ -38,7 +38,7 @@ from ..schemas import (
SaveDataOptionEnum,
)
from ..services import crawler_manager
from . import app_settings, covers, notify
from . import adapters, app_settings, covers, notify
from .db import get_session
from .ingest import IngestResult, ingest_run
from .models import (
@@ -68,30 +68,33 @@ _PLATFORM_ENUM = {
"zhihu": PlatformEnum.ZHIHU,
}
_XHS_WEB_BASE = "https://www.xiaohongshu.com"
_CREATOR_PATH = "/user/profile"
_NOTE_PATH = "/explore"
# Timeout used when the caller does not care; tasks carry their own.
DEFAULT_RUN_TIMEOUT_SECONDS = 3600
def build_target_url(value: str, kind: str) -> str:
def build_target_url(value: str, kind: str, platform: str) -> str:
"""Turn a stored target into a URL the crawler's parser accepts.
Always emits a full URL rather than a bare id: the XHS parser accepts a bare
24-hex id only, so the URL form is the safer universal input. The
``xsec_token`` is appended when present but is deliberately optional -- it
expires, and the id alone is what keeps a long-running task alive.
Always emits a full URL rather than a bare id: both platforms' parsers accept
a bare id only in a narrower form, so the URL is the safer universal input.
The shape itself is platform-specific and comes from ``adapters``.
"""
path = _CREATOR_PATH if kind == MODE_CREATOR else _NOTE_PATH
return f"{_XHS_WEB_BASE}{path}/{value}"
spec = adapters.adapter(platform)
path = spec.creator_path if kind == MODE_CREATOR else spec.note_path
return f"{spec.web_base}{path}/{value}"
def build_target_urls(mode: str, targets: Iterable[MonitorTarget]) -> List[str]:
def build_target_urls(
mode: str, targets: Iterable[MonitorTarget], platform: str
) -> List[str]:
"""存储的目标 -> 爬虫接受的 URL。
``xsec_token`` 只有小红书有,而且是会过期的刷新令牌 —— 有就带上,没有就算了。
抖音恒为空,所以这一段对它是天然的 no-op,不需要平台分支。
"""
urls = []
for target in targets:
url = build_target_url(target.external_id, target.kind)
url = build_target_url(target.external_id, target.kind, platform)
if target.xsec_token:
url = f"{url}?xsec_token={target.xsec_token}"
if target.xsec_source:
@@ -159,7 +162,7 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
raise ValueError(f"Monitor task {task_id} has no enabled targets")
platform = task.platform
urls = build_target_urls(task.mode, targets)
urls = build_target_urls(task.mode, targets, platform)
cookie = await get_cookie(session, platform)
strategy = await _strategy_settings(session, platform)
# System-wide switch. On a headless server the crawler must attach to the
@@ -249,7 +252,8 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
if run is None or task is None:
raise ValueError(f"Run {run_id} or task {task_id} vanished during execution")
if exit_code == -1 and not (out_dir / "xhs").exists():
# 目录名按平台解析 —— 抖音的平台 id 是 dy 而产物目录是 douyin,写死就永远判不准。
if exit_code == -1 and not (out_dir / adapters.artifact_dir(task.platform)).exists():
# run_and_wait returns -1 when the process could not start or timed out.
run.status = RUN_TIMEOUT
run.finished_at = get_current_timestamp()
+22 -14
View File
@@ -67,8 +67,9 @@ class MonitorScheduler:
def __init__(self) -> None:
self._loop_task: Optional[asyncio.Task] = None
self._stopping = asyncio.Event()
# Avoids logging "no cookie" on every single tick.
self._warned_no_cookie = False
# Avoids logging "no cookie" on every single tick. Per platform, because
# warning once for Xiaohongshu must not silence the warning for Douyin.
self._warned_no_cookie: set = set()
async def start(self) -> None:
if self._loop_task is not None and not self._loop_task.done():
@@ -203,19 +204,26 @@ class MonitorScheduler:
if task is None:
return
# No cookie means every run would report an auth failure. Leave the
# task due rather than advancing: it starts working the moment the
# user pastes one.
cookie = await get_cookie(session)
# 没有 cookie 就跳过,是为了不让任务每轮白跑一趟出个认证失败。任务留在
# due 状态而不推进 —— 用户一粘上 cookie 它就能自己跑起来。
#
# **但开着 CDP 时必须放行**:那种模式下登录态来自被接管的那个浏览器,
# 粘不粘 cookie 根本轮不到它决定成败。不放行的话,选了「接管已有 Chrome」
# 却没粘 cookie 的用户会发现任务永远不被触发,而且什么错都不报。
cookie = await get_cookie(session, task.platform)
if not cookie:
if not self._warned_no_cookie:
print(
"[monitor.scheduler] no XHS cookie configured; "
"scheduled tasks will not run until one is set"
)
self._warned_no_cookie = True
return
self._warned_no_cookie = False
cdp_enabled = await app_settings.get_value(
session, "cdp_enabled", fallback=False
)
if not cdp_enabled:
if task.platform not in self._warned_no_cookie:
print(
f"[monitor.scheduler] no {task.platform} cookie configured; "
"scheduled tasks will not run until one is set or CDP is enabled"
)
self._warned_no_cookie.add(task.platform)
return
self._warned_no_cookie.discard(task.platform)
# Advance before running so a crash mid-run cannot cause an immediate
# re-fire, and so a long outage coalesces into a single run instead
+24 -19
View File
@@ -19,7 +19,6 @@
"""Task CRUD and dashboard queries for the monitoring layer."""
import asyncio
import re
from typing import Any, Dict, List, Optional
from urllib.parse import parse_qs, urlparse
@@ -28,7 +27,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import app_settings, covers, platforms, schedule
from . import adapters, app_settings, covers, platforms, schedule
from .db import get_session
from .platforms import PLATFORM_XHS
from .models import (
@@ -75,11 +74,8 @@ def _require_clock_fields(mode: str, hours: list, days: list) -> None:
if mode == schedule.MODE_WEEKLY and not days:
raise ValueError("按周调度至少要选一个星期")
_CREATOR_URL_RE = re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)")
_NOTE_URL_RE = re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)")
# XHS user ids and note ids are 24-char hex; allow a slightly wider range so a
# format change degrades into "still accepted" rather than "rejected".
_BARE_ID_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
# 各平台的链接形态、id 形状、短链域名都在 adapters.py —— 那里是「平台之间不一样」
# 的东西的唯一出处,所以这里不再留任何平台字面量。
class TargetParseError(ValueError):
@@ -95,27 +91,36 @@ def parse_target_input(
Storing the id separately from the token is what keeps a long-running task
alive: tokens expire, ids do not.
URL shapes are platform-specific. Only Xiaohongshu is wired, so anything else
is rejected here as well as at task creation -- parsing a Douyin link as if it
were a Xiaohongshu one would be worse than refusing it.
URL shapes are platform-specific and come from ``adapters``. A platform with
no adapter is rejected here as well as at task creation -- parsing a Douyin
link as if it were a Xiaohongshu one would be worse than refusing it.
"""
if platform != PLATFORM_XHS:
if not adapters.has_adapter(platform):
raise TargetParseError(f"暂不支持解析该平台({platform})的目标链接")
spec = adapters.adapter(platform)
raw = (value or "").strip()
if not raw:
raise TargetParseError("Empty target")
creator_mode = mode == MODE_CREATOR
expected = "博主主页" if creator_mode else "笔记"
external_id = ""
if raw.startswith("http") or "/" in raw:
# xhslink.com and other short links are not resolvable without a network
# round-trip, so only the direct profile/explore forms are supported.
match = _CREATOR_URL_RE.search(raw) if mode == MODE_CREATOR else _NOTE_URL_RE.search(raw)
if not match:
expected = "博主主页" if mode == MODE_CREATOR else "笔记"
if any(host in raw for host in spec.short_link_hosts):
# 短链要联网跳一次才知道指向谁,而这里没有网络可跳。明确拒绝好过存一个
# 解析不出 id 的值 —— 那会变成一个永远抓不到东西、还不报错的任务。
raise TargetParseError(f"{expected}短链无法解析,请粘贴完整链接:{raw}")
patterns = spec.creator_url_res if creator_mode else spec.note_url_res
for pattern in patterns:
match = pattern.search(raw)
if match:
external_id = match.group(1)
break
if not external_id:
raise TargetParseError(f"无法从链接中解析出{expected} ID:{raw}")
external_id = match.group(1)
elif _BARE_ID_RE.match(raw):
elif (spec.creator_bare_re if creator_mode else spec.note_bare_re).match(raw):
external_id = raw
else:
raise TargetParseError(f"无法识别的目标:{raw}")
@@ -260,7 +265,7 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
now = get_current_timestamp()
seen: set[str] = set()
for value in payload["targets"]:
parsed = parse_target_input(value, task.mode)
parsed = parse_target_input(value, task.mode, task.platform)
if parsed["external_id"] in seen:
continue
seen.add(parsed["external_id"])
+1 -1
View File
@@ -287,7 +287,7 @@ set MC_PASSWORD=我的新密码 # Windows cmd
未接通的平台**可以选,但各页会显示明确的说明面板**,并且**创建任务会被直接拒绝**:
```
400 抖音的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。
400 B站的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。
```
而不是接受任务、然后让它永远跑不出数据 —— 那正是之前"博主主页解析失败被误报成登录失效"的同一种静默故障。
+31
View File
@@ -26,6 +26,37 @@ async def test_cmd_arg_crawler_max_notes_count():
config.CRAWLER_MAX_NOTES_COUNT = orig_notes
config.CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES = orig_comments
def test_douyin_monitor_command_uses_the_right_flags():
"""抖音监控任务拼出来的命令行。
与 runner 走的是同一条 _build_command 路径,所以这一条能守住「监控任务的参数
没拼错」—— 尤其是平台值必须是 dy(而不是 douyin),否则上游根本认不出平台。
"""
cm = CrawlerManager()
req = CrawlerStartRequest(
platform=PlatformEnum.DOUYIN,
login_type=LoginTypeEnum.COOKIE,
crawler_type=CrawlerTypeEnum.CREATOR,
creator_ids="https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X",
save_data_path="./data/monitor_runs/1/2",
enable_cdp_mode=True,
inject_all_cookies=True,
save_login_state=True,
max_notes_count=20,
max_comments_count=50,
)
cmd = cm._build_command(req)
idx = cmd.index("--platform")
assert cmd[idx + 1] == "dy"
idx = cmd.index("--type")
assert cmd[idx + 1] == "creator"
idx = cmd.index("--creator_id")
assert cmd[idx + 1] == "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X"
idx = cmd.index("--enable_cdp_mode")
assert cmd[idx + 1] == "true"
def test_crawler_manager_build_command():
cm = CrawlerManager()
+111
View File
@@ -78,6 +78,117 @@ class TestParseTargetInput:
with pytest.raises(TargetParseError):
parse_target_input("not a url at all !!", "creator")
# --- 抖音 -------------------------------------------------------------
# 链接形态由平台决定,所以每一个都要显式带上 "dy"。
def test_douyin_creator_url(self):
parsed = parse_target_input(
"https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
"?from_tab_name=main",
"creator",
"dy",
)
assert (
parsed["external_id"]
== "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
)
def test_douyin_video_url(self):
parsed = parse_target_input(
"https://www.douyin.com/video/7525082444551310602", "note", "dy"
)
assert parsed["external_id"] == "7525082444551310602"
def test_douyin_modal_id_url(self):
"""在别人主页或搜索结果里点开视频,拿到的就是带 modal_id 的链接。"""
parsed = parse_target_input(
"https://www.douyin.com/root/search/python?aid=b733a3b0&modal_id=7471165520058862848",
"note",
"dy",
)
assert parsed["external_id"] == "7471165520058862848"
def test_douyin_bare_sec_uid_is_accepted(self):
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
parsed = parse_target_input(sec_uid, "creator", "dy")
assert parsed["external_id"] == sec_uid
def test_douyin_bare_sec_uid_beyond_the_xhs_length_cap(self):
"""裸 id 的长度上限必须按平台分开。
小红书那条规则封顶 64 字符,而 sec_user_id 长过 64 是常态(实测样本 55,
但字段本身是变长的)。共用一条规则的话,长一点的 sec_uid 会被直接拒掉 ——
对用户来说就是「粘贴了一个完全正确的链接却报无法识别」。
"""
sec_uid = "MS4wLjABAAAA" + "aB3dEf6hIj9lMn2pQr5tUv8xYz1" * 3
assert len(sec_uid) > 64
parsed = parse_target_input(sec_uid, "creator", "dy")
assert parsed["external_id"] == sec_uid
# 同一条 id 拿小红书规则来解析会被拒 —— 这正是两条规则必须分开的原因。
with pytest.raises(TargetParseError):
parse_target_input(sec_uid, "creator", "xhs")
def test_douyin_bare_video_id(self):
parsed = parse_target_input("7525082444551310602", "note", "dy")
assert parsed["external_id"] == "7525082444551310602"
# 抖音不需要 xsec_token —— 和小红书不同,裸链接就能用。
assert parsed["xsec_token"] == ""
def test_douyin_short_link_is_rejected_with_a_reason(self):
"""短链要联网跳一次才知道指向谁。明确拒绝好过存一个永远抓不到东西的目标。"""
with pytest.raises(TargetParseError) as excinfo:
parse_target_input("https://v.douyin.com/drIPtQ_WPWY/", "note", "dy")
assert "短链" in str(excinfo.value)
def test_a_douyin_link_is_not_parsed_with_xhs_rules(self):
with pytest.raises(TargetParseError):
parse_target_input(
"https://www.douyin.com/video/7525082444551310602", "note", "xhs"
)
def test_an_xhs_link_is_not_parsed_for_douyin(self):
with pytest.raises(TargetParseError):
parse_target_input(NOTE_URL, "note", "dy")
def test_a_platform_without_an_adapter_is_rejected(self):
with pytest.raises(TargetParseError):
parse_target_input("whatever", "creator", "bili")
class TestTargetReplacement:
@pytest.mark.asyncio
async def test_replacing_targets_uses_the_tasks_own_platform(self, client):
"""改目标必须按任务**自己**的平台解析。
``update_task`` 原先漏传了 platform,解析回落到默认的小红书。只有小红书时
行为恰好正确,接上抖音就会拿小红书的正则去解析抖音链接 —— 建任务时对、
改任务时错,是最难注意到的那种不一致。
"""
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
created = await client.post(
"/api/monitor/tasks",
json={"name": "抖音", "mode": "creator", "platform": "dy", "targets": [sec_uid]},
)
assert created.status_code == 201
task_id = created.json()["id"]
updated = await client.patch(
f"/api/monitor/tasks/{task_id}",
json={"targets": [f"https://www.douyin.com/user/{sec_uid}"]},
)
assert updated.status_code == 200
tasks = (
await client.get("/api/monitor/tasks", params={"platform": "dy"})
).json()["tasks"]
target = tasks[0]["targets"][0]
assert target["external_id"] == sec_uid
assert target["raw_value"].startswith("https://www.douyin.com/user/")
class TestTaskCrud:
@pytest.mark.asyncio
+175 -2
View File
@@ -35,6 +35,7 @@ from sqlalchemy.pool import StaticPool
from tools.time_util import get_current_timestamp
from api.monitor import adapters
from api.monitor.ingest import describe_exit_code, ingest_run, parse_count
from api.monitor.models import (
EVENT_AUTH_FAILURE,
@@ -46,6 +47,7 @@ from api.monitor.models import (
EVENT_RUN_FAILED,
MODE_CREATOR,
MonitorBase,
MonitorComment,
MonitorEvent,
MonitorNote,
MonitorNoteMetric,
@@ -118,9 +120,14 @@ def _write_run_dir(
root: Path,
notes: List[Dict[str, Any]],
comments: Optional[List[Dict[str, Any]]] = None,
subdir: str = "xhs",
) -> Path:
"""Write a run's jsonl output in the crawler's own layout."""
jsonl_dir = root / "xhs" / "jsonl"
"""Write a run's jsonl output in the crawler's own layout.
``subdir`` 是**爬虫**落盘的目录名,不是监控层的平台 id —— 抖音那边这两者不同
(平台 id 是 ``dy``、目录是 ``douyin``),所以必须能分开指定,否则测不出那个差异。
"""
jsonl_dir = root / subdir / "jsonl"
jsonl_dir.mkdir(parents=True, exist_ok=True)
contents = jsonl_dir / "creator_contents_2026-01-01.jsonl"
@@ -170,6 +177,50 @@ def _comment(comment_id: str, note_id: str, create_time: int, **extra) -> Dict[s
return record
def _dy_note(aweme_id: str, liked: Any = "10", **extra) -> Dict[str, Any]:
"""抖音作品记录 —— 键名照抄 store/douyin/__init__.py 的落盘字段。
重点在于**没有** ``note_id``:抖音叫 ``aweme_id``。这一条差异没映射好,就是
每条记录都被 ingest 悄悄 continue 掉、一条不剩。
"""
record = {
"aweme_id": aweme_id,
"aweme_type": "0",
"title": f"title-{aweme_id}",
"desc": f"title-{aweme_id}",
"create_time": 1700000000000,
"creator_hash": "hash",
"nickname": "u***r",
"liked_count": liked,
"collected_count": "1",
"comment_count": "1",
"share_count": "1",
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
"cover_url": "https://img/cover.jpg",
}
record.update(extra)
return record
def _dy_comment(
comment_id: str, aweme_id: str, create_time: int, **extra
) -> Dict[str, Any]:
record = {
"comment_id": comment_id,
"create_time": create_time,
"aweme_id": aweme_id,
"content": f"content-{comment_id}",
"creator_hash": "hash",
"nickname": "u***r",
"sub_comment_count": "0",
"like_count": "0",
# 抖音顶层评论的父 id 是字符串 "0",不是空串。
"parent_comment_id": "0",
}
record.update(extra)
return record
async def _events(db: AsyncSession, event_type: Optional[str] = None) -> List[MonitorEvent]:
stmt = select(MonitorEvent)
if event_type:
@@ -525,3 +576,125 @@ class TestIdempotency:
assert result.new_notes == 0
assert result.new_comments == 0
assert len(list((await db.scalars(select(MonitorNote))).all())) == notes_after_first
class TestDouyinIngest:
"""抖音的产物形状与小红书不同 —— 这里钉住「不会被静默丢掉」。
这一组存在的理由,是这个改动最危险的失败模式:字段名或目录名没对上时,ingest
不报错,只是**一条都不入库**,然后被当成「疑似登录失效」报出去。
"""
async def _ingest(
self,
db,
tmp_path,
notes,
comments=None,
platform="dy",
subdir="douyin",
):
task = await _make_task(db, platform=platform)
run = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, notes, comments, subdir=subdir)
result = await ingest_run(db, run, task, tmp_path)
return task, run, result
@pytest.mark.asyncio
async def test_notes_are_ingested_under_their_douyin_field_names(self, db, tmp_path):
aweme_id = "7525082444551310602"
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note(aweme_id)])
note = await db.scalar(select(MonitorNote))
assert note is not None, "抖音作品被静默丢弃了 —— 多半是 aweme_id 没映射到 note_id"
assert note.note_id == aweme_id
assert note.note_url == f"https://www.douyin.com/video/{aweme_id}"
assert note.cover == "https://img/cover.jpg"
assert note.source_kind == "0"
assert note.published_at == 1700000000000
assert result.notes_fetched == 1
@pytest.mark.asyncio
async def test_the_artifact_directory_is_not_the_platform_id(self, db, tmp_path):
"""目录名与平台 id 不一致,是这套适配里最反直觉的一条。
抖音的平台 id 是 ``dy``,而爬虫把产物写在 ``douyin/`` 下。把它钉在这里,
是为了让「顺手改成一致」这件事会在测试里红掉,而不是让 ingest 悄悄读 0 条。
"""
assert adapters.artifact_dir("dy") == "douyin"
@pytest.mark.asyncio
async def test_writing_into_the_platform_id_directory_reads_nothing(self, db, tmp_path):
"""反面:产物落在 ``dy/`` 下时一条都读不到 —— 这正是映射要解决的问题。"""
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note("1")], subdir="dy")
assert result.notes_fetched == 0
@pytest.mark.asyncio
async def test_misplaced_output_is_blamed_on_the_directory_not_the_login(
self, db, tmp_path
):
"""产物其实抓到了,只是目录名不对 —— 不该报成「疑似登录失效」。
这是最难查的一类故障:登录是好的、数据也抓到了,但报出来的现象和登录失效
一模一样,会把人指去查完全错误的方向。
"""
_task, run, _result = await self._ingest(
db, tmp_path, [_dy_note("1")], subdir="dy"
)
assert any("目录" in event.title for event in await _events(db, EVENT_NO_DATA))
assert await _events(db, EVENT_AUTH_FAILURE) == []
assert run.error_message and "dy" in run.error_message
@pytest.mark.asyncio
async def test_comments_are_linked_through_aweme_id(self, db, tmp_path):
aweme_id = "7525082444551310602"
_task, _run, result = await self._ingest(
db,
tmp_path,
[_dy_note(aweme_id)],
comments=[_dy_comment("c1", aweme_id, 500)],
)
comment = await db.scalar(select(MonitorComment))
assert comment is not None, "抖音评论被静默丢弃了 —— 多半是 aweme_id 没映射"
assert comment.note_id == aweme_id
assert result.comments_fetched == 1
@pytest.mark.asyncio
async def test_a_top_level_parent_of_zero_becomes_empty(self, db, tmp_path):
"""抖音顶层评论的父 id 是 "0";原样存进去,前端会多出一堆悬空的父节点。"""
aweme_id = "7525082444551310602"
await self._ingest(
db,
tmp_path,
[_dy_note(aweme_id)],
comments=[
_dy_comment("c1", aweme_id, 500),
_dy_comment("c2", aweme_id, 600, parent_comment_id="c1"),
],
)
by_id = {c.comment_id: c for c in (await db.scalars(select(MonitorComment))).all()}
assert by_id["c1"].parent_comment_id == ""
assert by_id["c2"].parent_comment_id == "c1"
@pytest.mark.asyncio
async def test_the_four_metrics_need_no_mapping(self, db, tmp_path):
"""四个指标键两边同名 —— 抖音作品照样进 monitor_note_metric,差分照常。"""
aweme_id = "7525082444551310602"
task = await _make_task(db, platform="dy")
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="100")], subdir="douyin")
run1 = await _make_run(db, task, started_at=1)
await ingest_run(db, run1, task, tmp_path)
metric = await db.scalar(select(MonitorNoteMetric))
assert metric is not None and metric.liked_count == 100
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="150")], subdir="douyin")
run2 = await _make_run(db, task, started_at=2)
await ingest_run(db, run2, task, tmp_path)
assert len(await _events(db, EVENT_METRIC_DELTA)) == 1
+16
View File
@@ -118,6 +118,22 @@ class TestBuildRunMessage:
assert "标题A" in message
assert "https://www.xiaohongshu.com/explore/abc123" in message
@pytest.mark.asyncio
async def test_douyin_notes_link_to_douyin(self, db):
"""链接形状按平台走 —— 群里点进去该是能看的作品,不是 404。"""
task, run = await _seed(db)
task.platform = "dy"
_add_event(
db, task, run, EVENT_NEW_NOTE, "新作品:标题A",
payload={"note_id": "7525082444551310602", "title": "标题A"},
)
await db.flush()
message = await notify.build_run_message(db, task, run)
assert "https://www.douyin.com/video/7525082444551310602" in message
assert "xiaohongshu.com" not in message
@pytest.mark.asyncio
async def test_long_note_lists_are_truncated(self, db):
"""A first run can find dozens; a wall of text is worse than a count."""
+43 -3
View File
@@ -34,7 +34,7 @@ from api.monitor.models import (
RUN_SUCCESS,
)
from api.monitor.scheduler import MonitorScheduler
from api.monitor.settings import set_cookie
from api.monitor.settings import set_cookie, set_setting
from tools.time_util import get_current_timestamp
MS_PER_MINUTE = 60_000
@@ -72,12 +72,14 @@ async def executed(monkeypatch):
return calls
async def _make_task(next_run_at, enabled: bool = True, interval: int = 60) -> int:
async def _make_task(
next_run_at, enabled: bool = True, interval: int = 60, platform: str = "xhs"
) -> int:
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="t",
platform="xhs",
platform=platform,
mode=MODE_CREATOR,
enabled=enabled,
interval_minutes=interval,
@@ -145,6 +147,44 @@ class TestFiring:
assert executed == []
@pytest.mark.asyncio
async def test_the_cookie_gate_reads_the_tasks_own_platform(
self, monkeypatch, db, executed
):
"""cookie 闸门要按任务自己的平台取。
以前这里是 ``get_cookie(session)``(默认小红书)—— 只有小红书时看不出问题,
接上抖音后,抖音任务会因为读的是小红书那份 cookie 而永远不被触发,且不报错。
"""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
async with monitor_db.get_session() as session:
await set_cookie(session, "sessionid=dy-secret", "dy")
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
await MonitorScheduler().tick()
assert executed == [(task_id, "scheduled")]
@pytest.mark.asyncio
async def test_cdp_mode_frees_a_task_from_the_cookie_gate(
self, monkeypatch, db, executed
):
"""开着 CDP 时不该再要求先粘 cookie。
CDP 模式下登录态来自被接管的那台浏览器,粘不粘 cookie 都由不得它 —— 不放行的话,
选了「接管已有 Chrome」却没粘 cookie 的用户会发现任务永远不跑,而且什么错都不报。
"""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
async with monitor_db.get_session() as session:
await set_setting(session, "system.cdp_enabled", "true")
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
await MonitorScheduler().tick()
assert executed == [(task_id, "scheduled")]
@pytest.mark.asyncio
async def test_long_outage_coalesces_into_one_run(self, monkeypatch, db, executed):
"""A missed schedule fires once, not once per missed interval."""
+32 -7
View File
@@ -24,6 +24,7 @@ import pytest_asyncio
from sqlalchemy import text
from api.main import app
from api.monitor import adapters
from api.monitor import db as monitor_db
from api.monitor import platforms
from api.monitor.models import MonitorTask
@@ -54,7 +55,28 @@ class TestCapabilityMatrix:
# what stops the UI offering a platform that can never produce data.
assert all("monitor_wired" in p for p in body["platforms"])
assert by_value["xhs"]["monitor_wired"] is True
assert by_value["dy"]["monitor_wired"] is False
assert by_value["dy"]["monitor_wired"] is True
def test_every_wired_platform_has_an_adapter(self):
"""能力矩阵说「接通了」,就必须真的有一套适配管子。
两个注册表(platforms.PLATFORM_CAPABILITIES 与 adapters.ADAPTERS)分开是有意的
—— 前者是给前端看的能力描述,后者是爬虫的管道细节。代价是它们可能漂移,
所以在这里钉一条:凡声明接通的,必须能找到适配器。
"""
for platform in platforms.all_platforms():
if platforms.is_monitor_wired(platform):
assert adapters.has_adapter(platform), f"{platform} 声明接通但没有适配器"
@pytest.mark.asyncio
async def test_target_hints_are_exposed_for_wired_platforms(self, client):
"""前端的目标输入框拿它做 placeholder —— 让用户看到本平台该粘什么样的链接。"""
body = (await client.get("/api/config/platforms")).json()
by_value = {p["value"]: p for p in body["platforms"]}
assert "douyin.com/user/" in by_value["dy"]["target_hints"]["creator"]
assert "douyin.com/video/" in by_value["dy"]["target_hints"]["note"]
assert "xiaohongshu.com" in by_value["xhs"]["target_hints"]["creator"]
@pytest.mark.asyncio
async def test_metrics_are_per_platform_and_labelled(self, client):
@@ -80,14 +102,17 @@ class TestCapabilityMatrix:
class TestTaskCreationGuard:
@pytest.mark.asyncio
async def test_unwired_platform_is_rejected_with_an_explanation(self, client):
"""Accepting it would create a task that silently never produces data."""
"""Accepting it would create a task that silently never produces data.
用 B站 而不是抖音:抖音现在接通了,不再是「已知但未接通」的例子。
"""
response = await client.post(
"/api/monitor/tasks",
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": ["x"]},
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
)
assert response.status_code == 400
detail = response.json()["detail"]
assert "抖音" in detail
assert "B站" in detail
assert "尚未接通" in detail
@pytest.mark.asyncio
@@ -102,7 +127,7 @@ class TestTaskCreationGuard:
async def test_no_task_row_is_created_when_rejected(self, client):
await client.post(
"/api/monitor/tasks",
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": ["x"]},
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
)
assert (await client.get("/api/monitor/tasks")).json()["tasks"] == []
@@ -126,8 +151,8 @@ class TestTaskCreationGuard:
class TestPlatformScoping:
async def _seed_two_platforms(self, client):
"""One real XHS task plus a Douyin task inserted directly, since the API
refuses to create the latter."""
"""One XHS task created through the API, plus a Douyin task inserted
directly so its fields can be pinned exactly."""
await client.post(
"/api/monitor/tasks",
json={"name": "小红书任务", "mode": "creator", "targets": [XHS_TARGET]},
@@ -40,7 +40,7 @@ export function UnwiredPlatformNotice({ area }: { area: string }) {
{capability.label} 的{area}尚未接通
</h2>
<p className="text-[11px] font-mono text-cyber-text-muted">
{capability.label}的爬虫模块是支持的,但监控层目前只接通了小红书
{capability.label}的爬虫模块是支持的,但监控层目前只接通了小红书与抖音
</p>
</div>
</div>
@@ -94,9 +94,9 @@ export function UnwiredPlatformNotice({ area }: { area: string }) {
<div className="flex items-start gap-2 text-[10px] font-mono text-cyber-text-muted">
<MonitorSmartphone className="w-3.5 h-3.5 mt-0.5 flex-shrink-0" />
<p>
切换到小红书即可正常使用。若要接通该平台,需要在
<span className="text-cyber-neon-cyan"> runner / ingest / 目标解析 </span>
三处补上平台适配(目前这三处是硬编码小红书的)。
切换到已接通的平台即可正常使用。若要接通该平台,需要在
<span className="text-cyber-neon-cyan"> adapters.py </span>
里补一份适配:产物目录名、jsonl 字段名、目标链接形态。
</p>
</div>
</div>
@@ -20,6 +20,7 @@ import {
SelectValue,
} from '@/components/ui/select'
import { useCreateTask, useSettings, useUpdateTask } from '@/hooks/useMonitor'
import { useCurrentPlatform } from '@/hooks/usePlatform'
import { describeSchedule } from '@/lib/monitorFormat'
import type {
MonitorMode,
@@ -106,6 +107,10 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
const createTask = useCreateTask()
const updateTask = useUpdateTask()
const { data: settings } = useSettings()
// 示例链接和措辞都由服务端的能力矩阵给 —— 前端不自己判断平台,否则加一个平台
// 就要改这里一次,而且很容易漏。
const { capability } = useCurrentPlatform()
const hints = capability?.target_hints
const [name, setName] = useState('')
const [mode, setMode] = useState<MonitorMode>('creator')
@@ -361,23 +366,27 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
<div className="space-y-2">
<Label className="text-xs font-mono text-cyber-text-secondary">
{mode === 'creator' ? '博主主页链接或 ID' : '笔记链接或 ID'}
{mode === 'creator'
? `${hints?.creator_label ?? '博主主页'}链接或 ID`
: `${hints?.note_label ?? '笔记'}链接或 ID`}
</Label>
<textarea
value={targets}
onChange={(event) => setTargets(event.target.value)}
rows={5}
placeholder={
mode === 'creator'
? '每行一个,支持完整主页链接或纯 ID:\nhttps://www.xiaohongshu.com/user/profile/5f58bd99...\n5f58bd990000000001003753'
: '每行一个,支持完整笔记链接或纯 ID:\nhttps://www.xiaohongshu.com/explore/6aa3d827...'
}
placeholder={`每行一个,支持完整链接或纯 ID:\n${(mode === 'creator' ? hints?.creator : hints?.note) ?? ''}`}
className={TEXTAREA_CLASS}
/>
<p className="text-[10px] font-mono text-cyber-text-muted">
已识别 <span className="text-cyber-neon-cyan">{targetList.length}</span> 个目标。
<span className="text-cyber-neon-orange">建议只填纯 ID</span>
——链接里的 xsec_token 会过期,纯 ID 永久有效。
{/* 「只填纯 ID」是小红书专属的劝告:它链接里的 xsec_token 会过期。
抖音的链接不带令牌,永久有效,那句话对它没有意义。 */}
{hints?.token_expires && (
<>
<span className="text-cyber-neon-orange">建议只填纯 ID</span>
——链接里的 xsec_token 会过期,纯 ID 永久有效。
</>
)}
</p>
</div>
+12
View File
@@ -318,6 +318,17 @@ export interface ReportResult {
* layer has been hooked up for it. The UI must never conflate the two -- a
* platform can be fully crawlable upstream and still unusable here.
*/
/** 目标输入框的示例与措辞,随平台变。服务端给,前端不自己判断平台。 */
export interface TargetHints {
creator?: string
note?: string
/** 该平台怎么称呼这两样东西 —— 抖音叫「作品」,小红书叫「笔记」。 */
creator_label?: string
note_label?: string
/** 链接里是否带会过期的令牌(只有小红书有)。 */
token_expires?: boolean
}
export interface PlatformCapability {
value: string
label: string
@@ -327,6 +338,7 @@ export interface PlatformCapability {
comment_levels: number
media: boolean
monitor_wired: boolean
target_hints?: TargetHints
}
export interface WebhookStatus {