Files
MediaCrawler/api/monitor/platforms.py
T
butubb 06718a1351
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
feat(monitor): 抖音接入博主监控
上游爬虫本身不缺抖音能力(三模式、四项指标、二级评论都与小红书对等、指标还是同名同列),
缺的全在监控层的适配。这次把「平台之间不一样」的管子集中到一个新模块,再把散落的
xhs 硬编码接上去。

* 新增 api/monitor/adapters.py:产物目录名、jsonl 字段别名、目标链接形态与正则、
  通知链接模板。不放进 platforms.py 是因为那个模块被 describe_all() 整个序列化进
  /api/config/platforms 交给前端,塞进正则和目录名会让爬虫内部细节漏进 API 载荷。
  代价是两个注册表可能漂移,用一条测试钉住「声明接通就必须有适配器」。
* 两个必须知道的坑,都在这版里处理掉了:
  1) 抖音的平台 id 是 dy,而 store 把产物写在 douyin/ 下(store/douyin/_store_impl.py:47)。
     不改就是 ingest 一个文件都读不到 —— 不报错,只是 0 条,然后被冒充成「疑似登录失效」。
  2) 抖音的作品没有 note_id(叫 aweme_id)、评论也用 aweme_id 指作品。ingest 第一步是
     `if not note_id: continue`,不映射就逐条全丢。
  另外抖音顶层评论的 parent_comment_id 是字符串 "0",归一成空串,免得前端多出悬空的父节点。
* 顺带把「东西抓到了、只是没落在期望目录里」单独识别出来。这类故障的现象和登录失效
  一模一样,按登录失效报会把人指去查完全错误的方向。
* 修两个既有 bug(今天只有小红书所以无害,加抖音就踩响):
  - service.py update_task 换目标时漏传 task.platform,回落到默认小红书
  - scheduler.py 取 cookie 没传 platform,抖音任务会读着小红书那份 cookie 不动
* 行为变更(已与用户确认):cookie 闸门改成「没 cookie 且没开 CDP」才跳过。
  CDP 模式下登录态来自被接管的浏览器,粘不粘 cookie 由不得它决定;不放行的话,
  选了「接管已有 Chrome」却没粘 cookie 的用户会看到任务永远不触发,而且不报错。
  副作用是开启了 CDP 的小红书任务也不再被该闸门拦住 —— 语义上是对的。
* 目标输入框的示例链接与措辞改由能力矩阵提供(notes_label 抖音说「作品」、小红书说
  「笔记」;「建议只填纯 ID」是小红书专属劝告,抖音链接不带令牌,不再显示)。

测试 +22 条(858 通过),其中最关键的是「抖音作品/评论不被静默丢弃」与「产物目录名
不等于平台 id」两条 —— 都是把最难查的失败模式钉死在回归网里。

注意:抖音这条路的**端到端尚未验证**,需要一份可用的抖音登录态(CDP 那台 Chrome 里
登录,或导出一份 cookie)。单测覆盖的是解析与入库,真实抓取还没跑过。
2026-10-10 14:55:28 +08:00

217 lines
7.8 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/platforms.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Platform capability matrix.
The single source of truth for what each platform can do. The UI renders its
platform switcher and metric columns from this, and the API validates against
it.
Two distinct things are recorded here, and conflating them would be misleading:
* ``crawler_modes`` / ``metrics`` / ``comment_levels`` / ``media`` describe what
the upstream crawler module actually supports. These were read out of the
platform modules, not assumed -- all seven implement search/detail/creator;
the real differences are in which interaction metrics they capture.
* ``monitor_wired`` says whether the *monitoring layer* has been hooked up. It
covers Xiaohongshu and Douyin. The parts where those two differ -- which
directory the crawler writes into, what the jsonl fields are called, what a
target URL looks like -- live in ``adapters.py``; the rest of the layer is
platform-neutral.
A platform can therefore be fully crawlable by upstream and still not usable for
monitoring, which is exactly the state of the other five today.
"""
from typing import Any, Dict, List, Optional
PLATFORM_XHS = "xhs"
PLATFORM_LABELS = {
"xhs": "小红书",
"dy": "抖音",
"ks": "快手",
"bili": "B站",
"wb": "微博",
"tieba": "贴吧",
"zhihu": "知乎",
}
# Interaction metrics each platform's store actually persists. Xiaohongshu has no
# play count or danmaku; Bilibili has both and the widest set; Kuaishou carries
# no comment/share/collect at all; Tieba stores only reply counts.
PLATFORM_CAPABILITIES: Dict[str, Dict[str, Any]] = {
"xhs": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": True,
"target_hints": {
"creator": "https://www.xiaohongshu.com/user/profile/5f58bd990000000001003753",
"note": "https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=...",
"creator_label": "博主主页",
"note_label": "笔记",
# 只有小红书的链接带会过期的 xsec_token。抖音的链接不带令牌,永久有效,
# 那句「建议只填纯 ID」的劝告对它没有意义。
"token_expires": True,
},
},
"dy": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": True,
# 用户可见的示例链接(前端的目标输入框用它做 placeholder)。放这里是因为
# 它属于「这个平台长什么样」的能力描述;真正干活的管子(正则、目录名、
# 字段别名)在 adapters.py。
"target_hints": {
"creator": "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE",
"note": "https://www.douyin.com/video/7525082444551310602",
"creator_label": "博主主页",
"note_label": "作品",
},
},
"ks": {
"crawler_modes": ["search", "detail", "creator"],
# No comment/share/collect in the Kuaishou store; sub-comments are stored
# flat with no parent link and carry no like count.
"metrics": ["liked_count", "view_count"],
"comment_levels": 1,
"media": True,
"monitor_wired": False,
},
"bili": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": [
"liked_count",
"video_play_count",
"video_danmaku",
"comment_count",
"video_favorite_count",
"video_coin_count",
"video_share_count",
],
"comment_levels": 2,
"media": True,
"monitor_wired": False,
},
"wb": {
"crawler_modes": ["search", "detail", "creator"],
# Weibo has no collect count, and its comment count field is named
# differently in the model.
"metrics": ["liked_count", "comments_count", "shared_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": False,
},
"tieba": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["total_replay_num", "total_replay_page"],
"comment_levels": 2,
"media": False,
"monitor_wired": False,
},
"zhihu": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["voteup_count", "comment_count"],
"comment_levels": 2,
"media": False,
"monitor_wired": False,
},
}
METRIC_LABELS = {
"liked_count": "点赞",
"comment_count": "评论",
"collected_count": "收藏",
"share_count": "分享",
"view_count": "播放",
"video_play_count": "播放",
"video_danmaku": "弹幕",
"video_favorite_count": "收藏",
"video_coin_count": "投币",
"video_share_count": "分享",
"comments_count": "评论",
"shared_count": "转发",
"total_replay_num": "回复数",
"total_replay_page": "回复页数",
"voteup_count": "赞同",
}
# Monitoring modes, mapped to the CLI crawler types upstream understands.
MONITOR_MODE_CREATOR = "creator"
MONITOR_MODE_NOTE = "note"
CLI_TYPE_FOR_MODE = {
MONITOR_MODE_CREATOR: "creator",
MONITOR_MODE_NOTE: "detail",
}
class UnsupportedPlatformError(ValueError):
"""Raised for an unknown platform, or one the monitor layer cannot run."""
def all_platforms() -> List[str]:
return list(PLATFORM_CAPABILITIES)
def is_known(platform: str) -> bool:
return platform in PLATFORM_CAPABILITIES
def is_monitor_wired(platform: str) -> bool:
return bool(PLATFORM_CAPABILITIES.get(platform, {}).get("monitor_wired"))
def describe(platform: str) -> Optional[Dict[str, Any]]:
capability = PLATFORM_CAPABILITIES.get(platform)
if capability is None:
return None
return {
"value": platform,
"label": PLATFORM_LABELS.get(platform, platform),
**capability,
"metric_labels": {
metric: METRIC_LABELS.get(metric, metric) for metric in capability["metrics"]
},
}
def describe_all() -> List[Dict[str, Any]]:
return [describe(platform) for platform in all_platforms()]
def ensure_runnable(platform: str) -> None:
"""Validate a platform for a monitoring task.
An unwired platform is rejected outright rather than accepted and left to
silently produce nothing -- the same silent-failure shape that made a valid
creator look like an expired login earlier.
"""
if not is_known(platform):
raise UnsupportedPlatformError(
f"未知平台:{platform}(支持:{', '.join(all_platforms())})"
)
if not is_monitor_wired(platform):
label = PLATFORM_LABELS.get(platform, platform)
raise UnsupportedPlatformError(
f"{label}的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。"
)