Compare commits

..
59 Commits
Author SHA1 Message Date
butubb 61808444ad feat(monitor): 评论接口带上 a_bogus 签名;修好作品导出的空列
Deploy VitePress site to Pages / build (push) Waiting to run
Deploy VitePress site to Pages / Deploy (push) Blocked by required conditions
**评论能采了。** 之前 comment/list 一直回 200 + 空 body,被读成「这条没评论」——
而它其实只是被网关挡了。缺的就是 a_bogus 签名,仓库里本来就有
(libs/douyin.js + execjs)。签上之后实测 200 / 9960 字节真评论。

只给评论接口签:作品、详情、博主资料三个不带签名也照常返回,而给它们加签名是
没验证过的改动。签名按需 import —— 那个模块 import 时就把 JS 喂给 execjs,
不该拖进监控层热路径。

**作品导出那几列一直是空的。** 列名写的是裸键 liked_count,而作品行的指标嵌在
metrics / deltas 里,row.get() 永远取到 None —— 导出来的表有「点赞/评论/收藏/
分享」四列,每一格都没有数。原来的测试只断言了「作品ID」,所以没发现。

顺手补上:导出带上 博主备注/昵称、作品备注、发布时间,时间戳格式化成人能读的
形态(原来是一串 13 位毫秒,Excel 里没法看也没法排序)。
2026-10-10 21:09:52 +08:00
butubb e0581682e1 fix(monitor): 没有作品的博主在作品栏里也要看得见
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
分组是从作品推出来的(按作品的 creator_hash 归组),于是没有作品的博主根本
不进列表:目标加了、资料采到了、粉丝数就躺在库里,界面上什么都看不见。

而这类博主恰恰是最该看见的 —— 还在涨粉,只是最近没发东西。藏起来正好藏反了。
线上就有一个:5 个目标里 3 个没作品,那 3 个连同已采到的粉丝数一起消失了。

改成 **账号快照 ∪ 作品** 两个来源:

* 有快照没作品 → 一个 0 篇的组,备注和粉丝数照常显示,组里写「暂无作品」;
* 有作品没快照 → 一个没有账号指标的组(小红书那条路不产生快照)。

/notes 因此多返回一份 `creators`,而不是让前端从作品里推 —— 作品推不出上面
第一类人。组头改读它,`MonitorNote` 上那几个字段降级成「顺着作品问作者」用。
2026-10-10 20:57:07 +08:00
butubb 3c4daae0ac chore: 忽略部署用的临时脚本
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
.remote_run.py 是我在部署机上跑命令的工具,不是这个项目的一部分。
2026-10-10 18:27:47 +08:00
butubb d02fec5914 fix(monitor): 组头取账号指标时,别被组里第一条作品带偏
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
分组只按博主,而账号指标是挂在 (任务, 博主) 上的 —— 同一个博主被两个任务
监控时,组里可能只有一部分作品带指标。取第一条的话,「第一个任务还没采过」
就会让整组看起来没有指标。
2026-10-10 18:24:42 +08:00
butubb 20e672834c feat(monitor): 博主的粉丝数,以及给作品起备注
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
两件都是「作品栏里把这东西认出来」的延伸:

* **账号级指标**:作品列表只会说「这条涨了多少赞」,说不了「这个人整个
  账号的粉丝在涨还是在掉」。抖音的资料接口本来就有粉丝数/总获赞/作品数,
  每轮顺手记一条快照(`monitor_creator_stat`,粒度 = 任务×博主×轮次,
  和作品指标同形)。组头显示最近一条。

  快照在「一条作品都没采到」的早退**之前**落:作品列表被风控挡住的那一轮,
  正是「粉丝还在涨、但新作品没在发现」最该被看见的时刻。

* **作品备注**:博主备注回答「这个账号是谁」,这条回答「这条我要盯着」。
  一个博主底下常常只有一两件值得盯的作品,所以不能合并成一条。键取
  (platform, note_id),跨任务共用一份。

两边都守住同一条口径:**不知道就是 null,不写成 0** —— 0 在趋势图上是一条
砸到底的线,和「还没采到」是两回事。
2026-10-10 18:09:03 +08:00
butubb 0a88474c92 feat(monitor): 给博主起备注 —— 作品栏里才认得出「这是谁」
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
作品栏按 creator_hash 把作品归到博主的组里,可那是个哈希;creator_name 是平台昵称,
粉丝少的号常常认不出。两样都没法把账号对上人。

* 新表 monitor_creator_alias,键取 **(platform, creator_hash)** 而不是按任务:
  哈希对同一个 uid 是稳定的,所以同一个博主出现在多个任务里时,备注只填一次。
* list_notes 带上 creator_alias(整体查一次再在内存里取,不是每条作品查一次)。
* `PUT /monitor/creators/{creator_hash}`,空串即清掉那条备注。
* 作品栏的博主组头:**优先显示备注**,起过备注之后平台昵称降成副标题(它仍是有用的
  对照);组头上一个铅笔按钮就地编辑,回车保存、失焦保存、Esc 取消。

Esc 那条要单独处理:取消之后紧接着的失焦会把刚放弃的内容存进去,所以用一个标记让那次
失焦闭嘴。

测试 +6:没起过时是空串、起了会跟着作品返回、**跨任务共用一条**、**不串到别的平台**、
空串清掉、前后空格会被去掉。
2026-10-10 18:00:13 +08:00
butubb 5552e2a2b8 fix(monitor): 抖音任务永远「运行中」—— page.evaluate 卡在一个死掉的标签页上,而我没给超时
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户报「一直在运行中」。库里那条 run 是真的卡住了:日志里一条采集输出都没有,
说明它卡在 collect 里、还没走到任何日志。探针定位到:

    标签页: ['']                        ← 一个 URL 为空的标签页,渲染进程已卡死
    context.cookies(): OK, 90 个          ← cookie 读得到
    page.evaluate('navigator.userAgent'): **永远不返回**

而问浏览器要身份(UA + client hints)是采集的**第一步**,`page.evaluate` 又**没设超时** ——
于是整轮挂在那儿,run 永远停在「运行中」。

三处修复,各挡一层:

1. `page.evaluate` / `context.cookies()` 全部加超时(8 秒)。卡住就跳过,不再无限等。
2. 不假设第一个标签页是好的:逐个试、优先抖音页;全都不行就临时开一个干净页问完关掉。
   拿不到就退回库里那份 cookie —— **不编造指纹**,那比没有更糟。
3. **进程内那条路补上整体超时**:爬虫那条靠 `run_and_wait(timeout=...)` 兜底,这条路
   没有子进程、没人管,里面任何一次卡住都会变成永久的「运行中」。

测试 +6:卡死的页会被跳过(真 sleep,验的正是超时)、没 UA 的页跳过、全不行时开临时页
并关掉它、优先抖音页;以及整轮卡住时 run 不会停在 running(含超时原因)。
2026-10-10 17:51:54 +08:00
butubb 9e13a7f686 fix(monitor): 抖音任务的 run 永远停在「排队中」—— 我上一版把状态标记缩进错了
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户报的现象:任务一直显示「排队中」。查库确认有两批 run 卡在 pending(任务 7 的 44/45、
任务 15 的 62/63)。两个原因,一个是我上一版改坏的:

1) **`RUN_RUNNING` 被我缩进进了爬虫那条分支。** 抖音走的是另一条路,于是它**从不标记
   「运行中」** —— 建完 pending 那一行就直接进采集,中途一旦出事(异常、进程被重启),
   状态就永远停在 pending。这是我加平台分岔时把原本在两条路公共位置的一行挪进去了。

2) **`recover()` 只收 `running`,够不着 `pending`。** 那行是上一轮建的、后面的采集却
   根本没机会开始(进程重启),它永远不会自己往前走。于是重启也救不回来,界面上就是
   一个永远「排队中」的幽灵。现在 pending 一起收。

两处都补了测试:抖音路的 run 必须在**采集开始之前**就已经是 running(这条改回去就会
失败);recover 要把 pending 也标成 interrupted。
2026-10-10 17:47:41 +08:00
butubb 9f70cd0924 fix(monitor): 抖音「作品」模式的目标被当成博主去查,白废一条本来能用的路
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
两种模式的**目标是不同的东西**,原来的 fetcher 却一条路走到底:

* 「作品」模式(粘贴作品链接)—— 目标本身就是作品 id,直接取详情即可。**这个接口没被
  那道真校验挡,今天就能用。**
* 「博主」模式 —— 目标是主页 sec_uid,要先拉作品列表;那个接口被挡,退化成刷新已知作品。

原来两种都去调 author_videos(它要的是博主 sec_uid),于是「作品」模式的监控拿作品号当
sec_uid 去查,必然失败 —— 而且失败原因说得很难懂(接口回你「未登录/不是浏览器」)。
结果就是:**新建「作品」模式的抖音监控永远抓不到东西**,而那本来是现有条件下唯一能用的。

现在按 mode 分岔。测试 +2:作品模式必须走 detail 且**不得**去调列表接口(走错了会
直接抛断言);一件作品坏掉不连累其他作品。

顺带记一条排查结论:博主主页的 HTML 里**没有**作品列表(RENDER_DATA 解出来只有
{isLogin, statusCode, isSpider}),所以「走页面 HTML 免接口」那条路也是死的。
2026-10-10 17:30:28 +08:00
butubb 95b1be2c2e fix(monitor): 一轮产物里重复的作品会让指标快照撞唯一键,整个 run 崩掉
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
跑真任务时踩到的:

    sqlalchemy.exc.IntegrityError: (1062, "Duplicate entry
    '7-7690458980574358513-45' for key 'uq_note_metric'")

根因是我上一版写错了一处作用域:退化路径里那个「遍历已知作品」的循环写在了**目标循环内部**,
所以任务有多个目标时,同一批已知作品会被拉两遍 → 同一件作品在一轮里出现两条记录 →
ingest 给同一件作品写两份本轮快照 → 撞 (task_id, note_id, run_id) 唯一键。

两处都修,各挡一层:

* douyin_fetch:去重集合挪到 collect 的最外层,**跨目标**只算一次;退化时也先查一遍
  已知作品是否已刷过。
* ingest:`_ingest_notes` 对「一轮里重复出现的 note_id」免疫。一层在源头、一层在入口,
  因为产物里重复并不罕见(多个目标指向同一个人、上游重跑、退化路径),不该靠上游自觉。

测试 +2:多目标时已知作品只刷一次;同一轮里重复的作品只落一份快照(这条会崩在
唯一键上,所以它测的正是运行时的那个崩法)。
2026-10-10 17:24:08 +08:00
butubb 3b6a437e6c feat(monitor): 抖音监控改走新的 Web 接口客户端(接上上一版的移植)
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上一版只把客户端写出来、验证了它单独可用,**但没有接进任何地方** —— 所以你建的任务跑起来
仍然在调爬虫子进程,报的仍然是那句自己编的「account blocked」。这一步把它接上。

* runner 的 Phase 2 按平台分岔:dy 走进程内 HTTP 客户端(douyin_fetch),其余平台照旧走
  爬虫子进程。抖音那条不再起 Playwright、不再构造那串自相矛盾的浏览器指纹参数。
* 新增 douyin_fetch:把采到的东西写成 store/douyin 那套 jsonl 形状 —— **ingest 完全不知道
  数据是从哪来的**,重采样/差分/事件/通知/报表全都照旧,一个字没改。
* 失败不再假装:一条都没采到就以非零退出码 + **真实原因**交给 ingest,落成
  「采集进程异常退出(code=1):…」。绝不会再掉进「疑似登录失效」那个分支。
* 已知作品列表接口(aweme/post)被抖音单独加了真校验(200 + 空 body),所以加了退化:
  拿不到列表就用库里已知的 aweme_id 逐条走 detail 刷新。**边界是:已知作品的指标能继续
  更新,新作品发现不了** —— 这个边界会以一条 warning 日志留下痕迹,不让它看起来一切正常。
* 顺带给客户端补上 video_detail(实测可用:200 / 45425 字节),退化路径靠它。

测试 +6:产物目录与文件名、评论文件即使为空也要建(ingest 靠它区分「没评论」和
「什么都没抓到」)、重复作品只写一次、列表被挡时的退化、彻底失败仍写出产物与原因、
评论失败不连累作品。
2026-10-10 17:21:12 +08:00
butubb 4f83a075f6 chore: 忽略 webui/.npm(服务器重建前端时容器写入的缓存,属构建产物)
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
2026-10-10 17:15:32 +08:00
butubb 7442bc10b8 feat(monitor): 抖音 Web 接口客户端 —— 绕开爬虫子进程,直接发 HTTP
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
移植自 mac-agent-os 的 mediacrawler_adapter:不起子进程、不开页面,用浏览器里那份
登录态直接调抖音 Web 接口。产物键名照抄 store/douyin,所以 ingest 那条链路一个字不用改。

**目前能用的(真环境实测,非推断)**:

    profile/other : 200, 7075 字节  —— 博主主页指标(粉丝/获赞/作品数/昵称)
    aweme/detail  : 200, 45425 字节 —— 单条作品详情(含点赞/评论/收藏/分享)

**目前不能用的:作品列表 `aweme/post`。** 两个互相独立的原因:
  1. 这个接口被抖音单独升级成了真校验:不带 x-tt-argus 回 403「Uifid Not Found」,
     带上 dummy 值回 200 + **空 body**。也就是说「头在不在」骗得过,「真校验」过不了。
     同一套头打 profile/other 和 aweme/detail 都是通的 —— 抖音是挑着接口加保护的,
     挑中的恰好是「批量拉作品列表」这个最敏感的动作。
  2. 改走页面截获也不行:CDP 浏览器打开博主主页会落到「验证码中间页」(当天大量探测的
     代价,过几小时要重测)。

  所以现在的边界是:**已知作品的指标刷新能做,自动发现新作品做不了**。

**排查中控住变量后得到的两条事实**(都写进注释了):
  · `Accept` / `Accept-Language` / `Referer` 才是主页接口能返回真数据的原因 —— 只有
    UA+client hints+Cookie 时是 200 但仅 121 字节的空壳,补上这三个头变 7074 字节。
    (我先前猜的 sec-ch-ua 不是关键。)
  · 因此 UA 与 client hints 必须**成套地取自同一个浏览器**,所以 BrowserIdentity 一次
    从 CDP 取齐 cookie + UA + hints,而不是各自写死。

「200 + 空 body 必须当场报错」也是刻意写死的:放过去它会在下游变成「这个博主没作品」,
把一次失败伪装成一条正常结果 —— 爬虫那条路正是这么栽的,还被翻译成「账号被封」。

顺带:把参考项目目录加进 .gitignore。上一次 `git add -A` 把 mac-agent-os-main 整个
(1429 个文件)带进了提交,已从历史里清掉。

测试 +11:cookie 解析、请求头成套性(含 uifid 缺失/回退)、产物键名与 store 对齐、
以及 _get 的三条失败路径(空 body / 403 带网关原话 / 正常返回)。
2026-10-10 17:13:04 +08:00
butubb 21ff01b894 fix(douyin): 缺 sec-ch-ua 请求头,网关回 200 + 空 body(不是「账号被封」)
先纠正一个我上一轮给错的结论:日志里的 `account blocked` **不是抖音说的**,是
MediaCrawler 自己编的:

    if response.text == "" or response.text == "blocked":
        raise Exception("account blocked")

真实情况只是**抖音返回了空 body**。我把它读成了「账号被风控」,还写进了运行历史和
给用户的结论里 —— 用户质疑「我网页版和手机版都能正常登录」,一查,他是对的。

实测定位(同一 URL、同一 cookie、同一参数):

  浏览器页面内 fetch : 200, 7077 字节  ✓
  httpx              : 200,    0 字节  ✗
    带 a_bogus       : 0 字节
    不带 a_bogus     : 0 字节
    四种 msToken 变体 : 全部 200 有数据(所以不是它)
  用浏览器那次的完整头重放 httpx : 200, 7077 字节 ✓

浏览器那次请求比爬虫多的,只有这三个头:

    sec-ch-ua: "Google Chrome";v="155", "Chromium";v="155", "Not(A:Brand";v=24
    sec-ch-ua-mobile: ?0
    sec-ch-ua-platform: "Linux"

爬虫的 UA 是从页面读的(声称是 Chrome 155)却不带 sec-ch-ua —— 「Chrome 的 UA +
没有 sec-ch-ua」是最典型的机器人特征。网关的回应方式是不报错、不给原因,回一个
200 + 空 body,HTTP 状态还写在成功那一栏。

修:media_platform/douyin/help.py 新增 client_hint_headers(),从
navigator.userAgentData 现算这三个头(现算而不是写死 —— 写死的版本号一旦和 UA 里的
对不上,就又是一个可疑特征);core.py 建客户端时带上。

诚实说明:我无法解释**为什么之前能跑**(run 34 还是成功的,40 分钟后同样的代码就
不行了)。最可能是字节那边收紧了这道校验,但我没有证据,别当结论。

测试 +8:还原出的头与真实浏览器抓到的值逐字一致;拿不到 userAgentData 时返回空而不
凭空编造(编一组和 UA 对不上的头比不带头更糟);mobile 标志;platform 缺失时仍发另两个。
2026-10-10 16:47:49 +08:00
butubb e4affe9170 feat(monitor): 运行历史要写清楚失败原因,不能只写「退出码 1」
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户的要求:运行历史的说明要写清楚。现在确实写不清楚 —— 抖音那次失败,运行历史里
只有一句 `Crawler exited with code 1`,而真正的报错 `DataFetchError: account blocked`
埋在子进程的 stderr 里,谁也看不到。

那两者本来是断开的两条路:子进程的输出只流向日志 WebSocket(前端 Terminal 看得到),
而监控层调 run_and_wait() 只拿得到一个退出码。

* crawler_manager 在 _push_log() 里留一份输出尾巴(80 行,每次 start 清空)——
  那是所有输出的唯一出口,挂这儿不会漏。新增 get_output_tail()。
* ingest 新增 diagnose_failure():倒着找第一行像异常的行(traceback 的末行),
  认不出就退回最后一行;管理器自己补的「Crawler exited with code」不是原因,排除掉。
* describe_exit_code() 接受这个原因并附在消息里;失败事件的标题也带上,这样企业微信
  通知和事件流不用翻日志就能看懂。
* runner 把尾巴交给 ingest;「超时/没起来」那条分支同样带上原因 —— -1 同时代表两种
  情况,而要查的东西完全不同。
* 运行历史那一格是截断的(240px),而失败原因现在有一整行 —— 补上 title 悬停显示,
  并放宽到 320px。没有悬停提示等于把最要紧的半句藏起来。

测试 +5:能挑出异常行、不会把管理器自己的话当成原因、没有输出时不报错、认不出时退回
最后一行、以及失败运行同时记下退出码与真因(含事件标题)。
2026-10-10 16:14:54 +08:00
butubb 1118d466be fix(monitor): 抖音的时间戳是秒、小红书是毫秒,不换算会把 2026 年显示成 1970 年
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上一个提交把作品发布日期透到界面之后,抖音那一行显示成 1970-01-22。查原始产物:

  小红书 time        = 1790923011000  → 毫秒 → 2026-10-02   ✓
  抖音   create_time = 1790574515     → 秒   → 2026-09-28   ✗(被当毫秒 → 1970-01-22)

抖音给的是**秒**。而且这条路径不只影响新增的发布日期 —— **评论的 create_time 走的是
同一条路**,所以抖音评论的时间一直是错的,只是之前界面上没显示出来,没人发现。

时间单位的换算正是 adapters 该管的事,所以加在那边:

* `PlatformAdapter.time_scale`(小红书 1、抖音 1000)+ `to_ms()`,解析不出来返回 None
  而不是伪造 0。
* ingest 用它换算作品的 published_at 和评论的 create_time。落库统一毫秒,展示层不必
  关心来源。
* 已入库的数据要能自愈:published_at 和评论 create_time 原先都是**只写一次**的,换算
  改对了老数据也修不回来。现在它们会在重采时跟着刷新(昵称早就是这么做的)。

测试 +1:抖音记录落库后 published_at 是 1790574515 * 1000,且年份是 2026 不是 1970。
dy 的 fixture 也改成用真实的秒值(原来写的是毫秒形态,所以测不出这个 bug)。

注意:库里那条抖音记录**仍带着错的值**,要等抖音下一次成功采集才会被修回来 —— 而它
现在正被平台风控挡着(account blocked),见下一条说明。
2026-10-10 16:08:42 +08:00
butubb 3486c7f524 feat(monitor): 作品栏和评论栏显示作品的发布日期
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
需求:抖音和小红书都要能看到作品的发布日期。

`MonitorNote.published_at` 其实**一直在库里**(ingest 早就按平台取:小红书 time、
抖音 create_time),只是从来没往 API 和界面上透 —— 后端序列化没这个键,前端类型里
也没有,所以界面上只有「首次发现」。

* 后端:list_notes 的序列化补上 published_at;_note_meta_map 也带上,于是
  list_comments 多一个 note_published_at,分组接口的桶多一个 published_at。
* 前端:NotesTable 新增「发布日期」列(要让分组表头的 colspan 从 +4 变 +5);
  评论栏作品那一层在标题旁显示日期 —— 同名作品不少,日期能帮着认。
* 新增 formatDate:发布日期问的是「哪一天发的」,绝对日期比「3天前」好认,也不会
  每天看都在变。具体到分钟的版本放在 title 里,悬停可见。

刻意和「首次发现」分开:前者是作者发布的那天,后者是我们第一次看到它的那天。把一个
早就存在的作品加进监控时,两者能差好几个月 —— 测试里就用不同的值把这两者钉住。

顺带修正一处过时注释:前端类型里还写着 creator_name 是「已脱敏的昵称」,
脱敏已经在上一个提交里关掉了(config.MASK_NICKNAME)。

测试 +2:桶要带发布日期;作品列表接口要带,且它不等于 first_seen_at。
2026-10-10 16:06:22 +08:00
butubb cd85587f00 feat(privacy): 关掉昵称脱敏 —— 脱敏有损,撞名就分不出博主
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
需求:评论栏/作品栏要能分清是哪个博主。修好分组字段之后名字仍带星号,因为上游作为
教学版默认对昵称做中间脱敏(首尾各留 1 字,中间星号)。这个脱敏是**有损**的:

  「张三」和「张四」都变成「张*」
  「小明老师」和「小刚老师」都变成「小***师」

而本仓库的用途是监控一批公开创作者账号,分清谁是谁正是这一层要干的事。所以关掉它。

* config/base_config.py 新增 MASK_NICKNAME = False(和 INJECT_ALL_COOKIES 一样是个
  开关,不是删代码 —— 改回 True 就恢复上游行为)。
* tools/user_hash.py 的 mask_nickname 读这个开关,关闭时原样返回。读的是模块属性而
  不是导入值,测试才能 monkeypatch。脱敏实现本身一字未动。
* 顺带修一个数据陈旧问题:评论是去重后直接 continue 的,昵称只在首次入库时写一次,
  于是开关一改(或评论者改名)老评论永远停在旧值 —— 而重采是唯一能拿到新值的途径。
  现在已存在的评论会跟着刷新昵称(作品那边的 creator_name 早就是这么做的)。
* anonymous 的 creator_hash 保持不变:那是分组用的稳定键,不是显示名。

测试:
* 三个隐私套件 + weibo 的 autouse fixture 强制把开关打开 —— 它们验的是**脱敏机制
  本身**,机制仍然必须正确,所以显式打开来测,而不是让它们随部署配置漂。
* test_mask_and_hash_tools 改成两个方向都覆盖(开着脱敏 / 关着脱敏)。
* test_tieba_extractor.py 里 8 处字面量的脱敏期望值换成真实昵称 —— 提取器现在就是
  返回原文的,期望值理应跟着改(这一条是行为变更的直接后果,不是测试放宽)。
* 新增一条:已入库的评论昵称会随重采刷新(且不会因刷新而重复插入)。
2026-10-10 15:48:15 +08:00
butubb 67837b407e fix(comments): 评论栏所有博主都显示成「未知博主」
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
现象:小红书的和抖音的评论栏,最外层分组全是「未知博主」,分不出谁是谁。

根因是分组接口漏了两个字段。api/routers/monitor.py 里按作品构造桶时只放了
note_id/note_title/note_cover/note_url/comments,而前端 groupByCreator 是用
bucket.creator_hash / bucket.creator_name 分组的 —— 两个都是 undefined,于是所有
博主塌成同一个 key,标签取空串回退成「未知博主」。

数据一直都在:每条评论上都带着 note_creator_hash / note_creator_name
(service.py:540-541),只是没往桶上搬。TS 的 CommentBucket 里也声明了这两个字段,
所以是后端没兑现自己的契约,不是前端写错。

修:构造桶时把作品的创作者一并放上去(同一个桶里的评论必然同属一个作品,取哪条都一样)。

测试:种子数据改成「两个作品属于不同博主」(原来是同一个 hash,测不出这个 bug),
新增一条断言每个桶带上自己那个博主、且两个博主的 hash 确实不同。
2026-10-10 15:41:42 +08:00
butubb 42f2209534 fix(douyin): 两个让抖音采集根本跑不起来的上游缺陷
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户建了个抖音博主任务,两次都是 exit 1、产物目录空空。查下来是上游两处缺陷,
与我们那层监控无关 —— 但它们在别的网络上未必复现,所以社区里没人报。

1) media_platform/douyin/core.py:101 —— 首页 goto 永远等不到 load

    await self.context_page.goto(self.index_url)   # 默认 wait_until="load"

   抖音首页的 load 事件不会触发(有长连接/埋点类请求一直挂着)。实测同一台 Chrome、
   同一个地址:domcontentloaded 0.7 秒返回,load 等满 90 秒仍超时。后果是整个采集
   一步没走就崩,退出码 1,看起来像「抖音不能用」。改成显式 domcontentloaded ——
   上游的贴吧和知乎本来就是这么写的,抖音这个页面只是恰好属于「永远不 load」那类。

2) media_platform/douyin/login.py:266 —— 注入 cookie 后页面是陈旧的

   login_by_cookies 把 cookie 塞进 context,但页面是在这之前加载的;SPA 只在加载时
   读一次登录态,localStorage.HasUserLogin 还停在"未登录",紧接着 check_login_state
   会对着这个陈旧值轮询到超时(600×1 秒=十分钟)再 sys.exit()。下一轮才正常,因为
   那时 cookie 已在 profile 里 —— 表现是"第一次白等十分钟、第二次才行",很容易被当
   偶发。修复:注入后 reload(domcontentloaded)。

   与 xhs 那个 __INITIAL_STATE__ 快照问题是同一类:页面状态是加载那一刻的快照。
   本仓库扫码登录与运营模块也各自踩过。

UPSTREAM.md 的第二类「上游 bug 修复」表补上这两条与成因说明(原来只有两条 xhs 的)。

验证:在服务器容器里手工复现,改完后真实采到数据 ——
  Parsed sec_user_id: MS4wLjABAAAArLubjxXEcLeiqehxgk4il4AHMu0hXu6_qlmD8z5WmMs
  get_all_user_aweme_posts ... video len : 1
  douyin aweme id:7690458980574358513, title:中秋哪儿都堵...
产物落在 douyin/jsonl/ 下,顺带把适配器里「平台 id 是 dy、目录是 douyin」的映射
用真实输出证实了。
2026-10-10 15:29:02 +08:00
butubb f3ea088c75 fix(monitor): 从界面上建的任务永远是小红书任务,平台从没被传上去
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
现象:站在抖音页面上新建任务,任务跑到小红书列表里去了;界面上还写着「笔记」。

根因在前端:TaskCreatePayload 里**根本没有 platform 字段**,handleSubmit 拼的 payload
自然也不带它,而后端是 `payload.get("platform") or PLATFORM_XHS` —— 于是不管在哪个
平台标签下建任务,落下来的都是小红书任务。以前只支持小红书,两边都看不出问题。

* TaskCreatePayload 补上 platform(并在注释里写明为什么它是必填),handleSubmit 带上
  当前平台。
* 更新任务时不带 platform:平台创建后不可更改,带着会让「平台能改」看起来像真的。
* 三处写死的「笔记」改成按平台取措辞(能力矩阵的 target_hints.note_label):
  · 任务编辑器的类型选择项「笔记(批量监控指定内容)」
  · 「笔记模式下此项不生效」那句提示
  · 任务卡片上的类型徽章 —— 它按**任务自己的**平台取词,不是当前平台,因为卡片未必
    只出现在同平台的列表里
  抖音管它们叫「作品」,小红书叫「笔记」,写死一个对另一个就是错的。

测试 +1:后端这一半也守住 —— 建任务时显式给了平台,就必须落到那个平台,且不得出现在
另一个平台的列表里。前端那半边是 UI,测不了,但后端守住能挡住「给了不用」这类退化。

注意:这次是纯前端漏传,后端那个「缺省回退小红书」的行为本身没变(有测试断言它是
有意为之的兼容行为)。要彻底消灭这类静默错误,可以把缺省值去掉、让 platform 必填 ——
那会破坏 API 兼容性,目前没有任何别的调用方,需要的话说一声。
2026-10-10 15:13:37 +08:00
butubb e77e5e2f15 fix(report): 空的任务集合被当成了「不限制平台」,导致报表串平台数据
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
现象:切到抖音,报表里显示的是小红书的数据。

根因是报告聚合里的一行真值判断:

    scope = list(task_ids) if task_ids else None

空列表是假值,而空列表在这里的含义是「这个平台一个任务都没有」,不是「不限制平台」。
于是 platform=dy 且抖音还没有任务时,_resolve_scope 返回的 [] 被翻译成了 None,
聚合范围从「抖音的任务」变成了**全部任务** —— 小红书的数字就这么显示在了抖音页面上。
顺带 task_ids 也回成 None,界面会显示成「全部任务」。

改成 `is not None`。空列表进去就让 in_([]) 恒假,结果为空,这才是对的。

排查时把所有同类写法过了一遍,只有这一处错,其余(service.py 的 10 处作用域judgement、
_resolve_scope、export)用的都是 `is not None`。

测试:新增两条,并且**验证过它们在修复前会红**(失败信息就是 assert 42 == 0 ——
查一个没有任何任务的平台,却返回了小红书那条作品的 42 个赞)。

同时修掉一条空跑的测试:test_a_platform_with_no_tasks_yields_empty_not_everything
原先种了任务却没有作品/指标数据,于是过滤生效与否结果都是 0,什么都测不出来 ——
这正是这个 bug 能活下来的原因。现在它会真的塞一条作品+快照进去,并在末尾断言
「小红书自己的报表看得到那条数据」,用来证明前面那两个 0 是过滤出来的而不是没数据。
2026-10-10 15:00:12 +08:00
butubb 06718a1351 feat(monitor): 抖音接入博主监控
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上游爬虫本身不缺抖音能力(三模式、四项指标、二级评论都与小红书对等、指标还是同名同列),
缺的全在监控层的适配。这次把「平台之间不一样」的管子集中到一个新模块,再把散落的
xhs 硬编码接上去。

* 新增 api/monitor/adapters.py:产物目录名、jsonl 字段别名、目标链接形态与正则、
  通知链接模板。不放进 platforms.py 是因为那个模块被 describe_all() 整个序列化进
  /api/config/platforms 交给前端,塞进正则和目录名会让爬虫内部细节漏进 API 载荷。
  代价是两个注册表可能漂移,用一条测试钉住「声明接通就必须有适配器」。
* 两个必须知道的坑,都在这版里处理掉了:
  1) 抖音的平台 id 是 dy,而 store 把产物写在 douyin/ 下(store/douyin/_store_impl.py:47)。
     不改就是 ingest 一个文件都读不到 —— 不报错,只是 0 条,然后被冒充成「疑似登录失效」。
  2) 抖音的作品没有 note_id(叫 aweme_id)、评论也用 aweme_id 指作品。ingest 第一步是
     `if not note_id: continue`,不映射就逐条全丢。
  另外抖音顶层评论的 parent_comment_id 是字符串 "0",归一成空串,免得前端多出悬空的父节点。
* 顺带把「东西抓到了、只是没落在期望目录里」单独识别出来。这类故障的现象和登录失效
  一模一样,按登录失效报会把人指去查完全错误的方向。
* 修两个既有 bug(今天只有小红书所以无害,加抖音就踩响):
  - service.py update_task 换目标时漏传 task.platform,回落到默认小红书
  - scheduler.py 取 cookie 没传 platform,抖音任务会读着小红书那份 cookie 不动
* 行为变更(已与用户确认):cookie 闸门改成「没 cookie 且没开 CDP」才跳过。
  CDP 模式下登录态来自被接管的浏览器,粘不粘 cookie 由不得它决定;不放行的话,
  选了「接管已有 Chrome」却没粘 cookie 的用户会看到任务永远不触发,而且不报错。
  副作用是开启了 CDP 的小红书任务也不再被该闸门拦住 —— 语义上是对的。
* 目标输入框的示例链接与措辞改由能力矩阵提供(notes_label 抖音说「作品」、小红书说
  「笔记」;「建议只填纯 ID」是小红书专属劝告,抖音链接不带令牌,不再显示)。

测试 +22 条(858 通过),其中最关键的是「抖音作品/评论不被静默丢弃」与「产物目录名
不等于平台 id」两条 —— 都是把最难查的失败模式钉死在回归网里。

注意:抖音这条路的**端到端尚未验证**,需要一份可用的抖音登录态(CDP 那台 Chrome 里
登录,或导出一份 cookie)。单测覆盖的是解析与入库,真实抓取还没跑过。
2026-10-10 14:55:28 +08:00
butubb e348de48d3 feat(upstream): 上游更新检查——定期比上游、落后了推企业微信
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
本仓库在上游 MediaCrawler 之上加了一整层(见 UPSTREAM.md),可合并流程默认
「有人知道上游动了」。而部署是 git pull --ff-only,只从自己的 Gitea 拉——上游的
提交不主动 fetch 就永远看不见。拖着不合并的代价是复利的:越久越难合。

于是把「上游动了没有」变成一条会自己跑、会推企业微信的通知:

* api/monitor/upstream.py:git fetch <url> <branch> 到 FETCH_HEAD,用
  rev-list --count HEAD..FETCH_HEAD 算落后数、FETCH_HEAD..HEAD 算领先数。
  用 git 而非托管商 API,因为只有 git 知道共同祖先在哪——本仓库含有上游没有的
  提交,直接比 tip 会得出错误结论。增量 fetch 只传几个新提交,不会遇到
  UPSTREAM.md 里说的「大包必断」。
* 只 fetch 到 FETCH_HEAD:不配 remote、不写 refs/remotes、不碰索引与工作区,
  所以不打断正在跑的采集,也不和 deploy.sh 的 git pull 抢锁。
* 挂在调度器 tick 上(不是采集,所以不看 is_busy、不受活跃时段限制——定时检查
  放在半夜反而最合适),按 checked_at + 间隔 到期才跑;失败也写 checked_at,
  于是 GitHub 不通时是每间隔重试一次,而不是每个 tick 撞一次墙。
* 同一个 tip 只推一次(记 tip 而不是「推过没」),上游真又动了会再推。
* 两个接口:GET /monitor/upstream 只读缓存;POST /monitor/upstream/check 手动
  查一次且刻意不推通知——点按钮的人正看着结果。
* 默认关闭,间隔默认一天。

Dockerfile 显式装 git(python:slim 不带,而这是唯一的依赖);deploy.sh 顺带补上
一个真 bug 的提示:Dockerfile/requirements.txt 变了只 up -d 用的还是旧镜像。
2026-10-10 09:17:06 +08:00
butubb 44cbe8e2aa fix(monitor): 已登录时点「同步登录态为 Cookie」要干等 30 秒
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
顺序反了:原先先开页面、读二维码,读完才发现「已经登录、没有二维码」。而读二维码内部
会 wait_for_selector 等满 30 秒才放弃 —— 用户点一下按钮要干等半分钟,还白开一个标签页。
实测日志里就是 `Page.wait_for_selector: Timeout 30000ms exceeded`。

改成先问登录状态(那是一次接口调用,很快),已登录就直接返回,根本不碰页面。
新增测试守住这个顺序:已登录时 context.new_page 不得被调用。
2026-10-09 13:49:11 +08:00
butubb de9ff58371 feat(monitor): 扫码同时存一份 Cookie,并把登录判定换成权威判据
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户问「现在是不是扫码就能自动获取 cookie」—— 不是,而且这正是那个面板显得没用的根源:
它只干了半件事,扫码**只写浏览器 profile,完全不提取 cookie**(qrlogin.py 里连一行
取 cookie 的代码都没有)。于是:
- CDP 开着时任务能用(复用 profile),但 Cookie 面板始终显示「未配置」
- CDP 一关,任务立刻断,因为库里那份 cookie 从来没被填过

现在扫码把两件事一起做了:写 profile(CDP 用)+ 存一份到库(Cookie 注入用)。
两种机制同时填上,开关怎么切都不断。cookie 只在内存里从 qrlogin 传到路由,不进响应体。

同时修掉一个同类 bug:监控侧的登录判定还在用页面里的 window.__INITIAL_STATE__,
而那是**页面加载那一刻的快照** —— 浏览器本来就登录着时它是对的,但扫码是加载之后
才登录的,快照不会翻转,表现为「扫了码却一直停在二维码上」。运营模块踩过同一个坑,
当时只修了那一处。现在两边统一为:拿 cookie 问后台接口「我是谁」。顺带不再需要页面导航,
检测变轻了。

前端:已登录时按钮原先被我藏起来了,面板于是变成一块只能看、不能操作的区域 ——
用户的原话是「没用」。现在两种状态都给按钮,含义不同:未登录=取二维码,
已登录=把当前登录态同步成 Cookie。

测试:tests/test_qrlogin.py 重写(stub 从页面探针换成后台接口),新增「成功会话必须
交出 cookie」「只能取一次」两例。
2026-10-09 13:46:48 +08:00
butubb f6ddc46d62 feat(notify): 通知拆成「新作品」与「异常」两个开关,异常默认开
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
问题:cookie 过期导致任务失败,但没有任何通知。查下来不是代码问题 ——
notify_enabled 在两个任务上都是 False,而它默认就是关的,事件(run_failed)
也确实生成了,只卡在最后一道闸门。

但那个默认值是错的。代码里的理由是「一条任务列表都推到一个群会很快变吵,所以默认静默」,
这个理由对新作品成立(可能每轮都有),对失败不成立:一次登录态失效意味着这个任务事实上
已经死了,而你不会知道,直到某天发现数据停在几周前。最该被告知的就是这种情况。

现在拆开:
- notify_enabled  —— 推送新作品,可能每轮都有,默认关
- notify_failures —— 推送异常(登录失效/运行失败/没抓到数据),默认开

事件按开关过滤(build_run_message):只勾了「新作品」的任务不该因为一次失败被推消息,
反之亦然,否则拆开开关就没有意义。已有任务由 _ensure_columns 补上 notify_failures=1,
所以会自动开始收到异常推送。

列名 notify_enabled 是历史遗留(它早先是唯一的通知开关),语义已收窄为「新作品」,
用注释写明,不做列重命名 —— 那需要单独的迁移,不值为一个内部工具做。
2026-10-09 13:37:02 +08:00
butubb 2e6fa955b0 feat(monitor): 博主组头补上指标合计与本轮增量
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
折叠之后组头只有「N 篇」,等于把信息藏了起来 —— 折叠应该意味着「收起来但仍可一眼读到」,
而不是「看不见」。现在组头与数据行列对齐,每个指标列给出该博主的合计,下面再带本轮增量
(那才是监控真正要看的)。

合计口径与项目一致:某个指标在所有作品上都是 null 时,合计是 null 而不是 0 ——
「0」是真实值、「null」是不知道,合计成 0 会让「还没采到」看起来像「互动为零」。
2026-10-08 13:58:26 +08:00
butubb 0eb6ba31c5 feat(monitor): 作品栏按博主分组折叠,评论栏改为 博主/作品/评论 三级
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
先回答现状:跳转作品的按钮本来就有(作品栏每行末尾的外链图标);
评论栏原本是**两级**(作品 -> 评论),第一级是作品不是博主,所以并不是三级。

- MonitorNote 新增 creator_name(爬虫产出的 jsonl 里本来就有昵称,只是没存)。
  存的是**已脱敏**的值(张***三),与项目一贯的匿名化姿态一致 —— 爬虫刻意不落原始
  user_id(tools/user_hash.py),所以 creator_hash 是唯一稳定的分组依据。
  实测该哈希是无盐 sha256,能用任务目标的 external_id 反算配对。
- 入库时刷新 creator_name:作者改昵称是常事,只在首次写一次会一直显示旧的
- 作品栏:按 creator_hash 分组,组头可折叠(默认展开 —— 折叠的默认值不该藏数据),
  行内跳转按钮加了 title 说明
- 评论栏:一级博主、二级作品、三级评论。_note_meta_map 补上博主维度,
  评论流和作品分组都带上它
- 认不出博主的作品归到「未知博主」,不丢

顺带修一个语法错误:JSX 注释放在三元表达式分支里是非法的(那是子节点语法不是表达式),
移进 div 内。
2026-10-08 09:43:19 +08:00
butubb b11bbf771a fix(covers): 封面地址是签名过期而非防盗链,改为本地缓存
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
实测推翻了之前的诊断。同一批图:

  当天签发的地址 /202610080841/...  -> 200,带不带 Referer 都一样
  隔天的地址     /202610070837/...  -> 403,带不带 Referer 都一样

路径里那段时间戳就是签发时刻。所以这是**过期**,Referer 根本不是那个维度 ——
上一轮加 referrerPolicy 是照着错误结论改的,白改。

修法:
- 采集入库时每轮刷新 cover 地址。原先只在首次入库写一次,旧作品的地址烂在库里,
  而且再怎么重跑也修不回来
- 新增 api/monitor/covers.py:把图下载落盘。图一旦落盘就与签名无关,永远可读
- 下载放在 runner 的 Phase 5(事务已提交之后),不放 ingest —— ingest 的文档写明
  No network,往里塞网络请求会毁掉它可离线测试这一点
- 新增 GET /api/monitor/covers/{note_id} 取图。这条路由带鉴权,封面不会被匿名读走
- service 返回本地地址优先,没有缓存时才退回远程
- 每轮只补一批(60 张):一次跑几百张既慢又会给图床压力,而旧地址本来就在陆续过期,
  分摊到几轮反而更稳

顺带修正 NoteCover 的注释 —— 它写着防盗链,而那个结论已被推翻,留个错的注释比没有更糟。
2026-10-08 08:45:30 +08:00
butubb fa227600fe fix(creator): 权限开通后的状态显示 + 同步范围可选可查
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- 权限显示:实测今天 display/status 已变成 1,但 tip 是「数据正在更新中,请耐心等待」,
  作品仍是 0 条。原逻辑只在未开通时显示提示条,于是会一边写着「数据已开通」一边列不出
  作品,自相矛盾。现在只要后台有话要说就显示,并按状态区分措辞。
- 同步范围:原本写死 90 天,界面上既看不见也改不了。现在可选 7/30/90/180/365/730 天,
  与后端校验上限一致。
- 新增 creator_account.last_sync_days,记录上次实际使用的范围 —— 否则界面只能说
  「同步过了」,说不清覆盖的是哪一段。选择器也会对齐到它,避免上次同步 1 年、这次
  点一下悄悄缩回 90 天。
2026-10-08 07:52:28 +08:00
butubb bef0a4fbde fix(creator): 扫码成功却一直停在二维码上 + 运营改为左右布局
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
【扫码不完成】根因是判据本身。原先读页面里的 window.__INITIAL_STATE__,而那是
页面加载那一刻的快照:监控那边的同一探针能用,是因为那台浏览器的页面加载时就已经
登录了;而扫码是加载之后才登录的 —— SPA 内部确实登进去了,但初始快照不会翻转,
于是检测永远等不到。

改成拿 cookie 直接问创作者后台 /api/galaxy/user/info 我是谁。实测这个判据很干净:
游客也会拿到 a1(所以签名算得出来),但接口直接回 401 无登录信息;只有真正登录了
才返回 user_id。所以「有 a1」什么都证明不了,后台认了才算。
顺带按 5 秒节流 —— 前端每 2 秒问一次,没必要每次都打后台接口。

【弹窗不关】成功后不自动关闭,停在二维码上会让人以为没成功。现在显示账号卡片与原话
提示,1.6 秒后自动关闭并提供一个「完成」按钮。

【已完成结果会残留】take_cookie 取走 cookie 就拆会话,而在飞的轮询会看到 _current 为空
回报 idle,把已显示的成功能擦掉。现在把结果记在模块里重复返回,关闭弹窗时清掉 ——
否则下次打开会立刻显示上次的成功。

【布局】按用户要求改成左右两栏(左账号列表、右数据面板),与监控统一,取消二级菜单。
未选过时默认选中第一个,右栏不会一开始就是空的。

测试:tests/test_creator_login.py 新增 7 例,含「游客会话永不完成」「接口不打满每次轮询」
「临时上下文用完必须关掉」。
2026-10-07 16:40:24 +08:00
butubb 2f5852e311 fix(ui): 运营与扫码登录态没跟着平台走,且两个登录面板分不清
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
两个 bug 是同一类问题:没做平台门禁。

1) 切到抖音等平台,「运营」仍列出小红书账号。运营读的是小红书创作者后台,
   其它平台根本没有对应的后台接口。现在按平台门禁换掉整个视图并说明原因 ——
   与 MonitorDashboard / ReportView 用 isWired 的做法一致。

2) 扫码登录面板把小红书的会话当成任意平台的登录态报。它的检测读的是小红书页面的
   __INITIAL_STATE__,后端 qrlogin.LOGIN_URL 里也只有 xhs 一项;而标题写的是
   「{当前平台}登录态」,于是切到抖音照样显示已登录—— 这是个具体的谎。
   现在非小红书直接换掉整个面板(只改标题不够,下面的块读的仍是小红书的状态)。

3) 两个面板的标题都是「登录态」,看不出区别。它们其实是不同机制:
   - Cookie 面板:存进库、每轮以 --cookies_file 注入子进程 → 改名为「Cookie(定时任务用)」
   - 扫码面板:写进浏览器 profile、CDP 模式复用          → 改名为「浏览器登录态(扫码)」
2026-10-07 16:35:12 +08:00
butubb 5484e5a3ae fix(creator): 详情接口用了 MySQL 5.7 不支持的 NULLS LAST
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
1064 语法错误。nullslast() 是 PostgreSQL 语法,MySQL 5.7 不认;而 MySQL 把 NULL 视为比任何值都小,
所以 DESC 本身就把它排在最后,不需要额外声明。

这个 bug 只在真机上暴露:SQLite 从 3.30 起支持 NULLS LAST,本地测试环境测不出来。
2026-10-07 16:32:53 +08:00
butubb c2b310c7bf feat(creator): 新增「运营」模块 —— 多账号扫码登录与创作者后台数据
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
侧边栏在「监控」右边加了「运营」:账号列表 → 点进二级详情看该账号的数据。

【为什么是独立模块而不是监控的子视图】两者形状不同:监控是公开数据(点赞/收藏/评论/分享)的每轮快照+差分;运营是创作者后台按日期给出的曝光/观看/完播率/涨粉。凭据不同、采集方式也不同 —— 那边要浏览器登录态,这边是纯请求。硬塞进同一个模型会同时污染两边。

【扫码登录的关键差异】监控的扫码把登录态写进浏览器默认 profile(爬虫要复用)。运营要的是 cookie 字符串(纯请求够用),所以每次登录开一个**临时上下文**,扫完取出 cookie 就丢弃 —— 登第二个账号不会把第一个顶掉,也不影响监控那个登录态,十个账号互不干扰。

【决策依据】tools/probe_creator_api.py 的 Phase 0 实测:签名可自造(XYW_:MD5 → base64 → AES-128-CBC,与 xhshow 内置实现常量逐字节一致);主站 cookie 即可认证创作者后台;接口与参数已与真实页面对齐。

后端:
- api/creator/models.py: creator_account / creator_note_stat。**复用 MonitorBase**,这样 create_all 与上一轮改成元数据驱动的 _ensure_columns 会自动覆盖新表
- api/creator/signing.py: XYW_ 签名,带三条实测结论(url= 前缀、appId=ugc、401 与 406 的区别)
- api/creator/client.py: 纯 httpx 客户端。字段名尚未亲眼验证过,所以写成多别名匹配;解析不出来存 None 而非 0
- api/creator/service.py: 账号 CRUD 与同步。cookie 绝不进入对外结构,只给 has_cookie
- api/creator/login.py: 临时上下文的扫码登录
- api/routers/creator.py: 8 条路由,全部带鉴权

前端:
- 侧边栏「运营」+ OperationView(账号列表 → 二级详情)+ AddAccountDialog
- 权限状态显眼呈现:pending 时照抄后台原话「已为您申请数据权限,次日可查看」,并说明此时同步返回 0 条是正常的,不是采集失败

测试:tests/test_creator_client.py 新增 48 例,含「cookie 不得出现在对外结构里」这条不变量,以及权限未生效时空壳响应的处理。
2026-10-07 16:30:45 +08:00
butubb 2613f7577f feat(creator): Phase 0 探针 —— 创作者后台数据可以纯请求拿到
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
结论:签名可自造、主站 cookie 即可认证、接口与参数已与真实页面对齐。不需要浏览器、不需要独立的创作者登录。

- tools/probe_creator_api.py: 纯 HTTP 探针。用 XYW_ 算法自签(MD5 → base64 → AES-128-CBC,
  密钥与 IV 与 xhshow/config/config.py 逐字节一致),对 note/analyze/list 发请求
- tools/probe_creator_page.py: 打开真实数据分析页,记录页面自己发的请求,作为地面真相

Phase 0 的三条实测结论:
1. 签名可伪造。三种写法里只有「url= + 路径 + 查询串」被接受(200);仅路径、或裸路径都 406。
   并且不带 cookie 时返回的是应用层 401「无登录信息」而非网关 406 —— 说明签名每次都已通过
2. 主站 .xiaohongshu.com 的 cookie 就能认证创作者后台,不需要单独的创作者会话
3. 当前账号 dfg 返回空数据不是技术问题:permission/query 的 tip_msg 是
   「已为您申请数据权限,次日可查看」,display/status 均为 0,即权限尚未生效

关键佐证:真实页面调 note/analyze/list 用的查询串与本探针生成的完全一致,且拿到同一份空响应。
2026-10-07 16:22:15 +08:00
butubb 2b9ebdad87 fix(ui): 「每轮最多采集作品数」标签是错的——它是每个博主的上限
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
爬虫里这个值是在 per-creator 的函数内比较的(client.py get_all_notes_by_creator 的 result 是局部变量),而外层 for 循环遍历全部目标。所以 100 个目标 × 20 篇 = 单轮最多 2000 篇,一篇都不会被丢弃。标签写成「每轮最多」会让人以为超出的会被截掉。

- creator 模式:标签改为「每个博主最多采集作品数」,并实时算出「N 个目标 × M 篇 → 单轮最多 X 篇」
- note 模式:禁用该输入并说明「此项不生效」——get_specified_notes 里没有任何 CRAWLER_MAX_NOTES_COUNT 引用,列出的每个链接都会被逐条抓
- 单轮估算超过 500 篇时给出警告:每篇还要抓最多 max_comments_count 条评论、并发为 1,容易触发限流,也可能跑不完就被默认 1 小时的任务超时中断
2026-10-07 15:46:11 +08:00
butubb 6eff6fcc83 fix: 扫码登录状态可独立查询 + 趋势图改回自适应纵轴
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
【登录反馈】实测这台浏览器 loggedIn=true,其实早就登录成功了;看不到反馈是判据和设计的问题:

1. 判据不可靠。原先靠 web_session 的值变化判断——对照组显示:一个全新的空 profile 首次访问小红书就会被发一个 web_session,所以「有这个 cookie」什么都证明不了。可信信号是页面自己的 __INITIAL_STATE__.user.loggedIn,但它是 Vue 响应式引用,必须 .value 解包(这就是前面探针读到 [object Object] 和 None 的原因)。
2. 状态绑死在临时会话上。扫码会话是内存状态,进程一重启就没(部署、崩溃都算),面板于是悄悄退回初始态——一次成功扫码看起来像什么都没发生。

改法不是让会话活得久,而是把「登没登录」变成随时可查、与会话无关:
- 新增 GET /api/monitor/login/state,直接问浏览器要答案,带 5 秒缓存;force=true 先重载页面再读,用于状态陈旧
- qrlogin 改为常驻 Playwright 客户端 + 复用同一个标签页,并在重启后认领浏览器里已存在的 xhs 标签页,避免堆孤儿页
- 把「读不到状态」与「未登录」分开——前者显示具体错误,不再悄悄显示成未登录
- 面板顶部常驻显示登录态与昵称,带「重新检测」按钮;扫码成功后自动翻转

【趋势图】上一轮改过头了。dataviz 规范里没有「折线图必须从 0 起」这条——基线相关的条文全是讲柱状图的(柱状图用长度编码数值,不从 0 起比例就是错的;折线图用位置编码,轴只需如实框住数据)。改回自适应,保留上一轮修好的左侧刻度栏让范围始终可见;步长收敛到 1/2/5×10ⁿ,全平序列撑开一档避免除零。
2026-10-07 15:42:58 +08:00
butubb e608b51210 test: 修正列渲染断言——SQLAlchemy 会对保留字加反引号
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
monitor_run.trigger 是 MySQL 保留字,渲染出来是 `trigger`。这比之前手写的列清单更正确:
清单里的裸 trigger 会直接语法错误。断言改为剥掉引号后比对列名。
2026-10-07 15:37:08 +08:00
butubb d937ff5fe6 fix(db): _ensure_columns 改为按模型元数据推导,并补上漏加的调度字段
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上一个提交加了 4 个调度字段,却没在 _ADDED_COLUMNS 里登记,后果是生产环境:
  - 任务列表接口报 Unknown column,UI 打不开任务列表
  - 调度器每 20 秒 tick 一次炸一次,定时任务完全不会触发
  最阴险的是启动完全正常——能连库、能起来,只是随后每条查询都失败。

- _ensure_columns 不再遍历手写清单,改为遍历 MonitorBase.metadata.sorted_tables,
  从根上消掉「加了字段忘了登记」这类漏
- 新增 _column_ddl:用 CreateColumn 渲染类型与可空性,并给 NOT NULL 列补 DEFAULT。
  模型的 default= 是 ORM 侧行为、不会进 DDL,而给已有数据的表加 NOT NULL 列必须有值,
  否则能否成功取决于服务端 sql_mode
- 主键列跳过:MySQL 不允许 AUTO_INCREMENT 与 DEFAULT 共存

新增 tests/test_monitor_column_migration.py 守住:每个 NOT NULL 列都必须能生成带
DEFAULT 的合法 ALTER。
2026-10-07 15:35:40 +08:00
butubb 8f4e5586e9 feat(schedule): 任务支持「每天定时 / 每周定时」,用选择器而不是手写 cron
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- schedule.py: 新增调度计算模块(纯函数,便于单测)。三种模式:interval(每 N 分钟)/ daily(选钟点)/ weekly(选星期 + 钟点)
- 钟点模式是「固定时刻」而非「固定延迟」——从日历重算,所以某轮跑晚了不会把之后每一轮都拖晚。interval 保持原语义:从上一轮开始计时
- 抖动只加给 interval。给「每天 9:00」也加抖动就成了 9:00–9:01 随机触发,操作者选的时间被悄悄改掉,只会像 bug
- 模型 / schema / service: 新增 schedule_mode / schedule_hours / schedule_days / schedule_minute。时钟字段存逗号分隔文本——几个小整数、永远整体读写,单开一张表只会换来 join。interval_minutes 保留且仍是默认值,已有任务不受影响
- service: 改动任何调度字段都按合并后的状态重算 next_run_at。重新启用也算改动,否则停用一个月再打开会带着一个月前的 next_run_at,一保存就立即触发
- 前端: 运行方式三选一 + 小时/星期胶囊多选 + 分钟下拉,并实时预览结果句子。任务卡片改显示后端拼好的 schedule_label,避免列表和编辑器对同一计划给出两种说法
- 校验: 钟点模式至少选一个时间,按周至少选一个星期

tests/test_schedule.py 新增 24 个用例,含「恰好等于当前时刻的档位归属下一天」这个会让调度器自循环的边界。
2026-10-07 15:33:17 +08:00
butubb 242e3a7837 fix(chart): 纵轴自 0 起、刻度归位、数据点恢复为正圆
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- 纵轴从 0 起,上界取整到好读的数(1/1.2/1.5/2/…)。此前按 [最小值,最大值] 自适应,491→506 这点变化被撑满整个图高、看着像暴涨,且上下两刻度就是 506 和 491 两个几乎一样的数,没有 0 做参照读不出量级
- 刻度线画在 0 / 中值 / 上界,标签贴在各自主线上、放进左侧刻度栏。原来是两个绝对定位的数字浮在图面上,会挤在一起读成一个数
- 数据点恢复为正圆:根因是 preserveAspectRatio="none" 把 600×160 的 viewBox 横向拉满容器,横纵缩放不一致,半径 4 的圆被压成椭圆。改为用 ResizeObserver 量出容器宽度、等比绘图
- 每个数据点都画出来(原先只画末点),点数多时自动缩小半径;悬停热区按点位间距铺满整段,不再固定 12px
- 当前值移到标题行,不再浮在图面上遮挡曲线
2026-10-07 15:26:42 +08:00
butubb c1068845a9 fix(deploy): 部署时必须强制重建容器
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
代码是 bind mount,容器配置与镜像都没变,所以裸的 `docker compose up -d` 会判定无需变更直接跳过,改了 .py 也不会生效。前端产物是磁盘上的静态文件、能即时生效,这一点很容易把问题盖住,直到有人改了后端代码才发现。

首次实测即命中:跑 deploy.sh 的输出是 'Container mediacrawler Running',没有重启。
2026-10-07 15:22:00 +08:00
butubb 89b7b1e825 fix: 封面图不再被图床拒绝,作品栏补上封面
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- NoteCover: 新增共用封面组件,核心是 referrerPolicy="no-referrer"。小红书图床对带外部 Referer 的请求一律 403,而浏览器对跨域 <img> 默认就会带上本站源作为 Referer —— 于是封面全显示成破图,而 URL 本身完全正常。实测同一张图:无 Referer 200 / 67968B,Referer 为本站 403 / 0B
- CommentsFeed: 改用 NoteCover,修掉评论栏封面全部加载失败
- NotesTable: 作品栏此前完全没有渲染封面,补上缩略图;封面缺失时用占位块,避免行高随封面陆续到达而跳动

放在一个组件里而不是就地加属性,是为了让下一个显示封面的页面不会漏掉。
2026-10-07 15:21:29 +08:00
butubb fe45442b01 chore: deploy.sh 标记为可执行(Windows 上创建的文件没有 exec 位)
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
2026-10-07 11:14:13 +08:00
butubb 74a592024c feat: 部署改为 git 驱动,容器以宿主用户运行
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
服务器实测可经 Cloudflare 443 访问 Gitea 且 git 协议正常(此前我只测了 13000 端口就断言不可达,是错的),因此不再需要 tar + SFTP。

- docker-compose: 增加 user: "1000:1000"。容器此前以 root 运行,写进挂载目录的每轮 jsonl 产物都是 root 属主,导致宿主用户连自己的部署目录都挪不动 —— 这在把部署迁到 /mnt/data 时实际发生了
- deploy.sh: 一条命令走完 拉代码 →(webui/ 有改动时)重建前端 → 重启容器。前端产物 api/webui 是 gitignore 的,git pull 带不过来,必须在服务器上重建一次
- Dockerfile: 补 npm 包。corepack 只管 yarn/pnpm 不管 npm,而前端要在服务器上重建;这样服务器只需要 Docker,不必另配 Node 环境
2026-10-07 11:13:43 +08:00
butubb cffb407d15 feat: 生产库克隆脚本 + 容器补 Node 运行时
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- tools/clone_database.py: 同服务器跨库克隆并建立独立应用账号的 provisioning 脚本。本机与服务器都没有 mysql 客户端,所以走 INSERT...SELECT 而非 dump/reload。刻意只读写命令行指定的两个 schema,且给应用单开账号而不是复用管理员
- Dockerfile: 补 nodejs。douyin/help.py 在模块导入阶段就执行 execjs.compile(libs/douyin.js),缺 JS 运行时会抛 RuntimeUnavailableError;而 main.py 要导入全部 7 个平台,于是整个应用连带环境自检一起挂掉

已在 192.168.20.220 上验证:7 个平台导入全部通过。
2026-10-07 11:08:11 +08:00
butubb bcc7361145 refactor: 镜像只装依赖,代码改为挂载
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- Dockerfile: 去掉 COPY . .,镜像退化为纯依赖层,改代码不再需要重建镜像
- Dockerfile: 补 libgl1 / libxcb1 / libglib2.0-0 等 X11 库。opencv-python 在导入时需要它们,而 tools/utils.py 会经 slider_util 导入 cv2 —— 缺了不是某个边角功能挂掉,是整个应用起不来(已在真机冒烟中验证)
- Dockerfile: 补 tzdata,否则 TZ=Asia/Shanghai 被静默忽略,所有时间戳落成 UTC
- Dockerfile: apt 与 pip 均改走国内源。deb.debian.org 实测约 13 kB/s,96 MB 构建依赖要跑半小时以上
- docker-compose: 挂载 ./ 到 /app,部署流程从「重建镜像」变成「重传 + 重启」
- main.py: 启动时若缺 api/webui/index.html 就打印警告,避免只返回一段 JSON 却看着一切正常
2026-10-07 10:57:44 +08:00
butubb 37ca1b1cd6 feat: CDP 接管开关 + 扫码登录面板 + Docker 部署
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- runner: enable_cdp_mode 从硬编码 False 改为系统设置 cdp_enabled。服务器部署下爬虫接管已开启远程调试的 Chrome(默认 9222),复用其 profile 登录态;本机桌面默认仍为关,行为不变
- qrlogin: 新增 CDP 扫码登录。Chrome 在服务器上跑于 Xvfb,show_qrcode 依赖的 PIL 桌面看图程序不存在,二维码无处可显示;改为经 CDP 从页面取出二维码交给 WebUI 渲染。刻意复用 browser.contexts[0](新建 context 是无痕 profile,扫了也白扫),且绝不调用 browser.close()(会连带关掉操作者自己的 Chrome)
- webui: 设置页新增扫码面板,替换原本跳到采集页看终端二维码的入口
- Dockerfile / .dockerignore / docker-compose.yml: 服务器部署。host 网络是必需而非图省事——容器里 127.0.0.1:9222 必须落到宿主机回环
- UPSTREAM.md: 补充 gitcode 镜像,用于 GitHub 大包传输必断时补历史
2026-10-07 10:41:11 +08:00
butubb 4e60524f37 feat: 监控面板 / 登录鉴权 / 多平台切换 / MySQL
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
在上游 MediaCrawler 之上新增一层:

- 监控层 api/monitor/ —— 多博主/多笔记的定时采集、指标快照差分、报表、
  企业微信通知。每轮采集写入独立目录,差分才成立。
- WebUI 登录鉴权 api/auth.py —— PBKDF2 口令 + 服务端会话,/api 全接口防护。
  WebSocket 单独加依赖:BaseHTTPMiddleware 对 ws 作用域直接放行,覆盖不到。
- 全局平台切换 + 能力矩阵 —— 如实区分「爬虫模块支持」与「监控层已接线」,
  未接通的平台直接拒绝建任务,而不是静默跑空。
- 监控库改用 MySQL 5.7(可回退 SQLite 供测试):逐表强制 utf8mb4
  (服务端与库默认都是 latin1),启动校验所连 schema 以防写错库,
  连接池 recycle + pre_ping 应对 MySQL 的 8 小时空闲断连。

修复上游缺陷:

- xhs/core.py: 主页抓取失败会跳掉整个博主,导致一条作品都抓不到,
  而那份资料只喂给一个空函数。改为尽力而为,失败不中断。
- xhs/login.py: cookie 登录只注入 web_session,冷启动签名会失败。
  新增 INJECT_ALL_COOKIES 开关(默认关闭,原有行为不变)。
- requirements.txt: 补上 websockets。它在上游 pyproject.toml 里有声明、
  这里漏了,导致 uvicorn 没有 WebSocket 能力,实时日志流从未工作。

改动过的上游文件清单及合并方式见 UPSTREAM.md。

测试:492 passed(另有 1 个既有的 Windows/gbk 上游测试失败,与本改动无关)
2026-10-07 09:58:40 +08:00
程序员阿江(Relakkes) 5d547f4586 update readme 2026-10-05 22:21:13 +08:00
程序员阿江(Relakkes) 83c638b9d1 docs: README 新增 OpenLux 赞助商,西语版补上 Atlas Cloud 2026-10-04 22:58:52 +08:00
程序员阿江(Relakkes) bf28178082 docs: 更新 README 中的 Pro 版介绍 2026-10-04 03:37:23 +08:00
程序员阿江(Relakkes) 380b426000 fix(dy): 补上 ArgusSecurityPlugin 要求的 x-tt-argus 请求头
抖音在边缘网关新挂了 ArgusSecurityPlugin,对 aweme/detail、aweme/post 这批接口做
业务前置校验:缺少 x-tt-argus 时直接 403,响应体为
"Blocked by ArgusSecurityPlugin Uifid Not Found";只补 uifid 参数但仍没有这个头
则是 "... Signature Not Found"——后者很容易被误判成 a_bogus / verifyFp 的问题。
网关当前不校验该头取值,传固定字符串即可(实测 "1" 就够)。
参考 https://github.com/Johnserf-Seed/f2/issues/443

注意这是权宜之计:网关哪天升级到真校验该值,会重新出现 Signature Not Found,
届时需要改为页面内注入 JS 让抖音自带 SDK 补齐 Argus 头。

验证:真实 cookie 下 get_video_by_id 恢复可用。

- client.py: 默认请求头补 x-tt-argus,cookie 有 UIFID 时带 uifid 头
  (两种都没有则不发,避免被当成「有但为空」)
- 新增 tests/test_douyin_argus_header.py(不发网络请求)
2026-09-19 13:25:59 +08:00
程序员阿江(Relakkes) 8ecfa31de2 chore(ks): 示例配置改用分享短链并补上支持的输入形式
- KS_SPECIFIED_ID_LIST 换成两条 /f/<token> 分享短链,用于验证短链解析链路
- 注释补上第三种输入形式(分享短链),与 help.py 的解析实现保持一致
2026-09-18 17:26:42 +08:00
程序员阿江(Relakkes) c7e6c9fdc0 feat(ks): 支持分享短链 /f/<token> 形式的视频输入
快手分享短链 https://www.kuaishou.com/f/X9Idt15MQb9L2cv 路径里是 share_token
而不是视频 ID,它只做 302 跳转,真实地址在 Location 里。原先的解析器只认
/short-video/<id> 和纯 ID,遇到短链会抛 ValueError 被 continue 掉——只有一行
ERROR 日志,看起来像"爬了但没数据"。

- help.py: 新增 /f/<token> 分支,返回 url_type="short"(token 不是视频 ID,
  必须跟随重定向),并补上第三种形式的文档
- client.py: 新增 resolve_short_url,GET 时 follow_redirects=False,读
  301/302/303/307/308 的 Location
- core.py: url_type == "short" 时先解析短链再解析一次,失败则跳过该条
- 新增 tests/test_kuaishou_url_parse.py(6 条,不发网络请求)

验证:两条真实短链分别解析到 3xyziwesje8e9jg / 3xbbkdxtxqm8sae,详情均成功;
带 query 的标准视频页不会被误判成短链。
2026-09-18 17:23:20 +08:00
程序员阿江(Relakkes) 1e1ae64cb4 fix(ks): 视频不可用时 photo 为 null 导致整轮爬取崩溃
快手对已删除/私密/不存在的视频,返回的 visionVideoDetail 里 photo/author 是
null——key 存在、值是 null。而 detail.get("photo", {}) 只在 key 缺失时给默认值,
key 存在且为 null 时拿到的仍是 None,紧接着的 photo.get(...) 抛 AttributeError。
该异常不在 get_video_info_task 的 except 列表里(只 catch DataFetchError /
KeyError),会穿过 asyncio.gather 直接把整轮爬取带崩。

- core.py: 用 (x or {}) 兜住 null;photo 为空时打 WARNING 并跳过该视频,
  不再返回半残的 detail 让下游存空记录、再去下载媒体
- 新增 tests/test_kuaishou_unavailable_video.py(不发网络请求)

验证:photo=null / photo 缺失 / visionVideoDetail 缺失 均返回 None;
真实 cookie 端到端确认不可用视频被跳过、正常视频照常返回详情。
2026-09-18 17:18:01 +08:00
程序员阿江(Relakkes) cf513e70a1 fix(dy): 补上 detail 接口的 uifid/verifyFp 风控参数
抖音 detail 接口的 Argus 风控要求 uifid / verifyFp / fp 三个参数,缺一则直接
403:响应体为 "Blocked by ArgusSecurityPlugin Uifid Not Found",补上 uifid 但
verifyFp 不对时换成 "... Signature Not Found"。原实现只传 aweme_id,于是
get_aweme_detail 全线失败,详情和媒体都拿不到。

- 三个参数取自浏览器 cookie:uifid 用 UIFID(缺失时退到 UIFID_TEMP),
  verifyFp/fp 用 s_v_web_id。必须用 cookie 里的 s_v_web_id——实测 uifid 搭配
  自生成的 verifyFp 会被判 Signature Not Found,两者需要同源
- 新增 tests/test_douyin_aweme_detail_params.py(不发网络请求)

验证:抓包比对真实浏览器发出的 detail 请求,参数逐项一致;真实爬取中详情与
74MB 视频均下载成功。
2026-09-18 17:00:11 +08:00
程序员阿江(Relakkes) 281b445f52 chore(xhs): 更新 XHS_SPECIFIED_NOTE_URL_LIST 示例笔记
换成当前可用的测试笔记,用于验证详情抓取与媒体下载链路。
2026-09-18 16:18:42 +08:00
155 changed files with 26814 additions and 469 deletions
+29
View File
@@ -0,0 +1,29 @@
# Keep the build context to what the server actually runs.
.git
.github
.venv
venv
# The WebUI sources are not needed -- only the bundle they produce, which lands
# in api/webui and is therefore NOT excluded.
webui/node_modules
webui/src
webui/dist
# Runtime state: per-run crawler output and the login browser profile. Mounted
# as a volume instead, so it survives image rebuilds.
data
browser_data
# Secrets come from the environment via compose, never baked into a layer.
.env
tests
docs
__pycache__
**/__pycache__
*.pyc
*.pyo
.pytest_cache
*.db
*.log
+12 -1
View File
@@ -183,4 +183,15 @@ agent_zone
debug_tools
database/*.db
.omx/
.omx/
# 别人放在这儿的参考项目(mac-agent-os)。它是独立仓库、21MB,不属于本项目 ——
# 一旦被 `git add -A` 扫进来就是永久留在历史里(踩过一次:1601 个文件里 1429 个是它)。
# 要读它就直接读磁盘上的目录,别提交。
mac-agent-os-main/
# 服务器上重建前端时,容器里的 npm 往挂载目录写的缓存。是构建产物,不该进仓库 ——
# 它长期以「未跟踪」状态躺在工作区,正是会被 `git add -A` 顺手带走的类型。
webui/.npm/
# 我用来在服务器上跑命令的一次性脚本(不属于这个项目)
.remote_run.py
+71
View File
@@ -0,0 +1,71 @@
# Server deployment image.
#
# No browser is bundled on purpose. On this deployment the crawler attaches over
# CDP to the Chrome already running on the host (see the 接管已有 Chrome setting),
# so shipping a second copy of Chromium would only add hundreds of megabytes and
# a login state that nothing uses. The Playwright Python package is still needed
# -- that is what speaks CDP -- hence PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD.
FROM python:3.11-slim
ENV PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
PIP_DISABLE_PIP_VERSION_CHECK=1 \
PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1 \
TZ=Asia/Shanghai
# asyncmy compiles a Cython extension, so a toolchain has to exist at build time.
# It is left installed: purging it risks taking libmysqlclient with it, and a
# slightly larger image is cheaper than a runtime that fails months later.
#
# deb.debian.org is effectively unusable from this network -- it was pulling the
# 96 MB of build dependencies at roughly 13 kB/s, which puts a build at well over
# half an hour. Point apt at a domestic mirror; override APT_MIRROR when building
# from somewhere that does not need it.
#
# The libgl1/libxcb1/... group is not for a GUI: opencv-python links against X11
# at import time, and tools/utils.py reaches cv2 through slider_util, so without
# them the *application* fails to import, not just some image utility. tzdata is
# here because TZ=Asia/Shanghai is silently ignored without it, which would put
# every stored timestamp in UTC. nodejs is for PyExecJS: douyin/help.py compiles
# libs/douyin.js at *import* time, and because main.py imports every platform,
# that single platform being importable-or-not decides whether the whole app
# (and the environment self-check) comes up. npm rides along so the WebUI can be
# rebuilt on the server (see deploy.sh) instead of only on a workstation --
# corepack is present but does not cover npm, only yarn and pnpm.
#
# The pip mirror is set for the same reason as the apt one: this host's route to
# the public index is slow.
#
# git is for the 上游更新检查 (api/monitor/upstream.py): it fetches the upstream
# repository into the mounted checkout to count how far behind this fork is.
# python:slim does not ship git, and nothing else here pulls it in.
ARG APT_MIRROR=mirrors.tuna.tsinghua.edu.cn
RUN set -eux; \
for f in /etc/apt/sources.list /etc/apt/sources.list.d/debian.sources; do \
if [ -f "$f" ]; then \
sed -i "s|deb.debian.org|${APT_MIRROR}|g; s|security.debian.org|${APT_MIRROR}|g" "$f"; \
fi; \
done; \
apt-get update; \
apt-get install -y --no-install-recommends \
build-essential pkg-config default-libmysqlclient-dev git \
libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 libxcb1 libgomp1 \
tzdata nodejs npm; \
rm -rf /var/lib/apt/lists/*
WORKDIR /app
# Requirements only -- this is the one layer that is expensive to build and
# changes rarely.
ARG PIP_INDEX=https://pypi.tuna.tsinghua.edu.cn/simple
COPY requirements.txt ./
RUN pip install --no-cache-dir -i "$PIP_INDEX" -r requirements.txt
# The application code is deliberately NOT copied in. compose mounts it at /app,
# so a code change is "re-upload the tarball, restart the container" instead of
# an image rebuild. Treat this image as the dependency layer and nothing else;
# rebuild it when, and only when, requirements.txt or this file changes.
EXPOSE 18051
# api.main reads MC_HOST / MC_PORT from the environment; compose supplies both.
CMD ["python", "-m", "api.main"]
+28 -6
View File
@@ -71,14 +71,16 @@
<strong>MediaCrawlerPro 重磅发布!开源不易,欢迎订阅支持</strong>
<details>
<summary>🚀 <b>开源版不够用?看看 MediaCrawlerPro 多了什么:断点续爬 · 多账号 + IP 代理池 · 去掉 Playwright · 多个 AI Agent 项目源码</b>(点击展开)</summary>
<br>
> 专注于学习成熟项目的架构设计,不仅仅是爬虫技术,Pro 版本的代码设计思路同样值得深入学习!
[MediaCrawlerPro](https://github.com/MediaCrawlerPro) 相较于开源版本的核心优势:
#### 🎯 核心功能升级
- ✅ **自媒体内容拆解Agent**(新增功能)
- ✅ **断点续爬功能**(重点特性)
- ✅ **多账号 + IP代理池支持**(重点特性)
- ✅ **去除 Playwright 依赖**,使用更简单
@@ -90,12 +92,16 @@
- ✅ **完美架构设计**,高扩展性,源码学习价值更大
#### 🎁 额外功能
- ✅ **AI Agent Skill 支持**(Codex / [OpenClaw](https://openclaw.ai/) 🦞 / [Hermes](https://github.com/NousResearch/hermes-agent) / Claude Code / [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) / [cc-haha](https://github.com/NanmiCoder/cc-haha) / WorkBuddy / 豆包 / Trae / Qoder / Cursor 一键安装,让 Agent 自动爬取数据)
- ✅ **评论分析 Agent**(🆕 新上线):输入关键词或链接,自动采集评论并生成研究报告
- ✅ **自媒体内容拆解 Agent**:解析内容、视频转文字、拆解爆款元素
- ✅ **多平台首页信息流推荐**(HomeFeed)和**热搜榜单**
- ✅ **自媒体视频下载器桌面端**(适合学习全栈开发)
- ✅ **多平台首页信息流推荐**(HomeFeed)
- ✅ **AI Agent Skill 支持**([OpenClaw](https://openclaw.ai/) 🦞 / Claude Code / Cursor 一键安装,让 Agent 自动爬取数据)
- [ ] **基于评论分析AI Agent正在开发中 🚀🚀**
- ✅ **AI 图片生成 Agent**:多轮迭代自动优化,内置精选模板库
点击查看:[MediaCrawlerPro 项目主页](https://github.com/MediaCrawlerPro) 更多介绍
开源不易,欢迎订阅支持!点击查看:[MediaCrawlerPro 项目主页](https://github.com/MediaCrawlerPro) 更多介绍
</details>
@@ -362,6 +368,22 @@ MediaCrawler 支持多种数据存储方式,包括 CSV、JSON、JSONL、Excel
<a href="https://go.nodemaven.com/MediaCrawlerSeptember">NodeMaven</a> 是面向网页抓取和自动化场景的高效代理服务商,提供市面上最高质量的 IP。主要优势包括 99.9% 可用性、ZIP 邮编定位、IP 过滤(所有代理的欺诈评分均低于 97%)、无需 KYC,以及代理带宽检测器、Meta 标签检测器、IP 查询等独家免费工具。MediaCrawler 用户使用优惠码 <code>CRAWLER35</code> 可享移动和住宅代理 35% 折扣,使用 <code>CRAWLER40</code> 可享 ISP(静态)代理 40% 折扣。👉 <a href="https://go.nodemaven.com/MediaCrawlerSeptember">访问 NodeMaven</a>
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://www.openlux.ai/register?channel=c_drxir46c"><img src="docs/static/images/openlux_logo.png" width="160" alt="OpenLux"></a>
</td>
<td valign="middle">
感谢 <a href="https://www.openlux.ai/register?channel=c_drxir46c">OpenLux</a> 对本项目的赞助!OpenLux 是一个面向企业的一站式 AI 聚合平台,汇集全球各大厂商主流大模型,平台提供高效、稳定的服务与及时的技术支持。Claude、OpenAI、Gemini 系列模型基准折扣分别低至官方的 0.882 折、0.4 折和 0.8 折。MediaCrawler 用户还可享受专属福利:通过<a href="https://www.openlux.ai/register?channel=c_drxir46c">专属链接注册</a>,充值最高可享 7.5% 优惠!👉 <a href="https://www.openlux.ai/register?channel=c_drxir46c">立即体验</a>
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://sx.org/c/CRAWLER3G"><img src="docs/static/images/sx_logo.png" width="180" alt="SX.ORG"></a>
</td>
<td valign="middle">
<a href="https://sx.org/c/CRAWLER3G">SX.ORG</a> 是专为高频数据采集与反爬对抗打造的高性能代理网络,完美适配 MediaCrawler 等多平台抓取工具。核心优势包括全球 190+ 地区真实住宅 IP 池、99.9% 稳定连通率、精准国家/城市及 ASN 运营商定位、全面支持 HTTP(S) 与 SOCKS5 协议,以及针对社交媒体风控优化的智能会话轮换。MediaCrawler 用户使用专属优惠码 <code>CRAWLER3G</code> 注册即可免费领取 <strong>3GB</strong> 优质测试流量。👉 <a href="https://sx.org/c/CRAWLER3G">访问 SX.ORG 领取 3GB 流量</a>
</td>
</tr>
</tbody>
</table>
+16
View File
@@ -310,6 +310,22 @@ MediaCrawler supports multiple data storage methods, including CSV, JSON, JSONL,
<a href="https://go.nodemaven.com/MediaCrawlerSeptember">NodeMaven</a> is an efficient proxy provider for web scraping and automation, offering the highest-quality IPs on the market. Key benefits include 99.9% uptime, ZIP targeting, IP filtering across all proxies (fraud score below 97%), no KYC, and unique free tools such as Proxy Bandwidth Checker, Meta Tag Checker, IP Lookup, and more. MediaCrawler users get 35% off mobile and residential proxies with code <code>CRAWLER35</code>, and 40% off ISP (static) proxies with code <code>CRAWLER40</code>. 👉 <a href="https://go.nodemaven.com/MediaCrawlerSeptember">Visit NodeMaven</a>
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://www.openlux.ai/register?channel=c_drxir46c"><img src="docs/static/images/openlux_logo.png" width="160" alt="OpenLux"></a>
</td>
<td valign="middle">
Thank you to <a href="https://www.openlux.ai/register?channel=c_drxir46c">OpenLux</a> for sponsoring this project! OpenLux is an all-in-one AI platform for businesses, bringing together leading AI models from major providers worldwide. With fast, reliable service and responsive technical support, OpenLux offers base pricing for Claude, OpenAI, and Gemini models as low as 8.82%, 4%, and 8% of official rates, respectively. Exclusive offer for MediaCrawler users: Sign up through our <a href="https://www.openlux.ai/register?channel=c_drxir46c">referral link</a> and enjoy up to 7.5% off credit top-ups! 👉 <a href="https://www.openlux.ai/register?channel=c_drxir46c">Get started with OpenLux</a>
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://sx.org/c/CRAWLER3G"><img src="docs/static/images/sx_logo.png" width="180" alt="SX.ORG"></a>
</td>
<td valign="middle">
<a href="https://sx.org/c/CRAWLER3G">SX.ORG</a> is a high-performance proxy network built for heavy web scraping and anti-bot bypass, fully compatible with MediaCrawler. Key advantages include global dynamic residential IP coverage across 190+ locations, 99.9% network uptime, precise country/city/ASN targeting, native HTTP(S) &amp; SOCKS5 support, and flexible session rotation for social media platforms. MediaCrawler users can use exclusive promo code <code>CRAWLER3G</code> at signup to get <strong>3 GB</strong> of free trial traffic. 👉 <a href="https://sx.org/c/CRAWLER3G">Claim 3 GB on SX.ORG</a>
</td>
</tr>
</tbody>
</table>
+24
View File
@@ -294,6 +294,14 @@ MediaCrawler soporta múltiples métodos de almacenamiento de datos, incluyendo
<a href="https://tikhub.io/?utm_source=github.com/NanmiCoder/MediaCrawler&utm_medium=marketing_social&utm_campaign=retargeting&utm_content=carousel_ad">TikHub.io</a> proporciona 900+ interfaces de datos altamente estables, cubriendo 14+ plataformas principales nacionales e internacionales incluyendo TK, DY, XHS, Y2B, Ins, X, etc. Soporta APIs de datos públicos multidimensionales para usuarios, contenido, productos, comentarios, etc., con 40M+ conjuntos de datos estructurados limpios. Use el código de invitación <code>cfzyejV9</code> para <a href="https://tikhub.io/?utm_source=github.com/NanmiCoder/MediaCrawler&utm_medium=marketing_social&utm_campaign=retargeting&utm_content=carousel_ad">registrarse y recargar</a>, y obtenga $2 adicionales de bonificación.
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://www.atlascloud.ai/?utm_source=github&utm_medium=link&utm_campaign=mei%27da%27c%27rmeidacrawler"><img width="160" alt="Atlas Cloud" src="docs/static/images/atlas_cloud_logo_black.png#gh-light-mode-only"><img width="160" alt="Atlas Cloud" src="docs/static/images/atlas_cloud_logo_white.png#gh-dark-mode-only"></a>
</td>
<td valign="middle">
<a href="https://www.atlascloud.ai/?utm_source=github&utm_medium=link&utm_campaign=mei%27da%27c%27rmeidacrawler">Atlas Cloud</a> es una plataforma de inferencia de IA multimodal que ofrece a los desarrolladores una única API de IA para acceder a APIs de generación de video, generación de imágenes y LLM. En lugar de gestionar integraciones con múltiples proveedores, se conecta una sola vez y obtiene acceso unificado a más de 300 modelos seleccionados de todas las modalidades. Descubra la nueva <a href="https://www.atlascloud.ai/console/coding-plan">promoción del coding plan</a> de Atlas Cloud para acceder a la API con un presupuesto más asequible.
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://go.nodemaven.com/MediaCrawlerSeptember"><img src="docs/static/images/nodemaven_banner_sep.png" width="180" alt="NodeMaven"></a>
@@ -302,6 +310,22 @@ MediaCrawler soporta múltiples métodos de almacenamiento de datos, incluyendo
<a href="https://go.nodemaven.com/MediaCrawlerSeptember">NodeMaven</a> es un proveedor eficiente de proxies para web scraping y automatización, con las IP de mayor calidad del mercado. Sus principales ventajas incluyen una disponibilidad del 99,9%, segmentación por código postal, filtrado de IP en todos los proxies (puntuación de fraude inferior al 97%), sin KYC y herramientas gratuitas exclusivas como Proxy Bandwidth Checker, Meta Tag Checker, IP Lookup y más. Los usuarios de MediaCrawler obtienen un 35% de descuento en proxies móviles y residenciales con el código <code>CRAWLER35</code>, y un 40% de descuento en proxies ISP (estáticos) con el código <code>CRAWLER40</code>. 👉 <a href="https://go.nodemaven.com/MediaCrawlerSeptember">Visita NodeMaven</a>
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://www.openlux.ai/register?channel=c_drxir46c"><img src="docs/static/images/openlux_logo.png" width="160" alt="OpenLux"></a>
</td>
<td valign="middle">
¡Gracias a <a href="https://www.openlux.ai/register?channel=c_drxir46c">OpenLux</a> por patrocinar este proyecto! OpenLux es una plataforma de IA todo en uno para empresas que reúne los principales modelos de IA de los grandes proveedores de todo el mundo. Con un servicio rápido y fiable y un soporte técnico ágil, OpenLux ofrece precios base para los modelos de Claude, OpenAI y Gemini desde tan solo el 8,82%, el 4% y el 8% de las tarifas oficiales, respectivamente. Oferta exclusiva para usuarios de MediaCrawler: regístrese a través de nuestro <a href="https://www.openlux.ai/register?channel=c_drxir46c">enlace de referido</a> y disfrute de hasta un 7,5% de descuento en las recargas de crédito. 👉 <a href="https://www.openlux.ai/register?channel=c_drxir46c">Empiece con OpenLux</a>
</td>
</tr>
<tr>
<td align="center" valign="middle">
<a href="https://sx.org/c/CRAWLER3G"><img src="docs/static/images/sx_logo.png" width="180" alt="SX.ORG"></a>
</td>
<td valign="middle">
<a href="https://sx.org/c/CRAWLER3G">SX.ORG</a> es una red de proxies de alto rendimiento diseñada para la extracción intensiva de datos web y la evasión de sistemas antibots, totalmente compatible con MediaCrawler. Sus principales ventajas incluyen cobertura de IP residenciales dinámicas en más de 190 ubicaciones, una disponibilidad de red del 99,9%, segmentación precisa por país, ciudad y ASN, soporte nativo de HTTP(S) y SOCKS5, y rotación flexible de sesiones para plataformas de redes sociales. Los usuarios de MediaCrawler pueden utilizar el código promocional exclusivo <code>CRAWLER3G</code> al registrarse para obtener <strong>3 GB</strong> de tráfico de prueba gratuito. 👉 <a href="https://sx.org/c/CRAWLER3G">Obtenga 3 GB en SX.ORG</a>
</td>
</tr>
</tbody>
</table>
+874
View File
@@ -0,0 +1,874 @@
# 内容抓取模块 · 开发技术说明
> 面向接手开发的团队 · 2026-10-10
> 全部内容基于**逐行读源码**整理,不是推测
> 范围:**只写内容抓取模块**,不涉及其他业务
---
## 一、这个模块是干什么的
从主流内容平台(抖音 / 小红书 / B站 / 知乎 / 任意网页)**采集内容数据**:
- 视频/笔记的元数据(标题、作者、发布时间、正文)
- 互动数据(点赞、评论、分享、播放)
- 评论列表
- 作者主页的全部作品列表
采到的数据落进本地 SQLite,供后续分析/运营使用。
---
## 二、整体架构(重要:三层降级是核心)
```
┌──────────────────────────────────────────────────────────┐
│ HTTP 层 routes/scrape.py(22 个端点) │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 引擎层 services/scrape_engine.py │
│ · URL 解析 → 标准化目标 │
│ · 按平台选适配器 │
│ · 同步 / 异步调度 │
│ · 结果落库(去重) │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 适配器层 services/adapters/*.py(每个平台一个) │
│ 基类 ScrapeAdapter 定义统一接口 + 工具降级 │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 工具层(★ 三层降级,这是本模块的核心设计) │
│ Level 1 OpenCLI 外部 Node CLI(主路径) │
│ Level 2 agent-browser Playwright + 真实 Chrome │
│ Level 3 web_crawler 通用网页兜底 │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 存储层 services/scrape_db.py(SQLite,4 张表) │
└──────────────────────────────────────────────────────────┘
```
### 为什么这么设计
抓取的最大风险是**单一方式失效**:目标平台改版、接口封禁、登录态过期。
所以**同一份数据有三条获取路径**,第一条失败自动降级到第二条,
**上层完全不感知**(对 engine 来说只是"拿到数据了")。
---
## 三、工具层详解(最关键的一层)
### 3.1 Level 1 — OpenCLI
**它是什么**:一个**第三方 Node.js CLI 工具**,包名 `@jackwener/opencli`。
```
实际安装位置(本机实测):
~/.workbuddy/binaries/node/versions/22.22.2/bin/opencli
→ 软链到 ../lib/node_modules/@jackwener/opencli/dist/src/main.js
```
**怎么调用**(`services/adapters/__init__.py` 的 `_run_opencli`):
```python
OPENCLI = os.environ.get("OPENCLI_PATH",
str(Path.home() / ".workbuddy" / "binaries" / "node" /
"versions" / "22.22.2" / "bin" / "opencli"))
async def _run_opencli(self, args: list, timeout: int = 60):
cmd = [self.OPENCLI] + args
proc = await asyncio.create_subprocess_exec(
*cmd, stdout=PIPE, stderr=PIPE)
stdout, stderr = await asyncio.wait_for(proc.communicate(), timeout=timeout)
...
return self._parse_output(stdout.decode().strip())
```
**调用示例**(抖音,`douyin_scrape.py`):
```bash
opencli douyin user-videos <sec_uid> --limit 20 --with_comments true -f json
opencli douyin stats <aweme_id> -f json
```
**⛔ 移植注意**:这个二进制**不在仓库里**,是外部依赖。移植时必须:
- 要么在目标机装 `npm i -g @jackwener/opencli`
- 要么改 `OPENCLI_PATH` 环境变量指向它的位置
### 3.2 Level 2 — agent-browser("套用真实浏览器"的做法)
**这是你问的重点。设计原则写在 `browser_helpers.py` 文件头**:
```
⛔ 绝不使用 Camoufox(养号专用,Firefox 内核 + 特殊指纹)
✅ 使用 Playwright 启动【真实 Chrome】(Chromium 内核,正常指纹)
```
**具体怎么"套真实浏览器"**(`browser_helpers.py` 的 `_get_browser`):
```python
_CHROME_PATHS = [
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome", # ← 系统真 Chrome
"/Applications/Chromium.app/Contents/MacOS/Chromium",
]
_CHROME_PATH = None
for p in _CHROME_PATHS:
if Path(p).exists():
_CHROME_PATH = p # 自动探测,找到就用系统已装的 Chrome
break
launch_kwargs = {
"headless": headless,
"args": [
"--disable-blink-features=AutomationControlled", # ★ 反检测关键
"--no-sandbox",
"--disable-dev-shm-usage",
"--disable-gpu",
"--window-size=1280,720",
],
}
if _CHROME_PATH:
launch_kwargs["executable_path"] = _CHROME_PATH # ★ 用系统 Chrome,不用 Playwright 自带
```
**四个关键设计点**:
| 点 | 做法 | 为什么 |
|---|---|---|
| **用什么内核** | 系统真实 Chrome(`executable_path` 指定) | Playwright 自带 Chromium 有明显特征;真实 Chrome 是正常用户指纹 |
| **怎么隐藏自动化** | `--disable-blink-features=AutomationControlled` | 这是最常被检测的自动化标志位 |
| **实例管理** | 模块级单例 `_browser` + `asyncio.Lock` | 避免每次请求都启动浏览器(启动 ~1-2 秒) |
| **会话隔离** | 每次 `browser.new_context()` | 每个任务独立 cookie 环境,互不污染 |
**页面加载策略**(`page_evaluate`):
```python
await page.goto(url, wait_until="domcontentloaded", timeout=timeout)
await page.wait_for_load_state("networkidle", timeout=timeout) # 等动态渲染
await asyncio.sleep(1) # 再等 1 秒保险
result = await page.evaluate(js_code) # 执行 JS 提取
```
**UA 伪装**(每次 context 都设置):
```python
context = await browser.new_context(
user_agent=("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/125.0.0.0 Safari/537.36"),
viewport={"width": 1280, "height": 720},
locale="zh-CN", # 中文环境,符合目标用户画像
)
```
**两个公开函数**:
```python
page_evaluate(url, js_code, timeout, headless) # 打开页面执行 JS,返回 dict
page_extract(url, selectors, timeout, headless) # 按 CSS 选择器提取文本
```
**Profile 持久化**(可选,环境变量控制):
```python
_USER_DATA_DIR = os.environ.get(
"SCRAPE_CHROME_USER_DATA",
str(Path.home() / "workbuddy-agent-os" / "agent-local" /
"runtime" / "scrape_chrome_profile"))
```
→ 想保持登录态就指定这个目录;不指定就用临时 context。
### 3.3 Level 3 — web_crawler
通用网页抓取兜底(`web_scrape.py`),用于非主流平台的页面。
### 3.4 降级怎么触发(`ScrapeAdapter._try_tools`)
```python
async def _try_tools(self, tool_level: int, funcs: list) -> tuple:
"""funcs: [(工具名, 可调用对象), ...],按顺序尝试"""
tools = [f for f in funcs[:tool_level]] # 按 level 截断
for name, fn in tools:
try:
result = await fn()
if result is not None:
return True, result, name # ★ 成功即返回,不再降级
except Exception as e:
logger.warning(f" ⚠️ [{self.platform}] 工具 {name} 失败: {e}")
return False, None, tools[-1][0] if tools else "none"
```
**关键语义**:
- `tool_level=1` → 只用 OpenCLI
- `tool_level=2` → OpenCLI → agent-browser(**默认**)
- `tool_level=3` → 三层全开
- **第一个成功就停**;返回 `(成功?, 结果, 用了哪个工具)`
---
## 四、适配器层(每平台一个)
### 4.1 统一接口(`services/adapters/__init__.py` 的基类)
**每个平台适配器必须实现 4 个方法**:
```python
class ScrapeAdapter:
platform = "" # 子类覆写,如 "douyin"
adapter_name = ""
async def collect_item(self, target, depth="light", tool_level=2) -> dict:
"""抓单条内容详情"""
async def collect_user(self, user_id, limit=20) -> list[dict]:
"""抓某个用户/作者的全部作品"""
async def collect_comments(self, item_id, limit=20) -> list[dict]:
"""抓评论"""
async def collect_search(self, keyword, limit=20) -> list[dict]:
"""按关键词搜索"""
```
**基类提供的公共能力**:
- `_try_tools(tool_level, funcs)` — 降级执行(见 3.4)
- `_run_opencli(args, timeout)` — 调 OpenCLI + 解析输出
- `_parse_output(text)` — **输出格式三层兜底解析**(见下)
- `_parse_lines(text)` — 纯文本兜底解析
### 4.2 输出格式三层兜底(细节,容易踩坑)
OpenCLI 的输出格式**不保证稳定**,所以解析做了三层:
```python
def _parse_output(self, text: str):
# 1. JSON 检测(以 [ 或 { 开头)→ json.loads
if text.startswith("[") or text.startswith("{"):
try:
return json.loads(text)
except json.JSONDecodeError:
logger.warning("JSON 解析失败,尝试 YAML 兜底")
# 2. YAML 解析(PyYAML 可用时)
try:
import yaml
parsed = yaml.safe_load(text)
if parsed is not None:
return parsed
except ImportError:
logger.debug("PyYAML 未安装,跳过 YAML")
# 3. 纯文本兜底:按行解析 key:value
return self._parse_lines(text)
```
`_parse_lines` 甚至**专门处理了 `top_comments` 块**(评论在纯文本里的多行结构)。
**⛔ 移植注意**:如果目标环境没有 PyYAML,会静默降级到第三层
(能跑,但嵌套结构会丢)。
### 4.3 各平台降级链(实测,**差异很大**)
```python
# 抖音(douyin_scrape.py)—— 两条路径
await self._try_tools(tool_level, [
("opencli", lambda: self._opencli_user_videos(...)),
("agent-browser", lambda: self._browser_user_profile(...)), # ← 有浏览器降级
])
# 小红书 / B站 / 知乎 —— ⚠️ 只有一条路径
await self._try_tools(2, [
("opencli", lambda: self._opencli_user_notes(...)), # ← 没有降级!
])
# 通用网页(web_scrape.py)—— 不用 OpenCLI
await self._try_tools(tool_level, [
("web_crawler", lambda: self._web_crawl(target)),
("agent-browser", lambda: self._browser_extract(target)),
])
```
**对照表**:
| 平台 | 降级链 | 有浏览器降级? | 备注 |
|---|---|---|---|
| **抖音** | `opencli` → `agent-browser` | ✅ | 唯一做了完整降级的平台 |
| **小红书** | `opencli`(单条) | ❌ | OpenCLI 挂了就抓不了 |
| **B站** | `opencli`(单条) | ❌ | 同上 |
| **知乎** | `opencli`(单条) | ❌ | 同上 |
| **通用网页** | `web_crawler` → `agent-browser` | ✅ | 走另一套(不用 OpenCLI) |
**⚠️ 这是模块的真实局限**:三个平台**只有一条路径**,没有降级能力。
接手时如果要提升健壮性,**最值得做的就是给它们补上浏览器降级**
(照抖音的 `_browser_*` 方法写即可)。
**另一个细节**:`collect_user` 传的是**硬编码的 `2`**(不是 `tool_level` 参数),
而 `collect_item` 才用传入的 `tool_level`:
```python
async def collect_user(self, user_id, limit=20): # 没有 tool_level 参数
await self._try_tools(2, [...]) # ← 写死 2
async def collect_item(self, target, depth, tool_level=2):
await self._try_tools(tool_level, [...]) # ← 用参数
```
→ **用户在前端设 `tool_level=1` 时,"抓用户主页"这条路仍会走到浏览器**。
### 4.35 登录态机制(★ 这是"如何用真实浏览器"的另一半)
**问题**:抖音的数据接口需要登录态(cookie),怎么拿到?
**答案**:`mediacrawler_adapter.py` 的做法 —— **用户在真实 Chrome 登录,程序通过 CDP 读解密后的 cookie**。
#### 为什么不能直接读 cookie 文件
代码注释原文(第 29 行):
```
# Chrome 新版把 cookie 值加密存在 SQLite 里,必须通过 CDP 读解密后的值。
```
Chrome v80+ 把 cookie **加密**存在 SQLite(`Cookies` 文件),
直接读文件拿到的是**密文**。必须让 **Chrome 自己解密** → 通过 **CDP 协议**问它。
#### 三步实现
**① 通过 CDP 读 cookie**(`_get_cookies`,第 68 行)
```python
async def _get_cookies() -> dict:
"""从 CDP 连接读取 Chrome cookie(解密后的值)"""
ctx = await _ensure_cdp() # 确保 CDP 连接
all_cookies = await ctx.cookies() # ← 让 Chrome 解密并返回
...
```
**② 转成 HTTP Header**(`_cookie_str`,第 86 行)
```python
def _cookie_str(cookies: dict) -> str:
return "; ".join(f"{k}={v}" for k, v in cookies.items())
```
**③ 带 cookie 直调抖音 API**(`_http_get`,第 93 行)
```python
def _http_get(url: str, cookies: dict, timeout: int = 15) -> dict:
headers = {
"User-Agent": "...Chrome/150.0.0.0 Safari/537.36",
"Cookie": _cookie_str(cookies), # ★
"Referer": "https://www.douyin.com/",
"Origin": "https://www.douyin.com",
}
...
```
#### 登录态判断
```python
cookies = await _get_cookies()
has_session = bool(cookies.get("sessionid")) # ← 有 sessionid 就算已登录
```
对应 HTTP 端点 `GET /api/scrape/check-login`。
#### 让用户登录(巧妙的做法)
代码注释(第 433 行):
```python
# ── 打开登录页(用 AppleScript 控制 Chrome,不需要 CDP) ──
"""在 Chrome 中打开抖音首页,让用户登录
登录后 cookie 自动保存到 Chrome profile,两个 Chrome 都会检测。"""
```
**→ 用 AppleScript 打开真实 Chrome**(不是 Playwright 控制的),
用户**在真实浏览器里手动登录** → cookie 存进 Chrome profile →
之后程序通过 CDP 读。
**这样最自然**:用户看到的是他熟悉的 Chrome,扫码登录,
✅ 不会被平台识别为"自动化登录"。
**⛔ 移植注意**:`open_login_page()` 用的是 **AppleScript**(`osascript`),
**macOS 专有**。Linux/Windows 要改实现(可以用 `open` 命令或直接 Playwright 打开)。
#### 一个隐藏技巧(避免暴露自动化)
代码注释(第 227-229 行):
```python
# 获取热评(复用已有 Chrome 页面,不创建新标签页)
# Chrome 已有 douyin.com 页面,直接用它的 JS 上下文执行 fetch
# ⚠️ 不要 new_page() — 那会在 Chrome 中闪出新标签页
```
**→ 复用用户已经打开的页面**执行 fetch,而不是新开标签页
(新标签页会闪一下,且更易被识别)。
---
### 4.4 抖音的特有实现(两份代码,别搞混)
**⚠️ 项目里有**两个**抖音相关模块**,职责不同:
**① `services/adapters/douyin_scrape.py`(181 行)**
- 走**工具降级**(OpenCLI → 浏览器)
- 浏览器路径用 JS 从页面 **DOM 文本**提取数据:
```javascript
// _browser_video_page 的 JS(正则从页面文本抓「获赞/粉丝/关注」)
const body = document.body.innerText || '';
const uidM = body.match(/抖音号[::]\s*(\S+)/);
const nickM = body.match(/@(\S+)/);
function extractNum(label) {
var m = body.match(new RegExp('(\\d+(?:\\.\\d+)?[万w]?)\\s*' + label));
...
}
return { aweme_id, title, author_nickname, douyin_id, digg_count, fans, following };
```
**② `services/mediacrawler_adapter.py`(457 行)**
- **不走工具降级**,完全独立的实现
- 文件头注释(原文):
> 全新架构:**不再依赖 CDP/Playwright/浏览器页面**。
> 直接从 Chrome profile 读取 cookie,通过 HTTP 请求调用抖音 API。
> 全程无窗口、无标签页、无闪烁。
- 用于**追踪视频/作者**这类需要高频刷新的场景(详见子代理报告)
**⛔ 关键澄清**:`mediacrawler_adapter.py` **虽然叫 mediacrawler,但不依赖
MediaCrawler 这个开源项目**——它只是借用名字,实际是"读 Chrome cookie + HTTP 调 API"。
---
## 五、引擎层(`services/scrape_engine.py`,361 行)
### 5.1 URL 解析(`resolve_target`,第 56 行)
**7 类目标自动识别**(实测代码):
```python
抖音短链 v.douyin.com/xxx → type=shortlink
抖音视频 douyin.com/video/{id} → type=video
抖音用户 douyin.com/user/{sec_uid} → type=user
小红书 xiaohongshu.com/explore/{id} → type=note
B站视频 bilibili.com/video/{BV} → type=video
B站用户 bilibili.com/space/{mid} → type=user
知乎 zhihu.com/answer/{id} | /question/ → type=item
通用网页 http(s)://... → type=page
纯 sec_uid MS4w 开头 或 len>20 → douyin/user
纯数字 len>=15 → douyin/video(aweme_id)
纯数字 其他 → zhihu/item
```
**短链解析**(`_resolve_shortlink`,第 136 行):用 `curl -sI` 拿 `Location` 头,
再用正则从跳转 URL 里抠出 `aweme_id`。
### 5.2 执行主流程(`run`,第 163 行)
```
run(request)
├─ 1. resolve_urls(targets) → 标准化目标列表
├─ 2. 短链逐个解析
├─ 3. 判断同步 / 异步
│ async_mode = request.async_mode 或 len(targets) > 50
│
├─ 【异步分支】
│ · run_id = uuid[:8]
│ · 内存状态 {status, total, completed, results, errors}
│ · asyncio.create_task(_run_async(...))
│ · 立即返回 {status:"async", run_id} ← 前端轮询
│
└─ 【同步分支】
· db.create_task("single", ...)
· for target: _scrape_one() → _save_item()
· db.update_task_status("completed", summary)
· 返回 {status, task_id, duration, total, success, errors, data}
```
### 5.3 单目标抓取(`_scrape_one`,第 261 行)
```python
adapter = self._get_adapter(platform) # 按平台取适配器(带缓存)
if target["type"] == "user":
return await adapter.collect_user(target["target_id"]) # 返回 list
elif target["type"] in ("video", "note"):
return await adapter.collect_item(target["target_id"], depth, tool_level)
else:
return None
```
**返回值语义**(重要):
- `dict` → 单条内容
- `list` → 多条(用户主页的所有作品)
- `None` → 失败
### 5.4 落库(`_save_item`,第 292 行)
```python
db_id = self.db.insert_item(
task_id=..., platform=..., item_id=..., url=..., title=...,
author_name=..., author_id=..., published_at=..., text_content=...,
tags=..., stats=..., extra=..., media=...)
comments = item.get("comments", [])
if comments and db_id:
self.db.insert_comments(db_id, comments) # ★ 评论独立表
```
### 5.5 异步模式(`_run_async`,第 315 行)
- ✅ **同时写内存 + 落库**(内存态供轮询,落库供持久化)
- 内存态在 `self._tasks[run_id]`(**进程重启即丢**)
- 查询用 `get_async_result(run_id)`
### 5.6 ⚠️ 已知的局限(接手时要清楚)
| 局限 | 说明 |
|---|---|
| **没有限流/并发控制** | `for target in ready:` 是**纯串行**,目标多时会慢;也没有请求间隔(可能触发平台风控) |
| **异步态存内存** | 重启 Dashboard 后 `_tasks` 丢失(但库里有记录,前端看不到进度) |
| **无重试** | 单目标失败只记 `errors`,不重试 |
| **adapter 实例缓存** | `self._adapters` 进程内缓存(无清理) |
---
### 5.7 一个完整请求的生命周期(跟着走一遍最快懂)
以"采集某抖音作者的全部视频"为例:
```
① 前端
POST /api/scrape/run
{"targets": ["r606391422378804368"], "tool_level": 2}
│
▼
② routes/scrape.py:44 api_scrape_run()
组装 request → engine.run(request)
│
▼
③ scrape_engine.py:163 run()
├─ resolve_urls(["r6063..."])
│ → resolve_target() 识别:以 MS4w 开头 → 抖音 sec_uid
│ → [{"platform":"douyin","type":"user","target_id":"r6063...","status":"resolved"}]
│
├─ 同步或异步?len(targets)=1,不大于 50 → 同步
│
├─ db.create_task("single","douyin",...) → task_id = 1
│
├─ for target: _scrape_one(target,"douyin","light",2)
│ │
│ ▼
│ scrape_engine.py:261
│ _get_adapter("douyin") → DouyinScrapeAdapter() (进程内缓存)
│ type=="user" → adapter.collect_user("r6063...")
│ │
│ ▼
│ douyin_scrape.py:21 collect_user()
│ _try_tools(2, [("opencli", ...), ("agent-browser", ...)])
│ │
│ ├─ 尝试 1:_opencli_user_videos()
│ │ _run_opencli(["douyin","user-videos","r6063...",
│ │ "--limit","20","--with_comments","true","-f","json"])
│ │ → subprocess 执行 opencli(Node CLI)
│ │ → _parse_output() ← JSON → YAML → 纯文本 三层兜底
│ │ → 成功返回 list[dict] → _try_tools 立刻返回,不再降级
│ │
│ └─ 尝试 1 失败(OpenCLI 没装/超时/报错)
│ → 尝试 2:_browser_user_profile()
│ → Playwright 启真实 Chrome → 打开页面 → JS 提取
│
│ → 每条数据 _to_schema() 转成统一格式
│ → 返回 list
│
├─ for item: _save_item(task_id=1, item)
│ db.insert_item(...) → db_id((platform,item_id) 唯一,重复则忽略)
│ db.insert_comments(db_id, comments) ← 评论另存
│
├─ db.update_task_status(1, "completed", summary={success:N, errors:0})
│
└─ return {status:"completed", task_id:1, duration, total, success, errors, data:[...]}
│
▼
④ 前端拿到 data,渲染列表
```
**异步分支的差异**(目标 > 50 个,或显式 `async_mode=true`):
```
run() 立即返回 {status:"async", run_id:"ab12cd34"}
↓(后台)
asyncio.create_task(_run_async(run_id, targets, ...))
↓
建持久化任务 → 逐个 _scrape_one + _save_item
↓
进度写内存 self._tasks[run_id](供轮询)
结果写 SQLite(供持久化)
↓
前端轮询 get_async_result(run_id) 看进度
```
**⚠️ 注意**:异步进度**只在内存**,Dashboard 重启就丢
(库里数据还在,但前端看不到进度了)。
---
## 六、存储层(`services/scrape_db.py`,415 行)
### 6.1 数据库位置
```python
DEFAULT_DB = AGENT_LOCAL / "data" / "scrape.db"
```
(`AGENT_LOCAL` 是环境变量;默认 `~/workbuddy-agent-os/agent-local`)
### 6.2 四张表(实测 `CREATE TABLE`)
```sql
-- ① 采集任务(一次 run 一条)
scrape_tasks(
id, type, -- single / batch / scheduled
platform, target, -- 目标(批量时是 JSON 数组)
depth, tool_level, machine,
status, -- pending / running / completed / failed
total_targets, completed_targets,
summary, -- 摘要 JSON
created_at
)
-- ② 采集到的内容
scrape_items(
id, task_id → scrape_tasks,
platform, item_id, -- 平台内唯一 ID
url, title, author_name, author_id,
published_at, collected_at, text_content, tags,
... -- 还有 stats / extra / media 等
)
-- ③ 评论
scrape_comments(id, item_db_id → scrape_items,
author_name, text, likes, replied_at)
-- ④ 采集源(长期跟踪)
scrape_sources(
id, platform, source_type, -- user / hashtag / keyword / url_list / api
target, display_name,
category, -- 自定义分类
notes,
schedule, -- CRON(定期采集)
depth, tool_level, last_collected,
status -- active / paused
)
```
### 6.3 方法清单(22 个,实测)
```
任务:create_task / update_task_status / get_task / list_tasks
内容:insert_item / get_item_id / item_exists / get_item / list_items
评论:insert_comments / get_comments
采集源:upsert_source / update_source / list_sources / get_due_sources /
update_source_collected / delete_source
统计:count_by_platform / count_today / sources_count / task_stats
```
### 6.4 去重机制
`insert_item` **依赖 `(platform, item_id)` 唯一约束** —— 重复插入时
用 `INSERT OR IGNORE` 模式(验证文档 L1-2 有测例)。
---
## 七、HTTP 层(`routes/scrape.py`,711 行 / 22 端点)
### 7.1 端点清单(实测)
```
采集
POST /api/scrape/run 发起采集(targets + depth + tool_level)
POST /api/scrape/resolve 只解析 URL,不采集
POST /api/scrape/douyin-stats 抖音数据查询
查询
GET /api/scrape/title 取标题
GET /api/scrape/result 结果
GET /api/scrape/tasks 任务列表
GET /api/scrape/items 内容列表
GET /api/scrape/items/{id} 单项详情
GET /api/scrape/stats 统计
采集源管理
POST /api/scrape/sources 新建源
GET /api/scrape/sources 源列表
DEL /api/scrape/sources/{id} 删源
追踪(视频 / 作者)
POST /api/scrape/track-video 追踪视频
GET /api/scrape/tracked-videos 已追踪视频
POST /api/scrape/delete-tracked/{id}
POST /api/scrape/refresh-video/{id} 刷新单个视频
POST /api/scrape/track-author 追踪作者
GET /api/scrape/tracked-authors 已追踪作者
POST /api/scrape/refresh-author/{id}
GET /api/scrape/author-history/{id}
主题 / 登录
POST /api/scrape/import-topics 批量导入主题
GET /api/scrape/check-login 检测登录态
GET /api/scrape/open-login 打开登录
```
### 7.2 关键实现(实测)
**`POST /api/scrape/run`**(第 43 行)—— 前端发起采集的唯一入口:
```python
@router.post("/run")
async def api_scrape_run(data: dict = {}):
targets = data.get("targets", data.get("target", []))
if isinstance(targets, str):
targets = [targets] # 兼容单个字符串
request = {
"targets": targets,
"platform": data.get("platform", "auto"),
"depth": data.get("depth", "light"),
"tool_level": data.get("tool_level", 2), # ← 默认 2(OpenCLI + 浏览器)
"machine": data.get("machine", ""),
"multi_machine": data.get("multi_machine", False),
"async_mode": data.get("async_mode", False),
}
engine = _get_engine() # 模块级单例
result = await engine.run(request)
return {"status": "ok", **result}
```
**登录态两端点**(第 692 / 703 行)—— 都委托给 `mediacrawler_adapter`:
```python
GET /api/scrape/check-login → mediacrawler_adapter.check_login_status()
POST /api/scrape/open-login → mediacrawler_adapter.open_login_page()
```
### 7.3 引擎实例
路由层用**模块级单例**拿 engine(`_get_engine()`),
所以 `engine._tasks`(异步态)和 `engine._adapters`(适配器缓存)
在整个 Dashboard 进程内共享。
---
## 七·五、Dashboard 插件(`plugins/crawl.py`,91 行)
抓取模块**作为 Dashboard 插件**注册(提供概览统计,不是核心逻辑):
```python
class CrawlDashboardPlugin(DashboardPlugin):
name = "crawl"
label = "内容抓取"
icon = "📡"
order = 35
```
**它做三件事**:
1. `summary()` — 概览:总抓取数 / 今日新增 / 抓取源(从 `ScrapeDB` 读)
2. `detail(machine)` — 指定机器的详情
3. `actions()` — 快捷操作(跳转 `scrape` 视图)
**注意**:它有个**兜底设计** —— 如果 `ScrapeDB` 不可用(数据库损坏/权限),
会退化成**统计知识库里的 md 文件数**,而不是报错。
**⛔ 移植注意**:如果目标项目没有这套插件框架,`plugins/crawl.py`
可以直接丢弃(它只是 Dashboard 的展示层,不影响抓取功能本身)。
---
## 八、前端(`frontend/src/views/scrape.js`,131 KB)
⚠️ **这是模块里最大的单文件**(131 KB)。功能覆盖:22 个端点的界面。
**建议接手团队**:不要照搬这个前端,按第七节的 HTTP 契约重写。
理由:131 KB 单文件难维护,且和本项目的视图框架耦合。
---
## 九、验证方案(项目里已有现成的)
`services/scrape_validation.md`(314 行)已经写了 **7 个 Level 的验证清单**:
```
L0 基础设施(3 项) Python import / SQLite 建库 / FastAPI 路由注册
L1 数据库层(3 项) 建任务 / 写入+去重 / 评论入库
L2 适配器 Mock(3 项) 工具降级逻辑 / 一级失败二级成功
L3 适配器真实(3 项) 抖音用户采集 / 小红书 / 详情+评论 ← 需 OpenCLI + 登录态
L4 引擎层(3 项) 解析 URL / 执行采集 / 异步轮询
L5 API 层(4 项) curl 打 4 个端点
L6 前端(4 项) 浏览器里操作
L7 异常(1+ 项) OpenCLI 不可用时应抛清晰错误
```
**⚠️ 但要注意**:该文档写于 2026-07-16,**里面的方法名已过时**:
```
文档写 resolve_targets ← 不存在
代码里是 resolve_urls ← 实际
文档写 get_result ← 不存在
代码里是 get_async_result ← 实际
```
**以代码为准**。
---
## 十、移植清单(换环境要改什么)
| # | 依赖 | 位置 | 处理 |
|---|---|---|---|
| 1 | **OpenCLI**(Node CLI) | `adapters/__init__.py` 的 `OPENCLI` 常量 | 目标机 `npm i -g @jackwener/opencli`,或设 `OPENCLI_PATH` |
| 2 | **真实 Chrome** | `browser_helpers.py` 的 `_CHROME_PATHS` | Linux/Windows 要改路径(如 `/usr/bin/google-chrome`) |
| 3 | **Playwright** | pip | `pip install playwright && playwright install chromium` |
| 4 | **PyYAML** | pip(可选但强烈建议) | 不装会降级到纯文本解析(丢嵌套结构) |
| 5 | **AGENT_LOCAL** 环境变量 | `scrape_db.py:18` | 定义了才能定位 `scrape.db` |
| 6 | **平台登录态** | Chrome profile | 目标机需手动登录一次目标平台 |
| 7 | `SCRAPE_CHROME_USER_DATA` | 环境变量(可选) | 要持久化登录态时指定 |
### 最小可运行子集(只要"能采集")
```
services/scrape_db.py 数据库(4 表)
services/adapters/__init__.py 基类 + 降级 + OpenCLI 调用 + 输出解析
services/adapters/browser_helpers.py 浏览器降级
services/adapters/<目标平台>_scrape.py 目标平台适配器
```
—— 这 4 个文件就能跑通单平台采集,不需要 engine/routes/前端。
---
## 十一、接手建议(按顺序)
```
第 1 步 装 OpenCLI + Playwright + 真实 Chrome,跑 validation.md 的 L0/L1
(这两级零外部依赖,能验证环境对不对)
第 2 步 跑 L2(Mock 测试)—— 验证降级逻辑,不需要真实平台
这时你已经能理解 _try_tools 的语义
第 3 步 登录目标平台,跑 L3(真实采集)—— 第一次真正拿数据
如果 OpenCLI 不通,会看到它降级到浏览器,日志里有 ⚠️
第 4 步 跑 L4/L5(引擎 + API)
第 5 步 替换前端(不要照搬 131 KB)
第 6 步 加你要的东西:限流 / 重试 / 并发控制(现在都没有)
```
---
## 十二、这个模块还没做的事(接手可以补)
```
① 限流与请求间隔 —— 现在纯串行、无间隔,目标多时可能触发平台风控
② 失败重试 —— 现在失败只记 errors
③ 并发控制 —— 没有信号量,大量目标只能串行
④ 异步态持久化 —— run 进度存内存,重启即丢
⑤ 登录态自动检测 —— check-login 端点有,但采集前没强制校验
⑥ 代理支持 —— 没有看到代理配置(多账号场景会需要)
```
---
## 附录:本说明的取证方式(可复现)
```bash
cd 05_tools/10_dashboard
# 架构
head -20 services/adapters/__init__.py
# 浏览器("套真实浏览器"的做法)
sed -n '1,70p' services/adapters/browser_helpers.py
# 降级机制
grep -n "_try_tools" -A 14 services/adapters/__init__.py
# 引擎流程
grep -nE "^ (async )?def " services/scrape_engine.py
# 表结构
grep -n "CREATE TABLE" -A 12 services/scrape_db.py
# 端点
grep -nE "^@router\." routes/scrape.py
```
+198
View File
@@ -0,0 +1,198 @@
# 与上游的差异管理
本仓库在 [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) 之上加了一层
监控/鉴权/多平台面板。这份文档记录**改了上游哪些文件、为什么**,以及**上游更新时怎么合并**。
---
## 一、改动分三类
冲突风险从低到高:
### 1. 纯新增文件(零冲突)
上游怎么改都不会碰到它们:
```
api/auth.py WebUI 登录鉴权
api/monitor/* 监控层整体(含 platforms.py 能力矩阵)
api/monitor/db.py MySQL 连接层(可回退 SQLite 供测试用)
api/monitor/migrate_from_sqlite.py SQLite → MySQL 一次性迁移脚本
api/monitor/upstream.py 上游更新检查(定时 fetch 上游并比对)
api/routers/{auth,monitor,settings}.py
api/schemas/{auth,monitor,settings}.py
api/services/interpreter.py 解释器探测(uv / .venv / 当前解释器)
webui/src/components/{monitor,settings,auth}/ 新视图
webui/src/components/layout/{PlatformSwitcher,UnwiredPlatformNotice}.tsx
webui/src/{hooks/useMonitor.ts,hooks/usePlatform.ts,store/platformStore.ts,lib/monitorFormat.ts,types/monitor.ts}
docs/监控功能使用说明.md
tests/test_{auth,settings,platforms,qrlogin,monitor_*,upstream}.py
Dockerfile / .dockerignore / docker-compose.yml 服务器部署用
```
> `api/monitor/upstream.py` 要调 `git`,而 `python:3.11-slim` 不带它 —— Dockerfile 里为此
> **显式装了 git**。改了 Dockerfile 就必须重建镜像(`docker compose build`),`./deploy.sh`
> 只重建前端,不重建镜像。
其中 `api/monitor/qrlogin.py` + `webui/.../QrLoginPanel.tsx` 是**服务器专用**的扫码登录:
那台机器上 Chrome 跑在 Xvfb 里,`show_qrcode` 调的 PIL `Image.show()` 需要桌面看图程序,
服务器没有,二维码会无处可去。所以改成用 CDP 把二维码从页面里读出来交给前端 `<img>` 显示。
### 2. 加法改动(低冲突)
只在既有文件里**新增**内容,不改动原有行:
| 文件 | 加了什么 |
|---|---|
| `cmd_arg/arg.py` | typer 选项:`--enable_cdp_mode`、`--inject_all_cookies`、`--save_login_state`、`--cookies_file`、`--crawler_max_sleep_sec`,以及对应的 `config.*` 回写 |
| `api/schemas/crawler.py` | `CrawlerStartRequest` 的若干**可选**字段(默认 `None`,不传则不加对应 CLI 参数) |
| `config/base_config.py` | `INJECT_ALL_COOKIES = False`;`MASK_NICKNAME = False`(关掉昵称脱敏,见第 3 节) |
| `api/routers/__init__.py` | 导出新增的 router |
| `requirements.txt` | 补上 `websockets`(上游 `pyproject.toml` 里有、`requirements.txt` 里漏了) |
| `tests/conftest.py` | 新增 `_bypass_auth_for_non_auth_suites` fixture |
### 3. 接线改动(中冲突,需要人看)
| 文件 | 改了什么 | 上游若在此处变动 |
|---|---|---|
| `api/main.py` | 注册 4 个 router 并加 `Depends(require_auth)`;`lifespan` 里初始化监控库、启动调度器、跑设置键迁移;`load_dotenv`;CORS 可配;`docs/redoc/openapi` 关闭;监听地址改 env | **最需要人工合并的文件**。留意 router 注册块、lifespan、`__main__` |
| `api/routers/websocket.py` | 两个 WS 路由加 `dependencies=[Depends(require_ws_auth)]` | 上游若新增 WS 路由,**必须同样加上**,否则那条流是裸奔的 |
| `api/services/crawler_manager.py` | 解释器探测替换硬编码 `uv run`;`_build_command` 转发新增参数;新增 `is_busy()` / `run_and_wait()` 与完成事件 | 留意 `_build_command` 的参数拼装 |
| `media_platform/xhs/login.py` | `login_by_cookies` 在 `INJECT_ALL_COOKIES` 打开时注入**全部** cookie(默认关闭,行为不变) | 小改动,好合并 |
| `tools/user_hash.py` | `mask_nickname` 改为读 `config.MASK_NICKNAME`,本仓库默认**不脱敏**(原样返回)。上游作为教学版默认脱敏,但那是有损的 ——「张三」「张四」都成「张*」,而分清谁是谁正是监控这一层要干的活。脱敏实现本身没删,改回 `True` 即恢复上游行为 | 与 `config/base_config.py` 一起改,两处不同步会不一致 |
### 4. 上游 bug 修复(建议回馈上游)
| 文件 | 修的问题 |
|---|---|
| `media_platform/xhs/core.py` | 见下节 1 |
| `media_platform/xhs/login.py` | 见下节 2(cookie 加固) |
| `media_platform/douyin/core.py` | 见下节 3(首页 `goto` 永远超时,采集根本起不来) |
| `media_platform/douyin/login.py` | 见下节 4(注入 cookie 后页面陈旧,白等十分钟) |
---
## 二、应该给上游提 PR 的四个修复
这四处都是**上游自身的缺陷**,提上去以后就不用自己背着:
### 1. 博主主页抓取失败会跳掉整个博主(`xhs/core.py`)
`get_creator_info()` 抓主页 HTML 解析 `window.__INITIAL_STATE__`,解析失败抛 `JSONDecodeError`——
它是 `ValueError` 的子类,被 `except ValueError` 误捕获,日志报成
"Failed to parse creator URL"(**误导**,URL 根本没解析错),然后 `continue` **跳过整个博主**。
而那份资料只喂给 `save_creator()`,它在教学版里是**空函数**。也就是说:
一个喂给空函数的抓取失败,让真正要抓的作品一条都没抓到,表现为"0 篇作品",
和"登录失效"长得一模一样。
修复:把资料抓取改成**尽力而为**,失败只警告、继续抓作品。
### 2. cookie 登录只注入 `web_session`(`xhs/login.py`)
`a1` / `webId` 等签名所需 cookie 只能靠持久化 profile 补,冷启动时签名会失败。
默认行为保持不变,用 `INJECT_ALL_COOKIES` 开关控制。
### 3. 抖音首页的 `goto` 永远等不到 `load`(`douyin/core.py:101`)
```python
await self.context_page.goto(self.index_url) # 默认 wait_until="load"
```
抖音首页的 `load` 事件**不会触发**(有长连接/埋点类请求一直挂着)。实测:同一台
Chrome、同一个地址,`domcontentloaded` 0.7 秒返回,而 `load` 等满 90 秒仍然超时。
后果是整个采集**一步都没走就崩**,退出码 1 —— 看起来像"抖音不能用"。
修复:显式 `wait_until="domcontentloaded"`。上游的贴吧(`tieba/core.py`)和知乎
(`zhihu/core.py`)本来就是这么写的,抖音这个页面只是恰好属于"永远不 load"的那类。
> 这个缺陷在本机可能复现不出来(换个网络/有缓存时 `load` 也许能触发),所以社区里
> 没人报。它和网络快慢无关:不是"慢",是那个事件根本不会发生。
### 4. 注入 cookie 后页面是陈旧的(`douyin/login.py:266`)
`login_by_cookies()` 把 cookie 塞进 browser context,但**页面是在这之前加载的** ——
SPA 只在加载时读一次登录态,`localStorage.HasUserLogin` 于是还停在"未登录",
紧接着的 `check_login_state()` 会对着这个陈旧的值轮询到超时(600 次 × 1 秒 = 十分钟),
然后 `sys.exit()`。**下一轮**才正常,因为那时 cookie 已经在 profile 里了。
表现是"第一次跑白等十分钟、第二次才行",很容易被当成偶发。
修复:注入完 cookie 后 `reload(wait_until="domcontentloaded")`,让站点立刻重新判定会话。
> 与第 1 条同源:都是"页面状态是加载那一刻的快照"。本仓库的扫码登录(`api/monitor/qrlogin.py`)
> 和运营模块也各自踩过这个坑,那里的判据改成了拿 cookie 问后台接口,而不是读页面快照。
---
## 三、上游更新时怎么操作
### 先让机器替你盯着
「上游更新检查」(`api/monitor/upstream.py`,开关在 WebUI 的**系统设置 → 上游更新**)会按
间隔 `git fetch` 上游、算出落后几个提交,有更新就推企业微信。它是这份文档的自动化版:
没有它,「上游动了」这件事只取决于谁偶尔想起来去 fetch 一次。
两个细节决定了它为什么是安全的:它只 fetch 到 `FETCH_HEAD`,**不写工作区、不建 remote、不碰
`refs/remotes`**,所以和正在跑的采集、和下面的 `git pull` 都不冲突;默认**关闭**,因为要联网,
且需要镜像里有 git。
### 日常流程
```bash
git stash # 或先 commit 到自己的分支(推荐)
git fetch origin main
git rebase origin/main # 冲突只会出现在上表第 3、4 类文件里
./.venv/Scripts/python.exe -m pytest tests/ -q # 502 个测试就是回归网
```
### 直连 GitHub 不通时(本机常见)
本机到 `github.com` 时通时不通,**大包传输必断**(`Recv failure: Connection was reset`
或 `unexpected disconnect while reading sideband packet`),所以 `git clone` / `--unshallow`
这类一次性拉全量的操作基本必失败。可用的替代源:
```bash
# gitcode 的 GitHub 镜像,国内直连,比 GitHub 本身还新一天以内
git remote add gitcode https://gitcode.com/gh_mirrors/me/MediaCrawler.git
git fetch --no-tags --unshallow gitcode # 本仓库当初就是这样补全历史的,约 2 秒
```
注意 `git fetch` 只写 `refs/remotes/*`,**不会动本地 `main`**;
但拉镜像会把 `upstream/main` 指到镜像的 tip(可能比 GitHub 晚一天),
等 GitHub 通了再 `git fetch upstream` 正回来即可。
### 强烈建议:先把改动提交掉
当前状态是**未提交**的(25 个上游文件被改 + 31 个新文件)。在 `main` 分支上裸着工作区,
一次 `git checkout .` 就全没了,而且没法 rebase。
```bash
git checkout -b local/monitor-panel
git add -A && git commit -m "监控面板 / 鉴权 / 多平台"
```
### 如果改动持续增长:fork
把本仓库 fork 到自己名下,上游设为 remote:
```bash
git remote rename origin upstream
git remote add origin <你的 fork>
git push -u origin local/monitor-panel
```
之后同步上游用 `git fetch upstream && git rebase upstream/main`。
---
## 四、合并时最容易忘的三件事
1. **新增的 `/api` 路由必须带鉴权**。跑一下 `tests/test_auth.py`——
里面有个测试会遍历 `app.routes`,断言除豁免集外每个 `/api` 路由无凭据都返回 401。
上游新增接口忘了加鉴权,这个测试会直接失败。
2. **新增的 WebSocket 路由必须加 `require_ws_auth`**。
`BaseHTTPMiddleware` 对 WS 完全不生效(`scope["type"] != "http"` 直接放行),
只靠中间件会漏。同样有测试守着。
3. **上游若改动 `AsyncFileWriter` 的输出路径规则**,`api/monitor/ingest.py::find_run_files`
会跟着失效——它靠 glob `{out_dir}/{platform}/jsonl/*_contents_*.jsonl` 定位每轮的产物。
+369
View File
@@ -0,0 +1,369 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/auth.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Authentication for the WebUI.
Design constraints that drove this, all verified against the codebase:
* **Cookies, not bearer headers, are the primary transport.** Browser
WebSockets cannot set custom headers on the handshake, and the data-export
downloads use ``window.open`` (a navigation, also header-less). Only a cookie
is carried on both. The same opaque token is *also* accepted from an
``Authorization: Bearer`` header so ``curl`` and scripts remain usable.
* **Enforcement is a ``Depends``, not middleware.** ``BaseHTTPMiddleware``
returns early for any non-``http`` scope, so it never sees a WebSocket --
a middleware-only gate would leave the live log stream wide open. It is also
overridable per-test via ``app.dependency_overrides``.
* **Sessions are server-side** so logout and password-change revoke immediately.
Only the environment variable ``MC_PASSWORD`` can bypass the stored hash. That is
the documented way back in if the password is forgotten, which is why it is
never persisted.
"""
import asyncio
import base64
import binascii
import hashlib
import hmac
import os
import secrets
import time
from typing import Optional
from anyio import to_thread
from fastapi import HTTPException, Request, WebSocket, WebSocketException, status
from sqlalchemy import delete
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from .monitor.db import get_session
from .monitor.models import (
SETTING_AUTH_PASSWORD_HASH,
SETTING_AUTH_PASSWORD_UPDATED_AT,
AuthSession,
)
from .monitor.settings import get_setting, set_setting
SESSION_COOKIE_NAME = "mc_session"
# OWASP's current PBKDF2-HMAC-SHA256 guidance. Deliberately slow -- see
# verify_password() for why that cost must not land on the event loop.
PBKDF2_ITERATIONS = 600_000
PBKDF2_ALGO = "pbkdf2_sha256"
# A single generic message for every failure mode, so the response never
# reveals whether a password is set, wrong, or empty.
INVALID_CREDENTIALS = "用户名或密码错误"
# Brute-force throttle. In-process is sufficient: this is a single-user tool and
# uvicorn runs one worker. Documented as reset-on-restart.
THROTTLE_THRESHOLD = 5
THROTTLE_WINDOW_SECONDS = 900
THROTTLE_MAX_LOCKOUT_SECONDS = 900
_failures: dict[str, list[float]] = {}
_throttle_lock = asyncio.Lock()
def _now() -> float:
"""Monotonic clock, indirected so tests can drive it without sleeping."""
return time.monotonic()
# ---------------------------------------------------------------------------
# Environment configuration (read at call time so tests can set it per-case)
# ---------------------------------------------------------------------------
def env_password() -> str:
return os.getenv("MC_PASSWORD", "").strip()
def cookie_secure() -> bool:
return os.getenv("MC_COOKIE_SECURE", "").strip().lower() in ("1", "true", "yes", "y")
def session_ttl_ms() -> int:
try:
hours = int(os.getenv("MC_SESSION_TTL_HOURS", "336"))
except ValueError:
hours = 336
return max(hours, 1) * 3_600_000
# ---------------------------------------------------------------------------
# Password hashing
# ---------------------------------------------------------------------------
def _b64(raw: bytes) -> str:
return base64.b64encode(raw).decode("ascii")
def hash_password(password: str, *, iterations: Optional[int] = None) -> str:
"""Return a self-describing hash so the iteration count can be raised later
without a migration: ``pbkdf2_sha256$<iterations>$<salt>$<hash>``.
``iterations`` is resolved at call time (not bound as a default) so tests can
lower it; the production value stays the module constant.
"""
iterations = iterations or PBKDF2_ITERATIONS
salt = secrets.token_bytes(16)
digest = hashlib.pbkdf2_hmac("sha256", password.encode("utf-8"), salt, iterations)
return f"{PBKDF2_ALGO}${iterations}${_b64(salt)}${_b64(digest)}"
def _verify_password_sync(password: str, stored: str) -> bool:
try:
algo, iterations_raw, salt_raw, digest_raw = stored.split("$")
if algo != PBKDF2_ALGO:
return False
salt = base64.b64decode(salt_raw)
expected = base64.b64decode(digest_raw)
actual = hashlib.pbkdf2_hmac("sha256", password.encode("utf-8"), salt, int(iterations_raw))
except (ValueError, TypeError, binascii.Error):
return False
return hmac.compare_digest(actual, expected)
async def verify_password(password: str, stored: str) -> bool:
"""Verify off the event loop.
At 600k iterations this takes a few hundred milliseconds. Running it inline
in an async handler would block the loop entirely -- stalling the monitor
scheduler and every websocket ping -- and present as "the whole UI freezes
when I click login".
"""
return await to_thread.run_sync(_verify_password_sync, password, stored)
async def current_password_hash(session: AsyncSession) -> str:
return (await get_setting(session, SETTING_AUTH_PASSWORD_HASH)) or ""
async def set_password(session: AsyncSession, password: str) -> None:
await set_setting(session, SETTING_AUTH_PASSWORD_HASH, hash_password(password))
await set_setting(
session, SETTING_AUTH_PASSWORD_UPDATED_AT, str(get_current_timestamp())
)
async def check_password(session: AsyncSession, password: str) -> bool:
"""The environment override wins over the stored hash, always.
That is the escape hatch: forgetting the password is recoverable by setting
MC_PASSWORD and restarting, without touching the database.
"""
override = env_password()
if override:
return hmac.compare_digest(password, override)
stored = await current_password_hash(session)
if not stored:
return False
return await verify_password(password, stored)
async def ensure_initial_credential() -> Optional[str]:
"""Seed a password on first run; returns it once so main() can print it.
Deliberately NOT an unauthenticated "set your password" endpoint: on a
LAN-exposed bind that is a claim-the-instance race where whoever reaches the
page first becomes the administrator. Generating and printing a random
password avoids the race and also avoids locking the operator out.
"""
if env_password():
return None
async with get_session() as session:
if await current_password_hash(session):
return None
generated = secrets.token_urlsafe(12)
await set_password(session, generated)
return generated
# ---------------------------------------------------------------------------
# Sessions
# ---------------------------------------------------------------------------
def _hash_token(token: str) -> str:
return hashlib.sha256(token.encode("utf-8")).hexdigest()
async def create_session(session: AsyncSession) -> tuple[str, int]:
"""Issue a session. Returns (token, expires_at_ms).
The caller receives the raw token; only its hash is stored.
"""
token = secrets.token_urlsafe(32)
now = get_current_timestamp()
expires_at = now + session_ttl_ms()
session.add(
AuthSession(
token_hash=_hash_token(token),
created_at=now,
expires_at=expires_at,
last_seen_at=now,
)
)
return token, expires_at
async def resolve_session(session: AsyncSession, token: str) -> Optional[AuthSession]:
if not token:
return None
row = await session.get(AuthSession, _hash_token(token))
if row is None:
return None
now = get_current_timestamp()
if row.expires_at <= now:
await session.delete(row)
return None
row.last_seen_at = now
return row
async def revoke_session(session: AsyncSession, token: str) -> None:
row = await session.get(AuthSession, _hash_token(token))
if row is not None:
await session.delete(row)
async def revoke_all_sessions(session: AsyncSession) -> int:
"""Used on password change, which is what makes "all devices logged out"
take effect immediately rather than at token expiry."""
result = await session.execute(delete(AuthSession))
return result.rowcount or 0
async def purge_expired_sessions(session: AsyncSession) -> None:
await session.execute(delete(AuthSession).where(AuthSession.expires_at <= get_current_timestamp()))
# ---------------------------------------------------------------------------
# Credential extraction and enforcement
# ---------------------------------------------------------------------------
def token_from_request(request: Request) -> str:
"""Cookie first (browsers, websockets, navigations), then Bearer (scripts)."""
token = request.cookies.get(SESSION_COOKIE_NAME, "")
if token:
return token
header = request.headers.get("authorization", "")
if header.lower().startswith("bearer "):
return header[7:].strip()
return ""
def _unauthorized() -> HTTPException:
return HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED,
detail=INVALID_CREDENTIALS,
headers={"WWW-Authenticate": "Bearer"},
)
async def require_auth(request: Request) -> None:
"""FastAPI dependency guarding the protected routers.
Applied per-router via ``include_router(..., dependencies=[Depends(...)])``
rather than as app-wide middleware, so it appears in the OpenAPI schema,
returns a correct 401, and can be overridden in tests.
"""
token = token_from_request(request)
if not token:
raise _unauthorized()
async with get_session() as session:
if await resolve_session(session, token) is None:
raise _unauthorized()
async def require_ws_auth(websocket: WebSocket) -> None:
"""Guard for WebSocket routes.
These need their own dependency: ``BaseHTTPMiddleware`` passes any non-http
scope straight through, and router-level HTTP dependencies do not apply to
websocket routes. Raising ``WebSocketException`` closes the handshake with
the given code; ``HTTPException`` would be meaningless here.
"""
token = websocket.cookies.get(SESSION_COOKIE_NAME, "")
if not token:
raise WebSocketException(code=status.WS_1008_POLICY_VIOLATION)
async with get_session() as session:
if await resolve_session(session, token) is None:
raise WebSocketException(code=status.WS_1008_POLICY_VIOLATION)
# ---------------------------------------------------------------------------
# Brute-force throttle
# ---------------------------------------------------------------------------
def client_key(request: Request) -> str:
"""Identify the caller for throttling.
``X-Forwarded-For`` is only consulted when the operator explicitly opts in,
because otherwise any client could spoof the header and throttle someone
else (or evade its own throttle).
"""
if os.getenv("MC_TRUST_PROXY", "").strip() == "1":
forwarded = request.headers.get("x-forwarded-for", "")
if forwarded:
return forwarded.split(",")[0].strip()
return request.client.host if request.client else "unknown"
def _recent_failures(key: str) -> list[float]:
cutoff = _now() - THROTTLE_WINDOW_SECONDS
return [ts for ts in _failures.get(key, []) if ts >= cutoff]
async def retry_after_seconds(key: str) -> int:
"""0 when not throttled, otherwise how long the caller must wait."""
async with _throttle_lock:
recent = _recent_failures(key)
_failures[key] = recent
if len(recent) < THROTTLE_THRESHOLD:
return 0
# Lockout doubles per failure past the threshold, capped.
extra = len(recent) - THROTTLE_THRESHOLD
lockout = min(2 ** extra, THROTTLE_MAX_LOCKOUT_SECONDS)
elapsed = _now() - recent[-1]
remaining = int(lockout - elapsed)
return max(remaining, 1)
async def record_failure(key: str) -> None:
async with _throttle_lock:
_failures.setdefault(key, []).append(_now())
async def clear_failures(key: str) -> None:
async with _throttle_lock:
_failures.pop(key, None)
def reset_throttle_state() -> None:
"""Test hook: drop all throttle state."""
_failures.clear()
+27
View File
@@ -0,0 +1,27 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/__init__.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营模块:管理自己的小红书账号,读取创作者后台的数据。
与 `api.monitor` 是**并列关系**,不是它的扩展。两者数据形状不同:监控是「每轮
快照 + 差分」的公开互动数据,这里是创作者后台按日期给出的曝光/观看/完播等运营
指标。硬塞进同一个模型会同时污染两边。
路线是**纯请求**(无浏览器),依据见 tools/probe_creator_api.py 的 Phase 0 实测:
签名可自造、主站 cookie 即可认证、接口与参数已与真实页面对齐。
"""
+345
View File
@@ -0,0 +1,345 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/client.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""小红书创作者后台的纯请求客户端。
无浏览器:签名在本地算(见 signing.py),请求走 httpx。
**关于字段名的谨慎**:Phase 0 抓到的响应里,列表接口因为账号权限未生效而返回空壳
(`data.result` 只有 `{success, code, message}` 没有数据),所以**真实字段名尚未亲眼
见过**。因此每个指标都写成**多别名匹配**,并且解析不出来时存 `None` 而不是 0 ——
0 是真实值,None 是"不知道",两者混淆会让报表说谎。
"""
import re
from typing import Any, Dict, List, Optional
import httpx
from .signing import sign_xyw, signed_api
CREATOR_ORIGIN = "https://creator.xiaohongshu.com"
DATA_ANALYSIS_PAGE = f"{CREATOR_ORIGIN}/statistics/data-analysis"
USER_INFO_PATH = "/api/galaxy/user/info"
PERMISSION_PATH = "/api/galaxy/creator/datacenter/permission/query"
NOTE_LIST_PATH = "/api/galaxy/creator/datacenter/note/analyze/list"
USER_AGENT = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
)
# 签名被网关拒绝时返回的响应体。区分它和普通业务错误很重要:406 说明签名写错了,
# 而应用层的 code/-100 说明签名没问题、只是没有登录态。
SIGNATURE_REJECTED_MARKERS = ("code", -1)
class CreatorApiError(RuntimeError):
"""调用创作者后台失败。``status`` 用于区分是网关拒绝还是业务错误。"""
def __init__(self, message: str, status: int = 0, payload: Any = None):
super().__init__(message)
self.status = status
self.payload = payload
def trans_cookies(cookie_str: str) -> Dict[str, str]:
"""把 cookie 字符串解析成字典。容忍末尾分号、换行和零散空格。"""
jar: Dict[str, str] = {}
for chunk in re.split(r"[;\n]", cookie_str or ""):
chunk = chunk.strip()
if not chunk or "=" not in chunk:
continue
name, value = chunk.split("=", 1)
name = name.strip()
if name:
jar[name] = value.strip()
return jar
def _pick(item: Dict[str, Any], *names: str) -> Any:
"""按别名顺序取第一个存在的键。
字段名来自二手资料,未亲眼验证,所以不赌单一命名。
"""
for name in names:
if name in item and item[name] is not None:
return item[name]
return None
_COUNT_UNITS = {"万": 10_000, "w": 10_000, "W": 10_000, "亿": 100_000_000, "k": 1_000, "K": 1_000}
def as_int(value: Any) -> Optional[int]:
"""解析计数。处理 "1.2万"、"1,234"、"123" 与已经是数字的情况。"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, (int, float)):
return int(value)
text = str(value).strip().replace(",", "").replace(" ", "")
if not text or text in ("-", "--", "暂无"):
return None
unit = 1
suffix = text[-1]
if suffix in _COUNT_UNITS:
unit = _COUNT_UNITS[suffix]
text = text[:-1]
try:
return int(float(text) * unit)
except ValueError:
return None
def as_float(value: Any) -> Optional[float]:
"""解析比率。``"12.3%"`` -> 12.3;``"0.123"`` 原样返回数字。"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, (int, float)):
return float(value)
text = str(value).strip().replace("%", "")
try:
return float(text)
except ValueError:
return None
_DURATION_RE = re.compile(r"(?:(\d+)\s*分)?\s*(?:(\d+)\s*秒)?")
def as_seconds(value: Any) -> Optional[float]:
"""解析时长。处理 ``"1分30秒"``、``"01:30"``、``"45"``(秒)。"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, (int, float)):
return float(value)
text = str(value).strip()
if not text or text in ("-", "--"):
return None
if ":" in text:
parts = text.split(":")
try:
total = 0.0
for part in parts:
total = total * 60 + float(part)
return total
except ValueError:
return None
if "分" in text or "秒" in text:
minutes = re.search(r"(\d+)\s*分", text)
seconds = re.search(r"(\d+)\s*秒", text)
if not minutes and not seconds:
return None
return float(minutes.group(1) if minutes else 0) * 60 + float(
seconds.group(1) if seconds else 0
)
try:
return float(text)
except ValueError:
return None
# 指标 -> 候选原始字段名。别名来自公开资料,权威与否只能等真实响应来验证。
_FIELD_ALIASES: Dict[str, tuple] = {
"exposure": ("exposure", "exposure_count", "imp", "impression", "impression_count"),
"views": ("views", "view", "view_count", "watch", "watch_count", "read_count"),
"likes": ("likes", "like", "like_count", "liked_count"),
"comments": ("comments", "comment", "comment_count", "comments_count"),
"favorites": ("favorites", "favorite", "favorite_count", "collect", "collect_count", "collected_count"),
"shares": ("shares", "share", "share_count", "shared_count"),
"new_followers": ("new_followers", "fans_growth", "follower_growth", "increase_fans"),
"danmaku": ("danmaku", "danmaku_count", "barrage"),
"cover_ctr": ("cover_ctr", "cover_click_rate", "cover_click_ratio", "ctr"),
"avg_watch_seconds": ("avg_watch_seconds", "avg_watch_time", "average_watch_time"),
"two_second_exit_rate": ("two_second_exit_rate", "2s_exit_rate", "exit_rate_2s"),
"completion_rate": ("completion_rate", "finish_rate", "complete_rate"),
}
_COUNT_FIELDS = {
"exposure", "views", "likes", "comments", "favorites", "shares", "new_followers", "danmaku",
}
_DURATION_FIELDS = {"avg_watch_seconds"}
_RATE_FIELDS = {"two_second_exit_rate", "completion_rate", "cover_ctr"}
NOTE_ID_ALIASES = ("note_id", "noteId", "id", "content_id", "item_id")
TITLE_ALIASES = ("title", "content", "note_title", "display_title")
PUBLISH_TIME_ALIASES = ("publish_time", "publishTime", "post_time", "create_time", "time")
def _normalize_metric(name: str, raw: Any) -> Any:
if name in _COUNT_FIELDS:
return as_int(raw)
if name in _DURATION_FIELDS:
return as_seconds(raw)
if name in _RATE_FIELDS:
return as_float(raw)
return raw
def normalize_note(item: Dict[str, Any]) -> Dict[str, Any]:
"""把一条原始记录规范化成落库用的字段。"""
note: Dict[str, Any] = {
"note_id": str(_pick(item, *NOTE_ID_ALIASES) or ""),
"title": str(_pick(item, *TITLE_ALIASES) or ""),
"publish_time": as_int(_pick(item, *PUBLISH_TIME_ALIASES)),
}
for name, aliases in _FIELD_ALIASES.items():
note[name] = _normalize_metric(name, _pick(item, *aliases))
return note
def find_note_list(payload: Any) -> List[Dict[str, Any]]:
"""在响应里找出笔记数组。
接口的**确切结构还没亲眼见过**(权限未生效时 `data.result` 里没有数据),
所以不写死路径:遍历 JSON,挑出"看起来像一批笔记记录"的那个列表 ——
元素是 dict,且至少带一个指标字段。找不到就返回空列表,让上层如实报"没数据"。
"""
best: List[Dict[str, Any]] = []
metric_keys = {alias for aliases in _FIELD_ALIASES.values() for alias in aliases}
def walk(value: Any) -> None:
nonlocal best
if isinstance(value, dict):
for child in value.values():
walk(child)
elif isinstance(value, list):
if value and isinstance(value[0], dict):
keys = set(value[0].keys())
if keys & metric_keys and len(value) > len(best):
best = value
for child in value:
walk(child)
walk(payload)
return best
class CreatorClient:
"""一个账号的客户端。``cookie`` 就是它的全部身份。"""
def __init__(self, cookie: str, timeout: float = 25.0):
self.cookies = trans_cookies(cookie)
self.a1 = self.cookies.get("a1", "")
self._timeout = timeout
@property
def looks_authenticated(self) -> bool:
"""签名需要 a1;没有它连请求都签不出来。"""
return bool(self.a1)
def _headers(self, api: str, body: dict | None = None) -> Dict[str, str]:
cookie_header = "; ".join(f"{k}={v}" for k, v in self.cookies.items())
return {
"user-agent": USER_AGENT,
"accept": "application/json, text/plain, */*",
"accept-language": "zh-CN,zh;q=0.9",
"origin": CREATOR_ORIGIN,
"referer": DATA_ANALYSIS_PAGE,
"cookie": cookie_header,
**sign_xyw(api, self.a1, body=body),
}
async def _get(self, path: str, params: Dict[str, Any]) -> Dict[str, Any]:
if not self.looks_authenticated:
raise CreatorApiError("cookie 里没有 a1,无法完成签名", status=0)
query = "&".join(f"{k}={v}" for k, v in params.items())
api = signed_api(path, query)
url = f"{CREATOR_ORIGIN}{path}?{query}" if query else f"{CREATOR_ORIGIN}{path}"
async with httpx.AsyncClient(timeout=self._timeout, follow_redirects=False) as client:
response = await client.get(url, headers=self._headers(api))
return self._unwrap(response)
@staticmethod
def _unwrap(response: httpx.Response) -> Dict[str, Any]:
if response.status_code == 406:
raise CreatorApiError(
"签名被网关拒绝(406)—— 待签字符串的拼法不对", status=406
)
try:
payload = response.json()
except Exception as exc: # noqa: BLE001
raise CreatorApiError(
f"响应不是 JSON(HTTP {response.status_code})", status=response.status_code
) from exc
if response.status_code == 401 or payload.get("code") == -100:
raise CreatorApiError("登录态无效或已过期", status=401, payload=payload)
if not payload.get("success", True):
raise CreatorApiError(
str(payload.get("msg") or "接口返回失败"),
status=response.status_code,
payload=payload,
)
return payload
async def fetch_user_info(self) -> Dict[str, Any]:
"""当前 cookie 属于哪个账号。登录后用它取名与去重。"""
payload = await self._get(USER_INFO_PATH, {})
data = payload.get("data") or {}
return {
"user_id": str(data.get("userId") or ""),
"nickname": str(data.get("userName") or ""),
"avatar": str(data.get("userAvatar") or ""),
"red_id": str(data.get("redId") or ""),
"role": str(data.get("role") or ""),
"permissions": list(data.get("permissions") or []),
}
async def fetch_permission(self) -> Dict[str, Any]:
"""数据权限状态。
``tip_msg`` 是后台原话(实测是"已为您申请数据权限,次日可查看"),照抄不改写 ——
这条信息必须原样交给用户,它解释了"为什么没有数据"。
"""
payload = await self._get(PERMISSION_PATH, {})
data = payload.get("data") or {}
return {
"display": data.get("display"),
"status": data.get("status"),
"tip": str(data.get("tip_msg") or ""),
}
async def fetch_note_list(
self, start_ms: int, end_ms: int, page_num: int = 1, page_size: int = 10, note_type: int = 0
) -> List[Dict[str, Any]]:
"""按发布时间区间取笔记列表。
参数与顺序**照抄真实页面的请求**(见 Phase 0 抓包),不要凭感觉改。
"""
payload = await self._get(
NOTE_LIST_PATH,
{
"post_begin_time": start_ms,
"post_end_time": end_ms,
"type": note_type,
"page_size": page_size,
"page_num": page_num,
},
)
return [normalize_note(item) for item in find_note_list(payload)]
+300
View File
@@ -0,0 +1,300 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/login.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营账号的扫码登录。
**与监控的扫码登录(`api.monitor.qrlogin`)有一处决定性差异**:那边把登录态写进
浏览器**默认 profile**,因为爬虫要复用它;这边要的是 **cookie 字符串**,因为采集
走纯 HTTP。所以这里每次登录都开一个**临时上下文**,扫完取出 cookie 就丢弃 ——
* 登第二个账号不会把第一个顶掉(默认 profile 只能装一个登录态);
* 完全不影响监控那个登录态;
* 十个账号互不干扰。
扫码入口仍是主站(`www.xiaohongshu.com`):Phase 0 实测证明**主站的 cookie 就能
认证创作者后台**,不需要单独的创作者登录。
"""
import asyncio
import time
from typing import Any, Dict, Optional
from playwright.async_api import async_playwright
from tools import utils
from ..monitor.platforms import PLATFORM_XHS
from .client import CreatorApiError, CreatorClient
QR_TTL_SECONDS = 300
STATUS_IDLE = "idle"
STATUS_WAITING = "waiting"
STATUS_SUCCESS = "success"
STATUS_EXPIRED = "expired"
STATUS_ERROR = "error"
LOGIN_URL = "https://www.xiaohongshu.com"
QR_SELECTOR = "xpath=//img[@class='qrcode-img']"
LOGIN_BUTTON_SELECTOR = "xpath=//*[@id='app']/div[1]/div[2]/div[1]/ul/div[1]/button"
# 判据不读页面状态,而是拿 cookie 直接问创作者后台"我是谁"。
#
# **为什么不用页面状态**:`window.__INITIAL_STATE__` 是**页面加载那一刻的快照**。
# 监控那边的同一个探针能用,是因为那台浏览器的页面加载时就已经登录了,快照里
# loggedIn 就是 true。而扫码是"页面加载之后才登录的"—— SPA 内部确实登进去了,
# 但那个初始快照不会翻转,于是检测永远等不到,界面就一直停在二维码上。
#
# `user/info` 则是权威的:实测**游客也会拿到 a1**(所以签名算得出来),但接口直接
# 回 401「无登录信息」;只有真正登录了才返回 user_id。所以"有 a1"什么都证明不了,
# "后台认这份身份"才是。
LOGIN_CHECK_INTERVAL_SECONDS = 5.0
_lock = asyncio.Lock()
_current: Optional["AccountLoginSession"] = None
# 完成后的快照。会话一旦被取走 cookie 就会拆掉,而前端可能还有一个在飞的轮询——
# 那个请求若看到 _current 为空就会回报 idle,把已经显示出来的成功状态又擦掉。
# 把结果留在这里,重复轮询就稳定得多。
_last_result: Optional[Dict[str, Any]] = None
_playwright: Any = None
def _cdp_url() -> str:
import os
import config
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
async def _connect():
global _playwright
if _playwright is None:
_playwright = await async_playwright().start()
return await _playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
async def _disconnect() -> None:
global _playwright
if _playwright is not None:
try:
await _playwright.stop()
except Exception:
pass
_playwright = None
async def _read_qr(page: Any) -> str:
image = await utils.find_login_qrcode(page, selector=QR_SELECTOR)
if image:
return image
# 登录框不一定自己弹出来,这是爬虫自身扫码流程的同款兜底。
await asyncio.sleep(0.5)
try:
await page.locator(LOGIN_BUTTON_SELECTOR).click(timeout=5000)
except Exception:
return ""
return await utils.find_login_qrcode(page, selector=QR_SELECTOR)
class AccountLoginSession:
"""一次针对**临时上下文**的扫码尝试。"""
def __init__(self, context: Any, page: Any) -> None:
self.status = STATUS_WAITING
self.message = "请用手机扫描二维码"
self.image = ""
self.started_at = time.time()
self.account: Optional[Dict[str, Any]] = None
self.cookie: str = ""
self.platform = PLATFORM_XHS
self._context = context
self._page = page
self._last_login_check = 0.0
@property
def elapsed(self) -> float:
return time.time() - self.started_at
async def refresh(self) -> None:
if self.status != STATUS_WAITING:
return
if self.elapsed > QR_TTL_SECONDS:
self.status = STATUS_EXPIRED
self.message = "二维码已超时,请重新获取"
return
# 前端每 2 秒问一次,但没必要每次都去打后台接口 —— 一次真实的网络往返
# 去确认一个通常还没发生的事件是浪费。
now = time.time()
if now - self._last_login_check < LOGIN_CHECK_INTERVAL_SECONDS:
return
self._last_login_check = now
try:
cookies = await self._context.cookies()
except Exception:
self.status = STATUS_ERROR
self.message = "登录窗口已被关闭,请重新获取"
return
cookie = "; ".join(f"{c['name']}={c['value']}" for c in cookies)
try:
info = await CreatorClient(cookie).fetch_user_info()
except CreatorApiError:
# 还没登录(或者刚扫、后端还没认),继续等。
return
if not info.get("user_id"):
return
# 登录成功:cookie 取自**这个临时上下文**,取完上下文就丢弃,
# 所以不会残留、也不会影响别的账号。
self.cookie = cookie
self.account = info
self.status = STATUS_SUCCESS
self.message = f"登录成功:{info.get('nickname') or info['user_id']}"
def snapshot(self) -> Dict[str, Any]:
return {
"status": self.status,
"message": self.message,
"image": self.image,
"elapsed": int(self.elapsed),
"expires_in": max(0, int(QR_TTL_SECONDS - self.elapsed)),
"account": self.account,
}
async def close(self) -> None:
"""关掉临时上下文。这是它存在的全部意义 —— 用完即弃。"""
for closer in (self._page.close, self._context.close):
try:
await closer()
except Exception:
pass
async def _teardown_locked() -> None:
global _current
if _current is not None:
await _current.close()
_current = None
async def start() -> Dict[str, Any]:
"""开一个临时上下文,打开登录页,取回二维码。"""
global _current, _last_result
async with _lock:
_last_result = None
await _teardown_locked()
try:
browser = await _connect()
except Exception as exc:
await _disconnect()
raise RuntimeError(
f"连接浏览器失败({_cdp_url()})。请确认服务器上的 Chrome 以 "
f"--remote-debugging-port 启动。原始错误:{exc}"
) from exc
# 临时上下文,不是 contexts[0]。这里刻意要一个干净的身份 ——
# 借用操作者自己的登录态会让"新增账号"变成"再读一遍当前账号"。
context = await browser.new_context()
page = await context.new_page()
try:
await page.goto(LOGIN_URL, wait_until="domcontentloaded", timeout=45000)
image = await _read_qr(page)
except Exception as exc:
try:
await page.close()
await context.close()
except Exception:
pass
raise RuntimeError(f"打开登录页失败:{exc}") from exc
session = AccountLoginSession(context, page)
session.image = image
if not image:
session.status = STATUS_ERROR
session.message = "页面上没找到二维码,请确认站点结构没有变化"
_current = session
return session.snapshot()
async def status() -> Dict[str, Any]:
async with _lock:
if _current is None:
if _last_result is not None:
return _last_result
return {
"status": STATUS_IDLE,
"message": "",
"image": "",
"elapsed": 0,
"expires_in": 0,
"account": None,
}
await _current.refresh()
return _current.snapshot()
async def remember_result(snapshot: Dict[str, Any]) -> None:
"""记住已完成的扫码结果,供后续轮询重复返回。"""
global _last_result
async with _lock:
_last_result = snapshot
async def take_cookie() -> Optional[str]:
"""取走已登录的 cookie 并结束会话。
由路由层在落库时调用。cookie 只经内存传递,**不进响应体** —— 它是凭证,
前端没有任何理由看到它。
"""
global _current
async with _lock:
if _current is None or _current.status != STATUS_SUCCESS:
return None
cookie = _current.cookie
await _teardown_locked()
return cookie
async def cancel() -> Dict[str, Any]:
global _last_result
async with _lock:
_last_result = None
await _teardown_locked()
return {
"status": STATUS_IDLE,
"message": "已取消",
"image": "",
"elapsed": 0,
"expires_in": 0,
"account": None,
}
async def shutdown() -> None:
async with _lock:
await _teardown_locked()
await _disconnect()
+132
View File
@@ -0,0 +1,132 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/models.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营模块的数据模型。
**刻意复用 `MonitorBase`**:这样 `init_db` 的 `create_all` 会顺手建出新表,而
`_ensure_columns`(已改为按模型元数据推导)也会自动给新表补字段 —— 不必再维护一份
建表语句。表落在同一个库里,与监控互不干扰。
"""
from typing import Optional
from sqlalchemy import BigInteger, Float, ForeignKey, Index, Integer, String, Text
from sqlalchemy.orm import Mapped, mapped_column, relationship
from ..monitor.models import MonitorBase
# 创作者后台的数据权限是「首次访问时自动申请、次日生效」。这个状态必须如实呈现:
# 显示成"没数据"会让人以为采集坏了,实际是在等审批。
PERMISSION_UNKNOWN = "unknown"
PERMISSION_PENDING = "pending" # 已申请,未生效(提示语:"次日可查看")
PERMISSION_ACTIVE = "active"
PERMISSION_MISSING = "missing" # 接口明确说没有权限
# 账号自身的可用性。
ACCOUNT_OK = "ok"
ACCOUNT_EXPIRED = "expired" # cookie 失效,需要重新扫码
ACCOUNT_ERROR = "error"
class CreatorAccount(MonitorBase):
"""一个自己的小红书账号。
纯请求路线下,**一个账号的全部身份就是一份 cookie** —— 没有浏览器 profile、
没有独立目录。所以"多账号"在这里只是表里的多行,不是多套运行环境。
`cookie` 是凭证:与监控的 cookie 同样对待,只存库、绝不回显接口。
"""
__tablename__ = "creator_account"
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
# 展示名。优先用后台返回的昵称,用户可以改。
nickname: Mapped[str] = mapped_column(String(128), nullable=False, default="")
# 创作者后台的账号标识,由 /api/galaxy/user/info 返回,用于去重。
user_id: Mapped[str] = mapped_column(String(64), nullable=False, default="", index=True)
red_id: Mapped[str] = mapped_column(String(64), nullable=False, default="")
avatar: Mapped[str] = mapped_column(Text, nullable=False, default="")
cookie: Mapped[str] = mapped_column(Text, nullable=False, default="")
status: Mapped[str] = mapped_column(String(16), nullable=False, default=ACCOUNT_OK)
permission_status: Mapped[str] = mapped_column(
String(16), nullable=False, default=PERMISSION_UNKNOWN
)
# 后台原话,例如"已为您申请数据权限,次日可查看"。照抄,不改写。
permission_tip: Mapped[str] = mapped_column(Text, nullable=False, default="")
last_checked_at: Mapped[Optional[int]] = mapped_column(BigInteger)
last_synced_at: Mapped[Optional[int]] = mapped_column(BigInteger)
# 上次同步用的时间范围(天,按发布时间)。当前展示的数据就是这个范围的产物 ——
# 不记下来的话,界面只能说明"同步过了",说不清是哪一段。
last_sync_days: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
last_error: Mapped[Optional[str]] = mapped_column(Text)
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
notes: Mapped[list["CreatorNoteStat"]] = relationship(
back_populates="account", cascade="all, delete-orphan"
)
class CreatorNoteStat(MonitorBase):
"""一篇作品在某个采集时点的运营数据。
创作者后台给的是**累计值**(截至查询时点),所以反复采集天然形成时间序列 ——
与监控的"快照 + 差分"是同一个思路,因此这里保留 `captured_at` 而不是覆盖写。
"""
__tablename__ = "creator_note_stat"
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
account_id: Mapped[int] = mapped_column(
ForeignKey("creator_account.id", ondelete="CASCADE"), nullable=False, index=True
)
note_id: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
# 发布时间(毫秒)。后台按发布时间筛选,这是它的主时间轴。
publish_time: Mapped[Optional[int]] = mapped_column(BigInteger)
# --- 运营指标 ---------------------------------------------------------
# 计数用 BigInteger:曝光量可以很大,用 INT 迟早溢出。
exposure: Mapped[Optional[int]] = mapped_column(BigInteger)
views: Mapped[Optional[int]] = mapped_column(BigInteger)
likes: Mapped[Optional[int]] = mapped_column(BigInteger)
comments: Mapped[Optional[int]] = mapped_column(BigInteger)
favorites: Mapped[Optional[int]] = mapped_column(BigInteger)
shares: Mapped[Optional[int]] = mapped_column(BigInteger)
new_followers: Mapped[Optional[int]] = mapped_column(BigInteger)
danmaku: Mapped[Optional[int]] = mapped_column(BigInteger)
# 比率与时长。后台返回的可能是 "12.3%"/"1分30秒" 这类字符串,解析不了的存 NULL
# 而不是 0 —— 与监控层的口径一致:0 是真实值,NULL 是"不知道"。
cover_ctr: Mapped[Optional[float]] = mapped_column(Float)
avg_watch_seconds: Mapped[Optional[float]] = mapped_column(Float)
two_second_exit_rate: Mapped[Optional[float]] = mapped_column(Float)
completion_rate: Mapped[Optional[float]] = mapped_column(Float)
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
account: Mapped["CreatorAccount"] = relationship(back_populates="notes")
__table_args__ = (
# 同一个时点同一篇只留一行,重复同步不会堆积。
Index("ix_creator_note_stat_unique", "account_id", "note_id", "captured_at", unique=True),
)
+360
View File
@@ -0,0 +1,360 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/service.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营账号的增删查与数据同步。
一条贯穿全文件的规则:**cookie 是凭证,永远不出现在返回给上层的结构里。**
对外只给 `has_cookie` 这样的布尔量,与监控层对 cookie 的处理保持一致。
"""
import asyncio
from datetime import datetime, time, timedelta
from typing import Any, Dict, List, Optional
from sqlalchemy import delete, func, select
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from .client import CreatorApiError, CreatorClient
from .models import (
ACCOUNT_ERROR,
ACCOUNT_EXPIRED,
ACCOUNT_OK,
PERMISSION_ACTIVE,
PERMISSION_MISSING,
PERMISSION_PENDING,
PERMISSION_UNKNOWN,
CreatorAccount,
CreatorNoteStat,
)
# 同步一次最多翻多少页。后台默认一页 10 条,200 页足以覆盖任何正常账号,
# 同时防止"接口不返回 has_more"时无限翻下去。
MAX_SYNC_PAGES = 200
PAGE_SIZE = 10
def _account_dict(account: CreatorAccount, note_count: int = 0) -> Dict[str, Any]:
"""账号的对外表示。**刻意不含 cookie。**"""
return {
"id": account.id,
"nickname": account.nickname,
"user_id": account.user_id,
"red_id": account.red_id,
"avatar": account.avatar,
"status": account.status,
"permission_status": account.permission_status,
# 后台原话照抄。"次日可查看"这类信息只能由它自己说,改写就失真了。
"permission_tip": account.permission_tip,
"last_checked_at": account.last_checked_at,
"last_synced_at": account.last_synced_at,
"last_sync_days": account.last_sync_days,
"last_error": account.last_error,
"has_cookie": bool(account.cookie),
"created_at": account.created_at,
"note_count": note_count,
}
async def list_accounts(session: AsyncSession) -> List[Dict[str, Any]]:
accounts = list(
(await session.scalars(select(CreatorAccount).order_by(CreatorAccount.id))).all()
)
counts = dict(
(
await session.execute(
select(CreatorNoteStat.account_id, func.count(func.distinct(CreatorNoteStat.note_id)))
.group_by(CreatorNoteStat.account_id)
)
).all()
)
return [_account_dict(account, counts.get(account.id, 0)) for account in accounts]
async def get_account(session: AsyncSession, account_id: int) -> CreatorAccount:
account = await session.get(CreatorAccount, account_id)
if account is None:
raise ValueError(f"账号 {account_id} 不存在")
return account
async def account_detail(session: AsyncSession, account_id: int) -> Dict[str, Any]:
account = await get_account(session, account_id)
notes = await latest_notes(session, account_id)
return {
"account": _account_dict(account, len(notes)),
"notes": notes,
"summary": _summarize(notes),
}
def _summarize(notes: List[Dict[str, Any]]) -> Dict[str, Any]:
"""账号级汇总。取最后一轮快照的累计值之和。"""
totals = {
key: 0
for key in ("exposure", "views", "likes", "comments", "favorites", "shares", "new_followers")
}
for note in notes:
for key in totals:
value = note.get(key)
if isinstance(value, (int, float)):
totals[key] += int(value)
return totals
async def latest_notes(session: AsyncSession, account_id: int) -> List[Dict[str, Any]]:
"""每个作品取**最近一次**快照。
表里保留全部历史(换个时点就是一条新行),但列表只该展示"现在",否则同一个
作品会在列表里出现多次。
"""
newest = (
select(
CreatorNoteStat.note_id,
func.max(CreatorNoteStat.captured_at).label("captured_at"),
)
.where(CreatorNoteStat.account_id == account_id)
.group_by(CreatorNoteStat.note_id)
.subquery()
)
rows = (
await session.scalars(
select(CreatorNoteStat)
.join(
newest,
(CreatorNoteStat.note_id == newest.c.note_id)
& (CreatorNoteStat.captured_at == newest.c.captured_at),
)
.where(CreatorNoteStat.account_id == account_id)
# 不要用 nullslast():那是 PostgreSQL 语法,MySQL 5.7 会直接抛 1064 语法错误。
# MySQL 把 NULL 视为比任何值都小,所以 DESC 天然把未解析出发布时间的排在最后。
# 这个 bug 只在真机上才暴露 —— SQLite 从 3.30 起支持 NULLS LAST,测试环境测不出来。
.order_by(CreatorNoteStat.publish_time.desc())
)
).all()
return [_note_dict(row) for row in rows]
def _note_dict(row: CreatorNoteStat) -> Dict[str, Any]:
return {
"note_id": row.note_id,
"title": row.title,
"publish_time": row.publish_time,
"exposure": row.exposure,
"views": row.views,
"likes": row.likes,
"comments": row.comments,
"favorites": row.favorites,
"shares": row.shares,
"new_followers": row.new_followers,
"danmaku": row.danmaku,
"cover_ctr": row.cover_ctr,
"avg_watch_seconds": row.avg_watch_seconds,
"two_second_exit_rate": row.two_second_exit_rate,
"completion_rate": row.completion_rate,
"captured_at": row.captured_at,
}
async def upsert_account_from_cookie(session: AsyncSession, cookie: str) -> Dict[str, Any]:
"""用一份 cookie 识别并保存账号。
识别靠 `user/info` 而不是让用户填名字 —— 填错名字只会让后面所有数据对不上号。
已有同 `user_id` 的账号则更新它的 cookie(重新登录)。
"""
client = CreatorClient(cookie)
if not client.looks_authenticated:
raise ValueError("这份 cookie 里没有 a1,无法签名,请重新扫码")
try:
info = await client.fetch_user_info()
except CreatorApiError as exc:
raise ValueError(f"登录态无法使用:{exc}") from exc
if not info.get("user_id"):
raise ValueError("接口没有返回账号标识,可能登录态无效")
now = get_current_timestamp()
account = await session.scalar(
select(CreatorAccount).where(CreatorAccount.user_id == info["user_id"])
)
if account is None:
account = CreatorAccount(created_at=now)
session.add(account)
account.nickname = info.get("nickname") or account.nickname or "未命名账号"
account.user_id = info["user_id"]
account.red_id = info.get("red_id") or ""
account.avatar = info.get("avatar") or ""
account.cookie = cookie
account.status = ACCOUNT_OK
account.last_error = None
account.last_checked_at = now
account.updated_at = now
await session.flush()
# 顺手把权限状态也拉一次:新账号几乎必然处于"已申请、次日生效",
# 当场告诉用户,比让他明天再回来问要好。
await refresh_permission(session, account)
return _account_dict(account)
async def refresh_permission(session: AsyncSession, account: CreatorAccount) -> None:
"""查询并记录数据权限状态。失败不影响账号本身可用。"""
try:
permission = await CreatorClient(account.cookie).fetch_permission()
except CreatorApiError as exc:
if exc.status == 401:
account.status = ACCOUNT_EXPIRED
account.last_error = str(exc)
else:
account.last_error = str(exc)
account.updated_at = get_current_timestamp()
return
display = permission.get("display")
status = permission.get("status")
account.permission_tip = permission.get("tip") or ""
if display or status:
account.permission_status = PERMISSION_ACTIVE
elif account.permission_tip:
# 有提示语但未开通 —— 实测就是"已为您申请数据权限,次日可查看"。
account.permission_status = PERMISSION_PENDING
else:
account.permission_status = PERMISSION_MISSING
account.status = ACCOUNT_OK
account.last_error = None
account.last_checked_at = get_current_timestamp()
account.updated_at = account.last_checked_at
async def check_account(session: AsyncSession, account_id: int) -> Dict[str, Any]:
"""重新检测一个账号:登录态还在不在、权限开通没有。"""
account = await get_account(session, account_id)
await refresh_permission(session, account)
count = (
await session.execute(
select(func.count(func.distinct(CreatorNoteStat.note_id))).where(
CreatorNoteStat.account_id == account_id
)
)
).scalar() or 0
return _account_dict(account, count)
async def delete_account(session: AsyncSession, account_id: int) -> None:
account = await get_account(session, account_id)
await session.execute(
delete(CreatorNoteStat).where(CreatorNoteStat.account_id == account_id)
)
await session.delete(account)
def _day_bounds(days: int) -> tuple[int, int]:
"""最近 N 天的起止(毫秒)。与后台的按发布时间筛选对齐。"""
today = datetime.now()
end = int(datetime.combine(today.date(), time(23, 59, 59)).timestamp() * 1000)
start = int(
datetime.combine((today - timedelta(days=days)).date(), time(0, 0, 0)).timestamp() * 1000
)
return start, end
async def sync_account(
session: AsyncSession, account_id: int, days: int = 90
) -> Dict[str, Any]:
"""拉取一个账号的作品运营数据并落库。
权限未生效时接口返回的是**空壳成功**(`data.result` 里没有数据),不是错误 ——
所以"同步成功但 0 条"是正常结果,必须如实回报,不能让用户以为采集坏了。
"""
account = await get_account(session, account_id)
if not account.cookie:
raise ValueError("该账号没有可用的登录态,请重新扫码")
client = CreatorClient(account.cookie)
start_ms, end_ms = _day_bounds(days)
now = get_current_timestamp()
collected: List[Dict[str, Any]] = []
for page in range(1, MAX_SYNC_PAGES + 1):
try:
batch = await client.fetch_note_list(start_ms, end_ms, page_num=page, page_size=PAGE_SIZE)
except CreatorApiError as exc:
account.last_error = str(exc)
if exc.status == 401:
account.status = ACCOUNT_EXPIRED
account.updated_at = get_current_timestamp()
raise ValueError(f"同步失败:{exc}") from exc
collected.extend(batch)
if len(batch) < PAGE_SIZE:
break
await asyncio.sleep(0.6) # 对后台客气一点,这是自己的账号但仍是自动化访问
# 先删掉本时点可能存在的重复行,再写入 —— 表上有 (account, note, captured_at)
# 唯一索引,重复同步不该报错。
await session.execute(
delete(CreatorNoteStat).where(
CreatorNoteStat.account_id == account_id, CreatorNoteStat.captured_at == now
)
)
for note in collected:
if not note.get("note_id"):
continue
session.add(
CreatorNoteStat(
account_id=account_id,
note_id=note["note_id"],
title=note.get("title") or "",
publish_time=note.get("publish_time"),
exposure=note.get("exposure"),
views=note.get("views"),
likes=note.get("likes"),
comments=note.get("comments"),
favorites=note.get("favorites"),
shares=note.get("shares"),
new_followers=note.get("new_followers"),
danmaku=note.get("danmaku"),
cover_ctr=note.get("cover_ctr"),
avg_watch_seconds=note.get("avg_watch_seconds"),
two_second_exit_rate=note.get("two_second_exit_rate"),
completion_rate=note.get("completion_rate"),
captured_at=now,
)
)
account.last_synced_at = now
account.last_sync_days = days
account.last_error = None
account.updated_at = now
await refresh_permission(session, account)
await session.flush()
return {
"account_id": account_id,
"fetched": len(collected),
"days": days,
"permission_status": account.permission_status,
"permission_tip": account.permission_tip,
}
+107
View File
@@ -0,0 +1,107 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/signing.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""创作者后台的请求签名(XYW_ 方案)。
主站与创作者后台用的是**两套不同的签名**:主站是 VMP 的 `XYS_`,创作者后台是
`XYW_`。后者简单得多 —— MD5 → base64 → AES-128-CBC,密钥与 IV 都是硬编码常量,
纯 Python 可算,不需要浏览器。
常量与 `xhshow/config/config.py` 逐字节一致(该库也据此实现了 `sign_xyw`),
并与独立的逆向实现 xiaohongshu-cli/creator_signing.py 互相印证。
**三条实测结论**(tools/probe_creator_api.py 的 Phase 0 输出):
1. 待签字符串必须是 `url=` + 路径 + 查询串 的形式。只给路径、或去掉 `url=` 前缀,
网关一律返回 **406**;写法正确时签名通过。
2. `appId` 用 `ugc`(创作者平台的取值),不是主站的 `xhs-pc-web`。
3. 不带 cookie 时返回的是应用层的 401「无登录信息」而非 406 —— 说明签名每次都过了,
认证是独立的一层。
"""
import base64
import hashlib
import json
from datetime import datetime
XYW_AES_KEY = b"7cc4adla5ay0701v"
XYW_AES_IV = b"4uzjr7mbsibcaldp"
# 与 xhshow 的 XYW_ENV_FLAGS_DEFAULT 一致。含义未知,但改了签名就不被接受。
XYW_ENV_FLAGS = "0|0|0|1|0|0|1|0|0|0|1|0|0|0|0|1|0|0|0"
XYW_PREFIX = "XYW_"
XYW_SIGN_SVN = "56"
XYW_SIGN_TYPE = "x2"
XYW_SIGN_VERSION = "1"
# 创作者平台的 appId。用主站的 xhs-pc-web 会被拒。
CREATOR_APP_ID = "ugc"
def _aes_encrypt_hex(plaintext: str) -> str:
from Crypto.Cipher import AES
from Crypto.Util.Padding import pad
cipher = AES.new(XYW_AES_KEY, AES.MODE_CBC, XYW_AES_IV)
return cipher.encrypt(pad(plaintext.encode("utf-8"), AES.block_size)).hex()
def sign_xyw(
api: str,
a1: str,
app_id: str = CREATOR_APP_ID,
body: dict | None = None,
timestamp_ms: int | None = None,
) -> dict[str, str]:
"""为一次创作者后台请求生成 ``x-s`` / ``x-t`` 请求头。
``api`` 必须是待签的完整字符串:``url=`` 加路径,GET 请求还要带上查询串。
POST 的 JSON body 追加在其后(紧凑分隔符、不转义非 ASCII),与参考实现一致。
"""
content = api
if body is not None:
content += json.dumps(body, separators=(",", ":"), ensure_ascii=False)
if timestamp_ms is None:
timestamp_ms = int(datetime.now().timestamp() * 1000)
digest = hashlib.md5(content.encode("utf-8")).hexdigest()
plaintext = f"x1={digest};x2={XYW_ENV_FLAGS};x3={a1};x4={timestamp_ms};"
encoded = base64.b64encode(plaintext.encode("utf-8")).decode("utf-8")
envelope = {
"signSvn": XYW_SIGN_SVN,
"signType": XYW_SIGN_TYPE,
"appId": app_id,
"signVersion": XYW_SIGN_VERSION,
"payload": _aes_encrypt_hex(encoded),
}
x_s = XYW_PREFIX + base64.b64encode(
json.dumps(envelope, separators=(",", ":")).encode("utf-8")
).decode("utf-8")
return {"x-s": x_s, "x-t": str(timestamp_ms)}
def signed_api(url_path: str, query: str = "") -> str:
"""把路径与查询串拼成待签字符串。
单独抽出来是因为这个格式**没有文档**,只能靠实测固定下来 —— 写错就是 406,
而 406 的响应体 ``{"code":-1,"success":false}`` 完全看不出错在哪。
"""
return f"url={url_path}?{query}" if query else f"url={url_path}"
+159 -43
View File
@@ -17,7 +17,7 @@
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""
MediaCrawler WebUI API Server
综合采集平台 API Server
Start command: uvicorn api.main:app --port 8080 --reload
Or: python -m api.main
"""
@@ -25,28 +25,114 @@ import asyncio
import os
import sys
import subprocess
from contextlib import asynccontextmanager
from pathlib import Path
import uvicorn
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from fastapi.staticfiles import StaticFiles
from fastapi.responses import FileResponse
from .routers import crawler_router, data_router, websocket_router
# Project root directory (used for running subprocesses like uv run main.py)
PROJECT_ROOT = Path(__file__).parent.parent
# Load .env before importing anything that reads os.getenv at module import time
# (config/db_config.py does). python-dotenv was already a declared dependency but
# nothing ever called it, so the shipped .env.example had no effect.
from dotenv import load_dotenv
load_dotenv(PROJECT_ROOT / ".env")
import uvicorn
from fastapi import Depends, FastAPI
from fastapi.middleware.cors import CORSMiddleware
from fastapi.staticfiles import StaticFiles
from fastapi.responses import FileResponse
from .auth import ensure_initial_credential, require_auth
from .routers import (
auth_router,
crawler_router,
creator_router,
data_router,
monitor_router,
settings_router,
websocket_router,
)
from .services.interpreter import describe_interpreter, resolve_python_cmd
@asynccontextmanager
async def lifespan(_app: FastAPI):
"""Start the monitor scheduler with the server, and shut it down cleanly.
The scheduler is a background asyncio task, so it must not be tied to a
browser session the way the log broadcaster is -- a scheduled run has to
happen whether or not anyone has the UI open.
"""
from .creator.login import shutdown as shutdown_creator_login
from .monitor.db import dispose_engine, init_db
from .monitor.qrlogin import shutdown as shutdown_qrlogin
from .monitor.scheduler import monitor_scheduler
await init_db()
# The WebUI bundle is gitignored and built separately, so a deployment that
# forgot it would otherwise come up looking healthy and serve a bare JSON
# stub at "/" -- worth one loud line at boot rather than a puzzled operator.
if not os.path.exists(os.path.join(WEBUI_DIR, "index.html")):
print(
"[综合采集平台] 警告:未找到前端产物 api/webui/index.html,"
"根路径只会返回一段 JSON。请先在 webui/ 下执行 npm run build。",
flush=True,
)
generated = await ensure_initial_credential()
if generated:
# Printed once, on the run that creates it. There is no unauthenticated
# "set your password" endpoint on purpose: on a LAN bind that would be a
# claim-the-instance race.
rule = "=" * 68
print(
f"\n{rule}\n"
" WebUI 首次启动,已生成登录密码:\n"
f"\n {generated}\n"
"\n 请立即登录并修改。忘记密码时可设置环境变量 MC_PASSWORD 后重启。\n"
f"{rule}\n",
flush=True,
)
await monitor_scheduler.start()
try:
yield
finally:
await monitor_scheduler.stop()
# Drops the tab a QR login may have opened and stops the Playwright
# client; leaving them would strand a driver process on every restart.
await shutdown_qrlogin()
# Same for the operator's account logins, which run in throwaway browser
# contexts -- those would otherwise be left open in the operator's Chrome.
await shutdown_creator_login()
await dispose_engine()
# Docs are disabled deliberately: /docs, /redoc and /openapi.json are
# unauthenticated by default, which would hand out a complete map of the API
# (and a "Try it out" console that 401s anyway).
app = FastAPI(
title="MediaCrawler WebUI API",
description="API for controlling MediaCrawler from WebUI",
version="1.0.0"
title="综合采集平台 API",
description="API for controlling 综合采集平台 from WebUI",
version="1.0.0",
lifespan=lifespan,
docs_url=None,
redoc_url=None,
openapi_url=None,
)
# Get webui static files directory
WEBUI_DIR = os.path.join(os.path.dirname(__file__), "webui")
# CORS configuration - allow frontend dev server access
# CORS only matters for a split-origin setup. In production this app serves the
# SPA itself, and in development Vite proxies /api here (see webui/vite.config.ts),
# so the browser always sees a single origin and CORS never actually triggers.
# Kept as an explicit allowlist -- never "*", which is invalid next to
# allow_credentials -- and extensible via env for a dev server reached over LAN.
_extra_origins = [o.strip() for o in os.getenv("MC_CORS_ORIGINS", "").split(",") if o.strip()]
app.add_middleware(
CORSMiddleware,
allow_origins=[
@@ -54,15 +140,25 @@ app.add_middleware(
"http://localhost:3000", # Backup port
"http://127.0.0.1:5173",
"http://127.0.0.1:3000",
*_extra_origins,
],
allow_origin_regex=os.getenv("MC_CORS_ORIGIN_REGEX") or None,
allow_credentials=True,
allow_methods=["*"],
allow_headers=["*"],
)
# Register routers
app.include_router(crawler_router, prefix="/api")
app.include_router(data_router, prefix="/api")
# Register routers.
# The auth router stays open -- it is the way in. Everything else under /api
# requires a session. Enforcement is a Depends applied per router rather than
# app-wide middleware, because middleware needs a hand-rolled path allowlist and,
# more importantly, never sees WebSocket scopes at all.
app.include_router(auth_router, prefix="/api")
app.include_router(crawler_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(creator_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(data_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(monitor_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(settings_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(websocket_router, prefix="/api")
@@ -73,9 +169,8 @@ async def serve_frontend():
if os.path.exists(index_path):
return FileResponse(index_path)
return {
"message": "MediaCrawler WebUI API",
"message": "综合采集平台 API",
"version": "1.0.0",
"docs": "/docs",
"note": "WebUI not found, please build it first: cd webui && npm run build"
}
@@ -85,18 +180,21 @@ async def health_check():
return {"status": "ok"}
@app.get("/api/env/check")
@app.get("/api/env/check", dependencies=[Depends(require_auth)])
async def check_environment():
"""Check if MediaCrawler environment is configured correctly"""
"""Check whether the crawler environment is configured correctly"""
try:
# Run uv run main.py --help command to check environment
# Use PROJECT_ROOT so it works regardless of where uvicorn was started
# Run `main.py --help` to check the environment.
# Resolve the interpreter the same way the crawler manager does, so this
# check can never disagree with how main.py is actually executed.
# Use PROJECT_ROOT so it works regardless of where uvicorn was started.
python_cmd = resolve_python_cmd()
if sys.platform == "win32":
loop = asyncio.get_running_loop()
process = await loop.run_in_executor(
None,
lambda: subprocess.run(
["uv", "run", "main.py", "--help"],
[*python_cmd, "main.py", "--help"],
capture_output=True,
timeout=30.0,
cwd=str(PROJECT_ROOT)
@@ -105,7 +203,7 @@ async def check_environment():
stdout, stderr = process.stdout, process.stderr # bytes
else:
process = await asyncio.create_subprocess_exec(
"uv", "run", "main.py", "--help",
*python_cmd, "main.py", "--help",
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
cwd=str(PROJECT_ROOT) # Project root directory
@@ -117,7 +215,8 @@ async def check_environment():
if process.returncode == 0:
return {
"success": True,
"message": "MediaCrawler environment configured correctly",
"message": "环境配置正确",
"interpreter": describe_interpreter(),
"output": stdout.decode("utf-8", errors="ignore")[:500] # Truncate to first 500 characters
}
else:
@@ -136,8 +235,11 @@ async def check_environment():
except FileNotFoundError:
return {
"success": False,
"message": "uv command not found",
"error": "Please ensure uv is installed and configured in system PATH"
"message": "Python interpreter not found",
"error": (
"Neither uv nor a usable interpreter was found. Install uv, or create a "
"project virtualenv (.venv) with the requirements installed."
)
}
except Exception as e:
return {
@@ -147,29 +249,29 @@ async def check_environment():
}
@app.get("/api/config/platforms")
@app.get("/api/config/platforms", dependencies=[Depends(require_auth)])
async def get_platforms():
"""Get list of supported platforms"""
return {
"platforms": [
{"value": "xhs", "label": "Xiaohongshu", "icon": "book-open"},
{"value": "dy", "label": "Douyin", "icon": "music"},
{"value": "ks", "label": "Kuaishou", "icon": "video"},
{"value": "bili", "label": "Bilibili", "icon": "tv"},
{"value": "wb", "label": "Weibo", "icon": "message-circle"},
{"value": "tieba", "label": "Baidu Tieba", "icon": "messages-square"},
{"value": "zhihu", "label": "Zhihu", "icon": "help-circle"},
]
}
"""Platform capability matrix.
Returns what each platform's crawler supports (modes, metrics, comment
levels, media) *and* whether the monitoring layer has been wired up for it.
The UI renders its platform switcher and metric columns from this, so the
two are never allowed to drift apart.
"""
from .monitor.platforms import describe_all
return {"platforms": describe_all()}
@app.get("/api/config/options")
@app.get("/api/config/options", dependencies=[Depends(require_auth)])
async def get_config_options():
"""Get all configuration options"""
return {
"login_types": [
{"value": "qrcode", "label": "QR Code Login"},
{"value": "cookie", "label": "Cookie Login"},
{"value": "qrcode", "label": "扫码登录"},
# Named for what it now does: the value itself is no longer typed
# here, it is reused from Settings.
{"value": "cookie", "label": "复用已保存的 Cookie"},
],
"crawler_types": [
{"value": "search", "label": "Search Mode"},
@@ -202,4 +304,18 @@ if os.path.exists(WEBUI_DIR):
if __name__ == "__main__":
uvicorn.run(app, host="0.0.0.0", port=8080)
# Loopback by default: the safe choice for anyone who has not thought about
# exposure. Set MC_HOST=0.0.0.0 (e.g. in .env) for LAN access. Before this,
# `python -m api.main` bound 0.0.0.0 while the documented `uvicorn api.main:app`
# bound loopback -- two launch paths with different exposure.
host = os.getenv("MC_HOST", "127.0.0.1")
port = int(os.getenv("MC_PORT", "8080"))
if host not in ("127.0.0.1", "localhost", "::1"):
print(
f"[综合采集平台] 监听 {host}:{port},局域网内其他机器可访问。\n"
f"[综合采集平台] 已启用密码鉴权;如需暴露到可信网络之外,请走 HTTPS 反向代理。",
flush=True,
)
uvicorn.run(app, host=host, port=port)
+19
View File
@@ -0,0 +1,19 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/__init__.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Scheduled monitoring layer: repeated crawls with change detection."""
+249
View File
@@ -0,0 +1,249 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/adapters.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""平台适配:两个平台之间**不一样**的那些管子。
监控层的大部分是平台中立的 —— 调度、入库、差分、报表、封面缓存、通知发送都与平台无关。
真正随平台变化的只有四样东西:
1. 爬虫把产物**落在哪个目录**(这里有个坑,见 ``artifact_dir``)
2. jsonl 里**字段叫什么**(抖音的作品没有 ``note_id``,叫 ``aweme_id``)
3. **目标链接**长什么样(怎么拼、怎么从链接里抠出 id)
4. 通知里的作品链接怎么拼
集中在这里,是为了让「加一个平台」变成在一处补一份数据,而不是去五个文件里找硬编码。
**为什么不放进 platforms.py**:那个模块被 ``describe_all()`` 整个序列化进
``GET /api/config/platforms`` 交给前端(连 ``**capability`` 一起),把正则、目录名、字段别名
塞进去会让爬虫的内部细节漏进 API 载荷,也会让「改适配」有动到接口形状的风险。
分工与既有的 schedule.py(算术)↔ scheduler.py(循环)一致。
"""
import re
from dataclasses import dataclass
from typing import Any, Dict, Mapping, Optional, Pattern, Tuple
from .platforms import PLATFORM_XHS
PLATFORM_DY = "dy"
def _first_cover(record: Dict[str, Any], fields: Tuple[str, ...]) -> str:
"""封面地址:取第一个非空字段,再取逗号分隔的第一段。
一条规则同时适配两边,所以不需要 per-platform 的函数:
小红书的 ``image_list`` 是 ``"url1,url2,..."``(要切第一段),
抖音的 ``cover_url`` 本身就是单个地址(切了等于没切)。
"""
for name in fields:
raw = record.get(name)
if raw:
return str(raw).split(",")[0].strip()
return ""
@dataclass(frozen=True)
class PlatformAdapter:
"""一个平台的全部「管子」。
字段别名的方向是**规范名 -> 该平台 jsonl 里的键**,读作「我们的列 ← 他们的键」。
"""
# 爬虫落盘用的目录名。**不等于平台 id**:抖音的平台 id 是 ``dy`` 而目录是 ``douyin``。
# 这不是笔误,是上游 store 里写死的(store/douyin/_store_impl.py:47)。改错这里的
# 后果是 ingest 一个文件都找不到 —— 它不会报错,只会落进「没抓到数据」分支,
# 然后被误报成「疑似登录失效」。
artifact_dir: str
web_base: str
creator_path: str
note_path: str
# 从链接里抠 id。是元组而不是单个正则,因为同一个平台可能有多种链接形态
# (抖音的作品链接还带 ?modal_id= 那种),按顺序试,第一个匹配的胜出。
# 每个正则必须恰好有一个捕获组。
creator_url_res: Tuple[Pattern, ...]
note_url_res: Tuple[Pattern, ...]
# 也允许直接粘贴裸 id —— 但两边的 id 形状不同,所以分开。
creator_bare_re: Pattern
note_bare_re: Pattern
# 短链(v.douyin.com 这种)无法在不发请求的情况下还原出 id,解析时单独报错,
# 好过存一个聚不出目标的值进去。
short_link_hosts: Tuple[str, ...]
note_fields: Mapping[str, str]
comment_fields: Mapping[str, str]
cover_fields: Tuple[str, ...]
# 时间戳换算成毫秒要乘的数。**小红书给毫秒、抖音给秒**,差 1000 倍;不换算的话
# 2026 年的作品会显示成 1970 年(实测踩到过:抖音作品发布日期显示 1970-01-22,
# 抖音评论的时间同理)。库里统一存毫秒,展示层才不用关心来源。
time_scale: int
def to_ms(self, value: Any) -> Optional[int]:
"""把平台的时间戳换算成毫秒;解析不出来返回 None(不伪造 0)。"""
try:
return int(value) * self.time_scale
except (TypeError, ValueError):
return None
def note_url(self, note_id: str) -> str:
"""作品的可点击链接。拼法与监控目标的链接是同一个形状 —— 通知里给的就是
人能直接点开看的那一个。"""
return f"{self.web_base}{self.note_path}/{note_id}"
def note_field(self, record: Dict[str, Any], name: str) -> Any:
"""按规范名读作品记录里的原始值(没有就是 None)。"""
return record.get(self.note_fields.get(name, name))
def comment_field(self, record: Dict[str, Any], name: str) -> Any:
return record.get(self.comment_fields.get(name, name))
def cover(self, record: Dict[str, Any]) -> str:
return _first_cover(record, self.cover_fields)
def parent_comment_id(self, record: Dict[str, Any]) -> str:
"""父评论 id,顶层评论一律归一成空串。
抖音顶层评论的 ``reply_id`` 是字符串 ``"0"``,小红书是 ``""`` —— 把 "0" 原样
存进去,前端就会多出一堆指向不存在的父评论的边。
"""
raw = self.comment_field(record, "parent_comment_id")
if raw is None:
return ""
raw = str(raw).strip()
return "" if raw in ("", "0") else raw
# 小红书 id 是 24 位 hex,允许稍宽一点,让格式变化退化成「仍然接受」而不是「拒绝」。
_XHS_BARE_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
XHS = PlatformAdapter(
artifact_dir="xhs",
web_base="https://www.xiaohongshu.com",
creator_path="/user/profile",
note_path="/explore",
creator_url_res=(re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)"),
),
creator_bare_re=_XHS_BARE_RE,
note_bare_re=_XHS_BARE_RE,
short_link_hosts=(),
note_fields={
"note_id": "note_id",
"title": "title",
"note_url": "note_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "type",
"published_at": "time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "note_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("image_list",),
# 小红书的时间戳本来就是毫秒(实测 time=1790923011000)。
time_scale=1,
)
# 抖音的 id 形状与小红书完全不同(见 media_platform/douyin/help.py:101-164):
# 作品 aweme_id 纯数字,如 7525082444551310602
# 博主 sec_user_id 形如 MS4wLjABAAAA...,含 - 和 _,**变长**(实测样本 55 字符,更长的也常见),
# 而小红书那条裸 id 规则封顶 64 —— 所以两条规则必须分开,否则长一点的 sec_uid
# 会被拒,表现为「粘贴了一个完全正确的链接却说无法识别」。
# 另外抖音**不需要 xsec_token**,裸链接就能用,比小红书简单。
DY = PlatformAdapter(
artifact_dir="douyin",
web_base="https://www.douyin.com",
creator_path="/user",
note_path="/video",
creator_url_res=(re.compile(r"douyin\.com/user/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"douyin\.com/video/(\d+)"),
# 带 modal_id 的链接:在别人主页或搜索结果里点开视频就是这个形态。
re.compile(r"[?&]modal_id=(\d+)"),
),
# 用长度而不是前缀来区分两者:sec_uid 是 20 字符以上的变长串,作品 id 是 19 位数字。
# 用前缀(MS4wLjABAAAA)更精确,但上游的 parse_creator_info_from_url 对裸 id 一律
# 照单全收,万一有别的前缀就会被我这里挡掉 —— 门槛设在长度上,两边都放得进,
# 又不会把 19 位的作品号误当成博主。
creator_bare_re=re.compile(r"^[A-Za-z0-9_-]{20,128}$"),
note_bare_re=re.compile(r"^\d{8,25}$"),
short_link_hosts=("v.douyin.com",),
note_fields={
"note_id": "aweme_id",
"title": "title",
"note_url": "aweme_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "aweme_type",
"published_at": "create_time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "aweme_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("cover_url",),
# 抖音给的是**秒**(实测 create_time=1790574515,即 2026-09-28)。
time_scale=1000,
)
ADAPTERS: Dict[str, PlatformAdapter] = {
PLATFORM_XHS: XHS,
PLATFORM_DY: DY,
}
class UnknownPlatformError(ValueError):
"""平台还没有适配器。"""
def adapter(platform: str) -> PlatformAdapter:
try:
return ADAPTERS[platform]
except KeyError as exc:
raise UnknownPlatformError(f"平台 {platform} 还没有适配器") from exc
def has_adapter(platform: str) -> bool:
return platform in ADAPTERS
def artifact_dir(platform: str) -> str:
"""该平台的爬虫会把 jsonl 落在哪个子目录下。
``runner`` 用它判断产物是否真的出现过,``ingest`` 用它定位文件 —— 两处必须用
同一个值,否则会出现「文件在,但两边找的目录不是同一个」这种最难查的错。
"""
return adapter(platform).artifact_dir
+463
View File
@@ -0,0 +1,463 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/app_settings.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Application settings, declared once and rendered from that declaration.
Every setting carries a **scope**, which is the whole reason this is not a flat
list:
* ``platform`` -- each platform keeps its own copy. A cookie obviously differs,
but so do crawl pacing and proxies: what is safe on one platform is a rate
limit on another. Stored as ``platform.<p>.<name>``.
* ``system`` -- one value for the whole instance. The notification webhook is
a single group chat, and the scheduler has a single active-hours window, so
scoping those per platform would be a fiction.
The registry is the single source of truth: the API returns it and the Settings
page builds its form from it, so adding a setting does not mean editing a
matching list on the frontend.
Two rules carry over from how the cookie and webhook were already handled:
* **Secrets are never returned.** A sensitive key comes back as
``{present, length, updated_at}``, never as a value.
* **Update is partial.** Only keys present in the request are written, so a form
that does not resubmit a secret cannot silently wipe it.
"""
from dataclasses import dataclass
from typing import Any, Dict, List, Optional
from sqlalchemy.ext.asyncio import AsyncSession
from .platforms import PLATFORM_XHS
from .settings import (
delete_setting,
get_setting,
platform_key,
set_setting,
system_key,
)
from .upstream import DEFAULT_BRANCH, DEFAULT_REMOTE_URL
SCOPE_PLATFORM = "platform"
SCOPE_SYSTEM = "system"
TYPE_BOOL = "bool"
TYPE_INT = "int"
TYPE_STR = "str"
TYPE_SECRET = "secret"
# Mirrors config/base_config.py. Nothing is written until the operator changes
# something; an unset value simply means "pass no CLI flag, so the config file's
# value applies".
_DEFAULT_SLEEP_SEC = 2
@dataclass
class SettingSpec:
name: str
scope: str
type: str
label: str
help: str = ""
default: Any = None
minimum: Optional[int] = None
maximum: Optional[int] = None
choices: Optional[List[str]] = None
affects_new_runs: bool = True
def key(self, platform: str = PLATFORM_XHS) -> str:
if self.scope == SCOPE_SYSTEM:
return system_key(self.name)
return platform_key(platform, self.name)
SETTING_SPECS: List[SettingSpec] = [
# --- 平台设置 -----------------------------------------------------------
SettingSpec(
name="cookie",
scope=SCOPE_PLATFORM,
type=TYPE_SECRET,
label="登录 Cookie",
help="定时监控必须持久化登录态。建议先手动登录一次再粘贴 Cookie。",
),
SettingSpec(
name="default_interval_minutes",
scope=SCOPE_PLATFORM,
type=TYPE_INT,
label="新任务默认采集间隔(分钟)",
help="仅影响新建任务时的默认值,不会改动已有任务。",
default=360,
minimum=30,
maximum=10080,
),
SettingSpec(
name="default_max_notes",
scope=SCOPE_PLATFORM,
type=TYPE_INT,
label="默认单轮作品上限",
default=20,
minimum=1,
maximum=500,
),
SettingSpec(
name="default_max_comments",
scope=SCOPE_PLATFORM,
type=TYPE_INT,
label="默认每篇评论抓取条数",
help="接口无时间排序,只取平台默认排序的前 N 条;N 越大越容易发现新评论。",
default=50,
minimum=1,
maximum=500,
),
SettingSpec(
name="enable_sub_comments",
scope=SCOPE_PLATFORM,
type=TYPE_BOOL,
label="抓取二级评论",
help="请求量显著增加,风控风险更高。",
default=False,
),
SettingSpec(
name="crawl_sleep_sec",
scope=SCOPE_PLATFORM,
type=TYPE_INT,
label="请求间隔(秒)",
help="调大更慢但更不容易触发平台限流。各平台风控容忍度不同,故分开配置。",
default=_DEFAULT_SLEEP_SEC,
minimum=0,
maximum=600,
),
SettingSpec(
name="enable_ip_proxy",
scope=SCOPE_PLATFORM,
type=TYPE_BOOL,
label="启用 IP 代理",
default=False,
),
SettingSpec(
name="proxy_provider",
scope=SCOPE_PLATFORM,
type=TYPE_STR,
label="代理提供方",
default="kuaidaili",
choices=["kuaidaili", "wandouhttp", "static"],
),
SettingSpec(
name="proxy_pool_count",
scope=SCOPE_PLATFORM,
type=TYPE_INT,
label="代理 IP 池大小",
default=2,
minimum=1,
maximum=100,
),
SettingSpec(
name="static_proxy_url",
scope=SCOPE_PLATFORM,
type=TYPE_STR,
label="静态代理地址",
help="仅当提供方选择 static 时使用,格式 http://host:port",
default="",
),
# --- 系统设置 -----------------------------------------------------------
SettingSpec(
name="wecom_webhook",
scope=SCOPE_SYSTEM,
type=TYPE_SECRET,
label="企业微信 Webhook",
help="企业微信群机器人地址。所有平台共用同一个群,只有开了推送开关的任务才会发消息。",
),
SettingSpec(
name="active_hours_start",
scope=SCOPE_SYSTEM,
type=TYPE_INT,
label="活跃时段开始(小时)",
help="只在此时段内触发定时采集。默认 0–23 即全天;支持跨午夜,如 22–6。",
default=0,
minimum=0,
maximum=23,
affects_new_runs=False,
),
SettingSpec(
name="active_hours_end",
scope=SCOPE_SYSTEM,
type=TYPE_INT,
label="活跃时段结束(小时)",
default=23,
minimum=0,
maximum=23,
affects_new_runs=False,
),
SettingSpec(
name="cdp_enabled",
scope=SCOPE_SYSTEM,
type=TYPE_BOOL,
label="接管已有 Chrome(CDP)",
help=(
"开启后爬虫不再自己启动浏览器,而是接管本机已开放远程调试端口的 Chrome"
"(默认 127.0.0.1:9222),复用它的登录态与扩展。"
"服务器部署请开启;本机桌面使用请保持关闭。"
),
default=False,
),
# --- 上游更新检查 -------------------------------------------------------
# 这几项不作用于采集,所以都标 affects_new_runs=False:改动它们不需要等下一轮,
# 也不影响采集命令的拼装。
SettingSpec(
name="upstream_check_enabled",
scope=SCOPE_SYSTEM,
type=TYPE_BOOL,
label="检查上游仓库更新",
help=(
"定期 fetch 上游仓库,看看它有没有新提交,并在有更新时推送通知。"
"本仓库在上游之上加了一整层(见 UPSTREAM.md),不定期看一眼就会越拖越难合并。"
),
default=False,
affects_new_runs=False,
),
SettingSpec(
name="upstream_check_interval_minutes",
scope=SCOPE_SYSTEM,
type=TYPE_INT,
label="上游检查间隔(分钟)",
help="默认 1440 分钟(每天一次)。检查只是 fetch,不需要太频繁。",
default=1440,
minimum=30,
maximum=10080,
affects_new_runs=False,
),
SettingSpec(
name="upstream_remote_url",
scope=SCOPE_SYSTEM,
type=TYPE_STR,
label="上游仓库地址",
help=(
"默认是 GitHub 上的上游。国内直连 GitHub 不稳时改成 gitcode 镜像"
"(见 UPSTREAM.md),或任意能访问到上游的地址。"
),
default=DEFAULT_REMOTE_URL,
affects_new_runs=False,
),
SettingSpec(
name="upstream_branch",
scope=SCOPE_SYSTEM,
type=TYPE_STR,
label="上游分支",
default=DEFAULT_BRANCH,
affects_new_runs=False,
),
SettingSpec(
name="upstream_notify",
scope=SCOPE_SYSTEM,
type=TYPE_BOOL,
label="上游有更新时推送通知",
help="只在出现此前没推过的上游提交时发一条,同一个更新不会反复推。",
default=True,
affects_new_runs=False,
),
]
SPECS_BY_NAME = {spec.name: spec for spec in SETTING_SPECS}
# Managed by their own endpoints; never writable through the settings API.
# Suffix-matched rather than enumerated, because the cookie bookkeeping keys
# exist once per platform.
_HIDDEN_KEY_SUFFIXES = (".cookie_updated_at", ".cookie_last_ok_at")
_HIDDEN_KEYS = {"auth_password_hash", "auth_password_updated_at"}
def _is_hidden(key: str) -> bool:
return key in _HIDDEN_KEYS or key.endswith(_HIDDEN_KEY_SUFFIXES)
class SettingValidationError(ValueError):
"""Raised for a value the registry will not accept."""
def _coerce(spec: SettingSpec, raw: Any) -> Any:
if spec.type == TYPE_SECRET:
return str(raw) if raw is not None else ""
if spec.type == TYPE_BOOL:
if isinstance(raw, bool):
return raw
text = str(raw).strip().lower()
if text in ("1", "true", "yes", "y", "on"):
return True
if text in ("0", "false", "no", "n", "off", ""):
return False
raise SettingValidationError(f"{spec.label}: 需要是/否")
if spec.type == TYPE_INT:
try:
value = int(raw)
except (TypeError, ValueError):
raise SettingValidationError(f"{spec.label}: 需要整数")
if spec.minimum is not None and value < spec.minimum:
raise SettingValidationError(f"{spec.label}: 不能小于 {spec.minimum}")
if spec.maximum is not None and value > spec.maximum:
raise SettingValidationError(f"{spec.label}: 不能大于 {spec.maximum}")
return value
value = str(raw) if raw is not None else ""
if spec.choices and value not in spec.choices:
raise SettingValidationError(f"{spec.label}: 只能是 {'/'.join(spec.choices)}")
return value
def _decode(spec: SettingSpec, raw: Optional[str]) -> Any:
if raw is None:
return spec.default
if spec.type == TYPE_BOOL:
return raw.strip().lower() in ("1", "true", "yes", "y", "on")
if spec.type == TYPE_INT:
try:
return int(raw)
except ValueError:
return spec.default
return raw
def _encode(spec: SettingSpec, value: Any) -> str:
if spec.type == TYPE_BOOL:
return "true" if value else "false"
return str(value)
def _describe(spec: SettingSpec, platform: str) -> Dict[str, Any]:
return {
"key": spec.key(platform),
"name": spec.name,
"scope": spec.scope,
"type": spec.type,
"label": spec.label,
"help": spec.help,
"default": spec.default,
"minimum": spec.minimum,
"maximum": spec.maximum,
"choices": spec.choices,
"affects_new_runs": spec.affects_new_runs,
}
async def get_all(session: AsyncSession, platform: str = PLATFORM_XHS) -> Dict[str, Any]:
"""Every editable setting for one platform, plus the system-wide ones.
Secrets come back masked, never in the clear.
"""
values: Dict[str, Any] = {}
secrets: Dict[str, Any] = {}
for spec in SETTING_SPECS:
key = spec.key(platform)
raw = await get_setting(session, key)
if spec.type == TYPE_SECRET:
secrets[key] = {"present": bool(raw), "length": len(raw or "")}
else:
values[key] = _decode(spec, raw)
return {
"platform": platform,
"values": values,
"secrets": secrets,
"specs": [_describe(spec, platform) for spec in SETTING_SPECS],
}
def _spec_for_key(key: str, platform: str) -> Optional[SettingSpec]:
"""Resolve a full key back to its spec, rejecting keys for another platform."""
for spec in SETTING_SPECS:
if spec.key(platform) == key:
return spec
return None
async def update(
session: AsyncSession, payload: Dict[str, Any], platform: str = PLATFORM_XHS
) -> List[str]:
"""Apply a partial update. Returns the keys that changed.
Only keys present in ``payload`` are touched: a form that omits a secret must
not blank it. Keys belonging to a different platform are rejected rather than
silently written somewhere unexpected.
"""
changed: List[str] = []
for key, raw in payload.items():
if _is_hidden(key):
continue
spec = _spec_for_key(key, platform)
if spec is None:
raise SettingValidationError(f"未知的设置项:{key}")
# An explicit empty string clears a secret -- that is how the UI removes
# one. For everything else it is just a value.
if spec.type == TYPE_SECRET and raw == "":
await delete_setting(session, key)
changed.append(key)
continue
value = _coerce(spec, raw)
await set_setting(session, key, _encode(spec, value))
changed.append(key)
return changed
async def get_value(
session: AsyncSession,
name: str,
platform: str = PLATFORM_XHS,
fallback: Any = None,
) -> Any:
"""Read one typed setting for internal callers (the runner, the scheduler)."""
spec = SPECS_BY_NAME.get(name)
if spec is None:
return fallback
raw = await get_setting(session, spec.key(platform))
if raw is None:
return spec.default if fallback is None else fallback
return _decode(spec, raw)
async def defaults(session: AsyncSession, platform: str = PLATFORM_XHS) -> Dict[str, Any]:
"""Defaults applied when creating a task on this platform.
This is what makes the Settings page govern new tasks: the create endpoint
falls back to these for anything the caller omits.
"""
return {
"interval_minutes": int(
await get_value(session, "default_interval_minutes", platform, 360)
),
"max_notes_count": int(await get_value(session, "default_max_notes", platform, 20)),
"max_comments_count": int(
await get_value(session, "default_max_comments", platform, 50)
),
}
async def active_hours(session: AsyncSession) -> tuple[int, int]:
"""The (start, end) hour window for scheduled runs. System-wide."""
start = await get_value(session, "active_hours_start", fallback=0)
end = await get_value(session, "active_hours_end", fallback=23)
return int(start), int(end)
+162
View File
@@ -0,0 +1,162 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/covers.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""作品封面本地缓存。
**为什么必须落盘**:小红书图床的地址是**带签名、会过期**的。路径里那段时间戳就是
签发时刻,实测:
/202610080841/... (当天签发) → 200,且带不带 Referer 都 200
/202610070837/... (隔天) → 403,且带不带 Referer 都 403
所以这是**过期**,不是防盗链 —— 改 Referer 那一类修法治不了本。图一旦下载到本地,
就与签名无关,永远可读。
下载失败**不能影响采集**:一张封面拿不到,不该让整轮数据丢失。
"""
import re
from pathlib import Path
from typing import Optional
import httpx
from .db import DATA_DIR
COVERS_DIR = DATA_DIR / "covers"
# 单张封面的上限。正常封面是几十 KB;超过这个数说明拿到的不是图,
# 或者该放弃这一张而不是把内存撑爆。
MAX_COVER_BYTES = 5 * 1024 * 1024
USER_AGENT = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
)
_EXTENSIONS = {
"image/jpeg": ".jpg",
"image/jpg": ".jpg",
"image/png": ".png",
"image/webp": ".webp",
"image/gif": ".gif",
"image/heic": ".heic",
}
# note_id 是平台的稳定标识,但仍要挡住路径穿越 —— 它会直接变成文件名。
_SAFE_ID = re.compile(r"^[A-Za-z0-9_-]{1,64}$")
def is_safe_note_id(note_id: str) -> bool:
return bool(note_id) and bool(_SAFE_ID.match(note_id))
def cache_dir() -> Path:
COVERS_DIR.mkdir(parents=True, exist_ok=True)
return COVERS_DIR
def find_cached(note_id: str) -> Optional[Path]:
"""已缓存的封面文件,没有则 None。扩展名按内容类型而定,所以逐一试。"""
if not is_safe_note_id(note_id):
return None
for extension in sorted(set(_EXTENSIONS.values())):
candidate = COVERS_DIR / f"{note_id}{extension}"
if candidate.is_file():
return candidate
return None
async def cache_cover(note_id: str, url: str) -> Optional[str]:
"""下载并保存一张封面,返回文件名;失败返回 None。
**从不抛异常**:调用方是采集入库流程,一张图拿不到不该让整轮数据出问题。
"""
if not url or not is_safe_note_id(note_id):
return None
existing = find_cached(note_id)
if existing is not None:
return existing.name
try:
async with httpx.AsyncClient(timeout=20, follow_redirects=True) as client:
response = await client.get(url, headers={"user-agent": USER_AGENT})
except Exception:
return None
if response.status_code != 200:
# 403 通常意味着签名已过期 —— 这一张就没了,等下一轮采集拿到新地址。
return None
content = response.content
if not content or len(content) > MAX_COVER_BYTES:
return None
content_type = (response.headers.get("content-type") or "").split(";")[0].strip().lower()
extension = _EXTENSIONS.get(content_type, ".jpg")
# 图床偶尔不报 content-type,那种情况下扩展名只能猜,但文件本身仍然是好的。
target = cache_dir() / f"{note_id}{extension}"
try:
target.write_bytes(content)
except OSError:
return None
return target.name
def cover_url(note_id: str, remote: str) -> str:
"""前端该用哪个地址。
本地有缓存就用自己的接口 —— 那是唯一不会过期的地址。没有就退回远程地址,
至少让图先显示出来(哪怕它很快会失效)。
"""
if find_cached(note_id) is not None:
return f"/api/monitor/covers/{note_id}"
return remote
async def cache_pending(session, task_id: int, limit: int = 60) -> int:
"""把还没有本地副本的封面补下来,返回本次下载成功的张数。
由 runner 在入库之后调用,**而不是在 ingest 里** —— ingest 是刻意保持离线的
(它的文档写明 No network),往里塞网络请求会毁掉这一点。
每轮只补一批:一次跑几百张图既慢又会给图床压力,而旧地址本来就在陆续过期,
分摊到几轮里补完反而更稳。
"""
from sqlalchemy import select
from .models import MonitorNote
notes = (
await session.scalars(
select(MonitorNote)
.where(MonitorNote.task_id == task_id, MonitorNote.cover != "")
.order_by(MonitorNote.last_seen_at.desc())
.limit(limit)
)
).all()
saved = 0
for note in notes:
if find_cached(note.note_id) is not None:
continue
if await cache_cover(note.note_id, note.cover):
saved += 1
return saved
+342
View File
@@ -0,0 +1,342 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/db.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Database engine for the monitoring layer.
**MySQL** by default (see ``config/db_config.py`` and ``.env``), with SQLite kept
as an option so the test suite can run without a reachable server.
Three things here exist because of specific MySQL 5.7 behaviour:
* **utf8mb4 is forced per table.** This instance's server *and* the target schema
default to ``latin1``; relying on either would mangle or reject Chinese text.
The charset is set on every table rather than on the database, so it holds no
matter what the schema default is.
* **Connections are recycled.** The monitor runs for weeks, and MySQL drops idle
connections after ``wait_timeout`` (8h by default). Without ``pool_recycle`` and
``pool_pre_ping`` the first query after a quiet night fails with "server has
gone away".
* **The connected schema is asserted at startup.** A misconfigured database name
is caught immediately instead of silently writing to the wrong schema.
Only the configured schema is ever touched: no ``CREATE DATABASE``, no ``USE``,
no cross-schema query.
"""
import os
import sys
from contextlib import asynccontextmanager
from pathlib import Path
from typing import AsyncIterator, Optional
from sqlalchemy import Column, event, text
from sqlalchemy.dialects import mysql
from sqlalchemy.schema import CreateColumn
from sqlalchemy.ext.asyncio import (
AsyncEngine,
AsyncSession,
async_sessionmaker,
create_async_engine,
)
from .models import MonitorBase
PROJECT_ROOT = Path(__file__).parent.parent.parent
DATA_DIR = PROJECT_ROOT / "data"
DEFAULT_SQLITE_PATH = DATA_DIR / "monitor.db"
# Load .env here as well as in api/main.py: this module is imported directly by
# scripts and tests, and a configuration that only applies when the server is the
# entry point is a trap. load_dotenv does not override real environment variables.
from dotenv import load_dotenv
load_dotenv(PROJECT_ROOT / ".env")
# Kept identical to config/db_config.py's defaults so one .env drives both the
# monitor database and the crawler's own DB output.
MYSQL_HOST = lambda: os.getenv("MYSQL_DB_HOST", "localhost") # noqa: E731
MYSQL_PORT = lambda: int(os.getenv("MYSQL_DB_PORT", "3306")) # noqa: E731
MYSQL_USER = lambda: os.getenv("MYSQL_DB_USER", "root") # noqa: E731
MYSQL_PWD = lambda: os.getenv("MYSQL_DB_PWD", "") # noqa: E731
MYSQL_DB_NAME = lambda: os.getenv("MYSQL_DB_NAME", "mediacrawler") # noqa: E731
_engine: Optional[AsyncEngine] = None
_session_factory: Optional[async_sessionmaker[AsyncSession]] = None
# None means "resolve from the environment" (MySQL). Tests set a SQLite URL.
_db_url: Optional[str] = None
_expected_schema: Optional[str] = None
def resolve_db_url() -> str:
"""Build the connection URL. MySQL unless overridden."""
if _db_url is not None:
return _db_url
from urllib.parse import quote_plus
user = quote_plus(MYSQL_USER())
password = quote_plus(MYSQL_PWD())
host = MYSQL_HOST()
port = MYSQL_PORT()
name = MYSQL_DB_NAME()
return f"mysql+aiomysql://{user}:{password}@{host}:{port}/{name}?charset=utf8mb4"
def is_mysql() -> bool:
return resolve_db_url().startswith("mysql")
def set_sqlite_path(path: Path) -> None:
"""Point the layer at SQLite. Used by the test suite only."""
global _db_url, _engine, _session_factory, _expected_schema
_db_url = f"sqlite+aiosqlite:///{Path(path)}"
_engine = None
_session_factory = None
_expected_schema = None
def set_db_url(url: str, expected_schema: Optional[str] = None) -> None:
"""Point the layer at an explicit URL. ``expected_schema`` enables the guard."""
global _db_url, _engine, _session_factory, _expected_schema
_db_url = url
_engine = None
_session_factory = None
_expected_schema = expected_schema
def expected_schema() -> Optional[str]:
"""The schema the connection must be using, if the guard applies."""
if _expected_schema is not None:
return _expected_schema
return MYSQL_DB_NAME() if is_mysql() else None
def get_engine() -> AsyncEngine:
global _engine
if _engine is None:
url = resolve_db_url()
kwargs: dict = {"future": True}
if url.startswith("mysql"):
# Recycle well inside MySQL's default 8h wait_timeout, and verify a
# pooled connection before handing it out.
kwargs.update(pool_recycle=3600, pool_pre_ping=True, pool_size=5, max_overflow=5)
kwargs["connect_args"] = {"charset": "utf8mb4"}
else:
Path(url.split("///", 1)[-1]).parent.mkdir(parents=True, exist_ok=True)
_engine = create_async_engine(url, **kwargs)
if url.startswith("sqlite"):
@event.listens_for(_engine.sync_engine, "connect")
def _set_sqlite_pragmas(dbapi_connection, _connection_record): # pragma: no cover
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA journal_mode=WAL")
cursor.execute("PRAGMA foreign_keys=ON")
cursor.close()
return _engine
def get_session_factory() -> async_sessionmaker[AsyncSession]:
global _session_factory
if _session_factory is None:
_session_factory = async_sessionmaker(
bind=get_engine(),
class_=AsyncSession,
expire_on_commit=False,
)
return _session_factory
@asynccontextmanager
async def get_session() -> AsyncIterator[AsyncSession]:
"""Transactional session. Commits on success, rolls back on error."""
factory = get_session_factory()
async with factory() as session:
try:
yield session
await session.commit()
except Exception:
await session.rollback()
raise
async def _assert_correct_schema(conn) -> None:
"""Refuse to run against anything but the configured schema.
A guard, not the guarantee: the real protection is a MySQL account scoped to
this one schema (see UPSTREAM.md). This catches the ordinary mistake of a
wrong database name in configuration, before a single row is written.
"""
if not is_mysql():
return
expected = expected_schema()
if not expected:
return
current = (await conn.execute(text("SELECT DATABASE()"))).scalar()
if current is None:
raise RuntimeError(
f"数据库连接未选定 schema,期望 {expected!r}。请检查 MYSQL_DB_NAME。"
)
# lower_case_table_names=1 makes names case-insensitive server-side.
if current.lower() != expected.lower():
raise RuntimeError(
f"连接的库是 {current!r},但配置要求 {expected!r}。"
f"为避免误写其它库,已拒绝启动。"
)
print(f"[monitor.db] 已连接 MySQL schema: {current}", flush=True)
async def init_db() -> None:
"""Create missing tables, then run the small in-place migrations."""
engine = get_engine()
async with engine.begin() as conn:
await _assert_correct_schema(conn)
await conn.run_sync(MonitorBase.metadata.create_all)
await _ensure_columns(conn)
await _migrate_setting_keys(conn)
# ``create_all`` creates missing *tables* but never adds *columns* to a table that
# already exists, so those need an explicit ALTER TABLE.
#
# Which columns those are is derived from the ORM metadata, not kept by hand. The
# hand-kept version was a trap: forgetting to register a new column there still let
# the app start -- it connects fine, then fails on every query and every scheduler
# tick. Which is exactly what happened when the scheduling columns were added.
def _implicit_default(column: Column) -> Optional[str]:
"""A literal to seed existing rows with when a NOT NULL column is added."""
default = column.default
if default is not None and getattr(default, "is_scalar", False):
value = default.arg
if isinstance(value, bool):
return "1" if value else "0"
if isinstance(value, (int, float)):
return str(value)
return "'" + str(value).replace("'", "''") + "'"
# No scalar default on the model. Fall back to the type's zero value, so that
# adding the column cannot depend on the server's sql_mode.
try:
python_type = column.type.python_type
except NotImplementedError:
return None
if python_type in (bool, int, float):
return "0"
if python_type is str:
return "''"
return None
def _column_ddl(column: Column) -> str:
"""One column as MySQL DDL for ``ALTER TABLE ... ADD COLUMN``.
``CreateColumn`` renders the name, type and nullability. The default is added
separately because a model's ``default=`` is applied by the ORM and never
reaches the DDL -- and a NOT NULL column added to a populated table needs a
value for the rows already sitting there.
"""
ddl = str(CreateColumn(column).compile(dialect=mysql.dialect()))
if not column.nullable and column.server_default is None:
seed = _implicit_default(column)
if seed is not None:
ddl += f" DEFAULT {seed}"
return ddl
async def _existing_columns(conn, table: str) -> set[str]:
if is_mysql():
rows = await conn.execute(
text(
"SELECT COLUMN_NAME FROM information_schema.COLUMNS "
"WHERE TABLE_SCHEMA = DATABASE() AND TABLE_NAME = :t"
),
{"t": table},
)
return {row[0] for row in rows}
rows = await conn.execute(text(f"PRAGMA table_info({table})"))
return {row[1] for row in rows}
async def _ensure_columns(conn) -> None:
"""Add every model column the live table is missing."""
for table in MonitorBase.metadata.sorted_tables:
existing = await _existing_columns(conn, table.name)
if not existing:
# Table did not exist before this run; create_all built it complete.
continue
for column in table.columns:
# Primary keys are always present, and MySQL rejects AUTO_INCREMENT
# alongside the DEFAULT this helper appends -- so skip them rather
# than emit DDL that could never run.
if column.name in existing or column.primary_key:
continue
print(f"[monitor.db] 补齐缺失字段 {table.name}.{column.name}", flush=True)
await conn.execute(
text(f"ALTER TABLE {table.name} ADD COLUMN {_column_ddl(column)}")
)
async def _migrate_setting_keys(conn) -> None:
"""Move pre-namespacing setting keys to their scoped names.
Idempotent: the legacy row is only renamed when the new key is absent, so an
operator's later value is never overwritten.
"""
from .models import LEGACY_SETTING_KEY_RENAMES
for legacy, scoped in LEGACY_SETTING_KEY_RENAMES.items():
exists = (
await conn.execute(
text("SELECT 1 FROM monitor_setting WHERE `key` = :k"), {"k": legacy}
)
).first()
if not exists:
continue
already = (
await conn.execute(
text("SELECT 1 FROM monitor_setting WHERE `key` = :k"), {"k": scoped}
)
).first()
if already:
# Both present: the scoped one is authoritative; drop the stale row.
await conn.execute(
text("DELETE FROM monitor_setting WHERE `key` = :k"), {"k": legacy}
)
continue
await conn.execute(
text("UPDATE monitor_setting SET `key` = :new WHERE `key` = :old"),
{"new": scoped, "old": legacy},
)
async def dispose_engine() -> None:
global _engine, _session_factory
if _engine is not None:
await _engine.dispose()
_engine = None
_session_factory = None
+561
View File
@@ -0,0 +1,561 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_api.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""抖音 Web 接口客户端 —— 直接发 HTTP,不起爬虫子进程。
**为什么另起一套。** 爬虫那条路(``media_platform/douyin``)会构造一大串浏览器指纹
参数:``browser_platform=MacIntel``、``os_name=Mac OS``、``browser_version=125.0.0.0``……
而 ``User-Agent`` 是从页面现读的(在服务器上是 Linux + Chrome 155)。参数说自己是 Mac,
UA 说自己是 Linux —— 抖音网关对这种自相矛盾的请求的处理方式是:**不报错、不给原因,
回一个 200 + 空 body**。爬虫那边把它翻译成 ``Exception("account blocked")``,看起来像
账号被封,其实什么都不是。
这份客户端只发必要参数(``device_platform`` / ``aid`` 那两三个),走浏览器自己也在用的
那条调用路径。它的做法来自 mac-agent-os 项目的 ``mediacrawler_adapter.py``,实测可用。
两个关键点:
* **cookie 走 CDP 现读。** Chrome 把 cookie 值加密存在 SQLite 里,只有 CDP 拿得到
解密后的值;而且浏览器里那份比库里存的旧快照新 —— 站点会自己轮换会话。
* **产物形状照抄 store。** ``aweme_id`` / ``aweme_url`` / ``cover_url`` / ``aweme_type`` /
``create_time``(**秒**,由 adapters 换算成毫秒)…… 这样 ingest 那条链路一个字都不用改。
"""
import asyncio
import os
import time
from dataclasses import dataclass
from typing import Any, Dict, List, Optional, Tuple
from urllib.parse import urlencode
import config
import httpx
from tools import utils
from tools.user_hash import anonymize_user_id
# 请求头。**要像一个浏览器**,而且必须是**同一个浏览器**:见 BrowserIdentity。
_BASE_HEADERS = {
"Accept": "application/json, text/plain, */*",
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
"Referer": "https://www.douyin.com/",
"Origin": "https://www.douyin.com",
}
# 网关的业务前置校验头。缺了它,抖音边缘网关的 ArgusSecurityPlugin 会直接回
# 403 并写明 "Blocked by ArgusSecurityPlugin Uifid Not Found" —— 难得一次它会说原因。
# 当前网关并不校验这个头的**值**,填什么都行;一旦升级到真校验,就得改成让页面里的
# SDK 自己生成(见 media_platform/douyin/client.py 里同一条注释)。
ARGUS_HEADER_VALUE = "1"
API_ORIGIN = "https://www.douyin.com"
PROFILE_PATH = "/aweme/v1/web/user/profile/other/"
POSTS_PATH = "/aweme/v1/web/aweme/post/"
DETAIL_PATH = "/aweme/v1/web/aweme/detail/"
COMMENT_PATH = "/aweme/v1/web/comment/list/"
# 一次请求的超时。抖音这两个接口正常都在一秒内返回。
REQUEST_TIMEOUT_SECONDS = 20.0
# 问浏览器要 UA / client hints 的超时。**这个必须有。**
# ``page.evaluate`` 打在一个渲染进程已经卡住的标签页上会**永远不返回**,而问身份是采集的
# 第一步 —— 它一挂,整个 run 就永远停在「运行中」(真踩过:标签页 URL 是空的,
# cookies() 正常,evaluate 一直不回来)。
EVALUATE_TIMEOUT_SECONDS = 8.0
# 单页最多要多少条。接口自己有上限,要多了也没用。
MAX_PAGE_SIZE = 20
class DouyinApiError(RuntimeError):
"""请求失败,或登录态不可用。"""
def _cdp_url() -> str:
"""浏览器 DevTools 端点。与扫码登录那边共用同一个开关。"""
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
@dataclass
class BrowserIdentity:
"""一个请求要像浏览器所需要的全部身份信息,**且必须来自同一个浏览器**。
只拿 cookie 是不够的。UA 声称自己是 Chrome 155、却不带 Chrome 155 该有的
``sec-ch-ua``,网关一眼就能看出这不是浏览器 —— 它的回应是 **200 + 空 body**:
不报错、不给原因,只看得到「抓到 0 条」。所以这三样必须成套地从同一处取。
"""
cookie: str
user_agent: str
client_hints: Dict[str, str]
def headers(self) -> Dict[str, str]:
headers = {
"User-Agent": self.user_agent,
**self.client_hints,
**_BASE_HEADERS,
"x-tt-argus": ARGUS_HEADER_VALUE,
"Cookie": self.cookie,
}
# uifid 是设备标识,网关要它;cookie 里没有就不带(送空值反而更像异常请求)。
uifid = _cookie_value(self.cookie, "UIFID") or _cookie_value(
self.cookie, "UIFID_TEMP"
)
if uifid:
headers["uifid"] = uifid
return headers
# 身份信息的短时缓存:一次采集要发好几个请求,没必要每次都连一遍 CDP。
_IDENTITY_TTL_SECONDS = 120.0
_identity_cache: Optional[Tuple[float, BrowserIdentity]] = None
async def _safe_evaluate(page: Any, expression: str) -> Any:
"""在页面上求值,带超时;任何失败都返回 None。
**不要直接调 ``page.evaluate``** —— 在渲染进程卡住的标签页上它会永远不返回(见
``EVALUATE_TIMEOUT_SECONDS`` 那段)。
"""
try:
return await asyncio.wait_for(
page.evaluate(expression), timeout=EVALUATE_TIMEOUT_SECONDS
)
except Exception:
return None
async def _identity_from_pages(context: Any) -> Tuple[str, Dict[str, str]]:
"""问出 UA 和 client hints。
不假设第一个标签页是好的 —— 它可能停在 URL 为空、渲染进程已卡住的状态(实测过)。
所以逐个试、每个都带超时;优先抖音页面,全都不行就临时开一个干净页问完关掉。
拿不到就返回空 —— 调用方据此退回库里那份 cookie,而不是拿一组编出来的指纹去请求
(那比没有更糟,见 BrowserIdentity 的说明)。
"""
from media_platform.douyin.help import client_hint_headers
pages = list(context.pages)
pages.sort(key=lambda page: 0 if "douyin" in (page.url or "") else 1)
for page in pages:
user_agent = await _safe_evaluate(page, "() => navigator.userAgent")
if user_agent:
hints = client_hint_headers(
await _safe_evaluate(page, "() => navigator.userAgentData || null")
)
return user_agent, hints or {}
temp = None
try:
temp = await asyncio.wait_for(
context.new_page(), timeout=EVALUATE_TIMEOUT_SECONDS
)
user_agent = await _safe_evaluate(temp, "() => navigator.userAgent")
hints = client_hint_headers(
await _safe_evaluate(temp, "() => navigator.userAgentData || null")
)
return user_agent or "", hints or {}
except Exception:
return "", {}
finally:
if temp is not None:
try:
await temp.close()
except Exception:
pass
async def _read_browser() -> Optional[BrowserIdentity]:
"""连上 CDP 浏览器,一次取齐 cookie、UA、client hints。
读不到返回 None(浏览器没开/没登录),由调用方决定怎么报 —— 不抛异常。
"""
from playwright.async_api import async_playwright
from media_platform.douyin.help import client_hint_headers
playwright = None
try:
playwright = await async_playwright().start()
browser = await playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
if not browser.contexts:
return None
# contexts[0] 是真实 profile。**不要 new_context()** —— 那是无痕式的,读不到登录态。
context = browser.contexts[0]
cookies = await asyncio.wait_for(
context.cookies(), timeout=EVALUATE_TIMEOUT_SECONDS
)
# UA 和 hints 要从页面里问 —— 它们是浏览器自己的事实,写死迟早对不上。
user_agent, hints = await _identity_from_pages(context)
except Exception as exc:
utils.logger.warning(f"[douyin_api] 读浏览器身份失败:{exc}")
return None
finally:
if playwright is not None:
# 只断开连接。**绝不能 browser.close()** —— 对这个 CDP 连接而言那会关掉
# 操作者自己的浏览器。
try:
await playwright.stop()
except Exception:
pass
douyin_cookies = {
cookie["name"]: cookie["value"]
for cookie in cookies
if "douyin" in cookie.get("domain", "") or "amemv" in cookie.get("domain", "")
}
return BrowserIdentity(
cookie=_cookie_from_dict(douyin_cookies),
user_agent=user_agent or "",
client_hints=hints,
)
async def browser_identity(cookie: str = "", force: bool = False) -> BrowserIdentity:
"""拿到一份可用的身份:**优先浏览器里那份**,其次退回传进来的 cookie(库里存的)。
优先浏览器的原因:站点会自己轮换会话,库里存的是粘贴那一刻的快照,浏览器里那份才是
当前有效的;而 UA/hints 更是只有浏览器自己知道。
"""
global _identity_cache
now = time.monotonic()
if not force and _identity_cache is not None:
cached_at, cached = _identity_cache
if now - cached_at < _IDENTITY_TTL_SECONDS:
return cached
identity = await _read_browser()
if identity is None or not _has_session(identity.cookie):
# 浏览器里没有可用会话,退回调用方给的那份。UA/hints 编不出来就不编 ——
# 一组和 UA 对不上的 hints 比没有更糟。
identity = BrowserIdentity(
cookie=_cookie_header(cookie), user_agent="", client_hints={}
)
_identity_cache = (now, identity)
return identity
def forget_identity() -> None:
"""丢掉缓存的身份。cookie 变了、或测试之间要隔离时调用。"""
global _identity_cache
_identity_cache = None
def _cookie_header(cookie: str) -> str:
"""把 ``a=1; b=2`` 形式的 cookie 串规整成请求头用的形状。"""
pairs = []
for part in (cookie or "").split(";"):
if "=" in part:
name, _, value = part.partition("=")
name = name.strip()
if name:
pairs.append(f"{name}={value.strip()}")
return "; ".join(pairs)
def _cookie_from_dict(cookies: Dict[str, str]) -> str:
return "; ".join(f"{name}={value}" for name, value in cookies.items())
def _cookie_value(cookie: str, name: str) -> str:
"""从一个 cookie 串里取某个键的值。"""
for part in (cookie or "").split(";"):
key, _, value = part.partition("=")
if key.strip() == name:
return value.strip()
return ""
def _sign(params: Dict[str, Any], path: str, user_agent: str) -> Dict[str, Any]:
"""给一组参数补上 ``a_bogus`` 签名,返回新 dict。
**按需 import**:那个模块在 import 的那一瞬间就把 ``libs/douyin.js`` 交给 execjs
编译(还要读相对路径),把它拖进监控层的热路径不合适。
签名算在**不含 a_bogus 的那串 query 上**,追加到末尾 —— 和爬虫那条路一致,也是
实测能过的形态。
"""
from media_platform.douyin.help import get_a_bogus_from_js
try:
return {
**params,
"a_bogus": get_a_bogus_from_js(path, urlencode(params), user_agent),
}
except Exception as exc: # execjs 起不来 / JS 抛错,都算签名失败
raise DouyinApiError(f"算 a_bogus 签名失败:{exc}") from exc
async def _get(
path: str,
params: Dict[str, Any],
identity: BrowserIdentity,
*,
signed: bool = False,
) -> Dict[str, Any]:
"""发一个 GET,返回 JSON。
只带调用方给的参数 —— **不要往里加 webid / msToken / browser_version 那一堆**,
那正是爬虫那条路失败的原因。
``signed=True`` 时补一个 ``a_bogus``。**只有评论接口需要它**:作品、详情、博主资料
三个不带签名也照常返回,而给它们加签名是没验证过的改动,不做。
"""
if signed:
params = _sign(params, path, identity.user_agent)
url = f"{API_ORIGIN}{path}"
async with httpx.AsyncClient(timeout=REQUEST_TIMEOUT_SECONDS) as client:
response = await client.get(
url,
params=params,
headers=identity.headers(),
)
if response.status_code != 200:
raise DouyinApiError(f"HTTP {response.status_code}:{response.text[:120]}")
# 「200 + 空 body」是抖音网关拒绝请求时的典型回应(见模块说明)。必须当成错误报出来,
# 否则会一路往下变成「这个博主没作品」。
if not response.text.strip():
raise DouyinApiError(
"接口返回了空内容 —— 通常是登录态失效,或请求被网关判成了非浏览器"
)
try:
return response.json()
except ValueError as exc:
raise DouyinApiError(f"返回的不是 JSON:{response.text[:120]}") from exc
def _as_int(value: Any) -> int:
try:
return int(value)
except (TypeError, ValueError):
return 0
def normalize_aweme(aweme: Dict[str, Any]) -> Dict[str, Any]:
"""把接口返回的一条作品,翻译成 store 落盘的那套键名。
键名必须和 ``store/douyin`` 一致 —— 跨过这一层之后,ingest 就不知道数据是从爬虫
来的还是从接口来的。
"""
author = aweme.get("author") or {}
statistics = aweme.get("statistics") or {}
aweme_id = str(aweme.get("aweme_id") or "")
cover = ((aweme.get("video") or {}).get("cover") or {}).get("url_list") or [""]
uid = str(author.get("uid") or "")
nickname = author.get("nickname") or ""
return {
"aweme_id": aweme_id,
"aweme_type": str(aweme.get("aweme_type") or ""),
# store 那边 title 取的是 desc。
"title": aweme.get("desc") or "",
"desc": aweme.get("desc") or "",
# **秒**。adapters.time_scale 会把它换成毫秒,和 store 写出来的形态一致。
"create_time": _as_int(aweme.get("create_time")),
"creator_hash": anonymize_user_id(uid or author.get("sec_uid") or ""),
"nickname": nickname,
"liked_count": str(_as_int(statistics.get("digg_count"))),
"comment_count": str(_as_int(statistics.get("comment_count"))),
"collected_count": str(_as_int(statistics.get("collect_count"))),
"share_count": str(_as_int(statistics.get("share_count"))),
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
"cover_url": cover[0] if cover else "",
"source_keyword": "",
}
async def author_videos(
sec_user_id: str, count: int = MAX_PAGE_SIZE, *, cookie: str = ""
) -> List[Dict[str, Any]]:
"""某个博主最新发布的作品(按发布时间倒序),已翻译成 store 的键名。
用 ``sec_user_id`` 而不是数字 uid:监控任务里存的就是主页链接里的那段 sec_uid,
而且这个接口两种都收(爬虫那边用的也是 sec_user_id)。
"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
payload = await _get(
POSTS_PATH,
{
"sec_user_id": sec_user_id,
"count": max(1, min(count, MAX_PAGE_SIZE)),
"max_cursor": 0,
"device_platform": "webapp",
"aid": 6383,
},
identity,
)
awemes = payload.get("aweme_list") or []
if not awemes and payload.get("status_code") not in (0, None):
raise DouyinApiError(
f"接口拒绝了请求(status_code={payload.get('status_code')})"
)
return [normalize_aweme(aweme) for aweme in awemes]
async def video_detail(aweme_id: str, *, cookie: str = "") -> Dict[str, Any]:
"""单条作品的详情,已翻译成 store 的键名。
这个接口**没有**被那道真校验挡着(实测 200 / 45425 字节),所以在拿不到作品列表时,
它是「刷新已知作品指标」的唯一途径。
"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
payload = await _get(
DETAIL_PATH,
{"aweme_id": aweme_id, "device_platform": "webapp", "aid": 6383},
identity,
)
aweme = payload.get("aweme_detail") or {}
if not aweme:
raise DouyinApiError(
f"接口没返回作品(status_code={payload.get('status_code')})"
)
return normalize_aweme(aweme)
async def author_profile(sec_user_id: str, *, cookie: str = "") -> Dict[str, Any]:
"""博主主页指标:昵称 / 粉丝数 / 总获赞 / 作品数。"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
payload = await _get(
PROFILE_PATH,
{"sec_user_id": sec_user_id, "device_platform": "webapp", "aid": 6383},
identity,
)
user = payload.get("user") or {}
if not user:
raise DouyinApiError(
f"接口没返回用户数据(status_code={payload.get('status_code')})"
)
return {
# 自报家门。快照表的唯一键是 (任务, creator_hash, 轮次),而作品是靠
# `anonymize_user_id(author.uid)` 得到这个哈希的 —— 这里走同一条路,两边才对得上,
# 否则快照会和作品分成两个人,界面上永远查不到。
"creator_hash": anonymize_user_id(
str(user.get("uid") or user.get("sec_uid") or "")
),
"nickname": user.get("nickname") or "",
"unique_id": user.get("unique_id") or "",
"fans": _as_int(user.get("follower_count")),
"total_favorited": _as_int(user.get("total_favorited")),
"works": _as_int(user.get("aweme_count")),
"following": _as_int(user.get("following_count")),
}
def normalize_comment(comment: Dict[str, Any], aweme_id: str) -> Dict[str, Any]:
"""把接口返回的一条评论,翻译成 store 落盘的那套键名。
HTTP 路线和页面路线共用它 —— 同一套键名,ingest 才不用关心数据是怎么来的。
刻意**不带** ``sub_comment_count`` / ``parent_comment_id`` 的猜测值:接口给了就用,
没给就留空,不编。
"""
user = comment.get("user") or {}
return {
"comment_id": str(comment.get("cid") or ""),
"aweme_id": aweme_id,
"content": comment.get("text") or "",
"nickname": user.get("nickname") or "",
"creator_hash": anonymize_user_id(
str(user.get("uid") or user.get("sec_uid") or "")
),
# 同为秒;adapters 会换算。
"create_time": _as_int(comment.get("create_time")),
"like_count": str(_as_int(comment.get("digg_count"))),
"sub_comment_count": str(_as_int(comment.get("reply_comment_total"))),
# 顶层评论在抖音里是 "0";adapters.parent_comment_id 会归一成空串。
"parent_comment_id": str(comment.get("reply_id") or "0"),
}
async def video_comments(
aweme_id: str, count: int = 20, *, cookie: str = ""
) -> List[Dict[str, Any]]:
"""一条作品的评论,翻译成 store 的评论键名。
刻意不带 ``sub_comment_count`` / ``parent_comment_id`` 的猜测值 —— 接口给了就用,
没给就留空,不编。
"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用")
payload = await _get(
COMMENT_PATH,
{
"aweme_id": aweme_id,
"count": max(1, min(count, MAX_PAGE_SIZE)),
"cursor": 0,
"device_platform": "webapp",
"aid": 6383,
},
identity,
# 这个接口**必须**签名。不签的话网关回 200 + 空 body,会被读成「这条没评论」,
# 而它其实只是被挡了 —— 和登录失效长得一模一样。(实测:带上签名 200/9960 字节
# 真评论,不带就是空的。)
signed=True,
)
records = [
normalize_comment(comment, aweme_id)
for comment in payload.get("comments") or []
]
return records
def _has_session(cookie: str) -> bool:
return "sessionid=" in (cookie or "")
async def check_login(cookie: str = "") -> Dict[str, Any]:
"""浏览器/库里现在有没有可用的抖音登录态。给设置页用。"""
identity = await browser_identity(cookie)
if _has_session(identity.cookie):
source = "browser" if identity.user_agent else "stored"
return {"ok": True, "source": source, "cookie_length": len(identity.cookie)}
return {"ok": False, "source": "", "cookie_length": 0}
async def main() -> None: # pragma: no cover - 手工排查用
"""``python -m api.monitor.douyin_api <sec_user_id>``"""
import sys
if len(sys.argv) < 2:
print(await check_login())
return
sec = sys.argv[1]
print(await author_profile(sec))
for record in await author_videos(sec, count=5):
print(record["create_time"], record["title"][:30], record["liked_count"])
if __name__ == "__main__": # pragma: no cover
asyncio.run(main())
+217
View File
@@ -0,0 +1,217 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_fetch.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""抖音采集:直接走 Web 接口,不起爬虫子进程。
与 ``media_platform/douyin`` 那条路的分工:
* 那边起一个 Playwright 子进程、构造一大串**自相矛盾的浏览器指纹参数**(参数说自己是
Mac + Chrome 125,UA 说自己是 Linux + Chrome 155),网关回一个 200 + 空 body,
然后被翻译成 ``Exception("account blocked")`` —— 看起来像账号被封,其实什么都不是。
* 这边不起子进程、只发必要参数,用浏览器里那份登录态。见 ``douyin_api``。
采到的东西写成 ``store/douyin`` 那套 jsonl 形状,所以 **ingest 完全不知道数据是怎么来的**
—— 重采样、差分、事件、通知、报表全都照旧。
"""
import json
from datetime import datetime
from pathlib import Path
from typing import Any, Dict, Iterable, List, Optional, Sequence
from tools import utils
from . import adapters, douyin_api
from .models import MODE_CREATOR, MODE_NOTE, MonitorTask
# 拉多少条作品。接口单页上限就是 20,要多了也没用。
DEFAULT_VIDEO_LIMIT = 20
async def collect(
out_dir: Path,
*,
platform: str,
mode: str,
limit: int,
want_comments: bool,
comment_limit: int,
targets: Sequence[Any],
known_aweme_ids: Iterable[str] = (),
cookie: str = "",
) -> Dict[str, Any]:
"""跑一轮抖音采集,把产物写进 ``out_dir`` 下 store 那个目录里。
参数是散的、不收 ORM 对象:调用方那边 ``task`` 在会话关掉之后就 detached 了,
传对象进来迟早会踩到「属性已过期」。
**不抛异常**:失败也把(可能为空的)产物落下去,并把原因放进 ``errors`` 交给调用方。
让调用方去决定这是「一轮正常但没数据」还是「一轮失败」—— 这个判断不该藏在这里。
"""
notes: List[Dict[str, Any]] = []
comments: List[Dict[str, Any]] = []
# 博主**账号级**指标(粉丝 / 总获赞 / 作品数)。作品列表之外单独要一次,
# 只有博主模式才有 —— 作品模式的目标是一件作品,没有"这个博主是谁"可问。
profiles: List[Dict[str, Any]] = []
errors: List[str] = []
# **整个 collect 只去重一次的、跨目标的集合**:退化路径会把「库里已知的全部作品」
# 在每个目标下都刷一遍,多个目标就会出现同一件作品好几条记录 —— 而一对一快照的
# 唯一键是 (task_id, note_id, run_id),同一条作品在一轮里出现两次会直接撞键。
seen_aweme: set = set()
for target in targets:
external_id = target.external_id
# 两种模式的目标是不同的东西,不能走同一条路:
# 作品模式 —— 目标本身就是作品 id,直接取详情(**这个接口没被挡,今天就能用**)。
# 博主模式 —— 目标是主页 sec_uid,要先拉作品列表;那个接口被真校验挡着,退化到
# 刷新库里已知的作品(新作品发现不了)。
if mode == MODE_NOTE:
try:
videos = [await douyin_api.video_detail(external_id, cookie=cookie)]
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取作品 {external_id} 失败:{exc}")
videos = []
else:
videos = await _creator_works(
external_id, limit, known_aweme_ids, seen_aweme, cookie, errors
)
profile = await _creator_profile(external_id, videos, cookie, errors)
if profile is not None:
profiles.append(profile)
for video in videos:
aweme_id = video.get("aweme_id")
if not aweme_id or aweme_id in seen_aweme:
continue
seen_aweme.add(aweme_id)
notes.append(video)
if want_comments:
try:
comments.extend(
await douyin_api.video_comments(
aweme_id, count=comment_limit, cookie=cookie
)
)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取作品 {aweme_id} 的评论失败:{exc}")
jsonl_dir = _write_artifacts(out_dir, platform, mode, notes, comments, profiles)
return {
"notes": len(notes),
"comments": len(comments),
"errors": errors,
"jsonl_dir": str(jsonl_dir),
}
async def _creator_profile(
sec_user_id: str,
videos: Sequence[Dict[str, Any]],
cookie: str,
errors: List[str],
) -> Optional[Dict[str, Any]]:
"""问一次博主的账号级指标。拿不到就算了 —— **不能因为顺手的附加信息失败,
就把这一轮本来采到的作品也判成失败。**
creator_hash 优先取作品自带的那个:作品是靠 ``anonymize_user_id(author.uid)`` 得到
哈希的,而快照表和作品必须对得上号,否则界面上永远查不出这个博主的粉丝数。只有当一件
作品都没采到时(列表被挡且没有已知作品可刷新),才退回资料接口自己算的哈希 ——
那种情况下也只剩它了。
"""
try:
profile = await douyin_api.author_profile(sec_user_id, cookie=cookie)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取博主 {sec_user_id} 的资料失败:{exc}")
return None
if videos:
profile["creator_hash"] = videos[0].get("creator_hash") or profile["creator_hash"]
if not profile.get("creator_hash"):
# 哈希都算不出来的快照没人能查到,落下去只是垃圾。
errors.append(f"博主 {sec_user_id} 的资料里没有可用的身份标识,跳过账号指标")
return None
return profile
async def _creator_works(
sec_user_id: str,
limit: int,
known_aweme_ids: Iterable[str],
seen_aweme: set,
cookie: str,
errors: List[str],
) -> List[Dict[str, Any]]:
"""一个博主的作品:先要列表,列表被挡时退化成刷新已知作品。
作品列表(``aweme/post``)被抖音单独加了真校验 —— 不带 ``x-tt-argus`` 回 403,
带上 dummy 值回 200 + 空 body。所以这里拿不到**新**作品,只能保住已知的。
"""
try:
return await douyin_api.author_videos(sec_user_id, count=limit, cookie=cookie)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取博主 {sec_user_id} 的作品列表失败:{exc}")
refreshed: List[Dict[str, Any]] = []
for aweme_id in known_aweme_ids:
if aweme_id in seen_aweme:
continue
try:
refreshed.append(await douyin_api.video_detail(aweme_id, cookie=cookie))
except douyin_api.DouyinApiError as detail_exc:
errors.append(f"刷新作品 {aweme_id} 失败:{detail_exc}")
return refreshed
def _write_artifacts(
out_dir: Path,
platform: str,
mode: str,
notes: List[Dict[str, Any]],
comments: List[Dict[str, Any]],
profiles: Sequence[Dict[str, Any]] = (),
) -> Path:
"""按爬虫那套目录与文件名写 jsonl。
目录名取 ``adapters.artifact_dir``(抖音是 ``douyin``,而不是平台 id ``dy``)——
和 ingest 找文件用的是同一个来源,两边不会走散。
"""
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
jsonl_dir.mkdir(parents=True, exist_ok=True)
kind = "creator" if mode == MODE_CREATOR else "detail"
date = datetime.now().strftime("%Y-%m-%d")
_write_jsonl(jsonl_dir / f"{kind}_contents_{date}.jsonl", notes)
# 评论文件即使没有评论也建出来:ingest 靠「文件在不在」区分「这一轮没评论」和
# 「这一轮什么都没抓到」,两种情况的含义完全不同。
_write_jsonl(jsonl_dir / f"{kind}_comments_{date}.jsonl", comments)
# 博主资料同样无条件写:空文件表示"问了但没问到",没有文件表示"这次根本没问"
# (作品模式)。两者在 ingest 那边走的是同一条路(都不落快照),但留空文件能让
# 事后翻 run 目录时看出到底问没问过。
_write_jsonl(jsonl_dir / f"{kind}_profile_{date}.jsonl", list(profiles))
return jsonl_dir
def _write_jsonl(path: Path, records: List[Dict[str, Any]]) -> None:
with path.open("w", encoding="utf-8") as handle:
for record in records:
handle.write(json.dumps(record, ensure_ascii=False) + "\n")
utils.logger.info(f"[douyin_fetch] 写出 {len(records)} 条 -> {path.name}")
+800
View File
@@ -0,0 +1,800 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/ingest.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Turn one run's crawled jsonl into snapshots and change events.
Pure-ish and offline testable: give it a directory of jsonl files, a run row and
a session, and it does the diffing. No network, no browser.
Correctness notes that drive the code below:
* Counts arrive as strings and may be abbreviated ("1.2万", "3亿"). A value that
cannot be parsed is stored as NULL, never 0 -- 0 would forge a large negative
delta on the next comparison.
* The comment endpoint has no time-sort, so only the platform's top-N window is
ever visible. A comment we have not seen before is therefore split into
"posted since last run" vs "seen for the first time", rather than claiming the
former always.
* A bad cookie does not make the crawler exit non-zero; it exits 0 having
fetched nothing. That is detected here as a suspected auth failure.
"""
import json
import re
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Dict, List, Optional, Sequence
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .platforms import PLATFORM_XHS
from .models import (
EVENT_AUTH_FAILURE,
EVENT_METRIC_DELTA,
EVENT_NEW_COMMENT_POSTED,
EVENT_NEW_COMMENT_SEEN,
EVENT_NEW_NOTE,
EVENT_NO_DATA,
EVENT_RUN_FAILED,
MonitorComment,
MonitorCreatorStat,
MonitorEvent,
MonitorNote,
MonitorNoteMetric,
MonitorRun,
MonitorTask,
RUN_FAILED,
RUN_PARTIAL,
RUN_SUCCESS,
)
_COUNT_UNITS = {
"": 1,
"万": 10_000,
"w": 10_000,
"W": 10_000,
"k": 1_000,
"K": 1_000,
"亿": 100_000_000,
}
_COUNT_RE = re.compile(r"^([\d.]+)\s*([万wWkK亿]?)$")
# Metric fields shared by the snapshot table and the delta comparison.
_METRIC_FIELDS = ("liked_count", "comment_count", "collected_count", "share_count")
def parse_count(value: Any) -> Optional[int]:
"""Parse an XHS interaction count into an int, or None if unintelligible.
Handles plain numbers, thousands separators, and the Chinese abbreviations
the platform actually returns ("1.2万" -> 12000, "3亿" -> 300000000).
"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, int):
return value
if isinstance(value, float):
return int(value)
text = str(value).strip().replace(",", "").replace(" ", "")
if not text:
return None
match = _COUNT_RE.match(text)
if not match:
return None
try:
number = float(match.group(1))
except ValueError:
return None
return int(number * _COUNT_UNITS.get(match.group(2), 1))
# Windows reports hard process failures as NTSTATUS values, which surface in the
# UI as meaningless large integers (e.g. 3221225794 = 0xC0000142). Translating
# the ones we actually see saves the reader a hex-decoding detour.
_WINDOWS_EXIT_REASONS = {
0xC0000005: "进程访问冲突 (ACCESS_VIOLATION)",
0xC00000FD: "栈溢出 (STACK_OVERFLOW)",
0xC000013A: "进程被中断(控制台关闭或 Ctrl+C)",
0xC0000142: "进程初始化失败 (STATUS_DLL_INIT_FAILED),属启动环境异常,重启服务后重试",
0xC0000409: "栈缓冲区溢出 (STACK_BUFFER_OVERRUN)",
}
def describe_exit_code(code: int, cause: Optional[str] = None) -> str:
"""Render an exit code so a human can act on it.
只有退出码时信息量约等于零(`code 1` 什么都能是),所以把从输出尾巴里认出来的
异常一并附上 —— 运行历史里那一格显示的正是这句话。
"""
unsigned = code & 0xFFFFFFFF if code < 0 else code
reason = _WINDOWS_EXIT_REASONS.get(unsigned)
base = f"Crawler exited with code {code}"
if reason:
base = f"{base} (0x{unsigned:08X}): {reason}"
return f"{base};原因:{cause}" if cause else base
# 从爬虫输出里认出一行「异常」。Python 的 traceback 末行形如
# ``media_platform.douyin.exception.DataFetchError: account blocked``。
_EXCEPTION_LINE_RE = re.compile(r"^[\w.]*[A-Za-z](?:Error|Exception|Timeout)\b")
def diagnose_failure(output_tail: Optional[Sequence[str]]) -> Optional[str]:
"""从爬虫输出的末尾挑出最能说明问题的一行。
「退出码 1」等于什么都没说:真正的报错埋在子进程的 stderr 里。倒着找第一行看起来
像异常的行(traceback 的末行),找不到就退回最后一行有效输出。
"""
if not output_tail:
return None
lines = [line.strip() for line in output_tail if line and line.strip()]
# 管理器自己补的那两句不是爬虫的报错,别被当成失败原因。
noise = ("Crawler exited with code", "Crawler completed successfully")
lines = [line for line in lines if not line.startswith(noise)]
if not lines:
return None
for line in reversed(lines):
if _EXCEPTION_LINE_RE.match(line):
return line[:300]
return lines[-1][:300]
@dataclass
class IngestResult:
status: str
notes_fetched: int = 0
comments_fetched: int = 0
new_notes: int = 0
new_comments: int = 0
is_baseline: bool = False
error: Optional[str] = None
events: List[str] = field(default_factory=list)
def _read_jsonl(path: Path) -> List[Dict[str, Any]]:
"""Read a jsonl file, skipping blank or malformed lines."""
records: List[Dict[str, Any]] = []
if not path.exists():
return records
with path.open("r", encoding="utf-8") as handle:
for line in handle:
line = line.strip()
if not line:
continue
try:
item = json.loads(line)
except json.JSONDecodeError:
continue
if isinstance(item, dict):
records.append(item)
return records
def find_run_files(
out_dir: Path, platform: str = PLATFORM_XHS
) -> tuple[List[Path], List[Path]]:
"""Locate the contents/comments jsonl files a run produced.
The crawler writes ``{save_data_path}/{platform}/jsonl/{type}_{item}_{date}.jsonl``.
Glob rather than reconstructing the name: both the crawler type and the date
are runtime-dependent. Returns lists because a crawl crossing midnight
produces one file per day.
``platform`` 是**监控层的平台 id**,而爬虫落盘的目录名未必同名(抖音的 id 是
``dy``、目录是 ``douyin``),所以这里经 adapters 解析 —— 调用方不必知道这个差异。
"""
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
if not jsonl_dir.is_dir():
return [], []
return (
sorted(jsonl_dir.glob("*_contents_*.jsonl")),
sorted(jsonl_dir.glob("*_comments_*.jsonl")),
)
def find_profile_files(out_dir: Path, platform: str = PLATFORM_XHS) -> List[Path]:
"""博主**账号级**指标那几行 jsonl(``creator_profile_*.jsonl``)。
单独一个函数而不是往 ``find_run_files`` 的返回值里塞第三个列表:那个返回值的两个
位置是有意义的(contents/comments),加一个会把所有调用点和解包语句都牵动一遍,
而这份产物是**可选**的 —— 小红书那条路(爬虫进程)根本不产生它。
"""
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
if not jsonl_dir.is_dir():
return []
return sorted(jsonl_dir.glob("*_profile_*.jsonl"))
def _misplaced_output_dirs(out_dir: Path, expected: str) -> List[str]:
"""在 out_dir 下找「有产物、但目录名不是期望的那个」的目录。
这是专门为**最难查的那类故障**准备的:产物目录名与平台对不上时,ingest 一个文件
都找不到,现象和「登录态失效」一模一样 —— 而实际上登录好好的、数据也抓到了,
只是没人去对的地方读。上游哪天改了 store 的目录名,这里能直接把实情说出来。
"""
found = []
try:
children = list(out_dir.iterdir())
except OSError:
return found
for child in children:
if child.name == expected or not child.is_dir():
continue
try:
if any(child.glob("jsonl/*_contents_*.jsonl")):
found.append(child.name)
except OSError:
continue
return sorted(found)
async def _emit(
session: AsyncSession,
run: MonitorRun,
event_type: str,
title: str,
*,
severity: str = "info",
target_kind: str = "",
target_id: str = "",
payload: Optional[Dict[str, Any]] = None,
) -> None:
session.add(
MonitorEvent(
task_id=run.task_id,
run_id=run.id,
type=event_type,
severity=severity,
target_kind=target_kind,
target_id=target_id,
title=title,
payload_json=json.dumps(payload or {}, ensure_ascii=False),
created_at=get_current_timestamp(),
)
)
async def _previous_run_started_at(
session: AsyncSession, task_id: int, run_id: int
) -> Optional[int]:
"""Started-at of the most recent earlier successful run, in ms."""
return await session.scalar(
select(MonitorRun.started_at)
.where(
MonitorRun.task_id == task_id,
MonitorRun.id != run_id,
MonitorRun.status.in_((RUN_SUCCESS, RUN_PARTIAL)),
MonitorRun.started_at.is_not(None),
# Same reasoning as _count_prior_successes: an empty run is a useless
# reference point for "was this comment posted since last time?".
MonitorRun.notes_fetched > 0,
)
.order_by(MonitorRun.id.desc())
.limit(1)
)
# How far back to look for proof that the stored login still works.
_AUTH_PROOF_WINDOW_MS = 6 * 60 * 60 * 1000
async def _another_task_succeeded_recently(session: AsyncSession, task_id: int) -> bool:
"""Whether a different task fetched data recently, proving the login is valid."""
since = get_current_timestamp() - _AUTH_PROOF_WINDOW_MS
count = await session.scalar(
select(func.count())
.select_from(MonitorRun)
.where(
MonitorRun.task_id != task_id,
MonitorRun.status == RUN_SUCCESS,
MonitorRun.started_at.is_not(None),
MonitorRun.started_at >= since,
)
)
return bool(count)
async def _count_prior_successes(session: AsyncSession, task_id: int, run_id: int) -> int:
return (
await session.scalar(
select(func.count())
.select_from(MonitorRun)
.where(
MonitorRun.task_id == task_id,
MonitorRun.id != run_id,
MonitorRun.status.in_((RUN_SUCCESS, RUN_PARTIAL)),
# A run that fetched nothing established no baseline. Without this
# check the first run that actually works after a failed one looks
# like a flood of newly discovered works.
MonitorRun.notes_fetched > 0,
)
)
) or 0
async def _ingest_notes(
session: AsyncSession,
run: MonitorRun,
records: List[Dict[str, Any]],
is_baseline: bool,
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert notes, write metric snapshots, and emit new-note/delta events.
记录里的字段一律经 ``adapter`` 读。抖音的作品没有 ``note_id``(叫 ``aweme_id``),
按名字硬取的话每条记录都会在下面第一行被 continue 掉 —— 一条都不报错地全丢。
"""
now = get_current_timestamp()
new_count = 0
# 同一轮里重复出现的作品只处理一次。**指标快照的唯一键是 (task_id, note_id, run_id)**,
# 同一件作品在一轮里进来两次会让第二次插入直接撞键、整个 run 崩掉 —— 产物里重复并不
# 罕见(多个目标指向同一个人、或退化路径重复刷新)。
seen_in_run: set = set()
for record in records:
note_id = adapter.note_field(record, "note_id")
if not note_id:
continue
if note_id in seen_in_run:
continue
seen_in_run.add(note_id)
note = await session.scalar(
select(MonitorNote).where(
MonitorNote.task_id == run.task_id,
MonitorNote.note_id == note_id,
)
)
title = (adapter.note_field(record, "title") or "")[:500]
cover = adapter.cover(record)
if note is None:
note = MonitorNote(
task_id=run.task_id,
note_id=note_id,
title=title,
note_url=adapter.note_field(record, "note_url") or "",
cover=cover,
creator_hash=adapter.note_field(record, "creator_hash") or "",
creator_name=adapter.note_field(record, "creator_name") or "",
source_kind=adapter.note_field(record, "source_kind") or "",
# 经 to_ms 换算:小红书给毫秒、抖音给秒,差 1000 倍。
published_at=adapter.to_ms(adapter.note_field(record, "published_at")),
first_seen_run_id=run.id,
first_seen_at=now,
last_seen_run_id=run.id,
last_seen_at=now,
)
session.add(note)
new_count += 1
if not is_baseline:
await _emit(
session,
run,
EVENT_NEW_NOTE,
f"新作品:{title or note_id}",
target_kind="note",
target_id=note_id,
payload={"note_id": note_id, "title": title},
)
else:
# Only refresh descriptive fields; seen-tracking is updated below.
if title:
note.title = title
# 封面地址**带签名、会过期**,所以每轮都用最新的覆盖它。原先只在首次入库
# 时写一次,结果旧作品的封面地址烂在库里 —— 隔天开始全是 403,而且再怎么
# 重跑也修不回来。落盘那份由 covers.cache_pending 负责(网络操作不在本模块)。
if cover:
note.cover = cover
# 昵称也要刷:作者改昵称是常事,只在首次入库写一次会一直显示旧的。
creator_name = adapter.note_field(record, "creator_name")
if creator_name:
note.creator_name = creator_name
# 发布时间也刷。正常情况下它不会变,但**换算单位改过之后**(抖音是秒、
# 小红书是毫秒),已经入库的那批只能靠重采修回来。
published = adapter.to_ms(adapter.note_field(record, "published_at"))
if published is not None:
note.published_at = published
note.last_seen_run_id = run.id
note.last_seen_at = now
await _snapshot_metrics(session, run, note_id, record, now, is_baseline)
return new_count
def _as_int(value: Any) -> Optional[int]:
try:
return int(value)
except (TypeError, ValueError):
return None
async def _snapshot_metrics(
session: AsyncSession,
run: MonitorRun,
note_id: str,
record: Dict[str, Any],
now: int,
is_baseline: bool,
) -> None:
"""Write this run's metric snapshot and report any change vs the previous one."""
previous = await session.scalar(
select(MonitorNoteMetric)
.where(
MonitorNoteMetric.task_id == run.task_id,
MonitorNoteMetric.note_id == note_id,
MonitorNoteMetric.run_id != run.id,
)
.order_by(MonitorNoteMetric.run_id.desc())
.limit(1)
)
parsed = {name: parse_count(record.get(name)) for name in _METRIC_FIELDS}
session.add(
MonitorNoteMetric(
task_id=run.task_id,
note_id=note_id,
run_id=run.id,
captured_at=now,
liked_count=parsed["liked_count"],
comment_count=parsed["comment_count"],
collected_count=parsed["collected_count"],
share_count=parsed["share_count"],
raw_liked_count=str(record.get("liked_count") or ""),
raw_comment_count=str(record.get("comment_count") or ""),
raw_collected_count=str(record.get("collected_count") or ""),
raw_share_count=str(record.get("share_count") or ""),
)
)
if previous is None or is_baseline:
return
deltas = {}
for name in _METRIC_FIELDS:
old, new = getattr(previous, name), parsed[name]
# A None on either side means the value was unparseable; skip rather
# than report a bogus change.
if old is None or new is None or old == new:
continue
deltas[name] = {"from": old, "to": new, "delta": new - old}
if deltas:
summary = "、".join(
f"{_metric_label(name)} {info['from']}→{info['to']}"
for name, info in deltas.items()
)
await _emit(
session,
run,
EVENT_METRIC_DELTA,
f"互动数据变化:{summary}",
target_kind="note",
target_id=note_id,
payload={"note_id": note_id, "deltas": deltas},
)
def _metric_label(name: str) -> str:
return {
"liked_count": "点赞",
"comment_count": "评论",
"collected_count": "收藏",
"share_count": "分享",
}.get(name, name)
async def _ingest_comments(
session: AsyncSession,
run: MonitorRun,
records: List[Dict[str, Any]],
is_baseline: bool,
previous_run_started_at: Optional[int],
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert comments and emit events for ones never seen before.
与作品同理,评论记录也要经 ``adapter`` 读:抖音的评论用 ``aweme_id`` 指作品。
"""
now = get_current_timestamp()
new_count = 0
for record in records:
comment_id = adapter.comment_field(record, "comment_id")
note_id = adapter.comment_field(record, "note_id")
if not comment_id or not note_id:
continue
existing = await session.scalar(
select(MonitorComment).where(
MonitorComment.task_id == run.task_id,
MonitorComment.note_id == note_id,
MonitorComment.comment_id == comment_id,
)
)
if existing is not None:
# 昵称要跟着刷,不能只写一次。评论是去重后直接 continue 的,若不刷新,
# 脱敏开关一改(或评论者改了昵称),已经入库的老评论会永远停在旧值上 ——
# 而重采是唯一能拿到新值的途径。作品那边的 creator_name 同理。
refreshed = adapter.comment_field(record, "creator_name")
if refreshed:
existing.nickname = refreshed
# 时间同理:单位换算修好之后,老数据要重采才能纠正。
created = adapter.to_ms(adapter.comment_field(record, "create_time"))
if created is not None:
existing.create_time = created
continue
create_time = adapter.to_ms(adapter.comment_field(record, "create_time"))
session.add(
MonitorComment(
task_id=run.task_id,
note_id=note_id,
comment_id=comment_id,
content=(adapter.comment_field(record, "content") or "")[:2000],
nickname=adapter.comment_field(record, "creator_name") or "",
creator_hash=adapter.comment_field(record, "creator_hash") or "",
create_time=create_time,
like_count=parse_count(adapter.comment_field(record, "like_count")),
sub_comment_count=_as_int(
adapter.comment_field(record, "sub_comment_count")
)
or 0,
parent_comment_id=adapter.parent_comment_id(record),
first_seen_run_id=run.id,
first_seen_at=now,
)
)
new_count += 1
if is_baseline:
continue
# Without a time-sorted comment API we can only observe the top-N window,
# so distinguish a genuinely new comment from one that just surfaced.
posted = (
create_time is not None
and previous_run_started_at is not None
and create_time > previous_run_started_at
)
await _emit(
session,
run,
EVENT_NEW_COMMENT_POSTED if posted else EVENT_NEW_COMMENT_SEEN,
f"{'新评论' if posted else '新出现评论'}:{(record.get('content') or '')[:60]}",
target_kind="note",
target_id=note_id,
payload={
"note_id": note_id,
"comment_id": comment_id,
"create_time": create_time,
"nickname": record.get("nickname") or "",
},
)
return new_count
async def _ingest_creator_stats(
session: AsyncSession,
run: MonitorRun,
records: Sequence[Dict[str, Any]],
) -> int:
"""把这一轮问到的博主账号级指标落成快照,返回条数。
和作品指标一样是**每轮一条**:账号级的粉丝数是缓慢变化的量,「今天比昨天多了 300」
才是有用的信号,单看一个绝对值没有意义 —— 所以这里只管记,分析交给查询端。
**没解析出来的值留 NULL,不写 0**:0 在趋势图上是一条砸到底的线,和「不知道」完全是
两回事(见 ``parse_count`` 的注释)。
"""
now = get_current_timestamp()
written = 0
seen: set = set()
for record in records:
creator_hash = str(record.get("creator_hash") or "").strip()
if not creator_hash or creator_hash in seen:
# 一个任务可以配多个目标,退化路径下它们可能指向同一个博主 —— 而唯一键是
# (task_id, creator_hash, run_id),重复插入会撞键把整轮炸掉。
continue
seen.add(creator_hash)
session.add(
MonitorCreatorStat(
task_id=run.task_id,
run_id=run.id,
creator_hash=creator_hash,
nickname=str(record.get("nickname") or "")[:128],
fans=parse_count(record.get("fans")),
total_favorited=parse_count(record.get("total_favorited")),
works_count=parse_count(record.get("works")),
following=parse_count(record.get("following")),
captured_at=now,
)
)
written += 1
return written
async def ingest_run(
session: AsyncSession,
run: MonitorRun,
task: MonitorTask,
out_dir: Path,
output_tail: Optional[Sequence[str]] = None,
) -> IngestResult:
"""Ingest one finished run and return what changed.
Sets ``run.status``, ``run.is_baseline`` and the counters on the run row.
On a failed or untrustworthy run nothing is diffed -- the "seen" sets only
ever grow, so a partial run must never be allowed to look like deletions.
"""
# A non-zero exit is a genuine crash: trust nothing this run produced.
if run.exit_code not in (0, None):
run.status = RUN_FAILED
cause = diagnose_failure(output_tail)
run.error_message = describe_exit_code(run.exit_code, cause)
title = f"采集进程异常退出(code={run.exit_code})"
if cause:
# 标题里也带上真因:企业微信通知和事件流都只看这一行,不写就还得去翻日志。
title = f"{title}:{cause}"
await _emit(
session,
run,
EVENT_RUN_FAILED,
title,
severity="error",
payload={
"exit_code": run.exit_code,
"detail": run.error_message,
"cause": cause,
},
)
return IngestResult(status=RUN_FAILED, error=run.error_message)
adapter = adapters.adapter(task.platform)
subdir = adapters.artifact_dir(task.platform)
contents_paths, comment_paths = find_run_files(out_dir, task.platform)
contents = [record for path in contents_paths for record in _read_jsonl(path)]
comments = [record for path in comment_paths for record in _read_jsonl(path)]
run.notes_fetched = len(contents)
run.comments_fetched = len(comments)
# 账号级快照**在「一条作品都没采到」的早退之前**落。博主的粉丝数并不会因为他这个
# 月的新作品列表被风控挡住就不存在 —— 那正是最该看到「粉丝还在涨、但新作品没在发现」
# 的时刻,跳过它等于在最需要它的那轮把数据丢掉。
profiles = [
record
for path in find_profile_files(out_dir, task.platform)
for record in _read_jsonl(path)
]
await _ingest_creator_stats(session, run, profiles)
# A bad cookie does NOT fail the process: XHS cookie login is never validated,
# so an unauthenticated session just returns zero notes with exit 0 -- and
# usually does not even create an output file. Treating that as "the creator
# posted nothing" would silently hide login outages, which is exactly what
# monitoring exists to catch.
if not contents:
run.status = RUN_PARTIAL
# 先排除「东西抓到了,只是没落在我们找的那个目录里」。这种故障的现象和登录失效
# 一模一样,但登录其实是好的 —— 按登录失效报会把人指到完全错的方向去查。
misplaced = _misplaced_output_dirs(out_dir, subdir)
if misplaced:
run.error_message = (
f"crawler wrote into {misplaced} but platform {task.platform} "
f"expects {subdir}"
)
await _emit(
session,
run,
EVENT_NO_DATA,
f"采集产物目录与平台不匹配(实际 {misplaced}、期望 {subdir}),本次未读到任何作品",
severity="error",
payload={
"out_dir": str(out_dir),
"expected": subdir,
"found": misplaced,
},
)
return IngestResult(
status=RUN_PARTIAL,
error=run.error_message,
comments_fetched=len(comments),
)
# Blaming the cookie is only honest if nothing else is authenticating.
# A sibling task that just succeeded proves the login works, so the
# fault is with this target (bad/expired per-creator token, an empty
# account, or a page-structure change).
if await _another_task_succeeded_recently(session, run.task_id):
run.error_message = (
"Crawler produced no notes for this target, but other tasks "
"succeeded recently, so the login is probably fine"
)
await _emit(
session,
run,
EVENT_NO_DATA,
"本次未抓到任何作品:其他任务近期采集正常,登录态应该没问题,请检查该目标是否有效",
severity="warning",
payload={"out_dir": str(out_dir)},
)
else:
run.error_message = "Crawler produced no notes; the login cookie may have expired"
await _emit(
session,
run,
EVENT_AUTH_FAILURE,
"疑似登录态失效:本次未抓到任何作品,请检查 Cookie",
severity="error",
payload={"out_dir": str(out_dir)},
)
return IngestResult(
status=RUN_PARTIAL,
error=run.error_message,
comments_fetched=len(comments),
)
is_baseline = await _count_prior_successes(session, run.task_id, run.id) == 0
run.is_baseline = is_baseline
run.status = RUN_SUCCESS
run.error_message = None
previous_started_at = (
None if is_baseline else await _previous_run_started_at(session, run.task_id, run.id)
)
result = IngestResult(
status=RUN_SUCCESS,
notes_fetched=len(contents),
comments_fetched=len(comments),
is_baseline=is_baseline,
)
result.new_notes = await _ingest_notes(session, run, contents, is_baseline, adapter)
if task.enable_comments:
result.new_comments = await _ingest_comments(
session, run, comments, is_baseline, previous_started_at, adapter
)
run.new_notes = result.new_notes
run.new_comments = result.new_comments
return result
+177
View File
@@ -0,0 +1,177 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/migrate_from_sqlite.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""One-off: copy the monitoring database from SQLite into MySQL.
python -m api.monitor.migrate_from_sqlite [--source data/monitor.db] [--dry-run]
Primary keys are preserved rather than reassigned, because rows in
``monitor_note`` / ``monitor_comment`` / ``monitor_run`` reference ``task_id``;
letting MySQL auto-assign new ids would silently break those links.
Refuses to run against a target that already holds data unless ``--force`` is
given, so a second accidental run cannot double everything up.
"""
import argparse
import sqlite3
import sys
from pathlib import Path
from typing import Any, Dict, List
PROJECT_ROOT = Path(__file__).parent.parent.parent
# Insert order matters: monitor_target and monitor_run carry real foreign keys to
# monitor_task, so the parent rows have to land first.
TABLES_IN_ORDER = [
"monitor_task",
"monitor_target",
"monitor_run",
"monitor_note",
"monitor_note_metric",
"monitor_comment",
"monitor_event",
"monitor_setting",
"auth_session",
]
def read_sqlite(path: Path) -> Dict[str, List[Dict[str, Any]]]:
if not path.exists():
raise SystemExit(f"找不到源库:{path}")
connection = sqlite3.connect(path)
connection.row_factory = sqlite3.Row
try:
existing = {
row[0]
for row in connection.execute(
"SELECT name FROM sqlite_master WHERE type='table'"
)
}
data: Dict[str, List[Dict[str, Any]]] = {}
for table in TABLES_IN_ORDER:
if table not in existing:
continue
rows = [dict(row) for row in connection.execute(f"SELECT * FROM {table}")]
if rows:
data[table] = rows
return data
finally:
connection.close()
def migrate(source: Path, dry_run: bool, force: bool) -> None:
import pymysql
from . import db as monitor_db
data = read_sqlite(source)
if not data:
print("源库里没有可迁移的数据。")
return
print("源库内容:")
for table, rows in data.items():
print(f" {table:22} {len(rows)} 行")
url = monitor_db.resolve_db_url()
if not url.startswith("mysql"):
raise SystemExit(f"目标不是 MySQL:{url}")
connection = pymysql.connect(
host=monitor_db.MYSQL_HOST(),
port=monitor_db.MYSQL_PORT(),
user=monitor_db.MYSQL_USER(),
password=monitor_db.MYSQL_PWD(),
database=monitor_db.MYSQL_DB_NAME(),
charset="utf8mb4",
autocommit=False,
)
try:
with connection.cursor() as cursor:
# Never write outside the configured schema.
cursor.execute("SELECT DATABASE()")
current = cursor.fetchone()[0]
expected = monitor_db.MYSQL_DB_NAME()
if current.lower() != expected.lower():
raise SystemExit(
f"当前连接的是 {current!r},配置要求 {expected!r};已中止。"
)
occupied = []
for table in data:
cursor.execute(f"SELECT COUNT(*) FROM `{table}`")
if cursor.fetchone()[0]:
occupied.append(table)
if occupied and not force:
raise SystemExit(
"目标库已有数据:" + ", ".join(occupied) + "\n"
"加 --force 才会继续(会与现有数据并存,造成重复)。"
)
if dry_run:
print("\n[试运行] 未写入任何数据。")
return
total = 0
for table, rows in data.items():
columns = list(rows[0].keys())
column_sql = ", ".join(f"`{c}`" for c in columns)
placeholders = ", ".join(["%s"] * len(columns))
statement = (
f"INSERT INTO `{table}` ({column_sql}) VALUES ({placeholders})"
)
cursor.executemany(
statement, [[row[c] for c in columns] for row in rows]
)
total += len(rows)
print(f" 已写入 {table:22} {len(rows)} 行")
connection.commit()
print(f"\n完成,共迁移 {total} 行。")
print("提示:源 SQLite 文件仍在原处,确认无误后自行删除。")
except Exception:
connection.rollback()
raise
finally:
connection.close()
def main(argv: List[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="把监控库从 SQLite 迁到 MySQL")
parser.add_argument(
"--source",
default=str(PROJECT_ROOT / "data" / "monitor.db"),
help="SQLite 源文件路径",
)
parser.add_argument("--dry-run", action="store_true", help="只检查,不写入")
parser.add_argument(
"--force", action="store_true", help="目标库已有数据时也继续"
)
args = parser.parse_args(argv)
migrate(Path(args.source), args.dry_run, args.force)
return 0
if __name__ == "__main__":
sys.exit(main())
+484
View File
@@ -0,0 +1,484 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/models.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Monitoring layer data model.
Lives in its own SQLite database (``data/monitor.db``) with its own declarative
Base, deliberately separate from the crawler's ``database/models.py``. The
crawler's DB store overwrites ``liked_count`` and friends in place on every
re-crawl, so it cannot answer "how did this note change?". These tables keep the
history the crawler throws away.
All timestamps are epoch **milliseconds** (BigInteger), matching the project's
own ``tools.time_util.get_current_timestamp()`` convention. Using ints
throughout avoids naive/aware datetime mixing bugs.
"""
from typing import Optional
from sqlalchemy import (
BigInteger,
Boolean,
ForeignKey,
Integer,
String,
Text,
UniqueConstraint,
)
from sqlalchemy.orm import DeclarativeBase, Mapped, mapped_column, relationship
class MonitorBase(DeclarativeBase):
"""Declarative base for the monitoring database."""
# Run statuses
RUN_PENDING = "pending"
RUN_RUNNING = "running"
RUN_SUCCESS = "success"
RUN_PARTIAL = "partial"
RUN_FAILED = "failed"
RUN_TIMEOUT = "timeout"
RUN_INTERRUPTED = "interrupted"
# Event types
EVENT_NEW_NOTE = "new_note"
EVENT_NEW_COMMENT_POSTED = "new_comment_posted"
EVENT_NEW_COMMENT_SEEN = "new_comment_seen"
EVENT_METRIC_DELTA = "metric_delta"
EVENT_RUN_FAILED = "run_failed"
EVENT_AUTH_FAILURE = "suspected_auth_failure"
# A run that completed cleanly yet fetched nothing, where the login is provably
# fine because another task just succeeded with it. The target, not the cookie,
# is what needs looking at.
EVENT_NO_DATA = "no_data_found"
# Task modes. One subprocess handles exactly one crawler type, so a task is
# either creator-driven or note-driven -- never both.
MODE_CREATOR = "creator"
MODE_NOTE = "note"
class MonitorTask(MonitorBase):
"""One monitored schedule: a set of targets, plus when to run them."""
__tablename__ = "monitor_task"
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
name: Mapped[str] = mapped_column(String(200), nullable=False)
platform: Mapped[str] = mapped_column(String(32), nullable=False, default="xhs")
mode: Mapped[str] = mapped_column(String(16), nullable=False)
enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
interval_minutes: Mapped[int] = mapped_column(Integer, nullable=False, default=360)
# How the task is scheduled. `interval` is the original "every N minutes" and
# stays the default; `daily` and `weekly` fire at chosen clock times instead
# (the arithmetic lives in schedule.py).
#
# The clock fields are comma-separated text rather than a child table: they
# are a handful of small integers, always read as a whole, and a table would
# buy nothing but joins.
schedule_mode: Mapped[str] = mapped_column(String(16), nullable=False, default="interval")
# 0-23, e.g. "9,12,18". Empty in interval mode.
schedule_hours: Mapped[str] = mapped_column(String(96), nullable=False, default="")
# 0-6 with Monday = 0, matching Python's date.weekday(). Weekly mode only.
schedule_days: Mapped[str] = mapped_column(String(32), nullable=False, default="")
# Minute past the hour, shared by every time in the schedule.
schedule_minute: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
# Crawl window knobs, mirrored onto each run's CLI flags.
max_notes_count: Mapped[int] = mapped_column(Integer, nullable=False, default=20)
enable_comments: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
max_comments_count: Mapped[int] = mapped_column(Integer, nullable=False, default=50)
run_timeout_seconds: Mapped[int] = mapped_column(Integer, nullable=False, default=3600)
# 通知分成两类,因为它们的性质完全不同:
#
# * `notify_enabled` —— **推送新作品**。可能每轮都有,一条任务列表都推到同一个群
# 会很快变吵,所以默认关。(列名是历史遗留:它早先是唯一的通知开关。)
# * `notify_failures` —— **推送异常**(登录失效 / 运行失败 / 没抓到数据)。频率低,
# 而且一旦发生就意味着这个任务从此**默默采不到任何东西**,你会一直不知道,
# 直到某天发现数据停在几周前。这正是最该被告知的情况,所以默认**开**。
notify_enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
notify_failures: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
# Scheduler state. Persisted so the schedule survives an API restart.
next_run_at: Mapped[Optional[int]] = mapped_column(BigInteger, index=True)
last_run_at: Mapped[Optional[int]] = mapped_column(BigInteger)
last_status: Mapped[str] = mapped_column(String(32), nullable=False, default="idle")
last_error: Mapped[Optional[str]] = mapped_column(Text)
# Lets the UI answer "why did I not get a push for this run?".
last_notified_at: Mapped[Optional[int]] = mapped_column(BigInteger)
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
targets: Mapped[list["MonitorTarget"]] = relationship(
back_populates="task",
cascade="all, delete-orphan",
lazy="selectin",
)
class MonitorTarget(MonitorBase):
"""One watched creator or note belonging to a task.
``external_id`` is the stable identity (XHS user_id / note_id). It is kept
separate from ``xsec_token`` on purpose: tokens expire within weeks, so
treating a tokenised URL as the primary key would make every long-running
task fail eventually.
"""
__tablename__ = "monitor_target"
__table_args__ = (
UniqueConstraint("task_id", "kind", "external_id", name="uq_monitor_target"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
)
kind: Mapped[str] = mapped_column(String(16), nullable=False)
external_id: Mapped[str] = mapped_column(String(128), nullable=False)
xsec_token: Mapped[str] = mapped_column(String(512), nullable=False, default="")
xsec_source: Mapped[str] = mapped_column(String(64), nullable=False, default="")
raw_value: Mapped[str] = mapped_column(Text, nullable=False, default="")
label: Mapped[str] = mapped_column(String(200), nullable=False, default="")
enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
task: Mapped["MonitorTask"] = relationship(back_populates="targets")
class MonitorRun(MonitorBase):
"""One subprocess execution. The run history in the UI is this table."""
__tablename__ = "monitor_run"
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
)
trigger: Mapped[str] = mapped_column(String(16), nullable=False, default="scheduled")
status: Mapped[str] = mapped_column(String(16), nullable=False, default=RUN_PENDING, index=True)
phase: Mapped[str] = mapped_column(String(16), nullable=False)
# Where this run's jsonl landed. Each run gets its own directory because the
# crawler's file writer names output by date only.
save_data_path: Mapped[str] = mapped_column(Text, nullable=False, default="")
queued_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
not_before: Mapped[int] = mapped_column(BigInteger, nullable=False, default=0)
started_at: Mapped[Optional[int]] = mapped_column(BigInteger)
finished_at: Mapped[Optional[int]] = mapped_column(BigInteger)
# BigInteger, not Integer: Windows reports failures as unsigned 32-bit
# NTSTATUS values (0xC0000142 = 3221225794), which overflow MySQL's signed
# INT. SQLite's dynamic typing hid this until the data was migrated.
exit_code: Mapped[Optional[int]] = mapped_column(BigInteger)
notes_fetched: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
comments_fetched: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
new_notes: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
new_comments: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
# The very first successful run of a task establishes the baseline: every
# note is "new" at that point, so emitting events would be pure noise.
is_baseline: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
# Window actually used, so the UI can be honest that comments are the top N
# in the platform's own ordering rather than a complete set.
max_comments_count: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
error_message: Mapped[Optional[str]] = mapped_column(Text)
class MonitorNote(MonitorBase):
"""A note ever seen by a task, plus when it was first/last seen.
Grain is (task, note) so the same note tracked by two tasks stays independent.
"""
__tablename__ = "monitor_note"
__table_args__ = (
UniqueConstraint("task_id", "note_id", name="uq_monitor_note"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
)
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
note_url: Mapped[str] = mapped_column(Text, nullable=False, default="")
cover: Mapped[str] = mapped_column(Text, nullable=False, default="")
# 创作者匿名哈希。爬虫刻意不落原始 user_id(见 tools/user_hash.py),
# 所以这是唯一稳定的创作者标识 —— 按博主分组就靠它。
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, default="")
# 创作者昵称,**已由爬虫脱敏**(张***三 这种)。存的是脱敏后的值,与项目一贯的
# 匿名化姿态一致;不存的话分组只能显示一串哈希,根本认不出是谁。
creator_name: Mapped[str] = mapped_column(String(200), nullable=False, default="")
source_kind: Mapped[str] = mapped_column(String(16), nullable=False, default="")
published_at: Mapped[Optional[int]] = mapped_column(BigInteger)
first_seen_run_id: Mapped[Optional[int]] = mapped_column(Integer)
first_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
last_seen_run_id: Mapped[Optional[int]] = mapped_column(Integer)
last_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
class MonitorNoteMetric(MonitorBase):
"""One metric snapshot per (task, note, run) -- the time series.
Raw strings are kept alongside the parsed integers so a mis-parsed "1.2万"
can always be audited after the fact.
"""
__tablename__ = "monitor_note_metric"
__table_args__ = (
UniqueConstraint("task_id", "note_id", "run_id", name="uq_note_metric"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
run_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
# NULL (not 0) when the platform value could not be parsed: storing 0 would
# forge a large negative delta on the next comparison.
liked_count: Mapped[Optional[int]] = mapped_column(Integer)
comment_count: Mapped[Optional[int]] = mapped_column(Integer)
collected_count: Mapped[Optional[int]] = mapped_column(Integer)
share_count: Mapped[Optional[int]] = mapped_column(Integer)
raw_liked_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
raw_comment_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
raw_collected_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
raw_share_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
class MonitorComment(MonitorBase):
"""A comment ever seen by a task.
The (task, note, comment) uniqueness gives idempotent dedup across runs for
free -- re-running the same crawl cannot double-count.
"""
__tablename__ = "monitor_comment"
__table_args__ = (
UniqueConstraint("task_id", "note_id", "comment_id", name="uq_monitor_comment"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
comment_id: Mapped[str] = mapped_column(String(128), nullable=False)
content: Mapped[str] = mapped_column(Text, nullable=False, default="")
nickname: Mapped[str] = mapped_column(String(200), nullable=False, default="")
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, default="")
# Platform-stated publish time. Used to distinguish a genuinely new comment
# from one that merely entered the visible top-N window this run.
create_time: Mapped[Optional[int]] = mapped_column(BigInteger)
like_count: Mapped[Optional[int]] = mapped_column(Integer)
sub_comment_count: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
parent_comment_id: Mapped[str] = mapped_column(String(128), nullable=False, default="")
first_seen_run_id: Mapped[Optional[int]] = mapped_column(Integer)
first_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
class MonitorEvent(MonitorBase):
"""Append-only change feed. This is what the dashboard reads."""
__tablename__ = "monitor_event"
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
run_id: Mapped[Optional[int]] = mapped_column(Integer, index=True)
type: Mapped[str] = mapped_column(String(32), nullable=False, index=True)
severity: Mapped[str] = mapped_column(String(16), nullable=False, default="info")
target_kind: Mapped[str] = mapped_column(String(16), nullable=False, default="")
target_id: Mapped[str] = mapped_column(String(128), nullable=False, default="")
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
payload_json: Mapped[str] = mapped_column(Text, nullable=False, default="{}")
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
is_read: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
class MonitorCreatorAlias(MonitorBase):
"""给博主起的备注。
作品栏和评论栏都按 ``creator_hash`` 把作品归到博主名下,可那是个哈希;
``creator_name`` 是平台上的昵称(而且粉丝少的号常常没有)。两样都认不出"这是谁"。
备注是**人自己起的名字**(「竞品A」「自家号-3」),用来把账号对上人。
键取 ``(platform, creator_hash)``:哈希对同一个 uid 是稳定的,所以同一个博主出现在
多个任务里时备注也是同一个,不用每个任务各填一遍。
"""
__tablename__ = "monitor_creator_alias"
__table_args__ = (
UniqueConstraint("platform", "creator_hash", name="uq_creator_alias"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
platform: Mapped[str] = mapped_column(String(16), nullable=False, index=True)
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
alias: Mapped[str] = mapped_column(String(128), nullable=False, default="")
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
class MonitorCreatorStat(MonitorBase):
"""博主的**账号级**快照:粉丝数 / 总获赞 / 作品数 / 关注数。
这是作品列表给不了的东西:作品级指标说"这一条视频涨了多少赞",账号级说"这个人
整个账号的粉丝是在涨还是在掉"。两者不互相替代。
粒度取 ``(任务, 博主, 轮次)``,和作品指标一样的形状 —— 于是趋势、差分、报表那套
现成的逻辑换个表就能用。
目前**只有抖音**会写它:小红书那条走的是爬虫子进程,而它的 ``save_creator()`` 在
教学版里是空函数,根本没落过创作者资料。所以表里只有抖音的博主。
"""
__tablename__ = "monitor_creator_stat"
__table_args__ = (
UniqueConstraint(
"task_id", "creator_hash", "run_id", name="uq_creator_stat"
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
)
run_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
nickname: Mapped[str] = mapped_column(String(128), nullable=False, default="")
# 都可能为 None:平台没给就留空,**不要伪造成 0** —— 0 是"掉到零",和"不知道"
# 在趋势图上是完全不同的两回事。
fans: Mapped[Optional[int]] = mapped_column(BigInteger)
total_favorited: Mapped[Optional[int]] = mapped_column(BigInteger)
works_count: Mapped[Optional[int]] = mapped_column(BigInteger)
following: Mapped[Optional[int]] = mapped_column(BigInteger)
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
class MonitorNoteAlias(MonitorBase):
"""给**作品**起的备注。
和 ``MonitorCreatorAlias`` 是一对:博主那条回答"这是谁",这条回答"这条我要盯着"。
键取 ``(platform, note_id)`` —— 作品 id 本身就带平台语义,但显式带上 platform 才能和
博主备注用同一套查询形状。
"""
__tablename__ = "monitor_note_alias"
__table_args__ = (
UniqueConstraint("platform", "note_id", name="uq_note_alias"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
platform: Mapped[str] = mapped_column(String(16), nullable=False, index=True)
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
alias: Mapped[str] = mapped_column(String(128), nullable=False, default="")
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
class MonitorSetting(MonitorBase):
"""Key/value store. Holds the XHS cookie for unattended runs."""
__tablename__ = "monitor_setting"
key: Mapped[str] = mapped_column(String(64), primary_key=True)
value: Mapped[str] = mapped_column(Text, nullable=False, default="")
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
class AuthSession(MonitorBase):
"""A WebUI login session.
Only the SHA-256 of the token is stored, never the token itself -- a leaked
database therefore does not hand over live sessions. This mirrors the
existing posture of never returning the XHS cookie or webhook value.
A stateful table (rather than a signed stateless token) is what makes "log
out" and "password changed" take effect immediately.
"""
__tablename__ = "auth_session"
token_hash: Mapped[str] = mapped_column(String(64), primary_key=True)
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
expires_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
last_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
SETTING_AUTH_PASSWORD_HASH = "auth_password_hash"
SETTING_AUTH_PASSWORD_UPDATED_AT = "auth_password_updated_at"
# Settings are namespaced by scope: `platform.<p>.<name>` for values each
# platform keeps its own copy of, `system.<name>` for values shared across all of
# them. Key builders live in settings.py.
SETTING_WECOM_WEBHOOK = "system.wecom_webhook"
# 上游更新检查的两条状态。都不是给用户编辑的设置项,所以不在 app_settings 的注册表里
# (那张表只列可编辑项,因此也不会被设置接口读出来)。
#
# 最近一次检查的结果整体存成一条 JSON:它总是被整体读写,拆成多个 key 只会带来
# 另一半没写完的不一致。
SETTING_UPSTREAM_STATE = "system.upstream_check_state"
# 已经推送过通知的那个上游 tip。换 tip 才再推 —— 否则每个检查周期都会把同样的
# 更新推一遍,直到有人去合并为止;而上游真又动了的时候应该再推一次。
SETTING_UPSTREAM_NOTIFIED_TIP = "system.upstream_notified_tip"
# Pre-namespacing keys, kept only so the startup migration can find and move
# them. Nothing should read these directly.
LEGACY_SETTING_KEY_RENAMES = {
# Pre-batch-2 flat keys.
"xhs_cookie": "platform.xhs.cookie",
"xhs_cookie_updated_at": "platform.xhs.cookie_updated_at",
"xhs_cookie_last_ok_at": "platform.xhs.cookie_last_ok_at",
"wecom_webhook": "system.wecom_webhook",
# Batch-2 keys, before settings gained a scope. Those values belonged to
# Xiaohongshu because it was the only platform, so they migrate to its scope;
# the two scheduling keys were always instance-wide.
"collect.default_interval_minutes": "platform.xhs.default_interval_minutes",
"collect.default_max_notes": "platform.xhs.default_max_notes",
"collect.default_max_comments": "platform.xhs.default_max_comments",
"collect.enable_sub_comments": "platform.xhs.enable_sub_comments",
"collect.crawl_sleep_sec": "platform.xhs.crawl_sleep_sec",
"collect.active_hours_start": "system.active_hours_start",
"collect.active_hours_end": "system.active_hours_end",
"proxy.enable_ip_proxy": "platform.xhs.enable_ip_proxy",
"proxy.provider": "platform.xhs.proxy_provider",
"proxy.pool_count": "platform.xhs.proxy_pool_count",
"proxy.static_proxy_url": "platform.xhs.static_proxy_url",
}
# utf8mb4 is forced on every table rather than left to the schema default: this
# deployment's MySQL server *and* the target database both default to latin1,
# which would mangle or reject Chinese text. Setting it per table means it holds
# regardless of what the schema default happens to be.
#
# Must run after every model is declared, hence the end of the module.
for _table in MonitorBase.metadata.tables.values():
_table.kwargs["mysql_charset"] = "utf8mb4"
_table.kwargs["mysql_collate"] = "utf8mb4_unicode_ci"
+217
View File
@@ -0,0 +1,217 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/notify.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Push notifications via a WeCom (企业微信) group robot webhook.
Two rules shape this module:
* **One message per run, not per event.** A run that finds twenty new notes must
produce one summary, not twenty pushes.
* **A failed push never fails the crawl.** Notification is best-effort: the run's
data is already committed by the time we get here, so every error is logged
and swallowed.
"""
import json
from typing import Optional
import httpx
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .models import (
EVENT_AUTH_FAILURE,
EVENT_NEW_NOTE,
EVENT_NO_DATA,
EVENT_RUN_FAILED,
MonitorEvent,
MonitorRun,
MonitorTask,
)
from .settings import get_setting
# Short on purpose: the scheduler awaits the run, so a hanging webhook would
# stall every other task behind it.
WEBHOOK_TIMEOUT_SECONDS = 10.0
# Only these event types are worth interrupting someone for. NO_DATA is included
# because a run that fetched nothing at all is always anomalous -- a creator
# always has *some* notes -- even when the login is not the culprit.
NOTIFIABLE_EVENT_TYPES = (
EVENT_AUTH_FAILURE,
EVENT_RUN_FAILED,
EVENT_NO_DATA,
EVENT_NEW_NOTE,
)
# WeCom markdown is a limited subset; coloured text is the one bit of flair it
# supports and it makes failures stand out in a busy group chat.
_COLOR_WARNING = "warning"
_COLOR_INFO = "info"
async def send_wecom(webhook_url: str, content: str) -> tuple[bool, str]:
"""Post a markdown message to a WeCom group robot.
Returns (ok, detail) rather than raising, so callers can surface the reason
in the UI when the user clicks "send test".
"""
if not webhook_url:
return False, "Webhook 未配置"
payload = {"msgtype": "markdown", "markdown": {"content": content}}
try:
async with httpx.AsyncClient(timeout=WEBHOOK_TIMEOUT_SECONDS) as client:
response = await client.post(webhook_url, json=payload)
response.raise_for_status()
body = response.json()
except httpx.HTTPError as exc:
return False, f"请求失败:{exc}"
except json.JSONDecodeError:
return False, "返回内容不是合法 JSON,请检查 Webhook 地址"
# WeCom answers 200 with a non-zero errcode on failure.
errcode = body.get("errcode")
if errcode != 0:
return False, f"企业微信返回 errcode={errcode} {body.get('errmsg', '')}"
return True, "发送成功"
async def get_webhook_url(session: AsyncSession) -> str:
from .models import SETTING_WECOM_WEBHOOK
return (await get_setting(session, SETTING_WECOM_WEBHOOK)) or ""
async def build_run_message(
session: AsyncSession,
task: MonitorTask,
run: MonitorRun,
) -> Optional[str]:
"""Compose one markdown summary for a finished run, or None if nothing to say.
事件按开关过滤:只勾了「新作品」的任务,不该因为一次失败被推消息,反之亦然 ——
否则拆开这两个开关就没有意义了。
"""
allowed = []
if task.notify_enabled:
allowed.append(EVENT_NEW_NOTE)
if task.notify_failures:
allowed.extend([EVENT_AUTH_FAILURE, EVENT_RUN_FAILED, EVENT_NO_DATA])
if not allowed:
return None
events = list(
(
await session.scalars(
select(MonitorEvent)
.where(
MonitorEvent.run_id == run.id,
MonitorEvent.type.in_(allowed),
)
.order_by(MonitorEvent.id)
)
).all()
)
if not events:
return None
failures = [
e for e in events if e.type in (EVENT_AUTH_FAILURE, EVENT_RUN_FAILED, EVENT_NO_DATA)
]
new_notes = [e for e in events if e.type == EVENT_NEW_NOTE]
lines: list[str] = []
if failures:
# Word the header from what actually happened, not from whether new notes
# accompanied it: a login outage usually brings no new notes either.
unavailable = any(e.type == EVENT_NO_DATA for e in failures) and not any(
e.type in (EVENT_AUTH_FAILURE, EVENT_RUN_FAILED) for e in failures
)
header = "监控任务未抓到数据" if unavailable else "监控任务异常"
lines.append(f"**⚠️ {header}:{task.name}**")
for event in failures:
lines.append(f'> <font color="{_COLOR_WARNING}">{event.title}</font>')
else:
lines.append(f"**📢 监控任务有新作品:{task.name}**")
if new_notes:
lines.append(f"> 新增作品 **{len(new_notes)}** 篇")
# Cap the listing: a first-ever run or a long gap can produce a lot, and
# a wall of text is worse than a count.
for event in new_notes[:10]:
payload = _load_payload(event.payload_json)
title = payload.get("title") or event.target_id
note_id = payload.get("note_id") or event.target_id
# 链接形状按平台来。抖音的作品是 /video/{id},写死小红书域名的话,
# 群里点进去会是一个 404 —— 而这正是通知唯一要它干的事。
url = adapters.adapter(task.platform).note_url(note_id)
lines.append(f"> [{title}]({url})")
if len(new_notes) > 10:
lines.append(f"> …等共 {len(new_notes)} 篇")
if run.is_baseline:
lines.append("> (首轮基线,未计入新增统计)")
return "\n".join(lines)
def _load_payload(raw: str) -> dict:
try:
payload = json.loads(raw or "{}")
except json.JSONDecodeError:
return {}
return payload if isinstance(payload, dict) else {}
async def notify_run(session: AsyncSession, task: MonitorTask, run: MonitorRun) -> Optional[str]:
"""Push a summary for a finished run if the task opted in.
Returns the message that was sent, or None. Never raises.
"""
try:
# 两个开关是分开的:只开「异常」不该因为新作品而发消息,反之亦然。
if not (task.notify_enabled or task.notify_failures):
return None
webhook_url = await get_webhook_url(session)
if not webhook_url:
return None
message = await build_run_message(session, task, run)
if not message:
return None
ok, detail = await send_wecom(webhook_url, message)
if not ok:
print(f"[monitor.notify] task {task.id} push failed: {detail}")
return None
task.last_notified_at = get_current_timestamp()
return message
except Exception as exc: # pragma: no cover - notification must never break a run
print(f"[monitor.notify] unexpected error: {exc}")
return None
+216
View File
@@ -0,0 +1,216 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/platforms.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Platform capability matrix.
The single source of truth for what each platform can do. The UI renders its
platform switcher and metric columns from this, and the API validates against
it.
Two distinct things are recorded here, and conflating them would be misleading:
* ``crawler_modes`` / ``metrics`` / ``comment_levels`` / ``media`` describe what
the upstream crawler module actually supports. These were read out of the
platform modules, not assumed -- all seven implement search/detail/creator;
the real differences are in which interaction metrics they capture.
* ``monitor_wired`` says whether the *monitoring layer* has been hooked up. It
covers Xiaohongshu and Douyin. The parts where those two differ -- which
directory the crawler writes into, what the jsonl fields are called, what a
target URL looks like -- live in ``adapters.py``; the rest of the layer is
platform-neutral.
A platform can therefore be fully crawlable by upstream and still not usable for
monitoring, which is exactly the state of the other five today.
"""
from typing import Any, Dict, List, Optional
PLATFORM_XHS = "xhs"
PLATFORM_LABELS = {
"xhs": "小红书",
"dy": "抖音",
"ks": "快手",
"bili": "B站",
"wb": "微博",
"tieba": "贴吧",
"zhihu": "知乎",
}
# Interaction metrics each platform's store actually persists. Xiaohongshu has no
# play count or danmaku; Bilibili has both and the widest set; Kuaishou carries
# no comment/share/collect at all; Tieba stores only reply counts.
PLATFORM_CAPABILITIES: Dict[str, Dict[str, Any]] = {
"xhs": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": True,
"target_hints": {
"creator": "https://www.xiaohongshu.com/user/profile/5f58bd990000000001003753",
"note": "https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=...",
"creator_label": "博主主页",
"note_label": "笔记",
# 只有小红书的链接带会过期的 xsec_token。抖音的链接不带令牌,永久有效,
# 那句「建议只填纯 ID」的劝告对它没有意义。
"token_expires": True,
},
},
"dy": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": True,
# 用户可见的示例链接(前端的目标输入框用它做 placeholder)。放这里是因为
# 它属于「这个平台长什么样」的能力描述;真正干活的管子(正则、目录名、
# 字段别名)在 adapters.py。
"target_hints": {
"creator": "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE",
"note": "https://www.douyin.com/video/7525082444551310602",
"creator_label": "博主主页",
"note_label": "作品",
},
},
"ks": {
"crawler_modes": ["search", "detail", "creator"],
# No comment/share/collect in the Kuaishou store; sub-comments are stored
# flat with no parent link and carry no like count.
"metrics": ["liked_count", "view_count"],
"comment_levels": 1,
"media": True,
"monitor_wired": False,
},
"bili": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": [
"liked_count",
"video_play_count",
"video_danmaku",
"comment_count",
"video_favorite_count",
"video_coin_count",
"video_share_count",
],
"comment_levels": 2,
"media": True,
"monitor_wired": False,
},
"wb": {
"crawler_modes": ["search", "detail", "creator"],
# Weibo has no collect count, and its comment count field is named
# differently in the model.
"metrics": ["liked_count", "comments_count", "shared_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": False,
},
"tieba": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["total_replay_num", "total_replay_page"],
"comment_levels": 2,
"media": False,
"monitor_wired": False,
},
"zhihu": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["voteup_count", "comment_count"],
"comment_levels": 2,
"media": False,
"monitor_wired": False,
},
}
METRIC_LABELS = {
"liked_count": "点赞",
"comment_count": "评论",
"collected_count": "收藏",
"share_count": "分享",
"view_count": "播放",
"video_play_count": "播放",
"video_danmaku": "弹幕",
"video_favorite_count": "收藏",
"video_coin_count": "投币",
"video_share_count": "分享",
"comments_count": "评论",
"shared_count": "转发",
"total_replay_num": "回复数",
"total_replay_page": "回复页数",
"voteup_count": "赞同",
}
# Monitoring modes, mapped to the CLI crawler types upstream understands.
MONITOR_MODE_CREATOR = "creator"
MONITOR_MODE_NOTE = "note"
CLI_TYPE_FOR_MODE = {
MONITOR_MODE_CREATOR: "creator",
MONITOR_MODE_NOTE: "detail",
}
class UnsupportedPlatformError(ValueError):
"""Raised for an unknown platform, or one the monitor layer cannot run."""
def all_platforms() -> List[str]:
return list(PLATFORM_CAPABILITIES)
def is_known(platform: str) -> bool:
return platform in PLATFORM_CAPABILITIES
def is_monitor_wired(platform: str) -> bool:
return bool(PLATFORM_CAPABILITIES.get(platform, {}).get("monitor_wired"))
def describe(platform: str) -> Optional[Dict[str, Any]]:
capability = PLATFORM_CAPABILITIES.get(platform)
if capability is None:
return None
return {
"value": platform,
"label": PLATFORM_LABELS.get(platform, platform),
**capability,
"metric_labels": {
metric: METRIC_LABELS.get(metric, metric) for metric in capability["metrics"]
},
}
def describe_all() -> List[Dict[str, Any]]:
return [describe(platform) for platform in all_platforms()]
def ensure_runnable(platform: str) -> None:
"""Validate a platform for a monitoring task.
An unwired platform is rejected outright rather than accepted and left to
silently produce nothing -- the same silent-failure shape that made a valid
creator look like an expired login earlier.
"""
if not is_known(platform):
raise UnsupportedPlatformError(
f"未知平台:{platform}(支持:{', '.join(all_platforms())})"
)
if not is_monitor_wired(platform):
label = PLATFORM_LABELS.get(platform, platform)
raise UnsupportedPlatformError(
f"{label}的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。"
)
+425
View File
@@ -0,0 +1,425 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/qrlogin.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""扫码登录,以及"现在到底登没登录"的查询。
**为什么需要这个模块**:服务器上 Chrome 跑在 Xvfb 里没有显示器,爬虫原本用
`show_qrcode`(PIL 的 `Image.show()`)弹窗展示二维码,那需要桌面看图程序,服务器上
没有。所以改成经 CDP 把二维码从页面里读出来交给 WebUI。
**为什么登录状态要能独立查询**:扫码会话是内存里的临时状态,进程一重启就没了
(部署、崩溃都算)。把"是否已登录"绑在它上面,就会出现"扫完了但界面没反应、
也不知道到底成没成"。所以状态查询是独立的、随时可调用的,二维码只是达成它的手段之一。
三个容易搞错的地方:
* **必须复用浏览器默认 context**。`browser.new_context()` 会造出一个无痕式的 profile,
扫了也白扫——爬虫读不到那份 cookie。真正的 profile 在 `browser.contexts[0]`。
* **绝不能调 `browser.close()`**。对 CDP 连接而言那会关掉操作者自己的 Chrome,
连带所有无关标签页。只能关本模块自己开的那一个。
* **不能靠 `web_session` 判断登录**。实测:一个全新的空 profile 首次访问小红书就会
被发一个 `web_session`,所以"有这个 cookie"什么都证明不了。可信信号是页面自己的
`__INITIAL_STATE__.user.loggedIn`。
"""
import asyncio
import os
import time
from typing import Any, Dict, Optional
import config
from playwright.async_api import async_playwright
from tools import utils
from ..creator.client import CreatorApiError, CreatorClient
from .platforms import PLATFORM_XHS
def _cookie_string(cookies) -> str:
"""把 CDP 拿到的 cookie 列表拼成请求头用的字符串。"""
return "; ".join(f"{c['name']}={c['value']}" for c in cookies)
# 二维码有效期。平台自己会更早轮换;这个上限只是为了让一次被放弃的尝试不会
# 永久占着一个标签页。
QR_TTL_SECONDS = 300
# 登录状态查询的缓存时长。轮询时不必每次都去问浏览器。
STATE_CACHE_SECONDS = 5
STATUS_IDLE = "idle"
STATUS_WAITING = "waiting"
STATUS_SUCCESS = "success"
STATUS_EXPIRED = "expired"
STATUS_ERROR = "error"
# 只有小红书接了监控流程,所以扫码也只对它开放。给别的平台显示一个按不动的按钮
# 是在假装功能存在。
LOGIN_URL: Dict[str, str] = {PLATFORM_XHS: "https://www.xiaohongshu.com"}
EXPLORE_URL: Dict[str, str] = {PLATFORM_XHS: "https://www.xiaohongshu.com/explore"}
QR_SELECTOR: Dict[str, str] = {PLATFORM_XHS: "xpath=//img[@class='qrcode-img']"}
LOGIN_BUTTON_SELECTOR: Dict[str, str] = {
PLATFORM_XHS: "xpath=//*[@id='app']/div[1]/div[2]/div[1]/ul/div[1]/button"
}
# 这里本来有一个读 window.__INITIAL_STATE__ 的 JS 探针,**已删除,不要加回来**。
#
# 它是页面加载那一刻的快照:浏览器本来就登录着时它是对的,但扫码是加载**之后**才
# 登录的,快照不会翻转,检测于是永远等不到 —— 表现为"扫了码却一直停在二维码上"。
# 运营模块踩过同一个坑。现在的判据是拿 cookie 问后台接口,见 check_login_state。
_lock = asyncio.Lock()
_current: Optional["QrLoginSession"] = None
# 常驻的 Playwright 客户端和本模块自己的标签页。长期持有是有意的:状态查询要能
# 随时回答,而每次都新建一个标签页会在操作者的浏览器里堆垃圾。
_playwright: Any = None
_page: Any = None
# (时间戳, 结果),避免轮询时反复问浏览器。
_state_cache: Optional[tuple[float, Dict[str, Any]]] = None
def _cdp_url() -> str:
"""浏览器 DevTools 端点。``MC_CDP_URL`` 优先,便于换主机而不用改代码。"""
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
def _login_url(platform: str) -> str:
if platform == PLATFORM_XHS and getattr(config, "XHS_INTERNATIONAL", False):
return "https://www.rednote.com"
return LOGIN_URL[platform]
async def _ensure_context() -> Any:
"""连上浏览器并返回它的默认 context。"""
global _playwright
if _playwright is None:
_playwright = await async_playwright().start()
try:
browser = await _playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
except Exception as exc:
await _reset_playwright()
raise RuntimeError(
f"连接浏览器失败({_cdp_url()})。请确认服务器上的 Chrome 以 "
f"--remote-debugging-port 启动。原始错误:{exc}"
) from exc
if not browser.contexts:
raise RuntimeError(
"浏览器没有可用上下文。CDP 已连上,但读不到 profile —— "
"请确认 Chrome 不是以无痕模式启动的。"
)
# contexts[0] 就是真实 profile,用它,不要 new_context()。
return browser.contexts[0]
async def _ensure_page(platform: str = PLATFORM_XHS, reload: bool = False) -> Any:
"""本模块在操作者浏览器里的那一个标签页,复用而不是反复新建。
若已有一个停在目标站点的标签页就认领它——进程重启后页柄会丢,但标签页还在,
认领可以避免在浏览器里留下一堆没人关的孤儿页。
"""
global _page
context = await _ensure_context()
if _page is not None:
try:
if _page.is_closed():
_page = None
except Exception:
_page = None
if _page is None:
for candidate in context.pages:
try:
if "xiaohongshu.com" in candidate.url or "rednote.com" in candidate.url:
_page = candidate
break
except Exception:
continue
if _page is None:
_page = await context.new_page()
try:
url = _page.url
except Exception:
url = ""
if reload or "xiaohongshu.com" not in url and "rednote.com" not in url:
await _page.goto(
EXPLORE_URL.get(platform, EXPLORE_URL[PLATFORM_XHS]),
wait_until="domcontentloaded",
timeout=45000,
)
return _page
async def check_login_state(force: bool = False) -> Dict[str, Any]:
"""问浏览器:现在登录了吗?
``force`` 会先重新加载页面。SPA 的状态会随登录实时更新,所以轮询时不必重载;
但若登录态是在别处失效的,页面上的副本可能是陈旧的,重新检测就该重载。
"""
global _state_cache
now = time.time()
if not force and _state_cache is not None:
cached_at, cached = _state_cache
if now - cached_at < STATE_CACHE_SECONDS:
return cached
try:
context = await _ensure_context()
cookies = await context.cookies()
except Exception as exc:
result = {
"known": False,
"logged_in": False,
"nickname": None,
"error": f"{exc.__class__.__name__}: {exc}",
}
_state_cache = (now, result)
return result
cookie = _cookie_string(cookies)
# 判据不再是页面里的 window.__INITIAL_STATE__ —— 那是**页面加载那一刻的快照**:
# 浏览器已登录时它是对的,但扫码是加载**之后**才登录的,快照不会翻转,检测就永远
# 等不到(运营模块踩过同一个坑)。改成拿 cookie 问后台「我是谁」,那是权威的:
# 实测游客也会被发一个 web_session,所以「有这个 cookie」什么都证明不了,
# 后台认了才算。
try:
info = await CreatorClient(cookie).fetch_user_info()
except CreatorApiError:
result = {"known": True, "logged_in": False, "nickname": None}
else:
result = {
"known": True,
"logged_in": bool(info.get("user_id")),
"nickname": info.get("nickname"),
}
_state_cache = (now, result)
return result
async def _current_cookie() -> str:
"""默认 profile 当前的小红书 cookie 串。
扫码面板要的不只是「登录了吗」,而是**把登录态拿出来存一份** —— 存进库之后,
即使 CDP 关掉、任务改用 --cookies_file 注入,也照样能跑。
"""
context = await _ensure_context()
return _cookie_string(await context.cookies())
async def _reset_playwright() -> None:
global _playwright, _page
_page = None
if _playwright is not None:
try:
await _playwright.stop()
except Exception:
pass
_playwright = None
class QrLoginSession:
"""一次进行中的扫码尝试。"""
def __init__(self, platform: str, page: Any) -> None:
self.platform = platform
self.status = STATUS_WAITING
self.message = "请用手机扫描二维码"
self.image = ""
self.started_at = time.time()
self.logged_in = False
self.nickname: Optional[str] = None
# 登录成功后从默认 profile 取出来的 cookie,供调用方存库。
self.cookie: str = ""
self.cookie_taken = False
self._page = page
@property
def elapsed(self) -> float:
return time.time() - self.started_at
async def refresh(self) -> None:
"""轮询一次,看扫码是否完成。"""
if self.status != STATUS_WAITING:
return
if self.elapsed > QR_TTL_SECONDS:
self.status = STATUS_EXPIRED
self.message = "二维码已超时,请重新获取"
return
state = await check_login_state()
if state.get("logged_in"):
self.cookie = await _current_cookie()
self.logged_in = True
self.nickname = state.get("nickname")
self.status = STATUS_SUCCESS
who = f"({self.nickname})" if self.nickname else ""
self.message = f"登录成功{who},登录态已写入浏览器 profile"
return
try:
if self._page.is_closed():
self.status = STATUS_ERROR
self.message = "二维码所在页面已被关闭,请重新获取"
except Exception:
pass
def snapshot(self) -> Dict[str, Any]:
return {
"status": self.status,
"platform": self.platform,
"image": self.image,
"message": self.message,
"elapsed": int(self.elapsed),
"expires_in": max(0, int(QR_TTL_SECONDS - self.elapsed)),
"logged_in": self.logged_in,
"nickname": self.nickname,
}
async def _read_qr(page: Any, platform: str) -> str:
"""把二维码从页面里取出来,必要时先点开登录框。"""
image = await utils.find_login_qrcode(page, selector=QR_SELECTOR[platform])
if image:
return image
# 登录框不一定自己弹出来。这是爬虫自身扫码流程里同款兜底。
await asyncio.sleep(0.5)
try:
await page.locator(LOGIN_BUTTON_SELECTOR[platform]).click(timeout=5000)
except Exception:
return ""
return await utils.find_login_qrcode(page, selector=QR_SELECTOR[platform])
async def _discard_current_locked() -> None:
global _current
_current = None
async def start(platform: str = PLATFORM_XHS) -> Dict[str, Any]:
"""在 CDP 浏览器里打开登录页,取回二维码。"""
global _current
if platform not in LOGIN_URL:
raise ValueError(f"平台 {platform} 尚未接入扫码登录(目前仅支持小红书)")
async with _lock:
await _discard_current_locked()
# **先问状态,再决定要不要开页面。** 顺序反过来是有代价的:读二维码内部会
# wait_for_selector 等满 30 秒才放弃,而已经登录时页面上根本没有二维码 ——
# 用户点一下按钮要干等半分钟,还白开一个标签页。
state = await check_login_state(force=True)
if state.get("logged_in"):
# 已经是登录状态时站点不显示二维码 —— 这本身就是成功,不是失败。
# 顺带把 cookie 取出来,让调用方可以存进库。
session = QrLoginSession(platform, None)
session.cookie = await _current_cookie()
session.status = STATUS_SUCCESS
session.logged_in = True
session.nickname = state.get("nickname")
who = f"({session.nickname})" if session.nickname else ""
session.message = f"浏览器已经是登录状态{who},无需扫码"
_current = session
return session.snapshot()
page = await _ensure_page(platform)
try:
await page.goto(
_login_url(platform), wait_until="domcontentloaded", timeout=45000
)
image = await _read_qr(page, platform)
except Exception as exc:
raise RuntimeError(f"打开登录页失败:{exc}") from exc
session = QrLoginSession(platform, page)
session.image = image
if not image:
session.status = STATUS_ERROR
session.message = "页面上没找到二维码,请确认站点结构没有变化"
_current = session
return session.snapshot()
async def status() -> Dict[str, Any]:
async with _lock:
if _current is None:
state = await check_login_state()
return {
"status": STATUS_IDLE,
"platform": None,
"image": "",
"message": "",
"elapsed": 0,
"expires_in": 0,
"logged_in": bool(state.get("logged_in")),
"nickname": state.get("nickname"),
}
await _current.refresh()
return _current.snapshot()
async def take_cookie() -> Optional[str]:
"""取走已登录会话的 cookie,且只给一次。
由路由层在落库时调用。**cookie 不进响应体** —— 它是凭证,前端没有理由看到它。
这里**刻意不结束会话**(与运营模块不同):那里取完即拆,因为临时上下文用完就该丢;
这里的浏览器 profile 是长期存在的,面板还该继续显示「已登录」。所以只标记已取过,
让重复轮询拿不到第二份、也就不会反复写库。
"""
async with _lock:
if _current is None or _current.status != STATUS_SUCCESS or _current.cookie_taken:
return None
_current.cookie_taken = True
return _current.cookie
async def cancel() -> Dict[str, Any]:
async with _lock:
await _discard_current_locked()
state = await check_login_state()
return {
"status": STATUS_IDLE,
"platform": None,
"image": "",
"message": "已取消",
"elapsed": 0,
"expires_in": 0,
"logged_in": bool(state.get("logged_in")),
"nickname": state.get("nickname"),
}
async def shutdown() -> None:
"""进程退出时断开连接。刻意不关那个标签页——它是操作者浏览器的一部分。"""
global _current
async with _lock:
_current = None
await _reset_playwright()
+207
View File
@@ -0,0 +1,207 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/report.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Cross-task reporting: what grew, and what is new, over a date range.
Two families of numbers that answer different questions and are therefore kept
as separate columns:
* **互动增量** — Σ(current − previous) across the selected notes. "How many likes
did this set of notes gain?"
* **新增内容** — count of newly discovered notes and comments. "How much new
material showed up?"
The per-day interaction delta is defined as *last value on the day* minus *last
value before the day* (0 when the note was first seen on that day). That keeps
growth from a note's first observation counted once, rather than smeared across
every later day.
Aggregation runs in Python over the snapshots rather than as one large SQL
query: the per-note-per-day baseline lookup is a windowed operation that SQLite
expresses awkwardly, and the row counts here are small enough that clarity is
worth more than the query planner.
"""
from bisect import bisect_right
from datetime import date, datetime, time, timedelta
from typing import Any, Dict, Iterable, List, Optional, Sequence
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from .models import MonitorComment, MonitorNote, MonitorNoteMetric
METRIC_FIELDS = ("liked_count", "comment_count", "collected_count", "share_count")
METRIC_LABELS = {
"liked_count": "点赞",
"comment_count": "评论",
"collected_count": "收藏",
"share_count": "分享",
}
def day_bounds(day: date) -> tuple[int, int]:
"""Inclusive epoch-millisecond bounds for a local calendar day."""
start = datetime.combine(day, time.min)
end = datetime.combine(day, time.max)
return int(start.timestamp() * 1000), int(end.timestamp() * 1000)
def iter_days(start: date, end: date) -> List[date]:
days = []
cursor = start
while cursor <= end:
days.append(cursor)
cursor += timedelta(days=1)
return days
def compute_daily_rows(
series_by_note: Dict[str, List[tuple[int, Dict[str, Optional[int]]]]],
notes_per_day: Dict[date, int],
comments_per_day: Dict[date, int],
days: Sequence[date],
) -> List[Dict[str, Any]]:
"""Pure aggregation. ``series_by_note`` must be sorted by timestamp ascending."""
prepared = {note_id: ([ts for ts, _ in points], points) for note_id, points in series_by_note.items()}
rows: List[Dict[str, Any]] = []
for day in days:
day_start, day_end = day_bounds(day)
totals = {field: 0 for field in METRIC_FIELDS}
# Records *which* metric could not be compared, not just that something
# could not. A blanket flag loses all value the moment one permanently
# unparseable field makes every row "incomplete".
partial_metrics: set[str] = set()
for times, points in prepared.values():
end_index = bisect_right(times, day_end) - 1
if end_index < 0:
# Not yet tracked on this day.
continue
end_values = points[end_index][1]
start_index = bisect_right(times, day_start - 1) - 1
# No earlier snapshot means the note first appeared in this window,
# so it starts from zero -- all of its count is genuinely new.
start_values = (
points[start_index][1] if start_index >= 0 else {f: 0 for f in METRIC_FIELDS}
)
for field in METRIC_FIELDS:
end_value, start_value = end_values.get(field), start_values.get(field)
if end_value is None or start_value is None:
# An unparseable count on either side makes the delta unknown;
# skipping beats reporting a fabricated number.
partial_metrics.add(field)
continue
totals[field] += end_value - start_value
row: Dict[str, Any] = {
"date": day.isoformat(),
"new_notes": notes_per_day.get(day, 0),
"new_comments": comments_per_day.get(day, 0),
"partial_metrics": sorted(partial_metrics),
}
row.update({f"{field}_delta": value for field, value in totals.items()})
rows.append(row)
return rows
async def build_report(
session: AsyncSession,
task_ids: Optional[Iterable[int]],
start_day: date,
end_day: date,
) -> Dict[str, Any]:
"""Daily rows plus totals for the selected tasks over the given date range."""
start_ms, _ = day_bounds(start_day)
_, end_ms = day_bounds(end_day)
# 必须是 `is not None`,不能写 `if task_ids` —— **空列表是假值**,而空列表在这里
# 的含义是「这个平台一个任务都没有」,不是「不限制平台」。用真值判断的话,
# 切到一个还没有任务的平台,报表会把**所有**任务的数据聚合出来(看起来就是
# 「抖音的报表里全是小红书的数据」)。
scope = list(task_ids) if task_ids is not None else None
days = iter_days(start_day, end_day)
# Fetch every snapshot up to the range end: the delta on the first day needs
# the last value from *before* the range, so a lower bound would be wrong.
metric_stmt = select(MonitorNoteMetric).where(MonitorNoteMetric.captured_at <= end_ms)
if scope is not None:
metric_stmt = metric_stmt.where(MonitorNoteMetric.task_id.in_(scope))
metric_stmt = metric_stmt.order_by(MonitorNoteMetric.note_id, MonitorNoteMetric.run_id)
series_by_note: Dict[str, List[tuple[int, Dict[str, Optional[int]]]]] = {}
included_note_ids: set[str] = set()
for snapshot in (await session.scalars(metric_stmt)).all():
included_note_ids.add(snapshot.note_id)
series_by_note.setdefault(snapshot.note_id, []).append(
(
snapshot.captured_at,
{field: getattr(snapshot, field) for field in METRIC_FIELDS},
)
)
note_stmt = select(MonitorNote.first_seen_at).where(
MonitorNote.first_seen_at >= start_ms, MonitorNote.first_seen_at <= end_ms
)
if scope is not None:
note_stmt = note_stmt.where(MonitorNote.task_id.in_(scope))
comment_stmt = select(MonitorComment.first_seen_at).where(
MonitorComment.first_seen_at >= start_ms, MonitorComment.first_seen_at <= end_ms
)
if scope is not None:
comment_stmt = comment_stmt.where(MonitorComment.task_id.in_(scope))
notes_per_day = _count_by_day((await session.scalars(note_stmt)).all())
comments_per_day = _count_by_day((await session.scalars(comment_stmt)).all())
rows = compute_daily_rows(series_by_note, notes_per_day, comments_per_day, days)
totals = {
"new_notes": sum(row["new_notes"] for row in rows),
"new_comments": sum(row["new_comments"] for row in rows),
}
for field in METRIC_FIELDS:
totals[f"{field}_delta"] = sum(row[f"{field}_delta"] for row in rows)
return {
"start_date": start_day.isoformat(),
"end_date": end_day.isoformat(),
"task_ids": scope,
"rows": rows,
"totals": totals,
"note_count": len(included_note_ids),
"has_partial_data": any(row["partial_metrics"] for row in rows),
"partial_metrics": sorted({field for row in rows for field in row["partial_metrics"]}),
"metric_labels": METRIC_LABELS,
}
def _count_by_day(timestamps: Iterable[Optional[int]]) -> Dict[date, int]:
counts: Dict[date, int] = {}
for ts in timestamps:
if ts is None:
continue
day = datetime.fromtimestamp(ts / 1000).date()
counts[day] = counts.get(day, 0) + 1
return counts
+372
View File
@@ -0,0 +1,372 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/runner.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Execute a single monitoring run: build the command, wait, then ingest.
Runs reuse ``CrawlerManager`` so that monitor crawls share the existing
single-subprocess guarantee and their logs stream to the existing Terminal
component over the existing log WebSocket.
"""
import asyncio
import os
from pathlib import Path
from typing import Iterable, List, Optional
from sqlalchemy import select
from tools.time_util import get_current_timestamp
from ..schemas import (
CrawlerStartRequest,
CrawlerTypeEnum,
LoginTypeEnum,
PlatformEnum,
SaveDataOptionEnum,
)
from ..services import crawler_manager
from . import adapters, app_settings, covers, douyin_fetch, notify
from .db import get_session
from .ingest import IngestResult, diagnose_failure, ingest_run
from .models import (
MODE_CREATOR,
RUN_FAILED,
RUN_PENDING,
RUN_RUNNING,
RUN_TIMEOUT,
MonitorNote,
MonitorRun,
MonitorTarget,
MonitorTask,
)
from .settings import get_cookie, mark_cookie_ok
PROJECT_ROOT = Path(__file__).parent.parent.parent
MONITOR_RUNS_DIR = PROJECT_ROOT / "data" / "monitor_runs"
# Monitor platform ids align with PlatformEnum's values, but mapping explicitly
# beats relying on that coincidence.
_PLATFORM_ENUM = {
"xhs": PlatformEnum.XHS,
"dy": PlatformEnum.DOUYIN,
"ks": PlatformEnum.KUAISHOU,
"bili": PlatformEnum.BILIBILI,
"wb": PlatformEnum.WEIBO,
"tieba": PlatformEnum.TIEBA,
"zhihu": PlatformEnum.ZHIHU,
}
# Timeout used when the caller does not care; tasks carry their own.
DEFAULT_RUN_TIMEOUT_SECONDS = 3600
def build_target_url(value: str, kind: str, platform: str) -> str:
"""Turn a stored target into a URL the crawler's parser accepts.
Always emits a full URL rather than a bare id: both platforms' parsers accept
a bare id only in a narrower form, so the URL is the safer universal input.
The shape itself is platform-specific and comes from ``adapters``.
"""
spec = adapters.adapter(platform)
path = spec.creator_path if kind == MODE_CREATOR else spec.note_path
return f"{spec.web_base}{path}/{value}"
def build_target_urls(
mode: str, targets: Iterable[MonitorTarget], platform: str
) -> List[str]:
"""存储的目标 -> 爬虫接受的 URL。
``xsec_token`` 只有小红书有,而且是会过期的刷新令牌 —— 有就带上,没有就算了。
抖音恒为空,所以这一段对它是天然的 no-op,不需要平台分支。
"""
urls = []
for target in targets:
url = build_target_url(target.external_id, target.kind, platform)
if target.xsec_token:
url = f"{url}?xsec_token={target.xsec_token}"
if target.xsec_source:
url = f"{url}&xsec_source={target.xsec_source}"
urls.append(url)
return urls
async def _strategy_settings(session, platform: str) -> dict:
"""Crawl-strategy and proxy settings for one platform.
Per-platform because the values genuinely differ: what is a safe request
interval on one site is a rate limit on another. Read per run rather than
cached, so a change takes effect on the next scheduled run.
"""
return {
"enable_sub_comments": bool(
await app_settings.get_value(session, "enable_sub_comments", platform, False)
),
"crawl_sleep_sec": int(
await app_settings.get_value(session, "crawl_sleep_sec", platform, 2)
),
"enable_ip_proxy": bool(
await app_settings.get_value(session, "enable_ip_proxy", platform, False)
),
"proxy_provider": await app_settings.get_value(
session, "proxy_provider", platform, "kuaidaili"
),
"proxy_pool_count": int(
await app_settings.get_value(session, "proxy_pool_count", platform, 2)
),
"static_proxy_url": await app_settings.get_value(
session, "static_proxy_url", platform, ""
),
}
def _write_cookie_file(path: Path, cookie: str) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(cookie, encoding="utf-8")
def _remove_cookie_file(path: Path) -> None:
"""Best-effort removal; the cookie is a credential, do not leave it around."""
try:
os.remove(path)
except OSError:
pass
async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
"""Run one monitoring cycle for ``task_id`` and ingest its output.
Split into three phases with separate short-lived DB sessions so no
transaction is held open across the multi-minute subprocess run.
"""
# --- Phase 1: book the run and work out where its output goes -------------
async with get_session() as session:
task = await session.get(MonitorTask, task_id)
if task is None:
raise ValueError(f"Monitor task {task_id} not found")
targets = [target for target in task.targets if target.enabled]
if not targets:
raise ValueError(f"Monitor task {task_id} has no enabled targets")
platform = task.platform
urls = build_target_urls(task.mode, targets, platform)
cookie = await get_cookie(session, platform)
strategy = await _strategy_settings(session, platform)
# System-wide switch. On a headless server the crawler must attach to the
# Chrome already listening on the debug port -- that browser is where the
# operator scanned the login QR, so its profile is the login. Default False
# keeps desktop runs launching a private browser exactly as before.
cdp_enabled = await app_settings.get_value(session, "cdp_enabled", fallback=False)
run = MonitorRun(
task_id=task.id,
trigger=trigger,
status=RUN_PENDING,
phase=task.mode,
save_data_path="",
queued_at=get_current_timestamp(),
not_before=0,
max_comments_count=task.max_comments_count if task.enable_comments else 0,
)
session.add(run)
await session.flush()
run_id = run.id
out_dir = MONITOR_RUNS_DIR / str(task.id) / str(run_id)
run.save_data_path = str(out_dir)
# Snapshot the values the subprocess needs; `task` is detached after commit.
mode = task.mode
enable_comments = task.enable_comments
max_notes_count = task.max_notes_count
max_comments_count = task.max_comments_count
timeout_seconds = task.run_timeout_seconds
# 抖音的作品列表接口被那道真校验挡着(见 douyin_fetch),拿不到列表时就靠这些
# 已知的 aweme_id 逐条刷新 —— 新作品发现不了,但已有作品的指标还能继续更新。
known_aweme_ids = list(
await session.scalars(
select(MonitorNote.note_id).where(MonitorNote.task_id == task.id)
)
)
# --- Phase 2: run the collection outside any transaction ------------------
# 抖音走**进程内 HTTP 客户端**,不起 Playwright 子进程:那边会构造一大串自相矛盾的
# 浏览器指纹参数(参数说 Mac + Chrome 125、UA 说 Linux + Chrome 155),网关回一个
# 200 + 空 body,然后被翻译成「account blocked」—— 看着像账号被封,其实什么都不是。
# 见 douyin_fetch / douyin_api。
# 先落「运行中」—— **两条路都要**。原先这一行只写在爬虫那条分支里,于是抖音那条路上
# run 一直停在 pending;一旦中途出事(异常、进程被重启),界面上就是一个永远
# 「排队中」的幽灵,而且 recover() 也只清理 running、收不到它。
async with get_session() as session:
run = await session.get(MonitorRun, run_id)
if run is not None:
run.status = RUN_RUNNING
run.started_at = get_current_timestamp()
in_process_tail: List[str] = []
if platform == adapters.PLATFORM_DY:
try:
# 也要有超时。爬虫那条路靠 run_and_wait(timeout=...) 兜底,这条路没有子进程、
# 没人管 —— 里面**任何一次卡住都会让 run 永远停在「运行中」**(真踩过:
# page.evaluate 打在一个卡死的标签页上不返回)。
fetched = await asyncio.wait_for(
douyin_fetch.collect(
out_dir,
platform=platform,
mode=mode,
limit=max_notes_count,
want_comments=enable_comments,
comment_limit=max_comments_count,
targets=targets,
known_aweme_ids=known_aweme_ids,
cookie=cookie,
),
timeout=timeout_seconds,
)
except asyncio.TimeoutError:
fetched = {
"notes": 0,
"comments": 0,
"errors": [f"抖音采集超过 {timeout_seconds} 秒仍未完成,已放弃这一轮"],
"jsonl_dir": "",
}
in_process_tail = list(fetched["errors"])
# 一条都没采到 = 这一轮失败,并把**真因**当作退出诊断传下去。否则它会掉进
# ingest 的「疑似登录失效」分支 —— 又骗人一次,正是这套东西一直在犯的毛病。
exit_code = 1 if (fetched["errors"] and not fetched["notes"]) else 0
if fetched["errors"] and fetched["notes"]:
# 有产物但带着错误,说明走了退化路径(比如作品列表被挡,只刷新了已知作品)。
# 这一轮状态是成功,但**不是**一切正常 —— 得留下痕迹,否则没人知道新作品
# 其实没在发现。
print(
"[monitor.runner] 抖音采集部分失败:"
+ ";".join(fetched["errors"])[:300]
)
else:
cookie_file = out_dir / ".cookies"
_write_cookie_file(cookie_file, cookie)
request = CrawlerStartRequest(
platform=_PLATFORM_ENUM[platform],
login_type=LoginTypeEnum.COOKIE,
crawler_type=CrawlerTypeEnum.CREATOR if mode == MODE_CREATOR else CrawlerTypeEnum.DETAIL,
creator_ids=",".join(urls) if mode == MODE_CREATOR else "",
specified_ids=",".join(urls) if mode != MODE_CREATOR else "",
start_page=1,
enable_comments=enable_comments,
enable_sub_comments=strategy["enable_sub_comments"],
enable_media=False,
save_option=SaveDataOptionEnum.JSONL,
cookies="",
headless=True,
max_notes_count=max_notes_count,
max_comments_count=max_comments_count,
# Isolate this run's output: the crawler names files by date only, so
# otherwise same-day runs would append into one shared file.
save_data_path=str(out_dir),
# Attach to the browser already running on CDP_DEBUG_PORT when the
# operator enabled it; otherwise launch a private, throwaway browser.
enable_cdp_mode=cdp_enabled,
# Only injecting web_session is not enough to sign requests from a cold
# browser profile.
inject_all_cookies=True,
save_login_state=True,
cookies_file=str(cookie_file),
max_concurrency_num=1,
# Strategy + proxy, surfaced on the Settings page.
crawler_max_sleep_sec=strategy["crawl_sleep_sec"],
enable_ip_proxy=strategy["enable_ip_proxy"],
ip_proxy_pool_count=strategy["proxy_pool_count"],
ip_proxy_provider_name=strategy["proxy_provider"],
static_proxy_url=strategy["static_proxy_url"] or None,
)
try:
exit_code = await crawler_manager.run_and_wait(request, timeout=timeout_seconds)
finally:
_remove_cookie_file(cookie_file)
# 失败时的诊断来源。抖音那条路没有子进程,尾巴就是它自己报的错 —— **别去读
# crawler_manager 的尾巴**,那里面是上一轮别的平台留下的东西,会张冠李戴。
output_tail = (
in_process_tail
if platform == adapters.PLATFORM_DY
else crawler_manager.get_output_tail()
)
# --- Phase 3: ingest ------------------------------------------------------
async with get_session() as session:
run = await session.get(MonitorRun, run_id)
task = await session.get(MonitorTask, task_id)
if run is None or task is None:
raise ValueError(f"Run {run_id} or task {task_id} vanished during execution")
# 目录名按平台解析 —— 抖音的平台 id 是 dy 而产物目录是 douyin,写死就永远判不准。
if exit_code == -1 and not (out_dir / adapters.artifact_dir(task.platform)).exists():
# run_and_wait returns -1 when the process could not start or timed out.
run.status = RUN_TIMEOUT
run.finished_at = get_current_timestamp()
run.exit_code = exit_code
run.error_message = "Run was killed by timeout or failed to start"
# -1 同时代表「超时」和「根本没起来」,两者要查的东西完全不同。输出尾巴里
# 有异常就带上它,否则运行历史里只能看到这句没有信息量的话。
cause = diagnose_failure(output_tail)
if cause:
run.error_message = f"{run.error_message};原因:{cause}"
result = IngestResult(status=RUN_TIMEOUT, error=run.error_message)
else:
run.exit_code = exit_code
run.finished_at = get_current_timestamp()
# 把爬虫输出的尾巴交给 ingest:退出码本身说明不了问题,运行历史里要显示的
# 是真正的报错(比如抖音的 DataFetchError: account blocked)。
result = await ingest_run(
session, run, task, out_dir, output_tail=output_tail
)
# A run that authenticated fine is the only useful signal that the
# stored cookie still works.
if result.notes_fetched > 0:
await mark_cookie_ok(session, task.platform)
task.last_run_at = run.finished_at
task.last_status = result.status
task.last_error = result.error
# --- Phase 4: notify ------------------------------------------------------
# Runs after the ingest transaction has committed, in its own session. A push
# failure must never roll back collected data, and notify_run() swallows its
# own errors for the same reason.
async with get_session() as session:
task = await session.get(MonitorTask, task_id)
run = await session.get(MonitorRun, run_id)
if task is not None and run is not None:
await notify.notify_run(session, task, run)
# --- Phase 5: 封面落盘 ------------------------------------------------------
# 也放在事务之外。封面地址带签名、会过期(实测隔天即 403),落盘之后才与签名无关。
# 下载慢且可能失败,占着一个入库事务是不合适的;失败也不影响本轮数据。
try:
async with get_session() as session:
saved = await covers.cache_pending(session, task_id)
if saved:
print(f"[monitor.runner] 缓存了 {saved} 张作品封面")
except Exception as exc: # noqa: BLE001 - 封面拿不到不该让整轮失败
print(f"[monitor.runner] 封面缓存失败:{exc}")
return result
+167
View File
@@ -0,0 +1,167 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/schedule.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Monitor task schedule arithmetic.
Three modes, all expressible by a picker. A raw cron string was ruled out on
purpose -- it is a small language, and the operator should not have to write one
to say "every day at nine":
* ``interval`` -- every N minutes.
* ``daily`` -- at chosen clock times, e.g. 09:00 and 18:30.
* ``weekly`` -- at chosen clock times on chosen weekdays, e.g. Mon-Fri 10:00.
The two clock modes are **fixed-time**, unlike ``interval``, which is fixed-delay.
The distinction matters: for an interval, measuring the next slot from when the
run starts is what stops a slow run from firing back-to-back. For a clock
schedule it would be wrong, because a run that starts at 09:07 would drag every
later run seven minutes late, compounding all day.
Fixed-time also means **no jitter is applied** to clock schedules. The operator
picked a time; quietly running at 09:04 instead of 09:00 is not a feature, it just
looks like a bug. ``interval`` keeps its jitter, where there is no stated time to
contradict.
All arithmetic is in the server's local timezone -- naive datetimes on purpose,
because the container is pinned to the operator's zone via TZ and pretending
otherwise would add a timezone concept nobody asked for.
"""
from datetime import datetime, time, timedelta
from typing import Optional, Sequence
MODE_INTERVAL = "interval"
MODE_DAILY = "daily"
MODE_WEEKLY = "weekly"
SCHEDULE_MODES = (MODE_INTERVAL, MODE_DAILY, MODE_WEEKLY)
CLOCK_MODES = (MODE_DAILY, MODE_WEEKLY)
_MS_PER_MINUTE = 60_000
_WEEKDAY_NAMES = "一二三四五六日"
def parse_hours(raw: Optional[str]) -> list[int]:
"""``"9,18"`` -> ``[9, 18]``. Sorted, de-duplicated, junk dropped."""
return _parse_int_list(raw, 0, 23)
def parse_days(raw: Optional[str]) -> list[int]:
"""``"0,2,4"`` -> ``[0, 2, 4]``. **0 is Monday**, matching ``date.weekday()``."""
return _parse_int_list(raw, 0, 6)
def _parse_int_list(raw: Optional[str], low: int, high: int) -> list[int]:
values: set[int] = set()
for chunk in (raw or "").split(","):
chunk = chunk.strip()
if not chunk:
continue
try:
number = int(chunk)
except ValueError:
# Stored values come from our own UI, but a hand-edited row must not
# be able to crash the scheduler loop.
continue
if low <= number <= high:
values.add(number)
return sorted(values)
def format_hours(hours: Sequence[int]) -> str:
return ",".join(str(hour) for hour in sorted(set(hours)))
def format_days(days: Sequence[int]) -> str:
return ",".join(str(day) for day in sorted(set(days)))
def describe(
*,
mode: str,
interval_minutes: int,
hours: Sequence[int],
days: Sequence[int],
minute: int,
) -> str:
"""One human sentence for the task card.
Lives here rather than in the frontend so the list view and the editor cannot
drift apart on what a schedule means.
"""
if mode == MODE_INTERVAL:
if interval_minutes % 1440 == 0:
return f"每 {interval_minutes // 1440} 天"
if interval_minutes % 60 == 0:
return f"每 {interval_minutes // 60} 小时"
return f"每 {interval_minutes} 分钟"
if not hours:
return "未设置时间"
clock = "、".join(f"{hour:02d}:{minute:02d}" for hour in sorted(set(hours)))
if mode == MODE_DAILY:
return f"每天 {clock}"
if not days:
return f"每天 {clock}"
labels = "、".join(f"周{_WEEKDAY_NAMES[day]}" for day in sorted(set(days)))
return f"{labels} {clock}"
def next_occurrence(
*,
mode: str,
interval_minutes: int,
hours: Sequence[int],
days: Sequence[int],
minute: int,
after_ms: int,
) -> Optional[int]:
"""The next fire time strictly after ``after_ms``, as epoch milliseconds.
``None`` means the schedule can never fire -- a clock mode with no hours
chosen. Callers store that as "no next run" rather than something in the past,
which would otherwise leave the task permanently due and re-running on every
tick.
"""
if mode == MODE_INTERVAL:
return after_ms + max(1, interval_minutes) * _MS_PER_MINUTE
if not hours:
return None
now = datetime.fromtimestamp(after_ms / 1000)
# No weekdays chosen means every day, matching describe(). Without the `days`
# guard an empty selection would produce an empty allowed set, no matching day,
# and a task that silently never runs.
allowed_days = set(days) if (mode == MODE_WEEKLY and days) else set(range(7))
# Eight days of lookahead covers today's remaining slots plus a full week,
# which is more than any weekday selection can need.
for offset in range(8):
day = (now + timedelta(days=offset)).date()
if day.weekday() not in allowed_days:
continue
for hour in sorted(set(hours)):
candidate = datetime.combine(day, time(hour=hour, minute=minute))
if candidate > now:
return int(candidate.timestamp() * 1000)
return None
+280
View File
@@ -0,0 +1,280 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/scheduler.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Background scheduler for monitor tasks.
One asyncio loop polls for due tasks and hands them to the runner. A plain loop
is enough here: there is exactly one process, one global crawler subprocess, and
therefore no concurrency to coordinate -- a cron-style library would add a
dependency without adding a capability.
Two families of schedule, and the difference matters:
* ``interval`` is **fixed-delay**, not fixed-rate -- ``next_run_at`` is measured
from the moment a run starts, so a slow run cannot make its task fire
back-to-back.
* the clock modes (``daily``/``weekly``) are **fixed-time** -- recomputed from the
calendar, so a run that starts late does not drag every later run with it.
The arithmetic for both lives in schedule.py.
The loop also carries the 上游更新检查: it is not a crawl, so it shares none of
the rules above (no subprocess, no active-hours gate) -- see
``_maybe_check_upstream``. It rides this loop rather than getting a thread of its
own because it is one HTTP-shaped fetch per day.
"""
import asyncio
import random
from datetime import datetime
from typing import Optional
from sqlalchemy import select
from tools.time_util import get_current_timestamp
from ..services import crawler_manager
from . import app_settings, schedule, upstream
from .db import get_session
from .models import (
RUN_INTERRUPTED,
RUN_PENDING,
RUN_RUNNING,
MonitorRun,
MonitorTask,
)
from .runner import execute_task
from .settings import get_cookie
POLL_INTERVAL_SECONDS = 20
# Spread tasks sharing an interval so they do not all come due on the same tick.
# Applied to interval mode only -- see the advance step below.
JITTER_SECONDS = 60
class MonitorScheduler:
"""Polls the task table and runs whatever is due."""
def __init__(self) -> None:
self._loop_task: Optional[asyncio.Task] = None
self._stopping = asyncio.Event()
# Avoids logging "no cookie" on every single tick. Per platform, because
# warning once for Xiaohongshu must not silence the warning for Douyin.
self._warned_no_cookie: set = set()
async def start(self) -> None:
if self._loop_task is not None and not self._loop_task.done():
return
self._stopping.clear()
self._loop_task = asyncio.create_task(self._run_loop())
async def stop(self) -> None:
self._stopping.set()
if self._loop_task is not None:
self._loop_task.cancel()
try:
await self._loop_task
except asyncio.CancelledError:
pass
self._loop_task = None
async def _run_loop(self) -> None:
try:
await self.recover()
except Exception as exc: # pragma: no cover - defensive
print(f"[monitor.scheduler] recovery failed: {exc}")
while not self._stopping.is_set():
try:
await self.tick()
except Exception as exc: # pragma: no cover - keep the loop alive
print(f"[monitor.scheduler] tick failed: {exc}")
# 独立于采集任务,因此单独一段 try:上游检查失败不该影响采集调度,
# 反过来也一样。
try:
await self._maybe_check_upstream()
except Exception as exc: # pragma: no cover - keep the loop alive
print(f"[monitor.scheduler] upstream check failed: {exc}")
await asyncio.sleep(POLL_INTERVAL_SECONDS)
async def _maybe_check_upstream(self) -> None:
"""到点就 fetch 一次上游仓库,看它有没有新提交。
与采集任务的三条规则都不同,各有理由:它不碰浏览器、也不占采集子进程,
所以不看 ``is_busy``;它只发一个 git 请求,没有被平台风控的风险,所以也不
受活跃时段限制 —— 定时检查放在半夜反而是最合适的。
"""
async with get_session() as session:
if not await app_settings.get_value(
session, "upstream_check_enabled", fallback=False
):
return
interval_minutes = int(
await app_settings.get_value(
session, "upstream_check_interval_minutes", fallback=1440
)
)
state = await upstream.load_state(session)
checked_at = int(state.get("checked_at") or 0)
now = get_current_timestamp()
# 失败也会写 checked_at,所以不通的时候同样是每个间隔重试一次,
# 而不是每个 tick(20 秒)都去撞一次墙。
if checked_at and now - checked_at < max(1, interval_minutes) * 60_000:
return
result = await upstream.run_check()
if result.get("behind"):
print(
f"[monitor.scheduler] 上游 {result.get('branch')} 领先 "
f"{result['behind']} 个提交"
)
elif not result.get("ok"):
print(f"[monitor.scheduler] 上游检查失败:{result.get('error')}")
async def recover(self) -> None:
"""Clean up state left behind by a server restart.
A run still marked ``running`` cannot be running -- its subprocess died
with the previous process. Marking it interrupted stops it from blocking
the UI as a phantom in-flight run.
**``pending`` 同样是残留**:那一行是上一轮建的,可它后面的采集根本没机会开始
(进程被重启,或者采集那条路抛了异常),所以它永远不会自己往前走。只清 running
的话,它会永远挂在界面上显示「排队中」—— 用户看到的就是任务卡住了。
"""
async with get_session() as session:
stale = (
await session.scalars(
select(MonitorRun).where(
MonitorRun.status.in_((RUN_RUNNING, RUN_PENDING))
)
)
).all()
for run in stale:
run.status = RUN_INTERRUPTED
run.finished_at = get_current_timestamp()
if stale:
print(
f"[monitor.scheduler] marked {len(stale)} interrupted run(s) "
f"left over from a previous process"
)
async def tick(self) -> None:
"""Run one due task, if the crawler is free and we are in the active window."""
# The crawler subprocess is a global singleton, so a manual crawl and a
# monitor run cannot overlap. Returning without advancing next_run_at
# leaves the task due, and it is picked up on a later tick.
if crawler_manager.is_busy():
return
async with get_session() as session:
if not await self._within_active_hours(session):
# Deliberately does not advance next_run_at: the task simply runs
# when the window next opens, rather than being skipped for a day.
return
await self._run_due_task()
async def _within_active_hours(self, session) -> bool:
"""Whether scheduled runs are allowed right now (local time)."""
start, end = await app_settings.active_hours(session)
hour = datetime.now().hour
if start <= end:
return start <= hour <= end
# Window wraps past midnight, e.g. 22 -> 6.
return hour >= start or hour <= end
async def _run_due_task(self) -> None:
async with get_session() as session:
task = await session.scalar(
select(MonitorTask)
.where(
MonitorTask.enabled.is_(True),
MonitorTask.next_run_at.is_not(None),
MonitorTask.next_run_at <= get_current_timestamp(),
)
.order_by(MonitorTask.next_run_at)
.limit(1)
)
if task is None:
return
# 没有 cookie 就跳过,是为了不让任务每轮白跑一趟出个认证失败。任务留在
# due 状态而不推进 —— 用户一粘上 cookie 它就能自己跑起来。
#
# **但开着 CDP 时必须放行**:那种模式下登录态来自被接管的那个浏览器,
# 粘不粘 cookie 根本轮不到它决定成败。不放行的话,选了「接管已有 Chrome」
# 却没粘 cookie 的用户会发现任务永远不被触发,而且什么错都不报。
cookie = await get_cookie(session, task.platform)
if not cookie:
cdp_enabled = await app_settings.get_value(
session, "cdp_enabled", fallback=False
)
if not cdp_enabled:
if task.platform not in self._warned_no_cookie:
print(
f"[monitor.scheduler] no {task.platform} cookie configured; "
"scheduled tasks will not run until one is set or CDP is enabled"
)
self._warned_no_cookie.add(task.platform)
return
self._warned_no_cookie.discard(task.platform)
# Advance before running so a crash mid-run cannot cause an immediate
# re-fire, and so a long outage coalesces into a single run instead
# of one run per missed interval.
now = get_current_timestamp()
following = schedule.next_occurrence(
mode=task.schedule_mode,
interval_minutes=task.interval_minutes,
hours=schedule.parse_hours(task.schedule_hours),
days=schedule.parse_days(task.schedule_days),
minute=task.schedule_minute,
after_ms=now,
)
if following is None:
# A clock schedule with no times can never fire. The API rejects
# that shape, so this guards against a hand-edited row: park the
# task with no next run rather than leaving it permanently due and
# re-running it on every tick.
task.next_run_at = None
print(
f"[monitor.scheduler] task {task.id} has no usable schedule "
f"and will not run until one is set"
)
elif task.schedule_mode == schedule.MODE_INTERVAL:
# Jitter belongs to the interval mode only. Spreading identical
# intervals apart is the point; nudging a time the operator
# explicitly picked is not -- it just looks like a broken clock.
task.next_run_at = following + random.randint(0, JITTER_SECONDS) * 1000
else:
task.next_run_at = following
task_id = task.id
try:
await execute_task(task_id, trigger="scheduled")
except Exception as exc:
print(f"[monitor.scheduler] task {task_id} failed: {exc}")
# Global singleton, mirroring the crawler_manager pattern.
monitor_scheduler = MonitorScheduler()
File diff suppressed because it is too large Load Diff
+108
View File
@@ -0,0 +1,108 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/settings.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Key/value settings for the monitoring layer, plus cookie health helpers.
The XHS cookie is what makes scheduled runs unattended. It expires every few
weeks, so alongside the value we track when it was last seen working -- that is
what lets the UI warn before a task silently stops collecting.
"""
from typing import Optional
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from .models import MonitorSetting
from .platforms import PLATFORM_XHS
def platform_key(platform: str, name: str) -> str:
"""Key for a setting that each platform keeps its own copy of."""
return f"platform.{platform}.{name}"
def system_key(name: str) -> str:
"""Key for a setting shared across every platform."""
return f"system.{name}"
def cookie_key(platform: str) -> str:
return platform_key(platform, "cookie")
def cookie_updated_key(platform: str) -> str:
return platform_key(platform, "cookie_updated_at")
def cookie_last_ok_key(platform: str) -> str:
return platform_key(platform, "cookie_last_ok_at")
async def get_setting(session: AsyncSession, key: str) -> Optional[str]:
return await session.scalar(select(MonitorSetting.value).where(MonitorSetting.key == key))
async def set_setting(session: AsyncSession, key: str, value: str) -> None:
row = await session.get(MonitorSetting, key)
now = get_current_timestamp()
if row is None:
session.add(MonitorSetting(key=key, value=value, updated_at=now))
else:
row.value = value
row.updated_at = now
async def delete_setting(session: AsyncSession, key: str) -> None:
row = await session.get(MonitorSetting, key)
if row is not None:
await session.delete(row)
async def get_cookie(session: AsyncSession, platform: str = PLATFORM_XHS) -> str:
return (await get_setting(session, cookie_key(platform))) or ""
async def set_cookie(
session: AsyncSession, cookie: str, platform: str = PLATFORM_XHS
) -> None:
await set_setting(session, cookie_key(platform), cookie)
await set_setting(session, cookie_updated_key(platform), str(get_current_timestamp()))
async def mark_cookie_ok(session: AsyncSession, platform: str = PLATFORM_XHS) -> None:
"""Record that a run authenticated successfully."""
await set_setting(session, cookie_last_ok_key(platform), str(get_current_timestamp()))
async def get_cookie_status(session: AsyncSession, platform: str = PLATFORM_XHS) -> dict:
"""Cookie health for the UI. Never returns the cookie value itself."""
cookie = await get_cookie(session, platform)
updated_at = await get_setting(session, cookie_updated_key(platform))
last_ok_at = await get_setting(session, cookie_last_ok_key(platform))
return {
"platform": platform,
"present": bool(cookie),
# Enough to eyeball whether the pasted value looks right, not enough to leak it.
"length": len(cookie),
"updated_at": int(updated_at) if updated_at else None,
"last_ok_at": int(last_ok_at) if last_ok_at else None,
}
+355
View File
@@ -0,0 +1,355 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/upstream.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""上游仓库更新检查。
本仓库在上游(NanmiCoder/MediaCrawler)之上加了一整层监控/鉴权/多平台面板,
合并流程写在 UPSTREAM.md 里。但那份流程默认**有人知道上游动了**——而部署脚本是
`git pull --ff-only`,只从我们自己的 Gitea 拉,上游的提交不主动去 fetch 就永远
看不见。拖着不合并的代价是复利的:越久越难合,最后只能放弃。这个模块把「上游动
了没有」变成一条可定时、会推到企业微信的通知。
三处刻意的取舍:
* **用 git 而不是托管商的 HTTP API。** 只有 git 算得出「落后几个提交」:托管商
API 能告诉你上游 tip 是什么,但它不知道我们与上游的共同祖先在哪,而分歧点恰恰
是真正要合的东西。本仓库还含有上游没有的提交,直接比 tip 会得出错误的结论。
* **按 URL fetch 到 FETCH_HEAD,不配置 remote、不写 refs/remotes。** 服务器上的
checkout 是从 Gitea 克隆的,本来就没有 upstream 这个 remote;用 URL 直取就不必
先去改它的 git 配置。顺带也避免往别人的部署里塞一个 remote。
* **只读不写工作区。** fetch 只落对象和 FETCH_HEAD,不碰索引与工作区,所以不会打断
正在跑的采集,也不会和 `./deploy.sh` 的 git pull 抢锁。
依赖一个外部命令:**git**。本机开发环境一定有;容器里是 Dockerfile 显式装的
(python:3.11-slim 默认不带)。
"""
import asyncio
import json
import os
import subprocess
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Dict, List, Tuple
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from .db import get_session
from .models import SETTING_UPSTREAM_NOTIFIED_TIP, SETTING_UPSTREAM_STATE
from .settings import get_setting, set_setting
PROJECT_ROOT = Path(__file__).parent.parent.parent
# 默认就是本仓库跟踪的那个上游。国内直连 GitHub 不稳时改成 gitcode 镜像即可
# (见 UPSTREAM.md「直连 GitHub 不通时」)。
DEFAULT_REMOTE_URL = "https://github.com/NanmiCoder/MediaCrawler.git"
DEFAULT_BRANCH = "main"
# fetch 要走网络,给宽松些;其余全是本地命令,慢到这个程度只能说明仓库坏了。
FETCH_TIMEOUT_SECONDS = 120
LOCAL_TIMEOUT_SECONDS = 20
# 通知里最多列几条提交。要传达的是「该动手了」,不是把 changelog 搬到群里。
MAX_LISTED_COMMITS = 10
# 状态里留几条给前端展示。比通知多留一些,界面上能看到更完整的列表。
MAX_STORED_COMMITS = 30
# git log 用 Unit Separator 分隔字段:它不可能出现在提交信息里,比制表符安全。
_RECORD_SEPARATOR = "\x1f"
_LOG_FORMAT = (
f"%h{_RECORD_SEPARATOR}%an{_RECORD_SEPARATOR}%ad{_RECORD_SEPARATOR}%s"
)
class GitError(RuntimeError):
"""git 不可用,或某条 git 命令失败了。"""
@dataclass(frozen=True)
class Commit:
"""一条上游提交,只留通知/展示需要的四个字段。"""
sha: str
author: str
date: str
subject: str
@dataclass
class CheckResult:
"""一次检查的结论。失败也是一种结论,用 ``ok``/``error`` 表达而不是抛异常。"""
ok: bool
# HEAD..FETCH_HEAD:上游有而我们没有的提交数 —— 要合的就是这些。
behind: int = 0
# FETCH_HEAD..HEAD:我们有自己的提交数 —— 也就是这一层的规模。
ahead: int = 0
tip: str = ""
head: str = ""
commits: List[Commit] = field(default_factory=list)
error: str = ""
def as_dict(self) -> Dict[str, Any]:
return {
"ok": self.ok,
"behind": self.behind,
"ahead": self.ahead,
"tip": self.tip,
"head": self.head,
"commits": [commit.__dict__ for commit in self.commits],
"error": self.error,
}
def _git(args: List[str], timeout: int) -> subprocess.CompletedProcess:
env = dict(os.environ)
# 远端要凭据时(地址写成了私有仓库),git 会停下来问密码,而这里没有终端可问,
# 于是挂到超时。关掉一切交互,让它立刻失败。
env["GIT_TERMINAL_PROMPT"] = "0"
env["GIT_ASKPASS"] = ""
env["SSH_ASKPASS"] = ""
# 容器里 uid 1000 没有 passwd 项,git 找不到 HOME 会抱怨。给一个存在且可写的。
env.setdefault("HOME", "/tmp")
return subprocess.run(
[
"git",
"-C",
str(PROJECT_ROOT),
# 只对自己这个 checkout 放行所有权检查。容器里 uid 一般与属主一致,
# 但 bind mount 的属主未必,一旦不一致 git 会直接拒绝干任何活。
"-c",
f"safe.directory={PROJECT_ROOT}",
# 忽略任何全局凭据助手:这是个只读的公开仓库,不该去翻钥匙串。
"-c",
"credential.helper=",
*args,
],
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
timeout=timeout,
env=env,
)
def _run(args: List[str], timeout: int) -> Tuple[int, str, str]:
"""跑一条 git 命令,返回 (returncode, stdout, stderr)。
只把「跑不起来」当异常;命令返回非零是正常结果,交给调用方处理。
"""
try:
proc = _git(args, timeout)
except FileNotFoundError as exc:
raise GitError("未找到 git 命令,请先安装 git") from exc
except subprocess.TimeoutExpired as exc:
raise GitError(f"git {args[0]} 超时({timeout} 秒)") from exc
return proc.returncode, (proc.stdout or "").strip(), (proc.stderr or "").strip()
def _require(args: List[str], timeout: int, what: str) -> str:
code, out, err = _run(args, timeout)
if code != 0:
# git 的报错通常是多行的,只留第一行;完整输出塞进日志反而更难读。
detail = err.splitlines()[0].strip() if err else "未知错误"
raise GitError(f"{what}:{detail}")
return out
def _to_int(raw: str) -> int:
try:
return int(raw)
except (TypeError, ValueError):
return 0
def _parse_log(raw: str) -> List[Commit]:
commits: List[Commit] = []
for line in raw.splitlines():
parts = line.split(_RECORD_SEPARATOR)
if len(parts) != 4:
# 格式不对就跳过这一条:一条读不出来的提交不该让整次检查失败。
continue
sha, author, date, subject = parts
commits.append(Commit(sha=sha, author=author, date=date, subject=subject))
return commits
def _check_sync(remote_url: str, branch: str) -> CheckResult:
"""阻塞实现,异步包装见 :func:`check`。"""
# 在 try 之前绑定:后面的失败结果也带上它 —— 「检查失败」时当前跑的是哪个
# 提交,正是排查时第一个想知道的。
head = ""
try:
head = _require(["rev-parse", "HEAD"], LOCAL_TIMEOUT_SECONDS, "读取本地 HEAD 失败")
# 增量 fetch:对象本地基本都已经有了,所以正常情况下只传几个新提交,
# 不会遇到 UPSTREAM.md 里说的「大包必断」。
_require(
["fetch", "--no-tags", remote_url, branch],
FETCH_TIMEOUT_SECONDS,
"从上游 fetch 失败",
)
tip = _require(["rev-parse", "FETCH_HEAD"], LOCAL_TIMEOUT_SECONDS, "读不到 FETCH_HEAD")
behind = _to_int(
_require(
["rev-list", "--count", "HEAD..FETCH_HEAD"],
LOCAL_TIMEOUT_SECONDS,
"统计落后提交数失败",
)
)
ahead = _to_int(
_require(
["rev-list", "--count", "FETCH_HEAD..HEAD"],
LOCAL_TIMEOUT_SECONDS,
"统计领先提交数失败",
)
)
# 只在确实落后时才读提交列表:已经是最新时这条 git log 毫无意义。
raw_log = ""
if behind:
raw_log = _require(
[
"log",
f"--max-count={MAX_STORED_COMMITS}",
"--date=short",
f"--format={_LOG_FORMAT}",
"HEAD..FETCH_HEAD",
],
LOCAL_TIMEOUT_SECONDS,
"读取新提交列表失败",
)
except GitError as exc:
return CheckResult(ok=False, head=head, error=str(exc))
return CheckResult(
ok=True,
behind=behind,
ahead=ahead,
tip=tip,
head=head,
commits=_parse_log(raw_log),
)
async def check(
remote_url: str = DEFAULT_REMOTE_URL, branch: str = DEFAULT_BRANCH
) -> CheckResult:
"""跑一次检查。网络与子进程都丢进线程,事件循环不被阻塞。
不抛异常:上游不通是常态(尤其是直连 GitHub),那也是一种要记录下来的结果。
"""
return await asyncio.to_thread(_check_sync, remote_url, branch)
def build_message(result: CheckResult, branch: str) -> str:
"""把一次「上游有新提交」的结果写成一条企业微信 markdown。"""
lines = [
"**🔔 上游 MediaCrawler 有更新**",
f"> 当前部署落后 `{branch}` **{result.behind}** 个提交",
]
if result.ahead:
lines.append(f"> (本仓库另有 {result.ahead} 个自己的提交,合并时注意保留)")
for commit in result.commits[:MAX_LISTED_COMMITS]:
lines.append(f"> `{commit.sha}` {commit.subject}")
# 落后数可能大于列出来的条数:状态里留的提交本身也是截断的(30 条),
# 所以这里比的是总数,不是 len(commits)。
if result.behind > MAX_LISTED_COMMITS:
lines.append(f"> …等共 {result.behind} 个提交")
lines.append("> 合并步骤见仓库根目录 `UPSTREAM.md`")
return "\n".join(lines)
async def load_state(session: AsyncSession) -> Dict[str, Any]:
"""最近一次检查的结果,从设置里读回来。没查过时是空字典。"""
raw = await get_setting(session, SETTING_UPSTREAM_STATE)
if not raw:
return {}
try:
state = json.loads(raw)
except json.JSONDecodeError:
# 手改坏了的行不该让接口 500,当作「没查过」即可。
return {}
return state if isinstance(state, dict) else {}
async def _save_state(session: AsyncSession, state: Dict[str, Any]) -> None:
# 整体读写,所以存成一条 JSON:拆成多个 key 只会带来写到一半的不一致。
await set_setting(session, SETTING_UPSTREAM_STATE, json.dumps(state, ensure_ascii=False))
async def run_check(notify_when_new: bool = True) -> Dict[str, Any]:
"""检查一次,落库,必要时推送。返回值可直接交给前端。
分三段各自的数据库会话:fetch 最长可能跑满两分钟,占着一个连接不合适 ——
理由与 runner.py 的分段完全相同。
"""
# 延迟导入:app_settings 在模块级 import 本模块(为了那个默认地址常量),
# 模块级反向 import 会成环。
from . import app_settings, notify
async with get_session() as session:
remote_url = str(
await app_settings.get_value(
session, "upstream_remote_url", fallback=DEFAULT_REMOTE_URL
)
or DEFAULT_REMOTE_URL
)
branch = str(
await app_settings.get_value(session, "upstream_branch", fallback=DEFAULT_BRANCH)
or DEFAULT_BRANCH
)
notify_enabled = bool(
await app_settings.get_value(session, "upstream_notify", fallback=True)
)
notified_tip = (await get_setting(session, SETTING_UPSTREAM_NOTIFIED_TIP)) or ""
webhook_url = await notify.get_webhook_url(session)
result = await check(remote_url, branch)
payload: Dict[str, Any] = {
"checked_at": get_current_timestamp(),
"remote_url": remote_url,
"branch": branch,
**result.as_dict(),
}
async with get_session() as session:
await _save_state(session, payload)
has_update = result.ok and result.behind > 0 and bool(result.tip)
# 同一个 tip 只推一次:否则每过一个检查周期就把同样的更新推到群里,
# 直到有人去合为止。上游真又动了(tip 变了)时应该再推。
if (
notify_when_new
and notify_enabled
and has_update
and webhook_url
and result.tip != notified_tip
):
ok, detail = await notify.send_wecom(webhook_url, build_message(result, branch))
if ok:
await set_setting(session, SETTING_UPSTREAM_NOTIFIED_TIP, result.tip)
payload["notified"] = True
else:
# 推送失败不该抹掉检查结果 —— 界面上仍然能看到「落后几个提交」。
payload["notify_error"] = detail
return payload
+13 -1
View File
@@ -16,8 +16,20 @@
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
from .auth import router as auth_router
from .crawler import router as crawler_router
from .creator import router as creator_router
from .data import router as data_router
from .monitor import router as monitor_router
from .settings import router as settings_router
from .websocket import router as websocket_router
__all__ = ["crawler_router", "data_router", "websocket_router"]
__all__ = [
"auth_router",
"crawler_router",
"creator_router",
"data_router",
"monitor_router",
"settings_router",
"websocket_router",
]
+173
View File
@@ -0,0 +1,173 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/auth.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Login / logout endpoints.
Deliberately exempt from ``require_auth``:
* ``/login`` -- it is the way in.
* ``/logout`` -- exempt so an already-expired session still gets a clean 200
and a cleared cookie instead of a confusing 401, which
would leave the browser holding a stale cookie.
"""
from fastapi import APIRouter, Depends, HTTPException, Request, Response, status
from ..auth import (
INVALID_CREDENTIALS,
SESSION_COOKIE_NAME,
check_password,
require_auth,
clear_failures,
client_key,
cookie_secure,
create_session,
purge_expired_sessions,
record_failure,
resolve_session,
retry_after_seconds,
revoke_all_sessions,
revoke_session,
set_password,
token_from_request,
)
from ..monitor.db import get_session
from ..schemas.auth import ChangePasswordPayload, LoginPayload
from tools.time_util import get_current_timestamp
router = APIRouter(prefix="/auth", tags=["auth"])
def _apply_session_cookie(response: Response, token: str, expires_at: int) -> None:
"""Attach the session cookie.
``secure`` is off by default because the panel is served over plain HTTP on
a LAN; setting it there means the browser silently discards the cookie and
the login page just loops with no error. ``SameSite=lax`` is also what
blocks cross-site POSTs, i.e. the CSRF defence for the write endpoints.
"""
max_age = max((expires_at - get_current_timestamp()) // 1000, 60)
response.set_cookie(
key=SESSION_COOKIE_NAME,
value=token,
max_age=max_age,
httponly=True,
secure=cookie_secure(),
samesite="lax",
path="/",
)
@router.post("/login")
async def login(payload: LoginPayload, request: Request, response: Response):
key = client_key(request)
wait = await retry_after_seconds(key)
if wait:
raise HTTPException(
status_code=status.HTTP_429_TOO_MANY_REQUESTS,
detail=f"尝试过于频繁,请 {wait} 秒后再试",
headers={"Retry-After": str(wait)},
)
async with get_session() as session:
valid = await check_password(session, payload.password)
token = ""
expires_at = 0
if valid:
await purge_expired_sessions(session)
token, expires_at = await create_session(session)
if not valid:
await record_failure(key)
# One generic message regardless of whether the password was wrong,
# empty, or simply not set yet -- no oracle.
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED, detail=INVALID_CREDENTIALS
)
await clear_failures(key)
_apply_session_cookie(response, token, expires_at)
return {"expires_at": expires_at}
@router.post("/logout")
async def logout(request: Request, response: Response):
token = token_from_request(request)
if token:
async with get_session() as session:
await revoke_session(session, token)
response.delete_cookie(SESSION_COOKIE_NAME, path="/")
return {"message": "已退出登录"}
@router.get("/me")
async def me(request: Request):
"""Identity probe. The SPA treats a 401 here as "show the login page".
Does its own resolution rather than using ``require_auth`` so it can also
report the expiry.
"""
token = token_from_request(request)
if not token:
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED, detail=INVALID_CREDENTIALS
)
async with get_session() as session:
row = await resolve_session(session, token)
if row is None:
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED, detail=INVALID_CREDENTIALS
)
return {"authenticated": True, "expires_at": row.expires_at}
# Auth as a route dependency, not only inside the handler: FastAPI validates the
# request body before the endpoint body runs, so an unauthenticated caller would
# otherwise get a 422 that confirms the endpoint and its schema exist.
@router.post("/password", dependencies=[Depends(require_auth)])
async def change_password(
payload: ChangePasswordPayload, request: Request, response: Response
):
"""Change the password and log every device out.
Revoking all sessions is the point: a password change is usually a response
to suspicion, and leaving other sessions alive would defeat it.
"""
token = token_from_request(request)
async with get_session() as session:
current = await resolve_session(session, token)
if current is None:
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED, detail=INVALID_CREDENTIALS
)
if not await check_password(session, payload.current):
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED, detail="当前密码不正确"
)
await set_password(session, payload.new)
await revoke_all_sessions(session)
# Issue a fresh session so the caller is not bounced mid-use.
new_token, expires_at = await create_session(session)
_apply_session_cookie(response, new_token, expires_at)
return {"message": "密码已更新,其他设备的登录已全部失效", "expires_at": expires_at}
+25 -2
View File
@@ -18,7 +18,9 @@
from fastapi import APIRouter, HTTPException
from ..schemas import CrawlerStartRequest, CrawlerStatusResponse
from ..monitor.db import get_session
from ..monitor.settings import get_cookie
from ..schemas import CrawlerStartRequest, CrawlerStatusResponse, LoginTypeEnum
from ..services import crawler_manager
router = APIRouter(prefix="/crawler", tags=["crawler"])
@@ -26,7 +28,28 @@ router = APIRouter(prefix="/crawler", tags=["crawler"])
@router.post("/start")
async def start_crawler(request: CrawlerStartRequest):
"""Start crawler task"""
"""Start crawler task.
A cookie login with no cookie supplied falls back to the one stored for the
selected platform. The manual crawl and the monitor therefore share a single
credential; keeping a second paste field on the crawl page meant it went
stale and could silently disagree with the monitor's.
"""
if (
request.login_type == LoginTypeEnum.COOKIE
and not request.cookies
and not request.cookies_file
):
async with get_session() as session:
stored = await get_cookie(session, request.platform.value)
if not stored:
raise HTTPException(
status_code=400,
detail="该平台尚未保存 Cookie,请到「设置 → 登录态」配置,或改用扫码登录",
)
request.cookies = stored
success = await crawler_manager.start(request)
if not success:
# Handle concurrent/duplicate requests: if process is already running, return 400 instead of 500
+150
View File
@@ -0,0 +1,150 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/creator.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营模块的 HTTP 接口。"""
import asyncio
from typing import Set
from fastapi import APIRouter, HTTPException, Query
from ..creator import login as creator_login
from ..creator import service
from ..monitor.db import get_session
router = APIRouter(prefix="/creator", tags=["creator"])
# 后台同步任务要留强引用:asyncio 只持弱引用,否则任务可能在跑完前被回收。
_sync_tasks: Set[asyncio.Task] = set()
def _bad_request(exc: ValueError) -> HTTPException:
return HTTPException(status_code=400, detail=str(exc))
@router.get("/accounts")
async def list_accounts():
"""账号列表。**不含 cookie**,只给 ``has_cookie``。"""
async with get_session() as session:
return {"accounts": await service.list_accounts(session)}
@router.get("/accounts/{account_id}")
async def get_account_detail(account_id: int):
async with get_session() as session:
try:
return await service.account_detail(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
@router.delete("/accounts/{account_id}")
async def delete_account(account_id: int):
async with get_session() as session:
try:
await service.delete_account(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
return {"message": "账号已删除"}
@router.post("/accounts/{account_id}/check")
async def check_account(account_id: int):
"""重测登录态与数据权限。"""
async with get_session() as session:
try:
return await service.check_account(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
@router.post("/accounts/{account_id}/sync")
async def sync_account(
account_id: int, days: int = Query(default=90, ge=1, le=730)
):
"""拉取作品数据。
**放后台跑**:要分页、还要按账号节流,几分钟很正常,而前端请求超时是 30 秒。
前端靠轮询账号列表里的 `last_synced_at` / `last_error` 看结果。
"""
async with get_session() as session:
try:
await service.get_account(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
task = asyncio.create_task(_sync_in_background(account_id, days))
_sync_tasks.add(task)
task.add_done_callback(_sync_tasks.discard)
return {"message": "同步已开始", "days": days}
async def _sync_in_background(account_id: int, days: int) -> None:
try:
async with get_session() as session:
result = await service.sync_account(session, account_id, days)
print(f"[creator] 账号 {account_id} 同步完成,取回 {result['fetched']} 条")
except Exception as exc: # noqa: BLE001 - 后台任务不能让异常逃逸成静默失败
print(f"[creator] 账号 {account_id} 同步失败: {exc}")
# ---------------------------------------------------------------------------
# 扫码新增账号
# ---------------------------------------------------------------------------
#
# 每次登录开一个**临时浏览器上下文**,扫完取出 cookie 就丢弃 —— 这样登第二个账号
# 不会把第一个顶掉,也不影响监控那个登录态。cookie 只在内存里从 login 模块传到
# 这里落库,**不进响应体**。
@router.post("/login")
async def start_login():
try:
return await creator_login.start()
except RuntimeError as exc:
raise HTTPException(status_code=502, detail=str(exc))
@router.get("/login")
async def poll_login():
"""轮询扫码结果;一旦成功就把账号落库并返回它。"""
snapshot = await creator_login.status()
if snapshot["status"] == creator_login.STATUS_SUCCESS:
# take_cookie 只在会话还在时返回 cookie,取走即拆会话;重复轮询拿到 None
# 就说明已经保存过了,直接返回上次的结果,不要退回 idle。
cookie = await creator_login.take_cookie()
if cookie:
try:
async with get_session() as session:
account = await service.upsert_account_from_cookie(session, cookie)
except ValueError as exc:
snapshot["status"] = creator_login.STATUS_ERROR
snapshot["message"] = f"扫码成功但保存账号失败:{exc}"
await creator_login.remember_result(snapshot)
return snapshot
snapshot["account"] = account
snapshot["message"] = f"已添加账号:{account['nickname']}"
await creator_login.remember_result(snapshot)
return snapshot
@router.delete("/login")
async def cancel_login():
return await creator_login.cancel()
+723
View File
@@ -0,0 +1,723 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/monitor.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""HTTP API for scheduled monitoring tasks."""
from datetime import date, datetime, timedelta
from typing import Any, Dict, List, Optional
from fastapi import APIRouter, HTTPException, Query, Response
from fastapi.responses import FileResponse
from ..monitor import covers, notify, qrlogin, report, service, upstream
from ..monitor.db import get_session
from ..monitor.platforms import PLATFORM_XHS
from ..monitor.settings import (
cookie_key,
delete_setting,
get_cookie_status,
get_setting,
set_cookie,
set_setting,
)
from ..monitor.models import SETTING_WECOM_WEBHOOK, MonitorTask
from ..schemas.monitor import (
CookiePayload,
CreatorAliasPayload,
NoteAliasPayload,
MonitorTaskCreate,
MonitorTaskUpdate,
WebhookPayload,
WebhookTestPayload,
)
router = APIRouter(prefix="/monitor", tags=["monitor"])
@router.get("/overview")
async def get_overview(platform: Optional[str] = None):
"""Headline numbers for the dashboard tiles, scoped to one platform."""
async with get_session() as session:
return await service.overview(session, platform)
# ---------------------------------------------------------------------------
# Tasks
# ---------------------------------------------------------------------------
@router.get("/tasks")
async def list_tasks(platform: Optional[str] = None):
async with get_session() as session:
return {"tasks": await service.list_tasks(session, platform)}
@router.post("/tasks", status_code=201)
async def create_task(payload: MonitorTaskCreate):
async with get_session() as session:
try:
task = await service.create_task(session, payload.model_dump())
except service.TargetParseError as exc:
raise HTTPException(status_code=400, detail=str(exc))
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc))
return {"id": task.id, "message": "Monitoring task created"}
@router.patch("/tasks/{task_id}")
async def update_task(task_id: int, payload: MonitorTaskUpdate):
async with get_session() as session:
try:
await service.update_task(session, task_id, payload.model_dump(exclude_unset=True))
except service.TargetParseError as exc:
raise HTTPException(status_code=400, detail=str(exc))
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
return {"message": "Monitoring task updated"}
@router.delete("/tasks/{task_id}")
async def delete_task(task_id: int):
async with get_session() as session:
try:
await service.delete_task(session, task_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
return {"message": "Monitoring task deleted"}
@router.post("/tasks/{task_id}/run")
async def run_task_now(task_id: int):
"""Queue a run immediately and return; the crawl itself takes minutes."""
async with get_session() as session:
task = await session.get(MonitorTask, task_id)
if task is None:
raise HTTPException(status_code=404, detail=f"Task {task_id} not found")
if not any(target.enabled for target in task.targets):
raise HTTPException(status_code=400, detail="Task has no enabled targets")
service.trigger_manual_run(task_id)
return {"message": "Run queued"}
@router.get("/tasks/{task_id}/runs")
async def list_runs(task_id: int, limit: int = Query(default=50, ge=1, le=500)):
async with get_session() as session:
return {"runs": await service.list_runs(session, task_id, limit=limit)}
# ---------------------------------------------------------------------------
# Collected data
# ---------------------------------------------------------------------------
@router.get("/notes")
async def list_notes(
task_id: Optional[int] = None,
only_new: bool = False,
limit: int = Query(default=200, ge=1, le=2000),
platform: Optional[str] = None,
):
async with get_session() as session:
return {
"notes": await service.list_notes(session, task_id, only_new, limit, platform),
# 博主**单独给一份**,而不是让前端从作品里推。作品推不出「一条作品都没有的
# 博主」—— 那正是最该显示的一类(还在涨粉,只是最近没发)。
"creators": await service.list_creators(session, task_id, platform),
}
@router.get("/notes/{note_id}/series")
async def note_series(note_id: str, task_id: Optional[int] = None):
"""Metric time series for a single note."""
async with get_session() as session:
return {"series": await service.note_series(session, note_id, task_id)}
@router.get("/comments")
async def list_comments(
task_id: Optional[int] = None,
note_id: Optional[str] = None,
group_by: Optional[str] = Query(
default=None, description="传 note 则按作品分组返回,便于阅读"
),
limit: int = Query(default=200, ge=1, le=2000),
platform: Optional[str] = None,
):
"""Comments, each carrying the work it belongs to.
``note_id`` filters to one work; ``group_by=note`` returns them bucketed per
work instead of as a flat stream.
"""
async with get_session() as session:
comments = await service.list_comments(session, task_id, note_id, limit, platform)
if group_by != "note":
return {"comments": comments, "total": len(comments)}
buckets: Dict[str, Dict[str, Any]] = {}
for comment in comments:
bucket = buckets.setdefault(
comment["note_id"],
{
"note_id": comment["note_id"],
"note_title": comment["note_title"],
"note_cover": comment["note_cover"],
"note_url": comment["note_url"],
# 作品所属的创作者。评论流按 博主 → 作品 → 评论 三级展开时,最外层
# 就是按这两个字段分组的 —— 少了它们,前端拿到的是 undefined,
# 于是所有博主塌成同一个分组、标签回退成「未知博主」。
# 同一个桶里的评论必然同属一个作品,所以取哪一条都一样。
"creator_hash": comment["note_creator_hash"],
"creator_name": comment["note_creator_name"],
# 作品的发布时间。评论流按 博主 → 作品 → 评论 展开时,作品那一层
# 光有标题不够 —— 同名作品不少,日期能帮着认。
"published_at": comment["note_published_at"],
"comments": [],
},
)
bucket["comments"].append(comment)
ordered = sorted(
buckets.values(),
key=lambda group: group["comments"][0]["first_seen_at"],
reverse=True,
)
return {"groups": ordered, "total": len(comments)}
@router.get("/comment-notes")
async def list_comment_notes(task_id: Optional[int] = None, platform: Optional[str] = None):
"""Works that have comments, newest first, with counts.
Feeds the comment filter dropdown so the operator can pick by title.
"""
async with get_session() as session:
return {"notes": await service.comment_note_groups(session, task_id, platform)}
@router.get("/events")
async def list_events(
task_id: Optional[int] = None,
type: Optional[str] = None,
since_id: Optional[int] = None,
limit: int = Query(default=200, ge=1, le=2000),
platform: Optional[str] = None,
):
async with get_session() as session:
events = await service.list_events(session, task_id, type, since_id, limit, platform)
return {"events": events, "latest_id": events[0]["id"] if events else since_id}
@router.post("/events/read")
async def mark_events_read(task_id: Optional[int] = None):
async with get_session() as session:
count = await service.mark_events_read(session, task_id)
return {"marked": count}
# ---------------------------------------------------------------------------
# Cookie / login health
# ---------------------------------------------------------------------------
@router.get("/cookie")
async def get_cookie_endpoint(platform: str = Query(default=PLATFORM_XHS)):
"""Cookie health only -- deliberately never returns the cookie value.
``platform`` defaults to Xiaohongshu so existing callers keep working; the
key it reads is the namespaced one.
"""
async with get_session() as session:
return await get_cookie_status(session, platform)
@router.post("/cookie")
async def set_cookie_endpoint(payload: CookiePayload, platform: str = Query(default=PLATFORM_XHS)):
async with get_session() as session:
await set_cookie(session, payload.cookie.strip(), platform)
return {"message": "Cookie saved"}
@router.delete("/cookie")
async def clear_cookie_endpoint(platform: str = Query(default=PLATFORM_XHS)):
async with get_session() as session:
await delete_setting(session, cookie_key(platform))
return {"message": "Cookie cleared"}
# ---------------------------------------------------------------------------
# 博主备注
# ---------------------------------------------------------------------------
@router.put("/creators/{creator_hash}")
async def set_creator_alias_endpoint(
creator_hash: str,
payload: CreatorAliasPayload,
platform: str = Query(default=PLATFORM_XHS),
):
"""给博主起个备注(界面上的「备注」)。
作品栏按 creator_hash 归组,可那是个哈希、昵称又常常认不出是谁 —— 备注是人自己起的
名字。空串表示清掉这条备注。
"""
async with get_session() as session:
await service.set_creator_alias(session, platform, creator_hash, payload.alias)
return {"creator_hash": creator_hash, "alias": payload.alias.strip()}
@router.put("/notes/{note_id}")
async def set_note_alias_endpoint(
note_id: str,
payload: NoteAliasPayload,
platform: str = Query(default=PLATFORM_XHS),
):
"""给**作品**起个备注。
和上面那条博主备注是一对:博主备注回答「这个账号是谁」,这条回答「这条作品我要盯着」。
两者不能合并 —— 一个博主底下常常只有一两件值得盯的作品。
"""
async with get_session() as session:
await service.set_note_alias(session, platform, note_id, payload.alias)
return {"note_id": note_id, "alias": payload.alias.strip()}
# ---------------------------------------------------------------------------
# QR login
# ---------------------------------------------------------------------------
# These drive the browser already listening on the CDP debug port, which is the
# same browser -- and therefore the same profile -- that monitor runs attach to.
# Scanning once is what makes later unattended runs logged in.
@router.get("/covers/{note_id}")
async def get_cover(note_id: str):
"""作品封面,从本地缓存读。
**为什么不让前端直连图床**:图床地址是带签名、会过期的 —— 实测隔天即 403,
而且带不带 Referer 都一样,所以那是过期而不是防盗链。本地那份与签名无关。
这个路由是带鉴权的(整条 monitor 路由都挂了 require_auth),所以封面不会被
匿名读走;前端用同源的 <img> 请求会自动带上会话 cookie。
"""
path = covers.find_cached(note_id)
if path is None:
raise HTTPException(status_code=404, detail="封面未缓存")
media_types = {
".jpg": "image/jpeg",
".png": "image/png",
".webp": "image/webp",
".gif": "image/gif",
".heic": "image/heic",
}
return FileResponse(
path,
media_type=media_types.get(path.suffix.lower(), "application/octet-stream"),
# 本地文件不会变(note_id 唯一),让浏览器自己缓存,省掉重复请求。
headers={"Cache-Control": "private, max-age=86400"},
)
@router.post("/login/qr")
async def start_qr_login(platform: str = Query(default=PLATFORM_XHS)):
"""Open the login page in the CDP browser and return its QR code.
A server deployment has no display (Chrome sits under Xvfb), so the code is
surfaced here for the operator to scan instead of in a desktop window that
does not exist.
"""
try:
return await qrlogin.start(platform)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc))
except RuntimeError as exc:
raise HTTPException(status_code=502, detail=str(exc))
@router.get("/login/qr")
async def get_qr_login():
"""Poll the live session: waiting -> success / expired / error.
扫码成功时**把 cookie 一并存进库**。扫码本来只写浏览器 profile,那只够 CDP 模式用;
存一份之后,CDP 关掉、任务改用 --cookies_file 注入也照样能跑 —— 两种机制同时填上,
开关怎么切都不会断。
"""
snapshot = await qrlogin.status()
if snapshot["status"] == qrlogin.STATUS_SUCCESS:
cookie = await qrlogin.take_cookie()
if cookie:
async with get_session() as session:
await set_cookie(session, cookie)
snapshot["cookie_saved"] = True
snapshot["message"] = f"{snapshot['message']};登录态已同时存入 Cookie"
return snapshot
@router.delete("/login/qr")
async def cancel_qr_login():
"""Drop our tab and stop polling."""
return await qrlogin.cancel()
@router.get("/login/state")
async def get_login_state(force: bool = Query(default=False)):
"""Ask the browser itself whether it is signed in.
Deliberately separate from the QR session above. That session is in-memory and
dies with the process -- a redeploy is enough -- so "am I logged in?" must not
hinge on it, or a successful scan looks like nothing happened.
``force`` reloads the page first, for when the login may have lapsed somewhere
else and the page's copy of the state is stale.
"""
return await qrlogin.check_login_state(force=force)
# ---------------------------------------------------------------------------
# Report
# ---------------------------------------------------------------------------
async def _resolve_scope(
session, task_ids: Optional[List[int]], platform: Optional[str]
) -> Optional[List[int]]:
"""Combine an explicit task selection with an optional platform filter.
``None`` means "no restriction"; an explicit list is intersected with the
platform's tasks so a stale selection cannot leak another platform's data
into a scoped report.
"""
if platform is None:
return task_ids
platform_ids = set(await service.platform_task_ids(session, platform))
if task_ids is None:
return list(platform_ids)
return [task for task in task_ids if task in platform_ids]
@router.get("/export")
async def export_data(
kind: str = Query(..., description="notes | comments | report"),
task_id: Optional[List[int]] = Query(default=None),
note_id: Optional[str] = None,
start_date: Optional[str] = Query(default=None, description="YYYY-MM-DD,report 用"),
end_date: Optional[str] = Query(default=None, description="YYYY-MM-DD,report 用"),
days: int = Query(default=7, ge=1, le=365),
file_format: str = Query(default="csv", alias="format", description="csv | xlsx"),
platform: Optional[str] = None,
):
"""Download collected data as CSV or Excel.
Reached by the browser as a navigation (``window.open``), which cannot carry
an Authorization header -- this is one of the reasons the session lives in a
cookie.
"""
if kind not in ("notes", "comments", "report"):
raise HTTPException(status_code=400, detail="kind 必须是 notes / comments / report")
if file_format not in ("csv", "xlsx"):
raise HTTPException(status_code=400, detail="format 必须是 csv 或 xlsx")
async with get_session() as session:
scoped = await _resolve_scope(session, task_id, platform)
if kind == "notes":
single_task = scoped[0] if scoped and len(scoped) == 1 else None
rows = await service.list_notes(session, single_task, False, 5000, platform)
elif kind == "comments":
single_task = scoped[0] if scoped and len(scoped) == 1 else None
rows = await service.list_comments(session, single_task, note_id, 5000, platform)
else:
try:
end_day = date.fromisoformat(end_date) if end_date else date.today()
start_day = (
date.fromisoformat(start_date)
if start_date
else end_day - timedelta(days=days - 1)
)
except ValueError:
raise HTTPException(status_code=400, detail="日期格式应为 YYYY-MM-DD")
# Not `report = ...`: that would make `report` a local name for the
# whole function and shadow the module import on this very line.
report_data = await report.build_report(session, scoped, start_day, end_day)
rows = report_data["rows"]
if not rows:
raise HTTPException(status_code=404, detail="该范围内没有数据可导出")
columns = _export_columns(kind)
stamp = date.today().isoformat()
# ASCII on purpose: a non-ASCII filename needs RFC 5987 encoding in
# Content-Disposition, and the plain `filename="..."` form used below would
# mangle it.
filename = f"export_{kind}_{stamp}.{file_format}"
if file_format == "xlsx":
payload = _to_xlsx(rows, columns, kind)
media_type = "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"
else:
payload = _to_csv(rows, columns)
media_type = "text/csv; charset=utf-8"
return Response(
content=payload,
media_type=media_type,
headers={"Content-Disposition": f'attachment; filename="{filename}"'},
)
def _export_columns(kind: str) -> List[tuple[str, str]]:
"""(key, header) pairs per export kind."""
if kind == "notes":
return [
# 先放「这是谁」:导出来是拿去比对和汇报的,一行只有作品 ID 没法用。
# 备注优先 —— 昵称常常认不出是谁(见 notes 表那一层的说明)。
("creator_alias", "博主备注"),
("creator_name", "博主昵称"),
("note_alias", "作品备注"),
("title", "标题"),
("note_id", "作品ID"),
("note_url", "链接"),
("published_at", "发布时间"),
# 指标嵌在 row["metrics"] 里,所以这里必须写成路径 —— 写成裸键名的话这几列
# 全空(见 _lookup)。
("metrics.liked_count", "点赞"),
("metrics.comment_count", "评论"),
("metrics.collected_count", "收藏"),
("metrics.share_count", "分享"),
("deltas.liked_count", "点赞增量"),
("deltas.comment_count", "评论增量"),
("first_seen_at", "首次发现"),
("last_seen_at", "最近采集"),
]
if kind == "comments":
return [
("note_creator_name", "博主昵称"),
("note_title", "所属作品"),
("note_id", "作品ID"),
("comment_id", "评论ID"),
("content", "内容"),
("nickname", "昵称"),
("like_count", "点赞"),
("sub_comment_count", "子评论数"),
("create_time", "发布时间"),
("first_seen_at", "首次发现"),
]
return [
("date", "日期"),
("new_notes", "新增作品"),
("new_comments", "新增评论"),
("liked_count_delta", "点赞增量"),
("comment_count_delta", "评论增量"),
("collected_count_delta", "收藏增量"),
("share_count_delta", "分享增量"),
]
# 表里存的是毫秒时间戳。直接倒进 CSV 就是一串 13 位数字 —— 打开 Excel 的人没法看,
# 也没法排序。这几个键统一格式化成人能读的形态。
_TIME_KEYS = {"published_at", "first_seen_at", "last_seen_at", "create_time"}
def _fmt_time(value: Any) -> str:
"""毫秒 → ``YYYY-MM-DD HH:MM``(服务器本地时区)。"""
try:
return datetime.fromtimestamp(int(value) / 1000).strftime("%Y-%m-%d %H:%M")
except (TypeError, ValueError, OSError, OverflowError):
return ""
def _cell_for(row: Dict[str, Any], key: str) -> Any:
"""一列的值:时间键格式化成人能读的,其余照原样(None 变空串)。"""
value = _lookup(row, key)
if key.rsplit(".", 1)[-1] in _TIME_KEYS:
return _fmt_time(value)
return _cell(value)
def _lookup(row: Dict[str, Any], key: str) -> Any:
"""取一列的值。键可以是 ``metrics.liked_count`` 这种路径。
作品行的指标是**嵌在** ``metrics`` / ``deltas`` 里的,而 ``_export_columns`` 里写的
是 ``liked_count`` —— 照顶层键直接 ``row.get()`` 的话,点赞/评论/收藏/分享四列连带
两个增量列**永远是空的**,导出来的表看着有这几列,其实一格都没有。
"""
value: Any = row
for part in key.split("."):
if not isinstance(value, dict):
return None
value = value.get(part)
return value
def _cell(value: Any) -> Any:
if value is None:
return ""
if isinstance(value, (list, dict)):
return ", ".join(str(v) for v in value) if isinstance(value, list) else str(value)
return value
def _to_csv(rows: List[Dict[str, Any]], columns: List[tuple[str, str]]) -> bytes:
import csv
import io
buffer = io.StringIO()
writer = csv.writer(buffer)
writer.writerow([header for _, header in columns])
for row in rows:
writer.writerow([_cell_for(row, key) for key, _ in columns])
# utf-8-sig: without the BOM Excel opens Chinese CSV as mojibake, which is
# the single most common complaint about CSV exports here.
return buffer.getvalue().encode("utf-8-sig")
def _to_xlsx(rows: List[Dict[str, Any]], columns: List[tuple[str, str]], sheet: str) -> bytes:
import io
from openpyxl import Workbook
workbook = Workbook()
worksheet = workbook.active
worksheet.title = {"notes": "作品", "comments": "评论"}.get(sheet, "报表")
worksheet.append([header for _, header in columns])
for row in rows:
worksheet.append([_cell_for(row, key) for key, _ in columns])
output = io.BytesIO()
workbook.save(output)
return output.getvalue()
@router.get("/report")
async def get_report(
task_id: Optional[List[int]] = Query(
default=None, description="Repeat to include several tasks; omit for all"
),
start_date: Optional[str] = Query(default=None, description="YYYY-MM-DD"),
end_date: Optional[str] = Query(default=None, description="YYYY-MM-DD"),
days: int = Query(default=7, ge=1, le=365, description="Window used when dates are omitted"),
platform: Optional[str] = None,
):
"""Daily new-content counts and interaction deltas for the selected tasks."""
try:
end_day = date.fromisoformat(end_date) if end_date else date.today()
start_day = date.fromisoformat(start_date) if start_date else end_day - timedelta(days=days - 1)
except ValueError:
raise HTTPException(status_code=400, detail="日期格式应为 YYYY-MM-DD")
if start_day > end_day:
raise HTTPException(status_code=400, detail="开始日期不能晚于结束日期")
async with get_session() as session:
scoped = await _resolve_scope(session, task_id, platform)
return await report.build_report(session, scoped, start_day, end_day)
# ---------------------------------------------------------------------------
# WeCom webhook
# ---------------------------------------------------------------------------
def _mask_webhook(url: str) -> str:
"""Show enough of the URL to recognise it, without exposing the robot key."""
if not url:
return ""
key_marker = "key="
index = url.find(key_marker)
if index == -1:
return url[:12] + "..." if len(url) > 12 else url
prefix = url[: index + len(key_marker)]
key = url[index + len(key_marker) :]
if len(key) <= 8:
return prefix + "*" * len(key)
return f"{prefix}{key[:4]}...{key[-4:]}"
@router.get("/webhook")
async def get_webhook():
async with get_session() as session:
url = (await get_setting(session, SETTING_WECOM_WEBHOOK)) or ""
return {"configured": bool(url), "masked": _mask_webhook(url)}
@router.post("/webhook")
async def set_webhook(payload: WebhookPayload):
url = payload.url.strip()
if url and "qyapi.weixin.qq.com" not in url:
# Catches the common mistake of pasting a group-chat invite or the app
# URL instead of the robot webhook.
raise HTTPException(
status_code=400,
detail="这不像企业微信机器人 Webhook 地址(应包含 qyapi.weixin.qq.com)",
)
async with get_session() as session:
await set_setting(session, SETTING_WECOM_WEBHOOK, url)
return {"message": "Webhook 已保存" if url else "Webhook 已清空", "configured": bool(url)}
@router.delete("/webhook")
async def clear_webhook():
async with get_session() as session:
await delete_setting(session, SETTING_WECOM_WEBHOOK)
return {"message": "Webhook 已删除"}
@router.post("/webhook/test")
async def test_webhook(payload: WebhookTestPayload):
"""Send a test message so the user can verify the robot works before relying on it."""
async with get_session() as session:
url = payload.url.strip() if payload.url else await notify.get_webhook_url(session)
ok, detail = await notify.send_wecom(
url, "**综合采集平台 通知测试**\n> 如果你看到这条消息,说明 Webhook 配置成功。"
)
if not ok:
raise HTTPException(status_code=400, detail=detail)
return {"message": detail}
# ---------------------------------------------------------------------------
# 上游更新
# ---------------------------------------------------------------------------
@router.get("/upstream")
async def get_upstream_status():
"""最近一次上游检查的结果。
只读缓存,不触发检查:fetch 要走网络、最长两分钟,不该由一个 GET 顺手发起。
没有查过时返回空对象,前端据此显示「尚未检查」。
"""
async with get_session() as session:
return await upstream.load_state(session)
@router.post("/upstream/check")
async def run_upstream_check():
"""立刻检查一次上游仓库,并把结果写回缓存。
即使「定期检查」开关是关的也照查 —— 手动点这一次的意义正在于此。这里会一直
等到 fetch 结束(前端给这条请求单独放长了超时),因为结果就是要给人看的。
``notify_when_new=False``:点这个按钮的人正看着结果,没必要再给自己推一条群消息。
没有推过的那批提交会留给下一次「定时检查」推 —— 推送状态记的是 tip,不是「推过没」。
"""
return await upstream.run_check(notify_when_new=False)
+70
View File
@@ -0,0 +1,70 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/settings.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Unified settings endpoint.
Consolidates what used to be scattered across the monitor router. The older
``/api/monitor/cookie`` and ``/api/monitor/webhook`` endpoints are deliberately
left in place -- they still work and removing them would be a breaking change for
no gain.
"""
from typing import Any, Dict
from fastapi import APIRouter, HTTPException, Query
from ..monitor import app_settings
from ..monitor.db import get_session
from ..monitor.platforms import PLATFORM_XHS
from ..schemas.settings import SettingsUpdatePayload
router = APIRouter(prefix="/settings", tags=["settings"])
@router.get("")
async def read_settings(platform: str = Query(default=PLATFORM_XHS)):
"""Settings for one platform, plus the system-wide ones.
Platform-scoped values are returned for the requested platform; system-scoped
values are the same regardless. Every spec carries its resolved ``key`` so the
UI can PUT changes straight back.
Sensitive values are returned as ``{present, length}`` only.
"""
async with get_session() as session:
return await app_settings.get_all(session, platform)
@router.put("")
async def write_settings(
payload: SettingsUpdatePayload, platform: str = Query(default=PLATFORM_XHS)
):
"""Partial update: only the keys present in the body are written.
A key belonging to a different platform is rejected rather than written
somewhere unexpected.
"""
values: Dict[str, Any] = payload.values()
async with get_session() as session:
try:
changed = await app_settings.update(session, values, platform)
except app_settings.SettingValidationError as exc:
raise HTTPException(status_code=400, detail=str(exc))
return {"message": f"已保存 {len(changed)} 项设置", "changed": changed}
+7 -3
View File
@@ -19,8 +19,9 @@
import asyncio
from typing import Set, Optional
from fastapi import APIRouter, WebSocket, WebSocketDisconnect
from fastapi import APIRouter, Depends, WebSocket, WebSocketDisconnect
from ..auth import require_ws_auth
from ..services import crawler_manager
router = APIRouter(tags=["websocket"])
@@ -86,7 +87,10 @@ def start_broadcaster():
_broadcaster_task = asyncio.create_task(log_broadcaster())
@router.websocket("/ws/logs")
# Websocket routes need their own auth dependency: BaseHTTPMiddleware returns
# early for any non-http scope, and HTTP router-level dependencies do not reach
# websocket routes. Without this the live crawl log stream would be wide open.
@router.websocket("/ws/logs", dependencies=[Depends(require_ws_auth)])
async def websocket_logs(websocket: WebSocket):
"""WebSocket log stream"""
print("[WS] New connection attempt")
@@ -134,7 +138,7 @@ async def websocket_logs(websocket: WebSocket):
print(f"[WS] Cleanup done, active connections: {len(manager.active_connections)}")
@router.websocket("/ws/status")
@router.websocket("/ws/status", dependencies=[Depends(require_ws_auth)])
async def websocket_status(websocket: WebSocket):
"""WebSocket status stream"""
await websocket.accept()
+34
View File
@@ -0,0 +1,34 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/auth.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Request models for authentication endpoints."""
from pydantic import BaseModel, Field
# A floor, not a policy: this is a single-operator internal panel, so the goal is
# only to reject obviously weak input.
MIN_PASSWORD_LENGTH = 8
class LoginPayload(BaseModel):
password: str = Field(min_length=1)
class ChangePasswordPayload(BaseModel):
current: str = Field(min_length=1)
new: str = Field(min_length=MIN_PASSWORD_LENGTH, max_length=256)
+28
View File
@@ -78,6 +78,34 @@ class CrawlerStartRequest(BaseModel):
max_notes_count: Optional[int] = Field(default=None, ge=1, le=MAX_API_LIMIT_COUNT)
max_comments_count: Optional[int] = Field(default=None, ge=1, le=MAX_API_LIMIT_COUNT)
# --- Options only used by scheduled monitor runs. Each defaults to None so
# the corresponding CLI flag is omitted entirely for manual Crawl-tab runs,
# which keeps their behaviour byte-identical to before.
# Isolate this run's output in its own directory. The crawler's own file
# writer names files by date only, so same-day runs would otherwise append
# into one shared file and could not be told apart.
save_data_path: Optional[str] = None
# Unattended runs must not try to attach to the user's desktop Chrome.
enable_cdp_mode: Optional[bool] = None
# XHS cookie login only injects `web_session` by default, which is not enough
# to sign API requests from a cold browser profile.
inject_all_cookies: Optional[bool] = None
save_login_state: Optional[bool] = None
# Preferred over `cookies`: a value on the command line is visible in the
# process list.
cookies_file: Optional[str] = None
max_concurrency_num: Optional[int] = Field(default=None, ge=1, le=MAX_API_LIMIT_COUNT)
# Crawl-strategy and proxy knobs surfaced on the Settings page. Like the
# fields above, each stays None unless the caller sets it, so the CLI flag is
# omitted entirely and the config-file default applies.
crawler_max_sleep_sec: Optional[int] = Field(default=None, ge=0, le=600)
enable_ip_proxy: Optional[bool] = None
ip_proxy_pool_count: Optional[int] = Field(default=None, ge=1, le=100)
ip_proxy_provider_name: Optional[str] = None
static_proxy_url: Optional[str] = None
class CrawlerStatusResponse(BaseModel):
"""Crawler status response"""
+128
View File
@@ -0,0 +1,128 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/monitor.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Request models for the monitoring API."""
from typing import List, Literal, Optional
from pydantic import BaseModel, Field, model_validator
# A floor on the interval is a correctness guard, not a nicety: every run
# launches a browser and hits XHS with several requests, so a short interval
# across many creators is the pattern that triggers rate limiting.
MIN_INTERVAL_MINUTES = 30
MAX_INTERVAL_MINUTES = 7 * 24 * 60
class MonitorTaskCreate(BaseModel):
name: str = Field(min_length=1, max_length=200)
platform: str = "xhs"
# One subprocess handles exactly one crawler type, so a task is either
# creator-driven or note-driven.
mode: Literal["creator", "note"]
# None means "use the value configured on the Settings page", which is what
# makes those defaults meaningful. Bounds still apply when a value is given.
interval_minutes: Optional[int] = Field(
default=None, ge=MIN_INTERVAL_MINUTES, le=MAX_INTERVAL_MINUTES
)
# --- Scheduling ---------------------------------------------------------
# All three modes are expressible with pickers; a raw cron string is
# deliberately not supported, since it is a small language to learn just to
# say "every day at nine".
schedule_mode: Literal["interval", "daily", "weekly"] = "interval"
# 0-23, e.g. [9, 12, 18]. Required for the two clock modes.
schedule_hours: List[int] = Field(default_factory=list)
# 0-6 with Monday = 0, matching Python's date.weekday(). Required for weekly.
schedule_days: List[int] = Field(default_factory=list)
# One minute for the whole schedule, so a task with three times is
# "09:30, 12:30, 18:30" rather than three separate minute choices.
schedule_minute: int = Field(default=0, ge=0, le=59)
max_notes_count: Optional[int] = Field(default=None, ge=1, le=500)
enable_comments: bool = True
# Raising this widens the comment window, which is the only lever available
# for noticing new comments -- the API has no time-sort.
max_comments_count: Optional[int] = Field(default=None, ge=1, le=500)
run_timeout_seconds: int = Field(default=3600, ge=60, le=86400)
enabled: bool = True
# 两类通知分开:新作品可能每轮都有(默认关,避免刷屏),
# 异常频率低且意味着任务已经停止工作(默认开,否则你会一直不知道)。
notify_enabled: bool = False
notify_failures: bool = True
# Raw pasted values: full URLs or bare ids, in either form.
targets: List[str] = Field(min_length=1)
@model_validator(mode="after")
def _validate_schedule(self) -> "MonitorTaskCreate":
if any(hour < 0 or hour > 23 for hour in self.schedule_hours):
raise ValueError("小时必须在 0-23 之间")
if any(day < 0 or day > 6 for day in self.schedule_days):
raise ValueError("星期必须在 0-6 之间(周一为 0)")
# A clock mode with no chosen time can never fire. Rejecting it here is
# what keeps next_occurrence()'s None branch unreachable in practice.
if self.schedule_mode in ("daily", "weekly") and not self.schedule_hours:
raise ValueError("按钟点调度至少要选一个时间")
if self.schedule_mode == "weekly" and not self.schedule_days:
raise ValueError("按周调度至少要选一个星期")
return self
class MonitorTaskUpdate(BaseModel):
name: Optional[str] = Field(default=None, min_length=1, max_length=200)
enabled: Optional[bool] = None
interval_minutes: Optional[int] = Field(
default=None, ge=MIN_INTERVAL_MINUTES, le=MAX_INTERVAL_MINUTES
)
# None means "leave alone". Cross-field validity depends on the merged state,
# so it is checked in the service rather than here.
schedule_mode: Optional[Literal["interval", "daily", "weekly"]] = None
schedule_hours: Optional[List[int]] = None
schedule_days: Optional[List[int]] = None
schedule_minute: Optional[int] = Field(default=None, ge=0, le=59)
max_notes_count: Optional[int] = Field(default=None, ge=1, le=500)
enable_comments: Optional[bool] = None
max_comments_count: Optional[int] = Field(default=None, ge=1, le=500)
run_timeout_seconds: Optional[int] = Field(default=None, ge=60, le=86400)
notify_enabled: Optional[bool] = None
notify_failures: Optional[bool] = None
# When present, replaces the whole target list.
targets: Optional[List[str]] = None
class CookiePayload(BaseModel):
cookie: str = Field(min_length=1)
class CreatorAliasPayload(BaseModel):
"""给博主起的备注。空串表示清掉这条备注。"""
alias: str = Field(default="", max_length=128)
class NoteAliasPayload(BaseModel):
"""给作品起的备注。空串表示清掉这条备注。"""
alias: str = Field(default="", max_length=128)
class WebhookPayload(BaseModel):
url: str = Field(default="", description="企业微信机器人 Webhook 地址,留空表示停用")
class WebhookTestPayload(BaseModel):
url: Optional[str] = Field(default=None, description="不传则使用已保存的地址")
+37
View File
@@ -0,0 +1,37 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/schemas/settings.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Request model for the settings endpoint."""
from typing import Any, Dict
from pydantic import BaseModel, ConfigDict
class SettingsUpdatePayload(BaseModel):
"""Partial update of arbitrary setting keys.
Fields are not declared here on purpose: ``api/monitor/app_settings.py``
owns the registry (key, type, bounds, choices) and validates against it, so
adding a setting does not mean editing a matching schema.
"""
model_config = ConfigDict(extra="allow")
def values(self) -> Dict[str, Any]:
return dict(self.model_extra or {})
+125 -11
View File
@@ -20,11 +20,19 @@ import asyncio
import subprocess
import signal
import os
from typing import Optional, List
from collections import deque
from typing import Deque, Optional, List
from datetime import datetime
from pathlib import Path
from ..schemas import CrawlerStartRequest, LogEntry
from .interpreter import resolve_python_cmd
# 留住多少行爬虫输出,供 run_and_wait 的调用方诊断失败原因。
# 子进程的输出本来只流向日志 WebSocket,监控层只看得到退出码 —— 于是「退出码 1」
# 成了运行历史里唯一的信息,真正的报错(比如抖音的 `DataFetchError: account blocked`)
# 谁也看不到。留个尾巴,让失败原因能被写进 run.error_message。
OUTPUT_TAIL_LINES = 80
class CrawlerManager:
@@ -43,6 +51,20 @@ class CrawlerManager:
self._project_root = Path(__file__).parent.parent.parent
# Log queue - for pushing to WebSocket
self._log_queue: Optional[asyncio.Queue] = None
# Completion signalling for run_and_wait(). Polling `status` is unreliable
# because stop() also resets it to "idle", and `self.process` gets replaced
# by any concurrent start(), so waiters need an explicit event instead.
self._done: asyncio.Event = asyncio.Event()
self.last_exit_code: Optional[int] = None
# 本次运行输出的末尾若干行。见 OUTPUT_TAIL_LINES。
self._output_tail: Deque[str] = deque(maxlen=OUTPUT_TAIL_LINES)
def get_output_tail(self) -> List[str]:
"""最近一次运行的输出尾巴(最早的排前面)。
只在 run_and_wait() 返回之后读才有意义 —— 它等到读输出的任务收尾才唤醒。
"""
return list(self._output_tail)
@property
def logs(self) -> List[LogEntry]:
@@ -54,6 +76,43 @@ class CrawlerManager:
self._log_queue = asyncio.Queue()
return self._log_queue
def is_busy(self) -> bool:
"""Whether a crawler process is currently alive.
This is the authoritative busy check -- `status` is a lagging indicator
that manual stop() also resets.
"""
return self.process is not None and self.process.poll() is None
async def run_and_wait(
self,
config: CrawlerStartRequest,
extra_args: Optional[List[str]] = None,
timeout: Optional[float] = None,
) -> int:
"""Start a crawler run and block until it exits, returning the exit code.
Used by the monitor scheduler. Returns a negative value if the run was
killed by `timeout` or if the process could not be started at all.
"""
started = await self.start(config, extra_args=extra_args)
if not started:
return -1
# Capture the process we just launched: a concurrent start() would
# replace self.process, so poll this reference rather than the attribute.
proc = self.process
if proc is None:
return -1
try:
await asyncio.wait_for(self._done.wait(), timeout=timeout)
except asyncio.TimeoutError:
await self.stop()
return -1
return self.last_exit_code if self.last_exit_code is not None else -1
def _create_log_entry(self, message: str, level: str = "info") -> LogEntry:
"""Create log entry"""
self._log_id += 1
@@ -71,6 +130,9 @@ class CrawlerManager:
async def _push_log(self, entry: LogEntry):
"""Push log to queue"""
# 这里是所有输出的唯一出口(读循环、收尾、以及管理器自己的提示都走它),
# 所以尾巴挂在这儿最省事,也不会漏。
self._output_tail.append(entry.message)
if self._log_queue is not None:
try:
self._log_queue.put_nowait(entry)
@@ -90,7 +152,11 @@ class CrawlerManager:
return "debug"
return "info"
async def start(self, config: CrawlerStartRequest) -> bool:
async def start(
self,
config: CrawlerStartRequest,
extra_args: Optional[List[str]] = None,
) -> bool:
"""Start crawler process"""
async with self._lock:
if self.process and self.process.poll() is None:
@@ -99,6 +165,11 @@ class CrawlerManager:
# Clear old logs
self._logs = []
self._log_id = 0
# Reset completion signalling for this run
self._done.clear()
self.last_exit_code = None
# 尾巴只属于本次运行,否则上一轮的报错会混进这一轮的诊断里。
self._output_tail.clear()
# Clear pending queue (don't replace object to avoid WebSocket broadcast coroutine holding old queue reference)
if self._log_queue is None:
@@ -111,7 +182,7 @@ class CrawlerManager:
pass
# Build command line arguments
cmd = self._build_command(config)
cmd = self._build_command(config, extra_args=extra_args)
# Log start information
entry = self._create_log_entry(f"Starting crawler: {' '.join(cmd)}", "info")
@@ -202,9 +273,13 @@ class CrawlerManager:
"error_message": None
}
def _build_command(self, config: CrawlerStartRequest) -> list:
def _build_command(
self,
config: CrawlerStartRequest,
extra_args: Optional[List[str]] = None,
) -> list:
"""Build main.py command line arguments"""
cmd = ["uv", "run", "python", "main.py"]
cmd = [*resolve_python_cmd(), "main.py"]
cmd.extend(["--platform", config.platform.value])
cmd.extend(["--lt", config.login_type.value])
@@ -232,22 +307,56 @@ class CrawlerManager:
if config.max_comments_count is not None:
cmd.extend(["--max_comments_count_singlenotes", str(config.max_comments_count)])
if config.cookies:
# Each of these is only appended when explicitly set, so manual runs from
# the Crawl tab keep exactly their previous behaviour.
if config.save_data_path:
cmd.extend(["--save_data_path", config.save_data_path])
if config.enable_cdp_mode is not None:
cmd.extend(["--enable_cdp_mode", "true" if config.enable_cdp_mode else "false"])
if config.inject_all_cookies is not None:
cmd.extend(["--inject_all_cookies", "true" if config.inject_all_cookies else "false"])
if config.save_login_state is not None:
cmd.extend(["--save_login_state", "true" if config.save_login_state else "false"])
if config.max_concurrency_num is not None:
cmd.extend(["--max_concurrency_num", str(config.max_concurrency_num)])
if config.crawler_max_sleep_sec is not None:
cmd.extend(["--crawler_max_sleep_sec", str(config.crawler_max_sleep_sec)])
if config.enable_ip_proxy is not None:
cmd.extend(["--enable_ip_proxy", "true" if config.enable_ip_proxy else "false"])
if config.ip_proxy_pool_count is not None:
cmd.extend(["--ip_proxy_pool_count", str(config.ip_proxy_pool_count)])
if config.ip_proxy_provider_name:
cmd.extend(["--ip_proxy_provider_name", config.ip_proxy_provider_name])
if config.static_proxy_url:
cmd.extend(["--static_proxy_url", config.static_proxy_url])
# Prefer a cookie file over passing the cookie on the command line, where
# it would be visible in the process list.
if config.cookies_file:
cmd.extend(["--cookies_file", config.cookies_file])
elif config.cookies:
cmd.extend(["--cookies", config.cookies])
cmd.extend(["--headless", "true" if config.headless else "false"])
if extra_args:
cmd.extend(extra_args)
return cmd
async def _read_output(self):
"""Asynchronously read process output"""
loop = asyncio.get_event_loop()
# Capture the process this reader was started for. self.process can be
# replaced by a subsequent start(), which would otherwise make us read
# the exit code of the wrong run.
proc = self.process
try:
while self.process and self.process.poll() is None:
while proc and proc.poll() is None:
# Read a line in thread pool
line = await loop.run_in_executor(
None, self.process.stdout.readline
None, proc.stdout.readline
)
if line:
line = line.strip()
@@ -257,9 +366,9 @@ class CrawlerManager:
await self._push_log(entry)
# Read remaining output
if self.process and self.process.stdout:
if proc and proc.stdout:
remaining = await loop.run_in_executor(
None, self.process.stdout.read
None, proc.stdout.read
)
if remaining:
for line in remaining.strip().split('\n'):
@@ -270,7 +379,7 @@ class CrawlerManager:
# Process ended
if self.status == "running":
exit_code = self.process.returncode if self.process else -1
exit_code = proc.returncode if proc else -1
if exit_code == 0:
entry = self._create_log_entry("Crawler completed successfully", "success")
else:
@@ -283,6 +392,11 @@ class CrawlerManager:
except Exception as e:
entry = self._create_log_entry(f"Error reading output: {str(e)}", "error")
await self._push_log(entry)
finally:
# Record the exit code and wake any run_and_wait() waiter. Runs in a
# finally so a cancelled read task still releases the waiter.
self.last_exit_code = proc.returncode if proc else None
self._done.set()
# Global singleton
+70
View File
@@ -0,0 +1,70 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/interpreter.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Interpreter resolution for spawning crawler subprocesses.
Historically both the crawler manager and the environment check hardcoded
``uv run``. ``uv`` is not guaranteed to be installed, so resolve the command
prefix in one place: prefer ``uv`` (matching upstream docs), fall back to a
project-local virtualenv, and finally to the interpreter running the server.
"""
import shutil
import sys
from pathlib import Path
# Project root: api/services/interpreter.py -> services -> api -> repo root
PROJECT_ROOT = Path(__file__).parent.parent.parent
def venv_python_path(project_root: Path | None = None) -> Path:
"""Return the path to the project venv's Python executable."""
root = project_root if project_root is not None else PROJECT_ROOT
if sys.platform == "win32":
return root / ".venv" / "Scripts" / "python.exe"
return root / ".venv" / "bin" / "python"
def resolve_python_cmd(project_root: Path | None = None) -> list[str]:
"""Resolve the command prefix used to run ``main.py``.
Order of preference:
1. ``uv`` if it is on PATH -- matches the upstream documented workflow.
2. The project-local ``.venv`` if it exists.
3. The interpreter currently running the API server.
Returns a list because the caller appends ``main.py`` and its flags.
"""
if shutil.which("uv"):
return ["uv", "run", "python"]
venv_python = venv_python_path(project_root)
if venv_python.exists():
return [str(venv_python)]
return [sys.executable]
def describe_interpreter(project_root: Path | None = None) -> str:
"""Human-readable description of what resolve_python_cmd() picks."""
cmd = resolve_python_cmd(project_root)
if cmd[0] == "uv":
return "uv run python"
if cmd[0] == sys.executable:
return f"current interpreter ({sys.executable})"
return f"project virtualenv ({cmd[0]})"
+59
View File
@@ -300,6 +300,14 @@ async def parse_cmd(argv: Optional[Sequence[str]] = None):
rich_help_panel="Performance Configuration",
),
] = config.MAX_CONCURRENCY_NUM,
crawler_max_sleep_sec: Annotated[
int,
typer.Option(
"--crawler_max_sleep_sec",
help="Seconds to wait between requests. Higher is slower but far less likely to trip platform rate limiting",
rich_help_panel="Performance Configuration",
),
] = config.CRAWLER_MAX_SLEEP_SEC,
save_data_path: Annotated[
str,
typer.Option(
@@ -308,6 +316,41 @@ async def parse_cmd(argv: Optional[Sequence[str]] = None):
rich_help_panel="Storage Configuration",
),
] = config.SAVE_DATA_PATH,
enable_cdp_mode: Annotated[
str,
typer.Option(
"--enable_cdp_mode",
help="Whether to drive the user's local Chrome over CDP instead of launching a browser, supports yes/true/t/y/1 or no/false/f/n/0. Set to false for unattended/server runs",
rich_help_panel="Runtime Configuration",
show_default=True,
),
] = str(config.ENABLE_CDP_MODE),
save_login_state: Annotated[
str,
typer.Option(
"--save_login_state",
help="Whether to persist the browser profile so a previous login can be reused, supports yes/true/t/y/1 or no/false/f/n/0",
rich_help_panel="Runtime Configuration",
show_default=True,
),
] = str(config.SAVE_LOGIN_STATE),
inject_all_cookies: Annotated[
str,
typer.Option(
"--inject_all_cookies",
help="Whether to inject every cookie supplied via --cookies/--cookies_file instead of only web_session, supports yes/true/t/y/1 or no/false/f/n/0",
rich_help_panel="Runtime Configuration",
show_default=True,
),
] = str(config.INJECT_ALL_COOKIES),
cookies_file: Annotated[
str,
typer.Option(
"--cookies_file",
help="Path to a file holding the cookie string. Preferred over --cookies, whose value is visible in the process list",
rich_help_panel="Runtime Configuration",
),
] = "",
enable_ip_proxy: Annotated[
str,
typer.Option(
@@ -350,6 +393,18 @@ async def parse_cmd(argv: Optional[Sequence[str]] = None):
enable_headless = _to_bool(headless)
enable_ip_proxy_value = _to_bool(enable_ip_proxy)
init_db_value = init_db.value if init_db else None
enable_cdp_mode_value = _to_bool(enable_cdp_mode)
save_login_state_value = _to_bool(save_login_state)
inject_all_cookies_value = _to_bool(inject_all_cookies)
# A file is preferred over --cookies: a literal value on the command line
# is visible to any other user on the machine via the process list.
if cookies_file:
try:
with open(cookies_file, "r", encoding="utf-8") as f:
cookies = f.read().strip()
except OSError as e:
raise typer.BadParameter(f"Unable to read --cookies_file: {e}")
# Parse specified_id and creator_id into lists
specified_id_list = [id.strip() for id in specified_id.split(",") if id.strip()] if specified_id else []
@@ -368,9 +423,13 @@ async def parse_cmd(argv: Optional[Sequence[str]] = None):
config.CDP_HEADLESS = enable_headless
config.SAVE_DATA_OPTION = save_data_option.value
config.COOKIES = cookies
config.ENABLE_CDP_MODE = enable_cdp_mode_value
config.SAVE_LOGIN_STATE = save_login_state_value
config.INJECT_ALL_COOKIES = inject_all_cookies_value
config.CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES = max_comments_count_singlenotes
config.CRAWLER_MAX_NOTES_COUNT = crawler_max_notes_count
config.MAX_CONCURRENCY_NUM = max_concurrency_num
config.CRAWLER_MAX_SLEEP_SEC = crawler_max_sleep_sec
config.SAVE_DATA_PATH = save_data_path
config.ENABLE_IP_PROXY = enable_ip_proxy_value
config.IP_PROXY_POOL_COUNT = ip_proxy_pool_count
+14
View File
@@ -52,6 +52,20 @@ HEADLESS = False
# Whether to save login status
SAVE_LOGIN_STATE = True
# 是否注入完整 cookie(默认 False,保持上游原有行为)。
# False 时 login_by_cookies 只写入 web_session;a1 / webId 等签名所需 cookie 只能靠
# browser_data 下的持久化 profile 补齐。无人值守场景(服务器上跑定时监控)应设为 True,
# 否则冷 profile 下 API 签名失败,且表现为「退出码 0 但抓到 0 条」的静默失败。
INJECT_ALL_COOKIES = False
# 是否对昵称做中间脱敏(默认 False —— 本仓库**关掉了**)。
# 上游作为教学版默认开启,保留首尾各 1 字、中间打星号,避免据昵称骚扰到真人。
# 但那是**有损**的:「张三」和「张四」都会变成「张*」,「小明老师」和「小刚老师」
# 都会变成「小***师」—— 而本仓库的用途是监控一批公开的创作者账号,分清谁是谁正是
# 这一层要干的事,撞名就等于看不出来。所以这里关掉,把原昵称原样落库。
# 想改回上游行为,把这一行改成 True 即可(脱敏机制本身没删)。
MASK_NICKNAME = False
# ==================== CDP (Chrome DevTools Protocol) 配置 ====================
# 是否启用 CDP 模式 - 使用用户本地的 Chrome/Edge 浏览器进行爬取,具有更好的反检测能力
# 开启后,会自动检测并启动用户的 Chrome/Edge 浏览器,通过 CDP 协议进行控制
+5 -2
View File
@@ -23,9 +23,12 @@
# Supported formats:
# 1. Full video URL: "https://www.kuaishou.com/short-video/3x3zxz4mjrsc8ke?authorId=3x84qugg4ch9zhs&streamSource=search"
# 2. Pure video ID: "3xf8enb8dbj6uig"
# 3. Share short link: "https://www.kuaishou.com/f/X9Idt15MQb9L2cv"
# (路径里是 share_token 不是视频 ID,会自动跟随 302 重定向解析)
KS_SPECIFIED_ID_LIST = [
"https://www.kuaishou.com/short-video/3x3zxz4mjrsc8ke?authorId=3x84qugg4ch9zhs&streamSource=search&area=searchxxnull&searchKey=python",
"3xf8enb8dbj6uig",
"https://www.kuaishou.com/f/X9Idt15MQb9L2cv",
"https://www.kuaishou.com/f/X-a8vLyTxvEvN2jg",
"a8vLyTxvEvN2jg",
# ........................
]
+1 -1
View File
@@ -25,7 +25,7 @@ SORT_TYPE = "popularity_descending"
# Specify the note URL list, which must carry the xsec_token parameter
XHS_SPECIFIED_NOTE_URL_LIST = [
"https://www.xiaohongshu.com/explore/64b95d01000000000c034587?xsec_token=AB0EFqJvINCkj6xOCKCQgfNNh8GdnBC_6XecG4QOddo3Q=&xsec_source=pc_cfeed"
"https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=YBIq8sY-0_K3BACQ7z57J9xJcdflrd8BAf5_zBeJtFMOQ=&xsec_source=pc_creatormng"
# ........................
]
Executable
+71
View File
@@ -0,0 +1,71 @@
#!/usr/bin/env bash
#
# 更新这台机器上的部署。用法:
#
# ./deploy.sh
#
# 为什么需要脚本而不是一句 `git pull && docker compose up -d`:
# 前端产物 api/webui 是 gitignore 的(它由 vite 生成),git pull 带不过来。
# 所以代码更新之后必须在服务器上重建一次前端,否则页面还是旧的。
# 这一步在容器里做,好处是服务器不需要装 Node —— 只有 Docker。
#
# node_modules 和 npm 缓存都留在挂载目录内,重复构建不会重新下载。
set -euo pipefail
cd "$(dirname "$0")"
# 用镜像里的解释器跑,宿主机的 Python 版本无关。
IMAGE=mediacrawler:latest
before=$(git rev-parse HEAD)
git pull --ff-only
after=$(git rev-parse HEAD)
if [ "$before" = "$after" ]; then
echo "== 代码已是最新($after)"
else
echo "== 代码更新 $before -> $after"
git --no-pager log --oneline "$before..$after" | sed 's/^/ /'
fi
# 镜像层(依赖)改动只能靠重建,而这一步不是自动的:Dockerfile 或 requirements.txt 变了,
# 下面那句 `docker compose up -d --force-recreate` 用的是旧镜像,改动根本不会生效。
# 至少要说出来,否则现象是「代码明明更新了,功能却报缺依赖」。
if [ "$before" != "$after" ] && ! git diff --quiet "$before" "$after" -- Dockerfile requirements.txt; then
echo "!! Dockerfile / requirements.txt 有改动,需要重建镜像后重跑本脚本:"
echo " docker compose build"
fi
# 前端重建的两种情况:产物根本不存在(首次部署),或 webui/ 有改动。
if [ ! -f api/webui/index.html ]; then
need_build=1
reason="前端产物不存在"
elif [ "$before" != "$after" ] && ! git diff --quiet "$before" "$after" -- webui/; then
need_build=1
reason="webui/ 有改动"
else
need_build=0
reason=""
fi
if [ "$need_build" = "1" ]; then
echo "== 重建前端($reason)"
# -u 1000:1000 而不是 root:这里产出的文件要留在这个目录里给后面用,
# 以 root 生成的 node_modules 会让下次构建和人工清理都变得别扭。
# HOME 指向挂载目录,这样 npm 的缓存在宿主机上,重建时能复用。
docker run --rm \
-u 1000:1000 \
-w /app/webui \
-v "$PWD:/app" \
-e HOME=/app/webui \
-e npm_config_registry=https://registry.npmmirror.com \
"$IMAGE" sh -c 'npm ci --no-audit --no-fund && npm run build'
else
echo "== 前端无改动,跳过构建"
fi
# --force-recreate,而不是裸的 `up -d`:代码是 bind mount,容器配置和镜像都没变,
# 所以 `up -d` 会判定"无需变更"直接跳过,Python 代码的改动根本不会生效。前端产物是
# 磁盘上的静态文件,能即时生效,这一点很容易掩盖上面那个问题,直到有人改了 .py 才发现。
echo "== 重启容器"
docker compose up -d --force-recreate
docker compose ps
+33
View File
@@ -0,0 +1,33 @@
services:
mediacrawler:
build: .
image: mediacrawler:latest
container_name: mediacrawler
restart: unless-stopped
# Run as the user that owns this checkout. Without it the container is root,
# and every file it writes into the mounted tree -- the crawler's per-run
# jsonl output above all -- comes out root-owned. That does not break the app,
# but it does lock the operator out of moving or deleting their own
# deployment, which is exactly what happened the first time this was deployed.
user: "1000:1000"
# host networking is a requirement, not a convenience: the crawler attaches
# to the operator's Chrome at 127.0.0.1:9222, and inside a bridge network
# that loopback is the container's own, where no browser is listening.
# It also puts the app port directly on the host, so `ports:` is not used.
network_mode: host
env_file:
- .env
environment:
MC_HOST: 0.0.0.0
MC_PORT: "18051"
TZ: Asia/Shanghai
volumes:
# The code is mounted rather than baked in, so shipping a change is
# "git pull, restart" instead of an image rebuild. Only the dependencies
# live in the image, because those are the expensive part and they change
# rarely -- rebuild only when requirements.txt or the Dockerfile changes.
- ./:/app
Binary file not shown.

After

Width:  |  Height:  |  Size: 6.2 KiB

BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

+527
View File
@@ -0,0 +1,527 @@
# 小红书监控功能使用说明
> 定时重复采集一批博主或笔记,与上一轮快照对比,产出**新增作品 / 新增评论 / 点赞收藏评论数涨跌**。
本功能是在 MediaCrawler 之上新增的一层,代码集中在 `api/monitor/`,不侵入原有的
`media_platform/`、`store/` 等目录。
---
## 一、为什么需要单独一层
原项目是**一次性采集**:跑完即退出,没有调度、没有历史、没有差分。直接复用会遇到三个硬伤:
1. **指标会被覆盖**。`store/xhs/_store_impl.py::XhsDbStoreImplement.update_content()` 对已存在的笔记执行
`UPDATE ... SET liked_count = ...`,历史值直接丢失。跑第二遍根本看不出"点赞从 100 涨到了 500"。
2. **单进程串行**。`api/services/crawler_manager.py` 是全局单例,同一时刻只能跑一个 `main.py` 子进程。
3. **运行输出无法区分**。`AsyncFileWriter` 的文件名只带日期(`creator_contents_2026-10-07.jsonl`),
同一天多次运行会追加进同一个文件。
监控层为此做了对应处理:独立的快照库(保留历史)、调度器与手动采集互斥排队、以及**每轮采集写入独立目录**
(复用早已存在、但 API 层从未转发的 `--save_data_path` 参数)。
---
## 二、快速开始
### 1. 准备环境
```bash
# 依赖(若未安装 uv,本项目的解释器探测会自动回退到 .venv)
python -m venv .venv
.venv/Scripts/python -m pip install -r requirements.txt
.venv/Scripts/python -m playwright install chromium # 非 CDP 模式必需
# 前端
cd webui && npm install && npm run build
```
### 2. 启动
```bash
.venv/Scripts/python -m api.main # 或 uvicorn api.main:app --port 8080
```
打开 <http://localhost:8080>,右上角切换到「监控」。
> 解释器探测顺序:`uv`(若在 PATH)→ 项目 `.venv` → 当前解释器。
> `/api/env/check` 使用同一套逻辑,不会出现"检测失败但其实能跑"的情况。
### 3. 配置登录态(**无人值守的前提**)
在监控页左下角「小红书登录态」粘贴 Cookie。定时监控不能每次都扫码,必须持久化登录态。
> **强烈建议先手动扫码登录一次**,以播种 `browser_data/xhs_user_data_dir`,
> 之后再粘贴 Cookie 才可靠。原因见下方「限制」。
### 4. 新建监控任务
- **类型**
- `博主`:监控其作品,填博主主页链接或纯 ID
- `笔记`:批量监控指定内容,填笔记链接或纯 ID
- **目标**:每行一个。**建议只填纯 ID** —— 链接里的 `xsec_token` 会过期,纯 ID 永久有效。
- **间隔**:最小 30 分钟。每次运行都要拉起一次浏览器并多次请求平台,间隔过短容易触发风控。
- **每篇评论抓取条数**:默认 50。这个值直接决定能发现多少新评论,见下方限制。
任务创建后立即生效,也可随时点「立即运行」手动触发一轮。
---
## 二·五、报表
「报表」视图是**跨任务**的统计,用来回答"这批账号这段时间表现如何"。
**筛选**:勾选参与统计的任务(默认全选),选日期区间(或点「近 7/30/90 天」)。
**两类指标,含义不同,所以分列展示**:
| 列 | 含义 |
|---|---|
| 新增作品 / 新增评论 | 该日**首次发现**的作品数 / 评论条数 |
| 点赞 Δ / 评论 Δ / 收藏 Δ / 分享 Δ | 该日**互动增量**:Σ(当日末值 − 当日之前最后一次采到的值) |
增量的口径有两个要点:
- **作品首次出现的那天从 0 起算**,所以新作品的全部点赞都计入其首次发现日。这样做是为了让"新作品带来了多少赞"这件事可见,而不是把它的既有数据丢掉。
- **某天没采到某篇作品,那天的增量算 0**,不会把跨天的增长平摊到每一天。
底部会标明两件事:一是**哪些指标无法解析**(小红书可能返回 `"1.2万"` 这类值,解析失败的不会被当成 0 计入,否则会伪造出一个大的负增长),二是评论数受接口限制只覆盖前 N 条。
> 实现上聚合是在 Python 里做的,不是一条大 SQL。原因:按笔记、按天的"上一个基线值"查询是窗口操作,SQLite 表达起来很别扭,而这里的数据量很小,可读性比压榨查询计划更值钱。
---
## 二·六、企业微信通知
在「监控」视图左下角配置 Webhook 地址(企业微信群 → 添加群机器人 → 复制 Webhook 地址)。
**两个设计取舍**:
1. **一轮只发一条汇总**,不是每条事件发一条。一次跑出 20 篇新作品时,你收到的是"新增作品 20 篇"加前 10 条标题,而不是 20 条消息。
2. **推送失败绝不影响采集**。通知是在数据提交之后、用独立会话发送的,任何网络错误只记日志。爬虫跑成功了不会因为 webhook 挂了而被回滚。
**触发时机**(仅这两类):
- 任务失败 / 疑似登录态失效
- 发现新增作品
指标变化和新增评论**不会**推送(指标变化太频繁,评论量可能很大)。
**任务范围**:每个任务在编辑弹窗里有「推送企业微信通知」开关,**默认关闭**。这样一个 webhook 不会被一堆无关任务刷屏。
- 配置好地址后可以点「发测试」验证,也可以「保存前先测」。
- 地址里的 key 等同凭据,**服务端只回传打码形式**,要换只能重新粘贴(和 Cookie 一致)。
- 任务卡片上的 `last_notified_at`(列表接口会返回)可以回答"为什么这轮没收到推送"。
---
## 二·七、评论视图与导出
「评论」页默认**按作品分组**:每篇作品一个可折叠区块,**默认只展开最新的一组**,
避免打开就是一屏文字。切到「平铺」则是一条流,每条评论下方标注它属于哪篇作品
(封面缩略图 + 标题 + 跳原文链接)。
顶部可按作品筛选,选项里带每篇的评论数:
```
全部作品
烤面筋热量计算 (33)
孜卷热量计算 (3)
```
> 评论与作品的关联是后端 JOIN 出来的(`note_title` / `note_cover` / `note_url`),
> 因为评论表本身只存 `note_id`,光看 ID 没有任何可读性。
### 导出
评论页和报表页都有「导出」按钮,走浏览器下载:
| 端点 | 内容 |
|---|---|
| `?kind=notes` | 作品表(含互动增量列) |
| `?kind=comments` | 评论(含所属作品标题) |
| `?kind=report` | 报表按天汇总 |
- 支持 `csv` 与 `xlsx`
- **CSV 带 UTF-8 BOM**(`utf-8-sig`)—— 否则 Excel 打开中文是乱码,这是最常见的投诉
- 下载是**页面导航**(`window.open`),带不了自定义请求头,所以导出依赖 Cookie 鉴权 ——
这也是会话必须存在 Cookie 里的原因之一
---
## 二·八、登录与访问控制
面板默认要求登录 —— `/api` 下的所有接口都需要会话,只有 `/api/health`、
`/api/auth/login`、`/api/auth/logout` 例外。静态资源(页面本身、JS/CSS)不受限制,
否则登录页自己都加载不出来。
### 首次启动
自动生成一个随机密码并**打印在启动日志里**(只打印一次):
```
====================================================================
WebUI 首次启动,已生成登录密码:
94Shn1fMa7dV0jqF
请立即登录并修改。
====================================================================
```
> 刻意**不做**"打开页面让你设置密码"的流程。在局域网监听下,任何能访问到的人
> 都能抢先设置密码成为管理员;自动生成 + 打印避免了这种抢占,也避免了把自己锁在外面。
### 忘记密码
设置环境变量 `MC_PASSWORD` 后重启即可:
```bash
MC_PASSWORD=我的新密码 # Linux/macOS
set MC_PASSWORD=我的新密码 # Windows cmd
```
该变量**优先级始终高于**数据库里的密码,且**不会被写入磁盘**。登录后到设置页改成正式密码即可。
### 环境变量
配置写在项目根目录的 `.env`(已被 gitignore)。
| 变量 | 默认 | 说明 |
|---|---|---|
| `MC_HOST` | `127.0.0.1` | 监听地址。**要局域网访问须设为 `0.0.0.0`** |
| `MC_PORT` | `8080` | 端口 |
| `MC_PASSWORD` | 空 | 覆盖数据库密码,忘记密码时的恢复通道 |
| `MC_COOKIE_SECURE` | 关 | **面板走 HTTPS 时才开**。局域网明文下开启会导致浏览器丢弃 Cookie,表现为**登录页反复刷新且无任何报错** |
| `MC_SESSION_TTL_HOURS` | `336` | 登录有效期(14 天) |
| `MC_TRUST_PROXY` | 关 | 仅在受信任的反向代理之后开启,否则 `X-Forwarded-For` 可被伪造以绕过登录节流 |
| `MC_CORS_ORIGINS` | 空 | 附加的允许来源,逗号分隔 |
| `MC_CORS_ORIGIN_REGEX` | 空 | 允许来源的正则,用于局域网里的 Vite 开发服务器 |
> 本项目的 `.env` 此前**从未被加载过**(代码里没有任何 `load_dotenv` 调用,尽管
> `python-dotenv` 一直是依赖、`.env.example` 也一直在仓库里)。现已修复。
### 安全边界(请务必了解)
- **局域网是明文 HTTP**,所以 Cookie 没开 `Secure`,`SameSite=Lax`。
这意味着**同网段抓包能看到会话令牌**。安全边界是"内网 + 密码",不是传输加密。
- **超出可信网络之外请走 HTTPS 反向代理**,不要把本服务直接暴露到公网。
- **登录节流是进程内的**:重启即清零。单用户单 worker 场景足够;
若日后多 worker,节流会按 worker 各算各的。反向代理下 `request.client.host` 是代理地址,
需配合 `MC_TRUST_PROXY` 才能正确识别来源。
- `/docs`、`/redoc`、`/openapi.json` **已关闭** —— 它们默认不鉴权,等于免费公开整个 API 地图。
### 会话与登出
- 会话存在服务端(`auth_session` 表),库里只存令牌的 **SHA-256**,不存令牌本身
- 退出登录、**修改密码**都会立即失效(改密码会踢掉所有设备,并给当前设备补发一个新会话)
- 令牌可放在 Cookie(浏览器自动携带,WebSocket 与文件下载都依赖它)
或 `Authorization: Bearer`(方便脚本调用)
---
## 二·九、设置页
原先挤在监控页左下角的 Cookie 与 Webhook 面板已迁到这里,并补齐了采集策略、代理与账号安全。
### 分区与生效方式
| 分区 | 内容 | 生效时机 |
|---|---|---|
| 登录态 | 小红书 Cookie | 下一轮采集 |
| 通知 | 企业微信 Webhook | 下一条推送 |
| 采集策略 | 新任务默认间隔、默认单轮上限、默认评论条数 | **仅影响新建任务** |
| 采集策略 | 请求间隔、抓二级评论 | 下一轮采集 |
| 采集策略 | 活跃时段 | 定时任务的下一次触发 |
| 代理 | 开关、提供方、池大小、静态地址 | 下一轮采集 |
| 账号安全 | 修改密码 | 立即(其他设备全部掉线) |
**活跃时段**:只在此时段内触发定时采集,窗口外任务保持到期状态、不会丢失,
窗口一开照常执行。默认 `0–23` 即全天;也支持跨午夜(如 `22–6`)。
**"仅影响新建任务"** 的那几项是刻意的:改了默认间隔不应该把已有任务的间隔一起改掉。
### 设计要点
- **敏感值永不回传**:Cookie 和 Webhook 的 `GET` 只返回「是否已配置」与长度,不返回值。
表单不会把没动过的敏感项覆盖掉。
- **部分更新**:只有请求里出现的 key 会被写入。表单一角改动不会清空其他设置。
- **设置项由后端声明**:`api/monitor/app_settings.py` 里的注册表(类型、范围、选项、默认值)
是唯一事实来源,前端**按它生成表单**。加一个设置项不需要改前端字段清单。
- **越界即拒绝**:超出范围、未知的 key、非法的枚举值都返回 400 而不是静默接受。
> 「扫码登录」入口**尚未实现**。它需要跑起爬虫子进程、捕获二维码并实时推流,
> 属于一个独立功能而非设置项,这里不做一个半成品。
---
## 二·十、平台切换与能力矩阵
**右上角的下拉框统一切换平台**,「采集 / 监控 / 报表 / 设置」全部跟着变。选择会记住,
刷新后不会跳回小红书。采集页原来那个平台下拉已移除,避免出现两个事实来源。
### 已接通 vs 未接通
矩阵里有两个**不同**的概念,混淆会误导:
| 字段 | 含义 |
|---|---|
| `crawler_modes` / `metrics` / `comment_levels` / `media` | **上游爬虫模块**能做什么 |
| `monitor_wired` | **监控层**是否已接线 |
**7 个平台的爬虫模块都实现了 search / detail / creator**,真正的差异在指标上:
| 平台 | 指标 | 评论层级 | 媒体 | 监控接线 |
|---|---|---|---|---|
| 小红书 | 点赞 / 评论 / 收藏 / 分享 | 2 | ✅ | ✅ |
| 抖音 | 点赞 / 评论 / 收藏 / 分享(**无播放量**) | 2 | ✅ | ❌ |
| 快手 | 点赞 / 播放(无评论、分享、收藏) | 1 | ✅ | ❌ |
| B站 | 点赞 / **播放** / **弹幕** / 评论 / 收藏 / 投币 / 分享(最全) | 2 | ✅ | ❌ |
| 微博 | 点赞 / 评论 / 转发(无收藏) | 2 | ✅ | ❌ |
| 贴吧 | 仅回复数 | 2 | ❌ | ❌ |
| 知乎 | 赞同 / 评论 | 2 | ❌ | ❌ |
> **要更正一个常见误解**:这个代码库里**抖音不存播放量**(只映射点赞/收藏/评论/分享)。
> 有播放量的是 **B 站**,它还有弹幕。
未接通的平台**可以选,但各页会显示明确的说明面板**,并且**创建任务会被直接拒绝**:
```
400 B站的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。
```
而不是接受任务、然后让它永远跑不出数据 —— 那正是之前"博主主页解析失败被误报成登录失效"的同一种静默故障。
### 设置的两层
| 位置 | 范围 | 内容 |
|---|---|---|
| 左侧导航「设置」 | **按平台** | 登录 Cookie、采集策略、代理 |
| 右上角「系统设置」 | **全局** | 通知、活跃时段、上游更新、账号安全 |
**这不是随便分的**:企业微信只有一个群、调度器只有一套时段规则、密码只有一份 ——
把它们放进"小红书专属"的页面里,会让人以为它们是按平台存的。
存储上键名带作用域前缀:`platform.<平台>.<项>` 与 `system.<项>`。
**旧键会在启动时自动迁移**(`xhs_cookie` → `platform.xhs.cookie`),
且是幂等的:新键已存在时以新键为准,不会覆盖你后来改的值。
---
## 二·十一、上游更新检查
本仓库在 [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) 之上加了一整层
(监控 / 鉴权 / 多平台面板),差异管理与合并流程在根目录 `UPSTREAM.md` 里。但那份流程有个
隐含前提:**得有人知道上游动了**。部署脚本只从我们自己的 Gitea `git pull`,上游的提交不主动
去 fetch 就永远看不见 —— 拖着不合并的代价是复利的,越久越难合。
这一项就是替你定时去 fetch 的:按间隔(默认每天一次)拉一次上游,算出「当前部署落后几个
提交」,有更新就推一条企业微信,并把结果与提交列表显示在**右上角「系统设置」→「上游更新」**。
### 配置
| 项 | 默认 | 说明 |
|---|---|---|
| 检查上游仓库更新 | **关** | 总开关。默认关:它要联网 fetch,且需要容器里有 git(见下) |
| 上游检查间隔(分钟) | 1440 | 每天一次。最小 30 分钟 |
| 上游仓库地址 | GitHub 上游 | 国内直连 GitHub 不稳时改成 gitcode 镜像,见 `UPSTREAM.md` |
| 上游分支 | `main` | |
| 上游有更新时推送通知 | 开 | 只在出现**此前没推过**的提交时发一条,同一个更新不会反复推 |
### 几个刻意的行为
- **只读,不写工作区**:只 `git fetch <地址> <分支>` 到 `FETCH_HEAD` —— 不建 remote、不写
`refs/remotes`、不碰索引与工作区。所以它不会打断正在跑的采集,也不会和 `./deploy.sh`
的 `git pull` 抢锁。
- **不受活跃时段限制**:活跃时段是给采集定的(避免半夜去抓平台)。检查只是 fetch 一个公开
仓库,半夜跑反而更合适。
- **失败也是一种结果**:上游不通(尤其直连 GitHub)很常见。界面会显示失败原因与上次检查
时间,失败不推送,也**不会**因此改变下一次检查的时间 —— 每个间隔重试一次,而不是每个
调度 tick(20 秒)都去撞一次。
- **同一个更新只推一次**:推送状态记的是上游 tip。推过之后,上下游没动就不会再推;上游又
有新提交(tip 变了)时会再推一条。
- **「立即检查」不发通知**:点这个按钮的人正看着结果,没必要再给自己推一条群消息。那次
检查只写结果,没推的那批提交留给下一次定时检查推。
### 部署前提:镜像里要有 git
`python:3.11-slim` 不带 git,`Dockerfile` 里已显式安装。**因此这次更新需要重建镜像**:
```bash
docker compose build && ./deploy.sh
```
`./deploy.sh` 只重建前端,不会重建镜像。漏了这步的话,检查会报「未找到 git 命令」——
界面上看得见,不会静默。
---
## 三、必须知道的限制
### 1. 「新增评论」是近似值 —— 最重要的一条
小红书评论接口 `/api/sns/web/v2/comment/page` **没有排序参数**,只能拿到平台默认排序(热评优先)的
前 N 条。因此:
- 我们只能"每次抓前 N 条做差集",**新发布但沉底的评论不会被发现**
- N 调大能提高发现率,但请求量线性增长,风控风险上升
- 评论事件区分两种,UI 上也分别标注:
- `new_comment_posted`(新评论):`create_time` 晚于上一轮开始时间,是真·新发布
- `new_comment_seen`(新出现评论):只是本轮才进入可见窗口的历史评论
**这条限制无法通过调参绕过**,是该接口的固有限制。
### 2. Cookie 失效是「静默失败」
`login_by_cookies()` 只注入 `web_session`,而 API 签名还需要 `a1` / `webId` 等;
更麻烦的是**cookie 登录不做任何校验** —— 坏 Cookie 不会让进程报错退出,而是
**退出码 0、抓到 0 条**。
监控层因此把「退出码 0 且 0 条作品」判定为 `suspected_auth_failure` 并在 UI 上标红,
而不是当成"该博主没发新作品"。这是无人值守场景最容易误报的地方。
本实现额外做了两件事:
- 通过 `--inject_all_cookies` 注入**完整** Cookie(默认关闭,保持上游行为不变)
- 通过 `--cookies_file` 传 Cookie,避免明文出现在进程列表里
### 3. 作品窗口被截断
`每轮最多采集作品数`(默认 20)限定了"该博主的作品"到底指多少条。
UI 会把该上限显示在作品表旁,避免误以为看到了全部。
### 4. 昵称与用户 ID 已被上游脱敏
`store/xhs/__init__.py` 落库前调用 `mask_nickname()` 与 `anonymize_user_id()`,
存储的是**打码昵称**与哈希后的 `creator_hash`,没有真实昵称和 user_id。
这是上游的隐私保护设计,监控层未做改动。
### 5. 不发「笔记被删」事件
`creator` 模式只取前 N 条,笔记"消失"多半只是掉出窗口;`detail` 模式遇到
`xsec_token` 过期也会失败。二者与"真被删"无法区分,因此不产生删除事件,
改为在作品表里展示 `last_seen_at`。
### 6. 定时任务与手动采集互斥
二者共用同一个爬虫子进程。监控任务运行期间点「采集」会被拒绝(返回 400);
反之若有手动采集在跑,到期的监控任务会**保持到期状态排队**,不会丢失,空闲后自动补上。
---
## 四、数据存放
| 内容 | 位置 |
|---|---|
| 监控库(任务/快照/事件/评论/设置/会话) | **MySQL**,库名由 `MYSQL_DB_NAME` 指定 |
| 每轮原始 jsonl | `data/monitor_runs/{task_id}/{run_id}/{platform}/jsonl/` |
> 爬虫每轮的原始产出**仍然写独立 jsonl 目录**,不进 MySQL。
> 这是差分机制的基础:每轮写在单独目录里,才能算出"这轮新增了什么"。
> 多轮数据混在同一批表里的话,这个判断就做不到了。
### MySQL 配置与安全边界
连接信息写在 `.env`(已被 gitignore,不会进版本库):
```ini
MYSQL_DB_HOST=<数据库地址>
MYSQL_DB_PORT=3306
MYSQL_DB_USER=<账号>
MYSQL_DB_PWD=<密码>
MYSQL_DB_NAME=mediacrawler
```
> 真实凭据只写在 `.env` 里(已被 gitignore),**不要写进这个文档或任何会提交的文件**。
**"只操作这个库"由两层保证,缺一不可**:
1. **数据库授权(真正的保证)**。账号应只被授予目标库的权限:
```sql
REVOKE ALL PRIVILEGES, GRANT OPTION FROM 'MediaCrawler'@'%';
GRANT ALL PRIVILEGES ON `mediacrawler`.* TO 'MediaCrawler'@'%';
FLUSH PRIVILEGES;
```
这样该账号 `SHOW DATABASES` 只能看到目标库,**代码就算写错也碰不到别的库**。
2. **启动自检(防配置写错)**。应用启动时会执行 `SELECT DATABASE()`,
与 `MYSQL_DB_NAME` 不符就**拒绝启动**,而不是往错误的库里写。
**字符集**:这台服务的服务端和库默认都是 `latin1`。代码在建表时**逐表强制
`utf8mb4`**,不依赖库默认值 —— 否则中文会被拒或变成问号。
**连接保活**:MySQL 默认 8 小时断开空闲连接,而监控服务是常驻的。
已配置 `pool_recycle=3600` + `pool_pre_ping`,避免"server has gone away"。
**表引擎**:全部 InnoDB(`monitor_run.exit_code` 用 `BIGINT` —— Windows 的退出码是
无符号 32 位,`0xC0000142` 会溢出有符号 `INT`)。
### 从 SQLite 迁移(如有旧数据)
```bash
python -m api.monitor.migrate_from_sqlite --dry-run # 先看要迁什么
python -m api.monitor.migrate_from_sqlite # 正式迁移
```
保留原主键(否则 `task_id` 关联会错位);目标库非空时会拒绝执行,除非加 `--force`。
监控库中的 Cookie 为明文存储,这是当前版本的已知取舍。
---
## 五、API
所有操作都有对应的 HTTP 接口,UI 只是其中一层封装:
```
GET /api/monitor/overview 看板汇总
GET /api/monitor/tasks 任务列表
POST /api/monitor/tasks 新建任务
PATCH /api/monitor/tasks/{id} 修改
DELETE /api/monitor/tasks/{id} 删除
POST /api/monitor/tasks/{id}/run 立即运行(后台执行,立即返回)
GET /api/monitor/tasks/{id}/runs 运行历史
GET /api/monitor/notes 作品表(含与上一轮的 Δ)
GET /api/monitor/notes/{id}/series 单篇指标时间序列
GET /api/monitor/comments 评论流(带所属作品;?note_id= 筛选,?group_by=note 按作品分组)
GET /api/monitor/comment-notes 有评论的作品及其条数(评论筛选下拉用)
GET /api/monitor/export 导出(?kind=notes|comments|report&format=csv|xlsx)
GET /api/monitor/events 变化事件流
POST /api/monitor/events/read 标记已读
GET /api/monitor/cookie 登录态健康度(**不返回 Cookie 值**)
POST /api/monitor/cookie 保存 Cookie
DELETE /api/monitor/cookie 清除 Cookie
GET /api/config/platforms 平台能力矩阵(含 monitor_wired,前端据此渲染切换器)
GET /api/settings 设置 + 表单描述(?platform=,敏感值只回状态)
PUT /api/settings 部分更新(只写请求里出现的 key)
GET /api/auth/me 身份探测(401 即未登录)
POST /api/auth/login 登录(发 HttpOnly Cookie)
POST /api/auth/logout 退出
POST /api/auth/password 修改密码(踢掉所有其他设备)
GET /api/monitor/report 报表(?task_id=1&task_id=2&start_date=&end_date=)
GET /api/monitor/webhook 通知配置状态(**只返回打码地址**)
POST /api/monitor/webhook 保存 Webhook 地址
DELETE /api/monitor/webhook 删除 Webhook
POST /api/monitor/webhook/test 发送测试消息
GET /api/monitor/upstream 最近一次上游检查的缓存结果(没查过返回 {})
POST /api/monitor/upstream/check 立刻检查一次(等 fetch 跑完才返回,**不发通知**)
```
> `task_id` 用**重复参数**而非逗号拼接(`?task_id=1&task_id=2`);不传表示统计全部任务。
---
## 六、故障排查
| 现象 | 原因 / 处理 |
|---|---|
| 任务一直不运行 | 未配置 Cookie(调度器会跳过并保持任务到期);或全局已有采集在跑 |
| 任务标红「疑似登录态失效」 | Cookie 过期。重新粘贴;若反复失败,先手动扫码登录一次播种浏览器 profile |
| 抓到的作品数长期为 0 | 同上;也可能是该博主确实没有作品 |
| 发现不了新评论 | 评论接口无时间排序所致,调大「每篇评论抓取条数」可缓解但无法根治 |
| 首轮没有任何"新增"事件 | 刻意设计:首轮建立基线,全部数据视为已有,不产生变化事件 |
+26 -1
View File
@@ -39,6 +39,10 @@ from .exception import *
from .field import *
from .help import *
# 抖音边缘网关 ArgusSecurityPlugin 要求的请求头。网关目前不校验取值,
# 传固定字符串即可;将来若开始真校验,会重新出现 "Signature Not Found"。
DOUYIN_ARGUS_HEADER_VALUE = "1"
class DouYinClient(AbstractApiClient, ProxyRefreshMixin):
@@ -55,6 +59,15 @@ class DouYinClient(AbstractApiClient, ProxyRefreshMixin):
self.proxy = proxy
self.timeout = timeout
self.headers = headers
# 抖音边缘网关的 ArgusSecurityPlugin 会对这批接口做业务前置校验,缺少
# x-tt-argus 头时直接 403,响应体为
# "Blocked by ArgusSecurityPlugin Uifid Not Found"(补了 uifid 但没这个头则是
# "... Signature Not Found")。当前网关尚未校验该头的值,可传任意字符串;
# 一旦升级到真校验,需要改为 WebView 内注入 JS 让页面自带 SDK 补齐。
self.headers.setdefault("x-tt-argus", DOUYIN_ARGUS_HEADER_VALUE)
uifid = cookie_dict.get("UIFID") or cookie_dict.get("UIFID_TEMP", "")
if uifid:
self.headers.setdefault("uifid", uifid)
self._host = "https://www.douyin.com"
self.cookie_urls = [
"https://douyin.com",
@@ -214,7 +227,19 @@ class DouYinClient(AbstractApiClient, ProxyRefreshMixin):
:param aweme_id:
:return:
"""
params = {"aweme_id": aweme_id}
# 抖音 detail 接口的 Argus 风控要求这两个参数成套出现,缺一则直接 403
# (响应体为 "Blocked by ArgusSecurityPlugin Uifid Not Found"):
# uifid = UIFID cookie,没有时退到 UIFID_TEMP
# verifyFp / fp = s_v_web_id cookie
# 必须用 cookie 里的 s_v_web_id:实测 uifid 搭配自生成的 verifyFp 会被判成
# "Signature Not Found",两者同源才能通过。
s_v_web_id = self.cookie_dict.get("s_v_web_id", "")
params = {
"aweme_id": aweme_id,
"uifid": self.cookie_dict.get("UIFID") or self.cookie_dict.get("UIFID_TEMP", ""),
"verifyFp": s_v_web_id,
"fp": s_v_web_id,
}
headers = copy.copy(self.headers)
del headers["Origin"]
res = await self.get("/aweme/v1/web/aweme/detail/", params, headers)
+18 -2
View File
@@ -44,7 +44,11 @@ from . import media as douyin_media
from .client import DouYinClient
from .exception import DataFetchError
from .field import PublishTimeType
from .help import parse_video_info_from_url, parse_creator_info_from_url
from .help import (
client_hint_headers,
parse_creator_info_from_url,
parse_video_info_from_url,
)
from .login import DouYinLogin
@@ -98,7 +102,12 @@ class DouYinCrawler(AbstractCrawler):
await self.browser_context.add_init_script(path="libs/stealth.min.js")
self.context_page = await self.browser_context.new_page()
await self.context_page.goto(self.index_url)
# wait_until="domcontentloaded" instead of the default "load": the douyin
# home page never fires the load event (some long-lived request keeps it
# pending), so the default burns the whole timeout and the crawl dies
# before it starts. Measured here: domcontentloaded returns in 0.7s while
# load still times out at 90s. Tieba and Zhihu already do the same.
await self.context_page.goto(self.index_url, wait_until="domcontentloaded")
self.dy_client = await self.create_douyin_client(httpx_proxy_format)
if not await self.dy_client.pong(browser_context=self.browser_context):
@@ -315,10 +324,17 @@ class DouYinCrawler(AbstractCrawler):
self.browser_context,
urls=self.cookie_urls,
) # type: ignore
# 声称自己是 Chrome,就得带上 sec-ch-ua 系列头 —— 浏览器一定会带,而缺了它们
# 的请求在抖音网关看来就是机器人:回一个 **200 + 空 body**,不报错、不给原因,
# 表现为采集抓到 0 条。见 help.client_hint_headers 的实测记录。
client_hints = client_hint_headers(
await self.context_page.evaluate("() => navigator.userAgentData || null")
)
douyin_client = DouYinClient(
proxy=httpx_proxy,
headers={
"User-Agent": await self.context_page.evaluate("() => navigator.userAgent"),
**client_hints,
"Cookie": cookie_str,
"Host": "www.douyin.com",
"Origin": "https://www.douyin.com/",
+33 -1
View File
@@ -26,7 +26,7 @@
import random
import re
from typing import Optional
from typing import Dict, Optional
import execjs
from playwright.async_api import Page
@@ -98,6 +98,38 @@ async def get_a_bogus_from_playwright(params: str, post_data: dict, user_agent:
return a_bogus
def client_hint_headers(user_agent_data) -> Dict[str, str]:
"""由 ``navigator.userAgentData`` 还原 ``sec-ch-ua`` 系列请求头。
浏览器只要声称自己是 Chrome,就**一定会**带这三个头。缺了它们,「Chrome 的 UA +
没有 sec-ch-ua」就是最典型的机器人特征 —— 抖音网关会因此返回 **200 + 空 body**:
不报错、不给原因、HTTP 状态还是成功的,表现为采集拿到 0 条。
实测(同一 URL、同一 cookie、同一参数):不带头 → 0 字节;补上这三个头 → 7077 字节。
从 ``userAgentData`` 现算而不是写死,是为了 Chrome 升级后不会悄悄失配 —— 写死的
版本号和 UA 里的版本号一旦对不上,就又是一个可疑特征。
"""
if not isinstance(user_agent_data, dict):
return {}
brands = user_agent_data.get("brands") or []
sec_ch_ua = ", ".join(
f'"{brand.get("brand", "")}";v="{brand.get("version", "")}"' for brand in brands
)
if not sec_ch_ua:
return {}
headers = {
"sec-ch-ua": sec_ch_ua,
"sec-ch-ua-mobile": "?1" if user_agent_data.get("mobile") else "?0",
}
platform = user_agent_data.get("platform")
if platform:
headers["sec-ch-ua-platform"] = f'"{platform}"'
return headers
def parse_video_info_from_url(url: str) -> VideoUrlInfo:
"""
Parse video ID from Douyin video URL
+11
View File
@@ -272,3 +272,14 @@ class DouYinLogin(AbstractLogin):
'domain': ".douyin.com",
'path': "/"
}])
# Reload after injecting. The page was loaded *before* these cookies existed,
# so its `localStorage.HasUserLogin` still holds the logged-out value and
# check_login_state() polls it until the timeout expires (10 minutes) without
# ever succeeding -- only the *next* run works, because by then the cookies
# are in the profile. Reloading makes the site re-evaluate the session now.
try:
await self.context_page.reload(wait_until="domcontentloaded")
except Exception as exc: # pragma: no cover - reload is best effort
utils.logger.warning(
f"[DouYinLogin.login_by_cookies] reload after cookie injection failed: {exc}"
)
+31
View File
@@ -240,6 +240,37 @@ class KuaiShouClient(AbstractApiClient, ProxyRefreshMixin):
}
return await self.post("", post_data)
async def resolve_short_url(self, short_url: str) -> str:
"""解析快手分享短链(/f/xxx),返回重定向后的真实 URL。
短链路径里的 share_token 不是视频 ID,只能靠 302 的 Location 拿到
/short-video/<id> 形式的真实地址。
"""
async with make_async_client(proxy=self.proxy, follow_redirects=False) as client:
try:
utils.logger.info(
f"[KuaiShouClient.resolve_short_url] Resolving short URL: {short_url}"
)
response = await client.get(short_url, timeout=10, headers=self.headers)
# 短链通常返回 302
if response.status_code in (301, 302, 303, 307, 308):
redirect_url = response.headers.get("Location", "")
utils.logger.info(
f"[KuaiShouClient.resolve_short_url] Resolved to: {redirect_url}"
)
return redirect_url
utils.logger.warning(
f"[KuaiShouClient.resolve_short_url] Unexpected status code: {response.status_code}"
)
return ""
except Exception as e:
utils.logger.error(
f"[KuaiShouClient.resolve_short_url] Failed to resolve short URL: {e}"
)
return ""
async def get_video_info(self, photo_id: str) -> Dict:
"""
Kuaishou web video detail api
+40 -8
View File
@@ -204,6 +204,24 @@ class KuaishouCrawler(AbstractCrawler):
for video_url in config.KS_SPECIFIED_ID_LIST:
try:
video_info = parse_video_info_from_url(video_url)
# 分享短链(/f/xxx)要先跟随 302 重定向,才能拿到 /short-video/<id>
if video_info.url_type == "short":
utils.logger.info(
f"[KuaishouCrawler.get_specified_videos] Resolving short link: {video_url}"
)
resolved_url = await self.ks_client.resolve_short_url(video_url)
if resolved_url:
video_info = parse_video_info_from_url(resolved_url)
utils.logger.info(
f"[KuaishouCrawler.get_specified_videos] Short link resolved to video ID: {video_info.video_id}"
)
else:
utils.logger.error(
f"[KuaishouCrawler.get_specified_videos] Failed to resolve short link: {video_url}"
)
continue
video_ids.append(video_info.video_id)
utils.logger.info(f"Parsed video ID: {video_info.video_id} from {video_url}")
except ValueError as e:
@@ -237,15 +255,29 @@ class KuaishouCrawler(AbstractCrawler):
utils.logger.info(f"[KuaishouCrawler.get_video_info_task] Sleeping for {sleep_sec:.1f} seconds after fetching video details {video_id}")
detail = result.get("visionVideoDetail")
if detail:
photo = detail.get("photo", {})
author = detail.get("author", {})
utils.logger.info(
f"[KuaishouCrawler.get_video_info_task] video detail: "
f"id={photo.get('id', video_id)} author={author.get('name', '')} "
f"likes={photo.get('likeCount', '')} views={photo.get('viewCount', '')} "
f"caption={str(photo.get('caption', ''))[:50]}"
if not detail:
return None
# 快手对不可用视频(已删除/私密/不存在)返回的是
# visionVideoDetail: {photo: null, author: null}——key 在、值是 null。
# 注意 .get("photo", {}) 只在 key **缺失** 时给默认值,key 存在且为 null
# 时拿到的仍是 None,接着 .get() 就抛 AttributeError,而
# asyncio.gather 不会拦住它,整轮爬取会直接带崩。
photo = detail.get("photo") or {}
if not photo:
utils.logger.warning(
f"[KuaishouCrawler.get_video_info_task] 视频不可用"
f"(photo 为空,可能已删除或私密),跳过 video_id={video_id}"
)
return None
author = detail.get("author") or {}
utils.logger.info(
f"[KuaishouCrawler.get_video_info_task] video detail: "
f"id={photo.get('id', video_id)} author={author.get('name', '')} "
f"likes={photo.get('likeCount', '')} views={photo.get('viewCount', '')} "
f"caption={str(photo.get('caption', ''))[:50]}"
)
return detail
except DataFetchError as ex:
utils.logger.error(
+10
View File
@@ -92,6 +92,8 @@ def parse_video_info_from_url(url: str) -> VideoUrlInfo:
Supports the following formats:
1. Full video URL: "https://www.kuaishou.com/short-video/3x3zxz4mjrsc8ke?authorId=3x84qugg4ch9zhs&streamSource=search"
2. Pure video ID: "3x3zxz4mjrsc8ke"
3. Share short link: "https://www.kuaishou.com/f/X9Idt15MQb9L2cv"
(路径里是 share_token 而非视频 ID,返回 url_type="short",由调用方跟随重定向)
Args:
url: Kuaishou video link or video ID
@@ -109,6 +111,14 @@ def parse_video_info_from_url(url: str) -> VideoUrlInfo:
video_id = match.group(1)
return VideoUrlInfo(video_id=video_id, url_type="normal")
# 分享短链:https://www.kuaishou.com/f/X9Idt15MQb9L2cv
# 路径里的 share_token 不是视频 ID,必须跟随 302 重定向才能拿到真实地址,
# 所以这里只标记类型,交给调用方解析(url_type="short")
share_pattern = r'kuaishou\.com/f/([a-zA-Z0-9_-]+)'
match = re.search(share_pattern, url)
if match:
return VideoUrlInfo(video_id=match.group(1), url_type="short")
raise ValueError(f"Unable to parse video ID from URL: {url}")
+20 -5
View File
@@ -201,9 +201,22 @@ class XiaoHongShuCrawler(AbstractCrawler):
# Parse creator URL to get user_id and security tokens
creator_info: CreatorUrlInfo = parse_creator_info_from_url(creator_url)
utils.logger.info(f"[XiaoHongShuCrawler.get_creators_and_notes] Parse creator URL info: {creator_info}")
user_id = creator_info.user_id
except ValueError as e:
utils.logger.error(f"[XiaoHongShuCrawler.get_creators_and_notes] Failed to parse creator URL: {e}")
continue
# get creator detail info from web html content
user_id = creator_info.user_id
# Fetching the profile page is best-effort and must not abort the run.
# It only feeds save_creator(), which is a no-op in this build, while
# the notes themselves come from a completely different endpoint.
# Scraping the profile means parsing window.__INITIAL_STATE__ out of
# HTML, which fails whenever the platform serves a different page --
# a JSONDecodeError there is especially misleading because it is a
# ValueError subclass, so it used to be reported as "failed to parse
# creator URL" and then skipped the creator entirely, yielding zero
# notes for a perfectly valid target.
try:
createor_info: Dict = await self.xhs_client.get_creator_info(
user_id=user_id,
xsec_token=creator_info.xsec_token,
@@ -211,9 +224,6 @@ class XiaoHongShuCrawler(AbstractCrawler):
)
if createor_info:
await xhs_store.save_creator(user_id, creator=createor_info)
except ValueError as e:
utils.logger.error(f"[XiaoHongShuCrawler.get_creators_and_notes] Failed to parse creator URL: {e}")
continue
except (IPBlockError, PlatformAccessError) as e:
# Access restricted on the creator homepage, skip this creator instead of crashing the run.
utils.logger.error(
@@ -221,6 +231,11 @@ class XiaoHongShuCrawler(AbstractCrawler):
f"建议降低采集频率、更换 IP 或检查账号状态"
)
continue
except Exception as e:
utils.logger.warning(
f"[XiaoHongShuCrawler.get_creators_and_notes] Could not fetch profile for {user_id} "
f"({type(e).__name__}: {e}); continuing to fetch the creator's notes anyway"
)
# Use fixed crawling interval
crawl_interval = config.CRAWLER_MAX_SLEEP_SEC
+10 -1
View File
@@ -213,8 +213,12 @@ class XiaoHongShuLogin(AbstractLogin):
async def login_by_cookies(self):
"""login xiaohongshu website by cookies"""
utils.logger.info("[XiaoHongShuLogin.login_by_cookies] Begin login xiaohongshu by cookie ...")
injected = 0
for key, value in utils.convert_str_cookie_to_dict(self.cookie_str).items():
if key != "web_session": # Only set web_session cookie attribute
# Default (upstream) behaviour injects only web_session. Unattended runs
# need a1 / webId as well, otherwise signed API calls fail and the run
# exits 0 having fetched nothing -- a silent failure.
if not config.INJECT_ALL_COOKIES and key != "web_session":
continue
await self.browser_context.add_cookies([{
'name': key,
@@ -222,3 +226,8 @@ class XiaoHongShuLogin(AbstractLogin):
'domain': ".rednote.com" if config.XHS_INTERNATIONAL else ".xiaohongshu.com",
'path': "/"
}])
injected += 1
utils.logger.info(
f"[XiaoHongShuLogin.login_by_cookies] Injected {injected} cookie(s), "
f"inject_all_cookies={config.INJECT_ALL_COOKIES}"
)
+5 -1
View File
@@ -28,4 +28,8 @@ motor>=3.3.0
openpyxl>=3.1.2
pytest>=7.4.0
pytest-asyncio>=0.21.0
xhshow>=0.2.0
xhshow>=0.2.0
# Required by uvicorn to handle WebSocket upgrades. Declared in pyproject.toml
# but previously missing here, so installing from this file left the live log
# stream silently non-functional (uvicorn answers every upgrade with 404).
websockets>=15.0.1
+29
View File
@@ -96,3 +96,32 @@ def sample_xhs_creator():
"interaction": 50000,
"tag_list": '{"profession": "Designer", "interest": "Photography"}'
}
@pytest.fixture(autouse=True)
def _bypass_auth_for_non_auth_suites(request):
"""Skip API authentication for suites that are not about authentication.
Adding auth to every /api route breaks any test that speaks HTTP, so those
suites override the dependency here. This uses FastAPI's own
``dependency_overrides`` mechanism rather than a production-visible
"test mode" switch, which could be shipped enabled by accident.
``tests/test_auth.py`` is deliberately excluded: it must exercise the real
enforcement path, including the route-enumeration guard that asserts every
other /api route really does return 401.
"""
if request.node.fspath.basename == "test_auth.py":
yield
return
from api.auth import require_auth, require_ws_auth
from api.main import app
app.dependency_overrides[require_auth] = lambda: None
app.dependency_overrides[require_ws_auth] = lambda: None
try:
yield
finally:
app.dependency_overrides.pop(require_auth, None)
app.dependency_overrides.pop(require_ws_auth, None)
+31
View File
@@ -26,6 +26,37 @@ async def test_cmd_arg_crawler_max_notes_count():
config.CRAWLER_MAX_NOTES_COUNT = orig_notes
config.CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES = orig_comments
def test_douyin_monitor_command_uses_the_right_flags():
"""抖音监控任务拼出来的命令行。
与 runner 走的是同一条 _build_command 路径,所以这一条能守住「监控任务的参数
没拼错」—— 尤其是平台值必须是 dy(而不是 douyin),否则上游根本认不出平台。
"""
cm = CrawlerManager()
req = CrawlerStartRequest(
platform=PlatformEnum.DOUYIN,
login_type=LoginTypeEnum.COOKIE,
crawler_type=CrawlerTypeEnum.CREATOR,
creator_ids="https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X",
save_data_path="./data/monitor_runs/1/2",
enable_cdp_mode=True,
inject_all_cookies=True,
save_login_state=True,
max_notes_count=20,
max_comments_count=50,
)
cmd = cm._build_command(req)
idx = cmd.index("--platform")
assert cmd[idx + 1] == "dy"
idx = cmd.index("--type")
assert cmd[idx + 1] == "creator"
idx = cmd.index("--creator_id")
assert cmd[idx + 1] == "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X"
idx = cmd.index("--enable_cdp_mode")
assert cmd[idx + 1] == "true"
def test_crawler_manager_build_command():
cm = CrawlerManager()
+503
View File
@@ -0,0 +1,503 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_auth.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Tests for WebUI authentication.
Deliberately does NOT install ``app.dependency_overrides``: the point of this
file is to exercise the real enforcement path. Other suites override
``require_auth`` so they can keep testing their own concerns.
"""
import asyncio
import time
import httpx
import pytest
import pytest_asyncio
from fastapi import WebSocketException
from sqlalchemy import func, select
from api import auth
from api.main import app
from api.monitor import db as monitor_db
from api.monitor.models import AuthSession
PASSWORD = "correct-horse-battery"
# Captured at import, i.e. before the autouse fixture patches the module global,
# so the guard test below checks the value that actually ships.
REAL_PBKDF2_ITERATIONS = auth.PBKDF2_ITERATIONS
# Every /api route that is allowed to answer without a session.
EXEMPT_PATHS = {"/api/health", "/api/auth/login", "/api/auth/logout"}
@pytest.fixture(autouse=True)
def cheap_hashing(monkeypatch):
"""600k iterations is right in production and unusable in a test suite.
hash_password() resolves the count at call time precisely so this works.
"""
monkeypatch.setattr(auth, "PBKDF2_ITERATIONS", 1_000)
monkeypatch.delenv("MC_PASSWORD", raising=False)
auth.reset_throttle_state()
yield
auth.reset_throttle_state()
@pytest_asyncio.fixture
async def db(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
yield monitor_db
await monitor_db.dispose_engine()
@pytest_asyncio.fixture
async def client(db):
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
async def _seed_password(password: str = PASSWORD) -> None:
async with monitor_db.get_session() as session:
await auth.set_password(session, password)
# --------------------------------------------------------------------------
# Password hashing
# --------------------------------------------------------------------------
class TestPasswordHashing:
def test_iteration_count_has_not_been_lowered(self):
"""Guard: someone trimming this for speed would weaken every install."""
assert REAL_PBKDF2_ITERATIONS >= 600_000
def test_round_trip(self):
stored = auth.hash_password(PASSWORD)
assert auth._verify_password_sync(PASSWORD, stored) is True
def test_wrong_password_rejected(self):
stored = auth.hash_password(PASSWORD)
assert auth._verify_password_sync("wrong", stored) is False
def test_same_password_hashes_differently(self):
"""A fixed salt would let one rainbow table crack every install."""
assert auth.hash_password(PASSWORD) != auth.hash_password(PASSWORD)
def test_format_is_self_describing(self):
algo, iterations, salt, digest = auth.hash_password(PASSWORD).split("$")
assert algo == "pbkdf2_sha256"
assert int(iterations) == auth.PBKDF2_ITERATIONS
assert salt and digest
@pytest.mark.parametrize("stored", ["", "garbage", "md5$1$a$b", "pbkdf2_sha256$x$a$b"])
def test_malformed_stored_hash_is_rejected_not_raised(self, stored):
assert auth._verify_password_sync(PASSWORD, stored) is False
# --------------------------------------------------------------------------
# Credentials
# --------------------------------------------------------------------------
class TestCredentials:
@pytest.mark.asyncio
async def test_check_password_against_stored_hash(self, db):
await _seed_password()
async with monitor_db.get_session() as session:
assert await auth.check_password(session, PASSWORD) is True
assert await auth.check_password(session, "nope") is False
@pytest.mark.asyncio
async def test_no_password_configured_denies_everything(self, db):
"""An unset credential must not mean "open"."""
async with monitor_db.get_session() as session:
assert await auth.check_password(session, "") is False
assert await auth.check_password(session, PASSWORD) is False
@pytest.mark.asyncio
async def test_env_override_wins_and_is_not_persisted(self, db, monkeypatch):
"""The documented way back in after forgetting the password."""
await _seed_password("stored-password")
monkeypatch.setenv("MC_PASSWORD", "env-password")
async with monitor_db.get_session() as session:
assert await auth.check_password(session, "env-password") is True
assert await auth.check_password(session, "stored-password") is False
# Override must never be written to disk.
assert await auth.current_password_hash(session) != ""
assert "env-password" not in (await auth.current_password_hash(session))
@pytest.mark.asyncio
async def test_first_run_generates_a_credential(self, db, monkeypatch):
monkeypatch.delenv("MC_PASSWORD", raising=False)
generated = await auth.ensure_initial_credential()
assert generated
# Second call is a no-op.
assert await auth.ensure_initial_credential() is None
async with monitor_db.get_session() as session:
assert await auth.check_password(session, generated) is True
@pytest.mark.asyncio
async def test_first_run_defers_to_env_password(self, db, monkeypatch):
monkeypatch.setenv("MC_PASSWORD", "env-password")
assert await auth.ensure_initial_credential() is None
async with monitor_db.get_session() as session:
assert await auth.current_password_hash(session) == ""
# --------------------------------------------------------------------------
# Sessions
# --------------------------------------------------------------------------
class TestSessions:
@pytest.mark.asyncio
async def test_round_trip(self, db):
async with monitor_db.get_session() as session:
token, expires_at = await auth.create_session(session)
async with monitor_db.get_session() as session:
assert await auth.resolve_session(session, token) is not None
assert expires_at > 0
@pytest.mark.asyncio
async def test_only_the_hash_is_stored(self, db):
"""A database leak must not hand over live sessions."""
async with monitor_db.get_session() as session:
token, _ = await auth.create_session(session)
async with monitor_db.get_session() as session:
stored = (await session.scalars(select(AuthSession.token_hash))).all()
assert token not in stored
assert auth._hash_token(token) in stored
@pytest.mark.asyncio
async def test_expired_session_is_rejected_and_removed(self, db):
async with monitor_db.get_session() as session:
token, _ = await auth.create_session(session)
row = await session.get(AuthSession, auth._hash_token(token))
row.expires_at = 1 # long past
async with monitor_db.get_session() as session:
assert await auth.resolve_session(session, token) is None
# Fresh session: the identity map in the one above still holds the
# pending-delete object, so it would answer as if the row were present.
async with monitor_db.get_session() as session:
assert await session.get(AuthSession, auth._hash_token(token)) is None
@pytest.mark.asyncio
async def test_unknown_token_is_rejected(self, db):
async with monitor_db.get_session() as session:
assert await auth.resolve_session(session, "never-issued") is None
assert await auth.resolve_session(session, "") is None
@pytest.mark.asyncio
async def test_logout_revokes_only_that_session(self, db):
async with monitor_db.get_session() as session:
first, _ = await auth.create_session(session)
second, _ = await auth.create_session(session)
async with monitor_db.get_session() as session:
await auth.revoke_session(session, first)
async with monitor_db.get_session() as session:
assert await auth.resolve_session(session, first) is None
assert await auth.resolve_session(session, second) is not None
@pytest.mark.asyncio
async def test_revoke_all_clears_every_session(self, db):
async with monitor_db.get_session() as session:
await auth.create_session(session)
await auth.create_session(session)
async with monitor_db.get_session() as session:
removed = await auth.revoke_all_sessions(session)
assert removed == 2
async with monitor_db.get_session() as session:
assert await session.scalar(select(func.count()).select_from(AuthSession)) == 0
# --------------------------------------------------------------------------
# Throttle
# --------------------------------------------------------------------------
class TestThrottle:
@pytest.mark.asyncio
async def test_below_threshold_is_not_throttled(self, db):
key = "1.2.3.4"
for _ in range(auth.THROTTLE_THRESHOLD - 1):
await auth.record_failure(key)
assert await auth.retry_after_seconds(key) == 0
@pytest.mark.asyncio
async def test_lockout_after_repeated_failures(self, db):
key = "1.2.3.4"
for _ in range(auth.THROTTLE_THRESHOLD):
await auth.record_failure(key)
assert await auth.retry_after_seconds(key) > 0
@pytest.mark.asyncio
async def test_success_clears_failures(self, db):
key = "1.2.3.4"
for _ in range(auth.THROTTLE_THRESHOLD):
await auth.record_failure(key)
await auth.clear_failures(key)
assert await auth.retry_after_seconds(key) == 0
@pytest.mark.asyncio
async def test_failures_age_out_of_the_window(self, db, monkeypatch):
"""Driven by a fake clock rather than sleeping 15 minutes."""
key = "1.2.3.4"
clock = {"now": 1000.0}
monkeypatch.setattr(auth, "_now", lambda: clock["now"])
for _ in range(auth.THROTTLE_THRESHOLD):
await auth.record_failure(key)
assert await auth.retry_after_seconds(key) > 0
clock["now"] += auth.THROTTLE_WINDOW_SECONDS + 1
assert await auth.retry_after_seconds(key) == 0
@pytest.mark.asyncio
async def test_keys_are_independent(self, db):
for _ in range(auth.THROTTLE_THRESHOLD):
await auth.record_failure("attacker")
assert await auth.retry_after_seconds("attacker") > 0
assert await auth.retry_after_seconds("innocent") == 0
# --------------------------------------------------------------------------
# HTTP enforcement — the acceptance criteria
# --------------------------------------------------------------------------
class TestEnforcement:
@pytest.mark.asyncio
async def test_health_is_reachable_without_a_session(self, client):
assert (await client.get("/api/health")).status_code == 200
@pytest.mark.asyncio
async def test_protected_endpoint_returns_401_without_a_session(self, client):
response = await client.get("/api/monitor/tasks")
assert response.status_code == 401
@pytest.mark.asyncio
async def test_wrong_password_is_401_and_generic(self, client):
await _seed_password()
response = await client.post("/api/auth/login", json={"password": "wrong"})
assert response.status_code == 401
# Must not reveal whether a password is even configured.
assert response.json()["detail"] == auth.INVALID_CREDENTIALS
@pytest.mark.asyncio
async def test_login_unlocks_the_api(self, client):
await _seed_password()
login = await client.post("/api/auth/login", json={"password": PASSWORD})
assert login.status_code == 200
assert auth.SESSION_COOKIE_NAME in client.cookies
assert (await client.get("/api/monitor/tasks")).status_code == 200
@pytest.mark.asyncio
async def test_cookie_is_httponly_and_lax(self, client):
await _seed_password()
login = await client.post("/api/auth/login", json={"password": PASSWORD})
raw = login.headers["set-cookie"].lower()
assert "httponly" in raw
assert "samesite=lax" in raw
# Secure must be OFF by default: the LAN bind is plain HTTP and a Secure
# cookie is silently dropped there, looping the login page.
assert "secure" not in raw
@pytest.mark.asyncio
async def test_bearer_token_also_works(self, client):
"""Scripts and curl cannot use a cookie jar conveniently."""
await _seed_password()
login = await client.post("/api/auth/login", json={"password": PASSWORD})
token = login.cookies[auth.SESSION_COOKIE_NAME]
async with httpx.AsyncClient(
transport=httpx.ASGITransport(app=app), base_url="http://test"
) as bare:
bare.headers["Authorization"] = f"Bearer {token}"
assert (await bare.get("/api/monitor/tasks")).status_code == 200
@pytest.mark.asyncio
async def test_tampered_token_is_rejected(self, client):
await _seed_password()
await client.post("/api/auth/login", json={"password": PASSWORD})
client.cookies.set(auth.SESSION_COOKIE_NAME, "not-a-real-token")
assert (await client.get("/api/monitor/tasks")).status_code == 401
@pytest.mark.asyncio
async def test_logout_invalidates_the_session(self, client):
await _seed_password()
await client.post("/api/auth/login", json={"password": PASSWORD})
assert (await client.get("/api/monitor/tasks")).status_code == 200
assert (await client.post("/api/auth/logout")).status_code == 200
assert (await client.get("/api/monitor/tasks")).status_code == 401
@pytest.mark.asyncio
async def test_me_reports_401_when_logged_out(self, client):
await _seed_password()
assert (await client.get("/api/auth/me")).status_code == 401
await client.post("/api/auth/login", json={"password": PASSWORD})
me = await client.get("/api/auth/me")
assert me.status_code == 200
assert me.json()["authenticated"] is True
@pytest.mark.asyncio
async def test_password_change_evicts_other_devices(self, client):
await _seed_password()
# A second "device" holds its own session.
login = await client.post("/api/auth/login", json={"password": PASSWORD})
other_token = login.cookies[auth.SESSION_COOKIE_NAME]
changed = await client.post(
"/api/auth/password",
json={"current": PASSWORD, "new": "brand-new-password"},
)
assert changed.status_code == 200
# The old token is dead.
async with httpx.AsyncClient(
transport=httpx.ASGITransport(app=app), base_url="http://test"
) as other:
other.cookies.set(auth.SESSION_COOKIE_NAME, other_token)
assert (await other.get("/api/monitor/tasks")).status_code == 401
# ...and the caller is still logged in.
assert (await client.get("/api/monitor/tasks")).status_code == 200
@pytest.mark.asyncio
async def test_password_change_requires_the_current_password(self, client):
await _seed_password()
await client.post("/api/auth/login", json={"password": PASSWORD})
response = await client.post(
"/api/auth/password", json={"current": "wrong", "new": "whatever-new"}
)
assert response.status_code == 401
@pytest.mark.asyncio
async def test_repeated_failures_get_throttled(self, client):
await _seed_password()
for _ in range(auth.THROTTLE_THRESHOLD):
await client.post("/api/auth/login", json={"password": "wrong"})
blocked = await client.post("/api/auth/login", json={"password": PASSWORD})
assert blocked.status_code == 429
assert "retry-after" in {k.lower() for k in blocked.headers}
@pytest.mark.asyncio
async def test_docs_are_not_exposed(self, client):
for path in ("/docs", "/redoc", "/openapi.json"):
assert (await client.get(path)).status_code == 404
class TestEveryRouteIsGuarded:
@pytest.mark.asyncio
async def test_no_api_route_is_accidentally_open(self, client):
"""The guard that stops the next endpoint from shipping unauthenticated."""
unguarded = []
for route in app.routes:
path = getattr(route, "path", "")
methods = getattr(route, "methods", None)
if not path.startswith("/api") or not methods or path in EXEMPT_PATHS:
continue
# Substitute dummy values for path params so we reach the auth check
# rather than a 404/422 on the parameter itself.
concrete = "/".join(
"1" if segment.startswith("{") else segment for segment in path.split("/")
)
for method in methods - {"HEAD", "OPTIONS"}:
response = await client.request(method, concrete, json={})
if response.status_code != 401:
unguarded.append(f"{method} {path} -> {response.status_code}")
assert not unguarded, f"以下 /api 路由未受鉴权保护:{unguarded}"
# --------------------------------------------------------------------------
# WebSocket enforcement
# --------------------------------------------------------------------------
class _FakeWebSocket:
"""Only `.cookies` is read by require_ws_auth."""
def __init__(self, cookies):
self.cookies = cookies
class TestWebSocketAuth:
"""Guarding websockets needs its own mechanism: BaseHTTPMiddleware returns
early for non-http scopes, and HTTP router dependencies never run for them.
Without this the live crawl log stream would be wide open.
"""
@pytest.mark.asyncio
async def test_missing_cookie_is_rejected(self, db):
with pytest.raises(WebSocketException) as excinfo:
await auth.require_ws_auth(_FakeWebSocket({}))
assert excinfo.value.code == 1008
@pytest.mark.asyncio
async def test_valid_cookie_is_accepted(self, db):
async with monitor_db.get_session() as session:
token, _ = await auth.create_session(session)
# No exception means accepted.
await auth.require_ws_auth(_FakeWebSocket({auth.SESSION_COOKIE_NAME: token}))
@pytest.mark.asyncio
async def test_unknown_cookie_is_rejected(self, db):
with pytest.raises(WebSocketException):
await auth.require_ws_auth(_FakeWebSocket({auth.SESSION_COOKIE_NAME: "bogus"}))
def test_every_websocket_route_carries_the_guard(self):
guarded = {
route.path
for route in app.routes
if route.__class__.__name__ == "APIWebSocketRoute"
and any(
getattr(dep.dependency, "__name__", "") == "require_ws_auth"
for dep in (route.dependencies or [])
)
}
every_ws = {
route.path
for route in app.routes
if route.__class__.__name__ == "APIWebSocketRoute"
}
assert every_ws, "expected at least one websocket route"
assert every_ws == guarded
+128
View File
@@ -0,0 +1,128 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_cmd_arg_monitor_flags.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Tests for the CLI flags added to support unattended monitoring runs.
Each flag must default to the existing config value, so a manual crawl that does
not pass them behaves exactly as before.
"""
import pytest
import config
from cmd_arg.arg import parse_cmd
BASE_ARGS = ["--platform", "xhs", "--type", "creator", "--creator_id", "abc123"]
@pytest.fixture(autouse=True)
def _isolate_config(monkeypatch):
monkeypatch.setattr(config, "ENABLE_CDP_MODE", True)
monkeypatch.setattr(config, "INJECT_ALL_COOKIES", False)
monkeypatch.setattr(config, "SAVE_LOGIN_STATE", True)
monkeypatch.setattr(config, "COOKIES", "")
monkeypatch.setattr(config, "SAVE_DATA_PATH", "")
monkeypatch.setattr(config, "CRAWLER_MAX_SLEEP_SEC", 2)
yield
class TestEnableCdpMode:
"""CDP attaches to the user's desktop Chrome, which cannot work on a server."""
@pytest.mark.asyncio
async def test_false_disables_cdp(self):
await parse_cmd([*BASE_ARGS, "--enable_cdp_mode", "false"])
assert config.ENABLE_CDP_MODE is False
@pytest.mark.asyncio
async def test_defaults_to_config_value(self):
await parse_cmd(BASE_ARGS)
assert config.ENABLE_CDP_MODE is True
class TestCookieFlags:
@pytest.mark.asyncio
async def test_inject_all_cookies_enables_switch(self):
await parse_cmd([*BASE_ARGS, "--inject_all_cookies", "true"])
assert config.INJECT_ALL_COOKIES is True
@pytest.mark.asyncio
async def test_inject_all_cookies_defaults_off(self):
await parse_cmd(BASE_ARGS)
assert config.INJECT_ALL_COOKIES is False
@pytest.mark.asyncio
async def test_cookies_file_is_read_into_config(self, tmp_path):
cookie_file = tmp_path / "cookies.txt"
cookie_file.write_text("web_session=abc; a1=def", encoding="utf-8")
await parse_cmd([*BASE_ARGS, "--cookies_file", str(cookie_file)])
assert config.COOKIES == "web_session=abc; a1=def"
@pytest.mark.asyncio
async def test_cookies_file_wins_over_inline_cookies(self, tmp_path):
cookie_file = tmp_path / "cookies.txt"
cookie_file.write_text("web_session=fromfile", encoding="utf-8")
await parse_cmd(
[*BASE_ARGS, "--cookies", "web_session=inline", "--cookies_file", str(cookie_file)]
)
assert config.COOKIES == "web_session=fromfile"
@pytest.mark.asyncio
async def test_missing_cookies_file_is_rejected(self, tmp_path):
missing = tmp_path / "nope.txt"
with pytest.raises(Exception) as excinfo:
await parse_cmd([*BASE_ARGS, "--cookies_file", str(missing)])
# A silently-ignored unreadable cookie file would produce a crawl that
# returns nothing, which is exactly the failure mode this flag exists
# to avoid.
assert "cookies_file" in str(excinfo.value)
class TestSaveDataPath:
@pytest.mark.asyncio
async def test_save_data_path_is_applied(self):
await parse_cmd([*BASE_ARGS, "--save_data_path", "data/monitor_runs/1/2"])
assert config.SAVE_DATA_PATH == "data/monitor_runs/1/2"
class TestSaveLoginState:
@pytest.mark.asyncio
async def test_save_login_state_can_be_disabled(self):
await parse_cmd([*BASE_ARGS, "--save_login_state", "false"])
assert config.SAVE_LOGIN_STATE is False
class TestCrawlSleepSec:
"""Exposed on the Settings page; previously had no CLI flag at all."""
@pytest.mark.asyncio
async def test_value_is_applied(self):
await parse_cmd([*BASE_ARGS, "--crawler_max_sleep_sec", "9"])
assert config.CRAWLER_MAX_SLEEP_SEC == 9
@pytest.mark.asyncio
async def test_defaults_to_config_value(self, monkeypatch):
monkeypatch.setattr(config, "CRAWLER_MAX_SLEEP_SEC", 4)
await parse_cmd(BASE_ARGS)
assert config.CRAWLER_MAX_SLEEP_SEC == 4
+241
View File
@@ -0,0 +1,241 @@
# -*- coding: utf-8 -*-
"""创作者后台客户端的解析与签名。
**字段名尚未亲眼验证过**:Phase 0 抓响应时账号的数据权限还没生效,列表接口返回的是
空壳(`data.result` 里只有 `{success, code, message}`)。所以这些解析写成多别名匹配,
而这份测试就是它的规格 —— 等真实响应到手,先跑这里看哪些假设破了。
"""
import pytest
from api.creator import signing
from api.creator.client import (
CreatorClient,
as_float,
as_int,
as_seconds,
find_note_list,
normalize_note,
trans_cookies,
)
from api.creator.models import CreatorAccount
from api.creator.service import _account_dict
# --- cookie 解析 -----------------------------------------------------------
@pytest.mark.parametrize(
"raw, expected",
[
("a1=abc; web_session=xyz", {"a1": "abc", "web_session": "xyz"}),
("a1=abc;web_session=xyz;", {"a1": "abc", "web_session": "xyz"}),
("a1=abc\nweb_session=xyz", {"a1": "abc", "web_session": "xyz"}),
(" a1 = abc ; ", {"a1": "abc"}),
("", {}),
(None, {}),
# 值里可以有等号,不能被截断
("a1=abc=def", {"a1": "abc=def"}),
],
)
def test_trans_cookies(raw, expected):
assert trans_cookies(raw) == expected
def test_client_without_a1_cannot_sign():
"""a1 参与签名,没有它连请求都发不出去 —— 要提前拦而不是发出去再猜。"""
assert CreatorClient("web_session=abc").looks_authenticated is False
assert CreatorClient("a1=abc").looks_authenticated is True
# --- 数值解析 --------------------------------------------------------------
#
# 后台返回的可能是数字,也可能是 "1.2万" / "12.3%" / "1分30秒" 这类展示值。
# 解析不出来一律 None —— 不是 0。0 是真实值,None 是"不知道"。
@pytest.mark.parametrize(
"raw, expected",
[
(123, 123),
("123", 123),
("1,234", 1234),
("1.2万", 12000),
("3万", 30000),
("1.5w", 15000),
("1亿", 100000000),
(0, 0),
("0", 0),
# 这些必须是 None 而不是 0 —— 把"没给"当成"是零"会让报表说谎
(None, None),
("", None),
("-", None),
("暂无", None),
("abc", None),
(True, None),
],
)
def test_as_int(raw, expected):
assert as_int(raw) == expected
@pytest.mark.parametrize(
"raw, expected",
[
(12.3, 12.3),
("12.3%", 12.3),
("12.3", 12.3),
(None, None),
("-", None),
("暂无数据", None),
],
)
def test_as_float(raw, expected):
assert as_float(raw) == expected
@pytest.mark.parametrize(
"raw, expected",
[
(45, 45.0),
("45", 45.0),
("1分30秒", 90.0),
("2分", 120.0),
("30秒", 30.0),
("01:30", 90.0),
("1:00:00", 3600.0),
(None, None),
("-", None),
("abc", None),
],
)
def test_as_seconds(raw, expected):
assert as_seconds(raw) == expected
# --- 字段归一化 ------------------------------------------------------------
def test_normalize_note_maps_aliases():
"""不同来源的记录用不同字段名,别名表要能都接住。"""
note = normalize_note(
{
"note_id": "abc123",
"title": "标题",
"publish_time": 1700000000000,
"view_count": "1.2万",
"like_count": 34,
"collected_count": 5,
"share_count": 2,
"comment_count": 7,
"cover_click_rate": "12.5%",
"avg_watch_time": "1分30秒",
}
)
assert note["note_id"] == "abc123"
assert note["title"] == "标题"
assert note["views"] == 12000
assert note["likes"] == 34
assert note["favorites"] == 5
assert note["shares"] == 2
assert note["comments"] == 7
assert note["cover_ctr"] == 12.5
assert note["avg_watch_seconds"] == 90.0
def test_normalize_note_leaves_missing_fields_as_none():
note = normalize_note({"note_id": "abc123"})
assert note["note_id"] == "abc123"
assert note["views"] is None
assert note["likes"] is None
def test_find_note_list_digs_the_array_out_of_a_nested_payload():
"""接口的确切结构没见过,所以按"像是一批笔记记录"来找,不写死路径。"""
payload = {
"code": 0,
"data": {
"result": {
"success": True,
"notes": [
{"note_id": "n1", "views": 10, "likes": 1},
{"note_id": "n2", "views": 20, "likes": 2},
],
}
},
}
found = find_note_list(payload)
assert [item["note_id"] for item in found] == ["n1", "n2"]
def test_find_note_list_returns_empty_for_the_permission_gated_envelope():
"""权限未生效时接口返回的就是这个 —— 必须安静地给出空列表,不是报错。"""
payload = {
"code": 0,
"success": True,
"msg": "成功",
"data": {"result": {"success": True, "code": 0, "message": "success"}},
}
assert find_note_list(payload) == []
# --- 签名 ------------------------------------------------------------------
def test_signed_api_carries_the_url_prefix():
"""待签字符串必须带 `url=`。少了它网关返回 406,而 406 的响应体看不出错在哪。"""
assert signing.signed_api("/api/galaxy/user/info") == "url=/api/galaxy/user/info"
assert (
signing.signed_api("/api/x", "a=1&b=2")
== "url=/api/x?a=1&b=2"
)
def test_sign_returns_xs_and_xt():
headers = signing.sign_xyw("url=/api/galaxy/user/info", "some-a1")
assert set(headers) == {"x-s", "x-t"}
assert headers["x-s"].startswith("XYW_")
assert headers["x-t"].isdigit()
def test_signature_is_stable_for_a_fixed_timestamp():
"""同一输入同一时间戳必须得到同一签名 —— 否则说明有隐藏的随机源。"""
first = signing.sign_xyw("url=/api/x", "a1", timestamp_ms=1700000000000)
second = signing.sign_xyw("url=/api/x", "a1", timestamp_ms=1700000000000)
assert first == second
def test_signature_changes_with_the_signed_string():
"""签名必须真的绑定待签内容,否则改参数不会被发现 —— 那这个签名就没意义了。"""
base = signing.sign_xyw("url=/api/x?a=1", "a1", timestamp_ms=1700000000000)
other = signing.sign_xyw("url=/api/x?a=2", "a1", timestamp_ms=1700000000000)
assert base["x-s"] != other["x-s"]
# --- 凭证不外泄 ------------------------------------------------------------
def test_account_dict_never_carries_the_cookie():
"""cookie 等于登录态。对外结构里只该有 `has_cookie`。"""
account = CreatorAccount(
id=1,
nickname="测试",
user_id="u1",
cookie="a1=SECRET; web_session=SECRET",
created_at=0,
updated_at=0,
)
payload = _account_dict(account)
assert payload["has_cookie"] is True
assert "cookie" not in payload
assert "SECRET" not in str(payload)
+160
View File
@@ -0,0 +1,160 @@
# -*- coding: utf-8 -*-
"""运营账号扫码登录的完成判据。
这里守的是一个具体的故障:判据原先读页面里的 `window.__INITIAL_STATE__`,而那是
**页面加载那一刻的快照** —— 扫码是加载之后才登录的,快照不会翻转,于是登录明明
成功了,界面却永远停在二维码上。
现在改成拿 cookie 去问创作者后台"我是谁"。重要的是**游客也有 a1**(所以签名算得
出来),所以"有 a1"什么都不能证明,只有后台认了才算数。
"""
from unittest.mock import AsyncMock, MagicMock
import pytest
from api.creator import login as creator_login
from api.creator.client import CreatorApiError
class _FakeContext:
def __init__(self, cookies):
self._cookies = cookies
self.closed = False
async def cookies(self):
return self._cookies
async def close(self):
self.closed = True
class _FakePage:
def __init__(self):
self.closed = False
async def close(self):
self.closed = True
def _session(cookies):
context = _FakeContext(cookies)
page = _FakePage()
session = creator_login.AccountLoginSession(context, page)
# 跳过节流,让每次 refresh 都真的去问一次。
session._last_login_check = 0.0
return session, context, page
def _patch_client(monkeypatch, behaviour):
"""behaviour(cookie) -> dict 或抛异常。"""
class _Client:
def __init__(self, cookie, **kwargs):
self.cookie = cookie
async def fetch_user_info(self):
return behaviour(self.cookie)
monkeypatch.setattr(creator_login, "CreatorClient", _Client)
GUEST_COOKIES = [
{"name": "a1", "value": "guest-a1"},
{"name": "web_session", "value": "guest-session"},
]
@pytest.mark.asyncio
async def test_a_guest_session_never_completes(monkeypatch):
"""游客也有 a1,但后台回 401 —— 必须继续等,不能当成登录成功。"""
def behaviour(_cookie):
raise CreatorApiError("登录态无效或已过期", status=401)
_patch_client(monkeypatch, behaviour)
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
assert session.status == creator_login.STATUS_WAITING
assert session.cookie == ""
@pytest.mark.asyncio
async def test_a_recognised_identity_completes_the_login(monkeypatch):
_patch_client(
monkeypatch,
lambda _cookie: {"user_id": "u123", "nickname": "小明", "red_id": "1"},
)
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
assert session.status == creator_login.STATUS_SUCCESS
assert "小明" in session.message
assert "a1=guest-a1" in session.cookie
assert session.account["user_id"] == "u123"
@pytest.mark.asyncio
async def test_an_empty_identity_does_not_complete(monkeypatch):
"""接口返回 200 但没有账号标识 —— 同样不能算成功。"""
_patch_client(monkeypatch, lambda _cookie: {"user_id": "", "nickname": ""})
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
assert session.status == creator_login.STATUS_WAITING
@pytest.mark.asyncio
async def test_the_api_is_not_hit_on_every_poll(monkeypatch):
"""前端每 2 秒轮询一次,但每次轮询都打一次后台接口是浪费。"""
calls = []
def behaviour(cookie):
calls.append(cookie)
raise CreatorApiError("还没登录", status=401)
_patch_client(monkeypatch, behaviour)
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
# 第一次之后 _last_login_check 已经是"现在",后面两次应落在节流窗口内。
await session.refresh()
await session.refresh()
assert len(calls) == 1
@pytest.mark.asyncio
async def test_expiry_beats_a_successful_scan(monkeypatch):
_patch_client(monkeypatch, lambda _cookie: {"user_id": "u1", "nickname": "x"})
session, _context, _page = _session(GUEST_COOKIES)
session.started_at -= creator_login.QR_TTL_SECONDS + 1
await session.refresh()
assert session.status == creator_login.STATUS_EXPIRED
@pytest.mark.asyncio
async def test_a_closed_window_is_reported(monkeypatch):
session, context, _page = _session(GUEST_COOKIES)
context.cookies = AsyncMock(side_effect=RuntimeError("Target closed"))
await session.refresh()
assert session.status == creator_login.STATUS_ERROR
@pytest.mark.asyncio
async def test_closing_discards_the_temporary_context():
"""临时上下文是这个设计的关键 —— 用完必须关掉,否则会挂在操作者的 Chrome 里。"""
session, context, page = _session(GUEST_COOKIES)
await session.close()
assert context.closed is True
assert page.closed is True
+401
View File
@@ -0,0 +1,401 @@
# -*- coding: utf-8 -*-
"""抖音 Web 接口客户端 —— 纯逻辑部分(不发网络请求、不连浏览器)。
发请求那半边只能在真环境里验(要 CDP 浏览器 + 登录态),所以这里钉住的是那些
「错了会一路错到入库」的地方:请求头的成套性、cookie 解析、以及产物键名。
"""
from tools.user_hash import anonymize_user_id
import asyncio
import httpx
import pytest
from api.monitor import douyin_api
class TestCookieParsing:
def test_cookie_header_is_normalised(self):
assert douyin_api._cookie_header(" a=1 ; b = 2 ;; c=3 ") == "a=1; b=2; c=3"
def test_cookie_value_lookup(self):
assert douyin_api._cookie_value("a=1; UIFID=xyz; b=2", "UIFID") == "xyz"
assert douyin_api._cookie_value("a=1", "UIFID") == ""
assert douyin_api._cookie_value("", "UIFID") == ""
def test_session_detection(self):
assert douyin_api._has_session("a=1; sessionid=abc") is True
assert douyin_api._has_session("a=1; sessionid_ss=abc") is False
assert douyin_api._has_session("") is False
class TestRequestHeaders:
"""请求头必须**成套**,而且成套地来自同一个浏览器。
实测:只有 UA + client hints + Cookie 时,主页接口回 200 但只有 121 字节(空壳);
补上 Accept / Accept-Language / Referer 才变成 7074 字节的真数据。
"""
def test_the_full_set_is_sent(self):
identity = douyin_api.BrowserIdentity(
cookie="sessionid=s; UIFID=u1",
user_agent="UA-of-this-browser",
client_hints={"sec-ch-ua": '"Chrome";v="155"'},
)
headers = identity.headers()
assert headers["User-Agent"] == "UA-of-this-browser"
assert headers["sec-ch-ua"] == '"Chrome";v="155"'
assert headers["Accept"], "Accept 系列是主页接口能不能返回真数据的必要条件"
assert headers["Accept-Language"]
assert headers["Referer"] == "https://www.douyin.com/"
assert headers["x-tt-argus"] == douyin_api.ARGUS_HEADER_VALUE
assert headers["uifid"] == "u1"
assert headers["Cookie"] == "sessionid=s; UIFID=u1"
def test_uifid_is_omitted_when_absent(self):
"""cookie 里没有 uifid 就别带 —— 送个空值反而更像异常请求。"""
identity = douyin_api.BrowserIdentity(
cookie="sessionid=s", user_agent="UA", client_hints={}
)
assert "uifid" not in identity.headers()
def test_uifid_temp_is_used_as_a_fallback(self):
identity = douyin_api.BrowserIdentity(
cookie="sessionid=s; UIFID_TEMP=temp-1", user_agent="UA", client_hints={}
)
assert identity.headers()["uifid"] == "temp-1"
class TestNormalizeAweme:
def test_keys_match_what_the_store_writes(self):
"""键名必须和 store/douyin 一模一样,否则 ingest 一条都读不到。"""
record = douyin_api.normalize_aweme(
{
"aweme_id": 7690458980574358513,
"desc": "中秋哪儿都堵",
"create_time": 1790574515,
"author": {"uid": "776719710825195", "nickname": "AA建材王总"},
"statistics": {
"digg_count": 3,
"comment_count": 1,
"collect_count": 2,
"share_count": 0,
},
"video": {"cover": {"url_list": ["https://img/cover.jpg"]}},
}
)
assert record["aweme_id"] == "7690458980574358513"
assert record["title"] == "中秋哪儿都堵"
assert record["nickname"] == "AA建材王总"
assert record["cover_url"] == "https://img/cover.jpg"
assert (
record["aweme_url"]
== "https://www.douyin.com/video/7690458980574358513"
)
# 与 store 一致:creator_hash 是 uid 的匿名哈希。
assert record["creator_hash"] == anonymize_user_id("776719710825195")
# **秒**。adapters 的 time_scale=1000 会把它换成毫秒 —— 这一层不算毫秒。
assert record["create_time"] == 1790574515
# 指标按 store 的形态落成字符串,交给 ingest 的 parse_count 解析。
assert record["liked_count"] == "3"
assert record["collected_count"] == "2"
def test_missing_fields_do_not_crash(self):
record = douyin_api.normalize_aweme({"aweme_id": "1"})
assert record["aweme_id"] == "1"
assert record["title"] == ""
assert record["cover_url"] == ""
assert record["liked_count"] == "0"
assert record["create_time"] == 0
class TestIdentityFromPages:
"""身份得从浏览器里问,但**不能被一个卡死的标签页拖住**。
实测过:标签页 URL 为空、渲染进程卡死,``page.evaluate`` 永远不返回;而问身份是采集的
第一步 —— 没超时的话整轮就挂在那儿,run 永远停在「运行中」。
"""
class _Page:
def __init__(self, url, *, user_agent=None, hang=False):
self.url = url
self._user_agent = user_agent
self._hang = hang
self.closed = False
async def evaluate(self, expression):
if self._hang:
await asyncio.sleep(30) # 模拟渲染进程卡死
if expression.startswith("() => navigator.userAgentData"):
return {
"brands": [{"brand": "Chrome", "version": "155"}],
"mobile": False,
"platform": "Linux",
}
return self._user_agent
async def close(self):
self.closed = True
class _Context:
def __init__(self, pages, temp=None):
self.pages = pages
self._temp = temp
self.made_temp = False
async def new_page(self):
self.made_temp = True
if self._temp is None:
raise AssertionError("这个用例不该走到临时页")
return self._temp
def _fast_timeout(self, monkeypatch):
monkeypatch.setattr(douyin_api, "EVALUATE_TIMEOUT_SECONDS", 0.05)
@pytest.mark.asyncio
async def test_a_hanging_page_is_skipped(self, monkeypatch):
self._fast_timeout(monkeypatch)
stuck = self._Page("", hang=True)
good = self._Page("https://example.com/", user_agent="UA-of-good-page")
user_agent, hints = await douyin_api._identity_from_pages(
self._Context([stuck, good])
)
assert user_agent == "UA-of-good-page"
assert hints["sec-ch-ua"] == '"Chrome";v="155"'
@pytest.mark.asyncio
async def test_a_page_without_a_user_agent_is_skipped(self, monkeypatch):
self._fast_timeout(monkeypatch)
blank = self._Page("", user_agent=None)
good = self._Page("https://example.com/", user_agent="UA-of-good-page")
user_agent, _ = await douyin_api._identity_from_pages(
self._Context([blank, good])
)
assert user_agent == "UA-of-good-page"
@pytest.mark.asyncio
async def test_it_opens_a_temporary_page_when_nothing_else_works(self, monkeypatch):
self._fast_timeout(monkeypatch)
stuck = self._Page("", hang=True)
temp = self._Page("about:blank", user_agent="UA-of-temp-page")
context = self._Context([stuck], temp=temp)
user_agent, _ = await douyin_api._identity_from_pages(context)
assert context.made_temp is True
assert user_agent == "UA-of-temp-page"
assert temp.closed is True, "临时页问完要关掉,别在操作者的浏览器里留垃圾"
@pytest.mark.asyncio
async def test_a_douyin_page_is_preferred(self, monkeypatch):
"""有抖音页面就先问它 —— 它才是我们要模仿的那个身份。"""
self._fast_timeout(monkeypatch)
other = self._Page("https://example.com/", user_agent="UA-of-other")
douyin = self._Page("https://www.douyin.com/explore", user_agent="UA-of-douyin")
user_agent, _ = await douyin_api._identity_from_pages(
self._Context([other, douyin])
)
assert user_agent == "UA-of-douyin"
class TestGet:
"""`_get` 的失败路径 —— 它们决定了失败会不会被伪装成「这个博主没作品」。"""
@staticmethod
def _client_returning(monkeypatch, status_code: int, text: str):
class _Response:
def json(self):
import json as _json
return _json.loads(self.text)
response = _Response()
response.status_code = status_code
response.text = text
class _Client:
async def __aenter__(self):
return self
async def __aexit__(self, *exc):
return False
async def get(self, *args, **kwargs):
return response
monkeypatch.setattr(httpx, "AsyncClient", lambda **kwargs: _Client())
def _identity(self):
return douyin_api.BrowserIdentity(
cookie="sessionid=s", user_agent="UA", client_hints={}
)
def test_an_empty_body_is_an_error_not_an_empty_result(self, monkeypatch):
"""「200 + 空 body」是网关拒绝请求的典型回应。
必须当场报错 —— 放过去的话,它会在下游变成「这个博主没作品」,把一次失败伪装成
一条正常的空结果。爬虫那条路就是这么栽的,还被翻译成「账号被封」。
"""
self._client_returning(monkeypatch, 200, "")
with pytest.raises(douyin_api.DouyinApiError) as excinfo:
asyncio.run(douyin_api._get("/x", {}, self._identity()))
assert "空内容" in str(excinfo.value)
def test_a_403_carries_the_gateways_own_message(self, monkeypatch):
"""抖音难得会说原因,把它带出来,别丢。"""
self._client_returning(
monkeypatch, 403, "Blocked by ArgusSecurityPlugin Uifid Not Found"
)
with pytest.raises(douyin_api.DouyinApiError) as excinfo:
asyncio.run(douyin_api._get("/x", {}, self._identity()))
assert "403" in str(excinfo.value)
assert "Uifid Not Found" in str(excinfo.value)
def test_a_200_with_data_is_returned_as_is(self, monkeypatch):
self._client_returning(monkeypatch, 200, '{"user": {"nickname": "x"}}')
assert asyncio.run(douyin_api._get("/x", {}, self._identity())) == {
"user": {"nickname": "x"}
}
@pytest.fixture
def fake_signer(monkeypatch):
"""把真正的 ``a_bogus`` 签名换成假的。
真的那个在 import 的瞬间就要把 ``libs/douyin.js`` 喂给 execjs(还得有 node 和正确的
相对路径),单元测试不该依赖这些。**签名本身是实测过的**:带上它是 200 + 真评论,
不带是 200 + 空 body。这里只负责钉住「有没有带上、传对了没有」。
"""
import sys
import types
module = types.ModuleType("media_platform.douyin.help")
calls: list = []
def get_a_bogus_from_js(url: str, params: str, user_agent: str) -> str:
calls.append({"url": url, "params": params, "user_agent": user_agent})
return "FAKE-BOGUS"
module.get_a_bogus_from_js = get_a_bogus_from_js
monkeypatch.setitem(sys.modules, "media_platform.douyin.help", module)
return calls
@pytest.fixture
def recording_client(monkeypatch):
"""记下实际发出去的那次请求。"""
sent: dict = {}
class _Response:
status_code = 200
text = '{"comments": []}'
def json(self):
return {"comments": []}
class _Client:
async def __aenter__(self):
return self
async def __aexit__(self, *exc):
return False
async def get(self, url, **kwargs):
sent["url"] = url
sent.update(kwargs)
return _Response()
monkeypatch.setattr(httpx, "AsyncClient", lambda **kwargs: _Client())
return sent
class TestCommentSigning:
"""评论接口**必须**带 a_bogus。
不带的话网关回 200 + 空 body —— 那在下游会变成「这条作品没有评论」,把一次被挡住的
请求伪装成一条正常的空结果。和登录失效长得一模一样,查起来能查半天。
"""
def _identity(self):
return douyin_api.BrowserIdentity(
cookie="sessionid=s", user_agent="UA", client_hints={}
)
def test_the_signature_is_computed_over_the_unsigned_params(
self, fake_signer, recording_client
):
asyncio.run(
douyin_api._get(
douyin_api.COMMENT_PATH,
{"aweme_id": "123", "count": 20},
self._identity(),
signed=True,
)
)
assert len(fake_signer) == 1
# 签名算在**不含 a_bogus** 的那串上 —— 把它自己也算进去是循环的。
assert fake_signer[0]["params"] == "aweme_id=123&count=20"
assert fake_signer[0]["url"] == douyin_api.COMMENT_PATH
assert fake_signer[0]["user_agent"] == "UA"
def test_the_signature_goes_out_with_the_request(self, fake_signer, recording_client):
asyncio.run(
douyin_api._get(
douyin_api.COMMENT_PATH, {"aweme_id": "123"}, self._identity(), signed=True
)
)
assert recording_client["params"]["a_bogus"] == "FAKE-BOGUS"
assert recording_client["params"]["aweme_id"] == "123"
def test_unsigned_calls_never_touch_the_signer(self, fake_signer, recording_client):
"""作品 / 详情 / 博主资料三个接口不带签名也照常返回。
给它们加签名是**没验证过的改动** —— 所以这里钉住「不签」,防止有人图省事把
signed=True 改成全局默认。
"""
asyncio.run(
douyin_api._get(douyin_api.POSTS_PATH, {"sec_user_id": "x"}, self._identity())
)
assert fake_signer == []
assert "a_bogus" not in recording_client["params"]
def test_a_broken_signer_is_reported_as_such(self, monkeypatch, recording_client):
"""execjs 起不来时要说出是签名失败,而不是让它变成「没评论」。"""
import sys
import types
module = types.ModuleType("media_platform.douyin.help")
def boom(url, params, user_agent):
raise RuntimeError("node 没装")
module.get_a_bogus_from_js = boom
monkeypatch.setitem(sys.modules, "media_platform.douyin.help", module)
with pytest.raises(douyin_api.DouyinApiError, match="a_bogus"):
asyncio.run(
douyin_api._get(
douyin_api.COMMENT_PATH, {"aweme_id": "1"}, self._identity(), signed=True
)
)
+84
View File
@@ -0,0 +1,84 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_douyin_argus_header.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
"""抖音 ArgusSecurityPlugin 请求头的回归测试。
回归背景:抖音在边缘网关挂了 ArgusSecurityPlugin,对一批接口做业务前置校验。
缺少 ``x-tt-argus`` 请求头时直接 403,响应体为
``Blocked by ArgusSecurityPlugin Uifid Not Found``;补上 uifid 参数但仍没有这个头
则是 ``... Signature Not Found``(容易误导成 a_bogus / verifyFp 的问题)。
网关当前不校验该头的取值,传固定字符串即可。
这里不发起任何网络请求,只断言客户端默认请求头带上了这两个头。
"""
from __future__ import annotations
import pytest
from media_platform.douyin.client import DOUYIN_ARGUS_HEADER_VALUE, DouYinClient
COOKIE_DICT = {
"sessionid": "fake-session",
"UIFID": "uifid-from-cookie",
"UIFID_TEMP": "uifid-temp-from-cookie",
}
class _StubPage:
async def evaluate(self, expression): # noqa: ANN001
return {}
def _make_client(cookie_dict: dict) -> DouYinClient:
return DouYinClient(
headers={"User-Agent": "test-user-agent", "Cookie": "a=1"},
playwright_page=_StubPage(),
cookie_dict=cookie_dict,
)
def test_argus_header_is_present_by_default():
"""x-tt-argus 必须在默认请求头里"""
client = _make_client(COOKIE_DICT)
assert client.headers.get("x-tt-argus") == DOUYIN_ARGUS_HEADER_VALUE
def test_uifid_header_comes_from_cookie():
"""uifid 头取自 cookie 里的 UIFID"""
client = _make_client(COOKIE_DICT)
assert client.headers.get("uifid") == "uifid-from-cookie"
def test_uifid_header_falls_back_to_uifid_temp():
"""没有 UIFID 时退到 UIFID_TEMP"""
client = _make_client({"sessionid": "s", "UIFID_TEMP": "temp-only"})
assert client.headers.get("uifid") == "temp-only"
def test_missing_uifid_omits_header():
"""cookie 里两种都没有时不发这个头(发空值可能被当成「有但为空」)"""
client = _make_client({"sessionid": "s"})
assert client.headers.get("uifid") is None
assert client.headers.get("x-tt-argus") == DOUYIN_ARGUS_HEADER_VALUE
def test_caller_supplied_values_win():
"""调用方显式传了同名头时不覆盖"""
client = DouYinClient(
headers={"User-Agent": "ua", "Cookie": "a=1", "x-tt-argus": "custom"},
playwright_page=_StubPage(),
cookie_dict=COOKIE_DICT,
)
assert client.headers.get("x-tt-argus") == "custom"
+107
View File
@@ -0,0 +1,107 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_douyin_aweme_detail_params.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
"""抖音 aweme detail 接口的风控参数回归测试。
回归背景:上游 detail 接口的 Argus 风控要求 ``uifid`` / ``verifyFp`` / ``fp``
三个参数,缺一时直接 403,响应体为
``Blocked by ArgusSecurityPlugin Uifid Not Found``(补上 uifid 但 verifyFp 不对时
换成 ``... Signature Not Found``)。原先只传 aweme_id,于是详情全线失败,
连带媒体也拿不到。
这里不发起任何网络请求,只锁定这三个参数确实被带上、且取值来自浏览器 cookie。
"""
from __future__ import annotations
import pytest
from media_platform.douyin.client import DouYinClient
AWEME_ID = "7525538910311632128"
class _StubPage:
"""占位 page,仅用于构造 client;本测试不会走到 playwright 调用"""
async def evaluate(self, expression): # noqa: ANN001
return {}
def _make_client(cookie_dict: dict) -> DouYinClient:
return DouYinClient(
headers={
"User-Agent": "test-user-agent",
"Cookie": "a=1",
"Origin": "https://www.douyin.com/",
},
playwright_page=_StubPage(),
cookie_dict=cookie_dict,
)
@pytest.mark.asyncio
async def test_aweme_detail_sends_uifid_and_verify_fp(monkeypatch):
"""uifid 与 verifyFp/fp 必须成套出现,且都取自 cookie"""
captured: dict = {}
async def fake_get(uri, params=None, headers=None): # noqa: ANN001
captured["uri"] = uri
captured["params"] = dict(params or {})
return {"aweme_detail": {"aweme_id": AWEME_ID}}
client = _make_client(
{"s_v_web_id": "verify_test_fp", "UIFID": "uifid-from-cookie"}
)
monkeypatch.setattr(client, "get", fake_get)
await client.get_video_by_id(AWEME_ID)
assert captured["uri"] == "/aweme/v1/web/aweme/detail/"
assert captured["params"]["aweme_id"] == AWEME_ID
assert captured["params"]["uifid"] == "uifid-from-cookie"
# verifyFp 与 fp 必须同源,且用 cookie 里的 s_v_web_id(自生成的会被判 Signature Not Found)
assert captured["params"]["verifyFp"] == "verify_test_fp"
assert captured["params"]["fp"] == "verify_test_fp"
@pytest.mark.asyncio
async def test_aweme_detail_falls_back_to_uifid_temp(monkeypatch):
"""没有 UIFID 时退到 UIFID_TEMP"""
captured: dict = {}
async def fake_get(uri, params=None, headers=None): # noqa: ANN001
captured["params"] = dict(params or {})
return {"aweme_detail": {}}
client = _make_client({"s_v_web_id": "fp", "UIFID_TEMP": "temp-only"})
monkeypatch.setattr(client, "get", fake_get)
await client.get_video_by_id(AWEME_ID)
assert captured["params"]["uifid"] == "temp-only"
@pytest.mark.asyncio
async def test_aweme_detail_without_cookies_still_requests(monkeypatch):
"""cookie 缺失时不能抛异常,参数退化为空串(由服务端决定是否放行)"""
captured: dict = {}
async def fake_get(uri, params=None, headers=None): # noqa: ANN001
captured["params"] = dict(params or {})
return {"aweme_detail": {}}
client = _make_client({})
monkeypatch.setattr(client, "get", fake_get)
await client.get_video_by_id(AWEME_ID)
assert captured["params"]["uifid"] == ""
assert captured["params"]["verifyFp"] == ""
assert captured["params"]["fp"] == ""
+55
View File
@@ -0,0 +1,55 @@
# -*- coding: utf-8 -*-
"""sec-ch-ua 系列请求头:为什么必须带、怎么还原。
抖音网关对「声称自己是 Chrome、却没带 sec-ch-ua」的请求会回 **200 + 空 body** ——
不报错、不给原因、HTTP 状态还是成功的,采集侧只看到 0 条。实测同一 URL、同一 cookie、
同一参数:不带头 0 字节,补上这三个头 7077 字节。
"""
import pytest
from media_platform.douyin.help import client_hint_headers
# 实测从真实浏览器抓到的 navigator.userAgentData
REAL_UA_DATA = {
"brands": [
{"brand": "Google Chrome", "version": "155"},
{"brand": "Chromium", "version": "155"},
{"brand": "Not(A:Brand", "version": "24"},
],
"mobile": False,
"platform": "Linux",
}
def test_headers_are_derived_from_user_agent_data():
hints = client_hint_headers(REAL_UA_DATA)
assert hints["sec-ch-ua"] == (
'"Google Chrome";v="155", "Chromium";v="155", "Not(A:Brand";v="24"'
)
assert hints["sec-ch-ua-mobile"] == "?0"
assert hints["sec-ch-ua-platform"] == '"Linux"'
def test_mobile_is_reflected():
hints = client_hint_headers({**REAL_UA_DATA, "mobile": True})
assert hints["sec-ch-ua-mobile"] == "?1"
@pytest.mark.parametrize("value", [None, [], "not-a-dict", {}, {"brands": []}])
def test_nothing_is_invented_when_user_agent_data_is_unavailable(value):
"""拿不到就返回空。
凭空造一组和 UA 对不上的头,只会变成**另一个**可疑特征 —— 那比不带头更糟。
"""
assert client_hint_headers(value) == {}
def test_missing_platform_still_sends_the_other_two():
hints = client_hint_headers({"brands": REAL_UA_DATA["brands"], "mobile": False})
assert "sec-ch-ua" in hints
assert "sec-ch-ua-mobile" in hints
assert "sec-ch-ua-platform" not in hints
+411
View File
@@ -0,0 +1,411 @@
# -*- coding: utf-8 -*-
"""抖音采集编排 —— 不碰网络,把 douyin_api 整个换掉。
验的是编排本身:产物落在正确的目录、文件名是 store 那套、以及**作品列表被挡时的退化**
(用已知 aweme_id 逐条刷新)—— 那条退化路径决定了今天这个功能是「完全没用」还是
「已知作品还能看」。
"""
import json
import pytest
from api.monitor import douyin_api, douyin_fetch
def _profile(**overrides) -> dict:
profile = {
"creator_hash": "hash",
"nickname": "博主",
"unique_id": "abc",
"fans": 12000,
"total_favorited": 83000,
"works": 42,
"following": 7,
}
profile.update(overrides)
return profile
@pytest.fixture(autouse=True)
def fake_profile(monkeypatch):
"""每个用例都挡住「问博主资料」这一跳。
它是附加信息,不在任何一条编排路径上,但真发出去就会去连 9222 那个浏览器 ——
于是所有 creator 用例都会多出一次连接失败、并把 ``errors`` 弄脏。想验它自己的
用例再单独覆盖这个 fixture。
"""
async def _profile_call(sec_user_id, *, cookie=""):
return _profile()
monkeypatch.setattr(douyin_api, "author_profile", _profile_call)
def _video(aweme_id: str, likes: str = "1") -> dict:
return {
"aweme_id": aweme_id,
"title": f"title-{aweme_id}",
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
"cover_url": "",
"aweme_type": "0",
"create_time": 1790574515,
"creator_hash": "hash",
"nickname": "博主",
"liked_count": likes,
"comment_count": "0",
"collected_count": "0",
"share_count": "0",
}
def _comment(aweme_id: str, index: int) -> dict:
return {
"comment_id": f"c{index}",
"aweme_id": aweme_id,
"content": f"评论{index}",
"nickname": "路人",
"creator_hash": "h2",
"create_time": 1790574600,
"like_count": "0",
"sub_comment_count": "0",
"parent_comment_id": "0",
}
class _Target:
def __init__(self, external_id: str) -> None:
self.external_id = external_id
def _read(path):
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line]
async def _collect(tmp_path, **overrides):
kwargs = dict(
platform="dy",
mode="creator",
limit=20,
want_comments=True,
comment_limit=20,
targets=[_Target("MS4w-sec")],
cookie="sessionid=x",
)
kwargs.update(overrides)
return await douyin_fetch.collect(tmp_path, **kwargs)
class TestHappyPath:
@pytest.mark.asyncio
async def test_writes_the_layout_ingest_expects(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111"), _video("222")]
async def fake_comments(aweme_id, count=20, *, cookie=""):
return [_comment(aweme_id, 1)]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "video_comments", fake_comments)
result = await _collect(tmp_path)
assert result["notes"] == 2
assert result["comments"] == 2
assert result["errors"] == []
# 目录名必须是 douyin(不是平台 id dy)—— ingest 找文件用的是同一个来源。
jsonl_dir = tmp_path / "douyin" / "jsonl"
assert jsonl_dir.is_dir()
contents = list(jsonl_dir.glob("*_contents_*.jsonl"))
comments = list(jsonl_dir.glob("*_comments_*.jsonl"))
assert len(contents) == 1
assert len(comments) == 1
notes = _read(contents[0])
assert [n["aweme_id"] for n in notes] == ["111", "222"]
# 键名照抄 store/douyin —— ingest 靠这个读出来。
assert notes[0]["aweme_url"] == "https://www.douyin.com/video/111"
# 秒,不是毫秒;换算交给 adapters。
assert notes[0]["create_time"] == 1790574515
assert _read(comments[0])[0]["aweme_id"] == "111"
@pytest.mark.asyncio
async def test_the_comment_file_exists_even_without_comments(self, monkeypatch, tmp_path):
"""评论文件必须建出来。
ingest 靠「文件在不在」区分「这一轮没评论」和「这一轮什么都没抓到」——
两种情况的含义完全不同。
"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
await _collect(tmp_path, want_comments=False)
assert len(list((tmp_path / "douyin" / "jsonl").glob("*_comments_*.jsonl"))) == 1
@pytest.mark.asyncio
async def test_duplicate_works_are_written_once(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111"), _video("111")]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
result = await _collect(tmp_path, want_comments=False)
assert result["notes"] == 1
class TestNoteMode:
@pytest.mark.asyncio
async def test_a_work_target_is_fetched_by_detail_not_by_creator_list(
self, monkeypatch, tmp_path
):
"""作品模式的目标**本身就是作品 id**,不能拿它当博主的 sec_uid 去查列表。
走错了会必然失败,而且失败原因很难看懂(接口说你没登录/不是浏览器)——
「粘贴作品链接的监控」今天本来是能用的,别让它因为这一处走错而废掉。
"""
async def must_not_be_called(*args, **kwargs):
raise AssertionError("作品模式不该去拉博主的作品列表")
async def fake_detail(aweme_id, *, cookie=""):
return _video(aweme_id)
monkeypatch.setattr(douyin_api, "author_videos", must_not_be_called)
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
result = await _collect(
tmp_path,
mode="note",
want_comments=False,
targets=[_Target("111"), _Target("222")],
)
assert result["notes"] == 2
assert result["errors"] == []
notes = _read(list((tmp_path / "douyin" / "jsonl").glob("*_contents_*.jsonl"))[0])
assert [n["aweme_id"] for n in notes] == ["111", "222"]
@pytest.mark.asyncio
async def test_a_broken_work_does_not_lose_the_others(self, monkeypatch, tmp_path):
async def flaky(aweme_id, *, cookie=""):
if aweme_id == "222":
raise douyin_api.DouyinApiError("作品已被删除")
return _video(aweme_id)
monkeypatch.setattr(douyin_api, "video_detail", flaky)
result = await _collect(
tmp_path,
mode="note",
want_comments=False,
targets=[_Target("111"), _Target("222")],
)
assert result["notes"] == 1
assert any("222" in error for error in result["errors"])
class TestDegradation:
"""作品列表被挡时的行为 —— 决定了这个功能今天有没有用。"""
@pytest.mark.asyncio
async def test_falls_back_to_refreshing_known_works(self, monkeypatch, tmp_path):
async def blocked(sec_user_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("接口返回了空内容")
refreshed = []
async def fake_detail(aweme_id, *, cookie=""):
refreshed.append(aweme_id)
return _video(aweme_id, likes="9")
monkeypatch.setattr(douyin_api, "author_videos", blocked)
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
result = await _collect(
tmp_path, want_comments=False, known_aweme_ids=["999", "888"]
)
assert refreshed == ["999", "888"]
assert result["notes"] == 2
# 但错误照样报出来 —— 这一轮是「部分可用」,不是「一切正常」,别粉饰。
assert any("作品列表失败" in error for error in result["errors"])
@pytest.mark.asyncio
async def test_known_works_are_not_refetched_for_every_target(self, monkeypatch, tmp_path):
"""退化路径不能每个目标都把同一批已知作品再刷一遍。
任务有多个目标时那会让同一件作品在一轮里出现两次,而指标快照的唯一键是
(task_id, note_id, run_id) —— 第二次插入直接撞键,整个 run 崩掉(踩过:
Duplicate entry for key 'uq_note_metric')。
"""
async def blocked(sec_user_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("接口返回了空内容")
calls = []
async def fake_detail(aweme_id, *, cookie=""):
calls.append(aweme_id)
return _video(aweme_id)
monkeypatch.setattr(douyin_api, "author_videos", blocked)
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
result = await _collect(
tmp_path,
want_comments=False,
targets=[_Target("sec-a"), _Target("sec-b")],
known_aweme_ids=["999"],
)
assert calls == ["999"], "同一件已知作品只该刷一次,而不是每个目标一次"
assert result["notes"] == 1
@pytest.mark.asyncio
async def test_nothing_at_all_still_reports_the_reason(self, monkeypatch, tmp_path):
async def blocked(sec_user_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("接口返回了空内容")
monkeypatch.setattr(douyin_api, "author_videos", blocked)
result = await _collect(tmp_path, want_comments=False)
assert result["notes"] == 0
assert result["errors"]
# 产物仍然写出来(空的),让调用方去判断这是失败而不是「这个博主没作品」。
assert (tmp_path / "douyin" / "jsonl").is_dir()
@pytest.mark.asyncio
async def test_a_failing_comment_fetch_does_not_lose_the_work(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
async def broken_comments(aweme_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("评论接口抽风")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "video_comments", broken_comments)
result = await _collect(tmp_path)
# 评论拿不到是小事,作品不能跟着丢。
assert result["notes"] == 1
assert any("评论失败" in error for error in result["errors"])
class TestCreatorProfile:
"""博主的**账号级**指标 —— 粉丝 / 总获赞 / 作品数。
作品列表给不了这个东西:它说的是一件作品涨了多少赞,不是这个人整个账号的粉丝
在涨还是在掉。单独问一次资料接口。
"""
@staticmethod
def _profiles(tmp_path):
files = list((tmp_path / "douyin" / "jsonl").glob("*_profile_*.jsonl"))
assert len(files) == 1
return _read(files[0])
@pytest.mark.asyncio
async def test_the_profile_lands_in_the_run_dir(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
await _collect(tmp_path, want_comments=False)
assert self._profiles(tmp_path) == [_profile()]
@pytest.mark.asyncio
async def test_the_profile_is_keyed_by_the_same_hash_as_the_works(
self, monkeypatch, tmp_path
):
"""**这条是关键。** 快照表的唯一键是 (任务, creator_hash, 轮次),而界面上是按
作品的 creator_hash 归组去查它的。两边只要差一个字符,粉丝数就永远查不出来 ——
而且是静默的:表里有数据,界面上什么都没有。
"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")] # 作品带的哈希是 "hash"
# 资料接口自己算出来的是另一个值(比如它那边 uid 缺字段、只能拿 sec_uid 算)。
async def off_hash_profile(sec_user_id, *, cookie=""):
return _profile(creator_hash="另一个哈希")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "author_profile", off_hash_profile)
await _collect(tmp_path, want_comments=False)
# 以作品为准:作品才是界面上的行,快照必须挂在能查到它的那个键上。
assert self._profiles(tmp_path)[0]["creator_hash"] == "hash"
@pytest.mark.asyncio
async def test_a_profile_without_an_identity_is_dropped(self, monkeypatch, tmp_path):
"""哈希都算不出来的快照,落下去只会是一条谁也查不到的垃圾。"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return []
async def anonymous_profile(sec_user_id, *, cookie=""):
return _profile(creator_hash="")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "author_profile", anonymous_profile)
result = await _collect(tmp_path, want_comments=False)
assert self._profiles(tmp_path) == []
assert any("身份标识" in error for error in result["errors"])
@pytest.mark.asyncio
async def test_a_failing_profile_does_not_lose_the_works(self, monkeypatch, tmp_path):
"""附加信息拿不到,这一轮采到的作品不能跟着判成失败。"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
async def broken_profile(sec_user_id, *, cookie=""):
raise douyin_api.DouyinApiError("资料接口抽风")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "author_profile", broken_profile)
result = await _collect(tmp_path, want_comments=False)
assert result["notes"] == 1
assert self._profiles(tmp_path) == []
assert any("资料失败" in error for error in result["errors"])
@pytest.mark.asyncio
async def test_a_work_target_never_asks_for_a_profile(self, monkeypatch, tmp_path):
"""作品模式的目标是一件作品,没有「这个博主是谁」可问 —— 不该白发一个请求。"""
asked = []
async def fake_detail(aweme_id, *, cookie=""):
return _video(aweme_id)
async def recording_profile(sec_user_id, *, cookie=""):
asked.append(sec_user_id)
return _profile()
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
monkeypatch.setattr(douyin_api, "author_profile", recording_profile)
await _collect(tmp_path, mode="note", want_comments=False)
assert asked == []
# 文件仍然建出来(空的):事后翻 run 目录能看出「这次根本没问过」。
assert self._profiles(tmp_path) == []
+11
View File
@@ -41,6 +41,17 @@ from database.models import Base, DouyinAweme, DouyinAwemeComment
from tools.user_hash import anonymize_user_id, mask_nickname
@pytest.fixture(autouse=True)
def _force_nickname_masking(monkeypatch):
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
所以这里显式打开来测。
"""
monkeypatch.setattr(config, "MASK_NICKNAME", True)
# 抖音教学版禁用字段(键):不得作为存储 dict 的 key 出现。
FORBIDDEN_KEYS = {
"user_id", "sec_uid", "short_user_id", "user_unique_id",
+67
View File
@@ -0,0 +1,67 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_interpreter.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Tests for the subprocess interpreter resolver."""
import sys
from pathlib import Path
from api.services.interpreter import (
describe_interpreter,
resolve_python_cmd,
venv_python_path,
)
def _make_venv(root: Path) -> Path:
"""Create a fake venv layout and return the expected python path."""
exe = venv_python_path(root)
exe.parent.mkdir(parents=True, exist_ok=True)
exe.write_text("", encoding="utf-8")
return exe
def test_prefers_uv_when_available(monkeypatch, tmp_path):
monkeypatch.setattr("shutil.which", lambda name: "/usr/bin/uv" if name == "uv" else None)
_make_venv(tmp_path)
# uv wins even when a venv exists, matching the upstream documented workflow.
assert resolve_python_cmd(tmp_path) == ["uv", "run", "python"]
def test_falls_back_to_project_venv(monkeypatch, tmp_path):
monkeypatch.setattr("shutil.which", lambda name: None)
exe = _make_venv(tmp_path)
assert resolve_python_cmd(tmp_path) == [str(exe)]
def test_falls_back_to_current_interpreter(monkeypatch, tmp_path):
monkeypatch.setattr("shutil.which", lambda name: None)
# No uv, no venv anywhere under the given root.
assert resolve_python_cmd(tmp_path) == [sys.executable]
def test_describe_is_human_readable(monkeypatch, tmp_path):
monkeypatch.setattr("shutil.which", lambda name: None)
_make_venv(tmp_path)
assert "virtualenv" in describe_interpreter(tmp_path)
monkeypatch.setattr("shutil.which", lambda name: "/usr/bin/uv" if name == "uv" else None)
assert describe_interpreter(tmp_path) == "uv run python"
+13
View File
@@ -18,10 +18,23 @@ import types
import pytest
import config
import store.kuaishou as ks
from store.kuaishou import update_kuaishou_video, update_ks_video_comment
from tools.user_hash import anonymize_user_id, mask_nickname
@pytest.fixture(autouse=True)
def _force_nickname_masking(monkeypatch):
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
所以这里显式打开来测。
"""
monkeypatch.setattr(config, "MASK_NICKNAME", True)
# 教学版禁用字段(键):一律不得出现在存储 dict 中。
# 昵称字段 nickname 允许保留,但值须脱敏。
FORBIDDEN_KEYS = {"user_id", "avatar", "signature", "ip_location", "gender"}
+128
View File
@@ -0,0 +1,128 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_kuaishou_unavailable_video.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
"""快手详情接口遇到「不可用视频」时的回归测试。
回归背景:快手对已删除/私密/不存在的视频返回的是
``visionVideoDetail: {photo: null, author: null}`` —— key 存在、值是 null。
而 ``detail.get("photo", {})`` 只在 key **缺失** 时给默认值,key 存在且为 null
时拿到的仍是 ``None``,紧接着的 ``photo.get(...)`` 抛 AttributeError;
该异常不在 ``get_video_info_task`` 的 except 列表里,又会穿过
``asyncio.gather``,把整轮爬取直接带崩。
这里不发起网络请求,只用 stub client 驱动真实的任务函数。
"""
from __future__ import annotations
import asyncio
import random
import sys
from pathlib import Path
from types import SimpleNamespace
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
import config # noqa: E402
from media_platform.kuaishou.core import KuaishouCrawler # noqa: E402
PHOTO_ID = "3x3zxz4mjrsc8ke"
UNAVAILABLE_PHOTO_ID = "3xf8enb8dbj6uig"
# 真实响应里 photo/author 为 null 的那个视频
UNAVAILABLE_DETAIL = {
"visionVideoDetail": {
"status": 1,
"type": "video",
"author": None,
"photo": None,
"tags": [],
}
}
NORMAL_DETAIL = {
"visionVideoDetail": {
"status": 1,
"type": "video",
"author": {"name": "余胜军说Java"},
"photo": {"id": PHOTO_ID, "caption": "我教你学Python", "likeCount": 167000},
"tags": [],
}
}
class _StubClient:
def __init__(self, payload):
self._payload = payload
async def get_video_info(self, photo_id): # noqa: ANN001
return self._payload
def _make_crawler(payload) -> SimpleNamespace:
"""只借 get_video_info_task 用到的 self.ks_client,不需要完整 crawler"""
return SimpleNamespace(ks_client=_StubClient(payload))
@pytest.fixture(autouse=True)
def _no_sleep(monkeypatch):
"""去掉任务里的固定延时与随机抖动,测试不应真的等待"""
monkeypatch.setattr(config, "CRAWLER_MAX_SLEEP_SEC", 0)
monkeypatch.setattr(random, "uniform", lambda a, b: 0)
@pytest.mark.asyncio
async def test_unavailable_video_is_skipped_instead_of_crashing():
"""photo 为 null 时返回 None(跳过),而不是抛 AttributeError"""
crawler = _make_crawler(UNAVAILABLE_DETAIL)
result = await KuaishouCrawler.get_video_info_task(
crawler, UNAVAILABLE_PHOTO_ID, asyncio.Semaphore(1)
)
assert result is None
@pytest.mark.asyncio
async def test_missing_photo_key_is_also_skipped():
"""photo 字段整个缺失时同样跳过(.get 的默认值路径)"""
crawler = _make_crawler({"visionVideoDetail": {"status": 1, "author": None}})
result = await KuaishouCrawler.get_video_info_task(
crawler, UNAVAILABLE_PHOTO_ID, asyncio.Semaphore(1)
)
assert result is None
@pytest.mark.asyncio
async def test_normal_video_still_returns_detail():
"""正常视频不受影响"""
crawler = _make_crawler(NORMAL_DETAIL)
result = await KuaishouCrawler.get_video_info_task(
crawler, PHOTO_ID, asyncio.Semaphore(1)
)
assert result is not None
assert result["photo"]["id"] == PHOTO_ID
@pytest.mark.asyncio
async def test_empty_vision_video_detail_returns_none():
"""visionVideoDetail 整体缺失时返回 None"""
crawler = _make_crawler({})
result = await KuaishouCrawler.get_video_info_task(
crawler, PHOTO_ID, asyncio.Semaphore(1)
)
assert result is None
+73
View File
@@ -0,0 +1,73 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_kuaishou_url_parse.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
"""快手视频输入解析的回归测试。
覆盖三种输入形态:
1. 纯视频 ID
2. 标准视频页 ``/short-video/<id>``
3. 分享短链 ``/f/<share_token>`` —— 注意路径里是 **share_token 而不是视频 ID**,
必须跟随 302 重定向才能拿到真实 ID,所以只标记 ``url_type="short"`` 交给调用方。
这里只测纯函数,不发网络请求。
"""
from __future__ import annotations
import pytest
from media_platform.kuaishou.help import parse_video_info_from_url
def test_pure_video_id_is_normal():
info = parse_video_info_from_url("3xf8enb8dbj6uig")
assert info.video_id == "3xf8enb8dbj6uig"
assert info.url_type == "normal"
def test_short_video_url_is_normal():
info = parse_video_info_from_url(
"https://www.kuaishou.com/short-video/3x3zxz4mjrsc8ke"
"?authorId=3x84qugg4ch9zhs&streamSource=search&area=searchxxnull&searchKey=python"
)
assert info.video_id == "3x3zxz4mjrsc8ke"
assert info.url_type == "normal"
def test_share_short_link_is_marked_short():
"""分享短链必须标记为 short —— 路径里的 token 不是视频 ID"""
info = parse_video_info_from_url("https://www.kuaishou.com/f/X9Idt15MQb9L2cv")
assert info.video_id == "X9Idt15MQb9L2cv"
assert info.url_type == "short"
def test_share_short_link_with_dash_in_token():
info = parse_video_info_from_url("https://www.kuaishou.com/f/X-a8vLyTxvEvN2jg")
assert info.video_id == "X-a8vLyTxvEvN2jg"
assert info.url_type == "short"
def test_short_video_url_is_not_confused_with_share_link():
"""带 query 的标准视频页不能被误判成短链"""
info = parse_video_info_from_url(
"https://www.kuaishou.com/short-video/3xyziwesje8e9jg"
"?shareToken=X9Idt15MQb9L2cv&shareObjectId=3xyziwesje8e9jg"
)
assert info.video_id == "3xyziwesje8e9jg"
assert info.url_type == "normal"
def test_unparsable_url_raises():
with pytest.raises(ValueError):
parse_video_info_from_url("https://www.kuaishou.com/profile/3x84qugg4ch9zhs")
+319
View File
@@ -0,0 +1,319 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_monitor_api.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""API-level tests for the monitoring endpoints.
Run against an ASGI transport with a temporary database, so no server, network
or login is required. Lifespan is deliberately not exercised: it would start the
scheduler, and these tests only cover routing, validation and persistence.
"""
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import db as monitor_db
from api.monitor.service import TargetParseError, parse_target_input
CREATOR_URL = (
"https://www.xiaohongshu.com/user/profile/5f58bd990000000001003753"
"?xsec_token=ABYVg1evluJZZzpMX-VWzchxQ1qSNVW3r-jOEnKqMcgZw=&xsec_source=pc_search"
)
NOTE_URL = "https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=TOKEN&xsec_source=pc_search"
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
class TestParseTargetInput:
def test_full_url_splits_id_from_token(self):
"""The id is the stable key; the token is a refreshable credential."""
parsed = parse_target_input(CREATOR_URL, "creator")
assert parsed["external_id"] == "5f58bd990000000001003753"
assert parsed["xsec_token"].startswith("ABYVg1evluJZZzpMX")
assert parsed["xsec_source"] == "pc_search"
def test_bare_id_is_accepted(self):
parsed = parse_target_input("5f58bd990000000001003753", "creator")
assert parsed["external_id"] == "5f58bd990000000001003753"
assert parsed["xsec_token"] == ""
def test_note_url_without_token_still_parses(self):
parsed = parse_target_input(
"https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8", "note"
)
assert parsed["external_id"] == "6aa3d827000000002802c5c8"
assert parsed["xsec_token"] == ""
def test_creator_url_rejected_in_note_mode(self):
with pytest.raises(TargetParseError):
parse_target_input(CREATOR_URL, "note")
def test_garbage_is_rejected(self):
with pytest.raises(TargetParseError):
parse_target_input("not a url at all !!", "creator")
# --- 抖音 -------------------------------------------------------------
# 链接形态由平台决定,所以每一个都要显式带上 "dy"。
def test_douyin_creator_url(self):
parsed = parse_target_input(
"https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
"?from_tab_name=main",
"creator",
"dy",
)
assert (
parsed["external_id"]
== "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
)
def test_douyin_video_url(self):
parsed = parse_target_input(
"https://www.douyin.com/video/7525082444551310602", "note", "dy"
)
assert parsed["external_id"] == "7525082444551310602"
def test_douyin_modal_id_url(self):
"""在别人主页或搜索结果里点开视频,拿到的就是带 modal_id 的链接。"""
parsed = parse_target_input(
"https://www.douyin.com/root/search/python?aid=b733a3b0&modal_id=7471165520058862848",
"note",
"dy",
)
assert parsed["external_id"] == "7471165520058862848"
def test_douyin_bare_sec_uid_is_accepted(self):
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
parsed = parse_target_input(sec_uid, "creator", "dy")
assert parsed["external_id"] == sec_uid
def test_douyin_bare_sec_uid_beyond_the_xhs_length_cap(self):
"""裸 id 的长度上限必须按平台分开。
小红书那条规则封顶 64 字符,而 sec_user_id 长过 64 是常态(实测样本 55,
但字段本身是变长的)。共用一条规则的话,长一点的 sec_uid 会被直接拒掉 ——
对用户来说就是「粘贴了一个完全正确的链接却报无法识别」。
"""
sec_uid = "MS4wLjABAAAA" + "aB3dEf6hIj9lMn2pQr5tUv8xYz1" * 3
assert len(sec_uid) > 64
parsed = parse_target_input(sec_uid, "creator", "dy")
assert parsed["external_id"] == sec_uid
# 同一条 id 拿小红书规则来解析会被拒 —— 这正是两条规则必须分开的原因。
with pytest.raises(TargetParseError):
parse_target_input(sec_uid, "creator", "xhs")
def test_douyin_bare_video_id(self):
parsed = parse_target_input("7525082444551310602", "note", "dy")
assert parsed["external_id"] == "7525082444551310602"
# 抖音不需要 xsec_token —— 和小红书不同,裸链接就能用。
assert parsed["xsec_token"] == ""
def test_douyin_short_link_is_rejected_with_a_reason(self):
"""短链要联网跳一次才知道指向谁。明确拒绝好过存一个永远抓不到东西的目标。"""
with pytest.raises(TargetParseError) as excinfo:
parse_target_input("https://v.douyin.com/drIPtQ_WPWY/", "note", "dy")
assert "短链" in str(excinfo.value)
def test_a_douyin_link_is_not_parsed_with_xhs_rules(self):
with pytest.raises(TargetParseError):
parse_target_input(
"https://www.douyin.com/video/7525082444551310602", "note", "xhs"
)
def test_an_xhs_link_is_not_parsed_for_douyin(self):
with pytest.raises(TargetParseError):
parse_target_input(NOTE_URL, "note", "dy")
def test_a_platform_without_an_adapter_is_rejected(self):
with pytest.raises(TargetParseError):
parse_target_input("whatever", "creator", "bili")
class TestTargetReplacement:
@pytest.mark.asyncio
async def test_replacing_targets_uses_the_tasks_own_platform(self, client):
"""改目标必须按任务**自己**的平台解析。
``update_task`` 原先漏传了 platform,解析回落到默认的小红书。只有小红书时
行为恰好正确,接上抖音就会拿小红书的正则去解析抖音链接 —— 建任务时对、
改任务时错,是最难注意到的那种不一致。
"""
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
created = await client.post(
"/api/monitor/tasks",
json={"name": "抖音", "mode": "creator", "platform": "dy", "targets": [sec_uid]},
)
assert created.status_code == 201
task_id = created.json()["id"]
updated = await client.patch(
f"/api/monitor/tasks/{task_id}",
json={"targets": [f"https://www.douyin.com/user/{sec_uid}"]},
)
assert updated.status_code == 200
tasks = (
await client.get("/api/monitor/tasks", params={"platform": "dy"})
).json()["tasks"]
target = tasks[0]["targets"][0]
assert target["external_id"] == sec_uid
assert target["raw_value"].startswith("https://www.douyin.com/user/")
class TestTaskCrud:
@pytest.mark.asyncio
async def test_create_and_list_task(self, client):
response = await client.post(
"/api/monitor/tasks",
json={
"name": "网文作者监控",
"mode": "creator",
"interval_minutes": 120,
"targets": [CREATOR_URL, "5f58bd990000000001003754"],
},
)
assert response.status_code == 201
task_id = response.json()["id"]
listing = await client.get("/api/monitor/tasks")
assert listing.status_code == 200
tasks = listing.json()["tasks"]
assert len(tasks) == 1
assert tasks[0]["id"] == task_id
assert tasks[0]["target_count"] == 2
# next_run_at is persisted so the schedule survives a restart.
assert tasks[0]["next_run_at"] is not None
@pytest.mark.asyncio
async def test_duplicate_targets_are_deduplicated(self, client):
response = await client.post(
"/api/monitor/tasks",
json={
"name": "dedup",
"mode": "creator",
"targets": [CREATOR_URL, CREATOR_URL],
},
)
assert response.status_code == 201
listing = await client.get("/api/monitor/tasks")
assert listing.json()["tasks"][0]["target_count"] == 1
@pytest.mark.asyncio
async def test_invalid_target_returns_400(self, client):
response = await client.post(
"/api/monitor/tasks",
json={"name": "bad", "mode": "creator", "targets": ["!!! nonsense !!!"]},
)
assert response.status_code == 400
@pytest.mark.asyncio
async def test_interval_floor_is_enforced(self, client):
"""A tight poll loop is the pattern that triggers platform rate limits."""
response = await client.post(
"/api/monitor/tasks",
json={"name": "too fast", "mode": "creator", "interval_minutes": 1, "targets": [CREATOR_URL]},
)
assert response.status_code == 422
@pytest.mark.asyncio
async def test_update_and_delete(self, client):
created = await client.post(
"/api/monitor/tasks",
json={"name": "t", "mode": "note", "targets": [NOTE_URL]},
)
task_id = created.json()["id"]
patched = await client.patch(f"/api/monitor/tasks/{task_id}", json={"enabled": False})
assert patched.status_code == 200
listing = await client.get("/api/monitor/tasks")
assert listing.json()["tasks"][0]["enabled"] is False
deleted = await client.delete(f"/api/monitor/tasks/{task_id}")
assert deleted.status_code == 200
assert (await client.get("/api/monitor/tasks")).json()["tasks"] == []
@pytest.mark.asyncio
async def test_run_now_on_missing_task_is_404(self, client):
response = await client.post("/api/monitor/tasks/9999/run")
assert response.status_code == 404
@pytest.mark.asyncio
async def test_run_history_starts_empty(self, client):
created = await client.post(
"/api/monitor/tasks",
json={"name": "t", "mode": "creator", "targets": [CREATOR_URL]},
)
task_id = created.json()["id"]
runs = await client.get(f"/api/monitor/tasks/{task_id}/runs")
assert runs.status_code == 200
assert runs.json()["runs"] == []
class TestCookieEndpoints:
@pytest.mark.asyncio
async def test_cookie_value_is_never_returned(self, client):
"""The GET must expose health only, never the credential."""
secret = "web_session=SUPERSECRETVALUE; a1=abc123"
saved = await client.post("/api/monitor/cookie", json={"cookie": secret})
assert saved.status_code == 200
status_response = await client.get("/api/monitor/cookie")
assert status_response.status_code == 200
body = status_response.json()
assert body["present"] is True
assert body["length"] == len(secret)
assert "SUPERSECRETVALUE" not in status_response.text
@pytest.mark.asyncio
async def test_cookie_initially_absent_and_clearable(self, client):
assert (await client.get("/api/monitor/cookie")).json()["present"] is False
await client.post("/api/monitor/cookie", json={"cookie": "web_session=x"})
assert (await client.get("/api/monitor/cookie")).json()["present"] is True
await client.delete("/api/monitor/cookie")
assert (await client.get("/api/monitor/cookie")).json()["present"] is False
class TestDashboardQueries:
@pytest.mark.asyncio
async def test_empty_dashboard_shapes(self, client):
assert (await client.get("/api/monitor/notes")).json()["notes"] == []
assert (await client.get("/api/monitor/comments")).json()["comments"] == []
assert (await client.get("/api/monitor/events")).json()["events"] == []
overview = (await client.get("/api/monitor/overview")).json()
assert overview["tasks"] == 0
assert overview["notes"] == 0
+83
View File
@@ -0,0 +1,83 @@
# -*- coding: utf-8 -*-
"""Guards for the in-place column migration.
``create_all`` creates missing tables but never adds columns to a table that
already exists, so new model columns are applied by ``_ensure_columns``. That step
was originally driven by a hand-kept list, and forgetting to update it did not
fail loudly -- the app still started, connected, and then failed on every query
and every scheduler tick. These tests pin down its replacement, which derives the
work from the model metadata.
"""
import pytest
from sqlalchemy import Boolean, Column, Integer, MetaData, String, Table
from api.monitor import db
from api.monitor.models import MonitorBase
def _migratable_columns():
for table in MonitorBase.metadata.sorted_tables:
for column in table.columns:
if column.primary_key:
continue
yield pytest.param(column, id=f"{table.name}.{column.name}")
def _ddl(column: Column) -> str:
"""Render a detached column, so the tests never mutate the real metadata."""
scratch = Table("scratch", MetaData(), column)
return db._column_ddl(scratch.columns[0])
@pytest.mark.parametrize("column", _migratable_columns())
def test_column_renders_as_ddl(column):
ddl = db._column_ddl(column)
# SQLAlchemy back-quotes an identifier only when it has to, so the rendered
# name matches the column's with the quoting stripped -- which is the case for
# reserved words like monitor_run.trigger. That quoting is a feature: the
# hand-kept list this replaced would have emitted bare `trigger` and died on a
# syntax error.
assert ddl.split(" ", 1)[0].strip("`") == column.name
# MySQL refuses AUTO_INCREMENT together with the DEFAULT this helper appends
# to NOT NULL columns.
assert "AUTO_INCREMENT" not in ddl
@pytest.mark.parametrize("column", _migratable_columns())
def test_not_null_columns_carry_a_default(column):
"""So ADD COLUMN cannot fail on a table that already holds rows.
Without a DEFAULT, whether the ALTER succeeds depends on the server's
sql_mode -- not something a deployment should hinge on.
"""
if column.nullable:
pytest.skip("nullable column needs no seed value")
assert "DEFAULT" in db._column_ddl(column)
def test_boolean_default_becomes_a_mysql_literal():
"""Python's True is not a SQL keyword; it has to become 1."""
assert "DEFAULT 1" in _ddl(Column("flag", Boolean, nullable=False, default=True))
def test_string_defaults_are_quoted():
assert "DEFAULT 'interval'" in _ddl(
Column("mode", String(16), nullable=False, default="interval")
)
def test_integer_defaults_are_not_quoted():
ddl = _ddl(Column("n", Integer, nullable=False, default=0))
assert "DEFAULT 0" in ddl
assert "DEFAULT '0'" not in ddl
def test_a_not_null_column_without_a_model_default_falls_back_to_zero():
"""Belt and braces: even a column the model gives no default still migrates."""
ddl = _ddl(Column("n", Integer, nullable=False))
assert "DEFAULT 0" in ddl
+372
View File
@@ -0,0 +1,372 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_monitor_comments.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Comment note-association, grouping, and the export endpoint."""
import csv
import io
import re
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import db as monitor_db
from api.monitor.models import (
MODE_CREATOR,
MonitorComment,
MonitorNote,
MonitorTask,
)
TASK_NAME = "评论归属测试"
# 作品的发布时间。和 first_seen_at(我们第一次看到它)刻意取不同的值 —— 两者混成
# 一个概念是最容易犯的错。
PUBLISHED_A = 1_699_000_000_000
PUBLISHED_B = 1_699_100_000_000
async def _seed():
"""Two works; three comments on the first, one on the second."""
async with monitor_db.get_session() as session:
task = MonitorTask(
name=TASK_NAME, platform="xhs", mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=True,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
session.add(task)
await session.flush()
# 两个作品**属于不同的博主** —— 评论流最外层按创作者分组,同一个人就没得测了。
for note_id, title, creator_hash, creator_name, published_at in (
("note-a", "作品甲", "hash-a", "博主甲", PUBLISHED_A),
("note-b", "作品乙", "hash-b", "博主乙", PUBLISHED_B),
):
session.add(
MonitorNote(
task_id=task.id, note_id=note_id, title=title,
note_url=f"https://www.xiaohongshu.com/explore/{note_id}",
cover=f"https://img/{note_id}.jpg", creator_hash=creator_hash,
creator_name=creator_name,
source_kind="video", published_at=published_at,
first_seen_run_id=1, first_seen_at=1_700_000_000_000,
last_seen_run_id=1, last_seen_at=1_700_000_000_000,
)
)
# note-a has three comments, note-b has one.
plan = [
("c1", "note-a", 1_700_000_001_000),
("c2", "note-a", 1_700_000_002_000),
("c3", "note-a", 1_700_000_003_000),
("c4", "note-b", 1_700_000_004_000),
]
for comment_id, note_id, seen_at in plan:
session.add(
MonitorComment(
task_id=task.id, note_id=note_id, comment_id=comment_id,
content=f"内容-{comment_id}", nickname="u***r", creator_hash="h",
create_time=seen_at, like_count=1, sub_comment_count=0,
parent_comment_id="", first_seen_run_id=1, first_seen_at=seen_at,
)
)
return task.id
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
await _seed()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
class TestCommentsCarryTheirNote:
@pytest.mark.asyncio
async def test_each_comment_names_its_work(self, client):
"""A bare note_id is unreadable -- the title is the whole point."""
response = await client.get("/api/monitor/comments")
assert response.status_code == 200
comments = response.json()["comments"]
assert len(comments) == 4
by_id = {c["comment_id"]: c for c in comments}
assert by_id["c1"]["note_title"] == "作品甲"
assert by_id["c1"]["note_url"].endswith("note-a")
assert by_id["c1"]["note_cover"].endswith("note-a.jpg")
assert by_id["c4"]["note_title"] == "作品乙"
@pytest.mark.asyncio
async def test_note_id_filters_the_stream(self, client):
response = await client.get("/api/monitor/comments", params={"note_id": "note-a"})
comments = response.json()["comments"]
assert {c["comment_id"] for c in comments} == {"c1", "c2", "c3"}
class TestGroupByNote:
@pytest.mark.asyncio
async def test_groups_bucket_by_work(self, client):
response = await client.get("/api/monitor/comments", params={"group_by": "note"})
body = response.json()
assert "groups" in body
assert body["total"] == 4
groups = {g["note_id"]: g for g in body["groups"]}
assert set(groups) == {"note-a", "note-b"}
assert len(groups["note-a"]["comments"]) == 3
assert len(groups["note-b"]["comments"]) == 1
assert groups["note-a"]["note_title"] == "作品甲"
@pytest.mark.asyncio
async def test_newest_group_comes_first(self, client):
"""The UI expands the first group by default, so it must be the newest."""
response = await client.get("/api/monitor/comments", params={"group_by": "note"})
groups = response.json()["groups"]
# note-b's only comment is the most recent overall.
assert groups[0]["note_id"] == "note-b"
@pytest.mark.asyncio
async def test_each_bucket_carries_its_creator(self, client):
"""桶上必须带作品的创作者 —— 评论流最外层就是按它分组的。
少了这两个字段,前端拿到的 creator_hash / creator_name 都是 undefined,
于是所有博主塌成同一个分组、标签回退成「未知博主」:一个人都分不出来。
"""
groups = {
group["note_id"]: group
for group in (
await client.get("/api/monitor/comments", params={"group_by": "note"})
).json()["groups"]
}
assert groups["note-a"]["creator_hash"] == "hash-a"
assert groups["note-a"]["creator_name"] == "博主甲"
assert groups["note-b"]["creator_hash"] == "hash-b"
assert groups["note-b"]["creator_name"] == "博主乙"
# 两个作品的创作者必须真的不同,否则界面上照样分不出来。
assert groups["note-a"]["creator_hash"] != groups["note-b"]["creator_hash"]
@pytest.mark.asyncio
async def test_each_bucket_carries_the_publish_date(self, client):
"""作品那一层要带发布日期:同名作品不少,日期能帮着认。
注意它和 first_seen_at 是两个概念 —— 前者是作者发布的那天,后者是我们第一次
看到它的那天。把老作品加进监控时两者能差好几个月。
"""
groups = {
group["note_id"]: group
for group in (
await client.get("/api/monitor/comments", params={"group_by": "note"})
).json()["groups"]
}
assert groups["note-a"]["published_at"] == PUBLISHED_A
assert groups["note-b"]["published_at"] == PUBLISHED_B
assert groups["note-a"]["published_at"] != groups["note-b"]["published_at"]
@pytest.mark.asyncio
async def test_the_notes_endpoint_exposes_the_publish_date(self, client):
"""作品列表也要带上它 —— 作品栏就是靠这个显示「发布日期」列的。"""
notes = {
note["note_id"]: note
for note in (await client.get("/api/monitor/notes")).json()["notes"]
}
assert notes["note-a"]["published_at"] == PUBLISHED_A
assert notes["note-b"]["published_at"] == PUBLISHED_B
# 和「首次发现」不是同一个值 —— 两者混了的话这个断言会抓到。
assert notes["note-a"]["first_seen_at"] != notes["note-a"]["published_at"]
@pytest.mark.asyncio
async def test_flat_shape_is_unchanged_without_the_flag(self, client):
body = (await client.get("/api/monitor/comments")).json()
assert "comments" in body and "groups" not in body
class TestCommentNoteFilterOptions:
@pytest.mark.asyncio
async def test_options_carry_counts_and_titles(self, client):
response = await client.get("/api/monitor/comment-notes")
assert response.status_code == 200
notes = {n["note_id"]: n for n in response.json()["notes"]}
assert notes["note-a"]["comment_count"] == 3
assert notes["note-b"]["comment_count"] == 1
assert notes["note-a"]["note_title"] == "作品甲"
@pytest.mark.asyncio
async def test_scoped_to_a_task(self, client):
tasks = (await client.get("/api/monitor/tasks")).json()["tasks"]
task_id = tasks[0]["id"]
scoped = await client.get("/api/monitor/comment-notes", params={"task_id": task_id})
assert len(scoped.json()["notes"]) == 2
# A task with no comments yields an empty list, not an error.
other = await client.get("/api/monitor/comment-notes", params={"task_id": 9999})
assert other.json()["notes"] == []
class TestExport:
@pytest.mark.asyncio
async def test_csv_has_a_bom_so_excel_does_not_mangle_chinese(self, client):
response = await client.get("/api/monitor/export", params={"kind": "comments"})
assert response.status_code == 200
assert response.content.startswith(b"\xef\xbb\xbf")
assert "attachment" in response.headers["content-disposition"]
text = response.content.decode("utf-8-sig")
rows = list(csv.DictReader(io.StringIO(text)))
assert len(rows) == 4
assert rows[0]["所属作品"] in ("作品甲", "作品乙")
@pytest.mark.asyncio
async def test_notes_export(self, client):
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
assert {r["作品ID"] for r in rows} == {"note-a", "note-b"}
@pytest.mark.asyncio
async def test_xlsx_is_a_readable_workbook(self, client):
from openpyxl import load_workbook
response = await client.get(
"/api/monitor/export", params={"kind": "comments", "format": "xlsx"}
)
assert response.status_code == 200
workbook = load_workbook(io.BytesIO(response.content))
sheet = workbook.active
assert sheet.max_row == 5 # header + four comments
assert sheet.cell(row=1, column=1).value == "博主昵称"
assert sheet.cell(row=1, column=2).value == "所属作品"
@pytest.mark.asyncio
async def test_report_export(self, client):
response = await client.get(
"/api/monitor/export",
params={"kind": "report", "days": 3},
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
assert len(rows) == 3
assert "日期" in rows[0]
@pytest.mark.asyncio
async def test_unknown_kind_and_format_are_rejected(self, client):
assert (
await client.get("/api/monitor/export", params={"kind": "nope"})
).status_code == 400
assert (
await client.get("/api/monitor/export", params={"kind": "notes", "format": "pdf"})
).status_code == 400
@pytest.mark.asyncio
async def test_empty_selection_is_a_404_not_an_empty_file(self, client):
"""An empty download looks like a bug; say so instead."""
response = await client.get(
"/api/monitor/export", params={"kind": "comments", "note_id": "no-such-note"}
)
assert response.status_code == 404
class TestNotesExportColumns:
"""作品导出的列 —— 「有列名」和「列里有数」是两回事。
原先这几列写的是裸键名 `liked_count`,而作品行的指标是嵌在 `metrics` 里的,
于是导出来的表有「点赞/评论/收藏/分享」四列,**每一格都是空的**,还没人发现 ——
因为原来的测试只断言了 `作品ID`。
"""
@pytest.mark.asyncio
async def test_the_metric_columns_actually_contain_numbers(self, client):
from api.monitor.models import MonitorNoteMetric
async with monitor_db.get_session() as session:
from sqlalchemy import select
task_id = (await session.scalar(select(MonitorTask.id))).__int__()
session.add(
MonitorNoteMetric(
task_id=task_id, note_id="note-a", run_id=1,
captured_at=1_700_000_000_000,
liked_count=123, comment_count=45,
collected_count=6, share_count=7,
)
)
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
by_id = {row["作品ID"]: row for row in rows}
assert by_id["note-a"]["点赞"] == "123"
assert by_id["note-a"]["评论"] == "45"
assert by_id["note-a"]["收藏"] == "6"
assert by_id["note-a"]["分享"] == "7"
@pytest.mark.asyncio
async def test_a_missing_metric_is_left_empty_not_zero(self, client):
"""没采到的指标留空。写 0 的话,导出来的表会声称这条作品零互动。"""
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
assert rows[0]["点赞"] == ""
@pytest.mark.asyncio
async def test_the_export_says_who_the_creator_is(self, client):
"""一行只有作品 ID 没法用 —— 导出来是拿去比对和汇报的。
备注优先:昵称常常认不出是谁,而备注是人自己起的名字。
"""
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
assert {row["博主昵称"] for row in rows} == {"博主甲", "博主乙"}
@pytest.mark.asyncio
async def test_times_are_readable_not_raw_milliseconds(self, client):
"""毫秒时间戳倒进 CSV 就是 13 位数字,打开 Excel 的人没法看、也没法排序。"""
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
by_id = {row["作品ID"]: row for row in rows}
# 断言**形状**而不是具体时刻:格式化用的是服务器本地时区,写死一个字符串的话
# 换个时区的机器上就会红。
assert re.fullmatch(r"\d{4}-\d{2}-\d{2} \d{2}:\d{2}", by_id["note-a"]["发布时间"])
assert re.fullmatch(r"\d{4}-\d{2}-\d{2} \d{2}:\d{2}", by_id["note-a"]["首次发现"])
# 而且不能是原始毫秒。
assert by_id["note-a"]["发布时间"] != str(PUBLISHED_A)
+322
View File
@@ -0,0 +1,322 @@
# -*- coding: utf-8 -*-
"""作品栏里的**博主**:备注(他到底是谁)与账号级指标(他现在多大)。
按 creator_hash 分组、显示 creator_name,两样都认不出人:一个是哈希,一个是平台昵称。
备注是人自己起的名字。账号级指标则是作品列表给不了的东西 —— 作品说的是"这条涨了多少赞",
粉丝数说的是"这个人整个账号在涨还是在掉"。
"""
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import db as monitor_db
from api.monitor.models import (
MODE_CREATOR,
MonitorCreatorStat,
MonitorNote,
MonitorTask,
)
CREATOR_HASH = "hash-a"
NICKNAME = "张三"
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
async def _seed(platform: str = "xhs", task_name: str = "任务") -> int:
"""一个任务 + 一条作品,博主固定用 CREATOR_HASH / NICKNAME。"""
async with monitor_db.get_session() as session:
task = MonitorTask(
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
session.add(task)
await session.flush()
session.add(
MonitorNote(
task_id=task.id, note_id=f"{platform}-n1", title="作品",
note_url="", cover="", creator_hash=CREATOR_HASH,
creator_name=NICKNAME, source_kind="", published_at=None,
first_seen_run_id=1, first_seen_at=0,
last_seen_run_id=1, last_seen_at=0,
)
)
return task.id
async def _notes(client, platform: str = "xhs"):
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["notes"]
async def _set_alias(client, alias: str, platform: str = "xhs", creator_hash=CREATOR_HASH):
return await client.put(
f"/api/monitor/creators/{creator_hash}",
params={"platform": platform},
json={"alias": alias},
)
class TestCreatorAlias:
@pytest.mark.asyncio
async def test_notes_start_without_an_alias(self, client):
await _seed()
assert (await _notes(client))[0]["creator_alias"] == ""
@pytest.mark.asyncio
async def test_an_alias_comes_back_with_the_notes(self, client):
await _seed()
response = await _set_alias(client, "竞品A")
assert response.status_code == 200
note = (await _notes(client))[0]
assert note["creator_alias"] == "竞品A"
# 备注是**叠加**在昵称之上的,不是替换 —— 昵称仍然是有用的对照。
assert note["creator_name"] == NICKNAME
@pytest.mark.asyncio
async def test_an_alias_is_shared_across_tasks(self, client):
"""同一个博主出现在两个任务里,备注只该填一次。
creator_hash 对同一个 uid 是稳定的,所以键取 (platform, creator_hash) 而不是
按任务存 —— 否则每加一个任务都要重新认一遍人。
"""
await _seed(task_name="任务甲")
await _seed(task_name="任务乙")
await _set_alias(client, "竞品A")
for note in await _notes(client):
assert note["creator_alias"] == "竞品A"
@pytest.mark.asyncio
async def test_the_alias_does_not_leak_to_another_platform(self, client):
"""同一个哈希在另一个平台上是另一个(或同一个)人 —— 别串味。"""
await _seed(platform="xhs")
await _seed(platform="dy")
await _set_alias(client, "小红书那边的", platform="xhs")
assert (await _notes(client, "xhs"))[0]["creator_alias"] == "小红书那边的"
assert (await _notes(client, "dy"))[0]["creator_alias"] == ""
@pytest.mark.asyncio
async def test_an_empty_alias_clears_it(self, client):
await _seed()
await _set_alias(client, "竞品A")
await _set_alias(client, "")
assert (await _notes(client))[0]["creator_alias"] == ""
@pytest.mark.asyncio
async def test_an_alias_is_trimmed(self, client):
await _seed()
await _set_alias(client, " 竞品A ")
assert (await _notes(client))[0]["creator_alias"] == "竞品A"
async def _add_stat(
task_id: int,
run_id: int,
fans: int | None,
*,
total_favorited: int | None = 83000,
works: int | None = 42,
creator_hash: str = CREATOR_HASH,
captured_at: int = 1,
) -> None:
async with monitor_db.get_session() as session:
session.add(
MonitorCreatorStat(
task_id=task_id,
run_id=run_id,
creator_hash=creator_hash,
nickname=NICKNAME,
fans=fans,
total_favorited=total_favorited,
works_count=works,
following=7,
captured_at=captured_at,
)
)
class TestCreatorStats:
"""账号级指标跟着作品一起返回 —— 界面上是按博主归组的,为了一个组头再发一轮请求
没道理。"""
@pytest.mark.asyncio
async def test_without_snapshots_the_fields_are_null(self, client):
"""null 而不是 0:0 会显示成「粉丝 0」,而事实是"还没采到"。"""
await _seed()
note = (await _notes(client))[0]
assert note["creator_fans"] is None
assert note["creator_total_favorited"] is None
assert note["creator_works"] is None
@pytest.mark.asyncio
async def test_a_snapshot_shows_up_on_every_work_of_that_creator(self, client):
task_id = await _seed()
await _add_stat(task_id, run_id=1, fans=12000)
for note in await _notes(client):
assert note["creator_fans"] == 12000
assert note["creator_total_favorited"] == 83000
assert note["creator_works"] == 42
@pytest.mark.asyncio
async def test_the_latest_run_wins(self, client):
"""一轮一条,所以总会有好几条 —— 给界面的必须是最近那条。"""
task_id = await _seed()
await _add_stat(task_id, run_id=1, fans=12000)
await _add_stat(task_id, run_id=2, fans=12300)
assert (await _notes(client))[0]["creator_fans"] == 12300
@pytest.mark.asyncio
async def test_another_platforms_snapshot_does_not_leak(self, client):
"""两个平台上恰好同名同哈希的博主是两个人 —— 快照挂在任务上,不该串。"""
await _seed(platform="xhs")
dy_task_id = await _seed(platform="dy")
await _add_stat(dy_task_id, run_id=1, fans=999)
assert (await _notes(client, "xhs"))[0]["creator_fans"] is None
assert (await _notes(client, "dy"))[0]["creator_fans"] == 999
@pytest.mark.asyncio
async def test_an_unparsed_count_stays_null(self, client):
"""快照在,但某一项没解析出来 —— 那一项必须是 null,不能变成 0。
和上一条的区别:那条是"根本没有快照",这条是"有快照、其中一项平台没给"。
界面上两种都该是「—」。
"""
task_id = await _seed()
await _add_stat(
task_id, run_id=1, fans=None, total_favorited=None, works=None
)
note = (await _notes(client))[0]
assert note["creator_fans"] is None
assert note["creator_works"] is None
# 快照本身是有的(有采集时间),只是值不知道 —— 前端要能分开这两件事。
assert note["creator_stats_at"] is not None
async def _seed_task(platform: str = "xhs", task_name: str = "空任务") -> int:
"""只有任务,**一条作品都没有**。"""
async with monitor_db.get_session() as session:
task = MonitorTask(
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
session.add(task)
await session.flush()
return task.id
async def _creators(client, platform: str = "xhs"):
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["creators"]
class TestCreatorList:
"""作品栏要显示谁 —— **包括一条作品都没有的博主**。
组原先是从作品推出来的,于是「目标加了、资料采到了、粉丝数在库里,界面上什么都
没有」。而这类博主恰恰最该看见:还在涨粉,只是最近没发东西。
"""
@pytest.mark.asyncio
async def test_a_creator_with_no_works_still_shows_up(self, client):
"""**这条就是这个改动的全部理由。**"""
task_id = await _seed_task()
await _add_stat(task_id, run_id=1, fans=12000)
creators = await _creators(client)
assert len(creators) == 1
assert creators[0]["creator_hash"] == CREATOR_HASH
assert creators[0]["creator_name"] == NICKNAME # 只剩快照这一个来源
assert creators[0]["note_count"] == 0
assert creators[0]["creator_fans"] == 12000
@pytest.mark.asyncio
async def test_a_creator_with_works_carries_their_count(self, client):
task_id = await _seed()
await _add_stat(task_id, run_id=1, fans=12000)
creators = await _creators(client)
assert len(creators) == 1
assert creators[0]["note_count"] == 1
assert creators[0]["creator_name"] == NICKNAME
assert creators[0]["creator_fans"] == 12000
@pytest.mark.asyncio
async def test_a_creator_with_neither_works_nor_a_snapshot_is_absent(self, client):
"""两个来源都没有 = 我们对他一无所知,不该凭空造一个组出来。"""
await _seed_task()
assert await _creators(client) == []
@pytest.mark.asyncio
async def test_a_work_without_a_snapshot_still_lists_its_creator(self, client):
"""小红书那条路不产生账号快照 —— 那边只能靠作品认出人来。"""
await _seed()
creators = await _creators(client)
assert len(creators) == 1
assert creators[0]["note_count"] == 1
assert creators[0]["creator_fans"] is None
@pytest.mark.asyncio
async def test_the_creator_remark_comes_along(self, client):
"""没有作品的博主也要能起备注 —— 否则「这是谁」在最需要的时候认不出来。"""
task_id = await _seed_task()
await _add_stat(task_id, run_id=1, fans=12000)
await _set_alias(client, "竞品A")
assert (await _creators(client))[0]["creator_alias"] == "竞品A"
@pytest.mark.asyncio
async def test_platforms_stay_apart(self, client):
dy_task = await _seed_task(platform="dy")
await _add_stat(dy_task, run_id=1, fans=999)
assert await _creators(client, "xhs") == []
assert (await _creators(client, "dy"))[0]["creator_fans"] == 999
@pytest.mark.asyncio
async def test_the_same_creator_under_two_tasks_is_two_rows_the_ui_merges(self, client):
"""服务端按 任务×博主 给(快照就是那么存的),合并交给界面 —— 因为备注跨任务
是同一条,合并之后的组才是人眼里的「一个博主」。"""
first = await _seed(task_name="任务甲")
second = await _seed(task_name="任务乙")
await _add_stat(first, run_id=1, fans=12000)
creators = await _creators(client)
assert len(creators) == 2
assert {row["creator_hash"] for row in creators} == {CREATOR_HASH}
assert sum(row["note_count"] for row in creators) == 2
+978
View File
@@ -0,0 +1,978 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_monitor_ingest.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Offline tests for the monitoring ingest/diff layer.
These run without network, browser or login and cover the correctness caveats
that matter most: baseline suppression, count parsing, NULL-vs-zero, the
posted/seen comment split, idempotency, and the silent-cookie-failure signal.
"""
import json
from datetime import date
from pathlib import Path
from typing import Any, Dict, List, Optional
import pytest
import pytest_asyncio
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_async_engine
from sqlalchemy.pool import StaticPool
from tools.time_util import get_current_timestamp
from api.monitor import adapters
from api.monitor.ingest import (
describe_exit_code,
diagnose_failure,
ingest_run,
parse_count,
)
from api.monitor.models import (
EVENT_AUTH_FAILURE,
EVENT_METRIC_DELTA,
EVENT_NEW_COMMENT_POSTED,
EVENT_NEW_COMMENT_SEEN,
EVENT_NEW_NOTE,
EVENT_NO_DATA,
EVENT_RUN_FAILED,
MODE_CREATOR,
MonitorBase,
MonitorComment,
MonitorCreatorStat,
MonitorEvent,
MonitorNote,
MonitorNoteMetric,
MonitorRun,
MonitorTask,
RUN_FAILED,
RUN_PARTIAL,
RUN_SUCCESS,
)
@pytest_asyncio.fixture
async def db():
"""An isolated in-memory monitoring database."""
engine = create_async_engine("sqlite+aiosqlite://", poolclass=StaticPool)
async with engine.begin() as conn:
await conn.run_sync(MonitorBase.metadata.create_all)
factory = async_sessionmaker(engine, class_=AsyncSession, expire_on_commit=False)
async with factory() as db_session:
yield db_session
await engine.dispose()
async def _make_task(db: AsyncSession, **overrides) -> MonitorTask:
defaults = dict(
name="test task",
platform="xhs",
mode=MODE_CREATOR,
enabled=True,
interval_minutes=60,
max_notes_count=20,
enable_comments=True,
max_comments_count=50,
run_timeout_seconds=3600,
created_at=0,
updated_at=0,
)
defaults.update(overrides)
task = MonitorTask(**defaults)
db.add(task)
await db.flush()
return task
async def _make_run(
db: AsyncSession,
task: MonitorTask,
started_at: int,
exit_code: Optional[int] = 0,
) -> MonitorRun:
run = MonitorRun(
task_id=task.id,
trigger="manual",
status=RUN_SUCCESS,
phase=task.mode,
save_data_path="",
queued_at=started_at,
not_before=0,
started_at=started_at,
exit_code=exit_code,
)
db.add(run)
await db.flush()
return run
def _write_run_dir(
root: Path,
notes: List[Dict[str, Any]],
comments: Optional[List[Dict[str, Any]]] = None,
subdir: str = "xhs",
profiles: Optional[List[Dict[str, Any]]] = None,
) -> Path:
"""Write a run's jsonl output in the crawler's own layout.
``subdir`` 是**爬虫**落盘的目录名,不是监控层的平台 id —— 抖音那边这两者不同
(平台 id 是 ``dy``、目录是 ``douyin``),所以必须能分开指定,否则测不出那个差异。
``profiles`` 为 None 时**不写**这个文件(小红书那条路根本不产生它),给列表时写
——包括空列表,那是「问了但没问到」。
"""
jsonl_dir = root / subdir / "jsonl"
jsonl_dir.mkdir(parents=True, exist_ok=True)
contents = jsonl_dir / "creator_contents_2026-01-01.jsonl"
contents.write_text(
"\n".join(json.dumps(n, ensure_ascii=False) for n in notes),
encoding="utf-8",
)
if comments is not None:
comment_file = jsonl_dir / "creator_comments_2026-01-01.jsonl"
comment_file.write_text(
"\n".join(json.dumps(c, ensure_ascii=False) for c in comments),
encoding="utf-8",
)
if profiles is not None:
profile_file = jsonl_dir / "creator_profile_2026-01-01.jsonl"
profile_file.write_text(
"\n".join(json.dumps(p, ensure_ascii=False) for p in profiles),
encoding="utf-8",
)
return root
def _note(note_id: str, liked: Any = "10", **extra) -> Dict[str, Any]:
record = {
"note_id": note_id,
"title": f"title-{note_id}",
"note_url": f"https://www.xiaohongshu.com/explore/{note_id}",
"image_list": "https://img/cover.jpg",
"creator_hash": "hash",
"time": 1700000000000,
"liked_count": liked,
"comment_count": "1",
"collected_count": "1",
"share_count": "1",
}
record.update(extra)
return record
def _comment(comment_id: str, note_id: str, create_time: int, **extra) -> Dict[str, Any]:
record = {
"comment_id": comment_id,
"note_id": note_id,
"content": f"content-{comment_id}",
"nickname": "u***r",
"creator_hash": "hash",
"create_time": create_time,
"like_count": "0",
"sub_comment_count": 0,
"parent_comment_id": "",
}
record.update(extra)
return record
def _dy_note(aweme_id: str, liked: Any = "10", **extra) -> Dict[str, Any]:
"""抖音作品记录 —— 键名照抄 store/douyin/__init__.py 的落盘字段。
重点在于**没有** ``note_id``:抖音叫 ``aweme_id``。这一条差异没映射好,就是
每条记录都被 ingest 悄悄 continue 掉、一条不剩。
"""
record = {
"aweme_id": aweme_id,
"aweme_type": "0",
"title": f"title-{aweme_id}",
"desc": f"title-{aweme_id}",
# 抖音给的是**秒**(实测 1790574515 = 2026-09-28),小红书给毫秒。落库统一
# 换算成毫秒,这个 fixture 必须照真实形态写,否则测不出单位问题。
"create_time": 1790574515,
"creator_hash": "hash",
"nickname": "u***r",
"liked_count": liked,
"collected_count": "1",
"comment_count": "1",
"share_count": "1",
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
"cover_url": "https://img/cover.jpg",
}
record.update(extra)
return record
def _dy_comment(
comment_id: str, aweme_id: str, create_time: int, **extra
) -> Dict[str, Any]:
record = {
"comment_id": comment_id,
"create_time": create_time,
"aweme_id": aweme_id,
"content": f"content-{comment_id}",
"creator_hash": "hash",
"nickname": "u***r",
"sub_comment_count": "0",
"like_count": "0",
# 抖音顶层评论的父 id 是字符串 "0",不是空串。
"parent_comment_id": "0",
}
record.update(extra)
return record
async def _events(db: AsyncSession, event_type: Optional[str] = None) -> List[MonitorEvent]:
stmt = select(MonitorEvent)
if event_type:
stmt = stmt.where(MonitorEvent.type == event_type)
return list((await db.scalars(stmt)).all())
# --------------------------------------------------------------------------
# parse_count
# --------------------------------------------------------------------------
class TestParseCount:
@pytest.mark.parametrize(
"raw,expected",
[
("1234", 1234),
("1.2万", 12000),
("1.2w", 12000),
("3亿", 300000000),
("1,234", 1234),
(42, 42),
],
)
def test_parses_platform_formats(self, raw, expected):
assert parse_count(raw) == expected
@pytest.mark.parametrize("raw", ["", None, "暂无", "-", "abc", True])
def test_unparseable_values_return_none(self, raw):
assert parse_count(raw) is None
# --------------------------------------------------------------------------
# Exit codes
# --------------------------------------------------------------------------
class TestExitCodeStorage:
"""Guards a bug that only showed up when the data moved to MySQL.
Windows reports process failures as unsigned 32-bit NTSTATUS values
(0xC0000142 = 3221225794). That overflows MySQL's signed INT, while SQLite's
dynamic typing accepted it happily -- so the column silently worked until a
real migration hit it with real data.
"""
def test_column_is_bigint_not_int(self):
from sqlalchemy import BigInteger
from api.monitor.models import MonitorRun
column_type = MonitorRun.__table__.c.exit_code.type
assert isinstance(column_type, BigInteger), (
f"exit_code must be BigInteger to hold unsigned 32-bit codes, got {column_type!r}"
)
@pytest.mark.asyncio
async def test_an_ntstatus_value_round_trips(self, db):
task = await _make_task(db)
run = await _make_run(db, task, started_at=1000, exit_code=3221225794)
await db.commit()
stored = await db.scalar(
select(MonitorRun.exit_code).where(MonitorRun.id == run.id)
)
assert stored == 3221225794
class TestDescribeExitCode:
def test_windows_status_code_is_decoded(self):
"""3221225794 is 0xC0000142, which is meaningless without decoding."""
message = describe_exit_code(3221225794)
assert "0xC0000142" in message
assert "DLL_INIT_FAILED" in message
def test_negative_signed_form_is_also_decoded(self):
# Python may hand back the signed form depending on how it was launched.
assert "0xC0000142" in describe_exit_code(-1073741502)
def test_unknown_code_degrades_to_the_raw_number(self):
assert describe_exit_code(1) == "Crawler exited with code 1"
# --------------------------------------------------------------------------
# Notes
# --------------------------------------------------------------------------
class TestNoteIngest:
@pytest.mark.asyncio
async def test_baseline_run_emits_no_new_note_events(self, db, tmp_path):
task = await _make_task(db)
run = await _make_run(db, task, started_at=1000)
_write_run_dir(tmp_path, [_note("n1"), _note("n2")], comments=[])
result = await ingest_run(db, run, task, tmp_path)
assert result.status == RUN_SUCCESS
assert result.is_baseline is True
assert result.new_notes == 2
# Everything is "new" on the first run; emitting that would be pure noise.
assert await _events(db, EVENT_NEW_NOTE) == []
assert len(list((await db.scalars(select(MonitorNote))).all())) == 2
@pytest.mark.asyncio
async def test_an_empty_run_does_not_establish_a_baseline(self, db, tmp_path):
"""A run that fetched nothing observed nothing, so it is not a baseline.
Otherwise the first crawl that actually works reports every work as
newly discovered.
"""
task = await _make_task(db)
empty_run = await _make_run(db, task, started_at=1000)
(tmp_path / "empty").mkdir(parents=True, exist_ok=True)
await ingest_run(db, empty_run, task, tmp_path / "empty")
real_run = await _make_run(db, task, started_at=2000)
result = await ingest_run(
db, real_run, task, _write_run_dir(tmp_path / "ok", [_note("n1")], comments=[])
)
assert result.is_baseline is True
assert await _events(db, EVENT_NEW_NOTE) == []
@pytest.mark.asyncio
async def test_second_run_reports_only_the_added_note(self, db, tmp_path):
task = await _make_task(db)
first_dir = _write_run_dir(tmp_path / "run1", [_note("n1")], comments=[])
run1 = await _make_run(db, task, started_at=1000)
await ingest_run(db, run1, task, first_dir)
second_dir = _write_run_dir(tmp_path / "run2", [_note("n1"), _note("n2")], comments=[])
run2 = await _make_run(db, task, started_at=2000)
result = await ingest_run(db, run2, task, second_dir)
assert result.is_baseline is False
assert result.new_notes == 1
events = await _events(db, EVENT_NEW_NOTE)
assert len(events) == 1
assert events[0].target_id == "n2"
assert events[0].run_id == run2.id
class TestMetricSnapshots:
@pytest.mark.asyncio
async def test_delta_event_emitted_when_like_count_changes(self, db, tmp_path):
task = await _make_task(db)
run1 = await _make_run(db, task, started_at=1000)
await ingest_run(db, run1, task, _write_run_dir(tmp_path / "r1", [_note("n1", "100")], comments=[]))
run2 = await _make_run(db, task, started_at=2000)
await ingest_run(db, run2, task, _write_run_dir(tmp_path / "r2", [_note("n1", "150")], comments=[]))
events = await _events(db, EVENT_METRIC_DELTA)
assert len(events) == 1
payload = json.loads(events[0].payload_json)
assert payload["deltas"]["liked_count"] == {"from": 100, "to": 150, "delta": 50}
@pytest.mark.asyncio
async def test_no_delta_when_nothing_changed(self, db, tmp_path):
task = await _make_task(db)
run1 = await _make_run(db, task, started_at=1000)
await ingest_run(db, run1, task, _write_run_dir(tmp_path / "r1", [_note("n1", "100")], comments=[]))
run2 = await _make_run(db, task, started_at=2000)
await ingest_run(db, run2, task, _write_run_dir(tmp_path / "r2", [_note("n1", "100")], comments=[]))
assert await _events(db, EVENT_METRIC_DELTA) == []
@pytest.mark.asyncio
async def test_unparseable_count_is_null_not_zero(self, db, tmp_path):
task = await _make_task(db)
run = await _make_run(db, task, started_at=1000)
await ingest_run(db, run, task, _write_run_dir(tmp_path, [_note("n1", "暂无")], comments=[]))
metric = await db.scalar(select(MonitorNoteMetric).where(MonitorNoteMetric.note_id == "n1"))
# Zero would forge a large negative delta on the next comparison.
assert metric.liked_count is None
assert metric.raw_liked_count == "暂无"
@pytest.mark.asyncio
async def test_no_delta_when_previous_value_was_unparseable(self, db, tmp_path):
task = await _make_task(db)
run1 = await _make_run(db, task, started_at=1000)
await ingest_run(db, run1, task, _write_run_dir(tmp_path / "r1", [_note("n1", "暂无")], comments=[]))
run2 = await _make_run(db, task, started_at=2000)
await ingest_run(db, run2, task, _write_run_dir(tmp_path / "r2", [_note("n1", "50")], comments=[]))
assert await _events(db, EVENT_METRIC_DELTA) == []
@pytest.mark.asyncio
async def test_metric_snapshot_survives_across_runs(self, db, tmp_path):
"""The crawler's own DB store overwrites metrics; ours must not."""
task = await _make_task(db)
for index, liked in enumerate(["100", "150", "300"]):
run = await _make_run(db, task, started_at=1000 * (index + 1))
await ingest_run(
db, run, task, _write_run_dir(tmp_path / f"r{index}", [_note("n1", liked)], comments=[])
)
snapshots = list(
(
await db.scalars(
select(MonitorNoteMetric)
.where(MonitorNoteMetric.note_id == "n1")
.order_by(MonitorNoteMetric.run_id)
)
).all()
)
assert [s.liked_count for s in snapshots] == [100, 150, 300]
# --------------------------------------------------------------------------
# Comments
# --------------------------------------------------------------------------
class TestCommentIngest:
@pytest.mark.asyncio
async def test_posted_vs_seen_split_by_create_time(self, db, tmp_path):
task = await _make_task(db)
# Baseline establishes the seen-set; no events on the first run.
run1 = await _make_run(db, task, started_at=1000)
await ingest_run(
db, run1, task,
_write_run_dir(tmp_path / "r1", [_note("n1")], comments=[_comment("c1", "n1", create_time=500)]),
)
assert await _events(db, EVENT_NEW_COMMENT_POSTED) == []
# c2 was published after run1 started -> genuinely new.
# c3 is old but only just surfaced in the top-N window -> seen, not posted.
run2 = await _make_run(db, task, started_at=2000)
await ingest_run(
db, run2, task,
_write_run_dir(
tmp_path / "r2",
[_note("n1")],
comments=[
_comment("c1", "n1", create_time=500),
_comment("c2", "n1", create_time=2500),
_comment("c3", "n1", create_time=100),
],
),
)
posted = await _events(db, EVENT_NEW_COMMENT_POSTED)
seen = await _events(db, EVENT_NEW_COMMENT_SEEN)
assert len(posted) == 1
assert json.loads(posted[0].payload_json)["comment_id"] == "c2"
assert len(seen) == 1
assert json.loads(seen[0].payload_json)["comment_id"] == "c3"
@pytest.mark.asyncio
async def test_comments_not_ingested_when_disabled(self, db, tmp_path):
task = await _make_task(db, enable_comments=False)
run = await _make_run(db, task, started_at=1000)
result = await ingest_run(
db, run, task,
_write_run_dir(tmp_path, [_note("n1")], comments=[_comment("c1", "n1", 500)]),
)
assert result.new_comments == 0
# --------------------------------------------------------------------------
# Failure handling
# --------------------------------------------------------------------------
class TestFailureHandling:
@pytest.mark.asyncio
async def test_nonzero_exit_is_a_failure(self, db, tmp_path):
task = await _make_task(db)
run = await _make_run(db, task, started_at=1000, exit_code=1)
_write_run_dir(tmp_path, [_note("n1")], comments=[])
result = await ingest_run(db, run, task, tmp_path)
assert result.status == RUN_FAILED
assert len(await _events(db, EVENT_RUN_FAILED)) == 1
# A crashed run must not touch the seen-set.
assert await db.scalar(select(MonitorNote.id)) is None
@pytest.mark.asyncio
async def test_zero_notes_with_exit_zero_is_a_suspected_auth_failure(self, db, tmp_path):
"""The silent-cookie-failure signature: exit 0 but nothing fetched.
A real bad-cookie run writes no output file at all, which is why the
exit code has to be checked before the files are.
"""
task = await _make_task(db)
run = await _make_run(db, task, started_at=1000, exit_code=0)
tmp_path.mkdir(parents=True, exist_ok=True)
result = await ingest_run(db, run, task, tmp_path)
assert result.status == RUN_PARTIAL
assert len(await _events(db, EVENT_AUTH_FAILURE)) == 1
assert await _events(db, EVENT_RUN_FAILED) == []
@pytest.mark.asyncio
async def test_no_data_is_not_blamed_on_the_cookie_when_a_sibling_succeeded(
self, db, tmp_path
):
"""A task that just worked proves the login is fine; do not cry wolf."""
healthy = await _make_task(db, name="healthy")
healthy_run = await _make_run(db, healthy, started_at=get_current_timestamp())
await ingest_run(
db, healthy_run, healthy,
_write_run_dir(tmp_path / "ok", [_note("n1")], comments=[]),
)
task = await _make_task(db, name="suspect")
run = await _make_run(db, task, started_at=get_current_timestamp())
(tmp_path / "empty").mkdir(parents=True, exist_ok=True)
result = await ingest_run(db, run, task, tmp_path / "empty")
assert result.status == RUN_PARTIAL
assert await _events(db, EVENT_NO_DATA) != []
assert await _events(db, EVENT_AUTH_FAILURE) == []
@pytest.mark.asyncio
async def test_empty_contents_file_is_also_an_auth_failure(self, db, tmp_path):
task = await _make_task(db)
run = await _make_run(db, task, started_at=1000, exit_code=0)
_write_run_dir(tmp_path, [], comments=[])
result = await ingest_run(db, run, task, tmp_path)
assert result.status == RUN_PARTIAL
assert len(await _events(db, EVENT_AUTH_FAILURE)) == 1
# --------------------------------------------------------------------------
# Idempotency
# --------------------------------------------------------------------------
class TestIdempotency:
@pytest.mark.asyncio
async def test_reingesting_the_same_data_adds_nothing(self, db, tmp_path):
task = await _make_task(db)
run_dir = _write_run_dir(
tmp_path, [_note("n1"), _note("n2")], comments=[_comment("c1", "n1", 500)]
)
run1 = await _make_run(db, task, started_at=1000)
await ingest_run(db, run1, task, run_dir)
notes_after_first = len(list((await db.scalars(select(MonitorNote))).all()))
# A retry of the same crawl content must not duplicate rows or events.
run2 = await _make_run(db, task, started_at=2000)
result = await ingest_run(db, run2, task, run_dir)
assert result.new_notes == 0
assert result.new_comments == 0
assert len(list((await db.scalars(select(MonitorNote))).all())) == notes_after_first
class TestDouyinIngest:
"""抖音的产物形状与小红书不同 —— 这里钉住「不会被静默丢掉」。
这一组存在的理由,是这个改动最危险的失败模式:字段名或目录名没对上时,ingest
不报错,只是**一条都不入库**,然后被当成「疑似登录失效」报出去。
"""
async def _ingest(
self,
db,
tmp_path,
notes,
comments=None,
platform="dy",
subdir="douyin",
):
task = await _make_task(db, platform=platform)
run = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, notes, comments, subdir=subdir)
result = await ingest_run(db, run, task, tmp_path)
return task, run, result
@pytest.mark.asyncio
async def test_notes_are_ingested_under_their_douyin_field_names(self, db, tmp_path):
aweme_id = "7525082444551310602"
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note(aweme_id)])
note = await db.scalar(select(MonitorNote))
assert note is not None, "抖音作品被静默丢弃了 —— 多半是 aweme_id 没映射到 note_id"
assert note.note_id == aweme_id
assert note.note_url == f"https://www.douyin.com/video/{aweme_id}"
assert note.cover == "https://img/cover.jpg"
assert note.source_kind == "0"
# 秒 → 毫秒,换算过才对。
assert note.published_at == 1790574515 * 1000
assert result.notes_fetched == 1
@pytest.mark.asyncio
async def test_timestamps_are_normalised_to_milliseconds(self, db, tmp_path):
"""抖音的时间戳是**秒**,小红书是毫秒 —— 差 1000 倍,必须换算。
不换算的话,2026 年的作品会显示成 1970 年。这是实测踩到的:抖音作品的
「发布日期」列显示成 1970-01-22(1790574515 被当成毫秒就是 21 天后)。
"""
aweme_id = "7525082444551310602"
await self._ingest(db, tmp_path, [_dy_note(aweme_id)])
note = await db.scalar(select(MonitorNote))
assert note.published_at == 1790574515 * 1000
# 落库的是毫秒,展示层才不用关心来源;乘完应该在 2026 年,不是 1970。
assert date.fromtimestamp(note.published_at / 1000).year == 2026
@pytest.mark.asyncio
async def test_the_artifact_directory_is_not_the_platform_id(self, db, tmp_path):
"""目录名与平台 id 不一致,是这套适配里最反直觉的一条。
抖音的平台 id 是 ``dy``,而爬虫把产物写在 ``douyin/`` 下。把它钉在这里,
是为了让「顺手改成一致」这件事会在测试里红掉,而不是让 ingest 悄悄读 0 条。
"""
assert adapters.artifact_dir("dy") == "douyin"
@pytest.mark.asyncio
async def test_writing_into_the_platform_id_directory_reads_nothing(self, db, tmp_path):
"""反面:产物落在 ``dy/`` 下时一条都读不到 —— 这正是映射要解决的问题。"""
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note("1")], subdir="dy")
assert result.notes_fetched == 0
@pytest.mark.asyncio
async def test_misplaced_output_is_blamed_on_the_directory_not_the_login(
self, db, tmp_path
):
"""产物其实抓到了,只是目录名不对 —— 不该报成「疑似登录失效」。
这是最难查的一类故障:登录是好的、数据也抓到了,但报出来的现象和登录失效
一模一样,会把人指去查完全错误的方向。
"""
_task, run, _result = await self._ingest(
db, tmp_path, [_dy_note("1")], subdir="dy"
)
assert any("目录" in event.title for event in await _events(db, EVENT_NO_DATA))
assert await _events(db, EVENT_AUTH_FAILURE) == []
assert run.error_message and "dy" in run.error_message
@pytest.mark.asyncio
async def test_comments_are_linked_through_aweme_id(self, db, tmp_path):
aweme_id = "7525082444551310602"
_task, _run, result = await self._ingest(
db,
tmp_path,
[_dy_note(aweme_id)],
comments=[_dy_comment("c1", aweme_id, 500)],
)
comment = await db.scalar(select(MonitorComment))
assert comment is not None, "抖音评论被静默丢弃了 —— 多半是 aweme_id 没映射"
assert comment.note_id == aweme_id
assert result.comments_fetched == 1
@pytest.mark.asyncio
async def test_a_top_level_parent_of_zero_becomes_empty(self, db, tmp_path):
"""抖音顶层评论的父 id 是 "0";原样存进去,前端会多出一堆悬空的父节点。"""
aweme_id = "7525082444551310602"
await self._ingest(
db,
tmp_path,
[_dy_note(aweme_id)],
comments=[
_dy_comment("c1", aweme_id, 500),
_dy_comment("c2", aweme_id, 600, parent_comment_id="c1"),
],
)
by_id = {c.comment_id: c for c in (await db.scalars(select(MonitorComment))).all()}
assert by_id["c1"].parent_comment_id == ""
assert by_id["c2"].parent_comment_id == "c1"
@pytest.mark.asyncio
async def test_the_four_metrics_need_no_mapping(self, db, tmp_path):
"""四个指标键两边同名 —— 抖音作品照样进 monitor_note_metric,差分照常。"""
aweme_id = "7525082444551310602"
task = await _make_task(db, platform="dy")
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="100")], subdir="douyin")
run1 = await _make_run(db, task, started_at=1)
await ingest_run(db, run1, task, tmp_path)
metric = await db.scalar(select(MonitorNoteMetric))
assert metric is not None and metric.liked_count == 100
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="150")], subdir="douyin")
run2 = await _make_run(db, task, started_at=2)
await ingest_run(db, run2, task, tmp_path)
assert len(await _events(db, EVENT_METRIC_DELTA)) == 1
class TestDuplicateRecordsInOneRun:
"""一轮产物里重复出现的作品只该算一次。
指标快照的唯一键是 ``(task_id, note_id, run_id)``:同一条作品在一轮里进来两次,
第二次插入会撞键并让整个 run 崩掉 —— 产物里重复并不罕见(多个目标指向同一个人、
或退化路径重复刷新)。
"""
@pytest.mark.asyncio
async def test_a_duplicated_note_is_processed_once(self, db, tmp_path):
task = await _make_task(db)
_write_run_dir(tmp_path, [_note("n1"), _note("n1")])
run = await _make_run(db, task, started_at=1)
result = await ingest_run(db, run, task, tmp_path)
assert result.new_notes == 1
assert len(list((await db.scalars(select(MonitorNote))).all())) == 1
# 崩就崩在这一句上:一条作品只能有一份本轮快照。
assert len(list((await db.scalars(select(MonitorNoteMetric))).all())) == 1
class TestNicknameRefresh:
"""已入库的评论,昵称要跟着重新采集的值走。
评论是去重后直接跳过的,若不刷新,脱敏开关一改(或评论者改了昵称),老数据就永远
停在旧值上 —— 而重采是唯一能拿到新值的途径。
"""
@pytest.mark.asyncio
async def test_an_existing_comment_gets_its_nickname_refreshed(self, db, tmp_path):
task = await _make_task(db)
_write_run_dir(tmp_path, [_note("n1")], comments=[_comment("c1", "n1", 500)])
run1 = await _make_run(db, task, started_at=1)
await ingest_run(db, run1, task, tmp_path)
assert (await db.scalar(select(MonitorComment))).nickname == "u***r"
_write_run_dir(
tmp_path,
[_note("n1")],
comments=[_comment("c1", "n1", 500, nickname="未脱敏的新昵称")],
)
run2 = await _make_run(db, task, started_at=2)
await ingest_run(db, run2, task, tmp_path)
comment = await db.scalar(select(MonitorComment))
assert comment.nickname == "未脱敏的新昵称"
# 去重的语义没变:同一条评论不该被插成两行。
assert (
len(list((await db.scalars(select(MonitorComment))).all())) == 1
)
class TestFailureDiagnosis:
"""失败原因要能被人看懂。
只写「退出码 1」等于什么都没说 —— 真正的报错埋在子进程的 stderr 里,而运行历史
里那一格显示的正是 run.error_message。
"""
TAIL = [
"2026-10-10 15:18:34 MediaCrawler INFO (core.py:385) - [DouYinCrawler] CDP浏览器信息",
"Traceback (most recent call last):",
' File "/app/main.py", line 114, in main',
" await crawler.start()",
"media_platform.douyin.exception.DataFetchError: account blocked, ",
]
def test_the_exception_line_is_picked_out_of_the_tail(self):
assert (
diagnose_failure(self.TAIL)
== "media_platform.douyin.exception.DataFetchError: account blocked,"
)
def test_the_managers_own_lines_are_not_mistaken_for_the_cause(self):
"""管理器自己补的那两句不是爬虫的报错,别被当成失败原因。"""
assert diagnose_failure(["Crawler exited with code: 1"]) is None
assert diagnose_failure(["Crawler completed successfully"]) is None
def test_nothing_to_say_is_not_an_error(self):
assert diagnose_failure(None) is None
assert diagnose_failure([]) is None
def test_it_falls_back_to_the_last_line(self):
assert (
diagnose_failure(["started fine", "then something odd"])
== "then something odd"
)
@pytest.mark.asyncio
async def test_a_failed_run_records_both_the_code_and_the_cause(self, db, tmp_path):
task = await _make_task(db)
run = await _make_run(db, task, started_at=1, exit_code=1)
result = await ingest_run(db, run, task, tmp_path, output_tail=self.TAIL)
assert result.status == RUN_FAILED
# 退出码和真因都要在,缺一个都还得去翻日志。
assert "code 1" in run.error_message
assert "account blocked" in run.error_message
events = await _events(db, EVENT_RUN_FAILED)
assert "account blocked" in events[0].title
# --------------------------------------------------------------------------
# 博主账号级指标(粉丝 / 总获赞 / 作品数)
# --------------------------------------------------------------------------
def _profile(creator_hash: str = "hash", **extra) -> Dict[str, Any]:
"""``creator_profile_*.jsonl`` 里的一行 —— 形状由 douyin_api.author_profile 决定。"""
record: Dict[str, Any] = {
"creator_hash": creator_hash,
"nickname": "博主",
"unique_id": "abc",
"fans": 12000,
"total_favorited": 83000,
"works": 42,
"following": 7,
}
record.update(extra)
return record
class TestCreatorStatSnapshots:
"""**账号级**指标和作品级指标是两回事:后者说"这条视频涨了多少赞",前者说
"这个人整个账号的粉丝在涨还是在掉"。作品列表给不了后者,所以单独存一张表。
"""
async def _ingest(self, db, tmp_path, notes, profiles, platform="dy", subdir="douyin"):
task = await _make_task(db, platform=platform)
run = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, notes, comments=[], subdir=subdir, profiles=profiles)
result = await ingest_run(db, run, task, tmp_path)
return task, run, result
@pytest.mark.asyncio
async def test_a_profile_becomes_a_snapshot(self, db, tmp_path):
task, run, _result = await self._ingest(db, tmp_path, [_dy_note("1")], [_profile()])
stat = await db.scalar(select(MonitorCreatorStat))
assert stat is not None
assert (stat.task_id, stat.run_id) == (task.id, run.id)
assert stat.creator_hash == "hash"
assert stat.nickname == "博主"
assert (stat.fans, stat.total_favorited, stat.works_count) == (12000, 83000, 42)
@pytest.mark.asyncio
async def test_a_missing_count_stays_null_not_zero(self, db, tmp_path):
"""0 是真实值(掉到零),null 是不知道。混起来趋势图就是在撒谎。"""
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(fans=None)])
stat = await db.scalar(select(MonitorCreatorStat))
assert stat.fans is None
# 同一个博主其它字段照常。
assert stat.works_count == 42
@pytest.mark.asyncio
async def test_abbreviated_counts_are_parsed(self, db, tmp_path):
"""走的是和作品指标同一个 parse_count —— 平台给你「1.2万」也得认。"""
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(fans="1.2万")])
assert (await db.scalar(select(MonitorCreatorStat))).fans == 12000
@pytest.mark.asyncio
async def test_the_same_creator_twice_in_one_run_yields_one_snapshot(self, db, tmp_path):
"""一个任务可以配多个目标,退化路径下它们可能落在同一个博主身上。
唯一键是 (task_id, creator_hash, run_id) —— 重复插入会撞键,把整轮炸掉。
(和作品重复那次是同一类事故。)
"""
await self._ingest(
db, tmp_path, [_dy_note("1")], [_profile(), _profile(nickname="另一条")]
)
stats = list((await db.scalars(select(MonitorCreatorStat))).all())
assert len(stats) == 1
@pytest.mark.asyncio
async def test_a_profile_without_a_hash_is_skipped(self, db, tmp_path):
"""哈希都算不出来,这条快照谁也查不到,落下去只是垃圾。"""
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(creator_hash="")])
assert (await db.scalars(select(MonitorCreatorStat))).all() == []
@pytest.mark.asyncio
async def test_no_profile_file_is_fine(self, db, tmp_path):
"""小红书那条路(爬虫进程)根本不产生这个文件 —— 不能因此报错。"""
task = await _make_task(db) # xhs
run = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, [_note("n1")], comments=[]) # 没有 profiles 参数
result = await ingest_run(db, run, task, tmp_path)
assert result.status == RUN_SUCCESS
assert (await db.scalars(select(MonitorCreatorStat))).all() == []
@pytest.mark.asyncio
async def test_the_stats_are_kept_even_when_no_works_were_fetched(self, db, tmp_path):
"""**这条是这里最值得留的一个。**
作品列表被风控挡住时,这一轮一条作品都拿不到、run 会被判成失败。但博主的粉丝数
并不因为这件事就不存在 —— 「粉丝还在涨,但新作品没在发现」恰恰是最该看见的时刻。
快照要是挂在「作品采到了」后面,就正好在最需要它的那一轮丢掉。
"""
_task, run, result = await self._ingest(db, tmp_path, [], [_profile()])
assert result.status == RUN_PARTIAL # 一条作品都没有,这轮确实不算成功
assert run.status == RUN_PARTIAL
stat = await db.scalar(select(MonitorCreatorStat))
assert stat is not None and stat.fans == 12000
@pytest.mark.asyncio
async def test_each_run_adds_its_own_snapshot(self, db, tmp_path):
"""趋势靠的就是这个:一条一轮,不要覆盖。"""
task = await _make_task(db, platform="dy")
first = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, [_dy_note("1")], comments=[], subdir="douyin",
profiles=[_profile(fans=12000)])
await ingest_run(db, first, task, tmp_path)
second = await _make_run(db, task, started_at=2000)
_write_run_dir(tmp_path, [_dy_note("1")], comments=[], subdir="douyin",
profiles=[_profile(fans=12300)])
await ingest_run(db, second, task, tmp_path)
stats = list(
(
await db.scalars(
select(MonitorCreatorStat).order_by(MonitorCreatorStat.run_id)
)
).all()
)
assert [s.fans for s in stats] == [12000, 12300]
+146
View File
@@ -0,0 +1,146 @@
# -*- coding: utf-8 -*-
"""作品备注 —— 一个博主底下,哪几条是真正要盯的。
和博主备注(test_monitor_creators.py)是一对,但回答的不是同一个问题:博主备注回答
「这个账号是谁」,作品备注回答「这条作品我要盯着」。一个博主底下常常只有一两件值得
盯的作品,所以不能合成一条。
"""
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import db as monitor_db
from api.monitor.models import MODE_CREATOR, MonitorNote, MonitorTask
NOTE_ID = "note-a"
TITLE = "中秋哪儿都堵"
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
async def _seed(platform: str = "xhs", task_name: str = "任务", note_id: str = NOTE_ID) -> int:
async with monitor_db.get_session() as session:
task = MonitorTask(
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
session.add(task)
await session.flush()
session.add(
MonitorNote(
task_id=task.id, note_id=note_id, title=TITLE,
note_url="", cover="", creator_hash="hash-a",
creator_name="张三", source_kind="", published_at=None,
first_seen_run_id=1, first_seen_at=0,
last_seen_run_id=1, last_seen_at=0,
)
)
return task.id
async def _notes(client, platform: str = "xhs"):
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["notes"]
async def _set_alias(client, alias: str, platform: str = "xhs", note_id: str = NOTE_ID):
return await client.put(
f"/api/monitor/notes/{note_id}",
params={"platform": platform},
json={"alias": alias},
)
class TestNoteAlias:
@pytest.mark.asyncio
async def test_notes_start_without_a_remark(self, client):
await _seed()
assert (await _notes(client))[0]["note_alias"] == ""
@pytest.mark.asyncio
async def test_a_remark_comes_back_with_the_notes(self, client):
await _seed()
response = await _set_alias(client, "重点")
assert response.status_code == 200
note = (await _notes(client))[0]
assert note["note_alias"] == "重点"
# 备注是**叠加**在标题之上的,不是替换 —— 标题仍然是这条作品本身。
assert note["title"] == TITLE
@pytest.mark.asyncio
async def test_a_remark_is_shared_across_tasks(self, client):
"""同一件作品被两个任务都监控时,备注只该填一次。"""
await _seed(task_name="任务甲")
await _seed(task_name="任务乙")
await _set_alias(client, "重点")
for note in await _notes(client):
assert note["note_alias"] == "重点"
@pytest.mark.asyncio
async def test_a_remark_does_not_leak_to_another_platform(self, client):
"""作品的 id 是平台各自的编号体系 —— 抖音的 123 和小红书的 123 是两条作品。"""
await _seed(platform="xhs")
await _seed(platform="dy")
await _set_alias(client, "小红书那边的", platform="xhs")
assert (await _notes(client, "xhs"))[0]["note_alias"] == "小红书那边的"
assert (await _notes(client, "dy"))[0]["note_alias"] == ""
@pytest.mark.asyncio
async def test_a_remark_does_not_leak_to_another_work(self, client):
"""钉住这里的**作用域**:键是 note_id。写错成按任务存的话,给一条起了备注,
同一个博主底下的其它作品会跟着一起变 —— 那这个功能就没用了。"""
await _seed(note_id="note-a")
await _seed(task_name="另一个任务", note_id="note-b")
await _set_alias(client, "重点", note_id="note-a")
by_id = {note["note_id"]: note["note_alias"] for note in await _notes(client)}
assert by_id == {"note-a": "重点", "note-b": ""}
@pytest.mark.asyncio
async def test_an_empty_remark_clears_it(self, client):
await _seed()
await _set_alias(client, "重点")
await _set_alias(client, "")
assert (await _notes(client))[0]["note_alias"] == ""
@pytest.mark.asyncio
async def test_a_remark_is_trimmed(self, client):
await _seed()
await _set_alias(client, " 重点 ")
assert (await _notes(client))[0]["note_alias"] == "重点"
@pytest.mark.asyncio
async def test_the_two_kinds_of_remark_stay_apart(self, client):
"""博主备注和作品备注是两张表、两个键 —— 一个不该把另一个盖掉。"""
await _seed()
await _set_alias(client, "重点")
await client.put(
"/api/monitor/creators/hash-a", params={"platform": "xhs"}, json={"alias": "竞品A"}
)
note = (await _notes(client))[0]
assert note["note_alias"] == "重点"
assert note["creator_alias"] == "竞品A"
+331
View File
@@ -0,0 +1,331 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_monitor_notify.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Tests for the WeCom notification layer.
The webhook is stubbed, so nothing here touches the network.
"""
import json
import pytest
import pytest_asyncio
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_async_engine
from sqlalchemy.pool import StaticPool
from api.monitor import notify
from api.monitor.models import (
EVENT_AUTH_FAILURE,
EVENT_METRIC_DELTA,
EVENT_NEW_NOTE,
EVENT_NEW_COMMENT_POSTED,
MODE_CREATOR,
SETTING_WECOM_WEBHOOK,
MonitorBase,
MonitorEvent,
MonitorRun,
MonitorTask,
RUN_SUCCESS,
)
from api.monitor.settings import set_setting
WEBHOOK = "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=abc123"
@pytest_asyncio.fixture
async def db():
engine = create_async_engine("sqlite+aiosqlite://", poolclass=StaticPool)
async with engine.begin() as conn:
await conn.run_sync(MonitorBase.metadata.create_all)
factory = async_sessionmaker(engine, class_=AsyncSession, expire_on_commit=False)
async with factory() as session:
yield session
await engine.dispose()
async def _seed(db: AsyncSession, notify_enabled: bool = True):
task = MonitorTask(
name="竞品监控", platform="xhs", mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=True,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=notify_enabled, created_at=0, updated_at=0,
)
db.add(task)
await db.flush()
run = MonitorRun(
task_id=task.id, trigger="scheduled", status=RUN_SUCCESS, phase=MODE_CREATOR,
save_data_path="", queued_at=0, not_before=0, max_comments_count=50,
)
db.add(run)
await db.flush()
return task, run
def _add_event(db, task, run, event_type, title, payload=None, severity="info"):
db.add(
MonitorEvent(
task_id=task.id, run_id=run.id, type=event_type, severity=severity,
target_kind="note", target_id="note-1", title=title,
payload_json=json.dumps(payload or {}, ensure_ascii=False), created_at=0,
)
)
# --------------------------------------------------------------------------
# Message building
# --------------------------------------------------------------------------
class TestBuildRunMessage:
@pytest.mark.asyncio
async def test_no_notifiable_events_means_no_message(self, db):
task, run = await _seed(db)
# Metric deltas are not something anyone wants pushed.
_add_event(db, task, run, EVENT_METRIC_DELTA, "点赞 10→20")
_add_event(db, task, run, EVENT_NEW_COMMENT_POSTED, "新评论")
await db.flush()
assert await notify.build_run_message(db, task, run) is None
@pytest.mark.asyncio
async def test_new_notes_are_listed_with_links(self, db):
task, run = await _seed(db)
_add_event(
db, task, run, EVENT_NEW_NOTE, "新作品:标题A",
payload={"note_id": "abc123", "title": "标题A"},
)
await db.flush()
message = await notify.build_run_message(db, task, run)
assert "竞品监控" in message
assert "新增作品 **1** 篇" in message
assert "标题A" in message
assert "https://www.xiaohongshu.com/explore/abc123" in message
@pytest.mark.asyncio
async def test_douyin_notes_link_to_douyin(self, db):
"""链接形状按平台走 —— 群里点进去该是能看的作品,不是 404。"""
task, run = await _seed(db)
task.platform = "dy"
_add_event(
db, task, run, EVENT_NEW_NOTE, "新作品:标题A",
payload={"note_id": "7525082444551310602", "title": "标题A"},
)
await db.flush()
message = await notify.build_run_message(db, task, run)
assert "https://www.douyin.com/video/7525082444551310602" in message
assert "xiaohongshu.com" not in message
@pytest.mark.asyncio
async def test_long_note_lists_are_truncated(self, db):
"""A first run can find dozens; a wall of text is worse than a count."""
task, run = await _seed(db)
for index in range(14):
_add_event(
db, task, run, EVENT_NEW_NOTE, f"新作品:{index}",
payload={"note_id": f"n{index}", "title": f"标题{index}"},
)
await db.flush()
message = await notify.build_run_message(db, task, run)
assert "新增作品 **14** 篇" in message
assert "标题0" in message
assert "标题13" not in message
assert "等共 14 篇" in message
@pytest.mark.asyncio
async def test_failure_is_reported_as_a_warning(self, db):
task, run = await _seed(db)
_add_event(
db, task, run, EVENT_AUTH_FAILURE,
"疑似登录态失效:本次未抓到任何作品", severity="error",
)
await db.flush()
message = await notify.build_run_message(db, task, run)
assert "异常" in message
assert "登录态失效" in message
assert notify._COLOR_WARNING in message
@pytest.mark.asyncio
async def test_baseline_runs_say_so(self, db):
task, run = await _seed(db)
run.is_baseline = True
_add_event(db, task, run, EVENT_NEW_NOTE, "新作品", payload={"note_id": "x", "title": "t"})
await db.flush()
message = await notify.build_run_message(db, task, run)
assert "基线" in message
# --------------------------------------------------------------------------
# notify_run gating
# --------------------------------------------------------------------------
class TestNotifyRunGating:
@pytest.mark.asyncio
async def test_disabled_task_is_skipped(self, db, monkeypatch):
task, run = await _seed(db, notify_enabled=False)
_add_event(db, task, run, EVENT_NEW_NOTE, "新作品", payload={"note_id": "x", "title": "t"})
await set_setting(db, SETTING_WECOM_WEBHOOK, WEBHOOK)
await db.flush()
called = []
monkeypatch.setattr(notify, "send_wecom", lambda *a, **k: called.append(a) or _ok())
assert await notify.notify_run(db, task, run) is None
assert called == []
@pytest.mark.asyncio
async def test_missing_webhook_is_skipped(self, db, monkeypatch):
task, run = await _seed(db, notify_enabled=True)
_add_event(db, task, run, EVENT_NEW_NOTE, "新作品", payload={"note_id": "x", "title": "t"})
await db.flush()
called = []
monkeypatch.setattr(notify, "send_wecom", lambda *a, **k: called.append(a) or _ok())
assert await notify.notify_run(db, task, run) is None
assert called == []
@pytest.mark.asyncio
async def test_successful_push_records_the_timestamp(self, db, monkeypatch):
task, run = await _seed(db, notify_enabled=True)
_add_event(db, task, run, EVENT_NEW_NOTE, "新作品", payload={"note_id": "x", "title": "t"})
await set_setting(db, SETTING_WECOM_WEBHOOK, WEBHOOK)
await db.flush()
monkeypatch.setattr(notify, "send_wecom", lambda *a, **k: _ok())
message = await notify.notify_run(db, task, run)
assert message is not None
# Lets the UI answer "why did I not get a push for this run?".
assert task.last_notified_at is not None
@pytest.mark.asyncio
async def test_push_failure_never_raises(self, db, monkeypatch):
"""A broken webhook must not take down the crawl that just succeeded."""
task, run = await _seed(db, notify_enabled=True)
_add_event(db, task, run, EVENT_NEW_NOTE, "新作品", payload={"note_id": "x", "title": "t"})
await set_setting(db, SETTING_WECOM_WEBHOOK, WEBHOOK)
await db.flush()
async def _boom(*args, **kwargs):
raise RuntimeError("network exploded")
monkeypatch.setattr(notify, "send_wecom", _boom)
assert await notify.notify_run(db, task, run) is None
async def _ok():
return True, "发送成功"
# --------------------------------------------------------------------------
# send_wecom
# --------------------------------------------------------------------------
class _FakeResponse:
def __init__(self, payload):
self._payload = payload
def raise_for_status(self):
return None
def json(self):
return self._payload
class _FakeClient:
"""Captures the request and replays a canned WeCom reply."""
last_payload = None
def __init__(self, reply=None, error=None):
self._reply = reply if reply is not None else {"errcode": 0, "errmsg": "ok"}
self._error = error
def __call__(self, *args, **kwargs):
return self
async def __aenter__(self):
return self
async def __aexit__(self, *exc):
return False
async def post(self, url, json=None):
if self._error:
raise self._error
type(self).last_payload = json
return _FakeResponse(self._reply)
class TestSendWecom:
@pytest.mark.asyncio
async def test_missing_url_is_reported(self):
ok, detail = await notify.send_wecom("", "hi")
assert ok is False
assert "未配置" in detail
@pytest.mark.asyncio
async def test_success(self, monkeypatch):
monkeypatch.setattr(notify.httpx, "AsyncClient", _FakeClient())
ok, detail = await notify.send_wecom(WEBHOOK, "**标题**\n> 内容")
assert ok is True
assert detail == "发送成功"
# WeCom expects a markdown message envelope.
assert _FakeClient.last_payload["msgtype"] == "markdown"
assert _FakeClient.last_payload["markdown"]["content"] == "**标题**\n> 内容"
@pytest.mark.asyncio
async def test_nonzero_errcode_is_a_failure(self, monkeypatch):
"""WeCom answers HTTP 200 even when it rejects the message."""
monkeypatch.setattr(
notify.httpx, "AsyncClient",
_FakeClient(reply={"errcode": 93000, "errmsg": "invalid webhook url"}),
)
ok, detail = await notify.send_wecom(WEBHOOK, "hi")
assert ok is False
assert "93000" in detail
@pytest.mark.asyncio
async def test_network_error_is_returned_not_raised(self, monkeypatch):
import httpx
monkeypatch.setattr(
notify.httpx, "AsyncClient",
_FakeClient(error=httpx.ConnectError("boom")),
)
ok, detail = await notify.send_wecom(WEBHOOK, "hi")
assert ok is False
assert "请求失败" in detail
+323
View File
@@ -0,0 +1,323 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_monitor_report.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Tests for the cross-task report aggregation.
The interaction delta is the part that is easy to get subtly wrong, so it is
covered directly against the pure aggregation function.
"""
from datetime import date, datetime
import pytest
import pytest_asyncio
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_async_engine
from sqlalchemy.pool import StaticPool
from api.monitor.models import (
MODE_CREATOR,
MonitorBase,
MonitorComment,
MonitorNote,
MonitorNoteMetric,
MonitorTask,
)
from api.monitor.report import build_report, compute_daily_rows, day_bounds, iter_days
def _ms(year: int, month: int, day: int, hour: int = 12) -> int:
return int(datetime(year, month, day, hour).timestamp() * 1000)
def _metrics(liked=0, comment=0, collected=0, share=0):
"""All four metrics default to parsed values; pass None to simulate a
platform value we could not parse."""
return {
"liked_count": liked,
"comment_count": comment,
"collected_count": collected,
"share_count": share,
}
class TestDayHelpers:
def test_day_bounds_cover_the_whole_local_day(self):
start, end = day_bounds(date(2026, 1, 10))
assert start < _ms(2026, 1, 10, 0) or start == _ms(2026, 1, 10, 0)
assert end > _ms(2026, 1, 10, 23)
def test_iter_days_is_inclusive(self):
days = iter_days(date(2026, 1, 10), date(2026, 1, 12))
assert days == [date(2026, 1, 10), date(2026, 1, 11), date(2026, 1, 12)]
class TestInteractionDelta:
def test_note_first_seen_counts_all_of_its_value(self):
"""A brand-new note has no earlier baseline, so it starts from zero."""
day = date(2026, 1, 10)
series = {"n1": [(_ms(2026, 1, 10, 10), _metrics(liked=100, comment=5))]}
rows = compute_daily_rows(series, {}, {}, [day])
assert rows[0]["liked_count_delta"] == 100
assert rows[0]["comment_count_delta"] == 5
def test_growth_is_split_across_days(self):
series = {
"n1": [
(_ms(2026, 1, 10, 10), _metrics(liked=100)),
(_ms(2026, 1, 11, 10), _metrics(liked=300)),
]
}
rows = compute_daily_rows(series, {}, {}, [date(2026, 1, 10), date(2026, 1, 11)])
# Day 1: 0 -> 100. Day 2: 100 -> 300.
assert [row["liked_count_delta"] for row in rows] == [100, 200]
def test_day_without_a_snapshot_reports_no_growth(self):
series = {
"n1": [
(_ms(2026, 1, 10, 10), _metrics(liked=100)),
(_ms(2026, 1, 12, 10), _metrics(liked=400)),
]
}
days = [date(2026, 1, 10), date(2026, 1, 11), date(2026, 1, 12)]
rows = compute_daily_rows(series, {}, {}, days)
# The note was not crawled on the 11th, so nothing is claimed for it.
assert [row["liked_count_delta"] for row in rows] == [100, 0, 300]
def test_deltas_aggregate_across_notes(self):
series = {
"n1": [
(_ms(2026, 1, 10, 10), _metrics(liked=100)),
(_ms(2026, 1, 11, 10), _metrics(liked=150)),
],
"n2": [
(_ms(2026, 1, 10, 10), _metrics(liked=10)),
(_ms(2026, 1, 11, 10), _metrics(liked=40)),
],
}
rows = compute_daily_rows(series, {}, {}, [date(2026, 1, 10), date(2026, 1, 11)])
assert [row["liked_count_delta"] for row in rows] == [110, 80]
def test_unparseable_metric_names_the_offending_field(self):
"""A NULL count makes the delta unknown; it must not be reported as 0."""
series = {
"n1": [
(_ms(2026, 1, 10, 10), _metrics(liked=100, comment=None)),
(_ms(2026, 1, 11, 10), _metrics(liked=200, comment=None)),
]
}
rows = compute_daily_rows(series, {}, {}, [date(2026, 1, 11)])
# Naming the field is actionable; a bare boolean is not.
assert rows[0]["partial_metrics"] == ["comment_count"]
# The parseable metric is still summed correctly.
assert rows[0]["liked_count_delta"] == 100
def test_unknown_value_only_taints_the_days_it_touches(self):
series = {
"n1": [
(_ms(2026, 1, 10, 10), _metrics(liked=None)),
(_ms(2026, 1, 11, 10), _metrics(liked=50)),
(_ms(2026, 1, 12, 10), _metrics(liked=90)),
]
}
days = [date(2026, 1, 10), date(2026, 1, 11), date(2026, 1, 12)]
rows = compute_daily_rows(series, {}, {}, days)
# Day 12 compares two known values, so it is clean.
assert [row["partial_metrics"] for row in rows] == [
["liked_count"],
["liked_count"],
[],
]
assert rows[2]["liked_count_delta"] == 40
def test_new_content_counts_come_from_the_day_maps(self):
rows = compute_daily_rows(
{},
{date(2026, 1, 10): 3},
{date(2026, 1, 10): 7},
[date(2026, 1, 10), date(2026, 1, 11)],
)
assert rows[0]["new_notes"] == 3
assert rows[0]["new_comments"] == 7
assert rows[1]["new_notes"] == 0
# --------------------------------------------------------------------------
# DB-backed report + task filtering
# --------------------------------------------------------------------------
@pytest_asyncio.fixture
async def db():
engine = create_async_engine("sqlite+aiosqlite://", poolclass=StaticPool)
async with engine.begin() as conn:
await conn.run_sync(MonitorBase.metadata.create_all)
factory = async_sessionmaker(engine, class_=AsyncSession, expire_on_commit=False)
async with factory() as session:
yield session
await engine.dispose()
async def _seed_task(db: AsyncSession, name: str, platform: str = "xhs") -> MonitorTask:
task = MonitorTask(
name=name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=True,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
db.add(task)
await db.flush()
return task
async def _seed_note_with_metrics(
db: AsyncSession, task: MonitorTask, note_id: str, samples
) -> None:
db.add(
MonitorNote(
task_id=task.id, note_id=note_id, title=note_id, note_url="",
cover="", creator_hash="", source_kind="", published_at=None,
first_seen_run_id=1, first_seen_at=samples[0][0],
last_seen_run_id=len(samples), last_seen_at=samples[-1][0],
)
)
for run_id, (ts, liked) in enumerate(samples, start=1):
db.add(
MonitorNoteMetric(
task_id=task.id, note_id=note_id, run_id=run_id, captured_at=ts,
liked_count=liked, comment_count=0, collected_count=0, share_count=0,
raw_liked_count=str(liked), raw_comment_count="0",
raw_collected_count="0", raw_share_count="0",
)
)
class TestBuildReport:
@pytest.mark.asyncio
async def test_totals_and_rows(self, db):
task = await _seed_task(db, "t1")
await _seed_note_with_metrics(
db, task, "n1",
[(_ms(2026, 1, 10, 10), 100), (_ms(2026, 1, 11, 10), 250)],
)
await db.commit()
result = await build_report(db, [task.id], date(2026, 1, 10), date(2026, 1, 11))
assert result["totals"]["liked_count_delta"] == 250
assert len(result["rows"]) == 2
assert result["note_count"] == 1
@pytest.mark.asyncio
async def test_task_selection_isolates_the_report(self, db):
"""The whole point: a report for a chosen subset must exclude the rest."""
kept = await _seed_task(db, "kept")
other = await _seed_task(db, "other")
await _seed_note_with_metrics(db, kept, "n1", [(_ms(2026, 1, 10, 10), 100)])
await _seed_note_with_metrics(db, other, "n2", [(_ms(2026, 1, 10, 10), 999)])
await db.commit()
only_kept = await build_report(db, [kept.id], date(2026, 1, 10), date(2026, 1, 10))
assert only_kept["totals"]["liked_count_delta"] == 100
assert only_kept["note_count"] == 1
both = await build_report(db, [kept.id, other.id], date(2026, 1, 10), date(2026, 1, 10))
assert both["totals"]["liked_count_delta"] == 1099
@pytest.mark.asyncio
async def test_no_task_filter_covers_everything(self, db):
first = await _seed_task(db, "a")
second = await _seed_task(db, "b")
await _seed_note_with_metrics(db, first, "n1", [(_ms(2026, 1, 10, 10), 10)])
await _seed_note_with_metrics(db, second, "n2", [(_ms(2026, 1, 10, 10), 20)])
await db.commit()
result = await build_report(db, None, date(2026, 1, 10), date(2026, 1, 10))
assert result["totals"]["liked_count_delta"] == 30
assert result["task_ids"] is None
@pytest.mark.asyncio
async def test_an_empty_selection_is_not_the_same_as_no_filter(self, db):
"""空列表 ≠ 不限制。
``_resolve_scope`` 在「这个平台一个任务都没有」时返回**空列表**。如果按真值
处理(``if task_ids``),报表就会退化成「不限制平台」,把**所有**任务的数据
聚合进来 —— 现象就是切到抖音,报表里却全是小红书的数据。
"""
other = await _seed_task(db, "xhs task", platform="xhs")
await _seed_note_with_metrics(db, other, "n1", [(_ms(2026, 1, 10, 10), 999)])
await db.commit()
result = await build_report(db, [], date(2026, 1, 10), date(2026, 1, 10))
assert result["totals"]["liked_count_delta"] == 0
assert result["note_count"] == 0
# 空列表要原样透出去;None 在 API 里的意思是「全部任务」,两者不能混。
assert result["task_ids"] == []
@pytest.mark.asyncio
async def test_baseline_from_before_the_range_is_used(self, db):
"""Growth is measured against the last value before the window opens."""
task = await _seed_task(db, "t")
await _seed_note_with_metrics(
db, task, "n1",
[(_ms(2026, 1, 5, 10), 1000), (_ms(2026, 1, 10, 10), 1050)],
)
await db.commit()
# Report only for the 10th: the delta must be 50, not 1050.
result = await build_report(db, [task.id], date(2026, 1, 10), date(2026, 1, 10))
assert result["totals"]["liked_count_delta"] == 50
@pytest.mark.asyncio
async def test_empty_range_returns_zeroed_rows(self, db):
result = await build_report(db, None, date(2026, 2, 1), date(2026, 2, 3))
assert len(result["rows"]) == 3
assert result["totals"]["liked_count_delta"] == 0
assert result["totals"]["new_notes"] == 0
@pytest.mark.asyncio
async def test_new_comments_are_counted_by_first_seen_day(self, db):
task = await _seed_task(db, "t")
db.add(
MonitorComment(
task_id=task.id, note_id="n1", comment_id="c1", content="x",
nickname="u", creator_hash="h", create_time=_ms(2026, 1, 9),
like_count=0, sub_comment_count=0, parent_comment_id="",
first_seen_run_id=1, first_seen_at=_ms(2026, 1, 10, 10),
)
)
await db.commit()
result = await build_report(db, [task.id], date(2026, 1, 10), date(2026, 1, 10))
assert result["totals"]["new_comments"] == 1
+110
View File
@@ -0,0 +1,110 @@
# -*- coding: utf-8 -*-
"""monitor runner —— 尤其是抖音那条(不走子进程的)路的运行状态流转。
这条路的地位特殊:它不经过 ``crawler_manager``,所以爬虫那套「退出码 / 日志尾巴」的
约定它一个都不沾。凡是写在那里面的东西,这条路都得单独有一份。
"""
import asyncio
import pytest
import pytest_asyncio
from sqlalchemy import select
from tools.time_util import get_current_timestamp
from api.monitor import db as monitor_db
from api.monitor import runner as runner_module
from api.monitor.models import (
MODE_CREATOR,
RUN_RUNNING,
MonitorRun,
MonitorTarget,
MonitorTask,
)
@pytest_asyncio.fixture
async def db(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
yield monitor_db
await monitor_db.dispose_engine()
async def _make_douyin_task() -> int:
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="dy", platform="dy", mode=MODE_CREATOR, enabled=True,
interval_minutes=360, max_notes_count=20, enable_comments=False,
max_comments_count=20, run_timeout_seconds=3600,
notify_enabled=False, notify_failures=False,
created_at=now, updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorTarget(
task_id=task.id, kind=MODE_CREATOR, external_id="MS4w-sec",
xsec_token="", xsec_source="", raw_value="MS4w-sec",
label="x", enabled=True, created_at=now,
)
)
return task.id
class TestDouyinRunStatus:
@pytest.mark.asyncio
async def test_the_run_is_marked_running_before_collecting(self, db, monkeypatch):
"""**采集开始之前**,run 就必须已经是 running。
这一行原先只写在爬虫那条分支里,于是抖音路上 run 一直停在 pending —— 一旦中途
出事(异常、或进程被重启),界面上就是一个永远「排队中」的幽灵,而且 recover()
当时也只收 running、够不着它。
"""
task_id = await _make_douyin_task()
seen = {}
async def fake_collect(out_dir, **kwargs):
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun).order_by(MonitorRun.id))
seen["status"] = run.status
return {
"notes": 0,
"comments": 0,
"errors": ["故意失败"],
"jsonl_dir": str(out_dir),
}
monkeypatch.setattr(runner_module.douyin_fetch, "collect", fake_collect)
await runner_module.execute_task(task_id, trigger="manual")
assert seen["status"] == RUN_RUNNING
@pytest.mark.asyncio
async def test_a_hanging_collect_does_not_leave_the_run_running(self, db, monkeypatch):
"""进程内那条路也要有超时。
爬虫那条靠 ``run_and_wait(timeout=...)`` 兜底,这条路没有子进程、没人管 ——
里面任何一次卡住(实测过 ``page.evaluate`` 打在一个卡死的标签页上不返回)都会让
run 永远停在「运行中」,界面上看起来就是任务卡死了。
"""
task_id = await _make_douyin_task()
async with monitor_db.get_session() as session:
task = await session.get(MonitorTask, task_id)
task.run_timeout_seconds = 1 # 把超时压到 1 秒,别让测试真等
async def hanging_collect(out_dir, **kwargs):
await asyncio.sleep(60)
raise AssertionError("不该走到这里")
monkeypatch.setattr(runner_module.douyin_fetch, "collect", hanging_collect)
await runner_module.execute_task(task_id, trigger="manual")
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun).order_by(MonitorRun.id))
assert run.status != RUN_RUNNING
assert "超时" in (run.error_message or "") or "超过" in (run.error_message or "")
+340
View File
@@ -0,0 +1,340 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_monitor_scheduler.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Tests for the monitor scheduler's firing, deferral and recovery rules."""
import pytest
import pytest_asyncio
from sqlalchemy import select
from api.monitor import db as monitor_db
from api.monitor import scheduler as scheduler_module
from api.monitor.models import (
MODE_CREATOR,
MonitorRun,
MonitorTarget,
MonitorTask,
RUN_INTERRUPTED,
RUN_PENDING,
RUN_RUNNING,
RUN_SUCCESS,
)
from api.monitor.scheduler import MonitorScheduler
from api.monitor.settings import set_cookie, set_setting
from tools.time_util import get_current_timestamp
MS_PER_MINUTE = 60_000
class FakeCrawlerManager:
"""Stands in for the global subprocess singleton."""
def __init__(self, busy: bool = False) -> None:
self.busy = busy
def is_busy(self) -> bool:
return self.busy
@pytest_asyncio.fixture
async def db(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
async with monitor_db.get_session() as session:
await set_cookie(session, "web_session=test")
yield monitor_db
await monitor_db.dispose_engine()
@pytest_asyncio.fixture
async def executed(monkeypatch):
"""Record execute_task calls instead of launching a real crawl."""
calls: list[tuple[int, str]] = []
async def _fake_execute(task_id: int, trigger: str = "manual"):
calls.append((task_id, trigger))
monkeypatch.setattr(scheduler_module, "execute_task", _fake_execute)
return calls
async def _make_task(
next_run_at, enabled: bool = True, interval: int = 60, platform: str = "xhs"
) -> int:
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="t",
platform=platform,
mode=MODE_CREATOR,
enabled=enabled,
interval_minutes=interval,
max_notes_count=20,
enable_comments=True,
max_comments_count=50,
run_timeout_seconds=3600,
next_run_at=next_run_at,
last_status="idle",
created_at=now,
updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorTarget(
task_id=task.id,
kind=MODE_CREATOR,
external_id="abc123",
xsec_token="",
xsec_source="",
raw_value="abc123",
label="abc123",
enabled=True,
created_at=now,
)
)
return task.id
async def _get_task(task_id: int) -> MonitorTask:
async with monitor_db.get_session() as session:
return await session.get(MonitorTask, task_id)
class TestFiring:
@pytest.mark.asyncio
async def test_due_task_runs_and_advances(self, monkeypatch, db, executed):
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
past = get_current_timestamp() - MS_PER_MINUTE
task_id = await _make_task(past)
await MonitorScheduler().tick()
assert executed == [(task_id, "scheduled")]
task = await _get_task(task_id)
# Fixed-delay: the next fire is measured from now, not from the missed slot.
assert task.next_run_at > get_current_timestamp()
@pytest.mark.asyncio
async def test_future_task_does_not_run(self, monkeypatch, db, executed):
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
await _make_task(get_current_timestamp() + 10 * MS_PER_MINUTE)
await MonitorScheduler().tick()
assert executed == []
@pytest.mark.asyncio
async def test_disabled_task_does_not_run(self, monkeypatch, db, executed):
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
await _make_task(get_current_timestamp() - MS_PER_MINUTE, enabled=False)
await MonitorScheduler().tick()
assert executed == []
@pytest.mark.asyncio
async def test_the_cookie_gate_reads_the_tasks_own_platform(
self, monkeypatch, db, executed
):
"""cookie 闸门要按任务自己的平台取。
以前这里是 ``get_cookie(session)``(默认小红书)—— 只有小红书时看不出问题,
接上抖音后,抖音任务会因为读的是小红书那份 cookie 而永远不被触发,且不报错。
"""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
async with monitor_db.get_session() as session:
await set_cookie(session, "sessionid=dy-secret", "dy")
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
await MonitorScheduler().tick()
assert executed == [(task_id, "scheduled")]
@pytest.mark.asyncio
async def test_cdp_mode_frees_a_task_from_the_cookie_gate(
self, monkeypatch, db, executed
):
"""开着 CDP 时不该再要求先粘 cookie。
CDP 模式下登录态来自被接管的那台浏览器,粘不粘 cookie 都由不得它 —— 不放行的话,
选了「接管已有 Chrome」却没粘 cookie 的用户会发现任务永远不跑,而且什么错都不报。
"""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
async with monitor_db.get_session() as session:
await set_setting(session, "system.cdp_enabled", "true")
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
await MonitorScheduler().tick()
assert executed == [(task_id, "scheduled")]
@pytest.mark.asyncio
async def test_long_outage_coalesces_into_one_run(self, monkeypatch, db, executed):
"""A missed schedule fires once, not once per missed interval."""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
# Due two days ago on a 1-hour interval.
await _make_task(get_current_timestamp() - 48 * 60 * MS_PER_MINUTE)
scheduler = MonitorScheduler()
await scheduler.tick()
await scheduler.tick()
assert len(executed) == 1
class TestDeferral:
@pytest.mark.asyncio
async def test_busy_crawler_defers_without_advancing(self, monkeypatch, db, executed):
"""A manual crawl must not consume the monitor task's slot or lose it."""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=True))
due_at = get_current_timestamp() - MS_PER_MINUTE
task_id = await _make_task(due_at)
await MonitorScheduler().tick()
assert executed == []
task = await _get_task(task_id)
# Still due, so the next free tick picks it up rather than skipping a cycle.
assert task.next_run_at == due_at
@pytest.mark.asyncio
async def test_deferred_task_runs_once_crawler_frees_up(self, monkeypatch, db, executed):
fake = FakeCrawlerManager(busy=True)
monkeypatch.setattr(scheduler_module, "crawler_manager", fake)
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE)
scheduler = MonitorScheduler()
await scheduler.tick()
assert executed == []
fake.busy = False
await scheduler.tick()
assert executed == [(task_id, "scheduled")]
class TestCookieGuard:
@pytest.mark.asyncio
async def test_no_cookie_blocks_run_and_keeps_task_due(self, monkeypatch, db, executed):
"""Without a cookie every run would be an auth failure; skip instead."""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
async with monitor_db.get_session() as session:
from api.monitor.settings import cookie_key, delete_setting
await delete_setting(session, cookie_key("xhs"))
due_at = get_current_timestamp() - MS_PER_MINUTE
task_id = await _make_task(due_at)
await MonitorScheduler().tick()
assert executed == []
task = await _get_task(task_id)
# Left due so it starts working the moment a cookie is pasted.
assert task.next_run_at == due_at
class TestRecovery:
@pytest.mark.asyncio
async def test_running_runs_are_marked_interrupted(self, db):
"""A run left 'running' cannot be alive -- its process died with the server."""
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="t", platform="xhs", mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=True,
max_comments_count=50, run_timeout_seconds=3600,
next_run_at=now, last_status="running", created_at=now, updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorRun(
task_id=task.id, trigger="scheduled", status=RUN_RUNNING,
phase=MODE_CREATOR, save_data_path="", queued_at=now, not_before=0,
started_at=now, max_comments_count=50,
)
)
await MonitorScheduler().recover()
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun))
assert run.status == RUN_INTERRUPTED
assert run.finished_at is not None
@pytest.mark.asyncio
async def test_pending_runs_are_also_cleaned_up(self, db):
"""挂在 ``pending`` 的 run 同样是残留,必须一起收。
那一行是上一轮建的,可它后面的采集根本没机会开始(进程被重启,或采集那条路抛了
异常)。只清 ``running`` 的话,它会永远挂在界面上显示「排队中」——
用户看到的就是任务卡死了。
"""
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="t", platform="dy", mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
next_run_at=now, last_status="pending", created_at=now, updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorRun(
task_id=task.id, trigger="manual", status=RUN_PENDING,
phase=MODE_CREATOR, save_data_path="", queued_at=now, not_before=0,
max_comments_count=50,
)
)
await MonitorScheduler().recover()
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun))
assert run.status == RUN_INTERRUPTED
assert run.finished_at is not None
@pytest.mark.asyncio
async def test_completed_runs_are_left_alone(self, db):
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="t", platform="xhs", mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=True,
max_comments_count=50, run_timeout_seconds=3600,
next_run_at=now, last_status="success", created_at=now, updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorRun(
task_id=task.id, trigger="scheduled", status=RUN_SUCCESS,
phase=MODE_CREATOR, save_data_path="", queued_at=now, not_before=0,
started_at=now, finished_at=now, max_comments_count=50,
)
)
await MonitorScheduler().recover()
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun))
assert run.status == RUN_SUCCESS
+24 -2
View File
@@ -14,6 +14,8 @@ import pathlib
import pytest
import config
ROOT = pathlib.Path(__file__).resolve().parent.parent
# 统一的禁用字段名(键)。昵称字段(nickname/user_nickname/screen_name/name/user_name)允许保留(值需脱敏)。
@@ -28,6 +30,17 @@ NICK_KEYS = {"nickname", "user_nickname", "screen_name", "name", "user_name"}
MASK_RE = re.compile(r"^.?\*{1,4}.?$")
@pytest.fixture(autouse=True)
def _force_nickname_masking(monkeypatch):
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
所以这里显式打开来测;开关两个方向的行为由 test_mask_and_hash_tools 覆盖。
"""
monkeypatch.setattr(config, "MASK_NICKNAME", True)
# ----------------------------- ORM 自省 -----------------------------
def test_orm_has_no_forbidden_columns():
@@ -84,17 +97,26 @@ def _check_nickname_masked(d: dict, raw: str, label: str):
assert MASK_RE.match(val) or "*" in val, f"[{label}] {k} 未脱敏: {val}"
def test_mask_and_hash_tools():
def test_mask_and_hash_tools(monkeypatch):
from tools.user_hash import anonymize_user_id, mask_nickname
h = anonymize_user_id("12345")
assert h and h != "12345" and re.fullmatch(r"[0-9a-f]{16}", h)
assert anonymize_user_id(None) == "" and anonymize_user_id("") == ""
# 昵称脱敏:首尾留1字、中间星号,且不等于原文
# 开关打开:首尾留 1 字、中间星号,且不等于原文。
monkeypatch.setattr(config, "MASK_NICKNAME", True)
assert mask_nickname("张三丰") != "张三丰"
assert "*" in mask_nickname("张三丰")
assert mask_nickname(None) == ""
assert mask_nickname("a") == "*"
# 开关关闭(本仓库的部署配置):原样返回。脱敏是有损的 —— 「张三」和「张四」
# 都会变成「张*」,而分清谁是谁正是监控这一层要干的事。
monkeypatch.setattr(config, "MASK_NICKNAME", False)
assert mask_nickname("张三丰") == "张三丰"
assert mask_nickname("a") == "a"
assert mask_nickname(None) == ""
def test_xhs_note_extraction_masks_user_info():
import asyncio
+391
View File
@@ -0,0 +1,391 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_platforms.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Platform capability matrix, platform scoping, and per-platform settings."""
import httpx
import pytest
import pytest_asyncio
from sqlalchemy import select, text
from tools.time_util import get_current_timestamp
from api.main import app
from api.monitor import adapters
from api.monitor import db as monitor_db
from api.monitor import platforms
from api.monitor.models import MonitorNote, MonitorNoteMetric, MonitorTask
XHS_TARGET = "5f58bd990000000001003753"
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
async def _seed_note_with_one_snapshot(task_name: str) -> None:
"""给某个任务塞一条作品和一次指标快照。
过滤类测试**必须有真数据**才有意义 —— 库里空着的话,过滤有没有生效结果都是 0,
测试就变成了空跑(这个坑踩过一次:一个报表串数据的 bug 因此没被拦住)。
"""
async with monitor_db.get_session() as session:
task = await session.scalar(select(MonitorTask).where(MonitorTask.name == task_name))
now = get_current_timestamp()
session.add(
MonitorNote(
task_id=task.id, note_id="seed-note", title="seed", note_url="",
cover="", creator_hash="", source_kind="", published_at=None,
first_seen_run_id=1, first_seen_at=now,
last_seen_run_id=1, last_seen_at=now,
)
)
session.add(
MonitorNoteMetric(
task_id=task.id, note_id="seed-note", run_id=1, captured_at=now,
liked_count=42, comment_count=0, collected_count=0, share_count=0,
raw_liked_count="42", raw_comment_count="0",
raw_collected_count="0", raw_share_count="0",
)
)
class TestCapabilityMatrix:
@pytest.mark.asyncio
async def test_matrix_is_exposed_to_the_ui(self, client):
body = (await client.get("/api/config/platforms")).json()
by_value = {p["value"]: p for p in body["platforms"]}
assert set(by_value) == {"xhs", "dy", "ks", "bili", "wb", "tieba", "zhihu"}
# Every entry must say whether monitoring is actually wired up -- this is
# what stops the UI offering a platform that can never produce data.
assert all("monitor_wired" in p for p in body["platforms"])
assert by_value["xhs"]["monitor_wired"] is True
assert by_value["dy"]["monitor_wired"] is True
def test_every_wired_platform_has_an_adapter(self):
"""能力矩阵说「接通了」,就必须真的有一套适配管子。
两个注册表(platforms.PLATFORM_CAPABILITIES 与 adapters.ADAPTERS)分开是有意的
—— 前者是给前端看的能力描述,后者是爬虫的管道细节。代价是它们可能漂移,
所以在这里钉一条:凡声明接通的,必须能找到适配器。
"""
for platform in platforms.all_platforms():
if platforms.is_monitor_wired(platform):
assert adapters.has_adapter(platform), f"{platform} 声明接通但没有适配器"
@pytest.mark.asyncio
async def test_target_hints_are_exposed_for_wired_platforms(self, client):
"""前端的目标输入框拿它做 placeholder —— 让用户看到本平台该粘什么样的链接。"""
body = (await client.get("/api/config/platforms")).json()
by_value = {p["value"]: p for p in body["platforms"]}
assert "douyin.com/user/" in by_value["dy"]["target_hints"]["creator"]
assert "douyin.com/video/" in by_value["dy"]["target_hints"]["note"]
assert "xiaohongshu.com" in by_value["xhs"]["target_hints"]["creator"]
@pytest.mark.asyncio
async def test_metrics_are_per_platform_and_labelled(self, client):
body = (await client.get("/api/config/platforms")).json()
by_value = {p["value"]: p for p in body["platforms"]}
# Bilibili has play count and danmaku; Xiaohongshu has neither.
assert "video_play_count" in by_value["bili"]["metrics"]
assert "video_danmaku" in by_value["bili"]["metrics"]
assert "video_play_count" not in by_value["xhs"]["metrics"]
# Every metric shown to a user must have a human label.
for capability in body["platforms"]:
for metric in capability["metrics"]:
assert capability["metric_labels"][metric]
def test_unknown_platform_is_not_monitor_wired(self):
assert platforms.is_known("xhs") is True
assert platforms.is_known("myspace") is False
assert platforms.is_monitor_wired("myspace") is False
class TestTaskCreationGuard:
@pytest.mark.asyncio
async def test_unwired_platform_is_rejected_with_an_explanation(self, client):
"""Accepting it would create a task that silently never produces data.
用 B站 而不是抖音:抖音现在接通了,不再是「已知但未接通」的例子。
"""
response = await client.post(
"/api/monitor/tasks",
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
)
assert response.status_code == 400
detail = response.json()["detail"]
assert "B站" in detail
assert "尚未接通" in detail
@pytest.mark.asyncio
async def test_unknown_platform_is_rejected(self, client):
response = await client.post(
"/api/monitor/tasks",
json={"name": "x", "mode": "creator", "platform": "myspace", "targets": ["x"]},
)
assert response.status_code == 400
@pytest.mark.asyncio
async def test_no_task_row_is_created_when_rejected(self, client):
await client.post(
"/api/monitor/tasks",
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
)
assert (await client.get("/api/monitor/tasks")).json()["tasks"] == []
@pytest.mark.asyncio
async def test_xhs_still_works_and_is_the_default(self, client):
explicit = await client.post(
"/api/monitor/tasks",
json={"name": "显式", "mode": "creator", "platform": "xhs", "targets": [XHS_TARGET]},
)
assert explicit.status_code == 201
defaulted = await client.post(
"/api/monitor/tasks",
json={"name": "默认", "mode": "creator", "targets": [XHS_TARGET]},
)
assert defaulted.status_code == 201
tasks = (await client.get("/api/monitor/tasks")).json()["tasks"]
assert {t["platform"] for t in tasks} == {"xhs"}
@pytest.mark.asyncio
async def test_an_explicit_platform_is_honoured_on_create(self, client):
"""建任务时给的平台必须落到那个平台。
缺省值是小红的(接口早期的兼容行为),所以「在抖音页面建任务」如果没有显式
带上 platform,就会安安静静地变成一个小红书任务 —— 不报错,只是出现在另一
个列表里。前端那半边已经改成必传;这里守住后端这一半:给了就必须用。
"""
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
created = await client.post(
"/api/monitor/tasks",
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": [sec_uid]},
)
assert created.status_code == 201
assert (await client.get("/api/monitor/tasks", params={"platform": "xhs"})).json()[
"tasks"
] == []
dy_tasks = (
await client.get("/api/monitor/tasks", params={"platform": "dy"})
).json()["tasks"]
assert [t["name"] for t in dy_tasks] == ["抖音任务"]
class TestPlatformScoping:
async def _seed_two_platforms(self, client):
"""One XHS task created through the API, plus a Douyin task inserted
directly so its fields can be pinned exactly."""
await client.post(
"/api/monitor/tasks",
json={"name": "小红书任务", "mode": "creator", "targets": [XHS_TARGET]},
)
async with monitor_db.get_session() as session:
session.add(
MonitorTask(
name="抖音任务", platform="dy", mode="creator", enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=True,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
)
@pytest.mark.asyncio
async def test_tasks_are_filtered_by_platform(self, client):
await self._seed_two_platforms(client)
all_tasks = (await client.get("/api/monitor/tasks")).json()["tasks"]
assert len(all_tasks) == 2
xhs_only = (await client.get("/api/monitor/tasks", params={"platform": "xhs"})).json()
assert [t["name"] for t in xhs_only["tasks"]] == ["小红书任务"]
dy_only = (await client.get("/api/monitor/tasks", params={"platform": "dy"})).json()
assert [t["name"] for t in dy_only["tasks"]] == ["抖音任务"]
@pytest.mark.asyncio
async def test_overview_is_scoped(self, client):
await self._seed_two_platforms(client)
assert (await client.get("/api/monitor/overview")).json()["tasks"] == 2
assert (
await client.get("/api/monitor/overview", params={"platform": "xhs"})
).json()["tasks"] == 1
@pytest.mark.asyncio
async def test_a_platform_with_no_tasks_yields_empty_not_everything(self, client):
"""空的任务集合不能退化成「不加过滤」。
**这里必须真的有数据。** 没有数据时,过滤生效与否结果都是 0 —— 这条测试原先
就栽在这个空跑上,所以没能拦下一个报表串数据的 bug(切到抖音,报表里却出现
小红书的数据)。最后那段「小红书自己的报表看得到」就是为了证明这些数据确实
存在、上面那两个 0 是过滤出来的。
"""
await self._seed_two_platforms(client)
await _seed_note_with_one_snapshot("小红书任务")
body = (await client.get("/api/monitor/notes", params={"platform": "bili"})).json()
assert body["notes"] == []
report = (
await client.get("/api/monitor/report", params={"platform": "bili"})
).json()
assert report["totals"]["liked_count_delta"] == 0
assert report["note_count"] == 0
xhs = (
await client.get("/api/monitor/report", params={"platform": "xhs"})
).json()
assert xhs["note_count"] == 1
class TestPerPlatformSettings:
@pytest.mark.asyncio
async def test_each_platform_keeps_its_own_values(self, client):
await client.put(
"/api/settings",
params={"platform": "xhs"},
json={"platform.xhs.crawl_sleep_sec": 3},
)
await client.put(
"/api/settings",
params={"platform": "dy"},
json={"platform.dy.crawl_sleep_sec": 9},
)
xhs = (await client.get("/api/settings", params={"platform": "xhs"})).json()
dy = (await client.get("/api/settings", params={"platform": "dy"})).json()
assert xhs["values"]["platform.xhs.crawl_sleep_sec"] == 3
assert dy["values"]["platform.dy.crawl_sleep_sec"] == 9
@pytest.mark.asyncio
async def test_system_settings_are_shared_across_platforms(self, client):
await client.put(
"/api/settings",
params={"platform": "xhs"},
json={"system.active_hours_start": 8},
)
dy = (await client.get("/api/settings", params={"platform": "dy"})).json()
assert dy["values"]["system.active_hours_start"] == 8
# ...and the system specs are present in every platform's response.
assert "system.active_hours_end" in dy["values"]
@pytest.mark.asyncio
async def test_a_key_for_another_platform_is_rejected(self, client):
"""Writing xhs's key while scoped to dy would land somewhere unexpected."""
response = await client.put(
"/api/settings",
params={"platform": "dy"},
json={"platform.xhs.crawl_sleep_sec": 5},
)
assert response.status_code == 400
@pytest.mark.asyncio
async def test_cookies_are_per_platform(self, client):
await client.post(
"/api/monitor/cookie",
params={"platform": "xhs"},
json={"cookie": "web_session=xhs-secret"},
)
xhs = (await client.get("/api/monitor/cookie", params={"platform": "xhs"})).json()
dy = (await client.get("/api/monitor/cookie", params={"platform": "dy"})).json()
assert xhs["present"] is True
assert dy["present"] is False
# The old endpoint still defaults to Xiaohongshu.
assert (await client.get("/api/monitor/cookie")).json()["present"] is True
class TestLegacyKeyMigration:
@pytest.mark.asyncio
async def test_old_flat_keys_are_moved_to_the_new_namespace(self, tmp_path):
"""Existing installs must not lose their cookie on upgrade."""
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
async with monitor_db.get_engine().begin() as conn:
await conn.execute(
text(
"INSERT INTO monitor_setting (key, value, updated_at) "
"VALUES ('xhs_cookie', 'web_session=legacy', 1)"
)
)
await conn.execute(
text(
"INSERT INTO monitor_setting (key, value, updated_at) "
"VALUES ('wecom_webhook', 'https://qyapi.weixin.qq.com/x', 1)"
)
)
# Re-running init performs the rename.
await monitor_db.init_db()
async with monitor_db.get_engine().begin() as conn:
rows = dict(
(await conn.execute(text("SELECT key, value FROM monitor_setting"))).all()
)
assert rows.get("platform.xhs.cookie") == "web_session=legacy"
assert rows.get("system.wecom_webhook") == "https://qyapi.weixin.qq.com/x"
assert "xhs_cookie" not in rows
assert "wecom_webhook" not in rows
await monitor_db.dispose_engine()
@pytest.mark.asyncio
async def test_migration_is_idempotent_and_keeps_the_newer_value(self, tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
async with monitor_db.get_engine().begin() as conn:
await conn.execute(
text(
"INSERT INTO monitor_setting (key, value, updated_at) VALUES "
"('platform.xhs.cookie', 'current', 2), ('xhs_cookie', 'stale', 1)"
)
)
await monitor_db.init_db()
async with monitor_db.get_engine().begin() as conn:
rows = dict(
(await conn.execute(text("SELECT key, value FROM monitor_setting"))).all()
)
assert rows.get("platform.xhs.cookie") == "current"
assert "xhs_cookie" not in rows
await monitor_db.dispose_engine()
+268
View File
@@ -0,0 +1,268 @@
# -*- coding: utf-8 -*-
"""监控侧扫码登录与登录态检测。
两个要点在这里被钉死:
* 二维码必须从浏览器**默认 context** 里读 —— 新建 context 是无痕式的 profile,
扫了也白扫,爬虫读不到那份 cookie;
* 「登录了吗」不能用页面里的 `window.__INITIAL_STATE__`。那是**页面加载那一刻的快照**:
浏览器本来就登录着时它是对的,但扫码是加载**之后**才登录的,快照不会翻转,
于是扫完码界面会一直停在二维码上。判据改成拿 cookie 问后台接口。
"""
from unittest.mock import AsyncMock, MagicMock
import pytest
from api.creator.client import CreatorApiError
from api.monitor import qrlogin
@pytest.fixture(autouse=True)
def _reset_module_state():
for attribute in ("_current", "_page", "_playwright", "_state_cache"):
setattr(qrlogin, attribute, None)
yield
for attribute in ("_current", "_page", "_playwright", "_state_cache"):
setattr(qrlogin, attribute, None)
XHS_COOKIES = [
{"name": "a1", "value": "an-a1-value"},
{"name": "web_session", "value": "a-session"},
]
def _fake_stack(cookies=None, qr="data:image/png;base64,AAAA"):
"""Chrome/Playwright 替身,行为与真实的一致。"""
page = MagicMock()
page.url = "https://www.xiaohongshu.com/explore"
page.is_closed = MagicMock(return_value=False)
page.goto = AsyncMock()
page.close = AsyncMock()
context = MagicMock()
context.pages = []
context.cookies = AsyncMock(return_value=list(cookies if cookies is not None else XHS_COOKIES))
context.new_page = AsyncMock(return_value=page)
browser = MagicMock()
browser.contexts = [context]
# 去新建 context 正是这里要防的 bug,所以让它直接炸,而不是悄悄返回一个无痕 profile。
browser.new_context = AsyncMock(
side_effect=AssertionError("must reuse browser.contexts[0], not a new context")
)
playwright = MagicMock()
playwright.chromium.connect_over_cdp = AsyncMock(return_value=browser)
playwright.stop = AsyncMock()
manager = MagicMock()
manager.start = AsyncMock(return_value=playwright)
return manager, playwright, browser, context, page, qr
def _patch(monkeypatch, manager, qr="data:image/png;base64,AAAA", resolver=None):
"""``resolver(cookie)`` 返回账号信息 dict,或抛 CreatorApiError。"""
if resolver is None:
resolver = lambda _cookie: {"user_id": "u1", "nickname": "小明"} # noqa: E731
class _Client:
def __init__(self, cookie, **kwargs):
self.cookie = cookie
async def fetch_user_info(self):
return resolver(self.cookie)
monkeypatch.setattr(qrlogin, "async_playwright", lambda: manager)
monkeypatch.setattr(qrlogin, "CreatorClient", _Client)
monkeypatch.setattr(qrlogin.utils, "find_login_qrcode", AsyncMock(return_value=qr))
def _signed_out(_cookie):
raise CreatorApiError("登录态无效或已过期", status=401)
@pytest.mark.asyncio
async def test_idle_reports_the_browsers_login_state(monkeypatch):
manager, *_ = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
snapshot = await qrlogin.status()
assert snapshot["status"] == qrlogin.STATUS_IDLE
assert snapshot["logged_in"] is True
assert snapshot["nickname"] == "小明"
@pytest.mark.asyncio
async def test_unwired_platform_is_rejected():
"""只有小红书接了扫码;别的平台必须直接报错,而不是给个按不动的按钮。"""
with pytest.raises(ValueError):
await qrlogin.start("dy")
@pytest.mark.asyncio
async def test_start_reads_the_qr_from_the_default_context(monkeypatch):
manager, _pw, browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=_signed_out)
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
assert snapshot["status"] == qrlogin.STATUS_WAITING
assert snapshot["image"] == "data:image/png;base64,AAAA"
context.new_page.assert_awaited_once()
browser.new_context.assert_not_called()
@pytest.mark.asyncio
async def test_start_does_not_open_a_second_tab(monkeypatch):
"""已有的 xhs 标签页会被认领,所以重启不会在浏览器里堆孤儿页。"""
manager, _pw, _browser, context, page, _qr = _fake_stack()
context.pages = [page]
_patch(monkeypatch, manager, resolver=_signed_out)
await qrlogin.start(qrlogin.PLATFORM_XHS)
context.new_page.assert_not_called()
@pytest.mark.asyncio
async def test_an_already_signed_in_profile_needs_no_scan(monkeypatch):
"""没二维码但 profile 已登录 —— 这是成功,不是失败。"""
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, qr="", resolver=lambda _c: {"user_id": "u9", "nickname": "老王"})
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
assert snapshot["status"] == qrlogin.STATUS_SUCCESS
assert snapshot["nickname"] == "老王"
@pytest.mark.asyncio
async def test_an_already_signed_in_profile_never_opens_a_page(monkeypatch):
"""已登录时**根本不该去开页面**。
读二维码内部会 wait_for_selector 等满 30 秒才放弃,而已经登录时页面上没有二维码 ——
顺序反了的话,用户点一下按钮要干等半分钟,还白开一个标签页。
"""
manager, _pw, _browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
await qrlogin.start(qrlogin.PLATFORM_XHS)
context.new_page.assert_not_called()
@pytest.mark.asyncio
async def test_no_qr_and_not_signed_in_is_an_error(monkeypatch):
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, qr="", resolver=_signed_out)
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
assert snapshot["status"] == qrlogin.STATUS_ERROR
@pytest.mark.asyncio
async def test_a_completed_scan_flips_the_session_to_success(monkeypatch):
"""**这个判据是重点**:扫码是页面加载之后才发生的,所以不能用页面快照来判断。"""
manager, *_rest = _fake_stack()
signed_in = {"value": False}
def resolver(cookie):
assert "a1=an-a1-value" in cookie # 判据必须真的用 cookie 去问
if not signed_in["value"]:
raise CreatorApiError("登录态无效或已过期", status=401)
return {"user_id": "u1", "nickname": "小红"}
_patch(monkeypatch, manager, resolver=resolver)
await qrlogin.start(qrlogin.PLATFORM_XHS)
assert qrlogin._current.status == qrlogin.STATUS_WAITING
# 操作者扫了码
signed_in["value"] = True
qrlogin._state_cache = None # 5 秒缓存否则会遮住这次变化
snapshot = await qrlogin.status()
assert snapshot["status"] == qrlogin.STATUS_SUCCESS
assert snapshot["nickname"] == "小红"
@pytest.mark.asyncio
async def test_the_successful_session_hands_over_a_cookie(monkeypatch):
"""扫码不该只写浏览器 profile —— 还要能把 cookie 交出来存库,
否则关掉 CDP 就断了。"""
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
await qrlogin.start(qrlogin.PLATFORM_XHS)
cookie = await qrlogin.take_cookie()
assert cookie is not None
assert "a1=an-a1-value" in cookie
# 只能取一次,否则每次轮询都会重复写库
assert await qrlogin.take_cookie() is None
@pytest.mark.asyncio
async def test_session_expires(monkeypatch):
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, resolver=_signed_out)
await qrlogin.start(qrlogin.PLATFORM_XHS)
qrlogin._current.started_at -= qrlogin.QR_TTL_SECONDS + 1
snapshot = await qrlogin.status()
assert snapshot["status"] == qrlogin.STATUS_EXPIRED
@pytest.mark.asyncio
async def test_check_login_state_reports_when_the_browser_cannot_answer(monkeypatch):
"""连不上浏览器时要说出来,不能悄悄报成「未登录」。"""
manager, _pw, _browser, context, _page, _qr = _fake_stack()
context.cookies = AsyncMock(side_effect=RuntimeError("Target closed"))
_patch(monkeypatch, manager)
state = await qrlogin.check_login_state()
assert state["known"] is False
assert state["logged_in"] is False
assert "Target closed" in state["error"]
@pytest.mark.asyncio
async def test_check_login_state_is_cached(monkeypatch):
manager, _pw, _browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "x"})
await qrlogin.check_login_state()
await qrlogin.check_login_state()
context.cookies.assert_awaited_once()
@pytest.mark.asyncio
async def test_force_bypasses_the_cache(monkeypatch):
manager, _pw, _browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "x"})
await qrlogin.check_login_state()
await qrlogin.check_login_state(force=True)
assert context.cookies.await_count == 2
@pytest.mark.asyncio
async def test_cancel_keeps_the_operators_tab(monkeypatch):
"""与运营模块不同:那里的上下文是临时的、用完即弃;这里的标签页属于操作者的浏览器。"""
manager, _pw, _browser, _context, page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=_signed_out)
await qrlogin.start(qrlogin.PLATFORM_XHS)
snapshot = await qrlogin.cancel()
assert snapshot["status"] == qrlogin.STATUS_IDLE
page.close.assert_not_called()
+194
View File
@@ -0,0 +1,194 @@
# -*- coding: utf-8 -*-
"""Tests for monitor task schedule arithmetic.
Everything here is timezone-local, matching the implementation: the container is
pinned to the operator's zone via TZ, so the tests build their expectations from
naive local datetimes too and stay correct wherever they run.
"""
from datetime import datetime, time, timedelta
import pytest
from api.monitor import schedule
def _ms(moment: datetime) -> int:
return int(moment.timestamp() * 1000)
def test_interval_is_now_plus_the_interval():
after = _ms(datetime(2026, 10, 7, 9, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_INTERVAL,
interval_minutes=120,
hours=[],
days=[],
minute=0,
after_ms=after,
)
assert nxt == after + 120 * 60_000
def test_daily_takes_the_soonest_remaining_time_today():
after = _ms(datetime(2026, 10, 7, 8, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[18, 9], # deliberately unsorted
days=[],
minute=30,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 7, 9, 30))
def test_daily_rolls_over_to_tomorrow_once_every_time_has_passed():
after = _ms(datetime(2026, 10, 7, 20, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[9, 18],
days=[],
minute=30,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 8, 9, 30))
def test_a_slot_exactly_now_belongs_to_the_next_day():
"""Strictly-after, so the run that just fired does not fire again.
The scheduler advances with ``after_ms`` set to the moment the run started,
which is at or just past the slot -- if the comparison were inclusive it would
pick the same slot back up and loop.
"""
after = _ms(datetime(2026, 10, 7, 9, 30))
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[9],
days=[],
minute=30,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 8, 9, 30))
def test_weekly_jumps_to_the_next_selected_weekday():
after_dt = datetime(2026, 10, 7, 8, 0)
target = (after_dt.weekday() + 2) % 7
nxt = schedule.next_occurrence(
mode=schedule.MODE_WEEKLY,
interval_minutes=60,
hours=[10],
days=[target],
minute=0,
after_ms=_ms(after_dt),
)
expected = datetime.combine((after_dt + timedelta(days=2)).date(), time(10, 0))
assert nxt == _ms(expected)
def test_weekly_can_fire_later_the_same_day():
after_dt = datetime(2026, 10, 7, 8, 0)
nxt = schedule.next_occurrence(
mode=schedule.MODE_WEEKLY,
interval_minutes=60,
hours=[21],
days=[after_dt.weekday()],
minute=15,
after_ms=_ms(after_dt),
)
assert nxt == _ms(datetime.combine(after_dt.date(), time(21, 15)))
def test_weekly_without_weekdays_means_every_day():
"""Otherwise an empty day selection would match nothing and never fire."""
after = _ms(datetime(2026, 10, 7, 8, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_WEEKLY,
interval_minutes=60,
hours=[9],
days=[],
minute=0,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 7, 9, 0))
def test_a_clock_schedule_with_no_times_can_never_fire():
"""Returned as None so the caller can park the task instead of leaving it due."""
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[],
days=[],
minute=0,
after_ms=_ms(datetime(2026, 10, 7, 8, 0)),
)
assert nxt is None
@pytest.mark.parametrize(
"raw, expected",
[
("9,18", [9, 18]),
("18,9", [9, 18]), # stored order is not guaranteed
("9,9,9", [9]),
("", []),
(None, []),
("9, 18 ", [9, 18]),
("9,99,-1,abc,", [9]), # junk is dropped, never raised
],
)
def test_parse_hours_is_forgiving(raw, expected):
assert schedule.parse_hours(raw) == expected
def test_parse_days_accepts_the_whole_week():
assert schedule.parse_days("0,1,2,3,4,5,6") == [0, 1, 2, 3, 4, 5, 6]
assert schedule.parse_days("7,-1") == []
@pytest.mark.parametrize(
"mode, interval, hours, days, minute, expected",
[
(schedule.MODE_INTERVAL, 360, [], [], 0, "每 6 小时"),
(schedule.MODE_INTERVAL, 1440, [], [], 0, "每 1 天"),
(schedule.MODE_INTERVAL, 45, [], [], 0, "每 45 分钟"),
(schedule.MODE_DAILY, 60, [9, 18], [], 30, "每天 09:30、18:30"),
(schedule.MODE_WEEKLY, 60, [10], [0, 1, 2, 3, 4], 0, "周一、周二、周三、周四、周五 10:00"),
(schedule.MODE_WEEKLY, 60, [10], [], 0, "每天 10:00"),
(schedule.MODE_DAILY, 60, [], [], 0, "未设置时间"),
],
)
def test_describe(mode, interval, hours, days, minute, expected):
assert (
schedule.describe(
mode=mode, interval_minutes=interval, hours=hours, days=days, minute=minute
)
== expected
)
def test_format_round_trips_through_parse():
hours = [9, 12, 18]
days = [0, 4]
assert schedule.parse_hours(schedule.format_hours(hours)) == hours
assert schedule.parse_days(schedule.format_days(days)) == days
+374
View File
@@ -0,0 +1,374 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_settings.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Unified settings endpoint, and the effect its values actually have."""
from datetime import datetime
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import app_settings, db as monitor_db
from api.monitor import scheduler as scheduler_module
from api.monitor.scheduler import MonitorScheduler
from api.monitor.settings import get_setting
SECRET_VALUE = "web_session=SUPERSECRET; a1=abc"
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
class TestReadSettings:
@pytest.mark.asyncio
async def test_returns_values_secrets_and_the_spec(self, client):
body = (await client.get("/api/settings")).json()
assert "values" in body and "secrets" in body and "specs" in body
# The spec drives the UI form, so every key must be described.
spec_keys = {spec["key"] for spec in body["specs"]}
assert "platform.xhs.default_interval_minutes" in spec_keys
assert "platform.xhs.enable_ip_proxy" in spec_keys
@pytest.mark.asyncio
async def test_unset_values_fall_back_to_spec_defaults(self, client):
values = (await client.get("/api/settings")).json()["values"]
assert values["platform.xhs.default_interval_minutes"] == 360
assert values["platform.xhs.enable_ip_proxy"] is False
@pytest.mark.asyncio
async def test_secrets_are_masked_never_returned(self, client):
await client.put("/api/settings", json={"platform.xhs.cookie": SECRET_VALUE})
response = await client.get("/api/settings")
assert SECRET_VALUE not in response.text
secret = response.json()["secrets"]["platform.xhs.cookie"]
assert secret["present"] is True
assert secret["length"] == len(SECRET_VALUE)
class TestUpdateSettings:
@pytest.mark.asyncio
async def test_partial_update_leaves_other_keys_alone(self, client):
await client.put(
"/api/settings",
json={"platform.xhs.default_interval_minutes": 120, "platform.xhs.cookie": SECRET_VALUE},
)
# A form that only submits the interval must not blank the cookie.
await client.put("/api/settings", json={"platform.xhs.default_interval_minutes": 240})
body = (await client.get("/api/settings")).json()
assert body["values"]["platform.xhs.default_interval_minutes"] == 240
assert body["secrets"]["platform.xhs.cookie"]["present"] is True
@pytest.mark.asyncio
async def test_empty_string_clears_a_secret(self, client):
await client.put("/api/settings", json={"platform.xhs.cookie": SECRET_VALUE})
await client.put("/api/settings", json={"platform.xhs.cookie": ""})
assert (await client.get("/api/settings")).json()["secrets"]["platform.xhs.cookie"][
"present"
] is False
@pytest.mark.asyncio
async def test_unknown_key_is_rejected(self, client):
response = await client.put("/api/settings", json={"nope.not.a.setting": 1})
assert response.status_code == 400
@pytest.mark.asyncio
async def test_out_of_range_is_rejected(self, client):
response = await client.put(
"/api/settings", json={"platform.xhs.default_interval_minutes": 1}
)
assert response.status_code == 400
@pytest.mark.asyncio
async def test_invalid_choice_is_rejected(self, client):
response = await client.put("/api/settings", json={"platform.xhs.proxy_provider": "nonsense"})
assert response.status_code == 400
@pytest.mark.asyncio
async def test_bools_accept_the_ui_shapes(self, client):
for raw in (True, "true", "1", "yes"):
response = await client.put("/api/settings", json={"platform.xhs.enable_ip_proxy": raw})
assert response.status_code == 200
assert (await client.get("/api/settings")).json()["values"][
"platform.xhs.enable_ip_proxy"
] is True
@pytest.mark.asyncio
async def test_password_hash_cannot_be_written_through_this_endpoint(self, client):
"""It has its own authenticated endpoint; this must not be a back door."""
await client.put("/api/settings", json={"auth_password_hash": "pbkdf2_sha256$1$a$b"})
async with monitor_db.get_session() as session:
assert await get_setting(session, "auth_password_hash") is None
class TestSettingsActuallyTakeEffect:
@pytest.mark.asyncio
async def test_new_tasks_use_the_configured_defaults(self, client):
await client.put(
"/api/settings",
json={
"platform.xhs.default_interval_minutes": 120,
"platform.xhs.default_max_notes": 7,
"platform.xhs.default_max_comments": 33,
},
)
await client.post(
"/api/monitor/tasks",
json={"name": "用默认值", "mode": "creator", "targets": ["5f58bd990000000001003753"]},
)
task = (await client.get("/api/monitor/tasks")).json()["tasks"][0]
assert task["interval_minutes"] == 120
assert task["max_notes_count"] == 7
assert task["max_comments_count"] == 33
@pytest.mark.asyncio
async def test_explicit_values_still_win_over_defaults(self, client):
await client.put("/api/settings", json={"platform.xhs.default_interval_minutes": 120})
await client.post(
"/api/monitor/tasks",
json={
"name": "显式值",
"mode": "creator",
"interval_minutes": 720,
"targets": ["5f58bd990000000001003753"],
},
)
task = (await client.get("/api/monitor/tasks")).json()["tasks"][0]
assert task["interval_minutes"] == 720
class TestRunnerAppliesStrategy:
@pytest.mark.asyncio
async def test_strategy_settings_reach_the_command(self, client):
"""Stored settings must actually change how the crawler is invoked."""
from api.services.crawler_manager import CrawlerManager
from api.schemas import CrawlerStartRequest, PlatformEnum, CrawlerTypeEnum
await client.put(
"/api/settings",
json={
"platform.xhs.crawl_sleep_sec": 7,
"platform.xhs.enable_sub_comments": True,
"platform.xhs.enable_ip_proxy": True,
"platform.xhs.proxy_provider": "static",
"platform.xhs.proxy_pool_count": 5,
"platform.xhs.static_proxy_url": "http://127.0.0.1:8888",
},
)
async with monitor_db.get_session() as session:
strategy = await scheduler_module.app_settings.get_value(
session, "crawl_sleep_sec", "xhs", 2
)
assert strategy == 7
# And the flag builder forwards them when present.
command = CrawlerManager()._build_command(
CrawlerStartRequest(
platform=PlatformEnum.XHS,
crawler_type=CrawlerTypeEnum.CREATOR,
creator_ids="abc",
crawler_max_sleep_sec=7,
enable_ip_proxy=True,
ip_proxy_provider_name="static",
ip_proxy_pool_count=5,
static_proxy_url="http://127.0.0.1:8888",
)
)
joined = " ".join(command)
assert "--crawler_max_sleep_sec 7" in joined
assert "--enable_ip_proxy true" in joined
assert "--ip_proxy_provider_name static" in joined
assert "--static_proxy_url http://127.0.0.1:8888" in joined
def _frozen_clock(hour: int):
"""Stand-in for the datetime class whose now() is pinned to a given hour.
Testing an hour window by sleeping is not an option; patching the class the
scheduler imported is the whole mechanism.
"""
class _Frozen:
@staticmethod
def now(tz=None):
return datetime(2026, 1, 1, hour)
return _Frozen
class TestManualCrawlCookieFallback:
"""The crawl page no longer has its own paste box; it reuses Settings."""
@pytest_asyncio.fixture
async def captured(self, monkeypatch):
# api.services re-exports the singleton instance, not the module.
from api.services import crawler_manager
seen: dict = {}
async def _fake_start(request, extra_args=None):
seen["cookies"] = request.cookies
return True
monkeypatch.setattr(crawler_manager, "start", _fake_start)
return seen
@pytest.mark.asyncio
async def test_falls_back_to_the_stored_cookie(self, client, captured):
await client.put(
"/api/settings", json={"platform.xhs.cookie": "web_session=stored"}
)
response = await client.post(
"/api/crawler/start",
json={
"platform": "xhs",
"login_type": "cookie",
"crawler_type": "creator",
"creator_ids": "abc",
},
)
assert response.status_code == 200
assert captured["cookies"] == "web_session=stored"
@pytest.mark.asyncio
async def test_an_explicit_cookie_still_wins(self, client, captured):
await client.put(
"/api/settings", json={"platform.xhs.cookie": "web_session=stored"}
)
await client.post(
"/api/crawler/start",
json={
"platform": "xhs",
"login_type": "cookie",
"crawler_type": "creator",
"creator_ids": "abc",
"cookies": "web_session=explicit",
},
)
assert captured["cookies"] == "web_session=explicit"
@pytest.mark.asyncio
async def test_missing_cookie_is_a_clear_error_not_a_silent_failure(
self, client, captured
):
"""Better a 400 that names the fix than a run that fetches nothing."""
response = await client.post(
"/api/crawler/start",
json={
"platform": "xhs",
"login_type": "cookie",
"crawler_type": "creator",
"creator_ids": "abc",
},
)
assert response.status_code == 400
assert "设置" in response.json()["detail"]
assert "cookies" not in captured
@pytest.mark.asyncio
async def test_the_cookie_is_read_per_platform(self, client, captured):
await client.put(
"/api/settings",
params={"platform": "xhs"},
json={"platform.xhs.cookie": "web_session=xhs-only"},
)
# Douyin has no stored cookie, so it must not borrow Xiaohongshu's.
response = await client.post(
"/api/crawler/start",
json={
"platform": "dy",
"login_type": "cookie",
"crawler_type": "creator",
"creator_ids": "abc",
},
)
assert response.status_code == 400
class TestActiveHours:
"""The window gate lives in the scheduler, not the crawler."""
@pytest.mark.asyncio
async def test_inside_a_daytime_window(self, client, monkeypatch):
await client.put(
"/api/settings",
json={"system.active_hours_start": 8, "system.active_hours_end": 22},
)
monkeypatch.setattr(scheduler_module, "datetime", _frozen_clock(12))
async with monitor_db.get_session() as session:
assert await MonitorScheduler()._within_active_hours(session) is True
@pytest.mark.asyncio
async def test_outside_a_daytime_window(self, client, monkeypatch):
await client.put(
"/api/settings",
json={"system.active_hours_start": 8, "system.active_hours_end": 22},
)
monkeypatch.setattr(scheduler_module, "datetime", _frozen_clock(3))
async with monitor_db.get_session() as session:
assert await MonitorScheduler()._within_active_hours(session) is False
@pytest.mark.asyncio
async def test_window_wrapping_past_midnight(self, client, monkeypatch):
await client.put(
"/api/settings",
json={"system.active_hours_start": 22, "system.active_hours_end": 6},
)
for hour, expected in ((23, True), (3, True), (12, False)):
monkeypatch.setattr(scheduler_module, "datetime", _frozen_clock(hour))
async with monitor_db.get_session() as session:
assert await MonitorScheduler()._within_active_hours(session) is expected
@pytest.mark.asyncio
async def test_default_window_covers_the_whole_day(self, client, monkeypatch):
for hour in (0, 12, 23):
monkeypatch.setattr(scheduler_module, "datetime", _frozen_clock(hour))
async with monitor_db.get_session() as session:
assert await MonitorScheduler()._within_active_hours(session) is True
+8 -8
View File
@@ -20,7 +20,7 @@ def test_extract_search_note_list_from_keyword_page():
assert notes[0].note_id == "9117888152"
assert notes[0].title.startswith("武汉交互空间科技")
assert notes[0].tieba_name == "武汉交互空间"
assert notes[0].user_nickname == "V***人"
assert notes[0].user_nickname == "VR虚拟达人"
def test_extract_search_note_list_from_current_pc_card_page():
@@ -56,7 +56,7 @@ def test_extract_search_note_list_from_current_pc_card_page():
assert notes[0].desc == "培训班需求,数学,英语,编程老师,专职兼职都可"
assert notes[0].tieba_name == "诸城吧"
assert notes[0].tieba_link.endswith("kw=%E8%AF%B8%E5%9F%8E")
assert notes[0].user_nickname == "7***7"
assert notes[0].user_nickname == "754023117"
assert notes[0].publish_time == "2026-3-15"
assert notes[0].total_replay_num == 19
@@ -147,7 +147,7 @@ def test_extract_note_detail_and_comments_from_current_pc_api():
assert note.note_id == "10451142633"
assert note.title == "这X尔斯对比巴尔斯,我只能说ID正确,允许居功自傲"
assert note.desc == "皮队败决处刑德国编程钢琴师兼职数学家"
assert note.user_nickname == "泰***克"
assert note.user_nickname == "泰高祖蒙斯克"
assert note.tieba_name == "dota2吧"
assert note.total_replay_num == 15
assert note.total_replay_page == 1
@@ -155,7 +155,7 @@ def test_extract_note_detail_and_comments_from_current_pc_api():
assert len(comments) == 1
assert comments[0].comment_id == "153154097267"
assert comments[0].content == "xg现在大树阵容另一个辅助不选控制"
assert comments[0].user_nickname == "期***3"
assert comments[0].user_nickname == "期胡希3"
assert comments[0].sub_comment_count == 4
# 教学版已移除 ip_location 等可定位真人字段
@@ -191,7 +191,7 @@ def test_extract_creator_info_and_threads_from_current_pc_api():
creator = extractor.extract_creator_info_from_api(creator_api)
thread_ids = extractor.extract_creator_thread_id_list_from_api(feed_api)
assert creator.user_nickname == "米***子"
assert creator.user_nickname == "米米世界大手子"
assert creator.fans == 58
assert creator.follows == 1
# 教学版已移除 user_id、user_name、ip_location 等可定位真人字段
@@ -223,7 +223,7 @@ def test_extract_tieba_note_list_from_bigpipe_thread_page():
assert len(notes) == 48
assert notes[0].note_id == "9079949995"
assert notes[0].title == "盗墓笔记全集+txt小说,已整理"
assert notes[0].user_nickname == "公***仲"
assert notes[0].user_nickname == "公子伯仲"
assert notes[0].tieba_name == "盗墓笔记吧"
assert notes[0].tieba_link.endswith("kw=%E7%9B%97%E5%A2%93%E7%AC%94%E8%AE%B0&ie=utf-8")
@@ -233,7 +233,7 @@ def test_extract_note_detail_from_post_page():
assert note.note_id == "9117905169"
assert note.title == "对于一个父亲来说,这个女儿14岁就死了"
assert note.user_nickname == "章***轩"
assert note.user_nickname == "章景轩"
assert note.tieba_name == "以太比特吧"
assert note.total_replay_num == 786
assert note.total_replay_page == 13
@@ -249,7 +249,7 @@ def test_extract_parent_comments_from_post_page():
assert len(comments) == 30
assert comments[0].comment_id == "150726491368"
assert comments[0].content == "中国队第22金!无悬念!"
assert comments[0].user_nickname == "h***n"
assert comments[0].user_nickname == "heinzfrentzen"
assert comments[0].tieba_name == "网球风云吧"
# 教学版已移除 ip_location 等可定位真人字段

Some files were not shown because too many files have changed in this diff Show More