Compare commits

...
49 Commits
Author SHA1 Message Date
butubb 61808444ad feat(monitor): 评论接口带上 a_bogus 签名;修好作品导出的空列
Deploy VitePress site to Pages / build (push) Waiting to run
Deploy VitePress site to Pages / Deploy (push) Blocked by required conditions
**评论能采了。** 之前 comment/list 一直回 200 + 空 body,被读成「这条没评论」——
而它其实只是被网关挡了。缺的就是 a_bogus 签名,仓库里本来就有
(libs/douyin.js + execjs)。签上之后实测 200 / 9960 字节真评论。

只给评论接口签:作品、详情、博主资料三个不带签名也照常返回,而给它们加签名是
没验证过的改动。签名按需 import —— 那个模块 import 时就把 JS 喂给 execjs,
不该拖进监控层热路径。

**作品导出那几列一直是空的。** 列名写的是裸键 liked_count,而作品行的指标嵌在
metrics / deltas 里,row.get() 永远取到 None —— 导出来的表有「点赞/评论/收藏/
分享」四列,每一格都没有数。原来的测试只断言了「作品ID」,所以没发现。

顺手补上:导出带上 博主备注/昵称、作品备注、发布时间,时间戳格式化成人能读的
形态(原来是一串 13 位毫秒,Excel 里没法看也没法排序)。
2026-10-10 21:09:52 +08:00
butubb e0581682e1 fix(monitor): 没有作品的博主在作品栏里也要看得见
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
分组是从作品推出来的(按作品的 creator_hash 归组),于是没有作品的博主根本
不进列表:目标加了、资料采到了、粉丝数就躺在库里,界面上什么都看不见。

而这类博主恰恰是最该看见的 —— 还在涨粉,只是最近没发东西。藏起来正好藏反了。
线上就有一个:5 个目标里 3 个没作品,那 3 个连同已采到的粉丝数一起消失了。

改成 **账号快照 ∪ 作品** 两个来源:

* 有快照没作品 → 一个 0 篇的组,备注和粉丝数照常显示,组里写「暂无作品」;
* 有作品没快照 → 一个没有账号指标的组(小红书那条路不产生快照)。

/notes 因此多返回一份 `creators`,而不是让前端从作品里推 —— 作品推不出上面
第一类人。组头改读它,`MonitorNote` 上那几个字段降级成「顺着作品问作者」用。
2026-10-10 20:57:07 +08:00
butubb 3c4daae0ac chore: 忽略部署用的临时脚本
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
.remote_run.py 是我在部署机上跑命令的工具,不是这个项目的一部分。
2026-10-10 18:27:47 +08:00
butubb d02fec5914 fix(monitor): 组头取账号指标时,别被组里第一条作品带偏
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
分组只按博主,而账号指标是挂在 (任务, 博主) 上的 —— 同一个博主被两个任务
监控时,组里可能只有一部分作品带指标。取第一条的话,「第一个任务还没采过」
就会让整组看起来没有指标。
2026-10-10 18:24:42 +08:00
butubb 20e672834c feat(monitor): 博主的粉丝数,以及给作品起备注
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
两件都是「作品栏里把这东西认出来」的延伸:

* **账号级指标**:作品列表只会说「这条涨了多少赞」,说不了「这个人整个
  账号的粉丝在涨还是在掉」。抖音的资料接口本来就有粉丝数/总获赞/作品数,
  每轮顺手记一条快照(`monitor_creator_stat`,粒度 = 任务×博主×轮次,
  和作品指标同形)。组头显示最近一条。

  快照在「一条作品都没采到」的早退**之前**落:作品列表被风控挡住的那一轮,
  正是「粉丝还在涨、但新作品没在发现」最该被看见的时刻。

* **作品备注**:博主备注回答「这个账号是谁」,这条回答「这条我要盯着」。
  一个博主底下常常只有一两件值得盯的作品,所以不能合并成一条。键取
  (platform, note_id),跨任务共用一份。

两边都守住同一条口径:**不知道就是 null,不写成 0** —— 0 在趋势图上是一条
砸到底的线,和「还没采到」是两回事。
2026-10-10 18:09:03 +08:00
butubb 0a88474c92 feat(monitor): 给博主起备注 —— 作品栏里才认得出「这是谁」
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
作品栏按 creator_hash 把作品归到博主的组里,可那是个哈希;creator_name 是平台昵称,
粉丝少的号常常认不出。两样都没法把账号对上人。

* 新表 monitor_creator_alias,键取 **(platform, creator_hash)** 而不是按任务:
  哈希对同一个 uid 是稳定的,所以同一个博主出现在多个任务里时,备注只填一次。
* list_notes 带上 creator_alias(整体查一次再在内存里取,不是每条作品查一次)。
* `PUT /monitor/creators/{creator_hash}`,空串即清掉那条备注。
* 作品栏的博主组头:**优先显示备注**,起过备注之后平台昵称降成副标题(它仍是有用的
  对照);组头上一个铅笔按钮就地编辑,回车保存、失焦保存、Esc 取消。

Esc 那条要单独处理:取消之后紧接着的失焦会把刚放弃的内容存进去,所以用一个标记让那次
失焦闭嘴。

测试 +6:没起过时是空串、起了会跟着作品返回、**跨任务共用一条**、**不串到别的平台**、
空串清掉、前后空格会被去掉。
2026-10-10 18:00:13 +08:00
butubb 5552e2a2b8 fix(monitor): 抖音任务永远「运行中」—— page.evaluate 卡在一个死掉的标签页上,而我没给超时
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户报「一直在运行中」。库里那条 run 是真的卡住了:日志里一条采集输出都没有,
说明它卡在 collect 里、还没走到任何日志。探针定位到:

    标签页: ['']                        ← 一个 URL 为空的标签页,渲染进程已卡死
    context.cookies(): OK, 90 个          ← cookie 读得到
    page.evaluate('navigator.userAgent'): **永远不返回**

而问浏览器要身份(UA + client hints)是采集的**第一步**,`page.evaluate` 又**没设超时** ——
于是整轮挂在那儿,run 永远停在「运行中」。

三处修复,各挡一层:

1. `page.evaluate` / `context.cookies()` 全部加超时(8 秒)。卡住就跳过,不再无限等。
2. 不假设第一个标签页是好的:逐个试、优先抖音页;全都不行就临时开一个干净页问完关掉。
   拿不到就退回库里那份 cookie —— **不编造指纹**,那比没有更糟。
3. **进程内那条路补上整体超时**:爬虫那条靠 `run_and_wait(timeout=...)` 兜底,这条路
   没有子进程、没人管,里面任何一次卡住都会变成永久的「运行中」。

测试 +6:卡死的页会被跳过(真 sleep,验的正是超时)、没 UA 的页跳过、全不行时开临时页
并关掉它、优先抖音页;以及整轮卡住时 run 不会停在 running(含超时原因)。
2026-10-10 17:51:54 +08:00
butubb 9e13a7f686 fix(monitor): 抖音任务的 run 永远停在「排队中」—— 我上一版把状态标记缩进错了
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户报的现象:任务一直显示「排队中」。查库确认有两批 run 卡在 pending(任务 7 的 44/45、
任务 15 的 62/63)。两个原因,一个是我上一版改坏的:

1) **`RUN_RUNNING` 被我缩进进了爬虫那条分支。** 抖音走的是另一条路,于是它**从不标记
   「运行中」** —— 建完 pending 那一行就直接进采集,中途一旦出事(异常、进程被重启),
   状态就永远停在 pending。这是我加平台分岔时把原本在两条路公共位置的一行挪进去了。

2) **`recover()` 只收 `running`,够不着 `pending`。** 那行是上一轮建的、后面的采集却
   根本没机会开始(进程重启),它永远不会自己往前走。于是重启也救不回来,界面上就是
   一个永远「排队中」的幽灵。现在 pending 一起收。

两处都补了测试:抖音路的 run 必须在**采集开始之前**就已经是 running(这条改回去就会
失败);recover 要把 pending 也标成 interrupted。
2026-10-10 17:47:41 +08:00
butubb 9f70cd0924 fix(monitor): 抖音「作品」模式的目标被当成博主去查,白废一条本来能用的路
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
两种模式的**目标是不同的东西**,原来的 fetcher 却一条路走到底:

* 「作品」模式(粘贴作品链接)—— 目标本身就是作品 id,直接取详情即可。**这个接口没被
  那道真校验挡,今天就能用。**
* 「博主」模式 —— 目标是主页 sec_uid,要先拉作品列表;那个接口被挡,退化成刷新已知作品。

原来两种都去调 author_videos(它要的是博主 sec_uid),于是「作品」模式的监控拿作品号当
sec_uid 去查,必然失败 —— 而且失败原因说得很难懂(接口回你「未登录/不是浏览器」)。
结果就是:**新建「作品」模式的抖音监控永远抓不到东西**,而那本来是现有条件下唯一能用的。

现在按 mode 分岔。测试 +2:作品模式必须走 detail 且**不得**去调列表接口(走错了会
直接抛断言);一件作品坏掉不连累其他作品。

顺带记一条排查结论:博主主页的 HTML 里**没有**作品列表(RENDER_DATA 解出来只有
{isLogin, statusCode, isSpider}),所以「走页面 HTML 免接口」那条路也是死的。
2026-10-10 17:30:28 +08:00
butubb 95b1be2c2e fix(monitor): 一轮产物里重复的作品会让指标快照撞唯一键,整个 run 崩掉
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
跑真任务时踩到的:

    sqlalchemy.exc.IntegrityError: (1062, "Duplicate entry
    '7-7690458980574358513-45' for key 'uq_note_metric'")

根因是我上一版写错了一处作用域:退化路径里那个「遍历已知作品」的循环写在了**目标循环内部**,
所以任务有多个目标时,同一批已知作品会被拉两遍 → 同一件作品在一轮里出现两条记录 →
ingest 给同一件作品写两份本轮快照 → 撞 (task_id, note_id, run_id) 唯一键。

两处都修,各挡一层:

* douyin_fetch:去重集合挪到 collect 的最外层,**跨目标**只算一次;退化时也先查一遍
  已知作品是否已刷过。
* ingest:`_ingest_notes` 对「一轮里重复出现的 note_id」免疫。一层在源头、一层在入口,
  因为产物里重复并不罕见(多个目标指向同一个人、上游重跑、退化路径),不该靠上游自觉。

测试 +2:多目标时已知作品只刷一次;同一轮里重复的作品只落一份快照(这条会崩在
唯一键上,所以它测的正是运行时的那个崩法)。
2026-10-10 17:24:08 +08:00
butubb 3b6a437e6c feat(monitor): 抖音监控改走新的 Web 接口客户端(接上上一版的移植)
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上一版只把客户端写出来、验证了它单独可用,**但没有接进任何地方** —— 所以你建的任务跑起来
仍然在调爬虫子进程,报的仍然是那句自己编的「account blocked」。这一步把它接上。

* runner 的 Phase 2 按平台分岔:dy 走进程内 HTTP 客户端(douyin_fetch),其余平台照旧走
  爬虫子进程。抖音那条不再起 Playwright、不再构造那串自相矛盾的浏览器指纹参数。
* 新增 douyin_fetch:把采到的东西写成 store/douyin 那套 jsonl 形状 —— **ingest 完全不知道
  数据是从哪来的**,重采样/差分/事件/通知/报表全都照旧,一个字没改。
* 失败不再假装:一条都没采到就以非零退出码 + **真实原因**交给 ingest,落成
  「采集进程异常退出(code=1):…」。绝不会再掉进「疑似登录失效」那个分支。
* 已知作品列表接口(aweme/post)被抖音单独加了真校验(200 + 空 body),所以加了退化:
  拿不到列表就用库里已知的 aweme_id 逐条走 detail 刷新。**边界是:已知作品的指标能继续
  更新,新作品发现不了** —— 这个边界会以一条 warning 日志留下痕迹,不让它看起来一切正常。
* 顺带给客户端补上 video_detail(实测可用:200 / 45425 字节),退化路径靠它。

测试 +6:产物目录与文件名、评论文件即使为空也要建(ingest 靠它区分「没评论」和
「什么都没抓到」)、重复作品只写一次、列表被挡时的退化、彻底失败仍写出产物与原因、
评论失败不连累作品。
2026-10-10 17:21:12 +08:00
butubb 4f83a075f6 chore: 忽略 webui/.npm(服务器重建前端时容器写入的缓存,属构建产物)
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
2026-10-10 17:15:32 +08:00
butubb 7442bc10b8 feat(monitor): 抖音 Web 接口客户端 —— 绕开爬虫子进程,直接发 HTTP
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
移植自 mac-agent-os 的 mediacrawler_adapter:不起子进程、不开页面,用浏览器里那份
登录态直接调抖音 Web 接口。产物键名照抄 store/douyin,所以 ingest 那条链路一个字不用改。

**目前能用的(真环境实测,非推断)**:

    profile/other : 200, 7075 字节  —— 博主主页指标(粉丝/获赞/作品数/昵称)
    aweme/detail  : 200, 45425 字节 —— 单条作品详情(含点赞/评论/收藏/分享)

**目前不能用的:作品列表 `aweme/post`。** 两个互相独立的原因:
  1. 这个接口被抖音单独升级成了真校验:不带 x-tt-argus 回 403「Uifid Not Found」,
     带上 dummy 值回 200 + **空 body**。也就是说「头在不在」骗得过,「真校验」过不了。
     同一套头打 profile/other 和 aweme/detail 都是通的 —— 抖音是挑着接口加保护的,
     挑中的恰好是「批量拉作品列表」这个最敏感的动作。
  2. 改走页面截获也不行:CDP 浏览器打开博主主页会落到「验证码中间页」(当天大量探测的
     代价,过几小时要重测)。

  所以现在的边界是:**已知作品的指标刷新能做,自动发现新作品做不了**。

**排查中控住变量后得到的两条事实**(都写进注释了):
  · `Accept` / `Accept-Language` / `Referer` 才是主页接口能返回真数据的原因 —— 只有
    UA+client hints+Cookie 时是 200 但仅 121 字节的空壳,补上这三个头变 7074 字节。
    (我先前猜的 sec-ch-ua 不是关键。)
  · 因此 UA 与 client hints 必须**成套地取自同一个浏览器**,所以 BrowserIdentity 一次
    从 CDP 取齐 cookie + UA + hints,而不是各自写死。

「200 + 空 body 必须当场报错」也是刻意写死的:放过去它会在下游变成「这个博主没作品」,
把一次失败伪装成一条正常结果 —— 爬虫那条路正是这么栽的,还被翻译成「账号被封」。

顺带:把参考项目目录加进 .gitignore。上一次 `git add -A` 把 mac-agent-os-main 整个
(1429 个文件)带进了提交,已从历史里清掉。

测试 +11:cookie 解析、请求头成套性(含 uifid 缺失/回退)、产物键名与 store 对齐、
以及 _get 的三条失败路径(空 body / 403 带网关原话 / 正常返回)。
2026-10-10 17:13:04 +08:00
butubb 21ff01b894 fix(douyin): 缺 sec-ch-ua 请求头,网关回 200 + 空 body(不是「账号被封」)
先纠正一个我上一轮给错的结论:日志里的 `account blocked` **不是抖音说的**,是
MediaCrawler 自己编的:

    if response.text == "" or response.text == "blocked":
        raise Exception("account blocked")

真实情况只是**抖音返回了空 body**。我把它读成了「账号被风控」,还写进了运行历史和
给用户的结论里 —— 用户质疑「我网页版和手机版都能正常登录」,一查,他是对的。

实测定位(同一 URL、同一 cookie、同一参数):

  浏览器页面内 fetch : 200, 7077 字节  ✓
  httpx              : 200,    0 字节  ✗
    带 a_bogus       : 0 字节
    不带 a_bogus     : 0 字节
    四种 msToken 变体 : 全部 200 有数据(所以不是它)
  用浏览器那次的完整头重放 httpx : 200, 7077 字节 ✓

浏览器那次请求比爬虫多的,只有这三个头:

    sec-ch-ua: "Google Chrome";v="155", "Chromium";v="155", "Not(A:Brand";v=24
    sec-ch-ua-mobile: ?0
    sec-ch-ua-platform: "Linux"

爬虫的 UA 是从页面读的(声称是 Chrome 155)却不带 sec-ch-ua —— 「Chrome 的 UA +
没有 sec-ch-ua」是最典型的机器人特征。网关的回应方式是不报错、不给原因,回一个
200 + 空 body,HTTP 状态还写在成功那一栏。

修:media_platform/douyin/help.py 新增 client_hint_headers(),从
navigator.userAgentData 现算这三个头(现算而不是写死 —— 写死的版本号一旦和 UA 里的
对不上,就又是一个可疑特征);core.py 建客户端时带上。

诚实说明:我无法解释**为什么之前能跑**(run 34 还是成功的,40 分钟后同样的代码就
不行了)。最可能是字节那边收紧了这道校验,但我没有证据,别当结论。

测试 +8:还原出的头与真实浏览器抓到的值逐字一致;拿不到 userAgentData 时返回空而不
凭空编造(编一组和 UA 对不上的头比不带头更糟);mobile 标志;platform 缺失时仍发另两个。
2026-10-10 16:47:49 +08:00
butubb e4affe9170 feat(monitor): 运行历史要写清楚失败原因,不能只写「退出码 1」
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户的要求:运行历史的说明要写清楚。现在确实写不清楚 —— 抖音那次失败,运行历史里
只有一句 `Crawler exited with code 1`,而真正的报错 `DataFetchError: account blocked`
埋在子进程的 stderr 里,谁也看不到。

那两者本来是断开的两条路:子进程的输出只流向日志 WebSocket(前端 Terminal 看得到),
而监控层调 run_and_wait() 只拿得到一个退出码。

* crawler_manager 在 _push_log() 里留一份输出尾巴(80 行,每次 start 清空)——
  那是所有输出的唯一出口,挂这儿不会漏。新增 get_output_tail()。
* ingest 新增 diagnose_failure():倒着找第一行像异常的行(traceback 的末行),
  认不出就退回最后一行;管理器自己补的「Crawler exited with code」不是原因,排除掉。
* describe_exit_code() 接受这个原因并附在消息里;失败事件的标题也带上,这样企业微信
  通知和事件流不用翻日志就能看懂。
* runner 把尾巴交给 ingest;「超时/没起来」那条分支同样带上原因 —— -1 同时代表两种
  情况,而要查的东西完全不同。
* 运行历史那一格是截断的(240px),而失败原因现在有一整行 —— 补上 title 悬停显示,
  并放宽到 320px。没有悬停提示等于把最要紧的半句藏起来。

测试 +5:能挑出异常行、不会把管理器自己的话当成原因、没有输出时不报错、认不出时退回
最后一行、以及失败运行同时记下退出码与真因(含事件标题)。
2026-10-10 16:14:54 +08:00
butubb 1118d466be fix(monitor): 抖音的时间戳是秒、小红书是毫秒,不换算会把 2026 年显示成 1970 年
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上一个提交把作品发布日期透到界面之后,抖音那一行显示成 1970-01-22。查原始产物:

  小红书 time        = 1790923011000  → 毫秒 → 2026-10-02   ✓
  抖音   create_time = 1790574515     → 秒   → 2026-09-28   ✗(被当毫秒 → 1970-01-22)

抖音给的是**秒**。而且这条路径不只影响新增的发布日期 —— **评论的 create_time 走的是
同一条路**,所以抖音评论的时间一直是错的,只是之前界面上没显示出来,没人发现。

时间单位的换算正是 adapters 该管的事,所以加在那边:

* `PlatformAdapter.time_scale`(小红书 1、抖音 1000)+ `to_ms()`,解析不出来返回 None
  而不是伪造 0。
* ingest 用它换算作品的 published_at 和评论的 create_time。落库统一毫秒,展示层不必
  关心来源。
* 已入库的数据要能自愈:published_at 和评论 create_time 原先都是**只写一次**的,换算
  改对了老数据也修不回来。现在它们会在重采时跟着刷新(昵称早就是这么做的)。

测试 +1:抖音记录落库后 published_at 是 1790574515 * 1000,且年份是 2026 不是 1970。
dy 的 fixture 也改成用真实的秒值(原来写的是毫秒形态,所以测不出这个 bug)。

注意:库里那条抖音记录**仍带着错的值**,要等抖音下一次成功采集才会被修回来 —— 而它
现在正被平台风控挡着(account blocked),见下一条说明。
2026-10-10 16:08:42 +08:00
butubb 3486c7f524 feat(monitor): 作品栏和评论栏显示作品的发布日期
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
需求:抖音和小红书都要能看到作品的发布日期。

`MonitorNote.published_at` 其实**一直在库里**(ingest 早就按平台取:小红书 time、
抖音 create_time),只是从来没往 API 和界面上透 —— 后端序列化没这个键,前端类型里
也没有,所以界面上只有「首次发现」。

* 后端:list_notes 的序列化补上 published_at;_note_meta_map 也带上,于是
  list_comments 多一个 note_published_at,分组接口的桶多一个 published_at。
* 前端:NotesTable 新增「发布日期」列(要让分组表头的 colspan 从 +4 变 +5);
  评论栏作品那一层在标题旁显示日期 —— 同名作品不少,日期能帮着认。
* 新增 formatDate:发布日期问的是「哪一天发的」,绝对日期比「3天前」好认,也不会
  每天看都在变。具体到分钟的版本放在 title 里,悬停可见。

刻意和「首次发现」分开:前者是作者发布的那天,后者是我们第一次看到它的那天。把一个
早就存在的作品加进监控时,两者能差好几个月 —— 测试里就用不同的值把这两者钉住。

顺带修正一处过时注释:前端类型里还写着 creator_name 是「已脱敏的昵称」,
脱敏已经在上一个提交里关掉了(config.MASK_NICKNAME)。

测试 +2:桶要带发布日期;作品列表接口要带,且它不等于 first_seen_at。
2026-10-10 16:06:22 +08:00
butubb cd85587f00 feat(privacy): 关掉昵称脱敏 —— 脱敏有损,撞名就分不出博主
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
需求:评论栏/作品栏要能分清是哪个博主。修好分组字段之后名字仍带星号,因为上游作为
教学版默认对昵称做中间脱敏(首尾各留 1 字,中间星号)。这个脱敏是**有损**的:

  「张三」和「张四」都变成「张*」
  「小明老师」和「小刚老师」都变成「小***师」

而本仓库的用途是监控一批公开创作者账号,分清谁是谁正是这一层要干的事。所以关掉它。

* config/base_config.py 新增 MASK_NICKNAME = False(和 INJECT_ALL_COOKIES 一样是个
  开关,不是删代码 —— 改回 True 就恢复上游行为)。
* tools/user_hash.py 的 mask_nickname 读这个开关,关闭时原样返回。读的是模块属性而
  不是导入值,测试才能 monkeypatch。脱敏实现本身一字未动。
* 顺带修一个数据陈旧问题:评论是去重后直接 continue 的,昵称只在首次入库时写一次,
  于是开关一改(或评论者改名)老评论永远停在旧值 —— 而重采是唯一能拿到新值的途径。
  现在已存在的评论会跟着刷新昵称(作品那边的 creator_name 早就是这么做的)。
* anonymous 的 creator_hash 保持不变:那是分组用的稳定键,不是显示名。

测试:
* 三个隐私套件 + weibo 的 autouse fixture 强制把开关打开 —— 它们验的是**脱敏机制
  本身**,机制仍然必须正确,所以显式打开来测,而不是让它们随部署配置漂。
* test_mask_and_hash_tools 改成两个方向都覆盖(开着脱敏 / 关着脱敏)。
* test_tieba_extractor.py 里 8 处字面量的脱敏期望值换成真实昵称 —— 提取器现在就是
  返回原文的,期望值理应跟着改(这一条是行为变更的直接后果,不是测试放宽)。
* 新增一条:已入库的评论昵称会随重采刷新(且不会因刷新而重复插入)。
2026-10-10 15:48:15 +08:00
butubb 67837b407e fix(comments): 评论栏所有博主都显示成「未知博主」
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
现象:小红书的和抖音的评论栏,最外层分组全是「未知博主」,分不出谁是谁。

根因是分组接口漏了两个字段。api/routers/monitor.py 里按作品构造桶时只放了
note_id/note_title/note_cover/note_url/comments,而前端 groupByCreator 是用
bucket.creator_hash / bucket.creator_name 分组的 —— 两个都是 undefined,于是所有
博主塌成同一个 key,标签取空串回退成「未知博主」。

数据一直都在:每条评论上都带着 note_creator_hash / note_creator_name
(service.py:540-541),只是没往桶上搬。TS 的 CommentBucket 里也声明了这两个字段,
所以是后端没兑现自己的契约,不是前端写错。

修:构造桶时把作品的创作者一并放上去(同一个桶里的评论必然同属一个作品,取哪条都一样)。

测试:种子数据改成「两个作品属于不同博主」(原来是同一个 hash,测不出这个 bug),
新增一条断言每个桶带上自己那个博主、且两个博主的 hash 确实不同。
2026-10-10 15:41:42 +08:00
butubb 42f2209534 fix(douyin): 两个让抖音采集根本跑不起来的上游缺陷
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户建了个抖音博主任务,两次都是 exit 1、产物目录空空。查下来是上游两处缺陷,
与我们那层监控无关 —— 但它们在别的网络上未必复现,所以社区里没人报。

1) media_platform/douyin/core.py:101 —— 首页 goto 永远等不到 load

    await self.context_page.goto(self.index_url)   # 默认 wait_until="load"

   抖音首页的 load 事件不会触发(有长连接/埋点类请求一直挂着)。实测同一台 Chrome、
   同一个地址:domcontentloaded 0.7 秒返回,load 等满 90 秒仍超时。后果是整个采集
   一步没走就崩,退出码 1,看起来像「抖音不能用」。改成显式 domcontentloaded ——
   上游的贴吧和知乎本来就是这么写的,抖音这个页面只是恰好属于「永远不 load」那类。

2) media_platform/douyin/login.py:266 —— 注入 cookie 后页面是陈旧的

   login_by_cookies 把 cookie 塞进 context,但页面是在这之前加载的;SPA 只在加载时
   读一次登录态,localStorage.HasUserLogin 还停在"未登录",紧接着 check_login_state
   会对着这个陈旧值轮询到超时(600×1 秒=十分钟)再 sys.exit()。下一轮才正常,因为
   那时 cookie 已在 profile 里 —— 表现是"第一次白等十分钟、第二次才行",很容易被当
   偶发。修复:注入后 reload(domcontentloaded)。

   与 xhs 那个 __INITIAL_STATE__ 快照问题是同一类:页面状态是加载那一刻的快照。
   本仓库扫码登录与运营模块也各自踩过。

UPSTREAM.md 的第二类「上游 bug 修复」表补上这两条与成因说明(原来只有两条 xhs 的)。

验证:在服务器容器里手工复现,改完后真实采到数据 ——
  Parsed sec_user_id: MS4wLjABAAAArLubjxXEcLeiqehxgk4il4AHMu0hXu6_qlmD8z5WmMs
  get_all_user_aweme_posts ... video len : 1
  douyin aweme id:7690458980574358513, title:中秋哪儿都堵...
产物落在 douyin/jsonl/ 下,顺带把适配器里「平台 id 是 dy、目录是 douyin」的映射
用真实输出证实了。
2026-10-10 15:29:02 +08:00
butubb f3ea088c75 fix(monitor): 从界面上建的任务永远是小红书任务,平台从没被传上去
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
现象:站在抖音页面上新建任务,任务跑到小红书列表里去了;界面上还写着「笔记」。

根因在前端:TaskCreatePayload 里**根本没有 platform 字段**,handleSubmit 拼的 payload
自然也不带它,而后端是 `payload.get("platform") or PLATFORM_XHS` —— 于是不管在哪个
平台标签下建任务,落下来的都是小红书任务。以前只支持小红书,两边都看不出问题。

* TaskCreatePayload 补上 platform(并在注释里写明为什么它是必填),handleSubmit 带上
  当前平台。
* 更新任务时不带 platform:平台创建后不可更改,带着会让「平台能改」看起来像真的。
* 三处写死的「笔记」改成按平台取措辞(能力矩阵的 target_hints.note_label):
  · 任务编辑器的类型选择项「笔记(批量监控指定内容)」
  · 「笔记模式下此项不生效」那句提示
  · 任务卡片上的类型徽章 —— 它按**任务自己的**平台取词,不是当前平台,因为卡片未必
    只出现在同平台的列表里
  抖音管它们叫「作品」,小红书叫「笔记」,写死一个对另一个就是错的。

测试 +1:后端这一半也守住 —— 建任务时显式给了平台,就必须落到那个平台,且不得出现在
另一个平台的列表里。前端那半边是 UI,测不了,但后端守住能挡住「给了不用」这类退化。

注意:这次是纯前端漏传,后端那个「缺省回退小红书」的行为本身没变(有测试断言它是
有意为之的兼容行为)。要彻底消灭这类静默错误,可以把缺省值去掉、让 platform 必填 ——
那会破坏 API 兼容性,目前没有任何别的调用方,需要的话说一声。
2026-10-10 15:13:37 +08:00
butubb e77e5e2f15 fix(report): 空的任务集合被当成了「不限制平台」,导致报表串平台数据
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
现象:切到抖音,报表里显示的是小红书的数据。

根因是报告聚合里的一行真值判断:

    scope = list(task_ids) if task_ids else None

空列表是假值,而空列表在这里的含义是「这个平台一个任务都没有」,不是「不限制平台」。
于是 platform=dy 且抖音还没有任务时,_resolve_scope 返回的 [] 被翻译成了 None,
聚合范围从「抖音的任务」变成了**全部任务** —— 小红书的数字就这么显示在了抖音页面上。
顺带 task_ids 也回成 None,界面会显示成「全部任务」。

改成 `is not None`。空列表进去就让 in_([]) 恒假,结果为空,这才是对的。

排查时把所有同类写法过了一遍,只有这一处错,其余(service.py 的 10 处作用域judgement、
_resolve_scope、export)用的都是 `is not None`。

测试:新增两条,并且**验证过它们在修复前会红**(失败信息就是 assert 42 == 0 ——
查一个没有任何任务的平台,却返回了小红书那条作品的 42 个赞)。

同时修掉一条空跑的测试:test_a_platform_with_no_tasks_yields_empty_not_everything
原先种了任务却没有作品/指标数据,于是过滤生效与否结果都是 0,什么都测不出来 ——
这正是这个 bug 能活下来的原因。现在它会真的塞一条作品+快照进去,并在末尾断言
「小红书自己的报表看得到那条数据」,用来证明前面那两个 0 是过滤出来的而不是没数据。
2026-10-10 15:00:12 +08:00
butubb 06718a1351 feat(monitor): 抖音接入博主监控
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上游爬虫本身不缺抖音能力(三模式、四项指标、二级评论都与小红书对等、指标还是同名同列),
缺的全在监控层的适配。这次把「平台之间不一样」的管子集中到一个新模块,再把散落的
xhs 硬编码接上去。

* 新增 api/monitor/adapters.py:产物目录名、jsonl 字段别名、目标链接形态与正则、
  通知链接模板。不放进 platforms.py 是因为那个模块被 describe_all() 整个序列化进
  /api/config/platforms 交给前端,塞进正则和目录名会让爬虫内部细节漏进 API 载荷。
  代价是两个注册表可能漂移,用一条测试钉住「声明接通就必须有适配器」。
* 两个必须知道的坑,都在这版里处理掉了:
  1) 抖音的平台 id 是 dy,而 store 把产物写在 douyin/ 下(store/douyin/_store_impl.py:47)。
     不改就是 ingest 一个文件都读不到 —— 不报错,只是 0 条,然后被冒充成「疑似登录失效」。
  2) 抖音的作品没有 note_id(叫 aweme_id)、评论也用 aweme_id 指作品。ingest 第一步是
     `if not note_id: continue`,不映射就逐条全丢。
  另外抖音顶层评论的 parent_comment_id 是字符串 "0",归一成空串,免得前端多出悬空的父节点。
* 顺带把「东西抓到了、只是没落在期望目录里」单独识别出来。这类故障的现象和登录失效
  一模一样,按登录失效报会把人指去查完全错误的方向。
* 修两个既有 bug(今天只有小红书所以无害,加抖音就踩响):
  - service.py update_task 换目标时漏传 task.platform,回落到默认小红书
  - scheduler.py 取 cookie 没传 platform,抖音任务会读着小红书那份 cookie 不动
* 行为变更(已与用户确认):cookie 闸门改成「没 cookie 且没开 CDP」才跳过。
  CDP 模式下登录态来自被接管的浏览器,粘不粘 cookie 由不得它决定;不放行的话,
  选了「接管已有 Chrome」却没粘 cookie 的用户会看到任务永远不触发,而且不报错。
  副作用是开启了 CDP 的小红书任务也不再被该闸门拦住 —— 语义上是对的。
* 目标输入框的示例链接与措辞改由能力矩阵提供(notes_label 抖音说「作品」、小红书说
  「笔记」;「建议只填纯 ID」是小红书专属劝告,抖音链接不带令牌,不再显示)。

测试 +22 条(858 通过),其中最关键的是「抖音作品/评论不被静默丢弃」与「产物目录名
不等于平台 id」两条 —— 都是把最难查的失败模式钉死在回归网里。

注意:抖音这条路的**端到端尚未验证**,需要一份可用的抖音登录态(CDP 那台 Chrome 里
登录,或导出一份 cookie)。单测覆盖的是解析与入库,真实抓取还没跑过。
2026-10-10 14:55:28 +08:00
butubb e348de48d3 feat(upstream): 上游更新检查——定期比上游、落后了推企业微信
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
本仓库在上游 MediaCrawler 之上加了一整层(见 UPSTREAM.md),可合并流程默认
「有人知道上游动了」。而部署是 git pull --ff-only,只从自己的 Gitea 拉——上游的
提交不主动 fetch 就永远看不见。拖着不合并的代价是复利的:越久越难合。

于是把「上游动了没有」变成一条会自己跑、会推企业微信的通知:

* api/monitor/upstream.py:git fetch <url> <branch> 到 FETCH_HEAD,用
  rev-list --count HEAD..FETCH_HEAD 算落后数、FETCH_HEAD..HEAD 算领先数。
  用 git 而非托管商 API,因为只有 git 知道共同祖先在哪——本仓库含有上游没有的
  提交,直接比 tip 会得出错误结论。增量 fetch 只传几个新提交,不会遇到
  UPSTREAM.md 里说的「大包必断」。
* 只 fetch 到 FETCH_HEAD:不配 remote、不写 refs/remotes、不碰索引与工作区,
  所以不打断正在跑的采集,也不和 deploy.sh 的 git pull 抢锁。
* 挂在调度器 tick 上(不是采集,所以不看 is_busy、不受活跃时段限制——定时检查
  放在半夜反而最合适),按 checked_at + 间隔 到期才跑;失败也写 checked_at,
  于是 GitHub 不通时是每间隔重试一次,而不是每个 tick 撞一次墙。
* 同一个 tip 只推一次(记 tip 而不是「推过没」),上游真又动了会再推。
* 两个接口:GET /monitor/upstream 只读缓存;POST /monitor/upstream/check 手动
  查一次且刻意不推通知——点按钮的人正看着结果。
* 默认关闭,间隔默认一天。

Dockerfile 显式装 git(python:slim 不带,而这是唯一的依赖);deploy.sh 顺带补上
一个真 bug 的提示:Dockerfile/requirements.txt 变了只 up -d 用的还是旧镜像。
2026-10-10 09:17:06 +08:00
butubb 44cbe8e2aa fix(monitor): 已登录时点「同步登录态为 Cookie」要干等 30 秒
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
顺序反了:原先先开页面、读二维码,读完才发现「已经登录、没有二维码」。而读二维码内部
会 wait_for_selector 等满 30 秒才放弃 —— 用户点一下按钮要干等半分钟,还白开一个标签页。
实测日志里就是 `Page.wait_for_selector: Timeout 30000ms exceeded`。

改成先问登录状态(那是一次接口调用,很快),已登录就直接返回,根本不碰页面。
新增测试守住这个顺序:已登录时 context.new_page 不得被调用。
2026-10-09 13:49:11 +08:00
butubb de9ff58371 feat(monitor): 扫码同时存一份 Cookie,并把登录判定换成权威判据
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
用户问「现在是不是扫码就能自动获取 cookie」—— 不是,而且这正是那个面板显得没用的根源:
它只干了半件事,扫码**只写浏览器 profile,完全不提取 cookie**(qrlogin.py 里连一行
取 cookie 的代码都没有)。于是:
- CDP 开着时任务能用(复用 profile),但 Cookie 面板始终显示「未配置」
- CDP 一关,任务立刻断,因为库里那份 cookie 从来没被填过

现在扫码把两件事一起做了:写 profile(CDP 用)+ 存一份到库(Cookie 注入用)。
两种机制同时填上,开关怎么切都不断。cookie 只在内存里从 qrlogin 传到路由,不进响应体。

同时修掉一个同类 bug:监控侧的登录判定还在用页面里的 window.__INITIAL_STATE__,
而那是**页面加载那一刻的快照** —— 浏览器本来就登录着时它是对的,但扫码是加载之后
才登录的,快照不会翻转,表现为「扫了码却一直停在二维码上」。运营模块踩过同一个坑,
当时只修了那一处。现在两边统一为:拿 cookie 问后台接口「我是谁」。顺带不再需要页面导航,
检测变轻了。

前端:已登录时按钮原先被我藏起来了,面板于是变成一块只能看、不能操作的区域 ——
用户的原话是「没用」。现在两种状态都给按钮,含义不同:未登录=取二维码,
已登录=把当前登录态同步成 Cookie。

测试:tests/test_qrlogin.py 重写(stub 从页面探针换成后台接口),新增「成功会话必须
交出 cookie」「只能取一次」两例。
2026-10-09 13:46:48 +08:00
butubb f6ddc46d62 feat(notify): 通知拆成「新作品」与「异常」两个开关,异常默认开
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
问题:cookie 过期导致任务失败,但没有任何通知。查下来不是代码问题 ——
notify_enabled 在两个任务上都是 False,而它默认就是关的,事件(run_failed)
也确实生成了,只卡在最后一道闸门。

但那个默认值是错的。代码里的理由是「一条任务列表都推到一个群会很快变吵,所以默认静默」,
这个理由对新作品成立(可能每轮都有),对失败不成立:一次登录态失效意味着这个任务事实上
已经死了,而你不会知道,直到某天发现数据停在几周前。最该被告知的就是这种情况。

现在拆开:
- notify_enabled  —— 推送新作品,可能每轮都有,默认关
- notify_failures —— 推送异常(登录失效/运行失败/没抓到数据),默认开

事件按开关过滤(build_run_message):只勾了「新作品」的任务不该因为一次失败被推消息,
反之亦然,否则拆开开关就没有意义。已有任务由 _ensure_columns 补上 notify_failures=1,
所以会自动开始收到异常推送。

列名 notify_enabled 是历史遗留(它早先是唯一的通知开关),语义已收窄为「新作品」,
用注释写明,不做列重命名 —— 那需要单独的迁移,不值为一个内部工具做。
2026-10-09 13:37:02 +08:00
butubb 2e6fa955b0 feat(monitor): 博主组头补上指标合计与本轮增量
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
折叠之后组头只有「N 篇」,等于把信息藏了起来 —— 折叠应该意味着「收起来但仍可一眼读到」,
而不是「看不见」。现在组头与数据行列对齐,每个指标列给出该博主的合计,下面再带本轮增量
(那才是监控真正要看的)。

合计口径与项目一致:某个指标在所有作品上都是 null 时,合计是 null 而不是 0 ——
「0」是真实值、「null」是不知道,合计成 0 会让「还没采到」看起来像「互动为零」。
2026-10-08 13:58:26 +08:00
butubb 0eb6ba31c5 feat(monitor): 作品栏按博主分组折叠,评论栏改为 博主/作品/评论 三级
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
先回答现状:跳转作品的按钮本来就有(作品栏每行末尾的外链图标);
评论栏原本是**两级**(作品 -> 评论),第一级是作品不是博主,所以并不是三级。

- MonitorNote 新增 creator_name(爬虫产出的 jsonl 里本来就有昵称,只是没存)。
  存的是**已脱敏**的值(张***三),与项目一贯的匿名化姿态一致 —— 爬虫刻意不落原始
  user_id(tools/user_hash.py),所以 creator_hash 是唯一稳定的分组依据。
  实测该哈希是无盐 sha256,能用任务目标的 external_id 反算配对。
- 入库时刷新 creator_name:作者改昵称是常事,只在首次写一次会一直显示旧的
- 作品栏:按 creator_hash 分组,组头可折叠(默认展开 —— 折叠的默认值不该藏数据),
  行内跳转按钮加了 title 说明
- 评论栏:一级博主、二级作品、三级评论。_note_meta_map 补上博主维度,
  评论流和作品分组都带上它
- 认不出博主的作品归到「未知博主」,不丢

顺带修一个语法错误:JSX 注释放在三元表达式分支里是非法的(那是子节点语法不是表达式),
移进 div 内。
2026-10-08 09:43:19 +08:00
butubb b11bbf771a fix(covers): 封面地址是签名过期而非防盗链,改为本地缓存
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
实测推翻了之前的诊断。同一批图:

  当天签发的地址 /202610080841/...  -> 200,带不带 Referer 都一样
  隔天的地址     /202610070837/...  -> 403,带不带 Referer 都一样

路径里那段时间戳就是签发时刻。所以这是**过期**,Referer 根本不是那个维度 ——
上一轮加 referrerPolicy 是照着错误结论改的,白改。

修法:
- 采集入库时每轮刷新 cover 地址。原先只在首次入库写一次,旧作品的地址烂在库里,
  而且再怎么重跑也修不回来
- 新增 api/monitor/covers.py:把图下载落盘。图一旦落盘就与签名无关,永远可读
- 下载放在 runner 的 Phase 5(事务已提交之后),不放 ingest —— ingest 的文档写明
  No network,往里塞网络请求会毁掉它可离线测试这一点
- 新增 GET /api/monitor/covers/{note_id} 取图。这条路由带鉴权,封面不会被匿名读走
- service 返回本地地址优先,没有缓存时才退回远程
- 每轮只补一批(60 张):一次跑几百张既慢又会给图床压力,而旧地址本来就在陆续过期,
  分摊到几轮反而更稳

顺带修正 NoteCover 的注释 —— 它写着防盗链,而那个结论已被推翻,留个错的注释比没有更糟。
2026-10-08 08:45:30 +08:00
butubb fa227600fe fix(creator): 权限开通后的状态显示 + 同步范围可选可查
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- 权限显示:实测今天 display/status 已变成 1,但 tip 是「数据正在更新中,请耐心等待」,
  作品仍是 0 条。原逻辑只在未开通时显示提示条,于是会一边写着「数据已开通」一边列不出
  作品,自相矛盾。现在只要后台有话要说就显示,并按状态区分措辞。
- 同步范围:原本写死 90 天,界面上既看不见也改不了。现在可选 7/30/90/180/365/730 天,
  与后端校验上限一致。
- 新增 creator_account.last_sync_days,记录上次实际使用的范围 —— 否则界面只能说
  「同步过了」,说不清覆盖的是哪一段。选择器也会对齐到它,避免上次同步 1 年、这次
  点一下悄悄缩回 90 天。
2026-10-08 07:52:28 +08:00
butubb bef0a4fbde fix(creator): 扫码成功却一直停在二维码上 + 运营改为左右布局
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
【扫码不完成】根因是判据本身。原先读页面里的 window.__INITIAL_STATE__,而那是
页面加载那一刻的快照:监控那边的同一探针能用,是因为那台浏览器的页面加载时就已经
登录了;而扫码是加载之后才登录的 —— SPA 内部确实登进去了,但初始快照不会翻转,
于是检测永远等不到。

改成拿 cookie 直接问创作者后台 /api/galaxy/user/info 我是谁。实测这个判据很干净:
游客也会拿到 a1(所以签名算得出来),但接口直接回 401 无登录信息;只有真正登录了
才返回 user_id。所以「有 a1」什么都证明不了,后台认了才算。
顺带按 5 秒节流 —— 前端每 2 秒问一次,没必要每次都打后台接口。

【弹窗不关】成功后不自动关闭,停在二维码上会让人以为没成功。现在显示账号卡片与原话
提示,1.6 秒后自动关闭并提供一个「完成」按钮。

【已完成结果会残留】take_cookie 取走 cookie 就拆会话,而在飞的轮询会看到 _current 为空
回报 idle,把已显示的成功能擦掉。现在把结果记在模块里重复返回,关闭弹窗时清掉 ——
否则下次打开会立刻显示上次的成功。

【布局】按用户要求改成左右两栏(左账号列表、右数据面板),与监控统一,取消二级菜单。
未选过时默认选中第一个,右栏不会一开始就是空的。

测试:tests/test_creator_login.py 新增 7 例,含「游客会话永不完成」「接口不打满每次轮询」
「临时上下文用完必须关掉」。
2026-10-07 16:40:24 +08:00
butubb 2f5852e311 fix(ui): 运营与扫码登录态没跟着平台走,且两个登录面板分不清
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
两个 bug 是同一类问题:没做平台门禁。

1) 切到抖音等平台,「运营」仍列出小红书账号。运营读的是小红书创作者后台,
   其它平台根本没有对应的后台接口。现在按平台门禁换掉整个视图并说明原因 ——
   与 MonitorDashboard / ReportView 用 isWired 的做法一致。

2) 扫码登录面板把小红书的会话当成任意平台的登录态报。它的检测读的是小红书页面的
   __INITIAL_STATE__,后端 qrlogin.LOGIN_URL 里也只有 xhs 一项;而标题写的是
   「{当前平台}登录态」,于是切到抖音照样显示已登录—— 这是个具体的谎。
   现在非小红书直接换掉整个面板(只改标题不够,下面的块读的仍是小红书的状态)。

3) 两个面板的标题都是「登录态」,看不出区别。它们其实是不同机制:
   - Cookie 面板:存进库、每轮以 --cookies_file 注入子进程 → 改名为「Cookie(定时任务用)」
   - 扫码面板:写进浏览器 profile、CDP 模式复用          → 改名为「浏览器登录态(扫码)」
2026-10-07 16:35:12 +08:00
butubb 5484e5a3ae fix(creator): 详情接口用了 MySQL 5.7 不支持的 NULLS LAST
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
1064 语法错误。nullslast() 是 PostgreSQL 语法,MySQL 5.7 不认;而 MySQL 把 NULL 视为比任何值都小,
所以 DESC 本身就把它排在最后,不需要额外声明。

这个 bug 只在真机上暴露:SQLite 从 3.30 起支持 NULLS LAST,本地测试环境测不出来。
2026-10-07 16:32:53 +08:00
butubb c2b310c7bf feat(creator): 新增「运营」模块 —— 多账号扫码登录与创作者后台数据
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
侧边栏在「监控」右边加了「运营」:账号列表 → 点进二级详情看该账号的数据。

【为什么是独立模块而不是监控的子视图】两者形状不同:监控是公开数据(点赞/收藏/评论/分享)的每轮快照+差分;运营是创作者后台按日期给出的曝光/观看/完播率/涨粉。凭据不同、采集方式也不同 —— 那边要浏览器登录态,这边是纯请求。硬塞进同一个模型会同时污染两边。

【扫码登录的关键差异】监控的扫码把登录态写进浏览器默认 profile(爬虫要复用)。运营要的是 cookie 字符串(纯请求够用),所以每次登录开一个**临时上下文**,扫完取出 cookie 就丢弃 —— 登第二个账号不会把第一个顶掉,也不影响监控那个登录态,十个账号互不干扰。

【决策依据】tools/probe_creator_api.py 的 Phase 0 实测:签名可自造(XYW_:MD5 → base64 → AES-128-CBC,与 xhshow 内置实现常量逐字节一致);主站 cookie 即可认证创作者后台;接口与参数已与真实页面对齐。

后端:
- api/creator/models.py: creator_account / creator_note_stat。**复用 MonitorBase**,这样 create_all 与上一轮改成元数据驱动的 _ensure_columns 会自动覆盖新表
- api/creator/signing.py: XYW_ 签名,带三条实测结论(url= 前缀、appId=ugc、401 与 406 的区别)
- api/creator/client.py: 纯 httpx 客户端。字段名尚未亲眼验证过,所以写成多别名匹配;解析不出来存 None 而非 0
- api/creator/service.py: 账号 CRUD 与同步。cookie 绝不进入对外结构,只给 has_cookie
- api/creator/login.py: 临时上下文的扫码登录
- api/routers/creator.py: 8 条路由,全部带鉴权

前端:
- 侧边栏「运营」+ OperationView(账号列表 → 二级详情)+ AddAccountDialog
- 权限状态显眼呈现:pending 时照抄后台原话「已为您申请数据权限,次日可查看」,并说明此时同步返回 0 条是正常的,不是采集失败

测试:tests/test_creator_client.py 新增 48 例,含「cookie 不得出现在对外结构里」这条不变量,以及权限未生效时空壳响应的处理。
2026-10-07 16:30:45 +08:00
butubb 2613f7577f feat(creator): Phase 0 探针 —— 创作者后台数据可以纯请求拿到
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
结论:签名可自造、主站 cookie 即可认证、接口与参数已与真实页面对齐。不需要浏览器、不需要独立的创作者登录。

- tools/probe_creator_api.py: 纯 HTTP 探针。用 XYW_ 算法自签(MD5 → base64 → AES-128-CBC,
  密钥与 IV 与 xhshow/config/config.py 逐字节一致),对 note/analyze/list 发请求
- tools/probe_creator_page.py: 打开真实数据分析页,记录页面自己发的请求,作为地面真相

Phase 0 的三条实测结论:
1. 签名可伪造。三种写法里只有「url= + 路径 + 查询串」被接受(200);仅路径、或裸路径都 406。
   并且不带 cookie 时返回的是应用层 401「无登录信息」而非网关 406 —— 说明签名每次都已通过
2. 主站 .xiaohongshu.com 的 cookie 就能认证创作者后台,不需要单独的创作者会话
3. 当前账号 dfg 返回空数据不是技术问题:permission/query 的 tip_msg 是
   「已为您申请数据权限,次日可查看」,display/status 均为 0,即权限尚未生效

关键佐证:真实页面调 note/analyze/list 用的查询串与本探针生成的完全一致,且拿到同一份空响应。
2026-10-07 16:22:15 +08:00
butubb 2b9ebdad87 fix(ui): 「每轮最多采集作品数」标签是错的——它是每个博主的上限
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
爬虫里这个值是在 per-creator 的函数内比较的(client.py get_all_notes_by_creator 的 result 是局部变量),而外层 for 循环遍历全部目标。所以 100 个目标 × 20 篇 = 单轮最多 2000 篇,一篇都不会被丢弃。标签写成「每轮最多」会让人以为超出的会被截掉。

- creator 模式:标签改为「每个博主最多采集作品数」,并实时算出「N 个目标 × M 篇 → 单轮最多 X 篇」
- note 模式:禁用该输入并说明「此项不生效」——get_specified_notes 里没有任何 CRAWLER_MAX_NOTES_COUNT 引用,列出的每个链接都会被逐条抓
- 单轮估算超过 500 篇时给出警告:每篇还要抓最多 max_comments_count 条评论、并发为 1,容易触发限流,也可能跑不完就被默认 1 小时的任务超时中断
2026-10-07 15:46:11 +08:00
butubb 6eff6fcc83 fix: 扫码登录状态可独立查询 + 趋势图改回自适应纵轴
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
【登录反馈】实测这台浏览器 loggedIn=true,其实早就登录成功了;看不到反馈是判据和设计的问题:

1. 判据不可靠。原先靠 web_session 的值变化判断——对照组显示:一个全新的空 profile 首次访问小红书就会被发一个 web_session,所以「有这个 cookie」什么都证明不了。可信信号是页面自己的 __INITIAL_STATE__.user.loggedIn,但它是 Vue 响应式引用,必须 .value 解包(这就是前面探针读到 [object Object] 和 None 的原因)。
2. 状态绑死在临时会话上。扫码会话是内存状态,进程一重启就没(部署、崩溃都算),面板于是悄悄退回初始态——一次成功扫码看起来像什么都没发生。

改法不是让会话活得久,而是把「登没登录」变成随时可查、与会话无关:
- 新增 GET /api/monitor/login/state,直接问浏览器要答案,带 5 秒缓存;force=true 先重载页面再读,用于状态陈旧
- qrlogin 改为常驻 Playwright 客户端 + 复用同一个标签页,并在重启后认领浏览器里已存在的 xhs 标签页,避免堆孤儿页
- 把「读不到状态」与「未登录」分开——前者显示具体错误,不再悄悄显示成未登录
- 面板顶部常驻显示登录态与昵称,带「重新检测」按钮;扫码成功后自动翻转

【趋势图】上一轮改过头了。dataviz 规范里没有「折线图必须从 0 起」这条——基线相关的条文全是讲柱状图的(柱状图用长度编码数值,不从 0 起比例就是错的;折线图用位置编码,轴只需如实框住数据)。改回自适应,保留上一轮修好的左侧刻度栏让范围始终可见;步长收敛到 1/2/5×10ⁿ,全平序列撑开一档避免除零。
2026-10-07 15:42:58 +08:00
butubb e608b51210 test: 修正列渲染断言——SQLAlchemy 会对保留字加反引号
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
monitor_run.trigger 是 MySQL 保留字,渲染出来是 `trigger`。这比之前手写的列清单更正确:
清单里的裸 trigger 会直接语法错误。断言改为剥掉引号后比对列名。
2026-10-07 15:37:08 +08:00
butubb d937ff5fe6 fix(db): _ensure_columns 改为按模型元数据推导,并补上漏加的调度字段
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
上一个提交加了 4 个调度字段,却没在 _ADDED_COLUMNS 里登记,后果是生产环境:
  - 任务列表接口报 Unknown column,UI 打不开任务列表
  - 调度器每 20 秒 tick 一次炸一次,定时任务完全不会触发
  最阴险的是启动完全正常——能连库、能起来,只是随后每条查询都失败。

- _ensure_columns 不再遍历手写清单,改为遍历 MonitorBase.metadata.sorted_tables,
  从根上消掉「加了字段忘了登记」这类漏
- 新增 _column_ddl:用 CreateColumn 渲染类型与可空性,并给 NOT NULL 列补 DEFAULT。
  模型的 default= 是 ORM 侧行为、不会进 DDL,而给已有数据的表加 NOT NULL 列必须有值,
  否则能否成功取决于服务端 sql_mode
- 主键列跳过:MySQL 不允许 AUTO_INCREMENT 与 DEFAULT 共存

新增 tests/test_monitor_column_migration.py 守住:每个 NOT NULL 列都必须能生成带
DEFAULT 的合法 ALTER。
2026-10-07 15:35:40 +08:00
butubb 8f4e5586e9 feat(schedule): 任务支持「每天定时 / 每周定时」,用选择器而不是手写 cron
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- schedule.py: 新增调度计算模块(纯函数,便于单测)。三种模式:interval(每 N 分钟)/ daily(选钟点)/ weekly(选星期 + 钟点)
- 钟点模式是「固定时刻」而非「固定延迟」——从日历重算,所以某轮跑晚了不会把之后每一轮都拖晚。interval 保持原语义:从上一轮开始计时
- 抖动只加给 interval。给「每天 9:00」也加抖动就成了 9:00–9:01 随机触发,操作者选的时间被悄悄改掉,只会像 bug
- 模型 / schema / service: 新增 schedule_mode / schedule_hours / schedule_days / schedule_minute。时钟字段存逗号分隔文本——几个小整数、永远整体读写,单开一张表只会换来 join。interval_minutes 保留且仍是默认值,已有任务不受影响
- service: 改动任何调度字段都按合并后的状态重算 next_run_at。重新启用也算改动,否则停用一个月再打开会带着一个月前的 next_run_at,一保存就立即触发
- 前端: 运行方式三选一 + 小时/星期胶囊多选 + 分钟下拉,并实时预览结果句子。任务卡片改显示后端拼好的 schedule_label,避免列表和编辑器对同一计划给出两种说法
- 校验: 钟点模式至少选一个时间,按周至少选一个星期

tests/test_schedule.py 新增 24 个用例,含「恰好等于当前时刻的档位归属下一天」这个会让调度器自循环的边界。
2026-10-07 15:33:17 +08:00
butubb 242e3a7837 fix(chart): 纵轴自 0 起、刻度归位、数据点恢复为正圆
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- 纵轴从 0 起,上界取整到好读的数(1/1.2/1.5/2/…)。此前按 [最小值,最大值] 自适应,491→506 这点变化被撑满整个图高、看着像暴涨,且上下两刻度就是 506 和 491 两个几乎一样的数,没有 0 做参照读不出量级
- 刻度线画在 0 / 中值 / 上界,标签贴在各自主线上、放进左侧刻度栏。原来是两个绝对定位的数字浮在图面上,会挤在一起读成一个数
- 数据点恢复为正圆:根因是 preserveAspectRatio="none" 把 600×160 的 viewBox 横向拉满容器,横纵缩放不一致,半径 4 的圆被压成椭圆。改为用 ResizeObserver 量出容器宽度、等比绘图
- 每个数据点都画出来(原先只画末点),点数多时自动缩小半径;悬停热区按点位间距铺满整段,不再固定 12px
- 当前值移到标题行,不再浮在图面上遮挡曲线
2026-10-07 15:26:42 +08:00
butubb c1068845a9 fix(deploy): 部署时必须强制重建容器
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
代码是 bind mount,容器配置与镜像都没变,所以裸的 `docker compose up -d` 会判定无需变更直接跳过,改了 .py 也不会生效。前端产物是磁盘上的静态文件、能即时生效,这一点很容易把问题盖住,直到有人改了后端代码才发现。

首次实测即命中:跑 deploy.sh 的输出是 'Container mediacrawler Running',没有重启。
2026-10-07 15:22:00 +08:00
butubb 89b7b1e825 fix: 封面图不再被图床拒绝,作品栏补上封面
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- NoteCover: 新增共用封面组件,核心是 referrerPolicy="no-referrer"。小红书图床对带外部 Referer 的请求一律 403,而浏览器对跨域 <img> 默认就会带上本站源作为 Referer —— 于是封面全显示成破图,而 URL 本身完全正常。实测同一张图:无 Referer 200 / 67968B,Referer 为本站 403 / 0B
- CommentsFeed: 改用 NoteCover,修掉评论栏封面全部加载失败
- NotesTable: 作品栏此前完全没有渲染封面,补上缩略图;封面缺失时用占位块,避免行高随封面陆续到达而跳动

放在一个组件里而不是就地加属性,是为了让下一个显示封面的页面不会漏掉。
2026-10-07 15:21:29 +08:00
butubb fe45442b01 chore: deploy.sh 标记为可执行(Windows 上创建的文件没有 exec 位)
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
2026-10-07 11:14:13 +08:00
butubb 74a592024c feat: 部署改为 git 驱动,容器以宿主用户运行
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
服务器实测可经 Cloudflare 443 访问 Gitea 且 git 协议正常(此前我只测了 13000 端口就断言不可达,是错的),因此不再需要 tar + SFTP。

- docker-compose: 增加 user: "1000:1000"。容器此前以 root 运行,写进挂载目录的每轮 jsonl 产物都是 root 属主,导致宿主用户连自己的部署目录都挪不动 —— 这在把部署迁到 /mnt/data 时实际发生了
- deploy.sh: 一条命令走完 拉代码 →(webui/ 有改动时)重建前端 → 重启容器。前端产物 api/webui 是 gitignore 的,git pull 带不过来,必须在服务器上重建一次
- Dockerfile: 补 npm 包。corepack 只管 yarn/pnpm 不管 npm,而前端要在服务器上重建;这样服务器只需要 Docker,不必另配 Node 环境
2026-10-07 11:13:43 +08:00
butubb cffb407d15 feat: 生产库克隆脚本 + 容器补 Node 运行时
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- tools/clone_database.py: 同服务器跨库克隆并建立独立应用账号的 provisioning 脚本。本机与服务器都没有 mysql 客户端,所以走 INSERT...SELECT 而非 dump/reload。刻意只读写命令行指定的两个 schema,且给应用单开账号而不是复用管理员
- Dockerfile: 补 nodejs。douyin/help.py 在模块导入阶段就执行 execjs.compile(libs/douyin.js),缺 JS 运行时会抛 RuntimeUnavailableError;而 main.py 要导入全部 7 个平台,于是整个应用连带环境自检一起挂掉

已在 192.168.20.220 上验证:7 个平台导入全部通过。
2026-10-07 11:08:11 +08:00
butubb bcc7361145 refactor: 镜像只装依赖,代码改为挂载
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- Dockerfile: 去掉 COPY . .,镜像退化为纯依赖层,改代码不再需要重建镜像
- Dockerfile: 补 libgl1 / libxcb1 / libglib2.0-0 等 X11 库。opencv-python 在导入时需要它们,而 tools/utils.py 会经 slider_util 导入 cv2 —— 缺了不是某个边角功能挂掉,是整个应用起不来(已在真机冒烟中验证)
- Dockerfile: 补 tzdata,否则 TZ=Asia/Shanghai 被静默忽略,所有时间戳落成 UTC
- Dockerfile: apt 与 pip 均改走国内源。deb.debian.org 实测约 13 kB/s,96 MB 构建依赖要跑半小时以上
- docker-compose: 挂载 ./ 到 /app,部署流程从「重建镜像」变成「重传 + 重启」
- main.py: 启动时若缺 api/webui/index.html 就打印警告,避免只返回一段 JSON 却看着一切正常
2026-10-07 10:57:44 +08:00
butubb 37ca1b1cd6 feat: CDP 接管开关 + 扫码登录面板 + Docker 部署
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
- runner: enable_cdp_mode 从硬编码 False 改为系统设置 cdp_enabled。服务器部署下爬虫接管已开启远程调试的 Chrome(默认 9222),复用其 profile 登录态;本机桌面默认仍为关,行为不变
- qrlogin: 新增 CDP 扫码登录。Chrome 在服务器上跑于 Xvfb,show_qrcode 依赖的 PIL 桌面看图程序不存在,二维码无处可显示;改为经 CDP 从页面取出二维码交给 WebUI 渲染。刻意复用 browser.contexts[0](新建 context 是无痕 profile,扫了也白扫),且绝不调用 browser.close()(会连带关掉操作者自己的 Chrome)
- webui: 设置页新增扫码面板,替换原本跳到采集页看终端二维码的入口
- Dockerfile / .dockerignore / docker-compose.yml: 服务器部署。host 网络是必需而非图省事——容器里 127.0.0.1:9222 必须落到宿主机回环
- UPSTREAM.md: 补充 gitcode 镜像,用于 GitHub 大包传输必断时补历史
2026-10-07 10:41:11 +08:00
94 changed files with 13472 additions and 470 deletions
+29
View File
@@ -0,0 +1,29 @@
# Keep the build context to what the server actually runs.
.git
.github
.venv
venv
# The WebUI sources are not needed -- only the bundle they produce, which lands
# in api/webui and is therefore NOT excluded.
webui/node_modules
webui/src
webui/dist
# Runtime state: per-run crawler output and the login browser profile. Mounted
# as a volume instead, so it survives image rebuilds.
data
browser_data
# Secrets come from the environment via compose, never baked into a layer.
.env
tests
docs
__pycache__
**/__pycache__
*.pyc
*.pyo
.pytest_cache
*.db
*.log
+12 -1
View File
@@ -183,4 +183,15 @@ agent_zone
debug_tools
database/*.db
.omx/
.omx/
# 别人放在这儿的参考项目(mac-agent-os)。它是独立仓库、21MB,不属于本项目 ——
# 一旦被 `git add -A` 扫进来就是永久留在历史里(踩过一次:1601 个文件里 1429 个是它)。
# 要读它就直接读磁盘上的目录,别提交。
mac-agent-os-main/
# 服务器上重建前端时,容器里的 npm 往挂载目录写的缓存。是构建产物,不该进仓库 ——
# 它长期以「未跟踪」状态躺在工作区,正是会被 `git add -A` 顺手带走的类型。
webui/.npm/
# 我用来在服务器上跑命令的一次性脚本(不属于这个项目)
.remote_run.py
+71
View File
@@ -0,0 +1,71 @@
# Server deployment image.
#
# No browser is bundled on purpose. On this deployment the crawler attaches over
# CDP to the Chrome already running on the host (see the 接管已有 Chrome setting),
# so shipping a second copy of Chromium would only add hundreds of megabytes and
# a login state that nothing uses. The Playwright Python package is still needed
# -- that is what speaks CDP -- hence PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD.
FROM python:3.11-slim
ENV PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
PIP_DISABLE_PIP_VERSION_CHECK=1 \
PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1 \
TZ=Asia/Shanghai
# asyncmy compiles a Cython extension, so a toolchain has to exist at build time.
# It is left installed: purging it risks taking libmysqlclient with it, and a
# slightly larger image is cheaper than a runtime that fails months later.
#
# deb.debian.org is effectively unusable from this network -- it was pulling the
# 96 MB of build dependencies at roughly 13 kB/s, which puts a build at well over
# half an hour. Point apt at a domestic mirror; override APT_MIRROR when building
# from somewhere that does not need it.
#
# The libgl1/libxcb1/... group is not for a GUI: opencv-python links against X11
# at import time, and tools/utils.py reaches cv2 through slider_util, so without
# them the *application* fails to import, not just some image utility. tzdata is
# here because TZ=Asia/Shanghai is silently ignored without it, which would put
# every stored timestamp in UTC. nodejs is for PyExecJS: douyin/help.py compiles
# libs/douyin.js at *import* time, and because main.py imports every platform,
# that single platform being importable-or-not decides whether the whole app
# (and the environment self-check) comes up. npm rides along so the WebUI can be
# rebuilt on the server (see deploy.sh) instead of only on a workstation --
# corepack is present but does not cover npm, only yarn and pnpm.
#
# The pip mirror is set for the same reason as the apt one: this host's route to
# the public index is slow.
#
# git is for the 上游更新检查 (api/monitor/upstream.py): it fetches the upstream
# repository into the mounted checkout to count how far behind this fork is.
# python:slim does not ship git, and nothing else here pulls it in.
ARG APT_MIRROR=mirrors.tuna.tsinghua.edu.cn
RUN set -eux; \
for f in /etc/apt/sources.list /etc/apt/sources.list.d/debian.sources; do \
if [ -f "$f" ]; then \
sed -i "s|deb.debian.org|${APT_MIRROR}|g; s|security.debian.org|${APT_MIRROR}|g" "$f"; \
fi; \
done; \
apt-get update; \
apt-get install -y --no-install-recommends \
build-essential pkg-config default-libmysqlclient-dev git \
libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 libxcb1 libgomp1 \
tzdata nodejs npm; \
rm -rf /var/lib/apt/lists/*
WORKDIR /app
# Requirements only -- this is the one layer that is expensive to build and
# changes rarely.
ARG PIP_INDEX=https://pypi.tuna.tsinghua.edu.cn/simple
COPY requirements.txt ./
RUN pip install --no-cache-dir -i "$PIP_INDEX" -r requirements.txt
# The application code is deliberately NOT copied in. compose mounts it at /app,
# so a code change is "re-upload the tarball, restart the container" instead of
# an image rebuild. Treat this image as the dependency layer and nothing else;
# rebuild it when, and only when, requirements.txt or this file changes.
EXPOSE 18051
# api.main reads MC_HOST / MC_PORT from the environment; compose supplies both.
CMD ["python", "-m", "api.main"]
+874
View File
@@ -0,0 +1,874 @@
# 内容抓取模块 · 开发技术说明
> 面向接手开发的团队 · 2026-10-10
> 全部内容基于**逐行读源码**整理,不是推测
> 范围:**只写内容抓取模块**,不涉及其他业务
---
## 一、这个模块是干什么的
从主流内容平台(抖音 / 小红书 / B站 / 知乎 / 任意网页)**采集内容数据**:
- 视频/笔记的元数据(标题、作者、发布时间、正文)
- 互动数据(点赞、评论、分享、播放)
- 评论列表
- 作者主页的全部作品列表
采到的数据落进本地 SQLite,供后续分析/运营使用。
---
## 二、整体架构(重要:三层降级是核心)
```
┌──────────────────────────────────────────────────────────┐
│ HTTP 层 routes/scrape.py(22 个端点) │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 引擎层 services/scrape_engine.py │
│ · URL 解析 → 标准化目标 │
│ · 按平台选适配器 │
│ · 同步 / 异步调度 │
│ · 结果落库(去重) │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 适配器层 services/adapters/*.py(每个平台一个) │
│ 基类 ScrapeAdapter 定义统一接口 + 工具降级 │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 工具层(★ 三层降级,这是本模块的核心设计) │
│ Level 1 OpenCLI 外部 Node CLI(主路径) │
│ Level 2 agent-browser Playwright + 真实 Chrome │
│ Level 3 web_crawler 通用网页兜底 │
└───────────────────────┬──────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ 存储层 services/scrape_db.py(SQLite,4 张表) │
└──────────────────────────────────────────────────────────┘
```
### 为什么这么设计
抓取的最大风险是**单一方式失效**:目标平台改版、接口封禁、登录态过期。
所以**同一份数据有三条获取路径**,第一条失败自动降级到第二条,
**上层完全不感知**(对 engine 来说只是"拿到数据了")。
---
## 三、工具层详解(最关键的一层)
### 3.1 Level 1 — OpenCLI
**它是什么**:一个**第三方 Node.js CLI 工具**,包名 `@jackwener/opencli`。
```
实际安装位置(本机实测):
~/.workbuddy/binaries/node/versions/22.22.2/bin/opencli
→ 软链到 ../lib/node_modules/@jackwener/opencli/dist/src/main.js
```
**怎么调用**(`services/adapters/__init__.py` 的 `_run_opencli`):
```python
OPENCLI = os.environ.get("OPENCLI_PATH",
str(Path.home() / ".workbuddy" / "binaries" / "node" /
"versions" / "22.22.2" / "bin" / "opencli"))
async def _run_opencli(self, args: list, timeout: int = 60):
cmd = [self.OPENCLI] + args
proc = await asyncio.create_subprocess_exec(
*cmd, stdout=PIPE, stderr=PIPE)
stdout, stderr = await asyncio.wait_for(proc.communicate(), timeout=timeout)
...
return self._parse_output(stdout.decode().strip())
```
**调用示例**(抖音,`douyin_scrape.py`):
```bash
opencli douyin user-videos <sec_uid> --limit 20 --with_comments true -f json
opencli douyin stats <aweme_id> -f json
```
**⛔ 移植注意**:这个二进制**不在仓库里**,是外部依赖。移植时必须:
- 要么在目标机装 `npm i -g @jackwener/opencli`
- 要么改 `OPENCLI_PATH` 环境变量指向它的位置
### 3.2 Level 2 — agent-browser("套用真实浏览器"的做法)
**这是你问的重点。设计原则写在 `browser_helpers.py` 文件头**:
```
⛔ 绝不使用 Camoufox(养号专用,Firefox 内核 + 特殊指纹)
✅ 使用 Playwright 启动【真实 Chrome】(Chromium 内核,正常指纹)
```
**具体怎么"套真实浏览器"**(`browser_helpers.py` 的 `_get_browser`):
```python
_CHROME_PATHS = [
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome", # ← 系统真 Chrome
"/Applications/Chromium.app/Contents/MacOS/Chromium",
]
_CHROME_PATH = None
for p in _CHROME_PATHS:
if Path(p).exists():
_CHROME_PATH = p # 自动探测,找到就用系统已装的 Chrome
break
launch_kwargs = {
"headless": headless,
"args": [
"--disable-blink-features=AutomationControlled", # ★ 反检测关键
"--no-sandbox",
"--disable-dev-shm-usage",
"--disable-gpu",
"--window-size=1280,720",
],
}
if _CHROME_PATH:
launch_kwargs["executable_path"] = _CHROME_PATH # ★ 用系统 Chrome,不用 Playwright 自带
```
**四个关键设计点**:
| 点 | 做法 | 为什么 |
|---|---|---|
| **用什么内核** | 系统真实 Chrome(`executable_path` 指定) | Playwright 自带 Chromium 有明显特征;真实 Chrome 是正常用户指纹 |
| **怎么隐藏自动化** | `--disable-blink-features=AutomationControlled` | 这是最常被检测的自动化标志位 |
| **实例管理** | 模块级单例 `_browser` + `asyncio.Lock` | 避免每次请求都启动浏览器(启动 ~1-2 秒) |
| **会话隔离** | 每次 `browser.new_context()` | 每个任务独立 cookie 环境,互不污染 |
**页面加载策略**(`page_evaluate`):
```python
await page.goto(url, wait_until="domcontentloaded", timeout=timeout)
await page.wait_for_load_state("networkidle", timeout=timeout) # 等动态渲染
await asyncio.sleep(1) # 再等 1 秒保险
result = await page.evaluate(js_code) # 执行 JS 提取
```
**UA 伪装**(每次 context 都设置):
```python
context = await browser.new_context(
user_agent=("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/125.0.0.0 Safari/537.36"),
viewport={"width": 1280, "height": 720},
locale="zh-CN", # 中文环境,符合目标用户画像
)
```
**两个公开函数**:
```python
page_evaluate(url, js_code, timeout, headless) # 打开页面执行 JS,返回 dict
page_extract(url, selectors, timeout, headless) # 按 CSS 选择器提取文本
```
**Profile 持久化**(可选,环境变量控制):
```python
_USER_DATA_DIR = os.environ.get(
"SCRAPE_CHROME_USER_DATA",
str(Path.home() / "workbuddy-agent-os" / "agent-local" /
"runtime" / "scrape_chrome_profile"))
```
→ 想保持登录态就指定这个目录;不指定就用临时 context。
### 3.3 Level 3 — web_crawler
通用网页抓取兜底(`web_scrape.py`),用于非主流平台的页面。
### 3.4 降级怎么触发(`ScrapeAdapter._try_tools`)
```python
async def _try_tools(self, tool_level: int, funcs: list) -> tuple:
"""funcs: [(工具名, 可调用对象), ...],按顺序尝试"""
tools = [f for f in funcs[:tool_level]] # 按 level 截断
for name, fn in tools:
try:
result = await fn()
if result is not None:
return True, result, name # ★ 成功即返回,不再降级
except Exception as e:
logger.warning(f" ⚠️ [{self.platform}] 工具 {name} 失败: {e}")
return False, None, tools[-1][0] if tools else "none"
```
**关键语义**:
- `tool_level=1` → 只用 OpenCLI
- `tool_level=2` → OpenCLI → agent-browser(**默认**)
- `tool_level=3` → 三层全开
- **第一个成功就停**;返回 `(成功?, 结果, 用了哪个工具)`
---
## 四、适配器层(每平台一个)
### 4.1 统一接口(`services/adapters/__init__.py` 的基类)
**每个平台适配器必须实现 4 个方法**:
```python
class ScrapeAdapter:
platform = "" # 子类覆写,如 "douyin"
adapter_name = ""
async def collect_item(self, target, depth="light", tool_level=2) -> dict:
"""抓单条内容详情"""
async def collect_user(self, user_id, limit=20) -> list[dict]:
"""抓某个用户/作者的全部作品"""
async def collect_comments(self, item_id, limit=20) -> list[dict]:
"""抓评论"""
async def collect_search(self, keyword, limit=20) -> list[dict]:
"""按关键词搜索"""
```
**基类提供的公共能力**:
- `_try_tools(tool_level, funcs)` — 降级执行(见 3.4)
- `_run_opencli(args, timeout)` — 调 OpenCLI + 解析输出
- `_parse_output(text)` — **输出格式三层兜底解析**(见下)
- `_parse_lines(text)` — 纯文本兜底解析
### 4.2 输出格式三层兜底(细节,容易踩坑)
OpenCLI 的输出格式**不保证稳定**,所以解析做了三层:
```python
def _parse_output(self, text: str):
# 1. JSON 检测(以 [ 或 { 开头)→ json.loads
if text.startswith("[") or text.startswith("{"):
try:
return json.loads(text)
except json.JSONDecodeError:
logger.warning("JSON 解析失败,尝试 YAML 兜底")
# 2. YAML 解析(PyYAML 可用时)
try:
import yaml
parsed = yaml.safe_load(text)
if parsed is not None:
return parsed
except ImportError:
logger.debug("PyYAML 未安装,跳过 YAML")
# 3. 纯文本兜底:按行解析 key:value
return self._parse_lines(text)
```
`_parse_lines` 甚至**专门处理了 `top_comments` 块**(评论在纯文本里的多行结构)。
**⛔ 移植注意**:如果目标环境没有 PyYAML,会静默降级到第三层
(能跑,但嵌套结构会丢)。
### 4.3 各平台降级链(实测,**差异很大**)
```python
# 抖音(douyin_scrape.py)—— 两条路径
await self._try_tools(tool_level, [
("opencli", lambda: self._opencli_user_videos(...)),
("agent-browser", lambda: self._browser_user_profile(...)), # ← 有浏览器降级
])
# 小红书 / B站 / 知乎 —— ⚠️ 只有一条路径
await self._try_tools(2, [
("opencli", lambda: self._opencli_user_notes(...)), # ← 没有降级!
])
# 通用网页(web_scrape.py)—— 不用 OpenCLI
await self._try_tools(tool_level, [
("web_crawler", lambda: self._web_crawl(target)),
("agent-browser", lambda: self._browser_extract(target)),
])
```
**对照表**:
| 平台 | 降级链 | 有浏览器降级? | 备注 |
|---|---|---|---|
| **抖音** | `opencli` → `agent-browser` | ✅ | 唯一做了完整降级的平台 |
| **小红书** | `opencli`(单条) | ❌ | OpenCLI 挂了就抓不了 |
| **B站** | `opencli`(单条) | ❌ | 同上 |
| **知乎** | `opencli`(单条) | ❌ | 同上 |
| **通用网页** | `web_crawler` → `agent-browser` | ✅ | 走另一套(不用 OpenCLI) |
**⚠️ 这是模块的真实局限**:三个平台**只有一条路径**,没有降级能力。
接手时如果要提升健壮性,**最值得做的就是给它们补上浏览器降级**
(照抖音的 `_browser_*` 方法写即可)。
**另一个细节**:`collect_user` 传的是**硬编码的 `2`**(不是 `tool_level` 参数),
而 `collect_item` 才用传入的 `tool_level`:
```python
async def collect_user(self, user_id, limit=20): # 没有 tool_level 参数
await self._try_tools(2, [...]) # ← 写死 2
async def collect_item(self, target, depth, tool_level=2):
await self._try_tools(tool_level, [...]) # ← 用参数
```
→ **用户在前端设 `tool_level=1` 时,"抓用户主页"这条路仍会走到浏览器**。
### 4.35 登录态机制(★ 这是"如何用真实浏览器"的另一半)
**问题**:抖音的数据接口需要登录态(cookie),怎么拿到?
**答案**:`mediacrawler_adapter.py` 的做法 —— **用户在真实 Chrome 登录,程序通过 CDP 读解密后的 cookie**。
#### 为什么不能直接读 cookie 文件
代码注释原文(第 29 行):
```
# Chrome 新版把 cookie 值加密存在 SQLite 里,必须通过 CDP 读解密后的值。
```
Chrome v80+ 把 cookie **加密**存在 SQLite(`Cookies` 文件),
直接读文件拿到的是**密文**。必须让 **Chrome 自己解密** → 通过 **CDP 协议**问它。
#### 三步实现
**① 通过 CDP 读 cookie**(`_get_cookies`,第 68 行)
```python
async def _get_cookies() -> dict:
"""从 CDP 连接读取 Chrome cookie(解密后的值)"""
ctx = await _ensure_cdp() # 确保 CDP 连接
all_cookies = await ctx.cookies() # ← 让 Chrome 解密并返回
...
```
**② 转成 HTTP Header**(`_cookie_str`,第 86 行)
```python
def _cookie_str(cookies: dict) -> str:
return "; ".join(f"{k}={v}" for k, v in cookies.items())
```
**③ 带 cookie 直调抖音 API**(`_http_get`,第 93 行)
```python
def _http_get(url: str, cookies: dict, timeout: int = 15) -> dict:
headers = {
"User-Agent": "...Chrome/150.0.0.0 Safari/537.36",
"Cookie": _cookie_str(cookies), # ★
"Referer": "https://www.douyin.com/",
"Origin": "https://www.douyin.com",
}
...
```
#### 登录态判断
```python
cookies = await _get_cookies()
has_session = bool(cookies.get("sessionid")) # ← 有 sessionid 就算已登录
```
对应 HTTP 端点 `GET /api/scrape/check-login`。
#### 让用户登录(巧妙的做法)
代码注释(第 433 行):
```python
# ── 打开登录页(用 AppleScript 控制 Chrome,不需要 CDP) ──
"""在 Chrome 中打开抖音首页,让用户登录
登录后 cookie 自动保存到 Chrome profile,两个 Chrome 都会检测。"""
```
**→ 用 AppleScript 打开真实 Chrome**(不是 Playwright 控制的),
用户**在真实浏览器里手动登录** → cookie 存进 Chrome profile →
之后程序通过 CDP 读。
**这样最自然**:用户看到的是他熟悉的 Chrome,扫码登录,
✅ 不会被平台识别为"自动化登录"。
**⛔ 移植注意**:`open_login_page()` 用的是 **AppleScript**(`osascript`),
**macOS 专有**。Linux/Windows 要改实现(可以用 `open` 命令或直接 Playwright 打开)。
#### 一个隐藏技巧(避免暴露自动化)
代码注释(第 227-229 行):
```python
# 获取热评(复用已有 Chrome 页面,不创建新标签页)
# Chrome 已有 douyin.com 页面,直接用它的 JS 上下文执行 fetch
# ⚠️ 不要 new_page() — 那会在 Chrome 中闪出新标签页
```
**→ 复用用户已经打开的页面**执行 fetch,而不是新开标签页
(新标签页会闪一下,且更易被识别)。
---
### 4.4 抖音的特有实现(两份代码,别搞混)
**⚠️ 项目里有**两个**抖音相关模块**,职责不同:
**① `services/adapters/douyin_scrape.py`(181 行)**
- 走**工具降级**(OpenCLI → 浏览器)
- 浏览器路径用 JS 从页面 **DOM 文本**提取数据:
```javascript
// _browser_video_page 的 JS(正则从页面文本抓「获赞/粉丝/关注」)
const body = document.body.innerText || '';
const uidM = body.match(/抖音号[::]\s*(\S+)/);
const nickM = body.match(/@(\S+)/);
function extractNum(label) {
var m = body.match(new RegExp('(\\d+(?:\\.\\d+)?[万w]?)\\s*' + label));
...
}
return { aweme_id, title, author_nickname, douyin_id, digg_count, fans, following };
```
**② `services/mediacrawler_adapter.py`(457 行)**
- **不走工具降级**,完全独立的实现
- 文件头注释(原文):
> 全新架构:**不再依赖 CDP/Playwright/浏览器页面**。
> 直接从 Chrome profile 读取 cookie,通过 HTTP 请求调用抖音 API。
> 全程无窗口、无标签页、无闪烁。
- 用于**追踪视频/作者**这类需要高频刷新的场景(详见子代理报告)
**⛔ 关键澄清**:`mediacrawler_adapter.py` **虽然叫 mediacrawler,但不依赖
MediaCrawler 这个开源项目**——它只是借用名字,实际是"读 Chrome cookie + HTTP 调 API"。
---
## 五、引擎层(`services/scrape_engine.py`,361 行)
### 5.1 URL 解析(`resolve_target`,第 56 行)
**7 类目标自动识别**(实测代码):
```python
抖音短链 v.douyin.com/xxx → type=shortlink
抖音视频 douyin.com/video/{id} → type=video
抖音用户 douyin.com/user/{sec_uid} → type=user
小红书 xiaohongshu.com/explore/{id} → type=note
B站视频 bilibili.com/video/{BV} → type=video
B站用户 bilibili.com/space/{mid} → type=user
知乎 zhihu.com/answer/{id} | /question/ → type=item
通用网页 http(s)://... → type=page
纯 sec_uid MS4w 开头 或 len>20 → douyin/user
纯数字 len>=15 → douyin/video(aweme_id)
纯数字 其他 → zhihu/item
```
**短链解析**(`_resolve_shortlink`,第 136 行):用 `curl -sI` 拿 `Location` 头,
再用正则从跳转 URL 里抠出 `aweme_id`。
### 5.2 执行主流程(`run`,第 163 行)
```
run(request)
├─ 1. resolve_urls(targets) → 标准化目标列表
├─ 2. 短链逐个解析
├─ 3. 判断同步 / 异步
│ async_mode = request.async_mode 或 len(targets) > 50
│
├─ 【异步分支】
│ · run_id = uuid[:8]
│ · 内存状态 {status, total, completed, results, errors}
│ · asyncio.create_task(_run_async(...))
│ · 立即返回 {status:"async", run_id} ← 前端轮询
│
└─ 【同步分支】
· db.create_task("single", ...)
· for target: _scrape_one() → _save_item()
· db.update_task_status("completed", summary)
· 返回 {status, task_id, duration, total, success, errors, data}
```
### 5.3 单目标抓取(`_scrape_one`,第 261 行)
```python
adapter = self._get_adapter(platform) # 按平台取适配器(带缓存)
if target["type"] == "user":
return await adapter.collect_user(target["target_id"]) # 返回 list
elif target["type"] in ("video", "note"):
return await adapter.collect_item(target["target_id"], depth, tool_level)
else:
return None
```
**返回值语义**(重要):
- `dict` → 单条内容
- `list` → 多条(用户主页的所有作品)
- `None` → 失败
### 5.4 落库(`_save_item`,第 292 行)
```python
db_id = self.db.insert_item(
task_id=..., platform=..., item_id=..., url=..., title=...,
author_name=..., author_id=..., published_at=..., text_content=...,
tags=..., stats=..., extra=..., media=...)
comments = item.get("comments", [])
if comments and db_id:
self.db.insert_comments(db_id, comments) # ★ 评论独立表
```
### 5.5 异步模式(`_run_async`,第 315 行)
- ✅ **同时写内存 + 落库**(内存态供轮询,落库供持久化)
- 内存态在 `self._tasks[run_id]`(**进程重启即丢**)
- 查询用 `get_async_result(run_id)`
### 5.6 ⚠️ 已知的局限(接手时要清楚)
| 局限 | 说明 |
|---|---|
| **没有限流/并发控制** | `for target in ready:` 是**纯串行**,目标多时会慢;也没有请求间隔(可能触发平台风控) |
| **异步态存内存** | 重启 Dashboard 后 `_tasks` 丢失(但库里有记录,前端看不到进度) |
| **无重试** | 单目标失败只记 `errors`,不重试 |
| **adapter 实例缓存** | `self._adapters` 进程内缓存(无清理) |
---
### 5.7 一个完整请求的生命周期(跟着走一遍最快懂)
以"采集某抖音作者的全部视频"为例:
```
① 前端
POST /api/scrape/run
{"targets": ["r606391422378804368"], "tool_level": 2}
│
▼
② routes/scrape.py:44 api_scrape_run()
组装 request → engine.run(request)
│
▼
③ scrape_engine.py:163 run()
├─ resolve_urls(["r6063..."])
│ → resolve_target() 识别:以 MS4w 开头 → 抖音 sec_uid
│ → [{"platform":"douyin","type":"user","target_id":"r6063...","status":"resolved"}]
│
├─ 同步或异步?len(targets)=1,不大于 50 → 同步
│
├─ db.create_task("single","douyin",...) → task_id = 1
│
├─ for target: _scrape_one(target,"douyin","light",2)
│ │
│ ▼
│ scrape_engine.py:261
│ _get_adapter("douyin") → DouyinScrapeAdapter() (进程内缓存)
│ type=="user" → adapter.collect_user("r6063...")
│ │
│ ▼
│ douyin_scrape.py:21 collect_user()
│ _try_tools(2, [("opencli", ...), ("agent-browser", ...)])
│ │
│ ├─ 尝试 1:_opencli_user_videos()
│ │ _run_opencli(["douyin","user-videos","r6063...",
│ │ "--limit","20","--with_comments","true","-f","json"])
│ │ → subprocess 执行 opencli(Node CLI)
│ │ → _parse_output() ← JSON → YAML → 纯文本 三层兜底
│ │ → 成功返回 list[dict] → _try_tools 立刻返回,不再降级
│ │
│ └─ 尝试 1 失败(OpenCLI 没装/超时/报错)
│ → 尝试 2:_browser_user_profile()
│ → Playwright 启真实 Chrome → 打开页面 → JS 提取
│
│ → 每条数据 _to_schema() 转成统一格式
│ → 返回 list
│
├─ for item: _save_item(task_id=1, item)
│ db.insert_item(...) → db_id((platform,item_id) 唯一,重复则忽略)
│ db.insert_comments(db_id, comments) ← 评论另存
│
├─ db.update_task_status(1, "completed", summary={success:N, errors:0})
│
└─ return {status:"completed", task_id:1, duration, total, success, errors, data:[...]}
│
▼
④ 前端拿到 data,渲染列表
```
**异步分支的差异**(目标 > 50 个,或显式 `async_mode=true`):
```
run() 立即返回 {status:"async", run_id:"ab12cd34"}
↓(后台)
asyncio.create_task(_run_async(run_id, targets, ...))
↓
建持久化任务 → 逐个 _scrape_one + _save_item
↓
进度写内存 self._tasks[run_id](供轮询)
结果写 SQLite(供持久化)
↓
前端轮询 get_async_result(run_id) 看进度
```
**⚠️ 注意**:异步进度**只在内存**,Dashboard 重启就丢
(库里数据还在,但前端看不到进度了)。
---
## 六、存储层(`services/scrape_db.py`,415 行)
### 6.1 数据库位置
```python
DEFAULT_DB = AGENT_LOCAL / "data" / "scrape.db"
```
(`AGENT_LOCAL` 是环境变量;默认 `~/workbuddy-agent-os/agent-local`)
### 6.2 四张表(实测 `CREATE TABLE`)
```sql
-- ① 采集任务(一次 run 一条)
scrape_tasks(
id, type, -- single / batch / scheduled
platform, target, -- 目标(批量时是 JSON 数组)
depth, tool_level, machine,
status, -- pending / running / completed / failed
total_targets, completed_targets,
summary, -- 摘要 JSON
created_at
)
-- ② 采集到的内容
scrape_items(
id, task_id → scrape_tasks,
platform, item_id, -- 平台内唯一 ID
url, title, author_name, author_id,
published_at, collected_at, text_content, tags,
... -- 还有 stats / extra / media 等
)
-- ③ 评论
scrape_comments(id, item_db_id → scrape_items,
author_name, text, likes, replied_at)
-- ④ 采集源(长期跟踪)
scrape_sources(
id, platform, source_type, -- user / hashtag / keyword / url_list / api
target, display_name,
category, -- 自定义分类
notes,
schedule, -- CRON(定期采集)
depth, tool_level, last_collected,
status -- active / paused
)
```
### 6.3 方法清单(22 个,实测)
```
任务:create_task / update_task_status / get_task / list_tasks
内容:insert_item / get_item_id / item_exists / get_item / list_items
评论:insert_comments / get_comments
采集源:upsert_source / update_source / list_sources / get_due_sources /
update_source_collected / delete_source
统计:count_by_platform / count_today / sources_count / task_stats
```
### 6.4 去重机制
`insert_item` **依赖 `(platform, item_id)` 唯一约束** —— 重复插入时
用 `INSERT OR IGNORE` 模式(验证文档 L1-2 有测例)。
---
## 七、HTTP 层(`routes/scrape.py`,711 行 / 22 端点)
### 7.1 端点清单(实测)
```
采集
POST /api/scrape/run 发起采集(targets + depth + tool_level)
POST /api/scrape/resolve 只解析 URL,不采集
POST /api/scrape/douyin-stats 抖音数据查询
查询
GET /api/scrape/title 取标题
GET /api/scrape/result 结果
GET /api/scrape/tasks 任务列表
GET /api/scrape/items 内容列表
GET /api/scrape/items/{id} 单项详情
GET /api/scrape/stats 统计
采集源管理
POST /api/scrape/sources 新建源
GET /api/scrape/sources 源列表
DEL /api/scrape/sources/{id} 删源
追踪(视频 / 作者)
POST /api/scrape/track-video 追踪视频
GET /api/scrape/tracked-videos 已追踪视频
POST /api/scrape/delete-tracked/{id}
POST /api/scrape/refresh-video/{id} 刷新单个视频
POST /api/scrape/track-author 追踪作者
GET /api/scrape/tracked-authors 已追踪作者
POST /api/scrape/refresh-author/{id}
GET /api/scrape/author-history/{id}
主题 / 登录
POST /api/scrape/import-topics 批量导入主题
GET /api/scrape/check-login 检测登录态
GET /api/scrape/open-login 打开登录
```
### 7.2 关键实现(实测)
**`POST /api/scrape/run`**(第 43 行)—— 前端发起采集的唯一入口:
```python
@router.post("/run")
async def api_scrape_run(data: dict = {}):
targets = data.get("targets", data.get("target", []))
if isinstance(targets, str):
targets = [targets] # 兼容单个字符串
request = {
"targets": targets,
"platform": data.get("platform", "auto"),
"depth": data.get("depth", "light"),
"tool_level": data.get("tool_level", 2), # ← 默认 2(OpenCLI + 浏览器)
"machine": data.get("machine", ""),
"multi_machine": data.get("multi_machine", False),
"async_mode": data.get("async_mode", False),
}
engine = _get_engine() # 模块级单例
result = await engine.run(request)
return {"status": "ok", **result}
```
**登录态两端点**(第 692 / 703 行)—— 都委托给 `mediacrawler_adapter`:
```python
GET /api/scrape/check-login → mediacrawler_adapter.check_login_status()
POST /api/scrape/open-login → mediacrawler_adapter.open_login_page()
```
### 7.3 引擎实例
路由层用**模块级单例**拿 engine(`_get_engine()`),
所以 `engine._tasks`(异步态)和 `engine._adapters`(适配器缓存)
在整个 Dashboard 进程内共享。
---
## 七·五、Dashboard 插件(`plugins/crawl.py`,91 行)
抓取模块**作为 Dashboard 插件**注册(提供概览统计,不是核心逻辑):
```python
class CrawlDashboardPlugin(DashboardPlugin):
name = "crawl"
label = "内容抓取"
icon = "📡"
order = 35
```
**它做三件事**:
1. `summary()` — 概览:总抓取数 / 今日新增 / 抓取源(从 `ScrapeDB` 读)
2. `detail(machine)` — 指定机器的详情
3. `actions()` — 快捷操作(跳转 `scrape` 视图)
**注意**:它有个**兜底设计** —— 如果 `ScrapeDB` 不可用(数据库损坏/权限),
会退化成**统计知识库里的 md 文件数**,而不是报错。
**⛔ 移植注意**:如果目标项目没有这套插件框架,`plugins/crawl.py`
可以直接丢弃(它只是 Dashboard 的展示层,不影响抓取功能本身)。
---
## 八、前端(`frontend/src/views/scrape.js`,131 KB)
⚠️ **这是模块里最大的单文件**(131 KB)。功能覆盖:22 个端点的界面。
**建议接手团队**:不要照搬这个前端,按第七节的 HTTP 契约重写。
理由:131 KB 单文件难维护,且和本项目的视图框架耦合。
---
## 九、验证方案(项目里已有现成的)
`services/scrape_validation.md`(314 行)已经写了 **7 个 Level 的验证清单**:
```
L0 基础设施(3 项) Python import / SQLite 建库 / FastAPI 路由注册
L1 数据库层(3 项) 建任务 / 写入+去重 / 评论入库
L2 适配器 Mock(3 项) 工具降级逻辑 / 一级失败二级成功
L3 适配器真实(3 项) 抖音用户采集 / 小红书 / 详情+评论 ← 需 OpenCLI + 登录态
L4 引擎层(3 项) 解析 URL / 执行采集 / 异步轮询
L5 API 层(4 项) curl 打 4 个端点
L6 前端(4 项) 浏览器里操作
L7 异常(1+ 项) OpenCLI 不可用时应抛清晰错误
```
**⚠️ 但要注意**:该文档写于 2026-07-16,**里面的方法名已过时**:
```
文档写 resolve_targets ← 不存在
代码里是 resolve_urls ← 实际
文档写 get_result ← 不存在
代码里是 get_async_result ← 实际
```
**以代码为准**。
---
## 十、移植清单(换环境要改什么)
| # | 依赖 | 位置 | 处理 |
|---|---|---|---|
| 1 | **OpenCLI**(Node CLI) | `adapters/__init__.py` 的 `OPENCLI` 常量 | 目标机 `npm i -g @jackwener/opencli`,或设 `OPENCLI_PATH` |
| 2 | **真实 Chrome** | `browser_helpers.py` 的 `_CHROME_PATHS` | Linux/Windows 要改路径(如 `/usr/bin/google-chrome`) |
| 3 | **Playwright** | pip | `pip install playwright && playwright install chromium` |
| 4 | **PyYAML** | pip(可选但强烈建议) | 不装会降级到纯文本解析(丢嵌套结构) |
| 5 | **AGENT_LOCAL** 环境变量 | `scrape_db.py:18` | 定义了才能定位 `scrape.db` |
| 6 | **平台登录态** | Chrome profile | 目标机需手动登录一次目标平台 |
| 7 | `SCRAPE_CHROME_USER_DATA` | 环境变量(可选) | 要持久化登录态时指定 |
### 最小可运行子集(只要"能采集")
```
services/scrape_db.py 数据库(4 表)
services/adapters/__init__.py 基类 + 降级 + OpenCLI 调用 + 输出解析
services/adapters/browser_helpers.py 浏览器降级
services/adapters/<目标平台>_scrape.py 目标平台适配器
```
—— 这 4 个文件就能跑通单平台采集,不需要 engine/routes/前端。
---
## 十一、接手建议(按顺序)
```
第 1 步 装 OpenCLI + Playwright + 真实 Chrome,跑 validation.md 的 L0/L1
(这两级零外部依赖,能验证环境对不对)
第 2 步 跑 L2(Mock 测试)—— 验证降级逻辑,不需要真实平台
这时你已经能理解 _try_tools 的语义
第 3 步 登录目标平台,跑 L3(真实采集)—— 第一次真正拿数据
如果 OpenCLI 不通,会看到它降级到浏览器,日志里有 ⚠️
第 4 步 跑 L4/L5(引擎 + API)
第 5 步 替换前端(不要照搬 131 KB)
第 6 步 加你要的东西:限流 / 重试 / 并发控制(现在都没有)
```
---
## 十二、这个模块还没做的事(接手可以补)
```
① 限流与请求间隔 —— 现在纯串行、无间隔,目标多时可能触发平台风控
② 失败重试 —— 现在失败只记 errors
③ 并发控制 —— 没有信号量,大量目标只能串行
④ 异步态持久化 —— run 进度存内存,重启即丢
⑤ 登录态自动检测 —— check-login 端点有,但采集前没强制校验
⑥ 代理支持 —— 没有看到代理配置(多账号场景会需要)
```
---
## 附录:本说明的取证方式(可复现)
```bash
cd 05_tools/10_dashboard
# 架构
head -20 services/adapters/__init__.py
# 浏览器("套真实浏览器"的做法)
sed -n '1,70p' services/adapters/browser_helpers.py
# 降级机制
grep -n "_try_tools" -A 14 services/adapters/__init__.py
# 引擎流程
grep -nE "^ (async )?def " services/scrape_engine.py
# 表结构
grep -n "CREATE TABLE" -A 12 services/scrape_db.py
# 端点
grep -nE "^@router\." routes/scrape.py
```
+76 -7
View File
@@ -18,6 +18,7 @@ api/auth.py WebUI 登录鉴权
api/monitor/* 监控层整体(含 platforms.py 能力矩阵)
api/monitor/db.py MySQL 连接层(可回退 SQLite 供测试用)
api/monitor/migrate_from_sqlite.py SQLite → MySQL 一次性迁移脚本
api/monitor/upstream.py 上游更新检查(定时 fetch 上游并比对)
api/routers/{auth,monitor,settings}.py
api/schemas/{auth,monitor,settings}.py
api/services/interpreter.py 解释器探测(uv / .venv / 当前解释器)
@@ -25,9 +26,18 @@ webui/src/components/{monitor,settings,auth}/ 新视图
webui/src/components/layout/{PlatformSwitcher,UnwiredPlatformNotice}.tsx
webui/src/{hooks/useMonitor.ts,hooks/usePlatform.ts,store/platformStore.ts,lib/monitorFormat.ts,types/monitor.ts}
docs/监控功能使用说明.md
tests/test_{auth,settings,platforms,monitor_*}.py
tests/test_{auth,settings,platforms,qrlogin,monitor_*,upstream}.py
Dockerfile / .dockerignore / docker-compose.yml 服务器部署用
```
> `api/monitor/upstream.py` 要调 `git`,而 `python:3.11-slim` 不带它 —— Dockerfile 里为此
> **显式装了 git**。改了 Dockerfile 就必须重建镜像(`docker compose build`),`./deploy.sh`
> 只重建前端,不重建镜像。
其中 `api/monitor/qrlogin.py` + `webui/.../QrLoginPanel.tsx` 是**服务器专用**的扫码登录:
那台机器上 Chrome 跑在 Xvfb 里,`show_qrcode` 调的 PIL `Image.show()` 需要桌面看图程序,
服务器没有,二维码会无处可去。所以改成用 CDP 把二维码从页面里读出来交给前端 `<img>` 显示。
### 2. 加法改动(低冲突)
只在既有文件里**新增**内容,不改动原有行:
@@ -36,7 +46,7 @@ tests/test_{auth,settings,platforms,monitor_*}.py
|---|---|
| `cmd_arg/arg.py` | typer 选项:`--enable_cdp_mode`、`--inject_all_cookies`、`--save_login_state`、`--cookies_file`、`--crawler_max_sleep_sec`,以及对应的 `config.*` 回写 |
| `api/schemas/crawler.py` | `CrawlerStartRequest` 的若干**可选**字段(默认 `None`,不传则不加对应 CLI 参数) |
| `config/base_config.py` | `INJECT_ALL_COOKIES = False` |
| `config/base_config.py` | `INJECT_ALL_COOKIES = False`;`MASK_NICKNAME = False`(关掉昵称脱敏,见第 3 节) |
| `api/routers/__init__.py` | 导出新增的 router |
| `requirements.txt` | 补上 `websockets`(上游 `pyproject.toml` 里有、`requirements.txt` 里漏了) |
| `tests/conftest.py` | 新增 `_bypass_auth_for_non_auth_suites` fixture |
@@ -49,19 +59,22 @@ tests/test_{auth,settings,platforms,monitor_*}.py
| `api/routers/websocket.py` | 两个 WS 路由加 `dependencies=[Depends(require_ws_auth)]` | 上游若新增 WS 路由,**必须同样加上**,否则那条流是裸奔的 |
| `api/services/crawler_manager.py` | 解释器探测替换硬编码 `uv run`;`_build_command` 转发新增参数;新增 `is_busy()` / `run_and_wait()` 与完成事件 | 留意 `_build_command` 的参数拼装 |
| `media_platform/xhs/login.py` | `login_by_cookies` 在 `INJECT_ALL_COOKIES` 打开时注入**全部** cookie(默认关闭,行为不变) | 小改动,好合并 |
| `tools/user_hash.py` | `mask_nickname` 改为读 `config.MASK_NICKNAME`,本仓库默认**不脱敏**(原样返回)。上游作为教学版默认脱敏,但那是有损的 ——「张三」「张四」都成「张*」,而分清谁是谁正是监控这一层要干的活。脱敏实现本身没删,改回 `True` 即恢复上游行为 | 与 `config/base_config.py` 一起改,两处不同步会不一致 |
### 4. 上游 bug 修复(建议回馈上游)
| 文件 | 修的问题 |
|---|---|
| `media_platform/xhs/core.py` | 见下节 |
| `media_platform/xhs/login.py` | 同上(cookie 加固) |
| `media_platform/xhs/core.py` | 见下节 1 |
| `media_platform/xhs/login.py` | 见下节 2(cookie 加固) |
| `media_platform/douyin/core.py` | 见下节 3(首页 `goto` 永远超时,采集根本起不来) |
| `media_platform/douyin/login.py` | 见下节 4(注入 cookie 后页面陈旧,白等十分钟) |
---
## 二、应该给上游提 PR 的两个修复
## 二、应该给上游提 PR 的四个修复
这两处是**上游自身的缺陷**,提上去以后就不用自己背着:
这四处都是**上游自身的缺陷**,提上去以后就不用自己背着:
### 1. 博主主页抓取失败会跳掉整个博主(`xhs/core.py`)
@@ -80,19 +93,75 @@ tests/test_{auth,settings,platforms,monitor_*}.py
`a1` / `webId` 等签名所需 cookie 只能靠持久化 profile 补,冷启动时签名会失败。
默认行为保持不变,用 `INJECT_ALL_COOKIES` 开关控制。
### 3. 抖音首页的 `goto` 永远等不到 `load`(`douyin/core.py:101`)
```python
await self.context_page.goto(self.index_url) # 默认 wait_until="load"
```
抖音首页的 `load` 事件**不会触发**(有长连接/埋点类请求一直挂着)。实测:同一台
Chrome、同一个地址,`domcontentloaded` 0.7 秒返回,而 `load` 等满 90 秒仍然超时。
后果是整个采集**一步都没走就崩**,退出码 1 —— 看起来像"抖音不能用"。
修复:显式 `wait_until="domcontentloaded"`。上游的贴吧(`tieba/core.py`)和知乎
(`zhihu/core.py`)本来就是这么写的,抖音这个页面只是恰好属于"永远不 load"的那类。
> 这个缺陷在本机可能复现不出来(换个网络/有缓存时 `load` 也许能触发),所以社区里
> 没人报。它和网络快慢无关:不是"慢",是那个事件根本不会发生。
### 4. 注入 cookie 后页面是陈旧的(`douyin/login.py:266`)
`login_by_cookies()` 把 cookie 塞进 browser context,但**页面是在这之前加载的** ——
SPA 只在加载时读一次登录态,`localStorage.HasUserLogin` 于是还停在"未登录",
紧接着的 `check_login_state()` 会对着这个陈旧的值轮询到超时(600 次 × 1 秒 = 十分钟),
然后 `sys.exit()`。**下一轮**才正常,因为那时 cookie 已经在 profile 里了。
表现是"第一次跑白等十分钟、第二次才行",很容易被当成偶发。
修复:注入完 cookie 后 `reload(wait_until="domcontentloaded")`,让站点立刻重新判定会话。
> 与第 1 条同源:都是"页面状态是加载那一刻的快照"。本仓库的扫码登录(`api/monitor/qrlogin.py`)
> 和运营模块也各自踩过这个坑,那里的判据改成了拿 cookie 问后台接口,而不是读页面快照。
---
## 三、上游更新时怎么操作
### 先让机器替你盯着
「上游更新检查」(`api/monitor/upstream.py`,开关在 WebUI 的**系统设置 → 上游更新**)会按
间隔 `git fetch` 上游、算出落后几个提交,有更新就推企业微信。它是这份文档的自动化版:
没有它,「上游动了」这件事只取决于谁偶尔想起来去 fetch 一次。
两个细节决定了它为什么是安全的:它只 fetch 到 `FETCH_HEAD`,**不写工作区、不建 remote、不碰
`refs/remotes`**,所以和正在跑的采集、和下面的 `git pull` 都不冲突;默认**关闭**,因为要联网,
且需要镜像里有 git。
### 日常流程
```bash
git stash # 或先 commit 到自己的分支(推荐)
git fetch origin main
git rebase origin/main # 冲突只会出现在上表第 3、4 类文件里
./.venv/Scripts/python.exe -m pytest tests/ -q # 486 个测试就是回归网
./.venv/Scripts/python.exe -m pytest tests/ -q # 502 个测试就是回归网
```
### 直连 GitHub 不通时(本机常见)
本机到 `github.com` 时通时不通,**大包传输必断**(`Recv failure: Connection was reset`
或 `unexpected disconnect while reading sideband packet`),所以 `git clone` / `--unshallow`
这类一次性拉全量的操作基本必失败。可用的替代源:
```bash
# gitcode 的 GitHub 镜像,国内直连,比 GitHub 本身还新一天以内
git remote add gitcode https://gitcode.com/gh_mirrors/me/MediaCrawler.git
git fetch --no-tags --unshallow gitcode # 本仓库当初就是这样补全历史的,约 2 秒
```
注意 `git fetch` 只写 `refs/remotes/*`,**不会动本地 `main`**;
但拉镜像会把 `upstream/main` 指到镜像的 tip(可能比 GitHub 晚一天),
等 GitHub 通了再 `git fetch upstream` 正回来即可。
### 强烈建议:先把改动提交掉
当前状态是**未提交**的(25 个上游文件被改 + 31 个新文件)。在 `main` 分支上裸着工作区,
+27
View File
@@ -0,0 +1,27 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/__init__.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营模块:管理自己的小红书账号,读取创作者后台的数据。
与 `api.monitor` 是**并列关系**,不是它的扩展。两者数据形状不同:监控是「每轮
快照 + 差分」的公开互动数据,这里是创作者后台按日期给出的曝光/观看/完播等运营
指标。硬塞进同一个模型会同时污染两边。
路线是**纯请求**(无浏览器),依据见 tools/probe_creator_api.py 的 Phase 0 实测:
签名可自造、主站 cookie 即可认证、接口与参数已与真实页面对齐。
"""
+345
View File
@@ -0,0 +1,345 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/client.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""小红书创作者后台的纯请求客户端。
无浏览器:签名在本地算(见 signing.py),请求走 httpx。
**关于字段名的谨慎**:Phase 0 抓到的响应里,列表接口因为账号权限未生效而返回空壳
(`data.result` 只有 `{success, code, message}` 没有数据),所以**真实字段名尚未亲眼
见过**。因此每个指标都写成**多别名匹配**,并且解析不出来时存 `None` 而不是 0 ——
0 是真实值,None 是"不知道",两者混淆会让报表说谎。
"""
import re
from typing import Any, Dict, List, Optional
import httpx
from .signing import sign_xyw, signed_api
CREATOR_ORIGIN = "https://creator.xiaohongshu.com"
DATA_ANALYSIS_PAGE = f"{CREATOR_ORIGIN}/statistics/data-analysis"
USER_INFO_PATH = "/api/galaxy/user/info"
PERMISSION_PATH = "/api/galaxy/creator/datacenter/permission/query"
NOTE_LIST_PATH = "/api/galaxy/creator/datacenter/note/analyze/list"
USER_AGENT = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
)
# 签名被网关拒绝时返回的响应体。区分它和普通业务错误很重要:406 说明签名写错了,
# 而应用层的 code/-100 说明签名没问题、只是没有登录态。
SIGNATURE_REJECTED_MARKERS = ("code", -1)
class CreatorApiError(RuntimeError):
"""调用创作者后台失败。``status`` 用于区分是网关拒绝还是业务错误。"""
def __init__(self, message: str, status: int = 0, payload: Any = None):
super().__init__(message)
self.status = status
self.payload = payload
def trans_cookies(cookie_str: str) -> Dict[str, str]:
"""把 cookie 字符串解析成字典。容忍末尾分号、换行和零散空格。"""
jar: Dict[str, str] = {}
for chunk in re.split(r"[;\n]", cookie_str or ""):
chunk = chunk.strip()
if not chunk or "=" not in chunk:
continue
name, value = chunk.split("=", 1)
name = name.strip()
if name:
jar[name] = value.strip()
return jar
def _pick(item: Dict[str, Any], *names: str) -> Any:
"""按别名顺序取第一个存在的键。
字段名来自二手资料,未亲眼验证,所以不赌单一命名。
"""
for name in names:
if name in item and item[name] is not None:
return item[name]
return None
_COUNT_UNITS = {"万": 10_000, "w": 10_000, "W": 10_000, "亿": 100_000_000, "k": 1_000, "K": 1_000}
def as_int(value: Any) -> Optional[int]:
"""解析计数。处理 "1.2万"、"1,234"、"123" 与已经是数字的情况。"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, (int, float)):
return int(value)
text = str(value).strip().replace(",", "").replace(" ", "")
if not text or text in ("-", "--", "暂无"):
return None
unit = 1
suffix = text[-1]
if suffix in _COUNT_UNITS:
unit = _COUNT_UNITS[suffix]
text = text[:-1]
try:
return int(float(text) * unit)
except ValueError:
return None
def as_float(value: Any) -> Optional[float]:
"""解析比率。``"12.3%"`` -> 12.3;``"0.123"`` 原样返回数字。"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, (int, float)):
return float(value)
text = str(value).strip().replace("%", "")
try:
return float(text)
except ValueError:
return None
_DURATION_RE = re.compile(r"(?:(\d+)\s*分)?\s*(?:(\d+)\s*秒)?")
def as_seconds(value: Any) -> Optional[float]:
"""解析时长。处理 ``"1分30秒"``、``"01:30"``、``"45"``(秒)。"""
if value is None or isinstance(value, bool):
return None
if isinstance(value, (int, float)):
return float(value)
text = str(value).strip()
if not text or text in ("-", "--"):
return None
if ":" in text:
parts = text.split(":")
try:
total = 0.0
for part in parts:
total = total * 60 + float(part)
return total
except ValueError:
return None
if "分" in text or "秒" in text:
minutes = re.search(r"(\d+)\s*分", text)
seconds = re.search(r"(\d+)\s*秒", text)
if not minutes and not seconds:
return None
return float(minutes.group(1) if minutes else 0) * 60 + float(
seconds.group(1) if seconds else 0
)
try:
return float(text)
except ValueError:
return None
# 指标 -> 候选原始字段名。别名来自公开资料,权威与否只能等真实响应来验证。
_FIELD_ALIASES: Dict[str, tuple] = {
"exposure": ("exposure", "exposure_count", "imp", "impression", "impression_count"),
"views": ("views", "view", "view_count", "watch", "watch_count", "read_count"),
"likes": ("likes", "like", "like_count", "liked_count"),
"comments": ("comments", "comment", "comment_count", "comments_count"),
"favorites": ("favorites", "favorite", "favorite_count", "collect", "collect_count", "collected_count"),
"shares": ("shares", "share", "share_count", "shared_count"),
"new_followers": ("new_followers", "fans_growth", "follower_growth", "increase_fans"),
"danmaku": ("danmaku", "danmaku_count", "barrage"),
"cover_ctr": ("cover_ctr", "cover_click_rate", "cover_click_ratio", "ctr"),
"avg_watch_seconds": ("avg_watch_seconds", "avg_watch_time", "average_watch_time"),
"two_second_exit_rate": ("two_second_exit_rate", "2s_exit_rate", "exit_rate_2s"),
"completion_rate": ("completion_rate", "finish_rate", "complete_rate"),
}
_COUNT_FIELDS = {
"exposure", "views", "likes", "comments", "favorites", "shares", "new_followers", "danmaku",
}
_DURATION_FIELDS = {"avg_watch_seconds"}
_RATE_FIELDS = {"two_second_exit_rate", "completion_rate", "cover_ctr"}
NOTE_ID_ALIASES = ("note_id", "noteId", "id", "content_id", "item_id")
TITLE_ALIASES = ("title", "content", "note_title", "display_title")
PUBLISH_TIME_ALIASES = ("publish_time", "publishTime", "post_time", "create_time", "time")
def _normalize_metric(name: str, raw: Any) -> Any:
if name in _COUNT_FIELDS:
return as_int(raw)
if name in _DURATION_FIELDS:
return as_seconds(raw)
if name in _RATE_FIELDS:
return as_float(raw)
return raw
def normalize_note(item: Dict[str, Any]) -> Dict[str, Any]:
"""把一条原始记录规范化成落库用的字段。"""
note: Dict[str, Any] = {
"note_id": str(_pick(item, *NOTE_ID_ALIASES) or ""),
"title": str(_pick(item, *TITLE_ALIASES) or ""),
"publish_time": as_int(_pick(item, *PUBLISH_TIME_ALIASES)),
}
for name, aliases in _FIELD_ALIASES.items():
note[name] = _normalize_metric(name, _pick(item, *aliases))
return note
def find_note_list(payload: Any) -> List[Dict[str, Any]]:
"""在响应里找出笔记数组。
接口的**确切结构还没亲眼见过**(权限未生效时 `data.result` 里没有数据),
所以不写死路径:遍历 JSON,挑出"看起来像一批笔记记录"的那个列表 ——
元素是 dict,且至少带一个指标字段。找不到就返回空列表,让上层如实报"没数据"。
"""
best: List[Dict[str, Any]] = []
metric_keys = {alias for aliases in _FIELD_ALIASES.values() for alias in aliases}
def walk(value: Any) -> None:
nonlocal best
if isinstance(value, dict):
for child in value.values():
walk(child)
elif isinstance(value, list):
if value and isinstance(value[0], dict):
keys = set(value[0].keys())
if keys & metric_keys and len(value) > len(best):
best = value
for child in value:
walk(child)
walk(payload)
return best
class CreatorClient:
"""一个账号的客户端。``cookie`` 就是它的全部身份。"""
def __init__(self, cookie: str, timeout: float = 25.0):
self.cookies = trans_cookies(cookie)
self.a1 = self.cookies.get("a1", "")
self._timeout = timeout
@property
def looks_authenticated(self) -> bool:
"""签名需要 a1;没有它连请求都签不出来。"""
return bool(self.a1)
def _headers(self, api: str, body: dict | None = None) -> Dict[str, str]:
cookie_header = "; ".join(f"{k}={v}" for k, v in self.cookies.items())
return {
"user-agent": USER_AGENT,
"accept": "application/json, text/plain, */*",
"accept-language": "zh-CN,zh;q=0.9",
"origin": CREATOR_ORIGIN,
"referer": DATA_ANALYSIS_PAGE,
"cookie": cookie_header,
**sign_xyw(api, self.a1, body=body),
}
async def _get(self, path: str, params: Dict[str, Any]) -> Dict[str, Any]:
if not self.looks_authenticated:
raise CreatorApiError("cookie 里没有 a1,无法完成签名", status=0)
query = "&".join(f"{k}={v}" for k, v in params.items())
api = signed_api(path, query)
url = f"{CREATOR_ORIGIN}{path}?{query}" if query else f"{CREATOR_ORIGIN}{path}"
async with httpx.AsyncClient(timeout=self._timeout, follow_redirects=False) as client:
response = await client.get(url, headers=self._headers(api))
return self._unwrap(response)
@staticmethod
def _unwrap(response: httpx.Response) -> Dict[str, Any]:
if response.status_code == 406:
raise CreatorApiError(
"签名被网关拒绝(406)—— 待签字符串的拼法不对", status=406
)
try:
payload = response.json()
except Exception as exc: # noqa: BLE001
raise CreatorApiError(
f"响应不是 JSON(HTTP {response.status_code})", status=response.status_code
) from exc
if response.status_code == 401 or payload.get("code") == -100:
raise CreatorApiError("登录态无效或已过期", status=401, payload=payload)
if not payload.get("success", True):
raise CreatorApiError(
str(payload.get("msg") or "接口返回失败"),
status=response.status_code,
payload=payload,
)
return payload
async def fetch_user_info(self) -> Dict[str, Any]:
"""当前 cookie 属于哪个账号。登录后用它取名与去重。"""
payload = await self._get(USER_INFO_PATH, {})
data = payload.get("data") or {}
return {
"user_id": str(data.get("userId") or ""),
"nickname": str(data.get("userName") or ""),
"avatar": str(data.get("userAvatar") or ""),
"red_id": str(data.get("redId") or ""),
"role": str(data.get("role") or ""),
"permissions": list(data.get("permissions") or []),
}
async def fetch_permission(self) -> Dict[str, Any]:
"""数据权限状态。
``tip_msg`` 是后台原话(实测是"已为您申请数据权限,次日可查看"),照抄不改写 ——
这条信息必须原样交给用户,它解释了"为什么没有数据"。
"""
payload = await self._get(PERMISSION_PATH, {})
data = payload.get("data") or {}
return {
"display": data.get("display"),
"status": data.get("status"),
"tip": str(data.get("tip_msg") or ""),
}
async def fetch_note_list(
self, start_ms: int, end_ms: int, page_num: int = 1, page_size: int = 10, note_type: int = 0
) -> List[Dict[str, Any]]:
"""按发布时间区间取笔记列表。
参数与顺序**照抄真实页面的请求**(见 Phase 0 抓包),不要凭感觉改。
"""
payload = await self._get(
NOTE_LIST_PATH,
{
"post_begin_time": start_ms,
"post_end_time": end_ms,
"type": note_type,
"page_size": page_size,
"page_num": page_num,
},
)
return [normalize_note(item) for item in find_note_list(payload)]
+300
View File
@@ -0,0 +1,300 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/login.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营账号的扫码登录。
**与监控的扫码登录(`api.monitor.qrlogin`)有一处决定性差异**:那边把登录态写进
浏览器**默认 profile**,因为爬虫要复用它;这边要的是 **cookie 字符串**,因为采集
走纯 HTTP。所以这里每次登录都开一个**临时上下文**,扫完取出 cookie 就丢弃 ——
* 登第二个账号不会把第一个顶掉(默认 profile 只能装一个登录态);
* 完全不影响监控那个登录态;
* 十个账号互不干扰。
扫码入口仍是主站(`www.xiaohongshu.com`):Phase 0 实测证明**主站的 cookie 就能
认证创作者后台**,不需要单独的创作者登录。
"""
import asyncio
import time
from typing import Any, Dict, Optional
from playwright.async_api import async_playwright
from tools import utils
from ..monitor.platforms import PLATFORM_XHS
from .client import CreatorApiError, CreatorClient
QR_TTL_SECONDS = 300
STATUS_IDLE = "idle"
STATUS_WAITING = "waiting"
STATUS_SUCCESS = "success"
STATUS_EXPIRED = "expired"
STATUS_ERROR = "error"
LOGIN_URL = "https://www.xiaohongshu.com"
QR_SELECTOR = "xpath=//img[@class='qrcode-img']"
LOGIN_BUTTON_SELECTOR = "xpath=//*[@id='app']/div[1]/div[2]/div[1]/ul/div[1]/button"
# 判据不读页面状态,而是拿 cookie 直接问创作者后台"我是谁"。
#
# **为什么不用页面状态**:`window.__INITIAL_STATE__` 是**页面加载那一刻的快照**。
# 监控那边的同一个探针能用,是因为那台浏览器的页面加载时就已经登录了,快照里
# loggedIn 就是 true。而扫码是"页面加载之后才登录的"—— SPA 内部确实登进去了,
# 但那个初始快照不会翻转,于是检测永远等不到,界面就一直停在二维码上。
#
# `user/info` 则是权威的:实测**游客也会拿到 a1**(所以签名算得出来),但接口直接
# 回 401「无登录信息」;只有真正登录了才返回 user_id。所以"有 a1"什么都证明不了,
# "后台认这份身份"才是。
LOGIN_CHECK_INTERVAL_SECONDS = 5.0
_lock = asyncio.Lock()
_current: Optional["AccountLoginSession"] = None
# 完成后的快照。会话一旦被取走 cookie 就会拆掉,而前端可能还有一个在飞的轮询——
# 那个请求若看到 _current 为空就会回报 idle,把已经显示出来的成功状态又擦掉。
# 把结果留在这里,重复轮询就稳定得多。
_last_result: Optional[Dict[str, Any]] = None
_playwright: Any = None
def _cdp_url() -> str:
import os
import config
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
async def _connect():
global _playwright
if _playwright is None:
_playwright = await async_playwright().start()
return await _playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
async def _disconnect() -> None:
global _playwright
if _playwright is not None:
try:
await _playwright.stop()
except Exception:
pass
_playwright = None
async def _read_qr(page: Any) -> str:
image = await utils.find_login_qrcode(page, selector=QR_SELECTOR)
if image:
return image
# 登录框不一定自己弹出来,这是爬虫自身扫码流程的同款兜底。
await asyncio.sleep(0.5)
try:
await page.locator(LOGIN_BUTTON_SELECTOR).click(timeout=5000)
except Exception:
return ""
return await utils.find_login_qrcode(page, selector=QR_SELECTOR)
class AccountLoginSession:
"""一次针对**临时上下文**的扫码尝试。"""
def __init__(self, context: Any, page: Any) -> None:
self.status = STATUS_WAITING
self.message = "请用手机扫描二维码"
self.image = ""
self.started_at = time.time()
self.account: Optional[Dict[str, Any]] = None
self.cookie: str = ""
self.platform = PLATFORM_XHS
self._context = context
self._page = page
self._last_login_check = 0.0
@property
def elapsed(self) -> float:
return time.time() - self.started_at
async def refresh(self) -> None:
if self.status != STATUS_WAITING:
return
if self.elapsed > QR_TTL_SECONDS:
self.status = STATUS_EXPIRED
self.message = "二维码已超时,请重新获取"
return
# 前端每 2 秒问一次,但没必要每次都去打后台接口 —— 一次真实的网络往返
# 去确认一个通常还没发生的事件是浪费。
now = time.time()
if now - self._last_login_check < LOGIN_CHECK_INTERVAL_SECONDS:
return
self._last_login_check = now
try:
cookies = await self._context.cookies()
except Exception:
self.status = STATUS_ERROR
self.message = "登录窗口已被关闭,请重新获取"
return
cookie = "; ".join(f"{c['name']}={c['value']}" for c in cookies)
try:
info = await CreatorClient(cookie).fetch_user_info()
except CreatorApiError:
# 还没登录(或者刚扫、后端还没认),继续等。
return
if not info.get("user_id"):
return
# 登录成功:cookie 取自**这个临时上下文**,取完上下文就丢弃,
# 所以不会残留、也不会影响别的账号。
self.cookie = cookie
self.account = info
self.status = STATUS_SUCCESS
self.message = f"登录成功:{info.get('nickname') or info['user_id']}"
def snapshot(self) -> Dict[str, Any]:
return {
"status": self.status,
"message": self.message,
"image": self.image,
"elapsed": int(self.elapsed),
"expires_in": max(0, int(QR_TTL_SECONDS - self.elapsed)),
"account": self.account,
}
async def close(self) -> None:
"""关掉临时上下文。这是它存在的全部意义 —— 用完即弃。"""
for closer in (self._page.close, self._context.close):
try:
await closer()
except Exception:
pass
async def _teardown_locked() -> None:
global _current
if _current is not None:
await _current.close()
_current = None
async def start() -> Dict[str, Any]:
"""开一个临时上下文,打开登录页,取回二维码。"""
global _current, _last_result
async with _lock:
_last_result = None
await _teardown_locked()
try:
browser = await _connect()
except Exception as exc:
await _disconnect()
raise RuntimeError(
f"连接浏览器失败({_cdp_url()})。请确认服务器上的 Chrome 以 "
f"--remote-debugging-port 启动。原始错误:{exc}"
) from exc
# 临时上下文,不是 contexts[0]。这里刻意要一个干净的身份 ——
# 借用操作者自己的登录态会让"新增账号"变成"再读一遍当前账号"。
context = await browser.new_context()
page = await context.new_page()
try:
await page.goto(LOGIN_URL, wait_until="domcontentloaded", timeout=45000)
image = await _read_qr(page)
except Exception as exc:
try:
await page.close()
await context.close()
except Exception:
pass
raise RuntimeError(f"打开登录页失败:{exc}") from exc
session = AccountLoginSession(context, page)
session.image = image
if not image:
session.status = STATUS_ERROR
session.message = "页面上没找到二维码,请确认站点结构没有变化"
_current = session
return session.snapshot()
async def status() -> Dict[str, Any]:
async with _lock:
if _current is None:
if _last_result is not None:
return _last_result
return {
"status": STATUS_IDLE,
"message": "",
"image": "",
"elapsed": 0,
"expires_in": 0,
"account": None,
}
await _current.refresh()
return _current.snapshot()
async def remember_result(snapshot: Dict[str, Any]) -> None:
"""记住已完成的扫码结果,供后续轮询重复返回。"""
global _last_result
async with _lock:
_last_result = snapshot
async def take_cookie() -> Optional[str]:
"""取走已登录的 cookie 并结束会话。
由路由层在落库时调用。cookie 只经内存传递,**不进响应体** —— 它是凭证,
前端没有任何理由看到它。
"""
global _current
async with _lock:
if _current is None or _current.status != STATUS_SUCCESS:
return None
cookie = _current.cookie
await _teardown_locked()
return cookie
async def cancel() -> Dict[str, Any]:
global _last_result
async with _lock:
_last_result = None
await _teardown_locked()
return {
"status": STATUS_IDLE,
"message": "已取消",
"image": "",
"elapsed": 0,
"expires_in": 0,
"account": None,
}
async def shutdown() -> None:
async with _lock:
await _teardown_locked()
await _disconnect()
+132
View File
@@ -0,0 +1,132 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/models.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营模块的数据模型。
**刻意复用 `MonitorBase`**:这样 `init_db` 的 `create_all` 会顺手建出新表,而
`_ensure_columns`(已改为按模型元数据推导)也会自动给新表补字段 —— 不必再维护一份
建表语句。表落在同一个库里,与监控互不干扰。
"""
from typing import Optional
from sqlalchemy import BigInteger, Float, ForeignKey, Index, Integer, String, Text
from sqlalchemy.orm import Mapped, mapped_column, relationship
from ..monitor.models import MonitorBase
# 创作者后台的数据权限是「首次访问时自动申请、次日生效」。这个状态必须如实呈现:
# 显示成"没数据"会让人以为采集坏了,实际是在等审批。
PERMISSION_UNKNOWN = "unknown"
PERMISSION_PENDING = "pending" # 已申请,未生效(提示语:"次日可查看")
PERMISSION_ACTIVE = "active"
PERMISSION_MISSING = "missing" # 接口明确说没有权限
# 账号自身的可用性。
ACCOUNT_OK = "ok"
ACCOUNT_EXPIRED = "expired" # cookie 失效,需要重新扫码
ACCOUNT_ERROR = "error"
class CreatorAccount(MonitorBase):
"""一个自己的小红书账号。
纯请求路线下,**一个账号的全部身份就是一份 cookie** —— 没有浏览器 profile、
没有独立目录。所以"多账号"在这里只是表里的多行,不是多套运行环境。
`cookie` 是凭证:与监控的 cookie 同样对待,只存库、绝不回显接口。
"""
__tablename__ = "creator_account"
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
# 展示名。优先用后台返回的昵称,用户可以改。
nickname: Mapped[str] = mapped_column(String(128), nullable=False, default="")
# 创作者后台的账号标识,由 /api/galaxy/user/info 返回,用于去重。
user_id: Mapped[str] = mapped_column(String(64), nullable=False, default="", index=True)
red_id: Mapped[str] = mapped_column(String(64), nullable=False, default="")
avatar: Mapped[str] = mapped_column(Text, nullable=False, default="")
cookie: Mapped[str] = mapped_column(Text, nullable=False, default="")
status: Mapped[str] = mapped_column(String(16), nullable=False, default=ACCOUNT_OK)
permission_status: Mapped[str] = mapped_column(
String(16), nullable=False, default=PERMISSION_UNKNOWN
)
# 后台原话,例如"已为您申请数据权限,次日可查看"。照抄,不改写。
permission_tip: Mapped[str] = mapped_column(Text, nullable=False, default="")
last_checked_at: Mapped[Optional[int]] = mapped_column(BigInteger)
last_synced_at: Mapped[Optional[int]] = mapped_column(BigInteger)
# 上次同步用的时间范围(天,按发布时间)。当前展示的数据就是这个范围的产物 ——
# 不记下来的话,界面只能说明"同步过了",说不清是哪一段。
last_sync_days: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
last_error: Mapped[Optional[str]] = mapped_column(Text)
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
notes: Mapped[list["CreatorNoteStat"]] = relationship(
back_populates="account", cascade="all, delete-orphan"
)
class CreatorNoteStat(MonitorBase):
"""一篇作品在某个采集时点的运营数据。
创作者后台给的是**累计值**(截至查询时点),所以反复采集天然形成时间序列 ——
与监控的"快照 + 差分"是同一个思路,因此这里保留 `captured_at` 而不是覆盖写。
"""
__tablename__ = "creator_note_stat"
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
account_id: Mapped[int] = mapped_column(
ForeignKey("creator_account.id", ondelete="CASCADE"), nullable=False, index=True
)
note_id: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
# 发布时间(毫秒)。后台按发布时间筛选,这是它的主时间轴。
publish_time: Mapped[Optional[int]] = mapped_column(BigInteger)
# --- 运营指标 ---------------------------------------------------------
# 计数用 BigInteger:曝光量可以很大,用 INT 迟早溢出。
exposure: Mapped[Optional[int]] = mapped_column(BigInteger)
views: Mapped[Optional[int]] = mapped_column(BigInteger)
likes: Mapped[Optional[int]] = mapped_column(BigInteger)
comments: Mapped[Optional[int]] = mapped_column(BigInteger)
favorites: Mapped[Optional[int]] = mapped_column(BigInteger)
shares: Mapped[Optional[int]] = mapped_column(BigInteger)
new_followers: Mapped[Optional[int]] = mapped_column(BigInteger)
danmaku: Mapped[Optional[int]] = mapped_column(BigInteger)
# 比率与时长。后台返回的可能是 "12.3%"/"1分30秒" 这类字符串,解析不了的存 NULL
# 而不是 0 —— 与监控层的口径一致:0 是真实值,NULL 是"不知道"。
cover_ctr: Mapped[Optional[float]] = mapped_column(Float)
avg_watch_seconds: Mapped[Optional[float]] = mapped_column(Float)
two_second_exit_rate: Mapped[Optional[float]] = mapped_column(Float)
completion_rate: Mapped[Optional[float]] = mapped_column(Float)
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
account: Mapped["CreatorAccount"] = relationship(back_populates="notes")
__table_args__ = (
# 同一个时点同一篇只留一行,重复同步不会堆积。
Index("ix_creator_note_stat_unique", "account_id", "note_id", "captured_at", unique=True),
)
+360
View File
@@ -0,0 +1,360 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/service.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营账号的增删查与数据同步。
一条贯穿全文件的规则:**cookie 是凭证,永远不出现在返回给上层的结构里。**
对外只给 `has_cookie` 这样的布尔量,与监控层对 cookie 的处理保持一致。
"""
import asyncio
from datetime import datetime, time, timedelta
from typing import Any, Dict, List, Optional
from sqlalchemy import delete, func, select
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from .client import CreatorApiError, CreatorClient
from .models import (
ACCOUNT_ERROR,
ACCOUNT_EXPIRED,
ACCOUNT_OK,
PERMISSION_ACTIVE,
PERMISSION_MISSING,
PERMISSION_PENDING,
PERMISSION_UNKNOWN,
CreatorAccount,
CreatorNoteStat,
)
# 同步一次最多翻多少页。后台默认一页 10 条,200 页足以覆盖任何正常账号,
# 同时防止"接口不返回 has_more"时无限翻下去。
MAX_SYNC_PAGES = 200
PAGE_SIZE = 10
def _account_dict(account: CreatorAccount, note_count: int = 0) -> Dict[str, Any]:
"""账号的对外表示。**刻意不含 cookie。**"""
return {
"id": account.id,
"nickname": account.nickname,
"user_id": account.user_id,
"red_id": account.red_id,
"avatar": account.avatar,
"status": account.status,
"permission_status": account.permission_status,
# 后台原话照抄。"次日可查看"这类信息只能由它自己说,改写就失真了。
"permission_tip": account.permission_tip,
"last_checked_at": account.last_checked_at,
"last_synced_at": account.last_synced_at,
"last_sync_days": account.last_sync_days,
"last_error": account.last_error,
"has_cookie": bool(account.cookie),
"created_at": account.created_at,
"note_count": note_count,
}
async def list_accounts(session: AsyncSession) -> List[Dict[str, Any]]:
accounts = list(
(await session.scalars(select(CreatorAccount).order_by(CreatorAccount.id))).all()
)
counts = dict(
(
await session.execute(
select(CreatorNoteStat.account_id, func.count(func.distinct(CreatorNoteStat.note_id)))
.group_by(CreatorNoteStat.account_id)
)
).all()
)
return [_account_dict(account, counts.get(account.id, 0)) for account in accounts]
async def get_account(session: AsyncSession, account_id: int) -> CreatorAccount:
account = await session.get(CreatorAccount, account_id)
if account is None:
raise ValueError(f"账号 {account_id} 不存在")
return account
async def account_detail(session: AsyncSession, account_id: int) -> Dict[str, Any]:
account = await get_account(session, account_id)
notes = await latest_notes(session, account_id)
return {
"account": _account_dict(account, len(notes)),
"notes": notes,
"summary": _summarize(notes),
}
def _summarize(notes: List[Dict[str, Any]]) -> Dict[str, Any]:
"""账号级汇总。取最后一轮快照的累计值之和。"""
totals = {
key: 0
for key in ("exposure", "views", "likes", "comments", "favorites", "shares", "new_followers")
}
for note in notes:
for key in totals:
value = note.get(key)
if isinstance(value, (int, float)):
totals[key] += int(value)
return totals
async def latest_notes(session: AsyncSession, account_id: int) -> List[Dict[str, Any]]:
"""每个作品取**最近一次**快照。
表里保留全部历史(换个时点就是一条新行),但列表只该展示"现在",否则同一个
作品会在列表里出现多次。
"""
newest = (
select(
CreatorNoteStat.note_id,
func.max(CreatorNoteStat.captured_at).label("captured_at"),
)
.where(CreatorNoteStat.account_id == account_id)
.group_by(CreatorNoteStat.note_id)
.subquery()
)
rows = (
await session.scalars(
select(CreatorNoteStat)
.join(
newest,
(CreatorNoteStat.note_id == newest.c.note_id)
& (CreatorNoteStat.captured_at == newest.c.captured_at),
)
.where(CreatorNoteStat.account_id == account_id)
# 不要用 nullslast():那是 PostgreSQL 语法,MySQL 5.7 会直接抛 1064 语法错误。
# MySQL 把 NULL 视为比任何值都小,所以 DESC 天然把未解析出发布时间的排在最后。
# 这个 bug 只在真机上才暴露 —— SQLite 从 3.30 起支持 NULLS LAST,测试环境测不出来。
.order_by(CreatorNoteStat.publish_time.desc())
)
).all()
return [_note_dict(row) for row in rows]
def _note_dict(row: CreatorNoteStat) -> Dict[str, Any]:
return {
"note_id": row.note_id,
"title": row.title,
"publish_time": row.publish_time,
"exposure": row.exposure,
"views": row.views,
"likes": row.likes,
"comments": row.comments,
"favorites": row.favorites,
"shares": row.shares,
"new_followers": row.new_followers,
"danmaku": row.danmaku,
"cover_ctr": row.cover_ctr,
"avg_watch_seconds": row.avg_watch_seconds,
"two_second_exit_rate": row.two_second_exit_rate,
"completion_rate": row.completion_rate,
"captured_at": row.captured_at,
}
async def upsert_account_from_cookie(session: AsyncSession, cookie: str) -> Dict[str, Any]:
"""用一份 cookie 识别并保存账号。
识别靠 `user/info` 而不是让用户填名字 —— 填错名字只会让后面所有数据对不上号。
已有同 `user_id` 的账号则更新它的 cookie(重新登录)。
"""
client = CreatorClient(cookie)
if not client.looks_authenticated:
raise ValueError("这份 cookie 里没有 a1,无法签名,请重新扫码")
try:
info = await client.fetch_user_info()
except CreatorApiError as exc:
raise ValueError(f"登录态无法使用:{exc}") from exc
if not info.get("user_id"):
raise ValueError("接口没有返回账号标识,可能登录态无效")
now = get_current_timestamp()
account = await session.scalar(
select(CreatorAccount).where(CreatorAccount.user_id == info["user_id"])
)
if account is None:
account = CreatorAccount(created_at=now)
session.add(account)
account.nickname = info.get("nickname") or account.nickname or "未命名账号"
account.user_id = info["user_id"]
account.red_id = info.get("red_id") or ""
account.avatar = info.get("avatar") or ""
account.cookie = cookie
account.status = ACCOUNT_OK
account.last_error = None
account.last_checked_at = now
account.updated_at = now
await session.flush()
# 顺手把权限状态也拉一次:新账号几乎必然处于"已申请、次日生效",
# 当场告诉用户,比让他明天再回来问要好。
await refresh_permission(session, account)
return _account_dict(account)
async def refresh_permission(session: AsyncSession, account: CreatorAccount) -> None:
"""查询并记录数据权限状态。失败不影响账号本身可用。"""
try:
permission = await CreatorClient(account.cookie).fetch_permission()
except CreatorApiError as exc:
if exc.status == 401:
account.status = ACCOUNT_EXPIRED
account.last_error = str(exc)
else:
account.last_error = str(exc)
account.updated_at = get_current_timestamp()
return
display = permission.get("display")
status = permission.get("status")
account.permission_tip = permission.get("tip") or ""
if display or status:
account.permission_status = PERMISSION_ACTIVE
elif account.permission_tip:
# 有提示语但未开通 —— 实测就是"已为您申请数据权限,次日可查看"。
account.permission_status = PERMISSION_PENDING
else:
account.permission_status = PERMISSION_MISSING
account.status = ACCOUNT_OK
account.last_error = None
account.last_checked_at = get_current_timestamp()
account.updated_at = account.last_checked_at
async def check_account(session: AsyncSession, account_id: int) -> Dict[str, Any]:
"""重新检测一个账号:登录态还在不在、权限开通没有。"""
account = await get_account(session, account_id)
await refresh_permission(session, account)
count = (
await session.execute(
select(func.count(func.distinct(CreatorNoteStat.note_id))).where(
CreatorNoteStat.account_id == account_id
)
)
).scalar() or 0
return _account_dict(account, count)
async def delete_account(session: AsyncSession, account_id: int) -> None:
account = await get_account(session, account_id)
await session.execute(
delete(CreatorNoteStat).where(CreatorNoteStat.account_id == account_id)
)
await session.delete(account)
def _day_bounds(days: int) -> tuple[int, int]:
"""最近 N 天的起止(毫秒)。与后台的按发布时间筛选对齐。"""
today = datetime.now()
end = int(datetime.combine(today.date(), time(23, 59, 59)).timestamp() * 1000)
start = int(
datetime.combine((today - timedelta(days=days)).date(), time(0, 0, 0)).timestamp() * 1000
)
return start, end
async def sync_account(
session: AsyncSession, account_id: int, days: int = 90
) -> Dict[str, Any]:
"""拉取一个账号的作品运营数据并落库。
权限未生效时接口返回的是**空壳成功**(`data.result` 里没有数据),不是错误 ——
所以"同步成功但 0 条"是正常结果,必须如实回报,不能让用户以为采集坏了。
"""
account = await get_account(session, account_id)
if not account.cookie:
raise ValueError("该账号没有可用的登录态,请重新扫码")
client = CreatorClient(account.cookie)
start_ms, end_ms = _day_bounds(days)
now = get_current_timestamp()
collected: List[Dict[str, Any]] = []
for page in range(1, MAX_SYNC_PAGES + 1):
try:
batch = await client.fetch_note_list(start_ms, end_ms, page_num=page, page_size=PAGE_SIZE)
except CreatorApiError as exc:
account.last_error = str(exc)
if exc.status == 401:
account.status = ACCOUNT_EXPIRED
account.updated_at = get_current_timestamp()
raise ValueError(f"同步失败:{exc}") from exc
collected.extend(batch)
if len(batch) < PAGE_SIZE:
break
await asyncio.sleep(0.6) # 对后台客气一点,这是自己的账号但仍是自动化访问
# 先删掉本时点可能存在的重复行,再写入 —— 表上有 (account, note, captured_at)
# 唯一索引,重复同步不该报错。
await session.execute(
delete(CreatorNoteStat).where(
CreatorNoteStat.account_id == account_id, CreatorNoteStat.captured_at == now
)
)
for note in collected:
if not note.get("note_id"):
continue
session.add(
CreatorNoteStat(
account_id=account_id,
note_id=note["note_id"],
title=note.get("title") or "",
publish_time=note.get("publish_time"),
exposure=note.get("exposure"),
views=note.get("views"),
likes=note.get("likes"),
comments=note.get("comments"),
favorites=note.get("favorites"),
shares=note.get("shares"),
new_followers=note.get("new_followers"),
danmaku=note.get("danmaku"),
cover_ctr=note.get("cover_ctr"),
avg_watch_seconds=note.get("avg_watch_seconds"),
two_second_exit_rate=note.get("two_second_exit_rate"),
completion_rate=note.get("completion_rate"),
captured_at=now,
)
)
account.last_synced_at = now
account.last_sync_days = days
account.last_error = None
account.updated_at = now
await refresh_permission(session, account)
await session.flush()
return {
"account_id": account_id,
"fetched": len(collected),
"days": days,
"permission_status": account.permission_status,
"permission_tip": account.permission_tip,
}
+107
View File
@@ -0,0 +1,107 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/signing.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""创作者后台的请求签名(XYW_ 方案)。
主站与创作者后台用的是**两套不同的签名**:主站是 VMP 的 `XYS_`,创作者后台是
`XYW_`。后者简单得多 —— MD5 → base64 → AES-128-CBC,密钥与 IV 都是硬编码常量,
纯 Python 可算,不需要浏览器。
常量与 `xhshow/config/config.py` 逐字节一致(该库也据此实现了 `sign_xyw`),
并与独立的逆向实现 xiaohongshu-cli/creator_signing.py 互相印证。
**三条实测结论**(tools/probe_creator_api.py 的 Phase 0 输出):
1. 待签字符串必须是 `url=` + 路径 + 查询串 的形式。只给路径、或去掉 `url=` 前缀,
网关一律返回 **406**;写法正确时签名通过。
2. `appId` 用 `ugc`(创作者平台的取值),不是主站的 `xhs-pc-web`。
3. 不带 cookie 时返回的是应用层的 401「无登录信息」而非 406 —— 说明签名每次都过了,
认证是独立的一层。
"""
import base64
import hashlib
import json
from datetime import datetime
XYW_AES_KEY = b"7cc4adla5ay0701v"
XYW_AES_IV = b"4uzjr7mbsibcaldp"
# 与 xhshow 的 XYW_ENV_FLAGS_DEFAULT 一致。含义未知,但改了签名就不被接受。
XYW_ENV_FLAGS = "0|0|0|1|0|0|1|0|0|0|1|0|0|0|0|1|0|0|0"
XYW_PREFIX = "XYW_"
XYW_SIGN_SVN = "56"
XYW_SIGN_TYPE = "x2"
XYW_SIGN_VERSION = "1"
# 创作者平台的 appId。用主站的 xhs-pc-web 会被拒。
CREATOR_APP_ID = "ugc"
def _aes_encrypt_hex(plaintext: str) -> str:
from Crypto.Cipher import AES
from Crypto.Util.Padding import pad
cipher = AES.new(XYW_AES_KEY, AES.MODE_CBC, XYW_AES_IV)
return cipher.encrypt(pad(plaintext.encode("utf-8"), AES.block_size)).hex()
def sign_xyw(
api: str,
a1: str,
app_id: str = CREATOR_APP_ID,
body: dict | None = None,
timestamp_ms: int | None = None,
) -> dict[str, str]:
"""为一次创作者后台请求生成 ``x-s`` / ``x-t`` 请求头。
``api`` 必须是待签的完整字符串:``url=`` 加路径,GET 请求还要带上查询串。
POST 的 JSON body 追加在其后(紧凑分隔符、不转义非 ASCII),与参考实现一致。
"""
content = api
if body is not None:
content += json.dumps(body, separators=(",", ":"), ensure_ascii=False)
if timestamp_ms is None:
timestamp_ms = int(datetime.now().timestamp() * 1000)
digest = hashlib.md5(content.encode("utf-8")).hexdigest()
plaintext = f"x1={digest};x2={XYW_ENV_FLAGS};x3={a1};x4={timestamp_ms};"
encoded = base64.b64encode(plaintext.encode("utf-8")).decode("utf-8")
envelope = {
"signSvn": XYW_SIGN_SVN,
"signType": XYW_SIGN_TYPE,
"appId": app_id,
"signVersion": XYW_SIGN_VERSION,
"payload": _aes_encrypt_hex(encoded),
}
x_s = XYW_PREFIX + base64.b64encode(
json.dumps(envelope, separators=(",", ":")).encode("utf-8")
).decode("utf-8")
return {"x-s": x_s, "x-t": str(timestamp_ms)}
def signed_api(url_path: str, query: str = "") -> str:
"""把路径与查询串拼成待签字符串。
单独抽出来是因为这个格式**没有文档**,只能靠实测固定下来 —— 写错就是 406,
而 406 的响应体 ``{"code":-1,"success":false}`` 完全看不出错在哪。
"""
return f"url={url_path}?{query}" if query else f"url={url_path}"
+20
View File
@@ -48,6 +48,7 @@ from .auth import ensure_initial_credential, require_auth
from .routers import (
auth_router,
crawler_router,
creator_router,
data_router,
monitor_router,
settings_router,
@@ -64,11 +65,23 @@ async def lifespan(_app: FastAPI):
browser session the way the log broadcaster is -- a scheduled run has to
happen whether or not anyone has the UI open.
"""
from .creator.login import shutdown as shutdown_creator_login
from .monitor.db import dispose_engine, init_db
from .monitor.qrlogin import shutdown as shutdown_qrlogin
from .monitor.scheduler import monitor_scheduler
await init_db()
# The WebUI bundle is gitignored and built separately, so a deployment that
# forgot it would otherwise come up looking healthy and serve a bare JSON
# stub at "/" -- worth one loud line at boot rather than a puzzled operator.
if not os.path.exists(os.path.join(WEBUI_DIR, "index.html")):
print(
"[综合采集平台] 警告:未找到前端产物 api/webui/index.html,"
"根路径只会返回一段 JSON。请先在 webui/ 下执行 npm run build。",
flush=True,
)
generated = await ensure_initial_credential()
if generated:
# Printed once, on the run that creates it. There is no unauthenticated
@@ -89,6 +102,12 @@ async def lifespan(_app: FastAPI):
yield
finally:
await monitor_scheduler.stop()
# Drops the tab a QR login may have opened and stops the Playwright
# client; leaving them would strand a driver process on every restart.
await shutdown_qrlogin()
# Same for the operator's account logins, which run in throwaway browser
# contexts -- those would otherwise be left open in the operator's Chrome.
await shutdown_creator_login()
await dispose_engine()
@@ -136,6 +155,7 @@ app.add_middleware(
# more importantly, never sees WebSocket scopes at all.
app.include_router(auth_router, prefix="/api")
app.include_router(crawler_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(creator_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(data_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(monitor_router, prefix="/api", dependencies=[Depends(require_auth)])
app.include_router(settings_router, prefix="/api", dependencies=[Depends(require_auth)])
+249
View File
@@ -0,0 +1,249 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/adapters.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""平台适配:两个平台之间**不一样**的那些管子。
监控层的大部分是平台中立的 —— 调度、入库、差分、报表、封面缓存、通知发送都与平台无关。
真正随平台变化的只有四样东西:
1. 爬虫把产物**落在哪个目录**(这里有个坑,见 ``artifact_dir``)
2. jsonl 里**字段叫什么**(抖音的作品没有 ``note_id``,叫 ``aweme_id``)
3. **目标链接**长什么样(怎么拼、怎么从链接里抠出 id)
4. 通知里的作品链接怎么拼
集中在这里,是为了让「加一个平台」变成在一处补一份数据,而不是去五个文件里找硬编码。
**为什么不放进 platforms.py**:那个模块被 ``describe_all()`` 整个序列化进
``GET /api/config/platforms`` 交给前端(连 ``**capability`` 一起),把正则、目录名、字段别名
塞进去会让爬虫的内部细节漏进 API 载荷,也会让「改适配」有动到接口形状的风险。
分工与既有的 schedule.py(算术)↔ scheduler.py(循环)一致。
"""
import re
from dataclasses import dataclass
from typing import Any, Dict, Mapping, Optional, Pattern, Tuple
from .platforms import PLATFORM_XHS
PLATFORM_DY = "dy"
def _first_cover(record: Dict[str, Any], fields: Tuple[str, ...]) -> str:
"""封面地址:取第一个非空字段,再取逗号分隔的第一段。
一条规则同时适配两边,所以不需要 per-platform 的函数:
小红书的 ``image_list`` 是 ``"url1,url2,..."``(要切第一段),
抖音的 ``cover_url`` 本身就是单个地址(切了等于没切)。
"""
for name in fields:
raw = record.get(name)
if raw:
return str(raw).split(",")[0].strip()
return ""
@dataclass(frozen=True)
class PlatformAdapter:
"""一个平台的全部「管子」。
字段别名的方向是**规范名 -> 该平台 jsonl 里的键**,读作「我们的列 ← 他们的键」。
"""
# 爬虫落盘用的目录名。**不等于平台 id**:抖音的平台 id 是 ``dy`` 而目录是 ``douyin``。
# 这不是笔误,是上游 store 里写死的(store/douyin/_store_impl.py:47)。改错这里的
# 后果是 ingest 一个文件都找不到 —— 它不会报错,只会落进「没抓到数据」分支,
# 然后被误报成「疑似登录失效」。
artifact_dir: str
web_base: str
creator_path: str
note_path: str
# 从链接里抠 id。是元组而不是单个正则,因为同一个平台可能有多种链接形态
# (抖音的作品链接还带 ?modal_id= 那种),按顺序试,第一个匹配的胜出。
# 每个正则必须恰好有一个捕获组。
creator_url_res: Tuple[Pattern, ...]
note_url_res: Tuple[Pattern, ...]
# 也允许直接粘贴裸 id —— 但两边的 id 形状不同,所以分开。
creator_bare_re: Pattern
note_bare_re: Pattern
# 短链(v.douyin.com 这种)无法在不发请求的情况下还原出 id,解析时单独报错,
# 好过存一个聚不出目标的值进去。
short_link_hosts: Tuple[str, ...]
note_fields: Mapping[str, str]
comment_fields: Mapping[str, str]
cover_fields: Tuple[str, ...]
# 时间戳换算成毫秒要乘的数。**小红书给毫秒、抖音给秒**,差 1000 倍;不换算的话
# 2026 年的作品会显示成 1970 年(实测踩到过:抖音作品发布日期显示 1970-01-22,
# 抖音评论的时间同理)。库里统一存毫秒,展示层才不用关心来源。
time_scale: int
def to_ms(self, value: Any) -> Optional[int]:
"""把平台的时间戳换算成毫秒;解析不出来返回 None(不伪造 0)。"""
try:
return int(value) * self.time_scale
except (TypeError, ValueError):
return None
def note_url(self, note_id: str) -> str:
"""作品的可点击链接。拼法与监控目标的链接是同一个形状 —— 通知里给的就是
人能直接点开看的那一个。"""
return f"{self.web_base}{self.note_path}/{note_id}"
def note_field(self, record: Dict[str, Any], name: str) -> Any:
"""按规范名读作品记录里的原始值(没有就是 None)。"""
return record.get(self.note_fields.get(name, name))
def comment_field(self, record: Dict[str, Any], name: str) -> Any:
return record.get(self.comment_fields.get(name, name))
def cover(self, record: Dict[str, Any]) -> str:
return _first_cover(record, self.cover_fields)
def parent_comment_id(self, record: Dict[str, Any]) -> str:
"""父评论 id,顶层评论一律归一成空串。
抖音顶层评论的 ``reply_id`` 是字符串 ``"0"``,小红书是 ``""`` —— 把 "0" 原样
存进去,前端就会多出一堆指向不存在的父评论的边。
"""
raw = self.comment_field(record, "parent_comment_id")
if raw is None:
return ""
raw = str(raw).strip()
return "" if raw in ("", "0") else raw
# 小红书 id 是 24 位 hex,允许稍宽一点,让格式变化退化成「仍然接受」而不是「拒绝」。
_XHS_BARE_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
XHS = PlatformAdapter(
artifact_dir="xhs",
web_base="https://www.xiaohongshu.com",
creator_path="/user/profile",
note_path="/explore",
creator_url_res=(re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)"),
),
creator_bare_re=_XHS_BARE_RE,
note_bare_re=_XHS_BARE_RE,
short_link_hosts=(),
note_fields={
"note_id": "note_id",
"title": "title",
"note_url": "note_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "type",
"published_at": "time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "note_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("image_list",),
# 小红书的时间戳本来就是毫秒(实测 time=1790923011000)。
time_scale=1,
)
# 抖音的 id 形状与小红书完全不同(见 media_platform/douyin/help.py:101-164):
# 作品 aweme_id 纯数字,如 7525082444551310602
# 博主 sec_user_id 形如 MS4wLjABAAAA...,含 - 和 _,**变长**(实测样本 55 字符,更长的也常见),
# 而小红书那条裸 id 规则封顶 64 —— 所以两条规则必须分开,否则长一点的 sec_uid
# 会被拒,表现为「粘贴了一个完全正确的链接却说无法识别」。
# 另外抖音**不需要 xsec_token**,裸链接就能用,比小红书简单。
DY = PlatformAdapter(
artifact_dir="douyin",
web_base="https://www.douyin.com",
creator_path="/user",
note_path="/video",
creator_url_res=(re.compile(r"douyin\.com/user/([A-Za-z0-9_-]+)"),),
note_url_res=(
re.compile(r"douyin\.com/video/(\d+)"),
# 带 modal_id 的链接:在别人主页或搜索结果里点开视频就是这个形态。
re.compile(r"[?&]modal_id=(\d+)"),
),
# 用长度而不是前缀来区分两者:sec_uid 是 20 字符以上的变长串,作品 id 是 19 位数字。
# 用前缀(MS4wLjABAAAA)更精确,但上游的 parse_creator_info_from_url 对裸 id 一律
# 照单全收,万一有别的前缀就会被我这里挡掉 —— 门槛设在长度上,两边都放得进,
# 又不会把 19 位的作品号误当成博主。
creator_bare_re=re.compile(r"^[A-Za-z0-9_-]{20,128}$"),
note_bare_re=re.compile(r"^\d{8,25}$"),
short_link_hosts=("v.douyin.com",),
note_fields={
"note_id": "aweme_id",
"title": "title",
"note_url": "aweme_url",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"source_kind": "aweme_type",
"published_at": "create_time",
},
comment_fields={
"comment_id": "comment_id",
"note_id": "aweme_id",
"content": "content",
"creator_hash": "creator_hash",
"creator_name": "nickname",
"create_time": "create_time",
"like_count": "like_count",
"sub_comment_count": "sub_comment_count",
"parent_comment_id": "parent_comment_id",
},
cover_fields=("cover_url",),
# 抖音给的是**秒**(实测 create_time=1790574515,即 2026-09-28)。
time_scale=1000,
)
ADAPTERS: Dict[str, PlatformAdapter] = {
PLATFORM_XHS: XHS,
PLATFORM_DY: DY,
}
class UnknownPlatformError(ValueError):
"""平台还没有适配器。"""
def adapter(platform: str) -> PlatformAdapter:
try:
return ADAPTERS[platform]
except KeyError as exc:
raise UnknownPlatformError(f"平台 {platform} 还没有适配器") from exc
def has_adapter(platform: str) -> bool:
return platform in ADAPTERS
def artifact_dir(platform: str) -> str:
"""该平台的爬虫会把 jsonl 落在哪个子目录下。
``runner`` 用它判断产物是否真的出现过,``ingest`` 用它定位文件 —— 两处必须用
同一个值,否则会出现「文件在,但两边找的目录不是同一个」这种最难查的错。
"""
return adapter(platform).artifact_dir
+68
View File
@@ -53,6 +53,7 @@ from .settings import (
set_setting,
system_key,
)
from .upstream import DEFAULT_BRANCH, DEFAULT_REMOTE_URL
SCOPE_PLATFORM = "platform"
SCOPE_SYSTEM = "system"
@@ -204,6 +205,73 @@ SETTING_SPECS: List[SettingSpec] = [
maximum=23,
affects_new_runs=False,
),
SettingSpec(
name="cdp_enabled",
scope=SCOPE_SYSTEM,
type=TYPE_BOOL,
label="接管已有 Chrome(CDP)",
help=(
"开启后爬虫不再自己启动浏览器,而是接管本机已开放远程调试端口的 Chrome"
"(默认 127.0.0.1:9222),复用它的登录态与扩展。"
"服务器部署请开启;本机桌面使用请保持关闭。"
),
default=False,
),
# --- 上游更新检查 -------------------------------------------------------
# 这几项不作用于采集,所以都标 affects_new_runs=False:改动它们不需要等下一轮,
# 也不影响采集命令的拼装。
SettingSpec(
name="upstream_check_enabled",
scope=SCOPE_SYSTEM,
type=TYPE_BOOL,
label="检查上游仓库更新",
help=(
"定期 fetch 上游仓库,看看它有没有新提交,并在有更新时推送通知。"
"本仓库在上游之上加了一整层(见 UPSTREAM.md),不定期看一眼就会越拖越难合并。"
),
default=False,
affects_new_runs=False,
),
SettingSpec(
name="upstream_check_interval_minutes",
scope=SCOPE_SYSTEM,
type=TYPE_INT,
label="上游检查间隔(分钟)",
help="默认 1440 分钟(每天一次)。检查只是 fetch,不需要太频繁。",
default=1440,
minimum=30,
maximum=10080,
affects_new_runs=False,
),
SettingSpec(
name="upstream_remote_url",
scope=SCOPE_SYSTEM,
type=TYPE_STR,
label="上游仓库地址",
help=(
"默认是 GitHub 上的上游。国内直连 GitHub 不稳时改成 gitcode 镜像"
"(见 UPSTREAM.md),或任意能访问到上游的地址。"
),
default=DEFAULT_REMOTE_URL,
affects_new_runs=False,
),
SettingSpec(
name="upstream_branch",
scope=SCOPE_SYSTEM,
type=TYPE_STR,
label="上游分支",
default=DEFAULT_BRANCH,
affects_new_runs=False,
),
SettingSpec(
name="upstream_notify",
scope=SCOPE_SYSTEM,
type=TYPE_BOOL,
label="上游有更新时推送通知",
help="只在出现此前没推过的上游提交时发一条,同一个更新不会反复推。",
default=True,
affects_new_runs=False,
),
]
SPECS_BY_NAME = {spec.name: spec for spec in SETTING_SPECS}
+162
View File
@@ -0,0 +1,162 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/covers.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""作品封面本地缓存。
**为什么必须落盘**:小红书图床的地址是**带签名、会过期**的。路径里那段时间戳就是
签发时刻,实测:
/202610080841/... (当天签发) → 200,且带不带 Referer 都 200
/202610070837/... (隔天) → 403,且带不带 Referer 都 403
所以这是**过期**,不是防盗链 —— 改 Referer 那一类修法治不了本。图一旦下载到本地,
就与签名无关,永远可读。
下载失败**不能影响采集**:一张封面拿不到,不该让整轮数据丢失。
"""
import re
from pathlib import Path
from typing import Optional
import httpx
from .db import DATA_DIR
COVERS_DIR = DATA_DIR / "covers"
# 单张封面的上限。正常封面是几十 KB;超过这个数说明拿到的不是图,
# 或者该放弃这一张而不是把内存撑爆。
MAX_COVER_BYTES = 5 * 1024 * 1024
USER_AGENT = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
)
_EXTENSIONS = {
"image/jpeg": ".jpg",
"image/jpg": ".jpg",
"image/png": ".png",
"image/webp": ".webp",
"image/gif": ".gif",
"image/heic": ".heic",
}
# note_id 是平台的稳定标识,但仍要挡住路径穿越 —— 它会直接变成文件名。
_SAFE_ID = re.compile(r"^[A-Za-z0-9_-]{1,64}$")
def is_safe_note_id(note_id: str) -> bool:
return bool(note_id) and bool(_SAFE_ID.match(note_id))
def cache_dir() -> Path:
COVERS_DIR.mkdir(parents=True, exist_ok=True)
return COVERS_DIR
def find_cached(note_id: str) -> Optional[Path]:
"""已缓存的封面文件,没有则 None。扩展名按内容类型而定,所以逐一试。"""
if not is_safe_note_id(note_id):
return None
for extension in sorted(set(_EXTENSIONS.values())):
candidate = COVERS_DIR / f"{note_id}{extension}"
if candidate.is_file():
return candidate
return None
async def cache_cover(note_id: str, url: str) -> Optional[str]:
"""下载并保存一张封面,返回文件名;失败返回 None。
**从不抛异常**:调用方是采集入库流程,一张图拿不到不该让整轮数据出问题。
"""
if not url or not is_safe_note_id(note_id):
return None
existing = find_cached(note_id)
if existing is not None:
return existing.name
try:
async with httpx.AsyncClient(timeout=20, follow_redirects=True) as client:
response = await client.get(url, headers={"user-agent": USER_AGENT})
except Exception:
return None
if response.status_code != 200:
# 403 通常意味着签名已过期 —— 这一张就没了,等下一轮采集拿到新地址。
return None
content = response.content
if not content or len(content) > MAX_COVER_BYTES:
return None
content_type = (response.headers.get("content-type") or "").split(";")[0].strip().lower()
extension = _EXTENSIONS.get(content_type, ".jpg")
# 图床偶尔不报 content-type,那种情况下扩展名只能猜,但文件本身仍然是好的。
target = cache_dir() / f"{note_id}{extension}"
try:
target.write_bytes(content)
except OSError:
return None
return target.name
def cover_url(note_id: str, remote: str) -> str:
"""前端该用哪个地址。
本地有缓存就用自己的接口 —— 那是唯一不会过期的地址。没有就退回远程地址,
至少让图先显示出来(哪怕它很快会失效)。
"""
if find_cached(note_id) is not None:
return f"/api/monitor/covers/{note_id}"
return remote
async def cache_pending(session, task_id: int, limit: int = 60) -> int:
"""把还没有本地副本的封面补下来,返回本次下载成功的张数。
由 runner 在入库之后调用,**而不是在 ingest 里** —— ingest 是刻意保持离线的
(它的文档写明 No network),往里塞网络请求会毁掉这一点。
每轮只补一批:一次跑几百张图既慢又会给图床压力,而旧地址本来就在陆续过期,
分摊到几轮里补完反而更稳。
"""
from sqlalchemy import select
from .models import MonitorNote
notes = (
await session.scalars(
select(MonitorNote)
.where(MonitorNote.task_id == task_id, MonitorNote.cover != "")
.order_by(MonitorNote.last_seen_at.desc())
.limit(limit)
)
).all()
saved = 0
for note in notes:
if find_cached(note.note_id) is not None:
continue
if await cache_cover(note.note_id, note.cover):
saved += 1
return saved
+63 -14
View File
@@ -44,7 +44,9 @@ from contextlib import asynccontextmanager
from pathlib import Path
from typing import AsyncIterator, Optional
from sqlalchemy import event, text
from sqlalchemy import Column, event, text
from sqlalchemy.dialects import mysql
from sqlalchemy.schema import CreateColumn
from sqlalchemy.ext.asyncio import (
AsyncEngine,
AsyncSession,
@@ -214,14 +216,53 @@ async def init_db() -> None:
await _migrate_setting_keys(conn)
# Columns added to a table after it may already exist. ``create_all`` only
# creates missing *tables*, so new columns need an explicit ALTER TABLE.
_ADDED_COLUMNS: dict[str, list[tuple[str, str]]] = {
"monitor_task": [
("notify_enabled", "BOOLEAN NOT NULL DEFAULT 0"),
("last_notified_at", "BIGINT NULL"),
],
}
# ``create_all`` creates missing *tables* but never adds *columns* to a table that
# already exists, so those need an explicit ALTER TABLE.
#
# Which columns those are is derived from the ORM metadata, not kept by hand. The
# hand-kept version was a trap: forgetting to register a new column there still let
# the app start -- it connects fine, then fails on every query and every scheduler
# tick. Which is exactly what happened when the scheduling columns were added.
def _implicit_default(column: Column) -> Optional[str]:
"""A literal to seed existing rows with when a NOT NULL column is added."""
default = column.default
if default is not None and getattr(default, "is_scalar", False):
value = default.arg
if isinstance(value, bool):
return "1" if value else "0"
if isinstance(value, (int, float)):
return str(value)
return "'" + str(value).replace("'", "''") + "'"
# No scalar default on the model. Fall back to the type's zero value, so that
# adding the column cannot depend on the server's sql_mode.
try:
python_type = column.type.python_type
except NotImplementedError:
return None
if python_type in (bool, int, float):
return "0"
if python_type is str:
return "''"
return None
def _column_ddl(column: Column) -> str:
"""One column as MySQL DDL for ``ALTER TABLE ... ADD COLUMN``.
``CreateColumn`` renders the name, type and nullability. The default is added
separately because a model's ``default=`` is applied by the ORM and never
reaches the DDL -- and a NOT NULL column added to a populated table needs a
value for the rows already sitting there.
"""
ddl = str(CreateColumn(column).compile(dialect=mysql.dialect()))
if not column.nullable and column.server_default is None:
seed = _implicit_default(column)
if seed is not None:
ddl += f" DEFAULT {seed}"
return ddl
async def _existing_columns(conn, table: str) -> set[str]:
@@ -240,14 +281,22 @@ async def _existing_columns(conn, table: str) -> set[str]:
async def _ensure_columns(conn) -> None:
for table, columns in _ADDED_COLUMNS.items():
existing = await _existing_columns(conn, table)
"""Add every model column the live table is missing."""
for table in MonitorBase.metadata.sorted_tables:
existing = await _existing_columns(conn, table.name)
if not existing:
# Table did not exist before this run; create_all built it complete.
continue
for name, ddl in columns:
if name not in existing:
await conn.execute(text(f"ALTER TABLE {table} ADD COLUMN {name} {ddl}"))
for column in table.columns:
# Primary keys are always present, and MySQL rejects AUTO_INCREMENT
# alongside the DEFAULT this helper appends -- so skip them rather
# than emit DDL that could never run.
if column.name in existing or column.primary_key:
continue
print(f"[monitor.db] 补齐缺失字段 {table.name}.{column.name}", flush=True)
await conn.execute(
text(f"ALTER TABLE {table.name} ADD COLUMN {_column_ddl(column)}")
)
async def _migrate_setting_keys(conn) -> None:
+561
View File
@@ -0,0 +1,561 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_api.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""抖音 Web 接口客户端 —— 直接发 HTTP,不起爬虫子进程。
**为什么另起一套。** 爬虫那条路(``media_platform/douyin``)会构造一大串浏览器指纹
参数:``browser_platform=MacIntel``、``os_name=Mac OS``、``browser_version=125.0.0.0``……
而 ``User-Agent`` 是从页面现读的(在服务器上是 Linux + Chrome 155)。参数说自己是 Mac,
UA 说自己是 Linux —— 抖音网关对这种自相矛盾的请求的处理方式是:**不报错、不给原因,
回一个 200 + 空 body**。爬虫那边把它翻译成 ``Exception("account blocked")``,看起来像
账号被封,其实什么都不是。
这份客户端只发必要参数(``device_platform`` / ``aid`` 那两三个),走浏览器自己也在用的
那条调用路径。它的做法来自 mac-agent-os 项目的 ``mediacrawler_adapter.py``,实测可用。
两个关键点:
* **cookie 走 CDP 现读。** Chrome 把 cookie 值加密存在 SQLite 里,只有 CDP 拿得到
解密后的值;而且浏览器里那份比库里存的旧快照新 —— 站点会自己轮换会话。
* **产物形状照抄 store。** ``aweme_id`` / ``aweme_url`` / ``cover_url`` / ``aweme_type`` /
``create_time``(**秒**,由 adapters 换算成毫秒)…… 这样 ingest 那条链路一个字都不用改。
"""
import asyncio
import os
import time
from dataclasses import dataclass
from typing import Any, Dict, List, Optional, Tuple
from urllib.parse import urlencode
import config
import httpx
from tools import utils
from tools.user_hash import anonymize_user_id
# 请求头。**要像一个浏览器**,而且必须是**同一个浏览器**:见 BrowserIdentity。
_BASE_HEADERS = {
"Accept": "application/json, text/plain, */*",
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
"Referer": "https://www.douyin.com/",
"Origin": "https://www.douyin.com",
}
# 网关的业务前置校验头。缺了它,抖音边缘网关的 ArgusSecurityPlugin 会直接回
# 403 并写明 "Blocked by ArgusSecurityPlugin Uifid Not Found" —— 难得一次它会说原因。
# 当前网关并不校验这个头的**值**,填什么都行;一旦升级到真校验,就得改成让页面里的
# SDK 自己生成(见 media_platform/douyin/client.py 里同一条注释)。
ARGUS_HEADER_VALUE = "1"
API_ORIGIN = "https://www.douyin.com"
PROFILE_PATH = "/aweme/v1/web/user/profile/other/"
POSTS_PATH = "/aweme/v1/web/aweme/post/"
DETAIL_PATH = "/aweme/v1/web/aweme/detail/"
COMMENT_PATH = "/aweme/v1/web/comment/list/"
# 一次请求的超时。抖音这两个接口正常都在一秒内返回。
REQUEST_TIMEOUT_SECONDS = 20.0
# 问浏览器要 UA / client hints 的超时。**这个必须有。**
# ``page.evaluate`` 打在一个渲染进程已经卡住的标签页上会**永远不返回**,而问身份是采集的
# 第一步 —— 它一挂,整个 run 就永远停在「运行中」(真踩过:标签页 URL 是空的,
# cookies() 正常,evaluate 一直不回来)。
EVALUATE_TIMEOUT_SECONDS = 8.0
# 单页最多要多少条。接口自己有上限,要多了也没用。
MAX_PAGE_SIZE = 20
class DouyinApiError(RuntimeError):
"""请求失败,或登录态不可用。"""
def _cdp_url() -> str:
"""浏览器 DevTools 端点。与扫码登录那边共用同一个开关。"""
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
@dataclass
class BrowserIdentity:
"""一个请求要像浏览器所需要的全部身份信息,**且必须来自同一个浏览器**。
只拿 cookie 是不够的。UA 声称自己是 Chrome 155、却不带 Chrome 155 该有的
``sec-ch-ua``,网关一眼就能看出这不是浏览器 —— 它的回应是 **200 + 空 body**:
不报错、不给原因,只看得到「抓到 0 条」。所以这三样必须成套地从同一处取。
"""
cookie: str
user_agent: str
client_hints: Dict[str, str]
def headers(self) -> Dict[str, str]:
headers = {
"User-Agent": self.user_agent,
**self.client_hints,
**_BASE_HEADERS,
"x-tt-argus": ARGUS_HEADER_VALUE,
"Cookie": self.cookie,
}
# uifid 是设备标识,网关要它;cookie 里没有就不带(送空值反而更像异常请求)。
uifid = _cookie_value(self.cookie, "UIFID") or _cookie_value(
self.cookie, "UIFID_TEMP"
)
if uifid:
headers["uifid"] = uifid
return headers
# 身份信息的短时缓存:一次采集要发好几个请求,没必要每次都连一遍 CDP。
_IDENTITY_TTL_SECONDS = 120.0
_identity_cache: Optional[Tuple[float, BrowserIdentity]] = None
async def _safe_evaluate(page: Any, expression: str) -> Any:
"""在页面上求值,带超时;任何失败都返回 None。
**不要直接调 ``page.evaluate``** —— 在渲染进程卡住的标签页上它会永远不返回(见
``EVALUATE_TIMEOUT_SECONDS`` 那段)。
"""
try:
return await asyncio.wait_for(
page.evaluate(expression), timeout=EVALUATE_TIMEOUT_SECONDS
)
except Exception:
return None
async def _identity_from_pages(context: Any) -> Tuple[str, Dict[str, str]]:
"""问出 UA 和 client hints。
不假设第一个标签页是好的 —— 它可能停在 URL 为空、渲染进程已卡住的状态(实测过)。
所以逐个试、每个都带超时;优先抖音页面,全都不行就临时开一个干净页问完关掉。
拿不到就返回空 —— 调用方据此退回库里那份 cookie,而不是拿一组编出来的指纹去请求
(那比没有更糟,见 BrowserIdentity 的说明)。
"""
from media_platform.douyin.help import client_hint_headers
pages = list(context.pages)
pages.sort(key=lambda page: 0 if "douyin" in (page.url or "") else 1)
for page in pages:
user_agent = await _safe_evaluate(page, "() => navigator.userAgent")
if user_agent:
hints = client_hint_headers(
await _safe_evaluate(page, "() => navigator.userAgentData || null")
)
return user_agent, hints or {}
temp = None
try:
temp = await asyncio.wait_for(
context.new_page(), timeout=EVALUATE_TIMEOUT_SECONDS
)
user_agent = await _safe_evaluate(temp, "() => navigator.userAgent")
hints = client_hint_headers(
await _safe_evaluate(temp, "() => navigator.userAgentData || null")
)
return user_agent or "", hints or {}
except Exception:
return "", {}
finally:
if temp is not None:
try:
await temp.close()
except Exception:
pass
async def _read_browser() -> Optional[BrowserIdentity]:
"""连上 CDP 浏览器,一次取齐 cookie、UA、client hints。
读不到返回 None(浏览器没开/没登录),由调用方决定怎么报 —— 不抛异常。
"""
from playwright.async_api import async_playwright
from media_platform.douyin.help import client_hint_headers
playwright = None
try:
playwright = await async_playwright().start()
browser = await playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
if not browser.contexts:
return None
# contexts[0] 是真实 profile。**不要 new_context()** —— 那是无痕式的,读不到登录态。
context = browser.contexts[0]
cookies = await asyncio.wait_for(
context.cookies(), timeout=EVALUATE_TIMEOUT_SECONDS
)
# UA 和 hints 要从页面里问 —— 它们是浏览器自己的事实,写死迟早对不上。
user_agent, hints = await _identity_from_pages(context)
except Exception as exc:
utils.logger.warning(f"[douyin_api] 读浏览器身份失败:{exc}")
return None
finally:
if playwright is not None:
# 只断开连接。**绝不能 browser.close()** —— 对这个 CDP 连接而言那会关掉
# 操作者自己的浏览器。
try:
await playwright.stop()
except Exception:
pass
douyin_cookies = {
cookie["name"]: cookie["value"]
for cookie in cookies
if "douyin" in cookie.get("domain", "") or "amemv" in cookie.get("domain", "")
}
return BrowserIdentity(
cookie=_cookie_from_dict(douyin_cookies),
user_agent=user_agent or "",
client_hints=hints,
)
async def browser_identity(cookie: str = "", force: bool = False) -> BrowserIdentity:
"""拿到一份可用的身份:**优先浏览器里那份**,其次退回传进来的 cookie(库里存的)。
优先浏览器的原因:站点会自己轮换会话,库里存的是粘贴那一刻的快照,浏览器里那份才是
当前有效的;而 UA/hints 更是只有浏览器自己知道。
"""
global _identity_cache
now = time.monotonic()
if not force and _identity_cache is not None:
cached_at, cached = _identity_cache
if now - cached_at < _IDENTITY_TTL_SECONDS:
return cached
identity = await _read_browser()
if identity is None or not _has_session(identity.cookie):
# 浏览器里没有可用会话,退回调用方给的那份。UA/hints 编不出来就不编 ——
# 一组和 UA 对不上的 hints 比没有更糟。
identity = BrowserIdentity(
cookie=_cookie_header(cookie), user_agent="", client_hints={}
)
_identity_cache = (now, identity)
return identity
def forget_identity() -> None:
"""丢掉缓存的身份。cookie 变了、或测试之间要隔离时调用。"""
global _identity_cache
_identity_cache = None
def _cookie_header(cookie: str) -> str:
"""把 ``a=1; b=2`` 形式的 cookie 串规整成请求头用的形状。"""
pairs = []
for part in (cookie or "").split(";"):
if "=" in part:
name, _, value = part.partition("=")
name = name.strip()
if name:
pairs.append(f"{name}={value.strip()}")
return "; ".join(pairs)
def _cookie_from_dict(cookies: Dict[str, str]) -> str:
return "; ".join(f"{name}={value}" for name, value in cookies.items())
def _cookie_value(cookie: str, name: str) -> str:
"""从一个 cookie 串里取某个键的值。"""
for part in (cookie or "").split(";"):
key, _, value = part.partition("=")
if key.strip() == name:
return value.strip()
return ""
def _sign(params: Dict[str, Any], path: str, user_agent: str) -> Dict[str, Any]:
"""给一组参数补上 ``a_bogus`` 签名,返回新 dict。
**按需 import**:那个模块在 import 的那一瞬间就把 ``libs/douyin.js`` 交给 execjs
编译(还要读相对路径),把它拖进监控层的热路径不合适。
签名算在**不含 a_bogus 的那串 query 上**,追加到末尾 —— 和爬虫那条路一致,也是
实测能过的形态。
"""
from media_platform.douyin.help import get_a_bogus_from_js
try:
return {
**params,
"a_bogus": get_a_bogus_from_js(path, urlencode(params), user_agent),
}
except Exception as exc: # execjs 起不来 / JS 抛错,都算签名失败
raise DouyinApiError(f"算 a_bogus 签名失败:{exc}") from exc
async def _get(
path: str,
params: Dict[str, Any],
identity: BrowserIdentity,
*,
signed: bool = False,
) -> Dict[str, Any]:
"""发一个 GET,返回 JSON。
只带调用方给的参数 —— **不要往里加 webid / msToken / browser_version 那一堆**,
那正是爬虫那条路失败的原因。
``signed=True`` 时补一个 ``a_bogus``。**只有评论接口需要它**:作品、详情、博主资料
三个不带签名也照常返回,而给它们加签名是没验证过的改动,不做。
"""
if signed:
params = _sign(params, path, identity.user_agent)
url = f"{API_ORIGIN}{path}"
async with httpx.AsyncClient(timeout=REQUEST_TIMEOUT_SECONDS) as client:
response = await client.get(
url,
params=params,
headers=identity.headers(),
)
if response.status_code != 200:
raise DouyinApiError(f"HTTP {response.status_code}:{response.text[:120]}")
# 「200 + 空 body」是抖音网关拒绝请求时的典型回应(见模块说明)。必须当成错误报出来,
# 否则会一路往下变成「这个博主没作品」。
if not response.text.strip():
raise DouyinApiError(
"接口返回了空内容 —— 通常是登录态失效,或请求被网关判成了非浏览器"
)
try:
return response.json()
except ValueError as exc:
raise DouyinApiError(f"返回的不是 JSON:{response.text[:120]}") from exc
def _as_int(value: Any) -> int:
try:
return int(value)
except (TypeError, ValueError):
return 0
def normalize_aweme(aweme: Dict[str, Any]) -> Dict[str, Any]:
"""把接口返回的一条作品,翻译成 store 落盘的那套键名。
键名必须和 ``store/douyin`` 一致 —— 跨过这一层之后,ingest 就不知道数据是从爬虫
来的还是从接口来的。
"""
author = aweme.get("author") or {}
statistics = aweme.get("statistics") or {}
aweme_id = str(aweme.get("aweme_id") or "")
cover = ((aweme.get("video") or {}).get("cover") or {}).get("url_list") or [""]
uid = str(author.get("uid") or "")
nickname = author.get("nickname") or ""
return {
"aweme_id": aweme_id,
"aweme_type": str(aweme.get("aweme_type") or ""),
# store 那边 title 取的是 desc。
"title": aweme.get("desc") or "",
"desc": aweme.get("desc") or "",
# **秒**。adapters.time_scale 会把它换成毫秒,和 store 写出来的形态一致。
"create_time": _as_int(aweme.get("create_time")),
"creator_hash": anonymize_user_id(uid or author.get("sec_uid") or ""),
"nickname": nickname,
"liked_count": str(_as_int(statistics.get("digg_count"))),
"comment_count": str(_as_int(statistics.get("comment_count"))),
"collected_count": str(_as_int(statistics.get("collect_count"))),
"share_count": str(_as_int(statistics.get("share_count"))),
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
"cover_url": cover[0] if cover else "",
"source_keyword": "",
}
async def author_videos(
sec_user_id: str, count: int = MAX_PAGE_SIZE, *, cookie: str = ""
) -> List[Dict[str, Any]]:
"""某个博主最新发布的作品(按发布时间倒序),已翻译成 store 的键名。
用 ``sec_user_id`` 而不是数字 uid:监控任务里存的就是主页链接里的那段 sec_uid,
而且这个接口两种都收(爬虫那边用的也是 sec_user_id)。
"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
payload = await _get(
POSTS_PATH,
{
"sec_user_id": sec_user_id,
"count": max(1, min(count, MAX_PAGE_SIZE)),
"max_cursor": 0,
"device_platform": "webapp",
"aid": 6383,
},
identity,
)
awemes = payload.get("aweme_list") or []
if not awemes and payload.get("status_code") not in (0, None):
raise DouyinApiError(
f"接口拒绝了请求(status_code={payload.get('status_code')})"
)
return [normalize_aweme(aweme) for aweme in awemes]
async def video_detail(aweme_id: str, *, cookie: str = "") -> Dict[str, Any]:
"""单条作品的详情,已翻译成 store 的键名。
这个接口**没有**被那道真校验挡着(实测 200 / 45425 字节),所以在拿不到作品列表时,
它是「刷新已知作品指标」的唯一途径。
"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
payload = await _get(
DETAIL_PATH,
{"aweme_id": aweme_id, "device_platform": "webapp", "aid": 6383},
identity,
)
aweme = payload.get("aweme_detail") or {}
if not aweme:
raise DouyinApiError(
f"接口没返回作品(status_code={payload.get('status_code')})"
)
return normalize_aweme(aweme)
async def author_profile(sec_user_id: str, *, cookie: str = "") -> Dict[str, Any]:
"""博主主页指标:昵称 / 粉丝数 / 总获赞 / 作品数。"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
payload = await _get(
PROFILE_PATH,
{"sec_user_id": sec_user_id, "device_platform": "webapp", "aid": 6383},
identity,
)
user = payload.get("user") or {}
if not user:
raise DouyinApiError(
f"接口没返回用户数据(status_code={payload.get('status_code')})"
)
return {
# 自报家门。快照表的唯一键是 (任务, creator_hash, 轮次),而作品是靠
# `anonymize_user_id(author.uid)` 得到这个哈希的 —— 这里走同一条路,两边才对得上,
# 否则快照会和作品分成两个人,界面上永远查不到。
"creator_hash": anonymize_user_id(
str(user.get("uid") or user.get("sec_uid") or "")
),
"nickname": user.get("nickname") or "",
"unique_id": user.get("unique_id") or "",
"fans": _as_int(user.get("follower_count")),
"total_favorited": _as_int(user.get("total_favorited")),
"works": _as_int(user.get("aweme_count")),
"following": _as_int(user.get("following_count")),
}
def normalize_comment(comment: Dict[str, Any], aweme_id: str) -> Dict[str, Any]:
"""把接口返回的一条评论,翻译成 store 落盘的那套键名。
HTTP 路线和页面路线共用它 —— 同一套键名,ingest 才不用关心数据是怎么来的。
刻意**不带** ``sub_comment_count`` / ``parent_comment_id`` 的猜测值:接口给了就用,
没给就留空,不编。
"""
user = comment.get("user") or {}
return {
"comment_id": str(comment.get("cid") or ""),
"aweme_id": aweme_id,
"content": comment.get("text") or "",
"nickname": user.get("nickname") or "",
"creator_hash": anonymize_user_id(
str(user.get("uid") or user.get("sec_uid") or "")
),
# 同为秒;adapters 会换算。
"create_time": _as_int(comment.get("create_time")),
"like_count": str(_as_int(comment.get("digg_count"))),
"sub_comment_count": str(_as_int(comment.get("reply_comment_total"))),
# 顶层评论在抖音里是 "0";adapters.parent_comment_id 会归一成空串。
"parent_comment_id": str(comment.get("reply_id") or "0"),
}
async def video_comments(
aweme_id: str, count: int = 20, *, cookie: str = ""
) -> List[Dict[str, Any]]:
"""一条作品的评论,翻译成 store 的评论键名。
刻意不带 ``sub_comment_count`` / ``parent_comment_id`` 的猜测值 —— 接口给了就用,
没给就留空,不编。
"""
identity = await browser_identity(cookie)
if not _has_session(identity.cookie):
raise DouyinApiError("抖音登录态不可用")
payload = await _get(
COMMENT_PATH,
{
"aweme_id": aweme_id,
"count": max(1, min(count, MAX_PAGE_SIZE)),
"cursor": 0,
"device_platform": "webapp",
"aid": 6383,
},
identity,
# 这个接口**必须**签名。不签的话网关回 200 + 空 body,会被读成「这条没评论」,
# 而它其实只是被挡了 —— 和登录失效长得一模一样。(实测:带上签名 200/9960 字节
# 真评论,不带就是空的。)
signed=True,
)
records = [
normalize_comment(comment, aweme_id)
for comment in payload.get("comments") or []
]
return records
def _has_session(cookie: str) -> bool:
return "sessionid=" in (cookie or "")
async def check_login(cookie: str = "") -> Dict[str, Any]:
"""浏览器/库里现在有没有可用的抖音登录态。给设置页用。"""
identity = await browser_identity(cookie)
if _has_session(identity.cookie):
source = "browser" if identity.user_agent else "stored"
return {"ok": True, "source": source, "cookie_length": len(identity.cookie)}
return {"ok": False, "source": "", "cookie_length": 0}
async def main() -> None: # pragma: no cover - 手工排查用
"""``python -m api.monitor.douyin_api <sec_user_id>``"""
import sys
if len(sys.argv) < 2:
print(await check_login())
return
sec = sys.argv[1]
print(await author_profile(sec))
for record in await author_videos(sec, count=5):
print(record["create_time"], record["title"][:30], record["liked_count"])
if __name__ == "__main__": # pragma: no cover
asyncio.run(main())
+217
View File
@@ -0,0 +1,217 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_fetch.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""抖音采集:直接走 Web 接口,不起爬虫子进程。
与 ``media_platform/douyin`` 那条路的分工:
* 那边起一个 Playwright 子进程、构造一大串**自相矛盾的浏览器指纹参数**(参数说自己是
Mac + Chrome 125,UA 说自己是 Linux + Chrome 155),网关回一个 200 + 空 body,
然后被翻译成 ``Exception("account blocked")`` —— 看起来像账号被封,其实什么都不是。
* 这边不起子进程、只发必要参数,用浏览器里那份登录态。见 ``douyin_api``。
采到的东西写成 ``store/douyin`` 那套 jsonl 形状,所以 **ingest 完全不知道数据是怎么来的**
—— 重采样、差分、事件、通知、报表全都照旧。
"""
import json
from datetime import datetime
from pathlib import Path
from typing import Any, Dict, Iterable, List, Optional, Sequence
from tools import utils
from . import adapters, douyin_api
from .models import MODE_CREATOR, MODE_NOTE, MonitorTask
# 拉多少条作品。接口单页上限就是 20,要多了也没用。
DEFAULT_VIDEO_LIMIT = 20
async def collect(
out_dir: Path,
*,
platform: str,
mode: str,
limit: int,
want_comments: bool,
comment_limit: int,
targets: Sequence[Any],
known_aweme_ids: Iterable[str] = (),
cookie: str = "",
) -> Dict[str, Any]:
"""跑一轮抖音采集,把产物写进 ``out_dir`` 下 store 那个目录里。
参数是散的、不收 ORM 对象:调用方那边 ``task`` 在会话关掉之后就 detached 了,
传对象进来迟早会踩到「属性已过期」。
**不抛异常**:失败也把(可能为空的)产物落下去,并把原因放进 ``errors`` 交给调用方。
让调用方去决定这是「一轮正常但没数据」还是「一轮失败」—— 这个判断不该藏在这里。
"""
notes: List[Dict[str, Any]] = []
comments: List[Dict[str, Any]] = []
# 博主**账号级**指标(粉丝 / 总获赞 / 作品数)。作品列表之外单独要一次,
# 只有博主模式才有 —— 作品模式的目标是一件作品,没有"这个博主是谁"可问。
profiles: List[Dict[str, Any]] = []
errors: List[str] = []
# **整个 collect 只去重一次的、跨目标的集合**:退化路径会把「库里已知的全部作品」
# 在每个目标下都刷一遍,多个目标就会出现同一件作品好几条记录 —— 而一对一快照的
# 唯一键是 (task_id, note_id, run_id),同一条作品在一轮里出现两次会直接撞键。
seen_aweme: set = set()
for target in targets:
external_id = target.external_id
# 两种模式的目标是不同的东西,不能走同一条路:
# 作品模式 —— 目标本身就是作品 id,直接取详情(**这个接口没被挡,今天就能用**)。
# 博主模式 —— 目标是主页 sec_uid,要先拉作品列表;那个接口被真校验挡着,退化到
# 刷新库里已知的作品(新作品发现不了)。
if mode == MODE_NOTE:
try:
videos = [await douyin_api.video_detail(external_id, cookie=cookie)]
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取作品 {external_id} 失败:{exc}")
videos = []
else:
videos = await _creator_works(
external_id, limit, known_aweme_ids, seen_aweme, cookie, errors
)
profile = await _creator_profile(external_id, videos, cookie, errors)
if profile is not None:
profiles.append(profile)
for video in videos:
aweme_id = video.get("aweme_id")
if not aweme_id or aweme_id in seen_aweme:
continue
seen_aweme.add(aweme_id)
notes.append(video)
if want_comments:
try:
comments.extend(
await douyin_api.video_comments(
aweme_id, count=comment_limit, cookie=cookie
)
)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取作品 {aweme_id} 的评论失败:{exc}")
jsonl_dir = _write_artifacts(out_dir, platform, mode, notes, comments, profiles)
return {
"notes": len(notes),
"comments": len(comments),
"errors": errors,
"jsonl_dir": str(jsonl_dir),
}
async def _creator_profile(
sec_user_id: str,
videos: Sequence[Dict[str, Any]],
cookie: str,
errors: List[str],
) -> Optional[Dict[str, Any]]:
"""问一次博主的账号级指标。拿不到就算了 —— **不能因为顺手的附加信息失败,
就把这一轮本来采到的作品也判成失败。**
creator_hash 优先取作品自带的那个:作品是靠 ``anonymize_user_id(author.uid)`` 得到
哈希的,而快照表和作品必须对得上号,否则界面上永远查不出这个博主的粉丝数。只有当一件
作品都没采到时(列表被挡且没有已知作品可刷新),才退回资料接口自己算的哈希 ——
那种情况下也只剩它了。
"""
try:
profile = await douyin_api.author_profile(sec_user_id, cookie=cookie)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取博主 {sec_user_id} 的资料失败:{exc}")
return None
if videos:
profile["creator_hash"] = videos[0].get("creator_hash") or profile["creator_hash"]
if not profile.get("creator_hash"):
# 哈希都算不出来的快照没人能查到,落下去只是垃圾。
errors.append(f"博主 {sec_user_id} 的资料里没有可用的身份标识,跳过账号指标")
return None
return profile
async def _creator_works(
sec_user_id: str,
limit: int,
known_aweme_ids: Iterable[str],
seen_aweme: set,
cookie: str,
errors: List[str],
) -> List[Dict[str, Any]]:
"""一个博主的作品:先要列表,列表被挡时退化成刷新已知作品。
作品列表(``aweme/post``)被抖音单独加了真校验 —— 不带 ``x-tt-argus`` 回 403,
带上 dummy 值回 200 + 空 body。所以这里拿不到**新**作品,只能保住已知的。
"""
try:
return await douyin_api.author_videos(sec_user_id, count=limit, cookie=cookie)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取博主 {sec_user_id} 的作品列表失败:{exc}")
refreshed: List[Dict[str, Any]] = []
for aweme_id in known_aweme_ids:
if aweme_id in seen_aweme:
continue
try:
refreshed.append(await douyin_api.video_detail(aweme_id, cookie=cookie))
except douyin_api.DouyinApiError as detail_exc:
errors.append(f"刷新作品 {aweme_id} 失败:{detail_exc}")
return refreshed
def _write_artifacts(
out_dir: Path,
platform: str,
mode: str,
notes: List[Dict[str, Any]],
comments: List[Dict[str, Any]],
profiles: Sequence[Dict[str, Any]] = (),
) -> Path:
"""按爬虫那套目录与文件名写 jsonl。
目录名取 ``adapters.artifact_dir``(抖音是 ``douyin``,而不是平台 id ``dy``)——
和 ingest 找文件用的是同一个来源,两边不会走散。
"""
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
jsonl_dir.mkdir(parents=True, exist_ok=True)
kind = "creator" if mode == MODE_CREATOR else "detail"
date = datetime.now().strftime("%Y-%m-%d")
_write_jsonl(jsonl_dir / f"{kind}_contents_{date}.jsonl", notes)
# 评论文件即使没有评论也建出来:ingest 靠「文件在不在」区分「这一轮没评论」和
# 「这一轮什么都没抓到」,两种情况的含义完全不同。
_write_jsonl(jsonl_dir / f"{kind}_comments_{date}.jsonl", comments)
# 博主资料同样无条件写:空文件表示"问了但没问到",没有文件表示"这次根本没问"
# (作品模式)。两者在 ingest 那边走的是同一条路(都不落快照),但留空文件能让
# 事后翻 run 目录时看出到底问没问过。
_write_jsonl(jsonl_dir / f"{kind}_profile_{date}.jsonl", list(profiles))
return jsonl_dir
def _write_jsonl(path: Path, records: List[Dict[str, Any]]) -> None:
with path.open("w", encoding="utf-8") as handle:
for record in records:
handle.write(json.dumps(record, ensure_ascii=False) + "\n")
utils.logger.info(f"[douyin_fetch] 写出 {len(records)} 条 -> {path.name}")
+244 -33
View File
@@ -38,13 +38,14 @@ import json
import re
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Dict, List, Optional
from typing import Any, Dict, List, Optional, Sequence
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .platforms import PLATFORM_XHS
from .models import (
EVENT_AUTH_FAILURE,
@@ -55,6 +56,7 @@ from .models import (
EVENT_NO_DATA,
EVENT_RUN_FAILED,
MonitorComment,
MonitorCreatorStat,
MonitorEvent,
MonitorNote,
MonitorNoteMetric,
@@ -121,13 +123,46 @@ _WINDOWS_EXIT_REASONS = {
}
def describe_exit_code(code: int) -> str:
"""Render an exit code so a human can act on it."""
def describe_exit_code(code: int, cause: Optional[str] = None) -> str:
"""Render an exit code so a human can act on it.
只有退出码时信息量约等于零(`code 1` 什么都能是),所以把从输出尾巴里认出来的
异常一并附上 —— 运行历史里那一格显示的正是这句话。
"""
unsigned = code & 0xFFFFFFFF if code < 0 else code
reason = _WINDOWS_EXIT_REASONS.get(unsigned)
base = f"Crawler exited with code {code}"
if reason:
return f"Crawler exited with code {code} (0x{unsigned:08X}): {reason}"
return f"Crawler exited with code {code}"
base = f"{base} (0x{unsigned:08X}): {reason}"
return f"{base};原因:{cause}" if cause else base
# 从爬虫输出里认出一行「异常」。Python 的 traceback 末行形如
# ``media_platform.douyin.exception.DataFetchError: account blocked``。
_EXCEPTION_LINE_RE = re.compile(r"^[\w.]*[A-Za-z](?:Error|Exception|Timeout)\b")
def diagnose_failure(output_tail: Optional[Sequence[str]]) -> Optional[str]:
"""从爬虫输出的末尾挑出最能说明问题的一行。
「退出码 1」等于什么都没说:真正的报错埋在子进程的 stderr 里。倒着找第一行看起来
像异常的行(traceback 的末行),找不到就退回最后一行有效输出。
"""
if not output_tail:
return None
lines = [line.strip() for line in output_tail if line and line.strip()]
# 管理器自己补的那两句不是爬虫的报错,别被当成失败原因。
noise = ("Crawler exited with code", "Crawler completed successfully")
lines = [line for line in lines if not line.startswith(noise)]
if not lines:
return None
for line in reversed(lines):
if _EXCEPTION_LINE_RE.match(line):
return line[:300]
return lines[-1][:300]
@dataclass
@@ -171,8 +206,11 @@ def find_run_files(
Glob rather than reconstructing the name: both the crawler type and the date
are runtime-dependent. Returns lists because a crawl crossing midnight
produces one file per day.
``platform`` 是**监控层的平台 id**,而爬虫落盘的目录名未必同名(抖音的 id 是
``dy``、目录是 ``douyin``),所以这里经 adapters 解析 —— 调用方不必知道这个差异。
"""
jsonl_dir = out_dir / platform / "jsonl"
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
if not jsonl_dir.is_dir():
return [], []
@@ -182,6 +220,43 @@ def find_run_files(
)
def find_profile_files(out_dir: Path, platform: str = PLATFORM_XHS) -> List[Path]:
"""博主**账号级**指标那几行 jsonl(``creator_profile_*.jsonl``)。
单独一个函数而不是往 ``find_run_files`` 的返回值里塞第三个列表:那个返回值的两个
位置是有意义的(contents/comments),加一个会把所有调用点和解包语句都牵动一遍,
而这份产物是**可选**的 —— 小红书那条路(爬虫进程)根本不产生它。
"""
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
if not jsonl_dir.is_dir():
return []
return sorted(jsonl_dir.glob("*_profile_*.jsonl"))
def _misplaced_output_dirs(out_dir: Path, expected: str) -> List[str]:
"""在 out_dir 下找「有产物、但目录名不是期望的那个」的目录。
这是专门为**最难查的那类故障**准备的:产物目录名与平台对不上时,ingest 一个文件
都找不到,现象和「登录态失效」一模一样 —— 而实际上登录好好的、数据也抓到了,
只是没人去对的地方读。上游哪天改了 store 的目录名,这里能直接把实情说出来。
"""
found = []
try:
children = list(out_dir.iterdir())
except OSError:
return found
for child in children:
if child.name == expected or not child.is_dir():
continue
try:
if any(child.glob("jsonl/*_contents_*.jsonl")):
found.append(child.name)
except OSError:
continue
return sorted(found)
async def _emit(
session: AsyncSession,
run: MonitorRun,
@@ -271,15 +346,27 @@ async def _ingest_notes(
run: MonitorRun,
records: List[Dict[str, Any]],
is_baseline: bool,
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert notes, write metric snapshots, and emit new-note/delta events."""
"""Upsert notes, write metric snapshots, and emit new-note/delta events.
记录里的字段一律经 ``adapter`` 读。抖音的作品没有 ``note_id``(叫 ``aweme_id``),
按名字硬取的话每条记录都会在下面第一行被 continue 掉 —— 一条都不报错地全丢。
"""
now = get_current_timestamp()
new_count = 0
# 同一轮里重复出现的作品只处理一次。**指标快照的唯一键是 (task_id, note_id, run_id)**,
# 同一件作品在一轮里进来两次会让第二次插入直接撞键、整个 run 崩掉 —— 产物里重复并不
# 罕见(多个目标指向同一个人、或退化路径重复刷新)。
seen_in_run: set = set()
for record in records:
note_id = record.get("note_id")
note_id = adapter.note_field(record, "note_id")
if not note_id:
continue
if note_id in seen_in_run:
continue
seen_in_run.add(note_id)
note = await session.scalar(
select(MonitorNote).where(
@@ -288,20 +375,21 @@ async def _ingest_notes(
)
)
title = (record.get("title") or "")[:500]
raw_images = record.get("image_list") or ""
cover = raw_images.split(",")[0] if raw_images else ""
title = (adapter.note_field(record, "title") or "")[:500]
cover = adapter.cover(record)
if note is None:
note = MonitorNote(
task_id=run.task_id,
note_id=note_id,
title=title,
note_url=record.get("note_url") or "",
note_url=adapter.note_field(record, "note_url") or "",
cover=cover,
creator_hash=record.get("creator_hash") or "",
source_kind=record.get("type") or "",
published_at=_as_int(record.get("time")),
creator_hash=adapter.note_field(record, "creator_hash") or "",
creator_name=adapter.note_field(record, "creator_name") or "",
source_kind=adapter.note_field(record, "source_kind") or "",
# 经 to_ms 换算:小红书给毫秒、抖音给秒,差 1000 倍。
published_at=adapter.to_ms(adapter.note_field(record, "published_at")),
first_seen_run_id=run.id,
first_seen_at=now,
last_seen_run_id=run.id,
@@ -323,6 +411,20 @@ async def _ingest_notes(
# Only refresh descriptive fields; seen-tracking is updated below.
if title:
note.title = title
# 封面地址**带签名、会过期**,所以每轮都用最新的覆盖它。原先只在首次入库
# 时写一次,结果旧作品的封面地址烂在库里 —— 隔天开始全是 403,而且再怎么
# 重跑也修不回来。落盘那份由 covers.cache_pending 负责(网络操作不在本模块)。
if cover:
note.cover = cover
# 昵称也要刷:作者改昵称是常事,只在首次入库写一次会一直显示旧的。
creator_name = adapter.note_field(record, "creator_name")
if creator_name:
note.creator_name = creator_name
# 发布时间也刷。正常情况下它不会变,但**换算单位改过之后**(抖音是秒、
# 小红书是毫秒),已经入库的那批只能靠重采修回来。
published = adapter.to_ms(adapter.note_field(record, "published_at"))
if published is not None:
note.published_at = published
note.last_seen_run_id = run.id
note.last_seen_at = now
@@ -420,40 +522,57 @@ async def _ingest_comments(
records: List[Dict[str, Any]],
is_baseline: bool,
previous_run_started_at: Optional[int],
adapter: adapters.PlatformAdapter,
) -> int:
"""Upsert comments and emit events for ones never seen before."""
"""Upsert comments and emit events for ones never seen before.
与作品同理,评论记录也要经 ``adapter`` 读:抖音的评论用 ``aweme_id`` 指作品。
"""
now = get_current_timestamp()
new_count = 0
for record in records:
comment_id = record.get("comment_id")
note_id = record.get("note_id")
comment_id = adapter.comment_field(record, "comment_id")
note_id = adapter.comment_field(record, "note_id")
if not comment_id or not note_id:
continue
exists = await session.scalar(
select(MonitorComment.id).where(
existing = await session.scalar(
select(MonitorComment).where(
MonitorComment.task_id == run.task_id,
MonitorComment.note_id == note_id,
MonitorComment.comment_id == comment_id,
)
)
if exists is not None:
if existing is not None:
# 昵称要跟着刷,不能只写一次。评论是去重后直接 continue 的,若不刷新,
# 脱敏开关一改(或评论者改了昵称),已经入库的老评论会永远停在旧值上 ——
# 而重采是唯一能拿到新值的途径。作品那边的 creator_name 同理。
refreshed = adapter.comment_field(record, "creator_name")
if refreshed:
existing.nickname = refreshed
# 时间同理:单位换算修好之后,老数据要重采才能纠正。
created = adapter.to_ms(adapter.comment_field(record, "create_time"))
if created is not None:
existing.create_time = created
continue
create_time = _as_int(record.get("create_time"))
create_time = adapter.to_ms(adapter.comment_field(record, "create_time"))
session.add(
MonitorComment(
task_id=run.task_id,
note_id=note_id,
comment_id=comment_id,
content=(record.get("content") or "")[:2000],
nickname=record.get("nickname") or "",
creator_hash=record.get("creator_hash") or "",
content=(adapter.comment_field(record, "content") or "")[:2000],
nickname=adapter.comment_field(record, "creator_name") or "",
creator_hash=adapter.comment_field(record, "creator_hash") or "",
create_time=create_time,
like_count=parse_count(record.get("like_count")),
sub_comment_count=_as_int(record.get("sub_comment_count")) or 0,
parent_comment_id=record.get("parent_comment_id") or "",
like_count=parse_count(adapter.comment_field(record, "like_count")),
sub_comment_count=_as_int(
adapter.comment_field(record, "sub_comment_count")
)
or 0,
parent_comment_id=adapter.parent_comment_id(record),
first_seen_run_id=run.id,
first_seen_at=now,
)
@@ -488,11 +607,55 @@ async def _ingest_comments(
return new_count
async def _ingest_creator_stats(
session: AsyncSession,
run: MonitorRun,
records: Sequence[Dict[str, Any]],
) -> int:
"""把这一轮问到的博主账号级指标落成快照,返回条数。
和作品指标一样是**每轮一条**:账号级的粉丝数是缓慢变化的量,「今天比昨天多了 300」
才是有用的信号,单看一个绝对值没有意义 —— 所以这里只管记,分析交给查询端。
**没解析出来的值留 NULL,不写 0**:0 在趋势图上是一条砸到底的线,和「不知道」完全是
两回事(见 ``parse_count`` 的注释)。
"""
now = get_current_timestamp()
written = 0
seen: set = set()
for record in records:
creator_hash = str(record.get("creator_hash") or "").strip()
if not creator_hash or creator_hash in seen:
# 一个任务可以配多个目标,退化路径下它们可能指向同一个博主 —— 而唯一键是
# (task_id, creator_hash, run_id),重复插入会撞键把整轮炸掉。
continue
seen.add(creator_hash)
session.add(
MonitorCreatorStat(
task_id=run.task_id,
run_id=run.id,
creator_hash=creator_hash,
nickname=str(record.get("nickname") or "")[:128],
fans=parse_count(record.get("fans")),
total_favorited=parse_count(record.get("total_favorited")),
works_count=parse_count(record.get("works")),
following=parse_count(record.get("following")),
captured_at=now,
)
)
written += 1
return written
async def ingest_run(
session: AsyncSession,
run: MonitorRun,
task: MonitorTask,
out_dir: Path,
output_tail: Optional[Sequence[str]] = None,
) -> IngestResult:
"""Ingest one finished run and return what changed.
@@ -503,17 +666,29 @@ async def ingest_run(
# A non-zero exit is a genuine crash: trust nothing this run produced.
if run.exit_code not in (0, None):
run.status = RUN_FAILED
run.error_message = describe_exit_code(run.exit_code)
cause = diagnose_failure(output_tail)
run.error_message = describe_exit_code(run.exit_code, cause)
title = f"采集进程异常退出(code={run.exit_code})"
if cause:
# 标题里也带上真因:企业微信通知和事件流都只看这一行,不写就还得去翻日志。
title = f"{title}:{cause}"
await _emit(
session,
run,
EVENT_RUN_FAILED,
f"采集进程异常退出(code={run.exit_code})",
title,
severity="error",
payload={"exit_code": run.exit_code, "detail": run.error_message},
payload={
"exit_code": run.exit_code,
"detail": run.error_message,
"cause": cause,
},
)
return IngestResult(status=RUN_FAILED, error=run.error_message)
adapter = adapters.adapter(task.platform)
subdir = adapters.artifact_dir(task.platform)
contents_paths, comment_paths = find_run_files(out_dir, task.platform)
contents = [record for path in contents_paths for record in _read_jsonl(path)]
comments = [record for path in comment_paths for record in _read_jsonl(path)]
@@ -521,6 +696,16 @@ async def ingest_run(
run.notes_fetched = len(contents)
run.comments_fetched = len(comments)
# 账号级快照**在「一条作品都没采到」的早退之前**落。博主的粉丝数并不会因为他这个
# 月的新作品列表被风控挡住就不存在 —— 那正是最该看到「粉丝还在涨、但新作品没在发现」
# 的时刻,跳过它等于在最需要它的那轮把数据丢掉。
profiles = [
record
for path in find_profile_files(out_dir, task.platform)
for record in _read_jsonl(path)
]
await _ingest_creator_stats(session, run, profiles)
# A bad cookie does NOT fail the process: XHS cookie login is never validated,
# so an unauthenticated session just returns zero notes with exit 0 -- and
# usually does not even create an output file. Treating that as "the creator
@@ -529,6 +714,32 @@ async def ingest_run(
if not contents:
run.status = RUN_PARTIAL
# 先排除「东西抓到了,只是没落在我们找的那个目录里」。这种故障的现象和登录失效
# 一模一样,但登录其实是好的 —— 按登录失效报会把人指到完全错的方向去查。
misplaced = _misplaced_output_dirs(out_dir, subdir)
if misplaced:
run.error_message = (
f"crawler wrote into {misplaced} but platform {task.platform} "
f"expects {subdir}"
)
await _emit(
session,
run,
EVENT_NO_DATA,
f"采集产物目录与平台不匹配(实际 {misplaced}、期望 {subdir}),本次未读到任何作品",
severity="error",
payload={
"out_dir": str(out_dir),
"expected": subdir,
"found": misplaced,
},
)
return IngestResult(
status=RUN_PARTIAL,
error=run.error_message,
comments_fetched=len(comments),
)
# Blaming the cookie is only honest if nothing else is authenticating.
# A sibling task that just succeeded proves the login works, so the
# fault is with this target (bad/expired per-creator token, an empty
@@ -578,10 +789,10 @@ async def ingest_run(
comments_fetched=len(comments),
is_baseline=is_baseline,
)
result.new_notes = await _ingest_notes(session, run, contents, is_baseline)
result.new_notes = await _ingest_notes(session, run, contents, is_baseline, adapter)
if task.enable_comments:
result.new_comments = await _ingest_comments(
session, run, comments, is_baseline, previous_started_at
session, run, comments, is_baseline, previous_started_at, adapter
)
run.new_notes = result.new_notes
+120 -3
View File
@@ -75,7 +75,7 @@ MODE_NOTE = "note"
class MonitorTask(MonitorBase):
"""One monitored schedule: a set of targets plus an interval."""
"""One monitored schedule: a set of targets, plus when to run them."""
__tablename__ = "monitor_task"
@@ -86,15 +86,36 @@ class MonitorTask(MonitorBase):
enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
interval_minutes: Mapped[int] = mapped_column(Integer, nullable=False, default=360)
# How the task is scheduled. `interval` is the original "every N minutes" and
# stays the default; `daily` and `weekly` fire at chosen clock times instead
# (the arithmetic lives in schedule.py).
#
# The clock fields are comma-separated text rather than a child table: they
# are a handful of small integers, always read as a whole, and a table would
# buy nothing but joins.
schedule_mode: Mapped[str] = mapped_column(String(16), nullable=False, default="interval")
# 0-23, e.g. "9,12,18". Empty in interval mode.
schedule_hours: Mapped[str] = mapped_column(String(96), nullable=False, default="")
# 0-6 with Monday = 0, matching Python's date.weekday(). Weekly mode only.
schedule_days: Mapped[str] = mapped_column(String(32), nullable=False, default="")
# Minute past the hour, shared by every time in the schedule.
schedule_minute: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
# Crawl window knobs, mirrored onto each run's CLI flags.
max_notes_count: Mapped[int] = mapped_column(Integer, nullable=False, default=20)
enable_comments: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
max_comments_count: Mapped[int] = mapped_column(Integer, nullable=False, default=50)
run_timeout_seconds: Mapped[int] = mapped_column(Integer, nullable=False, default=3600)
# Push notifications are opt-in per task. A task list that all pushes to one
# webhook turns noisy fast, so silence is the default.
# 通知分成两类,因为它们的性质完全不同:
#
# * `notify_enabled` —— **推送新作品**。可能每轮都有,一条任务列表都推到同一个群
# 会很快变吵,所以默认关。(列名是历史遗留:它早先是唯一的通知开关。)
# * `notify_failures` —— **推送异常**(登录失效 / 运行失败 / 没抓到数据)。频率低,
# 而且一旦发生就意味着这个任务从此**默默采不到任何东西**,你会一直不知道,
# 直到某天发现数据停在几周前。这正是最该被告知的情况,所以默认**开**。
notify_enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
notify_failures: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
# Scheduler state. Persisted so the schedule survives an API restart.
next_run_at: Mapped[Optional[int]] = mapped_column(BigInteger, index=True)
@@ -205,7 +226,12 @@ class MonitorNote(MonitorBase):
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
note_url: Mapped[str] = mapped_column(Text, nullable=False, default="")
cover: Mapped[str] = mapped_column(Text, nullable=False, default="")
# 创作者匿名哈希。爬虫刻意不落原始 user_id(见 tools/user_hash.py),
# 所以这是唯一稳定的创作者标识 —— 按博主分组就靠它。
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, default="")
# 创作者昵称,**已由爬虫脱敏**(张***三 这种)。存的是脱敏后的值,与项目一贯的
# 匿名化姿态一致;不存的话分组只能显示一串哈希,根本认不出是谁。
creator_name: Mapped[str] = mapped_column(String(200), nullable=False, default="")
source_kind: Mapped[str] = mapped_column(String(16), nullable=False, default="")
published_at: Mapped[Optional[int]] = mapped_column(BigInteger)
@@ -294,6 +320,87 @@ class MonitorEvent(MonitorBase):
is_read: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
class MonitorCreatorAlias(MonitorBase):
"""给博主起的备注。
作品栏和评论栏都按 ``creator_hash`` 把作品归到博主名下,可那是个哈希;
``creator_name`` 是平台上的昵称(而且粉丝少的号常常没有)。两样都认不出"这是谁"。
备注是**人自己起的名字**(「竞品A」「自家号-3」),用来把账号对上人。
键取 ``(platform, creator_hash)``:哈希对同一个 uid 是稳定的,所以同一个博主出现在
多个任务里时备注也是同一个,不用每个任务各填一遍。
"""
__tablename__ = "monitor_creator_alias"
__table_args__ = (
UniqueConstraint("platform", "creator_hash", name="uq_creator_alias"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
platform: Mapped[str] = mapped_column(String(16), nullable=False, index=True)
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
alias: Mapped[str] = mapped_column(String(128), nullable=False, default="")
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
class MonitorCreatorStat(MonitorBase):
"""博主的**账号级**快照:粉丝数 / 总获赞 / 作品数 / 关注数。
这是作品列表给不了的东西:作品级指标说"这一条视频涨了多少赞",账号级说"这个人
整个账号的粉丝是在涨还是在掉"。两者不互相替代。
粒度取 ``(任务, 博主, 轮次)``,和作品指标一样的形状 —— 于是趋势、差分、报表那套
现成的逻辑换个表就能用。
目前**只有抖音**会写它:小红书那条走的是爬虫子进程,而它的 ``save_creator()`` 在
教学版里是空函数,根本没落过创作者资料。所以表里只有抖音的博主。
"""
__tablename__ = "monitor_creator_stat"
__table_args__ = (
UniqueConstraint(
"task_id", "creator_hash", "run_id", name="uq_creator_stat"
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
task_id: Mapped[int] = mapped_column(
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
)
run_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
nickname: Mapped[str] = mapped_column(String(128), nullable=False, default="")
# 都可能为 None:平台没给就留空,**不要伪造成 0** —— 0 是"掉到零",和"不知道"
# 在趋势图上是完全不同的两回事。
fans: Mapped[Optional[int]] = mapped_column(BigInteger)
total_favorited: Mapped[Optional[int]] = mapped_column(BigInteger)
works_count: Mapped[Optional[int]] = mapped_column(BigInteger)
following: Mapped[Optional[int]] = mapped_column(BigInteger)
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
class MonitorNoteAlias(MonitorBase):
"""给**作品**起的备注。
和 ``MonitorCreatorAlias`` 是一对:博主那条回答"这是谁",这条回答"这条我要盯着"。
键取 ``(platform, note_id)`` —— 作品 id 本身就带平台语义,但显式带上 platform 才能和
博主备注用同一套查询形状。
"""
__tablename__ = "monitor_note_alias"
__table_args__ = (
UniqueConstraint("platform", "note_id", name="uq_note_alias"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
platform: Mapped[str] = mapped_column(String(16), nullable=False, index=True)
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
alias: Mapped[str] = mapped_column(String(128), nullable=False, default="")
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
class MonitorSetting(MonitorBase):
"""Key/value store. Holds the XHS cookie for unattended runs."""
@@ -331,6 +438,16 @@ SETTING_AUTH_PASSWORD_UPDATED_AT = "auth_password_updated_at"
# them. Key builders live in settings.py.
SETTING_WECOM_WEBHOOK = "system.wecom_webhook"
# 上游更新检查的两条状态。都不是给用户编辑的设置项,所以不在 app_settings 的注册表里
# (那张表只列可编辑项,因此也不会被设置接口读出来)。
#
# 最近一次检查的结果整体存成一条 JSON:它总是被整体读写,拆成多个 key 只会带来
# 另一半没写完的不一致。
SETTING_UPSTREAM_STATE = "system.upstream_check_state"
# 已经推送过通知的那个上游 tip。换 tip 才再推 —— 否则每个检查周期都会把同样的
# 更新推一遍,直到有人去合并为止;而上游真又动了的时候应该再推一次。
SETTING_UPSTREAM_NOTIFIED_TIP = "system.upstream_notified_tip"
# Pre-namespacing keys, kept only so the startup migration can find and move
# them. Nothing should read these directly.
LEGACY_SETTING_KEY_RENAMES = {
+21 -4
View File
@@ -36,6 +36,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import adapters
from .models import (
EVENT_AUTH_FAILURE,
EVENT_NEW_NOTE,
@@ -107,14 +108,27 @@ async def build_run_message(
task: MonitorTask,
run: MonitorRun,
) -> Optional[str]:
"""Compose one markdown summary for a finished run, or None if nothing to say."""
"""Compose one markdown summary for a finished run, or None if nothing to say.
事件按开关过滤:只勾了「新作品」的任务,不该因为一次失败被推消息,反之亦然 ——
否则拆开这两个开关就没有意义了。
"""
allowed = []
if task.notify_enabled:
allowed.append(EVENT_NEW_NOTE)
if task.notify_failures:
allowed.extend([EVENT_AUTH_FAILURE, EVENT_RUN_FAILED, EVENT_NO_DATA])
if not allowed:
return None
events = list(
(
await session.scalars(
select(MonitorEvent)
.where(
MonitorEvent.run_id == run.id,
MonitorEvent.type.in_(NOTIFIABLE_EVENT_TYPES),
MonitorEvent.type.in_(allowed),
)
.order_by(MonitorEvent.id)
)
@@ -151,7 +165,9 @@ async def build_run_message(
payload = _load_payload(event.payload_json)
title = payload.get("title") or event.target_id
note_id = payload.get("note_id") or event.target_id
url = f"https://www.xiaohongshu.com/explore/{note_id}"
# 链接形状按平台来。抖音的作品是 /video/{id},写死小红书域名的话,
# 群里点进去会是一个 404 —— 而这正是通知唯一要它干的事。
url = adapters.adapter(task.platform).note_url(note_id)
lines.append(f"> [{title}]({url})")
if len(new_notes) > 10:
lines.append(f"> …等共 {len(new_notes)} 篇")
@@ -176,7 +192,8 @@ async def notify_run(session: AsyncSession, task: MonitorTask, run: MonitorRun)
Returns the message that was sent, or None. Never raises.
"""
try:
if not task.notify_enabled:
# 两个开关是分开的:只开「异常」不该因为新作品而发消息,反之亦然。
if not (task.notify_enabled or task.notify_failures):
return None
webhook_url = await get_webhook_url(session)
+24 -5
View File
@@ -29,12 +29,13 @@ Two distinct things are recorded here, and conflating them would be misleading:
platform modules, not assumed -- all seven implement search/detail/creator;
the real differences are in which interaction metrics they capture.
* ``monitor_wired`` says whether the *monitoring layer* has been hooked up. It
currently covers only Xiaohongshu: ``runner.py`` pins the platform,
``ingest.py`` reads a fixed ``xhs/jsonl`` directory, and ``service.py`` only
parses Xiaohongshu target URLs.
covers Xiaohongshu and Douyin. The parts where those two differ -- which
directory the crawler writes into, what the jsonl fields are called, what a
target URL looks like -- live in ``adapters.py``; the rest of the layer is
platform-neutral.
A platform can therefore be fully crawlable by upstream and still not usable for
monitoring, which is exactly the state of the other six today.
monitoring, which is exactly the state of the other five today.
"""
from typing import Any, Dict, List, Optional
@@ -61,13 +62,31 @@ PLATFORM_CAPABILITIES: Dict[str, Dict[str, Any]] = {
"comment_levels": 2,
"media": True,
"monitor_wired": True,
"target_hints": {
"creator": "https://www.xiaohongshu.com/user/profile/5f58bd990000000001003753",
"note": "https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=...",
"creator_label": "博主主页",
"note_label": "笔记",
# 只有小红书的链接带会过期的 xsec_token。抖音的链接不带令牌,永久有效,
# 那句「建议只填纯 ID」的劝告对它没有意义。
"token_expires": True,
},
},
"dy": {
"crawler_modes": ["search", "detail", "creator"],
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
"comment_levels": 2,
"media": True,
"monitor_wired": False,
"monitor_wired": True,
# 用户可见的示例链接(前端的目标输入框用它做 placeholder)。放这里是因为
# 它属于「这个平台长什么样」的能力描述;真正干活的管子(正则、目录名、
# 字段别名)在 adapters.py。
"target_hints": {
"creator": "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE",
"note": "https://www.douyin.com/video/7525082444551310602",
"creator_label": "博主主页",
"note_label": "作品",
},
},
"ks": {
"crawler_modes": ["search", "detail", "creator"],
+425
View File
@@ -0,0 +1,425 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/qrlogin.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""扫码登录,以及"现在到底登没登录"的查询。
**为什么需要这个模块**:服务器上 Chrome 跑在 Xvfb 里没有显示器,爬虫原本用
`show_qrcode`(PIL 的 `Image.show()`)弹窗展示二维码,那需要桌面看图程序,服务器上
没有。所以改成经 CDP 把二维码从页面里读出来交给 WebUI。
**为什么登录状态要能独立查询**:扫码会话是内存里的临时状态,进程一重启就没了
(部署、崩溃都算)。把"是否已登录"绑在它上面,就会出现"扫完了但界面没反应、
也不知道到底成没成"。所以状态查询是独立的、随时可调用的,二维码只是达成它的手段之一。
三个容易搞错的地方:
* **必须复用浏览器默认 context**。`browser.new_context()` 会造出一个无痕式的 profile,
扫了也白扫——爬虫读不到那份 cookie。真正的 profile 在 `browser.contexts[0]`。
* **绝不能调 `browser.close()`**。对 CDP 连接而言那会关掉操作者自己的 Chrome,
连带所有无关标签页。只能关本模块自己开的那一个。
* **不能靠 `web_session` 判断登录**。实测:一个全新的空 profile 首次访问小红书就会
被发一个 `web_session`,所以"有这个 cookie"什么都证明不了。可信信号是页面自己的
`__INITIAL_STATE__.user.loggedIn`。
"""
import asyncio
import os
import time
from typing import Any, Dict, Optional
import config
from playwright.async_api import async_playwright
from tools import utils
from ..creator.client import CreatorApiError, CreatorClient
from .platforms import PLATFORM_XHS
def _cookie_string(cookies) -> str:
"""把 CDP 拿到的 cookie 列表拼成请求头用的字符串。"""
return "; ".join(f"{c['name']}={c['value']}" for c in cookies)
# 二维码有效期。平台自己会更早轮换;这个上限只是为了让一次被放弃的尝试不会
# 永久占着一个标签页。
QR_TTL_SECONDS = 300
# 登录状态查询的缓存时长。轮询时不必每次都去问浏览器。
STATE_CACHE_SECONDS = 5
STATUS_IDLE = "idle"
STATUS_WAITING = "waiting"
STATUS_SUCCESS = "success"
STATUS_EXPIRED = "expired"
STATUS_ERROR = "error"
# 只有小红书接了监控流程,所以扫码也只对它开放。给别的平台显示一个按不动的按钮
# 是在假装功能存在。
LOGIN_URL: Dict[str, str] = {PLATFORM_XHS: "https://www.xiaohongshu.com"}
EXPLORE_URL: Dict[str, str] = {PLATFORM_XHS: "https://www.xiaohongshu.com/explore"}
QR_SELECTOR: Dict[str, str] = {PLATFORM_XHS: "xpath=//img[@class='qrcode-img']"}
LOGIN_BUTTON_SELECTOR: Dict[str, str] = {
PLATFORM_XHS: "xpath=//*[@id='app']/div[1]/div[2]/div[1]/ul/div[1]/button"
}
# 这里本来有一个读 window.__INITIAL_STATE__ 的 JS 探针,**已删除,不要加回来**。
#
# 它是页面加载那一刻的快照:浏览器本来就登录着时它是对的,但扫码是加载**之后**才
# 登录的,快照不会翻转,检测于是永远等不到 —— 表现为"扫了码却一直停在二维码上"。
# 运营模块踩过同一个坑。现在的判据是拿 cookie 问后台接口,见 check_login_state。
_lock = asyncio.Lock()
_current: Optional["QrLoginSession"] = None
# 常驻的 Playwright 客户端和本模块自己的标签页。长期持有是有意的:状态查询要能
# 随时回答,而每次都新建一个标签页会在操作者的浏览器里堆垃圾。
_playwright: Any = None
_page: Any = None
# (时间戳, 结果),避免轮询时反复问浏览器。
_state_cache: Optional[tuple[float, Dict[str, Any]]] = None
def _cdp_url() -> str:
"""浏览器 DevTools 端点。``MC_CDP_URL`` 优先,便于换主机而不用改代码。"""
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
def _login_url(platform: str) -> str:
if platform == PLATFORM_XHS and getattr(config, "XHS_INTERNATIONAL", False):
return "https://www.rednote.com"
return LOGIN_URL[platform]
async def _ensure_context() -> Any:
"""连上浏览器并返回它的默认 context。"""
global _playwright
if _playwright is None:
_playwright = await async_playwright().start()
try:
browser = await _playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
except Exception as exc:
await _reset_playwright()
raise RuntimeError(
f"连接浏览器失败({_cdp_url()})。请确认服务器上的 Chrome 以 "
f"--remote-debugging-port 启动。原始错误:{exc}"
) from exc
if not browser.contexts:
raise RuntimeError(
"浏览器没有可用上下文。CDP 已连上,但读不到 profile —— "
"请确认 Chrome 不是以无痕模式启动的。"
)
# contexts[0] 就是真实 profile,用它,不要 new_context()。
return browser.contexts[0]
async def _ensure_page(platform: str = PLATFORM_XHS, reload: bool = False) -> Any:
"""本模块在操作者浏览器里的那一个标签页,复用而不是反复新建。
若已有一个停在目标站点的标签页就认领它——进程重启后页柄会丢,但标签页还在,
认领可以避免在浏览器里留下一堆没人关的孤儿页。
"""
global _page
context = await _ensure_context()
if _page is not None:
try:
if _page.is_closed():
_page = None
except Exception:
_page = None
if _page is None:
for candidate in context.pages:
try:
if "xiaohongshu.com" in candidate.url or "rednote.com" in candidate.url:
_page = candidate
break
except Exception:
continue
if _page is None:
_page = await context.new_page()
try:
url = _page.url
except Exception:
url = ""
if reload or "xiaohongshu.com" not in url and "rednote.com" not in url:
await _page.goto(
EXPLORE_URL.get(platform, EXPLORE_URL[PLATFORM_XHS]),
wait_until="domcontentloaded",
timeout=45000,
)
return _page
async def check_login_state(force: bool = False) -> Dict[str, Any]:
"""问浏览器:现在登录了吗?
``force`` 会先重新加载页面。SPA 的状态会随登录实时更新,所以轮询时不必重载;
但若登录态是在别处失效的,页面上的副本可能是陈旧的,重新检测就该重载。
"""
global _state_cache
now = time.time()
if not force and _state_cache is not None:
cached_at, cached = _state_cache
if now - cached_at < STATE_CACHE_SECONDS:
return cached
try:
context = await _ensure_context()
cookies = await context.cookies()
except Exception as exc:
result = {
"known": False,
"logged_in": False,
"nickname": None,
"error": f"{exc.__class__.__name__}: {exc}",
}
_state_cache = (now, result)
return result
cookie = _cookie_string(cookies)
# 判据不再是页面里的 window.__INITIAL_STATE__ —— 那是**页面加载那一刻的快照**:
# 浏览器已登录时它是对的,但扫码是加载**之后**才登录的,快照不会翻转,检测就永远
# 等不到(运营模块踩过同一个坑)。改成拿 cookie 问后台「我是谁」,那是权威的:
# 实测游客也会被发一个 web_session,所以「有这个 cookie」什么都证明不了,
# 后台认了才算。
try:
info = await CreatorClient(cookie).fetch_user_info()
except CreatorApiError:
result = {"known": True, "logged_in": False, "nickname": None}
else:
result = {
"known": True,
"logged_in": bool(info.get("user_id")),
"nickname": info.get("nickname"),
}
_state_cache = (now, result)
return result
async def _current_cookie() -> str:
"""默认 profile 当前的小红书 cookie 串。
扫码面板要的不只是「登录了吗」,而是**把登录态拿出来存一份** —— 存进库之后,
即使 CDP 关掉、任务改用 --cookies_file 注入,也照样能跑。
"""
context = await _ensure_context()
return _cookie_string(await context.cookies())
async def _reset_playwright() -> None:
global _playwright, _page
_page = None
if _playwright is not None:
try:
await _playwright.stop()
except Exception:
pass
_playwright = None
class QrLoginSession:
"""一次进行中的扫码尝试。"""
def __init__(self, platform: str, page: Any) -> None:
self.platform = platform
self.status = STATUS_WAITING
self.message = "请用手机扫描二维码"
self.image = ""
self.started_at = time.time()
self.logged_in = False
self.nickname: Optional[str] = None
# 登录成功后从默认 profile 取出来的 cookie,供调用方存库。
self.cookie: str = ""
self.cookie_taken = False
self._page = page
@property
def elapsed(self) -> float:
return time.time() - self.started_at
async def refresh(self) -> None:
"""轮询一次,看扫码是否完成。"""
if self.status != STATUS_WAITING:
return
if self.elapsed > QR_TTL_SECONDS:
self.status = STATUS_EXPIRED
self.message = "二维码已超时,请重新获取"
return
state = await check_login_state()
if state.get("logged_in"):
self.cookie = await _current_cookie()
self.logged_in = True
self.nickname = state.get("nickname")
self.status = STATUS_SUCCESS
who = f"({self.nickname})" if self.nickname else ""
self.message = f"登录成功{who},登录态已写入浏览器 profile"
return
try:
if self._page.is_closed():
self.status = STATUS_ERROR
self.message = "二维码所在页面已被关闭,请重新获取"
except Exception:
pass
def snapshot(self) -> Dict[str, Any]:
return {
"status": self.status,
"platform": self.platform,
"image": self.image,
"message": self.message,
"elapsed": int(self.elapsed),
"expires_in": max(0, int(QR_TTL_SECONDS - self.elapsed)),
"logged_in": self.logged_in,
"nickname": self.nickname,
}
async def _read_qr(page: Any, platform: str) -> str:
"""把二维码从页面里取出来,必要时先点开登录框。"""
image = await utils.find_login_qrcode(page, selector=QR_SELECTOR[platform])
if image:
return image
# 登录框不一定自己弹出来。这是爬虫自身扫码流程里同款兜底。
await asyncio.sleep(0.5)
try:
await page.locator(LOGIN_BUTTON_SELECTOR[platform]).click(timeout=5000)
except Exception:
return ""
return await utils.find_login_qrcode(page, selector=QR_SELECTOR[platform])
async def _discard_current_locked() -> None:
global _current
_current = None
async def start(platform: str = PLATFORM_XHS) -> Dict[str, Any]:
"""在 CDP 浏览器里打开登录页,取回二维码。"""
global _current
if platform not in LOGIN_URL:
raise ValueError(f"平台 {platform} 尚未接入扫码登录(目前仅支持小红书)")
async with _lock:
await _discard_current_locked()
# **先问状态,再决定要不要开页面。** 顺序反过来是有代价的:读二维码内部会
# wait_for_selector 等满 30 秒才放弃,而已经登录时页面上根本没有二维码 ——
# 用户点一下按钮要干等半分钟,还白开一个标签页。
state = await check_login_state(force=True)
if state.get("logged_in"):
# 已经是登录状态时站点不显示二维码 —— 这本身就是成功,不是失败。
# 顺带把 cookie 取出来,让调用方可以存进库。
session = QrLoginSession(platform, None)
session.cookie = await _current_cookie()
session.status = STATUS_SUCCESS
session.logged_in = True
session.nickname = state.get("nickname")
who = f"({session.nickname})" if session.nickname else ""
session.message = f"浏览器已经是登录状态{who},无需扫码"
_current = session
return session.snapshot()
page = await _ensure_page(platform)
try:
await page.goto(
_login_url(platform), wait_until="domcontentloaded", timeout=45000
)
image = await _read_qr(page, platform)
except Exception as exc:
raise RuntimeError(f"打开登录页失败:{exc}") from exc
session = QrLoginSession(platform, page)
session.image = image
if not image:
session.status = STATUS_ERROR
session.message = "页面上没找到二维码,请确认站点结构没有变化"
_current = session
return session.snapshot()
async def status() -> Dict[str, Any]:
async with _lock:
if _current is None:
state = await check_login_state()
return {
"status": STATUS_IDLE,
"platform": None,
"image": "",
"message": "",
"elapsed": 0,
"expires_in": 0,
"logged_in": bool(state.get("logged_in")),
"nickname": state.get("nickname"),
}
await _current.refresh()
return _current.snapshot()
async def take_cookie() -> Optional[str]:
"""取走已登录会话的 cookie,且只给一次。
由路由层在落库时调用。**cookie 不进响应体** —— 它是凭证,前端没有理由看到它。
这里**刻意不结束会话**(与运营模块不同):那里取完即拆,因为临时上下文用完就该丢;
这里的浏览器 profile 是长期存在的,面板还该继续显示「已登录」。所以只标记已取过,
让重复轮询拿不到第二份、也就不会反复写库。
"""
async with _lock:
if _current is None or _current.status != STATUS_SUCCESS or _current.cookie_taken:
return None
_current.cookie_taken = True
return _current.cookie
async def cancel() -> Dict[str, Any]:
async with _lock:
await _discard_current_locked()
state = await check_login_state()
return {
"status": STATUS_IDLE,
"platform": None,
"image": "",
"message": "已取消",
"elapsed": 0,
"expires_in": 0,
"logged_in": bool(state.get("logged_in")),
"nickname": state.get("nickname"),
}
async def shutdown() -> None:
"""进程退出时断开连接。刻意不关那个标签页——它是操作者浏览器的一部分。"""
global _current
async with _lock:
_current = None
await _reset_playwright()
+5 -1
View File
@@ -135,7 +135,11 @@ async def build_report(
start_ms, _ = day_bounds(start_day)
_, end_ms = day_bounds(end_day)
scope = list(task_ids) if task_ids else None
# 必须是 `is not None`,不能写 `if task_ids` —— **空列表是假值**,而空列表在这里
# 的含义是「这个平台一个任务都没有」,不是「不限制平台」。用真值判断的话,
# 切到一个还没有任务的平台,报表会把**所有**任务的数据聚合出来(看起来就是
# 「抖音的报表里全是小红书的数据」)。
scope = list(task_ids) if task_ids is not None else None
days = iter_days(start_day, end_day)
# Fetch every snapshot up to the range end: the delta on the first day needs
+155 -60
View File
@@ -28,6 +28,8 @@ import os
from pathlib import Path
from typing import Iterable, List, Optional
from sqlalchemy import select
from tools.time_util import get_current_timestamp
from ..schemas import (
@@ -38,15 +40,16 @@ from ..schemas import (
SaveDataOptionEnum,
)
from ..services import crawler_manager
from . import app_settings, notify
from . import adapters, app_settings, covers, douyin_fetch, notify
from .db import get_session
from .ingest import IngestResult, ingest_run
from .ingest import IngestResult, diagnose_failure, ingest_run
from .models import (
MODE_CREATOR,
RUN_FAILED,
RUN_PENDING,
RUN_RUNNING,
RUN_TIMEOUT,
MonitorNote,
MonitorRun,
MonitorTarget,
MonitorTask,
@@ -68,30 +71,33 @@ _PLATFORM_ENUM = {
"zhihu": PlatformEnum.ZHIHU,
}
_XHS_WEB_BASE = "https://www.xiaohongshu.com"
_CREATOR_PATH = "/user/profile"
_NOTE_PATH = "/explore"
# Timeout used when the caller does not care; tasks carry their own.
DEFAULT_RUN_TIMEOUT_SECONDS = 3600
def build_target_url(value: str, kind: str) -> str:
def build_target_url(value: str, kind: str, platform: str) -> str:
"""Turn a stored target into a URL the crawler's parser accepts.
Always emits a full URL rather than a bare id: the XHS parser accepts a bare
24-hex id only, so the URL form is the safer universal input. The
``xsec_token`` is appended when present but is deliberately optional -- it
expires, and the id alone is what keeps a long-running task alive.
Always emits a full URL rather than a bare id: both platforms' parsers accept
a bare id only in a narrower form, so the URL is the safer universal input.
The shape itself is platform-specific and comes from ``adapters``.
"""
path = _CREATOR_PATH if kind == MODE_CREATOR else _NOTE_PATH
return f"{_XHS_WEB_BASE}{path}/{value}"
spec = adapters.adapter(platform)
path = spec.creator_path if kind == MODE_CREATOR else spec.note_path
return f"{spec.web_base}{path}/{value}"
def build_target_urls(mode: str, targets: Iterable[MonitorTarget]) -> List[str]:
def build_target_urls(
mode: str, targets: Iterable[MonitorTarget], platform: str
) -> List[str]:
"""存储的目标 -> 爬虫接受的 URL。
``xsec_token`` 只有小红书有,而且是会过期的刷新令牌 —— 有就带上,没有就算了。
抖音恒为空,所以这一段对它是天然的 no-op,不需要平台分支。
"""
urls = []
for target in targets:
url = build_target_url(target.external_id, target.kind)
url = build_target_url(target.external_id, target.kind, platform)
if target.xsec_token:
url = f"{url}?xsec_token={target.xsec_token}"
if target.xsec_source:
@@ -159,9 +165,14 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
raise ValueError(f"Monitor task {task_id} has no enabled targets")
platform = task.platform
urls = build_target_urls(task.mode, targets)
urls = build_target_urls(task.mode, targets, platform)
cookie = await get_cookie(session, platform)
strategy = await _strategy_settings(session, platform)
# System-wide switch. On a headless server the crawler must attach to the
# Chrome already listening on the debug port -- that browser is where the
# operator scanned the login QR, so its profile is the login. Default False
# keeps desktop runs launching a private browser exactly as before.
cdp_enabled = await app_settings.get_value(session, "cdp_enabled", fallback=False)
run = MonitorRun(
task_id=task.id,
@@ -186,55 +197,118 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
max_notes_count = task.max_notes_count
max_comments_count = task.max_comments_count
timeout_seconds = task.run_timeout_seconds
# 抖音的作品列表接口被那道真校验挡着(见 douyin_fetch),拿不到列表时就靠这些
# 已知的 aweme_id 逐条刷新 —— 新作品发现不了,但已有作品的指标还能继续更新。
known_aweme_ids = list(
await session.scalars(
select(MonitorNote.note_id).where(MonitorNote.task_id == task.id)
)
)
# --- Phase 2: run the crawler outside any transaction ---------------------
cookie_file = out_dir / ".cookies"
_write_cookie_file(cookie_file, cookie)
request = CrawlerStartRequest(
platform=_PLATFORM_ENUM[platform],
login_type=LoginTypeEnum.COOKIE,
crawler_type=CrawlerTypeEnum.CREATOR if mode == MODE_CREATOR else CrawlerTypeEnum.DETAIL,
creator_ids=",".join(urls) if mode == MODE_CREATOR else "",
specified_ids=",".join(urls) if mode != MODE_CREATOR else "",
start_page=1,
enable_comments=enable_comments,
enable_sub_comments=strategy["enable_sub_comments"],
enable_media=False,
save_option=SaveDataOptionEnum.JSONL,
cookies="",
headless=True,
max_notes_count=max_notes_count,
max_comments_count=max_comments_count,
# Isolate this run's output: the crawler names files by date only, so
# otherwise same-day runs would append into one shared file.
save_data_path=str(out_dir),
# Unattended runs must not try to attach to the user's desktop Chrome.
enable_cdp_mode=False,
# Only injecting web_session is not enough to sign requests from a cold
# browser profile.
inject_all_cookies=True,
save_login_state=True,
cookies_file=str(cookie_file),
max_concurrency_num=1,
# Strategy + proxy, surfaced on the Settings page.
crawler_max_sleep_sec=strategy["crawl_sleep_sec"],
enable_ip_proxy=strategy["enable_ip_proxy"],
ip_proxy_pool_count=strategy["proxy_pool_count"],
ip_proxy_provider_name=strategy["proxy_provider"],
static_proxy_url=strategy["static_proxy_url"] or None,
)
# --- Phase 2: run the collection outside any transaction ------------------
# 抖音走**进程内 HTTP 客户端**,不起 Playwright 子进程:那边会构造一大串自相矛盾的
# 浏览器指纹参数(参数说 Mac + Chrome 125、UA 说 Linux + Chrome 155),网关回一个
# 200 + 空 body,然后被翻译成「account blocked」—— 看着像账号被封,其实什么都不是。
# 见 douyin_fetch / douyin_api。
# 先落「运行中」—— **两条路都要**。原先这一行只写在爬虫那条分支里,于是抖音那条路上
# run 一直停在 pending;一旦中途出事(异常、进程被重启),界面上就是一个永远
# 「排队中」的幽灵,而且 recover() 也只清理 running、收不到它。
async with get_session() as session:
run = await session.get(MonitorRun, run_id)
if run is not None:
run.status = RUN_RUNNING
run.started_at = get_current_timestamp()
try:
exit_code = await crawler_manager.run_and_wait(request, timeout=timeout_seconds)
finally:
_remove_cookie_file(cookie_file)
in_process_tail: List[str] = []
if platform == adapters.PLATFORM_DY:
try:
# 也要有超时。爬虫那条路靠 run_and_wait(timeout=...) 兜底,这条路没有子进程、
# 没人管 —— 里面**任何一次卡住都会让 run 永远停在「运行中」**(真踩过:
# page.evaluate 打在一个卡死的标签页上不返回)。
fetched = await asyncio.wait_for(
douyin_fetch.collect(
out_dir,
platform=platform,
mode=mode,
limit=max_notes_count,
want_comments=enable_comments,
comment_limit=max_comments_count,
targets=targets,
known_aweme_ids=known_aweme_ids,
cookie=cookie,
),
timeout=timeout_seconds,
)
except asyncio.TimeoutError:
fetched = {
"notes": 0,
"comments": 0,
"errors": [f"抖音采集超过 {timeout_seconds} 秒仍未完成,已放弃这一轮"],
"jsonl_dir": "",
}
in_process_tail = list(fetched["errors"])
# 一条都没采到 = 这一轮失败,并把**真因**当作退出诊断传下去。否则它会掉进
# ingest 的「疑似登录失效」分支 —— 又骗人一次,正是这套东西一直在犯的毛病。
exit_code = 1 if (fetched["errors"] and not fetched["notes"]) else 0
if fetched["errors"] and fetched["notes"]:
# 有产物但带着错误,说明走了退化路径(比如作品列表被挡,只刷新了已知作品)。
# 这一轮状态是成功,但**不是**一切正常 —— 得留下痕迹,否则没人知道新作品
# 其实没在发现。
print(
"[monitor.runner] 抖音采集部分失败:"
+ ";".join(fetched["errors"])[:300]
)
else:
cookie_file = out_dir / ".cookies"
_write_cookie_file(cookie_file, cookie)
request = CrawlerStartRequest(
platform=_PLATFORM_ENUM[platform],
login_type=LoginTypeEnum.COOKIE,
crawler_type=CrawlerTypeEnum.CREATOR if mode == MODE_CREATOR else CrawlerTypeEnum.DETAIL,
creator_ids=",".join(urls) if mode == MODE_CREATOR else "",
specified_ids=",".join(urls) if mode != MODE_CREATOR else "",
start_page=1,
enable_comments=enable_comments,
enable_sub_comments=strategy["enable_sub_comments"],
enable_media=False,
save_option=SaveDataOptionEnum.JSONL,
cookies="",
headless=True,
max_notes_count=max_notes_count,
max_comments_count=max_comments_count,
# Isolate this run's output: the crawler names files by date only, so
# otherwise same-day runs would append into one shared file.
save_data_path=str(out_dir),
# Attach to the browser already running on CDP_DEBUG_PORT when the
# operator enabled it; otherwise launch a private, throwaway browser.
enable_cdp_mode=cdp_enabled,
# Only injecting web_session is not enough to sign requests from a cold
# browser profile.
inject_all_cookies=True,
save_login_state=True,
cookies_file=str(cookie_file),
max_concurrency_num=1,
# Strategy + proxy, surfaced on the Settings page.
crawler_max_sleep_sec=strategy["crawl_sleep_sec"],
enable_ip_proxy=strategy["enable_ip_proxy"],
ip_proxy_pool_count=strategy["proxy_pool_count"],
ip_proxy_provider_name=strategy["proxy_provider"],
static_proxy_url=strategy["static_proxy_url"] or None,
)
try:
exit_code = await crawler_manager.run_and_wait(request, timeout=timeout_seconds)
finally:
_remove_cookie_file(cookie_file)
# 失败时的诊断来源。抖音那条路没有子进程,尾巴就是它自己报的错 —— **别去读
# crawler_manager 的尾巴**,那里面是上一轮别的平台留下的东西,会张冠李戴。
output_tail = (
in_process_tail
if platform == adapters.PLATFORM_DY
else crawler_manager.get_output_tail()
)
# --- Phase 3: ingest ------------------------------------------------------
async with get_session() as session:
@@ -243,17 +317,27 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
if run is None or task is None:
raise ValueError(f"Run {run_id} or task {task_id} vanished during execution")
if exit_code == -1 and not (out_dir / "xhs").exists():
# 目录名按平台解析 —— 抖音的平台 id 是 dy 而产物目录是 douyin,写死就永远判不准。
if exit_code == -1 and not (out_dir / adapters.artifact_dir(task.platform)).exists():
# run_and_wait returns -1 when the process could not start or timed out.
run.status = RUN_TIMEOUT
run.finished_at = get_current_timestamp()
run.exit_code = exit_code
run.error_message = "Run was killed by timeout or failed to start"
# -1 同时代表「超时」和「根本没起来」,两者要查的东西完全不同。输出尾巴里
# 有异常就带上它,否则运行历史里只能看到这句没有信息量的话。
cause = diagnose_failure(output_tail)
if cause:
run.error_message = f"{run.error_message};原因:{cause}"
result = IngestResult(status=RUN_TIMEOUT, error=run.error_message)
else:
run.exit_code = exit_code
run.finished_at = get_current_timestamp()
result = await ingest_run(session, run, task, out_dir)
# 把爬虫输出的尾巴交给 ingest:退出码本身说明不了问题,运行历史里要显示的
# 是真正的报错(比如抖音的 DataFetchError: account blocked)。
result = await ingest_run(
session, run, task, out_dir, output_tail=output_tail
)
# A run that authenticated fine is the only useful signal that the
# stored cookie still works.
@@ -274,4 +358,15 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
if task is not None and run is not None:
await notify.notify_run(session, task, run)
# --- Phase 5: 封面落盘 ------------------------------------------------------
# 也放在事务之外。封面地址带签名、会过期(实测隔天即 403),落盘之后才与签名无关。
# 下载慢且可能失败,占着一个入库事务是不合适的;失败也不影响本轮数据。
try:
async with get_session() as session:
saved = await covers.cache_pending(session, task_id)
if saved:
print(f"[monitor.runner] 缓存了 {saved} 张作品封面")
except Exception as exc: # noqa: BLE001 - 封面拿不到不该让整轮失败
print(f"[monitor.runner] 封面缓存失败:{exc}")
return result
+167
View File
@@ -0,0 +1,167 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/schedule.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""Monitor task schedule arithmetic.
Three modes, all expressible by a picker. A raw cron string was ruled out on
purpose -- it is a small language, and the operator should not have to write one
to say "every day at nine":
* ``interval`` -- every N minutes.
* ``daily`` -- at chosen clock times, e.g. 09:00 and 18:30.
* ``weekly`` -- at chosen clock times on chosen weekdays, e.g. Mon-Fri 10:00.
The two clock modes are **fixed-time**, unlike ``interval``, which is fixed-delay.
The distinction matters: for an interval, measuring the next slot from when the
run starts is what stops a slow run from firing back-to-back. For a clock
schedule it would be wrong, because a run that starts at 09:07 would drag every
later run seven minutes late, compounding all day.
Fixed-time also means **no jitter is applied** to clock schedules. The operator
picked a time; quietly running at 09:04 instead of 09:00 is not a feature, it just
looks like a bug. ``interval`` keeps its jitter, where there is no stated time to
contradict.
All arithmetic is in the server's local timezone -- naive datetimes on purpose,
because the container is pinned to the operator's zone via TZ and pretending
otherwise would add a timezone concept nobody asked for.
"""
from datetime import datetime, time, timedelta
from typing import Optional, Sequence
MODE_INTERVAL = "interval"
MODE_DAILY = "daily"
MODE_WEEKLY = "weekly"
SCHEDULE_MODES = (MODE_INTERVAL, MODE_DAILY, MODE_WEEKLY)
CLOCK_MODES = (MODE_DAILY, MODE_WEEKLY)
_MS_PER_MINUTE = 60_000
_WEEKDAY_NAMES = "一二三四五六日"
def parse_hours(raw: Optional[str]) -> list[int]:
"""``"9,18"`` -> ``[9, 18]``. Sorted, de-duplicated, junk dropped."""
return _parse_int_list(raw, 0, 23)
def parse_days(raw: Optional[str]) -> list[int]:
"""``"0,2,4"`` -> ``[0, 2, 4]``. **0 is Monday**, matching ``date.weekday()``."""
return _parse_int_list(raw, 0, 6)
def _parse_int_list(raw: Optional[str], low: int, high: int) -> list[int]:
values: set[int] = set()
for chunk in (raw or "").split(","):
chunk = chunk.strip()
if not chunk:
continue
try:
number = int(chunk)
except ValueError:
# Stored values come from our own UI, but a hand-edited row must not
# be able to crash the scheduler loop.
continue
if low <= number <= high:
values.add(number)
return sorted(values)
def format_hours(hours: Sequence[int]) -> str:
return ",".join(str(hour) for hour in sorted(set(hours)))
def format_days(days: Sequence[int]) -> str:
return ",".join(str(day) for day in sorted(set(days)))
def describe(
*,
mode: str,
interval_minutes: int,
hours: Sequence[int],
days: Sequence[int],
minute: int,
) -> str:
"""One human sentence for the task card.
Lives here rather than in the frontend so the list view and the editor cannot
drift apart on what a schedule means.
"""
if mode == MODE_INTERVAL:
if interval_minutes % 1440 == 0:
return f"每 {interval_minutes // 1440} 天"
if interval_minutes % 60 == 0:
return f"每 {interval_minutes // 60} 小时"
return f"每 {interval_minutes} 分钟"
if not hours:
return "未设置时间"
clock = "、".join(f"{hour:02d}:{minute:02d}" for hour in sorted(set(hours)))
if mode == MODE_DAILY:
return f"每天 {clock}"
if not days:
return f"每天 {clock}"
labels = "、".join(f"周{_WEEKDAY_NAMES[day]}" for day in sorted(set(days)))
return f"{labels} {clock}"
def next_occurrence(
*,
mode: str,
interval_minutes: int,
hours: Sequence[int],
days: Sequence[int],
minute: int,
after_ms: int,
) -> Optional[int]:
"""The next fire time strictly after ``after_ms``, as epoch milliseconds.
``None`` means the schedule can never fire -- a clock mode with no hours
chosen. Callers store that as "no next run" rather than something in the past,
which would otherwise leave the task permanently due and re-running on every
tick.
"""
if mode == MODE_INTERVAL:
return after_ms + max(1, interval_minutes) * _MS_PER_MINUTE
if not hours:
return None
now = datetime.fromtimestamp(after_ms / 1000)
# No weekdays chosen means every day, matching describe(). Without the `days`
# guard an empty selection would produce an empty allowed set, no matching day,
# and a task that silently never runs.
allowed_days = set(days) if (mode == MODE_WEEKLY and days) else set(range(7))
# Eight days of lookahead covers today's remaining slots plus a full week,
# which is more than any weekday selection can need.
for offset in range(8):
day = (now + timedelta(days=offset)).date()
if day.weekday() not in allowed_days:
continue
for hour in sorted(set(hours)):
candidate = datetime.combine(day, time(hour=hour, minute=minute))
if candidate > now:
return int(candidate.timestamp() * 1000)
return None
+120 -25
View File
@@ -23,8 +23,20 @@ is enough here: there is exactly one process, one global crawler subprocess, and
therefore no concurrency to coordinate -- a cron-style library would add a
dependency without adding a capability.
Scheduling is **fixed-delay**, not fixed-rate: ``next_run_at`` is set from the
moment a run starts, so a slow run cannot make its task fire back-to-back.
Two families of schedule, and the difference matters:
* ``interval`` is **fixed-delay**, not fixed-rate -- ``next_run_at`` is measured
from the moment a run starts, so a slow run cannot make its task fire
back-to-back.
* the clock modes (``daily``/``weekly``) are **fixed-time** -- recomputed from the
calendar, so a run that starts late does not drag every later run with it.
The arithmetic for both lives in schedule.py.
The loop also carries the 上游更新检查: it is not a crawl, so it shares none of
the rules above (no subprocess, no active-hours gate) -- see
``_maybe_check_upstream``. It rides this loop rather than getting a thread of its
own because it is one HTTP-shaped fetch per day.
"""
import asyncio
@@ -37,18 +49,23 @@ from sqlalchemy import select
from tools.time_util import get_current_timestamp
from ..services import crawler_manager
from . import app_settings
from . import app_settings, schedule, upstream
from .db import get_session
from .models import MonitorRun, MonitorTask, RUN_INTERRUPTED, RUN_RUNNING
from .models import (
RUN_INTERRUPTED,
RUN_PENDING,
RUN_RUNNING,
MonitorRun,
MonitorTask,
)
from .runner import execute_task
from .settings import get_cookie
POLL_INTERVAL_SECONDS = 20
# Spread tasks sharing an interval so they do not all come due on the same tick.
# Applied to interval mode only -- see the advance step below.
JITTER_SECONDS = 60
_MS_PER_MINUTE = 60_000
class MonitorScheduler:
"""Polls the task table and runs whatever is due."""
@@ -56,8 +73,9 @@ class MonitorScheduler:
def __init__(self) -> None:
self._loop_task: Optional[asyncio.Task] = None
self._stopping = asyncio.Event()
# Avoids logging "no cookie" on every single tick.
self._warned_no_cookie = False
# Avoids logging "no cookie" on every single tick. Per platform, because
# warning once for Xiaohongshu must not silence the warning for Douyin.
self._warned_no_cookie: set = set()
async def start(self) -> None:
if self._loop_task is not None and not self._loop_task.done():
@@ -86,19 +104,66 @@ class MonitorScheduler:
await self.tick()
except Exception as exc: # pragma: no cover - keep the loop alive
print(f"[monitor.scheduler] tick failed: {exc}")
# 独立于采集任务,因此单独一段 try:上游检查失败不该影响采集调度,
# 反过来也一样。
try:
await self._maybe_check_upstream()
except Exception as exc: # pragma: no cover - keep the loop alive
print(f"[monitor.scheduler] upstream check failed: {exc}")
await asyncio.sleep(POLL_INTERVAL_SECONDS)
async def _maybe_check_upstream(self) -> None:
"""到点就 fetch 一次上游仓库,看它有没有新提交。
与采集任务的三条规则都不同,各有理由:它不碰浏览器、也不占采集子进程,
所以不看 ``is_busy``;它只发一个 git 请求,没有被平台风控的风险,所以也不
受活跃时段限制 —— 定时检查放在半夜反而是最合适的。
"""
async with get_session() as session:
if not await app_settings.get_value(
session, "upstream_check_enabled", fallback=False
):
return
interval_minutes = int(
await app_settings.get_value(
session, "upstream_check_interval_minutes", fallback=1440
)
)
state = await upstream.load_state(session)
checked_at = int(state.get("checked_at") or 0)
now = get_current_timestamp()
# 失败也会写 checked_at,所以不通的时候同样是每个间隔重试一次,
# 而不是每个 tick(20 秒)都去撞一次墙。
if checked_at and now - checked_at < max(1, interval_minutes) * 60_000:
return
result = await upstream.run_check()
if result.get("behind"):
print(
f"[monitor.scheduler] 上游 {result.get('branch')} 领先 "
f"{result['behind']} 个提交"
)
elif not result.get("ok"):
print(f"[monitor.scheduler] 上游检查失败:{result.get('error')}")
async def recover(self) -> None:
"""Clean up state left behind by a server restart.
A run still marked ``running`` cannot be running -- its subprocess died
with the previous process. Marking it interrupted stops it from blocking
the UI as a phantom in-flight run.
**``pending`` 同样是残留**:那一行是上一轮建的,可它后面的采集根本没机会开始
(进程被重启,或者采集那条路抛了异常),所以它永远不会自己往前走。只清 running
的话,它会永远挂在界面上显示「排队中」—— 用户看到的就是任务卡住了。
"""
async with get_session() as session:
stale = (
await session.scalars(
select(MonitorRun).where(MonitorRun.status == RUN_RUNNING)
select(MonitorRun).where(
MonitorRun.status.in_((RUN_RUNNING, RUN_PENDING))
)
)
).all()
for run in stale:
@@ -151,28 +216,58 @@ class MonitorScheduler:
if task is None:
return
# No cookie means every run would report an auth failure. Leave the
# task due rather than advancing: it starts working the moment the
# user pastes one.
cookie = await get_cookie(session)
# 没有 cookie 就跳过,是为了不让任务每轮白跑一趟出个认证失败。任务留在
# due 状态而不推进 —— 用户一粘上 cookie 它就能自己跑起来。
#
# **但开着 CDP 时必须放行**:那种模式下登录态来自被接管的那个浏览器,
# 粘不粘 cookie 根本轮不到它决定成败。不放行的话,选了「接管已有 Chrome」
# 却没粘 cookie 的用户会发现任务永远不被触发,而且什么错都不报。
cookie = await get_cookie(session, task.platform)
if not cookie:
if not self._warned_no_cookie:
print(
"[monitor.scheduler] no XHS cookie configured; "
"scheduled tasks will not run until one is set"
)
self._warned_no_cookie = True
return
self._warned_no_cookie = False
cdp_enabled = await app_settings.get_value(
session, "cdp_enabled", fallback=False
)
if not cdp_enabled:
if task.platform not in self._warned_no_cookie:
print(
f"[monitor.scheduler] no {task.platform} cookie configured; "
"scheduled tasks will not run until one is set or CDP is enabled"
)
self._warned_no_cookie.add(task.platform)
return
self._warned_no_cookie.discard(task.platform)
# Advance before running so a crash mid-run cannot cause an immediate
# re-fire, and so a long outage coalesces into a single run instead
# of one run per missed interval.
task.next_run_at = (
get_current_timestamp()
+ task.interval_minutes * _MS_PER_MINUTE
+ random.randint(0, JITTER_SECONDS) * 1000
now = get_current_timestamp()
following = schedule.next_occurrence(
mode=task.schedule_mode,
interval_minutes=task.interval_minutes,
hours=schedule.parse_hours(task.schedule_hours),
days=schedule.parse_days(task.schedule_days),
minute=task.schedule_minute,
after_ms=now,
)
if following is None:
# A clock schedule with no times can never fire. The API rejects
# that shape, so this guards against a hand-edited row: park the
# task with no next run rather than leaving it permanently due and
# re-running it on every tick.
task.next_run_at = None
print(
f"[monitor.scheduler] task {task.id} has no usable schedule "
f"and will not run until one is set"
)
elif task.schedule_mode == schedule.MODE_INTERVAL:
# Jitter belongs to the interval mode only. Spreading identical
# intervals apart is the point; nudging a time the operator
# explicitly picked is not -- it just looks like a broken clock.
task.next_run_at = following + random.randint(0, JITTER_SECONDS) * 1000
else:
task.next_run_at = following
task_id = task.id
try:
+398 -26
View File
@@ -19,8 +19,7 @@
"""Task CRUD and dashboard queries for the monitoring layer."""
import asyncio
import re
from typing import Any, Dict, List, Optional
from typing import Any, Dict, List, Optional, Sequence
from urllib.parse import parse_qs, urlparse
from sqlalchemy import delete, func, select
@@ -28,15 +27,18 @@ from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from . import app_settings, platforms
from . import adapters, app_settings, covers, platforms, schedule
from .db import get_session
from .platforms import PLATFORM_XHS
from .models import (
MODE_CREATOR,
MODE_NOTE,
MonitorComment,
MonitorCreatorAlias,
MonitorCreatorStat,
MonitorEvent,
MonitorNote,
MonitorNoteAlias,
MonitorNoteMetric,
MonitorRun,
MonitorTarget,
@@ -53,11 +55,30 @@ _background_runs: set[asyncio.Task] = set()
MIN_INTERVAL_MINUTES = 30
MAX_INTERVAL_MINUTES = 7 * 24 * 60
_CREATOR_URL_RE = re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)")
_NOTE_URL_RE = re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)")
# XHS user ids and note ids are 24-char hex; allow a slightly wider range so a
# format change degrades into "still accepted" rather than "rejected".
_BARE_ID_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
# Changing any of these invalidates the pending run slot.
SCHEDULE_FIELDS = {
"interval_minutes",
"schedule_mode",
"schedule_hours",
"schedule_days",
"schedule_minute",
}
def _require_clock_fields(mode: str, hours: list, days: list) -> None:
"""A clock schedule with no clock time can never fire.
The create schema already rejects that shape, but the update path merges
partial fields and therefore has no schema-level view of the result -- so the
check lives here, where both paths meet.
"""
if mode in schedule.CLOCK_MODES and not hours:
raise ValueError("按钟点调度至少要选一个时间")
if mode == schedule.MODE_WEEKLY and not days:
raise ValueError("按周调度至少要选一个星期")
# 各平台的链接形态、id 形状、短链域名都在 adapters.py —— 那里是「平台之间不一样」
# 的东西的唯一出处,所以这里不再留任何平台字面量。
class TargetParseError(ValueError):
@@ -73,27 +94,36 @@ def parse_target_input(
Storing the id separately from the token is what keeps a long-running task
alive: tokens expire, ids do not.
URL shapes are platform-specific. Only Xiaohongshu is wired, so anything else
is rejected here as well as at task creation -- parsing a Douyin link as if it
were a Xiaohongshu one would be worse than refusing it.
URL shapes are platform-specific and come from ``adapters``. A platform with
no adapter is rejected here as well as at task creation -- parsing a Douyin
link as if it were a Xiaohongshu one would be worse than refusing it.
"""
if platform != PLATFORM_XHS:
if not adapters.has_adapter(platform):
raise TargetParseError(f"暂不支持解析该平台({platform})的目标链接")
spec = adapters.adapter(platform)
raw = (value or "").strip()
if not raw:
raise TargetParseError("Empty target")
creator_mode = mode == MODE_CREATOR
expected = "博主主页" if creator_mode else "笔记"
external_id = ""
if raw.startswith("http") or "/" in raw:
# xhslink.com and other short links are not resolvable without a network
# round-trip, so only the direct profile/explore forms are supported.
match = _CREATOR_URL_RE.search(raw) if mode == MODE_CREATOR else _NOTE_URL_RE.search(raw)
if not match:
expected = "博主主页" if mode == MODE_CREATOR else "笔记"
if any(host in raw for host in spec.short_link_hosts):
# 短链要联网跳一次才知道指向谁,而这里没有网络可跳。明确拒绝好过存一个
# 解析不出 id 的值 —— 那会变成一个永远抓不到东西、还不报错的任务。
raise TargetParseError(f"{expected}短链无法解析,请粘贴完整链接:{raw}")
patterns = spec.creator_url_res if creator_mode else spec.note_url_res
for pattern in patterns:
match = pattern.search(raw)
if match:
external_id = match.group(1)
break
if not external_id:
raise TargetParseError(f"无法从链接中解析出{expected} ID:{raw}")
external_id = match.group(1)
elif _BARE_ID_RE.match(raw):
elif (spec.creator_bare_re if creator_mode else spec.note_bare_re).match(raw):
external_id = raw
else:
raise TargetParseError(f"无法识别的目标:{raw}")
@@ -139,7 +169,12 @@ async def create_task(session: AsyncSession, payload: Dict[str, Any]) -> Monitor
# the Settings page actually governs new tasks.
defaults = await app_settings.defaults(session, platform)
interval_minutes = payload.get("interval_minutes") or defaults["interval_minutes"]
interval_ms = int(interval_minutes) * 60_000
schedule_mode = payload.get("schedule_mode") or schedule.MODE_INTERVAL
schedule_hours = list(payload.get("schedule_hours") or [])
schedule_days = list(payload.get("schedule_days") or [])
schedule_minute = int(payload.get("schedule_minute") or 0)
_require_clock_fields(schedule_mode, schedule_hours, schedule_days)
task = MonitorTask(
name=payload["name"],
@@ -147,12 +182,24 @@ async def create_task(session: AsyncSession, payload: Dict[str, Any]) -> Monitor
mode=mode,
enabled=payload.get("enabled", True),
interval_minutes=interval_minutes,
schedule_mode=schedule_mode,
schedule_hours=schedule.format_hours(schedule_hours),
schedule_days=schedule.format_days(schedule_days),
schedule_minute=schedule_minute,
max_notes_count=payload.get("max_notes_count") or defaults["max_notes_count"],
enable_comments=payload.get("enable_comments", True),
max_comments_count=payload.get("max_comments_count") or defaults["max_comments_count"],
run_timeout_seconds=payload.get("run_timeout_seconds", 3600),
notify_enabled=payload.get("notify_enabled", False),
next_run_at=now + interval_ms,
notify_failures=payload.get("notify_failures", True),
next_run_at=schedule.next_occurrence(
mode=schedule_mode,
interval_minutes=interval_minutes,
hours=schedule_hours,
days=schedule_days,
minute=schedule_minute,
after_ms=now,
),
last_status="idle",
created_at=now,
updated_at=now,
@@ -189,19 +236,31 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
if task is None:
raise ValueError(f"Task {task_id} not found")
was_enabled = task.enabled
for field in (
"name",
"enabled",
"interval_minutes",
"schedule_mode",
"schedule_minute",
"max_notes_count",
"enable_comments",
"max_comments_count",
"run_timeout_seconds",
"notify_enabled",
"notify_failures",
):
if field in payload and payload[field] is not None:
setattr(task, field, payload[field])
# The clock lists are stored as comma-separated text, so they cannot go
# through the generic loop above.
if payload.get("schedule_hours") is not None:
task.schedule_hours = schedule.format_hours(payload["schedule_hours"])
if payload.get("schedule_days") is not None:
task.schedule_days = schedule.format_days(payload["schedule_days"])
# Replacing targets resets the baseline implicitly: a note set that now
# includes new ids will simply report them as new on the next run.
if payload.get("targets") is not None:
@@ -209,7 +268,7 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
now = get_current_timestamp()
seen: set[str] = set()
for value in payload["targets"]:
parsed = parse_target_input(value, task.mode)
parsed = parse_target_input(value, task.mode, task.platform)
if parsed["external_id"] in seen:
continue
seen.add(parsed["external_id"])
@@ -227,8 +286,23 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
)
)
if "interval_minutes" in payload and payload["interval_minutes"]:
task.next_run_at = get_current_timestamp() + payload["interval_minutes"] * 60_000
# Any change to when the task runs invalidates the pending slot, so recompute
# it from the merged state rather than working out which field moved.
# Re-enabling counts as a change too: otherwise a task switched off for a
# month comes back holding a next_run_at a month in the past and fires the
# instant it is saved.
if (SCHEDULE_FIELDS & set(payload)) or (task.enabled and not was_enabled):
hours = schedule.parse_hours(task.schedule_hours)
days = schedule.parse_days(task.schedule_days)
_require_clock_fields(task.schedule_mode, hours, days)
task.next_run_at = schedule.next_occurrence(
mode=task.schedule_mode,
interval_minutes=task.interval_minutes,
hours=hours,
days=days,
minute=task.schedule_minute,
after_ms=get_current_timestamp(),
)
task.updated_at = get_current_timestamp()
await session.flush()
@@ -257,6 +331,48 @@ def trigger_manual_run(task_id: int) -> None:
# Dashboard queries
# ---------------------------------------------------------------------------
async def _creator_alias_map(session: AsyncSession) -> Dict[tuple, str]:
"""``(platform, creator_hash) -> 备注``。
整体读一次再在内存里取,而不是每条作品查一次 —— 作品列表动辄几十条。备注本身很少,
表不会大。
"""
rows = (await session.scalars(select(MonitorCreatorAlias))).all()
return {(row.platform, row.creator_hash): row.alias for row in rows if row.alias}
async def set_creator_alias(
session: AsyncSession, platform: str, creator_hash: str, alias: str
) -> None:
"""给博主起备注;传空串就是删掉这条备注(界面上的"清空")。"""
alias = (alias or "").strip()[:128]
existing = await session.scalar(
select(MonitorCreatorAlias).where(
MonitorCreatorAlias.platform == platform,
MonitorCreatorAlias.creator_hash == creator_hash,
)
)
if not alias:
if existing is not None:
await session.delete(existing)
return
if existing is None:
session.add(
MonitorCreatorAlias(
platform=platform,
creator_hash=creator_hash,
alias=alias,
updated_at=get_current_timestamp(),
)
)
return
existing.alias = alias
existing.updated_at = get_current_timestamp()
async def _latest_successful_run_id(session: AsyncSession, task_id: int) -> Optional[int]:
return await session.scalar(
select(MonitorRun.id)
@@ -275,6 +391,75 @@ def _delta(current: Optional[int], previous: Optional[int]) -> Optional[int]:
return current - previous
async def _note_alias_map(session: AsyncSession) -> Dict[tuple, str]:
"""``(platform, note_id) -> 作品备注``。和博主备注一个道理,整体读一次。"""
rows = (await session.scalars(select(MonitorNoteAlias))).all()
return {(row.platform, row.note_id): row.alias for row in rows if row.alias}
async def set_note_alias(
session: AsyncSession, platform: str, note_id: str, alias: str
) -> None:
"""给作品起备注;空串就是删掉这条备注。"""
alias = (alias or "").strip()[:128]
existing = await session.scalar(
select(MonitorNoteAlias).where(
MonitorNoteAlias.platform == platform,
MonitorNoteAlias.note_id == note_id,
)
)
if not alias:
if existing is not None:
await session.delete(existing)
return
if existing is None:
session.add(
MonitorNoteAlias(
platform=platform,
note_id=note_id,
alias=alias,
updated_at=get_current_timestamp(),
)
)
return
existing.alias = alias
existing.updated_at = get_current_timestamp()
async def _latest_creator_stats(
session: AsyncSession,
task_ids: Optional[Sequence[int]] = None,
creator_hashes: Optional[Sequence[str]] = None,
) -> Dict[tuple, "MonitorCreatorStat"]:
"""``(task_id, creator_hash) -> 最近一条``账号级快照。
两个参数都是**可选过滤**:传 None 就是不限。``list_notes`` 两个都给(只要手里
这批作品涉及的博主),``list_creators`` 只给任务(要的是全部博主,包括一条作品
都没有的那些)。
按 run_id 而不是 captured_at 取「最近」:和作品指标用的是同一个口径,两者放一起
看才不会出现「作品数据来自第 8 轮、粉丝数来自第 9 轮」这种对不上的情况。
"""
if task_ids is not None and not task_ids:
return {}
if creator_hashes is not None and not creator_hashes:
return {}
query = select(MonitorCreatorStat).order_by(MonitorCreatorStat.run_id.desc())
if task_ids is not None:
query = query.where(MonitorCreatorStat.task_id.in_(list(task_ids)))
if creator_hashes is not None:
query = query.where(MonitorCreatorStat.creator_hash.in_(list(creator_hashes)))
latest: Dict[tuple, MonitorCreatorStat] = {}
for row in (await session.scalars(query)).all():
latest.setdefault((row.task_id, row.creator_hash), row) # 已按 run_id 倒序
return latest
async def list_notes(
session: AsyncSession,
task_id: Optional[int] = None,
@@ -316,6 +501,29 @@ async def list_notes(
latest_run_ids: Dict[int, Optional[int]] = {}
result: List[Dict[str, Any]] = []
# 博主的备注。键是 (platform, creator_hash) —— 作品的 platform 挂在它的任务上。
task_platform = {
row.id: row.platform
for row in (
await session.execute(
select(MonitorTask.id, MonitorTask.platform).where(
MonitorTask.id.in_({note.task_id for note in notes})
)
)
).all()
}
aliases = await _creator_alias_map(session)
note_aliases = await _note_alias_map(session)
# 账号级指标(粉丝 / 总获赞 / 作品数)。**挂在作品上一起返回**,因为界面上就是按博主
# 归组显示的 —— 让前端为了一个组头再发一轮请求没道理。同一个博主的所有作品拿到的是
# 同一条(键里带 task_id,所以跨任务不会串)。没有的(小红书那条路不产生它)就是 null,
# 前端据此整块不显示,而不是显示一个 0。
creator_stats = await _latest_creator_stats(
session,
[note.task_id for note in notes],
[note.creator_hash for note in notes if note.creator_hash],
)
for note in notes:
series = by_note.get(note.note_id, [])
current = series[0] if series else None
@@ -327,13 +535,42 @@ async def list_notes(
if note.first_seen_run_id != latest_run_ids[note.task_id]:
continue
stat = creator_stats.get((note.task_id, note.creator_hash))
result.append(
{
"task_id": note.task_id,
"note_id": note.note_id,
"title": note.title,
"note_url": note.note_url,
"cover": note.cover,
# 作品的发布时间(爬虫侧:小红书叫 time、抖音叫 create_time)。
# 和 first_seen_at 不是一回事 —— 那是**我们第一次看到它**的时间;把一个
# 早就存在的作品加进监控时,两者能差好几个月。可能为 null(平台没给,
# 或者值解析不出来),所以前端要能显示成「—」。
"published_at": note.published_at,
# 按博主分组用。creator_hash 是唯一稳定的创作者标识(原始 user_id
# 被爬虫刻意匿名化了),creator_name 是昵称本身 —— 本仓库关掉了脱敏
# (见 config.MASK_NICKNAME),所以就是原文。
"creator_hash": note.creator_hash,
"creator_name": note.creator_name,
# 人自己起的备注,界面上优先显示它 —— 昵称认不出是谁,哈希更认不出。
"creator_alias": aliases.get(
(task_platform.get(note.task_id, ""), note.creator_hash), ""
),
# 这条作品自己的备注。和博主备注是两件事:博主备注回答"这是谁",它回答
# "这条我要盯着"。
"note_alias": note_aliases.get(
(task_platform.get(note.task_id, ""), note.note_id), ""
),
# 博主账号级指标 —— 作品列表给不了的东西。三个值都可能为 null(平台没采
# 到、或者这条作品来自不产生它的数据源),前端据此整块不画。
"creator_fans": stat.fans if stat else None,
"creator_total_favorited": stat.total_favorited if stat else None,
"creator_works": stat.works_count if stat else None,
"creator_stats_at": stat.captured_at if stat else None,
# 优先给本地缓存地址:远程地址带签名、会过期(实测隔天即 403),
# 本地那份不会。没有缓存时才退回远程,至少让图先显示出来。
"cover": covers.cover_url(note.note_id, note.cover),
"first_seen_at": note.first_seen_at,
"last_seen_at": note.last_seen_at,
"is_new": note.first_seen_run_id == latest_run_ids.get(note.task_id),
@@ -368,6 +605,107 @@ async def list_notes(
return result
async def list_creators(
session: AsyncSession,
task_id: Optional[int] = None,
platform: Optional[str] = None,
) -> List[Dict[str, Any]]:
"""作品栏里要显示的**博主** —— **包括一条作品都没有的**。
分组原先是从作品推出来的(按作品的 creator_hash 归组),于是没有作品的博主根本
不会出现在列表里:目标加了、资料也采到了、粉丝数就躺在库里,界面上什么都看不见。
而「这个号在涨粉、只是最近没发作品」恰恰是最该看见的一种情况 —— 藏起来正好藏反了。
所以来源换成 **账号快照 ∪ 作品**:
* 有快照没作品 → 一个 0 篇的组,粉丝数照常显示;
* 有作品没快照 → 一个没有账号指标的组(小红书那条路不产生快照,就是这种情况)。
``creator_alias`` 从作品备注那张表来;``note_count`` / ``last_activity_at`` 用来
排序,让最近还在动的博主排在前面。
"""
scope: Optional[List[int]] = None
if task_id is not None:
scope = [task_id]
elif platform is not None:
scope = await platform_task_ids(session, platform)
if not scope:
return []
# 作品一侧:谁有作品、有几篇、最后一次是什么时候。
work_query = (
select(
MonitorNote.task_id,
MonitorNote.creator_hash,
func.count().label("note_count"),
func.max(MonitorNote.last_seen_at).label("last_seen_at"),
func.max(MonitorNote.creator_name).label("creator_name"),
)
.where(MonitorNote.creator_hash != "")
.group_by(MonitorNote.task_id, MonitorNote.creator_hash)
)
if scope is not None:
work_query = work_query.where(MonitorNote.task_id.in_(scope))
work: Dict[tuple, Dict[str, Any]] = {}
for row in (await session.execute(work_query)).all():
work[(row.task_id, row.creator_hash)] = {
"note_count": row.note_count,
"last_seen_at": row.last_seen_at,
"creator_name": row.creator_name or "",
}
stats = await _latest_creator_stats(session, task_ids=scope)
if not work and not stats:
return []
# 备注是按 (platform, creator_hash) 存的,所以要知道每个博主属于哪个平台。
involved = {key[0] for key in set(work) | set(stats)}
task_platform = {
row.id: row.platform
for row in (
await session.execute(
select(MonitorTask.id, MonitorTask.platform).where(
MonitorTask.id.in_(list(involved))
)
)
).all()
}
aliases = await _creator_alias_map(session)
result: List[Dict[str, Any]] = []
for key in set(work) | set(stats):
row_task, creator_hash = key
work_row = work.get(key)
stat = stats.get(key)
result.append(
{
"task_id": row_task,
"creator_hash": creator_hash,
# 昵称优先取作品的(那是界面上本来就在用的),快照的兜底 —— 没有作品
# 的博主只剩快照这一个来源。
"creator_name": (work_row or {}).get("creator_name")
or (stat.nickname if stat else ""),
"creator_alias": aliases.get(
(task_platform.get(row_task, ""), creator_hash), ""
),
"note_count": (work_row or {}).get("note_count", 0),
"creator_fans": stat.fans if stat else None,
"creator_total_favorited": stat.total_favorited if stat else None,
"creator_works": stat.works_count if stat else None,
"creator_stats_at": stat.captured_at if stat else None,
# 排序用:作品最近出现的时间,或者账号指标的采集时间,取晚的那个。
"last_activity_at": max(
(work_row or {}).get("last_seen_at") or 0,
stat.captured_at if stat else 0,
),
}
)
result.sort(key=lambda row: row["last_activity_at"], reverse=True)
return result
async def note_series(session: AsyncSession, note_id: str, task_id: Optional[int] = None) -> List[Dict[str, Any]]:
"""Metric time series for one note."""
query = (
@@ -409,9 +747,13 @@ async def _note_meta_map(
return {
row.note_id: {
"note_title": row.title,
"note_cover": row.cover,
"note_cover": covers.cover_url(row.note_id, row.cover),
"note_url": row.note_url,
"published_at": row.published_at,
"task_id": row.task_id,
# 博主维度也带上,评论流才能按 博主 -> 作品 -> 评论 三级展开。
"creator_hash": row.creator_hash,
"creator_name": row.creator_name,
}
for row in rows
}
@@ -457,6 +799,9 @@ async def list_comments(
"note_title": meta.get(row.note_id, {}).get("note_title", ""),
"note_cover": meta.get(row.note_id, {}).get("note_cover", ""),
"note_url": meta.get(row.note_id, {}).get("note_url", ""),
"note_creator_hash": meta.get(row.note_id, {}).get("creator_hash", ""),
"note_creator_name": meta.get(row.note_id, {}).get("creator_name", ""),
"note_published_at": meta.get(row.note_id, {}).get("published_at"),
}
for row in comments
]
@@ -507,6 +852,8 @@ async def comment_note_groups(
"note_title": meta.get(note_id, {}).get("note_title", ""),
"note_cover": meta.get(note_id, {}).get("note_cover", ""),
"note_url": meta.get(note_id, {}).get("note_url", ""),
"creator_hash": meta.get(note_id, {}).get("creator_hash", ""),
"creator_name": meta.get(note_id, {}).get("creator_name", ""),
"comment_count": count,
"latest_at": latest.get(note_id, 0),
}
@@ -623,11 +970,13 @@ async def list_tasks(
"mode": task.mode,
"enabled": task.enabled,
"interval_minutes": task.interval_minutes,
**_schedule_fields(task),
"max_notes_count": task.max_notes_count,
"enable_comments": task.enable_comments,
"max_comments_count": task.max_comments_count,
"run_timeout_seconds": task.run_timeout_seconds,
"notify_enabled": task.notify_enabled,
"notify_failures": task.notify_failures,
"next_run_at": task.next_run_at,
"last_run_at": task.last_run_at,
"last_status": task.last_status,
@@ -644,6 +993,29 @@ async def list_tasks(
]
def _schedule_fields(task: MonitorTask) -> Dict[str, Any]:
"""The schedule columns, plus the sentence the task list renders.
The label is composed here rather than in the frontend so the list and the
editor cannot drift on what a given schedule means.
"""
hours = schedule.parse_hours(task.schedule_hours)
days = schedule.parse_days(task.schedule_days)
return {
"schedule_mode": task.schedule_mode,
"schedule_hours": hours,
"schedule_days": days,
"schedule_minute": task.schedule_minute,
"schedule_label": schedule.describe(
mode=task.schedule_mode,
interval_minutes=task.interval_minutes,
hours=hours,
days=days,
minute=task.schedule_minute,
),
}
async def overview(session: AsyncSession, platform: Optional[str] = None) -> Dict[str, Any]:
"""Headline numbers for the dashboard tiles, scoped to one platform."""
now = get_current_timestamp()
+355
View File
@@ -0,0 +1,355 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/upstream.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""上游仓库更新检查。
本仓库在上游(NanmiCoder/MediaCrawler)之上加了一整层监控/鉴权/多平台面板,
合并流程写在 UPSTREAM.md 里。但那份流程默认**有人知道上游动了**——而部署脚本是
`git pull --ff-only`,只从我们自己的 Gitea 拉,上游的提交不主动去 fetch 就永远
看不见。拖着不合并的代价是复利的:越久越难合,最后只能放弃。这个模块把「上游动
了没有」变成一条可定时、会推到企业微信的通知。
三处刻意的取舍:
* **用 git 而不是托管商的 HTTP API。** 只有 git 算得出「落后几个提交」:托管商
API 能告诉你上游 tip 是什么,但它不知道我们与上游的共同祖先在哪,而分歧点恰恰
是真正要合的东西。本仓库还含有上游没有的提交,直接比 tip 会得出错误的结论。
* **按 URL fetch 到 FETCH_HEAD,不配置 remote、不写 refs/remotes。** 服务器上的
checkout 是从 Gitea 克隆的,本来就没有 upstream 这个 remote;用 URL 直取就不必
先去改它的 git 配置。顺带也避免往别人的部署里塞一个 remote。
* **只读不写工作区。** fetch 只落对象和 FETCH_HEAD,不碰索引与工作区,所以不会打断
正在跑的采集,也不会和 `./deploy.sh` 的 git pull 抢锁。
依赖一个外部命令:**git**。本机开发环境一定有;容器里是 Dockerfile 显式装的
(python:3.11-slim 默认不带)。
"""
import asyncio
import json
import os
import subprocess
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Dict, List, Tuple
from sqlalchemy.ext.asyncio import AsyncSession
from tools.time_util import get_current_timestamp
from .db import get_session
from .models import SETTING_UPSTREAM_NOTIFIED_TIP, SETTING_UPSTREAM_STATE
from .settings import get_setting, set_setting
PROJECT_ROOT = Path(__file__).parent.parent.parent
# 默认就是本仓库跟踪的那个上游。国内直连 GitHub 不稳时改成 gitcode 镜像即可
# (见 UPSTREAM.md「直连 GitHub 不通时」)。
DEFAULT_REMOTE_URL = "https://github.com/NanmiCoder/MediaCrawler.git"
DEFAULT_BRANCH = "main"
# fetch 要走网络,给宽松些;其余全是本地命令,慢到这个程度只能说明仓库坏了。
FETCH_TIMEOUT_SECONDS = 120
LOCAL_TIMEOUT_SECONDS = 20
# 通知里最多列几条提交。要传达的是「该动手了」,不是把 changelog 搬到群里。
MAX_LISTED_COMMITS = 10
# 状态里留几条给前端展示。比通知多留一些,界面上能看到更完整的列表。
MAX_STORED_COMMITS = 30
# git log 用 Unit Separator 分隔字段:它不可能出现在提交信息里,比制表符安全。
_RECORD_SEPARATOR = "\x1f"
_LOG_FORMAT = (
f"%h{_RECORD_SEPARATOR}%an{_RECORD_SEPARATOR}%ad{_RECORD_SEPARATOR}%s"
)
class GitError(RuntimeError):
"""git 不可用,或某条 git 命令失败了。"""
@dataclass(frozen=True)
class Commit:
"""一条上游提交,只留通知/展示需要的四个字段。"""
sha: str
author: str
date: str
subject: str
@dataclass
class CheckResult:
"""一次检查的结论。失败也是一种结论,用 ``ok``/``error`` 表达而不是抛异常。"""
ok: bool
# HEAD..FETCH_HEAD:上游有而我们没有的提交数 —— 要合的就是这些。
behind: int = 0
# FETCH_HEAD..HEAD:我们有自己的提交数 —— 也就是这一层的规模。
ahead: int = 0
tip: str = ""
head: str = ""
commits: List[Commit] = field(default_factory=list)
error: str = ""
def as_dict(self) -> Dict[str, Any]:
return {
"ok": self.ok,
"behind": self.behind,
"ahead": self.ahead,
"tip": self.tip,
"head": self.head,
"commits": [commit.__dict__ for commit in self.commits],
"error": self.error,
}
def _git(args: List[str], timeout: int) -> subprocess.CompletedProcess:
env = dict(os.environ)
# 远端要凭据时(地址写成了私有仓库),git 会停下来问密码,而这里没有终端可问,
# 于是挂到超时。关掉一切交互,让它立刻失败。
env["GIT_TERMINAL_PROMPT"] = "0"
env["GIT_ASKPASS"] = ""
env["SSH_ASKPASS"] = ""
# 容器里 uid 1000 没有 passwd 项,git 找不到 HOME 会抱怨。给一个存在且可写的。
env.setdefault("HOME", "/tmp")
return subprocess.run(
[
"git",
"-C",
str(PROJECT_ROOT),
# 只对自己这个 checkout 放行所有权检查。容器里 uid 一般与属主一致,
# 但 bind mount 的属主未必,一旦不一致 git 会直接拒绝干任何活。
"-c",
f"safe.directory={PROJECT_ROOT}",
# 忽略任何全局凭据助手:这是个只读的公开仓库,不该去翻钥匙串。
"-c",
"credential.helper=",
*args,
],
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
timeout=timeout,
env=env,
)
def _run(args: List[str], timeout: int) -> Tuple[int, str, str]:
"""跑一条 git 命令,返回 (returncode, stdout, stderr)。
只把「跑不起来」当异常;命令返回非零是正常结果,交给调用方处理。
"""
try:
proc = _git(args, timeout)
except FileNotFoundError as exc:
raise GitError("未找到 git 命令,请先安装 git") from exc
except subprocess.TimeoutExpired as exc:
raise GitError(f"git {args[0]} 超时({timeout} 秒)") from exc
return proc.returncode, (proc.stdout or "").strip(), (proc.stderr or "").strip()
def _require(args: List[str], timeout: int, what: str) -> str:
code, out, err = _run(args, timeout)
if code != 0:
# git 的报错通常是多行的,只留第一行;完整输出塞进日志反而更难读。
detail = err.splitlines()[0].strip() if err else "未知错误"
raise GitError(f"{what}:{detail}")
return out
def _to_int(raw: str) -> int:
try:
return int(raw)
except (TypeError, ValueError):
return 0
def _parse_log(raw: str) -> List[Commit]:
commits: List[Commit] = []
for line in raw.splitlines():
parts = line.split(_RECORD_SEPARATOR)
if len(parts) != 4:
# 格式不对就跳过这一条:一条读不出来的提交不该让整次检查失败。
continue
sha, author, date, subject = parts
commits.append(Commit(sha=sha, author=author, date=date, subject=subject))
return commits
def _check_sync(remote_url: str, branch: str) -> CheckResult:
"""阻塞实现,异步包装见 :func:`check`。"""
# 在 try 之前绑定:后面的失败结果也带上它 —— 「检查失败」时当前跑的是哪个
# 提交,正是排查时第一个想知道的。
head = ""
try:
head = _require(["rev-parse", "HEAD"], LOCAL_TIMEOUT_SECONDS, "读取本地 HEAD 失败")
# 增量 fetch:对象本地基本都已经有了,所以正常情况下只传几个新提交,
# 不会遇到 UPSTREAM.md 里说的「大包必断」。
_require(
["fetch", "--no-tags", remote_url, branch],
FETCH_TIMEOUT_SECONDS,
"从上游 fetch 失败",
)
tip = _require(["rev-parse", "FETCH_HEAD"], LOCAL_TIMEOUT_SECONDS, "读不到 FETCH_HEAD")
behind = _to_int(
_require(
["rev-list", "--count", "HEAD..FETCH_HEAD"],
LOCAL_TIMEOUT_SECONDS,
"统计落后提交数失败",
)
)
ahead = _to_int(
_require(
["rev-list", "--count", "FETCH_HEAD..HEAD"],
LOCAL_TIMEOUT_SECONDS,
"统计领先提交数失败",
)
)
# 只在确实落后时才读提交列表:已经是最新时这条 git log 毫无意义。
raw_log = ""
if behind:
raw_log = _require(
[
"log",
f"--max-count={MAX_STORED_COMMITS}",
"--date=short",
f"--format={_LOG_FORMAT}",
"HEAD..FETCH_HEAD",
],
LOCAL_TIMEOUT_SECONDS,
"读取新提交列表失败",
)
except GitError as exc:
return CheckResult(ok=False, head=head, error=str(exc))
return CheckResult(
ok=True,
behind=behind,
ahead=ahead,
tip=tip,
head=head,
commits=_parse_log(raw_log),
)
async def check(
remote_url: str = DEFAULT_REMOTE_URL, branch: str = DEFAULT_BRANCH
) -> CheckResult:
"""跑一次检查。网络与子进程都丢进线程,事件循环不被阻塞。
不抛异常:上游不通是常态(尤其是直连 GitHub),那也是一种要记录下来的结果。
"""
return await asyncio.to_thread(_check_sync, remote_url, branch)
def build_message(result: CheckResult, branch: str) -> str:
"""把一次「上游有新提交」的结果写成一条企业微信 markdown。"""
lines = [
"**🔔 上游 MediaCrawler 有更新**",
f"> 当前部署落后 `{branch}` **{result.behind}** 个提交",
]
if result.ahead:
lines.append(f"> (本仓库另有 {result.ahead} 个自己的提交,合并时注意保留)")
for commit in result.commits[:MAX_LISTED_COMMITS]:
lines.append(f"> `{commit.sha}` {commit.subject}")
# 落后数可能大于列出来的条数:状态里留的提交本身也是截断的(30 条),
# 所以这里比的是总数,不是 len(commits)。
if result.behind > MAX_LISTED_COMMITS:
lines.append(f"> …等共 {result.behind} 个提交")
lines.append("> 合并步骤见仓库根目录 `UPSTREAM.md`")
return "\n".join(lines)
async def load_state(session: AsyncSession) -> Dict[str, Any]:
"""最近一次检查的结果,从设置里读回来。没查过时是空字典。"""
raw = await get_setting(session, SETTING_UPSTREAM_STATE)
if not raw:
return {}
try:
state = json.loads(raw)
except json.JSONDecodeError:
# 手改坏了的行不该让接口 500,当作「没查过」即可。
return {}
return state if isinstance(state, dict) else {}
async def _save_state(session: AsyncSession, state: Dict[str, Any]) -> None:
# 整体读写,所以存成一条 JSON:拆成多个 key 只会带来写到一半的不一致。
await set_setting(session, SETTING_UPSTREAM_STATE, json.dumps(state, ensure_ascii=False))
async def run_check(notify_when_new: bool = True) -> Dict[str, Any]:
"""检查一次,落库,必要时推送。返回值可直接交给前端。
分三段各自的数据库会话:fetch 最长可能跑满两分钟,占着一个连接不合适 ——
理由与 runner.py 的分段完全相同。
"""
# 延迟导入:app_settings 在模块级 import 本模块(为了那个默认地址常量),
# 模块级反向 import 会成环。
from . import app_settings, notify
async with get_session() as session:
remote_url = str(
await app_settings.get_value(
session, "upstream_remote_url", fallback=DEFAULT_REMOTE_URL
)
or DEFAULT_REMOTE_URL
)
branch = str(
await app_settings.get_value(session, "upstream_branch", fallback=DEFAULT_BRANCH)
or DEFAULT_BRANCH
)
notify_enabled = bool(
await app_settings.get_value(session, "upstream_notify", fallback=True)
)
notified_tip = (await get_setting(session, SETTING_UPSTREAM_NOTIFIED_TIP)) or ""
webhook_url = await notify.get_webhook_url(session)
result = await check(remote_url, branch)
payload: Dict[str, Any] = {
"checked_at": get_current_timestamp(),
"remote_url": remote_url,
"branch": branch,
**result.as_dict(),
}
async with get_session() as session:
await _save_state(session, payload)
has_update = result.ok and result.behind > 0 and bool(result.tip)
# 同一个 tip 只推一次:否则每过一个检查周期就把同样的更新推到群里,
# 直到有人去合为止。上游真又动了(tip 变了)时应该再推。
if (
notify_when_new
and notify_enabled
and has_update
and webhook_url
and result.tip != notified_tip
):
ok, detail = await notify.send_wecom(webhook_url, build_message(result, branch))
if ok:
await set_setting(session, SETTING_UPSTREAM_NOTIFIED_TIP, result.tip)
payload["notified"] = True
else:
# 推送失败不该抹掉检查结果 —— 界面上仍然能看到「落后几个提交」。
payload["notify_error"] = detail
return payload
+2
View File
@@ -18,6 +18,7 @@
from .auth import router as auth_router
from .crawler import router as crawler_router
from .creator import router as creator_router
from .data import router as data_router
from .monitor import router as monitor_router
from .settings import router as settings_router
@@ -26,6 +27,7 @@ from .websocket import router as websocket_router
__all__ = [
"auth_router",
"crawler_router",
"creator_router",
"data_router",
"monitor_router",
"settings_router",
+150
View File
@@ -0,0 +1,150 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/creator.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""运营模块的 HTTP 接口。"""
import asyncio
from typing import Set
from fastapi import APIRouter, HTTPException, Query
from ..creator import login as creator_login
from ..creator import service
from ..monitor.db import get_session
router = APIRouter(prefix="/creator", tags=["creator"])
# 后台同步任务要留强引用:asyncio 只持弱引用,否则任务可能在跑完前被回收。
_sync_tasks: Set[asyncio.Task] = set()
def _bad_request(exc: ValueError) -> HTTPException:
return HTTPException(status_code=400, detail=str(exc))
@router.get("/accounts")
async def list_accounts():
"""账号列表。**不含 cookie**,只给 ``has_cookie``。"""
async with get_session() as session:
return {"accounts": await service.list_accounts(session)}
@router.get("/accounts/{account_id}")
async def get_account_detail(account_id: int):
async with get_session() as session:
try:
return await service.account_detail(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
@router.delete("/accounts/{account_id}")
async def delete_account(account_id: int):
async with get_session() as session:
try:
await service.delete_account(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
return {"message": "账号已删除"}
@router.post("/accounts/{account_id}/check")
async def check_account(account_id: int):
"""重测登录态与数据权限。"""
async with get_session() as session:
try:
return await service.check_account(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
@router.post("/accounts/{account_id}/sync")
async def sync_account(
account_id: int, days: int = Query(default=90, ge=1, le=730)
):
"""拉取作品数据。
**放后台跑**:要分页、还要按账号节流,几分钟很正常,而前端请求超时是 30 秒。
前端靠轮询账号列表里的 `last_synced_at` / `last_error` 看结果。
"""
async with get_session() as session:
try:
await service.get_account(session, account_id)
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc))
task = asyncio.create_task(_sync_in_background(account_id, days))
_sync_tasks.add(task)
task.add_done_callback(_sync_tasks.discard)
return {"message": "同步已开始", "days": days}
async def _sync_in_background(account_id: int, days: int) -> None:
try:
async with get_session() as session:
result = await service.sync_account(session, account_id, days)
print(f"[creator] 账号 {account_id} 同步完成,取回 {result['fetched']} 条")
except Exception as exc: # noqa: BLE001 - 后台任务不能让异常逃逸成静默失败
print(f"[creator] 账号 {account_id} 同步失败: {exc}")
# ---------------------------------------------------------------------------
# 扫码新增账号
# ---------------------------------------------------------------------------
#
# 每次登录开一个**临时浏览器上下文**,扫完取出 cookie 就丢弃 —— 这样登第二个账号
# 不会把第一个顶掉,也不影响监控那个登录态。cookie 只在内存里从 login 模块传到
# 这里落库,**不进响应体**。
@router.post("/login")
async def start_login():
try:
return await creator_login.start()
except RuntimeError as exc:
raise HTTPException(status_code=502, detail=str(exc))
@router.get("/login")
async def poll_login():
"""轮询扫码结果;一旦成功就把账号落库并返回它。"""
snapshot = await creator_login.status()
if snapshot["status"] == creator_login.STATUS_SUCCESS:
# take_cookie 只在会话还在时返回 cookie,取走即拆会话;重复轮询拿到 None
# 就说明已经保存过了,直接返回上次的结果,不要退回 idle。
cookie = await creator_login.take_cookie()
if cookie:
try:
async with get_session() as session:
account = await service.upsert_account_from_cookie(session, cookie)
except ValueError as exc:
snapshot["status"] = creator_login.STATUS_ERROR
snapshot["message"] = f"扫码成功但保存账号失败:{exc}"
await creator_login.remember_result(snapshot)
return snapshot
snapshot["account"] = account
snapshot["message"] = f"已添加账号:{account['nickname']}"
await creator_login.remember_result(snapshot)
return snapshot
@router.delete("/login")
async def cancel_login():
return await creator_login.cancel()
+233 -12
View File
@@ -18,12 +18,13 @@
"""HTTP API for scheduled monitoring tasks."""
from datetime import date, timedelta
from datetime import date, datetime, timedelta
from typing import Any, Dict, List, Optional
from fastapi import APIRouter, HTTPException, Query, Response
from fastapi.responses import FileResponse
from ..monitor import notify, report, service
from ..monitor import covers, notify, qrlogin, report, service, upstream
from ..monitor.db import get_session
from ..monitor.platforms import PLATFORM_XHS
from ..monitor.settings import (
@@ -37,6 +38,8 @@ from ..monitor.settings import (
from ..monitor.models import SETTING_WECOM_WEBHOOK, MonitorTask
from ..schemas.monitor import (
CookiePayload,
CreatorAliasPayload,
NoteAliasPayload,
MonitorTaskCreate,
MonitorTaskUpdate,
WebhookPayload,
@@ -130,7 +133,10 @@ async def list_notes(
):
async with get_session() as session:
return {
"notes": await service.list_notes(session, task_id, only_new, limit, platform)
"notes": await service.list_notes(session, task_id, only_new, limit, platform),
# 博主**单独给一份**,而不是让前端从作品里推。作品推不出「一条作品都没有的
# 博主」—— 那正是最该显示的一类(还在涨粉,只是最近没发)。
"creators": await service.list_creators(session, task_id, platform),
}
@@ -171,6 +177,15 @@ async def list_comments(
"note_title": comment["note_title"],
"note_cover": comment["note_cover"],
"note_url": comment["note_url"],
# 作品所属的创作者。评论流按 博主 → 作品 → 评论 三级展开时,最外层
# 就是按这两个字段分组的 —— 少了它们,前端拿到的是 undefined,
# 于是所有博主塌成同一个分组、标签回退成「未知博主」。
# 同一个桶里的评论必然同属一个作品,所以取哪一条都一样。
"creator_hash": comment["note_creator_hash"],
"creator_name": comment["note_creator_name"],
# 作品的发布时间。评论流按 博主 → 作品 → 评论 展开时,作品那一层
# 光有标题不够 —— 同名作品不少,日期能帮着认。
"published_at": comment["note_published_at"],
"comments": [],
},
)
@@ -243,6 +258,138 @@ async def clear_cookie_endpoint(platform: str = Query(default=PLATFORM_XHS)):
return {"message": "Cookie cleared"}
# ---------------------------------------------------------------------------
# 博主备注
# ---------------------------------------------------------------------------
@router.put("/creators/{creator_hash}")
async def set_creator_alias_endpoint(
creator_hash: str,
payload: CreatorAliasPayload,
platform: str = Query(default=PLATFORM_XHS),
):
"""给博主起个备注(界面上的「备注」)。
作品栏按 creator_hash 归组,可那是个哈希、昵称又常常认不出是谁 —— 备注是人自己起的
名字。空串表示清掉这条备注。
"""
async with get_session() as session:
await service.set_creator_alias(session, platform, creator_hash, payload.alias)
return {"creator_hash": creator_hash, "alias": payload.alias.strip()}
@router.put("/notes/{note_id}")
async def set_note_alias_endpoint(
note_id: str,
payload: NoteAliasPayload,
platform: str = Query(default=PLATFORM_XHS),
):
"""给**作品**起个备注。
和上面那条博主备注是一对:博主备注回答「这个账号是谁」,这条回答「这条作品我要盯着」。
两者不能合并 —— 一个博主底下常常只有一两件值得盯的作品。
"""
async with get_session() as session:
await service.set_note_alias(session, platform, note_id, payload.alias)
return {"note_id": note_id, "alias": payload.alias.strip()}
# ---------------------------------------------------------------------------
# QR login
# ---------------------------------------------------------------------------
# These drive the browser already listening on the CDP debug port, which is the
# same browser -- and therefore the same profile -- that monitor runs attach to.
# Scanning once is what makes later unattended runs logged in.
@router.get("/covers/{note_id}")
async def get_cover(note_id: str):
"""作品封面,从本地缓存读。
**为什么不让前端直连图床**:图床地址是带签名、会过期的 —— 实测隔天即 403,
而且带不带 Referer 都一样,所以那是过期而不是防盗链。本地那份与签名无关。
这个路由是带鉴权的(整条 monitor 路由都挂了 require_auth),所以封面不会被
匿名读走;前端用同源的 <img> 请求会自动带上会话 cookie。
"""
path = covers.find_cached(note_id)
if path is None:
raise HTTPException(status_code=404, detail="封面未缓存")
media_types = {
".jpg": "image/jpeg",
".png": "image/png",
".webp": "image/webp",
".gif": "image/gif",
".heic": "image/heic",
}
return FileResponse(
path,
media_type=media_types.get(path.suffix.lower(), "application/octet-stream"),
# 本地文件不会变(note_id 唯一),让浏览器自己缓存,省掉重复请求。
headers={"Cache-Control": "private, max-age=86400"},
)
@router.post("/login/qr")
async def start_qr_login(platform: str = Query(default=PLATFORM_XHS)):
"""Open the login page in the CDP browser and return its QR code.
A server deployment has no display (Chrome sits under Xvfb), so the code is
surfaced here for the operator to scan instead of in a desktop window that
does not exist.
"""
try:
return await qrlogin.start(platform)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc))
except RuntimeError as exc:
raise HTTPException(status_code=502, detail=str(exc))
@router.get("/login/qr")
async def get_qr_login():
"""Poll the live session: waiting -> success / expired / error.
扫码成功时**把 cookie 一并存进库**。扫码本来只写浏览器 profile,那只够 CDP 模式用;
存一份之后,CDP 关掉、任务改用 --cookies_file 注入也照样能跑 —— 两种机制同时填上,
开关怎么切都不会断。
"""
snapshot = await qrlogin.status()
if snapshot["status"] == qrlogin.STATUS_SUCCESS:
cookie = await qrlogin.take_cookie()
if cookie:
async with get_session() as session:
await set_cookie(session, cookie)
snapshot["cookie_saved"] = True
snapshot["message"] = f"{snapshot['message']};登录态已同时存入 Cookie"
return snapshot
@router.delete("/login/qr")
async def cancel_qr_login():
"""Drop our tab and stop polling."""
return await qrlogin.cancel()
@router.get("/login/state")
async def get_login_state(force: bool = Query(default=False)):
"""Ask the browser itself whether it is signed in.
Deliberately separate from the QR session above. That session is in-memory and
dies with the process -- a redeploy is enough -- so "am I logged in?" must not
hinge on it, or a successful scan looks like nothing happened.
``force`` reloads the page first, for when the login may have lapsed somewhere
else and the page's copy of the state is stale.
"""
return await qrlogin.check_login_state(force=force)
# ---------------------------------------------------------------------------
# Report
# ---------------------------------------------------------------------------
@@ -339,20 +486,29 @@ def _export_columns(kind: str) -> List[tuple[str, str]]:
"""(key, header) pairs per export kind."""
if kind == "notes":
return [
("note_id", "作品ID"),
# 先放「这是谁」:导出来是拿去比对和汇报的,一行只有作品 ID 没法用。
# 备注优先 —— 昵称常常认不出是谁(见 notes 表那一层的说明)。
("creator_alias", "博主备注"),
("creator_name", "博主昵称"),
("note_alias", "作品备注"),
("title", "标题"),
("note_id", "作品ID"),
("note_url", "链接"),
("liked_count", "点赞"),
("comment_count", "评论"),
("collected_count", "收藏"),
("share_count", "分享"),
("liked_count_delta", "点赞增量"),
("comment_count_delta", "评论增量"),
("published_at", "发布时间"),
# 指标嵌在 row["metrics"] 里,所以这里必须写成路径 —— 写成裸键名的话这几列
# 全空(见 _lookup)。
("metrics.liked_count", "点赞"),
("metrics.comment_count", "评论"),
("metrics.collected_count", "收藏"),
("metrics.share_count", "分享"),
("deltas.liked_count", "点赞增量"),
("deltas.comment_count", "评论增量"),
("first_seen_at", "首次发现"),
("last_seen_at", "最近采集"),
]
if kind == "comments":
return [
("note_creator_name", "博主昵称"),
("note_title", "所属作品"),
("note_id", "作品ID"),
("comment_id", "评论ID"),
@@ -374,6 +530,42 @@ def _export_columns(kind: str) -> List[tuple[str, str]]:
]
# 表里存的是毫秒时间戳。直接倒进 CSV 就是一串 13 位数字 —— 打开 Excel 的人没法看,
# 也没法排序。这几个键统一格式化成人能读的形态。
_TIME_KEYS = {"published_at", "first_seen_at", "last_seen_at", "create_time"}
def _fmt_time(value: Any) -> str:
"""毫秒 → ``YYYY-MM-DD HH:MM``(服务器本地时区)。"""
try:
return datetime.fromtimestamp(int(value) / 1000).strftime("%Y-%m-%d %H:%M")
except (TypeError, ValueError, OSError, OverflowError):
return ""
def _cell_for(row: Dict[str, Any], key: str) -> Any:
"""一列的值:时间键格式化成人能读的,其余照原样(None 变空串)。"""
value = _lookup(row, key)
if key.rsplit(".", 1)[-1] in _TIME_KEYS:
return _fmt_time(value)
return _cell(value)
def _lookup(row: Dict[str, Any], key: str) -> Any:
"""取一列的值。键可以是 ``metrics.liked_count`` 这种路径。
作品行的指标是**嵌在** ``metrics`` / ``deltas`` 里的,而 ``_export_columns`` 里写的
是 ``liked_count`` —— 照顶层键直接 ``row.get()`` 的话,点赞/评论/收藏/分享四列连带
两个增量列**永远是空的**,导出来的表看着有这几列,其实一格都没有。
"""
value: Any = row
for part in key.split("."):
if not isinstance(value, dict):
return None
value = value.get(part)
return value
def _cell(value: Any) -> Any:
if value is None:
return ""
@@ -390,7 +582,7 @@ def _to_csv(rows: List[Dict[str, Any]], columns: List[tuple[str, str]]) -> bytes
writer = csv.writer(buffer)
writer.writerow([header for _, header in columns])
for row in rows:
writer.writerow([_cell(row.get(key)) for key, _ in columns])
writer.writerow([_cell_for(row, key) for key, _ in columns])
# utf-8-sig: without the BOM Excel opens Chinese CSV as mojibake, which is
# the single most common complaint about CSV exports here.
@@ -407,7 +599,7 @@ def _to_xlsx(rows: List[Dict[str, Any]], columns: List[tuple[str, str]], sheet:
worksheet.title = {"notes": "作品", "comments": "评论"}.get(sheet, "报表")
worksheet.append([header for _, header in columns])
for row in rows:
worksheet.append([_cell(row.get(key)) for key, _ in columns])
worksheet.append([_cell_for(row, key) for key, _ in columns])
output = io.BytesIO()
workbook.save(output)
@@ -500,3 +692,32 @@ async def test_webhook(payload: WebhookTestPayload):
if not ok:
raise HTTPException(status_code=400, detail=detail)
return {"message": detail}
# ---------------------------------------------------------------------------
# 上游更新
# ---------------------------------------------------------------------------
@router.get("/upstream")
async def get_upstream_status():
"""最近一次上游检查的结果。
只读缓存,不触发检查:fetch 要走网络、最长两分钟,不该由一个 GET 顺手发起。
没有查过时返回空对象,前端据此显示「尚未检查」。
"""
async with get_session() as session:
return await upstream.load_state(session)
@router.post("/upstream/check")
async def run_upstream_check():
"""立刻检查一次上游仓库,并把结果写回缓存。
即使「定期检查」开关是关的也照查 —— 手动点这一次的意义正在于此。这里会一直
等到 fetch 结束(前端给这条请求单独放长了超时),因为结果就是要给人看的。
``notify_when_new=False``:点这个按钮的人正看着结果,没必要再给自己推一条群消息。
没有推过的那批提交会留给下一次「定时检查」推 —— 推送状态记的是 tip,不是「推过没」。
"""
return await upstream.run_check(notify_when_new=False)
+50 -3
View File
@@ -20,7 +20,7 @@
from typing import List, Literal, Optional
from pydantic import BaseModel, Field
from pydantic import BaseModel, Field, model_validator
# A floor on the interval is a correctness guard, not a nicety: every run
# launches a browser and hits XHS with several requests, so a short interval
@@ -40,6 +40,19 @@ class MonitorTaskCreate(BaseModel):
interval_minutes: Optional[int] = Field(
default=None, ge=MIN_INTERVAL_MINUTES, le=MAX_INTERVAL_MINUTES
)
# --- Scheduling ---------------------------------------------------------
# All three modes are expressible with pickers; a raw cron string is
# deliberately not supported, since it is a small language to learn just to
# say "every day at nine".
schedule_mode: Literal["interval", "daily", "weekly"] = "interval"
# 0-23, e.g. [9, 12, 18]. Required for the two clock modes.
schedule_hours: List[int] = Field(default_factory=list)
# 0-6 with Monday = 0, matching Python's date.weekday(). Required for weekly.
schedule_days: List[int] = Field(default_factory=list)
# One minute for the whole schedule, so a task with three times is
# "09:30, 12:30, 18:30" rather than three separate minute choices.
schedule_minute: int = Field(default=0, ge=0, le=59)
max_notes_count: Optional[int] = Field(default=None, ge=1, le=500)
enable_comments: bool = True
# Raising this widens the comment window, which is the only lever available
@@ -47,12 +60,27 @@ class MonitorTaskCreate(BaseModel):
max_comments_count: Optional[int] = Field(default=None, ge=1, le=500)
run_timeout_seconds: int = Field(default=3600, ge=60, le=86400)
enabled: bool = True
# Push a WeCom summary for runs that failed or found new works. Opt-in per
# task so a single webhook does not get flooded.
# 两类通知分开:新作品可能每轮都有(默认关,避免刷屏),
# 异常频率低且意味着任务已经停止工作(默认开,否则你会一直不知道)。
notify_enabled: bool = False
notify_failures: bool = True
# Raw pasted values: full URLs or bare ids, in either form.
targets: List[str] = Field(min_length=1)
@model_validator(mode="after")
def _validate_schedule(self) -> "MonitorTaskCreate":
if any(hour < 0 or hour > 23 for hour in self.schedule_hours):
raise ValueError("小时必须在 0-23 之间")
if any(day < 0 or day > 6 for day in self.schedule_days):
raise ValueError("星期必须在 0-6 之间(周一为 0)")
# A clock mode with no chosen time can never fire. Rejecting it here is
# what keeps next_occurrence()'s None branch unreachable in practice.
if self.schedule_mode in ("daily", "weekly") and not self.schedule_hours:
raise ValueError("按钟点调度至少要选一个时间")
if self.schedule_mode == "weekly" and not self.schedule_days:
raise ValueError("按周调度至少要选一个星期")
return self
class MonitorTaskUpdate(BaseModel):
name: Optional[str] = Field(default=None, min_length=1, max_length=200)
@@ -60,11 +88,18 @@ class MonitorTaskUpdate(BaseModel):
interval_minutes: Optional[int] = Field(
default=None, ge=MIN_INTERVAL_MINUTES, le=MAX_INTERVAL_MINUTES
)
# None means "leave alone". Cross-field validity depends on the merged state,
# so it is checked in the service rather than here.
schedule_mode: Optional[Literal["interval", "daily", "weekly"]] = None
schedule_hours: Optional[List[int]] = None
schedule_days: Optional[List[int]] = None
schedule_minute: Optional[int] = Field(default=None, ge=0, le=59)
max_notes_count: Optional[int] = Field(default=None, ge=1, le=500)
enable_comments: Optional[bool] = None
max_comments_count: Optional[int] = Field(default=None, ge=1, le=500)
run_timeout_seconds: Optional[int] = Field(default=None, ge=60, le=86400)
notify_enabled: Optional[bool] = None
notify_failures: Optional[bool] = None
# When present, replaces the whole target list.
targets: Optional[List[str]] = None
@@ -73,6 +108,18 @@ class CookiePayload(BaseModel):
cookie: str = Field(min_length=1)
class CreatorAliasPayload(BaseModel):
"""给博主起的备注。空串表示清掉这条备注。"""
alias: str = Field(default="", max_length=128)
class NoteAliasPayload(BaseModel):
"""给作品起的备注。空串表示清掉这条备注。"""
alias: str = Field(default="", max_length=128)
class WebhookPayload(BaseModel):
url: str = Field(default="", description="企业微信机器人 Webhook 地址,留空表示停用")
+22 -1
View File
@@ -20,13 +20,20 @@ import asyncio
import subprocess
import signal
import os
from typing import Optional, List
from collections import deque
from typing import Deque, Optional, List
from datetime import datetime
from pathlib import Path
from ..schemas import CrawlerStartRequest, LogEntry
from .interpreter import resolve_python_cmd
# 留住多少行爬虫输出,供 run_and_wait 的调用方诊断失败原因。
# 子进程的输出本来只流向日志 WebSocket,监控层只看得到退出码 —— 于是「退出码 1」
# 成了运行历史里唯一的信息,真正的报错(比如抖音的 `DataFetchError: account blocked`)
# 谁也看不到。留个尾巴,让失败原因能被写进 run.error_message。
OUTPUT_TAIL_LINES = 80
class CrawlerManager:
"""Crawler process manager"""
@@ -49,6 +56,15 @@ class CrawlerManager:
# by any concurrent start(), so waiters need an explicit event instead.
self._done: asyncio.Event = asyncio.Event()
self.last_exit_code: Optional[int] = None
# 本次运行输出的末尾若干行。见 OUTPUT_TAIL_LINES。
self._output_tail: Deque[str] = deque(maxlen=OUTPUT_TAIL_LINES)
def get_output_tail(self) -> List[str]:
"""最近一次运行的输出尾巴(最早的排前面)。
只在 run_and_wait() 返回之后读才有意义 —— 它等到读输出的任务收尾才唤醒。
"""
return list(self._output_tail)
@property
def logs(self) -> List[LogEntry]:
@@ -114,6 +130,9 @@ class CrawlerManager:
async def _push_log(self, entry: LogEntry):
"""Push log to queue"""
# 这里是所有输出的唯一出口(读循环、收尾、以及管理器自己的提示都走它),
# 所以尾巴挂在这儿最省事,也不会漏。
self._output_tail.append(entry.message)
if self._log_queue is not None:
try:
self._log_queue.put_nowait(entry)
@@ -149,6 +168,8 @@ class CrawlerManager:
# Reset completion signalling for this run
self._done.clear()
self.last_exit_code = None
# 尾巴只属于本次运行,否则上一轮的报错会混进这一轮的诊断里。
self._output_tail.clear()
# Clear pending queue (don't replace object to avoid WebSocket broadcast coroutine holding old queue reference)
if self._log_queue is None:
+8
View File
@@ -58,6 +58,14 @@ SAVE_LOGIN_STATE = True
# 否则冷 profile 下 API 签名失败,且表现为「退出码 0 但抓到 0 条」的静默失败。
INJECT_ALL_COOKIES = False
# 是否对昵称做中间脱敏(默认 False —— 本仓库**关掉了**)。
# 上游作为教学版默认开启,保留首尾各 1 字、中间打星号,避免据昵称骚扰到真人。
# 但那是**有损**的:「张三」和「张四」都会变成「张*」,「小明老师」和「小刚老师」
# 都会变成「小***师」—— 而本仓库的用途是监控一批公开的创作者账号,分清谁是谁正是
# 这一层要干的事,撞名就等于看不出来。所以这里关掉,把原昵称原样落库。
# 想改回上游行为,把这一行改成 True 即可(脱敏机制本身没删)。
MASK_NICKNAME = False
# ==================== CDP (Chrome DevTools Protocol) 配置 ====================
# 是否启用 CDP 模式 - 使用用户本地的 Chrome/Edge 浏览器进行爬取,具有更好的反检测能力
# 开启后,会自动检测并启动用户的 Chrome/Edge 浏览器,通过 CDP 协议进行控制
Executable
+71
View File
@@ -0,0 +1,71 @@
#!/usr/bin/env bash
#
# 更新这台机器上的部署。用法:
#
# ./deploy.sh
#
# 为什么需要脚本而不是一句 `git pull && docker compose up -d`:
# 前端产物 api/webui 是 gitignore 的(它由 vite 生成),git pull 带不过来。
# 所以代码更新之后必须在服务器上重建一次前端,否则页面还是旧的。
# 这一步在容器里做,好处是服务器不需要装 Node —— 只有 Docker。
#
# node_modules 和 npm 缓存都留在挂载目录内,重复构建不会重新下载。
set -euo pipefail
cd "$(dirname "$0")"
# 用镜像里的解释器跑,宿主机的 Python 版本无关。
IMAGE=mediacrawler:latest
before=$(git rev-parse HEAD)
git pull --ff-only
after=$(git rev-parse HEAD)
if [ "$before" = "$after" ]; then
echo "== 代码已是最新($after)"
else
echo "== 代码更新 $before -> $after"
git --no-pager log --oneline "$before..$after" | sed 's/^/ /'
fi
# 镜像层(依赖)改动只能靠重建,而这一步不是自动的:Dockerfile 或 requirements.txt 变了,
# 下面那句 `docker compose up -d --force-recreate` 用的是旧镜像,改动根本不会生效。
# 至少要说出来,否则现象是「代码明明更新了,功能却报缺依赖」。
if [ "$before" != "$after" ] && ! git diff --quiet "$before" "$after" -- Dockerfile requirements.txt; then
echo "!! Dockerfile / requirements.txt 有改动,需要重建镜像后重跑本脚本:"
echo " docker compose build"
fi
# 前端重建的两种情况:产物根本不存在(首次部署),或 webui/ 有改动。
if [ ! -f api/webui/index.html ]; then
need_build=1
reason="前端产物不存在"
elif [ "$before" != "$after" ] && ! git diff --quiet "$before" "$after" -- webui/; then
need_build=1
reason="webui/ 有改动"
else
need_build=0
reason=""
fi
if [ "$need_build" = "1" ]; then
echo "== 重建前端($reason)"
# -u 1000:1000 而不是 root:这里产出的文件要留在这个目录里给后面用,
# 以 root 生成的 node_modules 会让下次构建和人工清理都变得别扭。
# HOME 指向挂载目录,这样 npm 的缓存在宿主机上,重建时能复用。
docker run --rm \
-u 1000:1000 \
-w /app/webui \
-v "$PWD:/app" \
-e HOME=/app/webui \
-e npm_config_registry=https://registry.npmmirror.com \
"$IMAGE" sh -c 'npm ci --no-audit --no-fund && npm run build'
else
echo "== 前端无改动,跳过构建"
fi
# --force-recreate,而不是裸的 `up -d`:代码是 bind mount,容器配置和镜像都没变,
# 所以 `up -d` 会判定"无需变更"直接跳过,Python 代码的改动根本不会生效。前端产物是
# 磁盘上的静态文件,能即时生效,这一点很容易掩盖上面那个问题,直到有人改了 .py 才发现。
echo "== 重启容器"
docker compose up -d --force-recreate
docker compose ps
+33
View File
@@ -0,0 +1,33 @@
services:
mediacrawler:
build: .
image: mediacrawler:latest
container_name: mediacrawler
restart: unless-stopped
# Run as the user that owns this checkout. Without it the container is root,
# and every file it writes into the mounted tree -- the crawler's per-run
# jsonl output above all -- comes out root-owned. That does not break the app,
# but it does lock the operator out of moving or deleting their own
# deployment, which is exactly what happened the first time this was deployed.
user: "1000:1000"
# host networking is a requirement, not a convenience: the crawler attaches
# to the operator's Chrome at 127.0.0.1:9222, and inside a bridge network
# that loopback is the container's own, where no browser is listening.
# It also puts the app port directly on the host, so `ports:` is not used.
network_mode: host
env_file:
- .env
environment:
MC_HOST: 0.0.0.0
MC_PORT: "18051"
TZ: Asia/Shanghai
volumes:
# The code is mounted rather than baked in, so shipping a change is
# "git pull, restart" instead of an image rebuild. Only the dependencies
# live in the image, because those are the expensive part and they change
# rarely -- rebuild only when requirements.txt or the Dockerfile changes.
- ./:/app
+53 -2
View File
@@ -287,7 +287,7 @@ set MC_PASSWORD=我的新密码 # Windows cmd
未接通的平台**可以选,但各页会显示明确的说明面板**,并且**创建任务会被直接拒绝**:
```
400 抖音的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。
400 B站的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。
```
而不是接受任务、然后让它永远跑不出数据 —— 那正是之前"博主主页解析失败被误报成登录失效"的同一种静默故障。
@@ -297,7 +297,7 @@ set MC_PASSWORD=我的新密码 # Windows cmd
| 位置 | 范围 | 内容 |
|---|---|---|
| 左侧导航「设置」 | **按平台** | 登录 Cookie、采集策略、代理 |
| 右上角「系统设置」 | **全局** | 通知、活跃时段、账号安全 |
| 右上角「系统设置」 | **全局** | 通知、活跃时段、上游更新、账号安全 |
**这不是随便分的**:企业微信只有一个群、调度器只有一套时段规则、密码只有一份 ——
把它们放进"小红书专属"的页面里,会让人以为它们是按平台存的。
@@ -308,6 +308,54 @@ set MC_PASSWORD=我的新密码 # Windows cmd
---
## 二·十一、上游更新检查
本仓库在 [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) 之上加了一整层
(监控 / 鉴权 / 多平台面板),差异管理与合并流程在根目录 `UPSTREAM.md` 里。但那份流程有个
隐含前提:**得有人知道上游动了**。部署脚本只从我们自己的 Gitea `git pull`,上游的提交不主动
去 fetch 就永远看不见 —— 拖着不合并的代价是复利的,越久越难合。
这一项就是替你定时去 fetch 的:按间隔(默认每天一次)拉一次上游,算出「当前部署落后几个
提交」,有更新就推一条企业微信,并把结果与提交列表显示在**右上角「系统设置」→「上游更新」**。
### 配置
| 项 | 默认 | 说明 |
|---|---|---|
| 检查上游仓库更新 | **关** | 总开关。默认关:它要联网 fetch,且需要容器里有 git(见下) |
| 上游检查间隔(分钟) | 1440 | 每天一次。最小 30 分钟 |
| 上游仓库地址 | GitHub 上游 | 国内直连 GitHub 不稳时改成 gitcode 镜像,见 `UPSTREAM.md` |
| 上游分支 | `main` | |
| 上游有更新时推送通知 | 开 | 只在出现**此前没推过**的提交时发一条,同一个更新不会反复推 |
### 几个刻意的行为
- **只读,不写工作区**:只 `git fetch <地址> <分支>` 到 `FETCH_HEAD` —— 不建 remote、不写
`refs/remotes`、不碰索引与工作区。所以它不会打断正在跑的采集,也不会和 `./deploy.sh`
的 `git pull` 抢锁。
- **不受活跃时段限制**:活跃时段是给采集定的(避免半夜去抓平台)。检查只是 fetch 一个公开
仓库,半夜跑反而更合适。
- **失败也是一种结果**:上游不通(尤其直连 GitHub)很常见。界面会显示失败原因与上次检查
时间,失败不推送,也**不会**因此改变下一次检查的时间 —— 每个间隔重试一次,而不是每个
调度 tick(20 秒)都去撞一次。
- **同一个更新只推一次**:推送状态记的是上游 tip。推过之后,上下游没动就不会再推;上游又
有新提交(tip 变了)时会再推一条。
- **「立即检查」不发通知**:点这个按钮的人正看着结果,没必要再给自己推一条群消息。那次
检查只写结果,没推的那批提交留给下一次定时检查推。
### 部署前提:镜像里要有 git
`python:3.11-slim` 不带 git,`Dockerfile` 里已显式安装。**因此这次更新需要重建镜像**:
```bash
docker compose build && ./deploy.sh
```
`./deploy.sh` 只重建前端,不会重建镜像。漏了这步的话,检查会报「未找到 git 命令」——
界面上看得见,不会静默。
---
## 三、必须知道的限制
### 1. 「新增评论」是近似值 —— 最重要的一条
@@ -459,6 +507,9 @@ GET /api/monitor/webhook 通知配置状态(**只返回打码
POST /api/monitor/webhook 保存 Webhook 地址
DELETE /api/monitor/webhook 删除 Webhook
POST /api/monitor/webhook/test 发送测试消息
GET /api/monitor/upstream 最近一次上游检查的缓存结果(没查过返回 {})
POST /api/monitor/upstream/check 立刻检查一次(等 fetch 跑完才返回,**不发通知**)
```
> `task_id` 用**重复参数**而非逗号拼接(`?task_id=1&task_id=2`);不传表示统计全部任务。
+18 -2
View File
@@ -44,7 +44,11 @@ from . import media as douyin_media
from .client import DouYinClient
from .exception import DataFetchError
from .field import PublishTimeType
from .help import parse_video_info_from_url, parse_creator_info_from_url
from .help import (
client_hint_headers,
parse_creator_info_from_url,
parse_video_info_from_url,
)
from .login import DouYinLogin
@@ -98,7 +102,12 @@ class DouYinCrawler(AbstractCrawler):
await self.browser_context.add_init_script(path="libs/stealth.min.js")
self.context_page = await self.browser_context.new_page()
await self.context_page.goto(self.index_url)
# wait_until="domcontentloaded" instead of the default "load": the douyin
# home page never fires the load event (some long-lived request keeps it
# pending), so the default burns the whole timeout and the crawl dies
# before it starts. Measured here: domcontentloaded returns in 0.7s while
# load still times out at 90s. Tieba and Zhihu already do the same.
await self.context_page.goto(self.index_url, wait_until="domcontentloaded")
self.dy_client = await self.create_douyin_client(httpx_proxy_format)
if not await self.dy_client.pong(browser_context=self.browser_context):
@@ -315,10 +324,17 @@ class DouYinCrawler(AbstractCrawler):
self.browser_context,
urls=self.cookie_urls,
) # type: ignore
# 声称自己是 Chrome,就得带上 sec-ch-ua 系列头 —— 浏览器一定会带,而缺了它们
# 的请求在抖音网关看来就是机器人:回一个 **200 + 空 body**,不报错、不给原因,
# 表现为采集抓到 0 条。见 help.client_hint_headers 的实测记录。
client_hints = client_hint_headers(
await self.context_page.evaluate("() => navigator.userAgentData || null")
)
douyin_client = DouYinClient(
proxy=httpx_proxy,
headers={
"User-Agent": await self.context_page.evaluate("() => navigator.userAgent"),
**client_hints,
"Cookie": cookie_str,
"Host": "www.douyin.com",
"Origin": "https://www.douyin.com/",
+33 -1
View File
@@ -26,7 +26,7 @@
import random
import re
from typing import Optional
from typing import Dict, Optional
import execjs
from playwright.async_api import Page
@@ -98,6 +98,38 @@ async def get_a_bogus_from_playwright(params: str, post_data: dict, user_agent:
return a_bogus
def client_hint_headers(user_agent_data) -> Dict[str, str]:
"""由 ``navigator.userAgentData`` 还原 ``sec-ch-ua`` 系列请求头。
浏览器只要声称自己是 Chrome,就**一定会**带这三个头。缺了它们,「Chrome 的 UA +
没有 sec-ch-ua」就是最典型的机器人特征 —— 抖音网关会因此返回 **200 + 空 body**:
不报错、不给原因、HTTP 状态还是成功的,表现为采集拿到 0 条。
实测(同一 URL、同一 cookie、同一参数):不带头 → 0 字节;补上这三个头 → 7077 字节。
从 ``userAgentData`` 现算而不是写死,是为了 Chrome 升级后不会悄悄失配 —— 写死的
版本号和 UA 里的版本号一旦对不上,就又是一个可疑特征。
"""
if not isinstance(user_agent_data, dict):
return {}
brands = user_agent_data.get("brands") or []
sec_ch_ua = ", ".join(
f'"{brand.get("brand", "")}";v="{brand.get("version", "")}"' for brand in brands
)
if not sec_ch_ua:
return {}
headers = {
"sec-ch-ua": sec_ch_ua,
"sec-ch-ua-mobile": "?1" if user_agent_data.get("mobile") else "?0",
}
platform = user_agent_data.get("platform")
if platform:
headers["sec-ch-ua-platform"] = f'"{platform}"'
return headers
def parse_video_info_from_url(url: str) -> VideoUrlInfo:
"""
Parse video ID from Douyin video URL
+11
View File
@@ -272,3 +272,14 @@ class DouYinLogin(AbstractLogin):
'domain': ".douyin.com",
'path': "/"
}])
# Reload after injecting. The page was loaded *before* these cookies existed,
# so its `localStorage.HasUserLogin` still holds the logged-out value and
# check_login_state() polls it until the timeout expires (10 minutes) without
# ever succeeding -- only the *next* run works, because by then the cookies
# are in the profile. Reloading makes the site re-evaluate the session now.
try:
await self.context_page.reload(wait_until="domcontentloaded")
except Exception as exc: # pragma: no cover - reload is best effort
utils.logger.warning(
f"[DouYinLogin.login_by_cookies] reload after cookie injection failed: {exc}"
)
+31
View File
@@ -26,6 +26,37 @@ async def test_cmd_arg_crawler_max_notes_count():
config.CRAWLER_MAX_NOTES_COUNT = orig_notes
config.CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES = orig_comments
def test_douyin_monitor_command_uses_the_right_flags():
"""抖音监控任务拼出来的命令行。
与 runner 走的是同一条 _build_command 路径,所以这一条能守住「监控任务的参数
没拼错」—— 尤其是平台值必须是 dy(而不是 douyin),否则上游根本认不出平台。
"""
cm = CrawlerManager()
req = CrawlerStartRequest(
platform=PlatformEnum.DOUYIN,
login_type=LoginTypeEnum.COOKIE,
crawler_type=CrawlerTypeEnum.CREATOR,
creator_ids="https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X",
save_data_path="./data/monitor_runs/1/2",
enable_cdp_mode=True,
inject_all_cookies=True,
save_login_state=True,
max_notes_count=20,
max_comments_count=50,
)
cmd = cm._build_command(req)
idx = cmd.index("--platform")
assert cmd[idx + 1] == "dy"
idx = cmd.index("--type")
assert cmd[idx + 1] == "creator"
idx = cmd.index("--creator_id")
assert cmd[idx + 1] == "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X"
idx = cmd.index("--enable_cdp_mode")
assert cmd[idx + 1] == "true"
def test_crawler_manager_build_command():
cm = CrawlerManager()
+241
View File
@@ -0,0 +1,241 @@
# -*- coding: utf-8 -*-
"""创作者后台客户端的解析与签名。
**字段名尚未亲眼验证过**:Phase 0 抓响应时账号的数据权限还没生效,列表接口返回的是
空壳(`data.result` 里只有 `{success, code, message}`)。所以这些解析写成多别名匹配,
而这份测试就是它的规格 —— 等真实响应到手,先跑这里看哪些假设破了。
"""
import pytest
from api.creator import signing
from api.creator.client import (
CreatorClient,
as_float,
as_int,
as_seconds,
find_note_list,
normalize_note,
trans_cookies,
)
from api.creator.models import CreatorAccount
from api.creator.service import _account_dict
# --- cookie 解析 -----------------------------------------------------------
@pytest.mark.parametrize(
"raw, expected",
[
("a1=abc; web_session=xyz", {"a1": "abc", "web_session": "xyz"}),
("a1=abc;web_session=xyz;", {"a1": "abc", "web_session": "xyz"}),
("a1=abc\nweb_session=xyz", {"a1": "abc", "web_session": "xyz"}),
(" a1 = abc ; ", {"a1": "abc"}),
("", {}),
(None, {}),
# 值里可以有等号,不能被截断
("a1=abc=def", {"a1": "abc=def"}),
],
)
def test_trans_cookies(raw, expected):
assert trans_cookies(raw) == expected
def test_client_without_a1_cannot_sign():
"""a1 参与签名,没有它连请求都发不出去 —— 要提前拦而不是发出去再猜。"""
assert CreatorClient("web_session=abc").looks_authenticated is False
assert CreatorClient("a1=abc").looks_authenticated is True
# --- 数值解析 --------------------------------------------------------------
#
# 后台返回的可能是数字,也可能是 "1.2万" / "12.3%" / "1分30秒" 这类展示值。
# 解析不出来一律 None —— 不是 0。0 是真实值,None 是"不知道"。
@pytest.mark.parametrize(
"raw, expected",
[
(123, 123),
("123", 123),
("1,234", 1234),
("1.2万", 12000),
("3万", 30000),
("1.5w", 15000),
("1亿", 100000000),
(0, 0),
("0", 0),
# 这些必须是 None 而不是 0 —— 把"没给"当成"是零"会让报表说谎
(None, None),
("", None),
("-", None),
("暂无", None),
("abc", None),
(True, None),
],
)
def test_as_int(raw, expected):
assert as_int(raw) == expected
@pytest.mark.parametrize(
"raw, expected",
[
(12.3, 12.3),
("12.3%", 12.3),
("12.3", 12.3),
(None, None),
("-", None),
("暂无数据", None),
],
)
def test_as_float(raw, expected):
assert as_float(raw) == expected
@pytest.mark.parametrize(
"raw, expected",
[
(45, 45.0),
("45", 45.0),
("1分30秒", 90.0),
("2分", 120.0),
("30秒", 30.0),
("01:30", 90.0),
("1:00:00", 3600.0),
(None, None),
("-", None),
("abc", None),
],
)
def test_as_seconds(raw, expected):
assert as_seconds(raw) == expected
# --- 字段归一化 ------------------------------------------------------------
def test_normalize_note_maps_aliases():
"""不同来源的记录用不同字段名,别名表要能都接住。"""
note = normalize_note(
{
"note_id": "abc123",
"title": "标题",
"publish_time": 1700000000000,
"view_count": "1.2万",
"like_count": 34,
"collected_count": 5,
"share_count": 2,
"comment_count": 7,
"cover_click_rate": "12.5%",
"avg_watch_time": "1分30秒",
}
)
assert note["note_id"] == "abc123"
assert note["title"] == "标题"
assert note["views"] == 12000
assert note["likes"] == 34
assert note["favorites"] == 5
assert note["shares"] == 2
assert note["comments"] == 7
assert note["cover_ctr"] == 12.5
assert note["avg_watch_seconds"] == 90.0
def test_normalize_note_leaves_missing_fields_as_none():
note = normalize_note({"note_id": "abc123"})
assert note["note_id"] == "abc123"
assert note["views"] is None
assert note["likes"] is None
def test_find_note_list_digs_the_array_out_of_a_nested_payload():
"""接口的确切结构没见过,所以按"像是一批笔记记录"来找,不写死路径。"""
payload = {
"code": 0,
"data": {
"result": {
"success": True,
"notes": [
{"note_id": "n1", "views": 10, "likes": 1},
{"note_id": "n2", "views": 20, "likes": 2},
],
}
},
}
found = find_note_list(payload)
assert [item["note_id"] for item in found] == ["n1", "n2"]
def test_find_note_list_returns_empty_for_the_permission_gated_envelope():
"""权限未生效时接口返回的就是这个 —— 必须安静地给出空列表,不是报错。"""
payload = {
"code": 0,
"success": True,
"msg": "成功",
"data": {"result": {"success": True, "code": 0, "message": "success"}},
}
assert find_note_list(payload) == []
# --- 签名 ------------------------------------------------------------------
def test_signed_api_carries_the_url_prefix():
"""待签字符串必须带 `url=`。少了它网关返回 406,而 406 的响应体看不出错在哪。"""
assert signing.signed_api("/api/galaxy/user/info") == "url=/api/galaxy/user/info"
assert (
signing.signed_api("/api/x", "a=1&b=2")
== "url=/api/x?a=1&b=2"
)
def test_sign_returns_xs_and_xt():
headers = signing.sign_xyw("url=/api/galaxy/user/info", "some-a1")
assert set(headers) == {"x-s", "x-t"}
assert headers["x-s"].startswith("XYW_")
assert headers["x-t"].isdigit()
def test_signature_is_stable_for_a_fixed_timestamp():
"""同一输入同一时间戳必须得到同一签名 —— 否则说明有隐藏的随机源。"""
first = signing.sign_xyw("url=/api/x", "a1", timestamp_ms=1700000000000)
second = signing.sign_xyw("url=/api/x", "a1", timestamp_ms=1700000000000)
assert first == second
def test_signature_changes_with_the_signed_string():
"""签名必须真的绑定待签内容,否则改参数不会被发现 —— 那这个签名就没意义了。"""
base = signing.sign_xyw("url=/api/x?a=1", "a1", timestamp_ms=1700000000000)
other = signing.sign_xyw("url=/api/x?a=2", "a1", timestamp_ms=1700000000000)
assert base["x-s"] != other["x-s"]
# --- 凭证不外泄 ------------------------------------------------------------
def test_account_dict_never_carries_the_cookie():
"""cookie 等于登录态。对外结构里只该有 `has_cookie`。"""
account = CreatorAccount(
id=1,
nickname="测试",
user_id="u1",
cookie="a1=SECRET; web_session=SECRET",
created_at=0,
updated_at=0,
)
payload = _account_dict(account)
assert payload["has_cookie"] is True
assert "cookie" not in payload
assert "SECRET" not in str(payload)
+160
View File
@@ -0,0 +1,160 @@
# -*- coding: utf-8 -*-
"""运营账号扫码登录的完成判据。
这里守的是一个具体的故障:判据原先读页面里的 `window.__INITIAL_STATE__`,而那是
**页面加载那一刻的快照** —— 扫码是加载之后才登录的,快照不会翻转,于是登录明明
成功了,界面却永远停在二维码上。
现在改成拿 cookie 去问创作者后台"我是谁"。重要的是**游客也有 a1**(所以签名算得
出来),所以"有 a1"什么都不能证明,只有后台认了才算数。
"""
from unittest.mock import AsyncMock, MagicMock
import pytest
from api.creator import login as creator_login
from api.creator.client import CreatorApiError
class _FakeContext:
def __init__(self, cookies):
self._cookies = cookies
self.closed = False
async def cookies(self):
return self._cookies
async def close(self):
self.closed = True
class _FakePage:
def __init__(self):
self.closed = False
async def close(self):
self.closed = True
def _session(cookies):
context = _FakeContext(cookies)
page = _FakePage()
session = creator_login.AccountLoginSession(context, page)
# 跳过节流,让每次 refresh 都真的去问一次。
session._last_login_check = 0.0
return session, context, page
def _patch_client(monkeypatch, behaviour):
"""behaviour(cookie) -> dict 或抛异常。"""
class _Client:
def __init__(self, cookie, **kwargs):
self.cookie = cookie
async def fetch_user_info(self):
return behaviour(self.cookie)
monkeypatch.setattr(creator_login, "CreatorClient", _Client)
GUEST_COOKIES = [
{"name": "a1", "value": "guest-a1"},
{"name": "web_session", "value": "guest-session"},
]
@pytest.mark.asyncio
async def test_a_guest_session_never_completes(monkeypatch):
"""游客也有 a1,但后台回 401 —— 必须继续等,不能当成登录成功。"""
def behaviour(_cookie):
raise CreatorApiError("登录态无效或已过期", status=401)
_patch_client(monkeypatch, behaviour)
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
assert session.status == creator_login.STATUS_WAITING
assert session.cookie == ""
@pytest.mark.asyncio
async def test_a_recognised_identity_completes_the_login(monkeypatch):
_patch_client(
monkeypatch,
lambda _cookie: {"user_id": "u123", "nickname": "小明", "red_id": "1"},
)
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
assert session.status == creator_login.STATUS_SUCCESS
assert "小明" in session.message
assert "a1=guest-a1" in session.cookie
assert session.account["user_id"] == "u123"
@pytest.mark.asyncio
async def test_an_empty_identity_does_not_complete(monkeypatch):
"""接口返回 200 但没有账号标识 —— 同样不能算成功。"""
_patch_client(monkeypatch, lambda _cookie: {"user_id": "", "nickname": ""})
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
assert session.status == creator_login.STATUS_WAITING
@pytest.mark.asyncio
async def test_the_api_is_not_hit_on_every_poll(monkeypatch):
"""前端每 2 秒轮询一次,但每次轮询都打一次后台接口是浪费。"""
calls = []
def behaviour(cookie):
calls.append(cookie)
raise CreatorApiError("还没登录", status=401)
_patch_client(monkeypatch, behaviour)
session, _context, _page = _session(GUEST_COOKIES)
await session.refresh()
# 第一次之后 _last_login_check 已经是"现在",后面两次应落在节流窗口内。
await session.refresh()
await session.refresh()
assert len(calls) == 1
@pytest.mark.asyncio
async def test_expiry_beats_a_successful_scan(monkeypatch):
_patch_client(monkeypatch, lambda _cookie: {"user_id": "u1", "nickname": "x"})
session, _context, _page = _session(GUEST_COOKIES)
session.started_at -= creator_login.QR_TTL_SECONDS + 1
await session.refresh()
assert session.status == creator_login.STATUS_EXPIRED
@pytest.mark.asyncio
async def test_a_closed_window_is_reported(monkeypatch):
session, context, _page = _session(GUEST_COOKIES)
context.cookies = AsyncMock(side_effect=RuntimeError("Target closed"))
await session.refresh()
assert session.status == creator_login.STATUS_ERROR
@pytest.mark.asyncio
async def test_closing_discards_the_temporary_context():
"""临时上下文是这个设计的关键 —— 用完必须关掉,否则会挂在操作者的 Chrome 里。"""
session, context, page = _session(GUEST_COOKIES)
await session.close()
assert context.closed is True
assert page.closed is True
+401
View File
@@ -0,0 +1,401 @@
# -*- coding: utf-8 -*-
"""抖音 Web 接口客户端 —— 纯逻辑部分(不发网络请求、不连浏览器)。
发请求那半边只能在真环境里验(要 CDP 浏览器 + 登录态),所以这里钉住的是那些
「错了会一路错到入库」的地方:请求头的成套性、cookie 解析、以及产物键名。
"""
from tools.user_hash import anonymize_user_id
import asyncio
import httpx
import pytest
from api.monitor import douyin_api
class TestCookieParsing:
def test_cookie_header_is_normalised(self):
assert douyin_api._cookie_header(" a=1 ; b = 2 ;; c=3 ") == "a=1; b=2; c=3"
def test_cookie_value_lookup(self):
assert douyin_api._cookie_value("a=1; UIFID=xyz; b=2", "UIFID") == "xyz"
assert douyin_api._cookie_value("a=1", "UIFID") == ""
assert douyin_api._cookie_value("", "UIFID") == ""
def test_session_detection(self):
assert douyin_api._has_session("a=1; sessionid=abc") is True
assert douyin_api._has_session("a=1; sessionid_ss=abc") is False
assert douyin_api._has_session("") is False
class TestRequestHeaders:
"""请求头必须**成套**,而且成套地来自同一个浏览器。
实测:只有 UA + client hints + Cookie 时,主页接口回 200 但只有 121 字节(空壳);
补上 Accept / Accept-Language / Referer 才变成 7074 字节的真数据。
"""
def test_the_full_set_is_sent(self):
identity = douyin_api.BrowserIdentity(
cookie="sessionid=s; UIFID=u1",
user_agent="UA-of-this-browser",
client_hints={"sec-ch-ua": '"Chrome";v="155"'},
)
headers = identity.headers()
assert headers["User-Agent"] == "UA-of-this-browser"
assert headers["sec-ch-ua"] == '"Chrome";v="155"'
assert headers["Accept"], "Accept 系列是主页接口能不能返回真数据的必要条件"
assert headers["Accept-Language"]
assert headers["Referer"] == "https://www.douyin.com/"
assert headers["x-tt-argus"] == douyin_api.ARGUS_HEADER_VALUE
assert headers["uifid"] == "u1"
assert headers["Cookie"] == "sessionid=s; UIFID=u1"
def test_uifid_is_omitted_when_absent(self):
"""cookie 里没有 uifid 就别带 —— 送个空值反而更像异常请求。"""
identity = douyin_api.BrowserIdentity(
cookie="sessionid=s", user_agent="UA", client_hints={}
)
assert "uifid" not in identity.headers()
def test_uifid_temp_is_used_as_a_fallback(self):
identity = douyin_api.BrowserIdentity(
cookie="sessionid=s; UIFID_TEMP=temp-1", user_agent="UA", client_hints={}
)
assert identity.headers()["uifid"] == "temp-1"
class TestNormalizeAweme:
def test_keys_match_what_the_store_writes(self):
"""键名必须和 store/douyin 一模一样,否则 ingest 一条都读不到。"""
record = douyin_api.normalize_aweme(
{
"aweme_id": 7690458980574358513,
"desc": "中秋哪儿都堵",
"create_time": 1790574515,
"author": {"uid": "776719710825195", "nickname": "AA建材王总"},
"statistics": {
"digg_count": 3,
"comment_count": 1,
"collect_count": 2,
"share_count": 0,
},
"video": {"cover": {"url_list": ["https://img/cover.jpg"]}},
}
)
assert record["aweme_id"] == "7690458980574358513"
assert record["title"] == "中秋哪儿都堵"
assert record["nickname"] == "AA建材王总"
assert record["cover_url"] == "https://img/cover.jpg"
assert (
record["aweme_url"]
== "https://www.douyin.com/video/7690458980574358513"
)
# 与 store 一致:creator_hash 是 uid 的匿名哈希。
assert record["creator_hash"] == anonymize_user_id("776719710825195")
# **秒**。adapters 的 time_scale=1000 会把它换成毫秒 —— 这一层不算毫秒。
assert record["create_time"] == 1790574515
# 指标按 store 的形态落成字符串,交给 ingest 的 parse_count 解析。
assert record["liked_count"] == "3"
assert record["collected_count"] == "2"
def test_missing_fields_do_not_crash(self):
record = douyin_api.normalize_aweme({"aweme_id": "1"})
assert record["aweme_id"] == "1"
assert record["title"] == ""
assert record["cover_url"] == ""
assert record["liked_count"] == "0"
assert record["create_time"] == 0
class TestIdentityFromPages:
"""身份得从浏览器里问,但**不能被一个卡死的标签页拖住**。
实测过:标签页 URL 为空、渲染进程卡死,``page.evaluate`` 永远不返回;而问身份是采集的
第一步 —— 没超时的话整轮就挂在那儿,run 永远停在「运行中」。
"""
class _Page:
def __init__(self, url, *, user_agent=None, hang=False):
self.url = url
self._user_agent = user_agent
self._hang = hang
self.closed = False
async def evaluate(self, expression):
if self._hang:
await asyncio.sleep(30) # 模拟渲染进程卡死
if expression.startswith("() => navigator.userAgentData"):
return {
"brands": [{"brand": "Chrome", "version": "155"}],
"mobile": False,
"platform": "Linux",
}
return self._user_agent
async def close(self):
self.closed = True
class _Context:
def __init__(self, pages, temp=None):
self.pages = pages
self._temp = temp
self.made_temp = False
async def new_page(self):
self.made_temp = True
if self._temp is None:
raise AssertionError("这个用例不该走到临时页")
return self._temp
def _fast_timeout(self, monkeypatch):
monkeypatch.setattr(douyin_api, "EVALUATE_TIMEOUT_SECONDS", 0.05)
@pytest.mark.asyncio
async def test_a_hanging_page_is_skipped(self, monkeypatch):
self._fast_timeout(monkeypatch)
stuck = self._Page("", hang=True)
good = self._Page("https://example.com/", user_agent="UA-of-good-page")
user_agent, hints = await douyin_api._identity_from_pages(
self._Context([stuck, good])
)
assert user_agent == "UA-of-good-page"
assert hints["sec-ch-ua"] == '"Chrome";v="155"'
@pytest.mark.asyncio
async def test_a_page_without_a_user_agent_is_skipped(self, monkeypatch):
self._fast_timeout(monkeypatch)
blank = self._Page("", user_agent=None)
good = self._Page("https://example.com/", user_agent="UA-of-good-page")
user_agent, _ = await douyin_api._identity_from_pages(
self._Context([blank, good])
)
assert user_agent == "UA-of-good-page"
@pytest.mark.asyncio
async def test_it_opens_a_temporary_page_when_nothing_else_works(self, monkeypatch):
self._fast_timeout(monkeypatch)
stuck = self._Page("", hang=True)
temp = self._Page("about:blank", user_agent="UA-of-temp-page")
context = self._Context([stuck], temp=temp)
user_agent, _ = await douyin_api._identity_from_pages(context)
assert context.made_temp is True
assert user_agent == "UA-of-temp-page"
assert temp.closed is True, "临时页问完要关掉,别在操作者的浏览器里留垃圾"
@pytest.mark.asyncio
async def test_a_douyin_page_is_preferred(self, monkeypatch):
"""有抖音页面就先问它 —— 它才是我们要模仿的那个身份。"""
self._fast_timeout(monkeypatch)
other = self._Page("https://example.com/", user_agent="UA-of-other")
douyin = self._Page("https://www.douyin.com/explore", user_agent="UA-of-douyin")
user_agent, _ = await douyin_api._identity_from_pages(
self._Context([other, douyin])
)
assert user_agent == "UA-of-douyin"
class TestGet:
"""`_get` 的失败路径 —— 它们决定了失败会不会被伪装成「这个博主没作品」。"""
@staticmethod
def _client_returning(monkeypatch, status_code: int, text: str):
class _Response:
def json(self):
import json as _json
return _json.loads(self.text)
response = _Response()
response.status_code = status_code
response.text = text
class _Client:
async def __aenter__(self):
return self
async def __aexit__(self, *exc):
return False
async def get(self, *args, **kwargs):
return response
monkeypatch.setattr(httpx, "AsyncClient", lambda **kwargs: _Client())
def _identity(self):
return douyin_api.BrowserIdentity(
cookie="sessionid=s", user_agent="UA", client_hints={}
)
def test_an_empty_body_is_an_error_not_an_empty_result(self, monkeypatch):
"""「200 + 空 body」是网关拒绝请求的典型回应。
必须当场报错 —— 放过去的话,它会在下游变成「这个博主没作品」,把一次失败伪装成
一条正常的空结果。爬虫那条路就是这么栽的,还被翻译成「账号被封」。
"""
self._client_returning(monkeypatch, 200, "")
with pytest.raises(douyin_api.DouyinApiError) as excinfo:
asyncio.run(douyin_api._get("/x", {}, self._identity()))
assert "空内容" in str(excinfo.value)
def test_a_403_carries_the_gateways_own_message(self, monkeypatch):
"""抖音难得会说原因,把它带出来,别丢。"""
self._client_returning(
monkeypatch, 403, "Blocked by ArgusSecurityPlugin Uifid Not Found"
)
with pytest.raises(douyin_api.DouyinApiError) as excinfo:
asyncio.run(douyin_api._get("/x", {}, self._identity()))
assert "403" in str(excinfo.value)
assert "Uifid Not Found" in str(excinfo.value)
def test_a_200_with_data_is_returned_as_is(self, monkeypatch):
self._client_returning(monkeypatch, 200, '{"user": {"nickname": "x"}}')
assert asyncio.run(douyin_api._get("/x", {}, self._identity())) == {
"user": {"nickname": "x"}
}
@pytest.fixture
def fake_signer(monkeypatch):
"""把真正的 ``a_bogus`` 签名换成假的。
真的那个在 import 的瞬间就要把 ``libs/douyin.js`` 喂给 execjs(还得有 node 和正确的
相对路径),单元测试不该依赖这些。**签名本身是实测过的**:带上它是 200 + 真评论,
不带是 200 + 空 body。这里只负责钉住「有没有带上、传对了没有」。
"""
import sys
import types
module = types.ModuleType("media_platform.douyin.help")
calls: list = []
def get_a_bogus_from_js(url: str, params: str, user_agent: str) -> str:
calls.append({"url": url, "params": params, "user_agent": user_agent})
return "FAKE-BOGUS"
module.get_a_bogus_from_js = get_a_bogus_from_js
monkeypatch.setitem(sys.modules, "media_platform.douyin.help", module)
return calls
@pytest.fixture
def recording_client(monkeypatch):
"""记下实际发出去的那次请求。"""
sent: dict = {}
class _Response:
status_code = 200
text = '{"comments": []}'
def json(self):
return {"comments": []}
class _Client:
async def __aenter__(self):
return self
async def __aexit__(self, *exc):
return False
async def get(self, url, **kwargs):
sent["url"] = url
sent.update(kwargs)
return _Response()
monkeypatch.setattr(httpx, "AsyncClient", lambda **kwargs: _Client())
return sent
class TestCommentSigning:
"""评论接口**必须**带 a_bogus。
不带的话网关回 200 + 空 body —— 那在下游会变成「这条作品没有评论」,把一次被挡住的
请求伪装成一条正常的空结果。和登录失效长得一模一样,查起来能查半天。
"""
def _identity(self):
return douyin_api.BrowserIdentity(
cookie="sessionid=s", user_agent="UA", client_hints={}
)
def test_the_signature_is_computed_over_the_unsigned_params(
self, fake_signer, recording_client
):
asyncio.run(
douyin_api._get(
douyin_api.COMMENT_PATH,
{"aweme_id": "123", "count": 20},
self._identity(),
signed=True,
)
)
assert len(fake_signer) == 1
# 签名算在**不含 a_bogus** 的那串上 —— 把它自己也算进去是循环的。
assert fake_signer[0]["params"] == "aweme_id=123&count=20"
assert fake_signer[0]["url"] == douyin_api.COMMENT_PATH
assert fake_signer[0]["user_agent"] == "UA"
def test_the_signature_goes_out_with_the_request(self, fake_signer, recording_client):
asyncio.run(
douyin_api._get(
douyin_api.COMMENT_PATH, {"aweme_id": "123"}, self._identity(), signed=True
)
)
assert recording_client["params"]["a_bogus"] == "FAKE-BOGUS"
assert recording_client["params"]["aweme_id"] == "123"
def test_unsigned_calls_never_touch_the_signer(self, fake_signer, recording_client):
"""作品 / 详情 / 博主资料三个接口不带签名也照常返回。
给它们加签名是**没验证过的改动** —— 所以这里钉住「不签」,防止有人图省事把
signed=True 改成全局默认。
"""
asyncio.run(
douyin_api._get(douyin_api.POSTS_PATH, {"sec_user_id": "x"}, self._identity())
)
assert fake_signer == []
assert "a_bogus" not in recording_client["params"]
def test_a_broken_signer_is_reported_as_such(self, monkeypatch, recording_client):
"""execjs 起不来时要说出是签名失败,而不是让它变成「没评论」。"""
import sys
import types
module = types.ModuleType("media_platform.douyin.help")
def boom(url, params, user_agent):
raise RuntimeError("node 没装")
module.get_a_bogus_from_js = boom
monkeypatch.setitem(sys.modules, "media_platform.douyin.help", module)
with pytest.raises(douyin_api.DouyinApiError, match="a_bogus"):
asyncio.run(
douyin_api._get(
douyin_api.COMMENT_PATH, {"aweme_id": "1"}, self._identity(), signed=True
)
)
+55
View File
@@ -0,0 +1,55 @@
# -*- coding: utf-8 -*-
"""sec-ch-ua 系列请求头:为什么必须带、怎么还原。
抖音网关对「声称自己是 Chrome、却没带 sec-ch-ua」的请求会回 **200 + 空 body** ——
不报错、不给原因、HTTP 状态还是成功的,采集侧只看到 0 条。实测同一 URL、同一 cookie、
同一参数:不带头 0 字节,补上这三个头 7077 字节。
"""
import pytest
from media_platform.douyin.help import client_hint_headers
# 实测从真实浏览器抓到的 navigator.userAgentData
REAL_UA_DATA = {
"brands": [
{"brand": "Google Chrome", "version": "155"},
{"brand": "Chromium", "version": "155"},
{"brand": "Not(A:Brand", "version": "24"},
],
"mobile": False,
"platform": "Linux",
}
def test_headers_are_derived_from_user_agent_data():
hints = client_hint_headers(REAL_UA_DATA)
assert hints["sec-ch-ua"] == (
'"Google Chrome";v="155", "Chromium";v="155", "Not(A:Brand";v="24"'
)
assert hints["sec-ch-ua-mobile"] == "?0"
assert hints["sec-ch-ua-platform"] == '"Linux"'
def test_mobile_is_reflected():
hints = client_hint_headers({**REAL_UA_DATA, "mobile": True})
assert hints["sec-ch-ua-mobile"] == "?1"
@pytest.mark.parametrize("value", [None, [], "not-a-dict", {}, {"brands": []}])
def test_nothing_is_invented_when_user_agent_data_is_unavailable(value):
"""拿不到就返回空。
凭空造一组和 UA 对不上的头,只会变成**另一个**可疑特征 —— 那比不带头更糟。
"""
assert client_hint_headers(value) == {}
def test_missing_platform_still_sends_the_other_two():
hints = client_hint_headers({"brands": REAL_UA_DATA["brands"], "mobile": False})
assert "sec-ch-ua" in hints
assert "sec-ch-ua-mobile" in hints
assert "sec-ch-ua-platform" not in hints
+411
View File
@@ -0,0 +1,411 @@
# -*- coding: utf-8 -*-
"""抖音采集编排 —— 不碰网络,把 douyin_api 整个换掉。
验的是编排本身:产物落在正确的目录、文件名是 store 那套、以及**作品列表被挡时的退化**
(用已知 aweme_id 逐条刷新)—— 那条退化路径决定了今天这个功能是「完全没用」还是
「已知作品还能看」。
"""
import json
import pytest
from api.monitor import douyin_api, douyin_fetch
def _profile(**overrides) -> dict:
profile = {
"creator_hash": "hash",
"nickname": "博主",
"unique_id": "abc",
"fans": 12000,
"total_favorited": 83000,
"works": 42,
"following": 7,
}
profile.update(overrides)
return profile
@pytest.fixture(autouse=True)
def fake_profile(monkeypatch):
"""每个用例都挡住「问博主资料」这一跳。
它是附加信息,不在任何一条编排路径上,但真发出去就会去连 9222 那个浏览器 ——
于是所有 creator 用例都会多出一次连接失败、并把 ``errors`` 弄脏。想验它自己的
用例再单独覆盖这个 fixture。
"""
async def _profile_call(sec_user_id, *, cookie=""):
return _profile()
monkeypatch.setattr(douyin_api, "author_profile", _profile_call)
def _video(aweme_id: str, likes: str = "1") -> dict:
return {
"aweme_id": aweme_id,
"title": f"title-{aweme_id}",
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
"cover_url": "",
"aweme_type": "0",
"create_time": 1790574515,
"creator_hash": "hash",
"nickname": "博主",
"liked_count": likes,
"comment_count": "0",
"collected_count": "0",
"share_count": "0",
}
def _comment(aweme_id: str, index: int) -> dict:
return {
"comment_id": f"c{index}",
"aweme_id": aweme_id,
"content": f"评论{index}",
"nickname": "路人",
"creator_hash": "h2",
"create_time": 1790574600,
"like_count": "0",
"sub_comment_count": "0",
"parent_comment_id": "0",
}
class _Target:
def __init__(self, external_id: str) -> None:
self.external_id = external_id
def _read(path):
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line]
async def _collect(tmp_path, **overrides):
kwargs = dict(
platform="dy",
mode="creator",
limit=20,
want_comments=True,
comment_limit=20,
targets=[_Target("MS4w-sec")],
cookie="sessionid=x",
)
kwargs.update(overrides)
return await douyin_fetch.collect(tmp_path, **kwargs)
class TestHappyPath:
@pytest.mark.asyncio
async def test_writes_the_layout_ingest_expects(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111"), _video("222")]
async def fake_comments(aweme_id, count=20, *, cookie=""):
return [_comment(aweme_id, 1)]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "video_comments", fake_comments)
result = await _collect(tmp_path)
assert result["notes"] == 2
assert result["comments"] == 2
assert result["errors"] == []
# 目录名必须是 douyin(不是平台 id dy)—— ingest 找文件用的是同一个来源。
jsonl_dir = tmp_path / "douyin" / "jsonl"
assert jsonl_dir.is_dir()
contents = list(jsonl_dir.glob("*_contents_*.jsonl"))
comments = list(jsonl_dir.glob("*_comments_*.jsonl"))
assert len(contents) == 1
assert len(comments) == 1
notes = _read(contents[0])
assert [n["aweme_id"] for n in notes] == ["111", "222"]
# 键名照抄 store/douyin —— ingest 靠这个读出来。
assert notes[0]["aweme_url"] == "https://www.douyin.com/video/111"
# 秒,不是毫秒;换算交给 adapters。
assert notes[0]["create_time"] == 1790574515
assert _read(comments[0])[0]["aweme_id"] == "111"
@pytest.mark.asyncio
async def test_the_comment_file_exists_even_without_comments(self, monkeypatch, tmp_path):
"""评论文件必须建出来。
ingest 靠「文件在不在」区分「这一轮没评论」和「这一轮什么都没抓到」——
两种情况的含义完全不同。
"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
await _collect(tmp_path, want_comments=False)
assert len(list((tmp_path / "douyin" / "jsonl").glob("*_comments_*.jsonl"))) == 1
@pytest.mark.asyncio
async def test_duplicate_works_are_written_once(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111"), _video("111")]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
result = await _collect(tmp_path, want_comments=False)
assert result["notes"] == 1
class TestNoteMode:
@pytest.mark.asyncio
async def test_a_work_target_is_fetched_by_detail_not_by_creator_list(
self, monkeypatch, tmp_path
):
"""作品模式的目标**本身就是作品 id**,不能拿它当博主的 sec_uid 去查列表。
走错了会必然失败,而且失败原因很难看懂(接口说你没登录/不是浏览器)——
「粘贴作品链接的监控」今天本来是能用的,别让它因为这一处走错而废掉。
"""
async def must_not_be_called(*args, **kwargs):
raise AssertionError("作品模式不该去拉博主的作品列表")
async def fake_detail(aweme_id, *, cookie=""):
return _video(aweme_id)
monkeypatch.setattr(douyin_api, "author_videos", must_not_be_called)
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
result = await _collect(
tmp_path,
mode="note",
want_comments=False,
targets=[_Target("111"), _Target("222")],
)
assert result["notes"] == 2
assert result["errors"] == []
notes = _read(list((tmp_path / "douyin" / "jsonl").glob("*_contents_*.jsonl"))[0])
assert [n["aweme_id"] for n in notes] == ["111", "222"]
@pytest.mark.asyncio
async def test_a_broken_work_does_not_lose_the_others(self, monkeypatch, tmp_path):
async def flaky(aweme_id, *, cookie=""):
if aweme_id == "222":
raise douyin_api.DouyinApiError("作品已被删除")
return _video(aweme_id)
monkeypatch.setattr(douyin_api, "video_detail", flaky)
result = await _collect(
tmp_path,
mode="note",
want_comments=False,
targets=[_Target("111"), _Target("222")],
)
assert result["notes"] == 1
assert any("222" in error for error in result["errors"])
class TestDegradation:
"""作品列表被挡时的行为 —— 决定了这个功能今天有没有用。"""
@pytest.mark.asyncio
async def test_falls_back_to_refreshing_known_works(self, monkeypatch, tmp_path):
async def blocked(sec_user_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("接口返回了空内容")
refreshed = []
async def fake_detail(aweme_id, *, cookie=""):
refreshed.append(aweme_id)
return _video(aweme_id, likes="9")
monkeypatch.setattr(douyin_api, "author_videos", blocked)
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
result = await _collect(
tmp_path, want_comments=False, known_aweme_ids=["999", "888"]
)
assert refreshed == ["999", "888"]
assert result["notes"] == 2
# 但错误照样报出来 —— 这一轮是「部分可用」,不是「一切正常」,别粉饰。
assert any("作品列表失败" in error for error in result["errors"])
@pytest.mark.asyncio
async def test_known_works_are_not_refetched_for_every_target(self, monkeypatch, tmp_path):
"""退化路径不能每个目标都把同一批已知作品再刷一遍。
任务有多个目标时那会让同一件作品在一轮里出现两次,而指标快照的唯一键是
(task_id, note_id, run_id) —— 第二次插入直接撞键,整个 run 崩掉(踩过:
Duplicate entry for key 'uq_note_metric')。
"""
async def blocked(sec_user_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("接口返回了空内容")
calls = []
async def fake_detail(aweme_id, *, cookie=""):
calls.append(aweme_id)
return _video(aweme_id)
monkeypatch.setattr(douyin_api, "author_videos", blocked)
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
result = await _collect(
tmp_path,
want_comments=False,
targets=[_Target("sec-a"), _Target("sec-b")],
known_aweme_ids=["999"],
)
assert calls == ["999"], "同一件已知作品只该刷一次,而不是每个目标一次"
assert result["notes"] == 1
@pytest.mark.asyncio
async def test_nothing_at_all_still_reports_the_reason(self, monkeypatch, tmp_path):
async def blocked(sec_user_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("接口返回了空内容")
monkeypatch.setattr(douyin_api, "author_videos", blocked)
result = await _collect(tmp_path, want_comments=False)
assert result["notes"] == 0
assert result["errors"]
# 产物仍然写出来(空的),让调用方去判断这是失败而不是「这个博主没作品」。
assert (tmp_path / "douyin" / "jsonl").is_dir()
@pytest.mark.asyncio
async def test_a_failing_comment_fetch_does_not_lose_the_work(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
async def broken_comments(aweme_id, count=20, *, cookie=""):
raise douyin_api.DouyinApiError("评论接口抽风")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "video_comments", broken_comments)
result = await _collect(tmp_path)
# 评论拿不到是小事,作品不能跟着丢。
assert result["notes"] == 1
assert any("评论失败" in error for error in result["errors"])
class TestCreatorProfile:
"""博主的**账号级**指标 —— 粉丝 / 总获赞 / 作品数。
作品列表给不了这个东西:它说的是一件作品涨了多少赞,不是这个人整个账号的粉丝
在涨还是在掉。单独问一次资料接口。
"""
@staticmethod
def _profiles(tmp_path):
files = list((tmp_path / "douyin" / "jsonl").glob("*_profile_*.jsonl"))
assert len(files) == 1
return _read(files[0])
@pytest.mark.asyncio
async def test_the_profile_lands_in_the_run_dir(self, monkeypatch, tmp_path):
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
await _collect(tmp_path, want_comments=False)
assert self._profiles(tmp_path) == [_profile()]
@pytest.mark.asyncio
async def test_the_profile_is_keyed_by_the_same_hash_as_the_works(
self, monkeypatch, tmp_path
):
"""**这条是关键。** 快照表的唯一键是 (任务, creator_hash, 轮次),而界面上是按
作品的 creator_hash 归组去查它的。两边只要差一个字符,粉丝数就永远查不出来 ——
而且是静默的:表里有数据,界面上什么都没有。
"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")] # 作品带的哈希是 "hash"
# 资料接口自己算出来的是另一个值(比如它那边 uid 缺字段、只能拿 sec_uid 算)。
async def off_hash_profile(sec_user_id, *, cookie=""):
return _profile(creator_hash="另一个哈希")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "author_profile", off_hash_profile)
await _collect(tmp_path, want_comments=False)
# 以作品为准:作品才是界面上的行,快照必须挂在能查到它的那个键上。
assert self._profiles(tmp_path)[0]["creator_hash"] == "hash"
@pytest.mark.asyncio
async def test_a_profile_without_an_identity_is_dropped(self, monkeypatch, tmp_path):
"""哈希都算不出来的快照,落下去只会是一条谁也查不到的垃圾。"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return []
async def anonymous_profile(sec_user_id, *, cookie=""):
return _profile(creator_hash="")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "author_profile", anonymous_profile)
result = await _collect(tmp_path, want_comments=False)
assert self._profiles(tmp_path) == []
assert any("身份标识" in error for error in result["errors"])
@pytest.mark.asyncio
async def test_a_failing_profile_does_not_lose_the_works(self, monkeypatch, tmp_path):
"""附加信息拿不到,这一轮采到的作品不能跟着判成失败。"""
async def fake_videos(sec_user_id, count=20, *, cookie=""):
return [_video("111")]
async def broken_profile(sec_user_id, *, cookie=""):
raise douyin_api.DouyinApiError("资料接口抽风")
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
monkeypatch.setattr(douyin_api, "author_profile", broken_profile)
result = await _collect(tmp_path, want_comments=False)
assert result["notes"] == 1
assert self._profiles(tmp_path) == []
assert any("资料失败" in error for error in result["errors"])
@pytest.mark.asyncio
async def test_a_work_target_never_asks_for_a_profile(self, monkeypatch, tmp_path):
"""作品模式的目标是一件作品,没有「这个博主是谁」可问 —— 不该白发一个请求。"""
asked = []
async def fake_detail(aweme_id, *, cookie=""):
return _video(aweme_id)
async def recording_profile(sec_user_id, *, cookie=""):
asked.append(sec_user_id)
return _profile()
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
monkeypatch.setattr(douyin_api, "author_profile", recording_profile)
await _collect(tmp_path, mode="note", want_comments=False)
assert asked == []
# 文件仍然建出来(空的):事后翻 run 目录能看出「这次根本没问过」。
assert self._profiles(tmp_path) == []
+11
View File
@@ -41,6 +41,17 @@ from database.models import Base, DouyinAweme, DouyinAwemeComment
from tools.user_hash import anonymize_user_id, mask_nickname
@pytest.fixture(autouse=True)
def _force_nickname_masking(monkeypatch):
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
所以这里显式打开来测。
"""
monkeypatch.setattr(config, "MASK_NICKNAME", True)
# 抖音教学版禁用字段(键):不得作为存储 dict 的 key 出现。
FORBIDDEN_KEYS = {
"user_id", "sec_uid", "short_user_id", "user_unique_id",
+13
View File
@@ -18,10 +18,23 @@ import types
import pytest
import config
import store.kuaishou as ks
from store.kuaishou import update_kuaishou_video, update_ks_video_comment
from tools.user_hash import anonymize_user_id, mask_nickname
@pytest.fixture(autouse=True)
def _force_nickname_masking(monkeypatch):
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
所以这里显式打开来测。
"""
monkeypatch.setattr(config, "MASK_NICKNAME", True)
# 教学版禁用字段(键):一律不得出现在存储 dict 中。
# 昵称字段 nickname 允许保留,但值须脱敏。
FORBIDDEN_KEYS = {"user_id", "avatar", "signature", "ip_location", "gender"}
+111
View File
@@ -78,6 +78,117 @@ class TestParseTargetInput:
with pytest.raises(TargetParseError):
parse_target_input("not a url at all !!", "creator")
# --- 抖音 -------------------------------------------------------------
# 链接形态由平台决定,所以每一个都要显式带上 "dy"。
def test_douyin_creator_url(self):
parsed = parse_target_input(
"https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
"?from_tab_name=main",
"creator",
"dy",
)
assert (
parsed["external_id"]
== "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
)
def test_douyin_video_url(self):
parsed = parse_target_input(
"https://www.douyin.com/video/7525082444551310602", "note", "dy"
)
assert parsed["external_id"] == "7525082444551310602"
def test_douyin_modal_id_url(self):
"""在别人主页或搜索结果里点开视频,拿到的就是带 modal_id 的链接。"""
parsed = parse_target_input(
"https://www.douyin.com/root/search/python?aid=b733a3b0&modal_id=7471165520058862848",
"note",
"dy",
)
assert parsed["external_id"] == "7471165520058862848"
def test_douyin_bare_sec_uid_is_accepted(self):
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
parsed = parse_target_input(sec_uid, "creator", "dy")
assert parsed["external_id"] == sec_uid
def test_douyin_bare_sec_uid_beyond_the_xhs_length_cap(self):
"""裸 id 的长度上限必须按平台分开。
小红书那条规则封顶 64 字符,而 sec_user_id 长过 64 是常态(实测样本 55,
但字段本身是变长的)。共用一条规则的话,长一点的 sec_uid 会被直接拒掉 ——
对用户来说就是「粘贴了一个完全正确的链接却报无法识别」。
"""
sec_uid = "MS4wLjABAAAA" + "aB3dEf6hIj9lMn2pQr5tUv8xYz1" * 3
assert len(sec_uid) > 64
parsed = parse_target_input(sec_uid, "creator", "dy")
assert parsed["external_id"] == sec_uid
# 同一条 id 拿小红书规则来解析会被拒 —— 这正是两条规则必须分开的原因。
with pytest.raises(TargetParseError):
parse_target_input(sec_uid, "creator", "xhs")
def test_douyin_bare_video_id(self):
parsed = parse_target_input("7525082444551310602", "note", "dy")
assert parsed["external_id"] == "7525082444551310602"
# 抖音不需要 xsec_token —— 和小红书不同,裸链接就能用。
assert parsed["xsec_token"] == ""
def test_douyin_short_link_is_rejected_with_a_reason(self):
"""短链要联网跳一次才知道指向谁。明确拒绝好过存一个永远抓不到东西的目标。"""
with pytest.raises(TargetParseError) as excinfo:
parse_target_input("https://v.douyin.com/drIPtQ_WPWY/", "note", "dy")
assert "短链" in str(excinfo.value)
def test_a_douyin_link_is_not_parsed_with_xhs_rules(self):
with pytest.raises(TargetParseError):
parse_target_input(
"https://www.douyin.com/video/7525082444551310602", "note", "xhs"
)
def test_an_xhs_link_is_not_parsed_for_douyin(self):
with pytest.raises(TargetParseError):
parse_target_input(NOTE_URL, "note", "dy")
def test_a_platform_without_an_adapter_is_rejected(self):
with pytest.raises(TargetParseError):
parse_target_input("whatever", "creator", "bili")
class TestTargetReplacement:
@pytest.mark.asyncio
async def test_replacing_targets_uses_the_tasks_own_platform(self, client):
"""改目标必须按任务**自己**的平台解析。
``update_task`` 原先漏传了 platform,解析回落到默认的小红书。只有小红书时
行为恰好正确,接上抖音就会拿小红书的正则去解析抖音链接 —— 建任务时对、
改任务时错,是最难注意到的那种不一致。
"""
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
created = await client.post(
"/api/monitor/tasks",
json={"name": "抖音", "mode": "creator", "platform": "dy", "targets": [sec_uid]},
)
assert created.status_code == 201
task_id = created.json()["id"]
updated = await client.patch(
f"/api/monitor/tasks/{task_id}",
json={"targets": [f"https://www.douyin.com/user/{sec_uid}"]},
)
assert updated.status_code == 200
tasks = (
await client.get("/api/monitor/tasks", params={"platform": "dy"})
).json()["tasks"]
target = tasks[0]["targets"][0]
assert target["external_id"] == sec_uid
assert target["raw_value"].startswith("https://www.douyin.com/user/")
class TestTaskCrud:
@pytest.mark.asyncio
+83
View File
@@ -0,0 +1,83 @@
# -*- coding: utf-8 -*-
"""Guards for the in-place column migration.
``create_all`` creates missing tables but never adds columns to a table that
already exists, so new model columns are applied by ``_ensure_columns``. That step
was originally driven by a hand-kept list, and forgetting to update it did not
fail loudly -- the app still started, connected, and then failed on every query
and every scheduler tick. These tests pin down its replacement, which derives the
work from the model metadata.
"""
import pytest
from sqlalchemy import Boolean, Column, Integer, MetaData, String, Table
from api.monitor import db
from api.monitor.models import MonitorBase
def _migratable_columns():
for table in MonitorBase.metadata.sorted_tables:
for column in table.columns:
if column.primary_key:
continue
yield pytest.param(column, id=f"{table.name}.{column.name}")
def _ddl(column: Column) -> str:
"""Render a detached column, so the tests never mutate the real metadata."""
scratch = Table("scratch", MetaData(), column)
return db._column_ddl(scratch.columns[0])
@pytest.mark.parametrize("column", _migratable_columns())
def test_column_renders_as_ddl(column):
ddl = db._column_ddl(column)
# SQLAlchemy back-quotes an identifier only when it has to, so the rendered
# name matches the column's with the quoting stripped -- which is the case for
# reserved words like monitor_run.trigger. That quoting is a feature: the
# hand-kept list this replaced would have emitted bare `trigger` and died on a
# syntax error.
assert ddl.split(" ", 1)[0].strip("`") == column.name
# MySQL refuses AUTO_INCREMENT together with the DEFAULT this helper appends
# to NOT NULL columns.
assert "AUTO_INCREMENT" not in ddl
@pytest.mark.parametrize("column", _migratable_columns())
def test_not_null_columns_carry_a_default(column):
"""So ADD COLUMN cannot fail on a table that already holds rows.
Without a DEFAULT, whether the ALTER succeeds depends on the server's
sql_mode -- not something a deployment should hinge on.
"""
if column.nullable:
pytest.skip("nullable column needs no seed value")
assert "DEFAULT" in db._column_ddl(column)
def test_boolean_default_becomes_a_mysql_literal():
"""Python's True is not a SQL keyword; it has to become 1."""
assert "DEFAULT 1" in _ddl(Column("flag", Boolean, nullable=False, default=True))
def test_string_defaults_are_quoted():
assert "DEFAULT 'interval'" in _ddl(
Column("mode", String(16), nullable=False, default="interval")
)
def test_integer_defaults_are_not_quoted():
ddl = _ddl(Column("n", Integer, nullable=False, default=0))
assert "DEFAULT 0" in ddl
assert "DEFAULT '0'" not in ddl
def test_a_not_null_column_without_a_model_default_falls_back_to_zero():
"""Belt and braces: even a column the model gives no default still migrates."""
ddl = _ddl(Column("n", Integer, nullable=False))
assert "DEFAULT 0" in ddl
+144 -4
View File
@@ -20,6 +20,7 @@
import csv
import io
import re
import httpx
import pytest
@@ -36,6 +37,11 @@ from api.monitor.models import (
TASK_NAME = "评论归属测试"
# 作品的发布时间。和 first_seen_at(我们第一次看到它)刻意取不同的值 —— 两者混成
# 一个概念是最容易犯的错。
PUBLISHED_A = 1_699_000_000_000
PUBLISHED_B = 1_699_100_000_000
async def _seed():
"""Two works; three comments on the first, one on the second."""
@@ -49,13 +55,18 @@ async def _seed():
session.add(task)
await session.flush()
for note_id, title in (("note-a", "作品甲"), ("note-b", "作品乙")):
# 两个作品**属于不同的博主** —— 评论流最外层按创作者分组,同一个人就没得测了。
for note_id, title, creator_hash, creator_name, published_at in (
("note-a", "作品甲", "hash-a", "博主甲", PUBLISHED_A),
("note-b", "作品乙", "hash-b", "博主乙", PUBLISHED_B),
):
session.add(
MonitorNote(
task_id=task.id, note_id=note_id, title=title,
note_url=f"https://www.xiaohongshu.com/explore/{note_id}",
cover=f"https://img/{note_id}.jpg", creator_hash="h",
source_kind="video", published_at=None,
cover=f"https://img/{note_id}.jpg", creator_hash=creator_hash,
creator_name=creator_name,
source_kind="video", published_at=published_at,
first_seen_run_id=1, first_seen_at=1_700_000_000_000,
last_seen_run_id=1, last_seen_at=1_700_000_000_000,
)
@@ -139,6 +150,58 @@ class TestGroupByNote:
# note-b's only comment is the most recent overall.
assert groups[0]["note_id"] == "note-b"
@pytest.mark.asyncio
async def test_each_bucket_carries_its_creator(self, client):
"""桶上必须带作品的创作者 —— 评论流最外层就是按它分组的。
少了这两个字段,前端拿到的 creator_hash / creator_name 都是 undefined,
于是所有博主塌成同一个分组、标签回退成「未知博主」:一个人都分不出来。
"""
groups = {
group["note_id"]: group
for group in (
await client.get("/api/monitor/comments", params={"group_by": "note"})
).json()["groups"]
}
assert groups["note-a"]["creator_hash"] == "hash-a"
assert groups["note-a"]["creator_name"] == "博主甲"
assert groups["note-b"]["creator_hash"] == "hash-b"
assert groups["note-b"]["creator_name"] == "博主乙"
# 两个作品的创作者必须真的不同,否则界面上照样分不出来。
assert groups["note-a"]["creator_hash"] != groups["note-b"]["creator_hash"]
@pytest.mark.asyncio
async def test_each_bucket_carries_the_publish_date(self, client):
"""作品那一层要带发布日期:同名作品不少,日期能帮着认。
注意它和 first_seen_at 是两个概念 —— 前者是作者发布的那天,后者是我们第一次
看到它的那天。把老作品加进监控时两者能差好几个月。
"""
groups = {
group["note_id"]: group
for group in (
await client.get("/api/monitor/comments", params={"group_by": "note"})
).json()["groups"]
}
assert groups["note-a"]["published_at"] == PUBLISHED_A
assert groups["note-b"]["published_at"] == PUBLISHED_B
assert groups["note-a"]["published_at"] != groups["note-b"]["published_at"]
@pytest.mark.asyncio
async def test_the_notes_endpoint_exposes_the_publish_date(self, client):
"""作品列表也要带上它 —— 作品栏就是靠这个显示「发布日期」列的。"""
notes = {
note["note_id"]: note
for note in (await client.get("/api/monitor/notes")).json()["notes"]
}
assert notes["note-a"]["published_at"] == PUBLISHED_A
assert notes["note-b"]["published_at"] == PUBLISHED_B
# 和「首次发现」不是同一个值 —— 两者混了的话这个断言会抓到。
assert notes["note-a"]["first_seen_at"] != notes["note-a"]["published_at"]
@pytest.mark.asyncio
async def test_flat_shape_is_unchanged_without_the_flag(self, client):
body = (await client.get("/api/monitor/comments")).json()
@@ -202,7 +265,8 @@ class TestExport:
workbook = load_workbook(io.BytesIO(response.content))
sheet = workbook.active
assert sheet.max_row == 5 # header + four comments
assert sheet.cell(row=1, column=1).value == "所属作品"
assert sheet.cell(row=1, column=1).value == "博主昵称"
assert sheet.cell(row=1, column=2).value == "所属作品"
@pytest.mark.asyncio
async def test_report_export(self, client):
@@ -230,3 +294,79 @@ class TestExport:
"/api/monitor/export", params={"kind": "comments", "note_id": "no-such-note"}
)
assert response.status_code == 404
class TestNotesExportColumns:
"""作品导出的列 —— 「有列名」和「列里有数」是两回事。
原先这几列写的是裸键名 `liked_count`,而作品行的指标是嵌在 `metrics` 里的,
于是导出来的表有「点赞/评论/收藏/分享」四列,**每一格都是空的**,还没人发现 ——
因为原来的测试只断言了 `作品ID`。
"""
@pytest.mark.asyncio
async def test_the_metric_columns_actually_contain_numbers(self, client):
from api.monitor.models import MonitorNoteMetric
async with monitor_db.get_session() as session:
from sqlalchemy import select
task_id = (await session.scalar(select(MonitorTask.id))).__int__()
session.add(
MonitorNoteMetric(
task_id=task_id, note_id="note-a", run_id=1,
captured_at=1_700_000_000_000,
liked_count=123, comment_count=45,
collected_count=6, share_count=7,
)
)
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
by_id = {row["作品ID"]: row for row in rows}
assert by_id["note-a"]["点赞"] == "123"
assert by_id["note-a"]["评论"] == "45"
assert by_id["note-a"]["收藏"] == "6"
assert by_id["note-a"]["分享"] == "7"
@pytest.mark.asyncio
async def test_a_missing_metric_is_left_empty_not_zero(self, client):
"""没采到的指标留空。写 0 的话,导出来的表会声称这条作品零互动。"""
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
assert rows[0]["点赞"] == ""
@pytest.mark.asyncio
async def test_the_export_says_who_the_creator_is(self, client):
"""一行只有作品 ID 没法用 —— 导出来是拿去比对和汇报的。
备注优先:昵称常常认不出是谁,而备注是人自己起的名字。
"""
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
assert {row["博主昵称"] for row in rows} == {"博主甲", "博主乙"}
@pytest.mark.asyncio
async def test_times_are_readable_not_raw_milliseconds(self, client):
"""毫秒时间戳倒进 CSV 就是 13 位数字,打开 Excel 的人没法看、也没法排序。"""
response = await client.get(
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
)
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
by_id = {row["作品ID"]: row for row in rows}
# 断言**形状**而不是具体时刻:格式化用的是服务器本地时区,写死一个字符串的话
# 换个时区的机器上就会红。
assert re.fullmatch(r"\d{4}-\d{2}-\d{2} \d{2}:\d{2}", by_id["note-a"]["发布时间"])
assert re.fullmatch(r"\d{4}-\d{2}-\d{2} \d{2}:\d{2}", by_id["note-a"]["首次发现"])
# 而且不能是原始毫秒。
assert by_id["note-a"]["发布时间"] != str(PUBLISHED_A)
+322
View File
@@ -0,0 +1,322 @@
# -*- coding: utf-8 -*-
"""作品栏里的**博主**:备注(他到底是谁)与账号级指标(他现在多大)。
按 creator_hash 分组、显示 creator_name,两样都认不出人:一个是哈希,一个是平台昵称。
备注是人自己起的名字。账号级指标则是作品列表给不了的东西 —— 作品说的是"这条涨了多少赞",
粉丝数说的是"这个人整个账号在涨还是在掉"。
"""
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import db as monitor_db
from api.monitor.models import (
MODE_CREATOR,
MonitorCreatorStat,
MonitorNote,
MonitorTask,
)
CREATOR_HASH = "hash-a"
NICKNAME = "张三"
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
async def _seed(platform: str = "xhs", task_name: str = "任务") -> int:
"""一个任务 + 一条作品,博主固定用 CREATOR_HASH / NICKNAME。"""
async with monitor_db.get_session() as session:
task = MonitorTask(
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
session.add(task)
await session.flush()
session.add(
MonitorNote(
task_id=task.id, note_id=f"{platform}-n1", title="作品",
note_url="", cover="", creator_hash=CREATOR_HASH,
creator_name=NICKNAME, source_kind="", published_at=None,
first_seen_run_id=1, first_seen_at=0,
last_seen_run_id=1, last_seen_at=0,
)
)
return task.id
async def _notes(client, platform: str = "xhs"):
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["notes"]
async def _set_alias(client, alias: str, platform: str = "xhs", creator_hash=CREATOR_HASH):
return await client.put(
f"/api/monitor/creators/{creator_hash}",
params={"platform": platform},
json={"alias": alias},
)
class TestCreatorAlias:
@pytest.mark.asyncio
async def test_notes_start_without_an_alias(self, client):
await _seed()
assert (await _notes(client))[0]["creator_alias"] == ""
@pytest.mark.asyncio
async def test_an_alias_comes_back_with_the_notes(self, client):
await _seed()
response = await _set_alias(client, "竞品A")
assert response.status_code == 200
note = (await _notes(client))[0]
assert note["creator_alias"] == "竞品A"
# 备注是**叠加**在昵称之上的,不是替换 —— 昵称仍然是有用的对照。
assert note["creator_name"] == NICKNAME
@pytest.mark.asyncio
async def test_an_alias_is_shared_across_tasks(self, client):
"""同一个博主出现在两个任务里,备注只该填一次。
creator_hash 对同一个 uid 是稳定的,所以键取 (platform, creator_hash) 而不是
按任务存 —— 否则每加一个任务都要重新认一遍人。
"""
await _seed(task_name="任务甲")
await _seed(task_name="任务乙")
await _set_alias(client, "竞品A")
for note in await _notes(client):
assert note["creator_alias"] == "竞品A"
@pytest.mark.asyncio
async def test_the_alias_does_not_leak_to_another_platform(self, client):
"""同一个哈希在另一个平台上是另一个(或同一个)人 —— 别串味。"""
await _seed(platform="xhs")
await _seed(platform="dy")
await _set_alias(client, "小红书那边的", platform="xhs")
assert (await _notes(client, "xhs"))[0]["creator_alias"] == "小红书那边的"
assert (await _notes(client, "dy"))[0]["creator_alias"] == ""
@pytest.mark.asyncio
async def test_an_empty_alias_clears_it(self, client):
await _seed()
await _set_alias(client, "竞品A")
await _set_alias(client, "")
assert (await _notes(client))[0]["creator_alias"] == ""
@pytest.mark.asyncio
async def test_an_alias_is_trimmed(self, client):
await _seed()
await _set_alias(client, " 竞品A ")
assert (await _notes(client))[0]["creator_alias"] == "竞品A"
async def _add_stat(
task_id: int,
run_id: int,
fans: int | None,
*,
total_favorited: int | None = 83000,
works: int | None = 42,
creator_hash: str = CREATOR_HASH,
captured_at: int = 1,
) -> None:
async with monitor_db.get_session() as session:
session.add(
MonitorCreatorStat(
task_id=task_id,
run_id=run_id,
creator_hash=creator_hash,
nickname=NICKNAME,
fans=fans,
total_favorited=total_favorited,
works_count=works,
following=7,
captured_at=captured_at,
)
)
class TestCreatorStats:
"""账号级指标跟着作品一起返回 —— 界面上是按博主归组的,为了一个组头再发一轮请求
没道理。"""
@pytest.mark.asyncio
async def test_without_snapshots_the_fields_are_null(self, client):
"""null 而不是 0:0 会显示成「粉丝 0」,而事实是"还没采到"。"""
await _seed()
note = (await _notes(client))[0]
assert note["creator_fans"] is None
assert note["creator_total_favorited"] is None
assert note["creator_works"] is None
@pytest.mark.asyncio
async def test_a_snapshot_shows_up_on_every_work_of_that_creator(self, client):
task_id = await _seed()
await _add_stat(task_id, run_id=1, fans=12000)
for note in await _notes(client):
assert note["creator_fans"] == 12000
assert note["creator_total_favorited"] == 83000
assert note["creator_works"] == 42
@pytest.mark.asyncio
async def test_the_latest_run_wins(self, client):
"""一轮一条,所以总会有好几条 —— 给界面的必须是最近那条。"""
task_id = await _seed()
await _add_stat(task_id, run_id=1, fans=12000)
await _add_stat(task_id, run_id=2, fans=12300)
assert (await _notes(client))[0]["creator_fans"] == 12300
@pytest.mark.asyncio
async def test_another_platforms_snapshot_does_not_leak(self, client):
"""两个平台上恰好同名同哈希的博主是两个人 —— 快照挂在任务上,不该串。"""
await _seed(platform="xhs")
dy_task_id = await _seed(platform="dy")
await _add_stat(dy_task_id, run_id=1, fans=999)
assert (await _notes(client, "xhs"))[0]["creator_fans"] is None
assert (await _notes(client, "dy"))[0]["creator_fans"] == 999
@pytest.mark.asyncio
async def test_an_unparsed_count_stays_null(self, client):
"""快照在,但某一项没解析出来 —— 那一项必须是 null,不能变成 0。
和上一条的区别:那条是"根本没有快照",这条是"有快照、其中一项平台没给"。
界面上两种都该是「—」。
"""
task_id = await _seed()
await _add_stat(
task_id, run_id=1, fans=None, total_favorited=None, works=None
)
note = (await _notes(client))[0]
assert note["creator_fans"] is None
assert note["creator_works"] is None
# 快照本身是有的(有采集时间),只是值不知道 —— 前端要能分开这两件事。
assert note["creator_stats_at"] is not None
async def _seed_task(platform: str = "xhs", task_name: str = "空任务") -> int:
"""只有任务,**一条作品都没有**。"""
async with monitor_db.get_session() as session:
task = MonitorTask(
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
session.add(task)
await session.flush()
return task.id
async def _creators(client, platform: str = "xhs"):
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["creators"]
class TestCreatorList:
"""作品栏要显示谁 —— **包括一条作品都没有的博主**。
组原先是从作品推出来的,于是「目标加了、资料采到了、粉丝数在库里,界面上什么都
没有」。而这类博主恰恰最该看见:还在涨粉,只是最近没发东西。
"""
@pytest.mark.asyncio
async def test_a_creator_with_no_works_still_shows_up(self, client):
"""**这条就是这个改动的全部理由。**"""
task_id = await _seed_task()
await _add_stat(task_id, run_id=1, fans=12000)
creators = await _creators(client)
assert len(creators) == 1
assert creators[0]["creator_hash"] == CREATOR_HASH
assert creators[0]["creator_name"] == NICKNAME # 只剩快照这一个来源
assert creators[0]["note_count"] == 0
assert creators[0]["creator_fans"] == 12000
@pytest.mark.asyncio
async def test_a_creator_with_works_carries_their_count(self, client):
task_id = await _seed()
await _add_stat(task_id, run_id=1, fans=12000)
creators = await _creators(client)
assert len(creators) == 1
assert creators[0]["note_count"] == 1
assert creators[0]["creator_name"] == NICKNAME
assert creators[0]["creator_fans"] == 12000
@pytest.mark.asyncio
async def test_a_creator_with_neither_works_nor_a_snapshot_is_absent(self, client):
"""两个来源都没有 = 我们对他一无所知,不该凭空造一个组出来。"""
await _seed_task()
assert await _creators(client) == []
@pytest.mark.asyncio
async def test_a_work_without_a_snapshot_still_lists_its_creator(self, client):
"""小红书那条路不产生账号快照 —— 那边只能靠作品认出人来。"""
await _seed()
creators = await _creators(client)
assert len(creators) == 1
assert creators[0]["note_count"] == 1
assert creators[0]["creator_fans"] is None
@pytest.mark.asyncio
async def test_the_creator_remark_comes_along(self, client):
"""没有作品的博主也要能起备注 —— 否则「这是谁」在最需要的时候认不出来。"""
task_id = await _seed_task()
await _add_stat(task_id, run_id=1, fans=12000)
await _set_alias(client, "竞品A")
assert (await _creators(client))[0]["creator_alias"] == "竞品A"
@pytest.mark.asyncio
async def test_platforms_stay_apart(self, client):
dy_task = await _seed_task(platform="dy")
await _add_stat(dy_task, run_id=1, fans=999)
assert await _creators(client, "xhs") == []
assert (await _creators(client, "dy"))[0]["creator_fans"] == 999
@pytest.mark.asyncio
async def test_the_same_creator_under_two_tasks_is_two_rows_the_ui_merges(self, client):
"""服务端按 任务×博主 给(快照就是那么存的),合并交给界面 —— 因为备注跨任务
是同一条,合并之后的组才是人眼里的「一个博主」。"""
first = await _seed(task_name="任务甲")
second = await _seed(task_name="任务乙")
await _add_stat(first, run_id=1, fans=12000)
creators = await _creators(client)
assert len(creators) == 2
assert {row["creator_hash"] for row in creators} == {CREATOR_HASH}
assert sum(row["note_count"] for row in creators) == 2
+454 -3
View File
@@ -24,6 +24,7 @@ posted/seen comment split, idempotency, and the silent-cookie-failure signal.
"""
import json
from datetime import date
from pathlib import Path
from typing import Any, Dict, List, Optional
@@ -35,7 +36,13 @@ from sqlalchemy.pool import StaticPool
from tools.time_util import get_current_timestamp
from api.monitor.ingest import describe_exit_code, ingest_run, parse_count
from api.monitor import adapters
from api.monitor.ingest import (
describe_exit_code,
diagnose_failure,
ingest_run,
parse_count,
)
from api.monitor.models import (
EVENT_AUTH_FAILURE,
EVENT_METRIC_DELTA,
@@ -46,6 +53,8 @@ from api.monitor.models import (
EVENT_RUN_FAILED,
MODE_CREATOR,
MonitorBase,
MonitorComment,
MonitorCreatorStat,
MonitorEvent,
MonitorNote,
MonitorNoteMetric,
@@ -118,9 +127,18 @@ def _write_run_dir(
root: Path,
notes: List[Dict[str, Any]],
comments: Optional[List[Dict[str, Any]]] = None,
subdir: str = "xhs",
profiles: Optional[List[Dict[str, Any]]] = None,
) -> Path:
"""Write a run's jsonl output in the crawler's own layout."""
jsonl_dir = root / "xhs" / "jsonl"
"""Write a run's jsonl output in the crawler's own layout.
``subdir`` 是**爬虫**落盘的目录名,不是监控层的平台 id —— 抖音那边这两者不同
(平台 id 是 ``dy``、目录是 ``douyin``),所以必须能分开指定,否则测不出那个差异。
``profiles`` 为 None 时**不写**这个文件(小红书那条路根本不产生它),给列表时写
——包括空列表,那是「问了但没问到」。
"""
jsonl_dir = root / subdir / "jsonl"
jsonl_dir.mkdir(parents=True, exist_ok=True)
contents = jsonl_dir / "creator_contents_2026-01-01.jsonl"
@@ -134,6 +152,12 @@ def _write_run_dir(
"\n".join(json.dumps(c, ensure_ascii=False) for c in comments),
encoding="utf-8",
)
if profiles is not None:
profile_file = jsonl_dir / "creator_profile_2026-01-01.jsonl"
profile_file.write_text(
"\n".join(json.dumps(p, ensure_ascii=False) for p in profiles),
encoding="utf-8",
)
return root
@@ -170,6 +194,52 @@ def _comment(comment_id: str, note_id: str, create_time: int, **extra) -> Dict[s
return record
def _dy_note(aweme_id: str, liked: Any = "10", **extra) -> Dict[str, Any]:
"""抖音作品记录 —— 键名照抄 store/douyin/__init__.py 的落盘字段。
重点在于**没有** ``note_id``:抖音叫 ``aweme_id``。这一条差异没映射好,就是
每条记录都被 ingest 悄悄 continue 掉、一条不剩。
"""
record = {
"aweme_id": aweme_id,
"aweme_type": "0",
"title": f"title-{aweme_id}",
"desc": f"title-{aweme_id}",
# 抖音给的是**秒**(实测 1790574515 = 2026-09-28),小红书给毫秒。落库统一
# 换算成毫秒,这个 fixture 必须照真实形态写,否则测不出单位问题。
"create_time": 1790574515,
"creator_hash": "hash",
"nickname": "u***r",
"liked_count": liked,
"collected_count": "1",
"comment_count": "1",
"share_count": "1",
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
"cover_url": "https://img/cover.jpg",
}
record.update(extra)
return record
def _dy_comment(
comment_id: str, aweme_id: str, create_time: int, **extra
) -> Dict[str, Any]:
record = {
"comment_id": comment_id,
"create_time": create_time,
"aweme_id": aweme_id,
"content": f"content-{comment_id}",
"creator_hash": "hash",
"nickname": "u***r",
"sub_comment_count": "0",
"like_count": "0",
# 抖音顶层评论的父 id 是字符串 "0",不是空串。
"parent_comment_id": "0",
}
record.update(extra)
return record
async def _events(db: AsyncSession, event_type: Optional[str] = None) -> List[MonitorEvent]:
stmt = select(MonitorEvent)
if event_type:
@@ -525,3 +595,384 @@ class TestIdempotency:
assert result.new_notes == 0
assert result.new_comments == 0
assert len(list((await db.scalars(select(MonitorNote))).all())) == notes_after_first
class TestDouyinIngest:
"""抖音的产物形状与小红书不同 —— 这里钉住「不会被静默丢掉」。
这一组存在的理由,是这个改动最危险的失败模式:字段名或目录名没对上时,ingest
不报错,只是**一条都不入库**,然后被当成「疑似登录失效」报出去。
"""
async def _ingest(
self,
db,
tmp_path,
notes,
comments=None,
platform="dy",
subdir="douyin",
):
task = await _make_task(db, platform=platform)
run = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, notes, comments, subdir=subdir)
result = await ingest_run(db, run, task, tmp_path)
return task, run, result
@pytest.mark.asyncio
async def test_notes_are_ingested_under_their_douyin_field_names(self, db, tmp_path):
aweme_id = "7525082444551310602"
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note(aweme_id)])
note = await db.scalar(select(MonitorNote))
assert note is not None, "抖音作品被静默丢弃了 —— 多半是 aweme_id 没映射到 note_id"
assert note.note_id == aweme_id
assert note.note_url == f"https://www.douyin.com/video/{aweme_id}"
assert note.cover == "https://img/cover.jpg"
assert note.source_kind == "0"
# 秒 → 毫秒,换算过才对。
assert note.published_at == 1790574515 * 1000
assert result.notes_fetched == 1
@pytest.mark.asyncio
async def test_timestamps_are_normalised_to_milliseconds(self, db, tmp_path):
"""抖音的时间戳是**秒**,小红书是毫秒 —— 差 1000 倍,必须换算。
不换算的话,2026 年的作品会显示成 1970 年。这是实测踩到的:抖音作品的
「发布日期」列显示成 1970-01-22(1790574515 被当成毫秒就是 21 天后)。
"""
aweme_id = "7525082444551310602"
await self._ingest(db, tmp_path, [_dy_note(aweme_id)])
note = await db.scalar(select(MonitorNote))
assert note.published_at == 1790574515 * 1000
# 落库的是毫秒,展示层才不用关心来源;乘完应该在 2026 年,不是 1970。
assert date.fromtimestamp(note.published_at / 1000).year == 2026
@pytest.mark.asyncio
async def test_the_artifact_directory_is_not_the_platform_id(self, db, tmp_path):
"""目录名与平台 id 不一致,是这套适配里最反直觉的一条。
抖音的平台 id 是 ``dy``,而爬虫把产物写在 ``douyin/`` 下。把它钉在这里,
是为了让「顺手改成一致」这件事会在测试里红掉,而不是让 ingest 悄悄读 0 条。
"""
assert adapters.artifact_dir("dy") == "douyin"
@pytest.mark.asyncio
async def test_writing_into_the_platform_id_directory_reads_nothing(self, db, tmp_path):
"""反面:产物落在 ``dy/`` 下时一条都读不到 —— 这正是映射要解决的问题。"""
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note("1")], subdir="dy")
assert result.notes_fetched == 0
@pytest.mark.asyncio
async def test_misplaced_output_is_blamed_on_the_directory_not_the_login(
self, db, tmp_path
):
"""产物其实抓到了,只是目录名不对 —— 不该报成「疑似登录失效」。
这是最难查的一类故障:登录是好的、数据也抓到了,但报出来的现象和登录失效
一模一样,会把人指去查完全错误的方向。
"""
_task, run, _result = await self._ingest(
db, tmp_path, [_dy_note("1")], subdir="dy"
)
assert any("目录" in event.title for event in await _events(db, EVENT_NO_DATA))
assert await _events(db, EVENT_AUTH_FAILURE) == []
assert run.error_message and "dy" in run.error_message
@pytest.mark.asyncio
async def test_comments_are_linked_through_aweme_id(self, db, tmp_path):
aweme_id = "7525082444551310602"
_task, _run, result = await self._ingest(
db,
tmp_path,
[_dy_note(aweme_id)],
comments=[_dy_comment("c1", aweme_id, 500)],
)
comment = await db.scalar(select(MonitorComment))
assert comment is not None, "抖音评论被静默丢弃了 —— 多半是 aweme_id 没映射"
assert comment.note_id == aweme_id
assert result.comments_fetched == 1
@pytest.mark.asyncio
async def test_a_top_level_parent_of_zero_becomes_empty(self, db, tmp_path):
"""抖音顶层评论的父 id 是 "0";原样存进去,前端会多出一堆悬空的父节点。"""
aweme_id = "7525082444551310602"
await self._ingest(
db,
tmp_path,
[_dy_note(aweme_id)],
comments=[
_dy_comment("c1", aweme_id, 500),
_dy_comment("c2", aweme_id, 600, parent_comment_id="c1"),
],
)
by_id = {c.comment_id: c for c in (await db.scalars(select(MonitorComment))).all()}
assert by_id["c1"].parent_comment_id == ""
assert by_id["c2"].parent_comment_id == "c1"
@pytest.mark.asyncio
async def test_the_four_metrics_need_no_mapping(self, db, tmp_path):
"""四个指标键两边同名 —— 抖音作品照样进 monitor_note_metric,差分照常。"""
aweme_id = "7525082444551310602"
task = await _make_task(db, platform="dy")
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="100")], subdir="douyin")
run1 = await _make_run(db, task, started_at=1)
await ingest_run(db, run1, task, tmp_path)
metric = await db.scalar(select(MonitorNoteMetric))
assert metric is not None and metric.liked_count == 100
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="150")], subdir="douyin")
run2 = await _make_run(db, task, started_at=2)
await ingest_run(db, run2, task, tmp_path)
assert len(await _events(db, EVENT_METRIC_DELTA)) == 1
class TestDuplicateRecordsInOneRun:
"""一轮产物里重复出现的作品只该算一次。
指标快照的唯一键是 ``(task_id, note_id, run_id)``:同一条作品在一轮里进来两次,
第二次插入会撞键并让整个 run 崩掉 —— 产物里重复并不罕见(多个目标指向同一个人、
或退化路径重复刷新)。
"""
@pytest.mark.asyncio
async def test_a_duplicated_note_is_processed_once(self, db, tmp_path):
task = await _make_task(db)
_write_run_dir(tmp_path, [_note("n1"), _note("n1")])
run = await _make_run(db, task, started_at=1)
result = await ingest_run(db, run, task, tmp_path)
assert result.new_notes == 1
assert len(list((await db.scalars(select(MonitorNote))).all())) == 1
# 崩就崩在这一句上:一条作品只能有一份本轮快照。
assert len(list((await db.scalars(select(MonitorNoteMetric))).all())) == 1
class TestNicknameRefresh:
"""已入库的评论,昵称要跟着重新采集的值走。
评论是去重后直接跳过的,若不刷新,脱敏开关一改(或评论者改了昵称),老数据就永远
停在旧值上 —— 而重采是唯一能拿到新值的途径。
"""
@pytest.mark.asyncio
async def test_an_existing_comment_gets_its_nickname_refreshed(self, db, tmp_path):
task = await _make_task(db)
_write_run_dir(tmp_path, [_note("n1")], comments=[_comment("c1", "n1", 500)])
run1 = await _make_run(db, task, started_at=1)
await ingest_run(db, run1, task, tmp_path)
assert (await db.scalar(select(MonitorComment))).nickname == "u***r"
_write_run_dir(
tmp_path,
[_note("n1")],
comments=[_comment("c1", "n1", 500, nickname="未脱敏的新昵称")],
)
run2 = await _make_run(db, task, started_at=2)
await ingest_run(db, run2, task, tmp_path)
comment = await db.scalar(select(MonitorComment))
assert comment.nickname == "未脱敏的新昵称"
# 去重的语义没变:同一条评论不该被插成两行。
assert (
len(list((await db.scalars(select(MonitorComment))).all())) == 1
)
class TestFailureDiagnosis:
"""失败原因要能被人看懂。
只写「退出码 1」等于什么都没说 —— 真正的报错埋在子进程的 stderr 里,而运行历史
里那一格显示的正是 run.error_message。
"""
TAIL = [
"2026-10-10 15:18:34 MediaCrawler INFO (core.py:385) - [DouYinCrawler] CDP浏览器信息",
"Traceback (most recent call last):",
' File "/app/main.py", line 114, in main',
" await crawler.start()",
"media_platform.douyin.exception.DataFetchError: account blocked, ",
]
def test_the_exception_line_is_picked_out_of_the_tail(self):
assert (
diagnose_failure(self.TAIL)
== "media_platform.douyin.exception.DataFetchError: account blocked,"
)
def test_the_managers_own_lines_are_not_mistaken_for_the_cause(self):
"""管理器自己补的那两句不是爬虫的报错,别被当成失败原因。"""
assert diagnose_failure(["Crawler exited with code: 1"]) is None
assert diagnose_failure(["Crawler completed successfully"]) is None
def test_nothing_to_say_is_not_an_error(self):
assert diagnose_failure(None) is None
assert diagnose_failure([]) is None
def test_it_falls_back_to_the_last_line(self):
assert (
diagnose_failure(["started fine", "then something odd"])
== "then something odd"
)
@pytest.mark.asyncio
async def test_a_failed_run_records_both_the_code_and_the_cause(self, db, tmp_path):
task = await _make_task(db)
run = await _make_run(db, task, started_at=1, exit_code=1)
result = await ingest_run(db, run, task, tmp_path, output_tail=self.TAIL)
assert result.status == RUN_FAILED
# 退出码和真因都要在,缺一个都还得去翻日志。
assert "code 1" in run.error_message
assert "account blocked" in run.error_message
events = await _events(db, EVENT_RUN_FAILED)
assert "account blocked" in events[0].title
# --------------------------------------------------------------------------
# 博主账号级指标(粉丝 / 总获赞 / 作品数)
# --------------------------------------------------------------------------
def _profile(creator_hash: str = "hash", **extra) -> Dict[str, Any]:
"""``creator_profile_*.jsonl`` 里的一行 —— 形状由 douyin_api.author_profile 决定。"""
record: Dict[str, Any] = {
"creator_hash": creator_hash,
"nickname": "博主",
"unique_id": "abc",
"fans": 12000,
"total_favorited": 83000,
"works": 42,
"following": 7,
}
record.update(extra)
return record
class TestCreatorStatSnapshots:
"""**账号级**指标和作品级指标是两回事:后者说"这条视频涨了多少赞",前者说
"这个人整个账号的粉丝在涨还是在掉"。作品列表给不了后者,所以单独存一张表。
"""
async def _ingest(self, db, tmp_path, notes, profiles, platform="dy", subdir="douyin"):
task = await _make_task(db, platform=platform)
run = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, notes, comments=[], subdir=subdir, profiles=profiles)
result = await ingest_run(db, run, task, tmp_path)
return task, run, result
@pytest.mark.asyncio
async def test_a_profile_becomes_a_snapshot(self, db, tmp_path):
task, run, _result = await self._ingest(db, tmp_path, [_dy_note("1")], [_profile()])
stat = await db.scalar(select(MonitorCreatorStat))
assert stat is not None
assert (stat.task_id, stat.run_id) == (task.id, run.id)
assert stat.creator_hash == "hash"
assert stat.nickname == "博主"
assert (stat.fans, stat.total_favorited, stat.works_count) == (12000, 83000, 42)
@pytest.mark.asyncio
async def test_a_missing_count_stays_null_not_zero(self, db, tmp_path):
"""0 是真实值(掉到零),null 是不知道。混起来趋势图就是在撒谎。"""
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(fans=None)])
stat = await db.scalar(select(MonitorCreatorStat))
assert stat.fans is None
# 同一个博主其它字段照常。
assert stat.works_count == 42
@pytest.mark.asyncio
async def test_abbreviated_counts_are_parsed(self, db, tmp_path):
"""走的是和作品指标同一个 parse_count —— 平台给你「1.2万」也得认。"""
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(fans="1.2万")])
assert (await db.scalar(select(MonitorCreatorStat))).fans == 12000
@pytest.mark.asyncio
async def test_the_same_creator_twice_in_one_run_yields_one_snapshot(self, db, tmp_path):
"""一个任务可以配多个目标,退化路径下它们可能落在同一个博主身上。
唯一键是 (task_id, creator_hash, run_id) —— 重复插入会撞键,把整轮炸掉。
(和作品重复那次是同一类事故。)
"""
await self._ingest(
db, tmp_path, [_dy_note("1")], [_profile(), _profile(nickname="另一条")]
)
stats = list((await db.scalars(select(MonitorCreatorStat))).all())
assert len(stats) == 1
@pytest.mark.asyncio
async def test_a_profile_without_a_hash_is_skipped(self, db, tmp_path):
"""哈希都算不出来,这条快照谁也查不到,落下去只是垃圾。"""
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(creator_hash="")])
assert (await db.scalars(select(MonitorCreatorStat))).all() == []
@pytest.mark.asyncio
async def test_no_profile_file_is_fine(self, db, tmp_path):
"""小红书那条路(爬虫进程)根本不产生这个文件 —— 不能因此报错。"""
task = await _make_task(db) # xhs
run = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, [_note("n1")], comments=[]) # 没有 profiles 参数
result = await ingest_run(db, run, task, tmp_path)
assert result.status == RUN_SUCCESS
assert (await db.scalars(select(MonitorCreatorStat))).all() == []
@pytest.mark.asyncio
async def test_the_stats_are_kept_even_when_no_works_were_fetched(self, db, tmp_path):
"""**这条是这里最值得留的一个。**
作品列表被风控挡住时,这一轮一条作品都拿不到、run 会被判成失败。但博主的粉丝数
并不因为这件事就不存在 —— 「粉丝还在涨,但新作品没在发现」恰恰是最该看见的时刻。
快照要是挂在「作品采到了」后面,就正好在最需要它的那一轮丢掉。
"""
_task, run, result = await self._ingest(db, tmp_path, [], [_profile()])
assert result.status == RUN_PARTIAL # 一条作品都没有,这轮确实不算成功
assert run.status == RUN_PARTIAL
stat = await db.scalar(select(MonitorCreatorStat))
assert stat is not None and stat.fans == 12000
@pytest.mark.asyncio
async def test_each_run_adds_its_own_snapshot(self, db, tmp_path):
"""趋势靠的就是这个:一条一轮,不要覆盖。"""
task = await _make_task(db, platform="dy")
first = await _make_run(db, task, started_at=1)
_write_run_dir(tmp_path, [_dy_note("1")], comments=[], subdir="douyin",
profiles=[_profile(fans=12000)])
await ingest_run(db, first, task, tmp_path)
second = await _make_run(db, task, started_at=2000)
_write_run_dir(tmp_path, [_dy_note("1")], comments=[], subdir="douyin",
profiles=[_profile(fans=12300)])
await ingest_run(db, second, task, tmp_path)
stats = list(
(
await db.scalars(
select(MonitorCreatorStat).order_by(MonitorCreatorStat.run_id)
)
).all()
)
assert [s.fans for s in stats] == [12000, 12300]
+146
View File
@@ -0,0 +1,146 @@
# -*- coding: utf-8 -*-
"""作品备注 —— 一个博主底下,哪几条是真正要盯的。
和博主备注(test_monitor_creators.py)是一对,但回答的不是同一个问题:博主备注回答
「这个账号是谁」,作品备注回答「这条作品我要盯着」。一个博主底下常常只有一两件值得
盯的作品,所以不能合成一条。
"""
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import db as monitor_db
from api.monitor.models import MODE_CREATOR, MonitorNote, MonitorTask
NOTE_ID = "note-a"
TITLE = "中秋哪儿都堵"
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
async def _seed(platform: str = "xhs", task_name: str = "任务", note_id: str = NOTE_ID) -> int:
async with monitor_db.get_session() as session:
task = MonitorTask(
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
)
session.add(task)
await session.flush()
session.add(
MonitorNote(
task_id=task.id, note_id=note_id, title=TITLE,
note_url="", cover="", creator_hash="hash-a",
creator_name="张三", source_kind="", published_at=None,
first_seen_run_id=1, first_seen_at=0,
last_seen_run_id=1, last_seen_at=0,
)
)
return task.id
async def _notes(client, platform: str = "xhs"):
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["notes"]
async def _set_alias(client, alias: str, platform: str = "xhs", note_id: str = NOTE_ID):
return await client.put(
f"/api/monitor/notes/{note_id}",
params={"platform": platform},
json={"alias": alias},
)
class TestNoteAlias:
@pytest.mark.asyncio
async def test_notes_start_without_a_remark(self, client):
await _seed()
assert (await _notes(client))[0]["note_alias"] == ""
@pytest.mark.asyncio
async def test_a_remark_comes_back_with_the_notes(self, client):
await _seed()
response = await _set_alias(client, "重点")
assert response.status_code == 200
note = (await _notes(client))[0]
assert note["note_alias"] == "重点"
# 备注是**叠加**在标题之上的,不是替换 —— 标题仍然是这条作品本身。
assert note["title"] == TITLE
@pytest.mark.asyncio
async def test_a_remark_is_shared_across_tasks(self, client):
"""同一件作品被两个任务都监控时,备注只该填一次。"""
await _seed(task_name="任务甲")
await _seed(task_name="任务乙")
await _set_alias(client, "重点")
for note in await _notes(client):
assert note["note_alias"] == "重点"
@pytest.mark.asyncio
async def test_a_remark_does_not_leak_to_another_platform(self, client):
"""作品的 id 是平台各自的编号体系 —— 抖音的 123 和小红书的 123 是两条作品。"""
await _seed(platform="xhs")
await _seed(platform="dy")
await _set_alias(client, "小红书那边的", platform="xhs")
assert (await _notes(client, "xhs"))[0]["note_alias"] == "小红书那边的"
assert (await _notes(client, "dy"))[0]["note_alias"] == ""
@pytest.mark.asyncio
async def test_a_remark_does_not_leak_to_another_work(self, client):
"""钉住这里的**作用域**:键是 note_id。写错成按任务存的话,给一条起了备注,
同一个博主底下的其它作品会跟着一起变 —— 那这个功能就没用了。"""
await _seed(note_id="note-a")
await _seed(task_name="另一个任务", note_id="note-b")
await _set_alias(client, "重点", note_id="note-a")
by_id = {note["note_id"]: note["note_alias"] for note in await _notes(client)}
assert by_id == {"note-a": "重点", "note-b": ""}
@pytest.mark.asyncio
async def test_an_empty_remark_clears_it(self, client):
await _seed()
await _set_alias(client, "重点")
await _set_alias(client, "")
assert (await _notes(client))[0]["note_alias"] == ""
@pytest.mark.asyncio
async def test_a_remark_is_trimmed(self, client):
await _seed()
await _set_alias(client, " 重点 ")
assert (await _notes(client))[0]["note_alias"] == "重点"
@pytest.mark.asyncio
async def test_the_two_kinds_of_remark_stay_apart(self, client):
"""博主备注和作品备注是两张表、两个键 —— 一个不该把另一个盖掉。"""
await _seed()
await _set_alias(client, "重点")
await client.put(
"/api/monitor/creators/hash-a", params={"platform": "xhs"}, json={"alias": "竞品A"}
)
note = (await _notes(client))[0]
assert note["note_alias"] == "重点"
assert note["creator_alias"] == "竞品A"
+16
View File
@@ -118,6 +118,22 @@ class TestBuildRunMessage:
assert "标题A" in message
assert "https://www.xiaohongshu.com/explore/abc123" in message
@pytest.mark.asyncio
async def test_douyin_notes_link_to_douyin(self, db):
"""链接形状按平台走 —— 群里点进去该是能看的作品,不是 404。"""
task, run = await _seed(db)
task.platform = "dy"
_add_event(
db, task, run, EVENT_NEW_NOTE, "新作品:标题A",
payload={"note_id": "7525082444551310602", "title": "标题A"},
)
await db.flush()
message = await notify.build_run_message(db, task, run)
assert "https://www.douyin.com/video/7525082444551310602" in message
assert "xiaohongshu.com" not in message
@pytest.mark.asyncio
async def test_long_note_lists_are_truncated(self, db):
"""A first run can find dozens; a wall of text is worse than a count."""
+21 -2
View File
@@ -184,9 +184,9 @@ async def db():
await engine.dispose()
async def _seed_task(db: AsyncSession, name: str) -> MonitorTask:
async def _seed_task(db: AsyncSession, name: str, platform: str = "xhs") -> MonitorTask:
task = MonitorTask(
name=name, platform="xhs", mode=MODE_CREATOR, enabled=True,
name=name, platform=platform, mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=True,
max_comments_count=50, run_timeout_seconds=3600,
notify_enabled=False, created_at=0, updated_at=0,
@@ -263,6 +263,25 @@ class TestBuildReport:
assert result["totals"]["liked_count_delta"] == 30
assert result["task_ids"] is None
@pytest.mark.asyncio
async def test_an_empty_selection_is_not_the_same_as_no_filter(self, db):
"""空列表 ≠ 不限制。
``_resolve_scope`` 在「这个平台一个任务都没有」时返回**空列表**。如果按真值
处理(``if task_ids``),报表就会退化成「不限制平台」,把**所有**任务的数据
聚合进来 —— 现象就是切到抖音,报表里却全是小红书的数据。
"""
other = await _seed_task(db, "xhs task", platform="xhs")
await _seed_note_with_metrics(db, other, "n1", [(_ms(2026, 1, 10, 10), 999)])
await db.commit()
result = await build_report(db, [], date(2026, 1, 10), date(2026, 1, 10))
assert result["totals"]["liked_count_delta"] == 0
assert result["note_count"] == 0
# 空列表要原样透出去;None 在 API 里的意思是「全部任务」,两者不能混。
assert result["task_ids"] == []
@pytest.mark.asyncio
async def test_baseline_from_before_the_range_is_used(self, db):
"""Growth is measured against the last value before the window opens."""
+110
View File
@@ -0,0 +1,110 @@
# -*- coding: utf-8 -*-
"""monitor runner —— 尤其是抖音那条(不走子进程的)路的运行状态流转。
这条路的地位特殊:它不经过 ``crawler_manager``,所以爬虫那套「退出码 / 日志尾巴」的
约定它一个都不沾。凡是写在那里面的东西,这条路都得单独有一份。
"""
import asyncio
import pytest
import pytest_asyncio
from sqlalchemy import select
from tools.time_util import get_current_timestamp
from api.monitor import db as monitor_db
from api.monitor import runner as runner_module
from api.monitor.models import (
MODE_CREATOR,
RUN_RUNNING,
MonitorRun,
MonitorTarget,
MonitorTask,
)
@pytest_asyncio.fixture
async def db(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
yield monitor_db
await monitor_db.dispose_engine()
async def _make_douyin_task() -> int:
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="dy", platform="dy", mode=MODE_CREATOR, enabled=True,
interval_minutes=360, max_notes_count=20, enable_comments=False,
max_comments_count=20, run_timeout_seconds=3600,
notify_enabled=False, notify_failures=False,
created_at=now, updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorTarget(
task_id=task.id, kind=MODE_CREATOR, external_id="MS4w-sec",
xsec_token="", xsec_source="", raw_value="MS4w-sec",
label="x", enabled=True, created_at=now,
)
)
return task.id
class TestDouyinRunStatus:
@pytest.mark.asyncio
async def test_the_run_is_marked_running_before_collecting(self, db, monkeypatch):
"""**采集开始之前**,run 就必须已经是 running。
这一行原先只写在爬虫那条分支里,于是抖音路上 run 一直停在 pending —— 一旦中途
出事(异常、或进程被重启),界面上就是一个永远「排队中」的幽灵,而且 recover()
当时也只收 running、够不着它。
"""
task_id = await _make_douyin_task()
seen = {}
async def fake_collect(out_dir, **kwargs):
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun).order_by(MonitorRun.id))
seen["status"] = run.status
return {
"notes": 0,
"comments": 0,
"errors": ["故意失败"],
"jsonl_dir": str(out_dir),
}
monkeypatch.setattr(runner_module.douyin_fetch, "collect", fake_collect)
await runner_module.execute_task(task_id, trigger="manual")
assert seen["status"] == RUN_RUNNING
@pytest.mark.asyncio
async def test_a_hanging_collect_does_not_leave_the_run_running(self, db, monkeypatch):
"""进程内那条路也要有超时。
爬虫那条靠 ``run_and_wait(timeout=...)`` 兜底,这条路没有子进程、没人管 ——
里面任何一次卡住(实测过 ``page.evaluate`` 打在一个卡死的标签页上不返回)都会让
run 永远停在「运行中」,界面上看起来就是任务卡死了。
"""
task_id = await _make_douyin_task()
async with monitor_db.get_session() as session:
task = await session.get(MonitorTask, task_id)
task.run_timeout_seconds = 1 # 把超时压到 1 秒,别让测试真等
async def hanging_collect(out_dir, **kwargs):
await asyncio.sleep(60)
raise AssertionError("不该走到这里")
monkeypatch.setattr(runner_module.douyin_fetch, "collect", hanging_collect)
await runner_module.execute_task(task_id, trigger="manual")
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun).order_by(MonitorRun.id))
assert run.status != RUN_RUNNING
assert "超时" in (run.error_message or "") or "超过" in (run.error_message or "")
+77 -3
View File
@@ -30,11 +30,12 @@ from api.monitor.models import (
MonitorTarget,
MonitorTask,
RUN_INTERRUPTED,
RUN_PENDING,
RUN_RUNNING,
RUN_SUCCESS,
)
from api.monitor.scheduler import MonitorScheduler
from api.monitor.settings import set_cookie
from api.monitor.settings import set_cookie, set_setting
from tools.time_util import get_current_timestamp
MS_PER_MINUTE = 60_000
@@ -72,12 +73,14 @@ async def executed(monkeypatch):
return calls
async def _make_task(next_run_at, enabled: bool = True, interval: int = 60) -> int:
async def _make_task(
next_run_at, enabled: bool = True, interval: int = 60, platform: str = "xhs"
) -> int:
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="t",
platform="xhs",
platform=platform,
mode=MODE_CREATOR,
enabled=enabled,
interval_minutes=interval,
@@ -145,6 +148,44 @@ class TestFiring:
assert executed == []
@pytest.mark.asyncio
async def test_the_cookie_gate_reads_the_tasks_own_platform(
self, monkeypatch, db, executed
):
"""cookie 闸门要按任务自己的平台取。
以前这里是 ``get_cookie(session)``(默认小红书)—— 只有小红书时看不出问题,
接上抖音后,抖音任务会因为读的是小红书那份 cookie 而永远不被触发,且不报错。
"""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
async with monitor_db.get_session() as session:
await set_cookie(session, "sessionid=dy-secret", "dy")
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
await MonitorScheduler().tick()
assert executed == [(task_id, "scheduled")]
@pytest.mark.asyncio
async def test_cdp_mode_frees_a_task_from_the_cookie_gate(
self, monkeypatch, db, executed
):
"""开着 CDP 时不该再要求先粘 cookie。
CDP 模式下登录态来自被接管的那台浏览器,粘不粘 cookie 都由不得它 —— 不放行的话,
选了「接管已有 Chrome」却没粘 cookie 的用户会发现任务永远不跑,而且什么错都不报。
"""
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
async with monitor_db.get_session() as session:
await set_setting(session, "system.cdp_enabled", "true")
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
await MonitorScheduler().tick()
assert executed == [(task_id, "scheduled")]
@pytest.mark.asyncio
async def test_long_outage_coalesces_into_one_run(self, monkeypatch, db, executed):
"""A missed schedule fires once, not once per missed interval."""
@@ -239,6 +280,39 @@ class TestRecovery:
assert run.status == RUN_INTERRUPTED
assert run.finished_at is not None
@pytest.mark.asyncio
async def test_pending_runs_are_also_cleaned_up(self, db):
"""挂在 ``pending`` 的 run 同样是残留,必须一起收。
那一行是上一轮建的,可它后面的采集根本没机会开始(进程被重启,或采集那条路抛了
异常)。只清 ``running`` 的话,它会永远挂在界面上显示「排队中」——
用户看到的就是任务卡死了。
"""
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="t", platform="dy", mode=MODE_CREATOR, enabled=True,
interval_minutes=60, max_notes_count=20, enable_comments=False,
max_comments_count=50, run_timeout_seconds=3600,
next_run_at=now, last_status="pending", created_at=now, updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorRun(
task_id=task.id, trigger="manual", status=RUN_PENDING,
phase=MODE_CREATOR, save_data_path="", queued_at=now, not_before=0,
max_comments_count=50,
)
)
await MonitorScheduler().recover()
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun))
assert run.status == RUN_INTERRUPTED
assert run.finished_at is not None
@pytest.mark.asyncio
async def test_completed_runs_are_left_alone(self, db):
async with monitor_db.get_session() as session:
+24 -2
View File
@@ -14,6 +14,8 @@ import pathlib
import pytest
import config
ROOT = pathlib.Path(__file__).resolve().parent.parent
# 统一的禁用字段名(键)。昵称字段(nickname/user_nickname/screen_name/name/user_name)允许保留(值需脱敏)。
@@ -28,6 +30,17 @@ NICK_KEYS = {"nickname", "user_nickname", "screen_name", "name", "user_name"}
MASK_RE = re.compile(r"^.?\*{1,4}.?$")
@pytest.fixture(autouse=True)
def _force_nickname_masking(monkeypatch):
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
所以这里显式打开来测;开关两个方向的行为由 test_mask_and_hash_tools 覆盖。
"""
monkeypatch.setattr(config, "MASK_NICKNAME", True)
# ----------------------------- ORM 自省 -----------------------------
def test_orm_has_no_forbidden_columns():
@@ -84,17 +97,26 @@ def _check_nickname_masked(d: dict, raw: str, label: str):
assert MASK_RE.match(val) or "*" in val, f"[{label}] {k} 未脱敏: {val}"
def test_mask_and_hash_tools():
def test_mask_and_hash_tools(monkeypatch):
from tools.user_hash import anonymize_user_id, mask_nickname
h = anonymize_user_id("12345")
assert h and h != "12345" and re.fullmatch(r"[0-9a-f]{16}", h)
assert anonymize_user_id(None) == "" and anonymize_user_id("") == ""
# 昵称脱敏:首尾留1字、中间星号,且不等于原文
# 开关打开:首尾留 1 字、中间星号,且不等于原文。
monkeypatch.setattr(config, "MASK_NICKNAME", True)
assert mask_nickname("张三丰") != "张三丰"
assert "*" in mask_nickname("张三丰")
assert mask_nickname(None) == ""
assert mask_nickname("a") == "*"
# 开关关闭(本仓库的部署配置):原样返回。脱敏是有损的 —— 「张三」和「张四」
# 都会变成「张*」,而分清谁是谁正是监控这一层要干的事。
monkeypatch.setattr(config, "MASK_NICKNAME", False)
assert mask_nickname("张三丰") == "张三丰"
assert mask_nickname("a") == "a"
assert mask_nickname(None) == ""
def test_xhs_note_extraction_masks_user_info():
import asyncio
+99 -10
View File
@@ -21,12 +21,15 @@
import httpx
import pytest
import pytest_asyncio
from sqlalchemy import text
from sqlalchemy import select, text
from tools.time_util import get_current_timestamp
from api.main import app
from api.monitor import adapters
from api.monitor import db as monitor_db
from api.monitor import platforms
from api.monitor.models import MonitorTask
from api.monitor.models import MonitorNote, MonitorNoteMetric, MonitorTask
XHS_TARGET = "5f58bd990000000001003753"
@@ -43,6 +46,33 @@ async def client(tmp_path):
await monitor_db.dispose_engine()
async def _seed_note_with_one_snapshot(task_name: str) -> None:
"""给某个任务塞一条作品和一次指标快照。
过滤类测试**必须有真数据**才有意义 —— 库里空着的话,过滤有没有生效结果都是 0,
测试就变成了空跑(这个坑踩过一次:一个报表串数据的 bug 因此没被拦住)。
"""
async with monitor_db.get_session() as session:
task = await session.scalar(select(MonitorTask).where(MonitorTask.name == task_name))
now = get_current_timestamp()
session.add(
MonitorNote(
task_id=task.id, note_id="seed-note", title="seed", note_url="",
cover="", creator_hash="", source_kind="", published_at=None,
first_seen_run_id=1, first_seen_at=now,
last_seen_run_id=1, last_seen_at=now,
)
)
session.add(
MonitorNoteMetric(
task_id=task.id, note_id="seed-note", run_id=1, captured_at=now,
liked_count=42, comment_count=0, collected_count=0, share_count=0,
raw_liked_count="42", raw_comment_count="0",
raw_collected_count="0", raw_share_count="0",
)
)
class TestCapabilityMatrix:
@pytest.mark.asyncio
async def test_matrix_is_exposed_to_the_ui(self, client):
@@ -54,7 +84,28 @@ class TestCapabilityMatrix:
# what stops the UI offering a platform that can never produce data.
assert all("monitor_wired" in p for p in body["platforms"])
assert by_value["xhs"]["monitor_wired"] is True
assert by_value["dy"]["monitor_wired"] is False
assert by_value["dy"]["monitor_wired"] is True
def test_every_wired_platform_has_an_adapter(self):
"""能力矩阵说「接通了」,就必须真的有一套适配管子。
两个注册表(platforms.PLATFORM_CAPABILITIES 与 adapters.ADAPTERS)分开是有意的
—— 前者是给前端看的能力描述,后者是爬虫的管道细节。代价是它们可能漂移,
所以在这里钉一条:凡声明接通的,必须能找到适配器。
"""
for platform in platforms.all_platforms():
if platforms.is_monitor_wired(platform):
assert adapters.has_adapter(platform), f"{platform} 声明接通但没有适配器"
@pytest.mark.asyncio
async def test_target_hints_are_exposed_for_wired_platforms(self, client):
"""前端的目标输入框拿它做 placeholder —— 让用户看到本平台该粘什么样的链接。"""
body = (await client.get("/api/config/platforms")).json()
by_value = {p["value"]: p for p in body["platforms"]}
assert "douyin.com/user/" in by_value["dy"]["target_hints"]["creator"]
assert "douyin.com/video/" in by_value["dy"]["target_hints"]["note"]
assert "xiaohongshu.com" in by_value["xhs"]["target_hints"]["creator"]
@pytest.mark.asyncio
async def test_metrics_are_per_platform_and_labelled(self, client):
@@ -80,14 +131,17 @@ class TestCapabilityMatrix:
class TestTaskCreationGuard:
@pytest.mark.asyncio
async def test_unwired_platform_is_rejected_with_an_explanation(self, client):
"""Accepting it would create a task that silently never produces data."""
"""Accepting it would create a task that silently never produces data.
用 B站 而不是抖音:抖音现在接通了,不再是「已知但未接通」的例子。
"""
response = await client.post(
"/api/monitor/tasks",
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": ["x"]},
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
)
assert response.status_code == 400
detail = response.json()["detail"]
assert "抖音" in detail
assert "B站" in detail
assert "尚未接通" in detail
@pytest.mark.asyncio
@@ -102,7 +156,7 @@ class TestTaskCreationGuard:
async def test_no_task_row_is_created_when_rejected(self, client):
await client.post(
"/api/monitor/tasks",
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": ["x"]},
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
)
assert (await client.get("/api/monitor/tasks")).json()["tasks"] == []
@@ -123,11 +177,34 @@ class TestTaskCreationGuard:
tasks = (await client.get("/api/monitor/tasks")).json()["tasks"]
assert {t["platform"] for t in tasks} == {"xhs"}
@pytest.mark.asyncio
async def test_an_explicit_platform_is_honoured_on_create(self, client):
"""建任务时给的平台必须落到那个平台。
缺省值是小红的(接口早期的兼容行为),所以「在抖音页面建任务」如果没有显式
带上 platform,就会安安静静地变成一个小红书任务 —— 不报错,只是出现在另一
个列表里。前端那半边已经改成必传;这里守住后端这一半:给了就必须用。
"""
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
created = await client.post(
"/api/monitor/tasks",
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": [sec_uid]},
)
assert created.status_code == 201
assert (await client.get("/api/monitor/tasks", params={"platform": "xhs"})).json()[
"tasks"
] == []
dy_tasks = (
await client.get("/api/monitor/tasks", params={"platform": "dy"})
).json()["tasks"]
assert [t["name"] for t in dy_tasks] == ["抖音任务"]
class TestPlatformScoping:
async def _seed_two_platforms(self, client):
"""One real XHS task plus a Douyin task inserted directly, since the API
refuses to create the latter."""
"""One XHS task created through the API, plus a Douyin task inserted
directly so its fields can be pinned exactly."""
await client.post(
"/api/monitor/tasks",
json={"name": "小红书任务", "mode": "creator", "targets": [XHS_TARGET]},
@@ -166,8 +243,15 @@ class TestPlatformScoping:
@pytest.mark.asyncio
async def test_a_platform_with_no_tasks_yields_empty_not_everything(self, client):
"""An empty task set must not degrade into "no filter"."""
"""空的任务集合不能退化成「不加过滤」。
**这里必须真的有数据。** 没有数据时,过滤生效与否结果都是 0 —— 这条测试原先
就栽在这个空跑上,所以没能拦下一个报表串数据的 bug(切到抖音,报表里却出现
小红书的数据)。最后那段「小红书自己的报表看得到」就是为了证明这些数据确实
存在、上面那两个 0 是过滤出来的。
"""
await self._seed_two_platforms(client)
await _seed_note_with_one_snapshot("小红书任务")
body = (await client.get("/api/monitor/notes", params={"platform": "bili"})).json()
assert body["notes"] == []
@@ -178,6 +262,11 @@ class TestPlatformScoping:
assert report["totals"]["liked_count_delta"] == 0
assert report["note_count"] == 0
xhs = (
await client.get("/api/monitor/report", params={"platform": "xhs"})
).json()
assert xhs["note_count"] == 1
class TestPerPlatformSettings:
@pytest.mark.asyncio
+268
View File
@@ -0,0 +1,268 @@
# -*- coding: utf-8 -*-
"""监控侧扫码登录与登录态检测。
两个要点在这里被钉死:
* 二维码必须从浏览器**默认 context** 里读 —— 新建 context 是无痕式的 profile,
扫了也白扫,爬虫读不到那份 cookie;
* 「登录了吗」不能用页面里的 `window.__INITIAL_STATE__`。那是**页面加载那一刻的快照**:
浏览器本来就登录着时它是对的,但扫码是加载**之后**才登录的,快照不会翻转,
于是扫完码界面会一直停在二维码上。判据改成拿 cookie 问后台接口。
"""
from unittest.mock import AsyncMock, MagicMock
import pytest
from api.creator.client import CreatorApiError
from api.monitor import qrlogin
@pytest.fixture(autouse=True)
def _reset_module_state():
for attribute in ("_current", "_page", "_playwright", "_state_cache"):
setattr(qrlogin, attribute, None)
yield
for attribute in ("_current", "_page", "_playwright", "_state_cache"):
setattr(qrlogin, attribute, None)
XHS_COOKIES = [
{"name": "a1", "value": "an-a1-value"},
{"name": "web_session", "value": "a-session"},
]
def _fake_stack(cookies=None, qr="data:image/png;base64,AAAA"):
"""Chrome/Playwright 替身,行为与真实的一致。"""
page = MagicMock()
page.url = "https://www.xiaohongshu.com/explore"
page.is_closed = MagicMock(return_value=False)
page.goto = AsyncMock()
page.close = AsyncMock()
context = MagicMock()
context.pages = []
context.cookies = AsyncMock(return_value=list(cookies if cookies is not None else XHS_COOKIES))
context.new_page = AsyncMock(return_value=page)
browser = MagicMock()
browser.contexts = [context]
# 去新建 context 正是这里要防的 bug,所以让它直接炸,而不是悄悄返回一个无痕 profile。
browser.new_context = AsyncMock(
side_effect=AssertionError("must reuse browser.contexts[0], not a new context")
)
playwright = MagicMock()
playwright.chromium.connect_over_cdp = AsyncMock(return_value=browser)
playwright.stop = AsyncMock()
manager = MagicMock()
manager.start = AsyncMock(return_value=playwright)
return manager, playwright, browser, context, page, qr
def _patch(monkeypatch, manager, qr="data:image/png;base64,AAAA", resolver=None):
"""``resolver(cookie)`` 返回账号信息 dict,或抛 CreatorApiError。"""
if resolver is None:
resolver = lambda _cookie: {"user_id": "u1", "nickname": "小明"} # noqa: E731
class _Client:
def __init__(self, cookie, **kwargs):
self.cookie = cookie
async def fetch_user_info(self):
return resolver(self.cookie)
monkeypatch.setattr(qrlogin, "async_playwright", lambda: manager)
monkeypatch.setattr(qrlogin, "CreatorClient", _Client)
monkeypatch.setattr(qrlogin.utils, "find_login_qrcode", AsyncMock(return_value=qr))
def _signed_out(_cookie):
raise CreatorApiError("登录态无效或已过期", status=401)
@pytest.mark.asyncio
async def test_idle_reports_the_browsers_login_state(monkeypatch):
manager, *_ = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
snapshot = await qrlogin.status()
assert snapshot["status"] == qrlogin.STATUS_IDLE
assert snapshot["logged_in"] is True
assert snapshot["nickname"] == "小明"
@pytest.mark.asyncio
async def test_unwired_platform_is_rejected():
"""只有小红书接了扫码;别的平台必须直接报错,而不是给个按不动的按钮。"""
with pytest.raises(ValueError):
await qrlogin.start("dy")
@pytest.mark.asyncio
async def test_start_reads_the_qr_from_the_default_context(monkeypatch):
manager, _pw, browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=_signed_out)
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
assert snapshot["status"] == qrlogin.STATUS_WAITING
assert snapshot["image"] == "data:image/png;base64,AAAA"
context.new_page.assert_awaited_once()
browser.new_context.assert_not_called()
@pytest.mark.asyncio
async def test_start_does_not_open_a_second_tab(monkeypatch):
"""已有的 xhs 标签页会被认领,所以重启不会在浏览器里堆孤儿页。"""
manager, _pw, _browser, context, page, _qr = _fake_stack()
context.pages = [page]
_patch(monkeypatch, manager, resolver=_signed_out)
await qrlogin.start(qrlogin.PLATFORM_XHS)
context.new_page.assert_not_called()
@pytest.mark.asyncio
async def test_an_already_signed_in_profile_needs_no_scan(monkeypatch):
"""没二维码但 profile 已登录 —— 这是成功,不是失败。"""
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, qr="", resolver=lambda _c: {"user_id": "u9", "nickname": "老王"})
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
assert snapshot["status"] == qrlogin.STATUS_SUCCESS
assert snapshot["nickname"] == "老王"
@pytest.mark.asyncio
async def test_an_already_signed_in_profile_never_opens_a_page(monkeypatch):
"""已登录时**根本不该去开页面**。
读二维码内部会 wait_for_selector 等满 30 秒才放弃,而已经登录时页面上没有二维码 ——
顺序反了的话,用户点一下按钮要干等半分钟,还白开一个标签页。
"""
manager, _pw, _browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
await qrlogin.start(qrlogin.PLATFORM_XHS)
context.new_page.assert_not_called()
@pytest.mark.asyncio
async def test_no_qr_and_not_signed_in_is_an_error(monkeypatch):
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, qr="", resolver=_signed_out)
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
assert snapshot["status"] == qrlogin.STATUS_ERROR
@pytest.mark.asyncio
async def test_a_completed_scan_flips_the_session_to_success(monkeypatch):
"""**这个判据是重点**:扫码是页面加载之后才发生的,所以不能用页面快照来判断。"""
manager, *_rest = _fake_stack()
signed_in = {"value": False}
def resolver(cookie):
assert "a1=an-a1-value" in cookie # 判据必须真的用 cookie 去问
if not signed_in["value"]:
raise CreatorApiError("登录态无效或已过期", status=401)
return {"user_id": "u1", "nickname": "小红"}
_patch(monkeypatch, manager, resolver=resolver)
await qrlogin.start(qrlogin.PLATFORM_XHS)
assert qrlogin._current.status == qrlogin.STATUS_WAITING
# 操作者扫了码
signed_in["value"] = True
qrlogin._state_cache = None # 5 秒缓存否则会遮住这次变化
snapshot = await qrlogin.status()
assert snapshot["status"] == qrlogin.STATUS_SUCCESS
assert snapshot["nickname"] == "小红"
@pytest.mark.asyncio
async def test_the_successful_session_hands_over_a_cookie(monkeypatch):
"""扫码不该只写浏览器 profile —— 还要能把 cookie 交出来存库,
否则关掉 CDP 就断了。"""
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
await qrlogin.start(qrlogin.PLATFORM_XHS)
cookie = await qrlogin.take_cookie()
assert cookie is not None
assert "a1=an-a1-value" in cookie
# 只能取一次,否则每次轮询都会重复写库
assert await qrlogin.take_cookie() is None
@pytest.mark.asyncio
async def test_session_expires(monkeypatch):
manager, *_rest = _fake_stack()
_patch(monkeypatch, manager, resolver=_signed_out)
await qrlogin.start(qrlogin.PLATFORM_XHS)
qrlogin._current.started_at -= qrlogin.QR_TTL_SECONDS + 1
snapshot = await qrlogin.status()
assert snapshot["status"] == qrlogin.STATUS_EXPIRED
@pytest.mark.asyncio
async def test_check_login_state_reports_when_the_browser_cannot_answer(monkeypatch):
"""连不上浏览器时要说出来,不能悄悄报成「未登录」。"""
manager, _pw, _browser, context, _page, _qr = _fake_stack()
context.cookies = AsyncMock(side_effect=RuntimeError("Target closed"))
_patch(monkeypatch, manager)
state = await qrlogin.check_login_state()
assert state["known"] is False
assert state["logged_in"] is False
assert "Target closed" in state["error"]
@pytest.mark.asyncio
async def test_check_login_state_is_cached(monkeypatch):
manager, _pw, _browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "x"})
await qrlogin.check_login_state()
await qrlogin.check_login_state()
context.cookies.assert_awaited_once()
@pytest.mark.asyncio
async def test_force_bypasses_the_cache(monkeypatch):
manager, _pw, _browser, context, _page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "x"})
await qrlogin.check_login_state()
await qrlogin.check_login_state(force=True)
assert context.cookies.await_count == 2
@pytest.mark.asyncio
async def test_cancel_keeps_the_operators_tab(monkeypatch):
"""与运营模块不同:那里的上下文是临时的、用完即弃;这里的标签页属于操作者的浏览器。"""
manager, _pw, _browser, _context, page, _qr = _fake_stack()
_patch(monkeypatch, manager, resolver=_signed_out)
await qrlogin.start(qrlogin.PLATFORM_XHS)
snapshot = await qrlogin.cancel()
assert snapshot["status"] == qrlogin.STATUS_IDLE
page.close.assert_not_called()
+194
View File
@@ -0,0 +1,194 @@
# -*- coding: utf-8 -*-
"""Tests for monitor task schedule arithmetic.
Everything here is timezone-local, matching the implementation: the container is
pinned to the operator's zone via TZ, so the tests build their expectations from
naive local datetimes too and stay correct wherever they run.
"""
from datetime import datetime, time, timedelta
import pytest
from api.monitor import schedule
def _ms(moment: datetime) -> int:
return int(moment.timestamp() * 1000)
def test_interval_is_now_plus_the_interval():
after = _ms(datetime(2026, 10, 7, 9, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_INTERVAL,
interval_minutes=120,
hours=[],
days=[],
minute=0,
after_ms=after,
)
assert nxt == after + 120 * 60_000
def test_daily_takes_the_soonest_remaining_time_today():
after = _ms(datetime(2026, 10, 7, 8, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[18, 9], # deliberately unsorted
days=[],
minute=30,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 7, 9, 30))
def test_daily_rolls_over_to_tomorrow_once_every_time_has_passed():
after = _ms(datetime(2026, 10, 7, 20, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[9, 18],
days=[],
minute=30,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 8, 9, 30))
def test_a_slot_exactly_now_belongs_to_the_next_day():
"""Strictly-after, so the run that just fired does not fire again.
The scheduler advances with ``after_ms`` set to the moment the run started,
which is at or just past the slot -- if the comparison were inclusive it would
pick the same slot back up and loop.
"""
after = _ms(datetime(2026, 10, 7, 9, 30))
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[9],
days=[],
minute=30,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 8, 9, 30))
def test_weekly_jumps_to_the_next_selected_weekday():
after_dt = datetime(2026, 10, 7, 8, 0)
target = (after_dt.weekday() + 2) % 7
nxt = schedule.next_occurrence(
mode=schedule.MODE_WEEKLY,
interval_minutes=60,
hours=[10],
days=[target],
minute=0,
after_ms=_ms(after_dt),
)
expected = datetime.combine((after_dt + timedelta(days=2)).date(), time(10, 0))
assert nxt == _ms(expected)
def test_weekly_can_fire_later_the_same_day():
after_dt = datetime(2026, 10, 7, 8, 0)
nxt = schedule.next_occurrence(
mode=schedule.MODE_WEEKLY,
interval_minutes=60,
hours=[21],
days=[after_dt.weekday()],
minute=15,
after_ms=_ms(after_dt),
)
assert nxt == _ms(datetime.combine(after_dt.date(), time(21, 15)))
def test_weekly_without_weekdays_means_every_day():
"""Otherwise an empty day selection would match nothing and never fire."""
after = _ms(datetime(2026, 10, 7, 8, 0))
nxt = schedule.next_occurrence(
mode=schedule.MODE_WEEKLY,
interval_minutes=60,
hours=[9],
days=[],
minute=0,
after_ms=after,
)
assert nxt == _ms(datetime(2026, 10, 7, 9, 0))
def test_a_clock_schedule_with_no_times_can_never_fire():
"""Returned as None so the caller can park the task instead of leaving it due."""
nxt = schedule.next_occurrence(
mode=schedule.MODE_DAILY,
interval_minutes=60,
hours=[],
days=[],
minute=0,
after_ms=_ms(datetime(2026, 10, 7, 8, 0)),
)
assert nxt is None
@pytest.mark.parametrize(
"raw, expected",
[
("9,18", [9, 18]),
("18,9", [9, 18]), # stored order is not guaranteed
("9,9,9", [9]),
("", []),
(None, []),
("9, 18 ", [9, 18]),
("9,99,-1,abc,", [9]), # junk is dropped, never raised
],
)
def test_parse_hours_is_forgiving(raw, expected):
assert schedule.parse_hours(raw) == expected
def test_parse_days_accepts_the_whole_week():
assert schedule.parse_days("0,1,2,3,4,5,6") == [0, 1, 2, 3, 4, 5, 6]
assert schedule.parse_days("7,-1") == []
@pytest.mark.parametrize(
"mode, interval, hours, days, minute, expected",
[
(schedule.MODE_INTERVAL, 360, [], [], 0, "每 6 小时"),
(schedule.MODE_INTERVAL, 1440, [], [], 0, "每 1 天"),
(schedule.MODE_INTERVAL, 45, [], [], 0, "每 45 分钟"),
(schedule.MODE_DAILY, 60, [9, 18], [], 30, "每天 09:30、18:30"),
(schedule.MODE_WEEKLY, 60, [10], [0, 1, 2, 3, 4], 0, "周一、周二、周三、周四、周五 10:00"),
(schedule.MODE_WEEKLY, 60, [10], [], 0, "每天 10:00"),
(schedule.MODE_DAILY, 60, [], [], 0, "未设置时间"),
],
)
def test_describe(mode, interval, hours, days, minute, expected):
assert (
schedule.describe(
mode=mode, interval_minutes=interval, hours=hours, days=days, minute=minute
)
== expected
)
def test_format_round_trips_through_parse():
hours = [9, 12, 18]
days = [0, 4]
assert schedule.parse_hours(schedule.format_hours(hours)) == hours
assert schedule.parse_days(schedule.format_days(days)) == days
+8 -8
View File
@@ -20,7 +20,7 @@ def test_extract_search_note_list_from_keyword_page():
assert notes[0].note_id == "9117888152"
assert notes[0].title.startswith("武汉交互空间科技")
assert notes[0].tieba_name == "武汉交互空间"
assert notes[0].user_nickname == "V***人"
assert notes[0].user_nickname == "VR虚拟达人"
def test_extract_search_note_list_from_current_pc_card_page():
@@ -56,7 +56,7 @@ def test_extract_search_note_list_from_current_pc_card_page():
assert notes[0].desc == "培训班需求,数学,英语,编程老师,专职兼职都可"
assert notes[0].tieba_name == "诸城吧"
assert notes[0].tieba_link.endswith("kw=%E8%AF%B8%E5%9F%8E")
assert notes[0].user_nickname == "7***7"
assert notes[0].user_nickname == "754023117"
assert notes[0].publish_time == "2026-3-15"
assert notes[0].total_replay_num == 19
@@ -147,7 +147,7 @@ def test_extract_note_detail_and_comments_from_current_pc_api():
assert note.note_id == "10451142633"
assert note.title == "这X尔斯对比巴尔斯,我只能说ID正确,允许居功自傲"
assert note.desc == "皮队败决处刑德国编程钢琴师兼职数学家"
assert note.user_nickname == "泰***克"
assert note.user_nickname == "泰高祖蒙斯克"
assert note.tieba_name == "dota2吧"
assert note.total_replay_num == 15
assert note.total_replay_page == 1
@@ -155,7 +155,7 @@ def test_extract_note_detail_and_comments_from_current_pc_api():
assert len(comments) == 1
assert comments[0].comment_id == "153154097267"
assert comments[0].content == "xg现在大树阵容另一个辅助不选控制"
assert comments[0].user_nickname == "期***3"
assert comments[0].user_nickname == "期胡希3"
assert comments[0].sub_comment_count == 4
# 教学版已移除 ip_location 等可定位真人字段
@@ -191,7 +191,7 @@ def test_extract_creator_info_and_threads_from_current_pc_api():
creator = extractor.extract_creator_info_from_api(creator_api)
thread_ids = extractor.extract_creator_thread_id_list_from_api(feed_api)
assert creator.user_nickname == "米***子"
assert creator.user_nickname == "米米世界大手子"
assert creator.fans == 58
assert creator.follows == 1
# 教学版已移除 user_id、user_name、ip_location 等可定位真人字段
@@ -223,7 +223,7 @@ def test_extract_tieba_note_list_from_bigpipe_thread_page():
assert len(notes) == 48
assert notes[0].note_id == "9079949995"
assert notes[0].title == "盗墓笔记全集+txt小说,已整理"
assert notes[0].user_nickname == "公***仲"
assert notes[0].user_nickname == "公子伯仲"
assert notes[0].tieba_name == "盗墓笔记吧"
assert notes[0].tieba_link.endswith("kw=%E7%9B%97%E5%A2%93%E7%AC%94%E8%AE%B0&ie=utf-8")
@@ -233,7 +233,7 @@ def test_extract_note_detail_from_post_page():
assert note.note_id == "9117905169"
assert note.title == "对于一个父亲来说,这个女儿14岁就死了"
assert note.user_nickname == "章***轩"
assert note.user_nickname == "章景轩"
assert note.tieba_name == "以太比特吧"
assert note.total_replay_num == 786
assert note.total_replay_page == 13
@@ -249,7 +249,7 @@ def test_extract_parent_comments_from_post_page():
assert len(comments) == 30
assert comments[0].comment_id == "150726491368"
assert comments[0].content == "中国队第22金!无悬念!"
assert comments[0].user_nickname == "h***n"
assert comments[0].user_nickname == "heinzfrentzen"
assert comments[0].tieba_name == "网球风云吧"
# 教学版已移除 ip_location 等可定位真人字段
+420
View File
@@ -0,0 +1,420 @@
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_upstream.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""上游更新检查:git 输出怎么解析,以及什么时候才推送。
所有会碰网络的路径都被替掉了 —— 测试里既没有上游仓库,也不该有。真正被测的是
解析、去重和调度到期这三件事,它们才是容易出错的部分。
"""
import subprocess
import httpx
import pytest
import pytest_asyncio
from api.main import app
from api.monitor import db as monitor_db
from api.monitor import notify
from api.monitor import scheduler as scheduler_module
from api.monitor import upstream
from api.monitor.models import SETTING_WECOM_WEBHOOK
from api.monitor.scheduler import MonitorScheduler
from api.monitor.settings import set_setting
from tools.time_util import get_current_timestamp
SEP = upstream._RECORD_SEPARATOR
HEAD_SHA = "a" * 40
TIP_SHA = "b" * 40
WEBHOOK = "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=test"
@pytest_asyncio.fixture
async def db(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
yield monitor_db
await monitor_db.dispose_engine()
@pytest_asyncio.fixture
async def client(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
transport = httpx.ASGITransport(app=app)
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
yield http_client
await monitor_db.dispose_engine()
def _signature(args: list[str]) -> str:
"""Map one git invocation onto a key a test can name.
``rev-parse`` and ``rev-list`` are each called more than once with different
arguments, so the subcommand alone is not enough to key on.
"""
if args[0] == "rev-parse":
return f"rev-parse {args[1]}"
if args[0] == "rev-list":
return f"rev-list {args[-1]}"
return args[0]
@pytest.fixture
def fake_git(monkeypatch):
"""Install canned git output. Returns the list of invocations made."""
def _install(responses: dict) -> list:
calls: list = []
def _run(args, timeout):
calls.append(args)
key = _signature(args)
if key not in responses:
raise AssertionError(f"测试没有为这条 git 调用准备输出:{args}")
return responses[key]
monkeypatch.setattr(upstream, "_run", _run)
return calls
return _install
def _log_line(sha: str, subject: str) -> str:
return f"{sha}{SEP}张三{SEP}2026-10-01{SEP}{subject}"
class TestCheckParsing:
def test_counts_commits_and_reads_the_list(self, fake_git):
calls = fake_git(
{
"rev-parse HEAD": (0, HEAD_SHA, ""),
"fetch": (0, "", ""),
"rev-parse FETCH_HEAD": (0, TIP_SHA, ""),
"rev-list HEAD..FETCH_HEAD": (0, "3", ""),
"rev-list FETCH_HEAD..HEAD": (0, "128", ""),
"log": (
0,
"\n".join(
[
_log_line("abc1234", "fix: 修了扫码过期"),
_log_line("def5678", "feat: 加了新平台"),
_log_line("9999999", "docs: 更新说明"),
]
),
"",
),
}
)
result = upstream._check_sync("https://example.invalid/repo.git", "main")
assert result.ok is True
assert result.behind == 3
# 领先数就是这一层的规模,合并时要一起保留,所以值得单独报出来。
assert result.ahead == 128
assert result.tip == TIP_SHA
assert result.head == HEAD_SHA
assert [commit.subject for commit in result.commits] == [
"fix: 修了扫码过期",
"feat: 加了新平台",
"docs: 更新说明",
]
assert result.commits[0].sha == "abc1234"
assert result.error == ""
def test_up_to_date_skips_reading_the_log(self, fake_git):
"""不落后时不该再去读提交列表 —— 那条 git log 没有意义。"""
calls = fake_git(
{
"rev-parse HEAD": (0, HEAD_SHA, ""),
"fetch": (0, "", ""),
"rev-parse FETCH_HEAD": (0, TIP_SHA, ""),
"rev-list HEAD..FETCH_HEAD": (0, "0", ""),
"rev-list FETCH_HEAD..HEAD": (0, "128", ""),
}
)
result = upstream._check_sync("https://example.invalid/repo.git", "main")
assert result.ok is True
assert result.behind == 0
assert result.commits == []
assert all(call[0] != "log" for call in calls)
def test_fetch_failure_is_reported_not_raised(self, fake_git):
"""GitHub 不通是常态,那也该是一条能显示出来的结论。"""
fake_git(
{
"rev-parse HEAD": (0, HEAD_SHA, ""),
"fetch": (128, "", "fatal: unable to access 'https://github.com/': 连接超时\n第二行"),
}
)
result = upstream._check_sync("https://example.invalid/repo.git", "main")
assert result.ok is False
assert result.head == HEAD_SHA
# 只留第一行 stderr:git 的报错常常跟一大段建议,塞进界面反而看不清。
assert "连接超时" in result.error
assert "第二行" not in result.error
def test_not_a_git_repository_is_reported(self, fake_git):
fake_git({"rev-parse HEAD": (128, "", "fatal: not a git repository (or any of the parent directories): .git")})
result = upstream._check_sync("https://example.invalid/repo.git", "main")
assert result.ok is False
assert "读取本地 HEAD 失败" in result.error
def test_missing_git_binary_is_reported(self, monkeypatch):
def _explode(args, timeout):
raise FileNotFoundError("git")
monkeypatch.setattr(upstream, "_git", _explode)
result = upstream._check_sync("https://example.invalid/repo.git", "main")
assert result.ok is False
assert "未找到 git" in result.error
def test_fetch_timeout_is_reported(self, monkeypatch):
def _explode(args, timeout):
raise subprocess.TimeoutExpired(cmd="git", timeout=timeout)
monkeypatch.setattr(upstream, "_git", _explode)
result = upstream._check_sync("https://example.invalid/repo.git", "main")
assert result.ok is False
assert "超时" in result.error
class TestNotification:
@pytest.fixture
def sent(self, monkeypatch) -> list:
messages: list = []
async def _fake_send(url, content):
messages.append(content)
return True, "发送成功"
monkeypatch.setattr(notify, "send_wecom", _fake_send)
return messages
@staticmethod
def _patch_check(monkeypatch, *, behind: int, tip: str, ahead: int = 0):
async def _fake_check(remote_url=upstream.DEFAULT_REMOTE_URL, branch=upstream.DEFAULT_BRANCH):
return upstream.CheckResult(
ok=True,
behind=behind,
ahead=ahead,
tip=tip,
head=HEAD_SHA,
commits=[upstream.Commit(sha="abc1234", author="张三", date="2026-10-01", subject="fix: 修了扫码过期")],
)
monkeypatch.setattr(upstream, "check", _fake_check)
@pytest.mark.asyncio
async def test_pushes_once_per_upstream_tip(self, db, monkeypatch, sent):
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA, ahead=128)
async with monitor_db.get_session() as session:
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
first = await upstream.run_check()
# 同一个 tip 再查一次:不该重复推。
second = await upstream.run_check()
assert len(sent) == 1
assert first["notified"] is True
assert "notified" not in second
assert "落后 `main` **2** 个提交" in sent[0]
assert "abc1234" in sent[0]
# 上游又动了:tip 变了就该再推一次。
self._patch_check(monkeypatch, behind=5, tip="c" * 40, ahead=128)
third = await upstream.run_check()
assert len(sent) == 2
assert third["notified"] is True
@pytest.mark.asyncio
async def test_state_is_persisted_for_the_ui(self, db, monkeypatch, sent):
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA, ahead=128)
async with monitor_db.get_session() as session:
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
await upstream.run_check()
async with monitor_db.get_session() as session:
state = await upstream.load_state(session)
assert state["behind"] == 2
assert state["ahead"] == 128
assert state["tip"] == TIP_SHA
assert state["branch"] == "main"
assert state["checked_at"] > 0
@pytest.mark.asyncio
async def test_nothing_new_does_not_push(self, db, monkeypatch, sent):
self._patch_check(monkeypatch, behind=0, tip=TIP_SHA)
async with monitor_db.get_session() as session:
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
await upstream.run_check()
assert sent == []
@pytest.mark.asyncio
async def test_switch_off_does_not_push(self, db, monkeypatch, sent):
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA)
async with monitor_db.get_session() as session:
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
await set_setting(session, "system.upstream_notify", "false")
await upstream.run_check()
assert sent == []
@pytest.mark.asyncio
async def test_manual_check_can_skip_the_push(self, db, monkeypatch, sent):
"""手动点「立即检查」只看结果,不因为它把群消息推一遍。"""
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA)
async with monitor_db.get_session() as session:
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
await upstream.run_check(notify_when_new=False)
assert sent == []
@pytest.mark.asyncio
async def test_no_webhook_configured_is_not_an_error(self, db, monkeypatch, sent):
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA)
result = await upstream.run_check()
assert sent == []
assert result["ok"] is True
assert result["behind"] == 2
class TestScheduler:
@pytest.fixture
def checks(self, monkeypatch) -> list:
calls: list = []
async def _fake_run_check(notify_when_new: bool = True):
calls.append(notify_when_new)
# 真实的 run_check 会把 checked_at 写进去,调度器的「到点了没有」
# 全靠这个字段,所以替身也必须写。
async with monitor_db.get_session() as session:
await upstream._save_state(
session,
{"checked_at": get_current_timestamp(), "ok": True, "behind": 0},
)
return {"ok": True, "behind": 0}
monkeypatch.setattr(upstream, "run_check", _fake_run_check)
return calls
@pytest.mark.asyncio
async def test_disabled_never_checks(self, db, checks):
await MonitorScheduler()._maybe_check_upstream()
assert checks == []
@pytest.mark.asyncio
async def test_checks_when_due_and_then_waits_out_the_interval(self, db, checks):
async with monitor_db.get_session() as session:
await set_setting(session, "system.upstream_check_enabled", "true")
scheduler = MonitorScheduler()
await scheduler._maybe_check_upstream()
assert checks == [True]
# 刚查过:间隔(默认一天)没到就不该再查。
await scheduler._maybe_check_upstream()
assert checks == [True]
@pytest.mark.asyncio
async def test_a_stale_timestamp_is_due_again(self, db, checks):
async with monitor_db.get_session() as session:
await set_setting(session, "system.upstream_check_enabled", "true")
await set_setting(session, "system.upstream_check_interval_minutes", "30")
await upstream._save_state(
session,
# 差一分钟就到期,用来卡住边界:31 分钟前那次已经算过期。
{"checked_at": get_current_timestamp() - 31 * 60_000, "ok": True, "behind": 0},
)
await MonitorScheduler()._maybe_check_upstream()
assert checks == [True]
class TestEndpoint:
@pytest.mark.asyncio
async def test_status_is_empty_before_the_first_check(self, client):
response = await client.get("/api/monitor/upstream")
assert response.status_code == 200
assert response.json() == {}
@pytest.mark.asyncio
async def test_manual_check_runs_and_is_readable_back(self, client, monkeypatch):
async def _fake_check(remote_url=upstream.DEFAULT_REMOTE_URL, branch=upstream.DEFAULT_BRANCH):
return upstream.CheckResult(ok=True, behind=1, tip=TIP_SHA, head=HEAD_SHA)
monkeypatch.setattr(upstream, "check", _fake_check)
checked = (await client.post("/api/monitor/upstream/check")).json()
assert checked["behind"] == 1
# 结果落库,随后的 GET 读的是同一份缓存(而不是再 fetch 一次)。
cached = (await client.get("/api/monitor/upstream")).json()
assert cached["behind"] == 1
assert cached["tip"] == TIP_SHA
@pytest.mark.asyncio
async def test_manual_check_does_not_push(self, client, monkeypatch):
"""点按钮的人正看着结果,不该再给自己推一条群消息。"""
sent: list = []
async def _fake_send(url, content):
sent.append(content)
return True, "发送成功"
async def _fake_check(remote_url=upstream.DEFAULT_REMOTE_URL, branch=upstream.DEFAULT_BRANCH):
return upstream.CheckResult(ok=True, behind=1, tip=TIP_SHA, head=HEAD_SHA)
monkeypatch.setattr(notify, "send_wecom", _fake_send)
monkeypatch.setattr(upstream, "check", _fake_check)
async with monitor_db.get_session() as session:
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
body = (await client.post("/api/monitor/upstream/check")).json()
assert sent == []
assert body["behind"] == 1
# 没推过的那批提交留给下一次定时检查,所以这里不该记成已推送。
async with monitor_db.get_session() as session:
state = await upstream.load_state(session)
assert "notified" not in state
+14
View File
@@ -22,6 +22,20 @@ import asyncio
import pytest
import config
@pytest.fixture(autouse=True)
def _force_nickname_masking(monkeypatch):
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
所以这里显式打开来测。
"""
monkeypatch.setattr(config, "MASK_NICKNAME", True)
# 原始(明文)测试数据
RAW_USER_ID = 7654321
RAW_NICKNAME = "微博达人"
+200
View File
@@ -0,0 +1,200 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""Clone one schema to a new database on the same MySQL server, and provision a
scoped account for it.
Written for the production cut-over: the monitor data lives in `mediacrawler` and
the server deployment reads `mediacrawler_prod`. Both sit on the same host, so the
copy is a cross-schema ``INSERT ... SELECT`` rather than a dump-and-reload -- and
since neither this workstation nor the server has a mysql client, that is also
the only option available.
Two things this deliberately does NOT do, both because the server hosts ~22
unrelated production databases and the provisioning account is a full admin:
* it never writes outside the two schemas named on the command line;
* it hands the application its own account, scoped to the new schema, rather than
reusing that admin account for the app.
Foreign keys are disabled only for the duration of the copy. Source and target
definitions are identical, so ordering is the sole thing at stake, and turning
the checks off is what makes an arbitrary table order safe.
"""
import argparse
import asyncio
import secrets
import string
import sys
from pathlib import Path
import aiomysql
PASSWORD_ALPHABET = string.ascii_letters + string.digits
def generate_password(length: int = 32) -> str:
"""Alphanumeric only: it ends up in a .env value and a SQL literal, and a
symbol that needs escaping in either place is a support ticket waiting."""
return "".join(secrets.choice(PASSWORD_ALPHABET) for _ in range(length))
async def _admin_conn(args, db=None):
return await aiomysql.connect(
host=args.host,
port=args.port,
user=args.admin_user,
password=args.admin_password,
db=db,
charset="utf8mb4",
autocommit=True,
)
async def create_schema(args) -> None:
conn = await _admin_conn(args)
try:
cur = await conn.cursor()
await cur.execute(
f"CREATE DATABASE IF NOT EXISTS `{args.target_db}` "
"DEFAULT CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci"
)
await cur.execute(
"SELECT DEFAULT_CHARACTER_SET_NAME, DEFAULT_COLLATION_NAME "
"FROM information_schema.SCHEMATA WHERE SCHEMA_NAME = %s",
(args.target_db,),
)
charset, collation = await cur.fetchone()
print(f"[库] {args.target_db} {charset} / {collation}")
await cur.close()
finally:
conn.close()
async def provision_app_user(args, password: str) -> None:
conn = await _admin_conn(args)
try:
cur = await conn.cursor()
# The host is bound as a parameter instead of written as a literal '%':
# aiomysql applies %-substitution whenever arguments are supplied, so a
# bare '%' in the SQL is read as a format specifier and raises
# "unsupported format character".
await cur.execute(
"CREATE USER IF NOT EXISTS %s@%s IDENTIFIED BY %s",
(args.app_user, "%", password),
)
# Idempotent re-runs: if the account already existed, make the password
# the one we are about to write into .env rather than a stale one.
await cur.execute(
"ALTER USER %s@%s IDENTIFIED BY %s", (args.app_user, "%", password)
)
await cur.execute(
f"GRANT ALL PRIVILEGES ON `{args.target_db}`.* TO %s@%s",
(args.app_user, "%"),
)
await cur.execute("FLUSH PRIVILEGES")
print(f"[账号] {args.app_user}@% 授权范围仅 `{args.target_db}`.*")
await cur.close()
finally:
conn.close()
async def copy_tables(args) -> None:
src = await _admin_conn(args, db=args.source_db)
dst = await _admin_conn(args, db=args.target_db)
try:
scur = await src.cursor()
await scur.execute("SHOW TABLES")
tables = [row[0] for row in await scur.fetchall()]
print(f"[表] 源库共 {len(tables)} 张")
dcur = await dst.cursor()
# Order is arbitrary on purpose -- checks are off for the whole copy, so
# a table may safely be created before the table it references.
await dcur.execute("SET FOREIGN_KEY_CHECKS=0")
for table in tables:
await scur.execute(f"SHOW CREATE TABLE `{table}`")
ddl = (await scur.fetchone())[1]
# `IF NOT EXISTS` makes the script re-runnable after a partial run.
await dcur.execute(ddl.replace("CREATE TABLE", "CREATE TABLE IF NOT EXISTS", 1))
await dcur.execute(f"DELETE FROM `{table}`")
await dcur.execute(
f"INSERT INTO `{table}` SELECT * FROM `{args.source_db}`.`{table}`"
)
copied = dcur.rowcount
await scur.execute(f"SELECT COUNT(*) FROM `{table}`")
expected = (await scur.fetchone())[0]
flag = "OK" if copied == expected else "!! 行数不符"
print(f" {table:<26} {copied:>6} / {expected:<6} {flag}")
await dcur.execute("SET FOREIGN_KEY_CHECKS=1")
await dcur.close()
await scur.close()
finally:
src.close()
dst.close()
def write_env(path: Path, args, password: str) -> None:
path.write_text(
f"""# 服务器部署环境配置(已被 .gitignore 忽略,不会提交)
#
# 由 tools/clone_database.py 生成。库名改了之后必须同时确认账号对该库有权限:
# 应用启动时会校验「实际连到的库」是否等于下面的 MYSQL_DB_NAME,不符会拒绝启动。
MC_HOST=0.0.0.0
MC_PORT=18051
# --- 数据库(监控层) ---
MYSQL_DB_HOST={args.host}
MYSQL_DB_PORT={args.port}
MYSQL_DB_USER={args.app_user}
MYSQL_DB_PWD={password}
MYSQL_DB_NAME={args.target_db}
# 登录鉴权:留空则首次启动自动生成随机密码并打印在启动日志里
# MC_PASSWORD=
# 面板走 HTTPS 时才打开;局域网明文 HTTP 下必须保持注释,
# 否则浏览器丢弃 Cookie,表现为登录页反复刷新且无任何报错
# MC_COOKIE_SECURE=1
""",
encoding="utf-8",
)
path.chmod(0o600)
async def main() -> int:
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("--host", default="192.168.2.27")
ap.add_argument("--port", type=int, default=3306)
ap.add_argument("--admin-user", required=True)
ap.add_argument("--admin-password", required=True)
ap.add_argument("--source-db", required=True)
ap.add_argument("--target-db", required=True)
ap.add_argument("--app-user", required=True)
ap.add_argument("--env-file", type=Path)
args = ap.parse_args()
if args.source_db == args.target_db:
print("源库与目标库相同,拒绝执行", file=sys.stderr)
return 2
password = generate_password()
await create_schema(args)
await provision_app_user(args, password)
await copy_tables(args)
if args.env_file:
write_env(args.env_file, args, password)
print(f"[.env] 已写入 {args.env_file}(权限 600,密码未回显)")
print("\n完成。")
return 0
if __name__ == "__main__":
sys.exit(asyncio.run(main()))
+210
View File
@@ -0,0 +1,210 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""Phase 0 probe: can the creator backend be reached with a plain signed request?
The question this answers, and why it is worth a whole script: two sources
disagree. The `xhshow` library ships `sign_xyw()` whose docstring says it exists
because the creator data APIs *reject* the main-site signature with HTTP 406 -- but
a field report claims the creator gateway rejects *any* self-made request with 406
regardless of signature. Only a live request settles it.
The response is read as a three-way verdict, because a bare "it failed" is not
useful here:
406 -> the gateway rejected the signature. The browser-interception route is
the only way forward.
401 / not-logged-in
-> the signature PASSED and only the creator session is missing. That is
good news: it means Phase 1 is pure-request after all.
200 -> we are through, and the payload is captured for field mapping.
Cookies are read out of the browser over CDP and never printed -- they are
credentials, and this script has no reason to echo them.
"""
import argparse
import base64
import hashlib
import json
import sys
import urllib.parse
from datetime import datetime, time, timedelta
# The XYW_ scheme, matching both xhshow/config/config.py and the independent
# reverse-engineering in xiaohongshu-cli. Constants are byte-identical in both.
XYW_AES_KEY = b"7cc4adla5ay0701v"
XYW_AES_IV = b"4uzjr7mbsibcaldp"
XYW_ENV_FLAGS = "0|0|0|1|0|0|1|0|0|0|1|0|0|0|0|1|0|0|0"
CREATOR_ORIGIN = "https://creator.xiaohongshu.com"
# The data-analysis note list. This is the page the creator console itself calls,
# and it is where exposure/views live.
NOTE_LIST_PATH = "/api/galaxy/creator/datacenter/note/analyze/list"
USER_AGENT = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
)
def _aes_encrypt_hex(plaintext: str) -> str:
from Crypto.Cipher import AES
from Crypto.Util.Padding import pad
cipher = AES.new(XYW_AES_KEY, AES.MODE_CBC, XYW_AES_IV)
return cipher.encrypt(pad(plaintext.encode("utf-8"), AES.block_size)).hex()
def sign_xyw(api: str, a1: str, app_id: str = "ugc", data: dict | None = None) -> dict[str, str]:
"""Mirror of xiaohongshu-cli's creator_signing.sign_creator.
Written out rather than imported so the probe can vary ``api`` and ``app_id``
independently -- figuring out which combination the gateway accepts is the
entire point of running it.
"""
content = api
if data is not None:
content += json.dumps(data, separators=(",", ":"), ensure_ascii=False)
digest = hashlib.md5(content.encode("utf-8")).hexdigest()
timestamp_ms = int(datetime.now().timestamp() * 1000)
plaintext = f"x1={digest};x2={XYW_ENV_FLAGS};x3={a1};x4={timestamp_ms};"
encoded = base64.b64encode(plaintext.encode("utf-8")).decode("utf-8")
envelope = {
"signSvn": "56",
"signType": "x2",
"appId": app_id,
"signVersion": "1",
"payload": _aes_encrypt_hex(encoded),
}
x_s = "XYW_" + base64.b64encode(
json.dumps(envelope, separators=(",", ":")).encode("utf-8")
).decode("utf-8")
return {"x-s": x_s, "x-t": str(timestamp_ms)}
def build_query(start_days_ago: int, end_days_ago: int, page_size: int = 10) -> str:
"""Query string exactly as the working collector builds it: epoch millis."""
today = datetime.now().replace(hour=0, minute=0, second=0, microsecond=0)
def ms(days_ago: int, at_end: bool) -> int:
day = (today - timedelta(days=days_ago)).date()
clock = time(23, 59, 59) if at_end else time(0, 0, 0)
return int(datetime.combine(day, clock).timestamp() * 1000)
return urllib.parse.urlencode(
{
"post_begin_time": ms(start_days_ago, False),
"post_end_time": ms(end_days_ago, True),
"type": 0,
"page_size": page_size,
"page_num": 1,
}
)
async def cookies_from_cdp() -> dict[str, str]:
"""Session cookies for creator.xiaohongshu.com, read out of the live browser."""
from playwright.async_api import async_playwright
playwright = await async_playwright().start()
try:
browser = await playwright.chromium.connect_over_cdp(
"http://127.0.0.1:9222", timeout=15000
)
jars = {}
for context in browser.contexts:
for cookie in await context.cookies():
jars[cookie["name"]] = cookie["value"]
return jars
finally:
await playwright.stop()
def classify(status: int, body: str) -> str:
if status == 406:
return "406 —— 网关拒了签名(回退浏览器拦截路线)"
if status == 200:
return "200 —— 通了,可以考虑解析数据"
if status in (401, 403):
return f"{status} —— 签名过了,只差创作者会话(好消息)"
return f"{status} —— 未知,需要看响应体"
async def main() -> int:
import httpx
ap = argparse.ArgumentParser()
ap.add_argument("--start-days-ago", type=int, default=30)
ap.add_argument("--end-days-ago", type=int, default=0)
ap.add_argument("--app-id", default="ugc", help="参考实现用 ugc;xhshow 默认 xhs-pc-web")
ap.add_argument("--cookie", default="", help="留空则从 CDP 浏览器读取")
ap.add_argument("--show-body", action="store_true", help="打印响应前 800 字符")
ap.add_argument(
"--no-cookie-header",
action="store_true",
help="签名照签(仍需 a1)但不发 cookie 头,用来分清"
"「空数据是缺会话」还是「接口本身就这样」",
)
args = ap.parse_args()
if args.cookie:
cookies = dict(
pair.split("=", 1) for pair in args.cookie.split("; ") if "=" in pair
)
else:
cookies = await cookies_from_cdp()
print(f" cookie 条数 {len(cookies)},名字: {sorted(cookies)}")
if not cookies.get("a1"):
print(" ✗ 没有 a1 —— 签名必须用它,无法继续")
return 2
query = build_query(args.start_days_ago, args.end_days_ago)
cookie_header = "; ".join(f"{k}={v}" for k, v in cookies.items())
headers_common = {
"user-agent": USER_AGENT,
"accept": "application/json, text/plain, */*",
"origin": CREATOR_ORIGIN,
"referer": f"{CREATOR_ORIGIN}/statistics/data-analysis",
"accept-language": "zh-CN,zh;q=0.9",
}
if not args.no_cookie_header:
headers_common["cookie"] = cookie_header
# Three signing variants: the reference bakes the query into the signed string,
# but the exact form is not documented beyond an example with no query at all.
variants = {
"path+q(参考实现写法)": f"url={NOTE_LIST_PATH}?{query}",
"path only": f"url={NOTE_LIST_PATH}",
"裸 path+query(无 url= 前缀)": f"{NOTE_LIST_PATH}?{query}",
}
url = f"{CREATOR_ORIGIN}{NOTE_LIST_PATH}?{query}"
async with httpx.AsyncClient(timeout=25, follow_redirects=False) as client:
for label, api in variants.items():
signature = sign_xyw(api, cookies["a1"], app_id=args.app_id)
headers = {**headers_common, **signature}
try:
response = await client.get(url, headers=headers)
except Exception as exc: # noqa: BLE001
print(f" [{label}] 请求异常: {exc.__class__.__name__}: {exc}")
continue
print(f"\n [{label}]")
print(f" HTTP {response.status_code} {classify(response.status_code, response.text)}")
body = response.text or ""
if body:
print(f" 响应前 160 字符: {body[:160]!r}")
if args.show_body and body:
print(f" 完整响应: {body[:800]}")
print("\n 提示:若三种都返回 406,再试 --app-id xhs-pc-web。")
return 0
if __name__ == "__main__":
import asyncio
sys.exit(asyncio.run(main()))
+144
View File
@@ -0,0 +1,144 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""Phase 0, second half: read the real request off the real page.
The signed-request probe proved the signature is forgeable and that the main-site
cookie authenticates the creator backend -- but the endpoint it called returned an
envelope with no payload, which no real list endpoint does. That path came from a
third-party repo and may simply be stale.
So stop guessing at paths and watch the page. This opens the data-analysis page in
the browser that is already signed in, records every creator API call the page
itself makes, and prints each request's URL, method and the *shape* of the
response. The page's own requests are the ground truth.
Read-only: it navigates and observes. Nothing is submitted.
"""
import argparse
import asyncio
import json
import sys
from collections import Counter
DEFAULT_URL = "https://creator.xiaohongshu.com/statistics/data-analysis"
def shape(value, depth: int = 0) -> str:
"""Describe a JSON value's structure without dumping its data."""
if depth > 3:
return "…"
if isinstance(value, dict):
if not value:
return "{}"
inner = ", ".join(f"{k}: {shape(v, depth + 1)}" for k, v in list(value.items())[:12])
return "{" + inner + "}"
if isinstance(value, list):
if not value:
return "[]"
return f"[{len(value)} × {shape(value[0], depth + 1)}]"
if isinstance(value, str):
# Short values are shown as themselves -- the interesting ones here are
# status codes, roles and permission names, and "str(14)" tells you
# nothing. Long ones are almost always ids or urls, so only their length.
return repr(value) if len(value) <= 40 else f"str({len(value)})"
if isinstance(value, bool):
return str(value)
if isinstance(value, (int, float)):
return str(value)
return type(value).__name__
async def main() -> int:
from playwright.async_api import async_playwright
ap = argparse.ArgumentParser()
ap.add_argument("--url", default=DEFAULT_URL)
ap.add_argument("--wait", type=int, default=25, help="观察窗口(秒)")
ap.add_argument("--filter", default="/api/galaxy", help="只记录 URL 含此串的请求")
args = ap.parse_args()
playwright = await async_playwright().start()
seen: Counter[str] = Counter()
details: list[str] = []
try:
browser = await playwright.chromium.connect_over_cdp(
"http://127.0.0.1:9222", timeout=15000
)
context = browser.contexts[0]
async def on_response(response):
url = response.url
if args.filter not in url:
return
key = url.split("?")[0]
seen[key] += 1
if seen[key] > 1:
return
request = response.request
line = [f"\n {request.method} {key}"]
line.append(f" HTTP {response.status}")
query = url.split("?", 1)[1] if "?" in url else ""
if query:
line.append(f" 查询串: {query[:300]}")
post = request.post_data
if post:
line.append(f" POST body: {post[:300]}")
try:
payload = await response.json()
line.append(f" 响应结构: {shape(payload)[:600]}")
except Exception:
try:
text = await response.text()
line.append(f" 响应(非JSON)前 200: {text[:200]!r}")
except Exception as exc: # noqa: BLE001
line.append(f" 响应不可读: {exc.__class__.__name__}")
details.append("\n".join(line))
context.on("response", on_response)
page = await context.new_page()
try:
await page.goto(args.url, wait_until="domcontentloaded", timeout=45000)
print(f" 落地 URL: {page.url[:120]}")
title = await page.title()
print(f" 标题: {title[:80]!r}")
# A creator console that is genuinely reachable renders its shell; a
# login gate does not. This is the cheapest "are we in?" signal.
for label, selector in (
("登录表单", "//input[@type='password']"),
("扫码登录", "//*[contains(@class,'qrcode') or contains(@class,'qr-code')]"),
):
if await page.locator(selector).count() > 0:
print(f" ★ 页面上出现「{label}」—— 这个账号似乎没有创作者后台会话")
print(f"\n 观察 {args.wait} 秒,记录页面自己发的请求…")
await asyncio.sleep(args.wait)
finally:
try:
await page.close()
except Exception:
pass
context.remove_listener("response", on_response)
if not details:
print("\n ★ 没有捕获到任何匹配的请求 —— 页面很可能停在登录页,没有发出数据请求")
for block in details:
print(block)
print(f"\n 捕获到的接口(去重): {len(seen)}")
for key, count in seen.most_common():
print(f" ×{count} {key}")
finally:
await playwright.stop()
return 0
if __name__ == "__main__":
sys.exit(asyncio.run(main()))
+8
View File
@@ -7,6 +7,8 @@
# 昵称保留但做中间脱敏)。本模块提供匿名化与脱敏工具。
import hashlib
import config
def anonymize_user_id(user_id) -> str:
"""把原始用户 ID 转成匿名哈希,用于内容/评论记录的创作者分组,
@@ -25,10 +27,16 @@ def mask_nickname(name) -> str:
- 长度 == 2:首字 + "*"
- 长度 >= 3:首字 + "***" + 尾字
这样既保留教学分析所需的内容归属语义,又无法据昵称定位到真人。
**是否启用由 config.MASK_NICKNAME 决定,本仓库默认关闭**(原样返回)。
脱敏是有损的,撞名很常见 —— 详见 base_config 里那一段的说明。
开关读的是模块属性而不是导入时的值,这样测试可以 monkeypatch 它。
"""
if name is None:
return ""
s = str(name)
if not getattr(config, "MASK_NICKNAME", False):
return s
if len(s) <= 1:
return "*"
if len(s) == 2:
+3 -1
View File
@@ -5,13 +5,14 @@ import { Sidebar } from '@/components/layout/Sidebar'
import { MainContent } from '@/components/layout/MainContent'
import { CrawlerConfigPanel } from '@/components/config/CrawlerConfigPanel'
import { MonitorDashboard } from '@/components/monitor/MonitorDashboard'
import { OperationView } from '@/components/creator/OperationView'
import { ReportView } from '@/components/monitor/ReportView'
import { SettingsView } from '@/components/settings/SettingsView'
import { Login } from '@/components/auth/Login'
import { EnvironmentCheck, isEnvChecked } from '@/components/env/EnvironmentCheck'
import { authApi, setUnauthorizedHandler } from '@/lib/api'
export type AppView = 'crawler' | 'monitor' | 'report' | 'settings'
export type AppView = 'crawler' | 'monitor' | 'operation' | 'report' | 'settings'
function App() {
// null = still probing. Rendering the app while unknown would briefly mount
@@ -91,6 +92,7 @@ function App() {
</>
)}
{view === 'monitor' && <MonitorDashboard />}
{view === 'operation' && <OperationView />}
{view === 'report' && <ReportView />}
{view === 'settings' && <SettingsView onNavigate={setView} />}
</div>
@@ -0,0 +1,184 @@
import { useEffect, useState } from 'react'
import { useQueryClient } from '@tanstack/react-query'
import { AlertTriangle, CheckCircle2, Loader2, QrCode, RefreshCw, X } from 'lucide-react'
import { Button } from '@/components/ui/button'
import {
Dialog,
DialogContent,
DialogDescription,
DialogHeader,
DialogTitle,
} from '@/components/ui/dialog'
import {
useCancelCreatorLogin,
useCreatorLoginStatus,
useStartCreatorLogin,
} from '@/hooks/useCreator'
/**
* 扫码新增一个运营账号。
*
* 后端给每个账号开一个**临时浏览器上下文**,扫完取出 cookie 就丢弃,所以:
* 登第二个账号不会把第一个顶掉,也不会影响监控那个登录态。
*/
export function AddAccountDialog({
open,
onOpenChange,
}: {
open: boolean
onOpenChange: (open: boolean) => void
}) {
const queryClient = useQueryClient()
const [polling, setPolling] = useState(false)
const { data: state } = useCreatorLoginStatus(polling)
const start = useStartCreatorLogin()
const cancel = useCancelCreatorLogin()
const status = state?.status ?? 'idle'
useEffect(() => {
if (!open) return
if (status === 'waiting' || status === 'idle') return
setPolling(false)
if (status !== 'success') return
queryClient.invalidateQueries({ queryKey: ['creatorAccounts'] })
// 自动关掉:停在二维码上会让人以为没成功,而账号其实已经进了列表。
const timer = setTimeout(() => onOpenChange(false), 1600)
return () => clearTimeout(timer)
}, [status, open, queryClient, onOpenChange])
// 关闭时统一收尾:拆掉后端那个临时上下文(别让它挂在操作者的 Chrome 里),
// 并清掉后端记住的结果 —— 否则下次打开弹窗会立刻显示上一次的成功状态。
useEffect(() => {
if (open) return
setPolling(false)
cancel.mutate()
// 只在开关变化时跑,cancel/依赖本身不该触发它。
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [open])
const begin = () => {
setPolling(true)
start.mutate()
}
const busy = start.isPending || cancel.isPending
return (
<Dialog open={open} onOpenChange={onOpenChange}>
<DialogContent className="max-w-md">
<DialogHeader>
<DialogTitle className="font-mono">新增运营账号</DialogTitle>
<DialogDescription className="font-mono text-xs">
用小红书 App 扫码登录要添加的账号。每个账号使用独立的临时会话,互不影响。
</DialogDescription>
</DialogHeader>
<div className="space-y-3 py-2">
{status === 'waiting' && state?.image && (
<div className="flex flex-col items-center gap-3">
{/* 二维码直接来自页面,是个 data: URL —— 不经过第三方,也不在服务器上落盘 */}
<img
src={state.image}
alt="登录二维码"
className="w-44 h-44 rounded-md border border-cyber-border-DEFAULT bg-white p-1"
/>
<p className="text-[11px] font-mono text-cyber-text-muted">
剩余 <span className="text-cyber-neon-cyan">{state.expires_in}</span> 秒
</p>
<Button
variant="ghost"
size="sm"
disabled={busy}
onClick={() => {
setPolling(false)
cancel.mutate()
}}
>
<X className="w-3 h-3 mr-1" />
取消
</Button>
</div>
)}
{status === 'waiting' && !state?.image && (
<p className="flex items-center justify-center gap-2 py-6 text-[11px] font-mono text-cyber-text-muted">
<Loader2 className="w-3 h-3 animate-spin" />
正在从浏览器取二维码…
</p>
)}
{status === 'success' && (
<div className="space-y-3">
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-green leading-relaxed">
<CheckCircle2 className="w-3 h-3 mt-0.5 shrink-0" />
{state?.message || '账号已添加'}
</p>
{state?.account && (
<div className="flex items-center gap-2 rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 px-3 py-2">
{state.account.avatar && (
<img
src={state.account.avatar}
alt=""
referrerPolicy="no-referrer"
className="w-8 h-8 rounded-full object-cover"
/>
)}
<div className="min-w-0">
<div className="truncate font-mono text-xs text-cyber-text-primary">
{state.account.nickname}
</div>
<div className="text-[10px] font-mono text-cyber-text-muted">
小红书号 {state.account.red_id || '—'}
</div>
</div>
</div>
)}
{state?.account?.permission_tip && (
<p className="text-[10px] font-mono text-cyber-neon-orange leading-relaxed">
{state.account.permission_tip}
</p>
)}
<Button size="sm" onClick={() => onOpenChange(false)}>
完成
</Button>
</div>
)}
{(status === 'error' || status === 'expired') && (
<div className="space-y-2">
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-orange leading-relaxed">
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
{state?.message || '获取二维码失败'}
</p>
<Button variant="outline" size="sm" disabled={busy} onClick={begin}>
<RefreshCw className="w-3 h-3 mr-1" />
重新获取
</Button>
</div>
)}
{status === 'idle' && (
<div className="space-y-2">
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
需要先在「系统设置」里打开 接管已有 Chrome(CDP),
并确保那台 Chrome 正以 9222 端口运行。
</p>
<Button size="sm" disabled={busy} onClick={begin}>
{start.isPending ? (
<Loader2 className="w-3 h-3 mr-1 animate-spin" />
) : (
<QrCode className="w-3 h-3 mr-1" />
)}
获取二维码
</Button>
</div>
)}
</div>
</DialogContent>
</Dialog>
)
}
@@ -0,0 +1,424 @@
import { useEffect, useState } from 'react'
import {
AlertTriangle,
Briefcase,
Clock,
Plus,
RefreshCw,
Trash2,
UserRound,
} from 'lucide-react'
import { Badge } from '@/components/ui/badge'
import { Button } from '@/components/ui/button'
import {
Select,
SelectContent,
SelectItem,
SelectTrigger,
SelectValue,
} from '@/components/ui/select'
import {
useCheckCreatorAccount,
useCreatorAccount,
useCreatorAccounts,
useDeleteCreatorAccount,
useSyncCreatorAccount,
} from '@/hooks/useCreator'
import { useCurrentPlatform } from '@/hooks/usePlatform'
import { formatCount, formatDateTime, formatRelative } from '@/lib/monitorFormat'
import type { CreatorAccount, CreatorNote, CreatorPermissionStatus } from '@/types/creator'
import { AddAccountDialog } from './AddAccountDialog'
const PERMISSION_LABEL: Record<CreatorPermissionStatus, { text: string; tone: string }> = {
active: { text: '数据已开通', tone: 'text-cyber-neon-green' },
// 实测中最常见的一种,必须和"没有权限"分开说 —— 否则用户以为采集坏了,
// 其实只是在等次日生效。
pending: { text: '权限待生效', tone: 'text-cyber-neon-orange' },
missing: { text: '无数据权限', tone: 'text-cyber-neon-pink' },
unknown: { text: '权限未知', tone: 'text-cyber-text-muted' },
}
// 后台按**发布时间**筛选,所以这是"作品发布距今多少天",不是"最近多少天的数据"。
// 上限 730 天与后端校验一致。
const RANGE_OPTIONS = [
{ value: '7', label: '近 7 天' },
{ value: '30', label: '近 30 天' },
{ value: '90', label: '近 90 天' },
{ value: '180', label: '近 180 天' },
{ value: '365', label: '近 1 年' },
{ value: '730', label: '近 2 年' },
]
const SUMMARY_TILES: Array<{ key: keyof CreatorNote; label: string }> = [
{ key: 'exposure', label: '曝光' },
{ key: 'views', label: '观看' },
{ key: 'likes', label: '点赞' },
{ key: 'comments', label: '评论' },
{ key: 'favorites', label: '收藏' },
{ key: 'shares', label: '分享' },
{ key: 'new_followers', label: '涨粉' },
]
function formatRate(value: number | null): string {
return value === null ? '—' : `${value.toFixed(1)}%`
}
function formatSeconds(value: number | null): string {
if (value === null) return '—'
if (value < 60) return `${Math.round(value)} 秒`
return `${Math.floor(value / 60)} 分 ${Math.round(value % 60)} 秒`
}
/** 左栏里的一行账号。选中项高亮,与监控的任务卡片同一套语言。 */
function AccountRow({
account,
selected,
onSelect,
}: {
account: CreatorAccount
selected: boolean
onSelect: () => void
}) {
const permission = PERMISSION_LABEL[account.permission_status]
return (
<button
onClick={onSelect}
className={`w-full flex items-center gap-2.5 rounded-lg border px-2.5 py-2 text-left transition-colors ${
selected
? 'border-cyber-neon-cyan/50 bg-cyber-neon-cyan/10'
: 'border-cyber-border-subtle bg-cyber-bg-tertiary/40 hover:border-cyber-border-DEFAULT'
}`}
>
{account.avatar ? (
<img
src={account.avatar}
alt=""
// 小红书图床对带外部 Referer 的请求返回 403,头像同理。
referrerPolicy="no-referrer"
className="w-8 h-8 rounded-full object-cover flex-shrink-0 bg-cyber-bg-tertiary"
/>
) : (
<div className="w-8 h-8 rounded-full bg-cyber-bg-tertiary flex-shrink-0" />
)}
<div className="min-w-0 flex-1">
<div className="flex items-center gap-1.5">
<span className="truncate font-mono text-[11px] text-cyber-text-primary">
{account.nickname || '未命名账号'}
</span>
{account.status === 'expired' && (
<Badge variant="warning" className="text-[9px] px-1 py-0 flex-shrink-0">
失效
</Badge>
)}
</div>
<div className="mt-0.5 flex items-center gap-2 text-[9px] font-mono text-cyber-text-muted">
<span className={permission.tone}>{permission.text}</span>
<span>作品 {account.note_count}</span>
</div>
</div>
</button>
)
}
/** 右栏:所选账号的数据。 */
function AccountPanel({ accountId }: { accountId: number }) {
const { data, isLoading } = useCreatorAccount(accountId)
const sync = useSyncCreatorAccount()
const check = useCheckCreatorAccount()
const remove = useDeleteCreatorAccount()
// 拉取范围是**同步参数**,不是展示筛选:数据按这个范围取回来存库。
// 默认 90 天,与后端的默认值一致。
const [days, setDays] = useState('90')
// 选择器对齐到上次实际用的范围。不对齐的话,上次同步了 1 年、这次点同步会悄悄
// 缩回 90 天,而界面上没有任何迹象。
const lastDays = data?.account.last_sync_days ?? 0
useEffect(() => {
if (lastDays > 0) setDays(String(lastDays))
}, [lastDays])
if (isLoading || !data) {
return <p className="py-10 text-center text-[11px] font-mono text-cyber-text-muted">加载中…</p>
}
const { account, notes, summary } = data
const permission = PERMISSION_LABEL[account.permission_status]
return (
<div className="flex flex-col overflow-hidden h-full">
<div className="flex items-center gap-2 px-3 py-2.5 border-b border-cyber-border-subtle flex-shrink-0">
{account.avatar && (
<img
src={account.avatar}
alt=""
referrerPolicy="no-referrer"
className="w-7 h-7 rounded-full object-cover"
/>
)}
<span className="font-mono text-xs text-cyber-text-primary">
{account.nickname || '未命名账号'}
</span>
<span className="font-mono text-[10px] text-cyber-text-muted">
小红书号 {account.red_id || '—'}
</span>
<span className="ml-1 text-[10px] font-mono text-cyber-text-muted">
上次同步 {account.last_synced_at ? formatRelative(account.last_synced_at) : '从未'}
{/* 当前列出的作品就是上一次同步范围里的产物,所以要说清是哪一段 ——
否则只看"同步过了"分不清覆盖了多少。 */}
{account.last_sync_days > 0 && ` · 覆盖近 ${account.last_sync_days} 天发布的`}
</span>
<div className="ml-auto flex items-center gap-2">
<Button
variant="outline"
size="sm"
disabled={check.isPending}
onClick={() => check.mutate(account.id)}
>
<RefreshCw className={`w-3 h-3 mr-1 ${check.isPending ? 'animate-spin' : ''}`} />
检测
</Button>
<Select value={days} onValueChange={setDays}>
<SelectTrigger className="h-7 w-[104px] text-[10px] font-mono">
<SelectValue />
</SelectTrigger>
<SelectContent>
{RANGE_OPTIONS.map((option) => (
<SelectItem key={option.value} value={option.value} className="text-[11px]">
{option.label}
</SelectItem>
))}
</SelectContent>
</Select>
<Button
size="sm"
disabled={sync.isPending}
onClick={() => sync.mutate({ id: account.id, days: Number(days) })}
>
<RefreshCw className={`w-3 h-3 mr-1 ${sync.isPending ? 'animate-spin' : ''}`} />
同步数据
</Button>
<Button variant="ghost" size="sm" onClick={() => remove.mutate(account.id)}>
<Trash2 className="w-3 h-3" />
</Button>
</div>
</div>
<div className="flex-1 overflow-y-auto terminal-scroll px-3 py-3 space-y-3">
{/* 只要后台有话要说就显示,**包括已经开通的时候**。
实测开通后的原话是「数据正在更新中,请耐心等待」—— 若只在未开通时才显示,
就会一边写着"数据已开通"、一边列不出作品,自相矛盾。 */}
{account.permission_tip && (
<div
className={`flex items-start gap-2 rounded-md border px-3 py-2 ${
account.permission_status === 'active'
? 'border-cyber-border-DEFAULT bg-cyber-bg-tertiary/40'
: 'border-cyber-neon-orange/30 bg-cyber-neon-orange/5'
}`}
>
<Clock
className={`w-3.5 h-3.5 mt-0.5 flex-shrink-0 ${
account.permission_status === 'active'
? 'text-cyber-text-muted'
: 'text-cyber-neon-orange'
}`}
/>
<div className="text-[11px] font-mono leading-relaxed">
<span className={permission.tone}>{permission.text}</span>
<span className="ml-2 text-cyber-text-secondary">{account.permission_tip}</span>
<div className="mt-0.5 text-cyber-text-muted">
{account.permission_status === 'active'
? '权限已开通,但后台的数据可能还在准备中 —— 稍后重新同步即可。'
: '数据权限由创作者后台按天开通。在此之前同步会成功但返回 0 条 —— 那是正常的,不是采集失败。'}
</div>
</div>
</div>
)}
{account.last_error && (
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-pink leading-relaxed">
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
{account.last_error}
</p>
)}
<div className="grid grid-cols-4 gap-2 xl:grid-cols-7">
{SUMMARY_TILES.map((tile) => (
<div
key={tile.key}
className="rounded-lg border border-cyber-border-subtle bg-cyber-bg-tertiary/40 px-3 py-2"
>
<div className="text-[10px] font-mono text-cyber-text-muted">{tile.label}</div>
<div className="mt-0.5 font-mono text-sm text-cyber-text-primary">
{formatCount(summary[tile.key as keyof typeof summary] ?? 0)}
</div>
</div>
))}
</div>
{notes.length === 0 ? (
<p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">
{account.permission_status === 'active'
? '这个账号在所选时间范围内没有作品,或还没同步过 —— 点「同步数据」试试'
: '数据权限生效后即可看到作品数据'}
</p>
) : (
<div className="overflow-x-auto">
<table className="w-full text-xs font-mono">
<thead className="sticky top-0 bg-cyber-bg-tertiary">
<tr className="text-[10px] text-cyber-text-secondary">
<th className="px-2 py-2 text-left font-normal">作品</th>
{SUMMARY_TILES.map((tile) => (
<th key={tile.key} className="px-2 py-2 text-right font-normal">
{tile.label}
</th>
))}
<th className="px-2 py-2 text-right font-normal">点击率</th>
<th className="px-2 py-2 text-right font-normal">均看时长</th>
<th className="px-2 py-2 text-right font-normal">完播率</th>
</tr>
</thead>
<tbody>
{notes.map((note) => (
<tr key={note.note_id} className="border-t border-cyber-border-subtle">
<td className="max-w-[280px] px-2 py-2">
<div className="truncate text-cyber-text-primary" title={note.title}>
{note.title || note.note_id}
</div>
<div className="text-[9px] text-cyber-text-muted">
{note.publish_time ? formatDateTime(note.publish_time) : '—'}
</div>
</td>
{SUMMARY_TILES.map((tile) => (
<td key={tile.key} className="px-2 py-2 text-right text-cyber-text-primary">
{formatCount(note[tile.key] as number | null)}
</td>
))}
<td className="px-2 py-2 text-right text-cyber-text-secondary">
{formatRate(note.cover_ctr)}
</td>
<td className="px-2 py-2 text-right text-cyber-text-secondary">
{formatSeconds(note.avg_watch_seconds)}
</td>
<td className="px-2 py-2 text-right text-cyber-text-secondary">
{formatRate(note.completion_rate)}
</td>
</tr>
))}
</tbody>
</table>
</div>
)}
</div>
</div>
)
}
/**
* 运营:管理自己的小红书账号。
*
* 与「监控」并列而非其子视图,布局也照搬它:左边账号列表、右边数据面板。
* 两者关注的东西不同 —— 监控抓公开数据(点赞/收藏/评论/分享)的每轮差分,
* 这里是创作者后台的运营指标(曝光/观看/完播率/涨粉),凭据与采集方式都不同。
*/
export function OperationView() {
const { platform, capability } = useCurrentPlatform()
const [selectedId, setSelectedId] = useState<number | null>(null)
const [addOpen, setAddOpen] = useState(false)
// 有同步在后端跑时列表要自己刷新,否则状态永远是旧的。
const { data: accounts, isLoading } = useCreatorAccounts(true)
// 运营只对小红书成立:它读的是小红书创作者后台。不挡住的话,切到抖音时这里照样
// 列出小红书的账号 —— 看起来像是"抖音账号",其实平台维度根本没参与。
if (platform !== 'xhs') {
const label = capability?.label ?? platform
return (
<div className="flex-1 flex items-center justify-center overflow-y-auto terminal-scroll">
<div className="max-w-lg w-full mx-4 rounded-lg glass-panel float-panel p-6 space-y-3">
<div className="flex items-center gap-3">
<div className="p-2 rounded-md border border-cyber-neon-orange/40 bg-cyber-neon-orange/10">
<Briefcase className="w-5 h-5 text-cyber-neon-orange" />
</div>
<div>
<h2 className="font-mono text-sm text-cyber-text-primary">{label} 没有运营模块</h2>
<p className="text-[11px] font-mono text-cyber-text-muted">
运营读的是创作者后台的数据,目前只接入了小红书
</p>
</div>
</div>
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
小红书之外,这里没有可用的后台接口 —— 能采的是公开数据,那属于「监控」。
切回小红书即可看到运营账号。
</p>
</div>
</div>
)
}
// 没选过就默认第一个,和监控一样 —— 右栏不该一开始是空的。
const selected =
accounts?.find((account) => account.id === selectedId) ??
(accounts && accounts.length > 0 ? accounts[0] : null)
return (
<div className="flex-1 flex flex-col gap-3 overflow-hidden min-h-0 relative z-10">
<div className="flex-1 flex gap-3 overflow-hidden min-h-0">
{/* 左:账号列表 */}
<div className="w-[300px] flex-shrink-0 flex flex-col gap-3 overflow-hidden">
<div className="flex items-center justify-between flex-shrink-0">
<span className="font-mono text-xs text-cyber-text-primary">运营账号</span>
<Button size="sm" onClick={() => setAddOpen(true)}>
<Plus className="w-3 h-3 mr-1" />
新增
</Button>
</div>
<div className="flex-1 overflow-y-auto terminal-scroll space-y-2 pr-1">
{isLoading ? (
<p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">
加载中…
</p>
) : !accounts || accounts.length === 0 ? (
<div className="flex flex-col items-center gap-2 py-10">
<UserRound className="w-7 h-7 text-cyber-text-muted" />
<p className="text-center text-[11px] font-mono text-cyber-text-muted">
还没有运营账号
<br />
扫码添加一个
</p>
</div>
) : (
accounts.map((account) => (
<AccountRow
key={account.id}
account={account}
selected={selected?.id === account.id}
onSelect={() => setSelectedId(account.id)}
/>
))
)}
</div>
</div>
{/* 右:所选账号的数据 */}
<div className="flex-1 flex flex-col overflow-hidden min-w-0 rounded-lg glass-panel float-panel">
{selected ? (
<AccountPanel accountId={selected.id} />
) : (
<div className="flex-1 flex items-center justify-center">
<p className="text-[11px] font-mono text-cyber-text-muted">
从左边选一个账号,或先扫码添加
</p>
</div>
)}
</div>
</div>
<AddAccountDialog open={addOpen} onOpenChange={setAddOpen} />
</div>
)
}
+2 -1
View File
@@ -1,5 +1,5 @@
import { useState } from 'react'
import { Bug, Wifi, BarChart3, Cog, LogOut, Radar, Settings, Terminal } from 'lucide-react'
import { Briefcase, Bug, Wifi, BarChart3, Cog, LogOut, Radar, Settings, Terminal } from 'lucide-react'
import { useTranslation } from 'react-i18next'
import { Badge } from '@/components/ui/badge'
import { SystemSettingsDialog } from '@/components/settings/SystemSettingsDialog'
@@ -19,6 +19,7 @@ interface SidebarProps {
const NAV_ITEMS: Array<{ value: AppView; label: string; icon: typeof Terminal }> = [
{ value: 'crawler', label: '采集', icon: Terminal },
{ value: 'monitor', label: '监控', icon: Radar },
{ value: 'operation', label: '运营', icon: Briefcase },
{ value: 'report', label: '报表', icon: BarChart3 },
{ value: 'settings', label: '设置', icon: Settings },
]
@@ -40,7 +40,7 @@ export function UnwiredPlatformNotice({ area }: { area: string }) {
{capability.label} 的{area}尚未接通
</h2>
<p className="text-[11px] font-mono text-cyber-text-muted">
{capability.label}的爬虫模块是支持的,但监控层目前只接通了小红书
{capability.label}的爬虫模块是支持的,但监控层目前只接通了小红书与抖音
</p>
</div>
</div>
@@ -94,9 +94,9 @@ export function UnwiredPlatformNotice({ area }: { area: string }) {
<div className="flex items-start gap-2 text-[10px] font-mono text-cyber-text-muted">
<MonitorSmartphone className="w-3.5 h-3.5 mt-0.5 flex-shrink-0" />
<p>
切换到小红书即可正常使用。若要接通该平台,需要在
<span className="text-cyber-neon-cyan"> runner / ingest / 目标解析 </span>
三处补上平台适配(目前这三处是硬编码小红书的)。
切换到已接通的平台即可正常使用。若要接通该平台,需要在
<span className="text-cyber-neon-cyan"> adapters.py </span>
里补一份适配:产物目录名、jsonl 字段名、目标链接形态。
</p>
</div>
</div>
+95 -23
View File
@@ -7,6 +7,7 @@ import {
LayoutList,
Rows3,
ThumbsUp,
Users,
} from 'lucide-react'
import { Badge } from '@/components/ui/badge'
@@ -24,9 +25,11 @@ import {
useMonitorCommentsGrouped,
} from '@/hooks/useMonitor'
import { monitorApi } from '@/lib/api'
import { formatRelative } from '@/lib/monitorFormat'
import { formatDate, formatRelative } from '@/lib/monitorFormat'
import type { CommentBucket, MonitorComment } from '@/types/monitor'
import { NoteCover } from './NoteCover'
interface CommentsFeedProps {
taskId: number | null
}
@@ -49,22 +52,7 @@ function NoteBadge({
}) {
return (
<div className="flex items-center gap-2 min-w-0">
{cover ? (
<img
src={cover}
alt=""
loading="lazy"
className={`rounded object-cover bg-cyber-bg-tertiary flex-shrink-0 ${
compact ? 'w-6 h-8' : 'w-8 h-10'
}`}
/>
) : (
<div
className={`rounded bg-cyber-bg-tertiary flex-shrink-0 ${
compact ? 'w-6 h-8' : 'w-8 h-10'
}`}
/>
)}
<NoteCover src={cover} size={compact ? 'sm' : 'md'} />
<span className="truncate text-[10px] font-mono text-cyber-text-secondary" title={title || noteId}>
{title || noteId}
</span>
@@ -149,6 +137,15 @@ function CollapsibleGroup({
noteId={bucket.note_id}
/>
</div>
{/* 作品那一层光有标题不够 —— 同名的作品不少,发布日期能帮着认。 */}
{bucket.published_at && (
<span
className="text-[10px] font-mono text-cyber-text-muted flex-shrink-0"
title="作品发布时间"
>
{formatDate(bucket.published_at)}
</span>
)}
<Badge variant="outline" className="text-[10px] px-1.5 py-0 flex-shrink-0">
{bucket.comments.length} 条
</Badge>
@@ -169,6 +166,78 @@ function CollapsibleGroup({
)
}
/** 把作品分组按博主归并 —— 三级视图的第一级。 */
function groupByCreator(buckets: CommentBucket[]) {
const byCreator = new Map<string, { key: string; name: string; buckets: CommentBucket[] }>()
for (const bucket of buckets) {
const key = bucket.creator_hash || '__unknown__'
if (!byCreator.has(key)) {
byCreator.set(key, { key, name: bucket.creator_name || '', buckets: [] })
}
byCreator.get(key)!.buckets.push(bucket)
}
return [...byCreator.values()]
}
/**
* 一级:博主。二级作品、三级评论都收在它下面。
*
* 分组键是 `creator_hash`(爬虫刻意不落原始 user_id,这是唯一稳定的创作者标识),
* 显示名用已脱敏的昵称。默认展开 —— 折叠的默认值不该把数据藏起来。
*/
function CreatorGroup({
name,
buckets,
expandedNotes,
onToggleNote,
}: {
name: string
buckets: CommentBucket[]
expandedNotes: Set<string>
onToggleNote: (noteId: string) => void
}) {
const [open, setOpen] = useState(true)
const commentCount = buckets.reduce((sum, bucket) => sum + bucket.comments.length, 0)
return (
<div className="rounded-lg border border-cyber-border-DEFAULT overflow-hidden">
<button
onClick={() => setOpen(!open)}
className="w-full flex items-center gap-2 px-3 py-2 bg-cyber-bg-elevated/50 hover:bg-cyber-bg-elevated/80 transition-colors text-left"
>
{open ? (
<ChevronDown className="w-3.5 h-3.5 text-cyber-text-muted flex-shrink-0" />
) : (
<ChevronRight className="w-3.5 h-3.5 text-cyber-text-muted flex-shrink-0" />
)}
<Users className="w-3.5 h-3.5 text-cyber-neon-cyan flex-shrink-0" />
<span className="font-mono text-xs text-cyber-text-primary truncate">
{name || '未知博主'}
</span>
<span className="text-[10px] font-mono text-cyber-text-muted flex-shrink-0">
{buckets.length} 篇作品
</span>
<Badge variant="outline" className="text-[10px] px-1.5 py-0 flex-shrink-0 ml-auto">
{commentCount} 条评论
</Badge>
</button>
{open && (
<div className="p-2 space-y-2 bg-cyber-bg-secondary/30">
{buckets.map((bucket) => (
<CollapsibleGroup
key={bucket.note_id}
bucket={bucket}
open={expandedNotes.has(bucket.note_id)}
onToggle={() => onToggleNote(bucket.note_id)}
/>
))}
</div>
)}
</div>
)
}
export function CommentsFeed({ taskId }: CommentsFeedProps) {
const [grouped, setGrouped] = useState(true)
const [noteFilter, setNoteFilter] = useState<string>(ALL_NOTES)
@@ -280,12 +349,15 @@ export function CommentsFeed({ taskId }: CommentsFeedProps) {
</p>
) : grouped ? (
<div className="space-y-2">
{groups?.map((bucket) => (
<CollapsibleGroup
key={bucket.note_id}
bucket={bucket}
open={expanded.has(bucket.note_id)}
onToggle={() => toggleGroup(bucket.note_id)}
{/* 三级:博主 -> 作品 -> 评论。一个任务可以配多个博主,平铺作品的话
看不出哪条评论属于谁。 */}
{groupByCreator(groups ?? []).map((creator) => (
<CreatorGroup
key={creator.key}
name={creator.name}
buckets={creator.buckets}
expandedNotes={expanded}
onToggleNote={toggleGroup}
/>
))}
</div>
+4 -1
View File
@@ -31,8 +31,11 @@ export function CookiePanel() {
<div className="flex items-center justify-between gap-3">
<div className="flex items-center gap-2">
<KeyRound className="w-4 h-4 text-cyber-neon-cyan" />
{/* 标题要和扫码面板区分开:这个是"存进库里、每轮注入子进程"的 cookie,
那个是"写进浏览器 profile、CDP 模式复用"的登录态。两者是不同机制、
也都能同时算"登录着",都叫「登录态」就分不清在说哪一个。 */}
<span className="font-mono text-xs tracking-wider text-cyber-text-primary">
{label}登录态
{label} Cookie(定时任务用)
</span>
{status?.present ? (
<Badge variant="success" className="text-[10px]">
@@ -1,10 +1,11 @@
import { useState } from 'react'
import { Activity, BellRing, FileText, MessageSquare, Plus } from 'lucide-react'
import { Activity, BellRing, Download, FileText, MessageSquare, Plus } from 'lucide-react'
import { Badge } from '@/components/ui/badge'
import { Button } from '@/components/ui/button'
import { Tabs, TabsContent, TabsList, TabsTrigger } from '@/components/ui/tabs'
import { useMonitorOverview, useMonitorTasks } from '@/hooks/useMonitor'
import { monitorApi } from '@/lib/api'
import type { MonitorTask } from '@/types/monitor'
import { useCookieStatus, useWebhookStatus } from '@/hooks/useMonitor'
import { useCurrentPlatform } from '@/hooks/usePlatform'
@@ -166,6 +167,23 @@ export function MonitorDashboard() {
>
只看新增
</Button>
{/* 导出的是**这个任务**的全部作品,不是屏幕上这 200 条 —— 屏幕上那份是
为了好看才截断的,导出跟着截断就成了「导出来的比看到的少」。
走 window.open 而不是 blob:鉴权在 Cookie 上,浏览器自己会带上。 */}
<Button
variant="outline"
size="sm"
className="ml-auto"
onClick={() =>
window.open(
monitorApi.getExportUrl({ kind: 'notes', taskId: selectedTask.id }),
'_blank',
)
}
>
<Download className="w-3 h-3 mr-1" />
导出 CSV
</Button>
</div>
<NotesTable taskId={selectedTask.id} onlyNew={onlyNew} />
</TabsContent>
@@ -0,0 +1,41 @@
/**
* 作品封面缩略图。
*
* **图为什么曾经全是破图**:不是防盗链。小红书图床的地址**带签名、会过期** ——
* 路径里那段时间戳就是签发时刻。实测同一批图:
*
* 当天签发的地址 -> 200,带不带 Referer 都一样
* 隔天的地址 -> 403,带不带 Referer 都一样
*
* 所以 Referer 根本不是那个维度,改它是白改。真正的解法是后端把图**下载到本地**,
* 前端拿到的是 `/api/monitor/covers/{note_id}` —— 与签名无关,不会过期。
*
* `referrerPolicy` 保留着,是因为它仍然是个合理的默认(外链图不该把自己的地址
* 泄露给第三方),但**它不再是这里能正常显示的原因**。
*/
const SIZES = {
sm: 'w-6 h-8',
md: 'w-8 h-10',
lg: 'w-12 h-16',
} as const
export function NoteCover({
src,
size = 'md',
className = '',
}: {
src?: string
size?: keyof typeof SIZES
className?: string
}) {
const base = `rounded object-cover bg-cyber-bg-tertiary flex-shrink-0 ${SIZES[size]} ${className}`
// 占位块是必要的:作品在首轮采集前没有封面,若此时不占位,行高会随着封面陆续
// 到达而跳动。
if (!src) return <div className={base} aria-hidden />
return (
<img src={src} alt="" loading="lazy" referrerPolicy="no-referrer" className={base} />
)
}
+184 -95
View File
@@ -1,4 +1,4 @@
import { useMemo, useState } from 'react'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { useNoteSeries } from '@/hooks/useMonitor'
import { formatCount, formatDateTime } from '@/lib/monitorFormat'
@@ -12,13 +12,62 @@ const METRICS: Array<{ key: MetricKey; label: string }> = [
{ key: 'share_count', label: '分享' },
]
const VIEW_W = 600
/** 图表高度固定;宽度由容器量出来,见下方 attachContainer。 */
const VIEW_H = 160
const PAD_LEFT = 10
const PAD_RIGHT = 56
const PAD_TOP = 18
const PAD_LEFT = 44 // 左侧刻度栏,要放得下 "1.2万"
const PAD_RIGHT = 16
const PAD_TOP = 14
const PAD_BOTTOM = 22
const TICK_COUNT = 3
/** 把步长收敛到 1 / 2 / 5 × 10ⁿ,刻度才会落在好读的数上。 */
function niceNumber(value: number, round: boolean): number {
if (value <= 0) return 1
const exponent = Math.floor(Math.log10(value))
const fraction = value / 10 ** exponent
let nice: number
if (round) {
nice = fraction < 1.5 ? 1 : fraction < 3 ? 2 : fraction < 7 ? 5 : 10
} else {
nice = fraction <= 1 ? 1 : fraction <= 2 ? 2 : fraction <= 5 ? 5 : 10
}
return nice * 10 ** exponent
}
/**
* 给折线图选一个纵轴范围。
*
* **刻意不从 0 起。** 柱状图用「长度」编码数值,基线不为 0 比例就是错的;折线图用
* 「位置」编码,轴只需要如实框住数据 —— 这正是让 491 → 506 这段变化看得见的原因,
* 否则它会贴着 0–600 的底边变成一条直线。
*
* 代价是纵轴不再是 0,所以刻度必须落在左侧栏里、显示真实数值,让读图的人随时知道
* 范围是多少。这一点在上一轮已经修好。
*/
function niceAxis(values: number[]): { min: number; max: number; ticks: number[] } {
const dataMin = Math.min(...values)
const dataMax = Math.max(...values)
// 全平的一组数没有跨度可缩放,给它一个名义区间,让线落在图中而不是贴边。
const span = dataMax - dataMin || Math.max(1, Math.abs(dataMax) * 0.05 || 1)
const step = niceNumber(span / (TICK_COUNT - 1), true)
let min = Math.floor(dataMin / step) * step
let max = Math.ceil(dataMax / step) * step
if (min === max) {
// 取整后塌成一点会除零,撑开一档。
min -= step
max += step
}
const ticks: number[] = []
for (let value = min; value <= max + step / 2; value += step) {
ticks.push(Math.round(value))
}
return { min, max, ticks }
}
interface NoteTrendChartProps {
noteId: string
taskId: number | null
@@ -36,8 +85,25 @@ interface NoteTrendChartProps {
export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProps) {
const [metric, setMetric] = useState<MetricKey>('liked_count')
const [hoverIndex, setHoverIndex] = useState<number | null>(null)
const [width, setWidth] = useState(0)
const observer = useRef<ResizeObserver | null>(null)
const { data: series, isLoading } = useNoteSeries(noteId, taskId)
// 宽度是量出来的,不是拉伸出来的。原先靠 preserveAspectRatio="none" 把 600×160 的
// viewBox 横向撑满容器 —— 横纵缩放不一致,半径 4 的圆就被压成了椭圆。只有等比绘图,
// 圆才真的是圆。
const attachContainer = useCallback((node: HTMLDivElement | null) => {
observer.current?.disconnect()
if (!node) return
const measure = () => setWidth(Math.round(node.getBoundingClientRect().width))
const resizeObserver = new ResizeObserver(measure)
resizeObserver.observe(node)
observer.current = resizeObserver
measure()
}, [])
useEffect(() => () => observer.current?.disconnect(), [])
// A point with no parsed value is a gap, not a zero.
const points = useMemo(
() => (series ?? []).filter((point) => point[metric] !== null),
@@ -45,40 +111,59 @@ export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProp
)
const geometry = useMemo(() => {
if (points.length < 2) return null
if (points.length < 2 || width <= 0) return null
const values = points.map((point) => point[metric] as number)
const min = Math.min(...values)
const max = Math.max(...values)
// A flat series would divide by zero; give it a nominal band.
const span = max - min || 1
const { min, max, ticks } = niceAxis(values)
const innerW = VIEW_W - PAD_LEFT - PAD_RIGHT
const innerW = Math.max(1, width - PAD_LEFT - PAD_RIGHT)
const innerH = VIEW_H - PAD_TOP - PAD_BOTTOM
const yFor = (value: number) =>
PAD_TOP + innerH * (1 - (value - min) / (max - min))
const xy = points.map((point, index) => {
const value = point[metric] as number
return {
x: PAD_LEFT + (index / (points.length - 1)) * innerW,
y: PAD_TOP + innerH - ((value - min) / span) * innerH,
y: yFor(value),
value,
point,
}
})
return { xy, min, max }
}, [points, metric])
return { xy, ticks, yFor, innerW, innerH }
}, [points, metric, width])
if (isLoading) {
return <div className="p-4 text-[11px] font-mono text-cyber-text-muted">加载中…</div>
}
const metricLabel = METRICS.find((option) => option.key === metric)?.label ?? ''
const latest = points.length > 0 ? (points[points.length - 1][metric] as number) : null
// Fewer points get a bigger dot; a dense series would otherwise turn into a
// string of overlapping beads.
const dotRadius = points.length > 12 ? 2.5 : 4
const hitWidth = geometry ? Math.max(12, geometry.innerW / Math.max(1, points.length - 1)) : 12
return (
<div className="space-y-2">
<div className="flex items-center justify-between gap-3 flex-wrap">
<span className="font-mono text-xs text-cyber-text-primary">
指标趋势 · <span className="text-cyber-text-secondary">{noteTitle || noteId}</span>
</span>
<div className="flex items-baseline gap-2 min-w-0">
<span className="font-mono text-xs text-cyber-text-primary flex-shrink-0">指标趋势</span>
<span
className="font-mono text-[11px] text-cyber-text-secondary truncate"
title={noteTitle || noteId}
>
{noteTitle || noteId}
</span>
{latest !== null && (
<span className="font-mono text-xs text-cyber-neon-cyan flex-shrink-0">
{formatCount(latest)}
<span className="text-cyber-text-muted text-[10px] ml-1">{metricLabel}</span>
</span>
)}
</div>
<div className="flex items-center gap-1">
{METRICS.map((option) => (
<button
@@ -96,97 +181,101 @@ export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProp
</div>
</div>
{!geometry ? (
{points.length < 2 ? (
<p className="py-6 text-center text-[11px] font-mono text-cyber-text-muted">
至少需要两轮采集才能画出趋势(当前 {points.length} 个有效数据点)
</p>
) : (
<div className="relative">
<svg
viewBox={`0 0 ${VIEW_W} ${VIEW_H}`}
className="w-full h-40"
preserveAspectRatio="none"
onMouseLeave={() => setHoverIndex(null)}
>
{/* Recessive solid hairlines - never dashed. */}
{[0, 0.5, 1].map((ratio) => {
const y = PAD_TOP + (VIEW_H - PAD_TOP - PAD_BOTTOM) * ratio
return (
<div className="relative" ref={attachContainer}>
{geometry && (
<svg
width={width}
height={VIEW_H}
viewBox={`0 0 ${width} ${VIEW_H}`}
className="block"
onMouseLeave={() => setHoverIndex(null)}
>
{/* 刻度线画在 0 / 中值 / 上界上,标签就贴在对应位置 —— 不再浮在图面上,
也就不会出现两个数挤在一起读成一个数的情况。 */}
{geometry.ticks.map((tick) => {
const y = geometry.yFor(tick)
return (
<g key={tick}>
<line
x1={PAD_LEFT}
x2={width - PAD_RIGHT}
y1={y}
y2={y}
stroke="rgb(var(--cyber-text-muted) / 0.25)"
strokeWidth={1}
/>
<text
x={PAD_LEFT - 8}
y={y}
textAnchor="end"
dominantBaseline="middle"
fontSize={9}
fill="rgb(var(--cyber-text-muted))"
style={{ fontFamily: 'ui-monospace, SFMono-Regular, monospace' }}
>
{formatCount(tick)}
</text>
</g>
)
})}
{hoverIndex !== null && geometry.xy[hoverIndex] && (
<line
key={ratio}
x1={PAD_LEFT}
x2={VIEW_W - PAD_RIGHT}
y1={y}
y2={y}
stroke="rgb(var(--cyber-text-muted) / 0.25)"
x1={geometry.xy[hoverIndex].x}
x2={geometry.xy[hoverIndex].x}
y1={PAD_TOP}
y2={PAD_TOP + geometry.innerH}
stroke="rgb(var(--cyber-neon-cyan) / 0.5)"
strokeWidth={1}
vectorEffect="non-scaling-stroke"
/>
)
})}
)}
{hoverIndex !== null && geometry.xy[hoverIndex] && (
<line
x1={geometry.xy[hoverIndex].x}
x2={geometry.xy[hoverIndex].x}
y1={PAD_TOP}
y2={VIEW_H - PAD_BOTTOM}
stroke="rgb(var(--cyber-neon-cyan) / 0.5)"
strokeWidth={1}
vectorEffect="non-scaling-stroke"
<polyline
points={geometry.xy.map((node) => `${node.x},${node.y}`).join(' ')}
fill="none"
stroke="rgb(var(--cyber-neon-cyan))"
strokeWidth={2}
strokeLinejoin="round"
strokeLinecap="round"
/>
)}
<polyline
points={geometry.xy.map((node) => `${node.x},${node.y}`).join(' ')}
fill="none"
stroke="rgb(var(--cyber-neon-cyan))"
strokeWidth={2}
strokeLinejoin="round"
strokeLinecap="round"
vectorEffect="non-scaling-stroke"
/>
{geometry.xy.map((node, index) => (
<circle
key={index}
cx={node.x}
cy={node.y}
r={dotRadius}
fill="rgb(var(--cyber-neon-cyan))"
// 与环境色同色的描边,让重叠的点之间留出间隙。
stroke="rgb(var(--cyber-bg-primary))"
strokeWidth={2}
/>
))}
{/* Only the endpoint is labelled - a number on every point is noise. */}
<circle
cx={geometry.xy[geometry.xy.length - 1].x}
cy={geometry.xy[geometry.xy.length - 1].y}
r={4}
fill="rgb(var(--cyber-neon-cyan))"
stroke="rgb(var(--cyber-bg-primary))"
strokeWidth={2}
vectorEffect="non-scaling-stroke"
/>
{geometry.xy.map((node, index) => (
<rect
key={index}
x={node.x - hitWidth / 2}
y={PAD_TOP}
width={hitWidth}
height={geometry.innerH}
fill="transparent"
onMouseEnter={() => setHoverIndex(index)}
/>
))}
</svg>
)}
{geometry.xy.map((node, index) => (
<rect
key={index}
x={node.x - 6}
y={PAD_TOP}
width={12}
height={VIEW_H - PAD_TOP - PAD_BOTTOM}
fill="transparent"
onMouseEnter={() => setHoverIndex(index)}
/>
))}
</svg>
{/* Axis extremes live in text tokens, never the series colour. */}
<span className="absolute left-0 top-0 text-[9px] font-mono text-cyber-text-muted">
{formatCount(geometry.max)}
</span>
<span className="absolute left-0 bottom-5 text-[9px] font-mono text-cyber-text-muted">
{formatCount(geometry.min)}
</span>
<span className="absolute right-0 top-1/2 -translate-y-1/2 text-[11px] font-mono text-cyber-text-primary">
{formatCount(geometry.xy[geometry.xy.length - 1].value)}
</span>
{hoverIndex !== null && geometry.xy[hoverIndex] && (
{hoverIndex !== null && geometry?.xy[hoverIndex] && (
<div
className="absolute -top-1 px-2 py-1 rounded border border-cyber-border-DEFAULT bg-cyber-bg-elevated text-[10px] font-mono text-cyber-text-primary pointer-events-none whitespace-nowrap"
className="absolute -top-1 px-2 py-1 rounded border border-cyber-border-DEFAULT bg-cyber-bg-elevated text-[10px] font-mono text-cyber-text-primary pointer-events-none whitespace-nowrap z-10"
style={{
left: `${(geometry.xy[hoverIndex].x / VIEW_W) * 100}%`,
left: `${(geometry.xy[hoverIndex].x / width) * 100}%`,
transform: 'translateX(-50%)',
}}
>
@@ -202,7 +291,7 @@ export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProp
)}
<p className="text-[10px] font-mono text-cyber-text-muted">
共 {points.length} 个数据点,每轮采集记录一次快照
共 {points.length} 个数据点,每轮采集记录一次快照 · 纵轴范围见左侧刻度
</p>
</div>
)
+415 -61
View File
@@ -1,10 +1,17 @@
import { Fragment, useState } from 'react'
import { ChevronDown, ChevronRight, ExternalLink } from 'lucide-react'
import { Fragment, useMemo, useRef, useState } from 'react'
import { ChevronDown, ChevronRight, ExternalLink, Pencil, Users } from 'lucide-react'
import { Badge } from '@/components/ui/badge'
import { useMonitorNotes } from '@/hooks/useMonitor'
import { formatCount, formatDelta, formatRelative } from '@/lib/monitorFormat'
import type { MonitorNote, NoteMetrics } from '@/types/monitor'
import { useMonitorNotes, useSetCreatorAlias, useSetNoteAlias } from '@/hooks/useMonitor'
import {
formatCount,
formatDate,
formatDateTime,
formatDelta,
formatRelative,
} from '@/lib/monitorFormat'
import type { MonitorCreator, MonitorNote, NoteMetrics } from '@/types/monitor'
import { NoteCover } from './NoteCover'
import { NoteTrendChart } from './NoteTrendChart'
interface NotesTableProps {
@@ -38,15 +45,197 @@ const METRIC_COLUMNS: Array<{ key: keyof NoteMetrics; label: string }> = [
{ key: 'share_count', label: '分享' },
]
/**
* 一个博主名下所有作品的指标合计。
*
* **全为 null 时结果是 null 而不是 0** —— 与项目一贯的口径一致:「0」是真实值,
* 「null」是不知道。合计成 0 会让"还没采到"看起来像"互动为零"。
*/
function sumMetrics(notes: MonitorNote[]) {
const values: Record<string, number | null> = {}
const deltas: Record<string, number | null> = {}
for (const column of METRIC_COLUMNS) {
let sum = 0
let sawValue = false
let deltaSum = 0
let sawDelta = false
for (const note of notes) {
const value = note.metrics[column.key]
if (value !== null && value !== undefined) {
sum += value
sawValue = true
}
const delta = note.deltas[column.key]
if (delta !== null && delta !== undefined) {
deltaSum += delta
sawDelta = true
}
}
values[column.key] = sawValue ? sum : null
deltas[column.key] = sawDelta ? deltaSum : null
}
return { values, deltas }
}
// 多出来的那几列:展开箭头、作品、发布日期、首次发现、跳转链接;分组表头会跨掉整行。
const COLUMN_COUNT = METRIC_COLUMNS.length + 5
/**
* 账号级指标那几个字段。
*
* 作品的 payload 和博主的 payload 都带这一组,值也一样(同一次采集写的同一行),
* 所以组件只认字段、不认来源。
*/
type AccountStats = Pick<
MonitorCreator,
'creator_fans' | 'creator_total_favorited' | 'creator_works' | 'creator_stats_at'
>
/**
* 博主的**账号级**指标:粉丝 / 总获赞 / 作品数。
*
* 作品列表给不了这个 —— 那几列说的是「这条作品涨了多少赞」,这里说的是「这个人整个
* 账号在涨还是在掉」。**一个都没采到时整块不画**:画成「粉丝 0」比不画糟得多,那是
* 一句假话。
*/
function CreatorStats({ stats }: { stats: AccountStats }) {
const parts = [
stats.creator_fans !== null && `粉丝 ${formatCount(stats.creator_fans)}`,
stats.creator_total_favorited !== null && `获赞 ${formatCount(stats.creator_total_favorited)}`,
stats.creator_works !== null && `${formatCount(stats.creator_works)} 作品`,
].filter(Boolean) as string[]
if (parts.length === 0) return null
return (
<span
className="text-cyber-text-muted flex-shrink-0"
title={`账号指标采集于 ${formatDateTime(stats.creator_stats_at)}`}
>
{parts.join(' · ')}
</span>
)
}
/**
* 作品表,**按博主分组**。
*
* 一个任务可以配多个博主,平铺的话根本看不出哪篇是谁的。分组键是
* `creator_hash` —— 爬虫刻意不落原始 user_id,所以这是唯一稳定的创作者标识;
* 显示名是已脱敏的昵称(张***三)。两者都认不出来时才退回"未知博主"。
*/
export function NotesTable({ taskId, onlyNew }: NotesTableProps) {
const { data: notes, isLoading } = useMonitorNotes(taskId, onlyNew)
const [expanded, setExpanded] = useState<string | null>(null)
const { data, isLoading } = useMonitorNotes(taskId, onlyNew)
const notes = data?.notes
const creators = data?.creators
const [expandedNote, setExpandedNote] = useState<string | null>(null)
// 折叠状态按博主记。默认全展开 —— 藏起来的数据比多滚两屏更糟。
const [collapsed, setCollapsed] = useState<Set<string>>(new Set())
// 正在改备注的那个博主(creator_hash);null = 没在改。
const [editingCreator, setEditingCreator] = useState<string | null>(null)
const [aliasDraft, setAliasDraft] = useState('')
// 按 Escape 是「取消」,而失焦是「保存」—— 取消时得让紧接着那次失焦闭嘴,
// 否则它会把你刚放弃的内容存进去。
const aliasCancelled = useRef(false)
const setAlias = useSetCreatorAlias()
// 正在改备注的那条**作品**。和博主备注是两个互不干扰的编辑态。
const [editingNote, setEditingNote] = useState<string | null>(null)
const [noteAliasDraft, setNoteAliasDraft] = useState('')
const noteAliasCancelled = useRef(false)
const setNoteAlias = useSetNoteAlias()
const commitAlias = (creatorHash: string) => {
if (aliasCancelled.current) {
aliasCancelled.current = false
setEditingCreator(null)
return
}
setEditingCreator(null)
setAlias.mutate({ creatorHash, alias: aliasDraft })
}
const commitNoteAlias = (noteId: string) => {
if (noteAliasCancelled.current) {
noteAliasCancelled.current = false
setEditingNote(null)
return
}
setEditingNote(null)
setNoteAlias.mutate({ noteId, alias: noteAliasDraft })
}
const groups = useMemo(() => {
type Group = {
key: string
name: string
alias: string
/** 服务端说的这条博主名下有多少作品 —— 列表被 `onlyNew` 滤过时和 `notes.length` 不等。 */
totalNotes: number
stats: MonitorCreator | null
notes: MonitorNote[]
}
const byCreator = new Map<string, Group>()
// **先放博主,不是先放作品。** 「一条作品都没有的博主」在作品里根本推不出来 ——
// 目标加了、资料也采到了、粉丝数就躺在库里,可界面上什么都看不见。而他恰恰是最该
// 看见的一个:还在涨粉,只是最近没发东西。
for (const creator of creators ?? []) {
const key = creator.creator_hash || '__unknown__'
const existing = byCreator.get(key)
if (existing) {
// 同一个博主可能挂在多个任务下(服务端按 任务×博主 给),合并成一组。
existing.totalNotes += creator.note_count
if (!existing.stats?.creator_stats_at && creator.creator_stats_at) {
existing.stats = creator
}
existing.name ||= creator.creator_name
// 备注按 (platform, creator_hash) 存,所以跨任务就是同一条,取到即可。
existing.alias ||= creator.creator_alias
continue
}
byCreator.set(key, {
key,
name: creator.creator_name || '',
alias: creator.creator_alias || '',
totalNotes: creator.note_count,
stats: creator,
notes: [],
})
}
for (const note of notes ?? []) {
// 作品模式下每条作品都会带 creator_hash;真丢了也要有个兜底分组,
// 否则那些作品会凭空消失。
const key = note.creator_hash || '__unknown__'
let group = byCreator.get(key)
if (!group) {
group = {
key,
name: note.creator_name || '',
alias: note.creator_alias || '',
totalNotes: 0,
stats: null,
notes: [],
}
byCreator.set(key, group)
}
group.notes.push(note)
group.name ||= note.creator_name || ''
group.alias ||= note.creator_alias || ''
}
return [...byCreator.values()]
}, [notes, creators])
if (isLoading) {
return <p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">加载中…</p>
}
if (!notes || notes.length === 0) {
// 博主有、作品没有也是**要画**的:那正是「这个号还没被删,只是没发东西」。
if (groups.length === 0) {
return (
<p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">
{taskId === null
@@ -58,7 +247,13 @@ export function NotesTable({ taskId, onlyNew }: NotesTableProps) {
)
}
const isExpanded = (note: MonitorNote) => expanded === `${note.task_id}:${note.note_id}`
const toggleCreator = (key: string) =>
setCollapsed((prev) => {
const next = new Set(prev)
if (next.has(key)) next.delete(key)
else next.add(key)
return next
})
return (
<div className="overflow-x-auto terminal-scroll">
@@ -72,70 +267,229 @@ export function NotesTable({ taskId, onlyNew }: NotesTableProps) {
{column.label}
</th>
))}
<th className="text-right font-normal py-2 px-2">发布日期</th>
<th className="text-right font-normal py-2 px-2">首次发现</th>
<th className="w-8" />
</tr>
</thead>
<tbody>
{notes.map((note) => {
const open = isExpanded(note)
{groups.map((group) => {
const isCollapsed = collapsed.has(group.key)
const totals = sumMetrics(group.notes)
return (
<Fragment key={`${note.task_id}:${note.note_id}`}>
<Fragment key={group.key}>
{/* 组头与数据行**列对齐**:每个指标列给出该博主的合计,下面再带本轮增量。
折叠之后如果什么都看不到,那折叠就只是把信息藏起来了。 */}
<tr
onClick={() => setExpanded(open ? null : `${note.task_id}:${note.note_id}`)}
className="border-t border-cyber-border-subtle hover:bg-cyber-bg-elevated/50 cursor-pointer"
onClick={() => toggleCreator(group.key)}
className="border-t border-cyber-border-subtle bg-cyber-bg-tertiary/50 cursor-pointer hover:bg-cyber-bg-elevated/50"
>
<td className="py-2 px-1 text-cyber-text-muted">
{open ? <ChevronDown className="w-3 h-3" /> : <ChevronRight className="w-3 h-3" />}
</td>
<td className="py-2 px-2 max-w-[320px]">
<div className="flex items-center gap-1.5">
{note.is_new && (
<Badge variant="success" className="text-[9px] px-1 py-0 flex-shrink-0">
NEW
</Badge>
)}
<span className="truncate text-cyber-text-primary" title={note.title}>
{note.title || note.note_id}
</span>
</div>
<div className="text-[9px] text-cyber-text-muted truncate">
{note.note_id} · {note.snapshot_count} 次快照
</div>
</td>
{METRIC_COLUMNS.map((column) => (
<td key={column.key} className="py-2 px-2">
<MetricCell value={note.metrics[column.key]} delta={note.deltas[column.key]} />
</td>
))}
<td className="py-2 px-2 text-right text-[10px] text-cyber-text-muted">
{formatRelative(note.first_seen_at)}
</td>
<td className="py-2 px-1">
{note.note_url && (
<a
href={note.note_url}
target="_blank"
rel="noopener noreferrer"
onClick={(event) => event.stopPropagation()}
className="text-cyber-text-muted hover:text-cyber-neon-cyan"
>
<ExternalLink className="w-3 h-3" />
</a>
<td className="py-1.5 px-1 text-cyber-text-muted">
{isCollapsed ? (
<ChevronRight className="w-3 h-3" />
) : (
<ChevronDown className="w-3 h-3" />
)}
</td>
</tr>
{open && (
<tr className="bg-cyber-bg-secondary/40">
<td colSpan={METRIC_COLUMNS.length + 4} className="px-4 py-3">
<NoteTrendChart
noteId={note.note_id}
taskId={note.task_id}
noteTitle={note.title}
<td className="py-1.5 px-2">
{editingCreator === group.key ? (
<input
autoFocus
value={aliasDraft}
placeholder="备注,比如「竞品A」"
onClick={(event) => event.stopPropagation()}
onChange={(event) => setAliasDraft(event.target.value)}
onKeyDown={(event) => {
event.stopPropagation()
if (event.key === 'Enter') commitAlias(group.key)
if (event.key === 'Escape') {
aliasCancelled.current = true
setEditingCreator(null)
}
}}
onBlur={() => commitAlias(group.key)}
className="w-full rounded border border-cyber-neon-cyan/40 bg-cyber-bg-tertiary px-1.5 py-0.5 text-[11px] font-mono text-cyber-text-primary outline-none"
/>
) : (
<span className="flex items-center gap-1.5">
<Users className="w-3 h-3 text-cyber-neon-cyan flex-shrink-0" />
<span className="text-cyber-text-primary truncate">
{group.alias || group.name || '未知博主'}
</span>
{/* 起了备注之后,平台昵称降成副标题 —— 它仍是有用的对照。 */}
{group.alias && group.name && (
<span className="text-cyber-text-muted truncate">{group.name}</span>
)}
<span className="text-cyber-text-muted flex-shrink-0">
{/* 被 onlyNew 滤过时,把「显示了几篇 / 一共几篇」都说出来,
否则「1 篇」会让人以为这个号只发过一条。 */}
{group.notes.length}
{group.totalNotes > group.notes.length && `/${group.totalNotes}`} 篇
</span>
{group.stats && <CreatorStats stats={group.stats} />}
{/* 博主在、作品一条都没有 —— 说清楚,别让人以为列表挂了。 */}
{group.notes.length === 0 && (
<span className="text-cyber-text-muted/70 flex-shrink-0">
暂无作品
</span>
)}
<button
onClick={(event) => {
event.stopPropagation()
setAliasDraft(group.alias)
setEditingCreator(group.key)
}}
title="给这个博主起个备注"
className="text-cyber-text-muted hover:text-cyber-neon-cyan flex-shrink-0"
>
<Pencil className="w-3 h-3" />
</button>
</span>
)}
</td>
{METRIC_COLUMNS.map((column) => (
<td key={column.key} className="py-1.5 px-2">
<MetricCell
value={totals.values[column.key]}
delta={totals.deltas[column.key]}
/>
</td>
</tr>
)}
))}
<td />
<td />
</tr>
{!isCollapsed &&
group.notes.map((note) => {
const rowKey = `${note.task_id}:${note.note_id}`
const open = expandedNote === rowKey
return (
<Fragment key={rowKey}>
<tr
onClick={() => setExpandedNote(open ? null : rowKey)}
className="border-t border-cyber-border-subtle hover:bg-cyber-bg-elevated/50 cursor-pointer"
>
<td />
<td className="py-2 px-2 max-w-[320px]">
<div className="flex items-center gap-2">
<NoteCover src={note.cover} size="md" />
{/* min-w-0 so the title can actually truncate inside the flex row */}
<div className="min-w-0">
<div className="flex items-center gap-1.5">
{note.is_new && (
<Badge
variant="success"
className="text-[9px] px-1 py-0 flex-shrink-0"
>
NEW
</Badge>
)}
{/* 备注是**加**在标题前面的一枚标记,不是替换 ——
标题才是这条作品本身,扫列表时两个都要看得见。 */}
{note.note_alias && (
<Badge
variant="default"
className="text-[9px] px-1 py-0 flex-shrink-0"
>
{note.note_alias}
</Badge>
)}
<span
className="truncate text-cyber-text-primary"
title={note.title}
>
{note.title || note.note_id}
</span>
</div>
{editingNote === rowKey ? (
<input
autoFocus
value={noteAliasDraft}
placeholder="备注,比如「重点跟拍」"
onClick={(event) => event.stopPropagation()}
onChange={(event) => setNoteAliasDraft(event.target.value)}
onKeyDown={(event) => {
event.stopPropagation()
if (event.key === 'Enter') commitNoteAlias(note.note_id)
if (event.key === 'Escape') {
noteAliasCancelled.current = true
setEditingNote(null)
}
}}
onBlur={() => commitNoteAlias(note.note_id)}
className="mt-1 w-full rounded border border-cyber-neon-cyan/40 bg-cyber-bg-tertiary px-1.5 py-0.5 text-[11px] font-mono text-cyber-text-primary outline-none"
/>
) : (
<div className="flex items-center gap-1.5 text-[9px] text-cyber-text-muted">
<span className="truncate">
{note.note_id} · {note.snapshot_count} 次快照
</span>
<button
onClick={(event) => {
event.stopPropagation()
setNoteAliasDraft(note.note_alias)
setEditingNote(rowKey)
}}
title="给这条作品起个备注"
className="flex items-center gap-0.5 flex-shrink-0 text-cyber-text-muted hover:text-cyber-neon-cyan"
>
<Pencil className="w-2.5 h-2.5" />
{note.note_alias ? '改备注' : '备注'}
</button>
</div>
)}
</div>
</div>
</td>
{METRIC_COLUMNS.map((column) => (
<td key={column.key} className="py-2 px-2">
<MetricCell
value={note.metrics[column.key]}
delta={note.deltas[column.key]}
/>
</td>
))}
{/* 发布日期和「首次发现」是两回事:把早就发过的作品加进监控时,
前者是作者发的那天,后者是我们第一次看到它的那天。 */}
<td
className="py-2 px-2 text-right text-[10px] text-cyber-text-muted whitespace-nowrap"
title={formatDateTime(note.published_at)}
>
{formatDate(note.published_at)}
</td>
<td className="py-2 px-2 text-right text-[10px] text-cyber-text-muted">
{formatRelative(note.first_seen_at)}
</td>
<td className="py-2 px-1">
{/* 直接跳转到该作品。stopPropagation,否则点它会连带展开趋势图。 */}
{note.note_url && (
<a
href={note.note_url}
target="_blank"
rel="noopener noreferrer"
onClick={(event) => event.stopPropagation()}
title="打开原作品"
className="text-cyber-text-muted hover:text-cyber-neon-cyan"
>
<ExternalLink className="w-3 h-3" />
</a>
)}
</td>
</tr>
{open && (
<tr className="bg-cyber-bg-secondary/40">
<td colSpan={COLUMN_COUNT} className="px-4 py-3">
<NoteTrendChart
noteId={note.note_id}
taskId={note.task_id}
noteTitle={note.title}
/>
</td>
</tr>
)}
</Fragment>
)
})}
</Fragment>
)
})}
@@ -0,0 +1,238 @@
import { useEffect, useState } from 'react'
import { useQueryClient } from '@tanstack/react-query'
import {
AlertTriangle,
CheckCircle2,
KeyRound,
Loader2,
QrCode,
RefreshCw,
X,
} from 'lucide-react'
import { Badge } from '@/components/ui/badge'
import { Button } from '@/components/ui/button'
import {
useCancelQrLogin,
useLoginState,
useQrLoginStatus,
useRecheckLogin,
useStartQrLogin,
} from '@/hooks/useMonitor'
import { useCurrentPlatform } from '@/hooks/usePlatform'
/**
* Scan-to-login, for a host with no display.
*
* The crawler's own QR flow prints the code into the terminal via PIL's
* `Image.show()`, which needs a desktop image viewer. On a server Chrome runs
* under Xvfb and there is no such viewer, so the backend reads the code out of
* that same browser over CDP and hands it here.
*
* **The login state is shown first and independently of the scan.** It comes from
* the browser's own profile, so it stays true across a server restart — the QR
* session does not, and reporting "did it work?" from a value that a redeploy
* silently erases is how a successful scan ends up looking like nothing happened.
*/
export function QrLoginPanel() {
const { capability, platform } = useCurrentPlatform()
const queryClient = useQueryClient()
const [polling, setPolling] = useState(false)
// 登录态检测目前只实现了小红书:它读的是小红书页面的 __INITIAL_STATE__。
// 后端 qrlogin.LOGIN_URL 里也只有 xhs 一项。
//
// 不挡住的话会撒一个具体的谎 —— 切到抖音时,面板照样报"已登录",因为小红书的
// 会话还在,而标题写的是「抖音登录态」。
const wired = platform === 'xhs'
const { data: state } = useQrLoginStatus(polling && wired)
const { data: login } = useLoginState(polling && wired)
const start = useStartQrLogin()
const cancel = useCancelQrLogin()
const recheck = useRecheckLogin()
const label = capability?.label ?? platform
const status = state?.status ?? 'idle'
// The browser's own answer wins over the session's: the session is memory, the
// profile is not.
const loggedIn = login?.logged_in ?? state?.logged_in ?? false
const nickname = login?.nickname ?? state?.nickname ?? null
// Stop polling the moment the outcome is known, and refresh the cookie panel:
// a completed scan is what makes it start reporting a healthy login.
useEffect(() => {
if (status === 'waiting' || status === 'idle') return
setPolling(false)
if (status === 'success') {
queryClient.invalidateQueries({ queryKey: ['monitorCookie'] })
queryClient.invalidateQueries({ queryKey: ['monitorLoginState'] })
}
}, [status, queryClient])
// 未接入的平台直接换掉整个面板。只改标题是不够的 —— 下面那些块读的是小红书的
// 状态,照渲染出来还是会显示"已登录"。
if (!wired) {
return (
<div className="rounded-lg glass-panel float-panel p-4 space-y-2">
<div className="flex items-center gap-2">
<KeyRound className="w-4 h-4 text-cyber-text-muted flex-shrink-0" />
<span className="font-mono text-xs tracking-wider text-cyber-text-primary">
浏览器登录态(扫码)
</span>
<Badge variant="outline" className="text-[10px]">
{label}未接入
</Badge>
</div>
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
扫码登录目前只实现了<span className="text-cyber-neon-cyan">小红书</span> ——
登录态检测读的是小红书页面的登录状态,其它平台还没有对应的实现。
切回小红书即可使用。
</p>
</div>
)
}
const begin = () => {
setPolling(true)
start.mutate()
}
const busy = start.isPending || cancel.isPending || recheck.isPending
return (
<div className="rounded-lg glass-panel float-panel p-4 space-y-3">
<div className="flex items-center justify-between gap-3">
<div className="flex items-center gap-2 min-w-0">
<KeyRound className="w-4 h-4 text-cyber-neon-cyan flex-shrink-0" />
{/* 标题必须和 Cookie 面板区分开:那个是"存进库、每轮注入"的 cookie,
这个是"写进浏览器 profile、CDP 模式复用"的浏览器登录态。两者是
不同的机制,都能算"登录着",只写「登录态」会看不出指的是哪一个。 */}
<span className="font-mono text-xs tracking-wider text-cyber-text-primary flex-shrink-0">
浏览器登录态(扫码)
</span>
{!wired ? (
<Badge variant="outline" className="text-[10px]">
{label}未接入
</Badge>
) : loggedIn ? (
<Badge variant="success" className="text-[10px]">
已登录
</Badge>
) : (
<Badge variant="warning" className="text-[10px]">
未登录
</Badge>
)}
{loggedIn && nickname && (
<span className="truncate text-[10px] font-mono text-cyber-text-secondary">
{nickname}
</span>
)}
</div>
<Button
variant="ghost"
size="sm"
disabled={busy}
onClick={() => recheck.mutate()}
title="重新加载页面并检测登录态"
>
<RefreshCw className={`w-3 h-3 mr-1 ${recheck.isPending ? 'animate-spin' : ''}`} />
重新检测
</Button>
</div>
{/* 读不到状态时要说清楚是"读不到",而不是悄悄显示成"未登录" */}
{login?.known === false && (
<p className="flex items-start gap-2 text-[10px] font-mono text-cyber-neon-orange leading-relaxed">
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
读不到浏览器状态{login.error ? `:${login.error}` : ''}。
请确认那台 Chrome 正常、且已打开小红书页面。
</p>
)}
{status === 'waiting' && state?.image && (
<div className="flex items-start gap-4">
{/* The code is a data: URL straight from the page, so nothing is
fetched from a third party and no file is written server-side. */}
<img
src={state.image}
alt="登录二维码"
className="w-44 h-44 rounded-md border border-cyber-border-DEFAULT bg-white p-1"
/>
<div className="space-y-2 pt-1">
<p className="text-[11px] font-mono text-cyber-text-secondary leading-relaxed">
用<span className="text-cyber-neon-cyan">{label} App</span>扫码。
扫完这个面板会自动变成「已登录」,无需手动刷新。
</p>
<p className="text-[11px] font-mono text-cyber-text-muted">
剩余 <span className="text-cyber-neon-cyan">{state.expires_in}</span> 秒
</p>
<Button
variant="ghost"
size="sm"
disabled={busy}
onClick={() => {
setPolling(false)
cancel.mutate()
}}
>
<X className="w-3 h-3 mr-1" />
取消
</Button>
</div>
</div>
)}
{status === 'waiting' && !state?.image && (
<p className="flex items-center gap-2 text-[11px] font-mono text-cyber-text-muted">
<Loader2 className="w-3 h-3 animate-spin" />
正在从浏览器取二维码…
</p>
)}
{status === 'success' && (
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-green leading-relaxed">
<CheckCircle2 className="w-3 h-3 mt-0.5 shrink-0" />
{state?.message || '登录成功'}
</p>
)}
{(status === 'error' || status === 'expired') && (
<div className="space-y-2">
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-orange leading-relaxed">
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
{state?.message || '获取二维码失败'}
</p>
<Button variant="outline" size="sm" disabled={busy} onClick={begin}>
<RefreshCw className="w-3 h-3 mr-1" />
重新获取
</Button>
</div>
)}
{status === 'idle' && (
<div className="space-y-2">
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
{loggedIn
? '这台浏览器已是登录状态。CDP 模式下定时任务直接复用它;点下面的按钮可以把这份登录态也存成 Cookie —— 那样即使关掉 CDP、任务改用 Cookie 注入也照样能跑。'
: '经 CDP 接管服务器上已开启远程调试的 Chrome,把二维码取回来显示在这里。需要先在「系统设置」里打开 接管已有 Chrome(CDP),并确保那台 Chrome 正以 9222 端口运行。'}
</p>
{/* 已登录时**也要给按钮**。原先这里把按钮藏了,于是面板变成一块只能看、
不能操作的区域 —— 用户的原话是「没用」。两种状态下点击是同一个动作,
只是含义不同:没登录就是取二维码,已登录就是把当前登录态同步成 Cookie。 */}
<Button size="sm" disabled={busy} onClick={begin}>
{start.isPending ? (
<Loader2 className="w-3 h-3 mr-1 animate-spin" />
) : (
<QrCode className="w-3 h-3 mr-1" />
)}
{loggedIn ? '同步登录态为 Cookie' : '获取二维码'}
</Button>
</div>
)}
</div>
)
}
+6 -1
View File
@@ -83,7 +83,12 @@ export function RunHistory({ taskId }: RunHistoryProps) {
{formatDateTime(run.started_at)}
<span className="ml-1">({formatRelative(run.started_at)})</span>
</td>
<td className="py-2 px-2 text-[10px] text-cyber-neon-orange max-w-[240px] truncate">
{/* 这一格是截断的,而失败原因现在会带上一整行异常 —— 没有 title
就等于把最要紧的那半句藏起来了。 */}
<td
className="py-2 px-2 text-[10px] text-cyber-neon-orange max-w-[320px] truncate"
title={run.error_message ?? undefined}
>
{run.error_message ?? ''}
</td>
</tr>
+11 -3
View File
@@ -3,7 +3,8 @@ import { Bell, CalendarClock, Pencil, Play, Trash2 } from 'lucide-react'
import { Badge } from '@/components/ui/badge'
import { Button } from '@/components/ui/button'
import { useDeleteTask, useRunTaskNow, useUpdateTask } from '@/hooks/useMonitor'
import { formatInterval, formatRelative } from '@/lib/monitorFormat'
import { useCurrentPlatform } from '@/hooks/usePlatform'
import { formatRelative } from '@/lib/monitorFormat'
import type { MonitorTask } from '@/types/monitor'
interface TaskCardProps {
@@ -46,6 +47,12 @@ export function TaskCard({ task, selected, onSelect, onEdit }: TaskCardProps) {
const updateTask = useUpdateTask()
const deleteTask = useDeleteTask()
const runNow = useRunTaskNow()
// 按**任务自己的**平台取措辞,而不是当前平台 —— 卡片未必只出现在同平台的列表里。
// 抖音管它们叫「作品」,小红书叫「笔记」,写死一个对另一个就是错的。
const { platforms } = useCurrentPlatform()
const noteNoun =
platforms.find((entry) => entry.value === task.platform)?.target_hints?.note_label ??
'笔记'
// A suspected cookie failure surfaces here so it is visible without opening
// the event feed.
@@ -65,7 +72,7 @@ export function TaskCard({ task, selected, onSelect, onEdit }: TaskCardProps) {
<div className="flex items-center gap-2 flex-wrap">
<span className="font-mono text-sm text-cyber-text-primary truncate">{task.name}</span>
<Badge variant="outline" className="text-[10px]">
{task.mode === 'creator' ? '博主' : '笔记'}
{task.mode === 'creator' ? '博主' : noteNoun}
</Badge>
<Badge variant={statusVariant(task.last_status)} className="text-[10px]">
{STATUS_LABEL[task.last_status] ?? task.last_status}
@@ -82,7 +89,8 @@ export function TaskCard({ task, selected, onSelect, onEdit }: TaskCardProps) {
<span>
目标 <span className="text-cyber-neon-cyan">{task.target_count}</span> 个
</span>
<span>间隔 {formatInterval(task.interval_minutes)}</span>
{/* 由后端拼好,列表和编辑弹窗因此不会对同一个计划给出两种说法 */}
<span>{task.schedule_label}</span>
<span className="flex items-center gap-1">
<CalendarClock className="w-3 h-3" />
{task.enabled ? formatRelative(task.next_run_at) : '已暂停'}
+295 -35
View File
@@ -1,4 +1,4 @@
import { useEffect, useState } from 'react'
import { useEffect, useState, type ReactNode } from 'react'
import { Button } from '@/components/ui/button'
import { Checkbox } from '@/components/ui/checkbox'
@@ -20,7 +20,14 @@ import {
SelectValue,
} from '@/components/ui/select'
import { useCreateTask, useSettings, useUpdateTask } from '@/hooks/useMonitor'
import type { MonitorMode, MonitorTask, TaskCreatePayload } from '@/types/monitor'
import { useCurrentPlatform } from '@/hooks/usePlatform'
import { describeSchedule } from '@/lib/monitorFormat'
import type {
MonitorMode,
MonitorTask,
ScheduleMode,
TaskCreatePayload,
} from '@/types/monitor'
interface TaskEditorDialogProps {
open: boolean
@@ -44,19 +51,80 @@ const INTERVAL_OPTIONS = [
{ value: '10080', label: '7 天' },
]
const SCHEDULE_MODE_OPTIONS: Array<{ value: ScheduleMode; label: string }> = [
{ value: 'interval', label: '固定间隔' },
{ value: 'daily', label: '每天定时' },
{ value: 'weekly', label: '每周定时' },
]
const HOURS = Array.from({ length: 24 }, (_, hour) => hour)
/**
* Above this many works per run, warn about the request volume.
*
* The runner serialises crawls (max_concurrency_num=1) and each note also pulls
* up to `max_comments_count` comments, so the cost is works × (1 + comments) and
* the default run timeout is an hour.
*/
const NOTE_VOLUME_WARN = 500
const MINUTES = Array.from({ length: 60 }, (_, minute) => minute)
const WEEKDAYS = ['周一', '周二', '周三', '周四', '周五', '周六', '周日']
const pad = (value: number) => String(value).padStart(2, '0')
function toggleNumber(values: number[], value: number): number[] {
return values.includes(value)
? values.filter((item) => item !== value)
: [...values, value].sort((a, b) => a - b)
}
/** A selectable pill. Used for the hours of the day and the weekdays. */
function Chip({
active,
onClick,
children,
}: {
active: boolean
onClick: () => void
children: ReactNode
}) {
return (
<button
type="button"
onClick={onClick}
className={`px-1.5 py-0.5 rounded border text-[10px] font-mono transition-colors ${
active
? 'border-cyber-neon-cyan/60 bg-cyber-neon-cyan/20 text-cyber-neon-cyan'
: 'border-cyber-border-DEFAULT text-cyber-text-muted hover:text-cyber-text-secondary'
}`}
>
{children}
</button>
)
}
export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogProps) {
const isEdit = Boolean(task)
const createTask = useCreateTask()
const updateTask = useUpdateTask()
const { data: settings } = useSettings()
// 示例链接和措辞都由服务端的能力矩阵给 —— 前端不自己判断平台,否则加一个平台
// 就要改这里一次,而且很容易漏。
const { capability, platform } = useCurrentPlatform()
const hints = capability?.target_hints
const [name, setName] = useState('')
const [mode, setMode] = useState<MonitorMode>('creator')
const [intervalMinutes, setIntervalMinutes] = useState('360')
const [scheduleMode, setScheduleMode] = useState<ScheduleMode>('interval')
const [scheduleHours, setScheduleHours] = useState<number[]>([])
const [scheduleDays, setScheduleDays] = useState<number[]>([])
const [scheduleMinute, setScheduleMinute] = useState('0')
const [maxNotes, setMaxNotes] = useState('20')
const [enableComments, setEnableComments] = useState(true)
const [maxComments, setMaxComments] = useState('50')
const [notifyEnabled, setNotifyEnabled] = useState(false)
// 异常推送默认开:失败意味着这个任务从此默默采不到东西,而你不会知道。
const [notifyFailures, setNotifyFailures] = useState(true)
const [targets, setTargets] = useState('')
// Reset the form whenever the dialog is (re)opened. For a new task the
@@ -69,6 +137,10 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
setIntervalMinutes(
String(task?.interval_minutes ?? settings?.values['collect.default_interval_minutes'] ?? 360),
)
setScheduleMode(task?.schedule_mode ?? 'interval')
setScheduleHours(task?.schedule_hours ?? [])
setScheduleDays(task?.schedule_days ?? [])
setScheduleMinute(String(task?.schedule_minute ?? 0))
setMaxNotes(
String(task?.max_notes_count ?? settings?.values['collect.default_max_notes'] ?? 20),
)
@@ -77,6 +149,7 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
String(task?.max_comments_count ?? settings?.values['collect.default_max_comments'] ?? 50),
)
setNotifyEnabled(task?.notify_enabled ?? false)
setNotifyFailures(task?.notify_failures ?? true)
setTargets(task ? task.targets.map((t) => t.raw_value || t.external_id).join('\n') : '')
}, [open, task, settings])
@@ -85,29 +158,55 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
.map((value) => value.trim())
.filter(Boolean)
// Switching into a clock mode with nothing chosen would be invalid, so seed a
// sensible starting point -- the operator then edits rather than fills a blank.
const selectScheduleMode = (next: ScheduleMode) => {
setScheduleMode(next)
if (next !== 'interval' && scheduleHours.length === 0) setScheduleHours([9])
if (next === 'weekly' && scheduleDays.length === 0) setScheduleDays([0, 1, 2, 3, 4])
}
const pending = createTask.isPending || updateTask.isPending
// Note-mode and creator-mode targets are different shapes, so switching mode
// would silently mis-parse the list. The backend decides parsing from the
// task's stored mode, hence mode is fixed once created.
const canSubmit = name.trim().length > 0 && targetList.length > 0 && !pending
const clockIncomplete =
scheduleMode !== 'interval' &&
(scheduleHours.length === 0 || (scheduleMode === 'weekly' && scheduleDays.length === 0))
const canSubmit =
name.trim().length > 0 && targetList.length > 0 && !pending && !clockIncomplete
const handleSubmit = () => {
const payload: TaskCreatePayload = {
name: name.trim(),
// 当前平台。漏掉这一项,后端会退回小红书 —— 表现是「在抖音页面建的任务
// 跑到小红书列表里去了」,而且不报任何错。
platform,
mode,
interval_minutes: Number(intervalMinutes),
schedule_mode: scheduleMode,
// Interval mode has no clock times, so send none rather than leave stale
// ones behind from a mode the operator tried and abandoned.
schedule_hours: scheduleMode === 'interval' ? [] : scheduleHours,
schedule_days: scheduleMode === 'weekly' ? scheduleDays : [],
schedule_minute: Number(scheduleMinute),
max_notes_count: Number(maxNotes),
enable_comments: enableComments,
max_comments_count: Number(maxComments),
run_timeout_seconds: 3600,
enabled: true,
notify_enabled: notifyEnabled,
notify_failures: notifyFailures,
targets: targetList,
}
const done = () => onOpenChange(false)
if (isEdit && task) {
updateTask.mutate({ id: task.id, payload }, { onSuccess: done })
// 平台创建后不可更改,更新请求里就不带它了 —— 带着会让「平台能被改」这件事
// 看起来像是真的。
const updatePayload: Partial<TaskCreatePayload> = { ...payload }
delete updatePayload.platform
updateTask.mutate({ id: task.id, payload: updatePayload }, { onSuccess: done })
} else {
createTask.mutate(payload, { onSuccess: done })
}
@@ -149,7 +248,7 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
</SelectTrigger>
<SelectContent>
<SelectItem value="creator">博主(监控其作品)</SelectItem>
<SelectItem value="note">笔记(批量监控指定内容)</SelectItem>
<SelectItem value="note">{`${hints?.note_label ?? '笔记'}(批量监控指定内容)`}</SelectItem>
</SelectContent>
</Select>
{isEdit && (
@@ -159,6 +258,28 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
)}
</div>
<div className="space-y-2">
<Label className="text-xs font-mono text-cyber-text-secondary">运行计划</Label>
<div className="flex gap-1">
{SCHEDULE_MODE_OPTIONS.map((option) => (
<button
key={option.value}
type="button"
onClick={() => selectScheduleMode(option.value)}
className={`flex-1 rounded border px-1 py-1.5 text-[10px] font-mono transition-colors ${
scheduleMode === option.value
? 'border-cyber-neon-cyan/60 bg-cyber-neon-cyan/20 text-cyber-neon-cyan'
: 'border-cyber-border-DEFAULT text-cyber-text-muted hover:text-cyber-text-secondary'
}`}
>
{option.label}
</button>
))}
</div>
</div>
</div>
{scheduleMode === 'interval' ? (
<div className="space-y-2">
<Label className="text-xs font-mono text-cyber-text-secondary">采集间隔</Label>
<Select value={intervalMinutes} onValueChange={setIntervalMinutes}>
@@ -173,46 +294,154 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
))}
</SelectContent>
</Select>
<p className="text-[10px] font-mono text-cyber-text-muted">
从上一轮<span className="text-cyber-text-secondary">开始</span>计时,
所以某轮跑得久也不会让下一轮紧接着触发。
</p>
</div>
</div>
) : (
<div className="space-y-3 rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3">
{scheduleMode === 'weekly' && (
<div className="space-y-1.5">
<div className="text-[10px] font-mono text-cyber-text-secondary">星期</div>
<div className="flex flex-wrap gap-1">
{WEEKDAYS.map((label, day) => (
<Chip
key={day}
active={scheduleDays.includes(day)}
onClick={() => setScheduleDays((prev) => toggleNumber(prev, day))}
>
{label}
</Chip>
))}
</div>
</div>
)}
<div className="space-y-1.5">
<div className="text-[10px] font-mono text-cyber-text-secondary">
时间(可多选,点一次选中、再点取消)
</div>
<div className="flex flex-wrap gap-1">
{HOURS.map((hour) => (
<Chip
key={hour}
active={scheduleHours.includes(hour)}
onClick={() => setScheduleHours((prev) => toggleNumber(prev, hour))}
>
{pad(hour)}
</Chip>
))}
</div>
</div>
<div className="flex items-center gap-2">
<span className="text-[10px] font-mono text-cyber-text-secondary">分钟</span>
<Select value={scheduleMinute} onValueChange={setScheduleMinute}>
<SelectTrigger className="h-7 w-20 text-xs">
<SelectValue />
</SelectTrigger>
<SelectContent className="max-h-56">
{MINUTES.map((minute) => (
<SelectItem key={minute} value={String(minute)}>
{pad(minute)}
</SelectItem>
))}
</SelectContent>
</Select>
<span className="text-[10px] font-mono text-cyber-text-muted">
所有时间点共用
</span>
</div>
<div className="flex items-center justify-between gap-3 border-t border-cyber-border-subtle pt-2">
<span className="text-[11px] font-mono text-cyber-neon-cyan">
{describeSchedule(
scheduleMode,
Number(intervalMinutes),
scheduleHours,
scheduleDays,
Number(scheduleMinute),
)}
</span>
<span className="text-[10px] font-mono text-cyber-text-muted">
按服务器本地时间
</span>
</div>
</div>
)}
<div className="space-y-2">
<Label className="text-xs font-mono text-cyber-text-secondary">
{mode === 'creator' ? '博主主页链接或 ID' : '笔记链接或 ID'}
{mode === 'creator'
? `${hints?.creator_label ?? '博主主页'}链接或 ID`
: `${hints?.note_label ?? '笔记'}链接或 ID`}
</Label>
<textarea
value={targets}
onChange={(event) => setTargets(event.target.value)}
rows={5}
placeholder={
mode === 'creator'
? '每行一个,支持完整主页链接或纯 ID:\nhttps://www.xiaohongshu.com/user/profile/5f58bd99...\n5f58bd990000000001003753'
: '每行一个,支持完整笔记链接或纯 ID:\nhttps://www.xiaohongshu.com/explore/6aa3d827...'
}
placeholder={`每行一个,支持完整链接或纯 ID:\n${(mode === 'creator' ? hints?.creator : hints?.note) ?? ''}`}
className={TEXTAREA_CLASS}
/>
<p className="text-[10px] font-mono text-cyber-text-muted">
已识别 <span className="text-cyber-neon-cyan">{targetList.length}</span> 个目标。
<span className="text-cyber-neon-orange">建议只填纯 ID</span>
——链接里的 xsec_token 会过期,纯 ID 永久有效。
{/* 「只填纯 ID」是小红书专属的劝告:它链接里的 xsec_token 会过期。
抖音的链接不带令牌,永久有效,那句话对它没有意义。 */}
{hints?.token_expires && (
<>
<span className="text-cyber-neon-orange">建议只填纯 ID</span>
——链接里的 xsec_token 会过期,纯 ID 永久有效。
</>
)}
</p>
</div>
<div className="grid grid-cols-2 gap-3">
<div className="space-y-2">
<Label className="text-xs font-mono text-cyber-text-secondary">
每轮最多采集作品数
{mode === 'creator' ? '每个博主最多采集作品数' : '最多采集作品数'}
</Label>
<Input
type="number"
min={1}
value={maxNotes}
onChange={(event) => setMaxNotes(event.target.value)}
disabled={mode === 'note'}
className="h-9 text-xs"
/>
<p className="text-[10px] font-mono text-cyber-text-muted">
只取最新的前 N 条,决定了"该博主的作品"覆盖范围
</p>
{mode === 'creator' ? (
<>
{/* 这是「每个博主」的上限,不是一轮的总量 —— 爬虫里这个值是在
per-creator 的函数内比较的(client.py get_all_notes_by_creator),
而外层 for 循环会遍历全部目标。所以 100 个目标 × 20 篇 = 单轮
最多 2000 篇。标签写成「每轮最多」会让人以为超出的会被丢弃。 */}
<p className="text-[10px] font-mono text-cyber-text-muted">
这是<span className="text-cyber-text-secondary">每个博主</span>的上限,
不是一轮的总量。只取该博主最新的前 N 条,超出的不会补抓。
</p>
<p className="text-[10px] font-mono text-cyber-text-secondary">
{targetList.length} 个目标 × {maxNotes} 篇 × 每人 1 次
→ 单轮最多{' '}
<span className="text-cyber-neon-cyan">
{targetList.length * (Number(maxNotes) || 0)}
</span>{' '}
篇
</p>
{targetList.length * (Number(maxNotes) || 0) > NOTE_VOLUME_WARN && (
<p className="text-[10px] font-mono text-cyber-neon-orange leading-relaxed">
单轮量偏大:每篇还要抓最多 {maxComments} 条评论,且并发为 1。
容易触发平台限流,也可能跑不完就被任务超时(默认 1 小时)中断。
建议调低这个数,或拆成几个任务。
</p>
)}
</>
) : (
<p className="text-[10px] font-mono text-cyber-neon-orange">
{hints?.note_label ?? '笔记'}模式下此项不生效:你列出的每个链接都会被逐条抓取。
</p>
)}
</div>
<div className="space-y-2">
@@ -246,24 +475,55 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
</div>
</div>
<div className="flex items-start gap-2 rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3">
<Checkbox
id="notify-enabled"
checked={notifyEnabled}
onCheckedChange={(checked) => setNotifyEnabled(checked === true)}
/>
<div className="space-y-0.5">
<label
htmlFor="notify-enabled"
className="text-xs font-mono text-cyber-text-primary cursor-pointer"
>
推送企业微信通知
</label>
<p className="text-[10px] font-mono text-cyber-text-muted">
仅在本任务**采集失败 / 登录态失效**或**发现新作品**时推送,
一轮只发一条汇总。需先在监控页配置 Webhook 地址。
</p>
<div className="rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3 space-y-3">
{/* 两类通知的性质完全不同,所以分成两个开关:
异常低频且意味着任务已经停止工作 —— 默认开;
新作品可能每轮都有 —— 默认关,否则会刷屏。 */}
<div className="flex items-start gap-2">
<Checkbox
id="notify-failures"
checked={notifyFailures}
onCheckedChange={(checked) => setNotifyFailures(checked === true)}
/>
<div className="space-y-0.5">
<label
htmlFor="notify-failures"
className="text-xs font-mono text-cyber-text-primary cursor-pointer"
>
推送异常通知(建议保持开启)
</label>
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
登录态失效、采集进程失败、一篇都没抓到时推送。
<span className="text-cyber-neon-cyan">
异常意味着这个任务从此默默采不到任何东西
</span>
—— 关掉的话你不会知道,直到某天发现数据停在几周前。
</p>
</div>
</div>
<div className="flex items-start gap-2 border-t border-cyber-border-subtle pt-3">
<Checkbox
id="notify-enabled"
checked={notifyEnabled}
onCheckedChange={(checked) => setNotifyEnabled(checked === true)}
/>
<div className="space-y-0.5">
<label
htmlFor="notify-enabled"
className="text-xs font-mono text-cyber-text-primary cursor-pointer"
>
推送新作品通知
</label>
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
发现新作品时推送。监控多个博主时可能每轮都有,容易刷屏,所以默认关闭。
</p>
</div>
</div>
<p className="border-t border-cyber-border-subtle pt-2 text-[10px] font-mono text-cyber-text-muted">
两者都是一轮只发一条汇总。需先在右上角「系统设置」里配置企业微信 Webhook 地址。
</p>
</div>
</div>
+11 -8
View File
@@ -3,6 +3,7 @@ import { KeyRound, QrCode, Save } from 'lucide-react'
import { Button } from '@/components/ui/button'
import { CookiePanel } from '@/components/monitor/CookiePanel'
import { QrLoginPanel } from '@/components/monitor/QrLoginPanel'
import { Section, SettingField } from '@/components/settings/SettingFields'
import { useSettings, useUpdateSettings } from '@/hooks/useMonitor'
import { useCurrentPlatform } from '@/hooks/usePlatform'
@@ -66,17 +67,19 @@ export function SettingsView({ onNavigate }: { onNavigate?: (view: AppView) => v
</p>
</div>
<Section title="登录态" description={`${label}的登录 Cookie,定时监控必须持久化登录态。`}>
<Section
title="登录态"
description={`${label}的登录态。定时监控必须持久化,扫码一次最省事。`}
>
{/* QR first: it is the one that keeps working unattended, since the
scan lands in the very browser profile the monitor runs reuse. */}
<QrLoginPanel />
<CookiePanel />
{/* Not a duplicate login flow: the crawler has no login-only mode, so a
QR login is a side effect of a real crawl -- which is exactly what
the 采集 page already does. This preselects it rather than
reimplementing it. */}
<div className="pt-3 mt-1 border-t border-cyber-border-subtle space-y-2">
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
Cookie 不好使时,可以走一次**扫码登录**:二维码会显示在「采集」页的终端里。
扫码成功后浏览器 profile 会被更新,Cookie 的可靠性也会显著提升。
在本机桌面运行时,也可以走「采集」页的扫码:二维码会打印在终端里。
服务器没有显示器,那条路走不通,用上面的面板。
</p>
<Button
variant="outline"
@@ -89,7 +92,7 @@ export function SettingsView({ onNavigate }: { onNavigate?: (view: AppView) => v
}}
>
<QrCode className="w-3 h-3 mr-1" />
去扫码登录
去采集页扫码
</Button>
</div>
</Section>
@@ -12,6 +12,7 @@ import {
} from '@/components/ui/dialog'
import { WebhookPanel } from '@/components/monitor/WebhookPanel'
import { ChangePassword, SettingField } from '@/components/settings/SettingFields'
import { UpstreamPanel } from '@/components/settings/UpstreamPanel'
import { useSettings, useUpdateSettings } from '@/hooks/useMonitor'
type Draft = Record<string, boolean | number | string>
@@ -44,6 +45,18 @@ export function SystemSettingsDialog({
[data],
)
// 上游那几项单独成块,其余(时段、CDP)仍归在「调度」下。按前缀分流就够了:
// 注册表是唯一的来源,新增一项上游设置不需要再动这个文件。写成「排除 upstream_」
// 而不是「只取 active_hours_」,这样以后再加系统设置也不会从界面上凭空消失。
const upstreamSpecs = useMemo(
() => systemSpecs.filter((spec) => spec.name.startsWith('upstream_')),
[systemSpecs],
)
const scheduleSpecs = useMemo(
() => systemSpecs.filter((spec) => !spec.name.startsWith('upstream_')),
[systemSpecs],
)
const dirty = useMemo(() => {
if (!data?.values) return {}
const changes: Draft = {}
@@ -86,7 +99,28 @@ export function SystemSettingsDialog({
调度器全局只有一套时段规则,因此不按平台区分。
</p>
</div>
{systemSpecs.map((spec) => (
{scheduleSpecs.map((spec) => (
<SettingField
key={spec.key}
spec={spec}
value={draft[spec.key] ?? (spec.default as boolean | number | string)}
onChange={(next) => setDraft((prev) => ({ ...prev, [spec.key]: next }))}
/>
))}
</div>
<div className="space-y-3 border-t border-cyber-border-subtle pt-4">
<div>
<h3 className="font-mono text-xs tracking-wider text-cyber-text-primary">
上游更新
</h3>
<p className="mt-0.5 text-[10px] font-mono text-cyber-text-muted">
本仓库在上游 MediaCrawler 之上加了一整层,这里定期看看上游有没有新提交。
改动立即生效,不需要等下一轮采集。
</p>
</div>
<UpstreamPanel />
{upstreamSpecs.map((spec) => (
<SettingField
key={spec.key}
spec={spec}
@@ -0,0 +1,75 @@
import { RefreshCw } from 'lucide-react'
import { Badge } from '@/components/ui/badge'
import { Button } from '@/components/ui/button'
import { useCheckUpstream, useUpstreamStatus } from '@/hooks/useMonitor'
import { formatDateTime, formatRelative } from '@/lib/monitorFormat'
import type { UpstreamStatus } from '@/types/monitor'
/**
* 最近一次上游检查的结果,外加一个「立即检查」。
*
* 开关与间隔都由设置项本身渲染(同一张注册表),这里只补状态显示:光有一个开关,
* 用户没法知道它到底跑没跑、上游到底动没动。后端只缓存结果,不在这里发 fetch ——
* 「立即检查」才发,而且那条请求要等 fetch 跑完,所以超时是单独放长的。
*/
export function UpstreamPanel() {
const { data, isLoading } = useUpstreamStatus()
const checkNow = useCheckUpstream()
return (
<div className="space-y-2 rounded-lg border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3">
<div className="flex items-center justify-between gap-3">
<div className="flex items-center gap-2">{statusBadge(data)}</div>
<Button
variant="outline"
size="sm"
disabled={checkNow.isPending}
onClick={() => checkNow.mutate()}
>
<RefreshCw className={`w-3 h-3 mr-1 ${checkNow.isPending ? 'animate-spin' : ''}`} />
{checkNow.isPending ? '检查中…' : '立即检查'}
</Button>
</div>
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
{isLoading ? '读取中…' : describe(data)}
</p>
{data?.commits && data.commits.length > 0 && (
<ul className="space-y-1 border-t border-cyber-border-subtle pt-2">
{data.commits.slice(0, 10).map((commit) => (
<li key={commit.sha} className="flex gap-2 text-[10px] font-mono leading-relaxed">
<span className="text-cyber-text-muted shrink-0">{commit.sha}</span>
<span className="text-cyber-text-secondary truncate" title={commit.subject}>
{commit.subject}
</span>
</li>
))}
</ul>
)}
</div>
)
}
function statusBadge(data?: UpstreamStatus) {
if (!data || !data.checked_at) return <Badge variant="idle">尚未检查</Badge>
if (!data.ok) return <Badge variant="destructive">检查失败</Badge>
if ((data.behind ?? 0) > 0) return <Badge variant="warning">落后 {data.behind} 个提交</Badge>
return <Badge variant="success">已是最新</Badge>
}
function describe(data?: UpstreamStatus): string {
if (!data || !data.checked_at) {
return '还没有检查过。开启上面的开关会按间隔自动查,也可以点「立即检查」。'
}
const when = `上次检查:${formatDateTime(data.checked_at)}(${formatRelative(data.checked_at)})`
if (!data.ok) return `${when};${data.error ?? '未知错误'}`
const parts = [when, `上游 ${data.branch ?? 'main'} 领先 ${data.behind ?? 0} 个提交`]
// 领先数就是我们自己这一层的规模;合并时要保留的东西,值得一并说出来。
if (data.ahead) parts.push(`本仓库另有 ${data.ahead} 个自己的提交`)
if (data.notify_error) parts.push(`通知发送失败:${data.notify_error}`)
return parts.join(';')
}
+99
View File
@@ -0,0 +1,99 @@
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query'
import { toast } from 'sonner'
import { creatorApi } from '@/lib/api'
const ACCOUNTS_KEY = ['creatorAccounts']
/** 账号列表。同步在后台跑,所以列表本身也顺带轮询,好让状态自己刷新出来。 */
export function useCreatorAccounts(polling = false) {
return useQuery({
queryKey: ACCOUNTS_KEY,
queryFn: async () => (await creatorApi.listAccounts()).data.accounts,
refetchInterval: polling ? 5000 : false,
})
}
export function useCreatorAccount(id: number | null) {
return useQuery({
queryKey: ['creatorAccount', id],
queryFn: async () => (await creatorApi.getAccount(id as number)).data,
enabled: id !== null,
})
}
export function useDeleteCreatorAccount() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: (id: number) => creatorApi.deleteAccount(id),
onSuccess: () => {
toast.success('账号已删除')
queryClient.invalidateQueries({ queryKey: ACCOUNTS_KEY })
},
onError: (error: Error) => toast.error(`删除失败:${error.message}`),
})
}
export function useCheckCreatorAccount() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: (id: number) => creatorApi.checkAccount(id),
onSuccess: () => {
queryClient.invalidateQueries({ queryKey: ACCOUNTS_KEY })
},
onError: (error: Error) => toast.error(`检测失败:${error.message}`),
})
}
export function useSyncCreatorAccount() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: ({ id, days }: { id: number; days?: number }) =>
creatorApi.syncAccount(id, days),
onSuccess: () => {
// 后端是后台任务,立刻重新拉一次只会看到旧状态;给用户一句"已开始",
// 列表的轮询会把结果带回来。
toast.success('同步已开始,稍后自动刷新')
queryClient.invalidateQueries({ queryKey: ACCOUNTS_KEY })
},
onError: (error: Error) => toast.error(`同步失败:${error.message}`),
})
}
// --- 扫码新增账号 ---------------------------------------------------------
/** 只在扫码中轮询;空闲时没必要一直问。 */
export function useCreatorLoginStatus(polling: boolean) {
return useQuery({
queryKey: ['creatorLogin'],
queryFn: async () => (await creatorApi.getLogin()).data,
refetchInterval: polling ? 2000 : false,
})
}
export function useStartCreatorLogin() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: () => creatorApi.startLogin(),
onSuccess: (response) => {
queryClient.setQueryData(['creatorLogin'], response.data)
},
onError: (error: Error) => {
// 后端给的原因是这里唯一有价值的信息(连不上 Chrome、页面上没有二维码),
// 而 axios 会把它压成 "Request failed with status code 502"。
const detail = (error as { response?: { data?: { detail?: string } } })?.response?.data
?.detail
toast.error(`获取二维码失败:${detail ?? error.message}`)
},
})
}
export function useCancelCreatorLogin() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: () => creatorApi.cancelLogin(),
onSuccess: (response) => {
queryClient.setQueryData(['creatorLogin'], response.data)
},
})
}
+140 -1
View File
@@ -97,12 +97,19 @@ export function useTaskRuns(taskId: number | null) {
})
}
/**
* 作品**和博主**一起取。
*
* 博主不能从作品推出来 —— 「一条作品都没有的博主」在作品表里根本不存在,而它恰恰是
* 最该显示的一类。所以服务端单独给一份 `creators`,两者必须来自同一次请求,否则
* 组头和组里的作品会对不上(比如刚删掉一个任务)。
*/
export function useMonitorNotes(taskId: number | null, onlyNew: boolean) {
const platform = usePlatformParam()
return useQuery({
queryKey: ['monitorNotes', platform, taskId, onlyNew],
queryFn: async () =>
(await monitorApi.getNotes(taskId ?? undefined, onlyNew, 200, platform)).data.notes,
(await monitorApi.getNotes(taskId ?? undefined, onlyNew, 200, platform)).data,
refetchInterval: POLL_MS,
})
}
@@ -204,6 +211,74 @@ export function useClearCookie() {
})
}
// --- QR login (over CDP) ---------------------------------------------------
/** Polls only while a code is on screen; an idle session has nothing to watch. */
export function useQrLoginStatus(polling: boolean) {
return useQuery({
queryKey: ['monitorQrLogin'],
queryFn: async () => (await monitorApi.getQrLogin()).data,
refetchInterval: polling ? 2000 : false,
})
}
export function useStartQrLogin() {
const queryClient = useQueryClient()
const platform = usePlatformParam()
return useMutation({
mutationFn: () => monitorApi.startQrLogin(platform),
onSuccess: (response) => {
queryClient.setQueryData(['monitorQrLogin'], response.data)
},
onError: (error: Error) => {
// The backend's reason -- "Chrome is not reachable on ...", "no QR on the
// page" -- is the entire value here, and axios would flatten it to
// "Request failed with status code 502".
const detail = (error as { response?: { data?: { detail?: string } } })?.response
?.data?.detail
toast.error(`获取二维码失败:${detail ?? error.message}`)
},
})
}
export function useCancelQrLogin() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: () => monitorApi.cancelQrLogin(),
onSuccess: (response) => {
queryClient.setQueryData(['monitorQrLogin'], response.data)
},
})
}
/**
* Whether the browser is actually signed in.
*
* Tracked separately from the QR session above: that session lives in the
* server's memory and a restart erases it, while the profile it wrote to does
* not. Polled slowly when idle and briskly while a code is on screen.
*/
export function useLoginState(polling: boolean) {
return useQuery({
queryKey: ['monitorLoginState'],
queryFn: async () => (await monitorApi.getLoginState()).data,
refetchInterval: polling ? 4000 : 30000,
refetchOnWindowFocus: false,
})
}
/** Re-ask after reloading the page, for when the state looks stale. */
export function useRecheckLogin() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: () => monitorApi.getLoginState(true),
onSuccess: (response) => {
queryClient.setQueryData(['monitorLoginState'], response.data)
},
onError: (error: Error) => toast.error(`检测失败:${error.message}`),
})
}
// --- Settings -------------------------------------------------------------
export function useSettings() {
@@ -284,3 +359,67 @@ export function useTestWebhook() {
onError: (error: Error) => toast.error(`发送失败:${error.message}`),
})
}
// --- 博主备注 --------------------------------------------------------------
export function useSetCreatorAlias() {
const queryClient = useQueryClient()
const platform = usePlatformParam()
return useMutation({
mutationFn: ({ creatorHash, alias }: { creatorHash: string; alias: string }) =>
monitorApi.setCreatorAlias(creatorHash, alias, platform),
onSuccess: () => {
toast.success('备注已保存')
queryClient.invalidateQueries({ queryKey: ['monitorNotes'] })
queryClient.invalidateQueries({ queryKey: ['monitorCommentsGrouped'] })
},
onError: (error: Error) => toast.error(`备注保存失败:${error.message}`),
})
}
// --- 作品备注 --------------------------------------------------------------
export function useSetNoteAlias() {
const queryClient = useQueryClient()
const platform = usePlatformParam()
return useMutation({
mutationFn: ({ noteId, alias }: { noteId: string; alias: string }) =>
monitorApi.setNoteAlias(noteId, alias, platform),
onSuccess: () => {
toast.success('备注已保存')
queryClient.invalidateQueries({ queryKey: ['monitorNotes'] })
queryClient.invalidateQueries({ queryKey: ['monitorCommentsGrouped'] })
},
onError: (error: Error) => toast.error(`备注保存失败:${error.message}`),
})
}
// --- 上游更新检查 -----------------------------------------------------------------------------------------------------------------
export function useUpstreamStatus() {
return useQuery({
queryKey: ['monitorUpstream'],
queryFn: async () => (await monitorApi.getUpstream()).data,
// 检查本身是每天一次的量级,页面开着时慢点刷就够了。
refetchInterval: 60_000,
})
}
export function useCheckUpstream() {
const queryClient = useQueryClient()
return useMutation({
mutationFn: () => monitorApi.checkUpstream(),
onSuccess: (response) => {
const data = response.data
if (!data.ok) {
toast.error(`检查失败:${data.error ?? '未知错误'}`)
} else if ((data.behind ?? 0) > 0) {
toast.success(`上游有 ${data.behind} 个新提交`)
} else {
toast.success('已是最新,上游没有新提交')
}
queryClient.invalidateQueries({ queryKey: ['monitorUpstream'] })
},
onError: (error: Error) => toast.error(`检查失败:${error.message}`),
})
}
+64 -1
View File
@@ -5,17 +5,27 @@ import type {
CookieStatus,
MetricPoint,
MonitorComment,
MonitorCreator,
MonitorEvent,
MonitorNote,
MonitorOverview,
MonitorRun,
MonitorTask,
LoginState,
PlatformCapability,
QrLoginState,
ReportResult,
SettingsResponse,
TaskCreatePayload,
UpstreamStatus,
WebhookStatus,
} from '@/types/monitor'
import type {
CreatorAccount,
CreatorAccountDetail,
CreatorLoginState,
CreatorSyncResult,
} from '@/types/creator'
const api = axios.create({
baseURL: '/api',
@@ -177,7 +187,7 @@ export const monitorApi = {
api.get<{ runs: MonitorRun[] }>(`/monitor/tasks/${id}/runs`, { params: { limit } }),
getNotes: (taskId?: number, onlyNew = false, limit = 200, platform?: string) =>
api.get<{ notes: MonitorNote[] }>('/monitor/notes', {
api.get<{ notes: MonitorNote[]; creators: MonitorCreator[] }>('/monitor/notes', {
params: { task_id: taskId, only_new: onlyNew, limit, platform },
}),
// Metric time series for one note; the chart reads this.
@@ -256,10 +266,63 @@ export const monitorApi = {
return api.get<ReportResult>(`/monitor/report?${params.toString()}`)
},
// QR login over CDP. Reads the code out of the browser the crawler attaches
// to, so the operator can scan even when the host has no display.
startQrLogin: (platform?: string) =>
api.post<QrLoginState>('/monitor/login/qr', null, { params: { platform } }),
getQrLogin: () => api.get<QrLoginState>('/monitor/login/qr'),
cancelQrLogin: () => api.delete<QrLoginState>('/monitor/login/qr'),
/** `force` reloads the page first, for a stale-looking state. */
getLoginState: (force = false) =>
api.get<LoginState>('/monitor/login/state', { params: { force } }),
getWebhook: () => api.get<WebhookStatus>('/monitor/webhook'),
setWebhook: (url: string) => api.post('/monitor/webhook', { url }),
clearWebhook: () => api.delete('/monitor/webhook'),
testWebhook: (url?: string) => api.post('/monitor/webhook/test', { url: url ?? null }),
/** 最近一次上游检查的缓存结果;没有查过时是空对象。 */
getUpstream: () => api.get<UpstreamStatus>('/monitor/upstream'),
/** 给博主起备注(空串 = 清掉)。键是 creator_hash,跨任务同一个博主共用一条。 */
setCreatorAlias: (creatorHash: string, alias: string, platform?: string) =>
api.put(
`/monitor/creators/${encodeURIComponent(creatorHash)}`,
{ alias },
{ params: { platform } },
),
/** 给作品起备注(空串 = 清掉)。键是 note_id —— 和博主备注是两回事。 */
setNoteAlias: (noteId: string, alias: string, platform?: string) =>
api.put(
`/monitor/notes/${encodeURIComponent(noteId)}`,
{ alias },
{ params: { platform } },
),
/**
* 立刻检查一次。服务端要等 fetch 跑完才回答,而 axios 默认 30 秒对此不够 ——
* 一条卡住的 git fetch 能拖到两分钟,这里必须单独放长超时,否则会误报失败。
*/
checkUpstream: () =>
api.post<UpstreamStatus>('/monitor/upstream/check', null, { timeout: 150_000 }),
}
/**
* 运营模块。与 `monitorApi` 并列 —— 它管的是自己的账号,走的是纯请求的创作者后台,
* 和公开数据监控不是一回事。
*/
export const creatorApi = {
listAccounts: () => api.get<{ accounts: CreatorAccount[] }>('/creator/accounts'),
getAccount: (id: number) => api.get<CreatorAccountDetail>(`/creator/accounts/${id}`),
deleteAccount: (id: number) => api.delete(`/creator/accounts/${id}`),
checkAccount: (id: number) => api.post<CreatorAccount>(`/creator/accounts/${id}/check`),
// 同步在后端后台跑(分页 + 节流可能几分钟),这里只负责触发。
syncAccount: (id: number, days = 90) =>
api.post<CreatorSyncResult>(`/creator/accounts/${id}/sync`, null, { params: { days } }),
// 扫码新增账号。cookie 只在后端内存里流转,不会出现在这些响应里。
startLogin: () => api.post<CreatorLoginState>('/creator/login'),
getLogin: () => api.get<CreatorLoginState>('/creator/login'),
cancelLogin: () => api.delete<CreatorLoginState>('/creator/login'),
}
export default api
+49
View File
@@ -1,5 +1,7 @@
/** Formatting helpers for the monitoring dashboard. */
import type { ScheduleMode } from '@/types/monitor'
/** Compact count for display: mirrors how the platform itself abbreviates. */
export function formatCount(value: number | null | undefined): string {
if (value === null || value === undefined) return '—'
@@ -36,9 +38,56 @@ export function formatDateTime(ms: number | null | undefined): string {
return `${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())} ${pad(d.getHours())}:${pad(d.getMinutes())}`
}
/**
* 只到日。
*
* 发布日期问的是「哪一天发的」,绝对日期比「3天前」好认 —— 后者每天看都在变,而且
* 没法跟平台上的日期对。具体到分钟的那份放 title 里,需要时悬停看。
*/
export function formatDate(ms: number | null | undefined): string {
if (!ms) return '—'
const d = new Date(ms)
const pad = (n: number) => String(n).padStart(2, '0')
return `${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())}`
}
/** Human interval label for task cards. */
export function formatInterval(minutes: number): string {
if (minutes % 1440 === 0) return `${minutes / 1440} 天`
if (minutes % 60 === 0) return `${minutes / 60} 小时`
return `${minutes} 分钟`
}
const WEEKDAY_NAMES = '一二三四五六日'
/**
* Preview of a schedule, for the editor only.
*
* The task list renders the backend's own `schedule_label` instead. This exists
* so the current selection can be read back before it is saved -- which means the
* two are the same sentence written twice, and a change to one belongs in both.
*/
export function describeSchedule(
mode: ScheduleMode,
intervalMinutes: number,
hours: number[],
days: number[],
minute: number,
): string {
const pad = (value: number) => String(value).padStart(2, '0')
const sorted = (values: number[]) => [...values].sort((a, b) => a - b)
if (mode === 'interval') return `每 ${formatInterval(intervalMinutes)}`
if (hours.length === 0) return '未设置时间'
const clock = sorted(hours)
.map((hour) => `${pad(hour)}:${pad(minute)}`)
.join('、')
if (mode === 'daily' || days.length === 0) return `每天 ${clock}`
const labels = sorted(days)
.map((day) => `周${WEEKDAY_NAMES[day]}`)
.join('、')
return `${labels} ${clock}`
}
+91
View File
@@ -0,0 +1,91 @@
/** 运营模块:自己的小红书账号,以及创作者后台给的数据。 */
/**
* 数据权限状态。
*
* `pending` 是实测中最常见的一种:后台原话是「已为您申请数据权限,次日可查看」。
* 它必须和「没有权限」分开显示 —— 否则用户会以为采集坏了,其实只是在等审批。
*/
export type CreatorPermissionStatus = 'unknown' | 'pending' | 'active' | 'missing'
export type CreatorAccountStatus = 'ok' | 'expired' | 'error'
/** 注意:接口**不下发 cookie**,只给 `has_cookie`。 */
export interface CreatorAccount {
id: number
nickname: string
user_id: string
red_id: string
avatar: string
status: CreatorAccountStatus
permission_status: CreatorPermissionStatus
/** 后台原话,照抄不改写。 */
permission_tip: string
last_checked_at: number | null
last_synced_at: number | null
/** 上次同步用的范围(天)。0 表示从未同步过。 */
last_sync_days: number
last_error: string | null
has_cookie: boolean
created_at: number
note_count: number
}
/**
* 一篇作品的运营数据。
*
* 除时间外全部可为 null:接口没给、或解析不出来,都存 null 而不是 0 ——
* 0 是真实值,null 是"不知道",混在一起报表会说谎。
*/
export interface CreatorNote {
note_id: string
title: string
publish_time: number | null
exposure: number | null
views: number | null
likes: number | null
comments: number | null
favorites: number | null
shares: number | null
new_followers: number | null
danmaku: number | null
cover_ctr: number | null
avg_watch_seconds: number | null
two_second_exit_rate: number | null
completion_rate: number | null
captured_at: number
}
export interface CreatorSummary {
exposure: number
views: number
likes: number
comments: number
favorites: number
shares: number
new_followers: number
}
export interface CreatorAccountDetail {
account: CreatorAccount
notes: CreatorNote[]
summary: CreatorSummary
}
export type CreatorLoginStatus = 'idle' | 'waiting' | 'success' | 'expired' | 'error'
export interface CreatorLoginState {
status: CreatorLoginStatus
message: string
/** `data:image/...;base64,...`, empty unless status is `waiting`. */
image: string
elapsed: number
expires_in: number
/** 扫码成功后后端已落库的账号。 */
account: CreatorAccount | null
}
export interface CreatorSyncResult {
message: string
days: number
}
+196 -1
View File
@@ -28,6 +28,14 @@ export interface MonitorTarget {
enabled: boolean
}
/**
* How a task is scheduled.
*
* All three are expressible with pickers. A cron string is deliberately not
* supported -- it is a small language to learn just to say "every day at nine".
*/
export type ScheduleMode = 'interval' | 'daily' | 'weekly'
export interface MonitorTask {
id: number
name: string
@@ -35,12 +43,23 @@ export interface MonitorTask {
mode: MonitorMode
enabled: boolean
interval_minutes: number
schedule_mode: ScheduleMode
/** 0-23. Empty in interval mode. */
schedule_hours: number[]
/** 0-6 with Monday = 0, matching Python's `date.weekday()`. Weekly only. */
schedule_days: number[]
/** Minute past the hour, shared by every time in the schedule. */
schedule_minute: number
/** Composed by the backend, so the list and the editor cannot disagree. */
schedule_label: string
max_notes_count: number
enable_comments: boolean
max_comments_count: number
run_timeout_seconds: number
/** Opt-in per task so one webhook does not get flooded. */
/** 推送**新作品**。可能每轮都有,默认关以免刷屏。 */
notify_enabled: boolean
/** 推送**异常**(登录失效/运行失败/没抓到数据)。默认开。 */
notify_failures: boolean
/** Epoch milliseconds. */
next_run_at: number | null
last_run_at: number | null
@@ -67,6 +86,47 @@ export interface MonitorNote {
title: string
note_url: string
cover: string
/**
* 创作者标识。
*
* `creator_hash` 是唯一稳定的分组依据 —— 爬虫刻意不落原始 user_id(见
* tools/user_hash.py),所以没有比它更具体的身份了。
* `creator_name` 是昵称本身(本仓库关掉了脱敏,见 config.MASK_NICKNAME)。
*/
creator_hash: string
creator_name: string
/**
* 人自己给这个博主起的备注。
*
* 界面上**优先显示它**:`creator_name` 是平台昵称(常常认不出是谁),`creator_hash`
* 更认不出。备注是唯一能把账号对上人的东西。空串表示没起过。
*/
creator_alias: string
/**
* 人给**这条作品**起的备注。
*
* 和 `creator_alias` 是两件事:那条回答「这个账号是谁」,这条回答「这条我要盯着」。
* 一个博主底下常常只有一两件值得盯的作品,所以不能合并成一条。空串表示没起过。
*/
note_alias: string
/**
* 博主**账号级**指标 —— 作品列表给不了的东西:作品说的是「这条涨了多少赞」,
* 它说的是「这个人整个账号在涨还是在掉」。
*
* 三者都可能为 null:平台没采到(小红书那条路根本不产生它),或者某一项平台没给。
* 为 null 时界面整块不画 —— 画成「粉丝 0」就是在撒谎。
*
* 作品栏的组头**不要读这里**,读 `MonitorCreator` —— 一条作品都没有的博主不会出现
* 在作品列表里,只有那份博主列表能覆盖他。这里是给「顺着一条作品问它的作者」用的。
*/
creator_fans: number | null
creator_total_favorited: number | null
creator_works: number | null
/** 上述指标是哪一轮采到的。null = 从来没有过。 */
creator_stats_at: number | null
/** 作品的**发布**时间(爬虫侧:小红书 time、抖音 create_time)。可能与
* `first_seen_at` 差很远 —— 后者是「我们第一次看到它」的时间。平台没给时为 null。 */
published_at: number | null
first_seen_at: number
last_seen_at: number
is_new: boolean
@@ -76,6 +136,30 @@ export interface MonitorNote {
snapshot_count: number
}
/**
* 作品栏里的一位博主 —— **不依赖于他有没有作品**。
*
* 作品表是按 creator_hash 从作品推出来的,所以「一条作品都没有的博主」在那边根本
* 不存在。可这类博主恰恰是最该看见的:还在涨粉,只是最近没发东西。服务端因此单独
* 给一份列表,来源是**账号快照 ∪ 作品**。
*
* 没有作品时 `note_count` 是 0,其余字段照常有;反过来,小红书那条路不产生账号
* 快照,于是 `creator_fans` 等为 null —— 两边各缺一块,界面要都能显示。
*/
export interface MonitorCreator {
task_id: number
creator_hash: string
creator_name: string
creator_alias: string
note_count: number
creator_fans: number | null
creator_total_favorited: number | null
creator_works: number | null
creator_stats_at: number | null
/** 作品最后出现、或账号指标最后采集的时间 —— 排序用。 */
last_activity_at: number
}
export interface MetricPoint {
run_id: number
captured_at: number
@@ -99,6 +183,11 @@ export interface MonitorComment {
note_title: string
note_cover: string
note_url: string
/** 所属作品的创作者 —— 评论流按 博主 -> 作品 -> 评论 三级展开时用。 */
note_creator_hash: string
note_creator_name: string
/** 所属作品的发布时间。 */
note_published_at: number | null
}
/** One work that has comments, for the filter dropdown. */
@@ -107,6 +196,8 @@ export interface CommentNoteOption {
note_title: string
note_cover: string
note_url: string
creator_hash: string
creator_name: string
comment_count: number
latest_at: number
}
@@ -117,6 +208,10 @@ export interface CommentBucket {
note_title: string
note_cover: string
note_url: string
creator_hash: string
creator_name: string
/** 作品的发布时间。 */
published_at: number | null
comments: MonitorComment[]
}
@@ -161,6 +256,48 @@ export interface CookieStatus {
last_ok_at: number | null
}
export type QrLoginStatus = 'idle' | 'waiting' | 'success' | 'expired' | 'error'
/**
* A QR login driven over CDP, for servers with no display to show one on.
*
* `success` also covers "the browser was already signed in" — the scan and the
* existing session both just mean the profile the crawler attaches to is usable.
*/
export interface QrLoginState {
status: QrLoginStatus
platform: string | null
/** `data:image/...;base64,...`. Empty unless status is `waiting`. */
image: string
message: string
/** Seconds since the code was fetched. */
elapsed: number
/** Seconds left before the code is written off. */
expires_in: number
/** Whether the browser reports being signed in right now. */
logged_in: boolean
nickname: string | null
/** 登录态是否已同时存进库(供非 CDP 模式的定时任务使用)。 */
cookie_saved?: boolean
}
/**
* The browser's own answer to "am I signed in?".
*
* Asked of the page, not derived from the QR session -- that session lives in the
* server's memory and dies on a restart, so tying the answer to it makes a
* successful scan look like nothing happened.
*
* `known` is false when the page could not report (not loaded, browser
* unreachable); `logged_in` is then meaningless rather than false.
*/
export interface LoginState {
known: boolean
logged_in: boolean
nickname: string | null
error?: string
}
export interface MonitorOverview {
tasks: number
enabled_tasks: number
@@ -173,14 +310,26 @@ export interface MonitorOverview {
export interface TaskCreatePayload {
name: string
/**
* 任务归属的平台。
*
* **必须显式带上。** 后端在缺省时会退回小红书 —— 那是接口早期的兼容行为,
* 于是「在抖音页面上建任务」会安安静静地建出一个小红书任务(遇到过)。
*/
platform: string
mode: MonitorMode
interval_minutes: number
schedule_mode: ScheduleMode
schedule_hours: number[]
schedule_days: number[]
schedule_minute: number
max_notes_count: number
enable_comments: boolean
max_comments_count: number
run_timeout_seconds: number
enabled: boolean
notify_enabled: boolean
notify_failures: boolean
targets: string[]
}
@@ -236,6 +385,17 @@ export interface ReportResult {
* layer has been hooked up for it. The UI must never conflate the two -- a
* platform can be fully crawlable upstream and still unusable here.
*/
/** 目标输入框的示例与措辞,随平台变。服务端给,前端不自己判断平台。 */
export interface TargetHints {
creator?: string
note?: string
/** 该平台怎么称呼这两样东西 —— 抖音叫「作品」,小红书叫「笔记」。 */
creator_label?: string
note_label?: string
/** 链接里是否带会过期的令牌(只有小红书有)。 */
token_expires?: boolean
}
export interface PlatformCapability {
value: string
label: string
@@ -245,6 +405,7 @@ export interface PlatformCapability {
comment_levels: number
media: boolean
monitor_wired: boolean
target_hints?: TargetHints
}
export interface WebhookStatus {
@@ -297,3 +458,37 @@ export interface SettingsResponse {
secrets: Record<string, SecretStatus>
specs: SettingSpec[]
}
// --- 上游更新检查 -----------------------------------------------------------
export interface UpstreamCommit {
sha: string
author: string
/** 提交日期,`YYYY-MM-DD`(git 侧已格式化,不做本地化)。 */
date: string
subject: string
}
/**
* 最近一次上游检查的结果,服务端缓存。
*
* `checked_at` 为空表示还没查过(`GET /monitor/upstream` 返回空对象);`ok` 为
* false 时 `error` 一定有值 —— 上游不通是常态,那也是一条要显示出来的结论。
*/
export interface UpstreamStatus {
checked_at?: number | null
remote_url?: string
branch?: string
ok?: boolean
/** 上游有、当前代码没有的提交数 —— 要合的就是这些。 */
behind?: number
/** 当前代码有、上游没有的提交数 —— 也就是这一层改动自己的规模。 */
ahead?: number
tip?: string
head?: string
commits?: UpstreamCommit[]
error?: string
/** 本次检查是否推送了通知。 */
notified?: boolean
notify_error?: string
}