Files
butubb 37ca1b1cd6
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
feat: CDP 接管开关 + 扫码登录面板 + Docker 部署
- runner: enable_cdp_mode 从硬编码 False 改为系统设置 cdp_enabled。服务器部署下爬虫接管已开启远程调试的 Chrome(默认 9222),复用其 profile 登录态;本机桌面默认仍为关,行为不变
- qrlogin: 新增 CDP 扫码登录。Chrome 在服务器上跑于 Xvfb,show_qrcode 依赖的 PIL 桌面看图程序不存在,二维码无处可显示;改为经 CDP 从页面取出二维码交给 WebUI 渲染。刻意复用 browser.contexts[0](新建 context 是无痕 profile,扫了也白扫),且绝不调用 browser.close()(会连带关掉操作者自己的 Chrome)
- webui: 设置页新增扫码面板,替换原本跳到采集页看终端二维码的入口
- Dockerfile / .dockerignore / docker-compose.yml: 服务器部署。host 网络是必需而非图省事——容器里 127.0.0.1:9222 必须落到宿主机回环
- UPSTREAM.md: 补充 gitcode 镜像,用于 GitHub 大包传输必断时补历史
2026-10-07 10:41:11 +08:00

151 lines
7.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 与上游的差异管理
本仓库在 [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) 之上加了一层
监控/鉴权/多平台面板。这份文档记录**改了上游哪些文件、为什么**,以及**上游更新时怎么合并**。
---
## 一、改动分三类
冲突风险从低到高:
### 1. 纯新增文件(零冲突)
上游怎么改都不会碰到它们:
```
api/auth.py WebUI 登录鉴权
api/monitor/* 监控层整体(含 platforms.py 能力矩阵)
api/monitor/db.py MySQL 连接层(可回退 SQLite 供测试用)
api/monitor/migrate_from_sqlite.py SQLite → MySQL 一次性迁移脚本
api/routers/{auth,monitor,settings}.py
api/schemas/{auth,monitor,settings}.py
api/services/interpreter.py 解释器探测(uv / .venv / 当前解释器)
webui/src/components/{monitor,settings,auth}/ 新视图
webui/src/components/layout/{PlatformSwitcher,UnwiredPlatformNotice}.tsx
webui/src/{hooks/useMonitor.ts,hooks/usePlatform.ts,store/platformStore.ts,lib/monitorFormat.ts,types/monitor.ts}
docs/监控功能使用说明.md
tests/test_{auth,settings,platforms,qrlogin,monitor_*}.py
Dockerfile / .dockerignore / docker-compose.yml 服务器部署用
```
其中 `api/monitor/qrlogin.py` + `webui/.../QrLoginPanel.tsx` 是**服务器专用**的扫码登录:
那台机器上 Chrome 跑在 Xvfb 里,`show_qrcode` 调的 PIL `Image.show()` 需要桌面看图程序,
服务器没有,二维码会无处可去。所以改成用 CDP 把二维码从页面里读出来交给前端 `<img>` 显示。
### 2. 加法改动(低冲突)
只在既有文件里**新增**内容,不改动原有行:
| 文件 | 加了什么 |
|---|---|
| `cmd_arg/arg.py` | typer 选项:`--enable_cdp_mode`、`--inject_all_cookies`、`--save_login_state`、`--cookies_file`、`--crawler_max_sleep_sec`,以及对应的 `config.*` 回写 |
| `api/schemas/crawler.py` | `CrawlerStartRequest` 的若干**可选**字段(默认 `None`,不传则不加对应 CLI 参数) |
| `config/base_config.py` | `INJECT_ALL_COOKIES = False` |
| `api/routers/__init__.py` | 导出新增的 router |
| `requirements.txt` | 补上 `websockets`(上游 `pyproject.toml` 里有、`requirements.txt` 里漏了) |
| `tests/conftest.py` | 新增 `_bypass_auth_for_non_auth_suites` fixture |
### 3. 接线改动(中冲突,需要人看)
| 文件 | 改了什么 | 上游若在此处变动 |
|---|---|---|
| `api/main.py` | 注册 4 个 router 并加 `Depends(require_auth)`;`lifespan` 里初始化监控库、启动调度器、跑设置键迁移;`load_dotenv`;CORS 可配;`docs/redoc/openapi` 关闭;监听地址改 env | **最需要人工合并的文件**。留意 router 注册块、lifespan、`__main__` |
| `api/routers/websocket.py` | 两个 WS 路由加 `dependencies=[Depends(require_ws_auth)]` | 上游若新增 WS 路由,**必须同样加上**,否则那条流是裸奔的 |
| `api/services/crawler_manager.py` | 解释器探测替换硬编码 `uv run`;`_build_command` 转发新增参数;新增 `is_busy()` / `run_and_wait()` 与完成事件 | 留意 `_build_command` 的参数拼装 |
| `media_platform/xhs/login.py` | `login_by_cookies` 在 `INJECT_ALL_COOKIES` 打开时注入**全部** cookie(默认关闭,行为不变) | 小改动,好合并 |
### 4. 上游 bug 修复(建议回馈上游)
| 文件 | 修的问题 |
|---|---|
| `media_platform/xhs/core.py` | 见下节 |
| `media_platform/xhs/login.py` | 同上(cookie 加固) |
---
## 二、应该给上游提 PR 的两个修复
这两处是**上游自身的缺陷**,提上去以后就不用自己背着:
### 1. 博主主页抓取失败会跳掉整个博主(`xhs/core.py`)
`get_creator_info()` 抓主页 HTML 解析 `window.__INITIAL_STATE__`,解析失败抛 `JSONDecodeError`——
它是 `ValueError` 的子类,被 `except ValueError` 误捕获,日志报成
"Failed to parse creator URL"(**误导**,URL 根本没解析错),然后 `continue` **跳过整个博主**。
而那份资料只喂给 `save_creator()`,它在教学版里是**空函数**。也就是说:
一个喂给空函数的抓取失败,让真正要抓的作品一条都没抓到,表现为"0 篇作品",
和"登录失效"长得一模一样。
修复:把资料抓取改成**尽力而为**,失败只警告、继续抓作品。
### 2. cookie 登录只注入 `web_session`(`xhs/login.py`)
`a1` / `webId` 等签名所需 cookie 只能靠持久化 profile 补,冷启动时签名会失败。
默认行为保持不变,用 `INJECT_ALL_COOKIES` 开关控制。
---
## 三、上游更新时怎么操作
### 日常流程
```bash
git stash # 或先 commit 到自己的分支(推荐)
git fetch origin main
git rebase origin/main # 冲突只会出现在上表第 3、4 类文件里
./.venv/Scripts/python.exe -m pytest tests/ -q # 502 个测试就是回归网
```
### 直连 GitHub 不通时(本机常见)
本机到 `github.com` 时通时不通,**大包传输必断**(`Recv failure: Connection was reset`
或 `unexpected disconnect while reading sideband packet`),所以 `git clone` / `--unshallow`
这类一次性拉全量的操作基本必失败。可用的替代源:
```bash
# gitcode 的 GitHub 镜像,国内直连,比 GitHub 本身还新一天以内
git remote add gitcode https://gitcode.com/gh_mirrors/me/MediaCrawler.git
git fetch --no-tags --unshallow gitcode # 本仓库当初就是这样补全历史的,约 2 秒
```
注意 `git fetch` 只写 `refs/remotes/*`,**不会动本地 `main`**;
但拉镜像会把 `upstream/main` 指到镜像的 tip(可能比 GitHub 晚一天),
等 GitHub 通了再 `git fetch upstream` 正回来即可。
### 强烈建议:先把改动提交掉
当前状态是**未提交**的(25 个上游文件被改 + 31 个新文件)。在 `main` 分支上裸着工作区,
一次 `git checkout .` 就全没了,而且没法 rebase。
```bash
git checkout -b local/monitor-panel
git add -A && git commit -m "监控面板 / 鉴权 / 多平台"
```
### 如果改动持续增长:fork
把本仓库 fork 到自己名下,上游设为 remote:
```bash
git remote rename origin upstream
git remote add origin <你的 fork>
git push -u origin local/monitor-panel
```
之后同步上游用 `git fetch upstream && git rebase upstream/main`。
---
## 四、合并时最容易忘的三件事
1. **新增的 `/api` 路由必须带鉴权**。跑一下 `tests/test_auth.py`——
里面有个测试会遍历 `app.routes`,断言除豁免集外每个 `/api` 路由无凭据都返回 401。
上游新增接口忘了加鉴权,这个测试会直接失败。
2. **新增的 WebSocket 路由必须加 `require_ws_auth`**。
`BaseHTTPMiddleware` 对 WS 完全不生效(`scope["type"] != "http"` 直接放行),
只靠中间件会漏。同样有测试守着。
3. **上游若改动 `AsyncFileWriter` 的输出路径规则**,`api/monitor/ingest.py::find_run_files`
会跟着失效——它靠 glob `{out_dir}/{platform}/jsonl/*_contents_*.jsonl` 定位每轮的产物。