Compare commits
49
Commits
4e60524f37
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
61808444ad | ||
|
|
e0581682e1 | ||
|
|
3c4daae0ac | ||
|
|
d02fec5914 | ||
|
|
20e672834c | ||
|
|
0a88474c92 | ||
|
|
5552e2a2b8 | ||
|
|
9e13a7f686 | ||
|
|
9f70cd0924 | ||
|
|
95b1be2c2e | ||
|
|
3b6a437e6c | ||
|
|
4f83a075f6 | ||
|
|
7442bc10b8 | ||
|
|
21ff01b894 | ||
|
|
e4affe9170 | ||
|
|
1118d466be | ||
|
|
3486c7f524 | ||
|
|
cd85587f00 | ||
|
|
67837b407e | ||
|
|
42f2209534 | ||
|
|
f3ea088c75 | ||
|
|
e77e5e2f15 | ||
|
|
06718a1351 | ||
|
|
e348de48d3 | ||
|
|
44cbe8e2aa | ||
|
|
de9ff58371 | ||
|
|
f6ddc46d62 | ||
|
|
2e6fa955b0 | ||
|
|
0eb6ba31c5 | ||
|
|
b11bbf771a | ||
|
|
fa227600fe | ||
|
|
bef0a4fbde | ||
|
|
2f5852e311 | ||
|
|
5484e5a3ae | ||
|
|
c2b310c7bf | ||
|
|
2613f7577f | ||
|
|
2b9ebdad87 | ||
|
|
6eff6fcc83 | ||
|
|
e608b51210 | ||
|
|
d937ff5fe6 | ||
|
|
8f4e5586e9 | ||
|
|
242e3a7837 | ||
|
|
c1068845a9 | ||
|
|
89b7b1e825 | ||
|
|
fe45442b01 | ||
|
|
74a592024c | ||
|
|
cffb407d15 | ||
|
|
bcc7361145 | ||
|
|
37ca1b1cd6 |
@@ -0,0 +1,29 @@
|
||||
# Keep the build context to what the server actually runs.
|
||||
.git
|
||||
.github
|
||||
.venv
|
||||
venv
|
||||
|
||||
# The WebUI sources are not needed -- only the bundle they produce, which lands
|
||||
# in api/webui and is therefore NOT excluded.
|
||||
webui/node_modules
|
||||
webui/src
|
||||
webui/dist
|
||||
|
||||
# Runtime state: per-run crawler output and the login browser profile. Mounted
|
||||
# as a volume instead, so it survives image rebuilds.
|
||||
data
|
||||
browser_data
|
||||
|
||||
# Secrets come from the environment via compose, never baked into a layer.
|
||||
.env
|
||||
|
||||
tests
|
||||
docs
|
||||
__pycache__
|
||||
**/__pycache__
|
||||
*.pyc
|
||||
*.pyo
|
||||
.pytest_cache
|
||||
*.db
|
||||
*.log
|
||||
+12
-1
@@ -183,4 +183,15 @@ agent_zone
|
||||
debug_tools
|
||||
|
||||
database/*.db
|
||||
.omx/
|
||||
.omx/
|
||||
|
||||
# 别人放在这儿的参考项目(mac-agent-os)。它是独立仓库、21MB,不属于本项目 ——
|
||||
# 一旦被 `git add -A` 扫进来就是永久留在历史里(踩过一次:1601 个文件里 1429 个是它)。
|
||||
# 要读它就直接读磁盘上的目录,别提交。
|
||||
mac-agent-os-main/
|
||||
|
||||
# 服务器上重建前端时,容器里的 npm 往挂载目录写的缓存。是构建产物,不该进仓库 ——
|
||||
# 它长期以「未跟踪」状态躺在工作区,正是会被 `git add -A` 顺手带走的类型。
|
||||
webui/.npm/
|
||||
# 我用来在服务器上跑命令的一次性脚本(不属于这个项目)
|
||||
.remote_run.py
|
||||
|
||||
+71
@@ -0,0 +1,71 @@
|
||||
# Server deployment image.
|
||||
#
|
||||
# No browser is bundled on purpose. On this deployment the crawler attaches over
|
||||
# CDP to the Chrome already running on the host (see the 接管已有 Chrome setting),
|
||||
# so shipping a second copy of Chromium would only add hundreds of megabytes and
|
||||
# a login state that nothing uses. The Playwright Python package is still needed
|
||||
# -- that is what speaks CDP -- hence PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD.
|
||||
FROM python:3.11-slim
|
||||
|
||||
ENV PYTHONUNBUFFERED=1 \
|
||||
PYTHONDONTWRITEBYTECODE=1 \
|
||||
PIP_DISABLE_PIP_VERSION_CHECK=1 \
|
||||
PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1 \
|
||||
TZ=Asia/Shanghai
|
||||
|
||||
# asyncmy compiles a Cython extension, so a toolchain has to exist at build time.
|
||||
# It is left installed: purging it risks taking libmysqlclient with it, and a
|
||||
# slightly larger image is cheaper than a runtime that fails months later.
|
||||
#
|
||||
# deb.debian.org is effectively unusable from this network -- it was pulling the
|
||||
# 96 MB of build dependencies at roughly 13 kB/s, which puts a build at well over
|
||||
# half an hour. Point apt at a domestic mirror; override APT_MIRROR when building
|
||||
# from somewhere that does not need it.
|
||||
#
|
||||
# The libgl1/libxcb1/... group is not for a GUI: opencv-python links against X11
|
||||
# at import time, and tools/utils.py reaches cv2 through slider_util, so without
|
||||
# them the *application* fails to import, not just some image utility. tzdata is
|
||||
# here because TZ=Asia/Shanghai is silently ignored without it, which would put
|
||||
# every stored timestamp in UTC. nodejs is for PyExecJS: douyin/help.py compiles
|
||||
# libs/douyin.js at *import* time, and because main.py imports every platform,
|
||||
# that single platform being importable-or-not decides whether the whole app
|
||||
# (and the environment self-check) comes up. npm rides along so the WebUI can be
|
||||
# rebuilt on the server (see deploy.sh) instead of only on a workstation --
|
||||
# corepack is present but does not cover npm, only yarn and pnpm.
|
||||
#
|
||||
# The pip mirror is set for the same reason as the apt one: this host's route to
|
||||
# the public index is slow.
|
||||
#
|
||||
# git is for the 上游更新检查 (api/monitor/upstream.py): it fetches the upstream
|
||||
# repository into the mounted checkout to count how far behind this fork is.
|
||||
# python:slim does not ship git, and nothing else here pulls it in.
|
||||
ARG APT_MIRROR=mirrors.tuna.tsinghua.edu.cn
|
||||
RUN set -eux; \
|
||||
for f in /etc/apt/sources.list /etc/apt/sources.list.d/debian.sources; do \
|
||||
if [ -f "$f" ]; then \
|
||||
sed -i "s|deb.debian.org|${APT_MIRROR}|g; s|security.debian.org|${APT_MIRROR}|g" "$f"; \
|
||||
fi; \
|
||||
done; \
|
||||
apt-get update; \
|
||||
apt-get install -y --no-install-recommends \
|
||||
build-essential pkg-config default-libmysqlclient-dev git \
|
||||
libgl1 libglib2.0-0 libsm6 libxext6 libxrender1 libxcb1 libgomp1 \
|
||||
tzdata nodejs npm; \
|
||||
rm -rf /var/lib/apt/lists/*
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
# Requirements only -- this is the one layer that is expensive to build and
|
||||
# changes rarely.
|
||||
ARG PIP_INDEX=https://pypi.tuna.tsinghua.edu.cn/simple
|
||||
COPY requirements.txt ./
|
||||
RUN pip install --no-cache-dir -i "$PIP_INDEX" -r requirements.txt
|
||||
|
||||
# The application code is deliberately NOT copied in. compose mounts it at /app,
|
||||
# so a code change is "re-upload the tarball, restart the container" instead of
|
||||
# an image rebuild. Treat this image as the dependency layer and nothing else;
|
||||
# rebuild it when, and only when, requirements.txt or this file changes.
|
||||
EXPOSE 18051
|
||||
|
||||
# api.main reads MC_HOST / MC_PORT from the environment; compose supplies both.
|
||||
CMD ["python", "-m", "api.main"]
|
||||
@@ -0,0 +1,874 @@
|
||||
# 内容抓取模块 · 开发技术说明
|
||||
|
||||
> 面向接手开发的团队 · 2026-10-10
|
||||
> 全部内容基于**逐行读源码**整理,不是推测
|
||||
> 范围:**只写内容抓取模块**,不涉及其他业务
|
||||
|
||||
---
|
||||
|
||||
## 一、这个模块是干什么的
|
||||
|
||||
从主流内容平台(抖音 / 小红书 / B站 / 知乎 / 任意网页)**采集内容数据**:
|
||||
- 视频/笔记的元数据(标题、作者、发布时间、正文)
|
||||
- 互动数据(点赞、评论、分享、播放)
|
||||
- 评论列表
|
||||
- 作者主页的全部作品列表
|
||||
|
||||
采到的数据落进本地 SQLite,供后续分析/运营使用。
|
||||
|
||||
---
|
||||
|
||||
## 二、整体架构(重要:三层降级是核心)
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
│ HTTP 层 routes/scrape.py(22 个端点) │
|
||||
└───────────────────────┬──────────────────────────────────┘
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
│ 引擎层 services/scrape_engine.py │
|
||||
│ · URL 解析 → 标准化目标 │
|
||||
│ · 按平台选适配器 │
|
||||
│ · 同步 / 异步调度 │
|
||||
│ · 结果落库(去重) │
|
||||
└───────────────────────┬──────────────────────────────────┘
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
│ 适配器层 services/adapters/*.py(每个平台一个) │
|
||||
│ 基类 ScrapeAdapter 定义统一接口 + 工具降级 │
|
||||
└───────────────────────┬──────────────────────────────────┘
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
│ 工具层(★ 三层降级,这是本模块的核心设计) │
|
||||
│ Level 1 OpenCLI 外部 Node CLI(主路径) │
|
||||
│ Level 2 agent-browser Playwright + 真实 Chrome │
|
||||
│ Level 3 web_crawler 通用网页兜底 │
|
||||
└───────────────────────┬──────────────────────────────────┘
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
│ 存储层 services/scrape_db.py(SQLite,4 张表) │
|
||||
└──────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### 为什么这么设计
|
||||
|
||||
抓取的最大风险是**单一方式失效**:目标平台改版、接口封禁、登录态过期。
|
||||
所以**同一份数据有三条获取路径**,第一条失败自动降级到第二条,
|
||||
**上层完全不感知**(对 engine 来说只是"拿到数据了")。
|
||||
|
||||
---
|
||||
|
||||
## 三、工具层详解(最关键的一层)
|
||||
|
||||
### 3.1 Level 1 — OpenCLI
|
||||
|
||||
**它是什么**:一个**第三方 Node.js CLI 工具**,包名 `@jackwener/opencli`。
|
||||
|
||||
```
|
||||
实际安装位置(本机实测):
|
||||
~/.workbuddy/binaries/node/versions/22.22.2/bin/opencli
|
||||
→ 软链到 ../lib/node_modules/@jackwener/opencli/dist/src/main.js
|
||||
```
|
||||
|
||||
**怎么调用**(`services/adapters/__init__.py` 的 `_run_opencli`):
|
||||
|
||||
```python
|
||||
OPENCLI = os.environ.get("OPENCLI_PATH",
|
||||
str(Path.home() / ".workbuddy" / "binaries" / "node" /
|
||||
"versions" / "22.22.2" / "bin" / "opencli"))
|
||||
|
||||
async def _run_opencli(self, args: list, timeout: int = 60):
|
||||
cmd = [self.OPENCLI] + args
|
||||
proc = await asyncio.create_subprocess_exec(
|
||||
*cmd, stdout=PIPE, stderr=PIPE)
|
||||
stdout, stderr = await asyncio.wait_for(proc.communicate(), timeout=timeout)
|
||||
...
|
||||
return self._parse_output(stdout.decode().strip())
|
||||
```
|
||||
|
||||
**调用示例**(抖音,`douyin_scrape.py`):
|
||||
```bash
|
||||
opencli douyin user-videos <sec_uid> --limit 20 --with_comments true -f json
|
||||
opencli douyin stats <aweme_id> -f json
|
||||
```
|
||||
|
||||
**⛔ 移植注意**:这个二进制**不在仓库里**,是外部依赖。移植时必须:
|
||||
- 要么在目标机装 `npm i -g @jackwener/opencli`
|
||||
- 要么改 `OPENCLI_PATH` 环境变量指向它的位置
|
||||
|
||||
### 3.2 Level 2 — agent-browser("套用真实浏览器"的做法)
|
||||
|
||||
**这是你问的重点。设计原则写在 `browser_helpers.py` 文件头**:
|
||||
|
||||
```
|
||||
⛔ 绝不使用 Camoufox(养号专用,Firefox 内核 + 特殊指纹)
|
||||
✅ 使用 Playwright 启动【真实 Chrome】(Chromium 内核,正常指纹)
|
||||
```
|
||||
|
||||
**具体怎么"套真实浏览器"**(`browser_helpers.py` 的 `_get_browser`):
|
||||
|
||||
```python
|
||||
_CHROME_PATHS = [
|
||||
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome", # ← 系统真 Chrome
|
||||
"/Applications/Chromium.app/Contents/MacOS/Chromium",
|
||||
]
|
||||
_CHROME_PATH = None
|
||||
for p in _CHROME_PATHS:
|
||||
if Path(p).exists():
|
||||
_CHROME_PATH = p # 自动探测,找到就用系统已装的 Chrome
|
||||
break
|
||||
|
||||
launch_kwargs = {
|
||||
"headless": headless,
|
||||
"args": [
|
||||
"--disable-blink-features=AutomationControlled", # ★ 反检测关键
|
||||
"--no-sandbox",
|
||||
"--disable-dev-shm-usage",
|
||||
"--disable-gpu",
|
||||
"--window-size=1280,720",
|
||||
],
|
||||
}
|
||||
if _CHROME_PATH:
|
||||
launch_kwargs["executable_path"] = _CHROME_PATH # ★ 用系统 Chrome,不用 Playwright 自带
|
||||
```
|
||||
|
||||
**四个关键设计点**:
|
||||
|
||||
| 点 | 做法 | 为什么 |
|
||||
|---|---|---|
|
||||
| **用什么内核** | 系统真实 Chrome(`executable_path` 指定) | Playwright 自带 Chromium 有明显特征;真实 Chrome 是正常用户指纹 |
|
||||
| **怎么隐藏自动化** | `--disable-blink-features=AutomationControlled` | 这是最常被检测的自动化标志位 |
|
||||
| **实例管理** | 模块级单例 `_browser` + `asyncio.Lock` | 避免每次请求都启动浏览器(启动 ~1-2 秒) |
|
||||
| **会话隔离** | 每次 `browser.new_context()` | 每个任务独立 cookie 环境,互不污染 |
|
||||
|
||||
**页面加载策略**(`page_evaluate`):
|
||||
```python
|
||||
await page.goto(url, wait_until="domcontentloaded", timeout=timeout)
|
||||
await page.wait_for_load_state("networkidle", timeout=timeout) # 等动态渲染
|
||||
await asyncio.sleep(1) # 再等 1 秒保险
|
||||
result = await page.evaluate(js_code) # 执行 JS 提取
|
||||
```
|
||||
|
||||
**UA 伪装**(每次 context 都设置):
|
||||
```python
|
||||
context = await browser.new_context(
|
||||
user_agent=("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
|
||||
"AppleWebKit/537.36 (KHTML, like Gecko) "
|
||||
"Chrome/125.0.0.0 Safari/537.36"),
|
||||
viewport={"width": 1280, "height": 720},
|
||||
locale="zh-CN", # 中文环境,符合目标用户画像
|
||||
)
|
||||
```
|
||||
|
||||
**两个公开函数**:
|
||||
```python
|
||||
page_evaluate(url, js_code, timeout, headless) # 打开页面执行 JS,返回 dict
|
||||
page_extract(url, selectors, timeout, headless) # 按 CSS 选择器提取文本
|
||||
```
|
||||
|
||||
**Profile 持久化**(可选,环境变量控制):
|
||||
```python
|
||||
_USER_DATA_DIR = os.environ.get(
|
||||
"SCRAPE_CHROME_USER_DATA",
|
||||
str(Path.home() / "workbuddy-agent-os" / "agent-local" /
|
||||
"runtime" / "scrape_chrome_profile"))
|
||||
```
|
||||
→ 想保持登录态就指定这个目录;不指定就用临时 context。
|
||||
|
||||
### 3.3 Level 3 — web_crawler
|
||||
|
||||
通用网页抓取兜底(`web_scrape.py`),用于非主流平台的页面。
|
||||
|
||||
### 3.4 降级怎么触发(`ScrapeAdapter._try_tools`)
|
||||
|
||||
```python
|
||||
async def _try_tools(self, tool_level: int, funcs: list) -> tuple:
|
||||
"""funcs: [(工具名, 可调用对象), ...],按顺序尝试"""
|
||||
tools = [f for f in funcs[:tool_level]] # 按 level 截断
|
||||
for name, fn in tools:
|
||||
try:
|
||||
result = await fn()
|
||||
if result is not None:
|
||||
return True, result, name # ★ 成功即返回,不再降级
|
||||
except Exception as e:
|
||||
logger.warning(f" ⚠️ [{self.platform}] 工具 {name} 失败: {e}")
|
||||
return False, None, tools[-1][0] if tools else "none"
|
||||
```
|
||||
|
||||
**关键语义**:
|
||||
- `tool_level=1` → 只用 OpenCLI
|
||||
- `tool_level=2` → OpenCLI → agent-browser(**默认**)
|
||||
- `tool_level=3` → 三层全开
|
||||
- **第一个成功就停**;返回 `(成功?, 结果, 用了哪个工具)`
|
||||
|
||||
---
|
||||
|
||||
## 四、适配器层(每平台一个)
|
||||
|
||||
### 4.1 统一接口(`services/adapters/__init__.py` 的基类)
|
||||
|
||||
**每个平台适配器必须实现 4 个方法**:
|
||||
|
||||
```python
|
||||
class ScrapeAdapter:
|
||||
platform = "" # 子类覆写,如 "douyin"
|
||||
adapter_name = ""
|
||||
|
||||
async def collect_item(self, target, depth="light", tool_level=2) -> dict:
|
||||
"""抓单条内容详情"""
|
||||
async def collect_user(self, user_id, limit=20) -> list[dict]:
|
||||
"""抓某个用户/作者的全部作品"""
|
||||
async def collect_comments(self, item_id, limit=20) -> list[dict]:
|
||||
"""抓评论"""
|
||||
async def collect_search(self, keyword, limit=20) -> list[dict]:
|
||||
"""按关键词搜索"""
|
||||
```
|
||||
|
||||
**基类提供的公共能力**:
|
||||
- `_try_tools(tool_level, funcs)` — 降级执行(见 3.4)
|
||||
- `_run_opencli(args, timeout)` — 调 OpenCLI + 解析输出
|
||||
- `_parse_output(text)` — **输出格式三层兜底解析**(见下)
|
||||
- `_parse_lines(text)` — 纯文本兜底解析
|
||||
|
||||
### 4.2 输出格式三层兜底(细节,容易踩坑)
|
||||
|
||||
OpenCLI 的输出格式**不保证稳定**,所以解析做了三层:
|
||||
|
||||
```python
|
||||
def _parse_output(self, text: str):
|
||||
# 1. JSON 检测(以 [ 或 { 开头)→ json.loads
|
||||
if text.startswith("[") or text.startswith("{"):
|
||||
try:
|
||||
return json.loads(text)
|
||||
except json.JSONDecodeError:
|
||||
logger.warning("JSON 解析失败,尝试 YAML 兜底")
|
||||
|
||||
# 2. YAML 解析(PyYAML 可用时)
|
||||
try:
|
||||
import yaml
|
||||
parsed = yaml.safe_load(text)
|
||||
if parsed is not None:
|
||||
return parsed
|
||||
except ImportError:
|
||||
logger.debug("PyYAML 未安装,跳过 YAML")
|
||||
|
||||
# 3. 纯文本兜底:按行解析 key:value
|
||||
return self._parse_lines(text)
|
||||
```
|
||||
|
||||
`_parse_lines` 甚至**专门处理了 `top_comments` 块**(评论在纯文本里的多行结构)。
|
||||
|
||||
**⛔ 移植注意**:如果目标环境没有 PyYAML,会静默降级到第三层
|
||||
(能跑,但嵌套结构会丢)。
|
||||
|
||||
### 4.3 各平台降级链(实测,**差异很大**)
|
||||
|
||||
```python
|
||||
# 抖音(douyin_scrape.py)—— 两条路径
|
||||
await self._try_tools(tool_level, [
|
||||
("opencli", lambda: self._opencli_user_videos(...)),
|
||||
("agent-browser", lambda: self._browser_user_profile(...)), # ← 有浏览器降级
|
||||
])
|
||||
|
||||
# 小红书 / B站 / 知乎 —— ⚠️ 只有一条路径
|
||||
await self._try_tools(2, [
|
||||
("opencli", lambda: self._opencli_user_notes(...)), # ← 没有降级!
|
||||
])
|
||||
|
||||
# 通用网页(web_scrape.py)—— 不用 OpenCLI
|
||||
await self._try_tools(tool_level, [
|
||||
("web_crawler", lambda: self._web_crawl(target)),
|
||||
("agent-browser", lambda: self._browser_extract(target)),
|
||||
])
|
||||
```
|
||||
|
||||
**对照表**:
|
||||
|
||||
| 平台 | 降级链 | 有浏览器降级? | 备注 |
|
||||
|---|---|---|---|
|
||||
| **抖音** | `opencli` → `agent-browser` | ✅ | 唯一做了完整降级的平台 |
|
||||
| **小红书** | `opencli`(单条) | ❌ | OpenCLI 挂了就抓不了 |
|
||||
| **B站** | `opencli`(单条) | ❌ | 同上 |
|
||||
| **知乎** | `opencli`(单条) | ❌ | 同上 |
|
||||
| **通用网页** | `web_crawler` → `agent-browser` | ✅ | 走另一套(不用 OpenCLI) |
|
||||
|
||||
**⚠️ 这是模块的真实局限**:三个平台**只有一条路径**,没有降级能力。
|
||||
接手时如果要提升健壮性,**最值得做的就是给它们补上浏览器降级**
|
||||
(照抖音的 `_browser_*` 方法写即可)。
|
||||
|
||||
**另一个细节**:`collect_user` 传的是**硬编码的 `2`**(不是 `tool_level` 参数),
|
||||
而 `collect_item` 才用传入的 `tool_level`:
|
||||
```python
|
||||
async def collect_user(self, user_id, limit=20): # 没有 tool_level 参数
|
||||
await self._try_tools(2, [...]) # ← 写死 2
|
||||
async def collect_item(self, target, depth, tool_level=2):
|
||||
await self._try_tools(tool_level, [...]) # ← 用参数
|
||||
```
|
||||
→ **用户在前端设 `tool_level=1` 时,"抓用户主页"这条路仍会走到浏览器**。
|
||||
|
||||
### 4.35 登录态机制(★ 这是"如何用真实浏览器"的另一半)
|
||||
|
||||
**问题**:抖音的数据接口需要登录态(cookie),怎么拿到?
|
||||
|
||||
**答案**:`mediacrawler_adapter.py` 的做法 —— **用户在真实 Chrome 登录,程序通过 CDP 读解密后的 cookie**。
|
||||
|
||||
#### 为什么不能直接读 cookie 文件
|
||||
|
||||
代码注释原文(第 29 行):
|
||||
```
|
||||
# Chrome 新版把 cookie 值加密存在 SQLite 里,必须通过 CDP 读解密后的值。
|
||||
```
|
||||
|
||||
Chrome v80+ 把 cookie **加密**存在 SQLite(`Cookies` 文件),
|
||||
直接读文件拿到的是**密文**。必须让 **Chrome 自己解密** → 通过 **CDP 协议**问它。
|
||||
|
||||
#### 三步实现
|
||||
|
||||
**① 通过 CDP 读 cookie**(`_get_cookies`,第 68 行)
|
||||
```python
|
||||
async def _get_cookies() -> dict:
|
||||
"""从 CDP 连接读取 Chrome cookie(解密后的值)"""
|
||||
ctx = await _ensure_cdp() # 确保 CDP 连接
|
||||
all_cookies = await ctx.cookies() # ← 让 Chrome 解密并返回
|
||||
...
|
||||
```
|
||||
|
||||
**② 转成 HTTP Header**(`_cookie_str`,第 86 行)
|
||||
```python
|
||||
def _cookie_str(cookies: dict) -> str:
|
||||
return "; ".join(f"{k}={v}" for k, v in cookies.items())
|
||||
```
|
||||
|
||||
**③ 带 cookie 直调抖音 API**(`_http_get`,第 93 行)
|
||||
```python
|
||||
def _http_get(url: str, cookies: dict, timeout: int = 15) -> dict:
|
||||
headers = {
|
||||
"User-Agent": "...Chrome/150.0.0.0 Safari/537.36",
|
||||
"Cookie": _cookie_str(cookies), # ★
|
||||
"Referer": "https://www.douyin.com/",
|
||||
"Origin": "https://www.douyin.com",
|
||||
}
|
||||
...
|
||||
```
|
||||
|
||||
#### 登录态判断
|
||||
|
||||
```python
|
||||
cookies = await _get_cookies()
|
||||
has_session = bool(cookies.get("sessionid")) # ← 有 sessionid 就算已登录
|
||||
```
|
||||
对应 HTTP 端点 `GET /api/scrape/check-login`。
|
||||
|
||||
#### 让用户登录(巧妙的做法)
|
||||
|
||||
代码注释(第 433 行):
|
||||
```python
|
||||
# ── 打开登录页(用 AppleScript 控制 Chrome,不需要 CDP) ──
|
||||
"""在 Chrome 中打开抖音首页,让用户登录
|
||||
登录后 cookie 自动保存到 Chrome profile,两个 Chrome 都会检测。"""
|
||||
```
|
||||
|
||||
**→ 用 AppleScript 打开真实 Chrome**(不是 Playwright 控制的),
|
||||
用户**在真实浏览器里手动登录** → cookie 存进 Chrome profile →
|
||||
之后程序通过 CDP 读。
|
||||
|
||||
**这样最自然**:用户看到的是他熟悉的 Chrome,扫码登录,
|
||||
✅ 不会被平台识别为"自动化登录"。
|
||||
|
||||
**⛔ 移植注意**:`open_login_page()` 用的是 **AppleScript**(`osascript`),
|
||||
**macOS 专有**。Linux/Windows 要改实现(可以用 `open` 命令或直接 Playwright 打开)。
|
||||
|
||||
#### 一个隐藏技巧(避免暴露自动化)
|
||||
|
||||
代码注释(第 227-229 行):
|
||||
```python
|
||||
# 获取热评(复用已有 Chrome 页面,不创建新标签页)
|
||||
# Chrome 已有 douyin.com 页面,直接用它的 JS 上下文执行 fetch
|
||||
# ⚠️ 不要 new_page() — 那会在 Chrome 中闪出新标签页
|
||||
```
|
||||
|
||||
**→ 复用用户已经打开的页面**执行 fetch,而不是新开标签页
|
||||
(新标签页会闪一下,且更易被识别)。
|
||||
|
||||
---
|
||||
|
||||
### 4.4 抖音的特有实现(两份代码,别搞混)
|
||||
|
||||
**⚠️ 项目里有**两个**抖音相关模块**,职责不同:
|
||||
|
||||
**① `services/adapters/douyin_scrape.py`(181 行)**
|
||||
- 走**工具降级**(OpenCLI → 浏览器)
|
||||
- 浏览器路径用 JS 从页面 **DOM 文本**提取数据:
|
||||
```javascript
|
||||
// _browser_video_page 的 JS(正则从页面文本抓「获赞/粉丝/关注」)
|
||||
const body = document.body.innerText || '';
|
||||
const uidM = body.match(/抖音号[::]\s*(\S+)/);
|
||||
const nickM = body.match(/@(\S+)/);
|
||||
function extractNum(label) {
|
||||
var m = body.match(new RegExp('(\\d+(?:\\.\\d+)?[万w]?)\\s*' + label));
|
||||
...
|
||||
}
|
||||
return { aweme_id, title, author_nickname, douyin_id, digg_count, fans, following };
|
||||
```
|
||||
|
||||
**② `services/mediacrawler_adapter.py`(457 行)**
|
||||
- **不走工具降级**,完全独立的实现
|
||||
- 文件头注释(原文):
|
||||
> 全新架构:**不再依赖 CDP/Playwright/浏览器页面**。
|
||||
> 直接从 Chrome profile 读取 cookie,通过 HTTP 请求调用抖音 API。
|
||||
> 全程无窗口、无标签页、无闪烁。
|
||||
- 用于**追踪视频/作者**这类需要高频刷新的场景(详见子代理报告)
|
||||
|
||||
**⛔ 关键澄清**:`mediacrawler_adapter.py` **虽然叫 mediacrawler,但不依赖
|
||||
MediaCrawler 这个开源项目**——它只是借用名字,实际是"读 Chrome cookie + HTTP 调 API"。
|
||||
|
||||
---
|
||||
|
||||
## 五、引擎层(`services/scrape_engine.py`,361 行)
|
||||
|
||||
### 5.1 URL 解析(`resolve_target`,第 56 行)
|
||||
|
||||
**7 类目标自动识别**(实测代码):
|
||||
|
||||
```python
|
||||
抖音短链 v.douyin.com/xxx → type=shortlink
|
||||
抖音视频 douyin.com/video/{id} → type=video
|
||||
抖音用户 douyin.com/user/{sec_uid} → type=user
|
||||
小红书 xiaohongshu.com/explore/{id} → type=note
|
||||
B站视频 bilibili.com/video/{BV} → type=video
|
||||
B站用户 bilibili.com/space/{mid} → type=user
|
||||
知乎 zhihu.com/answer/{id} | /question/ → type=item
|
||||
通用网页 http(s)://... → type=page
|
||||
纯 sec_uid MS4w 开头 或 len>20 → douyin/user
|
||||
纯数字 len>=15 → douyin/video(aweme_id)
|
||||
纯数字 其他 → zhihu/item
|
||||
```
|
||||
|
||||
**短链解析**(`_resolve_shortlink`,第 136 行):用 `curl -sI` 拿 `Location` 头,
|
||||
再用正则从跳转 URL 里抠出 `aweme_id`。
|
||||
|
||||
### 5.2 执行主流程(`run`,第 163 行)
|
||||
|
||||
```
|
||||
run(request)
|
||||
├─ 1. resolve_urls(targets) → 标准化目标列表
|
||||
├─ 2. 短链逐个解析
|
||||
├─ 3. 判断同步 / 异步
|
||||
│ async_mode = request.async_mode 或 len(targets) > 50
|
||||
│
|
||||
├─ 【异步分支】
|
||||
│ · run_id = uuid[:8]
|
||||
│ · 内存状态 {status, total, completed, results, errors}
|
||||
│ · asyncio.create_task(_run_async(...))
|
||||
│ · 立即返回 {status:"async", run_id} ← 前端轮询
|
||||
│
|
||||
└─ 【同步分支】
|
||||
· db.create_task("single", ...)
|
||||
· for target: _scrape_one() → _save_item()
|
||||
· db.update_task_status("completed", summary)
|
||||
· 返回 {status, task_id, duration, total, success, errors, data}
|
||||
```
|
||||
|
||||
### 5.3 单目标抓取(`_scrape_one`,第 261 行)
|
||||
|
||||
```python
|
||||
adapter = self._get_adapter(platform) # 按平台取适配器(带缓存)
|
||||
if target["type"] == "user":
|
||||
return await adapter.collect_user(target["target_id"]) # 返回 list
|
||||
elif target["type"] in ("video", "note"):
|
||||
return await adapter.collect_item(target["target_id"], depth, tool_level)
|
||||
else:
|
||||
return None
|
||||
```
|
||||
|
||||
**返回值语义**(重要):
|
||||
- `dict` → 单条内容
|
||||
- `list` → 多条(用户主页的所有作品)
|
||||
- `None` → 失败
|
||||
|
||||
### 5.4 落库(`_save_item`,第 292 行)
|
||||
|
||||
```python
|
||||
db_id = self.db.insert_item(
|
||||
task_id=..., platform=..., item_id=..., url=..., title=...,
|
||||
author_name=..., author_id=..., published_at=..., text_content=...,
|
||||
tags=..., stats=..., extra=..., media=...)
|
||||
|
||||
comments = item.get("comments", [])
|
||||
if comments and db_id:
|
||||
self.db.insert_comments(db_id, comments) # ★ 评论独立表
|
||||
```
|
||||
|
||||
### 5.5 异步模式(`_run_async`,第 315 行)
|
||||
|
||||
- ✅ **同时写内存 + 落库**(内存态供轮询,落库供持久化)
|
||||
- 内存态在 `self._tasks[run_id]`(**进程重启即丢**)
|
||||
- 查询用 `get_async_result(run_id)`
|
||||
|
||||
### 5.6 ⚠️ 已知的局限(接手时要清楚)
|
||||
|
||||
| 局限 | 说明 |
|
||||
|---|---|
|
||||
| **没有限流/并发控制** | `for target in ready:` 是**纯串行**,目标多时会慢;也没有请求间隔(可能触发平台风控) |
|
||||
| **异步态存内存** | 重启 Dashboard 后 `_tasks` 丢失(但库里有记录,前端看不到进度) |
|
||||
| **无重试** | 单目标失败只记 `errors`,不重试 |
|
||||
| **adapter 实例缓存** | `self._adapters` 进程内缓存(无清理) |
|
||||
|
||||
---
|
||||
|
||||
### 5.7 一个完整请求的生命周期(跟着走一遍最快懂)
|
||||
|
||||
以"采集某抖音作者的全部视频"为例:
|
||||
|
||||
```
|
||||
① 前端
|
||||
POST /api/scrape/run
|
||||
{"targets": ["r606391422378804368"], "tool_level": 2}
|
||||
│
|
||||
▼
|
||||
② routes/scrape.py:44 api_scrape_run()
|
||||
组装 request → engine.run(request)
|
||||
│
|
||||
▼
|
||||
③ scrape_engine.py:163 run()
|
||||
├─ resolve_urls(["r6063..."])
|
||||
│ → resolve_target() 识别:以 MS4w 开头 → 抖音 sec_uid
|
||||
│ → [{"platform":"douyin","type":"user","target_id":"r6063...","status":"resolved"}]
|
||||
│
|
||||
├─ 同步或异步?len(targets)=1,不大于 50 → 同步
|
||||
│
|
||||
├─ db.create_task("single","douyin",...) → task_id = 1
|
||||
│
|
||||
├─ for target: _scrape_one(target,"douyin","light",2)
|
||||
│ │
|
||||
│ ▼
|
||||
│ scrape_engine.py:261
|
||||
│ _get_adapter("douyin") → DouyinScrapeAdapter() (进程内缓存)
|
||||
│ type=="user" → adapter.collect_user("r6063...")
|
||||
│ │
|
||||
│ ▼
|
||||
│ douyin_scrape.py:21 collect_user()
|
||||
│ _try_tools(2, [("opencli", ...), ("agent-browser", ...)])
|
||||
│ │
|
||||
│ ├─ 尝试 1:_opencli_user_videos()
|
||||
│ │ _run_opencli(["douyin","user-videos","r6063...",
|
||||
│ │ "--limit","20","--with_comments","true","-f","json"])
|
||||
│ │ → subprocess 执行 opencli(Node CLI)
|
||||
│ │ → _parse_output() ← JSON → YAML → 纯文本 三层兜底
|
||||
│ │ → 成功返回 list[dict] → _try_tools 立刻返回,不再降级
|
||||
│ │
|
||||
│ └─ 尝试 1 失败(OpenCLI 没装/超时/报错)
|
||||
│ → 尝试 2:_browser_user_profile()
|
||||
│ → Playwright 启真实 Chrome → 打开页面 → JS 提取
|
||||
│
|
||||
│ → 每条数据 _to_schema() 转成统一格式
|
||||
│ → 返回 list
|
||||
│
|
||||
├─ for item: _save_item(task_id=1, item)
|
||||
│ db.insert_item(...) → db_id((platform,item_id) 唯一,重复则忽略)
|
||||
│ db.insert_comments(db_id, comments) ← 评论另存
|
||||
│
|
||||
├─ db.update_task_status(1, "completed", summary={success:N, errors:0})
|
||||
│
|
||||
└─ return {status:"completed", task_id:1, duration, total, success, errors, data:[...]}
|
||||
│
|
||||
▼
|
||||
④ 前端拿到 data,渲染列表
|
||||
```
|
||||
|
||||
**异步分支的差异**(目标 > 50 个,或显式 `async_mode=true`):
|
||||
```
|
||||
run() 立即返回 {status:"async", run_id:"ab12cd34"}
|
||||
↓(后台)
|
||||
asyncio.create_task(_run_async(run_id, targets, ...))
|
||||
↓
|
||||
建持久化任务 → 逐个 _scrape_one + _save_item
|
||||
↓
|
||||
进度写内存 self._tasks[run_id](供轮询)
|
||||
结果写 SQLite(供持久化)
|
||||
↓
|
||||
前端轮询 get_async_result(run_id) 看进度
|
||||
```
|
||||
|
||||
**⚠️ 注意**:异步进度**只在内存**,Dashboard 重启就丢
|
||||
(库里数据还在,但前端看不到进度了)。
|
||||
|
||||
---
|
||||
|
||||
## 六、存储层(`services/scrape_db.py`,415 行)
|
||||
|
||||
### 6.1 数据库位置
|
||||
|
||||
```python
|
||||
DEFAULT_DB = AGENT_LOCAL / "data" / "scrape.db"
|
||||
```
|
||||
(`AGENT_LOCAL` 是环境变量;默认 `~/workbuddy-agent-os/agent-local`)
|
||||
|
||||
### 6.2 四张表(实测 `CREATE TABLE`)
|
||||
|
||||
```sql
|
||||
-- ① 采集任务(一次 run 一条)
|
||||
scrape_tasks(
|
||||
id, type, -- single / batch / scheduled
|
||||
platform, target, -- 目标(批量时是 JSON 数组)
|
||||
depth, tool_level, machine,
|
||||
status, -- pending / running / completed / failed
|
||||
total_targets, completed_targets,
|
||||
summary, -- 摘要 JSON
|
||||
created_at
|
||||
)
|
||||
|
||||
-- ② 采集到的内容
|
||||
scrape_items(
|
||||
id, task_id → scrape_tasks,
|
||||
platform, item_id, -- 平台内唯一 ID
|
||||
url, title, author_name, author_id,
|
||||
published_at, collected_at, text_content, tags,
|
||||
... -- 还有 stats / extra / media 等
|
||||
)
|
||||
|
||||
-- ③ 评论
|
||||
scrape_comments(id, item_db_id → scrape_items,
|
||||
author_name, text, likes, replied_at)
|
||||
|
||||
-- ④ 采集源(长期跟踪)
|
||||
scrape_sources(
|
||||
id, platform, source_type, -- user / hashtag / keyword / url_list / api
|
||||
target, display_name,
|
||||
category, -- 自定义分类
|
||||
notes,
|
||||
schedule, -- CRON(定期采集)
|
||||
depth, tool_level, last_collected,
|
||||
status -- active / paused
|
||||
)
|
||||
```
|
||||
|
||||
### 6.3 方法清单(22 个,实测)
|
||||
|
||||
```
|
||||
任务:create_task / update_task_status / get_task / list_tasks
|
||||
内容:insert_item / get_item_id / item_exists / get_item / list_items
|
||||
评论:insert_comments / get_comments
|
||||
采集源:upsert_source / update_source / list_sources / get_due_sources /
|
||||
update_source_collected / delete_source
|
||||
统计:count_by_platform / count_today / sources_count / task_stats
|
||||
```
|
||||
|
||||
### 6.4 去重机制
|
||||
|
||||
`insert_item` **依赖 `(platform, item_id)` 唯一约束** —— 重复插入时
|
||||
用 `INSERT OR IGNORE` 模式(验证文档 L1-2 有测例)。
|
||||
|
||||
---
|
||||
|
||||
## 七、HTTP 层(`routes/scrape.py`,711 行 / 22 端点)
|
||||
|
||||
### 7.1 端点清单(实测)
|
||||
|
||||
```
|
||||
采集
|
||||
POST /api/scrape/run 发起采集(targets + depth + tool_level)
|
||||
POST /api/scrape/resolve 只解析 URL,不采集
|
||||
POST /api/scrape/douyin-stats 抖音数据查询
|
||||
|
||||
查询
|
||||
GET /api/scrape/title 取标题
|
||||
GET /api/scrape/result 结果
|
||||
GET /api/scrape/tasks 任务列表
|
||||
GET /api/scrape/items 内容列表
|
||||
GET /api/scrape/items/{id} 单项详情
|
||||
GET /api/scrape/stats 统计
|
||||
|
||||
采集源管理
|
||||
POST /api/scrape/sources 新建源
|
||||
GET /api/scrape/sources 源列表
|
||||
DEL /api/scrape/sources/{id} 删源
|
||||
|
||||
追踪(视频 / 作者)
|
||||
POST /api/scrape/track-video 追踪视频
|
||||
GET /api/scrape/tracked-videos 已追踪视频
|
||||
POST /api/scrape/delete-tracked/{id}
|
||||
POST /api/scrape/refresh-video/{id} 刷新单个视频
|
||||
POST /api/scrape/track-author 追踪作者
|
||||
GET /api/scrape/tracked-authors 已追踪作者
|
||||
POST /api/scrape/refresh-author/{id}
|
||||
GET /api/scrape/author-history/{id}
|
||||
|
||||
主题 / 登录
|
||||
POST /api/scrape/import-topics 批量导入主题
|
||||
GET /api/scrape/check-login 检测登录态
|
||||
GET /api/scrape/open-login 打开登录
|
||||
```
|
||||
|
||||
### 7.2 关键实现(实测)
|
||||
|
||||
**`POST /api/scrape/run`**(第 43 行)—— 前端发起采集的唯一入口:
|
||||
```python
|
||||
@router.post("/run")
|
||||
async def api_scrape_run(data: dict = {}):
|
||||
targets = data.get("targets", data.get("target", []))
|
||||
if isinstance(targets, str):
|
||||
targets = [targets] # 兼容单个字符串
|
||||
request = {
|
||||
"targets": targets,
|
||||
"platform": data.get("platform", "auto"),
|
||||
"depth": data.get("depth", "light"),
|
||||
"tool_level": data.get("tool_level", 2), # ← 默认 2(OpenCLI + 浏览器)
|
||||
"machine": data.get("machine", ""),
|
||||
"multi_machine": data.get("multi_machine", False),
|
||||
"async_mode": data.get("async_mode", False),
|
||||
}
|
||||
engine = _get_engine() # 模块级单例
|
||||
result = await engine.run(request)
|
||||
return {"status": "ok", **result}
|
||||
```
|
||||
|
||||
**登录态两端点**(第 692 / 703 行)—— 都委托给 `mediacrawler_adapter`:
|
||||
```python
|
||||
GET /api/scrape/check-login → mediacrawler_adapter.check_login_status()
|
||||
POST /api/scrape/open-login → mediacrawler_adapter.open_login_page()
|
||||
```
|
||||
|
||||
### 7.3 引擎实例
|
||||
|
||||
路由层用**模块级单例**拿 engine(`_get_engine()`),
|
||||
所以 `engine._tasks`(异步态)和 `engine._adapters`(适配器缓存)
|
||||
在整个 Dashboard 进程内共享。
|
||||
|
||||
---
|
||||
|
||||
## 七·五、Dashboard 插件(`plugins/crawl.py`,91 行)
|
||||
|
||||
抓取模块**作为 Dashboard 插件**注册(提供概览统计,不是核心逻辑):
|
||||
|
||||
```python
|
||||
class CrawlDashboardPlugin(DashboardPlugin):
|
||||
name = "crawl"
|
||||
label = "内容抓取"
|
||||
icon = "📡"
|
||||
order = 35
|
||||
```
|
||||
|
||||
**它做三件事**:
|
||||
1. `summary()` — 概览:总抓取数 / 今日新增 / 抓取源(从 `ScrapeDB` 读)
|
||||
2. `detail(machine)` — 指定机器的详情
|
||||
3. `actions()` — 快捷操作(跳转 `scrape` 视图)
|
||||
|
||||
**注意**:它有个**兜底设计** —— 如果 `ScrapeDB` 不可用(数据库损坏/权限),
|
||||
会退化成**统计知识库里的 md 文件数**,而不是报错。
|
||||
|
||||
**⛔ 移植注意**:如果目标项目没有这套插件框架,`plugins/crawl.py`
|
||||
可以直接丢弃(它只是 Dashboard 的展示层,不影响抓取功能本身)。
|
||||
|
||||
---
|
||||
|
||||
## 八、前端(`frontend/src/views/scrape.js`,131 KB)
|
||||
|
||||
⚠️ **这是模块里最大的单文件**(131 KB)。功能覆盖:22 个端点的界面。
|
||||
|
||||
**建议接手团队**:不要照搬这个前端,按第七节的 HTTP 契约重写。
|
||||
理由:131 KB 单文件难维护,且和本项目的视图框架耦合。
|
||||
|
||||
---
|
||||
|
||||
## 九、验证方案(项目里已有现成的)
|
||||
|
||||
`services/scrape_validation.md`(314 行)已经写了 **7 个 Level 的验证清单**:
|
||||
|
||||
```
|
||||
L0 基础设施(3 项) Python import / SQLite 建库 / FastAPI 路由注册
|
||||
L1 数据库层(3 项) 建任务 / 写入+去重 / 评论入库
|
||||
L2 适配器 Mock(3 项) 工具降级逻辑 / 一级失败二级成功
|
||||
L3 适配器真实(3 项) 抖音用户采集 / 小红书 / 详情+评论 ← 需 OpenCLI + 登录态
|
||||
L4 引擎层(3 项) 解析 URL / 执行采集 / 异步轮询
|
||||
L5 API 层(4 项) curl 打 4 个端点
|
||||
L6 前端(4 项) 浏览器里操作
|
||||
L7 异常(1+ 项) OpenCLI 不可用时应抛清晰错误
|
||||
```
|
||||
|
||||
**⚠️ 但要注意**:该文档写于 2026-07-16,**里面的方法名已过时**:
|
||||
```
|
||||
文档写 resolve_targets ← 不存在
|
||||
代码里是 resolve_urls ← 实际
|
||||
文档写 get_result ← 不存在
|
||||
代码里是 get_async_result ← 实际
|
||||
```
|
||||
**以代码为准**。
|
||||
|
||||
---
|
||||
|
||||
## 十、移植清单(换环境要改什么)
|
||||
|
||||
| # | 依赖 | 位置 | 处理 |
|
||||
|---|---|---|---|
|
||||
| 1 | **OpenCLI**(Node CLI) | `adapters/__init__.py` 的 `OPENCLI` 常量 | 目标机 `npm i -g @jackwener/opencli`,或设 `OPENCLI_PATH` |
|
||||
| 2 | **真实 Chrome** | `browser_helpers.py` 的 `_CHROME_PATHS` | Linux/Windows 要改路径(如 `/usr/bin/google-chrome`) |
|
||||
| 3 | **Playwright** | pip | `pip install playwright && playwright install chromium` |
|
||||
| 4 | **PyYAML** | pip(可选但强烈建议) | 不装会降级到纯文本解析(丢嵌套结构) |
|
||||
| 5 | **AGENT_LOCAL** 环境变量 | `scrape_db.py:18` | 定义了才能定位 `scrape.db` |
|
||||
| 6 | **平台登录态** | Chrome profile | 目标机需手动登录一次目标平台 |
|
||||
| 7 | `SCRAPE_CHROME_USER_DATA` | 环境变量(可选) | 要持久化登录态时指定 |
|
||||
|
||||
### 最小可运行子集(只要"能采集")
|
||||
|
||||
```
|
||||
services/scrape_db.py 数据库(4 表)
|
||||
services/adapters/__init__.py 基类 + 降级 + OpenCLI 调用 + 输出解析
|
||||
services/adapters/browser_helpers.py 浏览器降级
|
||||
services/adapters/<目标平台>_scrape.py 目标平台适配器
|
||||
```
|
||||
—— 这 4 个文件就能跑通单平台采集,不需要 engine/routes/前端。
|
||||
|
||||
---
|
||||
|
||||
## 十一、接手建议(按顺序)
|
||||
|
||||
```
|
||||
第 1 步 装 OpenCLI + Playwright + 真实 Chrome,跑 validation.md 的 L0/L1
|
||||
(这两级零外部依赖,能验证环境对不对)
|
||||
|
||||
第 2 步 跑 L2(Mock 测试)—— 验证降级逻辑,不需要真实平台
|
||||
这时你已经能理解 _try_tools 的语义
|
||||
|
||||
第 3 步 登录目标平台,跑 L3(真实采集)—— 第一次真正拿数据
|
||||
如果 OpenCLI 不通,会看到它降级到浏览器,日志里有 ⚠️
|
||||
|
||||
第 4 步 跑 L4/L5(引擎 + API)
|
||||
|
||||
第 5 步 替换前端(不要照搬 131 KB)
|
||||
|
||||
第 6 步 加你要的东西:限流 / 重试 / 并发控制(现在都没有)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 十二、这个模块还没做的事(接手可以补)
|
||||
|
||||
```
|
||||
① 限流与请求间隔 —— 现在纯串行、无间隔,目标多时可能触发平台风控
|
||||
② 失败重试 —— 现在失败只记 errors
|
||||
③ 并发控制 —— 没有信号量,大量目标只能串行
|
||||
④ 异步态持久化 —— run 进度存内存,重启即丢
|
||||
⑤ 登录态自动检测 —— check-login 端点有,但采集前没强制校验
|
||||
⑥ 代理支持 —— 没有看到代理配置(多账号场景会需要)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 附录:本说明的取证方式(可复现)
|
||||
|
||||
```bash
|
||||
cd 05_tools/10_dashboard
|
||||
# 架构
|
||||
head -20 services/adapters/__init__.py
|
||||
# 浏览器("套真实浏览器"的做法)
|
||||
sed -n '1,70p' services/adapters/browser_helpers.py
|
||||
# 降级机制
|
||||
grep -n "_try_tools" -A 14 services/adapters/__init__.py
|
||||
# 引擎流程
|
||||
grep -nE "^ (async )?def " services/scrape_engine.py
|
||||
# 表结构
|
||||
grep -n "CREATE TABLE" -A 12 services/scrape_db.py
|
||||
# 端点
|
||||
grep -nE "^@router\." routes/scrape.py
|
||||
```
|
||||
+76
-7
@@ -18,6 +18,7 @@ api/auth.py WebUI 登录鉴权
|
||||
api/monitor/* 监控层整体(含 platforms.py 能力矩阵)
|
||||
api/monitor/db.py MySQL 连接层(可回退 SQLite 供测试用)
|
||||
api/monitor/migrate_from_sqlite.py SQLite → MySQL 一次性迁移脚本
|
||||
api/monitor/upstream.py 上游更新检查(定时 fetch 上游并比对)
|
||||
api/routers/{auth,monitor,settings}.py
|
||||
api/schemas/{auth,monitor,settings}.py
|
||||
api/services/interpreter.py 解释器探测(uv / .venv / 当前解释器)
|
||||
@@ -25,9 +26,18 @@ webui/src/components/{monitor,settings,auth}/ 新视图
|
||||
webui/src/components/layout/{PlatformSwitcher,UnwiredPlatformNotice}.tsx
|
||||
webui/src/{hooks/useMonitor.ts,hooks/usePlatform.ts,store/platformStore.ts,lib/monitorFormat.ts,types/monitor.ts}
|
||||
docs/监控功能使用说明.md
|
||||
tests/test_{auth,settings,platforms,monitor_*}.py
|
||||
tests/test_{auth,settings,platforms,qrlogin,monitor_*,upstream}.py
|
||||
Dockerfile / .dockerignore / docker-compose.yml 服务器部署用
|
||||
```
|
||||
|
||||
> `api/monitor/upstream.py` 要调 `git`,而 `python:3.11-slim` 不带它 —— Dockerfile 里为此
|
||||
> **显式装了 git**。改了 Dockerfile 就必须重建镜像(`docker compose build`),`./deploy.sh`
|
||||
> 只重建前端,不重建镜像。
|
||||
|
||||
其中 `api/monitor/qrlogin.py` + `webui/.../QrLoginPanel.tsx` 是**服务器专用**的扫码登录:
|
||||
那台机器上 Chrome 跑在 Xvfb 里,`show_qrcode` 调的 PIL `Image.show()` 需要桌面看图程序,
|
||||
服务器没有,二维码会无处可去。所以改成用 CDP 把二维码从页面里读出来交给前端 `<img>` 显示。
|
||||
|
||||
### 2. 加法改动(低冲突)
|
||||
|
||||
只在既有文件里**新增**内容,不改动原有行:
|
||||
@@ -36,7 +46,7 @@ tests/test_{auth,settings,platforms,monitor_*}.py
|
||||
|---|---|
|
||||
| `cmd_arg/arg.py` | typer 选项:`--enable_cdp_mode`、`--inject_all_cookies`、`--save_login_state`、`--cookies_file`、`--crawler_max_sleep_sec`,以及对应的 `config.*` 回写 |
|
||||
| `api/schemas/crawler.py` | `CrawlerStartRequest` 的若干**可选**字段(默认 `None`,不传则不加对应 CLI 参数) |
|
||||
| `config/base_config.py` | `INJECT_ALL_COOKIES = False` |
|
||||
| `config/base_config.py` | `INJECT_ALL_COOKIES = False`;`MASK_NICKNAME = False`(关掉昵称脱敏,见第 3 节) |
|
||||
| `api/routers/__init__.py` | 导出新增的 router |
|
||||
| `requirements.txt` | 补上 `websockets`(上游 `pyproject.toml` 里有、`requirements.txt` 里漏了) |
|
||||
| `tests/conftest.py` | 新增 `_bypass_auth_for_non_auth_suites` fixture |
|
||||
@@ -49,19 +59,22 @@ tests/test_{auth,settings,platforms,monitor_*}.py
|
||||
| `api/routers/websocket.py` | 两个 WS 路由加 `dependencies=[Depends(require_ws_auth)]` | 上游若新增 WS 路由,**必须同样加上**,否则那条流是裸奔的 |
|
||||
| `api/services/crawler_manager.py` | 解释器探测替换硬编码 `uv run`;`_build_command` 转发新增参数;新增 `is_busy()` / `run_and_wait()` 与完成事件 | 留意 `_build_command` 的参数拼装 |
|
||||
| `media_platform/xhs/login.py` | `login_by_cookies` 在 `INJECT_ALL_COOKIES` 打开时注入**全部** cookie(默认关闭,行为不变) | 小改动,好合并 |
|
||||
| `tools/user_hash.py` | `mask_nickname` 改为读 `config.MASK_NICKNAME`,本仓库默认**不脱敏**(原样返回)。上游作为教学版默认脱敏,但那是有损的 ——「张三」「张四」都成「张*」,而分清谁是谁正是监控这一层要干的活。脱敏实现本身没删,改回 `True` 即恢复上游行为 | 与 `config/base_config.py` 一起改,两处不同步会不一致 |
|
||||
|
||||
### 4. 上游 bug 修复(建议回馈上游)
|
||||
|
||||
| 文件 | 修的问题 |
|
||||
|---|---|
|
||||
| `media_platform/xhs/core.py` | 见下节 |
|
||||
| `media_platform/xhs/login.py` | 同上(cookie 加固) |
|
||||
| `media_platform/xhs/core.py` | 见下节 1 |
|
||||
| `media_platform/xhs/login.py` | 见下节 2(cookie 加固) |
|
||||
| `media_platform/douyin/core.py` | 见下节 3(首页 `goto` 永远超时,采集根本起不来) |
|
||||
| `media_platform/douyin/login.py` | 见下节 4(注入 cookie 后页面陈旧,白等十分钟) |
|
||||
|
||||
---
|
||||
|
||||
## 二、应该给上游提 PR 的两个修复
|
||||
## 二、应该给上游提 PR 的四个修复
|
||||
|
||||
这两处是**上游自身的缺陷**,提上去以后就不用自己背着:
|
||||
这四处都是**上游自身的缺陷**,提上去以后就不用自己背着:
|
||||
|
||||
### 1. 博主主页抓取失败会跳掉整个博主(`xhs/core.py`)
|
||||
|
||||
@@ -80,19 +93,75 @@ tests/test_{auth,settings,platforms,monitor_*}.py
|
||||
`a1` / `webId` 等签名所需 cookie 只能靠持久化 profile 补,冷启动时签名会失败。
|
||||
默认行为保持不变,用 `INJECT_ALL_COOKIES` 开关控制。
|
||||
|
||||
### 3. 抖音首页的 `goto` 永远等不到 `load`(`douyin/core.py:101`)
|
||||
|
||||
```python
|
||||
await self.context_page.goto(self.index_url) # 默认 wait_until="load"
|
||||
```
|
||||
|
||||
抖音首页的 `load` 事件**不会触发**(有长连接/埋点类请求一直挂着)。实测:同一台
|
||||
Chrome、同一个地址,`domcontentloaded` 0.7 秒返回,而 `load` 等满 90 秒仍然超时。
|
||||
后果是整个采集**一步都没走就崩**,退出码 1 —— 看起来像"抖音不能用"。
|
||||
|
||||
修复:显式 `wait_until="domcontentloaded"`。上游的贴吧(`tieba/core.py`)和知乎
|
||||
(`zhihu/core.py`)本来就是这么写的,抖音这个页面只是恰好属于"永远不 load"的那类。
|
||||
|
||||
> 这个缺陷在本机可能复现不出来(换个网络/有缓存时 `load` 也许能触发),所以社区里
|
||||
> 没人报。它和网络快慢无关:不是"慢",是那个事件根本不会发生。
|
||||
|
||||
### 4. 注入 cookie 后页面是陈旧的(`douyin/login.py:266`)
|
||||
|
||||
`login_by_cookies()` 把 cookie 塞进 browser context,但**页面是在这之前加载的** ——
|
||||
SPA 只在加载时读一次登录态,`localStorage.HasUserLogin` 于是还停在"未登录",
|
||||
紧接着的 `check_login_state()` 会对着这个陈旧的值轮询到超时(600 次 × 1 秒 = 十分钟),
|
||||
然后 `sys.exit()`。**下一轮**才正常,因为那时 cookie 已经在 profile 里了。
|
||||
|
||||
表现是"第一次跑白等十分钟、第二次才行",很容易被当成偶发。
|
||||
|
||||
修复:注入完 cookie 后 `reload(wait_until="domcontentloaded")`,让站点立刻重新判定会话。
|
||||
|
||||
> 与第 1 条同源:都是"页面状态是加载那一刻的快照"。本仓库的扫码登录(`api/monitor/qrlogin.py`)
|
||||
> 和运营模块也各自踩过这个坑,那里的判据改成了拿 cookie 问后台接口,而不是读页面快照。
|
||||
|
||||
---
|
||||
|
||||
## 三、上游更新时怎么操作
|
||||
|
||||
### 先让机器替你盯着
|
||||
|
||||
「上游更新检查」(`api/monitor/upstream.py`,开关在 WebUI 的**系统设置 → 上游更新**)会按
|
||||
间隔 `git fetch` 上游、算出落后几个提交,有更新就推企业微信。它是这份文档的自动化版:
|
||||
没有它,「上游动了」这件事只取决于谁偶尔想起来去 fetch 一次。
|
||||
|
||||
两个细节决定了它为什么是安全的:它只 fetch 到 `FETCH_HEAD`,**不写工作区、不建 remote、不碰
|
||||
`refs/remotes`**,所以和正在跑的采集、和下面的 `git pull` 都不冲突;默认**关闭**,因为要联网,
|
||||
且需要镜像里有 git。
|
||||
|
||||
### 日常流程
|
||||
|
||||
```bash
|
||||
git stash # 或先 commit 到自己的分支(推荐)
|
||||
git fetch origin main
|
||||
git rebase origin/main # 冲突只会出现在上表第 3、4 类文件里
|
||||
./.venv/Scripts/python.exe -m pytest tests/ -q # 486 个测试就是回归网
|
||||
./.venv/Scripts/python.exe -m pytest tests/ -q # 502 个测试就是回归网
|
||||
```
|
||||
|
||||
### 直连 GitHub 不通时(本机常见)
|
||||
|
||||
本机到 `github.com` 时通时不通,**大包传输必断**(`Recv failure: Connection was reset`
|
||||
或 `unexpected disconnect while reading sideband packet`),所以 `git clone` / `--unshallow`
|
||||
这类一次性拉全量的操作基本必失败。可用的替代源:
|
||||
|
||||
```bash
|
||||
# gitcode 的 GitHub 镜像,国内直连,比 GitHub 本身还新一天以内
|
||||
git remote add gitcode https://gitcode.com/gh_mirrors/me/MediaCrawler.git
|
||||
git fetch --no-tags --unshallow gitcode # 本仓库当初就是这样补全历史的,约 2 秒
|
||||
```
|
||||
|
||||
注意 `git fetch` 只写 `refs/remotes/*`,**不会动本地 `main`**;
|
||||
但拉镜像会把 `upstream/main` 指到镜像的 tip(可能比 GitHub 晚一天),
|
||||
等 GitHub 通了再 `git fetch upstream` 正回来即可。
|
||||
|
||||
### 强烈建议:先把改动提交掉
|
||||
|
||||
当前状态是**未提交**的(25 个上游文件被改 + 31 个新文件)。在 `main` 分支上裸着工作区,
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/__init__.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""运营模块:管理自己的小红书账号,读取创作者后台的数据。
|
||||
|
||||
与 `api.monitor` 是**并列关系**,不是它的扩展。两者数据形状不同:监控是「每轮
|
||||
快照 + 差分」的公开互动数据,这里是创作者后台按日期给出的曝光/观看/完播等运营
|
||||
指标。硬塞进同一个模型会同时污染两边。
|
||||
|
||||
路线是**纯请求**(无浏览器),依据见 tools/probe_creator_api.py 的 Phase 0 实测:
|
||||
签名可自造、主站 cookie 即可认证、接口与参数已与真实页面对齐。
|
||||
"""
|
||||
@@ -0,0 +1,345 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/client.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""小红书创作者后台的纯请求客户端。
|
||||
|
||||
无浏览器:签名在本地算(见 signing.py),请求走 httpx。
|
||||
|
||||
**关于字段名的谨慎**:Phase 0 抓到的响应里,列表接口因为账号权限未生效而返回空壳
|
||||
(`data.result` 只有 `{success, code, message}` 没有数据),所以**真实字段名尚未亲眼
|
||||
见过**。因此每个指标都写成**多别名匹配**,并且解析不出来时存 `None` 而不是 0 ——
|
||||
0 是真实值,None 是"不知道",两者混淆会让报表说谎。
|
||||
"""
|
||||
|
||||
import re
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
import httpx
|
||||
|
||||
from .signing import sign_xyw, signed_api
|
||||
|
||||
CREATOR_ORIGIN = "https://creator.xiaohongshu.com"
|
||||
DATA_ANALYSIS_PAGE = f"{CREATOR_ORIGIN}/statistics/data-analysis"
|
||||
|
||||
USER_INFO_PATH = "/api/galaxy/user/info"
|
||||
PERMISSION_PATH = "/api/galaxy/creator/datacenter/permission/query"
|
||||
NOTE_LIST_PATH = "/api/galaxy/creator/datacenter/note/analyze/list"
|
||||
|
||||
USER_AGENT = (
|
||||
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
|
||||
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
|
||||
)
|
||||
|
||||
# 签名被网关拒绝时返回的响应体。区分它和普通业务错误很重要:406 说明签名写错了,
|
||||
# 而应用层的 code/-100 说明签名没问题、只是没有登录态。
|
||||
SIGNATURE_REJECTED_MARKERS = ("code", -1)
|
||||
|
||||
|
||||
class CreatorApiError(RuntimeError):
|
||||
"""调用创作者后台失败。``status`` 用于区分是网关拒绝还是业务错误。"""
|
||||
|
||||
def __init__(self, message: str, status: int = 0, payload: Any = None):
|
||||
super().__init__(message)
|
||||
self.status = status
|
||||
self.payload = payload
|
||||
|
||||
|
||||
def trans_cookies(cookie_str: str) -> Dict[str, str]:
|
||||
"""把 cookie 字符串解析成字典。容忍末尾分号、换行和零散空格。"""
|
||||
jar: Dict[str, str] = {}
|
||||
for chunk in re.split(r"[;\n]", cookie_str or ""):
|
||||
chunk = chunk.strip()
|
||||
if not chunk or "=" not in chunk:
|
||||
continue
|
||||
name, value = chunk.split("=", 1)
|
||||
name = name.strip()
|
||||
if name:
|
||||
jar[name] = value.strip()
|
||||
return jar
|
||||
|
||||
|
||||
def _pick(item: Dict[str, Any], *names: str) -> Any:
|
||||
"""按别名顺序取第一个存在的键。
|
||||
|
||||
字段名来自二手资料,未亲眼验证,所以不赌单一命名。
|
||||
"""
|
||||
for name in names:
|
||||
if name in item and item[name] is not None:
|
||||
return item[name]
|
||||
return None
|
||||
|
||||
|
||||
_COUNT_UNITS = {"万": 10_000, "w": 10_000, "W": 10_000, "亿": 100_000_000, "k": 1_000, "K": 1_000}
|
||||
|
||||
|
||||
def as_int(value: Any) -> Optional[int]:
|
||||
"""解析计数。处理 "1.2万"、"1,234"、"123" 与已经是数字的情况。"""
|
||||
if value is None or isinstance(value, bool):
|
||||
return None
|
||||
if isinstance(value, (int, float)):
|
||||
return int(value)
|
||||
|
||||
text = str(value).strip().replace(",", "").replace(" ", "")
|
||||
if not text or text in ("-", "--", "暂无"):
|
||||
return None
|
||||
|
||||
unit = 1
|
||||
suffix = text[-1]
|
||||
if suffix in _COUNT_UNITS:
|
||||
unit = _COUNT_UNITS[suffix]
|
||||
text = text[:-1]
|
||||
try:
|
||||
return int(float(text) * unit)
|
||||
except ValueError:
|
||||
return None
|
||||
|
||||
|
||||
def as_float(value: Any) -> Optional[float]:
|
||||
"""解析比率。``"12.3%"`` -> 12.3;``"0.123"`` 原样返回数字。"""
|
||||
if value is None or isinstance(value, bool):
|
||||
return None
|
||||
if isinstance(value, (int, float)):
|
||||
return float(value)
|
||||
|
||||
text = str(value).strip().replace("%", "")
|
||||
try:
|
||||
return float(text)
|
||||
except ValueError:
|
||||
return None
|
||||
|
||||
|
||||
_DURATION_RE = re.compile(r"(?:(\d+)\s*分)?\s*(?:(\d+)\s*秒)?")
|
||||
|
||||
|
||||
def as_seconds(value: Any) -> Optional[float]:
|
||||
"""解析时长。处理 ``"1分30秒"``、``"01:30"``、``"45"``(秒)。"""
|
||||
if value is None or isinstance(value, bool):
|
||||
return None
|
||||
if isinstance(value, (int, float)):
|
||||
return float(value)
|
||||
|
||||
text = str(value).strip()
|
||||
if not text or text in ("-", "--"):
|
||||
return None
|
||||
|
||||
if ":" in text:
|
||||
parts = text.split(":")
|
||||
try:
|
||||
total = 0.0
|
||||
for part in parts:
|
||||
total = total * 60 + float(part)
|
||||
return total
|
||||
except ValueError:
|
||||
return None
|
||||
|
||||
if "分" in text or "秒" in text:
|
||||
minutes = re.search(r"(\d+)\s*分", text)
|
||||
seconds = re.search(r"(\d+)\s*秒", text)
|
||||
if not minutes and not seconds:
|
||||
return None
|
||||
return float(minutes.group(1) if minutes else 0) * 60 + float(
|
||||
seconds.group(1) if seconds else 0
|
||||
)
|
||||
|
||||
try:
|
||||
return float(text)
|
||||
except ValueError:
|
||||
return None
|
||||
|
||||
|
||||
# 指标 -> 候选原始字段名。别名来自公开资料,权威与否只能等真实响应来验证。
|
||||
_FIELD_ALIASES: Dict[str, tuple] = {
|
||||
"exposure": ("exposure", "exposure_count", "imp", "impression", "impression_count"),
|
||||
"views": ("views", "view", "view_count", "watch", "watch_count", "read_count"),
|
||||
"likes": ("likes", "like", "like_count", "liked_count"),
|
||||
"comments": ("comments", "comment", "comment_count", "comments_count"),
|
||||
"favorites": ("favorites", "favorite", "favorite_count", "collect", "collect_count", "collected_count"),
|
||||
"shares": ("shares", "share", "share_count", "shared_count"),
|
||||
"new_followers": ("new_followers", "fans_growth", "follower_growth", "increase_fans"),
|
||||
"danmaku": ("danmaku", "danmaku_count", "barrage"),
|
||||
"cover_ctr": ("cover_ctr", "cover_click_rate", "cover_click_ratio", "ctr"),
|
||||
"avg_watch_seconds": ("avg_watch_seconds", "avg_watch_time", "average_watch_time"),
|
||||
"two_second_exit_rate": ("two_second_exit_rate", "2s_exit_rate", "exit_rate_2s"),
|
||||
"completion_rate": ("completion_rate", "finish_rate", "complete_rate"),
|
||||
}
|
||||
|
||||
_COUNT_FIELDS = {
|
||||
"exposure", "views", "likes", "comments", "favorites", "shares", "new_followers", "danmaku",
|
||||
}
|
||||
_DURATION_FIELDS = {"avg_watch_seconds"}
|
||||
_RATE_FIELDS = {"two_second_exit_rate", "completion_rate", "cover_ctr"}
|
||||
|
||||
NOTE_ID_ALIASES = ("note_id", "noteId", "id", "content_id", "item_id")
|
||||
TITLE_ALIASES = ("title", "content", "note_title", "display_title")
|
||||
PUBLISH_TIME_ALIASES = ("publish_time", "publishTime", "post_time", "create_time", "time")
|
||||
|
||||
|
||||
def _normalize_metric(name: str, raw: Any) -> Any:
|
||||
if name in _COUNT_FIELDS:
|
||||
return as_int(raw)
|
||||
if name in _DURATION_FIELDS:
|
||||
return as_seconds(raw)
|
||||
if name in _RATE_FIELDS:
|
||||
return as_float(raw)
|
||||
return raw
|
||||
|
||||
|
||||
def normalize_note(item: Dict[str, Any]) -> Dict[str, Any]:
|
||||
"""把一条原始记录规范化成落库用的字段。"""
|
||||
note: Dict[str, Any] = {
|
||||
"note_id": str(_pick(item, *NOTE_ID_ALIASES) or ""),
|
||||
"title": str(_pick(item, *TITLE_ALIASES) or ""),
|
||||
"publish_time": as_int(_pick(item, *PUBLISH_TIME_ALIASES)),
|
||||
}
|
||||
for name, aliases in _FIELD_ALIASES.items():
|
||||
note[name] = _normalize_metric(name, _pick(item, *aliases))
|
||||
return note
|
||||
|
||||
|
||||
def find_note_list(payload: Any) -> List[Dict[str, Any]]:
|
||||
"""在响应里找出笔记数组。
|
||||
|
||||
接口的**确切结构还没亲眼见过**(权限未生效时 `data.result` 里没有数据),
|
||||
所以不写死路径:遍历 JSON,挑出"看起来像一批笔记记录"的那个列表 ——
|
||||
元素是 dict,且至少带一个指标字段。找不到就返回空列表,让上层如实报"没数据"。
|
||||
"""
|
||||
best: List[Dict[str, Any]] = []
|
||||
metric_keys = {alias for aliases in _FIELD_ALIASES.values() for alias in aliases}
|
||||
|
||||
def walk(value: Any) -> None:
|
||||
nonlocal best
|
||||
if isinstance(value, dict):
|
||||
for child in value.values():
|
||||
walk(child)
|
||||
elif isinstance(value, list):
|
||||
if value and isinstance(value[0], dict):
|
||||
keys = set(value[0].keys())
|
||||
if keys & metric_keys and len(value) > len(best):
|
||||
best = value
|
||||
for child in value:
|
||||
walk(child)
|
||||
|
||||
walk(payload)
|
||||
return best
|
||||
|
||||
|
||||
class CreatorClient:
|
||||
"""一个账号的客户端。``cookie`` 就是它的全部身份。"""
|
||||
|
||||
def __init__(self, cookie: str, timeout: float = 25.0):
|
||||
self.cookies = trans_cookies(cookie)
|
||||
self.a1 = self.cookies.get("a1", "")
|
||||
self._timeout = timeout
|
||||
|
||||
@property
|
||||
def looks_authenticated(self) -> bool:
|
||||
"""签名需要 a1;没有它连请求都签不出来。"""
|
||||
return bool(self.a1)
|
||||
|
||||
def _headers(self, api: str, body: dict | None = None) -> Dict[str, str]:
|
||||
cookie_header = "; ".join(f"{k}={v}" for k, v in self.cookies.items())
|
||||
return {
|
||||
"user-agent": USER_AGENT,
|
||||
"accept": "application/json, text/plain, */*",
|
||||
"accept-language": "zh-CN,zh;q=0.9",
|
||||
"origin": CREATOR_ORIGIN,
|
||||
"referer": DATA_ANALYSIS_PAGE,
|
||||
"cookie": cookie_header,
|
||||
**sign_xyw(api, self.a1, body=body),
|
||||
}
|
||||
|
||||
async def _get(self, path: str, params: Dict[str, Any]) -> Dict[str, Any]:
|
||||
if not self.looks_authenticated:
|
||||
raise CreatorApiError("cookie 里没有 a1,无法完成签名", status=0)
|
||||
|
||||
query = "&".join(f"{k}={v}" for k, v in params.items())
|
||||
api = signed_api(path, query)
|
||||
url = f"{CREATOR_ORIGIN}{path}?{query}" if query else f"{CREATOR_ORIGIN}{path}"
|
||||
|
||||
async with httpx.AsyncClient(timeout=self._timeout, follow_redirects=False) as client:
|
||||
response = await client.get(url, headers=self._headers(api))
|
||||
return self._unwrap(response)
|
||||
|
||||
@staticmethod
|
||||
def _unwrap(response: httpx.Response) -> Dict[str, Any]:
|
||||
if response.status_code == 406:
|
||||
raise CreatorApiError(
|
||||
"签名被网关拒绝(406)—— 待签字符串的拼法不对", status=406
|
||||
)
|
||||
try:
|
||||
payload = response.json()
|
||||
except Exception as exc: # noqa: BLE001
|
||||
raise CreatorApiError(
|
||||
f"响应不是 JSON(HTTP {response.status_code})", status=response.status_code
|
||||
) from exc
|
||||
|
||||
if response.status_code == 401 or payload.get("code") == -100:
|
||||
raise CreatorApiError("登录态无效或已过期", status=401, payload=payload)
|
||||
if not payload.get("success", True):
|
||||
raise CreatorApiError(
|
||||
str(payload.get("msg") or "接口返回失败"),
|
||||
status=response.status_code,
|
||||
payload=payload,
|
||||
)
|
||||
return payload
|
||||
|
||||
async def fetch_user_info(self) -> Dict[str, Any]:
|
||||
"""当前 cookie 属于哪个账号。登录后用它取名与去重。"""
|
||||
payload = await self._get(USER_INFO_PATH, {})
|
||||
data = payload.get("data") or {}
|
||||
return {
|
||||
"user_id": str(data.get("userId") or ""),
|
||||
"nickname": str(data.get("userName") or ""),
|
||||
"avatar": str(data.get("userAvatar") or ""),
|
||||
"red_id": str(data.get("redId") or ""),
|
||||
"role": str(data.get("role") or ""),
|
||||
"permissions": list(data.get("permissions") or []),
|
||||
}
|
||||
|
||||
async def fetch_permission(self) -> Dict[str, Any]:
|
||||
"""数据权限状态。
|
||||
|
||||
``tip_msg`` 是后台原话(实测是"已为您申请数据权限,次日可查看"),照抄不改写 ——
|
||||
这条信息必须原样交给用户,它解释了"为什么没有数据"。
|
||||
"""
|
||||
payload = await self._get(PERMISSION_PATH, {})
|
||||
data = payload.get("data") or {}
|
||||
return {
|
||||
"display": data.get("display"),
|
||||
"status": data.get("status"),
|
||||
"tip": str(data.get("tip_msg") or ""),
|
||||
}
|
||||
|
||||
async def fetch_note_list(
|
||||
self, start_ms: int, end_ms: int, page_num: int = 1, page_size: int = 10, note_type: int = 0
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""按发布时间区间取笔记列表。
|
||||
|
||||
参数与顺序**照抄真实页面的请求**(见 Phase 0 抓包),不要凭感觉改。
|
||||
"""
|
||||
payload = await self._get(
|
||||
NOTE_LIST_PATH,
|
||||
{
|
||||
"post_begin_time": start_ms,
|
||||
"post_end_time": end_ms,
|
||||
"type": note_type,
|
||||
"page_size": page_size,
|
||||
"page_num": page_num,
|
||||
},
|
||||
)
|
||||
return [normalize_note(item) for item in find_note_list(payload)]
|
||||
@@ -0,0 +1,300 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/login.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""运营账号的扫码登录。
|
||||
|
||||
**与监控的扫码登录(`api.monitor.qrlogin`)有一处决定性差异**:那边把登录态写进
|
||||
浏览器**默认 profile**,因为爬虫要复用它;这边要的是 **cookie 字符串**,因为采集
|
||||
走纯 HTTP。所以这里每次登录都开一个**临时上下文**,扫完取出 cookie 就丢弃 ——
|
||||
|
||||
* 登第二个账号不会把第一个顶掉(默认 profile 只能装一个登录态);
|
||||
* 完全不影响监控那个登录态;
|
||||
* 十个账号互不干扰。
|
||||
|
||||
扫码入口仍是主站(`www.xiaohongshu.com`):Phase 0 实测证明**主站的 cookie 就能
|
||||
认证创作者后台**,不需要单独的创作者登录。
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import time
|
||||
from typing import Any, Dict, Optional
|
||||
|
||||
from playwright.async_api import async_playwright
|
||||
from tools import utils
|
||||
|
||||
from ..monitor.platforms import PLATFORM_XHS
|
||||
from .client import CreatorApiError, CreatorClient
|
||||
|
||||
QR_TTL_SECONDS = 300
|
||||
|
||||
STATUS_IDLE = "idle"
|
||||
STATUS_WAITING = "waiting"
|
||||
STATUS_SUCCESS = "success"
|
||||
STATUS_EXPIRED = "expired"
|
||||
STATUS_ERROR = "error"
|
||||
|
||||
LOGIN_URL = "https://www.xiaohongshu.com"
|
||||
QR_SELECTOR = "xpath=//img[@class='qrcode-img']"
|
||||
LOGIN_BUTTON_SELECTOR = "xpath=//*[@id='app']/div[1]/div[2]/div[1]/ul/div[1]/button"
|
||||
|
||||
# 判据不读页面状态,而是拿 cookie 直接问创作者后台"我是谁"。
|
||||
#
|
||||
# **为什么不用页面状态**:`window.__INITIAL_STATE__` 是**页面加载那一刻的快照**。
|
||||
# 监控那边的同一个探针能用,是因为那台浏览器的页面加载时就已经登录了,快照里
|
||||
# loggedIn 就是 true。而扫码是"页面加载之后才登录的"—— SPA 内部确实登进去了,
|
||||
# 但那个初始快照不会翻转,于是检测永远等不到,界面就一直停在二维码上。
|
||||
#
|
||||
# `user/info` 则是权威的:实测**游客也会拿到 a1**(所以签名算得出来),但接口直接
|
||||
# 回 401「无登录信息」;只有真正登录了才返回 user_id。所以"有 a1"什么都证明不了,
|
||||
# "后台认这份身份"才是。
|
||||
LOGIN_CHECK_INTERVAL_SECONDS = 5.0
|
||||
|
||||
_lock = asyncio.Lock()
|
||||
_current: Optional["AccountLoginSession"] = None
|
||||
|
||||
# 完成后的快照。会话一旦被取走 cookie 就会拆掉,而前端可能还有一个在飞的轮询——
|
||||
# 那个请求若看到 _current 为空就会回报 idle,把已经显示出来的成功状态又擦掉。
|
||||
# 把结果留在这里,重复轮询就稳定得多。
|
||||
_last_result: Optional[Dict[str, Any]] = None
|
||||
|
||||
_playwright: Any = None
|
||||
|
||||
|
||||
def _cdp_url() -> str:
|
||||
import os
|
||||
|
||||
import config
|
||||
|
||||
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
|
||||
|
||||
|
||||
async def _connect():
|
||||
global _playwright
|
||||
if _playwright is None:
|
||||
_playwright = await async_playwright().start()
|
||||
return await _playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
|
||||
|
||||
|
||||
async def _disconnect() -> None:
|
||||
global _playwright
|
||||
if _playwright is not None:
|
||||
try:
|
||||
await _playwright.stop()
|
||||
except Exception:
|
||||
pass
|
||||
_playwright = None
|
||||
|
||||
|
||||
async def _read_qr(page: Any) -> str:
|
||||
image = await utils.find_login_qrcode(page, selector=QR_SELECTOR)
|
||||
if image:
|
||||
return image
|
||||
# 登录框不一定自己弹出来,这是爬虫自身扫码流程的同款兜底。
|
||||
await asyncio.sleep(0.5)
|
||||
try:
|
||||
await page.locator(LOGIN_BUTTON_SELECTOR).click(timeout=5000)
|
||||
except Exception:
|
||||
return ""
|
||||
return await utils.find_login_qrcode(page, selector=QR_SELECTOR)
|
||||
|
||||
|
||||
class AccountLoginSession:
|
||||
"""一次针对**临时上下文**的扫码尝试。"""
|
||||
|
||||
def __init__(self, context: Any, page: Any) -> None:
|
||||
self.status = STATUS_WAITING
|
||||
self.message = "请用手机扫描二维码"
|
||||
self.image = ""
|
||||
self.started_at = time.time()
|
||||
self.account: Optional[Dict[str, Any]] = None
|
||||
self.cookie: str = ""
|
||||
self.platform = PLATFORM_XHS
|
||||
self._context = context
|
||||
self._page = page
|
||||
self._last_login_check = 0.0
|
||||
|
||||
@property
|
||||
def elapsed(self) -> float:
|
||||
return time.time() - self.started_at
|
||||
|
||||
async def refresh(self) -> None:
|
||||
if self.status != STATUS_WAITING:
|
||||
return
|
||||
if self.elapsed > QR_TTL_SECONDS:
|
||||
self.status = STATUS_EXPIRED
|
||||
self.message = "二维码已超时,请重新获取"
|
||||
return
|
||||
|
||||
# 前端每 2 秒问一次,但没必要每次都去打后台接口 —— 一次真实的网络往返
|
||||
# 去确认一个通常还没发生的事件是浪费。
|
||||
now = time.time()
|
||||
if now - self._last_login_check < LOGIN_CHECK_INTERVAL_SECONDS:
|
||||
return
|
||||
self._last_login_check = now
|
||||
|
||||
try:
|
||||
cookies = await self._context.cookies()
|
||||
except Exception:
|
||||
self.status = STATUS_ERROR
|
||||
self.message = "登录窗口已被关闭,请重新获取"
|
||||
return
|
||||
|
||||
cookie = "; ".join(f"{c['name']}={c['value']}" for c in cookies)
|
||||
try:
|
||||
info = await CreatorClient(cookie).fetch_user_info()
|
||||
except CreatorApiError:
|
||||
# 还没登录(或者刚扫、后端还没认),继续等。
|
||||
return
|
||||
|
||||
if not info.get("user_id"):
|
||||
return
|
||||
|
||||
# 登录成功:cookie 取自**这个临时上下文**,取完上下文就丢弃,
|
||||
# 所以不会残留、也不会影响别的账号。
|
||||
self.cookie = cookie
|
||||
self.account = info
|
||||
self.status = STATUS_SUCCESS
|
||||
self.message = f"登录成功:{info.get('nickname') or info['user_id']}"
|
||||
|
||||
def snapshot(self) -> Dict[str, Any]:
|
||||
return {
|
||||
"status": self.status,
|
||||
"message": self.message,
|
||||
"image": self.image,
|
||||
"elapsed": int(self.elapsed),
|
||||
"expires_in": max(0, int(QR_TTL_SECONDS - self.elapsed)),
|
||||
"account": self.account,
|
||||
}
|
||||
|
||||
async def close(self) -> None:
|
||||
"""关掉临时上下文。这是它存在的全部意义 —— 用完即弃。"""
|
||||
for closer in (self._page.close, self._context.close):
|
||||
try:
|
||||
await closer()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
async def _teardown_locked() -> None:
|
||||
global _current
|
||||
if _current is not None:
|
||||
await _current.close()
|
||||
_current = None
|
||||
|
||||
|
||||
async def start() -> Dict[str, Any]:
|
||||
"""开一个临时上下文,打开登录页,取回二维码。"""
|
||||
global _current, _last_result
|
||||
|
||||
async with _lock:
|
||||
_last_result = None
|
||||
await _teardown_locked()
|
||||
|
||||
try:
|
||||
browser = await _connect()
|
||||
except Exception as exc:
|
||||
await _disconnect()
|
||||
raise RuntimeError(
|
||||
f"连接浏览器失败({_cdp_url()})。请确认服务器上的 Chrome 以 "
|
||||
f"--remote-debugging-port 启动。原始错误:{exc}"
|
||||
) from exc
|
||||
|
||||
# 临时上下文,不是 contexts[0]。这里刻意要一个干净的身份 ——
|
||||
# 借用操作者自己的登录态会让"新增账号"变成"再读一遍当前账号"。
|
||||
context = await browser.new_context()
|
||||
page = await context.new_page()
|
||||
|
||||
try:
|
||||
await page.goto(LOGIN_URL, wait_until="domcontentloaded", timeout=45000)
|
||||
image = await _read_qr(page)
|
||||
except Exception as exc:
|
||||
try:
|
||||
await page.close()
|
||||
await context.close()
|
||||
except Exception:
|
||||
pass
|
||||
raise RuntimeError(f"打开登录页失败:{exc}") from exc
|
||||
|
||||
session = AccountLoginSession(context, page)
|
||||
session.image = image
|
||||
if not image:
|
||||
session.status = STATUS_ERROR
|
||||
session.message = "页面上没找到二维码,请确认站点结构没有变化"
|
||||
|
||||
_current = session
|
||||
return session.snapshot()
|
||||
|
||||
|
||||
async def status() -> Dict[str, Any]:
|
||||
async with _lock:
|
||||
if _current is None:
|
||||
if _last_result is not None:
|
||||
return _last_result
|
||||
return {
|
||||
"status": STATUS_IDLE,
|
||||
"message": "",
|
||||
"image": "",
|
||||
"elapsed": 0,
|
||||
"expires_in": 0,
|
||||
"account": None,
|
||||
}
|
||||
await _current.refresh()
|
||||
return _current.snapshot()
|
||||
|
||||
|
||||
async def remember_result(snapshot: Dict[str, Any]) -> None:
|
||||
"""记住已完成的扫码结果,供后续轮询重复返回。"""
|
||||
global _last_result
|
||||
async with _lock:
|
||||
_last_result = snapshot
|
||||
|
||||
|
||||
async def take_cookie() -> Optional[str]:
|
||||
"""取走已登录的 cookie 并结束会话。
|
||||
|
||||
由路由层在落库时调用。cookie 只经内存传递,**不进响应体** —— 它是凭证,
|
||||
前端没有任何理由看到它。
|
||||
"""
|
||||
global _current
|
||||
async with _lock:
|
||||
if _current is None or _current.status != STATUS_SUCCESS:
|
||||
return None
|
||||
cookie = _current.cookie
|
||||
await _teardown_locked()
|
||||
return cookie
|
||||
|
||||
|
||||
async def cancel() -> Dict[str, Any]:
|
||||
global _last_result
|
||||
async with _lock:
|
||||
_last_result = None
|
||||
await _teardown_locked()
|
||||
return {
|
||||
"status": STATUS_IDLE,
|
||||
"message": "已取消",
|
||||
"image": "",
|
||||
"elapsed": 0,
|
||||
"expires_in": 0,
|
||||
"account": None,
|
||||
}
|
||||
|
||||
|
||||
async def shutdown() -> None:
|
||||
async with _lock:
|
||||
await _teardown_locked()
|
||||
await _disconnect()
|
||||
@@ -0,0 +1,132 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/models.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""运营模块的数据模型。
|
||||
|
||||
**刻意复用 `MonitorBase`**:这样 `init_db` 的 `create_all` 会顺手建出新表,而
|
||||
`_ensure_columns`(已改为按模型元数据推导)也会自动给新表补字段 —— 不必再维护一份
|
||||
建表语句。表落在同一个库里,与监控互不干扰。
|
||||
"""
|
||||
|
||||
from typing import Optional
|
||||
|
||||
from sqlalchemy import BigInteger, Float, ForeignKey, Index, Integer, String, Text
|
||||
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
||||
|
||||
from ..monitor.models import MonitorBase
|
||||
|
||||
# 创作者后台的数据权限是「首次访问时自动申请、次日生效」。这个状态必须如实呈现:
|
||||
# 显示成"没数据"会让人以为采集坏了,实际是在等审批。
|
||||
PERMISSION_UNKNOWN = "unknown"
|
||||
PERMISSION_PENDING = "pending" # 已申请,未生效(提示语:"次日可查看")
|
||||
PERMISSION_ACTIVE = "active"
|
||||
PERMISSION_MISSING = "missing" # 接口明确说没有权限
|
||||
|
||||
# 账号自身的可用性。
|
||||
ACCOUNT_OK = "ok"
|
||||
ACCOUNT_EXPIRED = "expired" # cookie 失效,需要重新扫码
|
||||
ACCOUNT_ERROR = "error"
|
||||
|
||||
|
||||
class CreatorAccount(MonitorBase):
|
||||
"""一个自己的小红书账号。
|
||||
|
||||
纯请求路线下,**一个账号的全部身份就是一份 cookie** —— 没有浏览器 profile、
|
||||
没有独立目录。所以"多账号"在这里只是表里的多行,不是多套运行环境。
|
||||
|
||||
`cookie` 是凭证:与监控的 cookie 同样对待,只存库、绝不回显接口。
|
||||
"""
|
||||
|
||||
__tablename__ = "creator_account"
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
# 展示名。优先用后台返回的昵称,用户可以改。
|
||||
nickname: Mapped[str] = mapped_column(String(128), nullable=False, default="")
|
||||
# 创作者后台的账号标识,由 /api/galaxy/user/info 返回,用于去重。
|
||||
user_id: Mapped[str] = mapped_column(String(64), nullable=False, default="", index=True)
|
||||
red_id: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
avatar: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
|
||||
cookie: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
|
||||
status: Mapped[str] = mapped_column(String(16), nullable=False, default=ACCOUNT_OK)
|
||||
permission_status: Mapped[str] = mapped_column(
|
||||
String(16), nullable=False, default=PERMISSION_UNKNOWN
|
||||
)
|
||||
# 后台原话,例如"已为您申请数据权限,次日可查看"。照抄,不改写。
|
||||
permission_tip: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
|
||||
last_checked_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
last_synced_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
# 上次同步用的时间范围(天,按发布时间)。当前展示的数据就是这个范围的产物 ——
|
||||
# 不记下来的话,界面只能说明"同步过了",说不清是哪一段。
|
||||
last_sync_days: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
last_error: Mapped[Optional[str]] = mapped_column(Text)
|
||||
|
||||
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
notes: Mapped[list["CreatorNoteStat"]] = relationship(
|
||||
back_populates="account", cascade="all, delete-orphan"
|
||||
)
|
||||
|
||||
|
||||
class CreatorNoteStat(MonitorBase):
|
||||
"""一篇作品在某个采集时点的运营数据。
|
||||
|
||||
创作者后台给的是**累计值**(截至查询时点),所以反复采集天然形成时间序列 ——
|
||||
与监控的"快照 + 差分"是同一个思路,因此这里保留 `captured_at` 而不是覆盖写。
|
||||
"""
|
||||
|
||||
__tablename__ = "creator_note_stat"
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
account_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("creator_account.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
note_id: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
|
||||
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
# 发布时间(毫秒)。后台按发布时间筛选,这是它的主时间轴。
|
||||
publish_time: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
|
||||
# --- 运营指标 ---------------------------------------------------------
|
||||
# 计数用 BigInteger:曝光量可以很大,用 INT 迟早溢出。
|
||||
exposure: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
views: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
likes: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
comments: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
favorites: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
shares: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
new_followers: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
danmaku: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
|
||||
# 比率与时长。后台返回的可能是 "12.3%"/"1分30秒" 这类字符串,解析不了的存 NULL
|
||||
# 而不是 0 —— 与监控层的口径一致:0 是真实值,NULL 是"不知道"。
|
||||
cover_ctr: Mapped[Optional[float]] = mapped_column(Float)
|
||||
avg_watch_seconds: Mapped[Optional[float]] = mapped_column(Float)
|
||||
two_second_exit_rate: Mapped[Optional[float]] = mapped_column(Float)
|
||||
completion_rate: Mapped[Optional[float]] = mapped_column(Float)
|
||||
|
||||
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
|
||||
|
||||
account: Mapped["CreatorAccount"] = relationship(back_populates="notes")
|
||||
|
||||
__table_args__ = (
|
||||
# 同一个时点同一篇只留一行,重复同步不会堆积。
|
||||
Index("ix_creator_note_stat_unique", "account_id", "note_id", "captured_at", unique=True),
|
||||
)
|
||||
@@ -0,0 +1,360 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/service.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""运营账号的增删查与数据同步。
|
||||
|
||||
一条贯穿全文件的规则:**cookie 是凭证,永远不出现在返回给上层的结构里。**
|
||||
对外只给 `has_cookie` 这样的布尔量,与监控层对 cookie 的处理保持一致。
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
from datetime import datetime, time, timedelta
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from sqlalchemy import delete, func, select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from .client import CreatorApiError, CreatorClient
|
||||
from .models import (
|
||||
ACCOUNT_ERROR,
|
||||
ACCOUNT_EXPIRED,
|
||||
ACCOUNT_OK,
|
||||
PERMISSION_ACTIVE,
|
||||
PERMISSION_MISSING,
|
||||
PERMISSION_PENDING,
|
||||
PERMISSION_UNKNOWN,
|
||||
CreatorAccount,
|
||||
CreatorNoteStat,
|
||||
)
|
||||
|
||||
# 同步一次最多翻多少页。后台默认一页 10 条,200 页足以覆盖任何正常账号,
|
||||
# 同时防止"接口不返回 has_more"时无限翻下去。
|
||||
MAX_SYNC_PAGES = 200
|
||||
PAGE_SIZE = 10
|
||||
|
||||
|
||||
def _account_dict(account: CreatorAccount, note_count: int = 0) -> Dict[str, Any]:
|
||||
"""账号的对外表示。**刻意不含 cookie。**"""
|
||||
return {
|
||||
"id": account.id,
|
||||
"nickname": account.nickname,
|
||||
"user_id": account.user_id,
|
||||
"red_id": account.red_id,
|
||||
"avatar": account.avatar,
|
||||
"status": account.status,
|
||||
"permission_status": account.permission_status,
|
||||
# 后台原话照抄。"次日可查看"这类信息只能由它自己说,改写就失真了。
|
||||
"permission_tip": account.permission_tip,
|
||||
"last_checked_at": account.last_checked_at,
|
||||
"last_synced_at": account.last_synced_at,
|
||||
"last_sync_days": account.last_sync_days,
|
||||
"last_error": account.last_error,
|
||||
"has_cookie": bool(account.cookie),
|
||||
"created_at": account.created_at,
|
||||
"note_count": note_count,
|
||||
}
|
||||
|
||||
|
||||
async def list_accounts(session: AsyncSession) -> List[Dict[str, Any]]:
|
||||
accounts = list(
|
||||
(await session.scalars(select(CreatorAccount).order_by(CreatorAccount.id))).all()
|
||||
)
|
||||
counts = dict(
|
||||
(
|
||||
await session.execute(
|
||||
select(CreatorNoteStat.account_id, func.count(func.distinct(CreatorNoteStat.note_id)))
|
||||
.group_by(CreatorNoteStat.account_id)
|
||||
)
|
||||
).all()
|
||||
)
|
||||
return [_account_dict(account, counts.get(account.id, 0)) for account in accounts]
|
||||
|
||||
|
||||
async def get_account(session: AsyncSession, account_id: int) -> CreatorAccount:
|
||||
account = await session.get(CreatorAccount, account_id)
|
||||
if account is None:
|
||||
raise ValueError(f"账号 {account_id} 不存在")
|
||||
return account
|
||||
|
||||
|
||||
async def account_detail(session: AsyncSession, account_id: int) -> Dict[str, Any]:
|
||||
account = await get_account(session, account_id)
|
||||
notes = await latest_notes(session, account_id)
|
||||
return {
|
||||
"account": _account_dict(account, len(notes)),
|
||||
"notes": notes,
|
||||
"summary": _summarize(notes),
|
||||
}
|
||||
|
||||
|
||||
def _summarize(notes: List[Dict[str, Any]]) -> Dict[str, Any]:
|
||||
"""账号级汇总。取最后一轮快照的累计值之和。"""
|
||||
totals = {
|
||||
key: 0
|
||||
for key in ("exposure", "views", "likes", "comments", "favorites", "shares", "new_followers")
|
||||
}
|
||||
for note in notes:
|
||||
for key in totals:
|
||||
value = note.get(key)
|
||||
if isinstance(value, (int, float)):
|
||||
totals[key] += int(value)
|
||||
return totals
|
||||
|
||||
|
||||
async def latest_notes(session: AsyncSession, account_id: int) -> List[Dict[str, Any]]:
|
||||
"""每个作品取**最近一次**快照。
|
||||
|
||||
表里保留全部历史(换个时点就是一条新行),但列表只该展示"现在",否则同一个
|
||||
作品会在列表里出现多次。
|
||||
"""
|
||||
newest = (
|
||||
select(
|
||||
CreatorNoteStat.note_id,
|
||||
func.max(CreatorNoteStat.captured_at).label("captured_at"),
|
||||
)
|
||||
.where(CreatorNoteStat.account_id == account_id)
|
||||
.group_by(CreatorNoteStat.note_id)
|
||||
.subquery()
|
||||
)
|
||||
rows = (
|
||||
await session.scalars(
|
||||
select(CreatorNoteStat)
|
||||
.join(
|
||||
newest,
|
||||
(CreatorNoteStat.note_id == newest.c.note_id)
|
||||
& (CreatorNoteStat.captured_at == newest.c.captured_at),
|
||||
)
|
||||
.where(CreatorNoteStat.account_id == account_id)
|
||||
# 不要用 nullslast():那是 PostgreSQL 语法,MySQL 5.7 会直接抛 1064 语法错误。
|
||||
# MySQL 把 NULL 视为比任何值都小,所以 DESC 天然把未解析出发布时间的排在最后。
|
||||
# 这个 bug 只在真机上才暴露 —— SQLite 从 3.30 起支持 NULLS LAST,测试环境测不出来。
|
||||
.order_by(CreatorNoteStat.publish_time.desc())
|
||||
)
|
||||
).all()
|
||||
return [_note_dict(row) for row in rows]
|
||||
|
||||
|
||||
def _note_dict(row: CreatorNoteStat) -> Dict[str, Any]:
|
||||
return {
|
||||
"note_id": row.note_id,
|
||||
"title": row.title,
|
||||
"publish_time": row.publish_time,
|
||||
"exposure": row.exposure,
|
||||
"views": row.views,
|
||||
"likes": row.likes,
|
||||
"comments": row.comments,
|
||||
"favorites": row.favorites,
|
||||
"shares": row.shares,
|
||||
"new_followers": row.new_followers,
|
||||
"danmaku": row.danmaku,
|
||||
"cover_ctr": row.cover_ctr,
|
||||
"avg_watch_seconds": row.avg_watch_seconds,
|
||||
"two_second_exit_rate": row.two_second_exit_rate,
|
||||
"completion_rate": row.completion_rate,
|
||||
"captured_at": row.captured_at,
|
||||
}
|
||||
|
||||
|
||||
async def upsert_account_from_cookie(session: AsyncSession, cookie: str) -> Dict[str, Any]:
|
||||
"""用一份 cookie 识别并保存账号。
|
||||
|
||||
识别靠 `user/info` 而不是让用户填名字 —— 填错名字只会让后面所有数据对不上号。
|
||||
已有同 `user_id` 的账号则更新它的 cookie(重新登录)。
|
||||
"""
|
||||
client = CreatorClient(cookie)
|
||||
if not client.looks_authenticated:
|
||||
raise ValueError("这份 cookie 里没有 a1,无法签名,请重新扫码")
|
||||
|
||||
try:
|
||||
info = await client.fetch_user_info()
|
||||
except CreatorApiError as exc:
|
||||
raise ValueError(f"登录态无法使用:{exc}") from exc
|
||||
|
||||
if not info.get("user_id"):
|
||||
raise ValueError("接口没有返回账号标识,可能登录态无效")
|
||||
|
||||
now = get_current_timestamp()
|
||||
account = await session.scalar(
|
||||
select(CreatorAccount).where(CreatorAccount.user_id == info["user_id"])
|
||||
)
|
||||
if account is None:
|
||||
account = CreatorAccount(created_at=now)
|
||||
session.add(account)
|
||||
|
||||
account.nickname = info.get("nickname") or account.nickname or "未命名账号"
|
||||
account.user_id = info["user_id"]
|
||||
account.red_id = info.get("red_id") or ""
|
||||
account.avatar = info.get("avatar") or ""
|
||||
account.cookie = cookie
|
||||
account.status = ACCOUNT_OK
|
||||
account.last_error = None
|
||||
account.last_checked_at = now
|
||||
account.updated_at = now
|
||||
|
||||
await session.flush()
|
||||
|
||||
# 顺手把权限状态也拉一次:新账号几乎必然处于"已申请、次日生效",
|
||||
# 当场告诉用户,比让他明天再回来问要好。
|
||||
await refresh_permission(session, account)
|
||||
|
||||
return _account_dict(account)
|
||||
|
||||
|
||||
async def refresh_permission(session: AsyncSession, account: CreatorAccount) -> None:
|
||||
"""查询并记录数据权限状态。失败不影响账号本身可用。"""
|
||||
try:
|
||||
permission = await CreatorClient(account.cookie).fetch_permission()
|
||||
except CreatorApiError as exc:
|
||||
if exc.status == 401:
|
||||
account.status = ACCOUNT_EXPIRED
|
||||
account.last_error = str(exc)
|
||||
else:
|
||||
account.last_error = str(exc)
|
||||
account.updated_at = get_current_timestamp()
|
||||
return
|
||||
|
||||
display = permission.get("display")
|
||||
status = permission.get("status")
|
||||
account.permission_tip = permission.get("tip") or ""
|
||||
|
||||
if display or status:
|
||||
account.permission_status = PERMISSION_ACTIVE
|
||||
elif account.permission_tip:
|
||||
# 有提示语但未开通 —— 实测就是"已为您申请数据权限,次日可查看"。
|
||||
account.permission_status = PERMISSION_PENDING
|
||||
else:
|
||||
account.permission_status = PERMISSION_MISSING
|
||||
|
||||
account.status = ACCOUNT_OK
|
||||
account.last_error = None
|
||||
account.last_checked_at = get_current_timestamp()
|
||||
account.updated_at = account.last_checked_at
|
||||
|
||||
|
||||
async def check_account(session: AsyncSession, account_id: int) -> Dict[str, Any]:
|
||||
"""重新检测一个账号:登录态还在不在、权限开通没有。"""
|
||||
account = await get_account(session, account_id)
|
||||
await refresh_permission(session, account)
|
||||
count = (
|
||||
await session.execute(
|
||||
select(func.count(func.distinct(CreatorNoteStat.note_id))).where(
|
||||
CreatorNoteStat.account_id == account_id
|
||||
)
|
||||
)
|
||||
).scalar() or 0
|
||||
return _account_dict(account, count)
|
||||
|
||||
|
||||
async def delete_account(session: AsyncSession, account_id: int) -> None:
|
||||
account = await get_account(session, account_id)
|
||||
await session.execute(
|
||||
delete(CreatorNoteStat).where(CreatorNoteStat.account_id == account_id)
|
||||
)
|
||||
await session.delete(account)
|
||||
|
||||
|
||||
def _day_bounds(days: int) -> tuple[int, int]:
|
||||
"""最近 N 天的起止(毫秒)。与后台的按发布时间筛选对齐。"""
|
||||
today = datetime.now()
|
||||
end = int(datetime.combine(today.date(), time(23, 59, 59)).timestamp() * 1000)
|
||||
start = int(
|
||||
datetime.combine((today - timedelta(days=days)).date(), time(0, 0, 0)).timestamp() * 1000
|
||||
)
|
||||
return start, end
|
||||
|
||||
|
||||
async def sync_account(
|
||||
session: AsyncSession, account_id: int, days: int = 90
|
||||
) -> Dict[str, Any]:
|
||||
"""拉取一个账号的作品运营数据并落库。
|
||||
|
||||
权限未生效时接口返回的是**空壳成功**(`data.result` 里没有数据),不是错误 ——
|
||||
所以"同步成功但 0 条"是正常结果,必须如实回报,不能让用户以为采集坏了。
|
||||
"""
|
||||
account = await get_account(session, account_id)
|
||||
if not account.cookie:
|
||||
raise ValueError("该账号没有可用的登录态,请重新扫码")
|
||||
|
||||
client = CreatorClient(account.cookie)
|
||||
start_ms, end_ms = _day_bounds(days)
|
||||
now = get_current_timestamp()
|
||||
|
||||
collected: List[Dict[str, Any]] = []
|
||||
for page in range(1, MAX_SYNC_PAGES + 1):
|
||||
try:
|
||||
batch = await client.fetch_note_list(start_ms, end_ms, page_num=page, page_size=PAGE_SIZE)
|
||||
except CreatorApiError as exc:
|
||||
account.last_error = str(exc)
|
||||
if exc.status == 401:
|
||||
account.status = ACCOUNT_EXPIRED
|
||||
account.updated_at = get_current_timestamp()
|
||||
raise ValueError(f"同步失败:{exc}") from exc
|
||||
|
||||
collected.extend(batch)
|
||||
if len(batch) < PAGE_SIZE:
|
||||
break
|
||||
await asyncio.sleep(0.6) # 对后台客气一点,这是自己的账号但仍是自动化访问
|
||||
|
||||
# 先删掉本时点可能存在的重复行,再写入 —— 表上有 (account, note, captured_at)
|
||||
# 唯一索引,重复同步不该报错。
|
||||
await session.execute(
|
||||
delete(CreatorNoteStat).where(
|
||||
CreatorNoteStat.account_id == account_id, CreatorNoteStat.captured_at == now
|
||||
)
|
||||
)
|
||||
for note in collected:
|
||||
if not note.get("note_id"):
|
||||
continue
|
||||
session.add(
|
||||
CreatorNoteStat(
|
||||
account_id=account_id,
|
||||
note_id=note["note_id"],
|
||||
title=note.get("title") or "",
|
||||
publish_time=note.get("publish_time"),
|
||||
exposure=note.get("exposure"),
|
||||
views=note.get("views"),
|
||||
likes=note.get("likes"),
|
||||
comments=note.get("comments"),
|
||||
favorites=note.get("favorites"),
|
||||
shares=note.get("shares"),
|
||||
new_followers=note.get("new_followers"),
|
||||
danmaku=note.get("danmaku"),
|
||||
cover_ctr=note.get("cover_ctr"),
|
||||
avg_watch_seconds=note.get("avg_watch_seconds"),
|
||||
two_second_exit_rate=note.get("two_second_exit_rate"),
|
||||
completion_rate=note.get("completion_rate"),
|
||||
captured_at=now,
|
||||
)
|
||||
)
|
||||
|
||||
account.last_synced_at = now
|
||||
account.last_sync_days = days
|
||||
account.last_error = None
|
||||
account.updated_at = now
|
||||
await refresh_permission(session, account)
|
||||
await session.flush()
|
||||
|
||||
return {
|
||||
"account_id": account_id,
|
||||
"fetched": len(collected),
|
||||
"days": days,
|
||||
"permission_status": account.permission_status,
|
||||
"permission_tip": account.permission_tip,
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/creator/signing.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""创作者后台的请求签名(XYW_ 方案)。
|
||||
|
||||
主站与创作者后台用的是**两套不同的签名**:主站是 VMP 的 `XYS_`,创作者后台是
|
||||
`XYW_`。后者简单得多 —— MD5 → base64 → AES-128-CBC,密钥与 IV 都是硬编码常量,
|
||||
纯 Python 可算,不需要浏览器。
|
||||
|
||||
常量与 `xhshow/config/config.py` 逐字节一致(该库也据此实现了 `sign_xyw`),
|
||||
并与独立的逆向实现 xiaohongshu-cli/creator_signing.py 互相印证。
|
||||
|
||||
**三条实测结论**(tools/probe_creator_api.py 的 Phase 0 输出):
|
||||
1. 待签字符串必须是 `url=` + 路径 + 查询串 的形式。只给路径、或去掉 `url=` 前缀,
|
||||
网关一律返回 **406**;写法正确时签名通过。
|
||||
2. `appId` 用 `ugc`(创作者平台的取值),不是主站的 `xhs-pc-web`。
|
||||
3. 不带 cookie 时返回的是应用层的 401「无登录信息」而非 406 —— 说明签名每次都过了,
|
||||
认证是独立的一层。
|
||||
"""
|
||||
|
||||
import base64
|
||||
import hashlib
|
||||
import json
|
||||
from datetime import datetime
|
||||
|
||||
XYW_AES_KEY = b"7cc4adla5ay0701v"
|
||||
XYW_AES_IV = b"4uzjr7mbsibcaldp"
|
||||
|
||||
# 与 xhshow 的 XYW_ENV_FLAGS_DEFAULT 一致。含义未知,但改了签名就不被接受。
|
||||
XYW_ENV_FLAGS = "0|0|0|1|0|0|1|0|0|0|1|0|0|0|0|1|0|0|0"
|
||||
|
||||
XYW_PREFIX = "XYW_"
|
||||
XYW_SIGN_SVN = "56"
|
||||
XYW_SIGN_TYPE = "x2"
|
||||
XYW_SIGN_VERSION = "1"
|
||||
|
||||
# 创作者平台的 appId。用主站的 xhs-pc-web 会被拒。
|
||||
CREATOR_APP_ID = "ugc"
|
||||
|
||||
|
||||
def _aes_encrypt_hex(plaintext: str) -> str:
|
||||
from Crypto.Cipher import AES
|
||||
from Crypto.Util.Padding import pad
|
||||
|
||||
cipher = AES.new(XYW_AES_KEY, AES.MODE_CBC, XYW_AES_IV)
|
||||
return cipher.encrypt(pad(plaintext.encode("utf-8"), AES.block_size)).hex()
|
||||
|
||||
|
||||
def sign_xyw(
|
||||
api: str,
|
||||
a1: str,
|
||||
app_id: str = CREATOR_APP_ID,
|
||||
body: dict | None = None,
|
||||
timestamp_ms: int | None = None,
|
||||
) -> dict[str, str]:
|
||||
"""为一次创作者后台请求生成 ``x-s`` / ``x-t`` 请求头。
|
||||
|
||||
``api`` 必须是待签的完整字符串:``url=`` 加路径,GET 请求还要带上查询串。
|
||||
POST 的 JSON body 追加在其后(紧凑分隔符、不转义非 ASCII),与参考实现一致。
|
||||
"""
|
||||
content = api
|
||||
if body is not None:
|
||||
content += json.dumps(body, separators=(",", ":"), ensure_ascii=False)
|
||||
|
||||
if timestamp_ms is None:
|
||||
timestamp_ms = int(datetime.now().timestamp() * 1000)
|
||||
|
||||
digest = hashlib.md5(content.encode("utf-8")).hexdigest()
|
||||
plaintext = f"x1={digest};x2={XYW_ENV_FLAGS};x3={a1};x4={timestamp_ms};"
|
||||
encoded = base64.b64encode(plaintext.encode("utf-8")).decode("utf-8")
|
||||
|
||||
envelope = {
|
||||
"signSvn": XYW_SIGN_SVN,
|
||||
"signType": XYW_SIGN_TYPE,
|
||||
"appId": app_id,
|
||||
"signVersion": XYW_SIGN_VERSION,
|
||||
"payload": _aes_encrypt_hex(encoded),
|
||||
}
|
||||
x_s = XYW_PREFIX + base64.b64encode(
|
||||
json.dumps(envelope, separators=(",", ":")).encode("utf-8")
|
||||
).decode("utf-8")
|
||||
|
||||
return {"x-s": x_s, "x-t": str(timestamp_ms)}
|
||||
|
||||
|
||||
def signed_api(url_path: str, query: str = "") -> str:
|
||||
"""把路径与查询串拼成待签字符串。
|
||||
|
||||
单独抽出来是因为这个格式**没有文档**,只能靠实测固定下来 —— 写错就是 406,
|
||||
而 406 的响应体 ``{"code":-1,"success":false}`` 完全看不出错在哪。
|
||||
"""
|
||||
return f"url={url_path}?{query}" if query else f"url={url_path}"
|
||||
+20
@@ -48,6 +48,7 @@ from .auth import ensure_initial_credential, require_auth
|
||||
from .routers import (
|
||||
auth_router,
|
||||
crawler_router,
|
||||
creator_router,
|
||||
data_router,
|
||||
monitor_router,
|
||||
settings_router,
|
||||
@@ -64,11 +65,23 @@ async def lifespan(_app: FastAPI):
|
||||
browser session the way the log broadcaster is -- a scheduled run has to
|
||||
happen whether or not anyone has the UI open.
|
||||
"""
|
||||
from .creator.login import shutdown as shutdown_creator_login
|
||||
from .monitor.db import dispose_engine, init_db
|
||||
from .monitor.qrlogin import shutdown as shutdown_qrlogin
|
||||
from .monitor.scheduler import monitor_scheduler
|
||||
|
||||
await init_db()
|
||||
|
||||
# The WebUI bundle is gitignored and built separately, so a deployment that
|
||||
# forgot it would otherwise come up looking healthy and serve a bare JSON
|
||||
# stub at "/" -- worth one loud line at boot rather than a puzzled operator.
|
||||
if not os.path.exists(os.path.join(WEBUI_DIR, "index.html")):
|
||||
print(
|
||||
"[综合采集平台] 警告:未找到前端产物 api/webui/index.html,"
|
||||
"根路径只会返回一段 JSON。请先在 webui/ 下执行 npm run build。",
|
||||
flush=True,
|
||||
)
|
||||
|
||||
generated = await ensure_initial_credential()
|
||||
if generated:
|
||||
# Printed once, on the run that creates it. There is no unauthenticated
|
||||
@@ -89,6 +102,12 @@ async def lifespan(_app: FastAPI):
|
||||
yield
|
||||
finally:
|
||||
await monitor_scheduler.stop()
|
||||
# Drops the tab a QR login may have opened and stops the Playwright
|
||||
# client; leaving them would strand a driver process on every restart.
|
||||
await shutdown_qrlogin()
|
||||
# Same for the operator's account logins, which run in throwaway browser
|
||||
# contexts -- those would otherwise be left open in the operator's Chrome.
|
||||
await shutdown_creator_login()
|
||||
await dispose_engine()
|
||||
|
||||
|
||||
@@ -136,6 +155,7 @@ app.add_middleware(
|
||||
# more importantly, never sees WebSocket scopes at all.
|
||||
app.include_router(auth_router, prefix="/api")
|
||||
app.include_router(crawler_router, prefix="/api", dependencies=[Depends(require_auth)])
|
||||
app.include_router(creator_router, prefix="/api", dependencies=[Depends(require_auth)])
|
||||
app.include_router(data_router, prefix="/api", dependencies=[Depends(require_auth)])
|
||||
app.include_router(monitor_router, prefix="/api", dependencies=[Depends(require_auth)])
|
||||
app.include_router(settings_router, prefix="/api", dependencies=[Depends(require_auth)])
|
||||
|
||||
@@ -0,0 +1,249 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/adapters.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""平台适配:两个平台之间**不一样**的那些管子。
|
||||
|
||||
监控层的大部分是平台中立的 —— 调度、入库、差分、报表、封面缓存、通知发送都与平台无关。
|
||||
真正随平台变化的只有四样东西:
|
||||
|
||||
1. 爬虫把产物**落在哪个目录**(这里有个坑,见 ``artifact_dir``)
|
||||
2. jsonl 里**字段叫什么**(抖音的作品没有 ``note_id``,叫 ``aweme_id``)
|
||||
3. **目标链接**长什么样(怎么拼、怎么从链接里抠出 id)
|
||||
4. 通知里的作品链接怎么拼
|
||||
|
||||
集中在这里,是为了让「加一个平台」变成在一处补一份数据,而不是去五个文件里找硬编码。
|
||||
|
||||
**为什么不放进 platforms.py**:那个模块被 ``describe_all()`` 整个序列化进
|
||||
``GET /api/config/platforms`` 交给前端(连 ``**capability`` 一起),把正则、目录名、字段别名
|
||||
塞进去会让爬虫的内部细节漏进 API 载荷,也会让「改适配」有动到接口形状的风险。
|
||||
分工与既有的 schedule.py(算术)↔ scheduler.py(循环)一致。
|
||||
"""
|
||||
|
||||
import re
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Dict, Mapping, Optional, Pattern, Tuple
|
||||
|
||||
from .platforms import PLATFORM_XHS
|
||||
|
||||
PLATFORM_DY = "dy"
|
||||
|
||||
|
||||
def _first_cover(record: Dict[str, Any], fields: Tuple[str, ...]) -> str:
|
||||
"""封面地址:取第一个非空字段,再取逗号分隔的第一段。
|
||||
|
||||
一条规则同时适配两边,所以不需要 per-platform 的函数:
|
||||
小红书的 ``image_list`` 是 ``"url1,url2,..."``(要切第一段),
|
||||
抖音的 ``cover_url`` 本身就是单个地址(切了等于没切)。
|
||||
"""
|
||||
for name in fields:
|
||||
raw = record.get(name)
|
||||
if raw:
|
||||
return str(raw).split(",")[0].strip()
|
||||
return ""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PlatformAdapter:
|
||||
"""一个平台的全部「管子」。
|
||||
|
||||
字段别名的方向是**规范名 -> 该平台 jsonl 里的键**,读作「我们的列 ← 他们的键」。
|
||||
"""
|
||||
|
||||
# 爬虫落盘用的目录名。**不等于平台 id**:抖音的平台 id 是 ``dy`` 而目录是 ``douyin``。
|
||||
# 这不是笔误,是上游 store 里写死的(store/douyin/_store_impl.py:47)。改错这里的
|
||||
# 后果是 ingest 一个文件都找不到 —— 它不会报错,只会落进「没抓到数据」分支,
|
||||
# 然后被误报成「疑似登录失效」。
|
||||
artifact_dir: str
|
||||
|
||||
web_base: str
|
||||
creator_path: str
|
||||
note_path: str
|
||||
|
||||
# 从链接里抠 id。是元组而不是单个正则,因为同一个平台可能有多种链接形态
|
||||
# (抖音的作品链接还带 ?modal_id= 那种),按顺序试,第一个匹配的胜出。
|
||||
# 每个正则必须恰好有一个捕获组。
|
||||
creator_url_res: Tuple[Pattern, ...]
|
||||
note_url_res: Tuple[Pattern, ...]
|
||||
# 也允许直接粘贴裸 id —— 但两边的 id 形状不同,所以分开。
|
||||
creator_bare_re: Pattern
|
||||
note_bare_re: Pattern
|
||||
# 短链(v.douyin.com 这种)无法在不发请求的情况下还原出 id,解析时单独报错,
|
||||
# 好过存一个聚不出目标的值进去。
|
||||
short_link_hosts: Tuple[str, ...]
|
||||
|
||||
note_fields: Mapping[str, str]
|
||||
comment_fields: Mapping[str, str]
|
||||
cover_fields: Tuple[str, ...]
|
||||
|
||||
# 时间戳换算成毫秒要乘的数。**小红书给毫秒、抖音给秒**,差 1000 倍;不换算的话
|
||||
# 2026 年的作品会显示成 1970 年(实测踩到过:抖音作品发布日期显示 1970-01-22,
|
||||
# 抖音评论的时间同理)。库里统一存毫秒,展示层才不用关心来源。
|
||||
time_scale: int
|
||||
|
||||
def to_ms(self, value: Any) -> Optional[int]:
|
||||
"""把平台的时间戳换算成毫秒;解析不出来返回 None(不伪造 0)。"""
|
||||
try:
|
||||
return int(value) * self.time_scale
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
|
||||
def note_url(self, note_id: str) -> str:
|
||||
"""作品的可点击链接。拼法与监控目标的链接是同一个形状 —— 通知里给的就是
|
||||
人能直接点开看的那一个。"""
|
||||
return f"{self.web_base}{self.note_path}/{note_id}"
|
||||
|
||||
def note_field(self, record: Dict[str, Any], name: str) -> Any:
|
||||
"""按规范名读作品记录里的原始值(没有就是 None)。"""
|
||||
return record.get(self.note_fields.get(name, name))
|
||||
|
||||
def comment_field(self, record: Dict[str, Any], name: str) -> Any:
|
||||
return record.get(self.comment_fields.get(name, name))
|
||||
|
||||
def cover(self, record: Dict[str, Any]) -> str:
|
||||
return _first_cover(record, self.cover_fields)
|
||||
|
||||
def parent_comment_id(self, record: Dict[str, Any]) -> str:
|
||||
"""父评论 id,顶层评论一律归一成空串。
|
||||
|
||||
抖音顶层评论的 ``reply_id`` 是字符串 ``"0"``,小红书是 ``""`` —— 把 "0" 原样
|
||||
存进去,前端就会多出一堆指向不存在的父评论的边。
|
||||
"""
|
||||
raw = self.comment_field(record, "parent_comment_id")
|
||||
if raw is None:
|
||||
return ""
|
||||
raw = str(raw).strip()
|
||||
return "" if raw in ("", "0") else raw
|
||||
|
||||
|
||||
# 小红书 id 是 24 位 hex,允许稍宽一点,让格式变化退化成「仍然接受」而不是「拒绝」。
|
||||
_XHS_BARE_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
|
||||
|
||||
XHS = PlatformAdapter(
|
||||
artifact_dir="xhs",
|
||||
web_base="https://www.xiaohongshu.com",
|
||||
creator_path="/user/profile",
|
||||
note_path="/explore",
|
||||
creator_url_res=(re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)"),),
|
||||
note_url_res=(
|
||||
re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)"),
|
||||
),
|
||||
creator_bare_re=_XHS_BARE_RE,
|
||||
note_bare_re=_XHS_BARE_RE,
|
||||
short_link_hosts=(),
|
||||
note_fields={
|
||||
"note_id": "note_id",
|
||||
"title": "title",
|
||||
"note_url": "note_url",
|
||||
"creator_hash": "creator_hash",
|
||||
"creator_name": "nickname",
|
||||
"source_kind": "type",
|
||||
"published_at": "time",
|
||||
},
|
||||
comment_fields={
|
||||
"comment_id": "comment_id",
|
||||
"note_id": "note_id",
|
||||
"content": "content",
|
||||
"creator_hash": "creator_hash",
|
||||
"creator_name": "nickname",
|
||||
"create_time": "create_time",
|
||||
"like_count": "like_count",
|
||||
"sub_comment_count": "sub_comment_count",
|
||||
"parent_comment_id": "parent_comment_id",
|
||||
},
|
||||
cover_fields=("image_list",),
|
||||
# 小红书的时间戳本来就是毫秒(实测 time=1790923011000)。
|
||||
time_scale=1,
|
||||
)
|
||||
|
||||
# 抖音的 id 形状与小红书完全不同(见 media_platform/douyin/help.py:101-164):
|
||||
# 作品 aweme_id 纯数字,如 7525082444551310602
|
||||
# 博主 sec_user_id 形如 MS4wLjABAAAA...,含 - 和 _,**变长**(实测样本 55 字符,更长的也常见),
|
||||
# 而小红书那条裸 id 规则封顶 64 —— 所以两条规则必须分开,否则长一点的 sec_uid
|
||||
# 会被拒,表现为「粘贴了一个完全正确的链接却说无法识别」。
|
||||
# 另外抖音**不需要 xsec_token**,裸链接就能用,比小红书简单。
|
||||
DY = PlatformAdapter(
|
||||
artifact_dir="douyin",
|
||||
web_base="https://www.douyin.com",
|
||||
creator_path="/user",
|
||||
note_path="/video",
|
||||
creator_url_res=(re.compile(r"douyin\.com/user/([A-Za-z0-9_-]+)"),),
|
||||
note_url_res=(
|
||||
re.compile(r"douyin\.com/video/(\d+)"),
|
||||
# 带 modal_id 的链接:在别人主页或搜索结果里点开视频就是这个形态。
|
||||
re.compile(r"[?&]modal_id=(\d+)"),
|
||||
),
|
||||
# 用长度而不是前缀来区分两者:sec_uid 是 20 字符以上的变长串,作品 id 是 19 位数字。
|
||||
# 用前缀(MS4wLjABAAAA)更精确,但上游的 parse_creator_info_from_url 对裸 id 一律
|
||||
# 照单全收,万一有别的前缀就会被我这里挡掉 —— 门槛设在长度上,两边都放得进,
|
||||
# 又不会把 19 位的作品号误当成博主。
|
||||
creator_bare_re=re.compile(r"^[A-Za-z0-9_-]{20,128}$"),
|
||||
note_bare_re=re.compile(r"^\d{8,25}$"),
|
||||
short_link_hosts=("v.douyin.com",),
|
||||
note_fields={
|
||||
"note_id": "aweme_id",
|
||||
"title": "title",
|
||||
"note_url": "aweme_url",
|
||||
"creator_hash": "creator_hash",
|
||||
"creator_name": "nickname",
|
||||
"source_kind": "aweme_type",
|
||||
"published_at": "create_time",
|
||||
},
|
||||
comment_fields={
|
||||
"comment_id": "comment_id",
|
||||
"note_id": "aweme_id",
|
||||
"content": "content",
|
||||
"creator_hash": "creator_hash",
|
||||
"creator_name": "nickname",
|
||||
"create_time": "create_time",
|
||||
"like_count": "like_count",
|
||||
"sub_comment_count": "sub_comment_count",
|
||||
"parent_comment_id": "parent_comment_id",
|
||||
},
|
||||
cover_fields=("cover_url",),
|
||||
# 抖音给的是**秒**(实测 create_time=1790574515,即 2026-09-28)。
|
||||
time_scale=1000,
|
||||
)
|
||||
|
||||
ADAPTERS: Dict[str, PlatformAdapter] = {
|
||||
PLATFORM_XHS: XHS,
|
||||
PLATFORM_DY: DY,
|
||||
}
|
||||
|
||||
|
||||
class UnknownPlatformError(ValueError):
|
||||
"""平台还没有适配器。"""
|
||||
|
||||
|
||||
def adapter(platform: str) -> PlatformAdapter:
|
||||
try:
|
||||
return ADAPTERS[platform]
|
||||
except KeyError as exc:
|
||||
raise UnknownPlatformError(f"平台 {platform} 还没有适配器") from exc
|
||||
|
||||
|
||||
def has_adapter(platform: str) -> bool:
|
||||
return platform in ADAPTERS
|
||||
|
||||
|
||||
def artifact_dir(platform: str) -> str:
|
||||
"""该平台的爬虫会把 jsonl 落在哪个子目录下。
|
||||
|
||||
``runner`` 用它判断产物是否真的出现过,``ingest`` 用它定位文件 —— 两处必须用
|
||||
同一个值,否则会出现「文件在,但两边找的目录不是同一个」这种最难查的错。
|
||||
"""
|
||||
return adapter(platform).artifact_dir
|
||||
@@ -53,6 +53,7 @@ from .settings import (
|
||||
set_setting,
|
||||
system_key,
|
||||
)
|
||||
from .upstream import DEFAULT_BRANCH, DEFAULT_REMOTE_URL
|
||||
|
||||
SCOPE_PLATFORM = "platform"
|
||||
SCOPE_SYSTEM = "system"
|
||||
@@ -204,6 +205,73 @@ SETTING_SPECS: List[SettingSpec] = [
|
||||
maximum=23,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="cdp_enabled",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_BOOL,
|
||||
label="接管已有 Chrome(CDP)",
|
||||
help=(
|
||||
"开启后爬虫不再自己启动浏览器,而是接管本机已开放远程调试端口的 Chrome"
|
||||
"(默认 127.0.0.1:9222),复用它的登录态与扩展。"
|
||||
"服务器部署请开启;本机桌面使用请保持关闭。"
|
||||
),
|
||||
default=False,
|
||||
),
|
||||
# --- 上游更新检查 -------------------------------------------------------
|
||||
# 这几项不作用于采集,所以都标 affects_new_runs=False:改动它们不需要等下一轮,
|
||||
# 也不影响采集命令的拼装。
|
||||
SettingSpec(
|
||||
name="upstream_check_enabled",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_BOOL,
|
||||
label="检查上游仓库更新",
|
||||
help=(
|
||||
"定期 fetch 上游仓库,看看它有没有新提交,并在有更新时推送通知。"
|
||||
"本仓库在上游之上加了一整层(见 UPSTREAM.md),不定期看一眼就会越拖越难合并。"
|
||||
),
|
||||
default=False,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="upstream_check_interval_minutes",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_INT,
|
||||
label="上游检查间隔(分钟)",
|
||||
help="默认 1440 分钟(每天一次)。检查只是 fetch,不需要太频繁。",
|
||||
default=1440,
|
||||
minimum=30,
|
||||
maximum=10080,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="upstream_remote_url",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_STR,
|
||||
label="上游仓库地址",
|
||||
help=(
|
||||
"默认是 GitHub 上的上游。国内直连 GitHub 不稳时改成 gitcode 镜像"
|
||||
"(见 UPSTREAM.md),或任意能访问到上游的地址。"
|
||||
),
|
||||
default=DEFAULT_REMOTE_URL,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="upstream_branch",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_STR,
|
||||
label="上游分支",
|
||||
default=DEFAULT_BRANCH,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="upstream_notify",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_BOOL,
|
||||
label="上游有更新时推送通知",
|
||||
help="只在出现此前没推过的上游提交时发一条,同一个更新不会反复推。",
|
||||
default=True,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
]
|
||||
|
||||
SPECS_BY_NAME = {spec.name: spec for spec in SETTING_SPECS}
|
||||
|
||||
@@ -0,0 +1,162 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/covers.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""作品封面本地缓存。
|
||||
|
||||
**为什么必须落盘**:小红书图床的地址是**带签名、会过期**的。路径里那段时间戳就是
|
||||
签发时刻,实测:
|
||||
|
||||
/202610080841/... (当天签发) → 200,且带不带 Referer 都 200
|
||||
/202610070837/... (隔天) → 403,且带不带 Referer 都 403
|
||||
|
||||
所以这是**过期**,不是防盗链 —— 改 Referer 那一类修法治不了本。图一旦下载到本地,
|
||||
就与签名无关,永远可读。
|
||||
|
||||
下载失败**不能影响采集**:一张封面拿不到,不该让整轮数据丢失。
|
||||
"""
|
||||
|
||||
import re
|
||||
from pathlib import Path
|
||||
from typing import Optional
|
||||
|
||||
import httpx
|
||||
|
||||
from .db import DATA_DIR
|
||||
|
||||
COVERS_DIR = DATA_DIR / "covers"
|
||||
|
||||
# 单张封面的上限。正常封面是几十 KB;超过这个数说明拿到的不是图,
|
||||
# 或者该放弃这一张而不是把内存撑爆。
|
||||
MAX_COVER_BYTES = 5 * 1024 * 1024
|
||||
|
||||
USER_AGENT = (
|
||||
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
|
||||
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
|
||||
)
|
||||
|
||||
_EXTENSIONS = {
|
||||
"image/jpeg": ".jpg",
|
||||
"image/jpg": ".jpg",
|
||||
"image/png": ".png",
|
||||
"image/webp": ".webp",
|
||||
"image/gif": ".gif",
|
||||
"image/heic": ".heic",
|
||||
}
|
||||
|
||||
# note_id 是平台的稳定标识,但仍要挡住路径穿越 —— 它会直接变成文件名。
|
||||
_SAFE_ID = re.compile(r"^[A-Za-z0-9_-]{1,64}$")
|
||||
|
||||
|
||||
def is_safe_note_id(note_id: str) -> bool:
|
||||
return bool(note_id) and bool(_SAFE_ID.match(note_id))
|
||||
|
||||
|
||||
def cache_dir() -> Path:
|
||||
COVERS_DIR.mkdir(parents=True, exist_ok=True)
|
||||
return COVERS_DIR
|
||||
|
||||
|
||||
def find_cached(note_id: str) -> Optional[Path]:
|
||||
"""已缓存的封面文件,没有则 None。扩展名按内容类型而定,所以逐一试。"""
|
||||
if not is_safe_note_id(note_id):
|
||||
return None
|
||||
for extension in sorted(set(_EXTENSIONS.values())):
|
||||
candidate = COVERS_DIR / f"{note_id}{extension}"
|
||||
if candidate.is_file():
|
||||
return candidate
|
||||
return None
|
||||
|
||||
|
||||
async def cache_cover(note_id: str, url: str) -> Optional[str]:
|
||||
"""下载并保存一张封面,返回文件名;失败返回 None。
|
||||
|
||||
**从不抛异常**:调用方是采集入库流程,一张图拿不到不该让整轮数据出问题。
|
||||
"""
|
||||
if not url or not is_safe_note_id(note_id):
|
||||
return None
|
||||
|
||||
existing = find_cached(note_id)
|
||||
if existing is not None:
|
||||
return existing.name
|
||||
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=20, follow_redirects=True) as client:
|
||||
response = await client.get(url, headers={"user-agent": USER_AGENT})
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
if response.status_code != 200:
|
||||
# 403 通常意味着签名已过期 —— 这一张就没了,等下一轮采集拿到新地址。
|
||||
return None
|
||||
|
||||
content = response.content
|
||||
if not content or len(content) > MAX_COVER_BYTES:
|
||||
return None
|
||||
|
||||
content_type = (response.headers.get("content-type") or "").split(";")[0].strip().lower()
|
||||
extension = _EXTENSIONS.get(content_type, ".jpg")
|
||||
# 图床偶尔不报 content-type,那种情况下扩展名只能猜,但文件本身仍然是好的。
|
||||
target = cache_dir() / f"{note_id}{extension}"
|
||||
|
||||
try:
|
||||
target.write_bytes(content)
|
||||
except OSError:
|
||||
return None
|
||||
return target.name
|
||||
|
||||
|
||||
def cover_url(note_id: str, remote: str) -> str:
|
||||
"""前端该用哪个地址。
|
||||
|
||||
本地有缓存就用自己的接口 —— 那是唯一不会过期的地址。没有就退回远程地址,
|
||||
至少让图先显示出来(哪怕它很快会失效)。
|
||||
"""
|
||||
if find_cached(note_id) is not None:
|
||||
return f"/api/monitor/covers/{note_id}"
|
||||
return remote
|
||||
|
||||
|
||||
async def cache_pending(session, task_id: int, limit: int = 60) -> int:
|
||||
"""把还没有本地副本的封面补下来,返回本次下载成功的张数。
|
||||
|
||||
由 runner 在入库之后调用,**而不是在 ingest 里** —— ingest 是刻意保持离线的
|
||||
(它的文档写明 No network),往里塞网络请求会毁掉这一点。
|
||||
|
||||
每轮只补一批:一次跑几百张图既慢又会给图床压力,而旧地址本来就在陆续过期,
|
||||
分摊到几轮里补完反而更稳。
|
||||
"""
|
||||
from sqlalchemy import select
|
||||
|
||||
from .models import MonitorNote
|
||||
|
||||
notes = (
|
||||
await session.scalars(
|
||||
select(MonitorNote)
|
||||
.where(MonitorNote.task_id == task_id, MonitorNote.cover != "")
|
||||
.order_by(MonitorNote.last_seen_at.desc())
|
||||
.limit(limit)
|
||||
)
|
||||
).all()
|
||||
|
||||
saved = 0
|
||||
for note in notes:
|
||||
if find_cached(note.note_id) is not None:
|
||||
continue
|
||||
if await cache_cover(note.note_id, note.cover):
|
||||
saved += 1
|
||||
return saved
|
||||
+63
-14
@@ -44,7 +44,9 @@ from contextlib import asynccontextmanager
|
||||
from pathlib import Path
|
||||
from typing import AsyncIterator, Optional
|
||||
|
||||
from sqlalchemy import event, text
|
||||
from sqlalchemy import Column, event, text
|
||||
from sqlalchemy.dialects import mysql
|
||||
from sqlalchemy.schema import CreateColumn
|
||||
from sqlalchemy.ext.asyncio import (
|
||||
AsyncEngine,
|
||||
AsyncSession,
|
||||
@@ -214,14 +216,53 @@ async def init_db() -> None:
|
||||
await _migrate_setting_keys(conn)
|
||||
|
||||
|
||||
# Columns added to a table after it may already exist. ``create_all`` only
|
||||
# creates missing *tables*, so new columns need an explicit ALTER TABLE.
|
||||
_ADDED_COLUMNS: dict[str, list[tuple[str, str]]] = {
|
||||
"monitor_task": [
|
||||
("notify_enabled", "BOOLEAN NOT NULL DEFAULT 0"),
|
||||
("last_notified_at", "BIGINT NULL"),
|
||||
],
|
||||
}
|
||||
# ``create_all`` creates missing *tables* but never adds *columns* to a table that
|
||||
# already exists, so those need an explicit ALTER TABLE.
|
||||
#
|
||||
# Which columns those are is derived from the ORM metadata, not kept by hand. The
|
||||
# hand-kept version was a trap: forgetting to register a new column there still let
|
||||
# the app start -- it connects fine, then fails on every query and every scheduler
|
||||
# tick. Which is exactly what happened when the scheduling columns were added.
|
||||
|
||||
|
||||
def _implicit_default(column: Column) -> Optional[str]:
|
||||
"""A literal to seed existing rows with when a NOT NULL column is added."""
|
||||
default = column.default
|
||||
if default is not None and getattr(default, "is_scalar", False):
|
||||
value = default.arg
|
||||
if isinstance(value, bool):
|
||||
return "1" if value else "0"
|
||||
if isinstance(value, (int, float)):
|
||||
return str(value)
|
||||
return "'" + str(value).replace("'", "''") + "'"
|
||||
|
||||
# No scalar default on the model. Fall back to the type's zero value, so that
|
||||
# adding the column cannot depend on the server's sql_mode.
|
||||
try:
|
||||
python_type = column.type.python_type
|
||||
except NotImplementedError:
|
||||
return None
|
||||
if python_type in (bool, int, float):
|
||||
return "0"
|
||||
if python_type is str:
|
||||
return "''"
|
||||
return None
|
||||
|
||||
|
||||
def _column_ddl(column: Column) -> str:
|
||||
"""One column as MySQL DDL for ``ALTER TABLE ... ADD COLUMN``.
|
||||
|
||||
``CreateColumn`` renders the name, type and nullability. The default is added
|
||||
separately because a model's ``default=`` is applied by the ORM and never
|
||||
reaches the DDL -- and a NOT NULL column added to a populated table needs a
|
||||
value for the rows already sitting there.
|
||||
"""
|
||||
ddl = str(CreateColumn(column).compile(dialect=mysql.dialect()))
|
||||
if not column.nullable and column.server_default is None:
|
||||
seed = _implicit_default(column)
|
||||
if seed is not None:
|
||||
ddl += f" DEFAULT {seed}"
|
||||
return ddl
|
||||
|
||||
|
||||
async def _existing_columns(conn, table: str) -> set[str]:
|
||||
@@ -240,14 +281,22 @@ async def _existing_columns(conn, table: str) -> set[str]:
|
||||
|
||||
|
||||
async def _ensure_columns(conn) -> None:
|
||||
for table, columns in _ADDED_COLUMNS.items():
|
||||
existing = await _existing_columns(conn, table)
|
||||
"""Add every model column the live table is missing."""
|
||||
for table in MonitorBase.metadata.sorted_tables:
|
||||
existing = await _existing_columns(conn, table.name)
|
||||
if not existing:
|
||||
# Table did not exist before this run; create_all built it complete.
|
||||
continue
|
||||
for name, ddl in columns:
|
||||
if name not in existing:
|
||||
await conn.execute(text(f"ALTER TABLE {table} ADD COLUMN {name} {ddl}"))
|
||||
for column in table.columns:
|
||||
# Primary keys are always present, and MySQL rejects AUTO_INCREMENT
|
||||
# alongside the DEFAULT this helper appends -- so skip them rather
|
||||
# than emit DDL that could never run.
|
||||
if column.name in existing or column.primary_key:
|
||||
continue
|
||||
print(f"[monitor.db] 补齐缺失字段 {table.name}.{column.name}", flush=True)
|
||||
await conn.execute(
|
||||
text(f"ALTER TABLE {table.name} ADD COLUMN {_column_ddl(column)}")
|
||||
)
|
||||
|
||||
|
||||
async def _migrate_setting_keys(conn) -> None:
|
||||
|
||||
@@ -0,0 +1,561 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_api.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""抖音 Web 接口客户端 —— 直接发 HTTP,不起爬虫子进程。
|
||||
|
||||
**为什么另起一套。** 爬虫那条路(``media_platform/douyin``)会构造一大串浏览器指纹
|
||||
参数:``browser_platform=MacIntel``、``os_name=Mac OS``、``browser_version=125.0.0.0``……
|
||||
而 ``User-Agent`` 是从页面现读的(在服务器上是 Linux + Chrome 155)。参数说自己是 Mac,
|
||||
UA 说自己是 Linux —— 抖音网关对这种自相矛盾的请求的处理方式是:**不报错、不给原因,
|
||||
回一个 200 + 空 body**。爬虫那边把它翻译成 ``Exception("account blocked")``,看起来像
|
||||
账号被封,其实什么都不是。
|
||||
|
||||
这份客户端只发必要参数(``device_platform`` / ``aid`` 那两三个),走浏览器自己也在用的
|
||||
那条调用路径。它的做法来自 mac-agent-os 项目的 ``mediacrawler_adapter.py``,实测可用。
|
||||
|
||||
两个关键点:
|
||||
|
||||
* **cookie 走 CDP 现读。** Chrome 把 cookie 值加密存在 SQLite 里,只有 CDP 拿得到
|
||||
解密后的值;而且浏览器里那份比库里存的旧快照新 —— 站点会自己轮换会话。
|
||||
* **产物形状照抄 store。** ``aweme_id`` / ``aweme_url`` / ``cover_url`` / ``aweme_type`` /
|
||||
``create_time``(**秒**,由 adapters 换算成毫秒)…… 这样 ingest 那条链路一个字都不用改。
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import os
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Dict, List, Optional, Tuple
|
||||
from urllib.parse import urlencode
|
||||
|
||||
import config
|
||||
import httpx
|
||||
from tools import utils
|
||||
from tools.user_hash import anonymize_user_id
|
||||
|
||||
# 请求头。**要像一个浏览器**,而且必须是**同一个浏览器**:见 BrowserIdentity。
|
||||
_BASE_HEADERS = {
|
||||
"Accept": "application/json, text/plain, */*",
|
||||
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
|
||||
"Referer": "https://www.douyin.com/",
|
||||
"Origin": "https://www.douyin.com",
|
||||
}
|
||||
|
||||
# 网关的业务前置校验头。缺了它,抖音边缘网关的 ArgusSecurityPlugin 会直接回
|
||||
# 403 并写明 "Blocked by ArgusSecurityPlugin Uifid Not Found" —— 难得一次它会说原因。
|
||||
# 当前网关并不校验这个头的**值**,填什么都行;一旦升级到真校验,就得改成让页面里的
|
||||
# SDK 自己生成(见 media_platform/douyin/client.py 里同一条注释)。
|
||||
ARGUS_HEADER_VALUE = "1"
|
||||
|
||||
API_ORIGIN = "https://www.douyin.com"
|
||||
PROFILE_PATH = "/aweme/v1/web/user/profile/other/"
|
||||
POSTS_PATH = "/aweme/v1/web/aweme/post/"
|
||||
DETAIL_PATH = "/aweme/v1/web/aweme/detail/"
|
||||
COMMENT_PATH = "/aweme/v1/web/comment/list/"
|
||||
|
||||
# 一次请求的超时。抖音这两个接口正常都在一秒内返回。
|
||||
REQUEST_TIMEOUT_SECONDS = 20.0
|
||||
|
||||
# 问浏览器要 UA / client hints 的超时。**这个必须有。**
|
||||
# ``page.evaluate`` 打在一个渲染进程已经卡住的标签页上会**永远不返回**,而问身份是采集的
|
||||
# 第一步 —— 它一挂,整个 run 就永远停在「运行中」(真踩过:标签页 URL 是空的,
|
||||
# cookies() 正常,evaluate 一直不回来)。
|
||||
EVALUATE_TIMEOUT_SECONDS = 8.0
|
||||
# 单页最多要多少条。接口自己有上限,要多了也没用。
|
||||
MAX_PAGE_SIZE = 20
|
||||
|
||||
|
||||
class DouyinApiError(RuntimeError):
|
||||
"""请求失败,或登录态不可用。"""
|
||||
|
||||
|
||||
def _cdp_url() -> str:
|
||||
"""浏览器 DevTools 端点。与扫码登录那边共用同一个开关。"""
|
||||
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
|
||||
|
||||
|
||||
@dataclass
|
||||
class BrowserIdentity:
|
||||
"""一个请求要像浏览器所需要的全部身份信息,**且必须来自同一个浏览器**。
|
||||
|
||||
只拿 cookie 是不够的。UA 声称自己是 Chrome 155、却不带 Chrome 155 该有的
|
||||
``sec-ch-ua``,网关一眼就能看出这不是浏览器 —— 它的回应是 **200 + 空 body**:
|
||||
不报错、不给原因,只看得到「抓到 0 条」。所以这三样必须成套地从同一处取。
|
||||
"""
|
||||
|
||||
cookie: str
|
||||
user_agent: str
|
||||
client_hints: Dict[str, str]
|
||||
|
||||
def headers(self) -> Dict[str, str]:
|
||||
headers = {
|
||||
"User-Agent": self.user_agent,
|
||||
**self.client_hints,
|
||||
**_BASE_HEADERS,
|
||||
"x-tt-argus": ARGUS_HEADER_VALUE,
|
||||
"Cookie": self.cookie,
|
||||
}
|
||||
# uifid 是设备标识,网关要它;cookie 里没有就不带(送空值反而更像异常请求)。
|
||||
uifid = _cookie_value(self.cookie, "UIFID") or _cookie_value(
|
||||
self.cookie, "UIFID_TEMP"
|
||||
)
|
||||
if uifid:
|
||||
headers["uifid"] = uifid
|
||||
return headers
|
||||
|
||||
|
||||
# 身份信息的短时缓存:一次采集要发好几个请求,没必要每次都连一遍 CDP。
|
||||
_IDENTITY_TTL_SECONDS = 120.0
|
||||
_identity_cache: Optional[Tuple[float, BrowserIdentity]] = None
|
||||
|
||||
|
||||
async def _safe_evaluate(page: Any, expression: str) -> Any:
|
||||
"""在页面上求值,带超时;任何失败都返回 None。
|
||||
|
||||
**不要直接调 ``page.evaluate``** —— 在渲染进程卡住的标签页上它会永远不返回(见
|
||||
``EVALUATE_TIMEOUT_SECONDS`` 那段)。
|
||||
"""
|
||||
try:
|
||||
return await asyncio.wait_for(
|
||||
page.evaluate(expression), timeout=EVALUATE_TIMEOUT_SECONDS
|
||||
)
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
async def _identity_from_pages(context: Any) -> Tuple[str, Dict[str, str]]:
|
||||
"""问出 UA 和 client hints。
|
||||
|
||||
不假设第一个标签页是好的 —— 它可能停在 URL 为空、渲染进程已卡住的状态(实测过)。
|
||||
所以逐个试、每个都带超时;优先抖音页面,全都不行就临时开一个干净页问完关掉。
|
||||
|
||||
拿不到就返回空 —— 调用方据此退回库里那份 cookie,而不是拿一组编出来的指纹去请求
|
||||
(那比没有更糟,见 BrowserIdentity 的说明)。
|
||||
"""
|
||||
from media_platform.douyin.help import client_hint_headers
|
||||
|
||||
pages = list(context.pages)
|
||||
pages.sort(key=lambda page: 0 if "douyin" in (page.url or "") else 1)
|
||||
for page in pages:
|
||||
user_agent = await _safe_evaluate(page, "() => navigator.userAgent")
|
||||
if user_agent:
|
||||
hints = client_hint_headers(
|
||||
await _safe_evaluate(page, "() => navigator.userAgentData || null")
|
||||
)
|
||||
return user_agent, hints or {}
|
||||
|
||||
temp = None
|
||||
try:
|
||||
temp = await asyncio.wait_for(
|
||||
context.new_page(), timeout=EVALUATE_TIMEOUT_SECONDS
|
||||
)
|
||||
user_agent = await _safe_evaluate(temp, "() => navigator.userAgent")
|
||||
hints = client_hint_headers(
|
||||
await _safe_evaluate(temp, "() => navigator.userAgentData || null")
|
||||
)
|
||||
return user_agent or "", hints or {}
|
||||
except Exception:
|
||||
return "", {}
|
||||
finally:
|
||||
if temp is not None:
|
||||
try:
|
||||
await temp.close()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
async def _read_browser() -> Optional[BrowserIdentity]:
|
||||
"""连上 CDP 浏览器,一次取齐 cookie、UA、client hints。
|
||||
|
||||
读不到返回 None(浏览器没开/没登录),由调用方决定怎么报 —— 不抛异常。
|
||||
"""
|
||||
from playwright.async_api import async_playwright
|
||||
|
||||
from media_platform.douyin.help import client_hint_headers
|
||||
|
||||
playwright = None
|
||||
try:
|
||||
playwright = await async_playwright().start()
|
||||
browser = await playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
|
||||
if not browser.contexts:
|
||||
return None
|
||||
# contexts[0] 是真实 profile。**不要 new_context()** —— 那是无痕式的,读不到登录态。
|
||||
context = browser.contexts[0]
|
||||
cookies = await asyncio.wait_for(
|
||||
context.cookies(), timeout=EVALUATE_TIMEOUT_SECONDS
|
||||
)
|
||||
|
||||
# UA 和 hints 要从页面里问 —— 它们是浏览器自己的事实,写死迟早对不上。
|
||||
user_agent, hints = await _identity_from_pages(context)
|
||||
except Exception as exc:
|
||||
utils.logger.warning(f"[douyin_api] 读浏览器身份失败:{exc}")
|
||||
return None
|
||||
finally:
|
||||
if playwright is not None:
|
||||
# 只断开连接。**绝不能 browser.close()** —— 对这个 CDP 连接而言那会关掉
|
||||
# 操作者自己的浏览器。
|
||||
try:
|
||||
await playwright.stop()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
douyin_cookies = {
|
||||
cookie["name"]: cookie["value"]
|
||||
for cookie in cookies
|
||||
if "douyin" in cookie.get("domain", "") or "amemv" in cookie.get("domain", "")
|
||||
}
|
||||
return BrowserIdentity(
|
||||
cookie=_cookie_from_dict(douyin_cookies),
|
||||
user_agent=user_agent or "",
|
||||
client_hints=hints,
|
||||
)
|
||||
|
||||
|
||||
async def browser_identity(cookie: str = "", force: bool = False) -> BrowserIdentity:
|
||||
"""拿到一份可用的身份:**优先浏览器里那份**,其次退回传进来的 cookie(库里存的)。
|
||||
|
||||
优先浏览器的原因:站点会自己轮换会话,库里存的是粘贴那一刻的快照,浏览器里那份才是
|
||||
当前有效的;而 UA/hints 更是只有浏览器自己知道。
|
||||
"""
|
||||
global _identity_cache
|
||||
|
||||
now = time.monotonic()
|
||||
if not force and _identity_cache is not None:
|
||||
cached_at, cached = _identity_cache
|
||||
if now - cached_at < _IDENTITY_TTL_SECONDS:
|
||||
return cached
|
||||
|
||||
identity = await _read_browser()
|
||||
if identity is None or not _has_session(identity.cookie):
|
||||
# 浏览器里没有可用会话,退回调用方给的那份。UA/hints 编不出来就不编 ——
|
||||
# 一组和 UA 对不上的 hints 比没有更糟。
|
||||
identity = BrowserIdentity(
|
||||
cookie=_cookie_header(cookie), user_agent="", client_hints={}
|
||||
)
|
||||
_identity_cache = (now, identity)
|
||||
return identity
|
||||
|
||||
|
||||
def forget_identity() -> None:
|
||||
"""丢掉缓存的身份。cookie 变了、或测试之间要隔离时调用。"""
|
||||
global _identity_cache
|
||||
_identity_cache = None
|
||||
|
||||
|
||||
def _cookie_header(cookie: str) -> str:
|
||||
"""把 ``a=1; b=2`` 形式的 cookie 串规整成请求头用的形状。"""
|
||||
pairs = []
|
||||
for part in (cookie or "").split(";"):
|
||||
if "=" in part:
|
||||
name, _, value = part.partition("=")
|
||||
name = name.strip()
|
||||
if name:
|
||||
pairs.append(f"{name}={value.strip()}")
|
||||
return "; ".join(pairs)
|
||||
|
||||
|
||||
def _cookie_from_dict(cookies: Dict[str, str]) -> str:
|
||||
return "; ".join(f"{name}={value}" for name, value in cookies.items())
|
||||
|
||||
|
||||
def _cookie_value(cookie: str, name: str) -> str:
|
||||
"""从一个 cookie 串里取某个键的值。"""
|
||||
for part in (cookie or "").split(";"):
|
||||
key, _, value = part.partition("=")
|
||||
if key.strip() == name:
|
||||
return value.strip()
|
||||
return ""
|
||||
|
||||
|
||||
def _sign(params: Dict[str, Any], path: str, user_agent: str) -> Dict[str, Any]:
|
||||
"""给一组参数补上 ``a_bogus`` 签名,返回新 dict。
|
||||
|
||||
**按需 import**:那个模块在 import 的那一瞬间就把 ``libs/douyin.js`` 交给 execjs
|
||||
编译(还要读相对路径),把它拖进监控层的热路径不合适。
|
||||
|
||||
签名算在**不含 a_bogus 的那串 query 上**,追加到末尾 —— 和爬虫那条路一致,也是
|
||||
实测能过的形态。
|
||||
"""
|
||||
from media_platform.douyin.help import get_a_bogus_from_js
|
||||
|
||||
try:
|
||||
return {
|
||||
**params,
|
||||
"a_bogus": get_a_bogus_from_js(path, urlencode(params), user_agent),
|
||||
}
|
||||
except Exception as exc: # execjs 起不来 / JS 抛错,都算签名失败
|
||||
raise DouyinApiError(f"算 a_bogus 签名失败:{exc}") from exc
|
||||
|
||||
|
||||
async def _get(
|
||||
path: str,
|
||||
params: Dict[str, Any],
|
||||
identity: BrowserIdentity,
|
||||
*,
|
||||
signed: bool = False,
|
||||
) -> Dict[str, Any]:
|
||||
"""发一个 GET,返回 JSON。
|
||||
|
||||
只带调用方给的参数 —— **不要往里加 webid / msToken / browser_version 那一堆**,
|
||||
那正是爬虫那条路失败的原因。
|
||||
|
||||
``signed=True`` 时补一个 ``a_bogus``。**只有评论接口需要它**:作品、详情、博主资料
|
||||
三个不带签名也照常返回,而给它们加签名是没验证过的改动,不做。
|
||||
"""
|
||||
if signed:
|
||||
params = _sign(params, path, identity.user_agent)
|
||||
|
||||
url = f"{API_ORIGIN}{path}"
|
||||
async with httpx.AsyncClient(timeout=REQUEST_TIMEOUT_SECONDS) as client:
|
||||
response = await client.get(
|
||||
url,
|
||||
params=params,
|
||||
headers=identity.headers(),
|
||||
)
|
||||
|
||||
if response.status_code != 200:
|
||||
raise DouyinApiError(f"HTTP {response.status_code}:{response.text[:120]}")
|
||||
|
||||
# 「200 + 空 body」是抖音网关拒绝请求时的典型回应(见模块说明)。必须当成错误报出来,
|
||||
# 否则会一路往下变成「这个博主没作品」。
|
||||
if not response.text.strip():
|
||||
raise DouyinApiError(
|
||||
"接口返回了空内容 —— 通常是登录态失效,或请求被网关判成了非浏览器"
|
||||
)
|
||||
|
||||
try:
|
||||
return response.json()
|
||||
except ValueError as exc:
|
||||
raise DouyinApiError(f"返回的不是 JSON:{response.text[:120]}") from exc
|
||||
|
||||
|
||||
def _as_int(value: Any) -> int:
|
||||
try:
|
||||
return int(value)
|
||||
except (TypeError, ValueError):
|
||||
return 0
|
||||
|
||||
|
||||
def normalize_aweme(aweme: Dict[str, Any]) -> Dict[str, Any]:
|
||||
"""把接口返回的一条作品,翻译成 store 落盘的那套键名。
|
||||
|
||||
键名必须和 ``store/douyin`` 一致 —— 跨过这一层之后,ingest 就不知道数据是从爬虫
|
||||
来的还是从接口来的。
|
||||
"""
|
||||
author = aweme.get("author") or {}
|
||||
statistics = aweme.get("statistics") or {}
|
||||
aweme_id = str(aweme.get("aweme_id") or "")
|
||||
cover = ((aweme.get("video") or {}).get("cover") or {}).get("url_list") or [""]
|
||||
uid = str(author.get("uid") or "")
|
||||
nickname = author.get("nickname") or ""
|
||||
|
||||
return {
|
||||
"aweme_id": aweme_id,
|
||||
"aweme_type": str(aweme.get("aweme_type") or ""),
|
||||
# store 那边 title 取的是 desc。
|
||||
"title": aweme.get("desc") or "",
|
||||
"desc": aweme.get("desc") or "",
|
||||
# **秒**。adapters.time_scale 会把它换成毫秒,和 store 写出来的形态一致。
|
||||
"create_time": _as_int(aweme.get("create_time")),
|
||||
"creator_hash": anonymize_user_id(uid or author.get("sec_uid") or ""),
|
||||
"nickname": nickname,
|
||||
"liked_count": str(_as_int(statistics.get("digg_count"))),
|
||||
"comment_count": str(_as_int(statistics.get("comment_count"))),
|
||||
"collected_count": str(_as_int(statistics.get("collect_count"))),
|
||||
"share_count": str(_as_int(statistics.get("share_count"))),
|
||||
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
|
||||
"cover_url": cover[0] if cover else "",
|
||||
"source_keyword": "",
|
||||
}
|
||||
|
||||
|
||||
async def author_videos(
|
||||
sec_user_id: str, count: int = MAX_PAGE_SIZE, *, cookie: str = ""
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""某个博主最新发布的作品(按发布时间倒序),已翻译成 store 的键名。
|
||||
|
||||
用 ``sec_user_id`` 而不是数字 uid:监控任务里存的就是主页链接里的那段 sec_uid,
|
||||
而且这个接口两种都收(爬虫那边用的也是 sec_user_id)。
|
||||
"""
|
||||
identity = await browser_identity(cookie)
|
||||
if not _has_session(identity.cookie):
|
||||
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
|
||||
|
||||
payload = await _get(
|
||||
POSTS_PATH,
|
||||
{
|
||||
"sec_user_id": sec_user_id,
|
||||
"count": max(1, min(count, MAX_PAGE_SIZE)),
|
||||
"max_cursor": 0,
|
||||
"device_platform": "webapp",
|
||||
"aid": 6383,
|
||||
},
|
||||
identity,
|
||||
)
|
||||
|
||||
awemes = payload.get("aweme_list") or []
|
||||
if not awemes and payload.get("status_code") not in (0, None):
|
||||
raise DouyinApiError(
|
||||
f"接口拒绝了请求(status_code={payload.get('status_code')})"
|
||||
)
|
||||
return [normalize_aweme(aweme) for aweme in awemes]
|
||||
|
||||
|
||||
async def video_detail(aweme_id: str, *, cookie: str = "") -> Dict[str, Any]:
|
||||
"""单条作品的详情,已翻译成 store 的键名。
|
||||
|
||||
这个接口**没有**被那道真校验挡着(实测 200 / 45425 字节),所以在拿不到作品列表时,
|
||||
它是「刷新已知作品指标」的唯一途径。
|
||||
"""
|
||||
identity = await browser_identity(cookie)
|
||||
if not _has_session(identity.cookie):
|
||||
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
|
||||
|
||||
payload = await _get(
|
||||
DETAIL_PATH,
|
||||
{"aweme_id": aweme_id, "device_platform": "webapp", "aid": 6383},
|
||||
identity,
|
||||
)
|
||||
aweme = payload.get("aweme_detail") or {}
|
||||
if not aweme:
|
||||
raise DouyinApiError(
|
||||
f"接口没返回作品(status_code={payload.get('status_code')})"
|
||||
)
|
||||
return normalize_aweme(aweme)
|
||||
|
||||
|
||||
async def author_profile(sec_user_id: str, *, cookie: str = "") -> Dict[str, Any]:
|
||||
"""博主主页指标:昵称 / 粉丝数 / 总获赞 / 作品数。"""
|
||||
identity = await browser_identity(cookie)
|
||||
if not _has_session(identity.cookie):
|
||||
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
|
||||
|
||||
payload = await _get(
|
||||
PROFILE_PATH,
|
||||
{"sec_user_id": sec_user_id, "device_platform": "webapp", "aid": 6383},
|
||||
identity,
|
||||
)
|
||||
user = payload.get("user") or {}
|
||||
if not user:
|
||||
raise DouyinApiError(
|
||||
f"接口没返回用户数据(status_code={payload.get('status_code')})"
|
||||
)
|
||||
return {
|
||||
# 自报家门。快照表的唯一键是 (任务, creator_hash, 轮次),而作品是靠
|
||||
# `anonymize_user_id(author.uid)` 得到这个哈希的 —— 这里走同一条路,两边才对得上,
|
||||
# 否则快照会和作品分成两个人,界面上永远查不到。
|
||||
"creator_hash": anonymize_user_id(
|
||||
str(user.get("uid") or user.get("sec_uid") or "")
|
||||
),
|
||||
"nickname": user.get("nickname") or "",
|
||||
"unique_id": user.get("unique_id") or "",
|
||||
"fans": _as_int(user.get("follower_count")),
|
||||
"total_favorited": _as_int(user.get("total_favorited")),
|
||||
"works": _as_int(user.get("aweme_count")),
|
||||
"following": _as_int(user.get("following_count")),
|
||||
}
|
||||
|
||||
|
||||
def normalize_comment(comment: Dict[str, Any], aweme_id: str) -> Dict[str, Any]:
|
||||
"""把接口返回的一条评论,翻译成 store 落盘的那套键名。
|
||||
|
||||
HTTP 路线和页面路线共用它 —— 同一套键名,ingest 才不用关心数据是怎么来的。
|
||||
刻意**不带** ``sub_comment_count`` / ``parent_comment_id`` 的猜测值:接口给了就用,
|
||||
没给就留空,不编。
|
||||
"""
|
||||
user = comment.get("user") or {}
|
||||
return {
|
||||
"comment_id": str(comment.get("cid") or ""),
|
||||
"aweme_id": aweme_id,
|
||||
"content": comment.get("text") or "",
|
||||
"nickname": user.get("nickname") or "",
|
||||
"creator_hash": anonymize_user_id(
|
||||
str(user.get("uid") or user.get("sec_uid") or "")
|
||||
),
|
||||
# 同为秒;adapters 会换算。
|
||||
"create_time": _as_int(comment.get("create_time")),
|
||||
"like_count": str(_as_int(comment.get("digg_count"))),
|
||||
"sub_comment_count": str(_as_int(comment.get("reply_comment_total"))),
|
||||
# 顶层评论在抖音里是 "0";adapters.parent_comment_id 会归一成空串。
|
||||
"parent_comment_id": str(comment.get("reply_id") or "0"),
|
||||
}
|
||||
|
||||
|
||||
async def video_comments(
|
||||
aweme_id: str, count: int = 20, *, cookie: str = ""
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""一条作品的评论,翻译成 store 的评论键名。
|
||||
|
||||
刻意不带 ``sub_comment_count`` / ``parent_comment_id`` 的猜测值 —— 接口给了就用,
|
||||
没给就留空,不编。
|
||||
"""
|
||||
identity = await browser_identity(cookie)
|
||||
if not _has_session(identity.cookie):
|
||||
raise DouyinApiError("抖音登录态不可用")
|
||||
|
||||
payload = await _get(
|
||||
COMMENT_PATH,
|
||||
{
|
||||
"aweme_id": aweme_id,
|
||||
"count": max(1, min(count, MAX_PAGE_SIZE)),
|
||||
"cursor": 0,
|
||||
"device_platform": "webapp",
|
||||
"aid": 6383,
|
||||
},
|
||||
identity,
|
||||
# 这个接口**必须**签名。不签的话网关回 200 + 空 body,会被读成「这条没评论」,
|
||||
# 而它其实只是被挡了 —— 和登录失效长得一模一样。(实测:带上签名 200/9960 字节
|
||||
# 真评论,不带就是空的。)
|
||||
signed=True,
|
||||
)
|
||||
|
||||
records = [
|
||||
normalize_comment(comment, aweme_id)
|
||||
for comment in payload.get("comments") or []
|
||||
]
|
||||
return records
|
||||
|
||||
|
||||
def _has_session(cookie: str) -> bool:
|
||||
return "sessionid=" in (cookie or "")
|
||||
|
||||
|
||||
async def check_login(cookie: str = "") -> Dict[str, Any]:
|
||||
"""浏览器/库里现在有没有可用的抖音登录态。给设置页用。"""
|
||||
identity = await browser_identity(cookie)
|
||||
if _has_session(identity.cookie):
|
||||
source = "browser" if identity.user_agent else "stored"
|
||||
return {"ok": True, "source": source, "cookie_length": len(identity.cookie)}
|
||||
return {"ok": False, "source": "", "cookie_length": 0}
|
||||
|
||||
|
||||
async def main() -> None: # pragma: no cover - 手工排查用
|
||||
"""``python -m api.monitor.douyin_api <sec_user_id>``"""
|
||||
import sys
|
||||
|
||||
if len(sys.argv) < 2:
|
||||
print(await check_login())
|
||||
return
|
||||
sec = sys.argv[1]
|
||||
print(await author_profile(sec))
|
||||
for record in await author_videos(sec, count=5):
|
||||
print(record["create_time"], record["title"][:30], record["liked_count"])
|
||||
|
||||
|
||||
if __name__ == "__main__": # pragma: no cover
|
||||
asyncio.run(main())
|
||||
@@ -0,0 +1,217 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_fetch.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""抖音采集:直接走 Web 接口,不起爬虫子进程。
|
||||
|
||||
与 ``media_platform/douyin`` 那条路的分工:
|
||||
|
||||
* 那边起一个 Playwright 子进程、构造一大串**自相矛盾的浏览器指纹参数**(参数说自己是
|
||||
Mac + Chrome 125,UA 说自己是 Linux + Chrome 155),网关回一个 200 + 空 body,
|
||||
然后被翻译成 ``Exception("account blocked")`` —— 看起来像账号被封,其实什么都不是。
|
||||
* 这边不起子进程、只发必要参数,用浏览器里那份登录态。见 ``douyin_api``。
|
||||
|
||||
采到的东西写成 ``store/douyin`` 那套 jsonl 形状,所以 **ingest 完全不知道数据是怎么来的**
|
||||
—— 重采样、差分、事件、通知、报表全都照旧。
|
||||
"""
|
||||
|
||||
import json
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, Iterable, List, Optional, Sequence
|
||||
|
||||
from tools import utils
|
||||
|
||||
from . import adapters, douyin_api
|
||||
from .models import MODE_CREATOR, MODE_NOTE, MonitorTask
|
||||
|
||||
# 拉多少条作品。接口单页上限就是 20,要多了也没用。
|
||||
DEFAULT_VIDEO_LIMIT = 20
|
||||
|
||||
|
||||
async def collect(
|
||||
out_dir: Path,
|
||||
*,
|
||||
platform: str,
|
||||
mode: str,
|
||||
limit: int,
|
||||
want_comments: bool,
|
||||
comment_limit: int,
|
||||
targets: Sequence[Any],
|
||||
known_aweme_ids: Iterable[str] = (),
|
||||
cookie: str = "",
|
||||
) -> Dict[str, Any]:
|
||||
"""跑一轮抖音采集,把产物写进 ``out_dir`` 下 store 那个目录里。
|
||||
|
||||
参数是散的、不收 ORM 对象:调用方那边 ``task`` 在会话关掉之后就 detached 了,
|
||||
传对象进来迟早会踩到「属性已过期」。
|
||||
|
||||
**不抛异常**:失败也把(可能为空的)产物落下去,并把原因放进 ``errors`` 交给调用方。
|
||||
让调用方去决定这是「一轮正常但没数据」还是「一轮失败」—— 这个判断不该藏在这里。
|
||||
"""
|
||||
notes: List[Dict[str, Any]] = []
|
||||
comments: List[Dict[str, Any]] = []
|
||||
# 博主**账号级**指标(粉丝 / 总获赞 / 作品数)。作品列表之外单独要一次,
|
||||
# 只有博主模式才有 —— 作品模式的目标是一件作品,没有"这个博主是谁"可问。
|
||||
profiles: List[Dict[str, Any]] = []
|
||||
errors: List[str] = []
|
||||
|
||||
# **整个 collect 只去重一次的、跨目标的集合**:退化路径会把「库里已知的全部作品」
|
||||
# 在每个目标下都刷一遍,多个目标就会出现同一件作品好几条记录 —— 而一对一快照的
|
||||
# 唯一键是 (task_id, note_id, run_id),同一条作品在一轮里出现两次会直接撞键。
|
||||
seen_aweme: set = set()
|
||||
|
||||
for target in targets:
|
||||
external_id = target.external_id
|
||||
|
||||
# 两种模式的目标是不同的东西,不能走同一条路:
|
||||
# 作品模式 —— 目标本身就是作品 id,直接取详情(**这个接口没被挡,今天就能用**)。
|
||||
# 博主模式 —— 目标是主页 sec_uid,要先拉作品列表;那个接口被真校验挡着,退化到
|
||||
# 刷新库里已知的作品(新作品发现不了)。
|
||||
if mode == MODE_NOTE:
|
||||
try:
|
||||
videos = [await douyin_api.video_detail(external_id, cookie=cookie)]
|
||||
except douyin_api.DouyinApiError as exc:
|
||||
errors.append(f"拉取作品 {external_id} 失败:{exc}")
|
||||
videos = []
|
||||
else:
|
||||
videos = await _creator_works(
|
||||
external_id, limit, known_aweme_ids, seen_aweme, cookie, errors
|
||||
)
|
||||
profile = await _creator_profile(external_id, videos, cookie, errors)
|
||||
if profile is not None:
|
||||
profiles.append(profile)
|
||||
|
||||
for video in videos:
|
||||
aweme_id = video.get("aweme_id")
|
||||
if not aweme_id or aweme_id in seen_aweme:
|
||||
continue
|
||||
seen_aweme.add(aweme_id)
|
||||
notes.append(video)
|
||||
|
||||
if want_comments:
|
||||
try:
|
||||
comments.extend(
|
||||
await douyin_api.video_comments(
|
||||
aweme_id, count=comment_limit, cookie=cookie
|
||||
)
|
||||
)
|
||||
except douyin_api.DouyinApiError as exc:
|
||||
errors.append(f"拉取作品 {aweme_id} 的评论失败:{exc}")
|
||||
|
||||
jsonl_dir = _write_artifacts(out_dir, platform, mode, notes, comments, profiles)
|
||||
return {
|
||||
"notes": len(notes),
|
||||
"comments": len(comments),
|
||||
"errors": errors,
|
||||
"jsonl_dir": str(jsonl_dir),
|
||||
}
|
||||
|
||||
|
||||
async def _creator_profile(
|
||||
sec_user_id: str,
|
||||
videos: Sequence[Dict[str, Any]],
|
||||
cookie: str,
|
||||
errors: List[str],
|
||||
) -> Optional[Dict[str, Any]]:
|
||||
"""问一次博主的账号级指标。拿不到就算了 —— **不能因为顺手的附加信息失败,
|
||||
就把这一轮本来采到的作品也判成失败。**
|
||||
|
||||
creator_hash 优先取作品自带的那个:作品是靠 ``anonymize_user_id(author.uid)`` 得到
|
||||
哈希的,而快照表和作品必须对得上号,否则界面上永远查不出这个博主的粉丝数。只有当一件
|
||||
作品都没采到时(列表被挡且没有已知作品可刷新),才退回资料接口自己算的哈希 ——
|
||||
那种情况下也只剩它了。
|
||||
"""
|
||||
try:
|
||||
profile = await douyin_api.author_profile(sec_user_id, cookie=cookie)
|
||||
except douyin_api.DouyinApiError as exc:
|
||||
errors.append(f"拉取博主 {sec_user_id} 的资料失败:{exc}")
|
||||
return None
|
||||
|
||||
if videos:
|
||||
profile["creator_hash"] = videos[0].get("creator_hash") or profile["creator_hash"]
|
||||
if not profile.get("creator_hash"):
|
||||
# 哈希都算不出来的快照没人能查到,落下去只是垃圾。
|
||||
errors.append(f"博主 {sec_user_id} 的资料里没有可用的身份标识,跳过账号指标")
|
||||
return None
|
||||
return profile
|
||||
|
||||
|
||||
async def _creator_works(
|
||||
sec_user_id: str,
|
||||
limit: int,
|
||||
known_aweme_ids: Iterable[str],
|
||||
seen_aweme: set,
|
||||
cookie: str,
|
||||
errors: List[str],
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""一个博主的作品:先要列表,列表被挡时退化成刷新已知作品。
|
||||
|
||||
作品列表(``aweme/post``)被抖音单独加了真校验 —— 不带 ``x-tt-argus`` 回 403,
|
||||
带上 dummy 值回 200 + 空 body。所以这里拿不到**新**作品,只能保住已知的。
|
||||
"""
|
||||
try:
|
||||
return await douyin_api.author_videos(sec_user_id, count=limit, cookie=cookie)
|
||||
except douyin_api.DouyinApiError as exc:
|
||||
errors.append(f"拉取博主 {sec_user_id} 的作品列表失败:{exc}")
|
||||
|
||||
refreshed: List[Dict[str, Any]] = []
|
||||
for aweme_id in known_aweme_ids:
|
||||
if aweme_id in seen_aweme:
|
||||
continue
|
||||
try:
|
||||
refreshed.append(await douyin_api.video_detail(aweme_id, cookie=cookie))
|
||||
except douyin_api.DouyinApiError as detail_exc:
|
||||
errors.append(f"刷新作品 {aweme_id} 失败:{detail_exc}")
|
||||
return refreshed
|
||||
|
||||
|
||||
def _write_artifacts(
|
||||
out_dir: Path,
|
||||
platform: str,
|
||||
mode: str,
|
||||
notes: List[Dict[str, Any]],
|
||||
comments: List[Dict[str, Any]],
|
||||
profiles: Sequence[Dict[str, Any]] = (),
|
||||
) -> Path:
|
||||
"""按爬虫那套目录与文件名写 jsonl。
|
||||
|
||||
目录名取 ``adapters.artifact_dir``(抖音是 ``douyin``,而不是平台 id ``dy``)——
|
||||
和 ingest 找文件用的是同一个来源,两边不会走散。
|
||||
"""
|
||||
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
|
||||
jsonl_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
kind = "creator" if mode == MODE_CREATOR else "detail"
|
||||
date = datetime.now().strftime("%Y-%m-%d")
|
||||
|
||||
_write_jsonl(jsonl_dir / f"{kind}_contents_{date}.jsonl", notes)
|
||||
# 评论文件即使没有评论也建出来:ingest 靠「文件在不在」区分「这一轮没评论」和
|
||||
# 「这一轮什么都没抓到」,两种情况的含义完全不同。
|
||||
_write_jsonl(jsonl_dir / f"{kind}_comments_{date}.jsonl", comments)
|
||||
# 博主资料同样无条件写:空文件表示"问了但没问到",没有文件表示"这次根本没问"
|
||||
# (作品模式)。两者在 ingest 那边走的是同一条路(都不落快照),但留空文件能让
|
||||
# 事后翻 run 目录时看出到底问没问过。
|
||||
_write_jsonl(jsonl_dir / f"{kind}_profile_{date}.jsonl", list(profiles))
|
||||
return jsonl_dir
|
||||
|
||||
|
||||
def _write_jsonl(path: Path, records: List[Dict[str, Any]]) -> None:
|
||||
with path.open("w", encoding="utf-8") as handle:
|
||||
for record in records:
|
||||
handle.write(json.dumps(record, ensure_ascii=False) + "\n")
|
||||
utils.logger.info(f"[douyin_fetch] 写出 {len(records)} 条 -> {path.name}")
|
||||
+244
-33
@@ -38,13 +38,14 @@ import json
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List, Optional
|
||||
from typing import Any, Dict, List, Optional, Sequence
|
||||
|
||||
from sqlalchemy import func, select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from . import adapters
|
||||
from .platforms import PLATFORM_XHS
|
||||
from .models import (
|
||||
EVENT_AUTH_FAILURE,
|
||||
@@ -55,6 +56,7 @@ from .models import (
|
||||
EVENT_NO_DATA,
|
||||
EVENT_RUN_FAILED,
|
||||
MonitorComment,
|
||||
MonitorCreatorStat,
|
||||
MonitorEvent,
|
||||
MonitorNote,
|
||||
MonitorNoteMetric,
|
||||
@@ -121,13 +123,46 @@ _WINDOWS_EXIT_REASONS = {
|
||||
}
|
||||
|
||||
|
||||
def describe_exit_code(code: int) -> str:
|
||||
"""Render an exit code so a human can act on it."""
|
||||
def describe_exit_code(code: int, cause: Optional[str] = None) -> str:
|
||||
"""Render an exit code so a human can act on it.
|
||||
|
||||
只有退出码时信息量约等于零(`code 1` 什么都能是),所以把从输出尾巴里认出来的
|
||||
异常一并附上 —— 运行历史里那一格显示的正是这句话。
|
||||
"""
|
||||
unsigned = code & 0xFFFFFFFF if code < 0 else code
|
||||
reason = _WINDOWS_EXIT_REASONS.get(unsigned)
|
||||
base = f"Crawler exited with code {code}"
|
||||
if reason:
|
||||
return f"Crawler exited with code {code} (0x{unsigned:08X}): {reason}"
|
||||
return f"Crawler exited with code {code}"
|
||||
base = f"{base} (0x{unsigned:08X}): {reason}"
|
||||
return f"{base};原因:{cause}" if cause else base
|
||||
|
||||
|
||||
# 从爬虫输出里认出一行「异常」。Python 的 traceback 末行形如
|
||||
# ``media_platform.douyin.exception.DataFetchError: account blocked``。
|
||||
_EXCEPTION_LINE_RE = re.compile(r"^[\w.]*[A-Za-z](?:Error|Exception|Timeout)\b")
|
||||
|
||||
|
||||
def diagnose_failure(output_tail: Optional[Sequence[str]]) -> Optional[str]:
|
||||
"""从爬虫输出的末尾挑出最能说明问题的一行。
|
||||
|
||||
「退出码 1」等于什么都没说:真正的报错埋在子进程的 stderr 里。倒着找第一行看起来
|
||||
像异常的行(traceback 的末行),找不到就退回最后一行有效输出。
|
||||
"""
|
||||
if not output_tail:
|
||||
return None
|
||||
|
||||
lines = [line.strip() for line in output_tail if line and line.strip()]
|
||||
# 管理器自己补的那两句不是爬虫的报错,别被当成失败原因。
|
||||
noise = ("Crawler exited with code", "Crawler completed successfully")
|
||||
lines = [line for line in lines if not line.startswith(noise)]
|
||||
if not lines:
|
||||
return None
|
||||
|
||||
for line in reversed(lines):
|
||||
if _EXCEPTION_LINE_RE.match(line):
|
||||
return line[:300]
|
||||
|
||||
return lines[-1][:300]
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -171,8 +206,11 @@ def find_run_files(
|
||||
Glob rather than reconstructing the name: both the crawler type and the date
|
||||
are runtime-dependent. Returns lists because a crawl crossing midnight
|
||||
produces one file per day.
|
||||
|
||||
``platform`` 是**监控层的平台 id**,而爬虫落盘的目录名未必同名(抖音的 id 是
|
||||
``dy``、目录是 ``douyin``),所以这里经 adapters 解析 —— 调用方不必知道这个差异。
|
||||
"""
|
||||
jsonl_dir = out_dir / platform / "jsonl"
|
||||
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
|
||||
if not jsonl_dir.is_dir():
|
||||
return [], []
|
||||
|
||||
@@ -182,6 +220,43 @@ def find_run_files(
|
||||
)
|
||||
|
||||
|
||||
def find_profile_files(out_dir: Path, platform: str = PLATFORM_XHS) -> List[Path]:
|
||||
"""博主**账号级**指标那几行 jsonl(``creator_profile_*.jsonl``)。
|
||||
|
||||
单独一个函数而不是往 ``find_run_files`` 的返回值里塞第三个列表:那个返回值的两个
|
||||
位置是有意义的(contents/comments),加一个会把所有调用点和解包语句都牵动一遍,
|
||||
而这份产物是**可选**的 —— 小红书那条路(爬虫进程)根本不产生它。
|
||||
"""
|
||||
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
|
||||
if not jsonl_dir.is_dir():
|
||||
return []
|
||||
return sorted(jsonl_dir.glob("*_profile_*.jsonl"))
|
||||
|
||||
|
||||
def _misplaced_output_dirs(out_dir: Path, expected: str) -> List[str]:
|
||||
"""在 out_dir 下找「有产物、但目录名不是期望的那个」的目录。
|
||||
|
||||
这是专门为**最难查的那类故障**准备的:产物目录名与平台对不上时,ingest 一个文件
|
||||
都找不到,现象和「登录态失效」一模一样 —— 而实际上登录好好的、数据也抓到了,
|
||||
只是没人去对的地方读。上游哪天改了 store 的目录名,这里能直接把实情说出来。
|
||||
"""
|
||||
found = []
|
||||
try:
|
||||
children = list(out_dir.iterdir())
|
||||
except OSError:
|
||||
return found
|
||||
|
||||
for child in children:
|
||||
if child.name == expected or not child.is_dir():
|
||||
continue
|
||||
try:
|
||||
if any(child.glob("jsonl/*_contents_*.jsonl")):
|
||||
found.append(child.name)
|
||||
except OSError:
|
||||
continue
|
||||
return sorted(found)
|
||||
|
||||
|
||||
async def _emit(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
@@ -271,15 +346,27 @@ async def _ingest_notes(
|
||||
run: MonitorRun,
|
||||
records: List[Dict[str, Any]],
|
||||
is_baseline: bool,
|
||||
adapter: adapters.PlatformAdapter,
|
||||
) -> int:
|
||||
"""Upsert notes, write metric snapshots, and emit new-note/delta events."""
|
||||
"""Upsert notes, write metric snapshots, and emit new-note/delta events.
|
||||
|
||||
记录里的字段一律经 ``adapter`` 读。抖音的作品没有 ``note_id``(叫 ``aweme_id``),
|
||||
按名字硬取的话每条记录都会在下面第一行被 continue 掉 —— 一条都不报错地全丢。
|
||||
"""
|
||||
now = get_current_timestamp()
|
||||
new_count = 0
|
||||
# 同一轮里重复出现的作品只处理一次。**指标快照的唯一键是 (task_id, note_id, run_id)**,
|
||||
# 同一件作品在一轮里进来两次会让第二次插入直接撞键、整个 run 崩掉 —— 产物里重复并不
|
||||
# 罕见(多个目标指向同一个人、或退化路径重复刷新)。
|
||||
seen_in_run: set = set()
|
||||
|
||||
for record in records:
|
||||
note_id = record.get("note_id")
|
||||
note_id = adapter.note_field(record, "note_id")
|
||||
if not note_id:
|
||||
continue
|
||||
if note_id in seen_in_run:
|
||||
continue
|
||||
seen_in_run.add(note_id)
|
||||
|
||||
note = await session.scalar(
|
||||
select(MonitorNote).where(
|
||||
@@ -288,20 +375,21 @@ async def _ingest_notes(
|
||||
)
|
||||
)
|
||||
|
||||
title = (record.get("title") or "")[:500]
|
||||
raw_images = record.get("image_list") or ""
|
||||
cover = raw_images.split(",")[0] if raw_images else ""
|
||||
title = (adapter.note_field(record, "title") or "")[:500]
|
||||
cover = adapter.cover(record)
|
||||
|
||||
if note is None:
|
||||
note = MonitorNote(
|
||||
task_id=run.task_id,
|
||||
note_id=note_id,
|
||||
title=title,
|
||||
note_url=record.get("note_url") or "",
|
||||
note_url=adapter.note_field(record, "note_url") or "",
|
||||
cover=cover,
|
||||
creator_hash=record.get("creator_hash") or "",
|
||||
source_kind=record.get("type") or "",
|
||||
published_at=_as_int(record.get("time")),
|
||||
creator_hash=adapter.note_field(record, "creator_hash") or "",
|
||||
creator_name=adapter.note_field(record, "creator_name") or "",
|
||||
source_kind=adapter.note_field(record, "source_kind") or "",
|
||||
# 经 to_ms 换算:小红书给毫秒、抖音给秒,差 1000 倍。
|
||||
published_at=adapter.to_ms(adapter.note_field(record, "published_at")),
|
||||
first_seen_run_id=run.id,
|
||||
first_seen_at=now,
|
||||
last_seen_run_id=run.id,
|
||||
@@ -323,6 +411,20 @@ async def _ingest_notes(
|
||||
# Only refresh descriptive fields; seen-tracking is updated below.
|
||||
if title:
|
||||
note.title = title
|
||||
# 封面地址**带签名、会过期**,所以每轮都用最新的覆盖它。原先只在首次入库
|
||||
# 时写一次,结果旧作品的封面地址烂在库里 —— 隔天开始全是 403,而且再怎么
|
||||
# 重跑也修不回来。落盘那份由 covers.cache_pending 负责(网络操作不在本模块)。
|
||||
if cover:
|
||||
note.cover = cover
|
||||
# 昵称也要刷:作者改昵称是常事,只在首次入库写一次会一直显示旧的。
|
||||
creator_name = adapter.note_field(record, "creator_name")
|
||||
if creator_name:
|
||||
note.creator_name = creator_name
|
||||
# 发布时间也刷。正常情况下它不会变,但**换算单位改过之后**(抖音是秒、
|
||||
# 小红书是毫秒),已经入库的那批只能靠重采修回来。
|
||||
published = adapter.to_ms(adapter.note_field(record, "published_at"))
|
||||
if published is not None:
|
||||
note.published_at = published
|
||||
note.last_seen_run_id = run.id
|
||||
note.last_seen_at = now
|
||||
|
||||
@@ -420,40 +522,57 @@ async def _ingest_comments(
|
||||
records: List[Dict[str, Any]],
|
||||
is_baseline: bool,
|
||||
previous_run_started_at: Optional[int],
|
||||
adapter: adapters.PlatformAdapter,
|
||||
) -> int:
|
||||
"""Upsert comments and emit events for ones never seen before."""
|
||||
"""Upsert comments and emit events for ones never seen before.
|
||||
|
||||
与作品同理,评论记录也要经 ``adapter`` 读:抖音的评论用 ``aweme_id`` 指作品。
|
||||
"""
|
||||
now = get_current_timestamp()
|
||||
new_count = 0
|
||||
|
||||
for record in records:
|
||||
comment_id = record.get("comment_id")
|
||||
note_id = record.get("note_id")
|
||||
comment_id = adapter.comment_field(record, "comment_id")
|
||||
note_id = adapter.comment_field(record, "note_id")
|
||||
if not comment_id or not note_id:
|
||||
continue
|
||||
|
||||
exists = await session.scalar(
|
||||
select(MonitorComment.id).where(
|
||||
existing = await session.scalar(
|
||||
select(MonitorComment).where(
|
||||
MonitorComment.task_id == run.task_id,
|
||||
MonitorComment.note_id == note_id,
|
||||
MonitorComment.comment_id == comment_id,
|
||||
)
|
||||
)
|
||||
if exists is not None:
|
||||
if existing is not None:
|
||||
# 昵称要跟着刷,不能只写一次。评论是去重后直接 continue 的,若不刷新,
|
||||
# 脱敏开关一改(或评论者改了昵称),已经入库的老评论会永远停在旧值上 ——
|
||||
# 而重采是唯一能拿到新值的途径。作品那边的 creator_name 同理。
|
||||
refreshed = adapter.comment_field(record, "creator_name")
|
||||
if refreshed:
|
||||
existing.nickname = refreshed
|
||||
# 时间同理:单位换算修好之后,老数据要重采才能纠正。
|
||||
created = adapter.to_ms(adapter.comment_field(record, "create_time"))
|
||||
if created is not None:
|
||||
existing.create_time = created
|
||||
continue
|
||||
|
||||
create_time = _as_int(record.get("create_time"))
|
||||
create_time = adapter.to_ms(adapter.comment_field(record, "create_time"))
|
||||
session.add(
|
||||
MonitorComment(
|
||||
task_id=run.task_id,
|
||||
note_id=note_id,
|
||||
comment_id=comment_id,
|
||||
content=(record.get("content") or "")[:2000],
|
||||
nickname=record.get("nickname") or "",
|
||||
creator_hash=record.get("creator_hash") or "",
|
||||
content=(adapter.comment_field(record, "content") or "")[:2000],
|
||||
nickname=adapter.comment_field(record, "creator_name") or "",
|
||||
creator_hash=adapter.comment_field(record, "creator_hash") or "",
|
||||
create_time=create_time,
|
||||
like_count=parse_count(record.get("like_count")),
|
||||
sub_comment_count=_as_int(record.get("sub_comment_count")) or 0,
|
||||
parent_comment_id=record.get("parent_comment_id") or "",
|
||||
like_count=parse_count(adapter.comment_field(record, "like_count")),
|
||||
sub_comment_count=_as_int(
|
||||
adapter.comment_field(record, "sub_comment_count")
|
||||
)
|
||||
or 0,
|
||||
parent_comment_id=adapter.parent_comment_id(record),
|
||||
first_seen_run_id=run.id,
|
||||
first_seen_at=now,
|
||||
)
|
||||
@@ -488,11 +607,55 @@ async def _ingest_comments(
|
||||
return new_count
|
||||
|
||||
|
||||
async def _ingest_creator_stats(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
records: Sequence[Dict[str, Any]],
|
||||
) -> int:
|
||||
"""把这一轮问到的博主账号级指标落成快照,返回条数。
|
||||
|
||||
和作品指标一样是**每轮一条**:账号级的粉丝数是缓慢变化的量,「今天比昨天多了 300」
|
||||
才是有用的信号,单看一个绝对值没有意义 —— 所以这里只管记,分析交给查询端。
|
||||
|
||||
**没解析出来的值留 NULL,不写 0**:0 在趋势图上是一条砸到底的线,和「不知道」完全是
|
||||
两回事(见 ``parse_count`` 的注释)。
|
||||
"""
|
||||
now = get_current_timestamp()
|
||||
written = 0
|
||||
seen: set = set()
|
||||
|
||||
for record in records:
|
||||
creator_hash = str(record.get("creator_hash") or "").strip()
|
||||
if not creator_hash or creator_hash in seen:
|
||||
# 一个任务可以配多个目标,退化路径下它们可能指向同一个博主 —— 而唯一键是
|
||||
# (task_id, creator_hash, run_id),重复插入会撞键把整轮炸掉。
|
||||
continue
|
||||
seen.add(creator_hash)
|
||||
|
||||
session.add(
|
||||
MonitorCreatorStat(
|
||||
task_id=run.task_id,
|
||||
run_id=run.id,
|
||||
creator_hash=creator_hash,
|
||||
nickname=str(record.get("nickname") or "")[:128],
|
||||
fans=parse_count(record.get("fans")),
|
||||
total_favorited=parse_count(record.get("total_favorited")),
|
||||
works_count=parse_count(record.get("works")),
|
||||
following=parse_count(record.get("following")),
|
||||
captured_at=now,
|
||||
)
|
||||
)
|
||||
written += 1
|
||||
|
||||
return written
|
||||
|
||||
|
||||
async def ingest_run(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
task: MonitorTask,
|
||||
out_dir: Path,
|
||||
output_tail: Optional[Sequence[str]] = None,
|
||||
) -> IngestResult:
|
||||
"""Ingest one finished run and return what changed.
|
||||
|
||||
@@ -503,17 +666,29 @@ async def ingest_run(
|
||||
# A non-zero exit is a genuine crash: trust nothing this run produced.
|
||||
if run.exit_code not in (0, None):
|
||||
run.status = RUN_FAILED
|
||||
run.error_message = describe_exit_code(run.exit_code)
|
||||
cause = diagnose_failure(output_tail)
|
||||
run.error_message = describe_exit_code(run.exit_code, cause)
|
||||
title = f"采集进程异常退出(code={run.exit_code})"
|
||||
if cause:
|
||||
# 标题里也带上真因:企业微信通知和事件流都只看这一行,不写就还得去翻日志。
|
||||
title = f"{title}:{cause}"
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_RUN_FAILED,
|
||||
f"采集进程异常退出(code={run.exit_code})",
|
||||
title,
|
||||
severity="error",
|
||||
payload={"exit_code": run.exit_code, "detail": run.error_message},
|
||||
payload={
|
||||
"exit_code": run.exit_code,
|
||||
"detail": run.error_message,
|
||||
"cause": cause,
|
||||
},
|
||||
)
|
||||
return IngestResult(status=RUN_FAILED, error=run.error_message)
|
||||
|
||||
adapter = adapters.adapter(task.platform)
|
||||
subdir = adapters.artifact_dir(task.platform)
|
||||
|
||||
contents_paths, comment_paths = find_run_files(out_dir, task.platform)
|
||||
contents = [record for path in contents_paths for record in _read_jsonl(path)]
|
||||
comments = [record for path in comment_paths for record in _read_jsonl(path)]
|
||||
@@ -521,6 +696,16 @@ async def ingest_run(
|
||||
run.notes_fetched = len(contents)
|
||||
run.comments_fetched = len(comments)
|
||||
|
||||
# 账号级快照**在「一条作品都没采到」的早退之前**落。博主的粉丝数并不会因为他这个
|
||||
# 月的新作品列表被风控挡住就不存在 —— 那正是最该看到「粉丝还在涨、但新作品没在发现」
|
||||
# 的时刻,跳过它等于在最需要它的那轮把数据丢掉。
|
||||
profiles = [
|
||||
record
|
||||
for path in find_profile_files(out_dir, task.platform)
|
||||
for record in _read_jsonl(path)
|
||||
]
|
||||
await _ingest_creator_stats(session, run, profiles)
|
||||
|
||||
# A bad cookie does NOT fail the process: XHS cookie login is never validated,
|
||||
# so an unauthenticated session just returns zero notes with exit 0 -- and
|
||||
# usually does not even create an output file. Treating that as "the creator
|
||||
@@ -529,6 +714,32 @@ async def ingest_run(
|
||||
if not contents:
|
||||
run.status = RUN_PARTIAL
|
||||
|
||||
# 先排除「东西抓到了,只是没落在我们找的那个目录里」。这种故障的现象和登录失效
|
||||
# 一模一样,但登录其实是好的 —— 按登录失效报会把人指到完全错的方向去查。
|
||||
misplaced = _misplaced_output_dirs(out_dir, subdir)
|
||||
if misplaced:
|
||||
run.error_message = (
|
||||
f"crawler wrote into {misplaced} but platform {task.platform} "
|
||||
f"expects {subdir}"
|
||||
)
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_NO_DATA,
|
||||
f"采集产物目录与平台不匹配(实际 {misplaced}、期望 {subdir}),本次未读到任何作品",
|
||||
severity="error",
|
||||
payload={
|
||||
"out_dir": str(out_dir),
|
||||
"expected": subdir,
|
||||
"found": misplaced,
|
||||
},
|
||||
)
|
||||
return IngestResult(
|
||||
status=RUN_PARTIAL,
|
||||
error=run.error_message,
|
||||
comments_fetched=len(comments),
|
||||
)
|
||||
|
||||
# Blaming the cookie is only honest if nothing else is authenticating.
|
||||
# A sibling task that just succeeded proves the login works, so the
|
||||
# fault is with this target (bad/expired per-creator token, an empty
|
||||
@@ -578,10 +789,10 @@ async def ingest_run(
|
||||
comments_fetched=len(comments),
|
||||
is_baseline=is_baseline,
|
||||
)
|
||||
result.new_notes = await _ingest_notes(session, run, contents, is_baseline)
|
||||
result.new_notes = await _ingest_notes(session, run, contents, is_baseline, adapter)
|
||||
if task.enable_comments:
|
||||
result.new_comments = await _ingest_comments(
|
||||
session, run, comments, is_baseline, previous_started_at
|
||||
session, run, comments, is_baseline, previous_started_at, adapter
|
||||
)
|
||||
|
||||
run.new_notes = result.new_notes
|
||||
|
||||
+120
-3
@@ -75,7 +75,7 @@ MODE_NOTE = "note"
|
||||
|
||||
|
||||
class MonitorTask(MonitorBase):
|
||||
"""One monitored schedule: a set of targets plus an interval."""
|
||||
"""One monitored schedule: a set of targets, plus when to run them."""
|
||||
|
||||
__tablename__ = "monitor_task"
|
||||
|
||||
@@ -86,15 +86,36 @@ class MonitorTask(MonitorBase):
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
|
||||
interval_minutes: Mapped[int] = mapped_column(Integer, nullable=False, default=360)
|
||||
|
||||
# How the task is scheduled. `interval` is the original "every N minutes" and
|
||||
# stays the default; `daily` and `weekly` fire at chosen clock times instead
|
||||
# (the arithmetic lives in schedule.py).
|
||||
#
|
||||
# The clock fields are comma-separated text rather than a child table: they
|
||||
# are a handful of small integers, always read as a whole, and a table would
|
||||
# buy nothing but joins.
|
||||
schedule_mode: Mapped[str] = mapped_column(String(16), nullable=False, default="interval")
|
||||
# 0-23, e.g. "9,12,18". Empty in interval mode.
|
||||
schedule_hours: Mapped[str] = mapped_column(String(96), nullable=False, default="")
|
||||
# 0-6 with Monday = 0, matching Python's date.weekday(). Weekly mode only.
|
||||
schedule_days: Mapped[str] = mapped_column(String(32), nullable=False, default="")
|
||||
# Minute past the hour, shared by every time in the schedule.
|
||||
schedule_minute: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
|
||||
# Crawl window knobs, mirrored onto each run's CLI flags.
|
||||
max_notes_count: Mapped[int] = mapped_column(Integer, nullable=False, default=20)
|
||||
enable_comments: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
|
||||
max_comments_count: Mapped[int] = mapped_column(Integer, nullable=False, default=50)
|
||||
run_timeout_seconds: Mapped[int] = mapped_column(Integer, nullable=False, default=3600)
|
||||
|
||||
# Push notifications are opt-in per task. A task list that all pushes to one
|
||||
# webhook turns noisy fast, so silence is the default.
|
||||
# 通知分成两类,因为它们的性质完全不同:
|
||||
#
|
||||
# * `notify_enabled` —— **推送新作品**。可能每轮都有,一条任务列表都推到同一个群
|
||||
# 会很快变吵,所以默认关。(列名是历史遗留:它早先是唯一的通知开关。)
|
||||
# * `notify_failures` —— **推送异常**(登录失效 / 运行失败 / 没抓到数据)。频率低,
|
||||
# 而且一旦发生就意味着这个任务从此**默默采不到任何东西**,你会一直不知道,
|
||||
# 直到某天发现数据停在几周前。这正是最该被告知的情况,所以默认**开**。
|
||||
notify_enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
|
||||
notify_failures: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
|
||||
|
||||
# Scheduler state. Persisted so the schedule survives an API restart.
|
||||
next_run_at: Mapped[Optional[int]] = mapped_column(BigInteger, index=True)
|
||||
@@ -205,7 +226,12 @@ class MonitorNote(MonitorBase):
|
||||
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
note_url: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
cover: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
# 创作者匿名哈希。爬虫刻意不落原始 user_id(见 tools/user_hash.py),
|
||||
# 所以这是唯一稳定的创作者标识 —— 按博主分组就靠它。
|
||||
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
# 创作者昵称,**已由爬虫脱敏**(张***三 这种)。存的是脱敏后的值,与项目一贯的
|
||||
# 匿名化姿态一致;不存的话分组只能显示一串哈希,根本认不出是谁。
|
||||
creator_name: Mapped[str] = mapped_column(String(200), nullable=False, default="")
|
||||
source_kind: Mapped[str] = mapped_column(String(16), nullable=False, default="")
|
||||
published_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
|
||||
@@ -294,6 +320,87 @@ class MonitorEvent(MonitorBase):
|
||||
is_read: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
|
||||
|
||||
|
||||
class MonitorCreatorAlias(MonitorBase):
|
||||
"""给博主起的备注。
|
||||
|
||||
作品栏和评论栏都按 ``creator_hash`` 把作品归到博主名下,可那是个哈希;
|
||||
``creator_name`` 是平台上的昵称(而且粉丝少的号常常没有)。两样都认不出"这是谁"。
|
||||
备注是**人自己起的名字**(「竞品A」「自家号-3」),用来把账号对上人。
|
||||
|
||||
键取 ``(platform, creator_hash)``:哈希对同一个 uid 是稳定的,所以同一个博主出现在
|
||||
多个任务里时备注也是同一个,不用每个任务各填一遍。
|
||||
"""
|
||||
|
||||
__tablename__ = "monitor_creator_alias"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("platform", "creator_hash", name="uq_creator_alias"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
platform: Mapped[str] = mapped_column(String(16), nullable=False, index=True)
|
||||
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
|
||||
alias: Mapped[str] = mapped_column(String(128), nullable=False, default="")
|
||||
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
|
||||
class MonitorCreatorStat(MonitorBase):
|
||||
"""博主的**账号级**快照:粉丝数 / 总获赞 / 作品数 / 关注数。
|
||||
|
||||
这是作品列表给不了的东西:作品级指标说"这一条视频涨了多少赞",账号级说"这个人
|
||||
整个账号的粉丝是在涨还是在掉"。两者不互相替代。
|
||||
|
||||
粒度取 ``(任务, 博主, 轮次)``,和作品指标一样的形状 —— 于是趋势、差分、报表那套
|
||||
现成的逻辑换个表就能用。
|
||||
|
||||
目前**只有抖音**会写它:小红书那条走的是爬虫子进程,而它的 ``save_creator()`` 在
|
||||
教学版里是空函数,根本没落过创作者资料。所以表里只有抖音的博主。
|
||||
"""
|
||||
|
||||
__tablename__ = "monitor_creator_stat"
|
||||
__table_args__ = (
|
||||
UniqueConstraint(
|
||||
"task_id", "creator_hash", "run_id", name="uq_creator_stat"
|
||||
),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
task_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
run_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
|
||||
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
|
||||
nickname: Mapped[str] = mapped_column(String(128), nullable=False, default="")
|
||||
|
||||
# 都可能为 None:平台没给就留空,**不要伪造成 0** —— 0 是"掉到零",和"不知道"
|
||||
# 在趋势图上是完全不同的两回事。
|
||||
fans: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
total_favorited: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
works_count: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
following: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
|
||||
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
|
||||
|
||||
|
||||
class MonitorNoteAlias(MonitorBase):
|
||||
"""给**作品**起的备注。
|
||||
|
||||
和 ``MonitorCreatorAlias`` 是一对:博主那条回答"这是谁",这条回答"这条我要盯着"。
|
||||
键取 ``(platform, note_id)`` —— 作品 id 本身就带平台语义,但显式带上 platform 才能和
|
||||
博主备注用同一套查询形状。
|
||||
"""
|
||||
|
||||
__tablename__ = "monitor_note_alias"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("platform", "note_id", name="uq_note_alias"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
platform: Mapped[str] = mapped_column(String(16), nullable=False, index=True)
|
||||
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
|
||||
alias: Mapped[str] = mapped_column(String(128), nullable=False, default="")
|
||||
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
|
||||
class MonitorSetting(MonitorBase):
|
||||
"""Key/value store. Holds the XHS cookie for unattended runs."""
|
||||
|
||||
@@ -331,6 +438,16 @@ SETTING_AUTH_PASSWORD_UPDATED_AT = "auth_password_updated_at"
|
||||
# them. Key builders live in settings.py.
|
||||
SETTING_WECOM_WEBHOOK = "system.wecom_webhook"
|
||||
|
||||
# 上游更新检查的两条状态。都不是给用户编辑的设置项,所以不在 app_settings 的注册表里
|
||||
# (那张表只列可编辑项,因此也不会被设置接口读出来)。
|
||||
#
|
||||
# 最近一次检查的结果整体存成一条 JSON:它总是被整体读写,拆成多个 key 只会带来
|
||||
# 另一半没写完的不一致。
|
||||
SETTING_UPSTREAM_STATE = "system.upstream_check_state"
|
||||
# 已经推送过通知的那个上游 tip。换 tip 才再推 —— 否则每个检查周期都会把同样的
|
||||
# 更新推一遍,直到有人去合并为止;而上游真又动了的时候应该再推一次。
|
||||
SETTING_UPSTREAM_NOTIFIED_TIP = "system.upstream_notified_tip"
|
||||
|
||||
# Pre-namespacing keys, kept only so the startup migration can find and move
|
||||
# them. Nothing should read these directly.
|
||||
LEGACY_SETTING_KEY_RENAMES = {
|
||||
|
||||
+21
-4
@@ -36,6 +36,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from . import adapters
|
||||
from .models import (
|
||||
EVENT_AUTH_FAILURE,
|
||||
EVENT_NEW_NOTE,
|
||||
@@ -107,14 +108,27 @@ async def build_run_message(
|
||||
task: MonitorTask,
|
||||
run: MonitorRun,
|
||||
) -> Optional[str]:
|
||||
"""Compose one markdown summary for a finished run, or None if nothing to say."""
|
||||
"""Compose one markdown summary for a finished run, or None if nothing to say.
|
||||
|
||||
事件按开关过滤:只勾了「新作品」的任务,不该因为一次失败被推消息,反之亦然 ——
|
||||
否则拆开这两个开关就没有意义了。
|
||||
"""
|
||||
allowed = []
|
||||
if task.notify_enabled:
|
||||
allowed.append(EVENT_NEW_NOTE)
|
||||
if task.notify_failures:
|
||||
allowed.extend([EVENT_AUTH_FAILURE, EVENT_RUN_FAILED, EVENT_NO_DATA])
|
||||
|
||||
if not allowed:
|
||||
return None
|
||||
|
||||
events = list(
|
||||
(
|
||||
await session.scalars(
|
||||
select(MonitorEvent)
|
||||
.where(
|
||||
MonitorEvent.run_id == run.id,
|
||||
MonitorEvent.type.in_(NOTIFIABLE_EVENT_TYPES),
|
||||
MonitorEvent.type.in_(allowed),
|
||||
)
|
||||
.order_by(MonitorEvent.id)
|
||||
)
|
||||
@@ -151,7 +165,9 @@ async def build_run_message(
|
||||
payload = _load_payload(event.payload_json)
|
||||
title = payload.get("title") or event.target_id
|
||||
note_id = payload.get("note_id") or event.target_id
|
||||
url = f"https://www.xiaohongshu.com/explore/{note_id}"
|
||||
# 链接形状按平台来。抖音的作品是 /video/{id},写死小红书域名的话,
|
||||
# 群里点进去会是一个 404 —— 而这正是通知唯一要它干的事。
|
||||
url = adapters.adapter(task.platform).note_url(note_id)
|
||||
lines.append(f"> [{title}]({url})")
|
||||
if len(new_notes) > 10:
|
||||
lines.append(f"> …等共 {len(new_notes)} 篇")
|
||||
@@ -176,7 +192,8 @@ async def notify_run(session: AsyncSession, task: MonitorTask, run: MonitorRun)
|
||||
Returns the message that was sent, or None. Never raises.
|
||||
"""
|
||||
try:
|
||||
if not task.notify_enabled:
|
||||
# 两个开关是分开的:只开「异常」不该因为新作品而发消息,反之亦然。
|
||||
if not (task.notify_enabled or task.notify_failures):
|
||||
return None
|
||||
|
||||
webhook_url = await get_webhook_url(session)
|
||||
|
||||
@@ -29,12 +29,13 @@ Two distinct things are recorded here, and conflating them would be misleading:
|
||||
platform modules, not assumed -- all seven implement search/detail/creator;
|
||||
the real differences are in which interaction metrics they capture.
|
||||
* ``monitor_wired`` says whether the *monitoring layer* has been hooked up. It
|
||||
currently covers only Xiaohongshu: ``runner.py`` pins the platform,
|
||||
``ingest.py`` reads a fixed ``xhs/jsonl`` directory, and ``service.py`` only
|
||||
parses Xiaohongshu target URLs.
|
||||
covers Xiaohongshu and Douyin. The parts where those two differ -- which
|
||||
directory the crawler writes into, what the jsonl fields are called, what a
|
||||
target URL looks like -- live in ``adapters.py``; the rest of the layer is
|
||||
platform-neutral.
|
||||
|
||||
A platform can therefore be fully crawlable by upstream and still not usable for
|
||||
monitoring, which is exactly the state of the other six today.
|
||||
monitoring, which is exactly the state of the other five today.
|
||||
"""
|
||||
|
||||
from typing import Any, Dict, List, Optional
|
||||
@@ -61,13 +62,31 @@ PLATFORM_CAPABILITIES: Dict[str, Dict[str, Any]] = {
|
||||
"comment_levels": 2,
|
||||
"media": True,
|
||||
"monitor_wired": True,
|
||||
"target_hints": {
|
||||
"creator": "https://www.xiaohongshu.com/user/profile/5f58bd990000000001003753",
|
||||
"note": "https://www.xiaohongshu.com/explore/6aa3d827000000002802c5c8?xsec_token=...",
|
||||
"creator_label": "博主主页",
|
||||
"note_label": "笔记",
|
||||
# 只有小红书的链接带会过期的 xsec_token。抖音的链接不带令牌,永久有效,
|
||||
# 那句「建议只填纯 ID」的劝告对它没有意义。
|
||||
"token_expires": True,
|
||||
},
|
||||
},
|
||||
"dy": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
|
||||
"comment_levels": 2,
|
||||
"media": True,
|
||||
"monitor_wired": False,
|
||||
"monitor_wired": True,
|
||||
# 用户可见的示例链接(前端的目标输入框用它做 placeholder)。放这里是因为
|
||||
# 它属于「这个平台长什么样」的能力描述;真正干活的管子(正则、目录名、
|
||||
# 字段别名)在 adapters.py。
|
||||
"target_hints": {
|
||||
"creator": "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE",
|
||||
"note": "https://www.douyin.com/video/7525082444551310602",
|
||||
"creator_label": "博主主页",
|
||||
"note_label": "作品",
|
||||
},
|
||||
},
|
||||
"ks": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
|
||||
@@ -0,0 +1,425 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/qrlogin.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""扫码登录,以及"现在到底登没登录"的查询。
|
||||
|
||||
**为什么需要这个模块**:服务器上 Chrome 跑在 Xvfb 里没有显示器,爬虫原本用
|
||||
`show_qrcode`(PIL 的 `Image.show()`)弹窗展示二维码,那需要桌面看图程序,服务器上
|
||||
没有。所以改成经 CDP 把二维码从页面里读出来交给 WebUI。
|
||||
|
||||
**为什么登录状态要能独立查询**:扫码会话是内存里的临时状态,进程一重启就没了
|
||||
(部署、崩溃都算)。把"是否已登录"绑在它上面,就会出现"扫完了但界面没反应、
|
||||
也不知道到底成没成"。所以状态查询是独立的、随时可调用的,二维码只是达成它的手段之一。
|
||||
|
||||
三个容易搞错的地方:
|
||||
|
||||
* **必须复用浏览器默认 context**。`browser.new_context()` 会造出一个无痕式的 profile,
|
||||
扫了也白扫——爬虫读不到那份 cookie。真正的 profile 在 `browser.contexts[0]`。
|
||||
* **绝不能调 `browser.close()`**。对 CDP 连接而言那会关掉操作者自己的 Chrome,
|
||||
连带所有无关标签页。只能关本模块自己开的那一个。
|
||||
* **不能靠 `web_session` 判断登录**。实测:一个全新的空 profile 首次访问小红书就会
|
||||
被发一个 `web_session`,所以"有这个 cookie"什么都证明不了。可信信号是页面自己的
|
||||
`__INITIAL_STATE__.user.loggedIn`。
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import os
|
||||
import time
|
||||
from typing import Any, Dict, Optional
|
||||
|
||||
import config
|
||||
from playwright.async_api import async_playwright
|
||||
from tools import utils
|
||||
|
||||
from ..creator.client import CreatorApiError, CreatorClient
|
||||
from .platforms import PLATFORM_XHS
|
||||
|
||||
|
||||
def _cookie_string(cookies) -> str:
|
||||
"""把 CDP 拿到的 cookie 列表拼成请求头用的字符串。"""
|
||||
return "; ".join(f"{c['name']}={c['value']}" for c in cookies)
|
||||
|
||||
# 二维码有效期。平台自己会更早轮换;这个上限只是为了让一次被放弃的尝试不会
|
||||
# 永久占着一个标签页。
|
||||
QR_TTL_SECONDS = 300
|
||||
|
||||
# 登录状态查询的缓存时长。轮询时不必每次都去问浏览器。
|
||||
STATE_CACHE_SECONDS = 5
|
||||
|
||||
STATUS_IDLE = "idle"
|
||||
STATUS_WAITING = "waiting"
|
||||
STATUS_SUCCESS = "success"
|
||||
STATUS_EXPIRED = "expired"
|
||||
STATUS_ERROR = "error"
|
||||
|
||||
# 只有小红书接了监控流程,所以扫码也只对它开放。给别的平台显示一个按不动的按钮
|
||||
# 是在假装功能存在。
|
||||
LOGIN_URL: Dict[str, str] = {PLATFORM_XHS: "https://www.xiaohongshu.com"}
|
||||
EXPLORE_URL: Dict[str, str] = {PLATFORM_XHS: "https://www.xiaohongshu.com/explore"}
|
||||
QR_SELECTOR: Dict[str, str] = {PLATFORM_XHS: "xpath=//img[@class='qrcode-img']"}
|
||||
LOGIN_BUTTON_SELECTOR: Dict[str, str] = {
|
||||
PLATFORM_XHS: "xpath=//*[@id='app']/div[1]/div[2]/div[1]/ul/div[1]/button"
|
||||
}
|
||||
|
||||
# 这里本来有一个读 window.__INITIAL_STATE__ 的 JS 探针,**已删除,不要加回来**。
|
||||
#
|
||||
# 它是页面加载那一刻的快照:浏览器本来就登录着时它是对的,但扫码是加载**之后**才
|
||||
# 登录的,快照不会翻转,检测于是永远等不到 —— 表现为"扫了码却一直停在二维码上"。
|
||||
# 运营模块踩过同一个坑。现在的判据是拿 cookie 问后台接口,见 check_login_state。
|
||||
|
||||
_lock = asyncio.Lock()
|
||||
_current: Optional["QrLoginSession"] = None
|
||||
|
||||
# 常驻的 Playwright 客户端和本模块自己的标签页。长期持有是有意的:状态查询要能
|
||||
# 随时回答,而每次都新建一个标签页会在操作者的浏览器里堆垃圾。
|
||||
_playwright: Any = None
|
||||
_page: Any = None
|
||||
|
||||
# (时间戳, 结果),避免轮询时反复问浏览器。
|
||||
_state_cache: Optional[tuple[float, Dict[str, Any]]] = None
|
||||
|
||||
|
||||
def _cdp_url() -> str:
|
||||
"""浏览器 DevTools 端点。``MC_CDP_URL`` 优先,便于换主机而不用改代码。"""
|
||||
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
|
||||
|
||||
|
||||
def _login_url(platform: str) -> str:
|
||||
if platform == PLATFORM_XHS and getattr(config, "XHS_INTERNATIONAL", False):
|
||||
return "https://www.rednote.com"
|
||||
return LOGIN_URL[platform]
|
||||
|
||||
|
||||
async def _ensure_context() -> Any:
|
||||
"""连上浏览器并返回它的默认 context。"""
|
||||
global _playwright
|
||||
|
||||
if _playwright is None:
|
||||
_playwright = await async_playwright().start()
|
||||
try:
|
||||
browser = await _playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
|
||||
except Exception as exc:
|
||||
await _reset_playwright()
|
||||
raise RuntimeError(
|
||||
f"连接浏览器失败({_cdp_url()})。请确认服务器上的 Chrome 以 "
|
||||
f"--remote-debugging-port 启动。原始错误:{exc}"
|
||||
) from exc
|
||||
|
||||
if not browser.contexts:
|
||||
raise RuntimeError(
|
||||
"浏览器没有可用上下文。CDP 已连上,但读不到 profile —— "
|
||||
"请确认 Chrome 不是以无痕模式启动的。"
|
||||
)
|
||||
# contexts[0] 就是真实 profile,用它,不要 new_context()。
|
||||
return browser.contexts[0]
|
||||
|
||||
|
||||
async def _ensure_page(platform: str = PLATFORM_XHS, reload: bool = False) -> Any:
|
||||
"""本模块在操作者浏览器里的那一个标签页,复用而不是反复新建。
|
||||
|
||||
若已有一个停在目标站点的标签页就认领它——进程重启后页柄会丢,但标签页还在,
|
||||
认领可以避免在浏览器里留下一堆没人关的孤儿页。
|
||||
"""
|
||||
global _page
|
||||
|
||||
context = await _ensure_context()
|
||||
|
||||
if _page is not None:
|
||||
try:
|
||||
if _page.is_closed():
|
||||
_page = None
|
||||
except Exception:
|
||||
_page = None
|
||||
|
||||
if _page is None:
|
||||
for candidate in context.pages:
|
||||
try:
|
||||
if "xiaohongshu.com" in candidate.url or "rednote.com" in candidate.url:
|
||||
_page = candidate
|
||||
break
|
||||
except Exception:
|
||||
continue
|
||||
|
||||
if _page is None:
|
||||
_page = await context.new_page()
|
||||
|
||||
try:
|
||||
url = _page.url
|
||||
except Exception:
|
||||
url = ""
|
||||
|
||||
if reload or "xiaohongshu.com" not in url and "rednote.com" not in url:
|
||||
await _page.goto(
|
||||
EXPLORE_URL.get(platform, EXPLORE_URL[PLATFORM_XHS]),
|
||||
wait_until="domcontentloaded",
|
||||
timeout=45000,
|
||||
)
|
||||
|
||||
return _page
|
||||
|
||||
|
||||
async def check_login_state(force: bool = False) -> Dict[str, Any]:
|
||||
"""问浏览器:现在登录了吗?
|
||||
|
||||
``force`` 会先重新加载页面。SPA 的状态会随登录实时更新,所以轮询时不必重载;
|
||||
但若登录态是在别处失效的,页面上的副本可能是陈旧的,重新检测就该重载。
|
||||
"""
|
||||
global _state_cache
|
||||
|
||||
now = time.time()
|
||||
if not force and _state_cache is not None:
|
||||
cached_at, cached = _state_cache
|
||||
if now - cached_at < STATE_CACHE_SECONDS:
|
||||
return cached
|
||||
|
||||
try:
|
||||
context = await _ensure_context()
|
||||
cookies = await context.cookies()
|
||||
except Exception as exc:
|
||||
result = {
|
||||
"known": False,
|
||||
"logged_in": False,
|
||||
"nickname": None,
|
||||
"error": f"{exc.__class__.__name__}: {exc}",
|
||||
}
|
||||
_state_cache = (now, result)
|
||||
return result
|
||||
|
||||
cookie = _cookie_string(cookies)
|
||||
|
||||
# 判据不再是页面里的 window.__INITIAL_STATE__ —— 那是**页面加载那一刻的快照**:
|
||||
# 浏览器已登录时它是对的,但扫码是加载**之后**才登录的,快照不会翻转,检测就永远
|
||||
# 等不到(运营模块踩过同一个坑)。改成拿 cookie 问后台「我是谁」,那是权威的:
|
||||
# 实测游客也会被发一个 web_session,所以「有这个 cookie」什么都证明不了,
|
||||
# 后台认了才算。
|
||||
try:
|
||||
info = await CreatorClient(cookie).fetch_user_info()
|
||||
except CreatorApiError:
|
||||
result = {"known": True, "logged_in": False, "nickname": None}
|
||||
else:
|
||||
result = {
|
||||
"known": True,
|
||||
"logged_in": bool(info.get("user_id")),
|
||||
"nickname": info.get("nickname"),
|
||||
}
|
||||
|
||||
_state_cache = (now, result)
|
||||
return result
|
||||
|
||||
|
||||
async def _current_cookie() -> str:
|
||||
"""默认 profile 当前的小红书 cookie 串。
|
||||
|
||||
扫码面板要的不只是「登录了吗」,而是**把登录态拿出来存一份** —— 存进库之后,
|
||||
即使 CDP 关掉、任务改用 --cookies_file 注入,也照样能跑。
|
||||
"""
|
||||
context = await _ensure_context()
|
||||
return _cookie_string(await context.cookies())
|
||||
|
||||
|
||||
async def _reset_playwright() -> None:
|
||||
global _playwright, _page
|
||||
_page = None
|
||||
if _playwright is not None:
|
||||
try:
|
||||
await _playwright.stop()
|
||||
except Exception:
|
||||
pass
|
||||
_playwright = None
|
||||
|
||||
|
||||
class QrLoginSession:
|
||||
"""一次进行中的扫码尝试。"""
|
||||
|
||||
def __init__(self, platform: str, page: Any) -> None:
|
||||
self.platform = platform
|
||||
self.status = STATUS_WAITING
|
||||
self.message = "请用手机扫描二维码"
|
||||
self.image = ""
|
||||
self.started_at = time.time()
|
||||
self.logged_in = False
|
||||
self.nickname: Optional[str] = None
|
||||
# 登录成功后从默认 profile 取出来的 cookie,供调用方存库。
|
||||
self.cookie: str = ""
|
||||
self.cookie_taken = False
|
||||
self._page = page
|
||||
|
||||
@property
|
||||
def elapsed(self) -> float:
|
||||
return time.time() - self.started_at
|
||||
|
||||
async def refresh(self) -> None:
|
||||
"""轮询一次,看扫码是否完成。"""
|
||||
if self.status != STATUS_WAITING:
|
||||
return
|
||||
if self.elapsed > QR_TTL_SECONDS:
|
||||
self.status = STATUS_EXPIRED
|
||||
self.message = "二维码已超时,请重新获取"
|
||||
return
|
||||
|
||||
state = await check_login_state()
|
||||
if state.get("logged_in"):
|
||||
self.cookie = await _current_cookie()
|
||||
self.logged_in = True
|
||||
self.nickname = state.get("nickname")
|
||||
self.status = STATUS_SUCCESS
|
||||
who = f"({self.nickname})" if self.nickname else ""
|
||||
self.message = f"登录成功{who},登录态已写入浏览器 profile"
|
||||
return
|
||||
|
||||
try:
|
||||
if self._page.is_closed():
|
||||
self.status = STATUS_ERROR
|
||||
self.message = "二维码所在页面已被关闭,请重新获取"
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
def snapshot(self) -> Dict[str, Any]:
|
||||
return {
|
||||
"status": self.status,
|
||||
"platform": self.platform,
|
||||
"image": self.image,
|
||||
"message": self.message,
|
||||
"elapsed": int(self.elapsed),
|
||||
"expires_in": max(0, int(QR_TTL_SECONDS - self.elapsed)),
|
||||
"logged_in": self.logged_in,
|
||||
"nickname": self.nickname,
|
||||
}
|
||||
|
||||
|
||||
async def _read_qr(page: Any, platform: str) -> str:
|
||||
"""把二维码从页面里取出来,必要时先点开登录框。"""
|
||||
image = await utils.find_login_qrcode(page, selector=QR_SELECTOR[platform])
|
||||
if image:
|
||||
return image
|
||||
# 登录框不一定自己弹出来。这是爬虫自身扫码流程里同款兜底。
|
||||
await asyncio.sleep(0.5)
|
||||
try:
|
||||
await page.locator(LOGIN_BUTTON_SELECTOR[platform]).click(timeout=5000)
|
||||
except Exception:
|
||||
return ""
|
||||
return await utils.find_login_qrcode(page, selector=QR_SELECTOR[platform])
|
||||
|
||||
|
||||
async def _discard_current_locked() -> None:
|
||||
global _current
|
||||
_current = None
|
||||
|
||||
|
||||
async def start(platform: str = PLATFORM_XHS) -> Dict[str, Any]:
|
||||
"""在 CDP 浏览器里打开登录页,取回二维码。"""
|
||||
global _current
|
||||
|
||||
if platform not in LOGIN_URL:
|
||||
raise ValueError(f"平台 {platform} 尚未接入扫码登录(目前仅支持小红书)")
|
||||
|
||||
async with _lock:
|
||||
await _discard_current_locked()
|
||||
|
||||
# **先问状态,再决定要不要开页面。** 顺序反过来是有代价的:读二维码内部会
|
||||
# wait_for_selector 等满 30 秒才放弃,而已经登录时页面上根本没有二维码 ——
|
||||
# 用户点一下按钮要干等半分钟,还白开一个标签页。
|
||||
state = await check_login_state(force=True)
|
||||
if state.get("logged_in"):
|
||||
# 已经是登录状态时站点不显示二维码 —— 这本身就是成功,不是失败。
|
||||
# 顺带把 cookie 取出来,让调用方可以存进库。
|
||||
session = QrLoginSession(platform, None)
|
||||
session.cookie = await _current_cookie()
|
||||
session.status = STATUS_SUCCESS
|
||||
session.logged_in = True
|
||||
session.nickname = state.get("nickname")
|
||||
who = f"({session.nickname})" if session.nickname else ""
|
||||
session.message = f"浏览器已经是登录状态{who},无需扫码"
|
||||
_current = session
|
||||
return session.snapshot()
|
||||
|
||||
page = await _ensure_page(platform)
|
||||
try:
|
||||
await page.goto(
|
||||
_login_url(platform), wait_until="domcontentloaded", timeout=45000
|
||||
)
|
||||
image = await _read_qr(page, platform)
|
||||
except Exception as exc:
|
||||
raise RuntimeError(f"打开登录页失败:{exc}") from exc
|
||||
|
||||
session = QrLoginSession(platform, page)
|
||||
session.image = image
|
||||
if not image:
|
||||
session.status = STATUS_ERROR
|
||||
session.message = "页面上没找到二维码,请确认站点结构没有变化"
|
||||
|
||||
_current = session
|
||||
return session.snapshot()
|
||||
|
||||
|
||||
async def status() -> Dict[str, Any]:
|
||||
async with _lock:
|
||||
if _current is None:
|
||||
state = await check_login_state()
|
||||
return {
|
||||
"status": STATUS_IDLE,
|
||||
"platform": None,
|
||||
"image": "",
|
||||
"message": "",
|
||||
"elapsed": 0,
|
||||
"expires_in": 0,
|
||||
"logged_in": bool(state.get("logged_in")),
|
||||
"nickname": state.get("nickname"),
|
||||
}
|
||||
await _current.refresh()
|
||||
return _current.snapshot()
|
||||
|
||||
|
||||
async def take_cookie() -> Optional[str]:
|
||||
"""取走已登录会话的 cookie,且只给一次。
|
||||
|
||||
由路由层在落库时调用。**cookie 不进响应体** —— 它是凭证,前端没有理由看到它。
|
||||
|
||||
这里**刻意不结束会话**(与运营模块不同):那里取完即拆,因为临时上下文用完就该丢;
|
||||
这里的浏览器 profile 是长期存在的,面板还该继续显示「已登录」。所以只标记已取过,
|
||||
让重复轮询拿不到第二份、也就不会反复写库。
|
||||
"""
|
||||
async with _lock:
|
||||
if _current is None or _current.status != STATUS_SUCCESS or _current.cookie_taken:
|
||||
return None
|
||||
_current.cookie_taken = True
|
||||
return _current.cookie
|
||||
|
||||
|
||||
async def cancel() -> Dict[str, Any]:
|
||||
async with _lock:
|
||||
await _discard_current_locked()
|
||||
state = await check_login_state()
|
||||
return {
|
||||
"status": STATUS_IDLE,
|
||||
"platform": None,
|
||||
"image": "",
|
||||
"message": "已取消",
|
||||
"elapsed": 0,
|
||||
"expires_in": 0,
|
||||
"logged_in": bool(state.get("logged_in")),
|
||||
"nickname": state.get("nickname"),
|
||||
}
|
||||
|
||||
|
||||
async def shutdown() -> None:
|
||||
"""进程退出时断开连接。刻意不关那个标签页——它是操作者浏览器的一部分。"""
|
||||
global _current
|
||||
async with _lock:
|
||||
_current = None
|
||||
await _reset_playwright()
|
||||
@@ -135,7 +135,11 @@ async def build_report(
|
||||
start_ms, _ = day_bounds(start_day)
|
||||
_, end_ms = day_bounds(end_day)
|
||||
|
||||
scope = list(task_ids) if task_ids else None
|
||||
# 必须是 `is not None`,不能写 `if task_ids` —— **空列表是假值**,而空列表在这里
|
||||
# 的含义是「这个平台一个任务都没有」,不是「不限制平台」。用真值判断的话,
|
||||
# 切到一个还没有任务的平台,报表会把**所有**任务的数据聚合出来(看起来就是
|
||||
# 「抖音的报表里全是小红书的数据」)。
|
||||
scope = list(task_ids) if task_ids is not None else None
|
||||
days = iter_days(start_day, end_day)
|
||||
|
||||
# Fetch every snapshot up to the range end: the delta on the first day needs
|
||||
|
||||
+155
-60
@@ -28,6 +28,8 @@ import os
|
||||
from pathlib import Path
|
||||
from typing import Iterable, List, Optional
|
||||
|
||||
from sqlalchemy import select
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from ..schemas import (
|
||||
@@ -38,15 +40,16 @@ from ..schemas import (
|
||||
SaveDataOptionEnum,
|
||||
)
|
||||
from ..services import crawler_manager
|
||||
from . import app_settings, notify
|
||||
from . import adapters, app_settings, covers, douyin_fetch, notify
|
||||
from .db import get_session
|
||||
from .ingest import IngestResult, ingest_run
|
||||
from .ingest import IngestResult, diagnose_failure, ingest_run
|
||||
from .models import (
|
||||
MODE_CREATOR,
|
||||
RUN_FAILED,
|
||||
RUN_PENDING,
|
||||
RUN_RUNNING,
|
||||
RUN_TIMEOUT,
|
||||
MonitorNote,
|
||||
MonitorRun,
|
||||
MonitorTarget,
|
||||
MonitorTask,
|
||||
@@ -68,30 +71,33 @@ _PLATFORM_ENUM = {
|
||||
"zhihu": PlatformEnum.ZHIHU,
|
||||
}
|
||||
|
||||
_XHS_WEB_BASE = "https://www.xiaohongshu.com"
|
||||
_CREATOR_PATH = "/user/profile"
|
||||
_NOTE_PATH = "/explore"
|
||||
|
||||
# Timeout used when the caller does not care; tasks carry their own.
|
||||
DEFAULT_RUN_TIMEOUT_SECONDS = 3600
|
||||
|
||||
|
||||
def build_target_url(value: str, kind: str) -> str:
|
||||
def build_target_url(value: str, kind: str, platform: str) -> str:
|
||||
"""Turn a stored target into a URL the crawler's parser accepts.
|
||||
|
||||
Always emits a full URL rather than a bare id: the XHS parser accepts a bare
|
||||
24-hex id only, so the URL form is the safer universal input. The
|
||||
``xsec_token`` is appended when present but is deliberately optional -- it
|
||||
expires, and the id alone is what keeps a long-running task alive.
|
||||
Always emits a full URL rather than a bare id: both platforms' parsers accept
|
||||
a bare id only in a narrower form, so the URL is the safer universal input.
|
||||
The shape itself is platform-specific and comes from ``adapters``.
|
||||
"""
|
||||
path = _CREATOR_PATH if kind == MODE_CREATOR else _NOTE_PATH
|
||||
return f"{_XHS_WEB_BASE}{path}/{value}"
|
||||
spec = adapters.adapter(platform)
|
||||
path = spec.creator_path if kind == MODE_CREATOR else spec.note_path
|
||||
return f"{spec.web_base}{path}/{value}"
|
||||
|
||||
|
||||
def build_target_urls(mode: str, targets: Iterable[MonitorTarget]) -> List[str]:
|
||||
def build_target_urls(
|
||||
mode: str, targets: Iterable[MonitorTarget], platform: str
|
||||
) -> List[str]:
|
||||
"""存储的目标 -> 爬虫接受的 URL。
|
||||
|
||||
``xsec_token`` 只有小红书有,而且是会过期的刷新令牌 —— 有就带上,没有就算了。
|
||||
抖音恒为空,所以这一段对它是天然的 no-op,不需要平台分支。
|
||||
"""
|
||||
urls = []
|
||||
for target in targets:
|
||||
url = build_target_url(target.external_id, target.kind)
|
||||
url = build_target_url(target.external_id, target.kind, platform)
|
||||
if target.xsec_token:
|
||||
url = f"{url}?xsec_token={target.xsec_token}"
|
||||
if target.xsec_source:
|
||||
@@ -159,9 +165,14 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
|
||||
raise ValueError(f"Monitor task {task_id} has no enabled targets")
|
||||
|
||||
platform = task.platform
|
||||
urls = build_target_urls(task.mode, targets)
|
||||
urls = build_target_urls(task.mode, targets, platform)
|
||||
cookie = await get_cookie(session, platform)
|
||||
strategy = await _strategy_settings(session, platform)
|
||||
# System-wide switch. On a headless server the crawler must attach to the
|
||||
# Chrome already listening on the debug port -- that browser is where the
|
||||
# operator scanned the login QR, so its profile is the login. Default False
|
||||
# keeps desktop runs launching a private browser exactly as before.
|
||||
cdp_enabled = await app_settings.get_value(session, "cdp_enabled", fallback=False)
|
||||
|
||||
run = MonitorRun(
|
||||
task_id=task.id,
|
||||
@@ -186,55 +197,118 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
|
||||
max_notes_count = task.max_notes_count
|
||||
max_comments_count = task.max_comments_count
|
||||
timeout_seconds = task.run_timeout_seconds
|
||||
# 抖音的作品列表接口被那道真校验挡着(见 douyin_fetch),拿不到列表时就靠这些
|
||||
# 已知的 aweme_id 逐条刷新 —— 新作品发现不了,但已有作品的指标还能继续更新。
|
||||
known_aweme_ids = list(
|
||||
await session.scalars(
|
||||
select(MonitorNote.note_id).where(MonitorNote.task_id == task.id)
|
||||
)
|
||||
)
|
||||
|
||||
# --- Phase 2: run the crawler outside any transaction ---------------------
|
||||
cookie_file = out_dir / ".cookies"
|
||||
_write_cookie_file(cookie_file, cookie)
|
||||
|
||||
request = CrawlerStartRequest(
|
||||
platform=_PLATFORM_ENUM[platform],
|
||||
login_type=LoginTypeEnum.COOKIE,
|
||||
crawler_type=CrawlerTypeEnum.CREATOR if mode == MODE_CREATOR else CrawlerTypeEnum.DETAIL,
|
||||
creator_ids=",".join(urls) if mode == MODE_CREATOR else "",
|
||||
specified_ids=",".join(urls) if mode != MODE_CREATOR else "",
|
||||
start_page=1,
|
||||
enable_comments=enable_comments,
|
||||
enable_sub_comments=strategy["enable_sub_comments"],
|
||||
enable_media=False,
|
||||
save_option=SaveDataOptionEnum.JSONL,
|
||||
cookies="",
|
||||
headless=True,
|
||||
max_notes_count=max_notes_count,
|
||||
max_comments_count=max_comments_count,
|
||||
# Isolate this run's output: the crawler names files by date only, so
|
||||
# otherwise same-day runs would append into one shared file.
|
||||
save_data_path=str(out_dir),
|
||||
# Unattended runs must not try to attach to the user's desktop Chrome.
|
||||
enable_cdp_mode=False,
|
||||
# Only injecting web_session is not enough to sign requests from a cold
|
||||
# browser profile.
|
||||
inject_all_cookies=True,
|
||||
save_login_state=True,
|
||||
cookies_file=str(cookie_file),
|
||||
max_concurrency_num=1,
|
||||
# Strategy + proxy, surfaced on the Settings page.
|
||||
crawler_max_sleep_sec=strategy["crawl_sleep_sec"],
|
||||
enable_ip_proxy=strategy["enable_ip_proxy"],
|
||||
ip_proxy_pool_count=strategy["proxy_pool_count"],
|
||||
ip_proxy_provider_name=strategy["proxy_provider"],
|
||||
static_proxy_url=strategy["static_proxy_url"] or None,
|
||||
)
|
||||
|
||||
# --- Phase 2: run the collection outside any transaction ------------------
|
||||
# 抖音走**进程内 HTTP 客户端**,不起 Playwright 子进程:那边会构造一大串自相矛盾的
|
||||
# 浏览器指纹参数(参数说 Mac + Chrome 125、UA 说 Linux + Chrome 155),网关回一个
|
||||
# 200 + 空 body,然后被翻译成「account blocked」—— 看着像账号被封,其实什么都不是。
|
||||
# 见 douyin_fetch / douyin_api。
|
||||
# 先落「运行中」—— **两条路都要**。原先这一行只写在爬虫那条分支里,于是抖音那条路上
|
||||
# run 一直停在 pending;一旦中途出事(异常、进程被重启),界面上就是一个永远
|
||||
# 「排队中」的幽灵,而且 recover() 也只清理 running、收不到它。
|
||||
async with get_session() as session:
|
||||
run = await session.get(MonitorRun, run_id)
|
||||
if run is not None:
|
||||
run.status = RUN_RUNNING
|
||||
run.started_at = get_current_timestamp()
|
||||
|
||||
try:
|
||||
exit_code = await crawler_manager.run_and_wait(request, timeout=timeout_seconds)
|
||||
finally:
|
||||
_remove_cookie_file(cookie_file)
|
||||
in_process_tail: List[str] = []
|
||||
if platform == adapters.PLATFORM_DY:
|
||||
try:
|
||||
# 也要有超时。爬虫那条路靠 run_and_wait(timeout=...) 兜底,这条路没有子进程、
|
||||
# 没人管 —— 里面**任何一次卡住都会让 run 永远停在「运行中」**(真踩过:
|
||||
# page.evaluate 打在一个卡死的标签页上不返回)。
|
||||
fetched = await asyncio.wait_for(
|
||||
douyin_fetch.collect(
|
||||
out_dir,
|
||||
platform=platform,
|
||||
mode=mode,
|
||||
limit=max_notes_count,
|
||||
want_comments=enable_comments,
|
||||
comment_limit=max_comments_count,
|
||||
targets=targets,
|
||||
known_aweme_ids=known_aweme_ids,
|
||||
cookie=cookie,
|
||||
),
|
||||
timeout=timeout_seconds,
|
||||
)
|
||||
except asyncio.TimeoutError:
|
||||
fetched = {
|
||||
"notes": 0,
|
||||
"comments": 0,
|
||||
"errors": [f"抖音采集超过 {timeout_seconds} 秒仍未完成,已放弃这一轮"],
|
||||
"jsonl_dir": "",
|
||||
}
|
||||
in_process_tail = list(fetched["errors"])
|
||||
# 一条都没采到 = 这一轮失败,并把**真因**当作退出诊断传下去。否则它会掉进
|
||||
# ingest 的「疑似登录失效」分支 —— 又骗人一次,正是这套东西一直在犯的毛病。
|
||||
exit_code = 1 if (fetched["errors"] and not fetched["notes"]) else 0
|
||||
if fetched["errors"] and fetched["notes"]:
|
||||
# 有产物但带着错误,说明走了退化路径(比如作品列表被挡,只刷新了已知作品)。
|
||||
# 这一轮状态是成功,但**不是**一切正常 —— 得留下痕迹,否则没人知道新作品
|
||||
# 其实没在发现。
|
||||
print(
|
||||
"[monitor.runner] 抖音采集部分失败:"
|
||||
+ ";".join(fetched["errors"])[:300]
|
||||
)
|
||||
else:
|
||||
cookie_file = out_dir / ".cookies"
|
||||
_write_cookie_file(cookie_file, cookie)
|
||||
|
||||
request = CrawlerStartRequest(
|
||||
platform=_PLATFORM_ENUM[platform],
|
||||
login_type=LoginTypeEnum.COOKIE,
|
||||
crawler_type=CrawlerTypeEnum.CREATOR if mode == MODE_CREATOR else CrawlerTypeEnum.DETAIL,
|
||||
creator_ids=",".join(urls) if mode == MODE_CREATOR else "",
|
||||
specified_ids=",".join(urls) if mode != MODE_CREATOR else "",
|
||||
start_page=1,
|
||||
enable_comments=enable_comments,
|
||||
enable_sub_comments=strategy["enable_sub_comments"],
|
||||
enable_media=False,
|
||||
save_option=SaveDataOptionEnum.JSONL,
|
||||
cookies="",
|
||||
headless=True,
|
||||
max_notes_count=max_notes_count,
|
||||
max_comments_count=max_comments_count,
|
||||
# Isolate this run's output: the crawler names files by date only, so
|
||||
# otherwise same-day runs would append into one shared file.
|
||||
save_data_path=str(out_dir),
|
||||
# Attach to the browser already running on CDP_DEBUG_PORT when the
|
||||
# operator enabled it; otherwise launch a private, throwaway browser.
|
||||
enable_cdp_mode=cdp_enabled,
|
||||
# Only injecting web_session is not enough to sign requests from a cold
|
||||
# browser profile.
|
||||
inject_all_cookies=True,
|
||||
save_login_state=True,
|
||||
cookies_file=str(cookie_file),
|
||||
max_concurrency_num=1,
|
||||
# Strategy + proxy, surfaced on the Settings page.
|
||||
crawler_max_sleep_sec=strategy["crawl_sleep_sec"],
|
||||
enable_ip_proxy=strategy["enable_ip_proxy"],
|
||||
ip_proxy_pool_count=strategy["proxy_pool_count"],
|
||||
ip_proxy_provider_name=strategy["proxy_provider"],
|
||||
static_proxy_url=strategy["static_proxy_url"] or None,
|
||||
)
|
||||
|
||||
try:
|
||||
exit_code = await crawler_manager.run_and_wait(request, timeout=timeout_seconds)
|
||||
finally:
|
||||
_remove_cookie_file(cookie_file)
|
||||
|
||||
# 失败时的诊断来源。抖音那条路没有子进程,尾巴就是它自己报的错 —— **别去读
|
||||
# crawler_manager 的尾巴**,那里面是上一轮别的平台留下的东西,会张冠李戴。
|
||||
output_tail = (
|
||||
in_process_tail
|
||||
if platform == adapters.PLATFORM_DY
|
||||
else crawler_manager.get_output_tail()
|
||||
)
|
||||
|
||||
# --- Phase 3: ingest ------------------------------------------------------
|
||||
async with get_session() as session:
|
||||
@@ -243,17 +317,27 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
|
||||
if run is None or task is None:
|
||||
raise ValueError(f"Run {run_id} or task {task_id} vanished during execution")
|
||||
|
||||
if exit_code == -1 and not (out_dir / "xhs").exists():
|
||||
# 目录名按平台解析 —— 抖音的平台 id 是 dy 而产物目录是 douyin,写死就永远判不准。
|
||||
if exit_code == -1 and not (out_dir / adapters.artifact_dir(task.platform)).exists():
|
||||
# run_and_wait returns -1 when the process could not start or timed out.
|
||||
run.status = RUN_TIMEOUT
|
||||
run.finished_at = get_current_timestamp()
|
||||
run.exit_code = exit_code
|
||||
run.error_message = "Run was killed by timeout or failed to start"
|
||||
# -1 同时代表「超时」和「根本没起来」,两者要查的东西完全不同。输出尾巴里
|
||||
# 有异常就带上它,否则运行历史里只能看到这句没有信息量的话。
|
||||
cause = diagnose_failure(output_tail)
|
||||
if cause:
|
||||
run.error_message = f"{run.error_message};原因:{cause}"
|
||||
result = IngestResult(status=RUN_TIMEOUT, error=run.error_message)
|
||||
else:
|
||||
run.exit_code = exit_code
|
||||
run.finished_at = get_current_timestamp()
|
||||
result = await ingest_run(session, run, task, out_dir)
|
||||
# 把爬虫输出的尾巴交给 ingest:退出码本身说明不了问题,运行历史里要显示的
|
||||
# 是真正的报错(比如抖音的 DataFetchError: account blocked)。
|
||||
result = await ingest_run(
|
||||
session, run, task, out_dir, output_tail=output_tail
|
||||
)
|
||||
|
||||
# A run that authenticated fine is the only useful signal that the
|
||||
# stored cookie still works.
|
||||
@@ -274,4 +358,15 @@ async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
|
||||
if task is not None and run is not None:
|
||||
await notify.notify_run(session, task, run)
|
||||
|
||||
# --- Phase 5: 封面落盘 ------------------------------------------------------
|
||||
# 也放在事务之外。封面地址带签名、会过期(实测隔天即 403),落盘之后才与签名无关。
|
||||
# 下载慢且可能失败,占着一个入库事务是不合适的;失败也不影响本轮数据。
|
||||
try:
|
||||
async with get_session() as session:
|
||||
saved = await covers.cache_pending(session, task_id)
|
||||
if saved:
|
||||
print(f"[monitor.runner] 缓存了 {saved} 张作品封面")
|
||||
except Exception as exc: # noqa: BLE001 - 封面拿不到不该让整轮失败
|
||||
print(f"[monitor.runner] 封面缓存失败:{exc}")
|
||||
|
||||
return result
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/schedule.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Monitor task schedule arithmetic.
|
||||
|
||||
Three modes, all expressible by a picker. A raw cron string was ruled out on
|
||||
purpose -- it is a small language, and the operator should not have to write one
|
||||
to say "every day at nine":
|
||||
|
||||
* ``interval`` -- every N minutes.
|
||||
* ``daily`` -- at chosen clock times, e.g. 09:00 and 18:30.
|
||||
* ``weekly`` -- at chosen clock times on chosen weekdays, e.g. Mon-Fri 10:00.
|
||||
|
||||
The two clock modes are **fixed-time**, unlike ``interval``, which is fixed-delay.
|
||||
The distinction matters: for an interval, measuring the next slot from when the
|
||||
run starts is what stops a slow run from firing back-to-back. For a clock
|
||||
schedule it would be wrong, because a run that starts at 09:07 would drag every
|
||||
later run seven minutes late, compounding all day.
|
||||
|
||||
Fixed-time also means **no jitter is applied** to clock schedules. The operator
|
||||
picked a time; quietly running at 09:04 instead of 09:00 is not a feature, it just
|
||||
looks like a bug. ``interval`` keeps its jitter, where there is no stated time to
|
||||
contradict.
|
||||
|
||||
All arithmetic is in the server's local timezone -- naive datetimes on purpose,
|
||||
because the container is pinned to the operator's zone via TZ and pretending
|
||||
otherwise would add a timezone concept nobody asked for.
|
||||
"""
|
||||
|
||||
from datetime import datetime, time, timedelta
|
||||
from typing import Optional, Sequence
|
||||
|
||||
MODE_INTERVAL = "interval"
|
||||
MODE_DAILY = "daily"
|
||||
MODE_WEEKLY = "weekly"
|
||||
|
||||
SCHEDULE_MODES = (MODE_INTERVAL, MODE_DAILY, MODE_WEEKLY)
|
||||
CLOCK_MODES = (MODE_DAILY, MODE_WEEKLY)
|
||||
|
||||
_MS_PER_MINUTE = 60_000
|
||||
_WEEKDAY_NAMES = "一二三四五六日"
|
||||
|
||||
|
||||
def parse_hours(raw: Optional[str]) -> list[int]:
|
||||
"""``"9,18"`` -> ``[9, 18]``. Sorted, de-duplicated, junk dropped."""
|
||||
return _parse_int_list(raw, 0, 23)
|
||||
|
||||
|
||||
def parse_days(raw: Optional[str]) -> list[int]:
|
||||
"""``"0,2,4"`` -> ``[0, 2, 4]``. **0 is Monday**, matching ``date.weekday()``."""
|
||||
return _parse_int_list(raw, 0, 6)
|
||||
|
||||
|
||||
def _parse_int_list(raw: Optional[str], low: int, high: int) -> list[int]:
|
||||
values: set[int] = set()
|
||||
for chunk in (raw or "").split(","):
|
||||
chunk = chunk.strip()
|
||||
if not chunk:
|
||||
continue
|
||||
try:
|
||||
number = int(chunk)
|
||||
except ValueError:
|
||||
# Stored values come from our own UI, but a hand-edited row must not
|
||||
# be able to crash the scheduler loop.
|
||||
continue
|
||||
if low <= number <= high:
|
||||
values.add(number)
|
||||
return sorted(values)
|
||||
|
||||
|
||||
def format_hours(hours: Sequence[int]) -> str:
|
||||
return ",".join(str(hour) for hour in sorted(set(hours)))
|
||||
|
||||
|
||||
def format_days(days: Sequence[int]) -> str:
|
||||
return ",".join(str(day) for day in sorted(set(days)))
|
||||
|
||||
|
||||
def describe(
|
||||
*,
|
||||
mode: str,
|
||||
interval_minutes: int,
|
||||
hours: Sequence[int],
|
||||
days: Sequence[int],
|
||||
minute: int,
|
||||
) -> str:
|
||||
"""One human sentence for the task card.
|
||||
|
||||
Lives here rather than in the frontend so the list view and the editor cannot
|
||||
drift apart on what a schedule means.
|
||||
"""
|
||||
if mode == MODE_INTERVAL:
|
||||
if interval_minutes % 1440 == 0:
|
||||
return f"每 {interval_minutes // 1440} 天"
|
||||
if interval_minutes % 60 == 0:
|
||||
return f"每 {interval_minutes // 60} 小时"
|
||||
return f"每 {interval_minutes} 分钟"
|
||||
|
||||
if not hours:
|
||||
return "未设置时间"
|
||||
|
||||
clock = "、".join(f"{hour:02d}:{minute:02d}" for hour in sorted(set(hours)))
|
||||
|
||||
if mode == MODE_DAILY:
|
||||
return f"每天 {clock}"
|
||||
|
||||
if not days:
|
||||
return f"每天 {clock}"
|
||||
labels = "、".join(f"周{_WEEKDAY_NAMES[day]}" for day in sorted(set(days)))
|
||||
return f"{labels} {clock}"
|
||||
|
||||
|
||||
def next_occurrence(
|
||||
*,
|
||||
mode: str,
|
||||
interval_minutes: int,
|
||||
hours: Sequence[int],
|
||||
days: Sequence[int],
|
||||
minute: int,
|
||||
after_ms: int,
|
||||
) -> Optional[int]:
|
||||
"""The next fire time strictly after ``after_ms``, as epoch milliseconds.
|
||||
|
||||
``None`` means the schedule can never fire -- a clock mode with no hours
|
||||
chosen. Callers store that as "no next run" rather than something in the past,
|
||||
which would otherwise leave the task permanently due and re-running on every
|
||||
tick.
|
||||
"""
|
||||
if mode == MODE_INTERVAL:
|
||||
return after_ms + max(1, interval_minutes) * _MS_PER_MINUTE
|
||||
|
||||
if not hours:
|
||||
return None
|
||||
|
||||
now = datetime.fromtimestamp(after_ms / 1000)
|
||||
# No weekdays chosen means every day, matching describe(). Without the `days`
|
||||
# guard an empty selection would produce an empty allowed set, no matching day,
|
||||
# and a task that silently never runs.
|
||||
allowed_days = set(days) if (mode == MODE_WEEKLY and days) else set(range(7))
|
||||
|
||||
# Eight days of lookahead covers today's remaining slots plus a full week,
|
||||
# which is more than any weekday selection can need.
|
||||
for offset in range(8):
|
||||
day = (now + timedelta(days=offset)).date()
|
||||
if day.weekday() not in allowed_days:
|
||||
continue
|
||||
for hour in sorted(set(hours)):
|
||||
candidate = datetime.combine(day, time(hour=hour, minute=minute))
|
||||
if candidate > now:
|
||||
return int(candidate.timestamp() * 1000)
|
||||
|
||||
return None
|
||||
+120
-25
@@ -23,8 +23,20 @@ is enough here: there is exactly one process, one global crawler subprocess, and
|
||||
therefore no concurrency to coordinate -- a cron-style library would add a
|
||||
dependency without adding a capability.
|
||||
|
||||
Scheduling is **fixed-delay**, not fixed-rate: ``next_run_at`` is set from the
|
||||
moment a run starts, so a slow run cannot make its task fire back-to-back.
|
||||
Two families of schedule, and the difference matters:
|
||||
|
||||
* ``interval`` is **fixed-delay**, not fixed-rate -- ``next_run_at`` is measured
|
||||
from the moment a run starts, so a slow run cannot make its task fire
|
||||
back-to-back.
|
||||
* the clock modes (``daily``/``weekly``) are **fixed-time** -- recomputed from the
|
||||
calendar, so a run that starts late does not drag every later run with it.
|
||||
|
||||
The arithmetic for both lives in schedule.py.
|
||||
|
||||
The loop also carries the 上游更新检查: it is not a crawl, so it shares none of
|
||||
the rules above (no subprocess, no active-hours gate) -- see
|
||||
``_maybe_check_upstream``. It rides this loop rather than getting a thread of its
|
||||
own because it is one HTTP-shaped fetch per day.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
@@ -37,18 +49,23 @@ from sqlalchemy import select
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from ..services import crawler_manager
|
||||
from . import app_settings
|
||||
from . import app_settings, schedule, upstream
|
||||
from .db import get_session
|
||||
from .models import MonitorRun, MonitorTask, RUN_INTERRUPTED, RUN_RUNNING
|
||||
from .models import (
|
||||
RUN_INTERRUPTED,
|
||||
RUN_PENDING,
|
||||
RUN_RUNNING,
|
||||
MonitorRun,
|
||||
MonitorTask,
|
||||
)
|
||||
from .runner import execute_task
|
||||
from .settings import get_cookie
|
||||
|
||||
POLL_INTERVAL_SECONDS = 20
|
||||
# Spread tasks sharing an interval so they do not all come due on the same tick.
|
||||
# Applied to interval mode only -- see the advance step below.
|
||||
JITTER_SECONDS = 60
|
||||
|
||||
_MS_PER_MINUTE = 60_000
|
||||
|
||||
|
||||
class MonitorScheduler:
|
||||
"""Polls the task table and runs whatever is due."""
|
||||
@@ -56,8 +73,9 @@ class MonitorScheduler:
|
||||
def __init__(self) -> None:
|
||||
self._loop_task: Optional[asyncio.Task] = None
|
||||
self._stopping = asyncio.Event()
|
||||
# Avoids logging "no cookie" on every single tick.
|
||||
self._warned_no_cookie = False
|
||||
# Avoids logging "no cookie" on every single tick. Per platform, because
|
||||
# warning once for Xiaohongshu must not silence the warning for Douyin.
|
||||
self._warned_no_cookie: set = set()
|
||||
|
||||
async def start(self) -> None:
|
||||
if self._loop_task is not None and not self._loop_task.done():
|
||||
@@ -86,19 +104,66 @@ class MonitorScheduler:
|
||||
await self.tick()
|
||||
except Exception as exc: # pragma: no cover - keep the loop alive
|
||||
print(f"[monitor.scheduler] tick failed: {exc}")
|
||||
# 独立于采集任务,因此单独一段 try:上游检查失败不该影响采集调度,
|
||||
# 反过来也一样。
|
||||
try:
|
||||
await self._maybe_check_upstream()
|
||||
except Exception as exc: # pragma: no cover - keep the loop alive
|
||||
print(f"[monitor.scheduler] upstream check failed: {exc}")
|
||||
await asyncio.sleep(POLL_INTERVAL_SECONDS)
|
||||
|
||||
async def _maybe_check_upstream(self) -> None:
|
||||
"""到点就 fetch 一次上游仓库,看它有没有新提交。
|
||||
|
||||
与采集任务的三条规则都不同,各有理由:它不碰浏览器、也不占采集子进程,
|
||||
所以不看 ``is_busy``;它只发一个 git 请求,没有被平台风控的风险,所以也不
|
||||
受活跃时段限制 —— 定时检查放在半夜反而是最合适的。
|
||||
"""
|
||||
async with get_session() as session:
|
||||
if not await app_settings.get_value(
|
||||
session, "upstream_check_enabled", fallback=False
|
||||
):
|
||||
return
|
||||
interval_minutes = int(
|
||||
await app_settings.get_value(
|
||||
session, "upstream_check_interval_minutes", fallback=1440
|
||||
)
|
||||
)
|
||||
state = await upstream.load_state(session)
|
||||
|
||||
checked_at = int(state.get("checked_at") or 0)
|
||||
now = get_current_timestamp()
|
||||
# 失败也会写 checked_at,所以不通的时候同样是每个间隔重试一次,
|
||||
# 而不是每个 tick(20 秒)都去撞一次墙。
|
||||
if checked_at and now - checked_at < max(1, interval_minutes) * 60_000:
|
||||
return
|
||||
|
||||
result = await upstream.run_check()
|
||||
if result.get("behind"):
|
||||
print(
|
||||
f"[monitor.scheduler] 上游 {result.get('branch')} 领先 "
|
||||
f"{result['behind']} 个提交"
|
||||
)
|
||||
elif not result.get("ok"):
|
||||
print(f"[monitor.scheduler] 上游检查失败:{result.get('error')}")
|
||||
|
||||
async def recover(self) -> None:
|
||||
"""Clean up state left behind by a server restart.
|
||||
|
||||
A run still marked ``running`` cannot be running -- its subprocess died
|
||||
with the previous process. Marking it interrupted stops it from blocking
|
||||
the UI as a phantom in-flight run.
|
||||
|
||||
**``pending`` 同样是残留**:那一行是上一轮建的,可它后面的采集根本没机会开始
|
||||
(进程被重启,或者采集那条路抛了异常),所以它永远不会自己往前走。只清 running
|
||||
的话,它会永远挂在界面上显示「排队中」—— 用户看到的就是任务卡住了。
|
||||
"""
|
||||
async with get_session() as session:
|
||||
stale = (
|
||||
await session.scalars(
|
||||
select(MonitorRun).where(MonitorRun.status == RUN_RUNNING)
|
||||
select(MonitorRun).where(
|
||||
MonitorRun.status.in_((RUN_RUNNING, RUN_PENDING))
|
||||
)
|
||||
)
|
||||
).all()
|
||||
for run in stale:
|
||||
@@ -151,28 +216,58 @@ class MonitorScheduler:
|
||||
if task is None:
|
||||
return
|
||||
|
||||
# No cookie means every run would report an auth failure. Leave the
|
||||
# task due rather than advancing: it starts working the moment the
|
||||
# user pastes one.
|
||||
cookie = await get_cookie(session)
|
||||
# 没有 cookie 就跳过,是为了不让任务每轮白跑一趟出个认证失败。任务留在
|
||||
# due 状态而不推进 —— 用户一粘上 cookie 它就能自己跑起来。
|
||||
#
|
||||
# **但开着 CDP 时必须放行**:那种模式下登录态来自被接管的那个浏览器,
|
||||
# 粘不粘 cookie 根本轮不到它决定成败。不放行的话,选了「接管已有 Chrome」
|
||||
# 却没粘 cookie 的用户会发现任务永远不被触发,而且什么错都不报。
|
||||
cookie = await get_cookie(session, task.platform)
|
||||
if not cookie:
|
||||
if not self._warned_no_cookie:
|
||||
print(
|
||||
"[monitor.scheduler] no XHS cookie configured; "
|
||||
"scheduled tasks will not run until one is set"
|
||||
)
|
||||
self._warned_no_cookie = True
|
||||
return
|
||||
self._warned_no_cookie = False
|
||||
cdp_enabled = await app_settings.get_value(
|
||||
session, "cdp_enabled", fallback=False
|
||||
)
|
||||
if not cdp_enabled:
|
||||
if task.platform not in self._warned_no_cookie:
|
||||
print(
|
||||
f"[monitor.scheduler] no {task.platform} cookie configured; "
|
||||
"scheduled tasks will not run until one is set or CDP is enabled"
|
||||
)
|
||||
self._warned_no_cookie.add(task.platform)
|
||||
return
|
||||
self._warned_no_cookie.discard(task.platform)
|
||||
|
||||
# Advance before running so a crash mid-run cannot cause an immediate
|
||||
# re-fire, and so a long outage coalesces into a single run instead
|
||||
# of one run per missed interval.
|
||||
task.next_run_at = (
|
||||
get_current_timestamp()
|
||||
+ task.interval_minutes * _MS_PER_MINUTE
|
||||
+ random.randint(0, JITTER_SECONDS) * 1000
|
||||
now = get_current_timestamp()
|
||||
following = schedule.next_occurrence(
|
||||
mode=task.schedule_mode,
|
||||
interval_minutes=task.interval_minutes,
|
||||
hours=schedule.parse_hours(task.schedule_hours),
|
||||
days=schedule.parse_days(task.schedule_days),
|
||||
minute=task.schedule_minute,
|
||||
after_ms=now,
|
||||
)
|
||||
|
||||
if following is None:
|
||||
# A clock schedule with no times can never fire. The API rejects
|
||||
# that shape, so this guards against a hand-edited row: park the
|
||||
# task with no next run rather than leaving it permanently due and
|
||||
# re-running it on every tick.
|
||||
task.next_run_at = None
|
||||
print(
|
||||
f"[monitor.scheduler] task {task.id} has no usable schedule "
|
||||
f"and will not run until one is set"
|
||||
)
|
||||
elif task.schedule_mode == schedule.MODE_INTERVAL:
|
||||
# Jitter belongs to the interval mode only. Spreading identical
|
||||
# intervals apart is the point; nudging a time the operator
|
||||
# explicitly picked is not -- it just looks like a broken clock.
|
||||
task.next_run_at = following + random.randint(0, JITTER_SECONDS) * 1000
|
||||
else:
|
||||
task.next_run_at = following
|
||||
|
||||
task_id = task.id
|
||||
|
||||
try:
|
||||
|
||||
+398
-26
@@ -19,8 +19,7 @@
|
||||
"""Task CRUD and dashboard queries for the monitoring layer."""
|
||||
|
||||
import asyncio
|
||||
import re
|
||||
from typing import Any, Dict, List, Optional
|
||||
from typing import Any, Dict, List, Optional, Sequence
|
||||
from urllib.parse import parse_qs, urlparse
|
||||
|
||||
from sqlalchemy import delete, func, select
|
||||
@@ -28,15 +27,18 @@ from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from . import app_settings, platforms
|
||||
from . import adapters, app_settings, covers, platforms, schedule
|
||||
from .db import get_session
|
||||
from .platforms import PLATFORM_XHS
|
||||
from .models import (
|
||||
MODE_CREATOR,
|
||||
MODE_NOTE,
|
||||
MonitorComment,
|
||||
MonitorCreatorAlias,
|
||||
MonitorCreatorStat,
|
||||
MonitorEvent,
|
||||
MonitorNote,
|
||||
MonitorNoteAlias,
|
||||
MonitorNoteMetric,
|
||||
MonitorRun,
|
||||
MonitorTarget,
|
||||
@@ -53,11 +55,30 @@ _background_runs: set[asyncio.Task] = set()
|
||||
MIN_INTERVAL_MINUTES = 30
|
||||
MAX_INTERVAL_MINUTES = 7 * 24 * 60
|
||||
|
||||
_CREATOR_URL_RE = re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)")
|
||||
_NOTE_URL_RE = re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)")
|
||||
# XHS user ids and note ids are 24-char hex; allow a slightly wider range so a
|
||||
# format change degrades into "still accepted" rather than "rejected".
|
||||
_BARE_ID_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
|
||||
# Changing any of these invalidates the pending run slot.
|
||||
SCHEDULE_FIELDS = {
|
||||
"interval_minutes",
|
||||
"schedule_mode",
|
||||
"schedule_hours",
|
||||
"schedule_days",
|
||||
"schedule_minute",
|
||||
}
|
||||
|
||||
|
||||
def _require_clock_fields(mode: str, hours: list, days: list) -> None:
|
||||
"""A clock schedule with no clock time can never fire.
|
||||
|
||||
The create schema already rejects that shape, but the update path merges
|
||||
partial fields and therefore has no schema-level view of the result -- so the
|
||||
check lives here, where both paths meet.
|
||||
"""
|
||||
if mode in schedule.CLOCK_MODES and not hours:
|
||||
raise ValueError("按钟点调度至少要选一个时间")
|
||||
if mode == schedule.MODE_WEEKLY and not days:
|
||||
raise ValueError("按周调度至少要选一个星期")
|
||||
|
||||
# 各平台的链接形态、id 形状、短链域名都在 adapters.py —— 那里是「平台之间不一样」
|
||||
# 的东西的唯一出处,所以这里不再留任何平台字面量。
|
||||
|
||||
|
||||
class TargetParseError(ValueError):
|
||||
@@ -73,27 +94,36 @@ def parse_target_input(
|
||||
Storing the id separately from the token is what keeps a long-running task
|
||||
alive: tokens expire, ids do not.
|
||||
|
||||
URL shapes are platform-specific. Only Xiaohongshu is wired, so anything else
|
||||
is rejected here as well as at task creation -- parsing a Douyin link as if it
|
||||
were a Xiaohongshu one would be worse than refusing it.
|
||||
URL shapes are platform-specific and come from ``adapters``. A platform with
|
||||
no adapter is rejected here as well as at task creation -- parsing a Douyin
|
||||
link as if it were a Xiaohongshu one would be worse than refusing it.
|
||||
"""
|
||||
if platform != PLATFORM_XHS:
|
||||
if not adapters.has_adapter(platform):
|
||||
raise TargetParseError(f"暂不支持解析该平台({platform})的目标链接")
|
||||
spec = adapters.adapter(platform)
|
||||
|
||||
raw = (value or "").strip()
|
||||
if not raw:
|
||||
raise TargetParseError("Empty target")
|
||||
|
||||
creator_mode = mode == MODE_CREATOR
|
||||
expected = "博主主页" if creator_mode else "笔记"
|
||||
|
||||
external_id = ""
|
||||
if raw.startswith("http") or "/" in raw:
|
||||
# xhslink.com and other short links are not resolvable without a network
|
||||
# round-trip, so only the direct profile/explore forms are supported.
|
||||
match = _CREATOR_URL_RE.search(raw) if mode == MODE_CREATOR else _NOTE_URL_RE.search(raw)
|
||||
if not match:
|
||||
expected = "博主主页" if mode == MODE_CREATOR else "笔记"
|
||||
if any(host in raw for host in spec.short_link_hosts):
|
||||
# 短链要联网跳一次才知道指向谁,而这里没有网络可跳。明确拒绝好过存一个
|
||||
# 解析不出 id 的值 —— 那会变成一个永远抓不到东西、还不报错的任务。
|
||||
raise TargetParseError(f"{expected}短链无法解析,请粘贴完整链接:{raw}")
|
||||
patterns = spec.creator_url_res if creator_mode else spec.note_url_res
|
||||
for pattern in patterns:
|
||||
match = pattern.search(raw)
|
||||
if match:
|
||||
external_id = match.group(1)
|
||||
break
|
||||
if not external_id:
|
||||
raise TargetParseError(f"无法从链接中解析出{expected} ID:{raw}")
|
||||
external_id = match.group(1)
|
||||
elif _BARE_ID_RE.match(raw):
|
||||
elif (spec.creator_bare_re if creator_mode else spec.note_bare_re).match(raw):
|
||||
external_id = raw
|
||||
else:
|
||||
raise TargetParseError(f"无法识别的目标:{raw}")
|
||||
@@ -139,7 +169,12 @@ async def create_task(session: AsyncSession, payload: Dict[str, Any]) -> Monitor
|
||||
# the Settings page actually governs new tasks.
|
||||
defaults = await app_settings.defaults(session, platform)
|
||||
interval_minutes = payload.get("interval_minutes") or defaults["interval_minutes"]
|
||||
interval_ms = int(interval_minutes) * 60_000
|
||||
|
||||
schedule_mode = payload.get("schedule_mode") or schedule.MODE_INTERVAL
|
||||
schedule_hours = list(payload.get("schedule_hours") or [])
|
||||
schedule_days = list(payload.get("schedule_days") or [])
|
||||
schedule_minute = int(payload.get("schedule_minute") or 0)
|
||||
_require_clock_fields(schedule_mode, schedule_hours, schedule_days)
|
||||
|
||||
task = MonitorTask(
|
||||
name=payload["name"],
|
||||
@@ -147,12 +182,24 @@ async def create_task(session: AsyncSession, payload: Dict[str, Any]) -> Monitor
|
||||
mode=mode,
|
||||
enabled=payload.get("enabled", True),
|
||||
interval_minutes=interval_minutes,
|
||||
schedule_mode=schedule_mode,
|
||||
schedule_hours=schedule.format_hours(schedule_hours),
|
||||
schedule_days=schedule.format_days(schedule_days),
|
||||
schedule_minute=schedule_minute,
|
||||
max_notes_count=payload.get("max_notes_count") or defaults["max_notes_count"],
|
||||
enable_comments=payload.get("enable_comments", True),
|
||||
max_comments_count=payload.get("max_comments_count") or defaults["max_comments_count"],
|
||||
run_timeout_seconds=payload.get("run_timeout_seconds", 3600),
|
||||
notify_enabled=payload.get("notify_enabled", False),
|
||||
next_run_at=now + interval_ms,
|
||||
notify_failures=payload.get("notify_failures", True),
|
||||
next_run_at=schedule.next_occurrence(
|
||||
mode=schedule_mode,
|
||||
interval_minutes=interval_minutes,
|
||||
hours=schedule_hours,
|
||||
days=schedule_days,
|
||||
minute=schedule_minute,
|
||||
after_ms=now,
|
||||
),
|
||||
last_status="idle",
|
||||
created_at=now,
|
||||
updated_at=now,
|
||||
@@ -189,19 +236,31 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
|
||||
if task is None:
|
||||
raise ValueError(f"Task {task_id} not found")
|
||||
|
||||
was_enabled = task.enabled
|
||||
|
||||
for field in (
|
||||
"name",
|
||||
"enabled",
|
||||
"interval_minutes",
|
||||
"schedule_mode",
|
||||
"schedule_minute",
|
||||
"max_notes_count",
|
||||
"enable_comments",
|
||||
"max_comments_count",
|
||||
"run_timeout_seconds",
|
||||
"notify_enabled",
|
||||
"notify_failures",
|
||||
):
|
||||
if field in payload and payload[field] is not None:
|
||||
setattr(task, field, payload[field])
|
||||
|
||||
# The clock lists are stored as comma-separated text, so they cannot go
|
||||
# through the generic loop above.
|
||||
if payload.get("schedule_hours") is not None:
|
||||
task.schedule_hours = schedule.format_hours(payload["schedule_hours"])
|
||||
if payload.get("schedule_days") is not None:
|
||||
task.schedule_days = schedule.format_days(payload["schedule_days"])
|
||||
|
||||
# Replacing targets resets the baseline implicitly: a note set that now
|
||||
# includes new ids will simply report them as new on the next run.
|
||||
if payload.get("targets") is not None:
|
||||
@@ -209,7 +268,7 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
|
||||
now = get_current_timestamp()
|
||||
seen: set[str] = set()
|
||||
for value in payload["targets"]:
|
||||
parsed = parse_target_input(value, task.mode)
|
||||
parsed = parse_target_input(value, task.mode, task.platform)
|
||||
if parsed["external_id"] in seen:
|
||||
continue
|
||||
seen.add(parsed["external_id"])
|
||||
@@ -227,8 +286,23 @@ async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, An
|
||||
)
|
||||
)
|
||||
|
||||
if "interval_minutes" in payload and payload["interval_minutes"]:
|
||||
task.next_run_at = get_current_timestamp() + payload["interval_minutes"] * 60_000
|
||||
# Any change to when the task runs invalidates the pending slot, so recompute
|
||||
# it from the merged state rather than working out which field moved.
|
||||
# Re-enabling counts as a change too: otherwise a task switched off for a
|
||||
# month comes back holding a next_run_at a month in the past and fires the
|
||||
# instant it is saved.
|
||||
if (SCHEDULE_FIELDS & set(payload)) or (task.enabled and not was_enabled):
|
||||
hours = schedule.parse_hours(task.schedule_hours)
|
||||
days = schedule.parse_days(task.schedule_days)
|
||||
_require_clock_fields(task.schedule_mode, hours, days)
|
||||
task.next_run_at = schedule.next_occurrence(
|
||||
mode=task.schedule_mode,
|
||||
interval_minutes=task.interval_minutes,
|
||||
hours=hours,
|
||||
days=days,
|
||||
minute=task.schedule_minute,
|
||||
after_ms=get_current_timestamp(),
|
||||
)
|
||||
|
||||
task.updated_at = get_current_timestamp()
|
||||
await session.flush()
|
||||
@@ -257,6 +331,48 @@ def trigger_manual_run(task_id: int) -> None:
|
||||
# Dashboard queries
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
async def _creator_alias_map(session: AsyncSession) -> Dict[tuple, str]:
|
||||
"""``(platform, creator_hash) -> 备注``。
|
||||
|
||||
整体读一次再在内存里取,而不是每条作品查一次 —— 作品列表动辄几十条。备注本身很少,
|
||||
表不会大。
|
||||
"""
|
||||
rows = (await session.scalars(select(MonitorCreatorAlias))).all()
|
||||
return {(row.platform, row.creator_hash): row.alias for row in rows if row.alias}
|
||||
|
||||
|
||||
async def set_creator_alias(
|
||||
session: AsyncSession, platform: str, creator_hash: str, alias: str
|
||||
) -> None:
|
||||
"""给博主起备注;传空串就是删掉这条备注(界面上的"清空")。"""
|
||||
alias = (alias or "").strip()[:128]
|
||||
existing = await session.scalar(
|
||||
select(MonitorCreatorAlias).where(
|
||||
MonitorCreatorAlias.platform == platform,
|
||||
MonitorCreatorAlias.creator_hash == creator_hash,
|
||||
)
|
||||
)
|
||||
|
||||
if not alias:
|
||||
if existing is not None:
|
||||
await session.delete(existing)
|
||||
return
|
||||
|
||||
if existing is None:
|
||||
session.add(
|
||||
MonitorCreatorAlias(
|
||||
platform=platform,
|
||||
creator_hash=creator_hash,
|
||||
alias=alias,
|
||||
updated_at=get_current_timestamp(),
|
||||
)
|
||||
)
|
||||
return
|
||||
|
||||
existing.alias = alias
|
||||
existing.updated_at = get_current_timestamp()
|
||||
|
||||
|
||||
async def _latest_successful_run_id(session: AsyncSession, task_id: int) -> Optional[int]:
|
||||
return await session.scalar(
|
||||
select(MonitorRun.id)
|
||||
@@ -275,6 +391,75 @@ def _delta(current: Optional[int], previous: Optional[int]) -> Optional[int]:
|
||||
return current - previous
|
||||
|
||||
|
||||
async def _note_alias_map(session: AsyncSession) -> Dict[tuple, str]:
|
||||
"""``(platform, note_id) -> 作品备注``。和博主备注一个道理,整体读一次。"""
|
||||
rows = (await session.scalars(select(MonitorNoteAlias))).all()
|
||||
return {(row.platform, row.note_id): row.alias for row in rows if row.alias}
|
||||
|
||||
|
||||
async def set_note_alias(
|
||||
session: AsyncSession, platform: str, note_id: str, alias: str
|
||||
) -> None:
|
||||
"""给作品起备注;空串就是删掉这条备注。"""
|
||||
alias = (alias or "").strip()[:128]
|
||||
existing = await session.scalar(
|
||||
select(MonitorNoteAlias).where(
|
||||
MonitorNoteAlias.platform == platform,
|
||||
MonitorNoteAlias.note_id == note_id,
|
||||
)
|
||||
)
|
||||
|
||||
if not alias:
|
||||
if existing is not None:
|
||||
await session.delete(existing)
|
||||
return
|
||||
|
||||
if existing is None:
|
||||
session.add(
|
||||
MonitorNoteAlias(
|
||||
platform=platform,
|
||||
note_id=note_id,
|
||||
alias=alias,
|
||||
updated_at=get_current_timestamp(),
|
||||
)
|
||||
)
|
||||
return
|
||||
|
||||
existing.alias = alias
|
||||
existing.updated_at = get_current_timestamp()
|
||||
|
||||
|
||||
async def _latest_creator_stats(
|
||||
session: AsyncSession,
|
||||
task_ids: Optional[Sequence[int]] = None,
|
||||
creator_hashes: Optional[Sequence[str]] = None,
|
||||
) -> Dict[tuple, "MonitorCreatorStat"]:
|
||||
"""``(task_id, creator_hash) -> 最近一条``账号级快照。
|
||||
|
||||
两个参数都是**可选过滤**:传 None 就是不限。``list_notes`` 两个都给(只要手里
|
||||
这批作品涉及的博主),``list_creators`` 只给任务(要的是全部博主,包括一条作品
|
||||
都没有的那些)。
|
||||
|
||||
按 run_id 而不是 captured_at 取「最近」:和作品指标用的是同一个口径,两者放一起
|
||||
看才不会出现「作品数据来自第 8 轮、粉丝数来自第 9 轮」这种对不上的情况。
|
||||
"""
|
||||
if task_ids is not None and not task_ids:
|
||||
return {}
|
||||
if creator_hashes is not None and not creator_hashes:
|
||||
return {}
|
||||
|
||||
query = select(MonitorCreatorStat).order_by(MonitorCreatorStat.run_id.desc())
|
||||
if task_ids is not None:
|
||||
query = query.where(MonitorCreatorStat.task_id.in_(list(task_ids)))
|
||||
if creator_hashes is not None:
|
||||
query = query.where(MonitorCreatorStat.creator_hash.in_(list(creator_hashes)))
|
||||
|
||||
latest: Dict[tuple, MonitorCreatorStat] = {}
|
||||
for row in (await session.scalars(query)).all():
|
||||
latest.setdefault((row.task_id, row.creator_hash), row) # 已按 run_id 倒序
|
||||
return latest
|
||||
|
||||
|
||||
async def list_notes(
|
||||
session: AsyncSession,
|
||||
task_id: Optional[int] = None,
|
||||
@@ -316,6 +501,29 @@ async def list_notes(
|
||||
latest_run_ids: Dict[int, Optional[int]] = {}
|
||||
result: List[Dict[str, Any]] = []
|
||||
|
||||
# 博主的备注。键是 (platform, creator_hash) —— 作品的 platform 挂在它的任务上。
|
||||
task_platform = {
|
||||
row.id: row.platform
|
||||
for row in (
|
||||
await session.execute(
|
||||
select(MonitorTask.id, MonitorTask.platform).where(
|
||||
MonitorTask.id.in_({note.task_id for note in notes})
|
||||
)
|
||||
)
|
||||
).all()
|
||||
}
|
||||
aliases = await _creator_alias_map(session)
|
||||
note_aliases = await _note_alias_map(session)
|
||||
# 账号级指标(粉丝 / 总获赞 / 作品数)。**挂在作品上一起返回**,因为界面上就是按博主
|
||||
# 归组显示的 —— 让前端为了一个组头再发一轮请求没道理。同一个博主的所有作品拿到的是
|
||||
# 同一条(键里带 task_id,所以跨任务不会串)。没有的(小红书那条路不产生它)就是 null,
|
||||
# 前端据此整块不显示,而不是显示一个 0。
|
||||
creator_stats = await _latest_creator_stats(
|
||||
session,
|
||||
[note.task_id for note in notes],
|
||||
[note.creator_hash for note in notes if note.creator_hash],
|
||||
)
|
||||
|
||||
for note in notes:
|
||||
series = by_note.get(note.note_id, [])
|
||||
current = series[0] if series else None
|
||||
@@ -327,13 +535,42 @@ async def list_notes(
|
||||
if note.first_seen_run_id != latest_run_ids[note.task_id]:
|
||||
continue
|
||||
|
||||
stat = creator_stats.get((note.task_id, note.creator_hash))
|
||||
|
||||
result.append(
|
||||
{
|
||||
"task_id": note.task_id,
|
||||
"note_id": note.note_id,
|
||||
"title": note.title,
|
||||
"note_url": note.note_url,
|
||||
"cover": note.cover,
|
||||
# 作品的发布时间(爬虫侧:小红书叫 time、抖音叫 create_time)。
|
||||
# 和 first_seen_at 不是一回事 —— 那是**我们第一次看到它**的时间;把一个
|
||||
# 早就存在的作品加进监控时,两者能差好几个月。可能为 null(平台没给,
|
||||
# 或者值解析不出来),所以前端要能显示成「—」。
|
||||
"published_at": note.published_at,
|
||||
# 按博主分组用。creator_hash 是唯一稳定的创作者标识(原始 user_id
|
||||
# 被爬虫刻意匿名化了),creator_name 是昵称本身 —— 本仓库关掉了脱敏
|
||||
# (见 config.MASK_NICKNAME),所以就是原文。
|
||||
"creator_hash": note.creator_hash,
|
||||
"creator_name": note.creator_name,
|
||||
# 人自己起的备注,界面上优先显示它 —— 昵称认不出是谁,哈希更认不出。
|
||||
"creator_alias": aliases.get(
|
||||
(task_platform.get(note.task_id, ""), note.creator_hash), ""
|
||||
),
|
||||
# 这条作品自己的备注。和博主备注是两件事:博主备注回答"这是谁",它回答
|
||||
# "这条我要盯着"。
|
||||
"note_alias": note_aliases.get(
|
||||
(task_platform.get(note.task_id, ""), note.note_id), ""
|
||||
),
|
||||
# 博主账号级指标 —— 作品列表给不了的东西。三个值都可能为 null(平台没采
|
||||
# 到、或者这条作品来自不产生它的数据源),前端据此整块不画。
|
||||
"creator_fans": stat.fans if stat else None,
|
||||
"creator_total_favorited": stat.total_favorited if stat else None,
|
||||
"creator_works": stat.works_count if stat else None,
|
||||
"creator_stats_at": stat.captured_at if stat else None,
|
||||
# 优先给本地缓存地址:远程地址带签名、会过期(实测隔天即 403),
|
||||
# 本地那份不会。没有缓存时才退回远程,至少让图先显示出来。
|
||||
"cover": covers.cover_url(note.note_id, note.cover),
|
||||
"first_seen_at": note.first_seen_at,
|
||||
"last_seen_at": note.last_seen_at,
|
||||
"is_new": note.first_seen_run_id == latest_run_ids.get(note.task_id),
|
||||
@@ -368,6 +605,107 @@ async def list_notes(
|
||||
return result
|
||||
|
||||
|
||||
async def list_creators(
|
||||
session: AsyncSession,
|
||||
task_id: Optional[int] = None,
|
||||
platform: Optional[str] = None,
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""作品栏里要显示的**博主** —— **包括一条作品都没有的**。
|
||||
|
||||
分组原先是从作品推出来的(按作品的 creator_hash 归组),于是没有作品的博主根本
|
||||
不会出现在列表里:目标加了、资料也采到了、粉丝数就躺在库里,界面上什么都看不见。
|
||||
而「这个号在涨粉、只是最近没发作品」恰恰是最该看见的一种情况 —— 藏起来正好藏反了。
|
||||
|
||||
所以来源换成 **账号快照 ∪ 作品**:
|
||||
|
||||
* 有快照没作品 → 一个 0 篇的组,粉丝数照常显示;
|
||||
* 有作品没快照 → 一个没有账号指标的组(小红书那条路不产生快照,就是这种情况)。
|
||||
|
||||
``creator_alias`` 从作品备注那张表来;``note_count`` / ``last_activity_at`` 用来
|
||||
排序,让最近还在动的博主排在前面。
|
||||
"""
|
||||
scope: Optional[List[int]] = None
|
||||
if task_id is not None:
|
||||
scope = [task_id]
|
||||
elif platform is not None:
|
||||
scope = await platform_task_ids(session, platform)
|
||||
if not scope:
|
||||
return []
|
||||
|
||||
# 作品一侧:谁有作品、有几篇、最后一次是什么时候。
|
||||
work_query = (
|
||||
select(
|
||||
MonitorNote.task_id,
|
||||
MonitorNote.creator_hash,
|
||||
func.count().label("note_count"),
|
||||
func.max(MonitorNote.last_seen_at).label("last_seen_at"),
|
||||
func.max(MonitorNote.creator_name).label("creator_name"),
|
||||
)
|
||||
.where(MonitorNote.creator_hash != "")
|
||||
.group_by(MonitorNote.task_id, MonitorNote.creator_hash)
|
||||
)
|
||||
if scope is not None:
|
||||
work_query = work_query.where(MonitorNote.task_id.in_(scope))
|
||||
|
||||
work: Dict[tuple, Dict[str, Any]] = {}
|
||||
for row in (await session.execute(work_query)).all():
|
||||
work[(row.task_id, row.creator_hash)] = {
|
||||
"note_count": row.note_count,
|
||||
"last_seen_at": row.last_seen_at,
|
||||
"creator_name": row.creator_name or "",
|
||||
}
|
||||
|
||||
stats = await _latest_creator_stats(session, task_ids=scope)
|
||||
if not work and not stats:
|
||||
return []
|
||||
|
||||
# 备注是按 (platform, creator_hash) 存的,所以要知道每个博主属于哪个平台。
|
||||
involved = {key[0] for key in set(work) | set(stats)}
|
||||
task_platform = {
|
||||
row.id: row.platform
|
||||
for row in (
|
||||
await session.execute(
|
||||
select(MonitorTask.id, MonitorTask.platform).where(
|
||||
MonitorTask.id.in_(list(involved))
|
||||
)
|
||||
)
|
||||
).all()
|
||||
}
|
||||
aliases = await _creator_alias_map(session)
|
||||
|
||||
result: List[Dict[str, Any]] = []
|
||||
for key in set(work) | set(stats):
|
||||
row_task, creator_hash = key
|
||||
work_row = work.get(key)
|
||||
stat = stats.get(key)
|
||||
result.append(
|
||||
{
|
||||
"task_id": row_task,
|
||||
"creator_hash": creator_hash,
|
||||
# 昵称优先取作品的(那是界面上本来就在用的),快照的兜底 —— 没有作品
|
||||
# 的博主只剩快照这一个来源。
|
||||
"creator_name": (work_row or {}).get("creator_name")
|
||||
or (stat.nickname if stat else ""),
|
||||
"creator_alias": aliases.get(
|
||||
(task_platform.get(row_task, ""), creator_hash), ""
|
||||
),
|
||||
"note_count": (work_row or {}).get("note_count", 0),
|
||||
"creator_fans": stat.fans if stat else None,
|
||||
"creator_total_favorited": stat.total_favorited if stat else None,
|
||||
"creator_works": stat.works_count if stat else None,
|
||||
"creator_stats_at": stat.captured_at if stat else None,
|
||||
# 排序用:作品最近出现的时间,或者账号指标的采集时间,取晚的那个。
|
||||
"last_activity_at": max(
|
||||
(work_row or {}).get("last_seen_at") or 0,
|
||||
stat.captured_at if stat else 0,
|
||||
),
|
||||
}
|
||||
)
|
||||
|
||||
result.sort(key=lambda row: row["last_activity_at"], reverse=True)
|
||||
return result
|
||||
|
||||
|
||||
async def note_series(session: AsyncSession, note_id: str, task_id: Optional[int] = None) -> List[Dict[str, Any]]:
|
||||
"""Metric time series for one note."""
|
||||
query = (
|
||||
@@ -409,9 +747,13 @@ async def _note_meta_map(
|
||||
return {
|
||||
row.note_id: {
|
||||
"note_title": row.title,
|
||||
"note_cover": row.cover,
|
||||
"note_cover": covers.cover_url(row.note_id, row.cover),
|
||||
"note_url": row.note_url,
|
||||
"published_at": row.published_at,
|
||||
"task_id": row.task_id,
|
||||
# 博主维度也带上,评论流才能按 博主 -> 作品 -> 评论 三级展开。
|
||||
"creator_hash": row.creator_hash,
|
||||
"creator_name": row.creator_name,
|
||||
}
|
||||
for row in rows
|
||||
}
|
||||
@@ -457,6 +799,9 @@ async def list_comments(
|
||||
"note_title": meta.get(row.note_id, {}).get("note_title", ""),
|
||||
"note_cover": meta.get(row.note_id, {}).get("note_cover", ""),
|
||||
"note_url": meta.get(row.note_id, {}).get("note_url", ""),
|
||||
"note_creator_hash": meta.get(row.note_id, {}).get("creator_hash", ""),
|
||||
"note_creator_name": meta.get(row.note_id, {}).get("creator_name", ""),
|
||||
"note_published_at": meta.get(row.note_id, {}).get("published_at"),
|
||||
}
|
||||
for row in comments
|
||||
]
|
||||
@@ -507,6 +852,8 @@ async def comment_note_groups(
|
||||
"note_title": meta.get(note_id, {}).get("note_title", ""),
|
||||
"note_cover": meta.get(note_id, {}).get("note_cover", ""),
|
||||
"note_url": meta.get(note_id, {}).get("note_url", ""),
|
||||
"creator_hash": meta.get(note_id, {}).get("creator_hash", ""),
|
||||
"creator_name": meta.get(note_id, {}).get("creator_name", ""),
|
||||
"comment_count": count,
|
||||
"latest_at": latest.get(note_id, 0),
|
||||
}
|
||||
@@ -623,11 +970,13 @@ async def list_tasks(
|
||||
"mode": task.mode,
|
||||
"enabled": task.enabled,
|
||||
"interval_minutes": task.interval_minutes,
|
||||
**_schedule_fields(task),
|
||||
"max_notes_count": task.max_notes_count,
|
||||
"enable_comments": task.enable_comments,
|
||||
"max_comments_count": task.max_comments_count,
|
||||
"run_timeout_seconds": task.run_timeout_seconds,
|
||||
"notify_enabled": task.notify_enabled,
|
||||
"notify_failures": task.notify_failures,
|
||||
"next_run_at": task.next_run_at,
|
||||
"last_run_at": task.last_run_at,
|
||||
"last_status": task.last_status,
|
||||
@@ -644,6 +993,29 @@ async def list_tasks(
|
||||
]
|
||||
|
||||
|
||||
def _schedule_fields(task: MonitorTask) -> Dict[str, Any]:
|
||||
"""The schedule columns, plus the sentence the task list renders.
|
||||
|
||||
The label is composed here rather than in the frontend so the list and the
|
||||
editor cannot drift on what a given schedule means.
|
||||
"""
|
||||
hours = schedule.parse_hours(task.schedule_hours)
|
||||
days = schedule.parse_days(task.schedule_days)
|
||||
return {
|
||||
"schedule_mode": task.schedule_mode,
|
||||
"schedule_hours": hours,
|
||||
"schedule_days": days,
|
||||
"schedule_minute": task.schedule_minute,
|
||||
"schedule_label": schedule.describe(
|
||||
mode=task.schedule_mode,
|
||||
interval_minutes=task.interval_minutes,
|
||||
hours=hours,
|
||||
days=days,
|
||||
minute=task.schedule_minute,
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
async def overview(session: AsyncSession, platform: Optional[str] = None) -> Dict[str, Any]:
|
||||
"""Headline numbers for the dashboard tiles, scoped to one platform."""
|
||||
now = get_current_timestamp()
|
||||
|
||||
@@ -0,0 +1,355 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/upstream.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""上游仓库更新检查。
|
||||
|
||||
本仓库在上游(NanmiCoder/MediaCrawler)之上加了一整层监控/鉴权/多平台面板,
|
||||
合并流程写在 UPSTREAM.md 里。但那份流程默认**有人知道上游动了**——而部署脚本是
|
||||
`git pull --ff-only`,只从我们自己的 Gitea 拉,上游的提交不主动去 fetch 就永远
|
||||
看不见。拖着不合并的代价是复利的:越久越难合,最后只能放弃。这个模块把「上游动
|
||||
了没有」变成一条可定时、会推到企业微信的通知。
|
||||
|
||||
三处刻意的取舍:
|
||||
|
||||
* **用 git 而不是托管商的 HTTP API。** 只有 git 算得出「落后几个提交」:托管商
|
||||
API 能告诉你上游 tip 是什么,但它不知道我们与上游的共同祖先在哪,而分歧点恰恰
|
||||
是真正要合的东西。本仓库还含有上游没有的提交,直接比 tip 会得出错误的结论。
|
||||
* **按 URL fetch 到 FETCH_HEAD,不配置 remote、不写 refs/remotes。** 服务器上的
|
||||
checkout 是从 Gitea 克隆的,本来就没有 upstream 这个 remote;用 URL 直取就不必
|
||||
先去改它的 git 配置。顺带也避免往别人的部署里塞一个 remote。
|
||||
* **只读不写工作区。** fetch 只落对象和 FETCH_HEAD,不碰索引与工作区,所以不会打断
|
||||
正在跑的采集,也不会和 `./deploy.sh` 的 git pull 抢锁。
|
||||
|
||||
依赖一个外部命令:**git**。本机开发环境一定有;容器里是 Dockerfile 显式装的
|
||||
(python:3.11-slim 默认不带)。
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List, Tuple
|
||||
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from .db import get_session
|
||||
from .models import SETTING_UPSTREAM_NOTIFIED_TIP, SETTING_UPSTREAM_STATE
|
||||
from .settings import get_setting, set_setting
|
||||
|
||||
PROJECT_ROOT = Path(__file__).parent.parent.parent
|
||||
|
||||
# 默认就是本仓库跟踪的那个上游。国内直连 GitHub 不稳时改成 gitcode 镜像即可
|
||||
# (见 UPSTREAM.md「直连 GitHub 不通时」)。
|
||||
DEFAULT_REMOTE_URL = "https://github.com/NanmiCoder/MediaCrawler.git"
|
||||
DEFAULT_BRANCH = "main"
|
||||
|
||||
# fetch 要走网络,给宽松些;其余全是本地命令,慢到这个程度只能说明仓库坏了。
|
||||
FETCH_TIMEOUT_SECONDS = 120
|
||||
LOCAL_TIMEOUT_SECONDS = 20
|
||||
|
||||
# 通知里最多列几条提交。要传达的是「该动手了」,不是把 changelog 搬到群里。
|
||||
MAX_LISTED_COMMITS = 10
|
||||
# 状态里留几条给前端展示。比通知多留一些,界面上能看到更完整的列表。
|
||||
MAX_STORED_COMMITS = 30
|
||||
|
||||
# git log 用 Unit Separator 分隔字段:它不可能出现在提交信息里,比制表符安全。
|
||||
_RECORD_SEPARATOR = "\x1f"
|
||||
_LOG_FORMAT = (
|
||||
f"%h{_RECORD_SEPARATOR}%an{_RECORD_SEPARATOR}%ad{_RECORD_SEPARATOR}%s"
|
||||
)
|
||||
|
||||
|
||||
class GitError(RuntimeError):
|
||||
"""git 不可用,或某条 git 命令失败了。"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Commit:
|
||||
"""一条上游提交,只留通知/展示需要的四个字段。"""
|
||||
|
||||
sha: str
|
||||
author: str
|
||||
date: str
|
||||
subject: str
|
||||
|
||||
|
||||
@dataclass
|
||||
class CheckResult:
|
||||
"""一次检查的结论。失败也是一种结论,用 ``ok``/``error`` 表达而不是抛异常。"""
|
||||
|
||||
ok: bool
|
||||
# HEAD..FETCH_HEAD:上游有而我们没有的提交数 —— 要合的就是这些。
|
||||
behind: int = 0
|
||||
# FETCH_HEAD..HEAD:我们有自己的提交数 —— 也就是这一层的规模。
|
||||
ahead: int = 0
|
||||
tip: str = ""
|
||||
head: str = ""
|
||||
commits: List[Commit] = field(default_factory=list)
|
||||
error: str = ""
|
||||
|
||||
def as_dict(self) -> Dict[str, Any]:
|
||||
return {
|
||||
"ok": self.ok,
|
||||
"behind": self.behind,
|
||||
"ahead": self.ahead,
|
||||
"tip": self.tip,
|
||||
"head": self.head,
|
||||
"commits": [commit.__dict__ for commit in self.commits],
|
||||
"error": self.error,
|
||||
}
|
||||
|
||||
|
||||
def _git(args: List[str], timeout: int) -> subprocess.CompletedProcess:
|
||||
env = dict(os.environ)
|
||||
# 远端要凭据时(地址写成了私有仓库),git 会停下来问密码,而这里没有终端可问,
|
||||
# 于是挂到超时。关掉一切交互,让它立刻失败。
|
||||
env["GIT_TERMINAL_PROMPT"] = "0"
|
||||
env["GIT_ASKPASS"] = ""
|
||||
env["SSH_ASKPASS"] = ""
|
||||
# 容器里 uid 1000 没有 passwd 项,git 找不到 HOME 会抱怨。给一个存在且可写的。
|
||||
env.setdefault("HOME", "/tmp")
|
||||
|
||||
return subprocess.run(
|
||||
[
|
||||
"git",
|
||||
"-C",
|
||||
str(PROJECT_ROOT),
|
||||
# 只对自己这个 checkout 放行所有权检查。容器里 uid 一般与属主一致,
|
||||
# 但 bind mount 的属主未必,一旦不一致 git 会直接拒绝干任何活。
|
||||
"-c",
|
||||
f"safe.directory={PROJECT_ROOT}",
|
||||
# 忽略任何全局凭据助手:这是个只读的公开仓库,不该去翻钥匙串。
|
||||
"-c",
|
||||
"credential.helper=",
|
||||
*args,
|
||||
],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
encoding="utf-8",
|
||||
errors="replace",
|
||||
timeout=timeout,
|
||||
env=env,
|
||||
)
|
||||
|
||||
|
||||
def _run(args: List[str], timeout: int) -> Tuple[int, str, str]:
|
||||
"""跑一条 git 命令,返回 (returncode, stdout, stderr)。
|
||||
|
||||
只把「跑不起来」当异常;命令返回非零是正常结果,交给调用方处理。
|
||||
"""
|
||||
try:
|
||||
proc = _git(args, timeout)
|
||||
except FileNotFoundError as exc:
|
||||
raise GitError("未找到 git 命令,请先安装 git") from exc
|
||||
except subprocess.TimeoutExpired as exc:
|
||||
raise GitError(f"git {args[0]} 超时({timeout} 秒)") from exc
|
||||
return proc.returncode, (proc.stdout or "").strip(), (proc.stderr or "").strip()
|
||||
|
||||
|
||||
def _require(args: List[str], timeout: int, what: str) -> str:
|
||||
code, out, err = _run(args, timeout)
|
||||
if code != 0:
|
||||
# git 的报错通常是多行的,只留第一行;完整输出塞进日志反而更难读。
|
||||
detail = err.splitlines()[0].strip() if err else "未知错误"
|
||||
raise GitError(f"{what}:{detail}")
|
||||
return out
|
||||
|
||||
|
||||
def _to_int(raw: str) -> int:
|
||||
try:
|
||||
return int(raw)
|
||||
except (TypeError, ValueError):
|
||||
return 0
|
||||
|
||||
|
||||
def _parse_log(raw: str) -> List[Commit]:
|
||||
commits: List[Commit] = []
|
||||
for line in raw.splitlines():
|
||||
parts = line.split(_RECORD_SEPARATOR)
|
||||
if len(parts) != 4:
|
||||
# 格式不对就跳过这一条:一条读不出来的提交不该让整次检查失败。
|
||||
continue
|
||||
sha, author, date, subject = parts
|
||||
commits.append(Commit(sha=sha, author=author, date=date, subject=subject))
|
||||
return commits
|
||||
|
||||
|
||||
def _check_sync(remote_url: str, branch: str) -> CheckResult:
|
||||
"""阻塞实现,异步包装见 :func:`check`。"""
|
||||
# 在 try 之前绑定:后面的失败结果也带上它 —— 「检查失败」时当前跑的是哪个
|
||||
# 提交,正是排查时第一个想知道的。
|
||||
head = ""
|
||||
try:
|
||||
head = _require(["rev-parse", "HEAD"], LOCAL_TIMEOUT_SECONDS, "读取本地 HEAD 失败")
|
||||
# 增量 fetch:对象本地基本都已经有了,所以正常情况下只传几个新提交,
|
||||
# 不会遇到 UPSTREAM.md 里说的「大包必断」。
|
||||
_require(
|
||||
["fetch", "--no-tags", remote_url, branch],
|
||||
FETCH_TIMEOUT_SECONDS,
|
||||
"从上游 fetch 失败",
|
||||
)
|
||||
tip = _require(["rev-parse", "FETCH_HEAD"], LOCAL_TIMEOUT_SECONDS, "读不到 FETCH_HEAD")
|
||||
behind = _to_int(
|
||||
_require(
|
||||
["rev-list", "--count", "HEAD..FETCH_HEAD"],
|
||||
LOCAL_TIMEOUT_SECONDS,
|
||||
"统计落后提交数失败",
|
||||
)
|
||||
)
|
||||
ahead = _to_int(
|
||||
_require(
|
||||
["rev-list", "--count", "FETCH_HEAD..HEAD"],
|
||||
LOCAL_TIMEOUT_SECONDS,
|
||||
"统计领先提交数失败",
|
||||
)
|
||||
)
|
||||
# 只在确实落后时才读提交列表:已经是最新时这条 git log 毫无意义。
|
||||
raw_log = ""
|
||||
if behind:
|
||||
raw_log = _require(
|
||||
[
|
||||
"log",
|
||||
f"--max-count={MAX_STORED_COMMITS}",
|
||||
"--date=short",
|
||||
f"--format={_LOG_FORMAT}",
|
||||
"HEAD..FETCH_HEAD",
|
||||
],
|
||||
LOCAL_TIMEOUT_SECONDS,
|
||||
"读取新提交列表失败",
|
||||
)
|
||||
except GitError as exc:
|
||||
return CheckResult(ok=False, head=head, error=str(exc))
|
||||
|
||||
return CheckResult(
|
||||
ok=True,
|
||||
behind=behind,
|
||||
ahead=ahead,
|
||||
tip=tip,
|
||||
head=head,
|
||||
commits=_parse_log(raw_log),
|
||||
)
|
||||
|
||||
|
||||
async def check(
|
||||
remote_url: str = DEFAULT_REMOTE_URL, branch: str = DEFAULT_BRANCH
|
||||
) -> CheckResult:
|
||||
"""跑一次检查。网络与子进程都丢进线程,事件循环不被阻塞。
|
||||
|
||||
不抛异常:上游不通是常态(尤其是直连 GitHub),那也是一种要记录下来的结果。
|
||||
"""
|
||||
return await asyncio.to_thread(_check_sync, remote_url, branch)
|
||||
|
||||
|
||||
def build_message(result: CheckResult, branch: str) -> str:
|
||||
"""把一次「上游有新提交」的结果写成一条企业微信 markdown。"""
|
||||
lines = [
|
||||
"**🔔 上游 MediaCrawler 有更新**",
|
||||
f"> 当前部署落后 `{branch}` **{result.behind}** 个提交",
|
||||
]
|
||||
if result.ahead:
|
||||
lines.append(f"> (本仓库另有 {result.ahead} 个自己的提交,合并时注意保留)")
|
||||
for commit in result.commits[:MAX_LISTED_COMMITS]:
|
||||
lines.append(f"> `{commit.sha}` {commit.subject}")
|
||||
# 落后数可能大于列出来的条数:状态里留的提交本身也是截断的(30 条),
|
||||
# 所以这里比的是总数,不是 len(commits)。
|
||||
if result.behind > MAX_LISTED_COMMITS:
|
||||
lines.append(f"> …等共 {result.behind} 个提交")
|
||||
lines.append("> 合并步骤见仓库根目录 `UPSTREAM.md`")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
async def load_state(session: AsyncSession) -> Dict[str, Any]:
|
||||
"""最近一次检查的结果,从设置里读回来。没查过时是空字典。"""
|
||||
raw = await get_setting(session, SETTING_UPSTREAM_STATE)
|
||||
if not raw:
|
||||
return {}
|
||||
try:
|
||||
state = json.loads(raw)
|
||||
except json.JSONDecodeError:
|
||||
# 手改坏了的行不该让接口 500,当作「没查过」即可。
|
||||
return {}
|
||||
return state if isinstance(state, dict) else {}
|
||||
|
||||
|
||||
async def _save_state(session: AsyncSession, state: Dict[str, Any]) -> None:
|
||||
# 整体读写,所以存成一条 JSON:拆成多个 key 只会带来写到一半的不一致。
|
||||
await set_setting(session, SETTING_UPSTREAM_STATE, json.dumps(state, ensure_ascii=False))
|
||||
|
||||
|
||||
async def run_check(notify_when_new: bool = True) -> Dict[str, Any]:
|
||||
"""检查一次,落库,必要时推送。返回值可直接交给前端。
|
||||
|
||||
分三段各自的数据库会话:fetch 最长可能跑满两分钟,占着一个连接不合适 ——
|
||||
理由与 runner.py 的分段完全相同。
|
||||
"""
|
||||
# 延迟导入:app_settings 在模块级 import 本模块(为了那个默认地址常量),
|
||||
# 模块级反向 import 会成环。
|
||||
from . import app_settings, notify
|
||||
|
||||
async with get_session() as session:
|
||||
remote_url = str(
|
||||
await app_settings.get_value(
|
||||
session, "upstream_remote_url", fallback=DEFAULT_REMOTE_URL
|
||||
)
|
||||
or DEFAULT_REMOTE_URL
|
||||
)
|
||||
branch = str(
|
||||
await app_settings.get_value(session, "upstream_branch", fallback=DEFAULT_BRANCH)
|
||||
or DEFAULT_BRANCH
|
||||
)
|
||||
notify_enabled = bool(
|
||||
await app_settings.get_value(session, "upstream_notify", fallback=True)
|
||||
)
|
||||
notified_tip = (await get_setting(session, SETTING_UPSTREAM_NOTIFIED_TIP)) or ""
|
||||
webhook_url = await notify.get_webhook_url(session)
|
||||
|
||||
result = await check(remote_url, branch)
|
||||
|
||||
payload: Dict[str, Any] = {
|
||||
"checked_at": get_current_timestamp(),
|
||||
"remote_url": remote_url,
|
||||
"branch": branch,
|
||||
**result.as_dict(),
|
||||
}
|
||||
|
||||
async with get_session() as session:
|
||||
await _save_state(session, payload)
|
||||
|
||||
has_update = result.ok and result.behind > 0 and bool(result.tip)
|
||||
# 同一个 tip 只推一次:否则每过一个检查周期就把同样的更新推到群里,
|
||||
# 直到有人去合为止。上游真又动了(tip 变了)时应该再推。
|
||||
if (
|
||||
notify_when_new
|
||||
and notify_enabled
|
||||
and has_update
|
||||
and webhook_url
|
||||
and result.tip != notified_tip
|
||||
):
|
||||
ok, detail = await notify.send_wecom(webhook_url, build_message(result, branch))
|
||||
if ok:
|
||||
await set_setting(session, SETTING_UPSTREAM_NOTIFIED_TIP, result.tip)
|
||||
payload["notified"] = True
|
||||
else:
|
||||
# 推送失败不该抹掉检查结果 —— 界面上仍然能看到「落后几个提交」。
|
||||
payload["notify_error"] = detail
|
||||
|
||||
return payload
|
||||
@@ -18,6 +18,7 @@
|
||||
|
||||
from .auth import router as auth_router
|
||||
from .crawler import router as crawler_router
|
||||
from .creator import router as creator_router
|
||||
from .data import router as data_router
|
||||
from .monitor import router as monitor_router
|
||||
from .settings import router as settings_router
|
||||
@@ -26,6 +27,7 @@ from .websocket import router as websocket_router
|
||||
__all__ = [
|
||||
"auth_router",
|
||||
"crawler_router",
|
||||
"creator_router",
|
||||
"data_router",
|
||||
"monitor_router",
|
||||
"settings_router",
|
||||
|
||||
@@ -0,0 +1,150 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/creator.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""运营模块的 HTTP 接口。"""
|
||||
|
||||
import asyncio
|
||||
from typing import Set
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Query
|
||||
|
||||
from ..creator import login as creator_login
|
||||
from ..creator import service
|
||||
from ..monitor.db import get_session
|
||||
|
||||
router = APIRouter(prefix="/creator", tags=["creator"])
|
||||
|
||||
# 后台同步任务要留强引用:asyncio 只持弱引用,否则任务可能在跑完前被回收。
|
||||
_sync_tasks: Set[asyncio.Task] = set()
|
||||
|
||||
|
||||
def _bad_request(exc: ValueError) -> HTTPException:
|
||||
return HTTPException(status_code=400, detail=str(exc))
|
||||
|
||||
|
||||
@router.get("/accounts")
|
||||
async def list_accounts():
|
||||
"""账号列表。**不含 cookie**,只给 ``has_cookie``。"""
|
||||
async with get_session() as session:
|
||||
return {"accounts": await service.list_accounts(session)}
|
||||
|
||||
|
||||
@router.get("/accounts/{account_id}")
|
||||
async def get_account_detail(account_id: int):
|
||||
async with get_session() as session:
|
||||
try:
|
||||
return await service.account_detail(session, account_id)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(status_code=404, detail=str(exc))
|
||||
|
||||
|
||||
@router.delete("/accounts/{account_id}")
|
||||
async def delete_account(account_id: int):
|
||||
async with get_session() as session:
|
||||
try:
|
||||
await service.delete_account(session, account_id)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(status_code=404, detail=str(exc))
|
||||
return {"message": "账号已删除"}
|
||||
|
||||
|
||||
@router.post("/accounts/{account_id}/check")
|
||||
async def check_account(account_id: int):
|
||||
"""重测登录态与数据权限。"""
|
||||
async with get_session() as session:
|
||||
try:
|
||||
return await service.check_account(session, account_id)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(status_code=404, detail=str(exc))
|
||||
|
||||
|
||||
@router.post("/accounts/{account_id}/sync")
|
||||
async def sync_account(
|
||||
account_id: int, days: int = Query(default=90, ge=1, le=730)
|
||||
):
|
||||
"""拉取作品数据。
|
||||
|
||||
**放后台跑**:要分页、还要按账号节流,几分钟很正常,而前端请求超时是 30 秒。
|
||||
前端靠轮询账号列表里的 `last_synced_at` / `last_error` 看结果。
|
||||
"""
|
||||
async with get_session() as session:
|
||||
try:
|
||||
await service.get_account(session, account_id)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(status_code=404, detail=str(exc))
|
||||
|
||||
task = asyncio.create_task(_sync_in_background(account_id, days))
|
||||
_sync_tasks.add(task)
|
||||
task.add_done_callback(_sync_tasks.discard)
|
||||
return {"message": "同步已开始", "days": days}
|
||||
|
||||
|
||||
async def _sync_in_background(account_id: int, days: int) -> None:
|
||||
try:
|
||||
async with get_session() as session:
|
||||
result = await service.sync_account(session, account_id, days)
|
||||
print(f"[creator] 账号 {account_id} 同步完成,取回 {result['fetched']} 条")
|
||||
except Exception as exc: # noqa: BLE001 - 后台任务不能让异常逃逸成静默失败
|
||||
print(f"[creator] 账号 {account_id} 同步失败: {exc}")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 扫码新增账号
|
||||
# ---------------------------------------------------------------------------
|
||||
#
|
||||
# 每次登录开一个**临时浏览器上下文**,扫完取出 cookie 就丢弃 —— 这样登第二个账号
|
||||
# 不会把第一个顶掉,也不影响监控那个登录态。cookie 只在内存里从 login 模块传到
|
||||
# 这里落库,**不进响应体**。
|
||||
|
||||
|
||||
@router.post("/login")
|
||||
async def start_login():
|
||||
try:
|
||||
return await creator_login.start()
|
||||
except RuntimeError as exc:
|
||||
raise HTTPException(status_code=502, detail=str(exc))
|
||||
|
||||
|
||||
@router.get("/login")
|
||||
async def poll_login():
|
||||
"""轮询扫码结果;一旦成功就把账号落库并返回它。"""
|
||||
snapshot = await creator_login.status()
|
||||
|
||||
if snapshot["status"] == creator_login.STATUS_SUCCESS:
|
||||
# take_cookie 只在会话还在时返回 cookie,取走即拆会话;重复轮询拿到 None
|
||||
# 就说明已经保存过了,直接返回上次的结果,不要退回 idle。
|
||||
cookie = await creator_login.take_cookie()
|
||||
if cookie:
|
||||
try:
|
||||
async with get_session() as session:
|
||||
account = await service.upsert_account_from_cookie(session, cookie)
|
||||
except ValueError as exc:
|
||||
snapshot["status"] = creator_login.STATUS_ERROR
|
||||
snapshot["message"] = f"扫码成功但保存账号失败:{exc}"
|
||||
await creator_login.remember_result(snapshot)
|
||||
return snapshot
|
||||
snapshot["account"] = account
|
||||
snapshot["message"] = f"已添加账号:{account['nickname']}"
|
||||
await creator_login.remember_result(snapshot)
|
||||
|
||||
return snapshot
|
||||
|
||||
|
||||
@router.delete("/login")
|
||||
async def cancel_login():
|
||||
return await creator_login.cancel()
|
||||
+233
-12
@@ -18,12 +18,13 @@
|
||||
|
||||
"""HTTP API for scheduled monitoring tasks."""
|
||||
|
||||
from datetime import date, timedelta
|
||||
from datetime import date, datetime, timedelta
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Query, Response
|
||||
from fastapi.responses import FileResponse
|
||||
|
||||
from ..monitor import notify, report, service
|
||||
from ..monitor import covers, notify, qrlogin, report, service, upstream
|
||||
from ..monitor.db import get_session
|
||||
from ..monitor.platforms import PLATFORM_XHS
|
||||
from ..monitor.settings import (
|
||||
@@ -37,6 +38,8 @@ from ..monitor.settings import (
|
||||
from ..monitor.models import SETTING_WECOM_WEBHOOK, MonitorTask
|
||||
from ..schemas.monitor import (
|
||||
CookiePayload,
|
||||
CreatorAliasPayload,
|
||||
NoteAliasPayload,
|
||||
MonitorTaskCreate,
|
||||
MonitorTaskUpdate,
|
||||
WebhookPayload,
|
||||
@@ -130,7 +133,10 @@ async def list_notes(
|
||||
):
|
||||
async with get_session() as session:
|
||||
return {
|
||||
"notes": await service.list_notes(session, task_id, only_new, limit, platform)
|
||||
"notes": await service.list_notes(session, task_id, only_new, limit, platform),
|
||||
# 博主**单独给一份**,而不是让前端从作品里推。作品推不出「一条作品都没有的
|
||||
# 博主」—— 那正是最该显示的一类(还在涨粉,只是最近没发)。
|
||||
"creators": await service.list_creators(session, task_id, platform),
|
||||
}
|
||||
|
||||
|
||||
@@ -171,6 +177,15 @@ async def list_comments(
|
||||
"note_title": comment["note_title"],
|
||||
"note_cover": comment["note_cover"],
|
||||
"note_url": comment["note_url"],
|
||||
# 作品所属的创作者。评论流按 博主 → 作品 → 评论 三级展开时,最外层
|
||||
# 就是按这两个字段分组的 —— 少了它们,前端拿到的是 undefined,
|
||||
# 于是所有博主塌成同一个分组、标签回退成「未知博主」。
|
||||
# 同一个桶里的评论必然同属一个作品,所以取哪一条都一样。
|
||||
"creator_hash": comment["note_creator_hash"],
|
||||
"creator_name": comment["note_creator_name"],
|
||||
# 作品的发布时间。评论流按 博主 → 作品 → 评论 展开时,作品那一层
|
||||
# 光有标题不够 —— 同名作品不少,日期能帮着认。
|
||||
"published_at": comment["note_published_at"],
|
||||
"comments": [],
|
||||
},
|
||||
)
|
||||
@@ -243,6 +258,138 @@ async def clear_cookie_endpoint(platform: str = Query(default=PLATFORM_XHS)):
|
||||
return {"message": "Cookie cleared"}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 博主备注
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@router.put("/creators/{creator_hash}")
|
||||
async def set_creator_alias_endpoint(
|
||||
creator_hash: str,
|
||||
payload: CreatorAliasPayload,
|
||||
platform: str = Query(default=PLATFORM_XHS),
|
||||
):
|
||||
"""给博主起个备注(界面上的「备注」)。
|
||||
|
||||
作品栏按 creator_hash 归组,可那是个哈希、昵称又常常认不出是谁 —— 备注是人自己起的
|
||||
名字。空串表示清掉这条备注。
|
||||
"""
|
||||
async with get_session() as session:
|
||||
await service.set_creator_alias(session, platform, creator_hash, payload.alias)
|
||||
return {"creator_hash": creator_hash, "alias": payload.alias.strip()}
|
||||
|
||||
|
||||
@router.put("/notes/{note_id}")
|
||||
async def set_note_alias_endpoint(
|
||||
note_id: str,
|
||||
payload: NoteAliasPayload,
|
||||
platform: str = Query(default=PLATFORM_XHS),
|
||||
):
|
||||
"""给**作品**起个备注。
|
||||
|
||||
和上面那条博主备注是一对:博主备注回答「这个账号是谁」,这条回答「这条作品我要盯着」。
|
||||
两者不能合并 —— 一个博主底下常常只有一两件值得盯的作品。
|
||||
"""
|
||||
async with get_session() as session:
|
||||
await service.set_note_alias(session, platform, note_id, payload.alias)
|
||||
return {"note_id": note_id, "alias": payload.alias.strip()}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# QR login
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
# These drive the browser already listening on the CDP debug port, which is the
|
||||
# same browser -- and therefore the same profile -- that monitor runs attach to.
|
||||
# Scanning once is what makes later unattended runs logged in.
|
||||
|
||||
|
||||
@router.get("/covers/{note_id}")
|
||||
async def get_cover(note_id: str):
|
||||
"""作品封面,从本地缓存读。
|
||||
|
||||
**为什么不让前端直连图床**:图床地址是带签名、会过期的 —— 实测隔天即 403,
|
||||
而且带不带 Referer 都一样,所以那是过期而不是防盗链。本地那份与签名无关。
|
||||
|
||||
这个路由是带鉴权的(整条 monitor 路由都挂了 require_auth),所以封面不会被
|
||||
匿名读走;前端用同源的 <img> 请求会自动带上会话 cookie。
|
||||
"""
|
||||
path = covers.find_cached(note_id)
|
||||
if path is None:
|
||||
raise HTTPException(status_code=404, detail="封面未缓存")
|
||||
|
||||
media_types = {
|
||||
".jpg": "image/jpeg",
|
||||
".png": "image/png",
|
||||
".webp": "image/webp",
|
||||
".gif": "image/gif",
|
||||
".heic": "image/heic",
|
||||
}
|
||||
return FileResponse(
|
||||
path,
|
||||
media_type=media_types.get(path.suffix.lower(), "application/octet-stream"),
|
||||
# 本地文件不会变(note_id 唯一),让浏览器自己缓存,省掉重复请求。
|
||||
headers={"Cache-Control": "private, max-age=86400"},
|
||||
)
|
||||
|
||||
|
||||
@router.post("/login/qr")
|
||||
async def start_qr_login(platform: str = Query(default=PLATFORM_XHS)):
|
||||
"""Open the login page in the CDP browser and return its QR code.
|
||||
|
||||
A server deployment has no display (Chrome sits under Xvfb), so the code is
|
||||
surfaced here for the operator to scan instead of in a desktop window that
|
||||
does not exist.
|
||||
"""
|
||||
try:
|
||||
return await qrlogin.start(platform)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(status_code=400, detail=str(exc))
|
||||
except RuntimeError as exc:
|
||||
raise HTTPException(status_code=502, detail=str(exc))
|
||||
|
||||
|
||||
@router.get("/login/qr")
|
||||
async def get_qr_login():
|
||||
"""Poll the live session: waiting -> success / expired / error.
|
||||
|
||||
扫码成功时**把 cookie 一并存进库**。扫码本来只写浏览器 profile,那只够 CDP 模式用;
|
||||
存一份之后,CDP 关掉、任务改用 --cookies_file 注入也照样能跑 —— 两种机制同时填上,
|
||||
开关怎么切都不会断。
|
||||
"""
|
||||
snapshot = await qrlogin.status()
|
||||
|
||||
if snapshot["status"] == qrlogin.STATUS_SUCCESS:
|
||||
cookie = await qrlogin.take_cookie()
|
||||
if cookie:
|
||||
async with get_session() as session:
|
||||
await set_cookie(session, cookie)
|
||||
snapshot["cookie_saved"] = True
|
||||
snapshot["message"] = f"{snapshot['message']};登录态已同时存入 Cookie"
|
||||
|
||||
return snapshot
|
||||
|
||||
|
||||
@router.delete("/login/qr")
|
||||
async def cancel_qr_login():
|
||||
"""Drop our tab and stop polling."""
|
||||
return await qrlogin.cancel()
|
||||
|
||||
|
||||
@router.get("/login/state")
|
||||
async def get_login_state(force: bool = Query(default=False)):
|
||||
"""Ask the browser itself whether it is signed in.
|
||||
|
||||
Deliberately separate from the QR session above. That session is in-memory and
|
||||
dies with the process -- a redeploy is enough -- so "am I logged in?" must not
|
||||
hinge on it, or a successful scan looks like nothing happened.
|
||||
|
||||
``force`` reloads the page first, for when the login may have lapsed somewhere
|
||||
else and the page's copy of the state is stale.
|
||||
"""
|
||||
return await qrlogin.check_login_state(force=force)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Report
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -339,20 +486,29 @@ def _export_columns(kind: str) -> List[tuple[str, str]]:
|
||||
"""(key, header) pairs per export kind."""
|
||||
if kind == "notes":
|
||||
return [
|
||||
("note_id", "作品ID"),
|
||||
# 先放「这是谁」:导出来是拿去比对和汇报的,一行只有作品 ID 没法用。
|
||||
# 备注优先 —— 昵称常常认不出是谁(见 notes 表那一层的说明)。
|
||||
("creator_alias", "博主备注"),
|
||||
("creator_name", "博主昵称"),
|
||||
("note_alias", "作品备注"),
|
||||
("title", "标题"),
|
||||
("note_id", "作品ID"),
|
||||
("note_url", "链接"),
|
||||
("liked_count", "点赞"),
|
||||
("comment_count", "评论"),
|
||||
("collected_count", "收藏"),
|
||||
("share_count", "分享"),
|
||||
("liked_count_delta", "点赞增量"),
|
||||
("comment_count_delta", "评论增量"),
|
||||
("published_at", "发布时间"),
|
||||
# 指标嵌在 row["metrics"] 里,所以这里必须写成路径 —— 写成裸键名的话这几列
|
||||
# 全空(见 _lookup)。
|
||||
("metrics.liked_count", "点赞"),
|
||||
("metrics.comment_count", "评论"),
|
||||
("metrics.collected_count", "收藏"),
|
||||
("metrics.share_count", "分享"),
|
||||
("deltas.liked_count", "点赞增量"),
|
||||
("deltas.comment_count", "评论增量"),
|
||||
("first_seen_at", "首次发现"),
|
||||
("last_seen_at", "最近采集"),
|
||||
]
|
||||
if kind == "comments":
|
||||
return [
|
||||
("note_creator_name", "博主昵称"),
|
||||
("note_title", "所属作品"),
|
||||
("note_id", "作品ID"),
|
||||
("comment_id", "评论ID"),
|
||||
@@ -374,6 +530,42 @@ def _export_columns(kind: str) -> List[tuple[str, str]]:
|
||||
]
|
||||
|
||||
|
||||
# 表里存的是毫秒时间戳。直接倒进 CSV 就是一串 13 位数字 —— 打开 Excel 的人没法看,
|
||||
# 也没法排序。这几个键统一格式化成人能读的形态。
|
||||
_TIME_KEYS = {"published_at", "first_seen_at", "last_seen_at", "create_time"}
|
||||
|
||||
|
||||
def _fmt_time(value: Any) -> str:
|
||||
"""毫秒 → ``YYYY-MM-DD HH:MM``(服务器本地时区)。"""
|
||||
try:
|
||||
return datetime.fromtimestamp(int(value) / 1000).strftime("%Y-%m-%d %H:%M")
|
||||
except (TypeError, ValueError, OSError, OverflowError):
|
||||
return ""
|
||||
|
||||
|
||||
def _cell_for(row: Dict[str, Any], key: str) -> Any:
|
||||
"""一列的值:时间键格式化成人能读的,其余照原样(None 变空串)。"""
|
||||
value = _lookup(row, key)
|
||||
if key.rsplit(".", 1)[-1] in _TIME_KEYS:
|
||||
return _fmt_time(value)
|
||||
return _cell(value)
|
||||
|
||||
|
||||
def _lookup(row: Dict[str, Any], key: str) -> Any:
|
||||
"""取一列的值。键可以是 ``metrics.liked_count`` 这种路径。
|
||||
|
||||
作品行的指标是**嵌在** ``metrics`` / ``deltas`` 里的,而 ``_export_columns`` 里写的
|
||||
是 ``liked_count`` —— 照顶层键直接 ``row.get()`` 的话,点赞/评论/收藏/分享四列连带
|
||||
两个增量列**永远是空的**,导出来的表看着有这几列,其实一格都没有。
|
||||
"""
|
||||
value: Any = row
|
||||
for part in key.split("."):
|
||||
if not isinstance(value, dict):
|
||||
return None
|
||||
value = value.get(part)
|
||||
return value
|
||||
|
||||
|
||||
def _cell(value: Any) -> Any:
|
||||
if value is None:
|
||||
return ""
|
||||
@@ -390,7 +582,7 @@ def _to_csv(rows: List[Dict[str, Any]], columns: List[tuple[str, str]]) -> bytes
|
||||
writer = csv.writer(buffer)
|
||||
writer.writerow([header for _, header in columns])
|
||||
for row in rows:
|
||||
writer.writerow([_cell(row.get(key)) for key, _ in columns])
|
||||
writer.writerow([_cell_for(row, key) for key, _ in columns])
|
||||
|
||||
# utf-8-sig: without the BOM Excel opens Chinese CSV as mojibake, which is
|
||||
# the single most common complaint about CSV exports here.
|
||||
@@ -407,7 +599,7 @@ def _to_xlsx(rows: List[Dict[str, Any]], columns: List[tuple[str, str]], sheet:
|
||||
worksheet.title = {"notes": "作品", "comments": "评论"}.get(sheet, "报表")
|
||||
worksheet.append([header for _, header in columns])
|
||||
for row in rows:
|
||||
worksheet.append([_cell(row.get(key)) for key, _ in columns])
|
||||
worksheet.append([_cell_for(row, key) for key, _ in columns])
|
||||
|
||||
output = io.BytesIO()
|
||||
workbook.save(output)
|
||||
@@ -500,3 +692,32 @@ async def test_webhook(payload: WebhookTestPayload):
|
||||
if not ok:
|
||||
raise HTTPException(status_code=400, detail=detail)
|
||||
return {"message": detail}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 上游更新
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@router.get("/upstream")
|
||||
async def get_upstream_status():
|
||||
"""最近一次上游检查的结果。
|
||||
|
||||
只读缓存,不触发检查:fetch 要走网络、最长两分钟,不该由一个 GET 顺手发起。
|
||||
没有查过时返回空对象,前端据此显示「尚未检查」。
|
||||
"""
|
||||
async with get_session() as session:
|
||||
return await upstream.load_state(session)
|
||||
|
||||
|
||||
@router.post("/upstream/check")
|
||||
async def run_upstream_check():
|
||||
"""立刻检查一次上游仓库,并把结果写回缓存。
|
||||
|
||||
即使「定期检查」开关是关的也照查 —— 手动点这一次的意义正在于此。这里会一直
|
||||
等到 fetch 结束(前端给这条请求单独放长了超时),因为结果就是要给人看的。
|
||||
|
||||
``notify_when_new=False``:点这个按钮的人正看着结果,没必要再给自己推一条群消息。
|
||||
没有推过的那批提交会留给下一次「定时检查」推 —— 推送状态记的是 tip,不是「推过没」。
|
||||
"""
|
||||
return await upstream.run_check(notify_when_new=False)
|
||||
|
||||
+50
-3
@@ -20,7 +20,7 @@
|
||||
|
||||
from typing import List, Literal, Optional
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
from pydantic import BaseModel, Field, model_validator
|
||||
|
||||
# A floor on the interval is a correctness guard, not a nicety: every run
|
||||
# launches a browser and hits XHS with several requests, so a short interval
|
||||
@@ -40,6 +40,19 @@ class MonitorTaskCreate(BaseModel):
|
||||
interval_minutes: Optional[int] = Field(
|
||||
default=None, ge=MIN_INTERVAL_MINUTES, le=MAX_INTERVAL_MINUTES
|
||||
)
|
||||
# --- Scheduling ---------------------------------------------------------
|
||||
# All three modes are expressible with pickers; a raw cron string is
|
||||
# deliberately not supported, since it is a small language to learn just to
|
||||
# say "every day at nine".
|
||||
schedule_mode: Literal["interval", "daily", "weekly"] = "interval"
|
||||
# 0-23, e.g. [9, 12, 18]. Required for the two clock modes.
|
||||
schedule_hours: List[int] = Field(default_factory=list)
|
||||
# 0-6 with Monday = 0, matching Python's date.weekday(). Required for weekly.
|
||||
schedule_days: List[int] = Field(default_factory=list)
|
||||
# One minute for the whole schedule, so a task with three times is
|
||||
# "09:30, 12:30, 18:30" rather than three separate minute choices.
|
||||
schedule_minute: int = Field(default=0, ge=0, le=59)
|
||||
|
||||
max_notes_count: Optional[int] = Field(default=None, ge=1, le=500)
|
||||
enable_comments: bool = True
|
||||
# Raising this widens the comment window, which is the only lever available
|
||||
@@ -47,12 +60,27 @@ class MonitorTaskCreate(BaseModel):
|
||||
max_comments_count: Optional[int] = Field(default=None, ge=1, le=500)
|
||||
run_timeout_seconds: int = Field(default=3600, ge=60, le=86400)
|
||||
enabled: bool = True
|
||||
# Push a WeCom summary for runs that failed or found new works. Opt-in per
|
||||
# task so a single webhook does not get flooded.
|
||||
# 两类通知分开:新作品可能每轮都有(默认关,避免刷屏),
|
||||
# 异常频率低且意味着任务已经停止工作(默认开,否则你会一直不知道)。
|
||||
notify_enabled: bool = False
|
||||
notify_failures: bool = True
|
||||
# Raw pasted values: full URLs or bare ids, in either form.
|
||||
targets: List[str] = Field(min_length=1)
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _validate_schedule(self) -> "MonitorTaskCreate":
|
||||
if any(hour < 0 or hour > 23 for hour in self.schedule_hours):
|
||||
raise ValueError("小时必须在 0-23 之间")
|
||||
if any(day < 0 or day > 6 for day in self.schedule_days):
|
||||
raise ValueError("星期必须在 0-6 之间(周一为 0)")
|
||||
# A clock mode with no chosen time can never fire. Rejecting it here is
|
||||
# what keeps next_occurrence()'s None branch unreachable in practice.
|
||||
if self.schedule_mode in ("daily", "weekly") and not self.schedule_hours:
|
||||
raise ValueError("按钟点调度至少要选一个时间")
|
||||
if self.schedule_mode == "weekly" and not self.schedule_days:
|
||||
raise ValueError("按周调度至少要选一个星期")
|
||||
return self
|
||||
|
||||
|
||||
class MonitorTaskUpdate(BaseModel):
|
||||
name: Optional[str] = Field(default=None, min_length=1, max_length=200)
|
||||
@@ -60,11 +88,18 @@ class MonitorTaskUpdate(BaseModel):
|
||||
interval_minutes: Optional[int] = Field(
|
||||
default=None, ge=MIN_INTERVAL_MINUTES, le=MAX_INTERVAL_MINUTES
|
||||
)
|
||||
# None means "leave alone". Cross-field validity depends on the merged state,
|
||||
# so it is checked in the service rather than here.
|
||||
schedule_mode: Optional[Literal["interval", "daily", "weekly"]] = None
|
||||
schedule_hours: Optional[List[int]] = None
|
||||
schedule_days: Optional[List[int]] = None
|
||||
schedule_minute: Optional[int] = Field(default=None, ge=0, le=59)
|
||||
max_notes_count: Optional[int] = Field(default=None, ge=1, le=500)
|
||||
enable_comments: Optional[bool] = None
|
||||
max_comments_count: Optional[int] = Field(default=None, ge=1, le=500)
|
||||
run_timeout_seconds: Optional[int] = Field(default=None, ge=60, le=86400)
|
||||
notify_enabled: Optional[bool] = None
|
||||
notify_failures: Optional[bool] = None
|
||||
# When present, replaces the whole target list.
|
||||
targets: Optional[List[str]] = None
|
||||
|
||||
@@ -73,6 +108,18 @@ class CookiePayload(BaseModel):
|
||||
cookie: str = Field(min_length=1)
|
||||
|
||||
|
||||
class CreatorAliasPayload(BaseModel):
|
||||
"""给博主起的备注。空串表示清掉这条备注。"""
|
||||
|
||||
alias: str = Field(default="", max_length=128)
|
||||
|
||||
|
||||
class NoteAliasPayload(BaseModel):
|
||||
"""给作品起的备注。空串表示清掉这条备注。"""
|
||||
|
||||
alias: str = Field(default="", max_length=128)
|
||||
|
||||
|
||||
class WebhookPayload(BaseModel):
|
||||
url: str = Field(default="", description="企业微信机器人 Webhook 地址,留空表示停用")
|
||||
|
||||
|
||||
@@ -20,13 +20,20 @@ import asyncio
|
||||
import subprocess
|
||||
import signal
|
||||
import os
|
||||
from typing import Optional, List
|
||||
from collections import deque
|
||||
from typing import Deque, Optional, List
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
from ..schemas import CrawlerStartRequest, LogEntry
|
||||
from .interpreter import resolve_python_cmd
|
||||
|
||||
# 留住多少行爬虫输出,供 run_and_wait 的调用方诊断失败原因。
|
||||
# 子进程的输出本来只流向日志 WebSocket,监控层只看得到退出码 —— 于是「退出码 1」
|
||||
# 成了运行历史里唯一的信息,真正的报错(比如抖音的 `DataFetchError: account blocked`)
|
||||
# 谁也看不到。留个尾巴,让失败原因能被写进 run.error_message。
|
||||
OUTPUT_TAIL_LINES = 80
|
||||
|
||||
|
||||
class CrawlerManager:
|
||||
"""Crawler process manager"""
|
||||
@@ -49,6 +56,15 @@ class CrawlerManager:
|
||||
# by any concurrent start(), so waiters need an explicit event instead.
|
||||
self._done: asyncio.Event = asyncio.Event()
|
||||
self.last_exit_code: Optional[int] = None
|
||||
# 本次运行输出的末尾若干行。见 OUTPUT_TAIL_LINES。
|
||||
self._output_tail: Deque[str] = deque(maxlen=OUTPUT_TAIL_LINES)
|
||||
|
||||
def get_output_tail(self) -> List[str]:
|
||||
"""最近一次运行的输出尾巴(最早的排前面)。
|
||||
|
||||
只在 run_and_wait() 返回之后读才有意义 —— 它等到读输出的任务收尾才唤醒。
|
||||
"""
|
||||
return list(self._output_tail)
|
||||
|
||||
@property
|
||||
def logs(self) -> List[LogEntry]:
|
||||
@@ -114,6 +130,9 @@ class CrawlerManager:
|
||||
|
||||
async def _push_log(self, entry: LogEntry):
|
||||
"""Push log to queue"""
|
||||
# 这里是所有输出的唯一出口(读循环、收尾、以及管理器自己的提示都走它),
|
||||
# 所以尾巴挂在这儿最省事,也不会漏。
|
||||
self._output_tail.append(entry.message)
|
||||
if self._log_queue is not None:
|
||||
try:
|
||||
self._log_queue.put_nowait(entry)
|
||||
@@ -149,6 +168,8 @@ class CrawlerManager:
|
||||
# Reset completion signalling for this run
|
||||
self._done.clear()
|
||||
self.last_exit_code = None
|
||||
# 尾巴只属于本次运行,否则上一轮的报错会混进这一轮的诊断里。
|
||||
self._output_tail.clear()
|
||||
|
||||
# Clear pending queue (don't replace object to avoid WebSocket broadcast coroutine holding old queue reference)
|
||||
if self._log_queue is None:
|
||||
|
||||
@@ -58,6 +58,14 @@ SAVE_LOGIN_STATE = True
|
||||
# 否则冷 profile 下 API 签名失败,且表现为「退出码 0 但抓到 0 条」的静默失败。
|
||||
INJECT_ALL_COOKIES = False
|
||||
|
||||
# 是否对昵称做中间脱敏(默认 False —— 本仓库**关掉了**)。
|
||||
# 上游作为教学版默认开启,保留首尾各 1 字、中间打星号,避免据昵称骚扰到真人。
|
||||
# 但那是**有损**的:「张三」和「张四」都会变成「张*」,「小明老师」和「小刚老师」
|
||||
# 都会变成「小***师」—— 而本仓库的用途是监控一批公开的创作者账号,分清谁是谁正是
|
||||
# 这一层要干的事,撞名就等于看不出来。所以这里关掉,把原昵称原样落库。
|
||||
# 想改回上游行为,把这一行改成 True 即可(脱敏机制本身没删)。
|
||||
MASK_NICKNAME = False
|
||||
|
||||
# ==================== CDP (Chrome DevTools Protocol) 配置 ====================
|
||||
# 是否启用 CDP 模式 - 使用用户本地的 Chrome/Edge 浏览器进行爬取,具有更好的反检测能力
|
||||
# 开启后,会自动检测并启动用户的 Chrome/Edge 浏览器,通过 CDP 协议进行控制
|
||||
|
||||
@@ -0,0 +1,71 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# 更新这台机器上的部署。用法:
|
||||
#
|
||||
# ./deploy.sh
|
||||
#
|
||||
# 为什么需要脚本而不是一句 `git pull && docker compose up -d`:
|
||||
# 前端产物 api/webui 是 gitignore 的(它由 vite 生成),git pull 带不过来。
|
||||
# 所以代码更新之后必须在服务器上重建一次前端,否则页面还是旧的。
|
||||
# 这一步在容器里做,好处是服务器不需要装 Node —— 只有 Docker。
|
||||
#
|
||||
# node_modules 和 npm 缓存都留在挂载目录内,重复构建不会重新下载。
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")"
|
||||
|
||||
# 用镜像里的解释器跑,宿主机的 Python 版本无关。
|
||||
IMAGE=mediacrawler:latest
|
||||
|
||||
before=$(git rev-parse HEAD)
|
||||
git pull --ff-only
|
||||
after=$(git rev-parse HEAD)
|
||||
|
||||
if [ "$before" = "$after" ]; then
|
||||
echo "== 代码已是最新($after)"
|
||||
else
|
||||
echo "== 代码更新 $before -> $after"
|
||||
git --no-pager log --oneline "$before..$after" | sed 's/^/ /'
|
||||
fi
|
||||
|
||||
# 镜像层(依赖)改动只能靠重建,而这一步不是自动的:Dockerfile 或 requirements.txt 变了,
|
||||
# 下面那句 `docker compose up -d --force-recreate` 用的是旧镜像,改动根本不会生效。
|
||||
# 至少要说出来,否则现象是「代码明明更新了,功能却报缺依赖」。
|
||||
if [ "$before" != "$after" ] && ! git diff --quiet "$before" "$after" -- Dockerfile requirements.txt; then
|
||||
echo "!! Dockerfile / requirements.txt 有改动,需要重建镜像后重跑本脚本:"
|
||||
echo " docker compose build"
|
||||
fi
|
||||
|
||||
# 前端重建的两种情况:产物根本不存在(首次部署),或 webui/ 有改动。
|
||||
if [ ! -f api/webui/index.html ]; then
|
||||
need_build=1
|
||||
reason="前端产物不存在"
|
||||
elif [ "$before" != "$after" ] && ! git diff --quiet "$before" "$after" -- webui/; then
|
||||
need_build=1
|
||||
reason="webui/ 有改动"
|
||||
else
|
||||
need_build=0
|
||||
reason=""
|
||||
fi
|
||||
|
||||
if [ "$need_build" = "1" ]; then
|
||||
echo "== 重建前端($reason)"
|
||||
# -u 1000:1000 而不是 root:这里产出的文件要留在这个目录里给后面用,
|
||||
# 以 root 生成的 node_modules 会让下次构建和人工清理都变得别扭。
|
||||
# HOME 指向挂载目录,这样 npm 的缓存在宿主机上,重建时能复用。
|
||||
docker run --rm \
|
||||
-u 1000:1000 \
|
||||
-w /app/webui \
|
||||
-v "$PWD:/app" \
|
||||
-e HOME=/app/webui \
|
||||
-e npm_config_registry=https://registry.npmmirror.com \
|
||||
"$IMAGE" sh -c 'npm ci --no-audit --no-fund && npm run build'
|
||||
else
|
||||
echo "== 前端无改动,跳过构建"
|
||||
fi
|
||||
|
||||
# --force-recreate,而不是裸的 `up -d`:代码是 bind mount,容器配置和镜像都没变,
|
||||
# 所以 `up -d` 会判定"无需变更"直接跳过,Python 代码的改动根本不会生效。前端产物是
|
||||
# 磁盘上的静态文件,能即时生效,这一点很容易掩盖上面那个问题,直到有人改了 .py 才发现。
|
||||
echo "== 重启容器"
|
||||
docker compose up -d --force-recreate
|
||||
docker compose ps
|
||||
@@ -0,0 +1,33 @@
|
||||
services:
|
||||
mediacrawler:
|
||||
build: .
|
||||
image: mediacrawler:latest
|
||||
container_name: mediacrawler
|
||||
restart: unless-stopped
|
||||
|
||||
# Run as the user that owns this checkout. Without it the container is root,
|
||||
# and every file it writes into the mounted tree -- the crawler's per-run
|
||||
# jsonl output above all -- comes out root-owned. That does not break the app,
|
||||
# but it does lock the operator out of moving or deleting their own
|
||||
# deployment, which is exactly what happened the first time this was deployed.
|
||||
user: "1000:1000"
|
||||
|
||||
# host networking is a requirement, not a convenience: the crawler attaches
|
||||
# to the operator's Chrome at 127.0.0.1:9222, and inside a bridge network
|
||||
# that loopback is the container's own, where no browser is listening.
|
||||
# It also puts the app port directly on the host, so `ports:` is not used.
|
||||
network_mode: host
|
||||
|
||||
env_file:
|
||||
- .env
|
||||
environment:
|
||||
MC_HOST: 0.0.0.0
|
||||
MC_PORT: "18051"
|
||||
TZ: Asia/Shanghai
|
||||
|
||||
volumes:
|
||||
# The code is mounted rather than baked in, so shipping a change is
|
||||
# "git pull, restart" instead of an image rebuild. Only the dependencies
|
||||
# live in the image, because those are the expensive part and they change
|
||||
# rarely -- rebuild only when requirements.txt or the Dockerfile changes.
|
||||
- ./:/app
|
||||
+53
-2
@@ -287,7 +287,7 @@ set MC_PASSWORD=我的新密码 # Windows cmd
|
||||
未接通的平台**可以选,但各页会显示明确的说明面板**,并且**创建任务会被直接拒绝**:
|
||||
|
||||
```
|
||||
400 抖音的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。
|
||||
400 B站的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。
|
||||
```
|
||||
|
||||
而不是接受任务、然后让它永远跑不出数据 —— 那正是之前"博主主页解析失败被误报成登录失效"的同一种静默故障。
|
||||
@@ -297,7 +297,7 @@ set MC_PASSWORD=我的新密码 # Windows cmd
|
||||
| 位置 | 范围 | 内容 |
|
||||
|---|---|---|
|
||||
| 左侧导航「设置」 | **按平台** | 登录 Cookie、采集策略、代理 |
|
||||
| 右上角「系统设置」 | **全局** | 通知、活跃时段、账号安全 |
|
||||
| 右上角「系统设置」 | **全局** | 通知、活跃时段、上游更新、账号安全 |
|
||||
|
||||
**这不是随便分的**:企业微信只有一个群、调度器只有一套时段规则、密码只有一份 ——
|
||||
把它们放进"小红书专属"的页面里,会让人以为它们是按平台存的。
|
||||
@@ -308,6 +308,54 @@ set MC_PASSWORD=我的新密码 # Windows cmd
|
||||
|
||||
---
|
||||
|
||||
## 二·十一、上游更新检查
|
||||
|
||||
本仓库在 [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) 之上加了一整层
|
||||
(监控 / 鉴权 / 多平台面板),差异管理与合并流程在根目录 `UPSTREAM.md` 里。但那份流程有个
|
||||
隐含前提:**得有人知道上游动了**。部署脚本只从我们自己的 Gitea `git pull`,上游的提交不主动
|
||||
去 fetch 就永远看不见 —— 拖着不合并的代价是复利的,越久越难合。
|
||||
|
||||
这一项就是替你定时去 fetch 的:按间隔(默认每天一次)拉一次上游,算出「当前部署落后几个
|
||||
提交」,有更新就推一条企业微信,并把结果与提交列表显示在**右上角「系统设置」→「上游更新」**。
|
||||
|
||||
### 配置
|
||||
|
||||
| 项 | 默认 | 说明 |
|
||||
|---|---|---|
|
||||
| 检查上游仓库更新 | **关** | 总开关。默认关:它要联网 fetch,且需要容器里有 git(见下) |
|
||||
| 上游检查间隔(分钟) | 1440 | 每天一次。最小 30 分钟 |
|
||||
| 上游仓库地址 | GitHub 上游 | 国内直连 GitHub 不稳时改成 gitcode 镜像,见 `UPSTREAM.md` |
|
||||
| 上游分支 | `main` | |
|
||||
| 上游有更新时推送通知 | 开 | 只在出现**此前没推过**的提交时发一条,同一个更新不会反复推 |
|
||||
|
||||
### 几个刻意的行为
|
||||
|
||||
- **只读,不写工作区**:只 `git fetch <地址> <分支>` 到 `FETCH_HEAD` —— 不建 remote、不写
|
||||
`refs/remotes`、不碰索引与工作区。所以它不会打断正在跑的采集,也不会和 `./deploy.sh`
|
||||
的 `git pull` 抢锁。
|
||||
- **不受活跃时段限制**:活跃时段是给采集定的(避免半夜去抓平台)。检查只是 fetch 一个公开
|
||||
仓库,半夜跑反而更合适。
|
||||
- **失败也是一种结果**:上游不通(尤其直连 GitHub)很常见。界面会显示失败原因与上次检查
|
||||
时间,失败不推送,也**不会**因此改变下一次检查的时间 —— 每个间隔重试一次,而不是每个
|
||||
调度 tick(20 秒)都去撞一次。
|
||||
- **同一个更新只推一次**:推送状态记的是上游 tip。推过之后,上下游没动就不会再推;上游又
|
||||
有新提交(tip 变了)时会再推一条。
|
||||
- **「立即检查」不发通知**:点这个按钮的人正看着结果,没必要再给自己推一条群消息。那次
|
||||
检查只写结果,没推的那批提交留给下一次定时检查推。
|
||||
|
||||
### 部署前提:镜像里要有 git
|
||||
|
||||
`python:3.11-slim` 不带 git,`Dockerfile` 里已显式安装。**因此这次更新需要重建镜像**:
|
||||
|
||||
```bash
|
||||
docker compose build && ./deploy.sh
|
||||
```
|
||||
|
||||
`./deploy.sh` 只重建前端,不会重建镜像。漏了这步的话,检查会报「未找到 git 命令」——
|
||||
界面上看得见,不会静默。
|
||||
|
||||
---
|
||||
|
||||
## 三、必须知道的限制
|
||||
|
||||
### 1. 「新增评论」是近似值 —— 最重要的一条
|
||||
@@ -459,6 +507,9 @@ GET /api/monitor/webhook 通知配置状态(**只返回打码
|
||||
POST /api/monitor/webhook 保存 Webhook 地址
|
||||
DELETE /api/monitor/webhook 删除 Webhook
|
||||
POST /api/monitor/webhook/test 发送测试消息
|
||||
|
||||
GET /api/monitor/upstream 最近一次上游检查的缓存结果(没查过返回 {})
|
||||
POST /api/monitor/upstream/check 立刻检查一次(等 fetch 跑完才返回,**不发通知**)
|
||||
```
|
||||
|
||||
> `task_id` 用**重复参数**而非逗号拼接(`?task_id=1&task_id=2`);不传表示统计全部任务。
|
||||
|
||||
@@ -44,7 +44,11 @@ from . import media as douyin_media
|
||||
from .client import DouYinClient
|
||||
from .exception import DataFetchError
|
||||
from .field import PublishTimeType
|
||||
from .help import parse_video_info_from_url, parse_creator_info_from_url
|
||||
from .help import (
|
||||
client_hint_headers,
|
||||
parse_creator_info_from_url,
|
||||
parse_video_info_from_url,
|
||||
)
|
||||
from .login import DouYinLogin
|
||||
|
||||
|
||||
@@ -98,7 +102,12 @@ class DouYinCrawler(AbstractCrawler):
|
||||
await self.browser_context.add_init_script(path="libs/stealth.min.js")
|
||||
|
||||
self.context_page = await self.browser_context.new_page()
|
||||
await self.context_page.goto(self.index_url)
|
||||
# wait_until="domcontentloaded" instead of the default "load": the douyin
|
||||
# home page never fires the load event (some long-lived request keeps it
|
||||
# pending), so the default burns the whole timeout and the crawl dies
|
||||
# before it starts. Measured here: domcontentloaded returns in 0.7s while
|
||||
# load still times out at 90s. Tieba and Zhihu already do the same.
|
||||
await self.context_page.goto(self.index_url, wait_until="domcontentloaded")
|
||||
|
||||
self.dy_client = await self.create_douyin_client(httpx_proxy_format)
|
||||
if not await self.dy_client.pong(browser_context=self.browser_context):
|
||||
@@ -315,10 +324,17 @@ class DouYinCrawler(AbstractCrawler):
|
||||
self.browser_context,
|
||||
urls=self.cookie_urls,
|
||||
) # type: ignore
|
||||
# 声称自己是 Chrome,就得带上 sec-ch-ua 系列头 —— 浏览器一定会带,而缺了它们
|
||||
# 的请求在抖音网关看来就是机器人:回一个 **200 + 空 body**,不报错、不给原因,
|
||||
# 表现为采集抓到 0 条。见 help.client_hint_headers 的实测记录。
|
||||
client_hints = client_hint_headers(
|
||||
await self.context_page.evaluate("() => navigator.userAgentData || null")
|
||||
)
|
||||
douyin_client = DouYinClient(
|
||||
proxy=httpx_proxy,
|
||||
headers={
|
||||
"User-Agent": await self.context_page.evaluate("() => navigator.userAgent"),
|
||||
**client_hints,
|
||||
"Cookie": cookie_str,
|
||||
"Host": "www.douyin.com",
|
||||
"Origin": "https://www.douyin.com/",
|
||||
|
||||
@@ -26,7 +26,7 @@
|
||||
|
||||
import random
|
||||
import re
|
||||
from typing import Optional
|
||||
from typing import Dict, Optional
|
||||
|
||||
import execjs
|
||||
from playwright.async_api import Page
|
||||
@@ -98,6 +98,38 @@ async def get_a_bogus_from_playwright(params: str, post_data: dict, user_agent:
|
||||
return a_bogus
|
||||
|
||||
|
||||
def client_hint_headers(user_agent_data) -> Dict[str, str]:
|
||||
"""由 ``navigator.userAgentData`` 还原 ``sec-ch-ua`` 系列请求头。
|
||||
|
||||
浏览器只要声称自己是 Chrome,就**一定会**带这三个头。缺了它们,「Chrome 的 UA +
|
||||
没有 sec-ch-ua」就是最典型的机器人特征 —— 抖音网关会因此返回 **200 + 空 body**:
|
||||
不报错、不给原因、HTTP 状态还是成功的,表现为采集拿到 0 条。
|
||||
|
||||
实测(同一 URL、同一 cookie、同一参数):不带头 → 0 字节;补上这三个头 → 7077 字节。
|
||||
|
||||
从 ``userAgentData`` 现算而不是写死,是为了 Chrome 升级后不会悄悄失配 —— 写死的
|
||||
版本号和 UA 里的版本号一旦对不上,就又是一个可疑特征。
|
||||
"""
|
||||
if not isinstance(user_agent_data, dict):
|
||||
return {}
|
||||
|
||||
brands = user_agent_data.get("brands") or []
|
||||
sec_ch_ua = ", ".join(
|
||||
f'"{brand.get("brand", "")}";v="{brand.get("version", "")}"' for brand in brands
|
||||
)
|
||||
if not sec_ch_ua:
|
||||
return {}
|
||||
|
||||
headers = {
|
||||
"sec-ch-ua": sec_ch_ua,
|
||||
"sec-ch-ua-mobile": "?1" if user_agent_data.get("mobile") else "?0",
|
||||
}
|
||||
platform = user_agent_data.get("platform")
|
||||
if platform:
|
||||
headers["sec-ch-ua-platform"] = f'"{platform}"'
|
||||
return headers
|
||||
|
||||
|
||||
def parse_video_info_from_url(url: str) -> VideoUrlInfo:
|
||||
"""
|
||||
Parse video ID from Douyin video URL
|
||||
|
||||
@@ -272,3 +272,14 @@ class DouYinLogin(AbstractLogin):
|
||||
'domain': ".douyin.com",
|
||||
'path': "/"
|
||||
}])
|
||||
# Reload after injecting. The page was loaded *before* these cookies existed,
|
||||
# so its `localStorage.HasUserLogin` still holds the logged-out value and
|
||||
# check_login_state() polls it until the timeout expires (10 minutes) without
|
||||
# ever succeeding -- only the *next* run works, because by then the cookies
|
||||
# are in the profile. Reloading makes the site re-evaluate the session now.
|
||||
try:
|
||||
await self.context_page.reload(wait_until="domcontentloaded")
|
||||
except Exception as exc: # pragma: no cover - reload is best effort
|
||||
utils.logger.warning(
|
||||
f"[DouYinLogin.login_by_cookies] reload after cookie injection failed: {exc}"
|
||||
)
|
||||
|
||||
@@ -26,6 +26,37 @@ async def test_cmd_arg_crawler_max_notes_count():
|
||||
config.CRAWLER_MAX_NOTES_COUNT = orig_notes
|
||||
config.CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES = orig_comments
|
||||
|
||||
def test_douyin_monitor_command_uses_the_right_flags():
|
||||
"""抖音监控任务拼出来的命令行。
|
||||
|
||||
与 runner 走的是同一条 _build_command 路径,所以这一条能守住「监控任务的参数
|
||||
没拼错」—— 尤其是平台值必须是 dy(而不是 douyin),否则上游根本认不出平台。
|
||||
"""
|
||||
cm = CrawlerManager()
|
||||
req = CrawlerStartRequest(
|
||||
platform=PlatformEnum.DOUYIN,
|
||||
login_type=LoginTypeEnum.COOKIE,
|
||||
crawler_type=CrawlerTypeEnum.CREATOR,
|
||||
creator_ids="https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X",
|
||||
save_data_path="./data/monitor_runs/1/2",
|
||||
enable_cdp_mode=True,
|
||||
inject_all_cookies=True,
|
||||
save_login_state=True,
|
||||
max_notes_count=20,
|
||||
max_comments_count=50,
|
||||
)
|
||||
cmd = cm._build_command(req)
|
||||
|
||||
idx = cmd.index("--platform")
|
||||
assert cmd[idx + 1] == "dy"
|
||||
idx = cmd.index("--type")
|
||||
assert cmd[idx + 1] == "creator"
|
||||
idx = cmd.index("--creator_id")
|
||||
assert cmd[idx + 1] == "https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X"
|
||||
idx = cmd.index("--enable_cdp_mode")
|
||||
assert cmd[idx + 1] == "true"
|
||||
|
||||
|
||||
def test_crawler_manager_build_command():
|
||||
cm = CrawlerManager()
|
||||
|
||||
|
||||
@@ -0,0 +1,241 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""创作者后台客户端的解析与签名。
|
||||
|
||||
**字段名尚未亲眼验证过**:Phase 0 抓响应时账号的数据权限还没生效,列表接口返回的是
|
||||
空壳(`data.result` 里只有 `{success, code, message}`)。所以这些解析写成多别名匹配,
|
||||
而这份测试就是它的规格 —— 等真实响应到手,先跑这里看哪些假设破了。
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
from api.creator import signing
|
||||
from api.creator.client import (
|
||||
CreatorClient,
|
||||
as_float,
|
||||
as_int,
|
||||
as_seconds,
|
||||
find_note_list,
|
||||
normalize_note,
|
||||
trans_cookies,
|
||||
)
|
||||
from api.creator.models import CreatorAccount
|
||||
from api.creator.service import _account_dict
|
||||
|
||||
|
||||
# --- cookie 解析 -----------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"raw, expected",
|
||||
[
|
||||
("a1=abc; web_session=xyz", {"a1": "abc", "web_session": "xyz"}),
|
||||
("a1=abc;web_session=xyz;", {"a1": "abc", "web_session": "xyz"}),
|
||||
("a1=abc\nweb_session=xyz", {"a1": "abc", "web_session": "xyz"}),
|
||||
(" a1 = abc ; ", {"a1": "abc"}),
|
||||
("", {}),
|
||||
(None, {}),
|
||||
# 值里可以有等号,不能被截断
|
||||
("a1=abc=def", {"a1": "abc=def"}),
|
||||
],
|
||||
)
|
||||
def test_trans_cookies(raw, expected):
|
||||
assert trans_cookies(raw) == expected
|
||||
|
||||
|
||||
def test_client_without_a1_cannot_sign():
|
||||
"""a1 参与签名,没有它连请求都发不出去 —— 要提前拦而不是发出去再猜。"""
|
||||
assert CreatorClient("web_session=abc").looks_authenticated is False
|
||||
assert CreatorClient("a1=abc").looks_authenticated is True
|
||||
|
||||
|
||||
# --- 数值解析 --------------------------------------------------------------
|
||||
#
|
||||
# 后台返回的可能是数字,也可能是 "1.2万" / "12.3%" / "1分30秒" 这类展示值。
|
||||
# 解析不出来一律 None —— 不是 0。0 是真实值,None 是"不知道"。
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"raw, expected",
|
||||
[
|
||||
(123, 123),
|
||||
("123", 123),
|
||||
("1,234", 1234),
|
||||
("1.2万", 12000),
|
||||
("3万", 30000),
|
||||
("1.5w", 15000),
|
||||
("1亿", 100000000),
|
||||
(0, 0),
|
||||
("0", 0),
|
||||
# 这些必须是 None 而不是 0 —— 把"没给"当成"是零"会让报表说谎
|
||||
(None, None),
|
||||
("", None),
|
||||
("-", None),
|
||||
("暂无", None),
|
||||
("abc", None),
|
||||
(True, None),
|
||||
],
|
||||
)
|
||||
def test_as_int(raw, expected):
|
||||
assert as_int(raw) == expected
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"raw, expected",
|
||||
[
|
||||
(12.3, 12.3),
|
||||
("12.3%", 12.3),
|
||||
("12.3", 12.3),
|
||||
(None, None),
|
||||
("-", None),
|
||||
("暂无数据", None),
|
||||
],
|
||||
)
|
||||
def test_as_float(raw, expected):
|
||||
assert as_float(raw) == expected
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"raw, expected",
|
||||
[
|
||||
(45, 45.0),
|
||||
("45", 45.0),
|
||||
("1分30秒", 90.0),
|
||||
("2分", 120.0),
|
||||
("30秒", 30.0),
|
||||
("01:30", 90.0),
|
||||
("1:00:00", 3600.0),
|
||||
(None, None),
|
||||
("-", None),
|
||||
("abc", None),
|
||||
],
|
||||
)
|
||||
def test_as_seconds(raw, expected):
|
||||
assert as_seconds(raw) == expected
|
||||
|
||||
|
||||
# --- 字段归一化 ------------------------------------------------------------
|
||||
|
||||
|
||||
def test_normalize_note_maps_aliases():
|
||||
"""不同来源的记录用不同字段名,别名表要能都接住。"""
|
||||
note = normalize_note(
|
||||
{
|
||||
"note_id": "abc123",
|
||||
"title": "标题",
|
||||
"publish_time": 1700000000000,
|
||||
"view_count": "1.2万",
|
||||
"like_count": 34,
|
||||
"collected_count": 5,
|
||||
"share_count": 2,
|
||||
"comment_count": 7,
|
||||
"cover_click_rate": "12.5%",
|
||||
"avg_watch_time": "1分30秒",
|
||||
}
|
||||
)
|
||||
|
||||
assert note["note_id"] == "abc123"
|
||||
assert note["title"] == "标题"
|
||||
assert note["views"] == 12000
|
||||
assert note["likes"] == 34
|
||||
assert note["favorites"] == 5
|
||||
assert note["shares"] == 2
|
||||
assert note["comments"] == 7
|
||||
assert note["cover_ctr"] == 12.5
|
||||
assert note["avg_watch_seconds"] == 90.0
|
||||
|
||||
|
||||
def test_normalize_note_leaves_missing_fields_as_none():
|
||||
note = normalize_note({"note_id": "abc123"})
|
||||
|
||||
assert note["note_id"] == "abc123"
|
||||
assert note["views"] is None
|
||||
assert note["likes"] is None
|
||||
|
||||
|
||||
def test_find_note_list_digs_the_array_out_of_a_nested_payload():
|
||||
"""接口的确切结构没见过,所以按"像是一批笔记记录"来找,不写死路径。"""
|
||||
payload = {
|
||||
"code": 0,
|
||||
"data": {
|
||||
"result": {
|
||||
"success": True,
|
||||
"notes": [
|
||||
{"note_id": "n1", "views": 10, "likes": 1},
|
||||
{"note_id": "n2", "views": 20, "likes": 2},
|
||||
],
|
||||
}
|
||||
},
|
||||
}
|
||||
|
||||
found = find_note_list(payload)
|
||||
|
||||
assert [item["note_id"] for item in found] == ["n1", "n2"]
|
||||
|
||||
|
||||
def test_find_note_list_returns_empty_for_the_permission_gated_envelope():
|
||||
"""权限未生效时接口返回的就是这个 —— 必须安静地给出空列表,不是报错。"""
|
||||
payload = {
|
||||
"code": 0,
|
||||
"success": True,
|
||||
"msg": "成功",
|
||||
"data": {"result": {"success": True, "code": 0, "message": "success"}},
|
||||
}
|
||||
|
||||
assert find_note_list(payload) == []
|
||||
|
||||
|
||||
# --- 签名 ------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_signed_api_carries_the_url_prefix():
|
||||
"""待签字符串必须带 `url=`。少了它网关返回 406,而 406 的响应体看不出错在哪。"""
|
||||
assert signing.signed_api("/api/galaxy/user/info") == "url=/api/galaxy/user/info"
|
||||
assert (
|
||||
signing.signed_api("/api/x", "a=1&b=2")
|
||||
== "url=/api/x?a=1&b=2"
|
||||
)
|
||||
|
||||
|
||||
def test_sign_returns_xs_and_xt():
|
||||
headers = signing.sign_xyw("url=/api/galaxy/user/info", "some-a1")
|
||||
|
||||
assert set(headers) == {"x-s", "x-t"}
|
||||
assert headers["x-s"].startswith("XYW_")
|
||||
assert headers["x-t"].isdigit()
|
||||
|
||||
|
||||
def test_signature_is_stable_for_a_fixed_timestamp():
|
||||
"""同一输入同一时间戳必须得到同一签名 —— 否则说明有隐藏的随机源。"""
|
||||
first = signing.sign_xyw("url=/api/x", "a1", timestamp_ms=1700000000000)
|
||||
second = signing.sign_xyw("url=/api/x", "a1", timestamp_ms=1700000000000)
|
||||
|
||||
assert first == second
|
||||
|
||||
|
||||
def test_signature_changes_with_the_signed_string():
|
||||
"""签名必须真的绑定待签内容,否则改参数不会被发现 —— 那这个签名就没意义了。"""
|
||||
base = signing.sign_xyw("url=/api/x?a=1", "a1", timestamp_ms=1700000000000)
|
||||
other = signing.sign_xyw("url=/api/x?a=2", "a1", timestamp_ms=1700000000000)
|
||||
|
||||
assert base["x-s"] != other["x-s"]
|
||||
|
||||
|
||||
# --- 凭证不外泄 ------------------------------------------------------------
|
||||
|
||||
|
||||
def test_account_dict_never_carries_the_cookie():
|
||||
"""cookie 等于登录态。对外结构里只该有 `has_cookie`。"""
|
||||
account = CreatorAccount(
|
||||
id=1,
|
||||
nickname="测试",
|
||||
user_id="u1",
|
||||
cookie="a1=SECRET; web_session=SECRET",
|
||||
created_at=0,
|
||||
updated_at=0,
|
||||
)
|
||||
|
||||
payload = _account_dict(account)
|
||||
|
||||
assert payload["has_cookie"] is True
|
||||
assert "cookie" not in payload
|
||||
assert "SECRET" not in str(payload)
|
||||
@@ -0,0 +1,160 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""运营账号扫码登录的完成判据。
|
||||
|
||||
这里守的是一个具体的故障:判据原先读页面里的 `window.__INITIAL_STATE__`,而那是
|
||||
**页面加载那一刻的快照** —— 扫码是加载之后才登录的,快照不会翻转,于是登录明明
|
||||
成功了,界面却永远停在二维码上。
|
||||
|
||||
现在改成拿 cookie 去问创作者后台"我是谁"。重要的是**游客也有 a1**(所以签名算得
|
||||
出来),所以"有 a1"什么都不能证明,只有后台认了才算数。
|
||||
"""
|
||||
|
||||
from unittest.mock import AsyncMock, MagicMock
|
||||
|
||||
import pytest
|
||||
|
||||
from api.creator import login as creator_login
|
||||
from api.creator.client import CreatorApiError
|
||||
|
||||
|
||||
class _FakeContext:
|
||||
def __init__(self, cookies):
|
||||
self._cookies = cookies
|
||||
self.closed = False
|
||||
|
||||
async def cookies(self):
|
||||
return self._cookies
|
||||
|
||||
async def close(self):
|
||||
self.closed = True
|
||||
|
||||
|
||||
class _FakePage:
|
||||
def __init__(self):
|
||||
self.closed = False
|
||||
|
||||
async def close(self):
|
||||
self.closed = True
|
||||
|
||||
|
||||
def _session(cookies):
|
||||
context = _FakeContext(cookies)
|
||||
page = _FakePage()
|
||||
session = creator_login.AccountLoginSession(context, page)
|
||||
# 跳过节流,让每次 refresh 都真的去问一次。
|
||||
session._last_login_check = 0.0
|
||||
return session, context, page
|
||||
|
||||
|
||||
def _patch_client(monkeypatch, behaviour):
|
||||
"""behaviour(cookie) -> dict 或抛异常。"""
|
||||
|
||||
class _Client:
|
||||
def __init__(self, cookie, **kwargs):
|
||||
self.cookie = cookie
|
||||
|
||||
async def fetch_user_info(self):
|
||||
return behaviour(self.cookie)
|
||||
|
||||
monkeypatch.setattr(creator_login, "CreatorClient", _Client)
|
||||
|
||||
|
||||
GUEST_COOKIES = [
|
||||
{"name": "a1", "value": "guest-a1"},
|
||||
{"name": "web_session", "value": "guest-session"},
|
||||
]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_guest_session_never_completes(monkeypatch):
|
||||
"""游客也有 a1,但后台回 401 —— 必须继续等,不能当成登录成功。"""
|
||||
|
||||
def behaviour(_cookie):
|
||||
raise CreatorApiError("登录态无效或已过期", status=401)
|
||||
|
||||
_patch_client(monkeypatch, behaviour)
|
||||
session, _context, _page = _session(GUEST_COOKIES)
|
||||
|
||||
await session.refresh()
|
||||
|
||||
assert session.status == creator_login.STATUS_WAITING
|
||||
assert session.cookie == ""
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_recognised_identity_completes_the_login(monkeypatch):
|
||||
_patch_client(
|
||||
monkeypatch,
|
||||
lambda _cookie: {"user_id": "u123", "nickname": "小明", "red_id": "1"},
|
||||
)
|
||||
session, _context, _page = _session(GUEST_COOKIES)
|
||||
|
||||
await session.refresh()
|
||||
|
||||
assert session.status == creator_login.STATUS_SUCCESS
|
||||
assert "小明" in session.message
|
||||
assert "a1=guest-a1" in session.cookie
|
||||
assert session.account["user_id"] == "u123"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_empty_identity_does_not_complete(monkeypatch):
|
||||
"""接口返回 200 但没有账号标识 —— 同样不能算成功。"""
|
||||
_patch_client(monkeypatch, lambda _cookie: {"user_id": "", "nickname": ""})
|
||||
session, _context, _page = _session(GUEST_COOKIES)
|
||||
|
||||
await session.refresh()
|
||||
|
||||
assert session.status == creator_login.STATUS_WAITING
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_api_is_not_hit_on_every_poll(monkeypatch):
|
||||
"""前端每 2 秒轮询一次,但每次轮询都打一次后台接口是浪费。"""
|
||||
calls = []
|
||||
|
||||
def behaviour(cookie):
|
||||
calls.append(cookie)
|
||||
raise CreatorApiError("还没登录", status=401)
|
||||
|
||||
_patch_client(monkeypatch, behaviour)
|
||||
session, _context, _page = _session(GUEST_COOKIES)
|
||||
|
||||
await session.refresh()
|
||||
# 第一次之后 _last_login_check 已经是"现在",后面两次应落在节流窗口内。
|
||||
await session.refresh()
|
||||
await session.refresh()
|
||||
|
||||
assert len(calls) == 1
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_expiry_beats_a_successful_scan(monkeypatch):
|
||||
_patch_client(monkeypatch, lambda _cookie: {"user_id": "u1", "nickname": "x"})
|
||||
session, _context, _page = _session(GUEST_COOKIES)
|
||||
session.started_at -= creator_login.QR_TTL_SECONDS + 1
|
||||
|
||||
await session.refresh()
|
||||
|
||||
assert session.status == creator_login.STATUS_EXPIRED
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_closed_window_is_reported(monkeypatch):
|
||||
session, context, _page = _session(GUEST_COOKIES)
|
||||
context.cookies = AsyncMock(side_effect=RuntimeError("Target closed"))
|
||||
|
||||
await session.refresh()
|
||||
|
||||
assert session.status == creator_login.STATUS_ERROR
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_closing_discards_the_temporary_context():
|
||||
"""临时上下文是这个设计的关键 —— 用完必须关掉,否则会挂在操作者的 Chrome 里。"""
|
||||
session, context, page = _session(GUEST_COOKIES)
|
||||
|
||||
await session.close()
|
||||
|
||||
assert context.closed is True
|
||||
assert page.closed is True
|
||||
@@ -0,0 +1,401 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""抖音 Web 接口客户端 —— 纯逻辑部分(不发网络请求、不连浏览器)。
|
||||
|
||||
发请求那半边只能在真环境里验(要 CDP 浏览器 + 登录态),所以这里钉住的是那些
|
||||
「错了会一路错到入库」的地方:请求头的成套性、cookie 解析、以及产物键名。
|
||||
"""
|
||||
|
||||
from tools.user_hash import anonymize_user_id
|
||||
|
||||
import asyncio
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
|
||||
from api.monitor import douyin_api
|
||||
|
||||
|
||||
class TestCookieParsing:
|
||||
def test_cookie_header_is_normalised(self):
|
||||
assert douyin_api._cookie_header(" a=1 ; b = 2 ;; c=3 ") == "a=1; b=2; c=3"
|
||||
|
||||
def test_cookie_value_lookup(self):
|
||||
assert douyin_api._cookie_value("a=1; UIFID=xyz; b=2", "UIFID") == "xyz"
|
||||
assert douyin_api._cookie_value("a=1", "UIFID") == ""
|
||||
assert douyin_api._cookie_value("", "UIFID") == ""
|
||||
|
||||
def test_session_detection(self):
|
||||
assert douyin_api._has_session("a=1; sessionid=abc") is True
|
||||
assert douyin_api._has_session("a=1; sessionid_ss=abc") is False
|
||||
assert douyin_api._has_session("") is False
|
||||
|
||||
|
||||
class TestRequestHeaders:
|
||||
"""请求头必须**成套**,而且成套地来自同一个浏览器。
|
||||
|
||||
实测:只有 UA + client hints + Cookie 时,主页接口回 200 但只有 121 字节(空壳);
|
||||
补上 Accept / Accept-Language / Referer 才变成 7074 字节的真数据。
|
||||
"""
|
||||
|
||||
def test_the_full_set_is_sent(self):
|
||||
identity = douyin_api.BrowserIdentity(
|
||||
cookie="sessionid=s; UIFID=u1",
|
||||
user_agent="UA-of-this-browser",
|
||||
client_hints={"sec-ch-ua": '"Chrome";v="155"'},
|
||||
)
|
||||
|
||||
headers = identity.headers()
|
||||
|
||||
assert headers["User-Agent"] == "UA-of-this-browser"
|
||||
assert headers["sec-ch-ua"] == '"Chrome";v="155"'
|
||||
assert headers["Accept"], "Accept 系列是主页接口能不能返回真数据的必要条件"
|
||||
assert headers["Accept-Language"]
|
||||
assert headers["Referer"] == "https://www.douyin.com/"
|
||||
assert headers["x-tt-argus"] == douyin_api.ARGUS_HEADER_VALUE
|
||||
assert headers["uifid"] == "u1"
|
||||
assert headers["Cookie"] == "sessionid=s; UIFID=u1"
|
||||
|
||||
def test_uifid_is_omitted_when_absent(self):
|
||||
"""cookie 里没有 uifid 就别带 —— 送个空值反而更像异常请求。"""
|
||||
identity = douyin_api.BrowserIdentity(
|
||||
cookie="sessionid=s", user_agent="UA", client_hints={}
|
||||
)
|
||||
|
||||
assert "uifid" not in identity.headers()
|
||||
|
||||
def test_uifid_temp_is_used_as_a_fallback(self):
|
||||
identity = douyin_api.BrowserIdentity(
|
||||
cookie="sessionid=s; UIFID_TEMP=temp-1", user_agent="UA", client_hints={}
|
||||
)
|
||||
|
||||
assert identity.headers()["uifid"] == "temp-1"
|
||||
|
||||
|
||||
class TestNormalizeAweme:
|
||||
def test_keys_match_what_the_store_writes(self):
|
||||
"""键名必须和 store/douyin 一模一样,否则 ingest 一条都读不到。"""
|
||||
record = douyin_api.normalize_aweme(
|
||||
{
|
||||
"aweme_id": 7690458980574358513,
|
||||
"desc": "中秋哪儿都堵",
|
||||
"create_time": 1790574515,
|
||||
"author": {"uid": "776719710825195", "nickname": "AA建材王总"},
|
||||
"statistics": {
|
||||
"digg_count": 3,
|
||||
"comment_count": 1,
|
||||
"collect_count": 2,
|
||||
"share_count": 0,
|
||||
},
|
||||
"video": {"cover": {"url_list": ["https://img/cover.jpg"]}},
|
||||
}
|
||||
)
|
||||
|
||||
assert record["aweme_id"] == "7690458980574358513"
|
||||
assert record["title"] == "中秋哪儿都堵"
|
||||
assert record["nickname"] == "AA建材王总"
|
||||
assert record["cover_url"] == "https://img/cover.jpg"
|
||||
assert (
|
||||
record["aweme_url"]
|
||||
== "https://www.douyin.com/video/7690458980574358513"
|
||||
)
|
||||
# 与 store 一致:creator_hash 是 uid 的匿名哈希。
|
||||
assert record["creator_hash"] == anonymize_user_id("776719710825195")
|
||||
# **秒**。adapters 的 time_scale=1000 会把它换成毫秒 —— 这一层不算毫秒。
|
||||
assert record["create_time"] == 1790574515
|
||||
# 指标按 store 的形态落成字符串,交给 ingest 的 parse_count 解析。
|
||||
assert record["liked_count"] == "3"
|
||||
assert record["collected_count"] == "2"
|
||||
|
||||
def test_missing_fields_do_not_crash(self):
|
||||
record = douyin_api.normalize_aweme({"aweme_id": "1"})
|
||||
|
||||
assert record["aweme_id"] == "1"
|
||||
assert record["title"] == ""
|
||||
assert record["cover_url"] == ""
|
||||
assert record["liked_count"] == "0"
|
||||
assert record["create_time"] == 0
|
||||
|
||||
|
||||
class TestIdentityFromPages:
|
||||
"""身份得从浏览器里问,但**不能被一个卡死的标签页拖住**。
|
||||
|
||||
实测过:标签页 URL 为空、渲染进程卡死,``page.evaluate`` 永远不返回;而问身份是采集的
|
||||
第一步 —— 没超时的话整轮就挂在那儿,run 永远停在「运行中」。
|
||||
"""
|
||||
|
||||
class _Page:
|
||||
def __init__(self, url, *, user_agent=None, hang=False):
|
||||
self.url = url
|
||||
self._user_agent = user_agent
|
||||
self._hang = hang
|
||||
self.closed = False
|
||||
|
||||
async def evaluate(self, expression):
|
||||
if self._hang:
|
||||
await asyncio.sleep(30) # 模拟渲染进程卡死
|
||||
if expression.startswith("() => navigator.userAgentData"):
|
||||
return {
|
||||
"brands": [{"brand": "Chrome", "version": "155"}],
|
||||
"mobile": False,
|
||||
"platform": "Linux",
|
||||
}
|
||||
return self._user_agent
|
||||
|
||||
async def close(self):
|
||||
self.closed = True
|
||||
|
||||
class _Context:
|
||||
def __init__(self, pages, temp=None):
|
||||
self.pages = pages
|
||||
self._temp = temp
|
||||
self.made_temp = False
|
||||
|
||||
async def new_page(self):
|
||||
self.made_temp = True
|
||||
if self._temp is None:
|
||||
raise AssertionError("这个用例不该走到临时页")
|
||||
return self._temp
|
||||
|
||||
def _fast_timeout(self, monkeypatch):
|
||||
monkeypatch.setattr(douyin_api, "EVALUATE_TIMEOUT_SECONDS", 0.05)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_hanging_page_is_skipped(self, monkeypatch):
|
||||
self._fast_timeout(monkeypatch)
|
||||
stuck = self._Page("", hang=True)
|
||||
good = self._Page("https://example.com/", user_agent="UA-of-good-page")
|
||||
|
||||
user_agent, hints = await douyin_api._identity_from_pages(
|
||||
self._Context([stuck, good])
|
||||
)
|
||||
|
||||
assert user_agent == "UA-of-good-page"
|
||||
assert hints["sec-ch-ua"] == '"Chrome";v="155"'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_page_without_a_user_agent_is_skipped(self, monkeypatch):
|
||||
self._fast_timeout(monkeypatch)
|
||||
blank = self._Page("", user_agent=None)
|
||||
good = self._Page("https://example.com/", user_agent="UA-of-good-page")
|
||||
|
||||
user_agent, _ = await douyin_api._identity_from_pages(
|
||||
self._Context([blank, good])
|
||||
)
|
||||
|
||||
assert user_agent == "UA-of-good-page"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_it_opens_a_temporary_page_when_nothing_else_works(self, monkeypatch):
|
||||
self._fast_timeout(monkeypatch)
|
||||
stuck = self._Page("", hang=True)
|
||||
temp = self._Page("about:blank", user_agent="UA-of-temp-page")
|
||||
context = self._Context([stuck], temp=temp)
|
||||
|
||||
user_agent, _ = await douyin_api._identity_from_pages(context)
|
||||
|
||||
assert context.made_temp is True
|
||||
assert user_agent == "UA-of-temp-page"
|
||||
assert temp.closed is True, "临时页问完要关掉,别在操作者的浏览器里留垃圾"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_douyin_page_is_preferred(self, monkeypatch):
|
||||
"""有抖音页面就先问它 —— 它才是我们要模仿的那个身份。"""
|
||||
self._fast_timeout(monkeypatch)
|
||||
other = self._Page("https://example.com/", user_agent="UA-of-other")
|
||||
douyin = self._Page("https://www.douyin.com/explore", user_agent="UA-of-douyin")
|
||||
|
||||
user_agent, _ = await douyin_api._identity_from_pages(
|
||||
self._Context([other, douyin])
|
||||
)
|
||||
|
||||
assert user_agent == "UA-of-douyin"
|
||||
|
||||
|
||||
class TestGet:
|
||||
"""`_get` 的失败路径 —— 它们决定了失败会不会被伪装成「这个博主没作品」。"""
|
||||
|
||||
@staticmethod
|
||||
def _client_returning(monkeypatch, status_code: int, text: str):
|
||||
class _Response:
|
||||
def json(self):
|
||||
import json as _json
|
||||
|
||||
return _json.loads(self.text)
|
||||
|
||||
response = _Response()
|
||||
response.status_code = status_code
|
||||
response.text = text
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self):
|
||||
return self
|
||||
|
||||
async def __aexit__(self, *exc):
|
||||
return False
|
||||
|
||||
async def get(self, *args, **kwargs):
|
||||
return response
|
||||
|
||||
monkeypatch.setattr(httpx, "AsyncClient", lambda **kwargs: _Client())
|
||||
|
||||
def _identity(self):
|
||||
return douyin_api.BrowserIdentity(
|
||||
cookie="sessionid=s", user_agent="UA", client_hints={}
|
||||
)
|
||||
|
||||
def test_an_empty_body_is_an_error_not_an_empty_result(self, monkeypatch):
|
||||
"""「200 + 空 body」是网关拒绝请求的典型回应。
|
||||
|
||||
必须当场报错 —— 放过去的话,它会在下游变成「这个博主没作品」,把一次失败伪装成
|
||||
一条正常的空结果。爬虫那条路就是这么栽的,还被翻译成「账号被封」。
|
||||
"""
|
||||
self._client_returning(monkeypatch, 200, "")
|
||||
|
||||
with pytest.raises(douyin_api.DouyinApiError) as excinfo:
|
||||
asyncio.run(douyin_api._get("/x", {}, self._identity()))
|
||||
|
||||
assert "空内容" in str(excinfo.value)
|
||||
|
||||
def test_a_403_carries_the_gateways_own_message(self, monkeypatch):
|
||||
"""抖音难得会说原因,把它带出来,别丢。"""
|
||||
self._client_returning(
|
||||
monkeypatch, 403, "Blocked by ArgusSecurityPlugin Uifid Not Found"
|
||||
)
|
||||
|
||||
with pytest.raises(douyin_api.DouyinApiError) as excinfo:
|
||||
asyncio.run(douyin_api._get("/x", {}, self._identity()))
|
||||
|
||||
assert "403" in str(excinfo.value)
|
||||
assert "Uifid Not Found" in str(excinfo.value)
|
||||
|
||||
def test_a_200_with_data_is_returned_as_is(self, monkeypatch):
|
||||
self._client_returning(monkeypatch, 200, '{"user": {"nickname": "x"}}')
|
||||
|
||||
assert asyncio.run(douyin_api._get("/x", {}, self._identity())) == {
|
||||
"user": {"nickname": "x"}
|
||||
}
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fake_signer(monkeypatch):
|
||||
"""把真正的 ``a_bogus`` 签名换成假的。
|
||||
|
||||
真的那个在 import 的瞬间就要把 ``libs/douyin.js`` 喂给 execjs(还得有 node 和正确的
|
||||
相对路径),单元测试不该依赖这些。**签名本身是实测过的**:带上它是 200 + 真评论,
|
||||
不带是 200 + 空 body。这里只负责钉住「有没有带上、传对了没有」。
|
||||
"""
|
||||
import sys
|
||||
import types
|
||||
|
||||
module = types.ModuleType("media_platform.douyin.help")
|
||||
calls: list = []
|
||||
|
||||
def get_a_bogus_from_js(url: str, params: str, user_agent: str) -> str:
|
||||
calls.append({"url": url, "params": params, "user_agent": user_agent})
|
||||
return "FAKE-BOGUS"
|
||||
|
||||
module.get_a_bogus_from_js = get_a_bogus_from_js
|
||||
monkeypatch.setitem(sys.modules, "media_platform.douyin.help", module)
|
||||
return calls
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def recording_client(monkeypatch):
|
||||
"""记下实际发出去的那次请求。"""
|
||||
sent: dict = {}
|
||||
|
||||
class _Response:
|
||||
status_code = 200
|
||||
text = '{"comments": []}'
|
||||
|
||||
def json(self):
|
||||
return {"comments": []}
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self):
|
||||
return self
|
||||
|
||||
async def __aexit__(self, *exc):
|
||||
return False
|
||||
|
||||
async def get(self, url, **kwargs):
|
||||
sent["url"] = url
|
||||
sent.update(kwargs)
|
||||
return _Response()
|
||||
|
||||
monkeypatch.setattr(httpx, "AsyncClient", lambda **kwargs: _Client())
|
||||
return sent
|
||||
|
||||
|
||||
class TestCommentSigning:
|
||||
"""评论接口**必须**带 a_bogus。
|
||||
|
||||
不带的话网关回 200 + 空 body —— 那在下游会变成「这条作品没有评论」,把一次被挡住的
|
||||
请求伪装成一条正常的空结果。和登录失效长得一模一样,查起来能查半天。
|
||||
"""
|
||||
|
||||
def _identity(self):
|
||||
return douyin_api.BrowserIdentity(
|
||||
cookie="sessionid=s", user_agent="UA", client_hints={}
|
||||
)
|
||||
|
||||
def test_the_signature_is_computed_over_the_unsigned_params(
|
||||
self, fake_signer, recording_client
|
||||
):
|
||||
asyncio.run(
|
||||
douyin_api._get(
|
||||
douyin_api.COMMENT_PATH,
|
||||
{"aweme_id": "123", "count": 20},
|
||||
self._identity(),
|
||||
signed=True,
|
||||
)
|
||||
)
|
||||
|
||||
assert len(fake_signer) == 1
|
||||
# 签名算在**不含 a_bogus** 的那串上 —— 把它自己也算进去是循环的。
|
||||
assert fake_signer[0]["params"] == "aweme_id=123&count=20"
|
||||
assert fake_signer[0]["url"] == douyin_api.COMMENT_PATH
|
||||
assert fake_signer[0]["user_agent"] == "UA"
|
||||
|
||||
def test_the_signature_goes_out_with_the_request(self, fake_signer, recording_client):
|
||||
asyncio.run(
|
||||
douyin_api._get(
|
||||
douyin_api.COMMENT_PATH, {"aweme_id": "123"}, self._identity(), signed=True
|
||||
)
|
||||
)
|
||||
|
||||
assert recording_client["params"]["a_bogus"] == "FAKE-BOGUS"
|
||||
assert recording_client["params"]["aweme_id"] == "123"
|
||||
|
||||
def test_unsigned_calls_never_touch_the_signer(self, fake_signer, recording_client):
|
||||
"""作品 / 详情 / 博主资料三个接口不带签名也照常返回。
|
||||
|
||||
给它们加签名是**没验证过的改动** —— 所以这里钉住「不签」,防止有人图省事把
|
||||
signed=True 改成全局默认。
|
||||
"""
|
||||
asyncio.run(
|
||||
douyin_api._get(douyin_api.POSTS_PATH, {"sec_user_id": "x"}, self._identity())
|
||||
)
|
||||
|
||||
assert fake_signer == []
|
||||
assert "a_bogus" not in recording_client["params"]
|
||||
|
||||
def test_a_broken_signer_is_reported_as_such(self, monkeypatch, recording_client):
|
||||
"""execjs 起不来时要说出是签名失败,而不是让它变成「没评论」。"""
|
||||
import sys
|
||||
import types
|
||||
|
||||
module = types.ModuleType("media_platform.douyin.help")
|
||||
|
||||
def boom(url, params, user_agent):
|
||||
raise RuntimeError("node 没装")
|
||||
|
||||
module.get_a_bogus_from_js = boom
|
||||
monkeypatch.setitem(sys.modules, "media_platform.douyin.help", module)
|
||||
|
||||
with pytest.raises(douyin_api.DouyinApiError, match="a_bogus"):
|
||||
asyncio.run(
|
||||
douyin_api._get(
|
||||
douyin_api.COMMENT_PATH, {"aweme_id": "1"}, self._identity(), signed=True
|
||||
)
|
||||
)
|
||||
@@ -0,0 +1,55 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""sec-ch-ua 系列请求头:为什么必须带、怎么还原。
|
||||
|
||||
抖音网关对「声称自己是 Chrome、却没带 sec-ch-ua」的请求会回 **200 + 空 body** ——
|
||||
不报错、不给原因、HTTP 状态还是成功的,采集侧只看到 0 条。实测同一 URL、同一 cookie、
|
||||
同一参数:不带头 0 字节,补上这三个头 7077 字节。
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
from media_platform.douyin.help import client_hint_headers
|
||||
|
||||
# 实测从真实浏览器抓到的 navigator.userAgentData
|
||||
REAL_UA_DATA = {
|
||||
"brands": [
|
||||
{"brand": "Google Chrome", "version": "155"},
|
||||
{"brand": "Chromium", "version": "155"},
|
||||
{"brand": "Not(A:Brand", "version": "24"},
|
||||
],
|
||||
"mobile": False,
|
||||
"platform": "Linux",
|
||||
}
|
||||
|
||||
|
||||
def test_headers_are_derived_from_user_agent_data():
|
||||
hints = client_hint_headers(REAL_UA_DATA)
|
||||
|
||||
assert hints["sec-ch-ua"] == (
|
||||
'"Google Chrome";v="155", "Chromium";v="155", "Not(A:Brand";v="24"'
|
||||
)
|
||||
assert hints["sec-ch-ua-mobile"] == "?0"
|
||||
assert hints["sec-ch-ua-platform"] == '"Linux"'
|
||||
|
||||
|
||||
def test_mobile_is_reflected():
|
||||
hints = client_hint_headers({**REAL_UA_DATA, "mobile": True})
|
||||
|
||||
assert hints["sec-ch-ua-mobile"] == "?1"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("value", [None, [], "not-a-dict", {}, {"brands": []}])
|
||||
def test_nothing_is_invented_when_user_agent_data_is_unavailable(value):
|
||||
"""拿不到就返回空。
|
||||
|
||||
凭空造一组和 UA 对不上的头,只会变成**另一个**可疑特征 —— 那比不带头更糟。
|
||||
"""
|
||||
assert client_hint_headers(value) == {}
|
||||
|
||||
|
||||
def test_missing_platform_still_sends_the_other_two():
|
||||
hints = client_hint_headers({"brands": REAL_UA_DATA["brands"], "mobile": False})
|
||||
|
||||
assert "sec-ch-ua" in hints
|
||||
assert "sec-ch-ua-mobile" in hints
|
||||
assert "sec-ch-ua-platform" not in hints
|
||||
@@ -0,0 +1,411 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""抖音采集编排 —— 不碰网络,把 douyin_api 整个换掉。
|
||||
|
||||
验的是编排本身:产物落在正确的目录、文件名是 store 那套、以及**作品列表被挡时的退化**
|
||||
(用已知 aweme_id 逐条刷新)—— 那条退化路径决定了今天这个功能是「完全没用」还是
|
||||
「已知作品还能看」。
|
||||
"""
|
||||
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from api.monitor import douyin_api, douyin_fetch
|
||||
|
||||
|
||||
def _profile(**overrides) -> dict:
|
||||
profile = {
|
||||
"creator_hash": "hash",
|
||||
"nickname": "博主",
|
||||
"unique_id": "abc",
|
||||
"fans": 12000,
|
||||
"total_favorited": 83000,
|
||||
"works": 42,
|
||||
"following": 7,
|
||||
}
|
||||
profile.update(overrides)
|
||||
return profile
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def fake_profile(monkeypatch):
|
||||
"""每个用例都挡住「问博主资料」这一跳。
|
||||
|
||||
它是附加信息,不在任何一条编排路径上,但真发出去就会去连 9222 那个浏览器 ——
|
||||
于是所有 creator 用例都会多出一次连接失败、并把 ``errors`` 弄脏。想验它自己的
|
||||
用例再单独覆盖这个 fixture。
|
||||
"""
|
||||
|
||||
async def _profile_call(sec_user_id, *, cookie=""):
|
||||
return _profile()
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_profile", _profile_call)
|
||||
|
||||
|
||||
def _video(aweme_id: str, likes: str = "1") -> dict:
|
||||
return {
|
||||
"aweme_id": aweme_id,
|
||||
"title": f"title-{aweme_id}",
|
||||
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
|
||||
"cover_url": "",
|
||||
"aweme_type": "0",
|
||||
"create_time": 1790574515,
|
||||
"creator_hash": "hash",
|
||||
"nickname": "博主",
|
||||
"liked_count": likes,
|
||||
"comment_count": "0",
|
||||
"collected_count": "0",
|
||||
"share_count": "0",
|
||||
}
|
||||
|
||||
|
||||
def _comment(aweme_id: str, index: int) -> dict:
|
||||
return {
|
||||
"comment_id": f"c{index}",
|
||||
"aweme_id": aweme_id,
|
||||
"content": f"评论{index}",
|
||||
"nickname": "路人",
|
||||
"creator_hash": "h2",
|
||||
"create_time": 1790574600,
|
||||
"like_count": "0",
|
||||
"sub_comment_count": "0",
|
||||
"parent_comment_id": "0",
|
||||
}
|
||||
|
||||
|
||||
class _Target:
|
||||
def __init__(self, external_id: str) -> None:
|
||||
self.external_id = external_id
|
||||
|
||||
|
||||
def _read(path):
|
||||
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line]
|
||||
|
||||
|
||||
async def _collect(tmp_path, **overrides):
|
||||
kwargs = dict(
|
||||
platform="dy",
|
||||
mode="creator",
|
||||
limit=20,
|
||||
want_comments=True,
|
||||
comment_limit=20,
|
||||
targets=[_Target("MS4w-sec")],
|
||||
cookie="sessionid=x",
|
||||
)
|
||||
kwargs.update(overrides)
|
||||
return await douyin_fetch.collect(tmp_path, **kwargs)
|
||||
|
||||
|
||||
class TestHappyPath:
|
||||
@pytest.mark.asyncio
|
||||
async def test_writes_the_layout_ingest_expects(self, monkeypatch, tmp_path):
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return [_video("111"), _video("222")]
|
||||
|
||||
async def fake_comments(aweme_id, count=20, *, cookie=""):
|
||||
return [_comment(aweme_id, 1)]
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
monkeypatch.setattr(douyin_api, "video_comments", fake_comments)
|
||||
|
||||
result = await _collect(tmp_path)
|
||||
|
||||
assert result["notes"] == 2
|
||||
assert result["comments"] == 2
|
||||
assert result["errors"] == []
|
||||
|
||||
# 目录名必须是 douyin(不是平台 id dy)—— ingest 找文件用的是同一个来源。
|
||||
jsonl_dir = tmp_path / "douyin" / "jsonl"
|
||||
assert jsonl_dir.is_dir()
|
||||
|
||||
contents = list(jsonl_dir.glob("*_contents_*.jsonl"))
|
||||
comments = list(jsonl_dir.glob("*_comments_*.jsonl"))
|
||||
assert len(contents) == 1
|
||||
assert len(comments) == 1
|
||||
|
||||
notes = _read(contents[0])
|
||||
assert [n["aweme_id"] for n in notes] == ["111", "222"]
|
||||
# 键名照抄 store/douyin —— ingest 靠这个读出来。
|
||||
assert notes[0]["aweme_url"] == "https://www.douyin.com/video/111"
|
||||
# 秒,不是毫秒;换算交给 adapters。
|
||||
assert notes[0]["create_time"] == 1790574515
|
||||
|
||||
assert _read(comments[0])[0]["aweme_id"] == "111"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_comment_file_exists_even_without_comments(self, monkeypatch, tmp_path):
|
||||
"""评论文件必须建出来。
|
||||
|
||||
ingest 靠「文件在不在」区分「这一轮没评论」和「这一轮什么都没抓到」——
|
||||
两种情况的含义完全不同。
|
||||
"""
|
||||
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return [_video("111")]
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
|
||||
await _collect(tmp_path, want_comments=False)
|
||||
|
||||
assert len(list((tmp_path / "douyin" / "jsonl").glob("*_comments_*.jsonl"))) == 1
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_duplicate_works_are_written_once(self, monkeypatch, tmp_path):
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return [_video("111"), _video("111")]
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
|
||||
result = await _collect(tmp_path, want_comments=False)
|
||||
|
||||
assert result["notes"] == 1
|
||||
|
||||
|
||||
class TestNoteMode:
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_work_target_is_fetched_by_detail_not_by_creator_list(
|
||||
self, monkeypatch, tmp_path
|
||||
):
|
||||
"""作品模式的目标**本身就是作品 id**,不能拿它当博主的 sec_uid 去查列表。
|
||||
|
||||
走错了会必然失败,而且失败原因很难看懂(接口说你没登录/不是浏览器)——
|
||||
「粘贴作品链接的监控」今天本来是能用的,别让它因为这一处走错而废掉。
|
||||
"""
|
||||
|
||||
async def must_not_be_called(*args, **kwargs):
|
||||
raise AssertionError("作品模式不该去拉博主的作品列表")
|
||||
|
||||
async def fake_detail(aweme_id, *, cookie=""):
|
||||
return _video(aweme_id)
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", must_not_be_called)
|
||||
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
|
||||
|
||||
result = await _collect(
|
||||
tmp_path,
|
||||
mode="note",
|
||||
want_comments=False,
|
||||
targets=[_Target("111"), _Target("222")],
|
||||
)
|
||||
|
||||
assert result["notes"] == 2
|
||||
assert result["errors"] == []
|
||||
|
||||
notes = _read(list((tmp_path / "douyin" / "jsonl").glob("*_contents_*.jsonl"))[0])
|
||||
assert [n["aweme_id"] for n in notes] == ["111", "222"]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_broken_work_does_not_lose_the_others(self, monkeypatch, tmp_path):
|
||||
async def flaky(aweme_id, *, cookie=""):
|
||||
if aweme_id == "222":
|
||||
raise douyin_api.DouyinApiError("作品已被删除")
|
||||
return _video(aweme_id)
|
||||
|
||||
monkeypatch.setattr(douyin_api, "video_detail", flaky)
|
||||
|
||||
result = await _collect(
|
||||
tmp_path,
|
||||
mode="note",
|
||||
want_comments=False,
|
||||
targets=[_Target("111"), _Target("222")],
|
||||
)
|
||||
|
||||
assert result["notes"] == 1
|
||||
assert any("222" in error for error in result["errors"])
|
||||
|
||||
|
||||
class TestDegradation:
|
||||
"""作品列表被挡时的行为 —— 决定了这个功能今天有没有用。"""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_falls_back_to_refreshing_known_works(self, monkeypatch, tmp_path):
|
||||
async def blocked(sec_user_id, count=20, *, cookie=""):
|
||||
raise douyin_api.DouyinApiError("接口返回了空内容")
|
||||
|
||||
refreshed = []
|
||||
|
||||
async def fake_detail(aweme_id, *, cookie=""):
|
||||
refreshed.append(aweme_id)
|
||||
return _video(aweme_id, likes="9")
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", blocked)
|
||||
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
|
||||
|
||||
result = await _collect(
|
||||
tmp_path, want_comments=False, known_aweme_ids=["999", "888"]
|
||||
)
|
||||
|
||||
assert refreshed == ["999", "888"]
|
||||
assert result["notes"] == 2
|
||||
# 但错误照样报出来 —— 这一轮是「部分可用」,不是「一切正常」,别粉饰。
|
||||
assert any("作品列表失败" in error for error in result["errors"])
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_known_works_are_not_refetched_for_every_target(self, monkeypatch, tmp_path):
|
||||
"""退化路径不能每个目标都把同一批已知作品再刷一遍。
|
||||
|
||||
任务有多个目标时那会让同一件作品在一轮里出现两次,而指标快照的唯一键是
|
||||
(task_id, note_id, run_id) —— 第二次插入直接撞键,整个 run 崩掉(踩过:
|
||||
Duplicate entry for key 'uq_note_metric')。
|
||||
"""
|
||||
|
||||
async def blocked(sec_user_id, count=20, *, cookie=""):
|
||||
raise douyin_api.DouyinApiError("接口返回了空内容")
|
||||
|
||||
calls = []
|
||||
|
||||
async def fake_detail(aweme_id, *, cookie=""):
|
||||
calls.append(aweme_id)
|
||||
return _video(aweme_id)
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", blocked)
|
||||
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
|
||||
|
||||
result = await _collect(
|
||||
tmp_path,
|
||||
want_comments=False,
|
||||
targets=[_Target("sec-a"), _Target("sec-b")],
|
||||
known_aweme_ids=["999"],
|
||||
)
|
||||
|
||||
assert calls == ["999"], "同一件已知作品只该刷一次,而不是每个目标一次"
|
||||
assert result["notes"] == 1
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_nothing_at_all_still_reports_the_reason(self, monkeypatch, tmp_path):
|
||||
async def blocked(sec_user_id, count=20, *, cookie=""):
|
||||
raise douyin_api.DouyinApiError("接口返回了空内容")
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", blocked)
|
||||
|
||||
result = await _collect(tmp_path, want_comments=False)
|
||||
|
||||
assert result["notes"] == 0
|
||||
assert result["errors"]
|
||||
# 产物仍然写出来(空的),让调用方去判断这是失败而不是「这个博主没作品」。
|
||||
assert (tmp_path / "douyin" / "jsonl").is_dir()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_failing_comment_fetch_does_not_lose_the_work(self, monkeypatch, tmp_path):
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return [_video("111")]
|
||||
|
||||
async def broken_comments(aweme_id, count=20, *, cookie=""):
|
||||
raise douyin_api.DouyinApiError("评论接口抽风")
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
monkeypatch.setattr(douyin_api, "video_comments", broken_comments)
|
||||
|
||||
result = await _collect(tmp_path)
|
||||
|
||||
# 评论拿不到是小事,作品不能跟着丢。
|
||||
assert result["notes"] == 1
|
||||
assert any("评论失败" in error for error in result["errors"])
|
||||
|
||||
|
||||
class TestCreatorProfile:
|
||||
"""博主的**账号级**指标 —— 粉丝 / 总获赞 / 作品数。
|
||||
|
||||
作品列表给不了这个东西:它说的是一件作品涨了多少赞,不是这个人整个账号的粉丝
|
||||
在涨还是在掉。单独问一次资料接口。
|
||||
"""
|
||||
|
||||
@staticmethod
|
||||
def _profiles(tmp_path):
|
||||
files = list((tmp_path / "douyin" / "jsonl").glob("*_profile_*.jsonl"))
|
||||
assert len(files) == 1
|
||||
return _read(files[0])
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_profile_lands_in_the_run_dir(self, monkeypatch, tmp_path):
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return [_video("111")]
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
|
||||
await _collect(tmp_path, want_comments=False)
|
||||
|
||||
assert self._profiles(tmp_path) == [_profile()]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_profile_is_keyed_by_the_same_hash_as_the_works(
|
||||
self, monkeypatch, tmp_path
|
||||
):
|
||||
"""**这条是关键。** 快照表的唯一键是 (任务, creator_hash, 轮次),而界面上是按
|
||||
作品的 creator_hash 归组去查它的。两边只要差一个字符,粉丝数就永远查不出来 ——
|
||||
而且是静默的:表里有数据,界面上什么都没有。
|
||||
"""
|
||||
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return [_video("111")] # 作品带的哈希是 "hash"
|
||||
|
||||
# 资料接口自己算出来的是另一个值(比如它那边 uid 缺字段、只能拿 sec_uid 算)。
|
||||
async def off_hash_profile(sec_user_id, *, cookie=""):
|
||||
return _profile(creator_hash="另一个哈希")
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
monkeypatch.setattr(douyin_api, "author_profile", off_hash_profile)
|
||||
|
||||
await _collect(tmp_path, want_comments=False)
|
||||
|
||||
# 以作品为准:作品才是界面上的行,快照必须挂在能查到它的那个键上。
|
||||
assert self._profiles(tmp_path)[0]["creator_hash"] == "hash"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_profile_without_an_identity_is_dropped(self, monkeypatch, tmp_path):
|
||||
"""哈希都算不出来的快照,落下去只会是一条谁也查不到的垃圾。"""
|
||||
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return []
|
||||
|
||||
async def anonymous_profile(sec_user_id, *, cookie=""):
|
||||
return _profile(creator_hash="")
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
monkeypatch.setattr(douyin_api, "author_profile", anonymous_profile)
|
||||
|
||||
result = await _collect(tmp_path, want_comments=False)
|
||||
|
||||
assert self._profiles(tmp_path) == []
|
||||
assert any("身份标识" in error for error in result["errors"])
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_failing_profile_does_not_lose_the_works(self, monkeypatch, tmp_path):
|
||||
"""附加信息拿不到,这一轮采到的作品不能跟着判成失败。"""
|
||||
|
||||
async def fake_videos(sec_user_id, count=20, *, cookie=""):
|
||||
return [_video("111")]
|
||||
|
||||
async def broken_profile(sec_user_id, *, cookie=""):
|
||||
raise douyin_api.DouyinApiError("资料接口抽风")
|
||||
|
||||
monkeypatch.setattr(douyin_api, "author_videos", fake_videos)
|
||||
monkeypatch.setattr(douyin_api, "author_profile", broken_profile)
|
||||
|
||||
result = await _collect(tmp_path, want_comments=False)
|
||||
|
||||
assert result["notes"] == 1
|
||||
assert self._profiles(tmp_path) == []
|
||||
assert any("资料失败" in error for error in result["errors"])
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_work_target_never_asks_for_a_profile(self, monkeypatch, tmp_path):
|
||||
"""作品模式的目标是一件作品,没有「这个博主是谁」可问 —— 不该白发一个请求。"""
|
||||
|
||||
asked = []
|
||||
|
||||
async def fake_detail(aweme_id, *, cookie=""):
|
||||
return _video(aweme_id)
|
||||
|
||||
async def recording_profile(sec_user_id, *, cookie=""):
|
||||
asked.append(sec_user_id)
|
||||
return _profile()
|
||||
|
||||
monkeypatch.setattr(douyin_api, "video_detail", fake_detail)
|
||||
monkeypatch.setattr(douyin_api, "author_profile", recording_profile)
|
||||
|
||||
await _collect(tmp_path, mode="note", want_comments=False)
|
||||
|
||||
assert asked == []
|
||||
# 文件仍然建出来(空的):事后翻 run 目录能看出「这次根本没问过」。
|
||||
assert self._profiles(tmp_path) == []
|
||||
@@ -41,6 +41,17 @@ from database.models import Base, DouyinAweme, DouyinAwemeComment
|
||||
from tools.user_hash import anonymize_user_id, mask_nickname
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _force_nickname_masking(monkeypatch):
|
||||
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
|
||||
|
||||
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
|
||||
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
|
||||
所以这里显式打开来测。
|
||||
"""
|
||||
monkeypatch.setattr(config, "MASK_NICKNAME", True)
|
||||
|
||||
|
||||
# 抖音教学版禁用字段(键):不得作为存储 dict 的 key 出现。
|
||||
FORBIDDEN_KEYS = {
|
||||
"user_id", "sec_uid", "short_user_id", "user_unique_id",
|
||||
|
||||
@@ -18,10 +18,23 @@ import types
|
||||
|
||||
import pytest
|
||||
|
||||
import config
|
||||
import store.kuaishou as ks
|
||||
from store.kuaishou import update_kuaishou_video, update_ks_video_comment
|
||||
from tools.user_hash import anonymize_user_id, mask_nickname
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _force_nickname_masking(monkeypatch):
|
||||
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
|
||||
|
||||
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
|
||||
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
|
||||
所以这里显式打开来测。
|
||||
"""
|
||||
monkeypatch.setattr(config, "MASK_NICKNAME", True)
|
||||
|
||||
|
||||
# 教学版禁用字段(键):一律不得出现在存储 dict 中。
|
||||
# 昵称字段 nickname 允许保留,但值须脱敏。
|
||||
FORBIDDEN_KEYS = {"user_id", "avatar", "signature", "ip_location", "gender"}
|
||||
|
||||
@@ -78,6 +78,117 @@ class TestParseTargetInput:
|
||||
with pytest.raises(TargetParseError):
|
||||
parse_target_input("not a url at all !!", "creator")
|
||||
|
||||
# --- 抖音 -------------------------------------------------------------
|
||||
# 链接形态由平台决定,所以每一个都要显式带上 "dy"。
|
||||
|
||||
def test_douyin_creator_url(self):
|
||||
parsed = parse_target_input(
|
||||
"https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
|
||||
"?from_tab_name=main",
|
||||
"creator",
|
||||
"dy",
|
||||
)
|
||||
assert (
|
||||
parsed["external_id"]
|
||||
== "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
|
||||
)
|
||||
|
||||
def test_douyin_video_url(self):
|
||||
parsed = parse_target_input(
|
||||
"https://www.douyin.com/video/7525082444551310602", "note", "dy"
|
||||
)
|
||||
assert parsed["external_id"] == "7525082444551310602"
|
||||
|
||||
def test_douyin_modal_id_url(self):
|
||||
"""在别人主页或搜索结果里点开视频,拿到的就是带 modal_id 的链接。"""
|
||||
parsed = parse_target_input(
|
||||
"https://www.douyin.com/root/search/python?aid=b733a3b0&modal_id=7471165520058862848",
|
||||
"note",
|
||||
"dy",
|
||||
)
|
||||
assert parsed["external_id"] == "7471165520058862848"
|
||||
|
||||
def test_douyin_bare_sec_uid_is_accepted(self):
|
||||
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
|
||||
|
||||
parsed = parse_target_input(sec_uid, "creator", "dy")
|
||||
|
||||
assert parsed["external_id"] == sec_uid
|
||||
|
||||
def test_douyin_bare_sec_uid_beyond_the_xhs_length_cap(self):
|
||||
"""裸 id 的长度上限必须按平台分开。
|
||||
|
||||
小红书那条规则封顶 64 字符,而 sec_user_id 长过 64 是常态(实测样本 55,
|
||||
但字段本身是变长的)。共用一条规则的话,长一点的 sec_uid 会被直接拒掉 ——
|
||||
对用户来说就是「粘贴了一个完全正确的链接却报无法识别」。
|
||||
"""
|
||||
sec_uid = "MS4wLjABAAAA" + "aB3dEf6hIj9lMn2pQr5tUv8xYz1" * 3
|
||||
assert len(sec_uid) > 64
|
||||
|
||||
parsed = parse_target_input(sec_uid, "creator", "dy")
|
||||
assert parsed["external_id"] == sec_uid
|
||||
|
||||
# 同一条 id 拿小红书规则来解析会被拒 —— 这正是两条规则必须分开的原因。
|
||||
with pytest.raises(TargetParseError):
|
||||
parse_target_input(sec_uid, "creator", "xhs")
|
||||
|
||||
def test_douyin_bare_video_id(self):
|
||||
parsed = parse_target_input("7525082444551310602", "note", "dy")
|
||||
assert parsed["external_id"] == "7525082444551310602"
|
||||
# 抖音不需要 xsec_token —— 和小红书不同,裸链接就能用。
|
||||
assert parsed["xsec_token"] == ""
|
||||
|
||||
def test_douyin_short_link_is_rejected_with_a_reason(self):
|
||||
"""短链要联网跳一次才知道指向谁。明确拒绝好过存一个永远抓不到东西的目标。"""
|
||||
with pytest.raises(TargetParseError) as excinfo:
|
||||
parse_target_input("https://v.douyin.com/drIPtQ_WPWY/", "note", "dy")
|
||||
assert "短链" in str(excinfo.value)
|
||||
|
||||
def test_a_douyin_link_is_not_parsed_with_xhs_rules(self):
|
||||
with pytest.raises(TargetParseError):
|
||||
parse_target_input(
|
||||
"https://www.douyin.com/video/7525082444551310602", "note", "xhs"
|
||||
)
|
||||
|
||||
def test_an_xhs_link_is_not_parsed_for_douyin(self):
|
||||
with pytest.raises(TargetParseError):
|
||||
parse_target_input(NOTE_URL, "note", "dy")
|
||||
|
||||
def test_a_platform_without_an_adapter_is_rejected(self):
|
||||
with pytest.raises(TargetParseError):
|
||||
parse_target_input("whatever", "creator", "bili")
|
||||
|
||||
|
||||
class TestTargetReplacement:
|
||||
@pytest.mark.asyncio
|
||||
async def test_replacing_targets_uses_the_tasks_own_platform(self, client):
|
||||
"""改目标必须按任务**自己**的平台解析。
|
||||
|
||||
``update_task`` 原先漏传了 platform,解析回落到默认的小红书。只有小红书时
|
||||
行为恰好正确,接上抖音就会拿小红书的正则去解析抖音链接 —— 建任务时对、
|
||||
改任务时错,是最难注意到的那种不一致。
|
||||
"""
|
||||
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
|
||||
created = await client.post(
|
||||
"/api/monitor/tasks",
|
||||
json={"name": "抖音", "mode": "creator", "platform": "dy", "targets": [sec_uid]},
|
||||
)
|
||||
assert created.status_code == 201
|
||||
task_id = created.json()["id"]
|
||||
|
||||
updated = await client.patch(
|
||||
f"/api/monitor/tasks/{task_id}",
|
||||
json={"targets": [f"https://www.douyin.com/user/{sec_uid}"]},
|
||||
)
|
||||
assert updated.status_code == 200
|
||||
|
||||
tasks = (
|
||||
await client.get("/api/monitor/tasks", params={"platform": "dy"})
|
||||
).json()["tasks"]
|
||||
target = tasks[0]["targets"][0]
|
||||
assert target["external_id"] == sec_uid
|
||||
assert target["raw_value"].startswith("https://www.douyin.com/user/")
|
||||
|
||||
|
||||
class TestTaskCrud:
|
||||
@pytest.mark.asyncio
|
||||
|
||||
@@ -0,0 +1,83 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""Guards for the in-place column migration.
|
||||
|
||||
``create_all`` creates missing tables but never adds columns to a table that
|
||||
already exists, so new model columns are applied by ``_ensure_columns``. That step
|
||||
was originally driven by a hand-kept list, and forgetting to update it did not
|
||||
fail loudly -- the app still started, connected, and then failed on every query
|
||||
and every scheduler tick. These tests pin down its replacement, which derives the
|
||||
work from the model metadata.
|
||||
"""
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import Boolean, Column, Integer, MetaData, String, Table
|
||||
|
||||
from api.monitor import db
|
||||
from api.monitor.models import MonitorBase
|
||||
|
||||
|
||||
def _migratable_columns():
|
||||
for table in MonitorBase.metadata.sorted_tables:
|
||||
for column in table.columns:
|
||||
if column.primary_key:
|
||||
continue
|
||||
yield pytest.param(column, id=f"{table.name}.{column.name}")
|
||||
|
||||
|
||||
def _ddl(column: Column) -> str:
|
||||
"""Render a detached column, so the tests never mutate the real metadata."""
|
||||
scratch = Table("scratch", MetaData(), column)
|
||||
return db._column_ddl(scratch.columns[0])
|
||||
|
||||
|
||||
@pytest.mark.parametrize("column", _migratable_columns())
|
||||
def test_column_renders_as_ddl(column):
|
||||
ddl = db._column_ddl(column)
|
||||
|
||||
# SQLAlchemy back-quotes an identifier only when it has to, so the rendered
|
||||
# name matches the column's with the quoting stripped -- which is the case for
|
||||
# reserved words like monitor_run.trigger. That quoting is a feature: the
|
||||
# hand-kept list this replaced would have emitted bare `trigger` and died on a
|
||||
# syntax error.
|
||||
assert ddl.split(" ", 1)[0].strip("`") == column.name
|
||||
# MySQL refuses AUTO_INCREMENT together with the DEFAULT this helper appends
|
||||
# to NOT NULL columns.
|
||||
assert "AUTO_INCREMENT" not in ddl
|
||||
|
||||
|
||||
@pytest.mark.parametrize("column", _migratable_columns())
|
||||
def test_not_null_columns_carry_a_default(column):
|
||||
"""So ADD COLUMN cannot fail on a table that already holds rows.
|
||||
|
||||
Without a DEFAULT, whether the ALTER succeeds depends on the server's
|
||||
sql_mode -- not something a deployment should hinge on.
|
||||
"""
|
||||
if column.nullable:
|
||||
pytest.skip("nullable column needs no seed value")
|
||||
|
||||
assert "DEFAULT" in db._column_ddl(column)
|
||||
|
||||
|
||||
def test_boolean_default_becomes_a_mysql_literal():
|
||||
"""Python's True is not a SQL keyword; it has to become 1."""
|
||||
assert "DEFAULT 1" in _ddl(Column("flag", Boolean, nullable=False, default=True))
|
||||
|
||||
|
||||
def test_string_defaults_are_quoted():
|
||||
assert "DEFAULT 'interval'" in _ddl(
|
||||
Column("mode", String(16), nullable=False, default="interval")
|
||||
)
|
||||
|
||||
|
||||
def test_integer_defaults_are_not_quoted():
|
||||
ddl = _ddl(Column("n", Integer, nullable=False, default=0))
|
||||
|
||||
assert "DEFAULT 0" in ddl
|
||||
assert "DEFAULT '0'" not in ddl
|
||||
|
||||
|
||||
def test_a_not_null_column_without_a_model_default_falls_back_to_zero():
|
||||
"""Belt and braces: even a column the model gives no default still migrates."""
|
||||
ddl = _ddl(Column("n", Integer, nullable=False))
|
||||
|
||||
assert "DEFAULT 0" in ddl
|
||||
@@ -20,6 +20,7 @@
|
||||
|
||||
import csv
|
||||
import io
|
||||
import re
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
@@ -36,6 +37,11 @@ from api.monitor.models import (
|
||||
|
||||
TASK_NAME = "评论归属测试"
|
||||
|
||||
# 作品的发布时间。和 first_seen_at(我们第一次看到它)刻意取不同的值 —— 两者混成
|
||||
# 一个概念是最容易犯的错。
|
||||
PUBLISHED_A = 1_699_000_000_000
|
||||
PUBLISHED_B = 1_699_100_000_000
|
||||
|
||||
|
||||
async def _seed():
|
||||
"""Two works; three comments on the first, one on the second."""
|
||||
@@ -49,13 +55,18 @@ async def _seed():
|
||||
session.add(task)
|
||||
await session.flush()
|
||||
|
||||
for note_id, title in (("note-a", "作品甲"), ("note-b", "作品乙")):
|
||||
# 两个作品**属于不同的博主** —— 评论流最外层按创作者分组,同一个人就没得测了。
|
||||
for note_id, title, creator_hash, creator_name, published_at in (
|
||||
("note-a", "作品甲", "hash-a", "博主甲", PUBLISHED_A),
|
||||
("note-b", "作品乙", "hash-b", "博主乙", PUBLISHED_B),
|
||||
):
|
||||
session.add(
|
||||
MonitorNote(
|
||||
task_id=task.id, note_id=note_id, title=title,
|
||||
note_url=f"https://www.xiaohongshu.com/explore/{note_id}",
|
||||
cover=f"https://img/{note_id}.jpg", creator_hash="h",
|
||||
source_kind="video", published_at=None,
|
||||
cover=f"https://img/{note_id}.jpg", creator_hash=creator_hash,
|
||||
creator_name=creator_name,
|
||||
source_kind="video", published_at=published_at,
|
||||
first_seen_run_id=1, first_seen_at=1_700_000_000_000,
|
||||
last_seen_run_id=1, last_seen_at=1_700_000_000_000,
|
||||
)
|
||||
@@ -139,6 +150,58 @@ class TestGroupByNote:
|
||||
# note-b's only comment is the most recent overall.
|
||||
assert groups[0]["note_id"] == "note-b"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_each_bucket_carries_its_creator(self, client):
|
||||
"""桶上必须带作品的创作者 —— 评论流最外层就是按它分组的。
|
||||
|
||||
少了这两个字段,前端拿到的 creator_hash / creator_name 都是 undefined,
|
||||
于是所有博主塌成同一个分组、标签回退成「未知博主」:一个人都分不出来。
|
||||
"""
|
||||
groups = {
|
||||
group["note_id"]: group
|
||||
for group in (
|
||||
await client.get("/api/monitor/comments", params={"group_by": "note"})
|
||||
).json()["groups"]
|
||||
}
|
||||
|
||||
assert groups["note-a"]["creator_hash"] == "hash-a"
|
||||
assert groups["note-a"]["creator_name"] == "博主甲"
|
||||
assert groups["note-b"]["creator_hash"] == "hash-b"
|
||||
assert groups["note-b"]["creator_name"] == "博主乙"
|
||||
# 两个作品的创作者必须真的不同,否则界面上照样分不出来。
|
||||
assert groups["note-a"]["creator_hash"] != groups["note-b"]["creator_hash"]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_each_bucket_carries_the_publish_date(self, client):
|
||||
"""作品那一层要带发布日期:同名作品不少,日期能帮着认。
|
||||
|
||||
注意它和 first_seen_at 是两个概念 —— 前者是作者发布的那天,后者是我们第一次
|
||||
看到它的那天。把老作品加进监控时两者能差好几个月。
|
||||
"""
|
||||
groups = {
|
||||
group["note_id"]: group
|
||||
for group in (
|
||||
await client.get("/api/monitor/comments", params={"group_by": "note"})
|
||||
).json()["groups"]
|
||||
}
|
||||
|
||||
assert groups["note-a"]["published_at"] == PUBLISHED_A
|
||||
assert groups["note-b"]["published_at"] == PUBLISHED_B
|
||||
assert groups["note-a"]["published_at"] != groups["note-b"]["published_at"]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_notes_endpoint_exposes_the_publish_date(self, client):
|
||||
"""作品列表也要带上它 —— 作品栏就是靠这个显示「发布日期」列的。"""
|
||||
notes = {
|
||||
note["note_id"]: note
|
||||
for note in (await client.get("/api/monitor/notes")).json()["notes"]
|
||||
}
|
||||
|
||||
assert notes["note-a"]["published_at"] == PUBLISHED_A
|
||||
assert notes["note-b"]["published_at"] == PUBLISHED_B
|
||||
# 和「首次发现」不是同一个值 —— 两者混了的话这个断言会抓到。
|
||||
assert notes["note-a"]["first_seen_at"] != notes["note-a"]["published_at"]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_flat_shape_is_unchanged_without_the_flag(self, client):
|
||||
body = (await client.get("/api/monitor/comments")).json()
|
||||
@@ -202,7 +265,8 @@ class TestExport:
|
||||
workbook = load_workbook(io.BytesIO(response.content))
|
||||
sheet = workbook.active
|
||||
assert sheet.max_row == 5 # header + four comments
|
||||
assert sheet.cell(row=1, column=1).value == "所属作品"
|
||||
assert sheet.cell(row=1, column=1).value == "博主昵称"
|
||||
assert sheet.cell(row=1, column=2).value == "所属作品"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_report_export(self, client):
|
||||
@@ -230,3 +294,79 @@ class TestExport:
|
||||
"/api/monitor/export", params={"kind": "comments", "note_id": "no-such-note"}
|
||||
)
|
||||
assert response.status_code == 404
|
||||
|
||||
|
||||
class TestNotesExportColumns:
|
||||
"""作品导出的列 —— 「有列名」和「列里有数」是两回事。
|
||||
|
||||
原先这几列写的是裸键名 `liked_count`,而作品行的指标是嵌在 `metrics` 里的,
|
||||
于是导出来的表有「点赞/评论/收藏/分享」四列,**每一格都是空的**,还没人发现 ——
|
||||
因为原来的测试只断言了 `作品ID`。
|
||||
"""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_metric_columns_actually_contain_numbers(self, client):
|
||||
from api.monitor.models import MonitorNoteMetric
|
||||
|
||||
async with monitor_db.get_session() as session:
|
||||
from sqlalchemy import select
|
||||
|
||||
task_id = (await session.scalar(select(MonitorTask.id))).__int__()
|
||||
session.add(
|
||||
MonitorNoteMetric(
|
||||
task_id=task_id, note_id="note-a", run_id=1,
|
||||
captured_at=1_700_000_000_000,
|
||||
liked_count=123, comment_count=45,
|
||||
collected_count=6, share_count=7,
|
||||
)
|
||||
)
|
||||
|
||||
response = await client.get(
|
||||
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
|
||||
)
|
||||
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
|
||||
by_id = {row["作品ID"]: row for row in rows}
|
||||
|
||||
assert by_id["note-a"]["点赞"] == "123"
|
||||
assert by_id["note-a"]["评论"] == "45"
|
||||
assert by_id["note-a"]["收藏"] == "6"
|
||||
assert by_id["note-a"]["分享"] == "7"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_missing_metric_is_left_empty_not_zero(self, client):
|
||||
"""没采到的指标留空。写 0 的话,导出来的表会声称这条作品零互动。"""
|
||||
response = await client.get(
|
||||
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
|
||||
)
|
||||
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
|
||||
|
||||
assert rows[0]["点赞"] == ""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_export_says_who_the_creator_is(self, client):
|
||||
"""一行只有作品 ID 没法用 —— 导出来是拿去比对和汇报的。
|
||||
|
||||
备注优先:昵称常常认不出是谁,而备注是人自己起的名字。
|
||||
"""
|
||||
response = await client.get(
|
||||
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
|
||||
)
|
||||
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
|
||||
|
||||
assert {row["博主昵称"] for row in rows} == {"博主甲", "博主乙"}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_times_are_readable_not_raw_milliseconds(self, client):
|
||||
"""毫秒时间戳倒进 CSV 就是 13 位数字,打开 Excel 的人没法看、也没法排序。"""
|
||||
response = await client.get(
|
||||
"/api/monitor/export", params={"kind": "notes", "format": "csv"}
|
||||
)
|
||||
rows = list(csv.DictReader(io.StringIO(response.content.decode("utf-8-sig"))))
|
||||
by_id = {row["作品ID"]: row for row in rows}
|
||||
|
||||
# 断言**形状**而不是具体时刻:格式化用的是服务器本地时区,写死一个字符串的话
|
||||
# 换个时区的机器上就会红。
|
||||
assert re.fullmatch(r"\d{4}-\d{2}-\d{2} \d{2}:\d{2}", by_id["note-a"]["发布时间"])
|
||||
assert re.fullmatch(r"\d{4}-\d{2}-\d{2} \d{2}:\d{2}", by_id["note-a"]["首次发现"])
|
||||
# 而且不能是原始毫秒。
|
||||
assert by_id["note-a"]["发布时间"] != str(PUBLISHED_A)
|
||||
|
||||
@@ -0,0 +1,322 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""作品栏里的**博主**:备注(他到底是谁)与账号级指标(他现在多大)。
|
||||
|
||||
按 creator_hash 分组、显示 creator_name,两样都认不出人:一个是哈希,一个是平台昵称。
|
||||
备注是人自己起的名字。账号级指标则是作品列表给不了的东西 —— 作品说的是"这条涨了多少赞",
|
||||
粉丝数说的是"这个人整个账号在涨还是在掉"。
|
||||
"""
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
|
||||
from api.main import app
|
||||
from api.monitor import db as monitor_db
|
||||
from api.monitor.models import (
|
||||
MODE_CREATOR,
|
||||
MonitorCreatorStat,
|
||||
MonitorNote,
|
||||
MonitorTask,
|
||||
)
|
||||
|
||||
CREATOR_HASH = "hash-a"
|
||||
NICKNAME = "张三"
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def client(tmp_path):
|
||||
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
|
||||
await monitor_db.init_db()
|
||||
transport = httpx.ASGITransport(app=app)
|
||||
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
|
||||
yield http_client
|
||||
await monitor_db.dispose_engine()
|
||||
|
||||
|
||||
async def _seed(platform: str = "xhs", task_name: str = "任务") -> int:
|
||||
"""一个任务 + 一条作品,博主固定用 CREATOR_HASH / NICKNAME。"""
|
||||
async with monitor_db.get_session() as session:
|
||||
task = MonitorTask(
|
||||
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
|
||||
interval_minutes=60, max_notes_count=20, enable_comments=False,
|
||||
max_comments_count=50, run_timeout_seconds=3600,
|
||||
notify_enabled=False, created_at=0, updated_at=0,
|
||||
)
|
||||
session.add(task)
|
||||
await session.flush()
|
||||
session.add(
|
||||
MonitorNote(
|
||||
task_id=task.id, note_id=f"{platform}-n1", title="作品",
|
||||
note_url="", cover="", creator_hash=CREATOR_HASH,
|
||||
creator_name=NICKNAME, source_kind="", published_at=None,
|
||||
first_seen_run_id=1, first_seen_at=0,
|
||||
last_seen_run_id=1, last_seen_at=0,
|
||||
)
|
||||
)
|
||||
return task.id
|
||||
|
||||
|
||||
async def _notes(client, platform: str = "xhs"):
|
||||
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["notes"]
|
||||
|
||||
|
||||
async def _set_alias(client, alias: str, platform: str = "xhs", creator_hash=CREATOR_HASH):
|
||||
return await client.put(
|
||||
f"/api/monitor/creators/{creator_hash}",
|
||||
params={"platform": platform},
|
||||
json={"alias": alias},
|
||||
)
|
||||
|
||||
|
||||
class TestCreatorAlias:
|
||||
@pytest.mark.asyncio
|
||||
async def test_notes_start_without_an_alias(self, client):
|
||||
await _seed()
|
||||
|
||||
assert (await _notes(client))[0]["creator_alias"] == ""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_alias_comes_back_with_the_notes(self, client):
|
||||
await _seed()
|
||||
|
||||
response = await _set_alias(client, "竞品A")
|
||||
|
||||
assert response.status_code == 200
|
||||
note = (await _notes(client))[0]
|
||||
assert note["creator_alias"] == "竞品A"
|
||||
# 备注是**叠加**在昵称之上的,不是替换 —— 昵称仍然是有用的对照。
|
||||
assert note["creator_name"] == NICKNAME
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_alias_is_shared_across_tasks(self, client):
|
||||
"""同一个博主出现在两个任务里,备注只该填一次。
|
||||
|
||||
creator_hash 对同一个 uid 是稳定的,所以键取 (platform, creator_hash) 而不是
|
||||
按任务存 —— 否则每加一个任务都要重新认一遍人。
|
||||
"""
|
||||
await _seed(task_name="任务甲")
|
||||
await _seed(task_name="任务乙")
|
||||
await _set_alias(client, "竞品A")
|
||||
|
||||
for note in await _notes(client):
|
||||
assert note["creator_alias"] == "竞品A"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_alias_does_not_leak_to_another_platform(self, client):
|
||||
"""同一个哈希在另一个平台上是另一个(或同一个)人 —— 别串味。"""
|
||||
await _seed(platform="xhs")
|
||||
await _seed(platform="dy")
|
||||
|
||||
await _set_alias(client, "小红书那边的", platform="xhs")
|
||||
|
||||
assert (await _notes(client, "xhs"))[0]["creator_alias"] == "小红书那边的"
|
||||
assert (await _notes(client, "dy"))[0]["creator_alias"] == ""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_empty_alias_clears_it(self, client):
|
||||
await _seed()
|
||||
await _set_alias(client, "竞品A")
|
||||
|
||||
await _set_alias(client, "")
|
||||
|
||||
assert (await _notes(client))[0]["creator_alias"] == ""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_alias_is_trimmed(self, client):
|
||||
await _seed()
|
||||
|
||||
await _set_alias(client, " 竞品A ")
|
||||
|
||||
assert (await _notes(client))[0]["creator_alias"] == "竞品A"
|
||||
|
||||
|
||||
async def _add_stat(
|
||||
task_id: int,
|
||||
run_id: int,
|
||||
fans: int | None,
|
||||
*,
|
||||
total_favorited: int | None = 83000,
|
||||
works: int | None = 42,
|
||||
creator_hash: str = CREATOR_HASH,
|
||||
captured_at: int = 1,
|
||||
) -> None:
|
||||
async with monitor_db.get_session() as session:
|
||||
session.add(
|
||||
MonitorCreatorStat(
|
||||
task_id=task_id,
|
||||
run_id=run_id,
|
||||
creator_hash=creator_hash,
|
||||
nickname=NICKNAME,
|
||||
fans=fans,
|
||||
total_favorited=total_favorited,
|
||||
works_count=works,
|
||||
following=7,
|
||||
captured_at=captured_at,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
class TestCreatorStats:
|
||||
"""账号级指标跟着作品一起返回 —— 界面上是按博主归组的,为了一个组头再发一轮请求
|
||||
没道理。"""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_without_snapshots_the_fields_are_null(self, client):
|
||||
"""null 而不是 0:0 会显示成「粉丝 0」,而事实是"还没采到"。"""
|
||||
await _seed()
|
||||
|
||||
note = (await _notes(client))[0]
|
||||
|
||||
assert note["creator_fans"] is None
|
||||
assert note["creator_total_favorited"] is None
|
||||
assert note["creator_works"] is None
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_snapshot_shows_up_on_every_work_of_that_creator(self, client):
|
||||
task_id = await _seed()
|
||||
await _add_stat(task_id, run_id=1, fans=12000)
|
||||
|
||||
for note in await _notes(client):
|
||||
assert note["creator_fans"] == 12000
|
||||
assert note["creator_total_favorited"] == 83000
|
||||
assert note["creator_works"] == 42
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_latest_run_wins(self, client):
|
||||
"""一轮一条,所以总会有好几条 —— 给界面的必须是最近那条。"""
|
||||
task_id = await _seed()
|
||||
await _add_stat(task_id, run_id=1, fans=12000)
|
||||
await _add_stat(task_id, run_id=2, fans=12300)
|
||||
|
||||
assert (await _notes(client))[0]["creator_fans"] == 12300
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_another_platforms_snapshot_does_not_leak(self, client):
|
||||
"""两个平台上恰好同名同哈希的博主是两个人 —— 快照挂在任务上,不该串。"""
|
||||
await _seed(platform="xhs")
|
||||
dy_task_id = await _seed(platform="dy")
|
||||
await _add_stat(dy_task_id, run_id=1, fans=999)
|
||||
|
||||
assert (await _notes(client, "xhs"))[0]["creator_fans"] is None
|
||||
assert (await _notes(client, "dy"))[0]["creator_fans"] == 999
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_unparsed_count_stays_null(self, client):
|
||||
"""快照在,但某一项没解析出来 —— 那一项必须是 null,不能变成 0。
|
||||
|
||||
和上一条的区别:那条是"根本没有快照",这条是"有快照、其中一项平台没给"。
|
||||
界面上两种都该是「—」。
|
||||
"""
|
||||
task_id = await _seed()
|
||||
await _add_stat(
|
||||
task_id, run_id=1, fans=None, total_favorited=None, works=None
|
||||
)
|
||||
|
||||
note = (await _notes(client))[0]
|
||||
|
||||
assert note["creator_fans"] is None
|
||||
assert note["creator_works"] is None
|
||||
# 快照本身是有的(有采集时间),只是值不知道 —— 前端要能分开这两件事。
|
||||
assert note["creator_stats_at"] is not None
|
||||
|
||||
|
||||
async def _seed_task(platform: str = "xhs", task_name: str = "空任务") -> int:
|
||||
"""只有任务,**一条作品都没有**。"""
|
||||
async with monitor_db.get_session() as session:
|
||||
task = MonitorTask(
|
||||
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
|
||||
interval_minutes=60, max_notes_count=20, enable_comments=False,
|
||||
max_comments_count=50, run_timeout_seconds=3600,
|
||||
notify_enabled=False, created_at=0, updated_at=0,
|
||||
)
|
||||
session.add(task)
|
||||
await session.flush()
|
||||
return task.id
|
||||
|
||||
|
||||
async def _creators(client, platform: str = "xhs"):
|
||||
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["creators"]
|
||||
|
||||
|
||||
class TestCreatorList:
|
||||
"""作品栏要显示谁 —— **包括一条作品都没有的博主**。
|
||||
|
||||
组原先是从作品推出来的,于是「目标加了、资料采到了、粉丝数在库里,界面上什么都
|
||||
没有」。而这类博主恰恰最该看见:还在涨粉,只是最近没发东西。
|
||||
"""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_creator_with_no_works_still_shows_up(self, client):
|
||||
"""**这条就是这个改动的全部理由。**"""
|
||||
task_id = await _seed_task()
|
||||
await _add_stat(task_id, run_id=1, fans=12000)
|
||||
|
||||
creators = await _creators(client)
|
||||
|
||||
assert len(creators) == 1
|
||||
assert creators[0]["creator_hash"] == CREATOR_HASH
|
||||
assert creators[0]["creator_name"] == NICKNAME # 只剩快照这一个来源
|
||||
assert creators[0]["note_count"] == 0
|
||||
assert creators[0]["creator_fans"] == 12000
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_creator_with_works_carries_their_count(self, client):
|
||||
task_id = await _seed()
|
||||
await _add_stat(task_id, run_id=1, fans=12000)
|
||||
|
||||
creators = await _creators(client)
|
||||
|
||||
assert len(creators) == 1
|
||||
assert creators[0]["note_count"] == 1
|
||||
assert creators[0]["creator_name"] == NICKNAME
|
||||
assert creators[0]["creator_fans"] == 12000
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_creator_with_neither_works_nor_a_snapshot_is_absent(self, client):
|
||||
"""两个来源都没有 = 我们对他一无所知,不该凭空造一个组出来。"""
|
||||
await _seed_task()
|
||||
|
||||
assert await _creators(client) == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_work_without_a_snapshot_still_lists_its_creator(self, client):
|
||||
"""小红书那条路不产生账号快照 —— 那边只能靠作品认出人来。"""
|
||||
await _seed()
|
||||
|
||||
creators = await _creators(client)
|
||||
|
||||
assert len(creators) == 1
|
||||
assert creators[0]["note_count"] == 1
|
||||
assert creators[0]["creator_fans"] is None
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_creator_remark_comes_along(self, client):
|
||||
"""没有作品的博主也要能起备注 —— 否则「这是谁」在最需要的时候认不出来。"""
|
||||
task_id = await _seed_task()
|
||||
await _add_stat(task_id, run_id=1, fans=12000)
|
||||
|
||||
await _set_alias(client, "竞品A")
|
||||
|
||||
assert (await _creators(client))[0]["creator_alias"] == "竞品A"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_platforms_stay_apart(self, client):
|
||||
dy_task = await _seed_task(platform="dy")
|
||||
await _add_stat(dy_task, run_id=1, fans=999)
|
||||
|
||||
assert await _creators(client, "xhs") == []
|
||||
assert (await _creators(client, "dy"))[0]["creator_fans"] == 999
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_same_creator_under_two_tasks_is_two_rows_the_ui_merges(self, client):
|
||||
"""服务端按 任务×博主 给(快照就是那么存的),合并交给界面 —— 因为备注跨任务
|
||||
是同一条,合并之后的组才是人眼里的「一个博主」。"""
|
||||
first = await _seed(task_name="任务甲")
|
||||
second = await _seed(task_name="任务乙")
|
||||
await _add_stat(first, run_id=1, fans=12000)
|
||||
|
||||
creators = await _creators(client)
|
||||
|
||||
assert len(creators) == 2
|
||||
assert {row["creator_hash"] for row in creators} == {CREATOR_HASH}
|
||||
assert sum(row["note_count"] for row in creators) == 2
|
||||
@@ -24,6 +24,7 @@ posted/seen comment split, idempotency, and the silent-cookie-failure signal.
|
||||
"""
|
||||
|
||||
import json
|
||||
from datetime import date
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
@@ -35,7 +36,13 @@ from sqlalchemy.pool import StaticPool
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from api.monitor.ingest import describe_exit_code, ingest_run, parse_count
|
||||
from api.monitor import adapters
|
||||
from api.monitor.ingest import (
|
||||
describe_exit_code,
|
||||
diagnose_failure,
|
||||
ingest_run,
|
||||
parse_count,
|
||||
)
|
||||
from api.monitor.models import (
|
||||
EVENT_AUTH_FAILURE,
|
||||
EVENT_METRIC_DELTA,
|
||||
@@ -46,6 +53,8 @@ from api.monitor.models import (
|
||||
EVENT_RUN_FAILED,
|
||||
MODE_CREATOR,
|
||||
MonitorBase,
|
||||
MonitorComment,
|
||||
MonitorCreatorStat,
|
||||
MonitorEvent,
|
||||
MonitorNote,
|
||||
MonitorNoteMetric,
|
||||
@@ -118,9 +127,18 @@ def _write_run_dir(
|
||||
root: Path,
|
||||
notes: List[Dict[str, Any]],
|
||||
comments: Optional[List[Dict[str, Any]]] = None,
|
||||
subdir: str = "xhs",
|
||||
profiles: Optional[List[Dict[str, Any]]] = None,
|
||||
) -> Path:
|
||||
"""Write a run's jsonl output in the crawler's own layout."""
|
||||
jsonl_dir = root / "xhs" / "jsonl"
|
||||
"""Write a run's jsonl output in the crawler's own layout.
|
||||
|
||||
``subdir`` 是**爬虫**落盘的目录名,不是监控层的平台 id —— 抖音那边这两者不同
|
||||
(平台 id 是 ``dy``、目录是 ``douyin``),所以必须能分开指定,否则测不出那个差异。
|
||||
|
||||
``profiles`` 为 None 时**不写**这个文件(小红书那条路根本不产生它),给列表时写
|
||||
——包括空列表,那是「问了但没问到」。
|
||||
"""
|
||||
jsonl_dir = root / subdir / "jsonl"
|
||||
jsonl_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
contents = jsonl_dir / "creator_contents_2026-01-01.jsonl"
|
||||
@@ -134,6 +152,12 @@ def _write_run_dir(
|
||||
"\n".join(json.dumps(c, ensure_ascii=False) for c in comments),
|
||||
encoding="utf-8",
|
||||
)
|
||||
if profiles is not None:
|
||||
profile_file = jsonl_dir / "creator_profile_2026-01-01.jsonl"
|
||||
profile_file.write_text(
|
||||
"\n".join(json.dumps(p, ensure_ascii=False) for p in profiles),
|
||||
encoding="utf-8",
|
||||
)
|
||||
return root
|
||||
|
||||
|
||||
@@ -170,6 +194,52 @@ def _comment(comment_id: str, note_id: str, create_time: int, **extra) -> Dict[s
|
||||
return record
|
||||
|
||||
|
||||
def _dy_note(aweme_id: str, liked: Any = "10", **extra) -> Dict[str, Any]:
|
||||
"""抖音作品记录 —— 键名照抄 store/douyin/__init__.py 的落盘字段。
|
||||
|
||||
重点在于**没有** ``note_id``:抖音叫 ``aweme_id``。这一条差异没映射好,就是
|
||||
每条记录都被 ingest 悄悄 continue 掉、一条不剩。
|
||||
"""
|
||||
record = {
|
||||
"aweme_id": aweme_id,
|
||||
"aweme_type": "0",
|
||||
"title": f"title-{aweme_id}",
|
||||
"desc": f"title-{aweme_id}",
|
||||
# 抖音给的是**秒**(实测 1790574515 = 2026-09-28),小红书给毫秒。落库统一
|
||||
# 换算成毫秒,这个 fixture 必须照真实形态写,否则测不出单位问题。
|
||||
"create_time": 1790574515,
|
||||
"creator_hash": "hash",
|
||||
"nickname": "u***r",
|
||||
"liked_count": liked,
|
||||
"collected_count": "1",
|
||||
"comment_count": "1",
|
||||
"share_count": "1",
|
||||
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
|
||||
"cover_url": "https://img/cover.jpg",
|
||||
}
|
||||
record.update(extra)
|
||||
return record
|
||||
|
||||
|
||||
def _dy_comment(
|
||||
comment_id: str, aweme_id: str, create_time: int, **extra
|
||||
) -> Dict[str, Any]:
|
||||
record = {
|
||||
"comment_id": comment_id,
|
||||
"create_time": create_time,
|
||||
"aweme_id": aweme_id,
|
||||
"content": f"content-{comment_id}",
|
||||
"creator_hash": "hash",
|
||||
"nickname": "u***r",
|
||||
"sub_comment_count": "0",
|
||||
"like_count": "0",
|
||||
# 抖音顶层评论的父 id 是字符串 "0",不是空串。
|
||||
"parent_comment_id": "0",
|
||||
}
|
||||
record.update(extra)
|
||||
return record
|
||||
|
||||
|
||||
async def _events(db: AsyncSession, event_type: Optional[str] = None) -> List[MonitorEvent]:
|
||||
stmt = select(MonitorEvent)
|
||||
if event_type:
|
||||
@@ -525,3 +595,384 @@ class TestIdempotency:
|
||||
assert result.new_notes == 0
|
||||
assert result.new_comments == 0
|
||||
assert len(list((await db.scalars(select(MonitorNote))).all())) == notes_after_first
|
||||
|
||||
|
||||
class TestDouyinIngest:
|
||||
"""抖音的产物形状与小红书不同 —— 这里钉住「不会被静默丢掉」。
|
||||
|
||||
这一组存在的理由,是这个改动最危险的失败模式:字段名或目录名没对上时,ingest
|
||||
不报错,只是**一条都不入库**,然后被当成「疑似登录失效」报出去。
|
||||
"""
|
||||
|
||||
async def _ingest(
|
||||
self,
|
||||
db,
|
||||
tmp_path,
|
||||
notes,
|
||||
comments=None,
|
||||
platform="dy",
|
||||
subdir="douyin",
|
||||
):
|
||||
task = await _make_task(db, platform=platform)
|
||||
run = await _make_run(db, task, started_at=1)
|
||||
_write_run_dir(tmp_path, notes, comments, subdir=subdir)
|
||||
result = await ingest_run(db, run, task, tmp_path)
|
||||
return task, run, result
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_notes_are_ingested_under_their_douyin_field_names(self, db, tmp_path):
|
||||
aweme_id = "7525082444551310602"
|
||||
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note(aweme_id)])
|
||||
|
||||
note = await db.scalar(select(MonitorNote))
|
||||
assert note is not None, "抖音作品被静默丢弃了 —— 多半是 aweme_id 没映射到 note_id"
|
||||
assert note.note_id == aweme_id
|
||||
assert note.note_url == f"https://www.douyin.com/video/{aweme_id}"
|
||||
assert note.cover == "https://img/cover.jpg"
|
||||
assert note.source_kind == "0"
|
||||
# 秒 → 毫秒,换算过才对。
|
||||
assert note.published_at == 1790574515 * 1000
|
||||
assert result.notes_fetched == 1
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_timestamps_are_normalised_to_milliseconds(self, db, tmp_path):
|
||||
"""抖音的时间戳是**秒**,小红书是毫秒 —— 差 1000 倍,必须换算。
|
||||
|
||||
不换算的话,2026 年的作品会显示成 1970 年。这是实测踩到的:抖音作品的
|
||||
「发布日期」列显示成 1970-01-22(1790574515 被当成毫秒就是 21 天后)。
|
||||
"""
|
||||
aweme_id = "7525082444551310602"
|
||||
await self._ingest(db, tmp_path, [_dy_note(aweme_id)])
|
||||
|
||||
note = await db.scalar(select(MonitorNote))
|
||||
|
||||
assert note.published_at == 1790574515 * 1000
|
||||
# 落库的是毫秒,展示层才不用关心来源;乘完应该在 2026 年,不是 1970。
|
||||
assert date.fromtimestamp(note.published_at / 1000).year == 2026
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_artifact_directory_is_not_the_platform_id(self, db, tmp_path):
|
||||
"""目录名与平台 id 不一致,是这套适配里最反直觉的一条。
|
||||
|
||||
抖音的平台 id 是 ``dy``,而爬虫把产物写在 ``douyin/`` 下。把它钉在这里,
|
||||
是为了让「顺手改成一致」这件事会在测试里红掉,而不是让 ingest 悄悄读 0 条。
|
||||
"""
|
||||
assert adapters.artifact_dir("dy") == "douyin"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_writing_into_the_platform_id_directory_reads_nothing(self, db, tmp_path):
|
||||
"""反面:产物落在 ``dy/`` 下时一条都读不到 —— 这正是映射要解决的问题。"""
|
||||
_task, _run, result = await self._ingest(db, tmp_path, [_dy_note("1")], subdir="dy")
|
||||
|
||||
assert result.notes_fetched == 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_misplaced_output_is_blamed_on_the_directory_not_the_login(
|
||||
self, db, tmp_path
|
||||
):
|
||||
"""产物其实抓到了,只是目录名不对 —— 不该报成「疑似登录失效」。
|
||||
|
||||
这是最难查的一类故障:登录是好的、数据也抓到了,但报出来的现象和登录失效
|
||||
一模一样,会把人指去查完全错误的方向。
|
||||
"""
|
||||
_task, run, _result = await self._ingest(
|
||||
db, tmp_path, [_dy_note("1")], subdir="dy"
|
||||
)
|
||||
|
||||
assert any("目录" in event.title for event in await _events(db, EVENT_NO_DATA))
|
||||
assert await _events(db, EVENT_AUTH_FAILURE) == []
|
||||
assert run.error_message and "dy" in run.error_message
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_comments_are_linked_through_aweme_id(self, db, tmp_path):
|
||||
aweme_id = "7525082444551310602"
|
||||
_task, _run, result = await self._ingest(
|
||||
db,
|
||||
tmp_path,
|
||||
[_dy_note(aweme_id)],
|
||||
comments=[_dy_comment("c1", aweme_id, 500)],
|
||||
)
|
||||
|
||||
comment = await db.scalar(select(MonitorComment))
|
||||
assert comment is not None, "抖音评论被静默丢弃了 —— 多半是 aweme_id 没映射"
|
||||
assert comment.note_id == aweme_id
|
||||
assert result.comments_fetched == 1
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_top_level_parent_of_zero_becomes_empty(self, db, tmp_path):
|
||||
"""抖音顶层评论的父 id 是 "0";原样存进去,前端会多出一堆悬空的父节点。"""
|
||||
aweme_id = "7525082444551310602"
|
||||
await self._ingest(
|
||||
db,
|
||||
tmp_path,
|
||||
[_dy_note(aweme_id)],
|
||||
comments=[
|
||||
_dy_comment("c1", aweme_id, 500),
|
||||
_dy_comment("c2", aweme_id, 600, parent_comment_id="c1"),
|
||||
],
|
||||
)
|
||||
|
||||
by_id = {c.comment_id: c for c in (await db.scalars(select(MonitorComment))).all()}
|
||||
assert by_id["c1"].parent_comment_id == ""
|
||||
assert by_id["c2"].parent_comment_id == "c1"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_four_metrics_need_no_mapping(self, db, tmp_path):
|
||||
"""四个指标键两边同名 —— 抖音作品照样进 monitor_note_metric,差分照常。"""
|
||||
aweme_id = "7525082444551310602"
|
||||
task = await _make_task(db, platform="dy")
|
||||
|
||||
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="100")], subdir="douyin")
|
||||
run1 = await _make_run(db, task, started_at=1)
|
||||
await ingest_run(db, run1, task, tmp_path)
|
||||
|
||||
metric = await db.scalar(select(MonitorNoteMetric))
|
||||
assert metric is not None and metric.liked_count == 100
|
||||
|
||||
_write_run_dir(tmp_path, [_dy_note(aweme_id, liked="150")], subdir="douyin")
|
||||
run2 = await _make_run(db, task, started_at=2)
|
||||
await ingest_run(db, run2, task, tmp_path)
|
||||
|
||||
assert len(await _events(db, EVENT_METRIC_DELTA)) == 1
|
||||
|
||||
|
||||
class TestDuplicateRecordsInOneRun:
|
||||
"""一轮产物里重复出现的作品只该算一次。
|
||||
|
||||
指标快照的唯一键是 ``(task_id, note_id, run_id)``:同一条作品在一轮里进来两次,
|
||||
第二次插入会撞键并让整个 run 崩掉 —— 产物里重复并不罕见(多个目标指向同一个人、
|
||||
或退化路径重复刷新)。
|
||||
"""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_duplicated_note_is_processed_once(self, db, tmp_path):
|
||||
task = await _make_task(db)
|
||||
_write_run_dir(tmp_path, [_note("n1"), _note("n1")])
|
||||
run = await _make_run(db, task, started_at=1)
|
||||
|
||||
result = await ingest_run(db, run, task, tmp_path)
|
||||
|
||||
assert result.new_notes == 1
|
||||
assert len(list((await db.scalars(select(MonitorNote))).all())) == 1
|
||||
# 崩就崩在这一句上:一条作品只能有一份本轮快照。
|
||||
assert len(list((await db.scalars(select(MonitorNoteMetric))).all())) == 1
|
||||
|
||||
|
||||
class TestNicknameRefresh:
|
||||
"""已入库的评论,昵称要跟着重新采集的值走。
|
||||
|
||||
评论是去重后直接跳过的,若不刷新,脱敏开关一改(或评论者改了昵称),老数据就永远
|
||||
停在旧值上 —— 而重采是唯一能拿到新值的途径。
|
||||
"""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_existing_comment_gets_its_nickname_refreshed(self, db, tmp_path):
|
||||
task = await _make_task(db)
|
||||
_write_run_dir(tmp_path, [_note("n1")], comments=[_comment("c1", "n1", 500)])
|
||||
run1 = await _make_run(db, task, started_at=1)
|
||||
await ingest_run(db, run1, task, tmp_path)
|
||||
|
||||
assert (await db.scalar(select(MonitorComment))).nickname == "u***r"
|
||||
|
||||
_write_run_dir(
|
||||
tmp_path,
|
||||
[_note("n1")],
|
||||
comments=[_comment("c1", "n1", 500, nickname="未脱敏的新昵称")],
|
||||
)
|
||||
run2 = await _make_run(db, task, started_at=2)
|
||||
await ingest_run(db, run2, task, tmp_path)
|
||||
|
||||
comment = await db.scalar(select(MonitorComment))
|
||||
assert comment.nickname == "未脱敏的新昵称"
|
||||
# 去重的语义没变:同一条评论不该被插成两行。
|
||||
assert (
|
||||
len(list((await db.scalars(select(MonitorComment))).all())) == 1
|
||||
)
|
||||
|
||||
|
||||
class TestFailureDiagnosis:
|
||||
"""失败原因要能被人看懂。
|
||||
|
||||
只写「退出码 1」等于什么都没说 —— 真正的报错埋在子进程的 stderr 里,而运行历史
|
||||
里那一格显示的正是 run.error_message。
|
||||
"""
|
||||
|
||||
TAIL = [
|
||||
"2026-10-10 15:18:34 MediaCrawler INFO (core.py:385) - [DouYinCrawler] CDP浏览器信息",
|
||||
"Traceback (most recent call last):",
|
||||
' File "/app/main.py", line 114, in main',
|
||||
" await crawler.start()",
|
||||
"media_platform.douyin.exception.DataFetchError: account blocked, ",
|
||||
]
|
||||
|
||||
def test_the_exception_line_is_picked_out_of_the_tail(self):
|
||||
assert (
|
||||
diagnose_failure(self.TAIL)
|
||||
== "media_platform.douyin.exception.DataFetchError: account blocked,"
|
||||
)
|
||||
|
||||
def test_the_managers_own_lines_are_not_mistaken_for_the_cause(self):
|
||||
"""管理器自己补的那两句不是爬虫的报错,别被当成失败原因。"""
|
||||
assert diagnose_failure(["Crawler exited with code: 1"]) is None
|
||||
assert diagnose_failure(["Crawler completed successfully"]) is None
|
||||
|
||||
def test_nothing_to_say_is_not_an_error(self):
|
||||
assert diagnose_failure(None) is None
|
||||
assert diagnose_failure([]) is None
|
||||
|
||||
def test_it_falls_back_to_the_last_line(self):
|
||||
assert (
|
||||
diagnose_failure(["started fine", "then something odd"])
|
||||
== "then something odd"
|
||||
)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_failed_run_records_both_the_code_and_the_cause(self, db, tmp_path):
|
||||
task = await _make_task(db)
|
||||
run = await _make_run(db, task, started_at=1, exit_code=1)
|
||||
|
||||
result = await ingest_run(db, run, task, tmp_path, output_tail=self.TAIL)
|
||||
|
||||
assert result.status == RUN_FAILED
|
||||
# 退出码和真因都要在,缺一个都还得去翻日志。
|
||||
assert "code 1" in run.error_message
|
||||
assert "account blocked" in run.error_message
|
||||
|
||||
events = await _events(db, EVENT_RUN_FAILED)
|
||||
assert "account blocked" in events[0].title
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# 博主账号级指标(粉丝 / 总获赞 / 作品数)
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _profile(creator_hash: str = "hash", **extra) -> Dict[str, Any]:
|
||||
"""``creator_profile_*.jsonl`` 里的一行 —— 形状由 douyin_api.author_profile 决定。"""
|
||||
record: Dict[str, Any] = {
|
||||
"creator_hash": creator_hash,
|
||||
"nickname": "博主",
|
||||
"unique_id": "abc",
|
||||
"fans": 12000,
|
||||
"total_favorited": 83000,
|
||||
"works": 42,
|
||||
"following": 7,
|
||||
}
|
||||
record.update(extra)
|
||||
return record
|
||||
|
||||
|
||||
class TestCreatorStatSnapshots:
|
||||
"""**账号级**指标和作品级指标是两回事:后者说"这条视频涨了多少赞",前者说
|
||||
"这个人整个账号的粉丝在涨还是在掉"。作品列表给不了后者,所以单独存一张表。
|
||||
"""
|
||||
|
||||
async def _ingest(self, db, tmp_path, notes, profiles, platform="dy", subdir="douyin"):
|
||||
task = await _make_task(db, platform=platform)
|
||||
run = await _make_run(db, task, started_at=1)
|
||||
_write_run_dir(tmp_path, notes, comments=[], subdir=subdir, profiles=profiles)
|
||||
result = await ingest_run(db, run, task, tmp_path)
|
||||
return task, run, result
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_profile_becomes_a_snapshot(self, db, tmp_path):
|
||||
task, run, _result = await self._ingest(db, tmp_path, [_dy_note("1")], [_profile()])
|
||||
|
||||
stat = await db.scalar(select(MonitorCreatorStat))
|
||||
|
||||
assert stat is not None
|
||||
assert (stat.task_id, stat.run_id) == (task.id, run.id)
|
||||
assert stat.creator_hash == "hash"
|
||||
assert stat.nickname == "博主"
|
||||
assert (stat.fans, stat.total_favorited, stat.works_count) == (12000, 83000, 42)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_missing_count_stays_null_not_zero(self, db, tmp_path):
|
||||
"""0 是真实值(掉到零),null 是不知道。混起来趋势图就是在撒谎。"""
|
||||
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(fans=None)])
|
||||
|
||||
stat = await db.scalar(select(MonitorCreatorStat))
|
||||
|
||||
assert stat.fans is None
|
||||
# 同一个博主其它字段照常。
|
||||
assert stat.works_count == 42
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_abbreviated_counts_are_parsed(self, db, tmp_path):
|
||||
"""走的是和作品指标同一个 parse_count —— 平台给你「1.2万」也得认。"""
|
||||
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(fans="1.2万")])
|
||||
|
||||
assert (await db.scalar(select(MonitorCreatorStat))).fans == 12000
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_same_creator_twice_in_one_run_yields_one_snapshot(self, db, tmp_path):
|
||||
"""一个任务可以配多个目标,退化路径下它们可能落在同一个博主身上。
|
||||
|
||||
唯一键是 (task_id, creator_hash, run_id) —— 重复插入会撞键,把整轮炸掉。
|
||||
(和作品重复那次是同一类事故。)
|
||||
"""
|
||||
await self._ingest(
|
||||
db, tmp_path, [_dy_note("1")], [_profile(), _profile(nickname="另一条")]
|
||||
)
|
||||
|
||||
stats = list((await db.scalars(select(MonitorCreatorStat))).all())
|
||||
|
||||
assert len(stats) == 1
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_profile_without_a_hash_is_skipped(self, db, tmp_path):
|
||||
"""哈希都算不出来,这条快照谁也查不到,落下去只是垃圾。"""
|
||||
await self._ingest(db, tmp_path, [_dy_note("1")], [_profile(creator_hash="")])
|
||||
|
||||
assert (await db.scalars(select(MonitorCreatorStat))).all() == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_no_profile_file_is_fine(self, db, tmp_path):
|
||||
"""小红书那条路(爬虫进程)根本不产生这个文件 —— 不能因此报错。"""
|
||||
task = await _make_task(db) # xhs
|
||||
run = await _make_run(db, task, started_at=1)
|
||||
_write_run_dir(tmp_path, [_note("n1")], comments=[]) # 没有 profiles 参数
|
||||
|
||||
result = await ingest_run(db, run, task, tmp_path)
|
||||
|
||||
assert result.status == RUN_SUCCESS
|
||||
assert (await db.scalars(select(MonitorCreatorStat))).all() == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_stats_are_kept_even_when_no_works_were_fetched(self, db, tmp_path):
|
||||
"""**这条是这里最值得留的一个。**
|
||||
|
||||
作品列表被风控挡住时,这一轮一条作品都拿不到、run 会被判成失败。但博主的粉丝数
|
||||
并不因为这件事就不存在 —— 「粉丝还在涨,但新作品没在发现」恰恰是最该看见的时刻。
|
||||
快照要是挂在「作品采到了」后面,就正好在最需要它的那一轮丢掉。
|
||||
"""
|
||||
_task, run, result = await self._ingest(db, tmp_path, [], [_profile()])
|
||||
|
||||
assert result.status == RUN_PARTIAL # 一条作品都没有,这轮确实不算成功
|
||||
assert run.status == RUN_PARTIAL
|
||||
stat = await db.scalar(select(MonitorCreatorStat))
|
||||
assert stat is not None and stat.fans == 12000
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_each_run_adds_its_own_snapshot(self, db, tmp_path):
|
||||
"""趋势靠的就是这个:一条一轮,不要覆盖。"""
|
||||
task = await _make_task(db, platform="dy")
|
||||
first = await _make_run(db, task, started_at=1)
|
||||
_write_run_dir(tmp_path, [_dy_note("1")], comments=[], subdir="douyin",
|
||||
profiles=[_profile(fans=12000)])
|
||||
await ingest_run(db, first, task, tmp_path)
|
||||
|
||||
second = await _make_run(db, task, started_at=2000)
|
||||
_write_run_dir(tmp_path, [_dy_note("1")], comments=[], subdir="douyin",
|
||||
profiles=[_profile(fans=12300)])
|
||||
await ingest_run(db, second, task, tmp_path)
|
||||
|
||||
stats = list(
|
||||
(
|
||||
await db.scalars(
|
||||
select(MonitorCreatorStat).order_by(MonitorCreatorStat.run_id)
|
||||
)
|
||||
).all()
|
||||
)
|
||||
|
||||
assert [s.fans for s in stats] == [12000, 12300]
|
||||
|
||||
@@ -0,0 +1,146 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""作品备注 —— 一个博主底下,哪几条是真正要盯的。
|
||||
|
||||
和博主备注(test_monitor_creators.py)是一对,但回答的不是同一个问题:博主备注回答
|
||||
「这个账号是谁」,作品备注回答「这条作品我要盯着」。一个博主底下常常只有一两件值得
|
||||
盯的作品,所以不能合成一条。
|
||||
"""
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
|
||||
from api.main import app
|
||||
from api.monitor import db as monitor_db
|
||||
from api.monitor.models import MODE_CREATOR, MonitorNote, MonitorTask
|
||||
|
||||
NOTE_ID = "note-a"
|
||||
TITLE = "中秋哪儿都堵"
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def client(tmp_path):
|
||||
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
|
||||
await monitor_db.init_db()
|
||||
transport = httpx.ASGITransport(app=app)
|
||||
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
|
||||
yield http_client
|
||||
await monitor_db.dispose_engine()
|
||||
|
||||
|
||||
async def _seed(platform: str = "xhs", task_name: str = "任务", note_id: str = NOTE_ID) -> int:
|
||||
async with monitor_db.get_session() as session:
|
||||
task = MonitorTask(
|
||||
name=task_name, platform=platform, mode=MODE_CREATOR, enabled=True,
|
||||
interval_minutes=60, max_notes_count=20, enable_comments=False,
|
||||
max_comments_count=50, run_timeout_seconds=3600,
|
||||
notify_enabled=False, created_at=0, updated_at=0,
|
||||
)
|
||||
session.add(task)
|
||||
await session.flush()
|
||||
session.add(
|
||||
MonitorNote(
|
||||
task_id=task.id, note_id=note_id, title=TITLE,
|
||||
note_url="", cover="", creator_hash="hash-a",
|
||||
creator_name="张三", source_kind="", published_at=None,
|
||||
first_seen_run_id=1, first_seen_at=0,
|
||||
last_seen_run_id=1, last_seen_at=0,
|
||||
)
|
||||
)
|
||||
return task.id
|
||||
|
||||
|
||||
async def _notes(client, platform: str = "xhs"):
|
||||
return (await client.get("/api/monitor/notes", params={"platform": platform})).json()["notes"]
|
||||
|
||||
|
||||
async def _set_alias(client, alias: str, platform: str = "xhs", note_id: str = NOTE_ID):
|
||||
return await client.put(
|
||||
f"/api/monitor/notes/{note_id}",
|
||||
params={"platform": platform},
|
||||
json={"alias": alias},
|
||||
)
|
||||
|
||||
|
||||
class TestNoteAlias:
|
||||
@pytest.mark.asyncio
|
||||
async def test_notes_start_without_a_remark(self, client):
|
||||
await _seed()
|
||||
|
||||
assert (await _notes(client))[0]["note_alias"] == ""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_remark_comes_back_with_the_notes(self, client):
|
||||
await _seed()
|
||||
|
||||
response = await _set_alias(client, "重点")
|
||||
|
||||
assert response.status_code == 200
|
||||
note = (await _notes(client))[0]
|
||||
assert note["note_alias"] == "重点"
|
||||
# 备注是**叠加**在标题之上的,不是替换 —— 标题仍然是这条作品本身。
|
||||
assert note["title"] == TITLE
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_remark_is_shared_across_tasks(self, client):
|
||||
"""同一件作品被两个任务都监控时,备注只该填一次。"""
|
||||
await _seed(task_name="任务甲")
|
||||
await _seed(task_name="任务乙")
|
||||
await _set_alias(client, "重点")
|
||||
|
||||
for note in await _notes(client):
|
||||
assert note["note_alias"] == "重点"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_remark_does_not_leak_to_another_platform(self, client):
|
||||
"""作品的 id 是平台各自的编号体系 —— 抖音的 123 和小红书的 123 是两条作品。"""
|
||||
await _seed(platform="xhs")
|
||||
await _seed(platform="dy")
|
||||
|
||||
await _set_alias(client, "小红书那边的", platform="xhs")
|
||||
|
||||
assert (await _notes(client, "xhs"))[0]["note_alias"] == "小红书那边的"
|
||||
assert (await _notes(client, "dy"))[0]["note_alias"] == ""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_remark_does_not_leak_to_another_work(self, client):
|
||||
"""钉住这里的**作用域**:键是 note_id。写错成按任务存的话,给一条起了备注,
|
||||
同一个博主底下的其它作品会跟着一起变 —— 那这个功能就没用了。"""
|
||||
await _seed(note_id="note-a")
|
||||
await _seed(task_name="另一个任务", note_id="note-b")
|
||||
|
||||
await _set_alias(client, "重点", note_id="note-a")
|
||||
|
||||
by_id = {note["note_id"]: note["note_alias"] for note in await _notes(client)}
|
||||
assert by_id == {"note-a": "重点", "note-b": ""}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_empty_remark_clears_it(self, client):
|
||||
await _seed()
|
||||
await _set_alias(client, "重点")
|
||||
|
||||
await _set_alias(client, "")
|
||||
|
||||
assert (await _notes(client))[0]["note_alias"] == ""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_remark_is_trimmed(self, client):
|
||||
await _seed()
|
||||
|
||||
await _set_alias(client, " 重点 ")
|
||||
|
||||
assert (await _notes(client))[0]["note_alias"] == "重点"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_two_kinds_of_remark_stay_apart(self, client):
|
||||
"""博主备注和作品备注是两张表、两个键 —— 一个不该把另一个盖掉。"""
|
||||
await _seed()
|
||||
|
||||
await _set_alias(client, "重点")
|
||||
await client.put(
|
||||
"/api/monitor/creators/hash-a", params={"platform": "xhs"}, json={"alias": "竞品A"}
|
||||
)
|
||||
|
||||
note = (await _notes(client))[0]
|
||||
assert note["note_alias"] == "重点"
|
||||
assert note["creator_alias"] == "竞品A"
|
||||
@@ -118,6 +118,22 @@ class TestBuildRunMessage:
|
||||
assert "标题A" in message
|
||||
assert "https://www.xiaohongshu.com/explore/abc123" in message
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_douyin_notes_link_to_douyin(self, db):
|
||||
"""链接形状按平台走 —— 群里点进去该是能看的作品,不是 404。"""
|
||||
task, run = await _seed(db)
|
||||
task.platform = "dy"
|
||||
_add_event(
|
||||
db, task, run, EVENT_NEW_NOTE, "新作品:标题A",
|
||||
payload={"note_id": "7525082444551310602", "title": "标题A"},
|
||||
)
|
||||
await db.flush()
|
||||
|
||||
message = await notify.build_run_message(db, task, run)
|
||||
|
||||
assert "https://www.douyin.com/video/7525082444551310602" in message
|
||||
assert "xiaohongshu.com" not in message
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_long_note_lists_are_truncated(self, db):
|
||||
"""A first run can find dozens; a wall of text is worse than a count."""
|
||||
|
||||
@@ -184,9 +184,9 @@ async def db():
|
||||
await engine.dispose()
|
||||
|
||||
|
||||
async def _seed_task(db: AsyncSession, name: str) -> MonitorTask:
|
||||
async def _seed_task(db: AsyncSession, name: str, platform: str = "xhs") -> MonitorTask:
|
||||
task = MonitorTask(
|
||||
name=name, platform="xhs", mode=MODE_CREATOR, enabled=True,
|
||||
name=name, platform=platform, mode=MODE_CREATOR, enabled=True,
|
||||
interval_minutes=60, max_notes_count=20, enable_comments=True,
|
||||
max_comments_count=50, run_timeout_seconds=3600,
|
||||
notify_enabled=False, created_at=0, updated_at=0,
|
||||
@@ -263,6 +263,25 @@ class TestBuildReport:
|
||||
assert result["totals"]["liked_count_delta"] == 30
|
||||
assert result["task_ids"] is None
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_empty_selection_is_not_the_same_as_no_filter(self, db):
|
||||
"""空列表 ≠ 不限制。
|
||||
|
||||
``_resolve_scope`` 在「这个平台一个任务都没有」时返回**空列表**。如果按真值
|
||||
处理(``if task_ids``),报表就会退化成「不限制平台」,把**所有**任务的数据
|
||||
聚合进来 —— 现象就是切到抖音,报表里却全是小红书的数据。
|
||||
"""
|
||||
other = await _seed_task(db, "xhs task", platform="xhs")
|
||||
await _seed_note_with_metrics(db, other, "n1", [(_ms(2026, 1, 10, 10), 999)])
|
||||
await db.commit()
|
||||
|
||||
result = await build_report(db, [], date(2026, 1, 10), date(2026, 1, 10))
|
||||
|
||||
assert result["totals"]["liked_count_delta"] == 0
|
||||
assert result["note_count"] == 0
|
||||
# 空列表要原样透出去;None 在 API 里的意思是「全部任务」,两者不能混。
|
||||
assert result["task_ids"] == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_baseline_from_before_the_range_is_used(self, db):
|
||||
"""Growth is measured against the last value before the window opens."""
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""monitor runner —— 尤其是抖音那条(不走子进程的)路的运行状态流转。
|
||||
|
||||
这条路的地位特殊:它不经过 ``crawler_manager``,所以爬虫那套「退出码 / 日志尾巴」的
|
||||
约定它一个都不沾。凡是写在那里面的东西,这条路都得单独有一份。
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from sqlalchemy import select
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from api.monitor import db as monitor_db
|
||||
from api.monitor import runner as runner_module
|
||||
from api.monitor.models import (
|
||||
MODE_CREATOR,
|
||||
RUN_RUNNING,
|
||||
MonitorRun,
|
||||
MonitorTarget,
|
||||
MonitorTask,
|
||||
)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def db(tmp_path):
|
||||
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
|
||||
await monitor_db.init_db()
|
||||
yield monitor_db
|
||||
await monitor_db.dispose_engine()
|
||||
|
||||
|
||||
async def _make_douyin_task() -> int:
|
||||
async with monitor_db.get_session() as session:
|
||||
now = get_current_timestamp()
|
||||
task = MonitorTask(
|
||||
name="dy", platform="dy", mode=MODE_CREATOR, enabled=True,
|
||||
interval_minutes=360, max_notes_count=20, enable_comments=False,
|
||||
max_comments_count=20, run_timeout_seconds=3600,
|
||||
notify_enabled=False, notify_failures=False,
|
||||
created_at=now, updated_at=now,
|
||||
)
|
||||
session.add(task)
|
||||
await session.flush()
|
||||
session.add(
|
||||
MonitorTarget(
|
||||
task_id=task.id, kind=MODE_CREATOR, external_id="MS4w-sec",
|
||||
xsec_token="", xsec_source="", raw_value="MS4w-sec",
|
||||
label="x", enabled=True, created_at=now,
|
||||
)
|
||||
)
|
||||
return task.id
|
||||
|
||||
|
||||
class TestDouyinRunStatus:
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_run_is_marked_running_before_collecting(self, db, monkeypatch):
|
||||
"""**采集开始之前**,run 就必须已经是 running。
|
||||
|
||||
这一行原先只写在爬虫那条分支里,于是抖音路上 run 一直停在 pending —— 一旦中途
|
||||
出事(异常、或进程被重启),界面上就是一个永远「排队中」的幽灵,而且 recover()
|
||||
当时也只收 running、够不着它。
|
||||
"""
|
||||
task_id = await _make_douyin_task()
|
||||
seen = {}
|
||||
|
||||
async def fake_collect(out_dir, **kwargs):
|
||||
async with monitor_db.get_session() as session:
|
||||
run = await session.scalar(select(MonitorRun).order_by(MonitorRun.id))
|
||||
seen["status"] = run.status
|
||||
return {
|
||||
"notes": 0,
|
||||
"comments": 0,
|
||||
"errors": ["故意失败"],
|
||||
"jsonl_dir": str(out_dir),
|
||||
}
|
||||
|
||||
monkeypatch.setattr(runner_module.douyin_fetch, "collect", fake_collect)
|
||||
|
||||
await runner_module.execute_task(task_id, trigger="manual")
|
||||
|
||||
assert seen["status"] == RUN_RUNNING
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_hanging_collect_does_not_leave_the_run_running(self, db, monkeypatch):
|
||||
"""进程内那条路也要有超时。
|
||||
|
||||
爬虫那条靠 ``run_and_wait(timeout=...)`` 兜底,这条路没有子进程、没人管 ——
|
||||
里面任何一次卡住(实测过 ``page.evaluate`` 打在一个卡死的标签页上不返回)都会让
|
||||
run 永远停在「运行中」,界面上看起来就是任务卡死了。
|
||||
"""
|
||||
task_id = await _make_douyin_task()
|
||||
async with monitor_db.get_session() as session:
|
||||
task = await session.get(MonitorTask, task_id)
|
||||
task.run_timeout_seconds = 1 # 把超时压到 1 秒,别让测试真等
|
||||
|
||||
async def hanging_collect(out_dir, **kwargs):
|
||||
await asyncio.sleep(60)
|
||||
raise AssertionError("不该走到这里")
|
||||
|
||||
monkeypatch.setattr(runner_module.douyin_fetch, "collect", hanging_collect)
|
||||
|
||||
await runner_module.execute_task(task_id, trigger="manual")
|
||||
|
||||
async with monitor_db.get_session() as session:
|
||||
run = await session.scalar(select(MonitorRun).order_by(MonitorRun.id))
|
||||
assert run.status != RUN_RUNNING
|
||||
assert "超时" in (run.error_message or "") or "超过" in (run.error_message or "")
|
||||
@@ -30,11 +30,12 @@ from api.monitor.models import (
|
||||
MonitorTarget,
|
||||
MonitorTask,
|
||||
RUN_INTERRUPTED,
|
||||
RUN_PENDING,
|
||||
RUN_RUNNING,
|
||||
RUN_SUCCESS,
|
||||
)
|
||||
from api.monitor.scheduler import MonitorScheduler
|
||||
from api.monitor.settings import set_cookie
|
||||
from api.monitor.settings import set_cookie, set_setting
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
MS_PER_MINUTE = 60_000
|
||||
@@ -72,12 +73,14 @@ async def executed(monkeypatch):
|
||||
return calls
|
||||
|
||||
|
||||
async def _make_task(next_run_at, enabled: bool = True, interval: int = 60) -> int:
|
||||
async def _make_task(
|
||||
next_run_at, enabled: bool = True, interval: int = 60, platform: str = "xhs"
|
||||
) -> int:
|
||||
async with monitor_db.get_session() as session:
|
||||
now = get_current_timestamp()
|
||||
task = MonitorTask(
|
||||
name="t",
|
||||
platform="xhs",
|
||||
platform=platform,
|
||||
mode=MODE_CREATOR,
|
||||
enabled=enabled,
|
||||
interval_minutes=interval,
|
||||
@@ -145,6 +148,44 @@ class TestFiring:
|
||||
|
||||
assert executed == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_cookie_gate_reads_the_tasks_own_platform(
|
||||
self, monkeypatch, db, executed
|
||||
):
|
||||
"""cookie 闸门要按任务自己的平台取。
|
||||
|
||||
以前这里是 ``get_cookie(session)``(默认小红书)—— 只有小红书时看不出问题,
|
||||
接上抖音后,抖音任务会因为读的是小红书那份 cookie 而永远不被触发,且不报错。
|
||||
"""
|
||||
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_cookie(session, "sessionid=dy-secret", "dy")
|
||||
|
||||
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
|
||||
|
||||
await MonitorScheduler().tick()
|
||||
|
||||
assert executed == [(task_id, "scheduled")]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_cdp_mode_frees_a_task_from_the_cookie_gate(
|
||||
self, monkeypatch, db, executed
|
||||
):
|
||||
"""开着 CDP 时不该再要求先粘 cookie。
|
||||
|
||||
CDP 模式下登录态来自被接管的那台浏览器,粘不粘 cookie 都由不得它 —— 不放行的话,
|
||||
选了「接管已有 Chrome」却没粘 cookie 的用户会发现任务永远不跑,而且什么错都不报。
|
||||
"""
|
||||
monkeypatch.setattr(scheduler_module, "crawler_manager", FakeCrawlerManager(busy=False))
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, "system.cdp_enabled", "true")
|
||||
|
||||
task_id = await _make_task(get_current_timestamp() - MS_PER_MINUTE, platform="dy")
|
||||
|
||||
await MonitorScheduler().tick()
|
||||
|
||||
assert executed == [(task_id, "scheduled")]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_long_outage_coalesces_into_one_run(self, monkeypatch, db, executed):
|
||||
"""A missed schedule fires once, not once per missed interval."""
|
||||
@@ -239,6 +280,39 @@ class TestRecovery:
|
||||
assert run.status == RUN_INTERRUPTED
|
||||
assert run.finished_at is not None
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_pending_runs_are_also_cleaned_up(self, db):
|
||||
"""挂在 ``pending`` 的 run 同样是残留,必须一起收。
|
||||
|
||||
那一行是上一轮建的,可它后面的采集根本没机会开始(进程被重启,或采集那条路抛了
|
||||
异常)。只清 ``running`` 的话,它会永远挂在界面上显示「排队中」——
|
||||
用户看到的就是任务卡死了。
|
||||
"""
|
||||
async with monitor_db.get_session() as session:
|
||||
now = get_current_timestamp()
|
||||
task = MonitorTask(
|
||||
name="t", platform="dy", mode=MODE_CREATOR, enabled=True,
|
||||
interval_minutes=60, max_notes_count=20, enable_comments=False,
|
||||
max_comments_count=50, run_timeout_seconds=3600,
|
||||
next_run_at=now, last_status="pending", created_at=now, updated_at=now,
|
||||
)
|
||||
session.add(task)
|
||||
await session.flush()
|
||||
session.add(
|
||||
MonitorRun(
|
||||
task_id=task.id, trigger="manual", status=RUN_PENDING,
|
||||
phase=MODE_CREATOR, save_data_path="", queued_at=now, not_before=0,
|
||||
max_comments_count=50,
|
||||
)
|
||||
)
|
||||
|
||||
await MonitorScheduler().recover()
|
||||
|
||||
async with monitor_db.get_session() as session:
|
||||
run = await session.scalar(select(MonitorRun))
|
||||
assert run.status == RUN_INTERRUPTED
|
||||
assert run.finished_at is not None
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_completed_runs_are_left_alone(self, db):
|
||||
async with monitor_db.get_session() as session:
|
||||
|
||||
@@ -14,6 +14,8 @@ import pathlib
|
||||
|
||||
import pytest
|
||||
|
||||
import config
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parent.parent
|
||||
|
||||
# 统一的禁用字段名(键)。昵称字段(nickname/user_nickname/screen_name/name/user_name)允许保留(值需脱敏)。
|
||||
@@ -28,6 +30,17 @@ NICK_KEYS = {"nickname", "user_nickname", "screen_name", "name", "user_name"}
|
||||
MASK_RE = re.compile(r"^.?\*{1,4}.?$")
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _force_nickname_masking(monkeypatch):
|
||||
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
|
||||
|
||||
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
|
||||
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
|
||||
所以这里显式打开来测;开关两个方向的行为由 test_mask_and_hash_tools 覆盖。
|
||||
"""
|
||||
monkeypatch.setattr(config, "MASK_NICKNAME", True)
|
||||
|
||||
|
||||
# ----------------------------- ORM 自省 -----------------------------
|
||||
|
||||
def test_orm_has_no_forbidden_columns():
|
||||
@@ -84,17 +97,26 @@ def _check_nickname_masked(d: dict, raw: str, label: str):
|
||||
assert MASK_RE.match(val) or "*" in val, f"[{label}] {k} 未脱敏: {val}"
|
||||
|
||||
|
||||
def test_mask_and_hash_tools():
|
||||
def test_mask_and_hash_tools(monkeypatch):
|
||||
from tools.user_hash import anonymize_user_id, mask_nickname
|
||||
h = anonymize_user_id("12345")
|
||||
assert h and h != "12345" and re.fullmatch(r"[0-9a-f]{16}", h)
|
||||
assert anonymize_user_id(None) == "" and anonymize_user_id("") == ""
|
||||
# 昵称脱敏:首尾留1字、中间星号,且不等于原文
|
||||
|
||||
# 开关打开:首尾留 1 字、中间星号,且不等于原文。
|
||||
monkeypatch.setattr(config, "MASK_NICKNAME", True)
|
||||
assert mask_nickname("张三丰") != "张三丰"
|
||||
assert "*" in mask_nickname("张三丰")
|
||||
assert mask_nickname(None) == ""
|
||||
assert mask_nickname("a") == "*"
|
||||
|
||||
# 开关关闭(本仓库的部署配置):原样返回。脱敏是有损的 —— 「张三」和「张四」
|
||||
# 都会变成「张*」,而分清谁是谁正是监控这一层要干的事。
|
||||
monkeypatch.setattr(config, "MASK_NICKNAME", False)
|
||||
assert mask_nickname("张三丰") == "张三丰"
|
||||
assert mask_nickname("a") == "a"
|
||||
assert mask_nickname(None) == ""
|
||||
|
||||
|
||||
def test_xhs_note_extraction_masks_user_info():
|
||||
import asyncio
|
||||
|
||||
+99
-10
@@ -21,12 +21,15 @@
|
||||
import httpx
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from sqlalchemy import text
|
||||
from sqlalchemy import select, text
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from api.main import app
|
||||
from api.monitor import adapters
|
||||
from api.monitor import db as monitor_db
|
||||
from api.monitor import platforms
|
||||
from api.monitor.models import MonitorTask
|
||||
from api.monitor.models import MonitorNote, MonitorNoteMetric, MonitorTask
|
||||
|
||||
XHS_TARGET = "5f58bd990000000001003753"
|
||||
|
||||
@@ -43,6 +46,33 @@ async def client(tmp_path):
|
||||
await monitor_db.dispose_engine()
|
||||
|
||||
|
||||
async def _seed_note_with_one_snapshot(task_name: str) -> None:
|
||||
"""给某个任务塞一条作品和一次指标快照。
|
||||
|
||||
过滤类测试**必须有真数据**才有意义 —— 库里空着的话,过滤有没有生效结果都是 0,
|
||||
测试就变成了空跑(这个坑踩过一次:一个报表串数据的 bug 因此没被拦住)。
|
||||
"""
|
||||
async with monitor_db.get_session() as session:
|
||||
task = await session.scalar(select(MonitorTask).where(MonitorTask.name == task_name))
|
||||
now = get_current_timestamp()
|
||||
session.add(
|
||||
MonitorNote(
|
||||
task_id=task.id, note_id="seed-note", title="seed", note_url="",
|
||||
cover="", creator_hash="", source_kind="", published_at=None,
|
||||
first_seen_run_id=1, first_seen_at=now,
|
||||
last_seen_run_id=1, last_seen_at=now,
|
||||
)
|
||||
)
|
||||
session.add(
|
||||
MonitorNoteMetric(
|
||||
task_id=task.id, note_id="seed-note", run_id=1, captured_at=now,
|
||||
liked_count=42, comment_count=0, collected_count=0, share_count=0,
|
||||
raw_liked_count="42", raw_comment_count="0",
|
||||
raw_collected_count="0", raw_share_count="0",
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
class TestCapabilityMatrix:
|
||||
@pytest.mark.asyncio
|
||||
async def test_matrix_is_exposed_to_the_ui(self, client):
|
||||
@@ -54,7 +84,28 @@ class TestCapabilityMatrix:
|
||||
# what stops the UI offering a platform that can never produce data.
|
||||
assert all("monitor_wired" in p for p in body["platforms"])
|
||||
assert by_value["xhs"]["monitor_wired"] is True
|
||||
assert by_value["dy"]["monitor_wired"] is False
|
||||
assert by_value["dy"]["monitor_wired"] is True
|
||||
|
||||
def test_every_wired_platform_has_an_adapter(self):
|
||||
"""能力矩阵说「接通了」,就必须真的有一套适配管子。
|
||||
|
||||
两个注册表(platforms.PLATFORM_CAPABILITIES 与 adapters.ADAPTERS)分开是有意的
|
||||
—— 前者是给前端看的能力描述,后者是爬虫的管道细节。代价是它们可能漂移,
|
||||
所以在这里钉一条:凡声明接通的,必须能找到适配器。
|
||||
"""
|
||||
for platform in platforms.all_platforms():
|
||||
if platforms.is_monitor_wired(platform):
|
||||
assert adapters.has_adapter(platform), f"{platform} 声明接通但没有适配器"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_target_hints_are_exposed_for_wired_platforms(self, client):
|
||||
"""前端的目标输入框拿它做 placeholder —— 让用户看到本平台该粘什么样的链接。"""
|
||||
body = (await client.get("/api/config/platforms")).json()
|
||||
by_value = {p["value"]: p for p in body["platforms"]}
|
||||
|
||||
assert "douyin.com/user/" in by_value["dy"]["target_hints"]["creator"]
|
||||
assert "douyin.com/video/" in by_value["dy"]["target_hints"]["note"]
|
||||
assert "xiaohongshu.com" in by_value["xhs"]["target_hints"]["creator"]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_metrics_are_per_platform_and_labelled(self, client):
|
||||
@@ -80,14 +131,17 @@ class TestCapabilityMatrix:
|
||||
class TestTaskCreationGuard:
|
||||
@pytest.mark.asyncio
|
||||
async def test_unwired_platform_is_rejected_with_an_explanation(self, client):
|
||||
"""Accepting it would create a task that silently never produces data."""
|
||||
"""Accepting it would create a task that silently never produces data.
|
||||
|
||||
用 B站 而不是抖音:抖音现在接通了,不再是「已知但未接通」的例子。
|
||||
"""
|
||||
response = await client.post(
|
||||
"/api/monitor/tasks",
|
||||
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": ["x"]},
|
||||
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
|
||||
)
|
||||
assert response.status_code == 400
|
||||
detail = response.json()["detail"]
|
||||
assert "抖音" in detail
|
||||
assert "B站" in detail
|
||||
assert "尚未接通" in detail
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@@ -102,7 +156,7 @@ class TestTaskCreationGuard:
|
||||
async def test_no_task_row_is_created_when_rejected(self, client):
|
||||
await client.post(
|
||||
"/api/monitor/tasks",
|
||||
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": ["x"]},
|
||||
json={"name": "B站任务", "mode": "creator", "platform": "bili", "targets": ["x"]},
|
||||
)
|
||||
assert (await client.get("/api/monitor/tasks")).json()["tasks"] == []
|
||||
|
||||
@@ -123,11 +177,34 @@ class TestTaskCreationGuard:
|
||||
tasks = (await client.get("/api/monitor/tasks")).json()["tasks"]
|
||||
assert {t["platform"] for t in tasks} == {"xhs"}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_explicit_platform_is_honoured_on_create(self, client):
|
||||
"""建任务时给的平台必须落到那个平台。
|
||||
|
||||
缺省值是小红的(接口早期的兼容行为),所以「在抖音页面建任务」如果没有显式
|
||||
带上 platform,就会安安静静地变成一个小红书任务 —— 不报错,只是出现在另一
|
||||
个列表里。前端那半边已经改成必传;这里守住后端这一半:给了就必须用。
|
||||
"""
|
||||
sec_uid = "MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE"
|
||||
created = await client.post(
|
||||
"/api/monitor/tasks",
|
||||
json={"name": "抖音任务", "mode": "creator", "platform": "dy", "targets": [sec_uid]},
|
||||
)
|
||||
assert created.status_code == 201
|
||||
|
||||
assert (await client.get("/api/monitor/tasks", params={"platform": "xhs"})).json()[
|
||||
"tasks"
|
||||
] == []
|
||||
dy_tasks = (
|
||||
await client.get("/api/monitor/tasks", params={"platform": "dy"})
|
||||
).json()["tasks"]
|
||||
assert [t["name"] for t in dy_tasks] == ["抖音任务"]
|
||||
|
||||
|
||||
class TestPlatformScoping:
|
||||
async def _seed_two_platforms(self, client):
|
||||
"""One real XHS task plus a Douyin task inserted directly, since the API
|
||||
refuses to create the latter."""
|
||||
"""One XHS task created through the API, plus a Douyin task inserted
|
||||
directly so its fields can be pinned exactly."""
|
||||
await client.post(
|
||||
"/api/monitor/tasks",
|
||||
json={"name": "小红书任务", "mode": "creator", "targets": [XHS_TARGET]},
|
||||
@@ -166,8 +243,15 @@ class TestPlatformScoping:
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_platform_with_no_tasks_yields_empty_not_everything(self, client):
|
||||
"""An empty task set must not degrade into "no filter"."""
|
||||
"""空的任务集合不能退化成「不加过滤」。
|
||||
|
||||
**这里必须真的有数据。** 没有数据时,过滤生效与否结果都是 0 —— 这条测试原先
|
||||
就栽在这个空跑上,所以没能拦下一个报表串数据的 bug(切到抖音,报表里却出现
|
||||
小红书的数据)。最后那段「小红书自己的报表看得到」就是为了证明这些数据确实
|
||||
存在、上面那两个 0 是过滤出来的。
|
||||
"""
|
||||
await self._seed_two_platforms(client)
|
||||
await _seed_note_with_one_snapshot("小红书任务")
|
||||
|
||||
body = (await client.get("/api/monitor/notes", params={"platform": "bili"})).json()
|
||||
assert body["notes"] == []
|
||||
@@ -178,6 +262,11 @@ class TestPlatformScoping:
|
||||
assert report["totals"]["liked_count_delta"] == 0
|
||||
assert report["note_count"] == 0
|
||||
|
||||
xhs = (
|
||||
await client.get("/api/monitor/report", params={"platform": "xhs"})
|
||||
).json()
|
||||
assert xhs["note_count"] == 1
|
||||
|
||||
|
||||
class TestPerPlatformSettings:
|
||||
@pytest.mark.asyncio
|
||||
|
||||
@@ -0,0 +1,268 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""监控侧扫码登录与登录态检测。
|
||||
|
||||
两个要点在这里被钉死:
|
||||
|
||||
* 二维码必须从浏览器**默认 context** 里读 —— 新建 context 是无痕式的 profile,
|
||||
扫了也白扫,爬虫读不到那份 cookie;
|
||||
* 「登录了吗」不能用页面里的 `window.__INITIAL_STATE__`。那是**页面加载那一刻的快照**:
|
||||
浏览器本来就登录着时它是对的,但扫码是加载**之后**才登录的,快照不会翻转,
|
||||
于是扫完码界面会一直停在二维码上。判据改成拿 cookie 问后台接口。
|
||||
"""
|
||||
|
||||
from unittest.mock import AsyncMock, MagicMock
|
||||
|
||||
import pytest
|
||||
|
||||
from api.creator.client import CreatorApiError
|
||||
from api.monitor import qrlogin
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _reset_module_state():
|
||||
for attribute in ("_current", "_page", "_playwright", "_state_cache"):
|
||||
setattr(qrlogin, attribute, None)
|
||||
yield
|
||||
for attribute in ("_current", "_page", "_playwright", "_state_cache"):
|
||||
setattr(qrlogin, attribute, None)
|
||||
|
||||
|
||||
XHS_COOKIES = [
|
||||
{"name": "a1", "value": "an-a1-value"},
|
||||
{"name": "web_session", "value": "a-session"},
|
||||
]
|
||||
|
||||
|
||||
def _fake_stack(cookies=None, qr="data:image/png;base64,AAAA"):
|
||||
"""Chrome/Playwright 替身,行为与真实的一致。"""
|
||||
page = MagicMock()
|
||||
page.url = "https://www.xiaohongshu.com/explore"
|
||||
page.is_closed = MagicMock(return_value=False)
|
||||
page.goto = AsyncMock()
|
||||
page.close = AsyncMock()
|
||||
|
||||
context = MagicMock()
|
||||
context.pages = []
|
||||
context.cookies = AsyncMock(return_value=list(cookies if cookies is not None else XHS_COOKIES))
|
||||
context.new_page = AsyncMock(return_value=page)
|
||||
|
||||
browser = MagicMock()
|
||||
browser.contexts = [context]
|
||||
# 去新建 context 正是这里要防的 bug,所以让它直接炸,而不是悄悄返回一个无痕 profile。
|
||||
browser.new_context = AsyncMock(
|
||||
side_effect=AssertionError("must reuse browser.contexts[0], not a new context")
|
||||
)
|
||||
|
||||
playwright = MagicMock()
|
||||
playwright.chromium.connect_over_cdp = AsyncMock(return_value=browser)
|
||||
playwright.stop = AsyncMock()
|
||||
|
||||
manager = MagicMock()
|
||||
manager.start = AsyncMock(return_value=playwright)
|
||||
|
||||
return manager, playwright, browser, context, page, qr
|
||||
|
||||
|
||||
def _patch(monkeypatch, manager, qr="data:image/png;base64,AAAA", resolver=None):
|
||||
"""``resolver(cookie)`` 返回账号信息 dict,或抛 CreatorApiError。"""
|
||||
if resolver is None:
|
||||
resolver = lambda _cookie: {"user_id": "u1", "nickname": "小明"} # noqa: E731
|
||||
|
||||
class _Client:
|
||||
def __init__(self, cookie, **kwargs):
|
||||
self.cookie = cookie
|
||||
|
||||
async def fetch_user_info(self):
|
||||
return resolver(self.cookie)
|
||||
|
||||
monkeypatch.setattr(qrlogin, "async_playwright", lambda: manager)
|
||||
monkeypatch.setattr(qrlogin, "CreatorClient", _Client)
|
||||
monkeypatch.setattr(qrlogin.utils, "find_login_qrcode", AsyncMock(return_value=qr))
|
||||
|
||||
|
||||
def _signed_out(_cookie):
|
||||
raise CreatorApiError("登录态无效或已过期", status=401)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_idle_reports_the_browsers_login_state(monkeypatch):
|
||||
manager, *_ = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
|
||||
|
||||
snapshot = await qrlogin.status()
|
||||
|
||||
assert snapshot["status"] == qrlogin.STATUS_IDLE
|
||||
assert snapshot["logged_in"] is True
|
||||
assert snapshot["nickname"] == "小明"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_unwired_platform_is_rejected():
|
||||
"""只有小红书接了扫码;别的平台必须直接报错,而不是给个按不动的按钮。"""
|
||||
with pytest.raises(ValueError):
|
||||
await qrlogin.start("dy")
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_start_reads_the_qr_from_the_default_context(monkeypatch):
|
||||
manager, _pw, browser, context, _page, _qr = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=_signed_out)
|
||||
|
||||
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
|
||||
assert snapshot["status"] == qrlogin.STATUS_WAITING
|
||||
assert snapshot["image"] == "data:image/png;base64,AAAA"
|
||||
context.new_page.assert_awaited_once()
|
||||
browser.new_context.assert_not_called()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_start_does_not_open_a_second_tab(monkeypatch):
|
||||
"""已有的 xhs 标签页会被认领,所以重启不会在浏览器里堆孤儿页。"""
|
||||
manager, _pw, _browser, context, page, _qr = _fake_stack()
|
||||
context.pages = [page]
|
||||
_patch(monkeypatch, manager, resolver=_signed_out)
|
||||
|
||||
await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
|
||||
context.new_page.assert_not_called()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_already_signed_in_profile_needs_no_scan(monkeypatch):
|
||||
"""没二维码但 profile 已登录 —— 这是成功,不是失败。"""
|
||||
manager, *_rest = _fake_stack()
|
||||
_patch(monkeypatch, manager, qr="", resolver=lambda _c: {"user_id": "u9", "nickname": "老王"})
|
||||
|
||||
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
|
||||
assert snapshot["status"] == qrlogin.STATUS_SUCCESS
|
||||
assert snapshot["nickname"] == "老王"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_already_signed_in_profile_never_opens_a_page(monkeypatch):
|
||||
"""已登录时**根本不该去开页面**。
|
||||
|
||||
读二维码内部会 wait_for_selector 等满 30 秒才放弃,而已经登录时页面上没有二维码 ——
|
||||
顺序反了的话,用户点一下按钮要干等半分钟,还白开一个标签页。
|
||||
"""
|
||||
manager, _pw, _browser, context, _page, _qr = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
|
||||
|
||||
await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
|
||||
context.new_page.assert_not_called()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_no_qr_and_not_signed_in_is_an_error(monkeypatch):
|
||||
manager, *_rest = _fake_stack()
|
||||
_patch(monkeypatch, manager, qr="", resolver=_signed_out)
|
||||
|
||||
snapshot = await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
|
||||
assert snapshot["status"] == qrlogin.STATUS_ERROR
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_completed_scan_flips_the_session_to_success(monkeypatch):
|
||||
"""**这个判据是重点**:扫码是页面加载之后才发生的,所以不能用页面快照来判断。"""
|
||||
manager, *_rest = _fake_stack()
|
||||
signed_in = {"value": False}
|
||||
|
||||
def resolver(cookie):
|
||||
assert "a1=an-a1-value" in cookie # 判据必须真的用 cookie 去问
|
||||
if not signed_in["value"]:
|
||||
raise CreatorApiError("登录态无效或已过期", status=401)
|
||||
return {"user_id": "u1", "nickname": "小红"}
|
||||
|
||||
_patch(monkeypatch, manager, resolver=resolver)
|
||||
await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
assert qrlogin._current.status == qrlogin.STATUS_WAITING
|
||||
|
||||
# 操作者扫了码
|
||||
signed_in["value"] = True
|
||||
qrlogin._state_cache = None # 5 秒缓存否则会遮住这次变化
|
||||
|
||||
snapshot = await qrlogin.status()
|
||||
|
||||
assert snapshot["status"] == qrlogin.STATUS_SUCCESS
|
||||
assert snapshot["nickname"] == "小红"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_successful_session_hands_over_a_cookie(monkeypatch):
|
||||
"""扫码不该只写浏览器 profile —— 还要能把 cookie 交出来存库,
|
||||
否则关掉 CDP 就断了。"""
|
||||
manager, *_rest = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "小明"})
|
||||
|
||||
await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
cookie = await qrlogin.take_cookie()
|
||||
|
||||
assert cookie is not None
|
||||
assert "a1=an-a1-value" in cookie
|
||||
# 只能取一次,否则每次轮询都会重复写库
|
||||
assert await qrlogin.take_cookie() is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_session_expires(monkeypatch):
|
||||
manager, *_rest = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=_signed_out)
|
||||
await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
|
||||
qrlogin._current.started_at -= qrlogin.QR_TTL_SECONDS + 1
|
||||
snapshot = await qrlogin.status()
|
||||
|
||||
assert snapshot["status"] == qrlogin.STATUS_EXPIRED
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_check_login_state_reports_when_the_browser_cannot_answer(monkeypatch):
|
||||
"""连不上浏览器时要说出来,不能悄悄报成「未登录」。"""
|
||||
manager, _pw, _browser, context, _page, _qr = _fake_stack()
|
||||
context.cookies = AsyncMock(side_effect=RuntimeError("Target closed"))
|
||||
_patch(monkeypatch, manager)
|
||||
|
||||
state = await qrlogin.check_login_state()
|
||||
|
||||
assert state["known"] is False
|
||||
assert state["logged_in"] is False
|
||||
assert "Target closed" in state["error"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_check_login_state_is_cached(monkeypatch):
|
||||
manager, _pw, _browser, context, _page, _qr = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "x"})
|
||||
|
||||
await qrlogin.check_login_state()
|
||||
await qrlogin.check_login_state()
|
||||
|
||||
context.cookies.assert_awaited_once()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_force_bypasses_the_cache(monkeypatch):
|
||||
manager, _pw, _browser, context, _page, _qr = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=lambda _c: {"user_id": "u1", "nickname": "x"})
|
||||
|
||||
await qrlogin.check_login_state()
|
||||
await qrlogin.check_login_state(force=True)
|
||||
|
||||
assert context.cookies.await_count == 2
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_cancel_keeps_the_operators_tab(monkeypatch):
|
||||
"""与运营模块不同:那里的上下文是临时的、用完即弃;这里的标签页属于操作者的浏览器。"""
|
||||
manager, _pw, _browser, _context, page, _qr = _fake_stack()
|
||||
_patch(monkeypatch, manager, resolver=_signed_out)
|
||||
await qrlogin.start(qrlogin.PLATFORM_XHS)
|
||||
|
||||
snapshot = await qrlogin.cancel()
|
||||
|
||||
assert snapshot["status"] == qrlogin.STATUS_IDLE
|
||||
page.close.assert_not_called()
|
||||
@@ -0,0 +1,194 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""Tests for monitor task schedule arithmetic.
|
||||
|
||||
Everything here is timezone-local, matching the implementation: the container is
|
||||
pinned to the operator's zone via TZ, so the tests build their expectations from
|
||||
naive local datetimes too and stay correct wherever they run.
|
||||
"""
|
||||
|
||||
from datetime import datetime, time, timedelta
|
||||
|
||||
import pytest
|
||||
|
||||
from api.monitor import schedule
|
||||
|
||||
|
||||
def _ms(moment: datetime) -> int:
|
||||
return int(moment.timestamp() * 1000)
|
||||
|
||||
|
||||
def test_interval_is_now_plus_the_interval():
|
||||
after = _ms(datetime(2026, 10, 7, 9, 0))
|
||||
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_INTERVAL,
|
||||
interval_minutes=120,
|
||||
hours=[],
|
||||
days=[],
|
||||
minute=0,
|
||||
after_ms=after,
|
||||
)
|
||||
|
||||
assert nxt == after + 120 * 60_000
|
||||
|
||||
|
||||
def test_daily_takes_the_soonest_remaining_time_today():
|
||||
after = _ms(datetime(2026, 10, 7, 8, 0))
|
||||
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_DAILY,
|
||||
interval_minutes=60,
|
||||
hours=[18, 9], # deliberately unsorted
|
||||
days=[],
|
||||
minute=30,
|
||||
after_ms=after,
|
||||
)
|
||||
|
||||
assert nxt == _ms(datetime(2026, 10, 7, 9, 30))
|
||||
|
||||
|
||||
def test_daily_rolls_over_to_tomorrow_once_every_time_has_passed():
|
||||
after = _ms(datetime(2026, 10, 7, 20, 0))
|
||||
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_DAILY,
|
||||
interval_minutes=60,
|
||||
hours=[9, 18],
|
||||
days=[],
|
||||
minute=30,
|
||||
after_ms=after,
|
||||
)
|
||||
|
||||
assert nxt == _ms(datetime(2026, 10, 8, 9, 30))
|
||||
|
||||
|
||||
def test_a_slot_exactly_now_belongs_to_the_next_day():
|
||||
"""Strictly-after, so the run that just fired does not fire again.
|
||||
|
||||
The scheduler advances with ``after_ms`` set to the moment the run started,
|
||||
which is at or just past the slot -- if the comparison were inclusive it would
|
||||
pick the same slot back up and loop.
|
||||
"""
|
||||
after = _ms(datetime(2026, 10, 7, 9, 30))
|
||||
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_DAILY,
|
||||
interval_minutes=60,
|
||||
hours=[9],
|
||||
days=[],
|
||||
minute=30,
|
||||
after_ms=after,
|
||||
)
|
||||
|
||||
assert nxt == _ms(datetime(2026, 10, 8, 9, 30))
|
||||
|
||||
|
||||
def test_weekly_jumps_to_the_next_selected_weekday():
|
||||
after_dt = datetime(2026, 10, 7, 8, 0)
|
||||
target = (after_dt.weekday() + 2) % 7
|
||||
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_WEEKLY,
|
||||
interval_minutes=60,
|
||||
hours=[10],
|
||||
days=[target],
|
||||
minute=0,
|
||||
after_ms=_ms(after_dt),
|
||||
)
|
||||
|
||||
expected = datetime.combine((after_dt + timedelta(days=2)).date(), time(10, 0))
|
||||
assert nxt == _ms(expected)
|
||||
|
||||
|
||||
def test_weekly_can_fire_later_the_same_day():
|
||||
after_dt = datetime(2026, 10, 7, 8, 0)
|
||||
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_WEEKLY,
|
||||
interval_minutes=60,
|
||||
hours=[21],
|
||||
days=[after_dt.weekday()],
|
||||
minute=15,
|
||||
after_ms=_ms(after_dt),
|
||||
)
|
||||
|
||||
assert nxt == _ms(datetime.combine(after_dt.date(), time(21, 15)))
|
||||
|
||||
|
||||
def test_weekly_without_weekdays_means_every_day():
|
||||
"""Otherwise an empty day selection would match nothing and never fire."""
|
||||
after = _ms(datetime(2026, 10, 7, 8, 0))
|
||||
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_WEEKLY,
|
||||
interval_minutes=60,
|
||||
hours=[9],
|
||||
days=[],
|
||||
minute=0,
|
||||
after_ms=after,
|
||||
)
|
||||
|
||||
assert nxt == _ms(datetime(2026, 10, 7, 9, 0))
|
||||
|
||||
|
||||
def test_a_clock_schedule_with_no_times_can_never_fire():
|
||||
"""Returned as None so the caller can park the task instead of leaving it due."""
|
||||
nxt = schedule.next_occurrence(
|
||||
mode=schedule.MODE_DAILY,
|
||||
interval_minutes=60,
|
||||
hours=[],
|
||||
days=[],
|
||||
minute=0,
|
||||
after_ms=_ms(datetime(2026, 10, 7, 8, 0)),
|
||||
)
|
||||
|
||||
assert nxt is None
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"raw, expected",
|
||||
[
|
||||
("9,18", [9, 18]),
|
||||
("18,9", [9, 18]), # stored order is not guaranteed
|
||||
("9,9,9", [9]),
|
||||
("", []),
|
||||
(None, []),
|
||||
("9, 18 ", [9, 18]),
|
||||
("9,99,-1,abc,", [9]), # junk is dropped, never raised
|
||||
],
|
||||
)
|
||||
def test_parse_hours_is_forgiving(raw, expected):
|
||||
assert schedule.parse_hours(raw) == expected
|
||||
|
||||
|
||||
def test_parse_days_accepts_the_whole_week():
|
||||
assert schedule.parse_days("0,1,2,3,4,5,6") == [0, 1, 2, 3, 4, 5, 6]
|
||||
assert schedule.parse_days("7,-1") == []
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"mode, interval, hours, days, minute, expected",
|
||||
[
|
||||
(schedule.MODE_INTERVAL, 360, [], [], 0, "每 6 小时"),
|
||||
(schedule.MODE_INTERVAL, 1440, [], [], 0, "每 1 天"),
|
||||
(schedule.MODE_INTERVAL, 45, [], [], 0, "每 45 分钟"),
|
||||
(schedule.MODE_DAILY, 60, [9, 18], [], 30, "每天 09:30、18:30"),
|
||||
(schedule.MODE_WEEKLY, 60, [10], [0, 1, 2, 3, 4], 0, "周一、周二、周三、周四、周五 10:00"),
|
||||
(schedule.MODE_WEEKLY, 60, [10], [], 0, "每天 10:00"),
|
||||
(schedule.MODE_DAILY, 60, [], [], 0, "未设置时间"),
|
||||
],
|
||||
)
|
||||
def test_describe(mode, interval, hours, days, minute, expected):
|
||||
assert (
|
||||
schedule.describe(
|
||||
mode=mode, interval_minutes=interval, hours=hours, days=days, minute=minute
|
||||
)
|
||||
== expected
|
||||
)
|
||||
|
||||
|
||||
def test_format_round_trips_through_parse():
|
||||
hours = [9, 12, 18]
|
||||
days = [0, 4]
|
||||
assert schedule.parse_hours(schedule.format_hours(hours)) == hours
|
||||
assert schedule.parse_days(schedule.format_days(days)) == days
|
||||
@@ -20,7 +20,7 @@ def test_extract_search_note_list_from_keyword_page():
|
||||
assert notes[0].note_id == "9117888152"
|
||||
assert notes[0].title.startswith("武汉交互空间科技")
|
||||
assert notes[0].tieba_name == "武汉交互空间"
|
||||
assert notes[0].user_nickname == "V***人"
|
||||
assert notes[0].user_nickname == "VR虚拟达人"
|
||||
|
||||
|
||||
def test_extract_search_note_list_from_current_pc_card_page():
|
||||
@@ -56,7 +56,7 @@ def test_extract_search_note_list_from_current_pc_card_page():
|
||||
assert notes[0].desc == "培训班需求,数学,英语,编程老师,专职兼职都可"
|
||||
assert notes[0].tieba_name == "诸城吧"
|
||||
assert notes[0].tieba_link.endswith("kw=%E8%AF%B8%E5%9F%8E")
|
||||
assert notes[0].user_nickname == "7***7"
|
||||
assert notes[0].user_nickname == "754023117"
|
||||
assert notes[0].publish_time == "2026-3-15"
|
||||
assert notes[0].total_replay_num == 19
|
||||
|
||||
@@ -147,7 +147,7 @@ def test_extract_note_detail_and_comments_from_current_pc_api():
|
||||
assert note.note_id == "10451142633"
|
||||
assert note.title == "这X尔斯对比巴尔斯,我只能说ID正确,允许居功自傲"
|
||||
assert note.desc == "皮队败决处刑德国编程钢琴师兼职数学家"
|
||||
assert note.user_nickname == "泰***克"
|
||||
assert note.user_nickname == "泰高祖蒙斯克"
|
||||
assert note.tieba_name == "dota2吧"
|
||||
assert note.total_replay_num == 15
|
||||
assert note.total_replay_page == 1
|
||||
@@ -155,7 +155,7 @@ def test_extract_note_detail_and_comments_from_current_pc_api():
|
||||
assert len(comments) == 1
|
||||
assert comments[0].comment_id == "153154097267"
|
||||
assert comments[0].content == "xg现在大树阵容另一个辅助不选控制"
|
||||
assert comments[0].user_nickname == "期***3"
|
||||
assert comments[0].user_nickname == "期胡希3"
|
||||
assert comments[0].sub_comment_count == 4
|
||||
# 教学版已移除 ip_location 等可定位真人字段
|
||||
|
||||
@@ -191,7 +191,7 @@ def test_extract_creator_info_and_threads_from_current_pc_api():
|
||||
creator = extractor.extract_creator_info_from_api(creator_api)
|
||||
thread_ids = extractor.extract_creator_thread_id_list_from_api(feed_api)
|
||||
|
||||
assert creator.user_nickname == "米***子"
|
||||
assert creator.user_nickname == "米米世界大手子"
|
||||
assert creator.fans == 58
|
||||
assert creator.follows == 1
|
||||
# 教学版已移除 user_id、user_name、ip_location 等可定位真人字段
|
||||
@@ -223,7 +223,7 @@ def test_extract_tieba_note_list_from_bigpipe_thread_page():
|
||||
assert len(notes) == 48
|
||||
assert notes[0].note_id == "9079949995"
|
||||
assert notes[0].title == "盗墓笔记全集+txt小说,已整理"
|
||||
assert notes[0].user_nickname == "公***仲"
|
||||
assert notes[0].user_nickname == "公子伯仲"
|
||||
assert notes[0].tieba_name == "盗墓笔记吧"
|
||||
assert notes[0].tieba_link.endswith("kw=%E7%9B%97%E5%A2%93%E7%AC%94%E8%AE%B0&ie=utf-8")
|
||||
|
||||
@@ -233,7 +233,7 @@ def test_extract_note_detail_from_post_page():
|
||||
|
||||
assert note.note_id == "9117905169"
|
||||
assert note.title == "对于一个父亲来说,这个女儿14岁就死了"
|
||||
assert note.user_nickname == "章***轩"
|
||||
assert note.user_nickname == "章景轩"
|
||||
assert note.tieba_name == "以太比特吧"
|
||||
assert note.total_replay_num == 786
|
||||
assert note.total_replay_page == 13
|
||||
@@ -249,7 +249,7 @@ def test_extract_parent_comments_from_post_page():
|
||||
assert len(comments) == 30
|
||||
assert comments[0].comment_id == "150726491368"
|
||||
assert comments[0].content == "中国队第22金!无悬念!"
|
||||
assert comments[0].user_nickname == "h***n"
|
||||
assert comments[0].user_nickname == "heinzfrentzen"
|
||||
assert comments[0].tieba_name == "网球风云吧"
|
||||
# 教学版已移除 ip_location 等可定位真人字段
|
||||
|
||||
|
||||
@@ -0,0 +1,420 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_upstream.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""上游更新检查:git 输出怎么解析,以及什么时候才推送。
|
||||
|
||||
所有会碰网络的路径都被替掉了 —— 测试里既没有上游仓库,也不该有。真正被测的是
|
||||
解析、去重和调度到期这三件事,它们才是容易出错的部分。
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
|
||||
from api.main import app
|
||||
from api.monitor import db as monitor_db
|
||||
from api.monitor import notify
|
||||
from api.monitor import scheduler as scheduler_module
|
||||
from api.monitor import upstream
|
||||
from api.monitor.models import SETTING_WECOM_WEBHOOK
|
||||
from api.monitor.scheduler import MonitorScheduler
|
||||
from api.monitor.settings import set_setting
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
SEP = upstream._RECORD_SEPARATOR
|
||||
HEAD_SHA = "a" * 40
|
||||
TIP_SHA = "b" * 40
|
||||
WEBHOOK = "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=test"
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def db(tmp_path):
|
||||
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
|
||||
await monitor_db.init_db()
|
||||
yield monitor_db
|
||||
await monitor_db.dispose_engine()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def client(tmp_path):
|
||||
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
|
||||
await monitor_db.init_db()
|
||||
|
||||
transport = httpx.ASGITransport(app=app)
|
||||
async with httpx.AsyncClient(transport=transport, base_url="http://test") as http_client:
|
||||
yield http_client
|
||||
|
||||
await monitor_db.dispose_engine()
|
||||
|
||||
|
||||
def _signature(args: list[str]) -> str:
|
||||
"""Map one git invocation onto a key a test can name.
|
||||
|
||||
``rev-parse`` and ``rev-list`` are each called more than once with different
|
||||
arguments, so the subcommand alone is not enough to key on.
|
||||
"""
|
||||
if args[0] == "rev-parse":
|
||||
return f"rev-parse {args[1]}"
|
||||
if args[0] == "rev-list":
|
||||
return f"rev-list {args[-1]}"
|
||||
return args[0]
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fake_git(monkeypatch):
|
||||
"""Install canned git output. Returns the list of invocations made."""
|
||||
|
||||
def _install(responses: dict) -> list:
|
||||
calls: list = []
|
||||
|
||||
def _run(args, timeout):
|
||||
calls.append(args)
|
||||
key = _signature(args)
|
||||
if key not in responses:
|
||||
raise AssertionError(f"测试没有为这条 git 调用准备输出:{args}")
|
||||
return responses[key]
|
||||
|
||||
monkeypatch.setattr(upstream, "_run", _run)
|
||||
return calls
|
||||
|
||||
return _install
|
||||
|
||||
|
||||
def _log_line(sha: str, subject: str) -> str:
|
||||
return f"{sha}{SEP}张三{SEP}2026-10-01{SEP}{subject}"
|
||||
|
||||
|
||||
class TestCheckParsing:
|
||||
def test_counts_commits_and_reads_the_list(self, fake_git):
|
||||
calls = fake_git(
|
||||
{
|
||||
"rev-parse HEAD": (0, HEAD_SHA, ""),
|
||||
"fetch": (0, "", ""),
|
||||
"rev-parse FETCH_HEAD": (0, TIP_SHA, ""),
|
||||
"rev-list HEAD..FETCH_HEAD": (0, "3", ""),
|
||||
"rev-list FETCH_HEAD..HEAD": (0, "128", ""),
|
||||
"log": (
|
||||
0,
|
||||
"\n".join(
|
||||
[
|
||||
_log_line("abc1234", "fix: 修了扫码过期"),
|
||||
_log_line("def5678", "feat: 加了新平台"),
|
||||
_log_line("9999999", "docs: 更新说明"),
|
||||
]
|
||||
),
|
||||
"",
|
||||
),
|
||||
}
|
||||
)
|
||||
|
||||
result = upstream._check_sync("https://example.invalid/repo.git", "main")
|
||||
|
||||
assert result.ok is True
|
||||
assert result.behind == 3
|
||||
# 领先数就是这一层的规模,合并时要一起保留,所以值得单独报出来。
|
||||
assert result.ahead == 128
|
||||
assert result.tip == TIP_SHA
|
||||
assert result.head == HEAD_SHA
|
||||
assert [commit.subject for commit in result.commits] == [
|
||||
"fix: 修了扫码过期",
|
||||
"feat: 加了新平台",
|
||||
"docs: 更新说明",
|
||||
]
|
||||
assert result.commits[0].sha == "abc1234"
|
||||
assert result.error == ""
|
||||
|
||||
def test_up_to_date_skips_reading_the_log(self, fake_git):
|
||||
"""不落后时不该再去读提交列表 —— 那条 git log 没有意义。"""
|
||||
calls = fake_git(
|
||||
{
|
||||
"rev-parse HEAD": (0, HEAD_SHA, ""),
|
||||
"fetch": (0, "", ""),
|
||||
"rev-parse FETCH_HEAD": (0, TIP_SHA, ""),
|
||||
"rev-list HEAD..FETCH_HEAD": (0, "0", ""),
|
||||
"rev-list FETCH_HEAD..HEAD": (0, "128", ""),
|
||||
}
|
||||
)
|
||||
|
||||
result = upstream._check_sync("https://example.invalid/repo.git", "main")
|
||||
|
||||
assert result.ok is True
|
||||
assert result.behind == 0
|
||||
assert result.commits == []
|
||||
assert all(call[0] != "log" for call in calls)
|
||||
|
||||
def test_fetch_failure_is_reported_not_raised(self, fake_git):
|
||||
"""GitHub 不通是常态,那也该是一条能显示出来的结论。"""
|
||||
fake_git(
|
||||
{
|
||||
"rev-parse HEAD": (0, HEAD_SHA, ""),
|
||||
"fetch": (128, "", "fatal: unable to access 'https://github.com/': 连接超时\n第二行"),
|
||||
}
|
||||
)
|
||||
|
||||
result = upstream._check_sync("https://example.invalid/repo.git", "main")
|
||||
|
||||
assert result.ok is False
|
||||
assert result.head == HEAD_SHA
|
||||
# 只留第一行 stderr:git 的报错常常跟一大段建议,塞进界面反而看不清。
|
||||
assert "连接超时" in result.error
|
||||
assert "第二行" not in result.error
|
||||
|
||||
def test_not_a_git_repository_is_reported(self, fake_git):
|
||||
fake_git({"rev-parse HEAD": (128, "", "fatal: not a git repository (or any of the parent directories): .git")})
|
||||
|
||||
result = upstream._check_sync("https://example.invalid/repo.git", "main")
|
||||
|
||||
assert result.ok is False
|
||||
assert "读取本地 HEAD 失败" in result.error
|
||||
|
||||
def test_missing_git_binary_is_reported(self, monkeypatch):
|
||||
def _explode(args, timeout):
|
||||
raise FileNotFoundError("git")
|
||||
|
||||
monkeypatch.setattr(upstream, "_git", _explode)
|
||||
|
||||
result = upstream._check_sync("https://example.invalid/repo.git", "main")
|
||||
|
||||
assert result.ok is False
|
||||
assert "未找到 git" in result.error
|
||||
|
||||
def test_fetch_timeout_is_reported(self, monkeypatch):
|
||||
def _explode(args, timeout):
|
||||
raise subprocess.TimeoutExpired(cmd="git", timeout=timeout)
|
||||
|
||||
monkeypatch.setattr(upstream, "_git", _explode)
|
||||
|
||||
result = upstream._check_sync("https://example.invalid/repo.git", "main")
|
||||
|
||||
assert result.ok is False
|
||||
assert "超时" in result.error
|
||||
|
||||
|
||||
class TestNotification:
|
||||
@pytest.fixture
|
||||
def sent(self, monkeypatch) -> list:
|
||||
messages: list = []
|
||||
|
||||
async def _fake_send(url, content):
|
||||
messages.append(content)
|
||||
return True, "发送成功"
|
||||
|
||||
monkeypatch.setattr(notify, "send_wecom", _fake_send)
|
||||
return messages
|
||||
|
||||
@staticmethod
|
||||
def _patch_check(monkeypatch, *, behind: int, tip: str, ahead: int = 0):
|
||||
async def _fake_check(remote_url=upstream.DEFAULT_REMOTE_URL, branch=upstream.DEFAULT_BRANCH):
|
||||
return upstream.CheckResult(
|
||||
ok=True,
|
||||
behind=behind,
|
||||
ahead=ahead,
|
||||
tip=tip,
|
||||
head=HEAD_SHA,
|
||||
commits=[upstream.Commit(sha="abc1234", author="张三", date="2026-10-01", subject="fix: 修了扫码过期")],
|
||||
)
|
||||
|
||||
monkeypatch.setattr(upstream, "check", _fake_check)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_pushes_once_per_upstream_tip(self, db, monkeypatch, sent):
|
||||
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA, ahead=128)
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
|
||||
|
||||
first = await upstream.run_check()
|
||||
# 同一个 tip 再查一次:不该重复推。
|
||||
second = await upstream.run_check()
|
||||
|
||||
assert len(sent) == 1
|
||||
assert first["notified"] is True
|
||||
assert "notified" not in second
|
||||
assert "落后 `main` **2** 个提交" in sent[0]
|
||||
assert "abc1234" in sent[0]
|
||||
|
||||
# 上游又动了:tip 变了就该再推一次。
|
||||
self._patch_check(monkeypatch, behind=5, tip="c" * 40, ahead=128)
|
||||
third = await upstream.run_check()
|
||||
|
||||
assert len(sent) == 2
|
||||
assert third["notified"] is True
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_state_is_persisted_for_the_ui(self, db, monkeypatch, sent):
|
||||
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA, ahead=128)
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
|
||||
|
||||
await upstream.run_check()
|
||||
|
||||
async with monitor_db.get_session() as session:
|
||||
state = await upstream.load_state(session)
|
||||
|
||||
assert state["behind"] == 2
|
||||
assert state["ahead"] == 128
|
||||
assert state["tip"] == TIP_SHA
|
||||
assert state["branch"] == "main"
|
||||
assert state["checked_at"] > 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_nothing_new_does_not_push(self, db, monkeypatch, sent):
|
||||
self._patch_check(monkeypatch, behind=0, tip=TIP_SHA)
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
|
||||
|
||||
await upstream.run_check()
|
||||
|
||||
assert sent == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_switch_off_does_not_push(self, db, monkeypatch, sent):
|
||||
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA)
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
|
||||
await set_setting(session, "system.upstream_notify", "false")
|
||||
|
||||
await upstream.run_check()
|
||||
|
||||
assert sent == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_manual_check_can_skip_the_push(self, db, monkeypatch, sent):
|
||||
"""手动点「立即检查」只看结果,不因为它把群消息推一遍。"""
|
||||
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA)
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
|
||||
|
||||
await upstream.run_check(notify_when_new=False)
|
||||
|
||||
assert sent == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_no_webhook_configured_is_not_an_error(self, db, monkeypatch, sent):
|
||||
self._patch_check(monkeypatch, behind=2, tip=TIP_SHA)
|
||||
|
||||
result = await upstream.run_check()
|
||||
|
||||
assert sent == []
|
||||
assert result["ok"] is True
|
||||
assert result["behind"] == 2
|
||||
|
||||
|
||||
class TestScheduler:
|
||||
@pytest.fixture
|
||||
def checks(self, monkeypatch) -> list:
|
||||
calls: list = []
|
||||
|
||||
async def _fake_run_check(notify_when_new: bool = True):
|
||||
calls.append(notify_when_new)
|
||||
# 真实的 run_check 会把 checked_at 写进去,调度器的「到点了没有」
|
||||
# 全靠这个字段,所以替身也必须写。
|
||||
async with monitor_db.get_session() as session:
|
||||
await upstream._save_state(
|
||||
session,
|
||||
{"checked_at": get_current_timestamp(), "ok": True, "behind": 0},
|
||||
)
|
||||
return {"ok": True, "behind": 0}
|
||||
|
||||
monkeypatch.setattr(upstream, "run_check", _fake_run_check)
|
||||
return calls
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_disabled_never_checks(self, db, checks):
|
||||
await MonitorScheduler()._maybe_check_upstream()
|
||||
|
||||
assert checks == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_checks_when_due_and_then_waits_out_the_interval(self, db, checks):
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, "system.upstream_check_enabled", "true")
|
||||
|
||||
scheduler = MonitorScheduler()
|
||||
await scheduler._maybe_check_upstream()
|
||||
assert checks == [True]
|
||||
|
||||
# 刚查过:间隔(默认一天)没到就不该再查。
|
||||
await scheduler._maybe_check_upstream()
|
||||
assert checks == [True]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_stale_timestamp_is_due_again(self, db, checks):
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, "system.upstream_check_enabled", "true")
|
||||
await set_setting(session, "system.upstream_check_interval_minutes", "30")
|
||||
await upstream._save_state(
|
||||
session,
|
||||
# 差一分钟就到期,用来卡住边界:31 分钟前那次已经算过期。
|
||||
{"checked_at": get_current_timestamp() - 31 * 60_000, "ok": True, "behind": 0},
|
||||
)
|
||||
|
||||
await MonitorScheduler()._maybe_check_upstream()
|
||||
|
||||
assert checks == [True]
|
||||
|
||||
|
||||
class TestEndpoint:
|
||||
@pytest.mark.asyncio
|
||||
async def test_status_is_empty_before_the_first_check(self, client):
|
||||
response = await client.get("/api/monitor/upstream")
|
||||
|
||||
assert response.status_code == 200
|
||||
assert response.json() == {}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_manual_check_runs_and_is_readable_back(self, client, monkeypatch):
|
||||
async def _fake_check(remote_url=upstream.DEFAULT_REMOTE_URL, branch=upstream.DEFAULT_BRANCH):
|
||||
return upstream.CheckResult(ok=True, behind=1, tip=TIP_SHA, head=HEAD_SHA)
|
||||
|
||||
monkeypatch.setattr(upstream, "check", _fake_check)
|
||||
|
||||
checked = (await client.post("/api/monitor/upstream/check")).json()
|
||||
assert checked["behind"] == 1
|
||||
|
||||
# 结果落库,随后的 GET 读的是同一份缓存(而不是再 fetch 一次)。
|
||||
cached = (await client.get("/api/monitor/upstream")).json()
|
||||
assert cached["behind"] == 1
|
||||
assert cached["tip"] == TIP_SHA
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_manual_check_does_not_push(self, client, monkeypatch):
|
||||
"""点按钮的人正看着结果,不该再给自己推一条群消息。"""
|
||||
sent: list = []
|
||||
|
||||
async def _fake_send(url, content):
|
||||
sent.append(content)
|
||||
return True, "发送成功"
|
||||
|
||||
async def _fake_check(remote_url=upstream.DEFAULT_REMOTE_URL, branch=upstream.DEFAULT_BRANCH):
|
||||
return upstream.CheckResult(ok=True, behind=1, tip=TIP_SHA, head=HEAD_SHA)
|
||||
|
||||
monkeypatch.setattr(notify, "send_wecom", _fake_send)
|
||||
monkeypatch.setattr(upstream, "check", _fake_check)
|
||||
async with monitor_db.get_session() as session:
|
||||
await set_setting(session, SETTING_WECOM_WEBHOOK, WEBHOOK)
|
||||
|
||||
body = (await client.post("/api/monitor/upstream/check")).json()
|
||||
|
||||
assert sent == []
|
||||
assert body["behind"] == 1
|
||||
# 没推过的那批提交留给下一次定时检查,所以这里不该记成已推送。
|
||||
async with monitor_db.get_session() as session:
|
||||
state = await upstream.load_state(session)
|
||||
assert "notified" not in state
|
||||
@@ -22,6 +22,20 @@ import asyncio
|
||||
|
||||
import pytest
|
||||
|
||||
import config
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _force_nickname_masking(monkeypatch):
|
||||
"""这一组验的是**脱敏机制本身**,所以强制把它打开。
|
||||
|
||||
本仓库的部署配置是关掉的(config.MASK_NICKNAME = False)—— 监控的是一批公开创作者
|
||||
账号,而脱敏是有损的(「张三」「张四」都成「张*」),分不出谁是谁。机制仍然必须正确,
|
||||
所以这里显式打开来测。
|
||||
"""
|
||||
monkeypatch.setattr(config, "MASK_NICKNAME", True)
|
||||
|
||||
|
||||
# 原始(明文)测试数据
|
||||
RAW_USER_ID = 7654321
|
||||
RAW_NICKNAME = "微博达人"
|
||||
|
||||
@@ -0,0 +1,200 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""Clone one schema to a new database on the same MySQL server, and provision a
|
||||
scoped account for it.
|
||||
|
||||
Written for the production cut-over: the monitor data lives in `mediacrawler` and
|
||||
the server deployment reads `mediacrawler_prod`. Both sit on the same host, so the
|
||||
copy is a cross-schema ``INSERT ... SELECT`` rather than a dump-and-reload -- and
|
||||
since neither this workstation nor the server has a mysql client, that is also
|
||||
the only option available.
|
||||
|
||||
Two things this deliberately does NOT do, both because the server hosts ~22
|
||||
unrelated production databases and the provisioning account is a full admin:
|
||||
|
||||
* it never writes outside the two schemas named on the command line;
|
||||
* it hands the application its own account, scoped to the new schema, rather than
|
||||
reusing that admin account for the app.
|
||||
|
||||
Foreign keys are disabled only for the duration of the copy. Source and target
|
||||
definitions are identical, so ordering is the sole thing at stake, and turning
|
||||
the checks off is what makes an arbitrary table order safe.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import secrets
|
||||
import string
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import aiomysql
|
||||
|
||||
PASSWORD_ALPHABET = string.ascii_letters + string.digits
|
||||
|
||||
|
||||
def generate_password(length: int = 32) -> str:
|
||||
"""Alphanumeric only: it ends up in a .env value and a SQL literal, and a
|
||||
symbol that needs escaping in either place is a support ticket waiting."""
|
||||
return "".join(secrets.choice(PASSWORD_ALPHABET) for _ in range(length))
|
||||
|
||||
|
||||
async def _admin_conn(args, db=None):
|
||||
return await aiomysql.connect(
|
||||
host=args.host,
|
||||
port=args.port,
|
||||
user=args.admin_user,
|
||||
password=args.admin_password,
|
||||
db=db,
|
||||
charset="utf8mb4",
|
||||
autocommit=True,
|
||||
)
|
||||
|
||||
|
||||
async def create_schema(args) -> None:
|
||||
conn = await _admin_conn(args)
|
||||
try:
|
||||
cur = await conn.cursor()
|
||||
await cur.execute(
|
||||
f"CREATE DATABASE IF NOT EXISTS `{args.target_db}` "
|
||||
"DEFAULT CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci"
|
||||
)
|
||||
await cur.execute(
|
||||
"SELECT DEFAULT_CHARACTER_SET_NAME, DEFAULT_COLLATION_NAME "
|
||||
"FROM information_schema.SCHEMATA WHERE SCHEMA_NAME = %s",
|
||||
(args.target_db,),
|
||||
)
|
||||
charset, collation = await cur.fetchone()
|
||||
print(f"[库] {args.target_db} {charset} / {collation}")
|
||||
await cur.close()
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
|
||||
async def provision_app_user(args, password: str) -> None:
|
||||
conn = await _admin_conn(args)
|
||||
try:
|
||||
cur = await conn.cursor()
|
||||
# The host is bound as a parameter instead of written as a literal '%':
|
||||
# aiomysql applies %-substitution whenever arguments are supplied, so a
|
||||
# bare '%' in the SQL is read as a format specifier and raises
|
||||
# "unsupported format character".
|
||||
await cur.execute(
|
||||
"CREATE USER IF NOT EXISTS %s@%s IDENTIFIED BY %s",
|
||||
(args.app_user, "%", password),
|
||||
)
|
||||
# Idempotent re-runs: if the account already existed, make the password
|
||||
# the one we are about to write into .env rather than a stale one.
|
||||
await cur.execute(
|
||||
"ALTER USER %s@%s IDENTIFIED BY %s", (args.app_user, "%", password)
|
||||
)
|
||||
await cur.execute(
|
||||
f"GRANT ALL PRIVILEGES ON `{args.target_db}`.* TO %s@%s",
|
||||
(args.app_user, "%"),
|
||||
)
|
||||
await cur.execute("FLUSH PRIVILEGES")
|
||||
print(f"[账号] {args.app_user}@% 授权范围仅 `{args.target_db}`.*")
|
||||
await cur.close()
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
|
||||
async def copy_tables(args) -> None:
|
||||
src = await _admin_conn(args, db=args.source_db)
|
||||
dst = await _admin_conn(args, db=args.target_db)
|
||||
try:
|
||||
scur = await src.cursor()
|
||||
await scur.execute("SHOW TABLES")
|
||||
tables = [row[0] for row in await scur.fetchall()]
|
||||
print(f"[表] 源库共 {len(tables)} 张")
|
||||
|
||||
dcur = await dst.cursor()
|
||||
# Order is arbitrary on purpose -- checks are off for the whole copy, so
|
||||
# a table may safely be created before the table it references.
|
||||
await dcur.execute("SET FOREIGN_KEY_CHECKS=0")
|
||||
|
||||
for table in tables:
|
||||
await scur.execute(f"SHOW CREATE TABLE `{table}`")
|
||||
ddl = (await scur.fetchone())[1]
|
||||
# `IF NOT EXISTS` makes the script re-runnable after a partial run.
|
||||
await dcur.execute(ddl.replace("CREATE TABLE", "CREATE TABLE IF NOT EXISTS", 1))
|
||||
|
||||
await dcur.execute(f"DELETE FROM `{table}`")
|
||||
await dcur.execute(
|
||||
f"INSERT INTO `{table}` SELECT * FROM `{args.source_db}`.`{table}`"
|
||||
)
|
||||
copied = dcur.rowcount
|
||||
|
||||
await scur.execute(f"SELECT COUNT(*) FROM `{table}`")
|
||||
expected = (await scur.fetchone())[0]
|
||||
flag = "OK" if copied == expected else "!! 行数不符"
|
||||
print(f" {table:<26} {copied:>6} / {expected:<6} {flag}")
|
||||
|
||||
await dcur.execute("SET FOREIGN_KEY_CHECKS=1")
|
||||
await dcur.close()
|
||||
await scur.close()
|
||||
finally:
|
||||
src.close()
|
||||
dst.close()
|
||||
|
||||
|
||||
def write_env(path: Path, args, password: str) -> None:
|
||||
path.write_text(
|
||||
f"""# 服务器部署环境配置(已被 .gitignore 忽略,不会提交)
|
||||
#
|
||||
# 由 tools/clone_database.py 生成。库名改了之后必须同时确认账号对该库有权限:
|
||||
# 应用启动时会校验「实际连到的库」是否等于下面的 MYSQL_DB_NAME,不符会拒绝启动。
|
||||
MC_HOST=0.0.0.0
|
||||
MC_PORT=18051
|
||||
|
||||
# --- 数据库(监控层) ---
|
||||
MYSQL_DB_HOST={args.host}
|
||||
MYSQL_DB_PORT={args.port}
|
||||
MYSQL_DB_USER={args.app_user}
|
||||
MYSQL_DB_PWD={password}
|
||||
MYSQL_DB_NAME={args.target_db}
|
||||
|
||||
# 登录鉴权:留空则首次启动自动生成随机密码并打印在启动日志里
|
||||
# MC_PASSWORD=
|
||||
|
||||
# 面板走 HTTPS 时才打开;局域网明文 HTTP 下必须保持注释,
|
||||
# 否则浏览器丢弃 Cookie,表现为登录页反复刷新且无任何报错
|
||||
# MC_COOKIE_SECURE=1
|
||||
""",
|
||||
encoding="utf-8",
|
||||
)
|
||||
path.chmod(0o600)
|
||||
|
||||
|
||||
async def main() -> int:
|
||||
ap = argparse.ArgumentParser(description=__doc__)
|
||||
ap.add_argument("--host", default="192.168.2.27")
|
||||
ap.add_argument("--port", type=int, default=3306)
|
||||
ap.add_argument("--admin-user", required=True)
|
||||
ap.add_argument("--admin-password", required=True)
|
||||
ap.add_argument("--source-db", required=True)
|
||||
ap.add_argument("--target-db", required=True)
|
||||
ap.add_argument("--app-user", required=True)
|
||||
ap.add_argument("--env-file", type=Path)
|
||||
args = ap.parse_args()
|
||||
|
||||
if args.source_db == args.target_db:
|
||||
print("源库与目标库相同,拒绝执行", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
password = generate_password()
|
||||
|
||||
await create_schema(args)
|
||||
await provision_app_user(args, password)
|
||||
await copy_tables(args)
|
||||
|
||||
if args.env_file:
|
||||
write_env(args.env_file, args, password)
|
||||
print(f"[.env] 已写入 {args.env_file}(权限 600,密码未回显)")
|
||||
|
||||
print("\n完成。")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(asyncio.run(main()))
|
||||
@@ -0,0 +1,210 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""Phase 0 probe: can the creator backend be reached with a plain signed request?
|
||||
|
||||
The question this answers, and why it is worth a whole script: two sources
|
||||
disagree. The `xhshow` library ships `sign_xyw()` whose docstring says it exists
|
||||
because the creator data APIs *reject* the main-site signature with HTTP 406 -- but
|
||||
a field report claims the creator gateway rejects *any* self-made request with 406
|
||||
regardless of signature. Only a live request settles it.
|
||||
|
||||
The response is read as a three-way verdict, because a bare "it failed" is not
|
||||
useful here:
|
||||
|
||||
406 -> the gateway rejected the signature. The browser-interception route is
|
||||
the only way forward.
|
||||
401 / not-logged-in
|
||||
-> the signature PASSED and only the creator session is missing. That is
|
||||
good news: it means Phase 1 is pure-request after all.
|
||||
200 -> we are through, and the payload is captured for field mapping.
|
||||
|
||||
Cookies are read out of the browser over CDP and never printed -- they are
|
||||
credentials, and this script has no reason to echo them.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import base64
|
||||
import hashlib
|
||||
import json
|
||||
import sys
|
||||
import urllib.parse
|
||||
from datetime import datetime, time, timedelta
|
||||
|
||||
# The XYW_ scheme, matching both xhshow/config/config.py and the independent
|
||||
# reverse-engineering in xiaohongshu-cli. Constants are byte-identical in both.
|
||||
XYW_AES_KEY = b"7cc4adla5ay0701v"
|
||||
XYW_AES_IV = b"4uzjr7mbsibcaldp"
|
||||
XYW_ENV_FLAGS = "0|0|0|1|0|0|1|0|0|0|1|0|0|0|0|1|0|0|0"
|
||||
|
||||
CREATOR_ORIGIN = "https://creator.xiaohongshu.com"
|
||||
# The data-analysis note list. This is the page the creator console itself calls,
|
||||
# and it is where exposure/views live.
|
||||
NOTE_LIST_PATH = "/api/galaxy/creator/datacenter/note/analyze/list"
|
||||
|
||||
USER_AGENT = (
|
||||
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
|
||||
"(KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36"
|
||||
)
|
||||
|
||||
|
||||
def _aes_encrypt_hex(plaintext: str) -> str:
|
||||
from Crypto.Cipher import AES
|
||||
from Crypto.Util.Padding import pad
|
||||
|
||||
cipher = AES.new(XYW_AES_KEY, AES.MODE_CBC, XYW_AES_IV)
|
||||
return cipher.encrypt(pad(plaintext.encode("utf-8"), AES.block_size)).hex()
|
||||
|
||||
|
||||
def sign_xyw(api: str, a1: str, app_id: str = "ugc", data: dict | None = None) -> dict[str, str]:
|
||||
"""Mirror of xiaohongshu-cli's creator_signing.sign_creator.
|
||||
|
||||
Written out rather than imported so the probe can vary ``api`` and ``app_id``
|
||||
independently -- figuring out which combination the gateway accepts is the
|
||||
entire point of running it.
|
||||
"""
|
||||
content = api
|
||||
if data is not None:
|
||||
content += json.dumps(data, separators=(",", ":"), ensure_ascii=False)
|
||||
|
||||
digest = hashlib.md5(content.encode("utf-8")).hexdigest()
|
||||
timestamp_ms = int(datetime.now().timestamp() * 1000)
|
||||
plaintext = f"x1={digest};x2={XYW_ENV_FLAGS};x3={a1};x4={timestamp_ms};"
|
||||
encoded = base64.b64encode(plaintext.encode("utf-8")).decode("utf-8")
|
||||
|
||||
envelope = {
|
||||
"signSvn": "56",
|
||||
"signType": "x2",
|
||||
"appId": app_id,
|
||||
"signVersion": "1",
|
||||
"payload": _aes_encrypt_hex(encoded),
|
||||
}
|
||||
x_s = "XYW_" + base64.b64encode(
|
||||
json.dumps(envelope, separators=(",", ":")).encode("utf-8")
|
||||
).decode("utf-8")
|
||||
return {"x-s": x_s, "x-t": str(timestamp_ms)}
|
||||
|
||||
|
||||
def build_query(start_days_ago: int, end_days_ago: int, page_size: int = 10) -> str:
|
||||
"""Query string exactly as the working collector builds it: epoch millis."""
|
||||
today = datetime.now().replace(hour=0, minute=0, second=0, microsecond=0)
|
||||
|
||||
def ms(days_ago: int, at_end: bool) -> int:
|
||||
day = (today - timedelta(days=days_ago)).date()
|
||||
clock = time(23, 59, 59) if at_end else time(0, 0, 0)
|
||||
return int(datetime.combine(day, clock).timestamp() * 1000)
|
||||
|
||||
return urllib.parse.urlencode(
|
||||
{
|
||||
"post_begin_time": ms(start_days_ago, False),
|
||||
"post_end_time": ms(end_days_ago, True),
|
||||
"type": 0,
|
||||
"page_size": page_size,
|
||||
"page_num": 1,
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
async def cookies_from_cdp() -> dict[str, str]:
|
||||
"""Session cookies for creator.xiaohongshu.com, read out of the live browser."""
|
||||
from playwright.async_api import async_playwright
|
||||
|
||||
playwright = await async_playwright().start()
|
||||
try:
|
||||
browser = await playwright.chromium.connect_over_cdp(
|
||||
"http://127.0.0.1:9222", timeout=15000
|
||||
)
|
||||
jars = {}
|
||||
for context in browser.contexts:
|
||||
for cookie in await context.cookies():
|
||||
jars[cookie["name"]] = cookie["value"]
|
||||
return jars
|
||||
finally:
|
||||
await playwright.stop()
|
||||
|
||||
|
||||
def classify(status: int, body: str) -> str:
|
||||
if status == 406:
|
||||
return "406 —— 网关拒了签名(回退浏览器拦截路线)"
|
||||
if status == 200:
|
||||
return "200 —— 通了,可以考虑解析数据"
|
||||
if status in (401, 403):
|
||||
return f"{status} —— 签名过了,只差创作者会话(好消息)"
|
||||
return f"{status} —— 未知,需要看响应体"
|
||||
|
||||
|
||||
async def main() -> int:
|
||||
import httpx
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--start-days-ago", type=int, default=30)
|
||||
ap.add_argument("--end-days-ago", type=int, default=0)
|
||||
ap.add_argument("--app-id", default="ugc", help="参考实现用 ugc;xhshow 默认 xhs-pc-web")
|
||||
ap.add_argument("--cookie", default="", help="留空则从 CDP 浏览器读取")
|
||||
ap.add_argument("--show-body", action="store_true", help="打印响应前 800 字符")
|
||||
ap.add_argument(
|
||||
"--no-cookie-header",
|
||||
action="store_true",
|
||||
help="签名照签(仍需 a1)但不发 cookie 头,用来分清"
|
||||
"「空数据是缺会话」还是「接口本身就这样」",
|
||||
)
|
||||
args = ap.parse_args()
|
||||
|
||||
if args.cookie:
|
||||
cookies = dict(
|
||||
pair.split("=", 1) for pair in args.cookie.split("; ") if "=" in pair
|
||||
)
|
||||
else:
|
||||
cookies = await cookies_from_cdp()
|
||||
|
||||
print(f" cookie 条数 {len(cookies)},名字: {sorted(cookies)}")
|
||||
if not cookies.get("a1"):
|
||||
print(" ✗ 没有 a1 —— 签名必须用它,无法继续")
|
||||
return 2
|
||||
|
||||
query = build_query(args.start_days_ago, args.end_days_ago)
|
||||
cookie_header = "; ".join(f"{k}={v}" for k, v in cookies.items())
|
||||
headers_common = {
|
||||
"user-agent": USER_AGENT,
|
||||
"accept": "application/json, text/plain, */*",
|
||||
"origin": CREATOR_ORIGIN,
|
||||
"referer": f"{CREATOR_ORIGIN}/statistics/data-analysis",
|
||||
"accept-language": "zh-CN,zh;q=0.9",
|
||||
}
|
||||
if not args.no_cookie_header:
|
||||
headers_common["cookie"] = cookie_header
|
||||
|
||||
# Three signing variants: the reference bakes the query into the signed string,
|
||||
# but the exact form is not documented beyond an example with no query at all.
|
||||
variants = {
|
||||
"path+q(参考实现写法)": f"url={NOTE_LIST_PATH}?{query}",
|
||||
"path only": f"url={NOTE_LIST_PATH}",
|
||||
"裸 path+query(无 url= 前缀)": f"{NOTE_LIST_PATH}?{query}",
|
||||
}
|
||||
|
||||
url = f"{CREATOR_ORIGIN}{NOTE_LIST_PATH}?{query}"
|
||||
async with httpx.AsyncClient(timeout=25, follow_redirects=False) as client:
|
||||
for label, api in variants.items():
|
||||
signature = sign_xyw(api, cookies["a1"], app_id=args.app_id)
|
||||
headers = {**headers_common, **signature}
|
||||
try:
|
||||
response = await client.get(url, headers=headers)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
print(f" [{label}] 请求异常: {exc.__class__.__name__}: {exc}")
|
||||
continue
|
||||
|
||||
print(f"\n [{label}]")
|
||||
print(f" HTTP {response.status_code} {classify(response.status_code, response.text)}")
|
||||
body = response.text or ""
|
||||
if body:
|
||||
print(f" 响应前 160 字符: {body[:160]!r}")
|
||||
if args.show_body and body:
|
||||
print(f" 完整响应: {body[:800]}")
|
||||
|
||||
print("\n 提示:若三种都返回 406,再试 --app-id xhs-pc-web。")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import asyncio
|
||||
|
||||
sys.exit(asyncio.run(main()))
|
||||
@@ -0,0 +1,144 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""Phase 0, second half: read the real request off the real page.
|
||||
|
||||
The signed-request probe proved the signature is forgeable and that the main-site
|
||||
cookie authenticates the creator backend -- but the endpoint it called returned an
|
||||
envelope with no payload, which no real list endpoint does. That path came from a
|
||||
third-party repo and may simply be stale.
|
||||
|
||||
So stop guessing at paths and watch the page. This opens the data-analysis page in
|
||||
the browser that is already signed in, records every creator API call the page
|
||||
itself makes, and prints each request's URL, method and the *shape* of the
|
||||
response. The page's own requests are the ground truth.
|
||||
|
||||
Read-only: it navigates and observes. Nothing is submitted.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import json
|
||||
import sys
|
||||
from collections import Counter
|
||||
|
||||
DEFAULT_URL = "https://creator.xiaohongshu.com/statistics/data-analysis"
|
||||
|
||||
|
||||
def shape(value, depth: int = 0) -> str:
|
||||
"""Describe a JSON value's structure without dumping its data."""
|
||||
if depth > 3:
|
||||
return "…"
|
||||
if isinstance(value, dict):
|
||||
if not value:
|
||||
return "{}"
|
||||
inner = ", ".join(f"{k}: {shape(v, depth + 1)}" for k, v in list(value.items())[:12])
|
||||
return "{" + inner + "}"
|
||||
if isinstance(value, list):
|
||||
if not value:
|
||||
return "[]"
|
||||
return f"[{len(value)} × {shape(value[0], depth + 1)}]"
|
||||
if isinstance(value, str):
|
||||
# Short values are shown as themselves -- the interesting ones here are
|
||||
# status codes, roles and permission names, and "str(14)" tells you
|
||||
# nothing. Long ones are almost always ids or urls, so only their length.
|
||||
return repr(value) if len(value) <= 40 else f"str({len(value)})"
|
||||
if isinstance(value, bool):
|
||||
return str(value)
|
||||
if isinstance(value, (int, float)):
|
||||
return str(value)
|
||||
return type(value).__name__
|
||||
|
||||
|
||||
async def main() -> int:
|
||||
from playwright.async_api import async_playwright
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--url", default=DEFAULT_URL)
|
||||
ap.add_argument("--wait", type=int, default=25, help="观察窗口(秒)")
|
||||
ap.add_argument("--filter", default="/api/galaxy", help="只记录 URL 含此串的请求")
|
||||
args = ap.parse_args()
|
||||
|
||||
playwright = await async_playwright().start()
|
||||
seen: Counter[str] = Counter()
|
||||
details: list[str] = []
|
||||
|
||||
try:
|
||||
browser = await playwright.chromium.connect_over_cdp(
|
||||
"http://127.0.0.1:9222", timeout=15000
|
||||
)
|
||||
context = browser.contexts[0]
|
||||
|
||||
async def on_response(response):
|
||||
url = response.url
|
||||
if args.filter not in url:
|
||||
return
|
||||
key = url.split("?")[0]
|
||||
seen[key] += 1
|
||||
if seen[key] > 1:
|
||||
return
|
||||
|
||||
request = response.request
|
||||
line = [f"\n {request.method} {key}"]
|
||||
line.append(f" HTTP {response.status}")
|
||||
|
||||
query = url.split("?", 1)[1] if "?" in url else ""
|
||||
if query:
|
||||
line.append(f" 查询串: {query[:300]}")
|
||||
|
||||
post = request.post_data
|
||||
if post:
|
||||
line.append(f" POST body: {post[:300]}")
|
||||
|
||||
try:
|
||||
payload = await response.json()
|
||||
line.append(f" 响应结构: {shape(payload)[:600]}")
|
||||
except Exception:
|
||||
try:
|
||||
text = await response.text()
|
||||
line.append(f" 响应(非JSON)前 200: {text[:200]!r}")
|
||||
except Exception as exc: # noqa: BLE001
|
||||
line.append(f" 响应不可读: {exc.__class__.__name__}")
|
||||
details.append("\n".join(line))
|
||||
|
||||
context.on("response", on_response)
|
||||
|
||||
page = await context.new_page()
|
||||
try:
|
||||
await page.goto(args.url, wait_until="domcontentloaded", timeout=45000)
|
||||
print(f" 落地 URL: {page.url[:120]}")
|
||||
title = await page.title()
|
||||
print(f" 标题: {title[:80]!r}")
|
||||
# A creator console that is genuinely reachable renders its shell; a
|
||||
# login gate does not. This is the cheapest "are we in?" signal.
|
||||
for label, selector in (
|
||||
("登录表单", "//input[@type='password']"),
|
||||
("扫码登录", "//*[contains(@class,'qrcode') or contains(@class,'qr-code')]"),
|
||||
):
|
||||
if await page.locator(selector).count() > 0:
|
||||
print(f" ★ 页面上出现「{label}」—— 这个账号似乎没有创作者后台会话")
|
||||
|
||||
print(f"\n 观察 {args.wait} 秒,记录页面自己发的请求…")
|
||||
await asyncio.sleep(args.wait)
|
||||
finally:
|
||||
try:
|
||||
await page.close()
|
||||
except Exception:
|
||||
pass
|
||||
context.remove_listener("response", on_response)
|
||||
|
||||
if not details:
|
||||
print("\n ★ 没有捕获到任何匹配的请求 —— 页面很可能停在登录页,没有发出数据请求")
|
||||
for block in details:
|
||||
print(block)
|
||||
|
||||
print(f"\n 捕获到的接口(去重): {len(seen)}")
|
||||
for key, count in seen.most_common():
|
||||
print(f" ×{count} {key}")
|
||||
finally:
|
||||
await playwright.stop()
|
||||
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(asyncio.run(main()))
|
||||
@@ -7,6 +7,8 @@
|
||||
# 昵称保留但做中间脱敏)。本模块提供匿名化与脱敏工具。
|
||||
import hashlib
|
||||
|
||||
import config
|
||||
|
||||
|
||||
def anonymize_user_id(user_id) -> str:
|
||||
"""把原始用户 ID 转成匿名哈希,用于内容/评论记录的创作者分组,
|
||||
@@ -25,10 +27,16 @@ def mask_nickname(name) -> str:
|
||||
- 长度 == 2:首字 + "*"
|
||||
- 长度 >= 3:首字 + "***" + 尾字
|
||||
这样既保留教学分析所需的内容归属语义,又无法据昵称定位到真人。
|
||||
|
||||
**是否启用由 config.MASK_NICKNAME 决定,本仓库默认关闭**(原样返回)。
|
||||
脱敏是有损的,撞名很常见 —— 详见 base_config 里那一段的说明。
|
||||
开关读的是模块属性而不是导入时的值,这样测试可以 monkeypatch 它。
|
||||
"""
|
||||
if name is None:
|
||||
return ""
|
||||
s = str(name)
|
||||
if not getattr(config, "MASK_NICKNAME", False):
|
||||
return s
|
||||
if len(s) <= 1:
|
||||
return "*"
|
||||
if len(s) == 2:
|
||||
|
||||
+3
-1
@@ -5,13 +5,14 @@ import { Sidebar } from '@/components/layout/Sidebar'
|
||||
import { MainContent } from '@/components/layout/MainContent'
|
||||
import { CrawlerConfigPanel } from '@/components/config/CrawlerConfigPanel'
|
||||
import { MonitorDashboard } from '@/components/monitor/MonitorDashboard'
|
||||
import { OperationView } from '@/components/creator/OperationView'
|
||||
import { ReportView } from '@/components/monitor/ReportView'
|
||||
import { SettingsView } from '@/components/settings/SettingsView'
|
||||
import { Login } from '@/components/auth/Login'
|
||||
import { EnvironmentCheck, isEnvChecked } from '@/components/env/EnvironmentCheck'
|
||||
import { authApi, setUnauthorizedHandler } from '@/lib/api'
|
||||
|
||||
export type AppView = 'crawler' | 'monitor' | 'report' | 'settings'
|
||||
export type AppView = 'crawler' | 'monitor' | 'operation' | 'report' | 'settings'
|
||||
|
||||
function App() {
|
||||
// null = still probing. Rendering the app while unknown would briefly mount
|
||||
@@ -91,6 +92,7 @@ function App() {
|
||||
</>
|
||||
)}
|
||||
{view === 'monitor' && <MonitorDashboard />}
|
||||
{view === 'operation' && <OperationView />}
|
||||
{view === 'report' && <ReportView />}
|
||||
{view === 'settings' && <SettingsView onNavigate={setView} />}
|
||||
</div>
|
||||
|
||||
@@ -0,0 +1,184 @@
|
||||
import { useEffect, useState } from 'react'
|
||||
import { useQueryClient } from '@tanstack/react-query'
|
||||
import { AlertTriangle, CheckCircle2, Loader2, QrCode, RefreshCw, X } from 'lucide-react'
|
||||
|
||||
import { Button } from '@/components/ui/button'
|
||||
import {
|
||||
Dialog,
|
||||
DialogContent,
|
||||
DialogDescription,
|
||||
DialogHeader,
|
||||
DialogTitle,
|
||||
} from '@/components/ui/dialog'
|
||||
import {
|
||||
useCancelCreatorLogin,
|
||||
useCreatorLoginStatus,
|
||||
useStartCreatorLogin,
|
||||
} from '@/hooks/useCreator'
|
||||
|
||||
/**
|
||||
* 扫码新增一个运营账号。
|
||||
*
|
||||
* 后端给每个账号开一个**临时浏览器上下文**,扫完取出 cookie 就丢弃,所以:
|
||||
* 登第二个账号不会把第一个顶掉,也不会影响监控那个登录态。
|
||||
*/
|
||||
export function AddAccountDialog({
|
||||
open,
|
||||
onOpenChange,
|
||||
}: {
|
||||
open: boolean
|
||||
onOpenChange: (open: boolean) => void
|
||||
}) {
|
||||
const queryClient = useQueryClient()
|
||||
const [polling, setPolling] = useState(false)
|
||||
|
||||
const { data: state } = useCreatorLoginStatus(polling)
|
||||
const start = useStartCreatorLogin()
|
||||
const cancel = useCancelCreatorLogin()
|
||||
|
||||
const status = state?.status ?? 'idle'
|
||||
|
||||
useEffect(() => {
|
||||
if (!open) return
|
||||
if (status === 'waiting' || status === 'idle') return
|
||||
setPolling(false)
|
||||
if (status !== 'success') return
|
||||
|
||||
queryClient.invalidateQueries({ queryKey: ['creatorAccounts'] })
|
||||
// 自动关掉:停在二维码上会让人以为没成功,而账号其实已经进了列表。
|
||||
const timer = setTimeout(() => onOpenChange(false), 1600)
|
||||
return () => clearTimeout(timer)
|
||||
}, [status, open, queryClient, onOpenChange])
|
||||
|
||||
// 关闭时统一收尾:拆掉后端那个临时上下文(别让它挂在操作者的 Chrome 里),
|
||||
// 并清掉后端记住的结果 —— 否则下次打开弹窗会立刻显示上一次的成功状态。
|
||||
useEffect(() => {
|
||||
if (open) return
|
||||
setPolling(false)
|
||||
cancel.mutate()
|
||||
// 只在开关变化时跑,cancel/依赖本身不该触发它。
|
||||
// eslint-disable-next-line react-hooks/exhaustive-deps
|
||||
}, [open])
|
||||
|
||||
const begin = () => {
|
||||
setPolling(true)
|
||||
start.mutate()
|
||||
}
|
||||
|
||||
const busy = start.isPending || cancel.isPending
|
||||
|
||||
return (
|
||||
<Dialog open={open} onOpenChange={onOpenChange}>
|
||||
<DialogContent className="max-w-md">
|
||||
<DialogHeader>
|
||||
<DialogTitle className="font-mono">新增运营账号</DialogTitle>
|
||||
<DialogDescription className="font-mono text-xs">
|
||||
用小红书 App 扫码登录要添加的账号。每个账号使用独立的临时会话,互不影响。
|
||||
</DialogDescription>
|
||||
</DialogHeader>
|
||||
|
||||
<div className="space-y-3 py-2">
|
||||
{status === 'waiting' && state?.image && (
|
||||
<div className="flex flex-col items-center gap-3">
|
||||
{/* 二维码直接来自页面,是个 data: URL —— 不经过第三方,也不在服务器上落盘 */}
|
||||
<img
|
||||
src={state.image}
|
||||
alt="登录二维码"
|
||||
className="w-44 h-44 rounded-md border border-cyber-border-DEFAULT bg-white p-1"
|
||||
/>
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted">
|
||||
剩余 <span className="text-cyber-neon-cyan">{state.expires_in}</span> 秒
|
||||
</p>
|
||||
<Button
|
||||
variant="ghost"
|
||||
size="sm"
|
||||
disabled={busy}
|
||||
onClick={() => {
|
||||
setPolling(false)
|
||||
cancel.mutate()
|
||||
}}
|
||||
>
|
||||
<X className="w-3 h-3 mr-1" />
|
||||
取消
|
||||
</Button>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{status === 'waiting' && !state?.image && (
|
||||
<p className="flex items-center justify-center gap-2 py-6 text-[11px] font-mono text-cyber-text-muted">
|
||||
<Loader2 className="w-3 h-3 animate-spin" />
|
||||
正在从浏览器取二维码…
|
||||
</p>
|
||||
)}
|
||||
|
||||
{status === 'success' && (
|
||||
<div className="space-y-3">
|
||||
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-green leading-relaxed">
|
||||
<CheckCircle2 className="w-3 h-3 mt-0.5 shrink-0" />
|
||||
{state?.message || '账号已添加'}
|
||||
</p>
|
||||
{state?.account && (
|
||||
<div className="flex items-center gap-2 rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 px-3 py-2">
|
||||
{state.account.avatar && (
|
||||
<img
|
||||
src={state.account.avatar}
|
||||
alt=""
|
||||
referrerPolicy="no-referrer"
|
||||
className="w-8 h-8 rounded-full object-cover"
|
||||
/>
|
||||
)}
|
||||
<div className="min-w-0">
|
||||
<div className="truncate font-mono text-xs text-cyber-text-primary">
|
||||
{state.account.nickname}
|
||||
</div>
|
||||
<div className="text-[10px] font-mono text-cyber-text-muted">
|
||||
小红书号 {state.account.red_id || '—'}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
{state?.account?.permission_tip && (
|
||||
<p className="text-[10px] font-mono text-cyber-neon-orange leading-relaxed">
|
||||
{state.account.permission_tip}
|
||||
</p>
|
||||
)}
|
||||
<Button size="sm" onClick={() => onOpenChange(false)}>
|
||||
完成
|
||||
</Button>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{(status === 'error' || status === 'expired') && (
|
||||
<div className="space-y-2">
|
||||
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-orange leading-relaxed">
|
||||
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
|
||||
{state?.message || '获取二维码失败'}
|
||||
</p>
|
||||
<Button variant="outline" size="sm" disabled={busy} onClick={begin}>
|
||||
<RefreshCw className="w-3 h-3 mr-1" />
|
||||
重新获取
|
||||
</Button>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{status === 'idle' && (
|
||||
<div className="space-y-2">
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
需要先在「系统设置」里打开 接管已有 Chrome(CDP),
|
||||
并确保那台 Chrome 正以 9222 端口运行。
|
||||
</p>
|
||||
<Button size="sm" disabled={busy} onClick={begin}>
|
||||
{start.isPending ? (
|
||||
<Loader2 className="w-3 h-3 mr-1 animate-spin" />
|
||||
) : (
|
||||
<QrCode className="w-3 h-3 mr-1" />
|
||||
)}
|
||||
获取二维码
|
||||
</Button>
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
</DialogContent>
|
||||
</Dialog>
|
||||
)
|
||||
}
|
||||
@@ -0,0 +1,424 @@
|
||||
import { useEffect, useState } from 'react'
|
||||
import {
|
||||
AlertTriangle,
|
||||
Briefcase,
|
||||
Clock,
|
||||
Plus,
|
||||
RefreshCw,
|
||||
Trash2,
|
||||
UserRound,
|
||||
} from 'lucide-react'
|
||||
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
import { Button } from '@/components/ui/button'
|
||||
import {
|
||||
Select,
|
||||
SelectContent,
|
||||
SelectItem,
|
||||
SelectTrigger,
|
||||
SelectValue,
|
||||
} from '@/components/ui/select'
|
||||
import {
|
||||
useCheckCreatorAccount,
|
||||
useCreatorAccount,
|
||||
useCreatorAccounts,
|
||||
useDeleteCreatorAccount,
|
||||
useSyncCreatorAccount,
|
||||
} from '@/hooks/useCreator'
|
||||
import { useCurrentPlatform } from '@/hooks/usePlatform'
|
||||
import { formatCount, formatDateTime, formatRelative } from '@/lib/monitorFormat'
|
||||
import type { CreatorAccount, CreatorNote, CreatorPermissionStatus } from '@/types/creator'
|
||||
|
||||
import { AddAccountDialog } from './AddAccountDialog'
|
||||
|
||||
const PERMISSION_LABEL: Record<CreatorPermissionStatus, { text: string; tone: string }> = {
|
||||
active: { text: '数据已开通', tone: 'text-cyber-neon-green' },
|
||||
// 实测中最常见的一种,必须和"没有权限"分开说 —— 否则用户以为采集坏了,
|
||||
// 其实只是在等次日生效。
|
||||
pending: { text: '权限待生效', tone: 'text-cyber-neon-orange' },
|
||||
missing: { text: '无数据权限', tone: 'text-cyber-neon-pink' },
|
||||
unknown: { text: '权限未知', tone: 'text-cyber-text-muted' },
|
||||
}
|
||||
|
||||
// 后台按**发布时间**筛选,所以这是"作品发布距今多少天",不是"最近多少天的数据"。
|
||||
// 上限 730 天与后端校验一致。
|
||||
const RANGE_OPTIONS = [
|
||||
{ value: '7', label: '近 7 天' },
|
||||
{ value: '30', label: '近 30 天' },
|
||||
{ value: '90', label: '近 90 天' },
|
||||
{ value: '180', label: '近 180 天' },
|
||||
{ value: '365', label: '近 1 年' },
|
||||
{ value: '730', label: '近 2 年' },
|
||||
]
|
||||
|
||||
const SUMMARY_TILES: Array<{ key: keyof CreatorNote; label: string }> = [
|
||||
{ key: 'exposure', label: '曝光' },
|
||||
{ key: 'views', label: '观看' },
|
||||
{ key: 'likes', label: '点赞' },
|
||||
{ key: 'comments', label: '评论' },
|
||||
{ key: 'favorites', label: '收藏' },
|
||||
{ key: 'shares', label: '分享' },
|
||||
{ key: 'new_followers', label: '涨粉' },
|
||||
]
|
||||
|
||||
function formatRate(value: number | null): string {
|
||||
return value === null ? '—' : `${value.toFixed(1)}%`
|
||||
}
|
||||
|
||||
function formatSeconds(value: number | null): string {
|
||||
if (value === null) return '—'
|
||||
if (value < 60) return `${Math.round(value)} 秒`
|
||||
return `${Math.floor(value / 60)} 分 ${Math.round(value % 60)} 秒`
|
||||
}
|
||||
|
||||
/** 左栏里的一行账号。选中项高亮,与监控的任务卡片同一套语言。 */
|
||||
function AccountRow({
|
||||
account,
|
||||
selected,
|
||||
onSelect,
|
||||
}: {
|
||||
account: CreatorAccount
|
||||
selected: boolean
|
||||
onSelect: () => void
|
||||
}) {
|
||||
const permission = PERMISSION_LABEL[account.permission_status]
|
||||
|
||||
return (
|
||||
<button
|
||||
onClick={onSelect}
|
||||
className={`w-full flex items-center gap-2.5 rounded-lg border px-2.5 py-2 text-left transition-colors ${
|
||||
selected
|
||||
? 'border-cyber-neon-cyan/50 bg-cyber-neon-cyan/10'
|
||||
: 'border-cyber-border-subtle bg-cyber-bg-tertiary/40 hover:border-cyber-border-DEFAULT'
|
||||
}`}
|
||||
>
|
||||
{account.avatar ? (
|
||||
<img
|
||||
src={account.avatar}
|
||||
alt=""
|
||||
// 小红书图床对带外部 Referer 的请求返回 403,头像同理。
|
||||
referrerPolicy="no-referrer"
|
||||
className="w-8 h-8 rounded-full object-cover flex-shrink-0 bg-cyber-bg-tertiary"
|
||||
/>
|
||||
) : (
|
||||
<div className="w-8 h-8 rounded-full bg-cyber-bg-tertiary flex-shrink-0" />
|
||||
)}
|
||||
|
||||
<div className="min-w-0 flex-1">
|
||||
<div className="flex items-center gap-1.5">
|
||||
<span className="truncate font-mono text-[11px] text-cyber-text-primary">
|
||||
{account.nickname || '未命名账号'}
|
||||
</span>
|
||||
{account.status === 'expired' && (
|
||||
<Badge variant="warning" className="text-[9px] px-1 py-0 flex-shrink-0">
|
||||
失效
|
||||
</Badge>
|
||||
)}
|
||||
</div>
|
||||
<div className="mt-0.5 flex items-center gap-2 text-[9px] font-mono text-cyber-text-muted">
|
||||
<span className={permission.tone}>{permission.text}</span>
|
||||
<span>作品 {account.note_count}</span>
|
||||
</div>
|
||||
</div>
|
||||
</button>
|
||||
)
|
||||
}
|
||||
|
||||
/** 右栏:所选账号的数据。 */
|
||||
function AccountPanel({ accountId }: { accountId: number }) {
|
||||
const { data, isLoading } = useCreatorAccount(accountId)
|
||||
const sync = useSyncCreatorAccount()
|
||||
const check = useCheckCreatorAccount()
|
||||
const remove = useDeleteCreatorAccount()
|
||||
// 拉取范围是**同步参数**,不是展示筛选:数据按这个范围取回来存库。
|
||||
// 默认 90 天,与后端的默认值一致。
|
||||
const [days, setDays] = useState('90')
|
||||
|
||||
// 选择器对齐到上次实际用的范围。不对齐的话,上次同步了 1 年、这次点同步会悄悄
|
||||
// 缩回 90 天,而界面上没有任何迹象。
|
||||
const lastDays = data?.account.last_sync_days ?? 0
|
||||
useEffect(() => {
|
||||
if (lastDays > 0) setDays(String(lastDays))
|
||||
}, [lastDays])
|
||||
|
||||
if (isLoading || !data) {
|
||||
return <p className="py-10 text-center text-[11px] font-mono text-cyber-text-muted">加载中…</p>
|
||||
}
|
||||
|
||||
const { account, notes, summary } = data
|
||||
const permission = PERMISSION_LABEL[account.permission_status]
|
||||
|
||||
return (
|
||||
<div className="flex flex-col overflow-hidden h-full">
|
||||
<div className="flex items-center gap-2 px-3 py-2.5 border-b border-cyber-border-subtle flex-shrink-0">
|
||||
{account.avatar && (
|
||||
<img
|
||||
src={account.avatar}
|
||||
alt=""
|
||||
referrerPolicy="no-referrer"
|
||||
className="w-7 h-7 rounded-full object-cover"
|
||||
/>
|
||||
)}
|
||||
<span className="font-mono text-xs text-cyber-text-primary">
|
||||
{account.nickname || '未命名账号'}
|
||||
</span>
|
||||
<span className="font-mono text-[10px] text-cyber-text-muted">
|
||||
小红书号 {account.red_id || '—'}
|
||||
</span>
|
||||
<span className="ml-1 text-[10px] font-mono text-cyber-text-muted">
|
||||
上次同步 {account.last_synced_at ? formatRelative(account.last_synced_at) : '从未'}
|
||||
{/* 当前列出的作品就是上一次同步范围里的产物,所以要说清是哪一段 ——
|
||||
否则只看"同步过了"分不清覆盖了多少。 */}
|
||||
{account.last_sync_days > 0 && ` · 覆盖近 ${account.last_sync_days} 天发布的`}
|
||||
</span>
|
||||
|
||||
<div className="ml-auto flex items-center gap-2">
|
||||
<Button
|
||||
variant="outline"
|
||||
size="sm"
|
||||
disabled={check.isPending}
|
||||
onClick={() => check.mutate(account.id)}
|
||||
>
|
||||
<RefreshCw className={`w-3 h-3 mr-1 ${check.isPending ? 'animate-spin' : ''}`} />
|
||||
检测
|
||||
</Button>
|
||||
<Select value={days} onValueChange={setDays}>
|
||||
<SelectTrigger className="h-7 w-[104px] text-[10px] font-mono">
|
||||
<SelectValue />
|
||||
</SelectTrigger>
|
||||
<SelectContent>
|
||||
{RANGE_OPTIONS.map((option) => (
|
||||
<SelectItem key={option.value} value={option.value} className="text-[11px]">
|
||||
{option.label}
|
||||
</SelectItem>
|
||||
))}
|
||||
</SelectContent>
|
||||
</Select>
|
||||
<Button
|
||||
size="sm"
|
||||
disabled={sync.isPending}
|
||||
onClick={() => sync.mutate({ id: account.id, days: Number(days) })}
|
||||
>
|
||||
<RefreshCw className={`w-3 h-3 mr-1 ${sync.isPending ? 'animate-spin' : ''}`} />
|
||||
同步数据
|
||||
</Button>
|
||||
<Button variant="ghost" size="sm" onClick={() => remove.mutate(account.id)}>
|
||||
<Trash2 className="w-3 h-3" />
|
||||
</Button>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="flex-1 overflow-y-auto terminal-scroll px-3 py-3 space-y-3">
|
||||
{/* 只要后台有话要说就显示,**包括已经开通的时候**。
|
||||
实测开通后的原话是「数据正在更新中,请耐心等待」—— 若只在未开通时才显示,
|
||||
就会一边写着"数据已开通"、一边列不出作品,自相矛盾。 */}
|
||||
{account.permission_tip && (
|
||||
<div
|
||||
className={`flex items-start gap-2 rounded-md border px-3 py-2 ${
|
||||
account.permission_status === 'active'
|
||||
? 'border-cyber-border-DEFAULT bg-cyber-bg-tertiary/40'
|
||||
: 'border-cyber-neon-orange/30 bg-cyber-neon-orange/5'
|
||||
}`}
|
||||
>
|
||||
<Clock
|
||||
className={`w-3.5 h-3.5 mt-0.5 flex-shrink-0 ${
|
||||
account.permission_status === 'active'
|
||||
? 'text-cyber-text-muted'
|
||||
: 'text-cyber-neon-orange'
|
||||
}`}
|
||||
/>
|
||||
<div className="text-[11px] font-mono leading-relaxed">
|
||||
<span className={permission.tone}>{permission.text}</span>
|
||||
<span className="ml-2 text-cyber-text-secondary">{account.permission_tip}</span>
|
||||
<div className="mt-0.5 text-cyber-text-muted">
|
||||
{account.permission_status === 'active'
|
||||
? '权限已开通,但后台的数据可能还在准备中 —— 稍后重新同步即可。'
|
||||
: '数据权限由创作者后台按天开通。在此之前同步会成功但返回 0 条 —— 那是正常的,不是采集失败。'}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{account.last_error && (
|
||||
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-pink leading-relaxed">
|
||||
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
|
||||
{account.last_error}
|
||||
</p>
|
||||
)}
|
||||
|
||||
<div className="grid grid-cols-4 gap-2 xl:grid-cols-7">
|
||||
{SUMMARY_TILES.map((tile) => (
|
||||
<div
|
||||
key={tile.key}
|
||||
className="rounded-lg border border-cyber-border-subtle bg-cyber-bg-tertiary/40 px-3 py-2"
|
||||
>
|
||||
<div className="text-[10px] font-mono text-cyber-text-muted">{tile.label}</div>
|
||||
<div className="mt-0.5 font-mono text-sm text-cyber-text-primary">
|
||||
{formatCount(summary[tile.key as keyof typeof summary] ?? 0)}
|
||||
</div>
|
||||
</div>
|
||||
))}
|
||||
</div>
|
||||
|
||||
{notes.length === 0 ? (
|
||||
<p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">
|
||||
{account.permission_status === 'active'
|
||||
? '这个账号在所选时间范围内没有作品,或还没同步过 —— 点「同步数据」试试'
|
||||
: '数据权限生效后即可看到作品数据'}
|
||||
</p>
|
||||
) : (
|
||||
<div className="overflow-x-auto">
|
||||
<table className="w-full text-xs font-mono">
|
||||
<thead className="sticky top-0 bg-cyber-bg-tertiary">
|
||||
<tr className="text-[10px] text-cyber-text-secondary">
|
||||
<th className="px-2 py-2 text-left font-normal">作品</th>
|
||||
{SUMMARY_TILES.map((tile) => (
|
||||
<th key={tile.key} className="px-2 py-2 text-right font-normal">
|
||||
{tile.label}
|
||||
</th>
|
||||
))}
|
||||
<th className="px-2 py-2 text-right font-normal">点击率</th>
|
||||
<th className="px-2 py-2 text-right font-normal">均看时长</th>
|
||||
<th className="px-2 py-2 text-right font-normal">完播率</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
{notes.map((note) => (
|
||||
<tr key={note.note_id} className="border-t border-cyber-border-subtle">
|
||||
<td className="max-w-[280px] px-2 py-2">
|
||||
<div className="truncate text-cyber-text-primary" title={note.title}>
|
||||
{note.title || note.note_id}
|
||||
</div>
|
||||
<div className="text-[9px] text-cyber-text-muted">
|
||||
{note.publish_time ? formatDateTime(note.publish_time) : '—'}
|
||||
</div>
|
||||
</td>
|
||||
{SUMMARY_TILES.map((tile) => (
|
||||
<td key={tile.key} className="px-2 py-2 text-right text-cyber-text-primary">
|
||||
{formatCount(note[tile.key] as number | null)}
|
||||
</td>
|
||||
))}
|
||||
<td className="px-2 py-2 text-right text-cyber-text-secondary">
|
||||
{formatRate(note.cover_ctr)}
|
||||
</td>
|
||||
<td className="px-2 py-2 text-right text-cyber-text-secondary">
|
||||
{formatSeconds(note.avg_watch_seconds)}
|
||||
</td>
|
||||
<td className="px-2 py-2 text-right text-cyber-text-secondary">
|
||||
{formatRate(note.completion_rate)}
|
||||
</td>
|
||||
</tr>
|
||||
))}
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* 运营:管理自己的小红书账号。
|
||||
*
|
||||
* 与「监控」并列而非其子视图,布局也照搬它:左边账号列表、右边数据面板。
|
||||
* 两者关注的东西不同 —— 监控抓公开数据(点赞/收藏/评论/分享)的每轮差分,
|
||||
* 这里是创作者后台的运营指标(曝光/观看/完播率/涨粉),凭据与采集方式都不同。
|
||||
*/
|
||||
export function OperationView() {
|
||||
const { platform, capability } = useCurrentPlatform()
|
||||
const [selectedId, setSelectedId] = useState<number | null>(null)
|
||||
const [addOpen, setAddOpen] = useState(false)
|
||||
// 有同步在后端跑时列表要自己刷新,否则状态永远是旧的。
|
||||
const { data: accounts, isLoading } = useCreatorAccounts(true)
|
||||
|
||||
// 运营只对小红书成立:它读的是小红书创作者后台。不挡住的话,切到抖音时这里照样
|
||||
// 列出小红书的账号 —— 看起来像是"抖音账号",其实平台维度根本没参与。
|
||||
if (platform !== 'xhs') {
|
||||
const label = capability?.label ?? platform
|
||||
return (
|
||||
<div className="flex-1 flex items-center justify-center overflow-y-auto terminal-scroll">
|
||||
<div className="max-w-lg w-full mx-4 rounded-lg glass-panel float-panel p-6 space-y-3">
|
||||
<div className="flex items-center gap-3">
|
||||
<div className="p-2 rounded-md border border-cyber-neon-orange/40 bg-cyber-neon-orange/10">
|
||||
<Briefcase className="w-5 h-5 text-cyber-neon-orange" />
|
||||
</div>
|
||||
<div>
|
||||
<h2 className="font-mono text-sm text-cyber-text-primary">{label} 没有运营模块</h2>
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted">
|
||||
运营读的是创作者后台的数据,目前只接入了小红书
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
小红书之外,这里没有可用的后台接口 —— 能采的是公开数据,那属于「监控」。
|
||||
切回小红书即可看到运营账号。
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
// 没选过就默认第一个,和监控一样 —— 右栏不该一开始是空的。
|
||||
const selected =
|
||||
accounts?.find((account) => account.id === selectedId) ??
|
||||
(accounts && accounts.length > 0 ? accounts[0] : null)
|
||||
|
||||
return (
|
||||
<div className="flex-1 flex flex-col gap-3 overflow-hidden min-h-0 relative z-10">
|
||||
<div className="flex-1 flex gap-3 overflow-hidden min-h-0">
|
||||
{/* 左:账号列表 */}
|
||||
<div className="w-[300px] flex-shrink-0 flex flex-col gap-3 overflow-hidden">
|
||||
<div className="flex items-center justify-between flex-shrink-0">
|
||||
<span className="font-mono text-xs text-cyber-text-primary">运营账号</span>
|
||||
<Button size="sm" onClick={() => setAddOpen(true)}>
|
||||
<Plus className="w-3 h-3 mr-1" />
|
||||
新增
|
||||
</Button>
|
||||
</div>
|
||||
|
||||
<div className="flex-1 overflow-y-auto terminal-scroll space-y-2 pr-1">
|
||||
{isLoading ? (
|
||||
<p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">
|
||||
加载中…
|
||||
</p>
|
||||
) : !accounts || accounts.length === 0 ? (
|
||||
<div className="flex flex-col items-center gap-2 py-10">
|
||||
<UserRound className="w-7 h-7 text-cyber-text-muted" />
|
||||
<p className="text-center text-[11px] font-mono text-cyber-text-muted">
|
||||
还没有运营账号
|
||||
<br />
|
||||
扫码添加一个
|
||||
</p>
|
||||
</div>
|
||||
) : (
|
||||
accounts.map((account) => (
|
||||
<AccountRow
|
||||
key={account.id}
|
||||
account={account}
|
||||
selected={selected?.id === account.id}
|
||||
onSelect={() => setSelectedId(account.id)}
|
||||
/>
|
||||
))
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{/* 右:所选账号的数据 */}
|
||||
<div className="flex-1 flex flex-col overflow-hidden min-w-0 rounded-lg glass-panel float-panel">
|
||||
{selected ? (
|
||||
<AccountPanel accountId={selected.id} />
|
||||
) : (
|
||||
<div className="flex-1 flex items-center justify-center">
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted">
|
||||
从左边选一个账号,或先扫码添加
|
||||
</p>
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<AddAccountDialog open={addOpen} onOpenChange={setAddOpen} />
|
||||
</div>
|
||||
)
|
||||
}
|
||||
@@ -1,5 +1,5 @@
|
||||
import { useState } from 'react'
|
||||
import { Bug, Wifi, BarChart3, Cog, LogOut, Radar, Settings, Terminal } from 'lucide-react'
|
||||
import { Briefcase, Bug, Wifi, BarChart3, Cog, LogOut, Radar, Settings, Terminal } from 'lucide-react'
|
||||
import { useTranslation } from 'react-i18next'
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
import { SystemSettingsDialog } from '@/components/settings/SystemSettingsDialog'
|
||||
@@ -19,6 +19,7 @@ interface SidebarProps {
|
||||
const NAV_ITEMS: Array<{ value: AppView; label: string; icon: typeof Terminal }> = [
|
||||
{ value: 'crawler', label: '采集', icon: Terminal },
|
||||
{ value: 'monitor', label: '监控', icon: Radar },
|
||||
{ value: 'operation', label: '运营', icon: Briefcase },
|
||||
{ value: 'report', label: '报表', icon: BarChart3 },
|
||||
{ value: 'settings', label: '设置', icon: Settings },
|
||||
]
|
||||
|
||||
@@ -40,7 +40,7 @@ export function UnwiredPlatformNotice({ area }: { area: string }) {
|
||||
{capability.label} 的{area}尚未接通
|
||||
</h2>
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted">
|
||||
{capability.label}的爬虫模块是支持的,但监控层目前只接通了小红书
|
||||
{capability.label}的爬虫模块是支持的,但监控层目前只接通了小红书与抖音
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
@@ -94,9 +94,9 @@ export function UnwiredPlatformNotice({ area }: { area: string }) {
|
||||
<div className="flex items-start gap-2 text-[10px] font-mono text-cyber-text-muted">
|
||||
<MonitorSmartphone className="w-3.5 h-3.5 mt-0.5 flex-shrink-0" />
|
||||
<p>
|
||||
切换到小红书即可正常使用。若要接通该平台,需要在
|
||||
<span className="text-cyber-neon-cyan"> runner / ingest / 目标解析 </span>
|
||||
三处补上平台适配(目前这三处是硬编码小红书的)。
|
||||
切换到已接通的平台即可正常使用。若要接通该平台,需要在
|
||||
<span className="text-cyber-neon-cyan"> adapters.py </span>
|
||||
里补一份适配:产物目录名、jsonl 字段名、目标链接形态。
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
@@ -7,6 +7,7 @@ import {
|
||||
LayoutList,
|
||||
Rows3,
|
||||
ThumbsUp,
|
||||
Users,
|
||||
} from 'lucide-react'
|
||||
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
@@ -24,9 +25,11 @@ import {
|
||||
useMonitorCommentsGrouped,
|
||||
} from '@/hooks/useMonitor'
|
||||
import { monitorApi } from '@/lib/api'
|
||||
import { formatRelative } from '@/lib/monitorFormat'
|
||||
import { formatDate, formatRelative } from '@/lib/monitorFormat'
|
||||
import type { CommentBucket, MonitorComment } from '@/types/monitor'
|
||||
|
||||
import { NoteCover } from './NoteCover'
|
||||
|
||||
interface CommentsFeedProps {
|
||||
taskId: number | null
|
||||
}
|
||||
@@ -49,22 +52,7 @@ function NoteBadge({
|
||||
}) {
|
||||
return (
|
||||
<div className="flex items-center gap-2 min-w-0">
|
||||
{cover ? (
|
||||
<img
|
||||
src={cover}
|
||||
alt=""
|
||||
loading="lazy"
|
||||
className={`rounded object-cover bg-cyber-bg-tertiary flex-shrink-0 ${
|
||||
compact ? 'w-6 h-8' : 'w-8 h-10'
|
||||
}`}
|
||||
/>
|
||||
) : (
|
||||
<div
|
||||
className={`rounded bg-cyber-bg-tertiary flex-shrink-0 ${
|
||||
compact ? 'w-6 h-8' : 'w-8 h-10'
|
||||
}`}
|
||||
/>
|
||||
)}
|
||||
<NoteCover src={cover} size={compact ? 'sm' : 'md'} />
|
||||
<span className="truncate text-[10px] font-mono text-cyber-text-secondary" title={title || noteId}>
|
||||
{title || noteId}
|
||||
</span>
|
||||
@@ -149,6 +137,15 @@ function CollapsibleGroup({
|
||||
noteId={bucket.note_id}
|
||||
/>
|
||||
</div>
|
||||
{/* 作品那一层光有标题不够 —— 同名的作品不少,发布日期能帮着认。 */}
|
||||
{bucket.published_at && (
|
||||
<span
|
||||
className="text-[10px] font-mono text-cyber-text-muted flex-shrink-0"
|
||||
title="作品发布时间"
|
||||
>
|
||||
{formatDate(bucket.published_at)}
|
||||
</span>
|
||||
)}
|
||||
<Badge variant="outline" className="text-[10px] px-1.5 py-0 flex-shrink-0">
|
||||
{bucket.comments.length} 条
|
||||
</Badge>
|
||||
@@ -169,6 +166,78 @@ function CollapsibleGroup({
|
||||
)
|
||||
}
|
||||
|
||||
/** 把作品分组按博主归并 —— 三级视图的第一级。 */
|
||||
function groupByCreator(buckets: CommentBucket[]) {
|
||||
const byCreator = new Map<string, { key: string; name: string; buckets: CommentBucket[] }>()
|
||||
for (const bucket of buckets) {
|
||||
const key = bucket.creator_hash || '__unknown__'
|
||||
if (!byCreator.has(key)) {
|
||||
byCreator.set(key, { key, name: bucket.creator_name || '', buckets: [] })
|
||||
}
|
||||
byCreator.get(key)!.buckets.push(bucket)
|
||||
}
|
||||
return [...byCreator.values()]
|
||||
}
|
||||
|
||||
/**
|
||||
* 一级:博主。二级作品、三级评论都收在它下面。
|
||||
*
|
||||
* 分组键是 `creator_hash`(爬虫刻意不落原始 user_id,这是唯一稳定的创作者标识),
|
||||
* 显示名用已脱敏的昵称。默认展开 —— 折叠的默认值不该把数据藏起来。
|
||||
*/
|
||||
function CreatorGroup({
|
||||
name,
|
||||
buckets,
|
||||
expandedNotes,
|
||||
onToggleNote,
|
||||
}: {
|
||||
name: string
|
||||
buckets: CommentBucket[]
|
||||
expandedNotes: Set<string>
|
||||
onToggleNote: (noteId: string) => void
|
||||
}) {
|
||||
const [open, setOpen] = useState(true)
|
||||
const commentCount = buckets.reduce((sum, bucket) => sum + bucket.comments.length, 0)
|
||||
|
||||
return (
|
||||
<div className="rounded-lg border border-cyber-border-DEFAULT overflow-hidden">
|
||||
<button
|
||||
onClick={() => setOpen(!open)}
|
||||
className="w-full flex items-center gap-2 px-3 py-2 bg-cyber-bg-elevated/50 hover:bg-cyber-bg-elevated/80 transition-colors text-left"
|
||||
>
|
||||
{open ? (
|
||||
<ChevronDown className="w-3.5 h-3.5 text-cyber-text-muted flex-shrink-0" />
|
||||
) : (
|
||||
<ChevronRight className="w-3.5 h-3.5 text-cyber-text-muted flex-shrink-0" />
|
||||
)}
|
||||
<Users className="w-3.5 h-3.5 text-cyber-neon-cyan flex-shrink-0" />
|
||||
<span className="font-mono text-xs text-cyber-text-primary truncate">
|
||||
{name || '未知博主'}
|
||||
</span>
|
||||
<span className="text-[10px] font-mono text-cyber-text-muted flex-shrink-0">
|
||||
{buckets.length} 篇作品
|
||||
</span>
|
||||
<Badge variant="outline" className="text-[10px] px-1.5 py-0 flex-shrink-0 ml-auto">
|
||||
{commentCount} 条评论
|
||||
</Badge>
|
||||
</button>
|
||||
|
||||
{open && (
|
||||
<div className="p-2 space-y-2 bg-cyber-bg-secondary/30">
|
||||
{buckets.map((bucket) => (
|
||||
<CollapsibleGroup
|
||||
key={bucket.note_id}
|
||||
bucket={bucket}
|
||||
open={expandedNotes.has(bucket.note_id)}
|
||||
onToggle={() => onToggleNote(bucket.note_id)}
|
||||
/>
|
||||
))}
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
export function CommentsFeed({ taskId }: CommentsFeedProps) {
|
||||
const [grouped, setGrouped] = useState(true)
|
||||
const [noteFilter, setNoteFilter] = useState<string>(ALL_NOTES)
|
||||
@@ -280,12 +349,15 @@ export function CommentsFeed({ taskId }: CommentsFeedProps) {
|
||||
</p>
|
||||
) : grouped ? (
|
||||
<div className="space-y-2">
|
||||
{groups?.map((bucket) => (
|
||||
<CollapsibleGroup
|
||||
key={bucket.note_id}
|
||||
bucket={bucket}
|
||||
open={expanded.has(bucket.note_id)}
|
||||
onToggle={() => toggleGroup(bucket.note_id)}
|
||||
{/* 三级:博主 -> 作品 -> 评论。一个任务可以配多个博主,平铺作品的话
|
||||
看不出哪条评论属于谁。 */}
|
||||
{groupByCreator(groups ?? []).map((creator) => (
|
||||
<CreatorGroup
|
||||
key={creator.key}
|
||||
name={creator.name}
|
||||
buckets={creator.buckets}
|
||||
expandedNotes={expanded}
|
||||
onToggleNote={toggleGroup}
|
||||
/>
|
||||
))}
|
||||
</div>
|
||||
|
||||
@@ -31,8 +31,11 @@ export function CookiePanel() {
|
||||
<div className="flex items-center justify-between gap-3">
|
||||
<div className="flex items-center gap-2">
|
||||
<KeyRound className="w-4 h-4 text-cyber-neon-cyan" />
|
||||
{/* 标题要和扫码面板区分开:这个是"存进库里、每轮注入子进程"的 cookie,
|
||||
那个是"写进浏览器 profile、CDP 模式复用"的登录态。两者是不同机制、
|
||||
也都能同时算"登录着",都叫「登录态」就分不清在说哪一个。 */}
|
||||
<span className="font-mono text-xs tracking-wider text-cyber-text-primary">
|
||||
{label}登录态
|
||||
{label} Cookie(定时任务用)
|
||||
</span>
|
||||
{status?.present ? (
|
||||
<Badge variant="success" className="text-[10px]">
|
||||
|
||||
@@ -1,10 +1,11 @@
|
||||
import { useState } from 'react'
|
||||
import { Activity, BellRing, FileText, MessageSquare, Plus } from 'lucide-react'
|
||||
import { Activity, BellRing, Download, FileText, MessageSquare, Plus } from 'lucide-react'
|
||||
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
import { Button } from '@/components/ui/button'
|
||||
import { Tabs, TabsContent, TabsList, TabsTrigger } from '@/components/ui/tabs'
|
||||
import { useMonitorOverview, useMonitorTasks } from '@/hooks/useMonitor'
|
||||
import { monitorApi } from '@/lib/api'
|
||||
import type { MonitorTask } from '@/types/monitor'
|
||||
import { useCookieStatus, useWebhookStatus } from '@/hooks/useMonitor'
|
||||
import { useCurrentPlatform } from '@/hooks/usePlatform'
|
||||
@@ -166,6 +167,23 @@ export function MonitorDashboard() {
|
||||
>
|
||||
只看新增
|
||||
</Button>
|
||||
{/* 导出的是**这个任务**的全部作品,不是屏幕上这 200 条 —— 屏幕上那份是
|
||||
为了好看才截断的,导出跟着截断就成了「导出来的比看到的少」。
|
||||
走 window.open 而不是 blob:鉴权在 Cookie 上,浏览器自己会带上。 */}
|
||||
<Button
|
||||
variant="outline"
|
||||
size="sm"
|
||||
className="ml-auto"
|
||||
onClick={() =>
|
||||
window.open(
|
||||
monitorApi.getExportUrl({ kind: 'notes', taskId: selectedTask.id }),
|
||||
'_blank',
|
||||
)
|
||||
}
|
||||
>
|
||||
<Download className="w-3 h-3 mr-1" />
|
||||
导出 CSV
|
||||
</Button>
|
||||
</div>
|
||||
<NotesTable taskId={selectedTask.id} onlyNew={onlyNew} />
|
||||
</TabsContent>
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
/**
|
||||
* 作品封面缩略图。
|
||||
*
|
||||
* **图为什么曾经全是破图**:不是防盗链。小红书图床的地址**带签名、会过期** ——
|
||||
* 路径里那段时间戳就是签发时刻。实测同一批图:
|
||||
*
|
||||
* 当天签发的地址 -> 200,带不带 Referer 都一样
|
||||
* 隔天的地址 -> 403,带不带 Referer 都一样
|
||||
*
|
||||
* 所以 Referer 根本不是那个维度,改它是白改。真正的解法是后端把图**下载到本地**,
|
||||
* 前端拿到的是 `/api/monitor/covers/{note_id}` —— 与签名无关,不会过期。
|
||||
*
|
||||
* `referrerPolicy` 保留着,是因为它仍然是个合理的默认(外链图不该把自己的地址
|
||||
* 泄露给第三方),但**它不再是这里能正常显示的原因**。
|
||||
*/
|
||||
|
||||
const SIZES = {
|
||||
sm: 'w-6 h-8',
|
||||
md: 'w-8 h-10',
|
||||
lg: 'w-12 h-16',
|
||||
} as const
|
||||
|
||||
export function NoteCover({
|
||||
src,
|
||||
size = 'md',
|
||||
className = '',
|
||||
}: {
|
||||
src?: string
|
||||
size?: keyof typeof SIZES
|
||||
className?: string
|
||||
}) {
|
||||
const base = `rounded object-cover bg-cyber-bg-tertiary flex-shrink-0 ${SIZES[size]} ${className}`
|
||||
|
||||
// 占位块是必要的:作品在首轮采集前没有封面,若此时不占位,行高会随着封面陆续
|
||||
// 到达而跳动。
|
||||
if (!src) return <div className={base} aria-hidden />
|
||||
|
||||
return (
|
||||
<img src={src} alt="" loading="lazy" referrerPolicy="no-referrer" className={base} />
|
||||
)
|
||||
}
|
||||
@@ -1,4 +1,4 @@
|
||||
import { useMemo, useState } from 'react'
|
||||
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
|
||||
import { useNoteSeries } from '@/hooks/useMonitor'
|
||||
import { formatCount, formatDateTime } from '@/lib/monitorFormat'
|
||||
@@ -12,13 +12,62 @@ const METRICS: Array<{ key: MetricKey; label: string }> = [
|
||||
{ key: 'share_count', label: '分享' },
|
||||
]
|
||||
|
||||
const VIEW_W = 600
|
||||
/** 图表高度固定;宽度由容器量出来,见下方 attachContainer。 */
|
||||
const VIEW_H = 160
|
||||
const PAD_LEFT = 10
|
||||
const PAD_RIGHT = 56
|
||||
const PAD_TOP = 18
|
||||
const PAD_LEFT = 44 // 左侧刻度栏,要放得下 "1.2万"
|
||||
const PAD_RIGHT = 16
|
||||
const PAD_TOP = 14
|
||||
const PAD_BOTTOM = 22
|
||||
|
||||
const TICK_COUNT = 3
|
||||
|
||||
/** 把步长收敛到 1 / 2 / 5 × 10ⁿ,刻度才会落在好读的数上。 */
|
||||
function niceNumber(value: number, round: boolean): number {
|
||||
if (value <= 0) return 1
|
||||
const exponent = Math.floor(Math.log10(value))
|
||||
const fraction = value / 10 ** exponent
|
||||
let nice: number
|
||||
if (round) {
|
||||
nice = fraction < 1.5 ? 1 : fraction < 3 ? 2 : fraction < 7 ? 5 : 10
|
||||
} else {
|
||||
nice = fraction <= 1 ? 1 : fraction <= 2 ? 2 : fraction <= 5 ? 5 : 10
|
||||
}
|
||||
return nice * 10 ** exponent
|
||||
}
|
||||
|
||||
/**
|
||||
* 给折线图选一个纵轴范围。
|
||||
*
|
||||
* **刻意不从 0 起。** 柱状图用「长度」编码数值,基线不为 0 比例就是错的;折线图用
|
||||
* 「位置」编码,轴只需要如实框住数据 —— 这正是让 491 → 506 这段变化看得见的原因,
|
||||
* 否则它会贴着 0–600 的底边变成一条直线。
|
||||
*
|
||||
* 代价是纵轴不再是 0,所以刻度必须落在左侧栏里、显示真实数值,让读图的人随时知道
|
||||
* 范围是多少。这一点在上一轮已经修好。
|
||||
*/
|
||||
function niceAxis(values: number[]): { min: number; max: number; ticks: number[] } {
|
||||
const dataMin = Math.min(...values)
|
||||
const dataMax = Math.max(...values)
|
||||
|
||||
// 全平的一组数没有跨度可缩放,给它一个名义区间,让线落在图中而不是贴边。
|
||||
const span = dataMax - dataMin || Math.max(1, Math.abs(dataMax) * 0.05 || 1)
|
||||
|
||||
const step = niceNumber(span / (TICK_COUNT - 1), true)
|
||||
let min = Math.floor(dataMin / step) * step
|
||||
let max = Math.ceil(dataMax / step) * step
|
||||
if (min === max) {
|
||||
// 取整后塌成一点会除零,撑开一档。
|
||||
min -= step
|
||||
max += step
|
||||
}
|
||||
|
||||
const ticks: number[] = []
|
||||
for (let value = min; value <= max + step / 2; value += step) {
|
||||
ticks.push(Math.round(value))
|
||||
}
|
||||
return { min, max, ticks }
|
||||
}
|
||||
|
||||
interface NoteTrendChartProps {
|
||||
noteId: string
|
||||
taskId: number | null
|
||||
@@ -36,8 +85,25 @@ interface NoteTrendChartProps {
|
||||
export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProps) {
|
||||
const [metric, setMetric] = useState<MetricKey>('liked_count')
|
||||
const [hoverIndex, setHoverIndex] = useState<number | null>(null)
|
||||
const [width, setWidth] = useState(0)
|
||||
const observer = useRef<ResizeObserver | null>(null)
|
||||
const { data: series, isLoading } = useNoteSeries(noteId, taskId)
|
||||
|
||||
// 宽度是量出来的,不是拉伸出来的。原先靠 preserveAspectRatio="none" 把 600×160 的
|
||||
// viewBox 横向撑满容器 —— 横纵缩放不一致,半径 4 的圆就被压成了椭圆。只有等比绘图,
|
||||
// 圆才真的是圆。
|
||||
const attachContainer = useCallback((node: HTMLDivElement | null) => {
|
||||
observer.current?.disconnect()
|
||||
if (!node) return
|
||||
const measure = () => setWidth(Math.round(node.getBoundingClientRect().width))
|
||||
const resizeObserver = new ResizeObserver(measure)
|
||||
resizeObserver.observe(node)
|
||||
observer.current = resizeObserver
|
||||
measure()
|
||||
}, [])
|
||||
|
||||
useEffect(() => () => observer.current?.disconnect(), [])
|
||||
|
||||
// A point with no parsed value is a gap, not a zero.
|
||||
const points = useMemo(
|
||||
() => (series ?? []).filter((point) => point[metric] !== null),
|
||||
@@ -45,40 +111,59 @@ export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProp
|
||||
)
|
||||
|
||||
const geometry = useMemo(() => {
|
||||
if (points.length < 2) return null
|
||||
if (points.length < 2 || width <= 0) return null
|
||||
|
||||
const values = points.map((point) => point[metric] as number)
|
||||
const min = Math.min(...values)
|
||||
const max = Math.max(...values)
|
||||
// A flat series would divide by zero; give it a nominal band.
|
||||
const span = max - min || 1
|
||||
const { min, max, ticks } = niceAxis(values)
|
||||
|
||||
const innerW = VIEW_W - PAD_LEFT - PAD_RIGHT
|
||||
const innerW = Math.max(1, width - PAD_LEFT - PAD_RIGHT)
|
||||
const innerH = VIEW_H - PAD_TOP - PAD_BOTTOM
|
||||
const yFor = (value: number) =>
|
||||
PAD_TOP + innerH * (1 - (value - min) / (max - min))
|
||||
|
||||
const xy = points.map((point, index) => {
|
||||
const value = point[metric] as number
|
||||
return {
|
||||
x: PAD_LEFT + (index / (points.length - 1)) * innerW,
|
||||
y: PAD_TOP + innerH - ((value - min) / span) * innerH,
|
||||
y: yFor(value),
|
||||
value,
|
||||
point,
|
||||
}
|
||||
})
|
||||
|
||||
return { xy, min, max }
|
||||
}, [points, metric])
|
||||
return { xy, ticks, yFor, innerW, innerH }
|
||||
}, [points, metric, width])
|
||||
|
||||
if (isLoading) {
|
||||
return <div className="p-4 text-[11px] font-mono text-cyber-text-muted">加载中…</div>
|
||||
}
|
||||
|
||||
const metricLabel = METRICS.find((option) => option.key === metric)?.label ?? ''
|
||||
const latest = points.length > 0 ? (points[points.length - 1][metric] as number) : null
|
||||
|
||||
// Fewer points get a bigger dot; a dense series would otherwise turn into a
|
||||
// string of overlapping beads.
|
||||
const dotRadius = points.length > 12 ? 2.5 : 4
|
||||
const hitWidth = geometry ? Math.max(12, geometry.innerW / Math.max(1, points.length - 1)) : 12
|
||||
|
||||
return (
|
||||
<div className="space-y-2">
|
||||
<div className="flex items-center justify-between gap-3 flex-wrap">
|
||||
<span className="font-mono text-xs text-cyber-text-primary">
|
||||
指标趋势 · <span className="text-cyber-text-secondary">{noteTitle || noteId}</span>
|
||||
</span>
|
||||
<div className="flex items-baseline gap-2 min-w-0">
|
||||
<span className="font-mono text-xs text-cyber-text-primary flex-shrink-0">指标趋势</span>
|
||||
<span
|
||||
className="font-mono text-[11px] text-cyber-text-secondary truncate"
|
||||
title={noteTitle || noteId}
|
||||
>
|
||||
{noteTitle || noteId}
|
||||
</span>
|
||||
{latest !== null && (
|
||||
<span className="font-mono text-xs text-cyber-neon-cyan flex-shrink-0">
|
||||
{formatCount(latest)}
|
||||
<span className="text-cyber-text-muted text-[10px] ml-1">{metricLabel}</span>
|
||||
</span>
|
||||
)}
|
||||
</div>
|
||||
<div className="flex items-center gap-1">
|
||||
{METRICS.map((option) => (
|
||||
<button
|
||||
@@ -96,97 +181,101 @@ export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProp
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{!geometry ? (
|
||||
{points.length < 2 ? (
|
||||
<p className="py-6 text-center text-[11px] font-mono text-cyber-text-muted">
|
||||
至少需要两轮采集才能画出趋势(当前 {points.length} 个有效数据点)
|
||||
</p>
|
||||
) : (
|
||||
<div className="relative">
|
||||
<svg
|
||||
viewBox={`0 0 ${VIEW_W} ${VIEW_H}`}
|
||||
className="w-full h-40"
|
||||
preserveAspectRatio="none"
|
||||
onMouseLeave={() => setHoverIndex(null)}
|
||||
>
|
||||
{/* Recessive solid hairlines - never dashed. */}
|
||||
{[0, 0.5, 1].map((ratio) => {
|
||||
const y = PAD_TOP + (VIEW_H - PAD_TOP - PAD_BOTTOM) * ratio
|
||||
return (
|
||||
<div className="relative" ref={attachContainer}>
|
||||
{geometry && (
|
||||
<svg
|
||||
width={width}
|
||||
height={VIEW_H}
|
||||
viewBox={`0 0 ${width} ${VIEW_H}`}
|
||||
className="block"
|
||||
onMouseLeave={() => setHoverIndex(null)}
|
||||
>
|
||||
{/* 刻度线画在 0 / 中值 / 上界上,标签就贴在对应位置 —— 不再浮在图面上,
|
||||
也就不会出现两个数挤在一起读成一个数的情况。 */}
|
||||
{geometry.ticks.map((tick) => {
|
||||
const y = geometry.yFor(tick)
|
||||
return (
|
||||
<g key={tick}>
|
||||
<line
|
||||
x1={PAD_LEFT}
|
||||
x2={width - PAD_RIGHT}
|
||||
y1={y}
|
||||
y2={y}
|
||||
stroke="rgb(var(--cyber-text-muted) / 0.25)"
|
||||
strokeWidth={1}
|
||||
/>
|
||||
<text
|
||||
x={PAD_LEFT - 8}
|
||||
y={y}
|
||||
textAnchor="end"
|
||||
dominantBaseline="middle"
|
||||
fontSize={9}
|
||||
fill="rgb(var(--cyber-text-muted))"
|
||||
style={{ fontFamily: 'ui-monospace, SFMono-Regular, monospace' }}
|
||||
>
|
||||
{formatCount(tick)}
|
||||
</text>
|
||||
</g>
|
||||
)
|
||||
})}
|
||||
|
||||
{hoverIndex !== null && geometry.xy[hoverIndex] && (
|
||||
<line
|
||||
key={ratio}
|
||||
x1={PAD_LEFT}
|
||||
x2={VIEW_W - PAD_RIGHT}
|
||||
y1={y}
|
||||
y2={y}
|
||||
stroke="rgb(var(--cyber-text-muted) / 0.25)"
|
||||
x1={geometry.xy[hoverIndex].x}
|
||||
x2={geometry.xy[hoverIndex].x}
|
||||
y1={PAD_TOP}
|
||||
y2={PAD_TOP + geometry.innerH}
|
||||
stroke="rgb(var(--cyber-neon-cyan) / 0.5)"
|
||||
strokeWidth={1}
|
||||
vectorEffect="non-scaling-stroke"
|
||||
/>
|
||||
)
|
||||
})}
|
||||
)}
|
||||
|
||||
{hoverIndex !== null && geometry.xy[hoverIndex] && (
|
||||
<line
|
||||
x1={geometry.xy[hoverIndex].x}
|
||||
x2={geometry.xy[hoverIndex].x}
|
||||
y1={PAD_TOP}
|
||||
y2={VIEW_H - PAD_BOTTOM}
|
||||
stroke="rgb(var(--cyber-neon-cyan) / 0.5)"
|
||||
strokeWidth={1}
|
||||
vectorEffect="non-scaling-stroke"
|
||||
<polyline
|
||||
points={geometry.xy.map((node) => `${node.x},${node.y}`).join(' ')}
|
||||
fill="none"
|
||||
stroke="rgb(var(--cyber-neon-cyan))"
|
||||
strokeWidth={2}
|
||||
strokeLinejoin="round"
|
||||
strokeLinecap="round"
|
||||
/>
|
||||
)}
|
||||
|
||||
<polyline
|
||||
points={geometry.xy.map((node) => `${node.x},${node.y}`).join(' ')}
|
||||
fill="none"
|
||||
stroke="rgb(var(--cyber-neon-cyan))"
|
||||
strokeWidth={2}
|
||||
strokeLinejoin="round"
|
||||
strokeLinecap="round"
|
||||
vectorEffect="non-scaling-stroke"
|
||||
/>
|
||||
{geometry.xy.map((node, index) => (
|
||||
<circle
|
||||
key={index}
|
||||
cx={node.x}
|
||||
cy={node.y}
|
||||
r={dotRadius}
|
||||
fill="rgb(var(--cyber-neon-cyan))"
|
||||
// 与环境色同色的描边,让重叠的点之间留出间隙。
|
||||
stroke="rgb(var(--cyber-bg-primary))"
|
||||
strokeWidth={2}
|
||||
/>
|
||||
))}
|
||||
|
||||
{/* Only the endpoint is labelled - a number on every point is noise. */}
|
||||
<circle
|
||||
cx={geometry.xy[geometry.xy.length - 1].x}
|
||||
cy={geometry.xy[geometry.xy.length - 1].y}
|
||||
r={4}
|
||||
fill="rgb(var(--cyber-neon-cyan))"
|
||||
stroke="rgb(var(--cyber-bg-primary))"
|
||||
strokeWidth={2}
|
||||
vectorEffect="non-scaling-stroke"
|
||||
/>
|
||||
{geometry.xy.map((node, index) => (
|
||||
<rect
|
||||
key={index}
|
||||
x={node.x - hitWidth / 2}
|
||||
y={PAD_TOP}
|
||||
width={hitWidth}
|
||||
height={geometry.innerH}
|
||||
fill="transparent"
|
||||
onMouseEnter={() => setHoverIndex(index)}
|
||||
/>
|
||||
))}
|
||||
</svg>
|
||||
)}
|
||||
|
||||
{geometry.xy.map((node, index) => (
|
||||
<rect
|
||||
key={index}
|
||||
x={node.x - 6}
|
||||
y={PAD_TOP}
|
||||
width={12}
|
||||
height={VIEW_H - PAD_TOP - PAD_BOTTOM}
|
||||
fill="transparent"
|
||||
onMouseEnter={() => setHoverIndex(index)}
|
||||
/>
|
||||
))}
|
||||
</svg>
|
||||
|
||||
{/* Axis extremes live in text tokens, never the series colour. */}
|
||||
<span className="absolute left-0 top-0 text-[9px] font-mono text-cyber-text-muted">
|
||||
{formatCount(geometry.max)}
|
||||
</span>
|
||||
<span className="absolute left-0 bottom-5 text-[9px] font-mono text-cyber-text-muted">
|
||||
{formatCount(geometry.min)}
|
||||
</span>
|
||||
<span className="absolute right-0 top-1/2 -translate-y-1/2 text-[11px] font-mono text-cyber-text-primary">
|
||||
{formatCount(geometry.xy[geometry.xy.length - 1].value)}
|
||||
</span>
|
||||
|
||||
{hoverIndex !== null && geometry.xy[hoverIndex] && (
|
||||
{hoverIndex !== null && geometry?.xy[hoverIndex] && (
|
||||
<div
|
||||
className="absolute -top-1 px-2 py-1 rounded border border-cyber-border-DEFAULT bg-cyber-bg-elevated text-[10px] font-mono text-cyber-text-primary pointer-events-none whitespace-nowrap"
|
||||
className="absolute -top-1 px-2 py-1 rounded border border-cyber-border-DEFAULT bg-cyber-bg-elevated text-[10px] font-mono text-cyber-text-primary pointer-events-none whitespace-nowrap z-10"
|
||||
style={{
|
||||
left: `${(geometry.xy[hoverIndex].x / VIEW_W) * 100}%`,
|
||||
left: `${(geometry.xy[hoverIndex].x / width) * 100}%`,
|
||||
transform: 'translateX(-50%)',
|
||||
}}
|
||||
>
|
||||
@@ -202,7 +291,7 @@ export function NoteTrendChart({ noteId, taskId, noteTitle }: NoteTrendChartProp
|
||||
)}
|
||||
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
共 {points.length} 个数据点,每轮采集记录一次快照
|
||||
共 {points.length} 个数据点,每轮采集记录一次快照 · 纵轴范围见左侧刻度
|
||||
</p>
|
||||
</div>
|
||||
)
|
||||
|
||||
@@ -1,10 +1,17 @@
|
||||
import { Fragment, useState } from 'react'
|
||||
import { ChevronDown, ChevronRight, ExternalLink } from 'lucide-react'
|
||||
import { Fragment, useMemo, useRef, useState } from 'react'
|
||||
import { ChevronDown, ChevronRight, ExternalLink, Pencil, Users } from 'lucide-react'
|
||||
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
import { useMonitorNotes } from '@/hooks/useMonitor'
|
||||
import { formatCount, formatDelta, formatRelative } from '@/lib/monitorFormat'
|
||||
import type { MonitorNote, NoteMetrics } from '@/types/monitor'
|
||||
import { useMonitorNotes, useSetCreatorAlias, useSetNoteAlias } from '@/hooks/useMonitor'
|
||||
import {
|
||||
formatCount,
|
||||
formatDate,
|
||||
formatDateTime,
|
||||
formatDelta,
|
||||
formatRelative,
|
||||
} from '@/lib/monitorFormat'
|
||||
import type { MonitorCreator, MonitorNote, NoteMetrics } from '@/types/monitor'
|
||||
import { NoteCover } from './NoteCover'
|
||||
import { NoteTrendChart } from './NoteTrendChart'
|
||||
|
||||
interface NotesTableProps {
|
||||
@@ -38,15 +45,197 @@ const METRIC_COLUMNS: Array<{ key: keyof NoteMetrics; label: string }> = [
|
||||
{ key: 'share_count', label: '分享' },
|
||||
]
|
||||
|
||||
/**
|
||||
* 一个博主名下所有作品的指标合计。
|
||||
*
|
||||
* **全为 null 时结果是 null 而不是 0** —— 与项目一贯的口径一致:「0」是真实值,
|
||||
* 「null」是不知道。合计成 0 会让"还没采到"看起来像"互动为零"。
|
||||
*/
|
||||
function sumMetrics(notes: MonitorNote[]) {
|
||||
const values: Record<string, number | null> = {}
|
||||
const deltas: Record<string, number | null> = {}
|
||||
|
||||
for (const column of METRIC_COLUMNS) {
|
||||
let sum = 0
|
||||
let sawValue = false
|
||||
let deltaSum = 0
|
||||
let sawDelta = false
|
||||
|
||||
for (const note of notes) {
|
||||
const value = note.metrics[column.key]
|
||||
if (value !== null && value !== undefined) {
|
||||
sum += value
|
||||
sawValue = true
|
||||
}
|
||||
const delta = note.deltas[column.key]
|
||||
if (delta !== null && delta !== undefined) {
|
||||
deltaSum += delta
|
||||
sawDelta = true
|
||||
}
|
||||
}
|
||||
|
||||
values[column.key] = sawValue ? sum : null
|
||||
deltas[column.key] = sawDelta ? deltaSum : null
|
||||
}
|
||||
return { values, deltas }
|
||||
}
|
||||
|
||||
// 多出来的那几列:展开箭头、作品、发布日期、首次发现、跳转链接;分组表头会跨掉整行。
|
||||
const COLUMN_COUNT = METRIC_COLUMNS.length + 5
|
||||
|
||||
/**
|
||||
* 账号级指标那几个字段。
|
||||
*
|
||||
* 作品的 payload 和博主的 payload 都带这一组,值也一样(同一次采集写的同一行),
|
||||
* 所以组件只认字段、不认来源。
|
||||
*/
|
||||
type AccountStats = Pick<
|
||||
MonitorCreator,
|
||||
'creator_fans' | 'creator_total_favorited' | 'creator_works' | 'creator_stats_at'
|
||||
>
|
||||
|
||||
/**
|
||||
* 博主的**账号级**指标:粉丝 / 总获赞 / 作品数。
|
||||
*
|
||||
* 作品列表给不了这个 —— 那几列说的是「这条作品涨了多少赞」,这里说的是「这个人整个
|
||||
* 账号在涨还是在掉」。**一个都没采到时整块不画**:画成「粉丝 0」比不画糟得多,那是
|
||||
* 一句假话。
|
||||
*/
|
||||
function CreatorStats({ stats }: { stats: AccountStats }) {
|
||||
const parts = [
|
||||
stats.creator_fans !== null && `粉丝 ${formatCount(stats.creator_fans)}`,
|
||||
stats.creator_total_favorited !== null && `获赞 ${formatCount(stats.creator_total_favorited)}`,
|
||||
stats.creator_works !== null && `${formatCount(stats.creator_works)} 作品`,
|
||||
].filter(Boolean) as string[]
|
||||
|
||||
if (parts.length === 0) return null
|
||||
|
||||
return (
|
||||
<span
|
||||
className="text-cyber-text-muted flex-shrink-0"
|
||||
title={`账号指标采集于 ${formatDateTime(stats.creator_stats_at)}`}
|
||||
>
|
||||
{parts.join(' · ')}
|
||||
</span>
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* 作品表,**按博主分组**。
|
||||
*
|
||||
* 一个任务可以配多个博主,平铺的话根本看不出哪篇是谁的。分组键是
|
||||
* `creator_hash` —— 爬虫刻意不落原始 user_id,所以这是唯一稳定的创作者标识;
|
||||
* 显示名是已脱敏的昵称(张***三)。两者都认不出来时才退回"未知博主"。
|
||||
*/
|
||||
export function NotesTable({ taskId, onlyNew }: NotesTableProps) {
|
||||
const { data: notes, isLoading } = useMonitorNotes(taskId, onlyNew)
|
||||
const [expanded, setExpanded] = useState<string | null>(null)
|
||||
const { data, isLoading } = useMonitorNotes(taskId, onlyNew)
|
||||
const notes = data?.notes
|
||||
const creators = data?.creators
|
||||
const [expandedNote, setExpandedNote] = useState<string | null>(null)
|
||||
// 折叠状态按博主记。默认全展开 —— 藏起来的数据比多滚两屏更糟。
|
||||
const [collapsed, setCollapsed] = useState<Set<string>>(new Set())
|
||||
// 正在改备注的那个博主(creator_hash);null = 没在改。
|
||||
const [editingCreator, setEditingCreator] = useState<string | null>(null)
|
||||
const [aliasDraft, setAliasDraft] = useState('')
|
||||
// 按 Escape 是「取消」,而失焦是「保存」—— 取消时得让紧接着那次失焦闭嘴,
|
||||
// 否则它会把你刚放弃的内容存进去。
|
||||
const aliasCancelled = useRef(false)
|
||||
const setAlias = useSetCreatorAlias()
|
||||
// 正在改备注的那条**作品**。和博主备注是两个互不干扰的编辑态。
|
||||
const [editingNote, setEditingNote] = useState<string | null>(null)
|
||||
const [noteAliasDraft, setNoteAliasDraft] = useState('')
|
||||
const noteAliasCancelled = useRef(false)
|
||||
const setNoteAlias = useSetNoteAlias()
|
||||
|
||||
const commitAlias = (creatorHash: string) => {
|
||||
if (aliasCancelled.current) {
|
||||
aliasCancelled.current = false
|
||||
setEditingCreator(null)
|
||||
return
|
||||
}
|
||||
setEditingCreator(null)
|
||||
setAlias.mutate({ creatorHash, alias: aliasDraft })
|
||||
}
|
||||
|
||||
const commitNoteAlias = (noteId: string) => {
|
||||
if (noteAliasCancelled.current) {
|
||||
noteAliasCancelled.current = false
|
||||
setEditingNote(null)
|
||||
return
|
||||
}
|
||||
setEditingNote(null)
|
||||
setNoteAlias.mutate({ noteId, alias: noteAliasDraft })
|
||||
}
|
||||
|
||||
const groups = useMemo(() => {
|
||||
type Group = {
|
||||
key: string
|
||||
name: string
|
||||
alias: string
|
||||
/** 服务端说的这条博主名下有多少作品 —— 列表被 `onlyNew` 滤过时和 `notes.length` 不等。 */
|
||||
totalNotes: number
|
||||
stats: MonitorCreator | null
|
||||
notes: MonitorNote[]
|
||||
}
|
||||
const byCreator = new Map<string, Group>()
|
||||
|
||||
// **先放博主,不是先放作品。** 「一条作品都没有的博主」在作品里根本推不出来 ——
|
||||
// 目标加了、资料也采到了、粉丝数就躺在库里,可界面上什么都看不见。而他恰恰是最该
|
||||
// 看见的一个:还在涨粉,只是最近没发东西。
|
||||
for (const creator of creators ?? []) {
|
||||
const key = creator.creator_hash || '__unknown__'
|
||||
const existing = byCreator.get(key)
|
||||
if (existing) {
|
||||
// 同一个博主可能挂在多个任务下(服务端按 任务×博主 给),合并成一组。
|
||||
existing.totalNotes += creator.note_count
|
||||
if (!existing.stats?.creator_stats_at && creator.creator_stats_at) {
|
||||
existing.stats = creator
|
||||
}
|
||||
existing.name ||= creator.creator_name
|
||||
// 备注按 (platform, creator_hash) 存,所以跨任务就是同一条,取到即可。
|
||||
existing.alias ||= creator.creator_alias
|
||||
continue
|
||||
}
|
||||
byCreator.set(key, {
|
||||
key,
|
||||
name: creator.creator_name || '',
|
||||
alias: creator.creator_alias || '',
|
||||
totalNotes: creator.note_count,
|
||||
stats: creator,
|
||||
notes: [],
|
||||
})
|
||||
}
|
||||
|
||||
for (const note of notes ?? []) {
|
||||
// 作品模式下每条作品都会带 creator_hash;真丢了也要有个兜底分组,
|
||||
// 否则那些作品会凭空消失。
|
||||
const key = note.creator_hash || '__unknown__'
|
||||
let group = byCreator.get(key)
|
||||
if (!group) {
|
||||
group = {
|
||||
key,
|
||||
name: note.creator_name || '',
|
||||
alias: note.creator_alias || '',
|
||||
totalNotes: 0,
|
||||
stats: null,
|
||||
notes: [],
|
||||
}
|
||||
byCreator.set(key, group)
|
||||
}
|
||||
group.notes.push(note)
|
||||
group.name ||= note.creator_name || ''
|
||||
group.alias ||= note.creator_alias || ''
|
||||
}
|
||||
|
||||
return [...byCreator.values()]
|
||||
}, [notes, creators])
|
||||
|
||||
if (isLoading) {
|
||||
return <p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">加载中…</p>
|
||||
}
|
||||
|
||||
if (!notes || notes.length === 0) {
|
||||
// 博主有、作品没有也是**要画**的:那正是「这个号还没被删,只是没发东西」。
|
||||
if (groups.length === 0) {
|
||||
return (
|
||||
<p className="py-8 text-center text-[11px] font-mono text-cyber-text-muted">
|
||||
{taskId === null
|
||||
@@ -58,7 +247,13 @@ export function NotesTable({ taskId, onlyNew }: NotesTableProps) {
|
||||
)
|
||||
}
|
||||
|
||||
const isExpanded = (note: MonitorNote) => expanded === `${note.task_id}:${note.note_id}`
|
||||
const toggleCreator = (key: string) =>
|
||||
setCollapsed((prev) => {
|
||||
const next = new Set(prev)
|
||||
if (next.has(key)) next.delete(key)
|
||||
else next.add(key)
|
||||
return next
|
||||
})
|
||||
|
||||
return (
|
||||
<div className="overflow-x-auto terminal-scroll">
|
||||
@@ -72,70 +267,229 @@ export function NotesTable({ taskId, onlyNew }: NotesTableProps) {
|
||||
{column.label}
|
||||
</th>
|
||||
))}
|
||||
<th className="text-right font-normal py-2 px-2">发布日期</th>
|
||||
<th className="text-right font-normal py-2 px-2">首次发现</th>
|
||||
<th className="w-8" />
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
{notes.map((note) => {
|
||||
const open = isExpanded(note)
|
||||
{groups.map((group) => {
|
||||
const isCollapsed = collapsed.has(group.key)
|
||||
const totals = sumMetrics(group.notes)
|
||||
return (
|
||||
<Fragment key={`${note.task_id}:${note.note_id}`}>
|
||||
<Fragment key={group.key}>
|
||||
{/* 组头与数据行**列对齐**:每个指标列给出该博主的合计,下面再带本轮增量。
|
||||
折叠之后如果什么都看不到,那折叠就只是把信息藏起来了。 */}
|
||||
<tr
|
||||
onClick={() => setExpanded(open ? null : `${note.task_id}:${note.note_id}`)}
|
||||
className="border-t border-cyber-border-subtle hover:bg-cyber-bg-elevated/50 cursor-pointer"
|
||||
onClick={() => toggleCreator(group.key)}
|
||||
className="border-t border-cyber-border-subtle bg-cyber-bg-tertiary/50 cursor-pointer hover:bg-cyber-bg-elevated/50"
|
||||
>
|
||||
<td className="py-2 px-1 text-cyber-text-muted">
|
||||
{open ? <ChevronDown className="w-3 h-3" /> : <ChevronRight className="w-3 h-3" />}
|
||||
</td>
|
||||
<td className="py-2 px-2 max-w-[320px]">
|
||||
<div className="flex items-center gap-1.5">
|
||||
{note.is_new && (
|
||||
<Badge variant="success" className="text-[9px] px-1 py-0 flex-shrink-0">
|
||||
NEW
|
||||
</Badge>
|
||||
)}
|
||||
<span className="truncate text-cyber-text-primary" title={note.title}>
|
||||
{note.title || note.note_id}
|
||||
</span>
|
||||
</div>
|
||||
<div className="text-[9px] text-cyber-text-muted truncate">
|
||||
{note.note_id} · {note.snapshot_count} 次快照
|
||||
</div>
|
||||
</td>
|
||||
{METRIC_COLUMNS.map((column) => (
|
||||
<td key={column.key} className="py-2 px-2">
|
||||
<MetricCell value={note.metrics[column.key]} delta={note.deltas[column.key]} />
|
||||
</td>
|
||||
))}
|
||||
<td className="py-2 px-2 text-right text-[10px] text-cyber-text-muted">
|
||||
{formatRelative(note.first_seen_at)}
|
||||
</td>
|
||||
<td className="py-2 px-1">
|
||||
{note.note_url && (
|
||||
<a
|
||||
href={note.note_url}
|
||||
target="_blank"
|
||||
rel="noopener noreferrer"
|
||||
onClick={(event) => event.stopPropagation()}
|
||||
className="text-cyber-text-muted hover:text-cyber-neon-cyan"
|
||||
>
|
||||
<ExternalLink className="w-3 h-3" />
|
||||
</a>
|
||||
<td className="py-1.5 px-1 text-cyber-text-muted">
|
||||
{isCollapsed ? (
|
||||
<ChevronRight className="w-3 h-3" />
|
||||
) : (
|
||||
<ChevronDown className="w-3 h-3" />
|
||||
)}
|
||||
</td>
|
||||
</tr>
|
||||
{open && (
|
||||
<tr className="bg-cyber-bg-secondary/40">
|
||||
<td colSpan={METRIC_COLUMNS.length + 4} className="px-4 py-3">
|
||||
<NoteTrendChart
|
||||
noteId={note.note_id}
|
||||
taskId={note.task_id}
|
||||
noteTitle={note.title}
|
||||
<td className="py-1.5 px-2">
|
||||
{editingCreator === group.key ? (
|
||||
<input
|
||||
autoFocus
|
||||
value={aliasDraft}
|
||||
placeholder="备注,比如「竞品A」"
|
||||
onClick={(event) => event.stopPropagation()}
|
||||
onChange={(event) => setAliasDraft(event.target.value)}
|
||||
onKeyDown={(event) => {
|
||||
event.stopPropagation()
|
||||
if (event.key === 'Enter') commitAlias(group.key)
|
||||
if (event.key === 'Escape') {
|
||||
aliasCancelled.current = true
|
||||
setEditingCreator(null)
|
||||
}
|
||||
}}
|
||||
onBlur={() => commitAlias(group.key)}
|
||||
className="w-full rounded border border-cyber-neon-cyan/40 bg-cyber-bg-tertiary px-1.5 py-0.5 text-[11px] font-mono text-cyber-text-primary outline-none"
|
||||
/>
|
||||
) : (
|
||||
<span className="flex items-center gap-1.5">
|
||||
<Users className="w-3 h-3 text-cyber-neon-cyan flex-shrink-0" />
|
||||
<span className="text-cyber-text-primary truncate">
|
||||
{group.alias || group.name || '未知博主'}
|
||||
</span>
|
||||
{/* 起了备注之后,平台昵称降成副标题 —— 它仍是有用的对照。 */}
|
||||
{group.alias && group.name && (
|
||||
<span className="text-cyber-text-muted truncate">{group.name}</span>
|
||||
)}
|
||||
<span className="text-cyber-text-muted flex-shrink-0">
|
||||
{/* 被 onlyNew 滤过时,把「显示了几篇 / 一共几篇」都说出来,
|
||||
否则「1 篇」会让人以为这个号只发过一条。 */}
|
||||
{group.notes.length}
|
||||
{group.totalNotes > group.notes.length && `/${group.totalNotes}`} 篇
|
||||
</span>
|
||||
{group.stats && <CreatorStats stats={group.stats} />}
|
||||
{/* 博主在、作品一条都没有 —— 说清楚,别让人以为列表挂了。 */}
|
||||
{group.notes.length === 0 && (
|
||||
<span className="text-cyber-text-muted/70 flex-shrink-0">
|
||||
暂无作品
|
||||
</span>
|
||||
)}
|
||||
<button
|
||||
onClick={(event) => {
|
||||
event.stopPropagation()
|
||||
setAliasDraft(group.alias)
|
||||
setEditingCreator(group.key)
|
||||
}}
|
||||
title="给这个博主起个备注"
|
||||
className="text-cyber-text-muted hover:text-cyber-neon-cyan flex-shrink-0"
|
||||
>
|
||||
<Pencil className="w-3 h-3" />
|
||||
</button>
|
||||
</span>
|
||||
)}
|
||||
</td>
|
||||
{METRIC_COLUMNS.map((column) => (
|
||||
<td key={column.key} className="py-1.5 px-2">
|
||||
<MetricCell
|
||||
value={totals.values[column.key]}
|
||||
delta={totals.deltas[column.key]}
|
||||
/>
|
||||
</td>
|
||||
</tr>
|
||||
)}
|
||||
))}
|
||||
<td />
|
||||
<td />
|
||||
</tr>
|
||||
|
||||
{!isCollapsed &&
|
||||
group.notes.map((note) => {
|
||||
const rowKey = `${note.task_id}:${note.note_id}`
|
||||
const open = expandedNote === rowKey
|
||||
return (
|
||||
<Fragment key={rowKey}>
|
||||
<tr
|
||||
onClick={() => setExpandedNote(open ? null : rowKey)}
|
||||
className="border-t border-cyber-border-subtle hover:bg-cyber-bg-elevated/50 cursor-pointer"
|
||||
>
|
||||
<td />
|
||||
<td className="py-2 px-2 max-w-[320px]">
|
||||
<div className="flex items-center gap-2">
|
||||
<NoteCover src={note.cover} size="md" />
|
||||
{/* min-w-0 so the title can actually truncate inside the flex row */}
|
||||
<div className="min-w-0">
|
||||
<div className="flex items-center gap-1.5">
|
||||
{note.is_new && (
|
||||
<Badge
|
||||
variant="success"
|
||||
className="text-[9px] px-1 py-0 flex-shrink-0"
|
||||
>
|
||||
NEW
|
||||
</Badge>
|
||||
)}
|
||||
{/* 备注是**加**在标题前面的一枚标记,不是替换 ——
|
||||
标题才是这条作品本身,扫列表时两个都要看得见。 */}
|
||||
{note.note_alias && (
|
||||
<Badge
|
||||
variant="default"
|
||||
className="text-[9px] px-1 py-0 flex-shrink-0"
|
||||
>
|
||||
{note.note_alias}
|
||||
</Badge>
|
||||
)}
|
||||
<span
|
||||
className="truncate text-cyber-text-primary"
|
||||
title={note.title}
|
||||
>
|
||||
{note.title || note.note_id}
|
||||
</span>
|
||||
</div>
|
||||
{editingNote === rowKey ? (
|
||||
<input
|
||||
autoFocus
|
||||
value={noteAliasDraft}
|
||||
placeholder="备注,比如「重点跟拍」"
|
||||
onClick={(event) => event.stopPropagation()}
|
||||
onChange={(event) => setNoteAliasDraft(event.target.value)}
|
||||
onKeyDown={(event) => {
|
||||
event.stopPropagation()
|
||||
if (event.key === 'Enter') commitNoteAlias(note.note_id)
|
||||
if (event.key === 'Escape') {
|
||||
noteAliasCancelled.current = true
|
||||
setEditingNote(null)
|
||||
}
|
||||
}}
|
||||
onBlur={() => commitNoteAlias(note.note_id)}
|
||||
className="mt-1 w-full rounded border border-cyber-neon-cyan/40 bg-cyber-bg-tertiary px-1.5 py-0.5 text-[11px] font-mono text-cyber-text-primary outline-none"
|
||||
/>
|
||||
) : (
|
||||
<div className="flex items-center gap-1.5 text-[9px] text-cyber-text-muted">
|
||||
<span className="truncate">
|
||||
{note.note_id} · {note.snapshot_count} 次快照
|
||||
</span>
|
||||
<button
|
||||
onClick={(event) => {
|
||||
event.stopPropagation()
|
||||
setNoteAliasDraft(note.note_alias)
|
||||
setEditingNote(rowKey)
|
||||
}}
|
||||
title="给这条作品起个备注"
|
||||
className="flex items-center gap-0.5 flex-shrink-0 text-cyber-text-muted hover:text-cyber-neon-cyan"
|
||||
>
|
||||
<Pencil className="w-2.5 h-2.5" />
|
||||
{note.note_alias ? '改备注' : '备注'}
|
||||
</button>
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
</td>
|
||||
{METRIC_COLUMNS.map((column) => (
|
||||
<td key={column.key} className="py-2 px-2">
|
||||
<MetricCell
|
||||
value={note.metrics[column.key]}
|
||||
delta={note.deltas[column.key]}
|
||||
/>
|
||||
</td>
|
||||
))}
|
||||
{/* 发布日期和「首次发现」是两回事:把早就发过的作品加进监控时,
|
||||
前者是作者发的那天,后者是我们第一次看到它的那天。 */}
|
||||
<td
|
||||
className="py-2 px-2 text-right text-[10px] text-cyber-text-muted whitespace-nowrap"
|
||||
title={formatDateTime(note.published_at)}
|
||||
>
|
||||
{formatDate(note.published_at)}
|
||||
</td>
|
||||
<td className="py-2 px-2 text-right text-[10px] text-cyber-text-muted">
|
||||
{formatRelative(note.first_seen_at)}
|
||||
</td>
|
||||
<td className="py-2 px-1">
|
||||
{/* 直接跳转到该作品。stopPropagation,否则点它会连带展开趋势图。 */}
|
||||
{note.note_url && (
|
||||
<a
|
||||
href={note.note_url}
|
||||
target="_blank"
|
||||
rel="noopener noreferrer"
|
||||
onClick={(event) => event.stopPropagation()}
|
||||
title="打开原作品"
|
||||
className="text-cyber-text-muted hover:text-cyber-neon-cyan"
|
||||
>
|
||||
<ExternalLink className="w-3 h-3" />
|
||||
</a>
|
||||
)}
|
||||
</td>
|
||||
</tr>
|
||||
{open && (
|
||||
<tr className="bg-cyber-bg-secondary/40">
|
||||
<td colSpan={COLUMN_COUNT} className="px-4 py-3">
|
||||
<NoteTrendChart
|
||||
noteId={note.note_id}
|
||||
taskId={note.task_id}
|
||||
noteTitle={note.title}
|
||||
/>
|
||||
</td>
|
||||
</tr>
|
||||
)}
|
||||
</Fragment>
|
||||
)
|
||||
})}
|
||||
</Fragment>
|
||||
)
|
||||
})}
|
||||
|
||||
@@ -0,0 +1,238 @@
|
||||
import { useEffect, useState } from 'react'
|
||||
import { useQueryClient } from '@tanstack/react-query'
|
||||
import {
|
||||
AlertTriangle,
|
||||
CheckCircle2,
|
||||
KeyRound,
|
||||
Loader2,
|
||||
QrCode,
|
||||
RefreshCw,
|
||||
X,
|
||||
} from 'lucide-react'
|
||||
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
import { Button } from '@/components/ui/button'
|
||||
import {
|
||||
useCancelQrLogin,
|
||||
useLoginState,
|
||||
useQrLoginStatus,
|
||||
useRecheckLogin,
|
||||
useStartQrLogin,
|
||||
} from '@/hooks/useMonitor'
|
||||
import { useCurrentPlatform } from '@/hooks/usePlatform'
|
||||
|
||||
/**
|
||||
* Scan-to-login, for a host with no display.
|
||||
*
|
||||
* The crawler's own QR flow prints the code into the terminal via PIL's
|
||||
* `Image.show()`, which needs a desktop image viewer. On a server Chrome runs
|
||||
* under Xvfb and there is no such viewer, so the backend reads the code out of
|
||||
* that same browser over CDP and hands it here.
|
||||
*
|
||||
* **The login state is shown first and independently of the scan.** It comes from
|
||||
* the browser's own profile, so it stays true across a server restart — the QR
|
||||
* session does not, and reporting "did it work?" from a value that a redeploy
|
||||
* silently erases is how a successful scan ends up looking like nothing happened.
|
||||
*/
|
||||
export function QrLoginPanel() {
|
||||
const { capability, platform } = useCurrentPlatform()
|
||||
const queryClient = useQueryClient()
|
||||
const [polling, setPolling] = useState(false)
|
||||
|
||||
// 登录态检测目前只实现了小红书:它读的是小红书页面的 __INITIAL_STATE__。
|
||||
// 后端 qrlogin.LOGIN_URL 里也只有 xhs 一项。
|
||||
//
|
||||
// 不挡住的话会撒一个具体的谎 —— 切到抖音时,面板照样报"已登录",因为小红书的
|
||||
// 会话还在,而标题写的是「抖音登录态」。
|
||||
const wired = platform === 'xhs'
|
||||
|
||||
const { data: state } = useQrLoginStatus(polling && wired)
|
||||
const { data: login } = useLoginState(polling && wired)
|
||||
const start = useStartQrLogin()
|
||||
const cancel = useCancelQrLogin()
|
||||
const recheck = useRecheckLogin()
|
||||
|
||||
const label = capability?.label ?? platform
|
||||
const status = state?.status ?? 'idle'
|
||||
|
||||
// The browser's own answer wins over the session's: the session is memory, the
|
||||
// profile is not.
|
||||
const loggedIn = login?.logged_in ?? state?.logged_in ?? false
|
||||
const nickname = login?.nickname ?? state?.nickname ?? null
|
||||
|
||||
// Stop polling the moment the outcome is known, and refresh the cookie panel:
|
||||
// a completed scan is what makes it start reporting a healthy login.
|
||||
useEffect(() => {
|
||||
if (status === 'waiting' || status === 'idle') return
|
||||
setPolling(false)
|
||||
if (status === 'success') {
|
||||
queryClient.invalidateQueries({ queryKey: ['monitorCookie'] })
|
||||
queryClient.invalidateQueries({ queryKey: ['monitorLoginState'] })
|
||||
}
|
||||
}, [status, queryClient])
|
||||
|
||||
// 未接入的平台直接换掉整个面板。只改标题是不够的 —— 下面那些块读的是小红书的
|
||||
// 状态,照渲染出来还是会显示"已登录"。
|
||||
if (!wired) {
|
||||
return (
|
||||
<div className="rounded-lg glass-panel float-panel p-4 space-y-2">
|
||||
<div className="flex items-center gap-2">
|
||||
<KeyRound className="w-4 h-4 text-cyber-text-muted flex-shrink-0" />
|
||||
<span className="font-mono text-xs tracking-wider text-cyber-text-primary">
|
||||
浏览器登录态(扫码)
|
||||
</span>
|
||||
<Badge variant="outline" className="text-[10px]">
|
||||
{label}未接入
|
||||
</Badge>
|
||||
</div>
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
扫码登录目前只实现了<span className="text-cyber-neon-cyan">小红书</span> ——
|
||||
登录态检测读的是小红书页面的登录状态,其它平台还没有对应的实现。
|
||||
切回小红书即可使用。
|
||||
</p>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
const begin = () => {
|
||||
setPolling(true)
|
||||
start.mutate()
|
||||
}
|
||||
|
||||
const busy = start.isPending || cancel.isPending || recheck.isPending
|
||||
|
||||
return (
|
||||
<div className="rounded-lg glass-panel float-panel p-4 space-y-3">
|
||||
<div className="flex items-center justify-between gap-3">
|
||||
<div className="flex items-center gap-2 min-w-0">
|
||||
<KeyRound className="w-4 h-4 text-cyber-neon-cyan flex-shrink-0" />
|
||||
{/* 标题必须和 Cookie 面板区分开:那个是"存进库、每轮注入"的 cookie,
|
||||
这个是"写进浏览器 profile、CDP 模式复用"的浏览器登录态。两者是
|
||||
不同的机制,都能算"登录着",只写「登录态」会看不出指的是哪一个。 */}
|
||||
<span className="font-mono text-xs tracking-wider text-cyber-text-primary flex-shrink-0">
|
||||
浏览器登录态(扫码)
|
||||
</span>
|
||||
{!wired ? (
|
||||
<Badge variant="outline" className="text-[10px]">
|
||||
{label}未接入
|
||||
</Badge>
|
||||
) : loggedIn ? (
|
||||
<Badge variant="success" className="text-[10px]">
|
||||
已登录
|
||||
</Badge>
|
||||
) : (
|
||||
<Badge variant="warning" className="text-[10px]">
|
||||
未登录
|
||||
</Badge>
|
||||
)}
|
||||
{loggedIn && nickname && (
|
||||
<span className="truncate text-[10px] font-mono text-cyber-text-secondary">
|
||||
{nickname}
|
||||
</span>
|
||||
)}
|
||||
</div>
|
||||
|
||||
<Button
|
||||
variant="ghost"
|
||||
size="sm"
|
||||
disabled={busy}
|
||||
onClick={() => recheck.mutate()}
|
||||
title="重新加载页面并检测登录态"
|
||||
>
|
||||
<RefreshCw className={`w-3 h-3 mr-1 ${recheck.isPending ? 'animate-spin' : ''}`} />
|
||||
重新检测
|
||||
</Button>
|
||||
</div>
|
||||
|
||||
{/* 读不到状态时要说清楚是"读不到",而不是悄悄显示成"未登录" */}
|
||||
{login?.known === false && (
|
||||
<p className="flex items-start gap-2 text-[10px] font-mono text-cyber-neon-orange leading-relaxed">
|
||||
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
|
||||
读不到浏览器状态{login.error ? `:${login.error}` : ''}。
|
||||
请确认那台 Chrome 正常、且已打开小红书页面。
|
||||
</p>
|
||||
)}
|
||||
|
||||
{status === 'waiting' && state?.image && (
|
||||
<div className="flex items-start gap-4">
|
||||
{/* The code is a data: URL straight from the page, so nothing is
|
||||
fetched from a third party and no file is written server-side. */}
|
||||
<img
|
||||
src={state.image}
|
||||
alt="登录二维码"
|
||||
className="w-44 h-44 rounded-md border border-cyber-border-DEFAULT bg-white p-1"
|
||||
/>
|
||||
<div className="space-y-2 pt-1">
|
||||
<p className="text-[11px] font-mono text-cyber-text-secondary leading-relaxed">
|
||||
用<span className="text-cyber-neon-cyan">{label} App</span>扫码。
|
||||
扫完这个面板会自动变成「已登录」,无需手动刷新。
|
||||
</p>
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted">
|
||||
剩余 <span className="text-cyber-neon-cyan">{state.expires_in}</span> 秒
|
||||
</p>
|
||||
<Button
|
||||
variant="ghost"
|
||||
size="sm"
|
||||
disabled={busy}
|
||||
onClick={() => {
|
||||
setPolling(false)
|
||||
cancel.mutate()
|
||||
}}
|
||||
>
|
||||
<X className="w-3 h-3 mr-1" />
|
||||
取消
|
||||
</Button>
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{status === 'waiting' && !state?.image && (
|
||||
<p className="flex items-center gap-2 text-[11px] font-mono text-cyber-text-muted">
|
||||
<Loader2 className="w-3 h-3 animate-spin" />
|
||||
正在从浏览器取二维码…
|
||||
</p>
|
||||
)}
|
||||
|
||||
{status === 'success' && (
|
||||
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-green leading-relaxed">
|
||||
<CheckCircle2 className="w-3 h-3 mt-0.5 shrink-0" />
|
||||
{state?.message || '登录成功'}
|
||||
</p>
|
||||
)}
|
||||
|
||||
{(status === 'error' || status === 'expired') && (
|
||||
<div className="space-y-2">
|
||||
<p className="flex items-start gap-2 text-[11px] font-mono text-cyber-neon-orange leading-relaxed">
|
||||
<AlertTriangle className="w-3 h-3 mt-0.5 shrink-0" />
|
||||
{state?.message || '获取二维码失败'}
|
||||
</p>
|
||||
<Button variant="outline" size="sm" disabled={busy} onClick={begin}>
|
||||
<RefreshCw className="w-3 h-3 mr-1" />
|
||||
重新获取
|
||||
</Button>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{status === 'idle' && (
|
||||
<div className="space-y-2">
|
||||
<p className="text-[11px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
{loggedIn
|
||||
? '这台浏览器已是登录状态。CDP 模式下定时任务直接复用它;点下面的按钮可以把这份登录态也存成 Cookie —— 那样即使关掉 CDP、任务改用 Cookie 注入也照样能跑。'
|
||||
: '经 CDP 接管服务器上已开启远程调试的 Chrome,把二维码取回来显示在这里。需要先在「系统设置」里打开 接管已有 Chrome(CDP),并确保那台 Chrome 正以 9222 端口运行。'}
|
||||
</p>
|
||||
{/* 已登录时**也要给按钮**。原先这里把按钮藏了,于是面板变成一块只能看、
|
||||
不能操作的区域 —— 用户的原话是「没用」。两种状态下点击是同一个动作,
|
||||
只是含义不同:没登录就是取二维码,已登录就是把当前登录态同步成 Cookie。 */}
|
||||
<Button size="sm" disabled={busy} onClick={begin}>
|
||||
{start.isPending ? (
|
||||
<Loader2 className="w-3 h-3 mr-1 animate-spin" />
|
||||
) : (
|
||||
<QrCode className="w-3 h-3 mr-1" />
|
||||
)}
|
||||
{loggedIn ? '同步登录态为 Cookie' : '获取二维码'}
|
||||
</Button>
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
@@ -83,7 +83,12 @@ export function RunHistory({ taskId }: RunHistoryProps) {
|
||||
{formatDateTime(run.started_at)}
|
||||
<span className="ml-1">({formatRelative(run.started_at)})</span>
|
||||
</td>
|
||||
<td className="py-2 px-2 text-[10px] text-cyber-neon-orange max-w-[240px] truncate">
|
||||
{/* 这一格是截断的,而失败原因现在会带上一整行异常 —— 没有 title
|
||||
就等于把最要紧的那半句藏起来了。 */}
|
||||
<td
|
||||
className="py-2 px-2 text-[10px] text-cyber-neon-orange max-w-[320px] truncate"
|
||||
title={run.error_message ?? undefined}
|
||||
>
|
||||
{run.error_message ?? ''}
|
||||
</td>
|
||||
</tr>
|
||||
|
||||
@@ -3,7 +3,8 @@ import { Bell, CalendarClock, Pencil, Play, Trash2 } from 'lucide-react'
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
import { Button } from '@/components/ui/button'
|
||||
import { useDeleteTask, useRunTaskNow, useUpdateTask } from '@/hooks/useMonitor'
|
||||
import { formatInterval, formatRelative } from '@/lib/monitorFormat'
|
||||
import { useCurrentPlatform } from '@/hooks/usePlatform'
|
||||
import { formatRelative } from '@/lib/monitorFormat'
|
||||
import type { MonitorTask } from '@/types/monitor'
|
||||
|
||||
interface TaskCardProps {
|
||||
@@ -46,6 +47,12 @@ export function TaskCard({ task, selected, onSelect, onEdit }: TaskCardProps) {
|
||||
const updateTask = useUpdateTask()
|
||||
const deleteTask = useDeleteTask()
|
||||
const runNow = useRunTaskNow()
|
||||
// 按**任务自己的**平台取措辞,而不是当前平台 —— 卡片未必只出现在同平台的列表里。
|
||||
// 抖音管它们叫「作品」,小红书叫「笔记」,写死一个对另一个就是错的。
|
||||
const { platforms } = useCurrentPlatform()
|
||||
const noteNoun =
|
||||
platforms.find((entry) => entry.value === task.platform)?.target_hints?.note_label ??
|
||||
'笔记'
|
||||
|
||||
// A suspected cookie failure surfaces here so it is visible without opening
|
||||
// the event feed.
|
||||
@@ -65,7 +72,7 @@ export function TaskCard({ task, selected, onSelect, onEdit }: TaskCardProps) {
|
||||
<div className="flex items-center gap-2 flex-wrap">
|
||||
<span className="font-mono text-sm text-cyber-text-primary truncate">{task.name}</span>
|
||||
<Badge variant="outline" className="text-[10px]">
|
||||
{task.mode === 'creator' ? '博主' : '笔记'}
|
||||
{task.mode === 'creator' ? '博主' : noteNoun}
|
||||
</Badge>
|
||||
<Badge variant={statusVariant(task.last_status)} className="text-[10px]">
|
||||
{STATUS_LABEL[task.last_status] ?? task.last_status}
|
||||
@@ -82,7 +89,8 @@ export function TaskCard({ task, selected, onSelect, onEdit }: TaskCardProps) {
|
||||
<span>
|
||||
目标 <span className="text-cyber-neon-cyan">{task.target_count}</span> 个
|
||||
</span>
|
||||
<span>间隔 {formatInterval(task.interval_minutes)}</span>
|
||||
{/* 由后端拼好,列表和编辑弹窗因此不会对同一个计划给出两种说法 */}
|
||||
<span>{task.schedule_label}</span>
|
||||
<span className="flex items-center gap-1">
|
||||
<CalendarClock className="w-3 h-3" />
|
||||
{task.enabled ? formatRelative(task.next_run_at) : '已暂停'}
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
import { useEffect, useState } from 'react'
|
||||
import { useEffect, useState, type ReactNode } from 'react'
|
||||
|
||||
import { Button } from '@/components/ui/button'
|
||||
import { Checkbox } from '@/components/ui/checkbox'
|
||||
@@ -20,7 +20,14 @@ import {
|
||||
SelectValue,
|
||||
} from '@/components/ui/select'
|
||||
import { useCreateTask, useSettings, useUpdateTask } from '@/hooks/useMonitor'
|
||||
import type { MonitorMode, MonitorTask, TaskCreatePayload } from '@/types/monitor'
|
||||
import { useCurrentPlatform } from '@/hooks/usePlatform'
|
||||
import { describeSchedule } from '@/lib/monitorFormat'
|
||||
import type {
|
||||
MonitorMode,
|
||||
MonitorTask,
|
||||
ScheduleMode,
|
||||
TaskCreatePayload,
|
||||
} from '@/types/monitor'
|
||||
|
||||
interface TaskEditorDialogProps {
|
||||
open: boolean
|
||||
@@ -44,19 +51,80 @@ const INTERVAL_OPTIONS = [
|
||||
{ value: '10080', label: '7 天' },
|
||||
]
|
||||
|
||||
const SCHEDULE_MODE_OPTIONS: Array<{ value: ScheduleMode; label: string }> = [
|
||||
{ value: 'interval', label: '固定间隔' },
|
||||
{ value: 'daily', label: '每天定时' },
|
||||
{ value: 'weekly', label: '每周定时' },
|
||||
]
|
||||
|
||||
const HOURS = Array.from({ length: 24 }, (_, hour) => hour)
|
||||
|
||||
/**
|
||||
* Above this many works per run, warn about the request volume.
|
||||
*
|
||||
* The runner serialises crawls (max_concurrency_num=1) and each note also pulls
|
||||
* up to `max_comments_count` comments, so the cost is works × (1 + comments) and
|
||||
* the default run timeout is an hour.
|
||||
*/
|
||||
const NOTE_VOLUME_WARN = 500
|
||||
const MINUTES = Array.from({ length: 60 }, (_, minute) => minute)
|
||||
const WEEKDAYS = ['周一', '周二', '周三', '周四', '周五', '周六', '周日']
|
||||
const pad = (value: number) => String(value).padStart(2, '0')
|
||||
|
||||
function toggleNumber(values: number[], value: number): number[] {
|
||||
return values.includes(value)
|
||||
? values.filter((item) => item !== value)
|
||||
: [...values, value].sort((a, b) => a - b)
|
||||
}
|
||||
|
||||
/** A selectable pill. Used for the hours of the day and the weekdays. */
|
||||
function Chip({
|
||||
active,
|
||||
onClick,
|
||||
children,
|
||||
}: {
|
||||
active: boolean
|
||||
onClick: () => void
|
||||
children: ReactNode
|
||||
}) {
|
||||
return (
|
||||
<button
|
||||
type="button"
|
||||
onClick={onClick}
|
||||
className={`px-1.5 py-0.5 rounded border text-[10px] font-mono transition-colors ${
|
||||
active
|
||||
? 'border-cyber-neon-cyan/60 bg-cyber-neon-cyan/20 text-cyber-neon-cyan'
|
||||
: 'border-cyber-border-DEFAULT text-cyber-text-muted hover:text-cyber-text-secondary'
|
||||
}`}
|
||||
>
|
||||
{children}
|
||||
</button>
|
||||
)
|
||||
}
|
||||
|
||||
export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogProps) {
|
||||
const isEdit = Boolean(task)
|
||||
const createTask = useCreateTask()
|
||||
const updateTask = useUpdateTask()
|
||||
const { data: settings } = useSettings()
|
||||
// 示例链接和措辞都由服务端的能力矩阵给 —— 前端不自己判断平台,否则加一个平台
|
||||
// 就要改这里一次,而且很容易漏。
|
||||
const { capability, platform } = useCurrentPlatform()
|
||||
const hints = capability?.target_hints
|
||||
|
||||
const [name, setName] = useState('')
|
||||
const [mode, setMode] = useState<MonitorMode>('creator')
|
||||
const [intervalMinutes, setIntervalMinutes] = useState('360')
|
||||
const [scheduleMode, setScheduleMode] = useState<ScheduleMode>('interval')
|
||||
const [scheduleHours, setScheduleHours] = useState<number[]>([])
|
||||
const [scheduleDays, setScheduleDays] = useState<number[]>([])
|
||||
const [scheduleMinute, setScheduleMinute] = useState('0')
|
||||
const [maxNotes, setMaxNotes] = useState('20')
|
||||
const [enableComments, setEnableComments] = useState(true)
|
||||
const [maxComments, setMaxComments] = useState('50')
|
||||
const [notifyEnabled, setNotifyEnabled] = useState(false)
|
||||
// 异常推送默认开:失败意味着这个任务从此默默采不到东西,而你不会知道。
|
||||
const [notifyFailures, setNotifyFailures] = useState(true)
|
||||
const [targets, setTargets] = useState('')
|
||||
|
||||
// Reset the form whenever the dialog is (re)opened. For a new task the
|
||||
@@ -69,6 +137,10 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
setIntervalMinutes(
|
||||
String(task?.interval_minutes ?? settings?.values['collect.default_interval_minutes'] ?? 360),
|
||||
)
|
||||
setScheduleMode(task?.schedule_mode ?? 'interval')
|
||||
setScheduleHours(task?.schedule_hours ?? [])
|
||||
setScheduleDays(task?.schedule_days ?? [])
|
||||
setScheduleMinute(String(task?.schedule_minute ?? 0))
|
||||
setMaxNotes(
|
||||
String(task?.max_notes_count ?? settings?.values['collect.default_max_notes'] ?? 20),
|
||||
)
|
||||
@@ -77,6 +149,7 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
String(task?.max_comments_count ?? settings?.values['collect.default_max_comments'] ?? 50),
|
||||
)
|
||||
setNotifyEnabled(task?.notify_enabled ?? false)
|
||||
setNotifyFailures(task?.notify_failures ?? true)
|
||||
setTargets(task ? task.targets.map((t) => t.raw_value || t.external_id).join('\n') : '')
|
||||
}, [open, task, settings])
|
||||
|
||||
@@ -85,29 +158,55 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
.map((value) => value.trim())
|
||||
.filter(Boolean)
|
||||
|
||||
// Switching into a clock mode with nothing chosen would be invalid, so seed a
|
||||
// sensible starting point -- the operator then edits rather than fills a blank.
|
||||
const selectScheduleMode = (next: ScheduleMode) => {
|
||||
setScheduleMode(next)
|
||||
if (next !== 'interval' && scheduleHours.length === 0) setScheduleHours([9])
|
||||
if (next === 'weekly' && scheduleDays.length === 0) setScheduleDays([0, 1, 2, 3, 4])
|
||||
}
|
||||
|
||||
const pending = createTask.isPending || updateTask.isPending
|
||||
// Note-mode and creator-mode targets are different shapes, so switching mode
|
||||
// would silently mis-parse the list. The backend decides parsing from the
|
||||
// task's stored mode, hence mode is fixed once created.
|
||||
const canSubmit = name.trim().length > 0 && targetList.length > 0 && !pending
|
||||
const clockIncomplete =
|
||||
scheduleMode !== 'interval' &&
|
||||
(scheduleHours.length === 0 || (scheduleMode === 'weekly' && scheduleDays.length === 0))
|
||||
const canSubmit =
|
||||
name.trim().length > 0 && targetList.length > 0 && !pending && !clockIncomplete
|
||||
|
||||
const handleSubmit = () => {
|
||||
const payload: TaskCreatePayload = {
|
||||
name: name.trim(),
|
||||
// 当前平台。漏掉这一项,后端会退回小红书 —— 表现是「在抖音页面建的任务
|
||||
// 跑到小红书列表里去了」,而且不报任何错。
|
||||
platform,
|
||||
mode,
|
||||
interval_minutes: Number(intervalMinutes),
|
||||
schedule_mode: scheduleMode,
|
||||
// Interval mode has no clock times, so send none rather than leave stale
|
||||
// ones behind from a mode the operator tried and abandoned.
|
||||
schedule_hours: scheduleMode === 'interval' ? [] : scheduleHours,
|
||||
schedule_days: scheduleMode === 'weekly' ? scheduleDays : [],
|
||||
schedule_minute: Number(scheduleMinute),
|
||||
max_notes_count: Number(maxNotes),
|
||||
enable_comments: enableComments,
|
||||
max_comments_count: Number(maxComments),
|
||||
run_timeout_seconds: 3600,
|
||||
enabled: true,
|
||||
notify_enabled: notifyEnabled,
|
||||
notify_failures: notifyFailures,
|
||||
targets: targetList,
|
||||
}
|
||||
|
||||
const done = () => onOpenChange(false)
|
||||
if (isEdit && task) {
|
||||
updateTask.mutate({ id: task.id, payload }, { onSuccess: done })
|
||||
// 平台创建后不可更改,更新请求里就不带它了 —— 带着会让「平台能被改」这件事
|
||||
// 看起来像是真的。
|
||||
const updatePayload: Partial<TaskCreatePayload> = { ...payload }
|
||||
delete updatePayload.platform
|
||||
updateTask.mutate({ id: task.id, payload: updatePayload }, { onSuccess: done })
|
||||
} else {
|
||||
createTask.mutate(payload, { onSuccess: done })
|
||||
}
|
||||
@@ -149,7 +248,7 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
</SelectTrigger>
|
||||
<SelectContent>
|
||||
<SelectItem value="creator">博主(监控其作品)</SelectItem>
|
||||
<SelectItem value="note">笔记(批量监控指定内容)</SelectItem>
|
||||
<SelectItem value="note">{`${hints?.note_label ?? '笔记'}(批量监控指定内容)`}</SelectItem>
|
||||
</SelectContent>
|
||||
</Select>
|
||||
{isEdit && (
|
||||
@@ -159,6 +258,28 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
)}
|
||||
</div>
|
||||
|
||||
<div className="space-y-2">
|
||||
<Label className="text-xs font-mono text-cyber-text-secondary">运行计划</Label>
|
||||
<div className="flex gap-1">
|
||||
{SCHEDULE_MODE_OPTIONS.map((option) => (
|
||||
<button
|
||||
key={option.value}
|
||||
type="button"
|
||||
onClick={() => selectScheduleMode(option.value)}
|
||||
className={`flex-1 rounded border px-1 py-1.5 text-[10px] font-mono transition-colors ${
|
||||
scheduleMode === option.value
|
||||
? 'border-cyber-neon-cyan/60 bg-cyber-neon-cyan/20 text-cyber-neon-cyan'
|
||||
: 'border-cyber-border-DEFAULT text-cyber-text-muted hover:text-cyber-text-secondary'
|
||||
}`}
|
||||
>
|
||||
{option.label}
|
||||
</button>
|
||||
))}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{scheduleMode === 'interval' ? (
|
||||
<div className="space-y-2">
|
||||
<Label className="text-xs font-mono text-cyber-text-secondary">采集间隔</Label>
|
||||
<Select value={intervalMinutes} onValueChange={setIntervalMinutes}>
|
||||
@@ -173,46 +294,154 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
))}
|
||||
</SelectContent>
|
||||
</Select>
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
从上一轮<span className="text-cyber-text-secondary">开始</span>计时,
|
||||
所以某轮跑得久也不会让下一轮紧接着触发。
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
) : (
|
||||
<div className="space-y-3 rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3">
|
||||
{scheduleMode === 'weekly' && (
|
||||
<div className="space-y-1.5">
|
||||
<div className="text-[10px] font-mono text-cyber-text-secondary">星期</div>
|
||||
<div className="flex flex-wrap gap-1">
|
||||
{WEEKDAYS.map((label, day) => (
|
||||
<Chip
|
||||
key={day}
|
||||
active={scheduleDays.includes(day)}
|
||||
onClick={() => setScheduleDays((prev) => toggleNumber(prev, day))}
|
||||
>
|
||||
{label}
|
||||
</Chip>
|
||||
))}
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
|
||||
<div className="space-y-1.5">
|
||||
<div className="text-[10px] font-mono text-cyber-text-secondary">
|
||||
时间(可多选,点一次选中、再点取消)
|
||||
</div>
|
||||
<div className="flex flex-wrap gap-1">
|
||||
{HOURS.map((hour) => (
|
||||
<Chip
|
||||
key={hour}
|
||||
active={scheduleHours.includes(hour)}
|
||||
onClick={() => setScheduleHours((prev) => toggleNumber(prev, hour))}
|
||||
>
|
||||
{pad(hour)}
|
||||
</Chip>
|
||||
))}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="flex items-center gap-2">
|
||||
<span className="text-[10px] font-mono text-cyber-text-secondary">分钟</span>
|
||||
<Select value={scheduleMinute} onValueChange={setScheduleMinute}>
|
||||
<SelectTrigger className="h-7 w-20 text-xs">
|
||||
<SelectValue />
|
||||
</SelectTrigger>
|
||||
<SelectContent className="max-h-56">
|
||||
{MINUTES.map((minute) => (
|
||||
<SelectItem key={minute} value={String(minute)}>
|
||||
{pad(minute)}
|
||||
</SelectItem>
|
||||
))}
|
||||
</SelectContent>
|
||||
</Select>
|
||||
<span className="text-[10px] font-mono text-cyber-text-muted">
|
||||
所有时间点共用
|
||||
</span>
|
||||
</div>
|
||||
|
||||
<div className="flex items-center justify-between gap-3 border-t border-cyber-border-subtle pt-2">
|
||||
<span className="text-[11px] font-mono text-cyber-neon-cyan">
|
||||
{describeSchedule(
|
||||
scheduleMode,
|
||||
Number(intervalMinutes),
|
||||
scheduleHours,
|
||||
scheduleDays,
|
||||
Number(scheduleMinute),
|
||||
)}
|
||||
</span>
|
||||
<span className="text-[10px] font-mono text-cyber-text-muted">
|
||||
按服务器本地时间
|
||||
</span>
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
|
||||
<div className="space-y-2">
|
||||
<Label className="text-xs font-mono text-cyber-text-secondary">
|
||||
{mode === 'creator' ? '博主主页链接或 ID' : '笔记链接或 ID'}
|
||||
{mode === 'creator'
|
||||
? `${hints?.creator_label ?? '博主主页'}链接或 ID`
|
||||
: `${hints?.note_label ?? '笔记'}链接或 ID`}
|
||||
</Label>
|
||||
<textarea
|
||||
value={targets}
|
||||
onChange={(event) => setTargets(event.target.value)}
|
||||
rows={5}
|
||||
placeholder={
|
||||
mode === 'creator'
|
||||
? '每行一个,支持完整主页链接或纯 ID:\nhttps://www.xiaohongshu.com/user/profile/5f58bd99...\n5f58bd990000000001003753'
|
||||
: '每行一个,支持完整笔记链接或纯 ID:\nhttps://www.xiaohongshu.com/explore/6aa3d827...'
|
||||
}
|
||||
placeholder={`每行一个,支持完整链接或纯 ID:\n${(mode === 'creator' ? hints?.creator : hints?.note) ?? ''}`}
|
||||
className={TEXTAREA_CLASS}
|
||||
/>
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
已识别 <span className="text-cyber-neon-cyan">{targetList.length}</span> 个目标。
|
||||
<span className="text-cyber-neon-orange">建议只填纯 ID</span>
|
||||
——链接里的 xsec_token 会过期,纯 ID 永久有效。
|
||||
{/* 「只填纯 ID」是小红书专属的劝告:它链接里的 xsec_token 会过期。
|
||||
抖音的链接不带令牌,永久有效,那句话对它没有意义。 */}
|
||||
{hints?.token_expires && (
|
||||
<>
|
||||
<span className="text-cyber-neon-orange">建议只填纯 ID</span>
|
||||
——链接里的 xsec_token 会过期,纯 ID 永久有效。
|
||||
</>
|
||||
)}
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<div className="grid grid-cols-2 gap-3">
|
||||
<div className="space-y-2">
|
||||
<Label className="text-xs font-mono text-cyber-text-secondary">
|
||||
每轮最多采集作品数
|
||||
{mode === 'creator' ? '每个博主最多采集作品数' : '最多采集作品数'}
|
||||
</Label>
|
||||
<Input
|
||||
type="number"
|
||||
min={1}
|
||||
value={maxNotes}
|
||||
onChange={(event) => setMaxNotes(event.target.value)}
|
||||
disabled={mode === 'note'}
|
||||
className="h-9 text-xs"
|
||||
/>
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
只取最新的前 N 条,决定了"该博主的作品"覆盖范围
|
||||
</p>
|
||||
|
||||
{mode === 'creator' ? (
|
||||
<>
|
||||
{/* 这是「每个博主」的上限,不是一轮的总量 —— 爬虫里这个值是在
|
||||
per-creator 的函数内比较的(client.py get_all_notes_by_creator),
|
||||
而外层 for 循环会遍历全部目标。所以 100 个目标 × 20 篇 = 单轮
|
||||
最多 2000 篇。标签写成「每轮最多」会让人以为超出的会被丢弃。 */}
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
这是<span className="text-cyber-text-secondary">每个博主</span>的上限,
|
||||
不是一轮的总量。只取该博主最新的前 N 条,超出的不会补抓。
|
||||
</p>
|
||||
<p className="text-[10px] font-mono text-cyber-text-secondary">
|
||||
{targetList.length} 个目标 × {maxNotes} 篇 × 每人 1 次
|
||||
→ 单轮最多{' '}
|
||||
<span className="text-cyber-neon-cyan">
|
||||
{targetList.length * (Number(maxNotes) || 0)}
|
||||
</span>{' '}
|
||||
篇
|
||||
</p>
|
||||
{targetList.length * (Number(maxNotes) || 0) > NOTE_VOLUME_WARN && (
|
||||
<p className="text-[10px] font-mono text-cyber-neon-orange leading-relaxed">
|
||||
单轮量偏大:每篇还要抓最多 {maxComments} 条评论,且并发为 1。
|
||||
容易触发平台限流,也可能跑不完就被任务超时(默认 1 小时)中断。
|
||||
建议调低这个数,或拆成几个任务。
|
||||
</p>
|
||||
)}
|
||||
</>
|
||||
) : (
|
||||
<p className="text-[10px] font-mono text-cyber-neon-orange">
|
||||
{hints?.note_label ?? '笔记'}模式下此项不生效:你列出的每个链接都会被逐条抓取。
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
|
||||
<div className="space-y-2">
|
||||
@@ -246,24 +475,55 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="flex items-start gap-2 rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3">
|
||||
<Checkbox
|
||||
id="notify-enabled"
|
||||
checked={notifyEnabled}
|
||||
onCheckedChange={(checked) => setNotifyEnabled(checked === true)}
|
||||
/>
|
||||
<div className="space-y-0.5">
|
||||
<label
|
||||
htmlFor="notify-enabled"
|
||||
className="text-xs font-mono text-cyber-text-primary cursor-pointer"
|
||||
>
|
||||
推送企业微信通知
|
||||
</label>
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
仅在本任务**采集失败 / 登录态失效**或**发现新作品**时推送,
|
||||
一轮只发一条汇总。需先在监控页配置 Webhook 地址。
|
||||
</p>
|
||||
<div className="rounded-md border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3 space-y-3">
|
||||
{/* 两类通知的性质完全不同,所以分成两个开关:
|
||||
异常低频且意味着任务已经停止工作 —— 默认开;
|
||||
新作品可能每轮都有 —— 默认关,否则会刷屏。 */}
|
||||
<div className="flex items-start gap-2">
|
||||
<Checkbox
|
||||
id="notify-failures"
|
||||
checked={notifyFailures}
|
||||
onCheckedChange={(checked) => setNotifyFailures(checked === true)}
|
||||
/>
|
||||
<div className="space-y-0.5">
|
||||
<label
|
||||
htmlFor="notify-failures"
|
||||
className="text-xs font-mono text-cyber-text-primary cursor-pointer"
|
||||
>
|
||||
推送异常通知(建议保持开启)
|
||||
</label>
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
登录态失效、采集进程失败、一篇都没抓到时推送。
|
||||
<span className="text-cyber-neon-cyan">
|
||||
异常意味着这个任务从此默默采不到任何东西
|
||||
</span>
|
||||
—— 关掉的话你不会知道,直到某天发现数据停在几周前。
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="flex items-start gap-2 border-t border-cyber-border-subtle pt-3">
|
||||
<Checkbox
|
||||
id="notify-enabled"
|
||||
checked={notifyEnabled}
|
||||
onCheckedChange={(checked) => setNotifyEnabled(checked === true)}
|
||||
/>
|
||||
<div className="space-y-0.5">
|
||||
<label
|
||||
htmlFor="notify-enabled"
|
||||
className="text-xs font-mono text-cyber-text-primary cursor-pointer"
|
||||
>
|
||||
推送新作品通知
|
||||
</label>
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
发现新作品时推送。监控多个博主时可能每轮都有,容易刷屏,所以默认关闭。
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<p className="border-t border-cyber-border-subtle pt-2 text-[10px] font-mono text-cyber-text-muted">
|
||||
两者都是一轮只发一条汇总。需先在右上角「系统设置」里配置企业微信 Webhook 地址。
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
|
||||
@@ -3,6 +3,7 @@ import { KeyRound, QrCode, Save } from 'lucide-react'
|
||||
|
||||
import { Button } from '@/components/ui/button'
|
||||
import { CookiePanel } from '@/components/monitor/CookiePanel'
|
||||
import { QrLoginPanel } from '@/components/monitor/QrLoginPanel'
|
||||
import { Section, SettingField } from '@/components/settings/SettingFields'
|
||||
import { useSettings, useUpdateSettings } from '@/hooks/useMonitor'
|
||||
import { useCurrentPlatform } from '@/hooks/usePlatform'
|
||||
@@ -66,17 +67,19 @@ export function SettingsView({ onNavigate }: { onNavigate?: (view: AppView) => v
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<Section title="登录态" description={`${label}的登录 Cookie,定时监控必须持久化登录态。`}>
|
||||
<Section
|
||||
title="登录态"
|
||||
description={`${label}的登录态。定时监控必须持久化,扫码一次最省事。`}
|
||||
>
|
||||
{/* QR first: it is the one that keeps working unattended, since the
|
||||
scan lands in the very browser profile the monitor runs reuse. */}
|
||||
<QrLoginPanel />
|
||||
<CookiePanel />
|
||||
|
||||
{/* Not a duplicate login flow: the crawler has no login-only mode, so a
|
||||
QR login is a side effect of a real crawl -- which is exactly what
|
||||
the 采集 page already does. This preselects it rather than
|
||||
reimplementing it. */}
|
||||
<div className="pt-3 mt-1 border-t border-cyber-border-subtle space-y-2">
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
Cookie 不好使时,可以走一次**扫码登录**:二维码会显示在「采集」页的终端里。
|
||||
扫码成功后浏览器 profile 会被更新,Cookie 的可靠性也会显著提升。
|
||||
在本机桌面运行时,也可以走「采集」页的扫码:二维码会打印在终端里。
|
||||
服务器没有显示器,那条路走不通,用上面的面板。
|
||||
</p>
|
||||
<Button
|
||||
variant="outline"
|
||||
@@ -89,7 +92,7 @@ export function SettingsView({ onNavigate }: { onNavigate?: (view: AppView) => v
|
||||
}}
|
||||
>
|
||||
<QrCode className="w-3 h-3 mr-1" />
|
||||
去扫码登录
|
||||
去采集页扫码
|
||||
</Button>
|
||||
</div>
|
||||
</Section>
|
||||
|
||||
@@ -12,6 +12,7 @@ import {
|
||||
} from '@/components/ui/dialog'
|
||||
import { WebhookPanel } from '@/components/monitor/WebhookPanel'
|
||||
import { ChangePassword, SettingField } from '@/components/settings/SettingFields'
|
||||
import { UpstreamPanel } from '@/components/settings/UpstreamPanel'
|
||||
import { useSettings, useUpdateSettings } from '@/hooks/useMonitor'
|
||||
|
||||
type Draft = Record<string, boolean | number | string>
|
||||
@@ -44,6 +45,18 @@ export function SystemSettingsDialog({
|
||||
[data],
|
||||
)
|
||||
|
||||
// 上游那几项单独成块,其余(时段、CDP)仍归在「调度」下。按前缀分流就够了:
|
||||
// 注册表是唯一的来源,新增一项上游设置不需要再动这个文件。写成「排除 upstream_」
|
||||
// 而不是「只取 active_hours_」,这样以后再加系统设置也不会从界面上凭空消失。
|
||||
const upstreamSpecs = useMemo(
|
||||
() => systemSpecs.filter((spec) => spec.name.startsWith('upstream_')),
|
||||
[systemSpecs],
|
||||
)
|
||||
const scheduleSpecs = useMemo(
|
||||
() => systemSpecs.filter((spec) => !spec.name.startsWith('upstream_')),
|
||||
[systemSpecs],
|
||||
)
|
||||
|
||||
const dirty = useMemo(() => {
|
||||
if (!data?.values) return {}
|
||||
const changes: Draft = {}
|
||||
@@ -86,7 +99,28 @@ export function SystemSettingsDialog({
|
||||
调度器全局只有一套时段规则,因此不按平台区分。
|
||||
</p>
|
||||
</div>
|
||||
{systemSpecs.map((spec) => (
|
||||
{scheduleSpecs.map((spec) => (
|
||||
<SettingField
|
||||
key={spec.key}
|
||||
spec={spec}
|
||||
value={draft[spec.key] ?? (spec.default as boolean | number | string)}
|
||||
onChange={(next) => setDraft((prev) => ({ ...prev, [spec.key]: next }))}
|
||||
/>
|
||||
))}
|
||||
</div>
|
||||
|
||||
<div className="space-y-3 border-t border-cyber-border-subtle pt-4">
|
||||
<div>
|
||||
<h3 className="font-mono text-xs tracking-wider text-cyber-text-primary">
|
||||
上游更新
|
||||
</h3>
|
||||
<p className="mt-0.5 text-[10px] font-mono text-cyber-text-muted">
|
||||
本仓库在上游 MediaCrawler 之上加了一整层,这里定期看看上游有没有新提交。
|
||||
改动立即生效,不需要等下一轮采集。
|
||||
</p>
|
||||
</div>
|
||||
<UpstreamPanel />
|
||||
{upstreamSpecs.map((spec) => (
|
||||
<SettingField
|
||||
key={spec.key}
|
||||
spec={spec}
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
import { RefreshCw } from 'lucide-react'
|
||||
|
||||
import { Badge } from '@/components/ui/badge'
|
||||
import { Button } from '@/components/ui/button'
|
||||
import { useCheckUpstream, useUpstreamStatus } from '@/hooks/useMonitor'
|
||||
import { formatDateTime, formatRelative } from '@/lib/monitorFormat'
|
||||
import type { UpstreamStatus } from '@/types/monitor'
|
||||
|
||||
/**
|
||||
* 最近一次上游检查的结果,外加一个「立即检查」。
|
||||
*
|
||||
* 开关与间隔都由设置项本身渲染(同一张注册表),这里只补状态显示:光有一个开关,
|
||||
* 用户没法知道它到底跑没跑、上游到底动没动。后端只缓存结果,不在这里发 fetch ——
|
||||
* 「立即检查」才发,而且那条请求要等 fetch 跑完,所以超时是单独放长的。
|
||||
*/
|
||||
export function UpstreamPanel() {
|
||||
const { data, isLoading } = useUpstreamStatus()
|
||||
const checkNow = useCheckUpstream()
|
||||
|
||||
return (
|
||||
<div className="space-y-2 rounded-lg border border-cyber-border-subtle bg-cyber-bg-tertiary/40 p-3">
|
||||
<div className="flex items-center justify-between gap-3">
|
||||
<div className="flex items-center gap-2">{statusBadge(data)}</div>
|
||||
<Button
|
||||
variant="outline"
|
||||
size="sm"
|
||||
disabled={checkNow.isPending}
|
||||
onClick={() => checkNow.mutate()}
|
||||
>
|
||||
<RefreshCw className={`w-3 h-3 mr-1 ${checkNow.isPending ? 'animate-spin' : ''}`} />
|
||||
{checkNow.isPending ? '检查中…' : '立即检查'}
|
||||
</Button>
|
||||
</div>
|
||||
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted leading-relaxed">
|
||||
{isLoading ? '读取中…' : describe(data)}
|
||||
</p>
|
||||
|
||||
{data?.commits && data.commits.length > 0 && (
|
||||
<ul className="space-y-1 border-t border-cyber-border-subtle pt-2">
|
||||
{data.commits.slice(0, 10).map((commit) => (
|
||||
<li key={commit.sha} className="flex gap-2 text-[10px] font-mono leading-relaxed">
|
||||
<span className="text-cyber-text-muted shrink-0">{commit.sha}</span>
|
||||
<span className="text-cyber-text-secondary truncate" title={commit.subject}>
|
||||
{commit.subject}
|
||||
</span>
|
||||
</li>
|
||||
))}
|
||||
</ul>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
function statusBadge(data?: UpstreamStatus) {
|
||||
if (!data || !data.checked_at) return <Badge variant="idle">尚未检查</Badge>
|
||||
if (!data.ok) return <Badge variant="destructive">检查失败</Badge>
|
||||
if ((data.behind ?? 0) > 0) return <Badge variant="warning">落后 {data.behind} 个提交</Badge>
|
||||
return <Badge variant="success">已是最新</Badge>
|
||||
}
|
||||
|
||||
function describe(data?: UpstreamStatus): string {
|
||||
if (!data || !data.checked_at) {
|
||||
return '还没有检查过。开启上面的开关会按间隔自动查,也可以点「立即检查」。'
|
||||
}
|
||||
|
||||
const when = `上次检查:${formatDateTime(data.checked_at)}(${formatRelative(data.checked_at)})`
|
||||
if (!data.ok) return `${when};${data.error ?? '未知错误'}`
|
||||
|
||||
const parts = [when, `上游 ${data.branch ?? 'main'} 领先 ${data.behind ?? 0} 个提交`]
|
||||
// 领先数就是我们自己这一层的规模;合并时要保留的东西,值得一并说出来。
|
||||
if (data.ahead) parts.push(`本仓库另有 ${data.ahead} 个自己的提交`)
|
||||
if (data.notify_error) parts.push(`通知发送失败:${data.notify_error}`)
|
||||
return parts.join(';')
|
||||
}
|
||||
@@ -0,0 +1,99 @@
|
||||
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query'
|
||||
import { toast } from 'sonner'
|
||||
|
||||
import { creatorApi } from '@/lib/api'
|
||||
|
||||
const ACCOUNTS_KEY = ['creatorAccounts']
|
||||
|
||||
/** 账号列表。同步在后台跑,所以列表本身也顺带轮询,好让状态自己刷新出来。 */
|
||||
export function useCreatorAccounts(polling = false) {
|
||||
return useQuery({
|
||||
queryKey: ACCOUNTS_KEY,
|
||||
queryFn: async () => (await creatorApi.listAccounts()).data.accounts,
|
||||
refetchInterval: polling ? 5000 : false,
|
||||
})
|
||||
}
|
||||
|
||||
export function useCreatorAccount(id: number | null) {
|
||||
return useQuery({
|
||||
queryKey: ['creatorAccount', id],
|
||||
queryFn: async () => (await creatorApi.getAccount(id as number)).data,
|
||||
enabled: id !== null,
|
||||
})
|
||||
}
|
||||
|
||||
export function useDeleteCreatorAccount() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: (id: number) => creatorApi.deleteAccount(id),
|
||||
onSuccess: () => {
|
||||
toast.success('账号已删除')
|
||||
queryClient.invalidateQueries({ queryKey: ACCOUNTS_KEY })
|
||||
},
|
||||
onError: (error: Error) => toast.error(`删除失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
export function useCheckCreatorAccount() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: (id: number) => creatorApi.checkAccount(id),
|
||||
onSuccess: () => {
|
||||
queryClient.invalidateQueries({ queryKey: ACCOUNTS_KEY })
|
||||
},
|
||||
onError: (error: Error) => toast.error(`检测失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
export function useSyncCreatorAccount() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: ({ id, days }: { id: number; days?: number }) =>
|
||||
creatorApi.syncAccount(id, days),
|
||||
onSuccess: () => {
|
||||
// 后端是后台任务,立刻重新拉一次只会看到旧状态;给用户一句"已开始",
|
||||
// 列表的轮询会把结果带回来。
|
||||
toast.success('同步已开始,稍后自动刷新')
|
||||
queryClient.invalidateQueries({ queryKey: ACCOUNTS_KEY })
|
||||
},
|
||||
onError: (error: Error) => toast.error(`同步失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
// --- 扫码新增账号 ---------------------------------------------------------
|
||||
|
||||
/** 只在扫码中轮询;空闲时没必要一直问。 */
|
||||
export function useCreatorLoginStatus(polling: boolean) {
|
||||
return useQuery({
|
||||
queryKey: ['creatorLogin'],
|
||||
queryFn: async () => (await creatorApi.getLogin()).data,
|
||||
refetchInterval: polling ? 2000 : false,
|
||||
})
|
||||
}
|
||||
|
||||
export function useStartCreatorLogin() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: () => creatorApi.startLogin(),
|
||||
onSuccess: (response) => {
|
||||
queryClient.setQueryData(['creatorLogin'], response.data)
|
||||
},
|
||||
onError: (error: Error) => {
|
||||
// 后端给的原因是这里唯一有价值的信息(连不上 Chrome、页面上没有二维码),
|
||||
// 而 axios 会把它压成 "Request failed with status code 502"。
|
||||
const detail = (error as { response?: { data?: { detail?: string } } })?.response?.data
|
||||
?.detail
|
||||
toast.error(`获取二维码失败:${detail ?? error.message}`)
|
||||
},
|
||||
})
|
||||
}
|
||||
|
||||
export function useCancelCreatorLogin() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: () => creatorApi.cancelLogin(),
|
||||
onSuccess: (response) => {
|
||||
queryClient.setQueryData(['creatorLogin'], response.data)
|
||||
},
|
||||
})
|
||||
}
|
||||
@@ -97,12 +97,19 @@ export function useTaskRuns(taskId: number | null) {
|
||||
})
|
||||
}
|
||||
|
||||
/**
|
||||
* 作品**和博主**一起取。
|
||||
*
|
||||
* 博主不能从作品推出来 —— 「一条作品都没有的博主」在作品表里根本不存在,而它恰恰是
|
||||
* 最该显示的一类。所以服务端单独给一份 `creators`,两者必须来自同一次请求,否则
|
||||
* 组头和组里的作品会对不上(比如刚删掉一个任务)。
|
||||
*/
|
||||
export function useMonitorNotes(taskId: number | null, onlyNew: boolean) {
|
||||
const platform = usePlatformParam()
|
||||
return useQuery({
|
||||
queryKey: ['monitorNotes', platform, taskId, onlyNew],
|
||||
queryFn: async () =>
|
||||
(await monitorApi.getNotes(taskId ?? undefined, onlyNew, 200, platform)).data.notes,
|
||||
(await monitorApi.getNotes(taskId ?? undefined, onlyNew, 200, platform)).data,
|
||||
refetchInterval: POLL_MS,
|
||||
})
|
||||
}
|
||||
@@ -204,6 +211,74 @@ export function useClearCookie() {
|
||||
})
|
||||
}
|
||||
|
||||
// --- QR login (over CDP) ---------------------------------------------------
|
||||
|
||||
/** Polls only while a code is on screen; an idle session has nothing to watch. */
|
||||
export function useQrLoginStatus(polling: boolean) {
|
||||
return useQuery({
|
||||
queryKey: ['monitorQrLogin'],
|
||||
queryFn: async () => (await monitorApi.getQrLogin()).data,
|
||||
refetchInterval: polling ? 2000 : false,
|
||||
})
|
||||
}
|
||||
|
||||
export function useStartQrLogin() {
|
||||
const queryClient = useQueryClient()
|
||||
const platform = usePlatformParam()
|
||||
return useMutation({
|
||||
mutationFn: () => monitorApi.startQrLogin(platform),
|
||||
onSuccess: (response) => {
|
||||
queryClient.setQueryData(['monitorQrLogin'], response.data)
|
||||
},
|
||||
onError: (error: Error) => {
|
||||
// The backend's reason -- "Chrome is not reachable on ...", "no QR on the
|
||||
// page" -- is the entire value here, and axios would flatten it to
|
||||
// "Request failed with status code 502".
|
||||
const detail = (error as { response?: { data?: { detail?: string } } })?.response
|
||||
?.data?.detail
|
||||
toast.error(`获取二维码失败:${detail ?? error.message}`)
|
||||
},
|
||||
})
|
||||
}
|
||||
|
||||
export function useCancelQrLogin() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: () => monitorApi.cancelQrLogin(),
|
||||
onSuccess: (response) => {
|
||||
queryClient.setQueryData(['monitorQrLogin'], response.data)
|
||||
},
|
||||
})
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether the browser is actually signed in.
|
||||
*
|
||||
* Tracked separately from the QR session above: that session lives in the
|
||||
* server's memory and a restart erases it, while the profile it wrote to does
|
||||
* not. Polled slowly when idle and briskly while a code is on screen.
|
||||
*/
|
||||
export function useLoginState(polling: boolean) {
|
||||
return useQuery({
|
||||
queryKey: ['monitorLoginState'],
|
||||
queryFn: async () => (await monitorApi.getLoginState()).data,
|
||||
refetchInterval: polling ? 4000 : 30000,
|
||||
refetchOnWindowFocus: false,
|
||||
})
|
||||
}
|
||||
|
||||
/** Re-ask after reloading the page, for when the state looks stale. */
|
||||
export function useRecheckLogin() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: () => monitorApi.getLoginState(true),
|
||||
onSuccess: (response) => {
|
||||
queryClient.setQueryData(['monitorLoginState'], response.data)
|
||||
},
|
||||
onError: (error: Error) => toast.error(`检测失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
// --- Settings -------------------------------------------------------------
|
||||
|
||||
export function useSettings() {
|
||||
@@ -284,3 +359,67 @@ export function useTestWebhook() {
|
||||
onError: (error: Error) => toast.error(`发送失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
// --- 博主备注 --------------------------------------------------------------
|
||||
|
||||
export function useSetCreatorAlias() {
|
||||
const queryClient = useQueryClient()
|
||||
const platform = usePlatformParam()
|
||||
return useMutation({
|
||||
mutationFn: ({ creatorHash, alias }: { creatorHash: string; alias: string }) =>
|
||||
monitorApi.setCreatorAlias(creatorHash, alias, platform),
|
||||
onSuccess: () => {
|
||||
toast.success('备注已保存')
|
||||
queryClient.invalidateQueries({ queryKey: ['monitorNotes'] })
|
||||
queryClient.invalidateQueries({ queryKey: ['monitorCommentsGrouped'] })
|
||||
},
|
||||
onError: (error: Error) => toast.error(`备注保存失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
// --- 作品备注 --------------------------------------------------------------
|
||||
|
||||
export function useSetNoteAlias() {
|
||||
const queryClient = useQueryClient()
|
||||
const platform = usePlatformParam()
|
||||
return useMutation({
|
||||
mutationFn: ({ noteId, alias }: { noteId: string; alias: string }) =>
|
||||
monitorApi.setNoteAlias(noteId, alias, platform),
|
||||
onSuccess: () => {
|
||||
toast.success('备注已保存')
|
||||
queryClient.invalidateQueries({ queryKey: ['monitorNotes'] })
|
||||
queryClient.invalidateQueries({ queryKey: ['monitorCommentsGrouped'] })
|
||||
},
|
||||
onError: (error: Error) => toast.error(`备注保存失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
// --- 上游更新检查 -----------------------------------------------------------------------------------------------------------------
|
||||
|
||||
export function useUpstreamStatus() {
|
||||
return useQuery({
|
||||
queryKey: ['monitorUpstream'],
|
||||
queryFn: async () => (await monitorApi.getUpstream()).data,
|
||||
// 检查本身是每天一次的量级,页面开着时慢点刷就够了。
|
||||
refetchInterval: 60_000,
|
||||
})
|
||||
}
|
||||
|
||||
export function useCheckUpstream() {
|
||||
const queryClient = useQueryClient()
|
||||
return useMutation({
|
||||
mutationFn: () => monitorApi.checkUpstream(),
|
||||
onSuccess: (response) => {
|
||||
const data = response.data
|
||||
if (!data.ok) {
|
||||
toast.error(`检查失败:${data.error ?? '未知错误'}`)
|
||||
} else if ((data.behind ?? 0) > 0) {
|
||||
toast.success(`上游有 ${data.behind} 个新提交`)
|
||||
} else {
|
||||
toast.success('已是最新,上游没有新提交')
|
||||
}
|
||||
queryClient.invalidateQueries({ queryKey: ['monitorUpstream'] })
|
||||
},
|
||||
onError: (error: Error) => toast.error(`检查失败:${error.message}`),
|
||||
})
|
||||
}
|
||||
|
||||
+64
-1
@@ -5,17 +5,27 @@ import type {
|
||||
CookieStatus,
|
||||
MetricPoint,
|
||||
MonitorComment,
|
||||
MonitorCreator,
|
||||
MonitorEvent,
|
||||
MonitorNote,
|
||||
MonitorOverview,
|
||||
MonitorRun,
|
||||
MonitorTask,
|
||||
LoginState,
|
||||
PlatformCapability,
|
||||
QrLoginState,
|
||||
ReportResult,
|
||||
SettingsResponse,
|
||||
TaskCreatePayload,
|
||||
UpstreamStatus,
|
||||
WebhookStatus,
|
||||
} from '@/types/monitor'
|
||||
import type {
|
||||
CreatorAccount,
|
||||
CreatorAccountDetail,
|
||||
CreatorLoginState,
|
||||
CreatorSyncResult,
|
||||
} from '@/types/creator'
|
||||
|
||||
const api = axios.create({
|
||||
baseURL: '/api',
|
||||
@@ -177,7 +187,7 @@ export const monitorApi = {
|
||||
api.get<{ runs: MonitorRun[] }>(`/monitor/tasks/${id}/runs`, { params: { limit } }),
|
||||
|
||||
getNotes: (taskId?: number, onlyNew = false, limit = 200, platform?: string) =>
|
||||
api.get<{ notes: MonitorNote[] }>('/monitor/notes', {
|
||||
api.get<{ notes: MonitorNote[]; creators: MonitorCreator[] }>('/monitor/notes', {
|
||||
params: { task_id: taskId, only_new: onlyNew, limit, platform },
|
||||
}),
|
||||
// Metric time series for one note; the chart reads this.
|
||||
@@ -256,10 +266,63 @@ export const monitorApi = {
|
||||
return api.get<ReportResult>(`/monitor/report?${params.toString()}`)
|
||||
},
|
||||
|
||||
// QR login over CDP. Reads the code out of the browser the crawler attaches
|
||||
// to, so the operator can scan even when the host has no display.
|
||||
startQrLogin: (platform?: string) =>
|
||||
api.post<QrLoginState>('/monitor/login/qr', null, { params: { platform } }),
|
||||
getQrLogin: () => api.get<QrLoginState>('/monitor/login/qr'),
|
||||
cancelQrLogin: () => api.delete<QrLoginState>('/monitor/login/qr'),
|
||||
/** `force` reloads the page first, for a stale-looking state. */
|
||||
getLoginState: (force = false) =>
|
||||
api.get<LoginState>('/monitor/login/state', { params: { force } }),
|
||||
|
||||
getWebhook: () => api.get<WebhookStatus>('/monitor/webhook'),
|
||||
setWebhook: (url: string) => api.post('/monitor/webhook', { url }),
|
||||
clearWebhook: () => api.delete('/monitor/webhook'),
|
||||
testWebhook: (url?: string) => api.post('/monitor/webhook/test', { url: url ?? null }),
|
||||
|
||||
/** 最近一次上游检查的缓存结果;没有查过时是空对象。 */
|
||||
getUpstream: () => api.get<UpstreamStatus>('/monitor/upstream'),
|
||||
|
||||
/** 给博主起备注(空串 = 清掉)。键是 creator_hash,跨任务同一个博主共用一条。 */
|
||||
setCreatorAlias: (creatorHash: string, alias: string, platform?: string) =>
|
||||
api.put(
|
||||
`/monitor/creators/${encodeURIComponent(creatorHash)}`,
|
||||
{ alias },
|
||||
{ params: { platform } },
|
||||
),
|
||||
/** 给作品起备注(空串 = 清掉)。键是 note_id —— 和博主备注是两回事。 */
|
||||
setNoteAlias: (noteId: string, alias: string, platform?: string) =>
|
||||
api.put(
|
||||
`/monitor/notes/${encodeURIComponent(noteId)}`,
|
||||
{ alias },
|
||||
{ params: { platform } },
|
||||
),
|
||||
/**
|
||||
* 立刻检查一次。服务端要等 fetch 跑完才回答,而 axios 默认 30 秒对此不够 ——
|
||||
* 一条卡住的 git fetch 能拖到两分钟,这里必须单独放长超时,否则会误报失败。
|
||||
*/
|
||||
checkUpstream: () =>
|
||||
api.post<UpstreamStatus>('/monitor/upstream/check', null, { timeout: 150_000 }),
|
||||
}
|
||||
|
||||
/**
|
||||
* 运营模块。与 `monitorApi` 并列 —— 它管的是自己的账号,走的是纯请求的创作者后台,
|
||||
* 和公开数据监控不是一回事。
|
||||
*/
|
||||
export const creatorApi = {
|
||||
listAccounts: () => api.get<{ accounts: CreatorAccount[] }>('/creator/accounts'),
|
||||
getAccount: (id: number) => api.get<CreatorAccountDetail>(`/creator/accounts/${id}`),
|
||||
deleteAccount: (id: number) => api.delete(`/creator/accounts/${id}`),
|
||||
checkAccount: (id: number) => api.post<CreatorAccount>(`/creator/accounts/${id}/check`),
|
||||
// 同步在后端后台跑(分页 + 节流可能几分钟),这里只负责触发。
|
||||
syncAccount: (id: number, days = 90) =>
|
||||
api.post<CreatorSyncResult>(`/creator/accounts/${id}/sync`, null, { params: { days } }),
|
||||
|
||||
// 扫码新增账号。cookie 只在后端内存里流转,不会出现在这些响应里。
|
||||
startLogin: () => api.post<CreatorLoginState>('/creator/login'),
|
||||
getLogin: () => api.get<CreatorLoginState>('/creator/login'),
|
||||
cancelLogin: () => api.delete<CreatorLoginState>('/creator/login'),
|
||||
}
|
||||
|
||||
export default api
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
/** Formatting helpers for the monitoring dashboard. */
|
||||
|
||||
import type { ScheduleMode } from '@/types/monitor'
|
||||
|
||||
/** Compact count for display: mirrors how the platform itself abbreviates. */
|
||||
export function formatCount(value: number | null | undefined): string {
|
||||
if (value === null || value === undefined) return '—'
|
||||
@@ -36,9 +38,56 @@ export function formatDateTime(ms: number | null | undefined): string {
|
||||
return `${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())} ${pad(d.getHours())}:${pad(d.getMinutes())}`
|
||||
}
|
||||
|
||||
/**
|
||||
* 只到日。
|
||||
*
|
||||
* 发布日期问的是「哪一天发的」,绝对日期比「3天前」好认 —— 后者每天看都在变,而且
|
||||
* 没法跟平台上的日期对。具体到分钟的那份放 title 里,需要时悬停看。
|
||||
*/
|
||||
export function formatDate(ms: number | null | undefined): string {
|
||||
if (!ms) return '—'
|
||||
const d = new Date(ms)
|
||||
const pad = (n: number) => String(n).padStart(2, '0')
|
||||
return `${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())}`
|
||||
}
|
||||
|
||||
/** Human interval label for task cards. */
|
||||
export function formatInterval(minutes: number): string {
|
||||
if (minutes % 1440 === 0) return `${minutes / 1440} 天`
|
||||
if (minutes % 60 === 0) return `${minutes / 60} 小时`
|
||||
return `${minutes} 分钟`
|
||||
}
|
||||
|
||||
const WEEKDAY_NAMES = '一二三四五六日'
|
||||
|
||||
/**
|
||||
* Preview of a schedule, for the editor only.
|
||||
*
|
||||
* The task list renders the backend's own `schedule_label` instead. This exists
|
||||
* so the current selection can be read back before it is saved -- which means the
|
||||
* two are the same sentence written twice, and a change to one belongs in both.
|
||||
*/
|
||||
export function describeSchedule(
|
||||
mode: ScheduleMode,
|
||||
intervalMinutes: number,
|
||||
hours: number[],
|
||||
days: number[],
|
||||
minute: number,
|
||||
): string {
|
||||
const pad = (value: number) => String(value).padStart(2, '0')
|
||||
const sorted = (values: number[]) => [...values].sort((a, b) => a - b)
|
||||
|
||||
if (mode === 'interval') return `每 ${formatInterval(intervalMinutes)}`
|
||||
if (hours.length === 0) return '未设置时间'
|
||||
|
||||
const clock = sorted(hours)
|
||||
.map((hour) => `${pad(hour)}:${pad(minute)}`)
|
||||
.join('、')
|
||||
|
||||
if (mode === 'daily' || days.length === 0) return `每天 ${clock}`
|
||||
|
||||
const labels = sorted(days)
|
||||
.map((day) => `周${WEEKDAY_NAMES[day]}`)
|
||||
.join('、')
|
||||
return `${labels} ${clock}`
|
||||
}
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
/** 运营模块:自己的小红书账号,以及创作者后台给的数据。 */
|
||||
|
||||
/**
|
||||
* 数据权限状态。
|
||||
*
|
||||
* `pending` 是实测中最常见的一种:后台原话是「已为您申请数据权限,次日可查看」。
|
||||
* 它必须和「没有权限」分开显示 —— 否则用户会以为采集坏了,其实只是在等审批。
|
||||
*/
|
||||
export type CreatorPermissionStatus = 'unknown' | 'pending' | 'active' | 'missing'
|
||||
|
||||
export type CreatorAccountStatus = 'ok' | 'expired' | 'error'
|
||||
|
||||
/** 注意:接口**不下发 cookie**,只给 `has_cookie`。 */
|
||||
export interface CreatorAccount {
|
||||
id: number
|
||||
nickname: string
|
||||
user_id: string
|
||||
red_id: string
|
||||
avatar: string
|
||||
status: CreatorAccountStatus
|
||||
permission_status: CreatorPermissionStatus
|
||||
/** 后台原话,照抄不改写。 */
|
||||
permission_tip: string
|
||||
last_checked_at: number | null
|
||||
last_synced_at: number | null
|
||||
/** 上次同步用的范围(天)。0 表示从未同步过。 */
|
||||
last_sync_days: number
|
||||
last_error: string | null
|
||||
has_cookie: boolean
|
||||
created_at: number
|
||||
note_count: number
|
||||
}
|
||||
|
||||
/**
|
||||
* 一篇作品的运营数据。
|
||||
*
|
||||
* 除时间外全部可为 null:接口没给、或解析不出来,都存 null 而不是 0 ——
|
||||
* 0 是真实值,null 是"不知道",混在一起报表会说谎。
|
||||
*/
|
||||
export interface CreatorNote {
|
||||
note_id: string
|
||||
title: string
|
||||
publish_time: number | null
|
||||
exposure: number | null
|
||||
views: number | null
|
||||
likes: number | null
|
||||
comments: number | null
|
||||
favorites: number | null
|
||||
shares: number | null
|
||||
new_followers: number | null
|
||||
danmaku: number | null
|
||||
cover_ctr: number | null
|
||||
avg_watch_seconds: number | null
|
||||
two_second_exit_rate: number | null
|
||||
completion_rate: number | null
|
||||
captured_at: number
|
||||
}
|
||||
|
||||
export interface CreatorSummary {
|
||||
exposure: number
|
||||
views: number
|
||||
likes: number
|
||||
comments: number
|
||||
favorites: number
|
||||
shares: number
|
||||
new_followers: number
|
||||
}
|
||||
|
||||
export interface CreatorAccountDetail {
|
||||
account: CreatorAccount
|
||||
notes: CreatorNote[]
|
||||
summary: CreatorSummary
|
||||
}
|
||||
|
||||
export type CreatorLoginStatus = 'idle' | 'waiting' | 'success' | 'expired' | 'error'
|
||||
|
||||
export interface CreatorLoginState {
|
||||
status: CreatorLoginStatus
|
||||
message: string
|
||||
/** `data:image/...;base64,...`, empty unless status is `waiting`. */
|
||||
image: string
|
||||
elapsed: number
|
||||
expires_in: number
|
||||
/** 扫码成功后后端已落库的账号。 */
|
||||
account: CreatorAccount | null
|
||||
}
|
||||
|
||||
export interface CreatorSyncResult {
|
||||
message: string
|
||||
days: number
|
||||
}
|
||||
+196
-1
@@ -28,6 +28,14 @@ export interface MonitorTarget {
|
||||
enabled: boolean
|
||||
}
|
||||
|
||||
/**
|
||||
* How a task is scheduled.
|
||||
*
|
||||
* All three are expressible with pickers. A cron string is deliberately not
|
||||
* supported -- it is a small language to learn just to say "every day at nine".
|
||||
*/
|
||||
export type ScheduleMode = 'interval' | 'daily' | 'weekly'
|
||||
|
||||
export interface MonitorTask {
|
||||
id: number
|
||||
name: string
|
||||
@@ -35,12 +43,23 @@ export interface MonitorTask {
|
||||
mode: MonitorMode
|
||||
enabled: boolean
|
||||
interval_minutes: number
|
||||
schedule_mode: ScheduleMode
|
||||
/** 0-23. Empty in interval mode. */
|
||||
schedule_hours: number[]
|
||||
/** 0-6 with Monday = 0, matching Python's `date.weekday()`. Weekly only. */
|
||||
schedule_days: number[]
|
||||
/** Minute past the hour, shared by every time in the schedule. */
|
||||
schedule_minute: number
|
||||
/** Composed by the backend, so the list and the editor cannot disagree. */
|
||||
schedule_label: string
|
||||
max_notes_count: number
|
||||
enable_comments: boolean
|
||||
max_comments_count: number
|
||||
run_timeout_seconds: number
|
||||
/** Opt-in per task so one webhook does not get flooded. */
|
||||
/** 推送**新作品**。可能每轮都有,默认关以免刷屏。 */
|
||||
notify_enabled: boolean
|
||||
/** 推送**异常**(登录失效/运行失败/没抓到数据)。默认开。 */
|
||||
notify_failures: boolean
|
||||
/** Epoch milliseconds. */
|
||||
next_run_at: number | null
|
||||
last_run_at: number | null
|
||||
@@ -67,6 +86,47 @@ export interface MonitorNote {
|
||||
title: string
|
||||
note_url: string
|
||||
cover: string
|
||||
/**
|
||||
* 创作者标识。
|
||||
*
|
||||
* `creator_hash` 是唯一稳定的分组依据 —— 爬虫刻意不落原始 user_id(见
|
||||
* tools/user_hash.py),所以没有比它更具体的身份了。
|
||||
* `creator_name` 是昵称本身(本仓库关掉了脱敏,见 config.MASK_NICKNAME)。
|
||||
*/
|
||||
creator_hash: string
|
||||
creator_name: string
|
||||
/**
|
||||
* 人自己给这个博主起的备注。
|
||||
*
|
||||
* 界面上**优先显示它**:`creator_name` 是平台昵称(常常认不出是谁),`creator_hash`
|
||||
* 更认不出。备注是唯一能把账号对上人的东西。空串表示没起过。
|
||||
*/
|
||||
creator_alias: string
|
||||
/**
|
||||
* 人给**这条作品**起的备注。
|
||||
*
|
||||
* 和 `creator_alias` 是两件事:那条回答「这个账号是谁」,这条回答「这条我要盯着」。
|
||||
* 一个博主底下常常只有一两件值得盯的作品,所以不能合并成一条。空串表示没起过。
|
||||
*/
|
||||
note_alias: string
|
||||
/**
|
||||
* 博主**账号级**指标 —— 作品列表给不了的东西:作品说的是「这条涨了多少赞」,
|
||||
* 它说的是「这个人整个账号在涨还是在掉」。
|
||||
*
|
||||
* 三者都可能为 null:平台没采到(小红书那条路根本不产生它),或者某一项平台没给。
|
||||
* 为 null 时界面整块不画 —— 画成「粉丝 0」就是在撒谎。
|
||||
*
|
||||
* 作品栏的组头**不要读这里**,读 `MonitorCreator` —— 一条作品都没有的博主不会出现
|
||||
* 在作品列表里,只有那份博主列表能覆盖他。这里是给「顺着一条作品问它的作者」用的。
|
||||
*/
|
||||
creator_fans: number | null
|
||||
creator_total_favorited: number | null
|
||||
creator_works: number | null
|
||||
/** 上述指标是哪一轮采到的。null = 从来没有过。 */
|
||||
creator_stats_at: number | null
|
||||
/** 作品的**发布**时间(爬虫侧:小红书 time、抖音 create_time)。可能与
|
||||
* `first_seen_at` 差很远 —— 后者是「我们第一次看到它」的时间。平台没给时为 null。 */
|
||||
published_at: number | null
|
||||
first_seen_at: number
|
||||
last_seen_at: number
|
||||
is_new: boolean
|
||||
@@ -76,6 +136,30 @@ export interface MonitorNote {
|
||||
snapshot_count: number
|
||||
}
|
||||
|
||||
/**
|
||||
* 作品栏里的一位博主 —— **不依赖于他有没有作品**。
|
||||
*
|
||||
* 作品表是按 creator_hash 从作品推出来的,所以「一条作品都没有的博主」在那边根本
|
||||
* 不存在。可这类博主恰恰是最该看见的:还在涨粉,只是最近没发东西。服务端因此单独
|
||||
* 给一份列表,来源是**账号快照 ∪ 作品**。
|
||||
*
|
||||
* 没有作品时 `note_count` 是 0,其余字段照常有;反过来,小红书那条路不产生账号
|
||||
* 快照,于是 `creator_fans` 等为 null —— 两边各缺一块,界面要都能显示。
|
||||
*/
|
||||
export interface MonitorCreator {
|
||||
task_id: number
|
||||
creator_hash: string
|
||||
creator_name: string
|
||||
creator_alias: string
|
||||
note_count: number
|
||||
creator_fans: number | null
|
||||
creator_total_favorited: number | null
|
||||
creator_works: number | null
|
||||
creator_stats_at: number | null
|
||||
/** 作品最后出现、或账号指标最后采集的时间 —— 排序用。 */
|
||||
last_activity_at: number
|
||||
}
|
||||
|
||||
export interface MetricPoint {
|
||||
run_id: number
|
||||
captured_at: number
|
||||
@@ -99,6 +183,11 @@ export interface MonitorComment {
|
||||
note_title: string
|
||||
note_cover: string
|
||||
note_url: string
|
||||
/** 所属作品的创作者 —— 评论流按 博主 -> 作品 -> 评论 三级展开时用。 */
|
||||
note_creator_hash: string
|
||||
note_creator_name: string
|
||||
/** 所属作品的发布时间。 */
|
||||
note_published_at: number | null
|
||||
}
|
||||
|
||||
/** One work that has comments, for the filter dropdown. */
|
||||
@@ -107,6 +196,8 @@ export interface CommentNoteOption {
|
||||
note_title: string
|
||||
note_cover: string
|
||||
note_url: string
|
||||
creator_hash: string
|
||||
creator_name: string
|
||||
comment_count: number
|
||||
latest_at: number
|
||||
}
|
||||
@@ -117,6 +208,10 @@ export interface CommentBucket {
|
||||
note_title: string
|
||||
note_cover: string
|
||||
note_url: string
|
||||
creator_hash: string
|
||||
creator_name: string
|
||||
/** 作品的发布时间。 */
|
||||
published_at: number | null
|
||||
comments: MonitorComment[]
|
||||
}
|
||||
|
||||
@@ -161,6 +256,48 @@ export interface CookieStatus {
|
||||
last_ok_at: number | null
|
||||
}
|
||||
|
||||
export type QrLoginStatus = 'idle' | 'waiting' | 'success' | 'expired' | 'error'
|
||||
|
||||
/**
|
||||
* A QR login driven over CDP, for servers with no display to show one on.
|
||||
*
|
||||
* `success` also covers "the browser was already signed in" — the scan and the
|
||||
* existing session both just mean the profile the crawler attaches to is usable.
|
||||
*/
|
||||
export interface QrLoginState {
|
||||
status: QrLoginStatus
|
||||
platform: string | null
|
||||
/** `data:image/...;base64,...`. Empty unless status is `waiting`. */
|
||||
image: string
|
||||
message: string
|
||||
/** Seconds since the code was fetched. */
|
||||
elapsed: number
|
||||
/** Seconds left before the code is written off. */
|
||||
expires_in: number
|
||||
/** Whether the browser reports being signed in right now. */
|
||||
logged_in: boolean
|
||||
nickname: string | null
|
||||
/** 登录态是否已同时存进库(供非 CDP 模式的定时任务使用)。 */
|
||||
cookie_saved?: boolean
|
||||
}
|
||||
|
||||
/**
|
||||
* The browser's own answer to "am I signed in?".
|
||||
*
|
||||
* Asked of the page, not derived from the QR session -- that session lives in the
|
||||
* server's memory and dies on a restart, so tying the answer to it makes a
|
||||
* successful scan look like nothing happened.
|
||||
*
|
||||
* `known` is false when the page could not report (not loaded, browser
|
||||
* unreachable); `logged_in` is then meaningless rather than false.
|
||||
*/
|
||||
export interface LoginState {
|
||||
known: boolean
|
||||
logged_in: boolean
|
||||
nickname: string | null
|
||||
error?: string
|
||||
}
|
||||
|
||||
export interface MonitorOverview {
|
||||
tasks: number
|
||||
enabled_tasks: number
|
||||
@@ -173,14 +310,26 @@ export interface MonitorOverview {
|
||||
|
||||
export interface TaskCreatePayload {
|
||||
name: string
|
||||
/**
|
||||
* 任务归属的平台。
|
||||
*
|
||||
* **必须显式带上。** 后端在缺省时会退回小红书 —— 那是接口早期的兼容行为,
|
||||
* 于是「在抖音页面上建任务」会安安静静地建出一个小红书任务(遇到过)。
|
||||
*/
|
||||
platform: string
|
||||
mode: MonitorMode
|
||||
interval_minutes: number
|
||||
schedule_mode: ScheduleMode
|
||||
schedule_hours: number[]
|
||||
schedule_days: number[]
|
||||
schedule_minute: number
|
||||
max_notes_count: number
|
||||
enable_comments: boolean
|
||||
max_comments_count: number
|
||||
run_timeout_seconds: number
|
||||
enabled: boolean
|
||||
notify_enabled: boolean
|
||||
notify_failures: boolean
|
||||
targets: string[]
|
||||
}
|
||||
|
||||
@@ -236,6 +385,17 @@ export interface ReportResult {
|
||||
* layer has been hooked up for it. The UI must never conflate the two -- a
|
||||
* platform can be fully crawlable upstream and still unusable here.
|
||||
*/
|
||||
/** 目标输入框的示例与措辞,随平台变。服务端给,前端不自己判断平台。 */
|
||||
export interface TargetHints {
|
||||
creator?: string
|
||||
note?: string
|
||||
/** 该平台怎么称呼这两样东西 —— 抖音叫「作品」,小红书叫「笔记」。 */
|
||||
creator_label?: string
|
||||
note_label?: string
|
||||
/** 链接里是否带会过期的令牌(只有小红书有)。 */
|
||||
token_expires?: boolean
|
||||
}
|
||||
|
||||
export interface PlatformCapability {
|
||||
value: string
|
||||
label: string
|
||||
@@ -245,6 +405,7 @@ export interface PlatformCapability {
|
||||
comment_levels: number
|
||||
media: boolean
|
||||
monitor_wired: boolean
|
||||
target_hints?: TargetHints
|
||||
}
|
||||
|
||||
export interface WebhookStatus {
|
||||
@@ -297,3 +458,37 @@ export interface SettingsResponse {
|
||||
secrets: Record<string, SecretStatus>
|
||||
specs: SettingSpec[]
|
||||
}
|
||||
|
||||
// --- 上游更新检查 -----------------------------------------------------------
|
||||
|
||||
export interface UpstreamCommit {
|
||||
sha: string
|
||||
author: string
|
||||
/** 提交日期,`YYYY-MM-DD`(git 侧已格式化,不做本地化)。 */
|
||||
date: string
|
||||
subject: string
|
||||
}
|
||||
|
||||
/**
|
||||
* 最近一次上游检查的结果,服务端缓存。
|
||||
*
|
||||
* `checked_at` 为空表示还没查过(`GET /monitor/upstream` 返回空对象);`ok` 为
|
||||
* false 时 `error` 一定有值 —— 上游不通是常态,那也是一条要显示出来的结论。
|
||||
*/
|
||||
export interface UpstreamStatus {
|
||||
checked_at?: number | null
|
||||
remote_url?: string
|
||||
branch?: string
|
||||
ok?: boolean
|
||||
/** 上游有、当前代码没有的提交数 —— 要合的就是这些。 */
|
||||
behind?: number
|
||||
/** 当前代码有、上游没有的提交数 —— 也就是这一层改动自己的规模。 */
|
||||
ahead?: number
|
||||
tip?: string
|
||||
head?: string
|
||||
commits?: UpstreamCommit[]
|
||||
error?: string
|
||||
/** 本次检查是否推送了通知。 */
|
||||
notified?: boolean
|
||||
notify_error?: string
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user