feat: 监控面板 / 登录鉴权 / 多平台切换 / MySQL
在上游 MediaCrawler 之上新增一层: - 监控层 api/monitor/ —— 多博主/多笔记的定时采集、指标快照差分、报表、 企业微信通知。每轮采集写入独立目录,差分才成立。 - WebUI 登录鉴权 api/auth.py —— PBKDF2 口令 + 服务端会话,/api 全接口防护。 WebSocket 单独加依赖:BaseHTTPMiddleware 对 ws 作用域直接放行,覆盖不到。 - 全局平台切换 + 能力矩阵 —— 如实区分「爬虫模块支持」与「监控层已接线」, 未接通的平台直接拒绝建任务,而不是静默跑空。 - 监控库改用 MySQL 5.7(可回退 SQLite 供测试):逐表强制 utf8mb4 (服务端与库默认都是 latin1),启动校验所连 schema 以防写错库, 连接池 recycle + pre_ping 应对 MySQL 的 8 小时空闲断连。 修复上游缺陷: - xhs/core.py: 主页抓取失败会跳掉整个博主,导致一条作品都抓不到, 而那份资料只喂给一个空函数。改为尽力而为,失败不中断。 - xhs/login.py: cookie 登录只注入 web_session,冷启动签名会失败。 新增 INJECT_ALL_COOKIES 开关(默认关闭,原有行为不变)。 - requirements.txt: 补上 websockets。它在上游 pyproject.toml 里有声明、 这里漏了,导致 uvicorn 没有 WebSocket 能力,实时日志流从未工作。 改动过的上游文件清单及合并方式见 UPSTREAM.md。 测试:492 passed(另有 1 个既有的 Windows/gbk 上游测试失败,与本改动无关)
This commit is contained in:
@@ -0,0 +1,19 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/__init__.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Scheduled monitoring layer: repeated crawls with change detection."""
|
||||
@@ -0,0 +1,395 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/app_settings.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Application settings, declared once and rendered from that declaration.
|
||||
|
||||
Every setting carries a **scope**, which is the whole reason this is not a flat
|
||||
list:
|
||||
|
||||
* ``platform`` -- each platform keeps its own copy. A cookie obviously differs,
|
||||
but so do crawl pacing and proxies: what is safe on one platform is a rate
|
||||
limit on another. Stored as ``platform.<p>.<name>``.
|
||||
* ``system`` -- one value for the whole instance. The notification webhook is
|
||||
a single group chat, and the scheduler has a single active-hours window, so
|
||||
scoping those per platform would be a fiction.
|
||||
|
||||
The registry is the single source of truth: the API returns it and the Settings
|
||||
page builds its form from it, so adding a setting does not mean editing a
|
||||
matching list on the frontend.
|
||||
|
||||
Two rules carry over from how the cookie and webhook were already handled:
|
||||
|
||||
* **Secrets are never returned.** A sensitive key comes back as
|
||||
``{present, length, updated_at}``, never as a value.
|
||||
* **Update is partial.** Only keys present in the request are written, so a form
|
||||
that does not resubmit a secret cannot silently wipe it.
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from .platforms import PLATFORM_XHS
|
||||
from .settings import (
|
||||
delete_setting,
|
||||
get_setting,
|
||||
platform_key,
|
||||
set_setting,
|
||||
system_key,
|
||||
)
|
||||
|
||||
SCOPE_PLATFORM = "platform"
|
||||
SCOPE_SYSTEM = "system"
|
||||
|
||||
TYPE_BOOL = "bool"
|
||||
TYPE_INT = "int"
|
||||
TYPE_STR = "str"
|
||||
TYPE_SECRET = "secret"
|
||||
|
||||
# Mirrors config/base_config.py. Nothing is written until the operator changes
|
||||
# something; an unset value simply means "pass no CLI flag, so the config file's
|
||||
# value applies".
|
||||
_DEFAULT_SLEEP_SEC = 2
|
||||
|
||||
|
||||
@dataclass
|
||||
class SettingSpec:
|
||||
name: str
|
||||
scope: str
|
||||
type: str
|
||||
label: str
|
||||
help: str = ""
|
||||
default: Any = None
|
||||
minimum: Optional[int] = None
|
||||
maximum: Optional[int] = None
|
||||
choices: Optional[List[str]] = None
|
||||
affects_new_runs: bool = True
|
||||
|
||||
def key(self, platform: str = PLATFORM_XHS) -> str:
|
||||
if self.scope == SCOPE_SYSTEM:
|
||||
return system_key(self.name)
|
||||
return platform_key(platform, self.name)
|
||||
|
||||
|
||||
SETTING_SPECS: List[SettingSpec] = [
|
||||
# --- 平台设置 -----------------------------------------------------------
|
||||
SettingSpec(
|
||||
name="cookie",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_SECRET,
|
||||
label="登录 Cookie",
|
||||
help="定时监控必须持久化登录态。建议先手动登录一次再粘贴 Cookie。",
|
||||
),
|
||||
SettingSpec(
|
||||
name="default_interval_minutes",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_INT,
|
||||
label="新任务默认采集间隔(分钟)",
|
||||
help="仅影响新建任务时的默认值,不会改动已有任务。",
|
||||
default=360,
|
||||
minimum=30,
|
||||
maximum=10080,
|
||||
),
|
||||
SettingSpec(
|
||||
name="default_max_notes",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_INT,
|
||||
label="默认单轮作品上限",
|
||||
default=20,
|
||||
minimum=1,
|
||||
maximum=500,
|
||||
),
|
||||
SettingSpec(
|
||||
name="default_max_comments",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_INT,
|
||||
label="默认每篇评论抓取条数",
|
||||
help="接口无时间排序,只取平台默认排序的前 N 条;N 越大越容易发现新评论。",
|
||||
default=50,
|
||||
minimum=1,
|
||||
maximum=500,
|
||||
),
|
||||
SettingSpec(
|
||||
name="enable_sub_comments",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_BOOL,
|
||||
label="抓取二级评论",
|
||||
help="请求量显著增加,风控风险更高。",
|
||||
default=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="crawl_sleep_sec",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_INT,
|
||||
label="请求间隔(秒)",
|
||||
help="调大更慢但更不容易触发平台限流。各平台风控容忍度不同,故分开配置。",
|
||||
default=_DEFAULT_SLEEP_SEC,
|
||||
minimum=0,
|
||||
maximum=600,
|
||||
),
|
||||
SettingSpec(
|
||||
name="enable_ip_proxy",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_BOOL,
|
||||
label="启用 IP 代理",
|
||||
default=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="proxy_provider",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_STR,
|
||||
label="代理提供方",
|
||||
default="kuaidaili",
|
||||
choices=["kuaidaili", "wandouhttp", "static"],
|
||||
),
|
||||
SettingSpec(
|
||||
name="proxy_pool_count",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_INT,
|
||||
label="代理 IP 池大小",
|
||||
default=2,
|
||||
minimum=1,
|
||||
maximum=100,
|
||||
),
|
||||
SettingSpec(
|
||||
name="static_proxy_url",
|
||||
scope=SCOPE_PLATFORM,
|
||||
type=TYPE_STR,
|
||||
label="静态代理地址",
|
||||
help="仅当提供方选择 static 时使用,格式 http://host:port",
|
||||
default="",
|
||||
),
|
||||
# --- 系统设置 -----------------------------------------------------------
|
||||
SettingSpec(
|
||||
name="wecom_webhook",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_SECRET,
|
||||
label="企业微信 Webhook",
|
||||
help="企业微信群机器人地址。所有平台共用同一个群,只有开了推送开关的任务才会发消息。",
|
||||
),
|
||||
SettingSpec(
|
||||
name="active_hours_start",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_INT,
|
||||
label="活跃时段开始(小时)",
|
||||
help="只在此时段内触发定时采集。默认 0–23 即全天;支持跨午夜,如 22–6。",
|
||||
default=0,
|
||||
minimum=0,
|
||||
maximum=23,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
SettingSpec(
|
||||
name="active_hours_end",
|
||||
scope=SCOPE_SYSTEM,
|
||||
type=TYPE_INT,
|
||||
label="活跃时段结束(小时)",
|
||||
default=23,
|
||||
minimum=0,
|
||||
maximum=23,
|
||||
affects_new_runs=False,
|
||||
),
|
||||
]
|
||||
|
||||
SPECS_BY_NAME = {spec.name: spec for spec in SETTING_SPECS}
|
||||
|
||||
# Managed by their own endpoints; never writable through the settings API.
|
||||
# Suffix-matched rather than enumerated, because the cookie bookkeeping keys
|
||||
# exist once per platform.
|
||||
_HIDDEN_KEY_SUFFIXES = (".cookie_updated_at", ".cookie_last_ok_at")
|
||||
_HIDDEN_KEYS = {"auth_password_hash", "auth_password_updated_at"}
|
||||
|
||||
|
||||
def _is_hidden(key: str) -> bool:
|
||||
return key in _HIDDEN_KEYS or key.endswith(_HIDDEN_KEY_SUFFIXES)
|
||||
|
||||
|
||||
class SettingValidationError(ValueError):
|
||||
"""Raised for a value the registry will not accept."""
|
||||
|
||||
|
||||
def _coerce(spec: SettingSpec, raw: Any) -> Any:
|
||||
if spec.type == TYPE_SECRET:
|
||||
return str(raw) if raw is not None else ""
|
||||
|
||||
if spec.type == TYPE_BOOL:
|
||||
if isinstance(raw, bool):
|
||||
return raw
|
||||
text = str(raw).strip().lower()
|
||||
if text in ("1", "true", "yes", "y", "on"):
|
||||
return True
|
||||
if text in ("0", "false", "no", "n", "off", ""):
|
||||
return False
|
||||
raise SettingValidationError(f"{spec.label}: 需要是/否")
|
||||
|
||||
if spec.type == TYPE_INT:
|
||||
try:
|
||||
value = int(raw)
|
||||
except (TypeError, ValueError):
|
||||
raise SettingValidationError(f"{spec.label}: 需要整数")
|
||||
if spec.minimum is not None and value < spec.minimum:
|
||||
raise SettingValidationError(f"{spec.label}: 不能小于 {spec.minimum}")
|
||||
if spec.maximum is not None and value > spec.maximum:
|
||||
raise SettingValidationError(f"{spec.label}: 不能大于 {spec.maximum}")
|
||||
return value
|
||||
|
||||
value = str(raw) if raw is not None else ""
|
||||
if spec.choices and value not in spec.choices:
|
||||
raise SettingValidationError(f"{spec.label}: 只能是 {'/'.join(spec.choices)}")
|
||||
return value
|
||||
|
||||
|
||||
def _decode(spec: SettingSpec, raw: Optional[str]) -> Any:
|
||||
if raw is None:
|
||||
return spec.default
|
||||
if spec.type == TYPE_BOOL:
|
||||
return raw.strip().lower() in ("1", "true", "yes", "y", "on")
|
||||
if spec.type == TYPE_INT:
|
||||
try:
|
||||
return int(raw)
|
||||
except ValueError:
|
||||
return spec.default
|
||||
return raw
|
||||
|
||||
|
||||
def _encode(spec: SettingSpec, value: Any) -> str:
|
||||
if spec.type == TYPE_BOOL:
|
||||
return "true" if value else "false"
|
||||
return str(value)
|
||||
|
||||
|
||||
def _describe(spec: SettingSpec, platform: str) -> Dict[str, Any]:
|
||||
return {
|
||||
"key": spec.key(platform),
|
||||
"name": spec.name,
|
||||
"scope": spec.scope,
|
||||
"type": spec.type,
|
||||
"label": spec.label,
|
||||
"help": spec.help,
|
||||
"default": spec.default,
|
||||
"minimum": spec.minimum,
|
||||
"maximum": spec.maximum,
|
||||
"choices": spec.choices,
|
||||
"affects_new_runs": spec.affects_new_runs,
|
||||
}
|
||||
|
||||
|
||||
async def get_all(session: AsyncSession, platform: str = PLATFORM_XHS) -> Dict[str, Any]:
|
||||
"""Every editable setting for one platform, plus the system-wide ones.
|
||||
|
||||
Secrets come back masked, never in the clear.
|
||||
"""
|
||||
values: Dict[str, Any] = {}
|
||||
secrets: Dict[str, Any] = {}
|
||||
|
||||
for spec in SETTING_SPECS:
|
||||
key = spec.key(platform)
|
||||
raw = await get_setting(session, key)
|
||||
|
||||
if spec.type == TYPE_SECRET:
|
||||
secrets[key] = {"present": bool(raw), "length": len(raw or "")}
|
||||
else:
|
||||
values[key] = _decode(spec, raw)
|
||||
|
||||
return {
|
||||
"platform": platform,
|
||||
"values": values,
|
||||
"secrets": secrets,
|
||||
"specs": [_describe(spec, platform) for spec in SETTING_SPECS],
|
||||
}
|
||||
|
||||
|
||||
def _spec_for_key(key: str, platform: str) -> Optional[SettingSpec]:
|
||||
"""Resolve a full key back to its spec, rejecting keys for another platform."""
|
||||
for spec in SETTING_SPECS:
|
||||
if spec.key(platform) == key:
|
||||
return spec
|
||||
return None
|
||||
|
||||
|
||||
async def update(
|
||||
session: AsyncSession, payload: Dict[str, Any], platform: str = PLATFORM_XHS
|
||||
) -> List[str]:
|
||||
"""Apply a partial update. Returns the keys that changed.
|
||||
|
||||
Only keys present in ``payload`` are touched: a form that omits a secret must
|
||||
not blank it. Keys belonging to a different platform are rejected rather than
|
||||
silently written somewhere unexpected.
|
||||
"""
|
||||
changed: List[str] = []
|
||||
|
||||
for key, raw in payload.items():
|
||||
if _is_hidden(key):
|
||||
continue
|
||||
|
||||
spec = _spec_for_key(key, platform)
|
||||
if spec is None:
|
||||
raise SettingValidationError(f"未知的设置项:{key}")
|
||||
|
||||
# An explicit empty string clears a secret -- that is how the UI removes
|
||||
# one. For everything else it is just a value.
|
||||
if spec.type == TYPE_SECRET and raw == "":
|
||||
await delete_setting(session, key)
|
||||
changed.append(key)
|
||||
continue
|
||||
|
||||
value = _coerce(spec, raw)
|
||||
await set_setting(session, key, _encode(spec, value))
|
||||
changed.append(key)
|
||||
|
||||
return changed
|
||||
|
||||
|
||||
async def get_value(
|
||||
session: AsyncSession,
|
||||
name: str,
|
||||
platform: str = PLATFORM_XHS,
|
||||
fallback: Any = None,
|
||||
) -> Any:
|
||||
"""Read one typed setting for internal callers (the runner, the scheduler)."""
|
||||
spec = SPECS_BY_NAME.get(name)
|
||||
if spec is None:
|
||||
return fallback
|
||||
raw = await get_setting(session, spec.key(platform))
|
||||
if raw is None:
|
||||
return spec.default if fallback is None else fallback
|
||||
return _decode(spec, raw)
|
||||
|
||||
|
||||
async def defaults(session: AsyncSession, platform: str = PLATFORM_XHS) -> Dict[str, Any]:
|
||||
"""Defaults applied when creating a task on this platform.
|
||||
|
||||
This is what makes the Settings page govern new tasks: the create endpoint
|
||||
falls back to these for anything the caller omits.
|
||||
"""
|
||||
return {
|
||||
"interval_minutes": int(
|
||||
await get_value(session, "default_interval_minutes", platform, 360)
|
||||
),
|
||||
"max_notes_count": int(await get_value(session, "default_max_notes", platform, 20)),
|
||||
"max_comments_count": int(
|
||||
await get_value(session, "default_max_comments", platform, 50)
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
async def active_hours(session: AsyncSession) -> tuple[int, int]:
|
||||
"""The (start, end) hour window for scheduled runs. System-wide."""
|
||||
start = await get_value(session, "active_hours_start", fallback=0)
|
||||
end = await get_value(session, "active_hours_end", fallback=23)
|
||||
return int(start), int(end)
|
||||
@@ -0,0 +1,293 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/db.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Database engine for the monitoring layer.
|
||||
|
||||
**MySQL** by default (see ``config/db_config.py`` and ``.env``), with SQLite kept
|
||||
as an option so the test suite can run without a reachable server.
|
||||
|
||||
Three things here exist because of specific MySQL 5.7 behaviour:
|
||||
|
||||
* **utf8mb4 is forced per table.** This instance's server *and* the target schema
|
||||
default to ``latin1``; relying on either would mangle or reject Chinese text.
|
||||
The charset is set on every table rather than on the database, so it holds no
|
||||
matter what the schema default is.
|
||||
* **Connections are recycled.** The monitor runs for weeks, and MySQL drops idle
|
||||
connections after ``wait_timeout`` (8h by default). Without ``pool_recycle`` and
|
||||
``pool_pre_ping`` the first query after a quiet night fails with "server has
|
||||
gone away".
|
||||
* **The connected schema is asserted at startup.** A misconfigured database name
|
||||
is caught immediately instead of silently writing to the wrong schema.
|
||||
|
||||
Only the configured schema is ever touched: no ``CREATE DATABASE``, no ``USE``,
|
||||
no cross-schema query.
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
from contextlib import asynccontextmanager
|
||||
from pathlib import Path
|
||||
from typing import AsyncIterator, Optional
|
||||
|
||||
from sqlalchemy import event, text
|
||||
from sqlalchemy.ext.asyncio import (
|
||||
AsyncEngine,
|
||||
AsyncSession,
|
||||
async_sessionmaker,
|
||||
create_async_engine,
|
||||
)
|
||||
|
||||
from .models import MonitorBase
|
||||
|
||||
PROJECT_ROOT = Path(__file__).parent.parent.parent
|
||||
DATA_DIR = PROJECT_ROOT / "data"
|
||||
DEFAULT_SQLITE_PATH = DATA_DIR / "monitor.db"
|
||||
|
||||
# Load .env here as well as in api/main.py: this module is imported directly by
|
||||
# scripts and tests, and a configuration that only applies when the server is the
|
||||
# entry point is a trap. load_dotenv does not override real environment variables.
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv(PROJECT_ROOT / ".env")
|
||||
|
||||
# Kept identical to config/db_config.py's defaults so one .env drives both the
|
||||
# monitor database and the crawler's own DB output.
|
||||
MYSQL_HOST = lambda: os.getenv("MYSQL_DB_HOST", "localhost") # noqa: E731
|
||||
MYSQL_PORT = lambda: int(os.getenv("MYSQL_DB_PORT", "3306")) # noqa: E731
|
||||
MYSQL_USER = lambda: os.getenv("MYSQL_DB_USER", "root") # noqa: E731
|
||||
MYSQL_PWD = lambda: os.getenv("MYSQL_DB_PWD", "") # noqa: E731
|
||||
MYSQL_DB_NAME = lambda: os.getenv("MYSQL_DB_NAME", "mediacrawler") # noqa: E731
|
||||
|
||||
_engine: Optional[AsyncEngine] = None
|
||||
_session_factory: Optional[async_sessionmaker[AsyncSession]] = None
|
||||
# None means "resolve from the environment" (MySQL). Tests set a SQLite URL.
|
||||
_db_url: Optional[str] = None
|
||||
_expected_schema: Optional[str] = None
|
||||
|
||||
|
||||
def resolve_db_url() -> str:
|
||||
"""Build the connection URL. MySQL unless overridden."""
|
||||
if _db_url is not None:
|
||||
return _db_url
|
||||
|
||||
from urllib.parse import quote_plus
|
||||
|
||||
user = quote_plus(MYSQL_USER())
|
||||
password = quote_plus(MYSQL_PWD())
|
||||
host = MYSQL_HOST()
|
||||
port = MYSQL_PORT()
|
||||
name = MYSQL_DB_NAME()
|
||||
return f"mysql+aiomysql://{user}:{password}@{host}:{port}/{name}?charset=utf8mb4"
|
||||
|
||||
|
||||
def is_mysql() -> bool:
|
||||
return resolve_db_url().startswith("mysql")
|
||||
|
||||
|
||||
def set_sqlite_path(path: Path) -> None:
|
||||
"""Point the layer at SQLite. Used by the test suite only."""
|
||||
global _db_url, _engine, _session_factory, _expected_schema
|
||||
_db_url = f"sqlite+aiosqlite:///{Path(path)}"
|
||||
_engine = None
|
||||
_session_factory = None
|
||||
_expected_schema = None
|
||||
|
||||
|
||||
def set_db_url(url: str, expected_schema: Optional[str] = None) -> None:
|
||||
"""Point the layer at an explicit URL. ``expected_schema`` enables the guard."""
|
||||
global _db_url, _engine, _session_factory, _expected_schema
|
||||
_db_url = url
|
||||
_engine = None
|
||||
_session_factory = None
|
||||
_expected_schema = expected_schema
|
||||
|
||||
|
||||
def expected_schema() -> Optional[str]:
|
||||
"""The schema the connection must be using, if the guard applies."""
|
||||
if _expected_schema is not None:
|
||||
return _expected_schema
|
||||
return MYSQL_DB_NAME() if is_mysql() else None
|
||||
|
||||
|
||||
def get_engine() -> AsyncEngine:
|
||||
global _engine
|
||||
if _engine is None:
|
||||
url = resolve_db_url()
|
||||
kwargs: dict = {"future": True}
|
||||
|
||||
if url.startswith("mysql"):
|
||||
# Recycle well inside MySQL's default 8h wait_timeout, and verify a
|
||||
# pooled connection before handing it out.
|
||||
kwargs.update(pool_recycle=3600, pool_pre_ping=True, pool_size=5, max_overflow=5)
|
||||
kwargs["connect_args"] = {"charset": "utf8mb4"}
|
||||
else:
|
||||
Path(url.split("///", 1)[-1]).parent.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
_engine = create_async_engine(url, **kwargs)
|
||||
|
||||
if url.startswith("sqlite"):
|
||||
|
||||
@event.listens_for(_engine.sync_engine, "connect")
|
||||
def _set_sqlite_pragmas(dbapi_connection, _connection_record): # pragma: no cover
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA foreign_keys=ON")
|
||||
cursor.close()
|
||||
|
||||
return _engine
|
||||
|
||||
|
||||
def get_session_factory() -> async_sessionmaker[AsyncSession]:
|
||||
global _session_factory
|
||||
if _session_factory is None:
|
||||
_session_factory = async_sessionmaker(
|
||||
bind=get_engine(),
|
||||
class_=AsyncSession,
|
||||
expire_on_commit=False,
|
||||
)
|
||||
return _session_factory
|
||||
|
||||
|
||||
@asynccontextmanager
|
||||
async def get_session() -> AsyncIterator[AsyncSession]:
|
||||
"""Transactional session. Commits on success, rolls back on error."""
|
||||
factory = get_session_factory()
|
||||
async with factory() as session:
|
||||
try:
|
||||
yield session
|
||||
await session.commit()
|
||||
except Exception:
|
||||
await session.rollback()
|
||||
raise
|
||||
|
||||
|
||||
async def _assert_correct_schema(conn) -> None:
|
||||
"""Refuse to run against anything but the configured schema.
|
||||
|
||||
A guard, not the guarantee: the real protection is a MySQL account scoped to
|
||||
this one schema (see UPSTREAM.md). This catches the ordinary mistake of a
|
||||
wrong database name in configuration, before a single row is written.
|
||||
"""
|
||||
if not is_mysql():
|
||||
return
|
||||
|
||||
expected = expected_schema()
|
||||
if not expected:
|
||||
return
|
||||
|
||||
current = (await conn.execute(text("SELECT DATABASE()"))).scalar()
|
||||
if current is None:
|
||||
raise RuntimeError(
|
||||
f"数据库连接未选定 schema,期望 {expected!r}。请检查 MYSQL_DB_NAME。"
|
||||
)
|
||||
# lower_case_table_names=1 makes names case-insensitive server-side.
|
||||
if current.lower() != expected.lower():
|
||||
raise RuntimeError(
|
||||
f"连接的库是 {current!r},但配置要求 {expected!r}。"
|
||||
f"为避免误写其它库,已拒绝启动。"
|
||||
)
|
||||
print(f"[monitor.db] 已连接 MySQL schema: {current}", flush=True)
|
||||
|
||||
|
||||
async def init_db() -> None:
|
||||
"""Create missing tables, then run the small in-place migrations."""
|
||||
engine = get_engine()
|
||||
async with engine.begin() as conn:
|
||||
await _assert_correct_schema(conn)
|
||||
await conn.run_sync(MonitorBase.metadata.create_all)
|
||||
await _ensure_columns(conn)
|
||||
await _migrate_setting_keys(conn)
|
||||
|
||||
|
||||
# Columns added to a table after it may already exist. ``create_all`` only
|
||||
# creates missing *tables*, so new columns need an explicit ALTER TABLE.
|
||||
_ADDED_COLUMNS: dict[str, list[tuple[str, str]]] = {
|
||||
"monitor_task": [
|
||||
("notify_enabled", "BOOLEAN NOT NULL DEFAULT 0"),
|
||||
("last_notified_at", "BIGINT NULL"),
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
async def _existing_columns(conn, table: str) -> set[str]:
|
||||
if is_mysql():
|
||||
rows = await conn.execute(
|
||||
text(
|
||||
"SELECT COLUMN_NAME FROM information_schema.COLUMNS "
|
||||
"WHERE TABLE_SCHEMA = DATABASE() AND TABLE_NAME = :t"
|
||||
),
|
||||
{"t": table},
|
||||
)
|
||||
return {row[0] for row in rows}
|
||||
|
||||
rows = await conn.execute(text(f"PRAGMA table_info({table})"))
|
||||
return {row[1] for row in rows}
|
||||
|
||||
|
||||
async def _ensure_columns(conn) -> None:
|
||||
for table, columns in _ADDED_COLUMNS.items():
|
||||
existing = await _existing_columns(conn, table)
|
||||
if not existing:
|
||||
# Table did not exist before this run; create_all built it complete.
|
||||
continue
|
||||
for name, ddl in columns:
|
||||
if name not in existing:
|
||||
await conn.execute(text(f"ALTER TABLE {table} ADD COLUMN {name} {ddl}"))
|
||||
|
||||
|
||||
async def _migrate_setting_keys(conn) -> None:
|
||||
"""Move pre-namespacing setting keys to their scoped names.
|
||||
|
||||
Idempotent: the legacy row is only renamed when the new key is absent, so an
|
||||
operator's later value is never overwritten.
|
||||
"""
|
||||
from .models import LEGACY_SETTING_KEY_RENAMES
|
||||
|
||||
for legacy, scoped in LEGACY_SETTING_KEY_RENAMES.items():
|
||||
exists = (
|
||||
await conn.execute(
|
||||
text("SELECT 1 FROM monitor_setting WHERE `key` = :k"), {"k": legacy}
|
||||
)
|
||||
).first()
|
||||
if not exists:
|
||||
continue
|
||||
|
||||
already = (
|
||||
await conn.execute(
|
||||
text("SELECT 1 FROM monitor_setting WHERE `key` = :k"), {"k": scoped}
|
||||
)
|
||||
).first()
|
||||
if already:
|
||||
# Both present: the scoped one is authoritative; drop the stale row.
|
||||
await conn.execute(
|
||||
text("DELETE FROM monitor_setting WHERE `key` = :k"), {"k": legacy}
|
||||
)
|
||||
continue
|
||||
|
||||
await conn.execute(
|
||||
text("UPDATE monitor_setting SET `key` = :new WHERE `key` = :old"),
|
||||
{"new": scoped, "old": legacy},
|
||||
)
|
||||
|
||||
|
||||
async def dispose_engine() -> None:
|
||||
global _engine, _session_factory
|
||||
if _engine is not None:
|
||||
await _engine.dispose()
|
||||
_engine = None
|
||||
_session_factory = None
|
||||
@@ -0,0 +1,589 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/ingest.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Turn one run's crawled jsonl into snapshots and change events.
|
||||
|
||||
Pure-ish and offline testable: give it a directory of jsonl files, a run row and
|
||||
a session, and it does the diffing. No network, no browser.
|
||||
|
||||
Correctness notes that drive the code below:
|
||||
|
||||
* Counts arrive as strings and may be abbreviated ("1.2万", "3亿"). A value that
|
||||
cannot be parsed is stored as NULL, never 0 -- 0 would forge a large negative
|
||||
delta on the next comparison.
|
||||
* The comment endpoint has no time-sort, so only the platform's top-N window is
|
||||
ever visible. A comment we have not seen before is therefore split into
|
||||
"posted since last run" vs "seen for the first time", rather than claiming the
|
||||
former always.
|
||||
* A bad cookie does not make the crawler exit non-zero; it exits 0 having
|
||||
fetched nothing. That is detected here as a suspected auth failure.
|
||||
"""
|
||||
|
||||
import json
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from sqlalchemy import func, select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from .platforms import PLATFORM_XHS
|
||||
from .models import (
|
||||
EVENT_AUTH_FAILURE,
|
||||
EVENT_METRIC_DELTA,
|
||||
EVENT_NEW_COMMENT_POSTED,
|
||||
EVENT_NEW_COMMENT_SEEN,
|
||||
EVENT_NEW_NOTE,
|
||||
EVENT_NO_DATA,
|
||||
EVENT_RUN_FAILED,
|
||||
MonitorComment,
|
||||
MonitorEvent,
|
||||
MonitorNote,
|
||||
MonitorNoteMetric,
|
||||
MonitorRun,
|
||||
MonitorTask,
|
||||
RUN_FAILED,
|
||||
RUN_PARTIAL,
|
||||
RUN_SUCCESS,
|
||||
)
|
||||
|
||||
_COUNT_UNITS = {
|
||||
"": 1,
|
||||
"万": 10_000,
|
||||
"w": 10_000,
|
||||
"W": 10_000,
|
||||
"k": 1_000,
|
||||
"K": 1_000,
|
||||
"亿": 100_000_000,
|
||||
}
|
||||
_COUNT_RE = re.compile(r"^([\d.]+)\s*([万wWkK亿]?)$")
|
||||
|
||||
# Metric fields shared by the snapshot table and the delta comparison.
|
||||
_METRIC_FIELDS = ("liked_count", "comment_count", "collected_count", "share_count")
|
||||
|
||||
|
||||
def parse_count(value: Any) -> Optional[int]:
|
||||
"""Parse an XHS interaction count into an int, or None if unintelligible.
|
||||
|
||||
Handles plain numbers, thousands separators, and the Chinese abbreviations
|
||||
the platform actually returns ("1.2万" -> 12000, "3亿" -> 300000000).
|
||||
"""
|
||||
if value is None or isinstance(value, bool):
|
||||
return None
|
||||
if isinstance(value, int):
|
||||
return value
|
||||
if isinstance(value, float):
|
||||
return int(value)
|
||||
|
||||
text = str(value).strip().replace(",", "").replace(" ", "")
|
||||
if not text:
|
||||
return None
|
||||
|
||||
match = _COUNT_RE.match(text)
|
||||
if not match:
|
||||
return None
|
||||
|
||||
try:
|
||||
number = float(match.group(1))
|
||||
except ValueError:
|
||||
return None
|
||||
|
||||
return int(number * _COUNT_UNITS.get(match.group(2), 1))
|
||||
|
||||
|
||||
# Windows reports hard process failures as NTSTATUS values, which surface in the
|
||||
# UI as meaningless large integers (e.g. 3221225794 = 0xC0000142). Translating
|
||||
# the ones we actually see saves the reader a hex-decoding detour.
|
||||
_WINDOWS_EXIT_REASONS = {
|
||||
0xC0000005: "进程访问冲突 (ACCESS_VIOLATION)",
|
||||
0xC00000FD: "栈溢出 (STACK_OVERFLOW)",
|
||||
0xC000013A: "进程被中断(控制台关闭或 Ctrl+C)",
|
||||
0xC0000142: "进程初始化失败 (STATUS_DLL_INIT_FAILED),属启动环境异常,重启服务后重试",
|
||||
0xC0000409: "栈缓冲区溢出 (STACK_BUFFER_OVERRUN)",
|
||||
}
|
||||
|
||||
|
||||
def describe_exit_code(code: int) -> str:
|
||||
"""Render an exit code so a human can act on it."""
|
||||
unsigned = code & 0xFFFFFFFF if code < 0 else code
|
||||
reason = _WINDOWS_EXIT_REASONS.get(unsigned)
|
||||
if reason:
|
||||
return f"Crawler exited with code {code} (0x{unsigned:08X}): {reason}"
|
||||
return f"Crawler exited with code {code}"
|
||||
|
||||
|
||||
@dataclass
|
||||
class IngestResult:
|
||||
status: str
|
||||
notes_fetched: int = 0
|
||||
comments_fetched: int = 0
|
||||
new_notes: int = 0
|
||||
new_comments: int = 0
|
||||
is_baseline: bool = False
|
||||
error: Optional[str] = None
|
||||
events: List[str] = field(default_factory=list)
|
||||
|
||||
|
||||
def _read_jsonl(path: Path) -> List[Dict[str, Any]]:
|
||||
"""Read a jsonl file, skipping blank or malformed lines."""
|
||||
records: List[Dict[str, Any]] = []
|
||||
if not path.exists():
|
||||
return records
|
||||
|
||||
with path.open("r", encoding="utf-8") as handle:
|
||||
for line in handle:
|
||||
line = line.strip()
|
||||
if not line:
|
||||
continue
|
||||
try:
|
||||
item = json.loads(line)
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
if isinstance(item, dict):
|
||||
records.append(item)
|
||||
return records
|
||||
|
||||
|
||||
def find_run_files(
|
||||
out_dir: Path, platform: str = PLATFORM_XHS
|
||||
) -> tuple[List[Path], List[Path]]:
|
||||
"""Locate the contents/comments jsonl files a run produced.
|
||||
|
||||
The crawler writes ``{save_data_path}/{platform}/jsonl/{type}_{item}_{date}.jsonl``.
|
||||
Glob rather than reconstructing the name: both the crawler type and the date
|
||||
are runtime-dependent. Returns lists because a crawl crossing midnight
|
||||
produces one file per day.
|
||||
"""
|
||||
jsonl_dir = out_dir / platform / "jsonl"
|
||||
if not jsonl_dir.is_dir():
|
||||
return [], []
|
||||
|
||||
return (
|
||||
sorted(jsonl_dir.glob("*_contents_*.jsonl")),
|
||||
sorted(jsonl_dir.glob("*_comments_*.jsonl")),
|
||||
)
|
||||
|
||||
|
||||
async def _emit(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
event_type: str,
|
||||
title: str,
|
||||
*,
|
||||
severity: str = "info",
|
||||
target_kind: str = "",
|
||||
target_id: str = "",
|
||||
payload: Optional[Dict[str, Any]] = None,
|
||||
) -> None:
|
||||
session.add(
|
||||
MonitorEvent(
|
||||
task_id=run.task_id,
|
||||
run_id=run.id,
|
||||
type=event_type,
|
||||
severity=severity,
|
||||
target_kind=target_kind,
|
||||
target_id=target_id,
|
||||
title=title,
|
||||
payload_json=json.dumps(payload or {}, ensure_ascii=False),
|
||||
created_at=get_current_timestamp(),
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
async def _previous_run_started_at(
|
||||
session: AsyncSession, task_id: int, run_id: int
|
||||
) -> Optional[int]:
|
||||
"""Started-at of the most recent earlier successful run, in ms."""
|
||||
return await session.scalar(
|
||||
select(MonitorRun.started_at)
|
||||
.where(
|
||||
MonitorRun.task_id == task_id,
|
||||
MonitorRun.id != run_id,
|
||||
MonitorRun.status.in_((RUN_SUCCESS, RUN_PARTIAL)),
|
||||
MonitorRun.started_at.is_not(None),
|
||||
# Same reasoning as _count_prior_successes: an empty run is a useless
|
||||
# reference point for "was this comment posted since last time?".
|
||||
MonitorRun.notes_fetched > 0,
|
||||
)
|
||||
.order_by(MonitorRun.id.desc())
|
||||
.limit(1)
|
||||
)
|
||||
|
||||
|
||||
# How far back to look for proof that the stored login still works.
|
||||
_AUTH_PROOF_WINDOW_MS = 6 * 60 * 60 * 1000
|
||||
|
||||
|
||||
async def _another_task_succeeded_recently(session: AsyncSession, task_id: int) -> bool:
|
||||
"""Whether a different task fetched data recently, proving the login is valid."""
|
||||
since = get_current_timestamp() - _AUTH_PROOF_WINDOW_MS
|
||||
count = await session.scalar(
|
||||
select(func.count())
|
||||
.select_from(MonitorRun)
|
||||
.where(
|
||||
MonitorRun.task_id != task_id,
|
||||
MonitorRun.status == RUN_SUCCESS,
|
||||
MonitorRun.started_at.is_not(None),
|
||||
MonitorRun.started_at >= since,
|
||||
)
|
||||
)
|
||||
return bool(count)
|
||||
|
||||
|
||||
async def _count_prior_successes(session: AsyncSession, task_id: int, run_id: int) -> int:
|
||||
return (
|
||||
await session.scalar(
|
||||
select(func.count())
|
||||
.select_from(MonitorRun)
|
||||
.where(
|
||||
MonitorRun.task_id == task_id,
|
||||
MonitorRun.id != run_id,
|
||||
MonitorRun.status.in_((RUN_SUCCESS, RUN_PARTIAL)),
|
||||
# A run that fetched nothing established no baseline. Without this
|
||||
# check the first run that actually works after a failed one looks
|
||||
# like a flood of newly discovered works.
|
||||
MonitorRun.notes_fetched > 0,
|
||||
)
|
||||
)
|
||||
) or 0
|
||||
|
||||
|
||||
async def _ingest_notes(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
records: List[Dict[str, Any]],
|
||||
is_baseline: bool,
|
||||
) -> int:
|
||||
"""Upsert notes, write metric snapshots, and emit new-note/delta events."""
|
||||
now = get_current_timestamp()
|
||||
new_count = 0
|
||||
|
||||
for record in records:
|
||||
note_id = record.get("note_id")
|
||||
if not note_id:
|
||||
continue
|
||||
|
||||
note = await session.scalar(
|
||||
select(MonitorNote).where(
|
||||
MonitorNote.task_id == run.task_id,
|
||||
MonitorNote.note_id == note_id,
|
||||
)
|
||||
)
|
||||
|
||||
title = (record.get("title") or "")[:500]
|
||||
raw_images = record.get("image_list") or ""
|
||||
cover = raw_images.split(",")[0] if raw_images else ""
|
||||
|
||||
if note is None:
|
||||
note = MonitorNote(
|
||||
task_id=run.task_id,
|
||||
note_id=note_id,
|
||||
title=title,
|
||||
note_url=record.get("note_url") or "",
|
||||
cover=cover,
|
||||
creator_hash=record.get("creator_hash") or "",
|
||||
source_kind=record.get("type") or "",
|
||||
published_at=_as_int(record.get("time")),
|
||||
first_seen_run_id=run.id,
|
||||
first_seen_at=now,
|
||||
last_seen_run_id=run.id,
|
||||
last_seen_at=now,
|
||||
)
|
||||
session.add(note)
|
||||
new_count += 1
|
||||
if not is_baseline:
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_NEW_NOTE,
|
||||
f"新作品:{title or note_id}",
|
||||
target_kind="note",
|
||||
target_id=note_id,
|
||||
payload={"note_id": note_id, "title": title},
|
||||
)
|
||||
else:
|
||||
# Only refresh descriptive fields; seen-tracking is updated below.
|
||||
if title:
|
||||
note.title = title
|
||||
note.last_seen_run_id = run.id
|
||||
note.last_seen_at = now
|
||||
|
||||
await _snapshot_metrics(session, run, note_id, record, now, is_baseline)
|
||||
|
||||
return new_count
|
||||
|
||||
|
||||
def _as_int(value: Any) -> Optional[int]:
|
||||
try:
|
||||
return int(value)
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
|
||||
|
||||
async def _snapshot_metrics(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
note_id: str,
|
||||
record: Dict[str, Any],
|
||||
now: int,
|
||||
is_baseline: bool,
|
||||
) -> None:
|
||||
"""Write this run's metric snapshot and report any change vs the previous one."""
|
||||
previous = await session.scalar(
|
||||
select(MonitorNoteMetric)
|
||||
.where(
|
||||
MonitorNoteMetric.task_id == run.task_id,
|
||||
MonitorNoteMetric.note_id == note_id,
|
||||
MonitorNoteMetric.run_id != run.id,
|
||||
)
|
||||
.order_by(MonitorNoteMetric.run_id.desc())
|
||||
.limit(1)
|
||||
)
|
||||
|
||||
parsed = {name: parse_count(record.get(name)) for name in _METRIC_FIELDS}
|
||||
|
||||
session.add(
|
||||
MonitorNoteMetric(
|
||||
task_id=run.task_id,
|
||||
note_id=note_id,
|
||||
run_id=run.id,
|
||||
captured_at=now,
|
||||
liked_count=parsed["liked_count"],
|
||||
comment_count=parsed["comment_count"],
|
||||
collected_count=parsed["collected_count"],
|
||||
share_count=parsed["share_count"],
|
||||
raw_liked_count=str(record.get("liked_count") or ""),
|
||||
raw_comment_count=str(record.get("comment_count") or ""),
|
||||
raw_collected_count=str(record.get("collected_count") or ""),
|
||||
raw_share_count=str(record.get("share_count") or ""),
|
||||
)
|
||||
)
|
||||
|
||||
if previous is None or is_baseline:
|
||||
return
|
||||
|
||||
deltas = {}
|
||||
for name in _METRIC_FIELDS:
|
||||
old, new = getattr(previous, name), parsed[name]
|
||||
# A None on either side means the value was unparseable; skip rather
|
||||
# than report a bogus change.
|
||||
if old is None or new is None or old == new:
|
||||
continue
|
||||
deltas[name] = {"from": old, "to": new, "delta": new - old}
|
||||
|
||||
if deltas:
|
||||
summary = "、".join(
|
||||
f"{_metric_label(name)} {info['from']}→{info['to']}"
|
||||
for name, info in deltas.items()
|
||||
)
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_METRIC_DELTA,
|
||||
f"互动数据变化:{summary}",
|
||||
target_kind="note",
|
||||
target_id=note_id,
|
||||
payload={"note_id": note_id, "deltas": deltas},
|
||||
)
|
||||
|
||||
|
||||
def _metric_label(name: str) -> str:
|
||||
return {
|
||||
"liked_count": "点赞",
|
||||
"comment_count": "评论",
|
||||
"collected_count": "收藏",
|
||||
"share_count": "分享",
|
||||
}.get(name, name)
|
||||
|
||||
|
||||
async def _ingest_comments(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
records: List[Dict[str, Any]],
|
||||
is_baseline: bool,
|
||||
previous_run_started_at: Optional[int],
|
||||
) -> int:
|
||||
"""Upsert comments and emit events for ones never seen before."""
|
||||
now = get_current_timestamp()
|
||||
new_count = 0
|
||||
|
||||
for record in records:
|
||||
comment_id = record.get("comment_id")
|
||||
note_id = record.get("note_id")
|
||||
if not comment_id or not note_id:
|
||||
continue
|
||||
|
||||
exists = await session.scalar(
|
||||
select(MonitorComment.id).where(
|
||||
MonitorComment.task_id == run.task_id,
|
||||
MonitorComment.note_id == note_id,
|
||||
MonitorComment.comment_id == comment_id,
|
||||
)
|
||||
)
|
||||
if exists is not None:
|
||||
continue
|
||||
|
||||
create_time = _as_int(record.get("create_time"))
|
||||
session.add(
|
||||
MonitorComment(
|
||||
task_id=run.task_id,
|
||||
note_id=note_id,
|
||||
comment_id=comment_id,
|
||||
content=(record.get("content") or "")[:2000],
|
||||
nickname=record.get("nickname") or "",
|
||||
creator_hash=record.get("creator_hash") or "",
|
||||
create_time=create_time,
|
||||
like_count=parse_count(record.get("like_count")),
|
||||
sub_comment_count=_as_int(record.get("sub_comment_count")) or 0,
|
||||
parent_comment_id=record.get("parent_comment_id") or "",
|
||||
first_seen_run_id=run.id,
|
||||
first_seen_at=now,
|
||||
)
|
||||
)
|
||||
new_count += 1
|
||||
|
||||
if is_baseline:
|
||||
continue
|
||||
|
||||
# Without a time-sorted comment API we can only observe the top-N window,
|
||||
# so distinguish a genuinely new comment from one that just surfaced.
|
||||
posted = (
|
||||
create_time is not None
|
||||
and previous_run_started_at is not None
|
||||
and create_time > previous_run_started_at
|
||||
)
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_NEW_COMMENT_POSTED if posted else EVENT_NEW_COMMENT_SEEN,
|
||||
f"{'新评论' if posted else '新出现评论'}:{(record.get('content') or '')[:60]}",
|
||||
target_kind="note",
|
||||
target_id=note_id,
|
||||
payload={
|
||||
"note_id": note_id,
|
||||
"comment_id": comment_id,
|
||||
"create_time": create_time,
|
||||
"nickname": record.get("nickname") or "",
|
||||
},
|
||||
)
|
||||
|
||||
return new_count
|
||||
|
||||
|
||||
async def ingest_run(
|
||||
session: AsyncSession,
|
||||
run: MonitorRun,
|
||||
task: MonitorTask,
|
||||
out_dir: Path,
|
||||
) -> IngestResult:
|
||||
"""Ingest one finished run and return what changed.
|
||||
|
||||
Sets ``run.status``, ``run.is_baseline`` and the counters on the run row.
|
||||
On a failed or untrustworthy run nothing is diffed -- the "seen" sets only
|
||||
ever grow, so a partial run must never be allowed to look like deletions.
|
||||
"""
|
||||
# A non-zero exit is a genuine crash: trust nothing this run produced.
|
||||
if run.exit_code not in (0, None):
|
||||
run.status = RUN_FAILED
|
||||
run.error_message = describe_exit_code(run.exit_code)
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_RUN_FAILED,
|
||||
f"采集进程异常退出(code={run.exit_code})",
|
||||
severity="error",
|
||||
payload={"exit_code": run.exit_code, "detail": run.error_message},
|
||||
)
|
||||
return IngestResult(status=RUN_FAILED, error=run.error_message)
|
||||
|
||||
contents_paths, comment_paths = find_run_files(out_dir, task.platform)
|
||||
contents = [record for path in contents_paths for record in _read_jsonl(path)]
|
||||
comments = [record for path in comment_paths for record in _read_jsonl(path)]
|
||||
|
||||
run.notes_fetched = len(contents)
|
||||
run.comments_fetched = len(comments)
|
||||
|
||||
# A bad cookie does NOT fail the process: XHS cookie login is never validated,
|
||||
# so an unauthenticated session just returns zero notes with exit 0 -- and
|
||||
# usually does not even create an output file. Treating that as "the creator
|
||||
# posted nothing" would silently hide login outages, which is exactly what
|
||||
# monitoring exists to catch.
|
||||
if not contents:
|
||||
run.status = RUN_PARTIAL
|
||||
|
||||
# Blaming the cookie is only honest if nothing else is authenticating.
|
||||
# A sibling task that just succeeded proves the login works, so the
|
||||
# fault is with this target (bad/expired per-creator token, an empty
|
||||
# account, or a page-structure change).
|
||||
if await _another_task_succeeded_recently(session, run.task_id):
|
||||
run.error_message = (
|
||||
"Crawler produced no notes for this target, but other tasks "
|
||||
"succeeded recently, so the login is probably fine"
|
||||
)
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_NO_DATA,
|
||||
"本次未抓到任何作品:其他任务近期采集正常,登录态应该没问题,请检查该目标是否有效",
|
||||
severity="warning",
|
||||
payload={"out_dir": str(out_dir)},
|
||||
)
|
||||
else:
|
||||
run.error_message = "Crawler produced no notes; the login cookie may have expired"
|
||||
await _emit(
|
||||
session,
|
||||
run,
|
||||
EVENT_AUTH_FAILURE,
|
||||
"疑似登录态失效:本次未抓到任何作品,请检查 Cookie",
|
||||
severity="error",
|
||||
payload={"out_dir": str(out_dir)},
|
||||
)
|
||||
|
||||
return IngestResult(
|
||||
status=RUN_PARTIAL,
|
||||
error=run.error_message,
|
||||
comments_fetched=len(comments),
|
||||
)
|
||||
|
||||
is_baseline = await _count_prior_successes(session, run.task_id, run.id) == 0
|
||||
run.is_baseline = is_baseline
|
||||
run.status = RUN_SUCCESS
|
||||
run.error_message = None
|
||||
|
||||
previous_started_at = (
|
||||
None if is_baseline else await _previous_run_started_at(session, run.task_id, run.id)
|
||||
)
|
||||
|
||||
result = IngestResult(
|
||||
status=RUN_SUCCESS,
|
||||
notes_fetched=len(contents),
|
||||
comments_fetched=len(comments),
|
||||
is_baseline=is_baseline,
|
||||
)
|
||||
result.new_notes = await _ingest_notes(session, run, contents, is_baseline)
|
||||
if task.enable_comments:
|
||||
result.new_comments = await _ingest_comments(
|
||||
session, run, comments, is_baseline, previous_started_at
|
||||
)
|
||||
|
||||
run.new_notes = result.new_notes
|
||||
run.new_comments = result.new_comments
|
||||
return result
|
||||
@@ -0,0 +1,177 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/migrate_from_sqlite.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""One-off: copy the monitoring database from SQLite into MySQL.
|
||||
|
||||
python -m api.monitor.migrate_from_sqlite [--source data/monitor.db] [--dry-run]
|
||||
|
||||
Primary keys are preserved rather than reassigned, because rows in
|
||||
``monitor_note`` / ``monitor_comment`` / ``monitor_run`` reference ``task_id``;
|
||||
letting MySQL auto-assign new ids would silently break those links.
|
||||
|
||||
Refuses to run against a target that already holds data unless ``--force`` is
|
||||
given, so a second accidental run cannot double everything up.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import sqlite3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List
|
||||
|
||||
PROJECT_ROOT = Path(__file__).parent.parent.parent
|
||||
|
||||
# Insert order matters: monitor_target and monitor_run carry real foreign keys to
|
||||
# monitor_task, so the parent rows have to land first.
|
||||
TABLES_IN_ORDER = [
|
||||
"monitor_task",
|
||||
"monitor_target",
|
||||
"monitor_run",
|
||||
"monitor_note",
|
||||
"monitor_note_metric",
|
||||
"monitor_comment",
|
||||
"monitor_event",
|
||||
"monitor_setting",
|
||||
"auth_session",
|
||||
]
|
||||
|
||||
|
||||
def read_sqlite(path: Path) -> Dict[str, List[Dict[str, Any]]]:
|
||||
if not path.exists():
|
||||
raise SystemExit(f"找不到源库:{path}")
|
||||
|
||||
connection = sqlite3.connect(path)
|
||||
connection.row_factory = sqlite3.Row
|
||||
try:
|
||||
existing = {
|
||||
row[0]
|
||||
for row in connection.execute(
|
||||
"SELECT name FROM sqlite_master WHERE type='table'"
|
||||
)
|
||||
}
|
||||
data: Dict[str, List[Dict[str, Any]]] = {}
|
||||
for table in TABLES_IN_ORDER:
|
||||
if table not in existing:
|
||||
continue
|
||||
rows = [dict(row) for row in connection.execute(f"SELECT * FROM {table}")]
|
||||
if rows:
|
||||
data[table] = rows
|
||||
return data
|
||||
finally:
|
||||
connection.close()
|
||||
|
||||
|
||||
def migrate(source: Path, dry_run: bool, force: bool) -> None:
|
||||
import pymysql
|
||||
|
||||
from . import db as monitor_db
|
||||
|
||||
data = read_sqlite(source)
|
||||
if not data:
|
||||
print("源库里没有可迁移的数据。")
|
||||
return
|
||||
|
||||
print("源库内容:")
|
||||
for table, rows in data.items():
|
||||
print(f" {table:22} {len(rows)} 行")
|
||||
|
||||
url = monitor_db.resolve_db_url()
|
||||
if not url.startswith("mysql"):
|
||||
raise SystemExit(f"目标不是 MySQL:{url}")
|
||||
|
||||
connection = pymysql.connect(
|
||||
host=monitor_db.MYSQL_HOST(),
|
||||
port=monitor_db.MYSQL_PORT(),
|
||||
user=monitor_db.MYSQL_USER(),
|
||||
password=monitor_db.MYSQL_PWD(),
|
||||
database=monitor_db.MYSQL_DB_NAME(),
|
||||
charset="utf8mb4",
|
||||
autocommit=False,
|
||||
)
|
||||
|
||||
try:
|
||||
with connection.cursor() as cursor:
|
||||
# Never write outside the configured schema.
|
||||
cursor.execute("SELECT DATABASE()")
|
||||
current = cursor.fetchone()[0]
|
||||
expected = monitor_db.MYSQL_DB_NAME()
|
||||
if current.lower() != expected.lower():
|
||||
raise SystemExit(
|
||||
f"当前连接的是 {current!r},配置要求 {expected!r};已中止。"
|
||||
)
|
||||
|
||||
occupied = []
|
||||
for table in data:
|
||||
cursor.execute(f"SELECT COUNT(*) FROM `{table}`")
|
||||
if cursor.fetchone()[0]:
|
||||
occupied.append(table)
|
||||
|
||||
if occupied and not force:
|
||||
raise SystemExit(
|
||||
"目标库已有数据:" + ", ".join(occupied) + "\n"
|
||||
"加 --force 才会继续(会与现有数据并存,造成重复)。"
|
||||
)
|
||||
|
||||
if dry_run:
|
||||
print("\n[试运行] 未写入任何数据。")
|
||||
return
|
||||
|
||||
total = 0
|
||||
for table, rows in data.items():
|
||||
columns = list(rows[0].keys())
|
||||
column_sql = ", ".join(f"`{c}`" for c in columns)
|
||||
placeholders = ", ".join(["%s"] * len(columns))
|
||||
statement = (
|
||||
f"INSERT INTO `{table}` ({column_sql}) VALUES ({placeholders})"
|
||||
)
|
||||
cursor.executemany(
|
||||
statement, [[row[c] for c in columns] for row in rows]
|
||||
)
|
||||
total += len(rows)
|
||||
print(f" 已写入 {table:22} {len(rows)} 行")
|
||||
|
||||
connection.commit()
|
||||
print(f"\n完成,共迁移 {total} 行。")
|
||||
print("提示:源 SQLite 文件仍在原处,确认无误后自行删除。")
|
||||
|
||||
except Exception:
|
||||
connection.rollback()
|
||||
raise
|
||||
finally:
|
||||
connection.close()
|
||||
|
||||
|
||||
def main(argv: List[str] | None = None) -> int:
|
||||
parser = argparse.ArgumentParser(description="把监控库从 SQLite 迁到 MySQL")
|
||||
parser.add_argument(
|
||||
"--source",
|
||||
default=str(PROJECT_ROOT / "data" / "monitor.db"),
|
||||
help="SQLite 源文件路径",
|
||||
)
|
||||
parser.add_argument("--dry-run", action="store_true", help="只检查,不写入")
|
||||
parser.add_argument(
|
||||
"--force", action="store_true", help="目标库已有数据时也继续"
|
||||
)
|
||||
args = parser.parse_args(argv)
|
||||
|
||||
migrate(Path(args.source), args.dry_run, args.force)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,367 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/models.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Monitoring layer data model.
|
||||
|
||||
Lives in its own SQLite database (``data/monitor.db``) with its own declarative
|
||||
Base, deliberately separate from the crawler's ``database/models.py``. The
|
||||
crawler's DB store overwrites ``liked_count`` and friends in place on every
|
||||
re-crawl, so it cannot answer "how did this note change?". These tables keep the
|
||||
history the crawler throws away.
|
||||
|
||||
All timestamps are epoch **milliseconds** (BigInteger), matching the project's
|
||||
own ``tools.time_util.get_current_timestamp()`` convention. Using ints
|
||||
throughout avoids naive/aware datetime mixing bugs.
|
||||
"""
|
||||
|
||||
from typing import Optional
|
||||
|
||||
from sqlalchemy import (
|
||||
BigInteger,
|
||||
Boolean,
|
||||
ForeignKey,
|
||||
Integer,
|
||||
String,
|
||||
Text,
|
||||
UniqueConstraint,
|
||||
)
|
||||
from sqlalchemy.orm import DeclarativeBase, Mapped, mapped_column, relationship
|
||||
|
||||
|
||||
class MonitorBase(DeclarativeBase):
|
||||
"""Declarative base for the monitoring database."""
|
||||
|
||||
|
||||
# Run statuses
|
||||
RUN_PENDING = "pending"
|
||||
RUN_RUNNING = "running"
|
||||
RUN_SUCCESS = "success"
|
||||
RUN_PARTIAL = "partial"
|
||||
RUN_FAILED = "failed"
|
||||
RUN_TIMEOUT = "timeout"
|
||||
RUN_INTERRUPTED = "interrupted"
|
||||
|
||||
# Event types
|
||||
EVENT_NEW_NOTE = "new_note"
|
||||
EVENT_NEW_COMMENT_POSTED = "new_comment_posted"
|
||||
EVENT_NEW_COMMENT_SEEN = "new_comment_seen"
|
||||
EVENT_METRIC_DELTA = "metric_delta"
|
||||
EVENT_RUN_FAILED = "run_failed"
|
||||
EVENT_AUTH_FAILURE = "suspected_auth_failure"
|
||||
# A run that completed cleanly yet fetched nothing, where the login is provably
|
||||
# fine because another task just succeeded with it. The target, not the cookie,
|
||||
# is what needs looking at.
|
||||
EVENT_NO_DATA = "no_data_found"
|
||||
|
||||
# Task modes. One subprocess handles exactly one crawler type, so a task is
|
||||
# either creator-driven or note-driven -- never both.
|
||||
MODE_CREATOR = "creator"
|
||||
MODE_NOTE = "note"
|
||||
|
||||
|
||||
class MonitorTask(MonitorBase):
|
||||
"""One monitored schedule: a set of targets plus an interval."""
|
||||
|
||||
__tablename__ = "monitor_task"
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
name: Mapped[str] = mapped_column(String(200), nullable=False)
|
||||
platform: Mapped[str] = mapped_column(String(32), nullable=False, default="xhs")
|
||||
mode: Mapped[str] = mapped_column(String(16), nullable=False)
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
|
||||
interval_minutes: Mapped[int] = mapped_column(Integer, nullable=False, default=360)
|
||||
|
||||
# Crawl window knobs, mirrored onto each run's CLI flags.
|
||||
max_notes_count: Mapped[int] = mapped_column(Integer, nullable=False, default=20)
|
||||
enable_comments: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
|
||||
max_comments_count: Mapped[int] = mapped_column(Integer, nullable=False, default=50)
|
||||
run_timeout_seconds: Mapped[int] = mapped_column(Integer, nullable=False, default=3600)
|
||||
|
||||
# Push notifications are opt-in per task. A task list that all pushes to one
|
||||
# webhook turns noisy fast, so silence is the default.
|
||||
notify_enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
|
||||
|
||||
# Scheduler state. Persisted so the schedule survives an API restart.
|
||||
next_run_at: Mapped[Optional[int]] = mapped_column(BigInteger, index=True)
|
||||
last_run_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
last_status: Mapped[str] = mapped_column(String(32), nullable=False, default="idle")
|
||||
last_error: Mapped[Optional[str]] = mapped_column(Text)
|
||||
# Lets the UI answer "why did I not get a push for this run?".
|
||||
last_notified_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
|
||||
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
targets: Mapped[list["MonitorTarget"]] = relationship(
|
||||
back_populates="task",
|
||||
cascade="all, delete-orphan",
|
||||
lazy="selectin",
|
||||
)
|
||||
|
||||
|
||||
class MonitorTarget(MonitorBase):
|
||||
"""One watched creator or note belonging to a task.
|
||||
|
||||
``external_id`` is the stable identity (XHS user_id / note_id). It is kept
|
||||
separate from ``xsec_token`` on purpose: tokens expire within weeks, so
|
||||
treating a tokenised URL as the primary key would make every long-running
|
||||
task fail eventually.
|
||||
"""
|
||||
|
||||
__tablename__ = "monitor_target"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("task_id", "kind", "external_id", name="uq_monitor_target"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
task_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
kind: Mapped[str] = mapped_column(String(16), nullable=False)
|
||||
external_id: Mapped[str] = mapped_column(String(128), nullable=False)
|
||||
xsec_token: Mapped[str] = mapped_column(String(512), nullable=False, default="")
|
||||
xsec_source: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
raw_value: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
label: Mapped[str] = mapped_column(String(200), nullable=False, default="")
|
||||
enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True)
|
||||
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
task: Mapped["MonitorTask"] = relationship(back_populates="targets")
|
||||
|
||||
|
||||
class MonitorRun(MonitorBase):
|
||||
"""One subprocess execution. The run history in the UI is this table."""
|
||||
|
||||
__tablename__ = "monitor_run"
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
task_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
trigger: Mapped[str] = mapped_column(String(16), nullable=False, default="scheduled")
|
||||
status: Mapped[str] = mapped_column(String(16), nullable=False, default=RUN_PENDING, index=True)
|
||||
phase: Mapped[str] = mapped_column(String(16), nullable=False)
|
||||
|
||||
# Where this run's jsonl landed. Each run gets its own directory because the
|
||||
# crawler's file writer names output by date only.
|
||||
save_data_path: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
|
||||
queued_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
not_before: Mapped[int] = mapped_column(BigInteger, nullable=False, default=0)
|
||||
started_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
finished_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
# BigInteger, not Integer: Windows reports failures as unsigned 32-bit
|
||||
# NTSTATUS values (0xC0000142 = 3221225794), which overflow MySQL's signed
|
||||
# INT. SQLite's dynamic typing hid this until the data was migrated.
|
||||
exit_code: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
|
||||
notes_fetched: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
comments_fetched: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
new_notes: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
new_comments: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
|
||||
# The very first successful run of a task establishes the baseline: every
|
||||
# note is "new" at that point, so emitting events would be pure noise.
|
||||
is_baseline: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
|
||||
|
||||
# Window actually used, so the UI can be honest that comments are the top N
|
||||
# in the platform's own ordering rather than a complete set.
|
||||
max_comments_count: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
|
||||
error_message: Mapped[Optional[str]] = mapped_column(Text)
|
||||
|
||||
|
||||
class MonitorNote(MonitorBase):
|
||||
"""A note ever seen by a task, plus when it was first/last seen.
|
||||
|
||||
Grain is (task, note) so the same note tracked by two tasks stays independent.
|
||||
"""
|
||||
|
||||
__tablename__ = "monitor_note"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("task_id", "note_id", name="uq_monitor_note"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
task_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("monitor_task.id", ondelete="CASCADE"), nullable=False, index=True
|
||||
)
|
||||
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
|
||||
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
note_url: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
cover: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
source_kind: Mapped[str] = mapped_column(String(16), nullable=False, default="")
|
||||
published_at: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
|
||||
first_seen_run_id: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
first_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
last_seen_run_id: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
last_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
|
||||
class MonitorNoteMetric(MonitorBase):
|
||||
"""One metric snapshot per (task, note, run) -- the time series.
|
||||
|
||||
Raw strings are kept alongside the parsed integers so a mis-parsed "1.2万"
|
||||
can always be audited after the fact.
|
||||
"""
|
||||
|
||||
__tablename__ = "monitor_note_metric"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("task_id", "note_id", "run_id", name="uq_note_metric"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
task_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
|
||||
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
|
||||
run_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
|
||||
captured_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
# NULL (not 0) when the platform value could not be parsed: storing 0 would
|
||||
# forge a large negative delta on the next comparison.
|
||||
liked_count: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
comment_count: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
collected_count: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
share_count: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
|
||||
raw_liked_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
raw_comment_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
raw_collected_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
raw_share_count: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
|
||||
|
||||
class MonitorComment(MonitorBase):
|
||||
"""A comment ever seen by a task.
|
||||
|
||||
The (task, note, comment) uniqueness gives idempotent dedup across runs for
|
||||
free -- re-running the same crawl cannot double-count.
|
||||
"""
|
||||
|
||||
__tablename__ = "monitor_comment"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("task_id", "note_id", "comment_id", name="uq_monitor_comment"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
task_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
|
||||
note_id: Mapped[str] = mapped_column(String(128), nullable=False, index=True)
|
||||
comment_id: Mapped[str] = mapped_column(String(128), nullable=False)
|
||||
content: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
nickname: Mapped[str] = mapped_column(String(200), nullable=False, default="")
|
||||
creator_hash: Mapped[str] = mapped_column(String(64), nullable=False, default="")
|
||||
# Platform-stated publish time. Used to distinguish a genuinely new comment
|
||||
# from one that merely entered the visible top-N window this run.
|
||||
create_time: Mapped[Optional[int]] = mapped_column(BigInteger)
|
||||
like_count: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
sub_comment_count: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
|
||||
parent_comment_id: Mapped[str] = mapped_column(String(128), nullable=False, default="")
|
||||
|
||||
first_seen_run_id: Mapped[Optional[int]] = mapped_column(Integer)
|
||||
first_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
|
||||
class MonitorEvent(MonitorBase):
|
||||
"""Append-only change feed. This is what the dashboard reads."""
|
||||
|
||||
__tablename__ = "monitor_event"
|
||||
|
||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
|
||||
task_id: Mapped[int] = mapped_column(Integer, nullable=False, index=True)
|
||||
run_id: Mapped[Optional[int]] = mapped_column(Integer, index=True)
|
||||
type: Mapped[str] = mapped_column(String(32), nullable=False, index=True)
|
||||
severity: Mapped[str] = mapped_column(String(16), nullable=False, default="info")
|
||||
target_kind: Mapped[str] = mapped_column(String(16), nullable=False, default="")
|
||||
target_id: Mapped[str] = mapped_column(String(128), nullable=False, default="")
|
||||
title: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
payload_json: Mapped[str] = mapped_column(Text, nullable=False, default="{}")
|
||||
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
|
||||
is_read: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
|
||||
|
||||
|
||||
class MonitorSetting(MonitorBase):
|
||||
"""Key/value store. Holds the XHS cookie for unattended runs."""
|
||||
|
||||
__tablename__ = "monitor_setting"
|
||||
|
||||
key: Mapped[str] = mapped_column(String(64), primary_key=True)
|
||||
value: Mapped[str] = mapped_column(Text, nullable=False, default="")
|
||||
updated_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
|
||||
class AuthSession(MonitorBase):
|
||||
"""A WebUI login session.
|
||||
|
||||
Only the SHA-256 of the token is stored, never the token itself -- a leaked
|
||||
database therefore does not hand over live sessions. This mirrors the
|
||||
existing posture of never returning the XHS cookie or webhook value.
|
||||
|
||||
A stateful table (rather than a signed stateless token) is what makes "log
|
||||
out" and "password changed" take effect immediately.
|
||||
"""
|
||||
|
||||
__tablename__ = "auth_session"
|
||||
|
||||
token_hash: Mapped[str] = mapped_column(String(64), primary_key=True)
|
||||
created_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
expires_at: Mapped[int] = mapped_column(BigInteger, nullable=False, index=True)
|
||||
last_seen_at: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
|
||||
|
||||
SETTING_AUTH_PASSWORD_HASH = "auth_password_hash"
|
||||
SETTING_AUTH_PASSWORD_UPDATED_AT = "auth_password_updated_at"
|
||||
|
||||
# Settings are namespaced by scope: `platform.<p>.<name>` for values each
|
||||
# platform keeps its own copy of, `system.<name>` for values shared across all of
|
||||
# them. Key builders live in settings.py.
|
||||
SETTING_WECOM_WEBHOOK = "system.wecom_webhook"
|
||||
|
||||
# Pre-namespacing keys, kept only so the startup migration can find and move
|
||||
# them. Nothing should read these directly.
|
||||
LEGACY_SETTING_KEY_RENAMES = {
|
||||
# Pre-batch-2 flat keys.
|
||||
"xhs_cookie": "platform.xhs.cookie",
|
||||
"xhs_cookie_updated_at": "platform.xhs.cookie_updated_at",
|
||||
"xhs_cookie_last_ok_at": "platform.xhs.cookie_last_ok_at",
|
||||
"wecom_webhook": "system.wecom_webhook",
|
||||
# Batch-2 keys, before settings gained a scope. Those values belonged to
|
||||
# Xiaohongshu because it was the only platform, so they migrate to its scope;
|
||||
# the two scheduling keys were always instance-wide.
|
||||
"collect.default_interval_minutes": "platform.xhs.default_interval_minutes",
|
||||
"collect.default_max_notes": "platform.xhs.default_max_notes",
|
||||
"collect.default_max_comments": "platform.xhs.default_max_comments",
|
||||
"collect.enable_sub_comments": "platform.xhs.enable_sub_comments",
|
||||
"collect.crawl_sleep_sec": "platform.xhs.crawl_sleep_sec",
|
||||
"collect.active_hours_start": "system.active_hours_start",
|
||||
"collect.active_hours_end": "system.active_hours_end",
|
||||
"proxy.enable_ip_proxy": "platform.xhs.enable_ip_proxy",
|
||||
"proxy.provider": "platform.xhs.proxy_provider",
|
||||
"proxy.pool_count": "platform.xhs.proxy_pool_count",
|
||||
"proxy.static_proxy_url": "platform.xhs.static_proxy_url",
|
||||
}
|
||||
|
||||
|
||||
# utf8mb4 is forced on every table rather than left to the schema default: this
|
||||
# deployment's MySQL server *and* the target database both default to latin1,
|
||||
# which would mangle or reject Chinese text. Setting it per table means it holds
|
||||
# regardless of what the schema default happens to be.
|
||||
#
|
||||
# Must run after every model is declared, hence the end of the module.
|
||||
for _table in MonitorBase.metadata.tables.values():
|
||||
_table.kwargs["mysql_charset"] = "utf8mb4"
|
||||
_table.kwargs["mysql_collate"] = "utf8mb4_unicode_ci"
|
||||
@@ -0,0 +1,200 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/notify.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Push notifications via a WeCom (企业微信) group robot webhook.
|
||||
|
||||
Two rules shape this module:
|
||||
|
||||
* **One message per run, not per event.** A run that finds twenty new notes must
|
||||
produce one summary, not twenty pushes.
|
||||
* **A failed push never fails the crawl.** Notification is best-effort: the run's
|
||||
data is already committed by the time we get here, so every error is logged
|
||||
and swallowed.
|
||||
"""
|
||||
|
||||
import json
|
||||
from typing import Optional
|
||||
|
||||
import httpx
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from .models import (
|
||||
EVENT_AUTH_FAILURE,
|
||||
EVENT_NEW_NOTE,
|
||||
EVENT_NO_DATA,
|
||||
EVENT_RUN_FAILED,
|
||||
MonitorEvent,
|
||||
MonitorRun,
|
||||
MonitorTask,
|
||||
)
|
||||
from .settings import get_setting
|
||||
|
||||
# Short on purpose: the scheduler awaits the run, so a hanging webhook would
|
||||
# stall every other task behind it.
|
||||
WEBHOOK_TIMEOUT_SECONDS = 10.0
|
||||
|
||||
# Only these event types are worth interrupting someone for. NO_DATA is included
|
||||
# because a run that fetched nothing at all is always anomalous -- a creator
|
||||
# always has *some* notes -- even when the login is not the culprit.
|
||||
NOTIFIABLE_EVENT_TYPES = (
|
||||
EVENT_AUTH_FAILURE,
|
||||
EVENT_RUN_FAILED,
|
||||
EVENT_NO_DATA,
|
||||
EVENT_NEW_NOTE,
|
||||
)
|
||||
|
||||
# WeCom markdown is a limited subset; coloured text is the one bit of flair it
|
||||
# supports and it makes failures stand out in a busy group chat.
|
||||
_COLOR_WARNING = "warning"
|
||||
_COLOR_INFO = "info"
|
||||
|
||||
|
||||
async def send_wecom(webhook_url: str, content: str) -> tuple[bool, str]:
|
||||
"""Post a markdown message to a WeCom group robot.
|
||||
|
||||
Returns (ok, detail) rather than raising, so callers can surface the reason
|
||||
in the UI when the user clicks "send test".
|
||||
"""
|
||||
if not webhook_url:
|
||||
return False, "Webhook 未配置"
|
||||
|
||||
payload = {"msgtype": "markdown", "markdown": {"content": content}}
|
||||
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=WEBHOOK_TIMEOUT_SECONDS) as client:
|
||||
response = await client.post(webhook_url, json=payload)
|
||||
response.raise_for_status()
|
||||
body = response.json()
|
||||
except httpx.HTTPError as exc:
|
||||
return False, f"请求失败:{exc}"
|
||||
except json.JSONDecodeError:
|
||||
return False, "返回内容不是合法 JSON,请检查 Webhook 地址"
|
||||
|
||||
# WeCom answers 200 with a non-zero errcode on failure.
|
||||
errcode = body.get("errcode")
|
||||
if errcode != 0:
|
||||
return False, f"企业微信返回 errcode={errcode} {body.get('errmsg', '')}"
|
||||
|
||||
return True, "发送成功"
|
||||
|
||||
|
||||
async def get_webhook_url(session: AsyncSession) -> str:
|
||||
from .models import SETTING_WECOM_WEBHOOK
|
||||
|
||||
return (await get_setting(session, SETTING_WECOM_WEBHOOK)) or ""
|
||||
|
||||
|
||||
async def build_run_message(
|
||||
session: AsyncSession,
|
||||
task: MonitorTask,
|
||||
run: MonitorRun,
|
||||
) -> Optional[str]:
|
||||
"""Compose one markdown summary for a finished run, or None if nothing to say."""
|
||||
events = list(
|
||||
(
|
||||
await session.scalars(
|
||||
select(MonitorEvent)
|
||||
.where(
|
||||
MonitorEvent.run_id == run.id,
|
||||
MonitorEvent.type.in_(NOTIFIABLE_EVENT_TYPES),
|
||||
)
|
||||
.order_by(MonitorEvent.id)
|
||||
)
|
||||
).all()
|
||||
)
|
||||
if not events:
|
||||
return None
|
||||
|
||||
failures = [
|
||||
e for e in events if e.type in (EVENT_AUTH_FAILURE, EVENT_RUN_FAILED, EVENT_NO_DATA)
|
||||
]
|
||||
new_notes = [e for e in events if e.type == EVENT_NEW_NOTE]
|
||||
|
||||
lines: list[str] = []
|
||||
|
||||
if failures:
|
||||
# Word the header from what actually happened, not from whether new notes
|
||||
# accompanied it: a login outage usually brings no new notes either.
|
||||
unavailable = any(e.type == EVENT_NO_DATA for e in failures) and not any(
|
||||
e.type in (EVENT_AUTH_FAILURE, EVENT_RUN_FAILED) for e in failures
|
||||
)
|
||||
header = "监控任务未抓到数据" if unavailable else "监控任务异常"
|
||||
lines.append(f"**⚠️ {header}:{task.name}**")
|
||||
for event in failures:
|
||||
lines.append(f'> <font color="{_COLOR_WARNING}">{event.title}</font>')
|
||||
else:
|
||||
lines.append(f"**📢 监控任务有新作品:{task.name}**")
|
||||
|
||||
if new_notes:
|
||||
lines.append(f"> 新增作品 **{len(new_notes)}** 篇")
|
||||
# Cap the listing: a first-ever run or a long gap can produce a lot, and
|
||||
# a wall of text is worse than a count.
|
||||
for event in new_notes[:10]:
|
||||
payload = _load_payload(event.payload_json)
|
||||
title = payload.get("title") or event.target_id
|
||||
note_id = payload.get("note_id") or event.target_id
|
||||
url = f"https://www.xiaohongshu.com/explore/{note_id}"
|
||||
lines.append(f"> [{title}]({url})")
|
||||
if len(new_notes) > 10:
|
||||
lines.append(f"> …等共 {len(new_notes)} 篇")
|
||||
|
||||
if run.is_baseline:
|
||||
lines.append("> (首轮基线,未计入新增统计)")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def _load_payload(raw: str) -> dict:
|
||||
try:
|
||||
payload = json.loads(raw or "{}")
|
||||
except json.JSONDecodeError:
|
||||
return {}
|
||||
return payload if isinstance(payload, dict) else {}
|
||||
|
||||
|
||||
async def notify_run(session: AsyncSession, task: MonitorTask, run: MonitorRun) -> Optional[str]:
|
||||
"""Push a summary for a finished run if the task opted in.
|
||||
|
||||
Returns the message that was sent, or None. Never raises.
|
||||
"""
|
||||
try:
|
||||
if not task.notify_enabled:
|
||||
return None
|
||||
|
||||
webhook_url = await get_webhook_url(session)
|
||||
if not webhook_url:
|
||||
return None
|
||||
|
||||
message = await build_run_message(session, task, run)
|
||||
if not message:
|
||||
return None
|
||||
|
||||
ok, detail = await send_wecom(webhook_url, message)
|
||||
if not ok:
|
||||
print(f"[monitor.notify] task {task.id} push failed: {detail}")
|
||||
return None
|
||||
|
||||
task.last_notified_at = get_current_timestamp()
|
||||
return message
|
||||
|
||||
except Exception as exc: # pragma: no cover - notification must never break a run
|
||||
print(f"[monitor.notify] unexpected error: {exc}")
|
||||
return None
|
||||
@@ -0,0 +1,197 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/platforms.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Platform capability matrix.
|
||||
|
||||
The single source of truth for what each platform can do. The UI renders its
|
||||
platform switcher and metric columns from this, and the API validates against
|
||||
it.
|
||||
|
||||
Two distinct things are recorded here, and conflating them would be misleading:
|
||||
|
||||
* ``crawler_modes`` / ``metrics`` / ``comment_levels`` / ``media`` describe what
|
||||
the upstream crawler module actually supports. These were read out of the
|
||||
platform modules, not assumed -- all seven implement search/detail/creator;
|
||||
the real differences are in which interaction metrics they capture.
|
||||
* ``monitor_wired`` says whether the *monitoring layer* has been hooked up. It
|
||||
currently covers only Xiaohongshu: ``runner.py`` pins the platform,
|
||||
``ingest.py`` reads a fixed ``xhs/jsonl`` directory, and ``service.py`` only
|
||||
parses Xiaohongshu target URLs.
|
||||
|
||||
A platform can therefore be fully crawlable by upstream and still not usable for
|
||||
monitoring, which is exactly the state of the other six today.
|
||||
"""
|
||||
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
PLATFORM_XHS = "xhs"
|
||||
|
||||
PLATFORM_LABELS = {
|
||||
"xhs": "小红书",
|
||||
"dy": "抖音",
|
||||
"ks": "快手",
|
||||
"bili": "B站",
|
||||
"wb": "微博",
|
||||
"tieba": "贴吧",
|
||||
"zhihu": "知乎",
|
||||
}
|
||||
|
||||
# Interaction metrics each platform's store actually persists. Xiaohongshu has no
|
||||
# play count or danmaku; Bilibili has both and the widest set; Kuaishou carries
|
||||
# no comment/share/collect at all; Tieba stores only reply counts.
|
||||
PLATFORM_CAPABILITIES: Dict[str, Dict[str, Any]] = {
|
||||
"xhs": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
|
||||
"comment_levels": 2,
|
||||
"media": True,
|
||||
"monitor_wired": True,
|
||||
},
|
||||
"dy": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
"metrics": ["liked_count", "comment_count", "collected_count", "share_count"],
|
||||
"comment_levels": 2,
|
||||
"media": True,
|
||||
"monitor_wired": False,
|
||||
},
|
||||
"ks": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
# No comment/share/collect in the Kuaishou store; sub-comments are stored
|
||||
# flat with no parent link and carry no like count.
|
||||
"metrics": ["liked_count", "view_count"],
|
||||
"comment_levels": 1,
|
||||
"media": True,
|
||||
"monitor_wired": False,
|
||||
},
|
||||
"bili": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
"metrics": [
|
||||
"liked_count",
|
||||
"video_play_count",
|
||||
"video_danmaku",
|
||||
"comment_count",
|
||||
"video_favorite_count",
|
||||
"video_coin_count",
|
||||
"video_share_count",
|
||||
],
|
||||
"comment_levels": 2,
|
||||
"media": True,
|
||||
"monitor_wired": False,
|
||||
},
|
||||
"wb": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
# Weibo has no collect count, and its comment count field is named
|
||||
# differently in the model.
|
||||
"metrics": ["liked_count", "comments_count", "shared_count"],
|
||||
"comment_levels": 2,
|
||||
"media": True,
|
||||
"monitor_wired": False,
|
||||
},
|
||||
"tieba": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
"metrics": ["total_replay_num", "total_replay_page"],
|
||||
"comment_levels": 2,
|
||||
"media": False,
|
||||
"monitor_wired": False,
|
||||
},
|
||||
"zhihu": {
|
||||
"crawler_modes": ["search", "detail", "creator"],
|
||||
"metrics": ["voteup_count", "comment_count"],
|
||||
"comment_levels": 2,
|
||||
"media": False,
|
||||
"monitor_wired": False,
|
||||
},
|
||||
}
|
||||
|
||||
METRIC_LABELS = {
|
||||
"liked_count": "点赞",
|
||||
"comment_count": "评论",
|
||||
"collected_count": "收藏",
|
||||
"share_count": "分享",
|
||||
"view_count": "播放",
|
||||
"video_play_count": "播放",
|
||||
"video_danmaku": "弹幕",
|
||||
"video_favorite_count": "收藏",
|
||||
"video_coin_count": "投币",
|
||||
"video_share_count": "分享",
|
||||
"comments_count": "评论",
|
||||
"shared_count": "转发",
|
||||
"total_replay_num": "回复数",
|
||||
"total_replay_page": "回复页数",
|
||||
"voteup_count": "赞同",
|
||||
}
|
||||
|
||||
# Monitoring modes, mapped to the CLI crawler types upstream understands.
|
||||
MONITOR_MODE_CREATOR = "creator"
|
||||
MONITOR_MODE_NOTE = "note"
|
||||
CLI_TYPE_FOR_MODE = {
|
||||
MONITOR_MODE_CREATOR: "creator",
|
||||
MONITOR_MODE_NOTE: "detail",
|
||||
}
|
||||
|
||||
|
||||
class UnsupportedPlatformError(ValueError):
|
||||
"""Raised for an unknown platform, or one the monitor layer cannot run."""
|
||||
|
||||
|
||||
def all_platforms() -> List[str]:
|
||||
return list(PLATFORM_CAPABILITIES)
|
||||
|
||||
|
||||
def is_known(platform: str) -> bool:
|
||||
return platform in PLATFORM_CAPABILITIES
|
||||
|
||||
|
||||
def is_monitor_wired(platform: str) -> bool:
|
||||
return bool(PLATFORM_CAPABILITIES.get(platform, {}).get("monitor_wired"))
|
||||
|
||||
|
||||
def describe(platform: str) -> Optional[Dict[str, Any]]:
|
||||
capability = PLATFORM_CAPABILITIES.get(platform)
|
||||
if capability is None:
|
||||
return None
|
||||
return {
|
||||
"value": platform,
|
||||
"label": PLATFORM_LABELS.get(platform, platform),
|
||||
**capability,
|
||||
"metric_labels": {
|
||||
metric: METRIC_LABELS.get(metric, metric) for metric in capability["metrics"]
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def describe_all() -> List[Dict[str, Any]]:
|
||||
return [describe(platform) for platform in all_platforms()]
|
||||
|
||||
|
||||
def ensure_runnable(platform: str) -> None:
|
||||
"""Validate a platform for a monitoring task.
|
||||
|
||||
An unwired platform is rejected outright rather than accepted and left to
|
||||
silently produce nothing -- the same silent-failure shape that made a valid
|
||||
creator look like an expired login earlier.
|
||||
"""
|
||||
if not is_known(platform):
|
||||
raise UnsupportedPlatformError(
|
||||
f"未知平台:{platform}(支持:{', '.join(all_platforms())})"
|
||||
)
|
||||
if not is_monitor_wired(platform):
|
||||
label = PLATFORM_LABELS.get(platform, platform)
|
||||
raise UnsupportedPlatformError(
|
||||
f"{label}的爬虫已支持,但监控层尚未接通,暂时无法创建监控任务。"
|
||||
)
|
||||
@@ -0,0 +1,203 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/report.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Cross-task reporting: what grew, and what is new, over a date range.
|
||||
|
||||
Two families of numbers that answer different questions and are therefore kept
|
||||
as separate columns:
|
||||
|
||||
* **互动增量** — Σ(current − previous) across the selected notes. "How many likes
|
||||
did this set of notes gain?"
|
||||
* **新增内容** — count of newly discovered notes and comments. "How much new
|
||||
material showed up?"
|
||||
|
||||
The per-day interaction delta is defined as *last value on the day* minus *last
|
||||
value before the day* (0 when the note was first seen on that day). That keeps
|
||||
growth from a note's first observation counted once, rather than smeared across
|
||||
every later day.
|
||||
|
||||
Aggregation runs in Python over the snapshots rather than as one large SQL
|
||||
query: the per-note-per-day baseline lookup is a windowed operation that SQLite
|
||||
expresses awkwardly, and the row counts here are small enough that clarity is
|
||||
worth more than the query planner.
|
||||
"""
|
||||
|
||||
from bisect import bisect_right
|
||||
from datetime import date, datetime, time, timedelta
|
||||
from typing import Any, Dict, Iterable, List, Optional, Sequence
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from .models import MonitorComment, MonitorNote, MonitorNoteMetric
|
||||
|
||||
METRIC_FIELDS = ("liked_count", "comment_count", "collected_count", "share_count")
|
||||
|
||||
METRIC_LABELS = {
|
||||
"liked_count": "点赞",
|
||||
"comment_count": "评论",
|
||||
"collected_count": "收藏",
|
||||
"share_count": "分享",
|
||||
}
|
||||
|
||||
|
||||
def day_bounds(day: date) -> tuple[int, int]:
|
||||
"""Inclusive epoch-millisecond bounds for a local calendar day."""
|
||||
start = datetime.combine(day, time.min)
|
||||
end = datetime.combine(day, time.max)
|
||||
return int(start.timestamp() * 1000), int(end.timestamp() * 1000)
|
||||
|
||||
|
||||
def iter_days(start: date, end: date) -> List[date]:
|
||||
days = []
|
||||
cursor = start
|
||||
while cursor <= end:
|
||||
days.append(cursor)
|
||||
cursor += timedelta(days=1)
|
||||
return days
|
||||
|
||||
|
||||
def compute_daily_rows(
|
||||
series_by_note: Dict[str, List[tuple[int, Dict[str, Optional[int]]]]],
|
||||
notes_per_day: Dict[date, int],
|
||||
comments_per_day: Dict[date, int],
|
||||
days: Sequence[date],
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Pure aggregation. ``series_by_note`` must be sorted by timestamp ascending."""
|
||||
prepared = {note_id: ([ts for ts, _ in points], points) for note_id, points in series_by_note.items()}
|
||||
|
||||
rows: List[Dict[str, Any]] = []
|
||||
for day in days:
|
||||
day_start, day_end = day_bounds(day)
|
||||
totals = {field: 0 for field in METRIC_FIELDS}
|
||||
# Records *which* metric could not be compared, not just that something
|
||||
# could not. A blanket flag loses all value the moment one permanently
|
||||
# unparseable field makes every row "incomplete".
|
||||
partial_metrics: set[str] = set()
|
||||
|
||||
for times, points in prepared.values():
|
||||
end_index = bisect_right(times, day_end) - 1
|
||||
if end_index < 0:
|
||||
# Not yet tracked on this day.
|
||||
continue
|
||||
|
||||
end_values = points[end_index][1]
|
||||
start_index = bisect_right(times, day_start - 1) - 1
|
||||
# No earlier snapshot means the note first appeared in this window,
|
||||
# so it starts from zero -- all of its count is genuinely new.
|
||||
start_values = (
|
||||
points[start_index][1] if start_index >= 0 else {f: 0 for f in METRIC_FIELDS}
|
||||
)
|
||||
|
||||
for field in METRIC_FIELDS:
|
||||
end_value, start_value = end_values.get(field), start_values.get(field)
|
||||
if end_value is None or start_value is None:
|
||||
# An unparseable count on either side makes the delta unknown;
|
||||
# skipping beats reporting a fabricated number.
|
||||
partial_metrics.add(field)
|
||||
continue
|
||||
totals[field] += end_value - start_value
|
||||
|
||||
row: Dict[str, Any] = {
|
||||
"date": day.isoformat(),
|
||||
"new_notes": notes_per_day.get(day, 0),
|
||||
"new_comments": comments_per_day.get(day, 0),
|
||||
"partial_metrics": sorted(partial_metrics),
|
||||
}
|
||||
row.update({f"{field}_delta": value for field, value in totals.items()})
|
||||
rows.append(row)
|
||||
|
||||
return rows
|
||||
|
||||
|
||||
async def build_report(
|
||||
session: AsyncSession,
|
||||
task_ids: Optional[Iterable[int]],
|
||||
start_day: date,
|
||||
end_day: date,
|
||||
) -> Dict[str, Any]:
|
||||
"""Daily rows plus totals for the selected tasks over the given date range."""
|
||||
start_ms, _ = day_bounds(start_day)
|
||||
_, end_ms = day_bounds(end_day)
|
||||
|
||||
scope = list(task_ids) if task_ids else None
|
||||
days = iter_days(start_day, end_day)
|
||||
|
||||
# Fetch every snapshot up to the range end: the delta on the first day needs
|
||||
# the last value from *before* the range, so a lower bound would be wrong.
|
||||
metric_stmt = select(MonitorNoteMetric).where(MonitorNoteMetric.captured_at <= end_ms)
|
||||
if scope is not None:
|
||||
metric_stmt = metric_stmt.where(MonitorNoteMetric.task_id.in_(scope))
|
||||
metric_stmt = metric_stmt.order_by(MonitorNoteMetric.note_id, MonitorNoteMetric.run_id)
|
||||
|
||||
series_by_note: Dict[str, List[tuple[int, Dict[str, Optional[int]]]]] = {}
|
||||
included_note_ids: set[str] = set()
|
||||
for snapshot in (await session.scalars(metric_stmt)).all():
|
||||
included_note_ids.add(snapshot.note_id)
|
||||
series_by_note.setdefault(snapshot.note_id, []).append(
|
||||
(
|
||||
snapshot.captured_at,
|
||||
{field: getattr(snapshot, field) for field in METRIC_FIELDS},
|
||||
)
|
||||
)
|
||||
|
||||
note_stmt = select(MonitorNote.first_seen_at).where(
|
||||
MonitorNote.first_seen_at >= start_ms, MonitorNote.first_seen_at <= end_ms
|
||||
)
|
||||
if scope is not None:
|
||||
note_stmt = note_stmt.where(MonitorNote.task_id.in_(scope))
|
||||
|
||||
comment_stmt = select(MonitorComment.first_seen_at).where(
|
||||
MonitorComment.first_seen_at >= start_ms, MonitorComment.first_seen_at <= end_ms
|
||||
)
|
||||
if scope is not None:
|
||||
comment_stmt = comment_stmt.where(MonitorComment.task_id.in_(scope))
|
||||
|
||||
notes_per_day = _count_by_day((await session.scalars(note_stmt)).all())
|
||||
comments_per_day = _count_by_day((await session.scalars(comment_stmt)).all())
|
||||
|
||||
rows = compute_daily_rows(series_by_note, notes_per_day, comments_per_day, days)
|
||||
|
||||
totals = {
|
||||
"new_notes": sum(row["new_notes"] for row in rows),
|
||||
"new_comments": sum(row["new_comments"] for row in rows),
|
||||
}
|
||||
for field in METRIC_FIELDS:
|
||||
totals[f"{field}_delta"] = sum(row[f"{field}_delta"] for row in rows)
|
||||
|
||||
return {
|
||||
"start_date": start_day.isoformat(),
|
||||
"end_date": end_day.isoformat(),
|
||||
"task_ids": scope,
|
||||
"rows": rows,
|
||||
"totals": totals,
|
||||
"note_count": len(included_note_ids),
|
||||
"has_partial_data": any(row["partial_metrics"] for row in rows),
|
||||
"partial_metrics": sorted({field for row in rows for field in row["partial_metrics"]}),
|
||||
"metric_labels": METRIC_LABELS,
|
||||
}
|
||||
|
||||
|
||||
def _count_by_day(timestamps: Iterable[Optional[int]]) -> Dict[date, int]:
|
||||
counts: Dict[date, int] = {}
|
||||
for ts in timestamps:
|
||||
if ts is None:
|
||||
continue
|
||||
day = datetime.fromtimestamp(ts / 1000).date()
|
||||
counts[day] = counts.get(day, 0) + 1
|
||||
return counts
|
||||
@@ -0,0 +1,277 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/runner.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Execute a single monitoring run: build the command, wait, then ingest.
|
||||
|
||||
Runs reuse ``CrawlerManager`` so that monitor crawls share the existing
|
||||
single-subprocess guarantee and their logs stream to the existing Terminal
|
||||
component over the existing log WebSocket.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import os
|
||||
from pathlib import Path
|
||||
from typing import Iterable, List, Optional
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from ..schemas import (
|
||||
CrawlerStartRequest,
|
||||
CrawlerTypeEnum,
|
||||
LoginTypeEnum,
|
||||
PlatformEnum,
|
||||
SaveDataOptionEnum,
|
||||
)
|
||||
from ..services import crawler_manager
|
||||
from . import app_settings, notify
|
||||
from .db import get_session
|
||||
from .ingest import IngestResult, ingest_run
|
||||
from .models import (
|
||||
MODE_CREATOR,
|
||||
RUN_FAILED,
|
||||
RUN_PENDING,
|
||||
RUN_RUNNING,
|
||||
RUN_TIMEOUT,
|
||||
MonitorRun,
|
||||
MonitorTarget,
|
||||
MonitorTask,
|
||||
)
|
||||
from .settings import get_cookie, mark_cookie_ok
|
||||
|
||||
PROJECT_ROOT = Path(__file__).parent.parent.parent
|
||||
MONITOR_RUNS_DIR = PROJECT_ROOT / "data" / "monitor_runs"
|
||||
|
||||
# Monitor platform ids align with PlatformEnum's values, but mapping explicitly
|
||||
# beats relying on that coincidence.
|
||||
_PLATFORM_ENUM = {
|
||||
"xhs": PlatformEnum.XHS,
|
||||
"dy": PlatformEnum.DOUYIN,
|
||||
"ks": PlatformEnum.KUAISHOU,
|
||||
"bili": PlatformEnum.BILIBILI,
|
||||
"wb": PlatformEnum.WEIBO,
|
||||
"tieba": PlatformEnum.TIEBA,
|
||||
"zhihu": PlatformEnum.ZHIHU,
|
||||
}
|
||||
|
||||
_XHS_WEB_BASE = "https://www.xiaohongshu.com"
|
||||
_CREATOR_PATH = "/user/profile"
|
||||
_NOTE_PATH = "/explore"
|
||||
|
||||
# Timeout used when the caller does not care; tasks carry their own.
|
||||
DEFAULT_RUN_TIMEOUT_SECONDS = 3600
|
||||
|
||||
|
||||
def build_target_url(value: str, kind: str) -> str:
|
||||
"""Turn a stored target into a URL the crawler's parser accepts.
|
||||
|
||||
Always emits a full URL rather than a bare id: the XHS parser accepts a bare
|
||||
24-hex id only, so the URL form is the safer universal input. The
|
||||
``xsec_token`` is appended when present but is deliberately optional -- it
|
||||
expires, and the id alone is what keeps a long-running task alive.
|
||||
"""
|
||||
path = _CREATOR_PATH if kind == MODE_CREATOR else _NOTE_PATH
|
||||
return f"{_XHS_WEB_BASE}{path}/{value}"
|
||||
|
||||
|
||||
def build_target_urls(mode: str, targets: Iterable[MonitorTarget]) -> List[str]:
|
||||
urls = []
|
||||
for target in targets:
|
||||
url = build_target_url(target.external_id, target.kind)
|
||||
if target.xsec_token:
|
||||
url = f"{url}?xsec_token={target.xsec_token}"
|
||||
if target.xsec_source:
|
||||
url = f"{url}&xsec_source={target.xsec_source}"
|
||||
urls.append(url)
|
||||
return urls
|
||||
|
||||
|
||||
async def _strategy_settings(session, platform: str) -> dict:
|
||||
"""Crawl-strategy and proxy settings for one platform.
|
||||
|
||||
Per-platform because the values genuinely differ: what is a safe request
|
||||
interval on one site is a rate limit on another. Read per run rather than
|
||||
cached, so a change takes effect on the next scheduled run.
|
||||
"""
|
||||
return {
|
||||
"enable_sub_comments": bool(
|
||||
await app_settings.get_value(session, "enable_sub_comments", platform, False)
|
||||
),
|
||||
"crawl_sleep_sec": int(
|
||||
await app_settings.get_value(session, "crawl_sleep_sec", platform, 2)
|
||||
),
|
||||
"enable_ip_proxy": bool(
|
||||
await app_settings.get_value(session, "enable_ip_proxy", platform, False)
|
||||
),
|
||||
"proxy_provider": await app_settings.get_value(
|
||||
session, "proxy_provider", platform, "kuaidaili"
|
||||
),
|
||||
"proxy_pool_count": int(
|
||||
await app_settings.get_value(session, "proxy_pool_count", platform, 2)
|
||||
),
|
||||
"static_proxy_url": await app_settings.get_value(
|
||||
session, "static_proxy_url", platform, ""
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def _write_cookie_file(path: Path, cookie: str) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(cookie, encoding="utf-8")
|
||||
|
||||
|
||||
def _remove_cookie_file(path: Path) -> None:
|
||||
"""Best-effort removal; the cookie is a credential, do not leave it around."""
|
||||
try:
|
||||
os.remove(path)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
async def execute_task(task_id: int, trigger: str = "manual") -> IngestResult:
|
||||
"""Run one monitoring cycle for ``task_id`` and ingest its output.
|
||||
|
||||
Split into three phases with separate short-lived DB sessions so no
|
||||
transaction is held open across the multi-minute subprocess run.
|
||||
"""
|
||||
# --- Phase 1: book the run and work out where its output goes -------------
|
||||
async with get_session() as session:
|
||||
task = await session.get(MonitorTask, task_id)
|
||||
if task is None:
|
||||
raise ValueError(f"Monitor task {task_id} not found")
|
||||
|
||||
targets = [target for target in task.targets if target.enabled]
|
||||
if not targets:
|
||||
raise ValueError(f"Monitor task {task_id} has no enabled targets")
|
||||
|
||||
platform = task.platform
|
||||
urls = build_target_urls(task.mode, targets)
|
||||
cookie = await get_cookie(session, platform)
|
||||
strategy = await _strategy_settings(session, platform)
|
||||
|
||||
run = MonitorRun(
|
||||
task_id=task.id,
|
||||
trigger=trigger,
|
||||
status=RUN_PENDING,
|
||||
phase=task.mode,
|
||||
save_data_path="",
|
||||
queued_at=get_current_timestamp(),
|
||||
not_before=0,
|
||||
max_comments_count=task.max_comments_count if task.enable_comments else 0,
|
||||
)
|
||||
session.add(run)
|
||||
await session.flush()
|
||||
|
||||
run_id = run.id
|
||||
out_dir = MONITOR_RUNS_DIR / str(task.id) / str(run_id)
|
||||
run.save_data_path = str(out_dir)
|
||||
|
||||
# Snapshot the values the subprocess needs; `task` is detached after commit.
|
||||
mode = task.mode
|
||||
enable_comments = task.enable_comments
|
||||
max_notes_count = task.max_notes_count
|
||||
max_comments_count = task.max_comments_count
|
||||
timeout_seconds = task.run_timeout_seconds
|
||||
|
||||
# --- Phase 2: run the crawler outside any transaction ---------------------
|
||||
cookie_file = out_dir / ".cookies"
|
||||
_write_cookie_file(cookie_file, cookie)
|
||||
|
||||
request = CrawlerStartRequest(
|
||||
platform=_PLATFORM_ENUM[platform],
|
||||
login_type=LoginTypeEnum.COOKIE,
|
||||
crawler_type=CrawlerTypeEnum.CREATOR if mode == MODE_CREATOR else CrawlerTypeEnum.DETAIL,
|
||||
creator_ids=",".join(urls) if mode == MODE_CREATOR else "",
|
||||
specified_ids=",".join(urls) if mode != MODE_CREATOR else "",
|
||||
start_page=1,
|
||||
enable_comments=enable_comments,
|
||||
enable_sub_comments=strategy["enable_sub_comments"],
|
||||
enable_media=False,
|
||||
save_option=SaveDataOptionEnum.JSONL,
|
||||
cookies="",
|
||||
headless=True,
|
||||
max_notes_count=max_notes_count,
|
||||
max_comments_count=max_comments_count,
|
||||
# Isolate this run's output: the crawler names files by date only, so
|
||||
# otherwise same-day runs would append into one shared file.
|
||||
save_data_path=str(out_dir),
|
||||
# Unattended runs must not try to attach to the user's desktop Chrome.
|
||||
enable_cdp_mode=False,
|
||||
# Only injecting web_session is not enough to sign requests from a cold
|
||||
# browser profile.
|
||||
inject_all_cookies=True,
|
||||
save_login_state=True,
|
||||
cookies_file=str(cookie_file),
|
||||
max_concurrency_num=1,
|
||||
# Strategy + proxy, surfaced on the Settings page.
|
||||
crawler_max_sleep_sec=strategy["crawl_sleep_sec"],
|
||||
enable_ip_proxy=strategy["enable_ip_proxy"],
|
||||
ip_proxy_pool_count=strategy["proxy_pool_count"],
|
||||
ip_proxy_provider_name=strategy["proxy_provider"],
|
||||
static_proxy_url=strategy["static_proxy_url"] or None,
|
||||
)
|
||||
|
||||
async with get_session() as session:
|
||||
run = await session.get(MonitorRun, run_id)
|
||||
if run is not None:
|
||||
run.status = RUN_RUNNING
|
||||
run.started_at = get_current_timestamp()
|
||||
|
||||
try:
|
||||
exit_code = await crawler_manager.run_and_wait(request, timeout=timeout_seconds)
|
||||
finally:
|
||||
_remove_cookie_file(cookie_file)
|
||||
|
||||
# --- Phase 3: ingest ------------------------------------------------------
|
||||
async with get_session() as session:
|
||||
run = await session.get(MonitorRun, run_id)
|
||||
task = await session.get(MonitorTask, task_id)
|
||||
if run is None or task is None:
|
||||
raise ValueError(f"Run {run_id} or task {task_id} vanished during execution")
|
||||
|
||||
if exit_code == -1 and not (out_dir / "xhs").exists():
|
||||
# run_and_wait returns -1 when the process could not start or timed out.
|
||||
run.status = RUN_TIMEOUT
|
||||
run.finished_at = get_current_timestamp()
|
||||
run.exit_code = exit_code
|
||||
run.error_message = "Run was killed by timeout or failed to start"
|
||||
result = IngestResult(status=RUN_TIMEOUT, error=run.error_message)
|
||||
else:
|
||||
run.exit_code = exit_code
|
||||
run.finished_at = get_current_timestamp()
|
||||
result = await ingest_run(session, run, task, out_dir)
|
||||
|
||||
# A run that authenticated fine is the only useful signal that the
|
||||
# stored cookie still works.
|
||||
if result.notes_fetched > 0:
|
||||
await mark_cookie_ok(session, task.platform)
|
||||
|
||||
task.last_run_at = run.finished_at
|
||||
task.last_status = result.status
|
||||
task.last_error = result.error
|
||||
|
||||
# --- Phase 4: notify ------------------------------------------------------
|
||||
# Runs after the ingest transaction has committed, in its own session. A push
|
||||
# failure must never roll back collected data, and notify_run() swallows its
|
||||
# own errors for the same reason.
|
||||
async with get_session() as session:
|
||||
task = await session.get(MonitorTask, task_id)
|
||||
run = await session.get(MonitorRun, run_id)
|
||||
if task is not None and run is not None:
|
||||
await notify.notify_run(session, task, run)
|
||||
|
||||
return result
|
||||
@@ -0,0 +1,185 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/scheduler.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Background scheduler for monitor tasks.
|
||||
|
||||
One asyncio loop polls for due tasks and hands them to the runner. A plain loop
|
||||
is enough here: there is exactly one process, one global crawler subprocess, and
|
||||
therefore no concurrency to coordinate -- a cron-style library would add a
|
||||
dependency without adding a capability.
|
||||
|
||||
Scheduling is **fixed-delay**, not fixed-rate: ``next_run_at`` is set from the
|
||||
moment a run starts, so a slow run cannot make its task fire back-to-back.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import random
|
||||
from datetime import datetime
|
||||
from typing import Optional
|
||||
|
||||
from sqlalchemy import select
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from ..services import crawler_manager
|
||||
from . import app_settings
|
||||
from .db import get_session
|
||||
from .models import MonitorRun, MonitorTask, RUN_INTERRUPTED, RUN_RUNNING
|
||||
from .runner import execute_task
|
||||
from .settings import get_cookie
|
||||
|
||||
POLL_INTERVAL_SECONDS = 20
|
||||
# Spread tasks sharing an interval so they do not all come due on the same tick.
|
||||
JITTER_SECONDS = 60
|
||||
|
||||
_MS_PER_MINUTE = 60_000
|
||||
|
||||
|
||||
class MonitorScheduler:
|
||||
"""Polls the task table and runs whatever is due."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._loop_task: Optional[asyncio.Task] = None
|
||||
self._stopping = asyncio.Event()
|
||||
# Avoids logging "no cookie" on every single tick.
|
||||
self._warned_no_cookie = False
|
||||
|
||||
async def start(self) -> None:
|
||||
if self._loop_task is not None and not self._loop_task.done():
|
||||
return
|
||||
self._stopping.clear()
|
||||
self._loop_task = asyncio.create_task(self._run_loop())
|
||||
|
||||
async def stop(self) -> None:
|
||||
self._stopping.set()
|
||||
if self._loop_task is not None:
|
||||
self._loop_task.cancel()
|
||||
try:
|
||||
await self._loop_task
|
||||
except asyncio.CancelledError:
|
||||
pass
|
||||
self._loop_task = None
|
||||
|
||||
async def _run_loop(self) -> None:
|
||||
try:
|
||||
await self.recover()
|
||||
except Exception as exc: # pragma: no cover - defensive
|
||||
print(f"[monitor.scheduler] recovery failed: {exc}")
|
||||
|
||||
while not self._stopping.is_set():
|
||||
try:
|
||||
await self.tick()
|
||||
except Exception as exc: # pragma: no cover - keep the loop alive
|
||||
print(f"[monitor.scheduler] tick failed: {exc}")
|
||||
await asyncio.sleep(POLL_INTERVAL_SECONDS)
|
||||
|
||||
async def recover(self) -> None:
|
||||
"""Clean up state left behind by a server restart.
|
||||
|
||||
A run still marked ``running`` cannot be running -- its subprocess died
|
||||
with the previous process. Marking it interrupted stops it from blocking
|
||||
the UI as a phantom in-flight run.
|
||||
"""
|
||||
async with get_session() as session:
|
||||
stale = (
|
||||
await session.scalars(
|
||||
select(MonitorRun).where(MonitorRun.status == RUN_RUNNING)
|
||||
)
|
||||
).all()
|
||||
for run in stale:
|
||||
run.status = RUN_INTERRUPTED
|
||||
run.finished_at = get_current_timestamp()
|
||||
if stale:
|
||||
print(
|
||||
f"[monitor.scheduler] marked {len(stale)} interrupted run(s) "
|
||||
f"left over from a previous process"
|
||||
)
|
||||
|
||||
async def tick(self) -> None:
|
||||
"""Run one due task, if the crawler is free and we are in the active window."""
|
||||
# The crawler subprocess is a global singleton, so a manual crawl and a
|
||||
# monitor run cannot overlap. Returning without advancing next_run_at
|
||||
# leaves the task due, and it is picked up on a later tick.
|
||||
if crawler_manager.is_busy():
|
||||
return
|
||||
|
||||
async with get_session() as session:
|
||||
if not await self._within_active_hours(session):
|
||||
# Deliberately does not advance next_run_at: the task simply runs
|
||||
# when the window next opens, rather than being skipped for a day.
|
||||
return
|
||||
|
||||
await self._run_due_task()
|
||||
|
||||
async def _within_active_hours(self, session) -> bool:
|
||||
"""Whether scheduled runs are allowed right now (local time)."""
|
||||
start, end = await app_settings.active_hours(session)
|
||||
hour = datetime.now().hour
|
||||
if start <= end:
|
||||
return start <= hour <= end
|
||||
# Window wraps past midnight, e.g. 22 -> 6.
|
||||
return hour >= start or hour <= end
|
||||
|
||||
async def _run_due_task(self) -> None:
|
||||
async with get_session() as session:
|
||||
task = await session.scalar(
|
||||
select(MonitorTask)
|
||||
.where(
|
||||
MonitorTask.enabled.is_(True),
|
||||
MonitorTask.next_run_at.is_not(None),
|
||||
MonitorTask.next_run_at <= get_current_timestamp(),
|
||||
)
|
||||
.order_by(MonitorTask.next_run_at)
|
||||
.limit(1)
|
||||
)
|
||||
|
||||
if task is None:
|
||||
return
|
||||
|
||||
# No cookie means every run would report an auth failure. Leave the
|
||||
# task due rather than advancing: it starts working the moment the
|
||||
# user pastes one.
|
||||
cookie = await get_cookie(session)
|
||||
if not cookie:
|
||||
if not self._warned_no_cookie:
|
||||
print(
|
||||
"[monitor.scheduler] no XHS cookie configured; "
|
||||
"scheduled tasks will not run until one is set"
|
||||
)
|
||||
self._warned_no_cookie = True
|
||||
return
|
||||
self._warned_no_cookie = False
|
||||
|
||||
# Advance before running so a crash mid-run cannot cause an immediate
|
||||
# re-fire, and so a long outage coalesces into a single run instead
|
||||
# of one run per missed interval.
|
||||
task.next_run_at = (
|
||||
get_current_timestamp()
|
||||
+ task.interval_minutes * _MS_PER_MINUTE
|
||||
+ random.randint(0, JITTER_SECONDS) * 1000
|
||||
)
|
||||
task_id = task.id
|
||||
|
||||
try:
|
||||
await execute_task(task_id, trigger="scheduled")
|
||||
except Exception as exc:
|
||||
print(f"[monitor.scheduler] task {task_id} failed: {exc}")
|
||||
|
||||
|
||||
# Global singleton, mirroring the crawler_manager pattern.
|
||||
monitor_scheduler = MonitorScheduler()
|
||||
@@ -0,0 +1,720 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/service.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Non-commercial learning license 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Task CRUD and dashboard queries for the monitoring layer."""
|
||||
|
||||
import asyncio
|
||||
import re
|
||||
from typing import Any, Dict, List, Optional
|
||||
from urllib.parse import parse_qs, urlparse
|
||||
|
||||
from sqlalchemy import delete, func, select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from . import app_settings, platforms
|
||||
from .db import get_session
|
||||
from .platforms import PLATFORM_XHS
|
||||
from .models import (
|
||||
MODE_CREATOR,
|
||||
MODE_NOTE,
|
||||
MonitorComment,
|
||||
MonitorEvent,
|
||||
MonitorNote,
|
||||
MonitorNoteMetric,
|
||||
MonitorRun,
|
||||
MonitorTarget,
|
||||
MonitorTask,
|
||||
RUN_SUCCESS,
|
||||
RUN_PARTIAL,
|
||||
)
|
||||
from .runner import execute_task
|
||||
|
||||
# Keep strong references to in-flight manual runs; asyncio only holds weak ones,
|
||||
# so without this a run can be garbage collected mid-flight.
|
||||
_background_runs: set[asyncio.Task] = set()
|
||||
|
||||
MIN_INTERVAL_MINUTES = 30
|
||||
MAX_INTERVAL_MINUTES = 7 * 24 * 60
|
||||
|
||||
_CREATOR_URL_RE = re.compile(r"xiaohongshu\.com/user/profile/([A-Za-z0-9_-]+)")
|
||||
_NOTE_URL_RE = re.compile(r"xiaohongshu\.com/(?:explore|discovery/item)/([A-Za-z0-9_-]+)")
|
||||
# XHS user ids and note ids are 24-char hex; allow a slightly wider range so a
|
||||
# format change degrades into "still accepted" rather than "rejected".
|
||||
_BARE_ID_RE = re.compile(r"^[A-Za-z0-9_-]{8,64}$")
|
||||
|
||||
|
||||
class TargetParseError(ValueError):
|
||||
"""Raised when a pasted monitoring target cannot be understood."""
|
||||
|
||||
|
||||
def parse_target_input(
|
||||
value: str, mode: str, platform: str = PLATFORM_XHS
|
||||
) -> Dict[str, str]:
|
||||
"""Parse a pasted creator/note value into a stable id plus a refreshable token.
|
||||
|
||||
Accepts either a full URL (with or without ``xsec_token``) or a bare id.
|
||||
Storing the id separately from the token is what keeps a long-running task
|
||||
alive: tokens expire, ids do not.
|
||||
|
||||
URL shapes are platform-specific. Only Xiaohongshu is wired, so anything else
|
||||
is rejected here as well as at task creation -- parsing a Douyin link as if it
|
||||
were a Xiaohongshu one would be worse than refusing it.
|
||||
"""
|
||||
if platform != PLATFORM_XHS:
|
||||
raise TargetParseError(f"暂不支持解析该平台({platform})的目标链接")
|
||||
|
||||
raw = (value or "").strip()
|
||||
if not raw:
|
||||
raise TargetParseError("Empty target")
|
||||
|
||||
external_id = ""
|
||||
if raw.startswith("http") or "/" in raw:
|
||||
# xhslink.com and other short links are not resolvable without a network
|
||||
# round-trip, so only the direct profile/explore forms are supported.
|
||||
match = _CREATOR_URL_RE.search(raw) if mode == MODE_CREATOR else _NOTE_URL_RE.search(raw)
|
||||
if not match:
|
||||
expected = "博主主页" if mode == MODE_CREATOR else "笔记"
|
||||
raise TargetParseError(f"无法从链接中解析出{expected} ID:{raw}")
|
||||
external_id = match.group(1)
|
||||
elif _BARE_ID_RE.match(raw):
|
||||
external_id = raw
|
||||
else:
|
||||
raise TargetParseError(f"无法识别的目标:{raw}")
|
||||
|
||||
params = parse_qs(urlparse(raw).query) if raw.startswith("http") else {}
|
||||
return {
|
||||
"external_id": external_id,
|
||||
"xsec_token": (params.get("xsec_token") or [""])[0],
|
||||
"xsec_source": (params.get("xsec_source") or [""])[0],
|
||||
"raw_value": raw,
|
||||
}
|
||||
|
||||
|
||||
async def platform_task_ids(session: AsyncSession, platform: str) -> List[int]:
|
||||
"""Ids of the tasks belonging to a platform.
|
||||
|
||||
Note/comment/event tables carry no platform column -- they hang off a task --
|
||||
so scoping a query to a platform means scoping it to that task set.
|
||||
"""
|
||||
return list(
|
||||
await session.scalars(select(MonitorTask.id).where(MonitorTask.platform == platform))
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Task CRUD
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
async def create_task(session: AsyncSession, payload: Dict[str, Any]) -> MonitorTask:
|
||||
mode = payload["mode"]
|
||||
if mode not in (MODE_CREATOR, MODE_NOTE):
|
||||
raise ValueError(f"Unsupported mode: {mode}")
|
||||
|
||||
# Rejects unknown platforms and, more importantly, platforms whose crawler
|
||||
# exists upstream but whose monitoring is not wired up -- accepting those
|
||||
# would create a task that can never produce data.
|
||||
platform = payload.get("platform") or PLATFORM_XHS
|
||||
platforms.ensure_runnable(platform)
|
||||
|
||||
now = get_current_timestamp()
|
||||
|
||||
# Fall back to the configured defaults for anything the caller left out, so
|
||||
# the Settings page actually governs new tasks.
|
||||
defaults = await app_settings.defaults(session, platform)
|
||||
interval_minutes = payload.get("interval_minutes") or defaults["interval_minutes"]
|
||||
interval_ms = int(interval_minutes) * 60_000
|
||||
|
||||
task = MonitorTask(
|
||||
name=payload["name"],
|
||||
platform=platform,
|
||||
mode=mode,
|
||||
enabled=payload.get("enabled", True),
|
||||
interval_minutes=interval_minutes,
|
||||
max_notes_count=payload.get("max_notes_count") or defaults["max_notes_count"],
|
||||
enable_comments=payload.get("enable_comments", True),
|
||||
max_comments_count=payload.get("max_comments_count") or defaults["max_comments_count"],
|
||||
run_timeout_seconds=payload.get("run_timeout_seconds", 3600),
|
||||
notify_enabled=payload.get("notify_enabled", False),
|
||||
next_run_at=now + interval_ms,
|
||||
last_status="idle",
|
||||
created_at=now,
|
||||
updated_at=now,
|
||||
)
|
||||
session.add(task)
|
||||
await session.flush()
|
||||
|
||||
seen: set[str] = set()
|
||||
for value in payload.get("targets", []):
|
||||
parsed = parse_target_input(value, mode, platform)
|
||||
if parsed["external_id"] in seen:
|
||||
continue
|
||||
seen.add(parsed["external_id"])
|
||||
session.add(
|
||||
MonitorTarget(
|
||||
task_id=task.id,
|
||||
kind=mode,
|
||||
external_id=parsed["external_id"],
|
||||
xsec_token=parsed["xsec_token"],
|
||||
xsec_source=parsed["xsec_source"],
|
||||
raw_value=parsed["raw_value"],
|
||||
label=parsed["external_id"],
|
||||
enabled=True,
|
||||
created_at=now,
|
||||
)
|
||||
)
|
||||
|
||||
await session.flush()
|
||||
return task
|
||||
|
||||
|
||||
async def update_task(session: AsyncSession, task_id: int, payload: Dict[str, Any]) -> MonitorTask:
|
||||
task = await session.get(MonitorTask, task_id)
|
||||
if task is None:
|
||||
raise ValueError(f"Task {task_id} not found")
|
||||
|
||||
for field in (
|
||||
"name",
|
||||
"enabled",
|
||||
"interval_minutes",
|
||||
"max_notes_count",
|
||||
"enable_comments",
|
||||
"max_comments_count",
|
||||
"run_timeout_seconds",
|
||||
"notify_enabled",
|
||||
):
|
||||
if field in payload and payload[field] is not None:
|
||||
setattr(task, field, payload[field])
|
||||
|
||||
# Replacing targets resets the baseline implicitly: a note set that now
|
||||
# includes new ids will simply report them as new on the next run.
|
||||
if payload.get("targets") is not None:
|
||||
await session.execute(delete(MonitorTarget).where(MonitorTarget.task_id == task_id))
|
||||
now = get_current_timestamp()
|
||||
seen: set[str] = set()
|
||||
for value in payload["targets"]:
|
||||
parsed = parse_target_input(value, task.mode)
|
||||
if parsed["external_id"] in seen:
|
||||
continue
|
||||
seen.add(parsed["external_id"])
|
||||
session.add(
|
||||
MonitorTarget(
|
||||
task_id=task_id,
|
||||
kind=task.mode,
|
||||
external_id=parsed["external_id"],
|
||||
xsec_token=parsed["xsec_token"],
|
||||
xsec_source=parsed["xsec_source"],
|
||||
raw_value=parsed["raw_value"],
|
||||
label=parsed["external_id"],
|
||||
enabled=True,
|
||||
created_at=now,
|
||||
)
|
||||
)
|
||||
|
||||
if "interval_minutes" in payload and payload["interval_minutes"]:
|
||||
task.next_run_at = get_current_timestamp() + payload["interval_minutes"] * 60_000
|
||||
|
||||
task.updated_at = get_current_timestamp()
|
||||
await session.flush()
|
||||
return task
|
||||
|
||||
|
||||
async def delete_task(session: AsyncSession, task_id: int) -> None:
|
||||
task = await session.get(MonitorTask, task_id)
|
||||
if task is None:
|
||||
raise ValueError(f"Task {task_id} not found")
|
||||
await session.delete(task)
|
||||
|
||||
|
||||
def trigger_manual_run(task_id: int) -> None:
|
||||
"""Fire a run in the background and return immediately.
|
||||
|
||||
A crawl takes minutes, so the HTTP request must not wait for it. The UI
|
||||
follows progress through the logs WebSocket and the run history.
|
||||
"""
|
||||
task = asyncio.create_task(execute_task(task_id, trigger="manual"))
|
||||
_background_runs.add(task)
|
||||
task.add_done_callback(_background_runs.discard)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Dashboard queries
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
async def _latest_successful_run_id(session: AsyncSession, task_id: int) -> Optional[int]:
|
||||
return await session.scalar(
|
||||
select(MonitorRun.id)
|
||||
.where(
|
||||
MonitorRun.task_id == task_id,
|
||||
MonitorRun.status.in_((RUN_SUCCESS, RUN_PARTIAL)),
|
||||
)
|
||||
.order_by(MonitorRun.id.desc())
|
||||
.limit(1)
|
||||
)
|
||||
|
||||
|
||||
def _delta(current: Optional[int], previous: Optional[int]) -> Optional[int]:
|
||||
if current is None or previous is None:
|
||||
return None
|
||||
return current - previous
|
||||
|
||||
|
||||
async def list_notes(
|
||||
session: AsyncSession,
|
||||
task_id: Optional[int] = None,
|
||||
only_new: bool = False,
|
||||
limit: int = 200,
|
||||
platform: Optional[str] = None,
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Tracked notes with their latest metrics and change vs the previous run."""
|
||||
query = select(MonitorNote).order_by(MonitorNote.last_seen_at.desc()).limit(limit)
|
||||
if task_id is not None:
|
||||
query = query.where(MonitorNote.task_id == task_id)
|
||||
if platform is not None:
|
||||
scoped = await platform_task_ids(session, platform)
|
||||
if not scoped:
|
||||
return []
|
||||
query = query.where(MonitorNote.task_id.in_(scoped))
|
||||
|
||||
notes = list((await session.scalars(query)).all())
|
||||
if not notes:
|
||||
return []
|
||||
|
||||
note_ids = [note.note_id for note in notes]
|
||||
|
||||
# Fetch every snapshot for these notes in one go and gather the two most
|
||||
# recent per note, rather than issuing two queries per note.
|
||||
snapshots = list(
|
||||
(
|
||||
await session.scalars(
|
||||
select(MonitorNoteMetric)
|
||||
.where(MonitorNoteMetric.note_id.in_(note_ids))
|
||||
.order_by(MonitorNoteMetric.note_id, MonitorNoteMetric.run_id.desc())
|
||||
)
|
||||
).all()
|
||||
)
|
||||
by_note: Dict[str, List[MonitorNoteMetric]] = {}
|
||||
for snapshot in snapshots:
|
||||
by_note.setdefault(snapshot.note_id, []).append(snapshot)
|
||||
|
||||
latest_run_ids: Dict[int, Optional[int]] = {}
|
||||
result: List[Dict[str, Any]] = []
|
||||
|
||||
for note in notes:
|
||||
series = by_note.get(note.note_id, [])
|
||||
current = series[0] if series else None
|
||||
previous = series[1] if len(series) > 1 else None
|
||||
|
||||
if only_new:
|
||||
if note.task_id not in latest_run_ids:
|
||||
latest_run_ids[note.task_id] = await _latest_successful_run_id(session, note.task_id)
|
||||
if note.first_seen_run_id != latest_run_ids[note.task_id]:
|
||||
continue
|
||||
|
||||
result.append(
|
||||
{
|
||||
"task_id": note.task_id,
|
||||
"note_id": note.note_id,
|
||||
"title": note.title,
|
||||
"note_url": note.note_url,
|
||||
"cover": note.cover,
|
||||
"first_seen_at": note.first_seen_at,
|
||||
"last_seen_at": note.last_seen_at,
|
||||
"is_new": note.first_seen_run_id == latest_run_ids.get(note.task_id),
|
||||
"metrics": {
|
||||
"liked_count": current.liked_count if current else None,
|
||||
"comment_count": current.comment_count if current else None,
|
||||
"collected_count": current.collected_count if current else None,
|
||||
"share_count": current.share_count if current else None,
|
||||
},
|
||||
"deltas": {
|
||||
"liked_count": _delta(
|
||||
current.liked_count if current else None,
|
||||
previous.liked_count if previous else None,
|
||||
),
|
||||
"comment_count": _delta(
|
||||
current.comment_count if current else None,
|
||||
previous.comment_count if previous else None,
|
||||
),
|
||||
"collected_count": _delta(
|
||||
current.collected_count if current else None,
|
||||
previous.collected_count if previous else None,
|
||||
),
|
||||
"share_count": _delta(
|
||||
current.share_count if current else None,
|
||||
previous.share_count if previous else None,
|
||||
),
|
||||
},
|
||||
"snapshot_count": len(series),
|
||||
}
|
||||
)
|
||||
|
||||
return result
|
||||
|
||||
|
||||
async def note_series(session: AsyncSession, note_id: str, task_id: Optional[int] = None) -> List[Dict[str, Any]]:
|
||||
"""Metric time series for one note."""
|
||||
query = (
|
||||
select(MonitorNoteMetric)
|
||||
.where(MonitorNoteMetric.note_id == note_id)
|
||||
.order_by(MonitorNoteMetric.run_id)
|
||||
)
|
||||
if task_id is not None:
|
||||
query = query.where(MonitorNoteMetric.task_id == task_id)
|
||||
|
||||
return [
|
||||
{
|
||||
"run_id": row.run_id,
|
||||
"captured_at": row.captured_at,
|
||||
"liked_count": row.liked_count,
|
||||
"comment_count": row.comment_count,
|
||||
"collected_count": row.collected_count,
|
||||
"share_count": row.share_count,
|
||||
}
|
||||
for row in (await session.scalars(query)).all()
|
||||
]
|
||||
|
||||
|
||||
async def _note_meta_map(
|
||||
session: AsyncSession, note_ids: List[str]
|
||||
) -> Dict[str, Dict[str, Any]]:
|
||||
"""Look up note title/cover/url for a set of note ids.
|
||||
|
||||
Fetched as one query and joined in Python rather than as a SQL join: the
|
||||
comment table has no foreign key to the note table (both are keyed by the
|
||||
platform's note id, per task), and a single IN() is easier to follow here.
|
||||
"""
|
||||
if not note_ids:
|
||||
return {}
|
||||
|
||||
rows = (
|
||||
await session.scalars(select(MonitorNote).where(MonitorNote.note_id.in_(set(note_ids))))
|
||||
).all()
|
||||
return {
|
||||
row.note_id: {
|
||||
"note_title": row.title,
|
||||
"note_cover": row.cover,
|
||||
"note_url": row.note_url,
|
||||
"task_id": row.task_id,
|
||||
}
|
||||
for row in rows
|
||||
}
|
||||
|
||||
|
||||
async def list_comments(
|
||||
session: AsyncSession,
|
||||
task_id: Optional[int] = None,
|
||||
note_id: Optional[str] = None,
|
||||
limit: int = 200,
|
||||
platform: Optional[str] = None,
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Comments, each carrying the note it belongs to.
|
||||
|
||||
The note association is the point: without it a comment stream is unreadable,
|
||||
since a bare note_id tells the operator nothing.
|
||||
"""
|
||||
query = select(MonitorComment).order_by(MonitorComment.first_seen_at.desc()).limit(limit)
|
||||
if task_id is not None:
|
||||
query = query.where(MonitorComment.task_id == task_id)
|
||||
if note_id is not None:
|
||||
query = query.where(MonitorComment.note_id == note_id)
|
||||
if platform is not None:
|
||||
scoped = await platform_task_ids(session, platform)
|
||||
if not scoped:
|
||||
return []
|
||||
query = query.where(MonitorComment.task_id.in_(scoped))
|
||||
|
||||
comments = list((await session.scalars(query)).all())
|
||||
meta = await _note_meta_map(session, [row.note_id for row in comments])
|
||||
|
||||
return [
|
||||
{
|
||||
"task_id": row.task_id,
|
||||
"note_id": row.note_id,
|
||||
"comment_id": row.comment_id,
|
||||
"content": row.content,
|
||||
"nickname": row.nickname,
|
||||
"create_time": row.create_time,
|
||||
"like_count": row.like_count,
|
||||
"sub_comment_count": row.sub_comment_count,
|
||||
"first_seen_at": row.first_seen_at,
|
||||
"note_title": meta.get(row.note_id, {}).get("note_title", ""),
|
||||
"note_cover": meta.get(row.note_id, {}).get("note_cover", ""),
|
||||
"note_url": meta.get(row.note_id, {}).get("note_url", ""),
|
||||
}
|
||||
for row in comments
|
||||
]
|
||||
|
||||
|
||||
async def comment_note_groups(
|
||||
session: AsyncSession, task_id: Optional[int] = None, platform: Optional[str] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Notes that have comments, newest first, with their comment counts.
|
||||
|
||||
Feeds the comment filter dropdown: the operator picks a work by title, so
|
||||
the counts need to be visible before choosing.
|
||||
"""
|
||||
scoped_ids: Optional[List[int]] = None
|
||||
if platform is not None:
|
||||
scoped_ids = await platform_task_ids(session, platform)
|
||||
if not scoped_ids:
|
||||
return []
|
||||
|
||||
count_query = select(
|
||||
MonitorComment.note_id, func.count().label("comment_count")
|
||||
).group_by(MonitorComment.note_id)
|
||||
if task_id is not None:
|
||||
count_query = count_query.where(MonitorComment.task_id == task_id)
|
||||
if scoped_ids is not None:
|
||||
count_query = count_query.where(MonitorComment.task_id.in_(scoped_ids))
|
||||
|
||||
counts = {row.note_id: row.comment_count for row in (await session.execute(count_query)).all()}
|
||||
if not counts:
|
||||
return []
|
||||
|
||||
latest_query = (
|
||||
select(MonitorComment.note_id, func.max(MonitorComment.first_seen_at).label("latest"))
|
||||
.where(MonitorComment.note_id.in_(set(counts)))
|
||||
.group_by(MonitorComment.note_id)
|
||||
)
|
||||
if task_id is not None:
|
||||
latest_query = latest_query.where(MonitorComment.task_id == task_id)
|
||||
if scoped_ids is not None:
|
||||
latest_query = latest_query.where(MonitorComment.task_id.in_(scoped_ids))
|
||||
latest = {row.note_id: row.latest for row in (await session.execute(latest_query)).all()}
|
||||
|
||||
meta = await _note_meta_map(session, list(counts))
|
||||
|
||||
groups = [
|
||||
{
|
||||
"note_id": note_id,
|
||||
"note_title": meta.get(note_id, {}).get("note_title", ""),
|
||||
"note_cover": meta.get(note_id, {}).get("note_cover", ""),
|
||||
"note_url": meta.get(note_id, {}).get("note_url", ""),
|
||||
"comment_count": count,
|
||||
"latest_at": latest.get(note_id, 0),
|
||||
}
|
||||
for note_id, count in counts.items()
|
||||
]
|
||||
groups.sort(key=lambda group: group["latest_at"], reverse=True)
|
||||
return groups
|
||||
|
||||
|
||||
async def list_events(
|
||||
session: AsyncSession,
|
||||
task_id: Optional[int] = None,
|
||||
event_type: Optional[str] = None,
|
||||
since_id: Optional[int] = None,
|
||||
limit: int = 200,
|
||||
platform: Optional[str] = None,
|
||||
) -> List[Dict[str, Any]]:
|
||||
query = select(MonitorEvent).order_by(MonitorEvent.id.desc()).limit(limit)
|
||||
if task_id is not None:
|
||||
query = query.where(MonitorEvent.task_id == task_id)
|
||||
if event_type is not None:
|
||||
query = query.where(MonitorEvent.type == event_type)
|
||||
if since_id is not None:
|
||||
query = query.where(MonitorEvent.id > since_id)
|
||||
if platform is not None:
|
||||
scoped = await platform_task_ids(session, platform)
|
||||
if not scoped:
|
||||
return []
|
||||
query = query.where(MonitorEvent.task_id.in_(scoped))
|
||||
|
||||
return [
|
||||
{
|
||||
"id": row.id,
|
||||
"task_id": row.task_id,
|
||||
"run_id": row.run_id,
|
||||
"type": row.type,
|
||||
"severity": row.severity,
|
||||
"target_kind": row.target_kind,
|
||||
"target_id": row.target_id,
|
||||
"title": row.title,
|
||||
"created_at": row.created_at,
|
||||
"is_read": row.is_read,
|
||||
}
|
||||
for row in (await session.scalars(query)).all()
|
||||
]
|
||||
|
||||
|
||||
async def list_runs(session: AsyncSession, task_id: int, limit: int = 50) -> List[Dict[str, Any]]:
|
||||
rows = (
|
||||
await session.scalars(
|
||||
select(MonitorRun)
|
||||
.where(MonitorRun.task_id == task_id)
|
||||
.order_by(MonitorRun.id.desc())
|
||||
.limit(limit)
|
||||
)
|
||||
).all()
|
||||
|
||||
return [
|
||||
{
|
||||
"id": row.id,
|
||||
"task_id": row.task_id,
|
||||
"status": row.status,
|
||||
"trigger": row.trigger,
|
||||
"queued_at": row.queued_at,
|
||||
"started_at": row.started_at,
|
||||
"finished_at": row.finished_at,
|
||||
"exit_code": row.exit_code,
|
||||
"notes_fetched": row.notes_fetched,
|
||||
"comments_fetched": row.comments_fetched,
|
||||
"new_notes": row.new_notes,
|
||||
"new_comments": row.new_comments,
|
||||
"is_baseline": row.is_baseline,
|
||||
"max_comments_count": row.max_comments_count,
|
||||
"error_message": row.error_message,
|
||||
}
|
||||
for row in rows
|
||||
]
|
||||
|
||||
|
||||
async def list_tasks(
|
||||
session: AsyncSession, platform: Optional[str] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
query = select(MonitorTask).order_by(MonitorTask.id)
|
||||
if platform is not None:
|
||||
query = query.where(MonitorTask.platform == platform)
|
||||
|
||||
tasks = list((await session.scalars(query)).all())
|
||||
if not tasks:
|
||||
return []
|
||||
|
||||
counts = dict(
|
||||
(
|
||||
await session.execute(
|
||||
select(MonitorTarget.task_id, func.count())
|
||||
.group_by(MonitorTarget.task_id)
|
||||
)
|
||||
).all()
|
||||
)
|
||||
unread = dict(
|
||||
(
|
||||
await session.execute(
|
||||
select(MonitorEvent.task_id, func.count())
|
||||
.where(MonitorEvent.is_read.is_(False))
|
||||
.group_by(MonitorEvent.task_id)
|
||||
)
|
||||
).all()
|
||||
)
|
||||
|
||||
return [
|
||||
{
|
||||
"id": task.id,
|
||||
"name": task.name,
|
||||
"platform": task.platform,
|
||||
"mode": task.mode,
|
||||
"enabled": task.enabled,
|
||||
"interval_minutes": task.interval_minutes,
|
||||
"max_notes_count": task.max_notes_count,
|
||||
"enable_comments": task.enable_comments,
|
||||
"max_comments_count": task.max_comments_count,
|
||||
"run_timeout_seconds": task.run_timeout_seconds,
|
||||
"notify_enabled": task.notify_enabled,
|
||||
"next_run_at": task.next_run_at,
|
||||
"last_run_at": task.last_run_at,
|
||||
"last_status": task.last_status,
|
||||
"last_error": task.last_error,
|
||||
"last_notified_at": task.last_notified_at,
|
||||
"target_count": counts.get(task.id, 0),
|
||||
"targets": [
|
||||
{"id": t.id, "external_id": t.external_id, "raw_value": t.raw_value, "enabled": t.enabled}
|
||||
for t in task.targets
|
||||
],
|
||||
"unread_events": unread.get(task.id, 0),
|
||||
}
|
||||
for task in tasks
|
||||
]
|
||||
|
||||
|
||||
async def overview(session: AsyncSession, platform: Optional[str] = None) -> Dict[str, Any]:
|
||||
"""Headline numbers for the dashboard tiles, scoped to one platform."""
|
||||
now = get_current_timestamp()
|
||||
day_ago = now - 24 * 60 * 60 * 1000
|
||||
|
||||
# Nothing but the task table carries a platform column, so the other counts
|
||||
# are scoped through the platform's task ids.
|
||||
scoped: Optional[List[int]] = None
|
||||
if platform is not None:
|
||||
scoped = await platform_task_ids(session, platform)
|
||||
|
||||
def by_task(stmt, column):
|
||||
return stmt if scoped is None else stmt.where(column.in_(scoped))
|
||||
|
||||
task_count = select(func.count()).select_from(MonitorTask)
|
||||
if platform is not None:
|
||||
task_count = task_count.where(MonitorTask.platform == platform)
|
||||
|
||||
enabled_count = select(func.count()).select_from(MonitorTask).where(
|
||||
MonitorTask.enabled.is_(True)
|
||||
)
|
||||
if platform is not None:
|
||||
enabled_count = enabled_count.where(MonitorTask.platform == platform)
|
||||
|
||||
return {
|
||||
"platform": platform,
|
||||
"tasks": await session.scalar(task_count) or 0,
|
||||
"enabled_tasks": await session.scalar(enabled_count) or 0,
|
||||
"notes": await session.scalar(
|
||||
by_task(select(func.count()).select_from(MonitorNote), MonitorNote.task_id)
|
||||
)
|
||||
or 0,
|
||||
"comments": await session.scalar(
|
||||
by_task(select(func.count()).select_from(MonitorComment), MonitorComment.task_id)
|
||||
)
|
||||
or 0,
|
||||
"events_24h": await session.scalar(
|
||||
by_task(
|
||||
select(func.count())
|
||||
.select_from(MonitorEvent)
|
||||
.where(MonitorEvent.created_at >= day_ago),
|
||||
MonitorEvent.task_id,
|
||||
)
|
||||
)
|
||||
or 0,
|
||||
"unread_events": await session.scalar(
|
||||
by_task(
|
||||
select(func.count())
|
||||
.select_from(MonitorEvent)
|
||||
.where(MonitorEvent.is_read.is_(False)),
|
||||
MonitorEvent.task_id,
|
||||
)
|
||||
)
|
||||
or 0,
|
||||
"running_runs": await session.scalar(
|
||||
by_task(
|
||||
select(func.count())
|
||||
.select_from(MonitorRun)
|
||||
.where(MonitorRun.status == "running"),
|
||||
MonitorRun.task_id,
|
||||
)
|
||||
)
|
||||
or 0,
|
||||
}
|
||||
|
||||
|
||||
async def mark_events_read(session: AsyncSession, task_id: Optional[int] = None) -> int:
|
||||
query = select(MonitorEvent).where(MonitorEvent.is_read.is_(False))
|
||||
if task_id is not None:
|
||||
query = query.where(MonitorEvent.task_id == task_id)
|
||||
rows = list((await session.scalars(query)).all())
|
||||
for row in rows:
|
||||
row.is_read = True
|
||||
return len(rows)
|
||||
@@ -0,0 +1,108 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
# Copyright (c) 2025 [email protected]
|
||||
#
|
||||
# This file is part of MediaCrawler project.
|
||||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/settings.py
|
||||
# GitHub: https://github.com/NanmiCoder
|
||||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||||
#
|
||||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||||
# 1. 不得用于任何商业用途。
|
||||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||||
# 5. 不得用于任何非法或不当的用途。
|
||||
#
|
||||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||||
|
||||
"""Key/value settings for the monitoring layer, plus cookie health helpers.
|
||||
|
||||
The XHS cookie is what makes scheduled runs unattended. It expires every few
|
||||
weeks, so alongside the value we track when it was last seen working -- that is
|
||||
what lets the UI warn before a task silently stops collecting.
|
||||
"""
|
||||
|
||||
from typing import Optional
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from .models import MonitorSetting
|
||||
from .platforms import PLATFORM_XHS
|
||||
|
||||
|
||||
def platform_key(platform: str, name: str) -> str:
|
||||
"""Key for a setting that each platform keeps its own copy of."""
|
||||
return f"platform.{platform}.{name}"
|
||||
|
||||
|
||||
def system_key(name: str) -> str:
|
||||
"""Key for a setting shared across every platform."""
|
||||
return f"system.{name}"
|
||||
|
||||
|
||||
def cookie_key(platform: str) -> str:
|
||||
return platform_key(platform, "cookie")
|
||||
|
||||
|
||||
def cookie_updated_key(platform: str) -> str:
|
||||
return platform_key(platform, "cookie_updated_at")
|
||||
|
||||
|
||||
def cookie_last_ok_key(platform: str) -> str:
|
||||
return platform_key(platform, "cookie_last_ok_at")
|
||||
|
||||
|
||||
async def get_setting(session: AsyncSession, key: str) -> Optional[str]:
|
||||
return await session.scalar(select(MonitorSetting.value).where(MonitorSetting.key == key))
|
||||
|
||||
|
||||
async def set_setting(session: AsyncSession, key: str, value: str) -> None:
|
||||
row = await session.get(MonitorSetting, key)
|
||||
now = get_current_timestamp()
|
||||
if row is None:
|
||||
session.add(MonitorSetting(key=key, value=value, updated_at=now))
|
||||
else:
|
||||
row.value = value
|
||||
row.updated_at = now
|
||||
|
||||
|
||||
async def delete_setting(session: AsyncSession, key: str) -> None:
|
||||
row = await session.get(MonitorSetting, key)
|
||||
if row is not None:
|
||||
await session.delete(row)
|
||||
|
||||
|
||||
async def get_cookie(session: AsyncSession, platform: str = PLATFORM_XHS) -> str:
|
||||
return (await get_setting(session, cookie_key(platform))) or ""
|
||||
|
||||
|
||||
async def set_cookie(
|
||||
session: AsyncSession, cookie: str, platform: str = PLATFORM_XHS
|
||||
) -> None:
|
||||
await set_setting(session, cookie_key(platform), cookie)
|
||||
await set_setting(session, cookie_updated_key(platform), str(get_current_timestamp()))
|
||||
|
||||
|
||||
async def mark_cookie_ok(session: AsyncSession, platform: str = PLATFORM_XHS) -> None:
|
||||
"""Record that a run authenticated successfully."""
|
||||
await set_setting(session, cookie_last_ok_key(platform), str(get_current_timestamp()))
|
||||
|
||||
|
||||
async def get_cookie_status(session: AsyncSession, platform: str = PLATFORM_XHS) -> dict:
|
||||
"""Cookie health for the UI. Never returns the cookie value itself."""
|
||||
cookie = await get_cookie(session, platform)
|
||||
updated_at = await get_setting(session, cookie_updated_key(platform))
|
||||
last_ok_at = await get_setting(session, cookie_last_ok_key(platform))
|
||||
|
||||
return {
|
||||
"platform": platform,
|
||||
"present": bool(cookie),
|
||||
# Enough to eyeball whether the pasted value looks right, not enough to leak it.
|
||||
"length": len(cookie),
|
||||
"updated_at": int(updated_at) if updated_at else None,
|
||||
"last_ok_at": int(last_ok_at) if last_ok_at else None,
|
||||
}
|
||||
Reference in New Issue
Block a user