feat(upstream): 上游更新检查——定期比上游、落后了推企业微信
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s

本仓库在上游 MediaCrawler 之上加了一整层(见 UPSTREAM.md),可合并流程默认
「有人知道上游动了」。而部署是 git pull --ff-only,只从自己的 Gitea 拉——上游的
提交不主动 fetch 就永远看不见。拖着不合并的代价是复利的:越久越难合。

于是把「上游动了没有」变成一条会自己跑、会推企业微信的通知:

* api/monitor/upstream.py:git fetch <url> <branch> 到 FETCH_HEAD,用
  rev-list --count HEAD..FETCH_HEAD 算落后数、FETCH_HEAD..HEAD 算领先数。
  用 git 而非托管商 API,因为只有 git 知道共同祖先在哪——本仓库含有上游没有的
  提交,直接比 tip 会得出错误结论。增量 fetch 只传几个新提交,不会遇到
  UPSTREAM.md 里说的「大包必断」。
* 只 fetch 到 FETCH_HEAD:不配 remote、不写 refs/remotes、不碰索引与工作区,
  所以不打断正在跑的采集,也不和 deploy.sh 的 git pull 抢锁。
* 挂在调度器 tick 上(不是采集,所以不看 is_busy、不受活跃时段限制——定时检查
  放在半夜反而最合适),按 checked_at + 间隔 到期才跑;失败也写 checked_at,
  于是 GitHub 不通时是每间隔重试一次,而不是每个 tick 撞一次墙。
* 同一个 tip 只推一次(记 tip 而不是「推过没」),上游真又动了会再推。
* 两个接口:GET /monitor/upstream 只读缓存;POST /monitor/upstream/check 手动
  查一次且刻意不推通知——点按钮的人正看着结果。
* 默认关闭,间隔默认一天。

Dockerfile 显式装 git(python:slim 不带,而这是唯一的依赖);deploy.sh 顺带补上
一个真 bug 的提示:Dockerfile/requirements.txt 变了只 up -d 用的还是旧镜像。
This commit is contained in:
2026-10-10 09:17:06 +08:00
parent 44cbe8e2aa
commit e348de48d3
15 changed files with 1183 additions and 6 deletions
+47 -1
View File
@@ -32,6 +32,11 @@ Two families of schedule, and the difference matters:
calendar, so a run that starts late does not drag every later run with it.
The arithmetic for both lives in schedule.py.
The loop also carries the 上游更新检查: it is not a crawl, so it shares none of
the rules above (no subprocess, no active-hours gate) -- see
``_maybe_check_upstream``. It rides this loop rather than getting a thread of its
own because it is one HTTP-shaped fetch per day.
"""
import asyncio
@@ -44,7 +49,7 @@ from sqlalchemy import select
from tools.time_util import get_current_timestamp
from ..services import crawler_manager
from . import app_settings, schedule
from . import app_settings, schedule, upstream
from .db import get_session
from .models import MonitorRun, MonitorTask, RUN_INTERRUPTED, RUN_RUNNING
from .runner import execute_task
@@ -92,8 +97,49 @@ class MonitorScheduler:
await self.tick()
except Exception as exc: # pragma: no cover - keep the loop alive
print(f"[monitor.scheduler] tick failed: {exc}")
# 独立于采集任务,因此单独一段 try:上游检查失败不该影响采集调度,
# 反过来也一样。
try:
await self._maybe_check_upstream()
except Exception as exc: # pragma: no cover - keep the loop alive
print(f"[monitor.scheduler] upstream check failed: {exc}")
await asyncio.sleep(POLL_INTERVAL_SECONDS)
async def _maybe_check_upstream(self) -> None:
"""到点就 fetch 一次上游仓库,看它有没有新提交。
与采集任务的三条规则都不同,各有理由:它不碰浏览器、也不占采集子进程,
所以不看 ``is_busy``;它只发一个 git 请求,没有被平台风控的风险,所以也不
受活跃时段限制 —— 定时检查放在半夜反而是最合适的。
"""
async with get_session() as session:
if not await app_settings.get_value(
session, "upstream_check_enabled", fallback=False
):
return
interval_minutes = int(
await app_settings.get_value(
session, "upstream_check_interval_minutes", fallback=1440
)
)
state = await upstream.load_state(session)
checked_at = int(state.get("checked_at") or 0)
now = get_current_timestamp()
# 失败也会写 checked_at,所以不通的时候同样是每个间隔重试一次,
# 而不是每个 tick(20 秒)都去撞一次墙。
if checked_at and now - checked_at < max(1, interval_minutes) * 60_000:
return
result = await upstream.run_check()
if result.get("behind"):
print(
f"[monitor.scheduler] 上游 {result.get('branch')} 领先 "
f"{result['behind']} 个提交"
)
elif not result.get("ok"):
print(f"[monitor.scheduler] 上游检查失败:{result.get('error')}")
async def recover(self) -> None:
"""Clean up state left behind by a server restart.