feat(upstream): 上游更新检查——定期比上游、落后了推企业微信
本仓库在上游 MediaCrawler 之上加了一整层(见 UPSTREAM.md),可合并流程默认 「有人知道上游动了」。而部署是 git pull --ff-only,只从自己的 Gitea 拉——上游的 提交不主动 fetch 就永远看不见。拖着不合并的代价是复利的:越久越难合。 于是把「上游动了没有」变成一条会自己跑、会推企业微信的通知: * api/monitor/upstream.py:git fetch <url> <branch> 到 FETCH_HEAD,用 rev-list --count HEAD..FETCH_HEAD 算落后数、FETCH_HEAD..HEAD 算领先数。 用 git 而非托管商 API,因为只有 git 知道共同祖先在哪——本仓库含有上游没有的 提交,直接比 tip 会得出错误结论。增量 fetch 只传几个新提交,不会遇到 UPSTREAM.md 里说的「大包必断」。 * 只 fetch 到 FETCH_HEAD:不配 remote、不写 refs/remotes、不碰索引与工作区, 所以不打断正在跑的采集,也不和 deploy.sh 的 git pull 抢锁。 * 挂在调度器 tick 上(不是采集,所以不看 is_busy、不受活跃时段限制——定时检查 放在半夜反而最合适),按 checked_at + 间隔 到期才跑;失败也写 checked_at, 于是 GitHub 不通时是每间隔重试一次,而不是每个 tick 撞一次墙。 * 同一个 tip 只推一次(记 tip 而不是「推过没」),上游真又动了会再推。 * 两个接口:GET /monitor/upstream 只读缓存;POST /monitor/upstream/check 手动 查一次且刻意不推通知——点按钮的人正看着结果。 * 默认关闭,间隔默认一天。 Dockerfile 显式装 git(python:slim 不带,而这是唯一的依赖);deploy.sh 顺带补上 一个真 bug 的提示:Dockerfile/requirements.txt 变了只 up -d 用的还是旧镜像。
This commit is contained in:
@@ -32,6 +32,11 @@ Two families of schedule, and the difference matters:
|
||||
calendar, so a run that starts late does not drag every later run with it.
|
||||
|
||||
The arithmetic for both lives in schedule.py.
|
||||
|
||||
The loop also carries the 上游更新检查: it is not a crawl, so it shares none of
|
||||
the rules above (no subprocess, no active-hours gate) -- see
|
||||
``_maybe_check_upstream``. It rides this loop rather than getting a thread of its
|
||||
own because it is one HTTP-shaped fetch per day.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
@@ -44,7 +49,7 @@ from sqlalchemy import select
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from ..services import crawler_manager
|
||||
from . import app_settings, schedule
|
||||
from . import app_settings, schedule, upstream
|
||||
from .db import get_session
|
||||
from .models import MonitorRun, MonitorTask, RUN_INTERRUPTED, RUN_RUNNING
|
||||
from .runner import execute_task
|
||||
@@ -92,8 +97,49 @@ class MonitorScheduler:
|
||||
await self.tick()
|
||||
except Exception as exc: # pragma: no cover - keep the loop alive
|
||||
print(f"[monitor.scheduler] tick failed: {exc}")
|
||||
# 独立于采集任务,因此单独一段 try:上游检查失败不该影响采集调度,
|
||||
# 反过来也一样。
|
||||
try:
|
||||
await self._maybe_check_upstream()
|
||||
except Exception as exc: # pragma: no cover - keep the loop alive
|
||||
print(f"[monitor.scheduler] upstream check failed: {exc}")
|
||||
await asyncio.sleep(POLL_INTERVAL_SECONDS)
|
||||
|
||||
async def _maybe_check_upstream(self) -> None:
|
||||
"""到点就 fetch 一次上游仓库,看它有没有新提交。
|
||||
|
||||
与采集任务的三条规则都不同,各有理由:它不碰浏览器、也不占采集子进程,
|
||||
所以不看 ``is_busy``;它只发一个 git 请求,没有被平台风控的风险,所以也不
|
||||
受活跃时段限制 —— 定时检查放在半夜反而是最合适的。
|
||||
"""
|
||||
async with get_session() as session:
|
||||
if not await app_settings.get_value(
|
||||
session, "upstream_check_enabled", fallback=False
|
||||
):
|
||||
return
|
||||
interval_minutes = int(
|
||||
await app_settings.get_value(
|
||||
session, "upstream_check_interval_minutes", fallback=1440
|
||||
)
|
||||
)
|
||||
state = await upstream.load_state(session)
|
||||
|
||||
checked_at = int(state.get("checked_at") or 0)
|
||||
now = get_current_timestamp()
|
||||
# 失败也会写 checked_at,所以不通的时候同样是每个间隔重试一次,
|
||||
# 而不是每个 tick(20 秒)都去撞一次墙。
|
||||
if checked_at and now - checked_at < max(1, interval_minutes) * 60_000:
|
||||
return
|
||||
|
||||
result = await upstream.run_check()
|
||||
if result.get("behind"):
|
||||
print(
|
||||
f"[monitor.scheduler] 上游 {result.get('branch')} 领先 "
|
||||
f"{result['behind']} 个提交"
|
||||
)
|
||||
elif not result.get("ok"):
|
||||
print(f"[monitor.scheduler] 上游检查失败:{result.get('error')}")
|
||||
|
||||
async def recover(self) -> None:
|
||||
"""Clean up state left behind by a server restart.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user