Files
MediaCrawler/tests/test_monitor_runner.py
T
butubb 9e13a7f686
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
fix(monitor): 抖音任务的 run 永远停在「排队中」—— 我上一版把状态标记缩进错了
用户报的现象:任务一直显示「排队中」。查库确认有两批 run 卡在 pending(任务 7 的 44/45、
任务 15 的 62/63)。两个原因,一个是我上一版改坏的:

1) **`RUN_RUNNING` 被我缩进进了爬虫那条分支。** 抖音走的是另一条路,于是它**从不标记
   「运行中」** —— 建完 pending 那一行就直接进采集,中途一旦出事(异常、进程被重启),
   状态就永远停在 pending。这是我加平台分岔时把原本在两条路公共位置的一行挪进去了。

2) **`recover()` 只收 `running`,够不着 `pending`。** 那行是上一轮建的、后面的采集却
   根本没机会开始(进程重启),它永远不会自己往前走。于是重启也救不回来,界面上就是
   一个永远「排队中」的幽灵。现在 pending 一起收。

两处都补了测试:抖音路的 run 必须在**采集开始之前**就已经是 running(这条改回去就会
失败);recover 要把 pending 也标成 interrupted。
2026-10-10 17:47:41 +08:00

83 lines
2.8 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# -*- coding: utf-8 -*-
"""monitor runner —— 尤其是抖音那条(不走子进程的)路的运行状态流转。
这条路的地位特殊:它不经过 ``crawler_manager``,所以爬虫那套「退出码 / 日志尾巴」的
约定它一个都不沾。凡是写在那里面的东西,这条路都得单独有一份。
"""
import pytest
import pytest_asyncio
from sqlalchemy import select
from tools.time_util import get_current_timestamp
from api.monitor import db as monitor_db
from api.monitor import runner as runner_module
from api.monitor.models import (
MODE_CREATOR,
RUN_RUNNING,
MonitorRun,
MonitorTarget,
MonitorTask,
)
@pytest_asyncio.fixture
async def db(tmp_path):
monitor_db.set_sqlite_path(tmp_path / "monitor.db")
await monitor_db.init_db()
yield monitor_db
await monitor_db.dispose_engine()
async def _make_douyin_task() -> int:
async with monitor_db.get_session() as session:
now = get_current_timestamp()
task = MonitorTask(
name="dy", platform="dy", mode=MODE_CREATOR, enabled=True,
interval_minutes=360, max_notes_count=20, enable_comments=False,
max_comments_count=20, run_timeout_seconds=3600,
notify_enabled=False, notify_failures=False,
created_at=now, updated_at=now,
)
session.add(task)
await session.flush()
session.add(
MonitorTarget(
task_id=task.id, kind=MODE_CREATOR, external_id="MS4w-sec",
xsec_token="", xsec_source="", raw_value="MS4w-sec",
label="x", enabled=True, created_at=now,
)
)
return task.id
class TestDouyinRunStatus:
@pytest.mark.asyncio
async def test_the_run_is_marked_running_before_collecting(self, db, monkeypatch):
"""**采集开始之前**,run 就必须已经是 running。
这一行原先只写在爬虫那条分支里,于是抖音路上 run 一直停在 pending —— 一旦中途
出事(异常、或进程被重启),界面上就是一个永远「排队中」的幽灵,而且 recover()
当时也只收 running、够不着它。
"""
task_id = await _make_douyin_task()
seen = {}
async def fake_collect(out_dir, **kwargs):
async with monitor_db.get_session() as session:
run = await session.scalar(select(MonitorRun).order_by(MonitorRun.id))
seen["status"] = run.status
return {
"notes": 0,
"comments": 0,
"errors": ["故意失败"],
"jsonl_dir": str(out_dir),
}
monkeypatch.setattr(runner_module.douyin_fetch, "collect", fake_collect)
await runner_module.execute_task(task_id, trigger="manual")
assert seen["status"] == RUN_RUNNING