fix(report): 空的任务集合被当成了「不限制平台」,导致报表串平台数据
现象:切到抖音,报表里显示的是小红书的数据。
根因是报告聚合里的一行真值判断:
scope = list(task_ids) if task_ids else None
空列表是假值,而空列表在这里的含义是「这个平台一个任务都没有」,不是「不限制平台」。
于是 platform=dy 且抖音还没有任务时,_resolve_scope 返回的 [] 被翻译成了 None,
聚合范围从「抖音的任务」变成了**全部任务** —— 小红书的数字就这么显示在了抖音页面上。
顺带 task_ids 也回成 None,界面会显示成「全部任务」。
改成 `is not None`。空列表进去就让 in_([]) 恒假,结果为空,这才是对的。
排查时把所有同类写法过了一遍,只有这一处错,其余(service.py 的 10 处作用域judgement、
_resolve_scope、export)用的都是 `is not None`。
测试:新增两条,并且**验证过它们在修复前会红**(失败信息就是 assert 42 == 0 ——
查一个没有任何任务的平台,却返回了小红书那条作品的 42 个赞)。
同时修掉一条空跑的测试:test_a_platform_with_no_tasks_yields_empty_not_everything
原先种了任务却没有作品/指标数据,于是过滤生效与否结果都是 0,什么都测不出来 ——
这正是这个 bug 能活下来的原因。现在它会真的塞一条作品+快照进去,并在末尾断言
「小红书自己的报表看得到那条数据」,用来证明前面那两个 0 是过滤出来的而不是没数据。
This commit is contained in:
+44
-3
@@ -21,13 +21,15 @@
|
||||
import httpx
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from sqlalchemy import text
|
||||
from sqlalchemy import select, text
|
||||
|
||||
from tools.time_util import get_current_timestamp
|
||||
|
||||
from api.main import app
|
||||
from api.monitor import adapters
|
||||
from api.monitor import db as monitor_db
|
||||
from api.monitor import platforms
|
||||
from api.monitor.models import MonitorTask
|
||||
from api.monitor.models import MonitorNote, MonitorNoteMetric, MonitorTask
|
||||
|
||||
XHS_TARGET = "5f58bd990000000001003753"
|
||||
|
||||
@@ -44,6 +46,33 @@ async def client(tmp_path):
|
||||
await monitor_db.dispose_engine()
|
||||
|
||||
|
||||
async def _seed_note_with_one_snapshot(task_name: str) -> None:
|
||||
"""给某个任务塞一条作品和一次指标快照。
|
||||
|
||||
过滤类测试**必须有真数据**才有意义 —— 库里空着的话,过滤有没有生效结果都是 0,
|
||||
测试就变成了空跑(这个坑踩过一次:一个报表串数据的 bug 因此没被拦住)。
|
||||
"""
|
||||
async with monitor_db.get_session() as session:
|
||||
task = await session.scalar(select(MonitorTask).where(MonitorTask.name == task_name))
|
||||
now = get_current_timestamp()
|
||||
session.add(
|
||||
MonitorNote(
|
||||
task_id=task.id, note_id="seed-note", title="seed", note_url="",
|
||||
cover="", creator_hash="", source_kind="", published_at=None,
|
||||
first_seen_run_id=1, first_seen_at=now,
|
||||
last_seen_run_id=1, last_seen_at=now,
|
||||
)
|
||||
)
|
||||
session.add(
|
||||
MonitorNoteMetric(
|
||||
task_id=task.id, note_id="seed-note", run_id=1, captured_at=now,
|
||||
liked_count=42, comment_count=0, collected_count=0, share_count=0,
|
||||
raw_liked_count="42", raw_comment_count="0",
|
||||
raw_collected_count="0", raw_share_count="0",
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
class TestCapabilityMatrix:
|
||||
@pytest.mark.asyncio
|
||||
async def test_matrix_is_exposed_to_the_ui(self, client):
|
||||
@@ -191,8 +220,15 @@ class TestPlatformScoping:
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_platform_with_no_tasks_yields_empty_not_everything(self, client):
|
||||
"""An empty task set must not degrade into "no filter"."""
|
||||
"""空的任务集合不能退化成「不加过滤」。
|
||||
|
||||
**这里必须真的有数据。** 没有数据时,过滤生效与否结果都是 0 —— 这条测试原先
|
||||
就栽在这个空跑上,所以没能拦下一个报表串数据的 bug(切到抖音,报表里却出现
|
||||
小红书的数据)。最后那段「小红书自己的报表看得到」就是为了证明这些数据确实
|
||||
存在、上面那两个 0 是过滤出来的。
|
||||
"""
|
||||
await self._seed_two_platforms(client)
|
||||
await _seed_note_with_one_snapshot("小红书任务")
|
||||
|
||||
body = (await client.get("/api/monitor/notes", params={"platform": "bili"})).json()
|
||||
assert body["notes"] == []
|
||||
@@ -203,6 +239,11 @@ class TestPlatformScoping:
|
||||
assert report["totals"]["liked_count_delta"] == 0
|
||||
assert report["note_count"] == 0
|
||||
|
||||
xhs = (
|
||||
await client.get("/api/monitor/report", params={"platform": "xhs"})
|
||||
).json()
|
||||
assert xhs["note_count"] == 1
|
||||
|
||||
|
||||
class TestPerPlatformSettings:
|
||||
@pytest.mark.asyncio
|
||||
|
||||
Reference in New Issue
Block a user