Files
MediaCrawler/api/monitor/douyin_fetch.py
T
butubb 20e672834c
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
feat(monitor): 博主的粉丝数,以及给作品起备注
两件都是「作品栏里把这东西认出来」的延伸:

* **账号级指标**:作品列表只会说「这条涨了多少赞」,说不了「这个人整个
  账号的粉丝在涨还是在掉」。抖音的资料接口本来就有粉丝数/总获赞/作品数,
  每轮顺手记一条快照(`monitor_creator_stat`,粒度 = 任务×博主×轮次,
  和作品指标同形)。组头显示最近一条。

  快照在「一条作品都没采到」的早退**之前**落:作品列表被风控挡住的那一轮,
  正是「粉丝还在涨、但新作品没在发现」最该被看见的时刻。

* **作品备注**:博主备注回答「这个账号是谁」,这条回答「这条我要盯着」。
  一个博主底下常常只有一两件值得盯的作品,所以不能合并成一条。键取
  (platform, note_id),跨任务共用一份。

两边都守住同一条口径:**不知道就是 null,不写成 0** —— 0 在趋势图上是一条
砸到底的线,和「还没采到」是两回事。
2026-10-10 18:09:03 +08:00

218 lines
9.3 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# -*- coding: utf-8 -*-
# Copyright (c) 2025 [email protected]
#
# This file is part of MediaCrawler project.
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_fetch.py
# GitHub: https://github.com/NanmiCoder
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
#
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
# 1. 不得用于任何商业用途。
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
# 3. 不得进行大规模爬取或对平台造成运营干扰。
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
# 5. 不得用于任何非法或不当的用途。
#
# 详细许可条款请参阅项目根目录下的LICENSE文件。
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
"""抖音采集:直接走 Web 接口,不起爬虫子进程。
与 ``media_platform/douyin`` 那条路的分工:
* 那边起一个 Playwright 子进程、构造一大串**自相矛盾的浏览器指纹参数**(参数说自己是
Mac + Chrome 125,UA 说自己是 Linux + Chrome 155),网关回一个 200 + 空 body,
然后被翻译成 ``Exception("account blocked")`` —— 看起来像账号被封,其实什么都不是。
* 这边不起子进程、只发必要参数,用浏览器里那份登录态。见 ``douyin_api``。
采到的东西写成 ``store/douyin`` 那套 jsonl 形状,所以 **ingest 完全不知道数据是怎么来的**
—— 重采样、差分、事件、通知、报表全都照旧。
"""
import json
from datetime import datetime
from pathlib import Path
from typing import Any, Dict, Iterable, List, Optional, Sequence
from tools import utils
from . import adapters, douyin_api
from .models import MODE_CREATOR, MODE_NOTE, MonitorTask
# 拉多少条作品。接口单页上限就是 20,要多了也没用。
DEFAULT_VIDEO_LIMIT = 20
async def collect(
out_dir: Path,
*,
platform: str,
mode: str,
limit: int,
want_comments: bool,
comment_limit: int,
targets: Sequence[Any],
known_aweme_ids: Iterable[str] = (),
cookie: str = "",
) -> Dict[str, Any]:
"""跑一轮抖音采集,把产物写进 ``out_dir`` 下 store 那个目录里。
参数是散的、不收 ORM 对象:调用方那边 ``task`` 在会话关掉之后就 detached 了,
传对象进来迟早会踩到「属性已过期」。
**不抛异常**:失败也把(可能为空的)产物落下去,并把原因放进 ``errors`` 交给调用方。
让调用方去决定这是「一轮正常但没数据」还是「一轮失败」—— 这个判断不该藏在这里。
"""
notes: List[Dict[str, Any]] = []
comments: List[Dict[str, Any]] = []
# 博主**账号级**指标(粉丝 / 总获赞 / 作品数)。作品列表之外单独要一次,
# 只有博主模式才有 —— 作品模式的目标是一件作品,没有"这个博主是谁"可问。
profiles: List[Dict[str, Any]] = []
errors: List[str] = []
# **整个 collect 只去重一次的、跨目标的集合**:退化路径会把「库里已知的全部作品」
# 在每个目标下都刷一遍,多个目标就会出现同一件作品好几条记录 —— 而一对一快照的
# 唯一键是 (task_id, note_id, run_id),同一条作品在一轮里出现两次会直接撞键。
seen_aweme: set = set()
for target in targets:
external_id = target.external_id
# 两种模式的目标是不同的东西,不能走同一条路:
# 作品模式 —— 目标本身就是作品 id,直接取详情(**这个接口没被挡,今天就能用**)。
# 博主模式 —— 目标是主页 sec_uid,要先拉作品列表;那个接口被真校验挡着,退化到
# 刷新库里已知的作品(新作品发现不了)。
if mode == MODE_NOTE:
try:
videos = [await douyin_api.video_detail(external_id, cookie=cookie)]
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取作品 {external_id} 失败:{exc}")
videos = []
else:
videos = await _creator_works(
external_id, limit, known_aweme_ids, seen_aweme, cookie, errors
)
profile = await _creator_profile(external_id, videos, cookie, errors)
if profile is not None:
profiles.append(profile)
for video in videos:
aweme_id = video.get("aweme_id")
if not aweme_id or aweme_id in seen_aweme:
continue
seen_aweme.add(aweme_id)
notes.append(video)
if want_comments:
try:
comments.extend(
await douyin_api.video_comments(
aweme_id, count=comment_limit, cookie=cookie
)
)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取作品 {aweme_id} 的评论失败:{exc}")
jsonl_dir = _write_artifacts(out_dir, platform, mode, notes, comments, profiles)
return {
"notes": len(notes),
"comments": len(comments),
"errors": errors,
"jsonl_dir": str(jsonl_dir),
}
async def _creator_profile(
sec_user_id: str,
videos: Sequence[Dict[str, Any]],
cookie: str,
errors: List[str],
) -> Optional[Dict[str, Any]]:
"""问一次博主的账号级指标。拿不到就算了 —— **不能因为顺手的附加信息失败,
就把这一轮本来采到的作品也判成失败。**
creator_hash 优先取作品自带的那个:作品是靠 ``anonymize_user_id(author.uid)`` 得到
哈希的,而快照表和作品必须对得上号,否则界面上永远查不出这个博主的粉丝数。只有当一件
作品都没采到时(列表被挡且没有已知作品可刷新),才退回资料接口自己算的哈希 ——
那种情况下也只剩它了。
"""
try:
profile = await douyin_api.author_profile(sec_user_id, cookie=cookie)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取博主 {sec_user_id} 的资料失败:{exc}")
return None
if videos:
profile["creator_hash"] = videos[0].get("creator_hash") or profile["creator_hash"]
if not profile.get("creator_hash"):
# 哈希都算不出来的快照没人能查到,落下去只是垃圾。
errors.append(f"博主 {sec_user_id} 的资料里没有可用的身份标识,跳过账号指标")
return None
return profile
async def _creator_works(
sec_user_id: str,
limit: int,
known_aweme_ids: Iterable[str],
seen_aweme: set,
cookie: str,
errors: List[str],
) -> List[Dict[str, Any]]:
"""一个博主的作品:先要列表,列表被挡时退化成刷新已知作品。
作品列表(``aweme/post``)被抖音单独加了真校验 —— 不带 ``x-tt-argus`` 回 403,
带上 dummy 值回 200 + 空 body。所以这里拿不到**新**作品,只能保住已知的。
"""
try:
return await douyin_api.author_videos(sec_user_id, count=limit, cookie=cookie)
except douyin_api.DouyinApiError as exc:
errors.append(f"拉取博主 {sec_user_id} 的作品列表失败:{exc}")
refreshed: List[Dict[str, Any]] = []
for aweme_id in known_aweme_ids:
if aweme_id in seen_aweme:
continue
try:
refreshed.append(await douyin_api.video_detail(aweme_id, cookie=cookie))
except douyin_api.DouyinApiError as detail_exc:
errors.append(f"刷新作品 {aweme_id} 失败:{detail_exc}")
return refreshed
def _write_artifacts(
out_dir: Path,
platform: str,
mode: str,
notes: List[Dict[str, Any]],
comments: List[Dict[str, Any]],
profiles: Sequence[Dict[str, Any]] = (),
) -> Path:
"""按爬虫那套目录与文件名写 jsonl。
目录名取 ``adapters.artifact_dir``(抖音是 ``douyin``,而不是平台 id ``dy``)——
和 ingest 找文件用的是同一个来源,两边不会走散。
"""
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
jsonl_dir.mkdir(parents=True, exist_ok=True)
kind = "creator" if mode == MODE_CREATOR else "detail"
date = datetime.now().strftime("%Y-%m-%d")
_write_jsonl(jsonl_dir / f"{kind}_contents_{date}.jsonl", notes)
# 评论文件即使没有评论也建出来:ingest 靠「文件在不在」区分「这一轮没评论」和
# 「这一轮什么都没抓到」,两种情况的含义完全不同。
_write_jsonl(jsonl_dir / f"{kind}_comments_{date}.jsonl", comments)
# 博主资料同样无条件写:空文件表示"问了但没问到",没有文件表示"这次根本没问"
# (作品模式)。两者在 ingest 那边走的是同一条路(都不落快照),但留空文件能让
# 事后翻 run 目录时看出到底问没问过。
_write_jsonl(jsonl_dir / f"{kind}_profile_{date}.jsonl", list(profiles))
return jsonl_dir
def _write_jsonl(path: Path, records: List[Dict[str, Any]]) -> None:
with path.open("w", encoding="utf-8") as handle:
for record in records:
handle.write(json.dumps(record, ensure_ascii=False) + "\n")
utils.logger.info(f"[douyin_fetch] 写出 {len(records)} 条 -> {path.name}")