两件都是「作品栏里把这东西认出来」的延伸: * **账号级指标**:作品列表只会说「这条涨了多少赞」,说不了「这个人整个 账号的粉丝在涨还是在掉」。抖音的资料接口本来就有粉丝数/总获赞/作品数, 每轮顺手记一条快照(`monitor_creator_stat`,粒度 = 任务×博主×轮次, 和作品指标同形)。组头显示最近一条。 快照在「一条作品都没采到」的早退**之前**落:作品列表被风控挡住的那一轮, 正是「粉丝还在涨、但新作品没在发现」最该被看见的时刻。 * **作品备注**:博主备注回答「这个账号是谁」,这条回答「这条我要盯着」。 一个博主底下常常只有一两件值得盯的作品,所以不能合并成一条。键取 (platform, note_id),跨任务共用一份。 两边都守住同一条口径:**不知道就是 null,不写成 0** —— 0 在趋势图上是一条 砸到底的线,和「还没采到」是两回事。
218 lines
9.3 KiB
Python
218 lines
9.3 KiB
Python
# -*- coding: utf-8 -*-
|
||
# Copyright (c) 2025 [email protected]
|
||
#
|
||
# This file is part of MediaCrawler project.
|
||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_fetch.py
|
||
# GitHub: https://github.com/NanmiCoder
|
||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||
#
|
||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||
# 1. 不得用于任何商业用途。
|
||
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
|
||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||
# 5. 不得用于任何非法或不当的用途。
|
||
#
|
||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||
|
||
"""抖音采集:直接走 Web 接口,不起爬虫子进程。
|
||
|
||
与 ``media_platform/douyin`` 那条路的分工:
|
||
|
||
* 那边起一个 Playwright 子进程、构造一大串**自相矛盾的浏览器指纹参数**(参数说自己是
|
||
Mac + Chrome 125,UA 说自己是 Linux + Chrome 155),网关回一个 200 + 空 body,
|
||
然后被翻译成 ``Exception("account blocked")`` —— 看起来像账号被封,其实什么都不是。
|
||
* 这边不起子进程、只发必要参数,用浏览器里那份登录态。见 ``douyin_api``。
|
||
|
||
采到的东西写成 ``store/douyin`` 那套 jsonl 形状,所以 **ingest 完全不知道数据是怎么来的**
|
||
—— 重采样、差分、事件、通知、报表全都照旧。
|
||
"""
|
||
|
||
import json
|
||
from datetime import datetime
|
||
from pathlib import Path
|
||
from typing import Any, Dict, Iterable, List, Optional, Sequence
|
||
|
||
from tools import utils
|
||
|
||
from . import adapters, douyin_api
|
||
from .models import MODE_CREATOR, MODE_NOTE, MonitorTask
|
||
|
||
# 拉多少条作品。接口单页上限就是 20,要多了也没用。
|
||
DEFAULT_VIDEO_LIMIT = 20
|
||
|
||
|
||
async def collect(
|
||
out_dir: Path,
|
||
*,
|
||
platform: str,
|
||
mode: str,
|
||
limit: int,
|
||
want_comments: bool,
|
||
comment_limit: int,
|
||
targets: Sequence[Any],
|
||
known_aweme_ids: Iterable[str] = (),
|
||
cookie: str = "",
|
||
) -> Dict[str, Any]:
|
||
"""跑一轮抖音采集,把产物写进 ``out_dir`` 下 store 那个目录里。
|
||
|
||
参数是散的、不收 ORM 对象:调用方那边 ``task`` 在会话关掉之后就 detached 了,
|
||
传对象进来迟早会踩到「属性已过期」。
|
||
|
||
**不抛异常**:失败也把(可能为空的)产物落下去,并把原因放进 ``errors`` 交给调用方。
|
||
让调用方去决定这是「一轮正常但没数据」还是「一轮失败」—— 这个判断不该藏在这里。
|
||
"""
|
||
notes: List[Dict[str, Any]] = []
|
||
comments: List[Dict[str, Any]] = []
|
||
# 博主**账号级**指标(粉丝 / 总获赞 / 作品数)。作品列表之外单独要一次,
|
||
# 只有博主模式才有 —— 作品模式的目标是一件作品,没有"这个博主是谁"可问。
|
||
profiles: List[Dict[str, Any]] = []
|
||
errors: List[str] = []
|
||
|
||
# **整个 collect 只去重一次的、跨目标的集合**:退化路径会把「库里已知的全部作品」
|
||
# 在每个目标下都刷一遍,多个目标就会出现同一件作品好几条记录 —— 而一对一快照的
|
||
# 唯一键是 (task_id, note_id, run_id),同一条作品在一轮里出现两次会直接撞键。
|
||
seen_aweme: set = set()
|
||
|
||
for target in targets:
|
||
external_id = target.external_id
|
||
|
||
# 两种模式的目标是不同的东西,不能走同一条路:
|
||
# 作品模式 —— 目标本身就是作品 id,直接取详情(**这个接口没被挡,今天就能用**)。
|
||
# 博主模式 —— 目标是主页 sec_uid,要先拉作品列表;那个接口被真校验挡着,退化到
|
||
# 刷新库里已知的作品(新作品发现不了)。
|
||
if mode == MODE_NOTE:
|
||
try:
|
||
videos = [await douyin_api.video_detail(external_id, cookie=cookie)]
|
||
except douyin_api.DouyinApiError as exc:
|
||
errors.append(f"拉取作品 {external_id} 失败:{exc}")
|
||
videos = []
|
||
else:
|
||
videos = await _creator_works(
|
||
external_id, limit, known_aweme_ids, seen_aweme, cookie, errors
|
||
)
|
||
profile = await _creator_profile(external_id, videos, cookie, errors)
|
||
if profile is not None:
|
||
profiles.append(profile)
|
||
|
||
for video in videos:
|
||
aweme_id = video.get("aweme_id")
|
||
if not aweme_id or aweme_id in seen_aweme:
|
||
continue
|
||
seen_aweme.add(aweme_id)
|
||
notes.append(video)
|
||
|
||
if want_comments:
|
||
try:
|
||
comments.extend(
|
||
await douyin_api.video_comments(
|
||
aweme_id, count=comment_limit, cookie=cookie
|
||
)
|
||
)
|
||
except douyin_api.DouyinApiError as exc:
|
||
errors.append(f"拉取作品 {aweme_id} 的评论失败:{exc}")
|
||
|
||
jsonl_dir = _write_artifacts(out_dir, platform, mode, notes, comments, profiles)
|
||
return {
|
||
"notes": len(notes),
|
||
"comments": len(comments),
|
||
"errors": errors,
|
||
"jsonl_dir": str(jsonl_dir),
|
||
}
|
||
|
||
|
||
async def _creator_profile(
|
||
sec_user_id: str,
|
||
videos: Sequence[Dict[str, Any]],
|
||
cookie: str,
|
||
errors: List[str],
|
||
) -> Optional[Dict[str, Any]]:
|
||
"""问一次博主的账号级指标。拿不到就算了 —— **不能因为顺手的附加信息失败,
|
||
就把这一轮本来采到的作品也判成失败。**
|
||
|
||
creator_hash 优先取作品自带的那个:作品是靠 ``anonymize_user_id(author.uid)`` 得到
|
||
哈希的,而快照表和作品必须对得上号,否则界面上永远查不出这个博主的粉丝数。只有当一件
|
||
作品都没采到时(列表被挡且没有已知作品可刷新),才退回资料接口自己算的哈希 ——
|
||
那种情况下也只剩它了。
|
||
"""
|
||
try:
|
||
profile = await douyin_api.author_profile(sec_user_id, cookie=cookie)
|
||
except douyin_api.DouyinApiError as exc:
|
||
errors.append(f"拉取博主 {sec_user_id} 的资料失败:{exc}")
|
||
return None
|
||
|
||
if videos:
|
||
profile["creator_hash"] = videos[0].get("creator_hash") or profile["creator_hash"]
|
||
if not profile.get("creator_hash"):
|
||
# 哈希都算不出来的快照没人能查到,落下去只是垃圾。
|
||
errors.append(f"博主 {sec_user_id} 的资料里没有可用的身份标识,跳过账号指标")
|
||
return None
|
||
return profile
|
||
|
||
|
||
async def _creator_works(
|
||
sec_user_id: str,
|
||
limit: int,
|
||
known_aweme_ids: Iterable[str],
|
||
seen_aweme: set,
|
||
cookie: str,
|
||
errors: List[str],
|
||
) -> List[Dict[str, Any]]:
|
||
"""一个博主的作品:先要列表,列表被挡时退化成刷新已知作品。
|
||
|
||
作品列表(``aweme/post``)被抖音单独加了真校验 —— 不带 ``x-tt-argus`` 回 403,
|
||
带上 dummy 值回 200 + 空 body。所以这里拿不到**新**作品,只能保住已知的。
|
||
"""
|
||
try:
|
||
return await douyin_api.author_videos(sec_user_id, count=limit, cookie=cookie)
|
||
except douyin_api.DouyinApiError as exc:
|
||
errors.append(f"拉取博主 {sec_user_id} 的作品列表失败:{exc}")
|
||
|
||
refreshed: List[Dict[str, Any]] = []
|
||
for aweme_id in known_aweme_ids:
|
||
if aweme_id in seen_aweme:
|
||
continue
|
||
try:
|
||
refreshed.append(await douyin_api.video_detail(aweme_id, cookie=cookie))
|
||
except douyin_api.DouyinApiError as detail_exc:
|
||
errors.append(f"刷新作品 {aweme_id} 失败:{detail_exc}")
|
||
return refreshed
|
||
|
||
|
||
def _write_artifacts(
|
||
out_dir: Path,
|
||
platform: str,
|
||
mode: str,
|
||
notes: List[Dict[str, Any]],
|
||
comments: List[Dict[str, Any]],
|
||
profiles: Sequence[Dict[str, Any]] = (),
|
||
) -> Path:
|
||
"""按爬虫那套目录与文件名写 jsonl。
|
||
|
||
目录名取 ``adapters.artifact_dir``(抖音是 ``douyin``,而不是平台 id ``dy``)——
|
||
和 ingest 找文件用的是同一个来源,两边不会走散。
|
||
"""
|
||
jsonl_dir = out_dir / adapters.artifact_dir(platform) / "jsonl"
|
||
jsonl_dir.mkdir(parents=True, exist_ok=True)
|
||
|
||
kind = "creator" if mode == MODE_CREATOR else "detail"
|
||
date = datetime.now().strftime("%Y-%m-%d")
|
||
|
||
_write_jsonl(jsonl_dir / f"{kind}_contents_{date}.jsonl", notes)
|
||
# 评论文件即使没有评论也建出来:ingest 靠「文件在不在」区分「这一轮没评论」和
|
||
# 「这一轮什么都没抓到」,两种情况的含义完全不同。
|
||
_write_jsonl(jsonl_dir / f"{kind}_comments_{date}.jsonl", comments)
|
||
# 博主资料同样无条件写:空文件表示"问了但没问到",没有文件表示"这次根本没问"
|
||
# (作品模式)。两者在 ingest 那边走的是同一条路(都不落快照),但留空文件能让
|
||
# 事后翻 run 目录时看出到底问没问过。
|
||
_write_jsonl(jsonl_dir / f"{kind}_profile_{date}.jsonl", list(profiles))
|
||
return jsonl_dir
|
||
|
||
|
||
def _write_jsonl(path: Path, records: List[Dict[str, Any]]) -> None:
|
||
with path.open("w", encoding="utf-8") as handle:
|
||
for record in records:
|
||
handle.write(json.dumps(record, ensure_ascii=False) + "\n")
|
||
utils.logger.info(f"[douyin_fetch] 写出 {len(records)} 条 -> {path.name}")
|