用户报「一直在运行中」。库里那条 run 是真的卡住了:日志里一条采集输出都没有,
说明它卡在 collect 里、还没走到任何日志。探针定位到:
标签页: [''] ← 一个 URL 为空的标签页,渲染进程已卡死
context.cookies(): OK, 90 个 ← cookie 读得到
page.evaluate('navigator.userAgent'): **永远不返回**
而问浏览器要身份(UA + client hints)是采集的**第一步**,`page.evaluate` 又**没设超时** ——
于是整轮挂在那儿,run 永远停在「运行中」。
三处修复,各挡一层:
1. `page.evaluate` / `context.cookies()` 全部加超时(8 秒)。卡住就跳过,不再无限等。
2. 不假设第一个标签页是好的:逐个试、优先抖音页;全都不行就临时开一个干净页问完关掉。
拿不到就退回库里那份 cookie —— **不编造指纹**,那比没有更糟。
3. **进程内那条路补上整体超时**:爬虫那条靠 `run_and_wait(timeout=...)` 兜底,这条路
没有子进程、没人管,里面任何一次卡住都会变成永久的「运行中」。
测试 +6:卡死的页会被跳过(真 sleep,验的正是超时)、没 UA 的页跳过、全不行时开临时页
并关掉它、优先抖音页;以及整轮卡住时 run 不会停在 running(含超时原因)。
521 lines
20 KiB
Python
521 lines
20 KiB
Python
# -*- coding: utf-8 -*-
|
||
# Copyright (c) 2025 [email protected]
|
||
#
|
||
# This file is part of MediaCrawler project.
|
||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_api.py
|
||
# GitHub: https://github.com/NanmiCoder
|
||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||
#
|
||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||
# 1. 不得用于任何商业用途。
|
||
# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。
|
||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||
# 5. 不得用于任何非法或不当的用途。
|
||
#
|
||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||
|
||
"""抖音 Web 接口客户端 —— 直接发 HTTP,不起爬虫子进程。
|
||
|
||
**为什么另起一套。** 爬虫那条路(``media_platform/douyin``)会构造一大串浏览器指纹
|
||
参数:``browser_platform=MacIntel``、``os_name=Mac OS``、``browser_version=125.0.0.0``……
|
||
而 ``User-Agent`` 是从页面现读的(在服务器上是 Linux + Chrome 155)。参数说自己是 Mac,
|
||
UA 说自己是 Linux —— 抖音网关对这种自相矛盾的请求的处理方式是:**不报错、不给原因,
|
||
回一个 200 + 空 body**。爬虫那边把它翻译成 ``Exception("account blocked")``,看起来像
|
||
账号被封,其实什么都不是。
|
||
|
||
这份客户端只发必要参数(``device_platform`` / ``aid`` 那两三个),走浏览器自己也在用的
|
||
那条调用路径。它的做法来自 mac-agent-os 项目的 ``mediacrawler_adapter.py``,实测可用。
|
||
|
||
两个关键点:
|
||
|
||
* **cookie 走 CDP 现读。** Chrome 把 cookie 值加密存在 SQLite 里,只有 CDP 拿得到
|
||
解密后的值;而且浏览器里那份比库里存的旧快照新 —— 站点会自己轮换会话。
|
||
* **产物形状照抄 store。** ``aweme_id`` / ``aweme_url`` / ``cover_url`` / ``aweme_type`` /
|
||
``create_time``(**秒**,由 adapters 换算成毫秒)…… 这样 ingest 那条链路一个字都不用改。
|
||
"""
|
||
|
||
import asyncio
|
||
import os
|
||
import time
|
||
from dataclasses import dataclass
|
||
from typing import Any, Dict, List, Optional, Tuple
|
||
|
||
import config
|
||
import httpx
|
||
from tools import utils
|
||
from tools.user_hash import anonymize_user_id
|
||
|
||
# 请求头。**要像一个浏览器**,而且必须是**同一个浏览器**:见 BrowserIdentity。
|
||
_BASE_HEADERS = {
|
||
"Accept": "application/json, text/plain, */*",
|
||
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
|
||
"Referer": "https://www.douyin.com/",
|
||
"Origin": "https://www.douyin.com",
|
||
}
|
||
|
||
# 网关的业务前置校验头。缺了它,抖音边缘网关的 ArgusSecurityPlugin 会直接回
|
||
# 403 并写明 "Blocked by ArgusSecurityPlugin Uifid Not Found" —— 难得一次它会说原因。
|
||
# 当前网关并不校验这个头的**值**,填什么都行;一旦升级到真校验,就得改成让页面里的
|
||
# SDK 自己生成(见 media_platform/douyin/client.py 里同一条注释)。
|
||
ARGUS_HEADER_VALUE = "1"
|
||
|
||
API_ORIGIN = "https://www.douyin.com"
|
||
PROFILE_PATH = "/aweme/v1/web/user/profile/other/"
|
||
POSTS_PATH = "/aweme/v1/web/aweme/post/"
|
||
DETAIL_PATH = "/aweme/v1/web/aweme/detail/"
|
||
COMMENT_PATH = "/aweme/v1/web/comment/list/"
|
||
|
||
# 一次请求的超时。抖音这两个接口正常都在一秒内返回。
|
||
REQUEST_TIMEOUT_SECONDS = 20.0
|
||
|
||
# 问浏览器要 UA / client hints 的超时。**这个必须有。**
|
||
# ``page.evaluate`` 打在一个渲染进程已经卡住的标签页上会**永远不返回**,而问身份是采集的
|
||
# 第一步 —— 它一挂,整个 run 就永远停在「运行中」(真踩过:标签页 URL 是空的,
|
||
# cookies() 正常,evaluate 一直不回来)。
|
||
EVALUATE_TIMEOUT_SECONDS = 8.0
|
||
# 单页最多要多少条。接口自己有上限,要多了也没用。
|
||
MAX_PAGE_SIZE = 20
|
||
|
||
|
||
class DouyinApiError(RuntimeError):
|
||
"""请求失败,或登录态不可用。"""
|
||
|
||
|
||
def _cdp_url() -> str:
|
||
"""浏览器 DevTools 端点。与扫码登录那边共用同一个开关。"""
|
||
return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}"
|
||
|
||
|
||
@dataclass
|
||
class BrowserIdentity:
|
||
"""一个请求要像浏览器所需要的全部身份信息,**且必须来自同一个浏览器**。
|
||
|
||
只拿 cookie 是不够的。UA 声称自己是 Chrome 155、却不带 Chrome 155 该有的
|
||
``sec-ch-ua``,网关一眼就能看出这不是浏览器 —— 它的回应是 **200 + 空 body**:
|
||
不报错、不给原因,只看得到「抓到 0 条」。所以这三样必须成套地从同一处取。
|
||
"""
|
||
|
||
cookie: str
|
||
user_agent: str
|
||
client_hints: Dict[str, str]
|
||
|
||
def headers(self) -> Dict[str, str]:
|
||
headers = {
|
||
"User-Agent": self.user_agent,
|
||
**self.client_hints,
|
||
**_BASE_HEADERS,
|
||
"x-tt-argus": ARGUS_HEADER_VALUE,
|
||
"Cookie": self.cookie,
|
||
}
|
||
# uifid 是设备标识,网关要它;cookie 里没有就不带(送空值反而更像异常请求)。
|
||
uifid = _cookie_value(self.cookie, "UIFID") or _cookie_value(
|
||
self.cookie, "UIFID_TEMP"
|
||
)
|
||
if uifid:
|
||
headers["uifid"] = uifid
|
||
return headers
|
||
|
||
|
||
# 身份信息的短时缓存:一次采集要发好几个请求,没必要每次都连一遍 CDP。
|
||
_IDENTITY_TTL_SECONDS = 120.0
|
||
_identity_cache: Optional[Tuple[float, BrowserIdentity]] = None
|
||
|
||
|
||
async def _safe_evaluate(page: Any, expression: str) -> Any:
|
||
"""在页面上求值,带超时;任何失败都返回 None。
|
||
|
||
**不要直接调 ``page.evaluate``** —— 在渲染进程卡住的标签页上它会永远不返回(见
|
||
``EVALUATE_TIMEOUT_SECONDS`` 那段)。
|
||
"""
|
||
try:
|
||
return await asyncio.wait_for(
|
||
page.evaluate(expression), timeout=EVALUATE_TIMEOUT_SECONDS
|
||
)
|
||
except Exception:
|
||
return None
|
||
|
||
|
||
async def _identity_from_pages(context: Any) -> Tuple[str, Dict[str, str]]:
|
||
"""问出 UA 和 client hints。
|
||
|
||
不假设第一个标签页是好的 —— 它可能停在 URL 为空、渲染进程已卡住的状态(实测过)。
|
||
所以逐个试、每个都带超时;优先抖音页面,全都不行就临时开一个干净页问完关掉。
|
||
|
||
拿不到就返回空 —— 调用方据此退回库里那份 cookie,而不是拿一组编出来的指纹去请求
|
||
(那比没有更糟,见 BrowserIdentity 的说明)。
|
||
"""
|
||
from media_platform.douyin.help import client_hint_headers
|
||
|
||
pages = list(context.pages)
|
||
pages.sort(key=lambda page: 0 if "douyin" in (page.url or "") else 1)
|
||
for page in pages:
|
||
user_agent = await _safe_evaluate(page, "() => navigator.userAgent")
|
||
if user_agent:
|
||
hints = client_hint_headers(
|
||
await _safe_evaluate(page, "() => navigator.userAgentData || null")
|
||
)
|
||
return user_agent, hints or {}
|
||
|
||
temp = None
|
||
try:
|
||
temp = await asyncio.wait_for(
|
||
context.new_page(), timeout=EVALUATE_TIMEOUT_SECONDS
|
||
)
|
||
user_agent = await _safe_evaluate(temp, "() => navigator.userAgent")
|
||
hints = client_hint_headers(
|
||
await _safe_evaluate(temp, "() => navigator.userAgentData || null")
|
||
)
|
||
return user_agent or "", hints or {}
|
||
except Exception:
|
||
return "", {}
|
||
finally:
|
||
if temp is not None:
|
||
try:
|
||
await temp.close()
|
||
except Exception:
|
||
pass
|
||
|
||
|
||
async def _read_browser() -> Optional[BrowserIdentity]:
|
||
"""连上 CDP 浏览器,一次取齐 cookie、UA、client hints。
|
||
|
||
读不到返回 None(浏览器没开/没登录),由调用方决定怎么报 —— 不抛异常。
|
||
"""
|
||
from playwright.async_api import async_playwright
|
||
|
||
from media_platform.douyin.help import client_hint_headers
|
||
|
||
playwright = None
|
||
try:
|
||
playwright = await async_playwright().start()
|
||
browser = await playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000)
|
||
if not browser.contexts:
|
||
return None
|
||
# contexts[0] 是真实 profile。**不要 new_context()** —— 那是无痕式的,读不到登录态。
|
||
context = browser.contexts[0]
|
||
cookies = await asyncio.wait_for(
|
||
context.cookies(), timeout=EVALUATE_TIMEOUT_SECONDS
|
||
)
|
||
|
||
# UA 和 hints 要从页面里问 —— 它们是浏览器自己的事实,写死迟早对不上。
|
||
user_agent, hints = await _identity_from_pages(context)
|
||
except Exception as exc:
|
||
utils.logger.warning(f"[douyin_api] 读浏览器身份失败:{exc}")
|
||
return None
|
||
finally:
|
||
if playwright is not None:
|
||
# 只断开连接。**绝不能 browser.close()** —— 对这个 CDP 连接而言那会关掉
|
||
# 操作者自己的浏览器。
|
||
try:
|
||
await playwright.stop()
|
||
except Exception:
|
||
pass
|
||
|
||
douyin_cookies = {
|
||
cookie["name"]: cookie["value"]
|
||
for cookie in cookies
|
||
if "douyin" in cookie.get("domain", "") or "amemv" in cookie.get("domain", "")
|
||
}
|
||
return BrowserIdentity(
|
||
cookie=_cookie_from_dict(douyin_cookies),
|
||
user_agent=user_agent or "",
|
||
client_hints=hints,
|
||
)
|
||
|
||
|
||
async def browser_identity(cookie: str = "", force: bool = False) -> BrowserIdentity:
|
||
"""拿到一份可用的身份:**优先浏览器里那份**,其次退回传进来的 cookie(库里存的)。
|
||
|
||
优先浏览器的原因:站点会自己轮换会话,库里存的是粘贴那一刻的快照,浏览器里那份才是
|
||
当前有效的;而 UA/hints 更是只有浏览器自己知道。
|
||
"""
|
||
global _identity_cache
|
||
|
||
now = time.monotonic()
|
||
if not force and _identity_cache is not None:
|
||
cached_at, cached = _identity_cache
|
||
if now - cached_at < _IDENTITY_TTL_SECONDS:
|
||
return cached
|
||
|
||
identity = await _read_browser()
|
||
if identity is None or not _has_session(identity.cookie):
|
||
# 浏览器里没有可用会话,退回调用方给的那份。UA/hints 编不出来就不编 ——
|
||
# 一组和 UA 对不上的 hints 比没有更糟。
|
||
identity = BrowserIdentity(
|
||
cookie=_cookie_header(cookie), user_agent="", client_hints={}
|
||
)
|
||
_identity_cache = (now, identity)
|
||
return identity
|
||
|
||
|
||
def forget_identity() -> None:
|
||
"""丢掉缓存的身份。cookie 变了、或测试之间要隔离时调用。"""
|
||
global _identity_cache
|
||
_identity_cache = None
|
||
|
||
|
||
def _cookie_header(cookie: str) -> str:
|
||
"""把 ``a=1; b=2`` 形式的 cookie 串规整成请求头用的形状。"""
|
||
pairs = []
|
||
for part in (cookie or "").split(";"):
|
||
if "=" in part:
|
||
name, _, value = part.partition("=")
|
||
name = name.strip()
|
||
if name:
|
||
pairs.append(f"{name}={value.strip()}")
|
||
return "; ".join(pairs)
|
||
|
||
|
||
def _cookie_from_dict(cookies: Dict[str, str]) -> str:
|
||
return "; ".join(f"{name}={value}" for name, value in cookies.items())
|
||
|
||
|
||
def _cookie_value(cookie: str, name: str) -> str:
|
||
"""从一个 cookie 串里取某个键的值。"""
|
||
for part in (cookie or "").split(";"):
|
||
key, _, value = part.partition("=")
|
||
if key.strip() == name:
|
||
return value.strip()
|
||
return ""
|
||
|
||
|
||
async def _get(
|
||
path: str, params: Dict[str, Any], identity: BrowserIdentity
|
||
) -> Dict[str, Any]:
|
||
"""发一个 GET,返回 JSON。
|
||
|
||
只带调用方给的参数 —— **不要往里加 webid / msToken / browser_version 那一堆**,
|
||
那正是爬虫那条路失败的原因。
|
||
"""
|
||
url = f"{API_ORIGIN}{path}"
|
||
async with httpx.AsyncClient(timeout=REQUEST_TIMEOUT_SECONDS) as client:
|
||
response = await client.get(
|
||
url,
|
||
params=params,
|
||
headers=identity.headers(),
|
||
)
|
||
|
||
if response.status_code != 200:
|
||
raise DouyinApiError(f"HTTP {response.status_code}:{response.text[:120]}")
|
||
|
||
# 「200 + 空 body」是抖音网关拒绝请求时的典型回应(见模块说明)。必须当成错误报出来,
|
||
# 否则会一路往下变成「这个博主没作品」。
|
||
if not response.text.strip():
|
||
raise DouyinApiError(
|
||
"接口返回了空内容 —— 通常是登录态失效,或请求被网关判成了非浏览器"
|
||
)
|
||
|
||
try:
|
||
return response.json()
|
||
except ValueError as exc:
|
||
raise DouyinApiError(f"返回的不是 JSON:{response.text[:120]}") from exc
|
||
|
||
|
||
def _as_int(value: Any) -> int:
|
||
try:
|
||
return int(value)
|
||
except (TypeError, ValueError):
|
||
return 0
|
||
|
||
|
||
def normalize_aweme(aweme: Dict[str, Any]) -> Dict[str, Any]:
|
||
"""把接口返回的一条作品,翻译成 store 落盘的那套键名。
|
||
|
||
键名必须和 ``store/douyin`` 一致 —— 跨过这一层之后,ingest 就不知道数据是从爬虫
|
||
来的还是从接口来的。
|
||
"""
|
||
author = aweme.get("author") or {}
|
||
statistics = aweme.get("statistics") or {}
|
||
aweme_id = str(aweme.get("aweme_id") or "")
|
||
cover = ((aweme.get("video") or {}).get("cover") or {}).get("url_list") or [""]
|
||
uid = str(author.get("uid") or "")
|
||
nickname = author.get("nickname") or ""
|
||
|
||
return {
|
||
"aweme_id": aweme_id,
|
||
"aweme_type": str(aweme.get("aweme_type") or ""),
|
||
# store 那边 title 取的是 desc。
|
||
"title": aweme.get("desc") or "",
|
||
"desc": aweme.get("desc") or "",
|
||
# **秒**。adapters.time_scale 会把它换成毫秒,和 store 写出来的形态一致。
|
||
"create_time": _as_int(aweme.get("create_time")),
|
||
"creator_hash": anonymize_user_id(uid or author.get("sec_uid") or ""),
|
||
"nickname": nickname,
|
||
"liked_count": str(_as_int(statistics.get("digg_count"))),
|
||
"comment_count": str(_as_int(statistics.get("comment_count"))),
|
||
"collected_count": str(_as_int(statistics.get("collect_count"))),
|
||
"share_count": str(_as_int(statistics.get("share_count"))),
|
||
"aweme_url": f"https://www.douyin.com/video/{aweme_id}",
|
||
"cover_url": cover[0] if cover else "",
|
||
"source_keyword": "",
|
||
}
|
||
|
||
|
||
async def author_videos(
|
||
sec_user_id: str, count: int = MAX_PAGE_SIZE, *, cookie: str = ""
|
||
) -> List[Dict[str, Any]]:
|
||
"""某个博主最新发布的作品(按发布时间倒序),已翻译成 store 的键名。
|
||
|
||
用 ``sec_user_id`` 而不是数字 uid:监控任务里存的就是主页链接里的那段 sec_uid,
|
||
而且这个接口两种都收(爬虫那边用的也是 sec_user_id)。
|
||
"""
|
||
identity = await browser_identity(cookie)
|
||
if not _has_session(identity.cookie):
|
||
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
|
||
|
||
payload = await _get(
|
||
POSTS_PATH,
|
||
{
|
||
"sec_user_id": sec_user_id,
|
||
"count": max(1, min(count, MAX_PAGE_SIZE)),
|
||
"max_cursor": 0,
|
||
"device_platform": "webapp",
|
||
"aid": 6383,
|
||
},
|
||
identity,
|
||
)
|
||
|
||
awemes = payload.get("aweme_list") or []
|
||
if not awemes and payload.get("status_code") not in (0, None):
|
||
raise DouyinApiError(
|
||
f"接口拒绝了请求(status_code={payload.get('status_code')})"
|
||
)
|
||
return [normalize_aweme(aweme) for aweme in awemes]
|
||
|
||
|
||
async def video_detail(aweme_id: str, *, cookie: str = "") -> Dict[str, Any]:
|
||
"""单条作品的详情,已翻译成 store 的键名。
|
||
|
||
这个接口**没有**被那道真校验挡着(实测 200 / 45425 字节),所以在拿不到作品列表时,
|
||
它是「刷新已知作品指标」的唯一途径。
|
||
"""
|
||
identity = await browser_identity(cookie)
|
||
if not _has_session(identity.cookie):
|
||
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
|
||
|
||
payload = await _get(
|
||
DETAIL_PATH,
|
||
{"aweme_id": aweme_id, "device_platform": "webapp", "aid": 6383},
|
||
identity,
|
||
)
|
||
aweme = payload.get("aweme_detail") or {}
|
||
if not aweme:
|
||
raise DouyinApiError(
|
||
f"接口没返回作品(status_code={payload.get('status_code')})"
|
||
)
|
||
return normalize_aweme(aweme)
|
||
|
||
|
||
async def author_profile(sec_user_id: str, *, cookie: str = "") -> Dict[str, Any]:
|
||
"""博主主页指标:昵称 / 粉丝数 / 总获赞 / 作品数。"""
|
||
identity = await browser_identity(cookie)
|
||
if not _has_session(identity.cookie):
|
||
raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有")
|
||
|
||
payload = await _get(
|
||
PROFILE_PATH,
|
||
{"sec_user_id": sec_user_id, "device_platform": "webapp", "aid": 6383},
|
||
identity,
|
||
)
|
||
user = payload.get("user") or {}
|
||
if not user:
|
||
raise DouyinApiError(
|
||
f"接口没返回用户数据(status_code={payload.get('status_code')})"
|
||
)
|
||
return {
|
||
"nickname": user.get("nickname") or "",
|
||
"unique_id": user.get("unique_id") or "",
|
||
"fans": _as_int(user.get("follower_count")),
|
||
"total_favorited": _as_int(user.get("total_favorited")),
|
||
"works": _as_int(user.get("aweme_count")),
|
||
"following": _as_int(user.get("following_count")),
|
||
}
|
||
|
||
|
||
def normalize_comment(comment: Dict[str, Any], aweme_id: str) -> Dict[str, Any]:
|
||
"""把接口返回的一条评论,翻译成 store 落盘的那套键名。
|
||
|
||
HTTP 路线和页面路线共用它 —— 同一套键名,ingest 才不用关心数据是怎么来的。
|
||
刻意**不带** ``sub_comment_count`` / ``parent_comment_id`` 的猜测值:接口给了就用,
|
||
没给就留空,不编。
|
||
"""
|
||
user = comment.get("user") or {}
|
||
return {
|
||
"comment_id": str(comment.get("cid") or ""),
|
||
"aweme_id": aweme_id,
|
||
"content": comment.get("text") or "",
|
||
"nickname": user.get("nickname") or "",
|
||
"creator_hash": anonymize_user_id(
|
||
str(user.get("uid") or user.get("sec_uid") or "")
|
||
),
|
||
# 同为秒;adapters 会换算。
|
||
"create_time": _as_int(comment.get("create_time")),
|
||
"like_count": str(_as_int(comment.get("digg_count"))),
|
||
"sub_comment_count": str(_as_int(comment.get("reply_comment_total"))),
|
||
# 顶层评论在抖音里是 "0";adapters.parent_comment_id 会归一成空串。
|
||
"parent_comment_id": str(comment.get("reply_id") or "0"),
|
||
}
|
||
|
||
|
||
async def video_comments(
|
||
aweme_id: str, count: int = 20, *, cookie: str = ""
|
||
) -> List[Dict[str, Any]]:
|
||
"""一条作品的评论,翻译成 store 的评论键名。
|
||
|
||
刻意不带 ``sub_comment_count`` / ``parent_comment_id`` 的猜测值 —— 接口给了就用,
|
||
没给就留空,不编。
|
||
"""
|
||
identity = await browser_identity(cookie)
|
||
if not _has_session(identity.cookie):
|
||
raise DouyinApiError("抖音登录态不可用")
|
||
|
||
payload = await _get(
|
||
COMMENT_PATH,
|
||
{
|
||
"aweme_id": aweme_id,
|
||
"count": max(1, min(count, MAX_PAGE_SIZE)),
|
||
"cursor": 0,
|
||
"device_platform": "webapp",
|
||
"aid": 6383,
|
||
},
|
||
identity,
|
||
)
|
||
|
||
records = [
|
||
normalize_comment(comment, aweme_id)
|
||
for comment in payload.get("comments") or []
|
||
]
|
||
return records
|
||
|
||
|
||
def _has_session(cookie: str) -> bool:
|
||
return "sessionid=" in (cookie or "")
|
||
|
||
|
||
async def check_login(cookie: str = "") -> Dict[str, Any]:
|
||
"""浏览器/库里现在有没有可用的抖音登录态。给设置页用。"""
|
||
identity = await browser_identity(cookie)
|
||
if _has_session(identity.cookie):
|
||
source = "browser" if identity.user_agent else "stored"
|
||
return {"ok": True, "source": source, "cookie_length": len(identity.cookie)}
|
||
return {"ok": False, "source": "", "cookie_length": 0}
|
||
|
||
|
||
async def main() -> None: # pragma: no cover - 手工排查用
|
||
"""``python -m api.monitor.douyin_api <sec_user_id>``"""
|
||
import sys
|
||
|
||
if len(sys.argv) < 2:
|
||
print(await check_login())
|
||
return
|
||
sec = sys.argv[1]
|
||
print(await author_profile(sec))
|
||
for record in await author_videos(sec, count=5):
|
||
print(record["create_time"], record["title"][:30], record["liked_count"])
|
||
|
||
|
||
if __name__ == "__main__": # pragma: no cover
|
||
asyncio.run(main())
|