先纠正一个我上一轮给错的结论:日志里的 `account blocked` **不是抖音说的**,是
MediaCrawler 自己编的:
if response.text == "" or response.text == "blocked":
raise Exception("account blocked")
真实情况只是**抖音返回了空 body**。我把它读成了「账号被风控」,还写进了运行历史和
给用户的结论里 —— 用户质疑「我网页版和手机版都能正常登录」,一查,他是对的。
实测定位(同一 URL、同一 cookie、同一参数):
浏览器页面内 fetch : 200, 7077 字节 ✓
httpx : 200, 0 字节 ✗
带 a_bogus : 0 字节
不带 a_bogus : 0 字节
四种 msToken 变体 : 全部 200 有数据(所以不是它)
用浏览器那次的完整头重放 httpx : 200, 7077 字节 ✓
浏览器那次请求比爬虫多的,只有这三个头:
sec-ch-ua: "Google Chrome";v="155", "Chromium";v="155", "Not(A:Brand";v=24
sec-ch-ua-mobile: ?0
sec-ch-ua-platform: "Linux"
爬虫的 UA 是从页面读的(声称是 Chrome 155)却不带 sec-ch-ua —— 「Chrome 的 UA +
没有 sec-ch-ua」是最典型的机器人特征。网关的回应方式是不报错、不给原因,回一个
200 + 空 body,HTTP 状态还写在成功那一栏。
修:media_platform/douyin/help.py 新增 client_hint_headers(),从
navigator.userAgentData 现算这三个头(现算而不是写死 —— 写死的版本号一旦和 UA 里的
对不上,就又是一个可疑特征);core.py 建客户端时带上。
诚实说明:我无法解释**为什么之前能跑**(run 34 还是成功的,40 分钟后同样的代码就
不行了)。最可能是字节那边收紧了这道校验,但我没有证据,别当结论。
测试 +8:还原出的头与真实浏览器抓到的值逐字一致;拿不到 userAgentData 时返回空而不
凭空编造(编一组和 UA 对不上的头比不带头更糟);mobile 标志;platform 缺失时仍发另两个。
231 lines
8.2 KiB
Python
231 lines
8.2 KiB
Python
# -*- coding: utf-8 -*-
|
||
# Copyright (c) 2025 [email protected]
|
||
#
|
||
# This file is part of MediaCrawler project.
|
||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/help.py
|
||
# GitHub: https://github.com/NanmiCoder
|
||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||
#
|
||
|
||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||
# 1. 不得用于任何商业用途。
|
||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||
# 5. 不得用于任何非法或不当的用途。
|
||
#
|
||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||
|
||
|
||
# -*- coding: utf-8 -*-
|
||
# @Author : [email protected]
|
||
# @Name: Programmer Ajiang-Relakkes
|
||
# @Time : 2024/6/10 02:24
|
||
# @Desc : Get a_bogus parameter, for learning and communication only, do not use for commercial purposes, contact author to delete if infringement
|
||
|
||
import random
|
||
import re
|
||
from typing import Dict, Optional
|
||
|
||
import execjs
|
||
from playwright.async_api import Page
|
||
|
||
from model.m_douyin import VideoUrlInfo, CreatorUrlInfo
|
||
from tools.crawler_util import extract_url_params_to_dict
|
||
|
||
douyin_sign_obj = execjs.compile(open('libs/douyin.js', encoding='utf-8-sig').read())
|
||
|
||
def get_web_id():
|
||
"""
|
||
Generate random webid
|
||
Returns:
|
||
|
||
"""
|
||
|
||
def e(t):
|
||
if t is not None:
|
||
return str(t ^ (int(16 * random.random()) >> (t // 4)))
|
||
else:
|
||
return ''.join(
|
||
[str(int(1e7)), '-', str(int(1e3)), '-', str(int(4e3)), '-', str(int(8e3)), '-', str(int(1e11))]
|
||
)
|
||
|
||
web_id = ''.join(
|
||
e(int(x)) if x in '018' else x for x in e(None)
|
||
)
|
||
return web_id.replace('-', '')[:19]
|
||
|
||
|
||
|
||
async def get_a_bogus(url: str, params: str, post_data: dict, user_agent: str, page: Page = None):
|
||
"""
|
||
Get a_bogus parameter, currently does not support POST request type signature
|
||
"""
|
||
return get_a_bogus_from_js(url, params, user_agent)
|
||
|
||
def get_a_bogus_from_js(url: str, params: str, user_agent: str):
|
||
"""
|
||
Get a_bogus parameter through js
|
||
Args:
|
||
url:
|
||
params:
|
||
user_agent:
|
||
|
||
Returns:
|
||
|
||
"""
|
||
sign_js_name = "sign_datail"
|
||
if "/reply" in url:
|
||
sign_js_name = "sign_reply"
|
||
return douyin_sign_obj.call(sign_js_name, params, user_agent)
|
||
|
||
|
||
|
||
async def get_a_bogus_from_playwright(params: str, post_data: dict, user_agent: str, page: Page):
|
||
"""
|
||
Get a_bogus parameter through playwright
|
||
playwright version is deprecated
|
||
Returns:
|
||
|
||
"""
|
||
if not post_data:
|
||
post_data = ""
|
||
a_bogus = await page.evaluate(
|
||
"([params, post_data, ua]) => window.bdms.init._v[2].p[42].apply(null, [0, 1, 8, params, post_data, ua])",
|
||
[params, post_data, user_agent])
|
||
|
||
return a_bogus
|
||
|
||
|
||
def client_hint_headers(user_agent_data) -> Dict[str, str]:
|
||
"""由 ``navigator.userAgentData`` 还原 ``sec-ch-ua`` 系列请求头。
|
||
|
||
浏览器只要声称自己是 Chrome,就**一定会**带这三个头。缺了它们,「Chrome 的 UA +
|
||
没有 sec-ch-ua」就是最典型的机器人特征 —— 抖音网关会因此返回 **200 + 空 body**:
|
||
不报错、不给原因、HTTP 状态还是成功的,表现为采集拿到 0 条。
|
||
|
||
实测(同一 URL、同一 cookie、同一参数):不带头 → 0 字节;补上这三个头 → 7077 字节。
|
||
|
||
从 ``userAgentData`` 现算而不是写死,是为了 Chrome 升级后不会悄悄失配 —— 写死的
|
||
版本号和 UA 里的版本号一旦对不上,就又是一个可疑特征。
|
||
"""
|
||
if not isinstance(user_agent_data, dict):
|
||
return {}
|
||
|
||
brands = user_agent_data.get("brands") or []
|
||
sec_ch_ua = ", ".join(
|
||
f'"{brand.get("brand", "")}";v="{brand.get("version", "")}"' for brand in brands
|
||
)
|
||
if not sec_ch_ua:
|
||
return {}
|
||
|
||
headers = {
|
||
"sec-ch-ua": sec_ch_ua,
|
||
"sec-ch-ua-mobile": "?1" if user_agent_data.get("mobile") else "?0",
|
||
}
|
||
platform = user_agent_data.get("platform")
|
||
if platform:
|
||
headers["sec-ch-ua-platform"] = f'"{platform}"'
|
||
return headers
|
||
|
||
|
||
def parse_video_info_from_url(url: str) -> VideoUrlInfo:
|
||
"""
|
||
Parse video ID from Douyin video URL
|
||
Supports the following formats:
|
||
1. Normal video link: https://www.douyin.com/video/7525082444551310602
|
||
2. Link with modal_id parameter:
|
||
- https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE?modal_id=7525082444551310602
|
||
- https://www.douyin.com/root/search/python?modal_id=7471165520058862848
|
||
3. Short link: https://v.douyin.com/iF12345ABC/ (requires client parsing)
|
||
4. Pure ID: 7525082444551310602
|
||
|
||
Args:
|
||
url: Douyin video link or ID
|
||
Returns:
|
||
VideoUrlInfo: Object containing video ID
|
||
"""
|
||
# If it's a pure numeric ID, return directly
|
||
if url.isdigit():
|
||
return VideoUrlInfo(aweme_id=url, url_type="normal")
|
||
|
||
# Check if it's a short link (v.douyin.com)
|
||
if "v.douyin.com" in url or url.startswith("http") and len(url) < 50 and "video" not in url:
|
||
return VideoUrlInfo(aweme_id="", url_type="short") # Requires client parsing
|
||
|
||
# Try to extract modal_id from URL parameters
|
||
params = extract_url_params_to_dict(url)
|
||
modal_id = params.get("modal_id")
|
||
if modal_id:
|
||
return VideoUrlInfo(aweme_id=modal_id, url_type="modal")
|
||
|
||
# Extract ID from standard video URL: /video/number
|
||
video_pattern = r'/video/(\d+)'
|
||
match = re.search(video_pattern, url)
|
||
if match:
|
||
aweme_id = match.group(1)
|
||
return VideoUrlInfo(aweme_id=aweme_id, url_type="normal")
|
||
|
||
raise ValueError(f"Unable to parse video ID from URL: {url}")
|
||
|
||
|
||
def parse_creator_info_from_url(url: str) -> CreatorUrlInfo:
|
||
"""
|
||
Parse creator ID (sec_user_id) from Douyin creator homepage URL
|
||
Supports the following formats:
|
||
1. Creator homepage: https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE?from_tab_name=main
|
||
2. Pure ID: MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE
|
||
|
||
Args:
|
||
url: Douyin creator homepage link or sec_user_id
|
||
Returns:
|
||
CreatorUrlInfo: Object containing creator ID
|
||
"""
|
||
# If it's a pure ID format (usually starts with MS4wLjABAAAA), return directly
|
||
if url.startswith("MS4wLjABAAAA") or (not url.startswith("http") and "douyin.com" not in url):
|
||
return CreatorUrlInfo(sec_user_id=url)
|
||
|
||
# Extract sec_user_id from creator homepage URL: /user/xxx
|
||
user_pattern = r'/user/([^/?]+)'
|
||
match = re.search(user_pattern, url)
|
||
if match:
|
||
sec_user_id = match.group(1)
|
||
return CreatorUrlInfo(sec_user_id=sec_user_id)
|
||
|
||
raise ValueError(f"Unable to parse creator ID from URL: {url}")
|
||
|
||
|
||
if __name__ == '__main__':
|
||
# Test video URL parsing
|
||
print("=== Video URL Parsing Test ===")
|
||
test_urls = [
|
||
"https://www.douyin.com/video/7525082444551310602",
|
||
"https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE?from_tab_name=main&modal_id=7525082444551310602",
|
||
"https://www.douyin.com/root/search/python?aid=b733a3b0-4662-4639-9a72-c2318fba9f3f&modal_id=7471165520058862848&type=general",
|
||
"7525082444551310602",
|
||
]
|
||
for url in test_urls:
|
||
try:
|
||
result = parse_video_info_from_url(url)
|
||
print(f"✓ URL: {url[:80]}...")
|
||
print(f" Result: {result}\n")
|
||
except Exception as e:
|
||
print(f"✗ URL: {url}")
|
||
print(f" Error: {e}\n")
|
||
|
||
# Test creator URL parsing
|
||
print("=== Creator URL Parsing Test ===")
|
||
test_creator_urls = [
|
||
"https://www.douyin.com/user/MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE?from_tab_name=main",
|
||
"MS4wLjABAAAATJPY7LAlaa5X-c8uNdWkvz0jUGgpw4eeXIwu_8BhvqE",
|
||
]
|
||
for url in test_creator_urls:
|
||
try:
|
||
result = parse_creator_info_from_url(url)
|
||
print(f"✓ URL: {url[:80]}...")
|
||
print(f" Result: {result}\n")
|
||
except Exception as e:
|
||
print(f"✗ URL: {url}")
|
||
print(f" Error: {e}\n")
|