在上游 MediaCrawler 之上新增一层: - 监控层 api/monitor/ —— 多博主/多笔记的定时采集、指标快照差分、报表、 企业微信通知。每轮采集写入独立目录,差分才成立。 - WebUI 登录鉴权 api/auth.py —— PBKDF2 口令 + 服务端会话,/api 全接口防护。 WebSocket 单独加依赖:BaseHTTPMiddleware 对 ws 作用域直接放行,覆盖不到。 - 全局平台切换 + 能力矩阵 —— 如实区分「爬虫模块支持」与「监控层已接线」, 未接通的平台直接拒绝建任务,而不是静默跑空。 - 监控库改用 MySQL 5.7(可回退 SQLite 供测试):逐表强制 utf8mb4 (服务端与库默认都是 latin1),启动校验所连 schema 以防写错库, 连接池 recycle + pre_ping 应对 MySQL 的 8 小时空闲断连。 修复上游缺陷: - xhs/core.py: 主页抓取失败会跳掉整个博主,导致一条作品都抓不到, 而那份资料只喂给一个空函数。改为尽力而为,失败不中断。 - xhs/login.py: cookie 登录只注入 web_session,冷启动签名会失败。 新增 INJECT_ALL_COOKIES 开关(默认关闭,原有行为不变)。 - requirements.txt: 补上 websockets。它在上游 pyproject.toml 里有声明、 这里漏了,导致 uvicorn 没有 WebSocket 能力,实时日志流从未工作。 改动过的上游文件清单及合并方式见 UPSTREAM.md。 测试:492 passed(另有 1 个既有的 Windows/gbk 上游测试失败,与本改动无关)
159 lines
6.9 KiB
Python
159 lines
6.9 KiB
Python
# -*- coding: utf-8 -*-
|
||
# Copyright (c) 2025 [email protected]
|
||
#
|
||
# This file is part of MediaCrawler project.
|
||
# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py
|
||
# GitHub: https://github.com/NanmiCoder
|
||
# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1
|
||
#
|
||
|
||
# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则:
|
||
# 1. 不得用于任何商业用途。
|
||
# 2. 使用时应遵守目标平台的使用条款和robots.txt规则。
|
||
# 3. 不得进行大规模爬取或对平台造成运营干扰。
|
||
# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。
|
||
# 5. 不得用于任何非法或不当的用途。
|
||
#
|
||
# 详细许可条款请参阅项目根目录下的LICENSE文件。
|
||
# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。
|
||
|
||
# Basic configuration
|
||
PLATFORM = "xhs" # Platform, xhs | dy | ks | bili | wb | tieba | zhihu
|
||
|
||
# 是否使用海外版小红书 (rednote.com)
|
||
# 开启后 API 走 webapi.rednote.com,cookie 域使用 .rednote.com
|
||
XHS_INTERNATIONAL = False
|
||
|
||
KEYWORDS = "编程副业,编程兼职" # Keyword search configuration, separated by English commas
|
||
LOGIN_TYPE = "qrcode" # qrcode or phone or cookie
|
||
COOKIES = ""
|
||
CRAWLER_TYPE = (
|
||
"search" # Crawling type, search (keyword search) | detail (post details) | creator (creator homepage data)
|
||
)
|
||
# Whether to enable IP proxy
|
||
ENABLE_IP_PROXY = False
|
||
|
||
# Number of proxy IP pools
|
||
IP_PROXY_POOL_COUNT = 2
|
||
|
||
# Proxy IP provider name
|
||
IP_PROXY_PROVIDER_NAME = "kuaidaili" # kuaidaili | wandouhttp | static
|
||
|
||
# Static proxy configuration (used when IP_PROXY_PROVIDER_NAME is set to "static")
|
||
# Format: "http://your_home_domain:port" or "http://user:password@your_home_domain:port"
|
||
STATIC_PROXY_URL = ""
|
||
|
||
# Setting to True will not open the browser (headless browser)
|
||
# Setting False will open a browser
|
||
# If Xiaohongshu keeps scanning the code to log in but fails, open the browser and manually pass the sliding verification code.
|
||
# If Douyin keeps prompting failure, open the browser and see if mobile phone number verification appears after scanning the QR code to log in. If it does, manually go through it and try again.
|
||
HEADLESS = False
|
||
|
||
# Whether to save login status
|
||
SAVE_LOGIN_STATE = True
|
||
|
||
# 是否注入完整 cookie(默认 False,保持上游原有行为)。
|
||
# False 时 login_by_cookies 只写入 web_session;a1 / webId 等签名所需 cookie 只能靠
|
||
# browser_data 下的持久化 profile 补齐。无人值守场景(服务器上跑定时监控)应设为 True,
|
||
# 否则冷 profile 下 API 签名失败,且表现为「退出码 0 但抓到 0 条」的静默失败。
|
||
INJECT_ALL_COOKIES = False
|
||
|
||
# ==================== CDP (Chrome DevTools Protocol) 配置 ====================
|
||
# 是否启用 CDP 模式 - 使用用户本地的 Chrome/Edge 浏览器进行爬取,具有更好的反检测能力
|
||
# 开启后,会自动检测并启动用户的 Chrome/Edge 浏览器,通过 CDP 协议进行控制
|
||
# 该方式使用真实浏览器环境,包括用户的扩展、Cookie 和设置,大幅降低被风控检测的风险
|
||
ENABLE_CDP_MODE = True
|
||
|
||
# CDP 调试端口,用于与浏览器通信
|
||
# 如果端口被占用,系统会自动尝试下一个可用端口
|
||
CDP_DEBUG_PORT = 9222
|
||
|
||
# 自定义浏览器路径(可选)
|
||
# 如果为空,系统会自动检测 Chrome/Edge 的安装路径
|
||
# Windows 示例: "C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe"
|
||
# macOS 示例: "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome"
|
||
CUSTOM_BROWSER_PATH = ""
|
||
|
||
# 是否在 CDP 模式下启用无头模式
|
||
# 注意:即使设置为 True,某些反检测功能在无头模式下可能无法正常工作
|
||
CDP_HEADLESS = False
|
||
|
||
# 浏览器启动超时时间(秒)
|
||
BROWSER_LAUNCH_TIMEOUT = 60
|
||
|
||
# 是否连接用户已打开的浏览器,而不是启动新的浏览器
|
||
# 开启后,程序会连接一个已经启用了远程调试的浏览器
|
||
# 用户需要在 Chrome 中开启远程调试:chrome://inspect/#remote-debugging
|
||
# 或者使用命令行参数启动 Chrome:--remote-debugging-port=9222
|
||
# 这种方式反检测效果最好,因为直接使用用户真实浏览器的所有 Cookie、扩展和浏览历史
|
||
CDP_CONNECT_EXISTING = True
|
||
|
||
# 程序结束时是否自动关闭浏览器
|
||
# 设置为 False 可以保持浏览器运行,方便调试
|
||
AUTO_CLOSE_BROWSER = True
|
||
|
||
# Data saving type option configuration, supports: csv, db, json, jsonl, sqlite, excel, postgres. It is best to save to DB, with deduplication function.
|
||
SAVE_DATA_OPTION = "jsonl" # csv or db or json or jsonl or sqlite or excel or postgres
|
||
|
||
# Data saving path, if not specified by default, it will be saved to the data folder.
|
||
SAVE_DATA_PATH = ""
|
||
|
||
# Browser file configuration cached by the user's browser
|
||
USER_DATA_DIR = "%s_user_data_dir" # %s will be replaced by platform name
|
||
|
||
# The number of pages to start crawling starts from the first page by default
|
||
START_PAGE = 1
|
||
|
||
# Control the number of crawled videos/posts
|
||
CRAWLER_MAX_NOTES_COUNT = 15
|
||
|
||
# Controlling the number of concurrent crawlers
|
||
MAX_CONCURRENCY_NUM = 1
|
||
|
||
# 是否启用媒体下载(封面、视频,以及图文帖的图片),默认关闭。
|
||
# 开启后媒体文件按 {SAVE_DATA_PATH 或 data}/{platform}/media/{内容ID}/ 目录聚合存放。
|
||
# 支持的平台:xhs / dy / ks / bili / wb(tieba、zhihu 的数据结构中没有媒体字段,不支持)。
|
||
# 命令行开关:--get_media
|
||
ENABLE_GET_MEDIA = False
|
||
|
||
# Whether to enable comment crawling mode. Comment crawling is enabled by default.
|
||
ENABLE_GET_COMMENTS = True
|
||
|
||
# Control the number of crawled first-level comments (single video/post)
|
||
CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES = 10
|
||
|
||
# Whether to enable the mode of crawling second-level comments. By default, crawling of second-level comments is not enabled.
|
||
# If the old version of the project uses db, you need to refer to schema/tables.sql line 287 to add table fields.
|
||
ENABLE_GET_SUB_COMMENTS = False
|
||
|
||
# word cloud related
|
||
# Whether to enable generating comment word clouds
|
||
ENABLE_GET_WORDCLOUD = False
|
||
# Custom words and their groups
|
||
# Add rule: xx:yy where xx is a custom-added phrase, and yy is the group name to which the phrase xx is assigned.
|
||
CUSTOM_WORDS = {
|
||
"零几": "年份", # Recognize "zero points" as a whole
|
||
"高频词": "专业术语", # Example custom words
|
||
}
|
||
|
||
# Deactivate (disabled) word file path
|
||
STOP_WORDS_FILE = "./docs/hit_stopwords.txt"
|
||
|
||
# Chinese font file path
|
||
FONT_PATH = "./docs/STZHONGS.TTF"
|
||
|
||
# Crawl interval
|
||
CRAWLER_MAX_SLEEP_SEC = 2
|
||
|
||
# 是否禁用 SSL 证书验证。仅在使用企业代理、Burp Suite、mitmproxy 等会注入自签名证书的中间人代理时设为 True。
|
||
# 警告:禁用 SSL 验证将使所有流量暴露于中间人攻击风险,请勿在生产环境中开启。
|
||
DISABLE_SSL_VERIFY = False
|
||
|
||
from .bilibili_config import *
|
||
from .xhs_config import *
|
||
from .dy_config import *
|
||
from .ks_config import *
|
||
from .weibo_config import *
|
||
from .tieba_config import *
|
||
from .zhihu_config import *
|