butubb c2b310c7bf
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
feat(creator): 新增「运营」模块 —— 多账号扫码登录与创作者后台数据
侧边栏在「监控」右边加了「运营」:账号列表 → 点进二级详情看该账号的数据。

【为什么是独立模块而不是监控的子视图】两者形状不同:监控是公开数据(点赞/收藏/评论/分享)的每轮快照+差分;运营是创作者后台按日期给出的曝光/观看/完播率/涨粉。凭据不同、采集方式也不同 —— 那边要浏览器登录态,这边是纯请求。硬塞进同一个模型会同时污染两边。

【扫码登录的关键差异】监控的扫码把登录态写进浏览器默认 profile(爬虫要复用)。运营要的是 cookie 字符串(纯请求够用),所以每次登录开一个**临时上下文**,扫完取出 cookie 就丢弃 —— 登第二个账号不会把第一个顶掉,也不影响监控那个登录态,十个账号互不干扰。

【决策依据】tools/probe_creator_api.py 的 Phase 0 实测:签名可自造(XYW_:MD5 → base64 → AES-128-CBC,与 xhshow 内置实现常量逐字节一致);主站 cookie 即可认证创作者后台;接口与参数已与真实页面对齐。

后端:
- api/creator/models.py: creator_account / creator_note_stat。**复用 MonitorBase**,这样 create_all 与上一轮改成元数据驱动的 _ensure_columns 会自动覆盖新表
- api/creator/signing.py: XYW_ 签名,带三条实测结论(url= 前缀、appId=ugc、401 与 406 的区别)
- api/creator/client.py: 纯 httpx 客户端。字段名尚未亲眼验证过,所以写成多别名匹配;解析不出来存 None 而非 0
- api/creator/service.py: 账号 CRUD 与同步。cookie 绝不进入对外结构,只给 has_cookie
- api/creator/login.py: 临时上下文的扫码登录
- api/routers/creator.py: 8 条路由,全部带鉴权

前端:
- 侧边栏「运营」+ OperationView(账号列表 → 二级详情)+ AddAccountDialog
- 权限状态显眼呈现:pending 时照抄后台原话「已为您申请数据权限,次日可查看」,并说明此时同步返回 0 条是正常的,不是采集失败

测试:tests/test_creator_client.py 新增 48 例,含「cookie 不得出现在对外结构里」这条不变量,以及权限未生效时空壳响应的处理。
2026-10-07 16:30:45 +08:00
2026-04-03 16:07:19 +08:00
2024-09-27 14:58:10 +08:00
2026-04-07 12:54:39 +08:00
2025-09-26 18:07:57 +08:00
2025-06-01 23:20:11 +08:00
2026-10-05 22:21:13 +08:00
2026-10-05 22:21:13 +08:00
2026-10-05 22:21:13 +08:00
2025-11-18 12:24:02 +08:00
2025-11-18 12:24:02 +08:00

🔥 MediaCrawler - Social Media Platform Crawler 🕷️


Disclaimer:

Please use this repository for learning purposes only ⚠️⚠️⚠️⚠️, Web scraping illegal cases

All content in this repository is for learning and reference purposes only, and commercial use is prohibited. No person or organization may use the content of this repository for illegal purposes or infringe upon the legitimate rights and interests of others. The web scraping technology involved in this repository is only for learning and research, and may not be used for large-scale crawling of other platforms or other illegal activities. This repository assumes no legal responsibility for any legal liability arising from the use of the content of this repository. By using the content of this repository, you agree to all terms and conditions of this disclaimer.

Click to view a more detailed disclaimer. Click to jump

📖 Project Introduction

A powerful multi-platform social media data collection tool that supports crawling public information from mainstream platforms including Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu, and more.

🔧 Technical Principles

  • Core Technology: Based on Playwright browser automation framework for login and maintaining login state
  • No JS Reverse Engineering Required: Uses browser context environment with preserved login state to obtain signature parameters through JS expressions
  • Advantages: No need to reverse complex encryption algorithms, significantly lowering the technical barrier

✨ Features

Platform Keyword Search Specific Post ID Crawling Secondary Comments Specific Creator Homepage Login State Cache IP Proxy Pool Generate Comment Word Cloud
Xiaohongshu ✅ ✅ ✅ ✅ ✅ ✅ ✅
Douyin ✅ ✅ ✅ ✅ ✅ ✅ ✅
Kuaishou ✅ ✅ ✅ ✅ ✅ ✅ ✅
Bilibili ✅ ✅ ✅ ✅ ✅ ✅ ✅
Weibo ✅ ✅ ✅ ✅ ✅ ✅ ✅
Tieba ✅ ✅ ✅ ✅ ✅ ✅ ✅
Zhihu ✅ ✅ ✅ ✅ ✅ ✅ ✅

MediaCrawlerPro Major Release! Open source is not easy, welcome to subscribe and support!

Focus on learning mature project architectural design, not just crawling technology. The code design philosophy of the Pro version is equally worth in-depth study!

MediaCrawlerPro core advantages over the open-source version:

🎯 Core Feature Upgrades

  • ✅ Content Deconstruction Agent (New feature)
  • ✅ Resume crawling functionality (Key feature)
  • ✅ Multi-account + IP proxy pool support (Key feature)
  • ✅ Remove Playwright dependency, easier to use
  • ✅ Complete Linux environment support

🏗️ Architectural Design Optimization

  • ✅ Code refactoring optimization, more readable and maintainable (decoupled JS signature logic)
  • ✅ Enterprise-level code quality, suitable for building large-scale crawler projects
  • ✅ Perfect architectural design, high scalability, greater source code learning value

🎁 Additional Features

  • ✅ Social media video downloader desktop app (suitable for learning full-stack development)
  • ✅ Multi-platform homepage feed recommendations (HomeFeed)
  • AI Agent based on comment analysis is under development 🚀🚀

Click to view: MediaCrawlerPro Project Homepage for more information

🚀 Quick Start

💡 Open source is not easy, if this project helps you, please give a ⭐ Star to support!

📋 Prerequisites

Before proceeding with the next steps, please ensure that uv is installed on your computer:

  • Installation Guide: uv Official Installation Guide
  • Verify Installation: Enter the command uv --version in the terminal. If the version number is displayed normally, the installation was successful
  • Recommendation Reason: uv is currently the most powerful Python package management tool, with fast speed and accurate dependency resolution

🟢 Node.js Installation

The project depends on Node.js, please download and install from the official website:

📦 Python Package Installation

# Enter project directory
cd MediaCrawler

# Use uv sync command to ensure consistency of python version and related dependency packages
uv sync

🌐 Browser Driver Installation

# Install browser driver
uv run playwright install

💡 Tip: MediaCrawler now supports using playwright to connect to your local Chrome browser, solving some issues caused by Webdriver.

Currently, xhs and dy are available using CDP mode to connect to local browsers. If needed, check the configuration items in config/base_config.py.

🚀 Run Crawler Program

# The project does not enable comment crawling mode by default. If you need comments, please modify the ENABLE_GET_COMMENTS variable in config/base_config.py
# Other supported options can also be viewed in config/base_config.py with Chinese comments

# Read keywords from configuration file to search related posts and crawl post information and comments
uv run main.py --platform xhs --lt qrcode --type search

# Read specified post ID list from configuration file to get information and comment information of specified posts
uv run main.py --platform xhs --lt qrcode --type detail

# Open corresponding APP to scan QR code for login

# For other platform crawler usage examples, execute the following command to view
uv run main.py --help

WebUI Support

🖥️ WebUI Visual Operation Interface

MediaCrawler provides a web-based visual operation interface, allowing you to easily use crawler features without command line.

For development, you need to start both the backend API service and the frontend Vite dev server:

# Terminal 1: start API server (default port 8080)
uv run uvicorn api.main:app --port 8080 --reload

# Terminal 2: start frontend dev server
cd webui
npm install
npm run dev        # starts on port 5173 by default and proxies /api to 8080

After successful startup, visit http://localhost:5173/ to open the WebUI interface.

On first launch, an environment check is performed (calls /api/env/check), so make sure the backend service is running. If the check fails, you can click "Skip Check" to bypass it temporarily.

Build for Production

If you want the API server to serve the WebUI static assets directly, build the frontend first:

cd webui
npm install
npm run build      # outputs to api/webui/

Then start only the API server:

uv run uvicorn api.main:app --port 8080 --reload

After successful startup, visit http://localhost:8080 to open the WebUI interface.

WebUI Features

  • Visualize crawler parameter configuration (platform, login method, crawling type, etc.)
  • Real-time view of crawler running status and logs
  • Data preview and export

Interface Preview

WebUI Interface Preview
🔗 Using Python native venv environment management (Not recommended)

Create and activate Python virtual environment

If crawling Douyin and Zhihu, you need to install nodejs environment in advance, version greater than or equal to: 16

# Enter project root directory
cd MediaCrawler

# Create virtual environment
# My python version is: 3.9.6, the libraries in requirements.txt are based on this version
# If using other python versions, the libraries in requirements.txt may not be compatible, please resolve on your own
python -m venv venv

# macOS & Linux activate virtual environment
source venv/bin/activate

# Windows activate virtual environment
venv\Scripts\activate

Install dependency libraries

pip install -r requirements.txt

Install playwright browser driver

playwright install

Run crawler program (native environment)

# The project does not enable comment crawling mode by default. If you need comments, please modify the ENABLE_GET_COMMENTS variable in config/base_config.py
# Other supported options can also be viewed in config/base_config.py with Chinese comments

# Read keywords from configuration file to search related posts and crawl post information and comments
python main.py --platform xhs --lt qrcode --type search

# Read specified post ID list from configuration file to get information and comment information of specified posts
python main.py --platform xhs --lt qrcode --type detail

# Open corresponding APP to scan QR code for login

# For other platform crawler usage examples, execute the following command to view
python main.py --help

💾 Data Storage

MediaCrawler supports multiple data storage methods, including CSV, JSON, JSONL, Excel, SQLite, and MySQL databases.

📖 For detailed usage instructions, please see: Data Storage Guide


🚀 MediaCrawlerPro Major Release 🚀! More features, better architectural design!

💬 Discussion Groups

  • WeChat Discussion Group: Click to join
  • Bilibili Account: Follow me, sharing AI and crawler technology knowledge

💰 Sponsor Display

Sponsor Introduction
TikHub TikHub.io provides 900+ highly stable data interfaces, covering 14+ mainstream domestic and international platforms including TK, DY, XHS, Y2B, Ins, X, etc. Supports multi-dimensional public data APIs for users, content, products, comments, etc., with 40M+ cleaned structured datasets. Use invitation code cfzyejV9 to register and recharge, and get an additional $2 bonus.
Atlas CloudAtlas Cloud Atlas Cloud is a full-modal AI inference platform that gives developers a single AI API to access video generation, image generation, and LLM APIs. Instead of managing multiple vendor integrations, you connect once and get unified access to 300+ curated models across all modalities. Check out Atlas Cloud's new coding plan promotion for more budget-friendly API access.
NodeMaven NodeMaven is an efficient proxy provider for web scraping and automation, offering the highest-quality IPs on the market. Key benefits include 99.9% uptime, ZIP targeting, IP filtering across all proxies (fraud score below 97%), no KYC, and unique free tools such as Proxy Bandwidth Checker, Meta Tag Checker, IP Lookup, and more. MediaCrawler users get 35% off mobile and residential proxies with code CRAWLER35, and 40% off ISP (static) proxies with code CRAWLER40. 👉 Visit NodeMaven
OpenLux Thank you to OpenLux for sponsoring this project! OpenLux is an all-in-one AI platform for businesses, bringing together leading AI models from major providers worldwide. With fast, reliable service and responsive technical support, OpenLux offers base pricing for Claude, OpenAI, and Gemini models as low as 8.82%, 4%, and 8% of official rates, respectively. Exclusive offer for MediaCrawler users: Sign up through our referral link and enjoy up to 7.5% off credit top-ups! 👉 Get started with OpenLux
SX.ORG SX.ORG is a high-performance proxy network built for heavy web scraping and anti-bot bypass, fully compatible with MediaCrawler. Key advantages include global dynamic residential IP coverage across 190+ locations, 99.9% network uptime, precise country/city/ASN targeting, native HTTP(S) & SOCKS5 support, and flexible session rotation for social media platforms. MediaCrawler users can use exclusive promo code CRAWLER3G at signup to get 3 GB of free trial traffic. 👉 Claim 3 GB on SX.ORG

🤝 Become a Sponsor

Become a sponsor and showcase your product here, getting massive exposure daily!

Contact Information:


📚 Other

⭐ Star Trend Chart

If this project helps you, please give a ⭐ Star to support and let more people see MediaCrawler!

Star History Chart

📚 References

Disclaimer

1. Project Purpose and Nature

This project (hereinafter referred to as "this project") was created as a technical research and learning tool, aimed at exploring and learning network data collection technologies. This project focuses on research of data crawling technologies for social media platforms, intended to provide learners and researchers with technical exchange purposes.

The project developer (hereinafter referred to as "developer") solemnly reminds users to strictly comply with relevant laws and regulations of the People's Republic of China when downloading, installing and using this project, including but not limited to the "Cybersecurity Law of the People's Republic of China", "Counter-Espionage Law of the People's Republic of China" and all applicable national laws and policies. Users shall bear all legal responsibilities that may arise from using this project.

3. Usage Purpose Restrictions

This project is strictly prohibited from being used for any illegal purposes or non-learning, non-research commercial activities. This project may not be used for any form of illegal intrusion into other people's computer systems, nor may it be used for any activities that infringe upon others' intellectual property rights or other legitimate rights and interests. Users should ensure that their use of this project is purely for personal learning and technical research, and may not be used for any form of illegal activities.

4. Disclaimer

The developer has made every effort to ensure the legitimacy and security of this project, but assumes no responsibility for any form of direct or indirect losses that may arise from users' use of this project. Including but not limited to any data loss, equipment damage, legal litigation, etc. caused by using this project.

5. Intellectual Property Statement

The intellectual property rights of this project belong to the developer. This project is protected by copyright law and international copyright treaties as well as other intellectual property laws and treaties. Users may download and use this project under the premise of complying with this statement and relevant laws and regulations.

6. Final Interpretation Rights

The developer has the final interpretation rights regarding this project. The developer reserves the right to change or update this disclaimer at any time without further notice.

🙏 Acknowledgments

JetBrains Open Source License Support

Thanks to JetBrains for providing free open source license support for this project!

JetBrains
S
Description
No description provided
Readme
30 MiB
Languages
Python 91.6%
TypeScript 8.2%