Commit Graph
3 Commits
Author SHA1 Message Date
butubb cd85587f00 feat(privacy): 关掉昵称脱敏 —— 脱敏有损,撞名就分不出博主
Deploy VitePress site to Pages / build (push) Canceled after 0s
Deploy VitePress site to Pages / Deploy (push) Canceled after 0s
需求:评论栏/作品栏要能分清是哪个博主。修好分组字段之后名字仍带星号,因为上游作为
教学版默认对昵称做中间脱敏(首尾各留 1 字,中间星号)。这个脱敏是**有损**的:

  「张三」和「张四」都变成「张*」
  「小明老师」和「小刚老师」都变成「小***师」

而本仓库的用途是监控一批公开创作者账号,分清谁是谁正是这一层要干的事。所以关掉它。

* config/base_config.py 新增 MASK_NICKNAME = False(和 INJECT_ALL_COOKIES 一样是个
  开关,不是删代码 —— 改回 True 就恢复上游行为)。
* tools/user_hash.py 的 mask_nickname 读这个开关,关闭时原样返回。读的是模块属性而
  不是导入值,测试才能 monkeypatch。脱敏实现本身一字未动。
* 顺带修一个数据陈旧问题:评论是去重后直接 continue 的,昵称只在首次入库时写一次,
  于是开关一改(或评论者改名)老评论永远停在旧值 —— 而重采是唯一能拿到新值的途径。
  现在已存在的评论会跟着刷新昵称(作品那边的 creator_name 早就是这么做的)。
* anonymous 的 creator_hash 保持不变:那是分组用的稳定键,不是显示名。

测试:
* 三个隐私套件 + weibo 的 autouse fixture 强制把开关打开 —— 它们验的是**脱敏机制
  本身**,机制仍然必须正确,所以显式打开来测,而不是让它们随部署配置漂。
* test_mask_and_hash_tools 改成两个方向都覆盖(开着脱敏 / 关着脱敏)。
* test_tieba_extractor.py 里 8 处字面量的脱敏期望值换成真实昵称 —— 提取器现在就是
  返回原文的,期望值理应跟着改(这一条是行为变更的直接后果,不是测试放宽)。
* 新增一条:已入库的评论昵称会随重采刷新(且不会因刷新而重复插入)。
2026-10-10 15:48:15 +08:00
程序员阿江(Relakkes) a06273ea6c fix: 将平台业务ID统一改为String并自动建表
- 抖音、B站、快手、微博的帖子/视频/评论ID从BigInteger改为String,
  避免PostgreSQL下字符串写入BIGINT报错及未来ID溢出风险
- B站dynamic_id改为String,修复API返回id_str被强转int导致的精度丢失
- 知乎提取器对content_id/question_id显式str()转换
- main.py启动数据库保存模式时自动建表,无需手动--init_db
- 同步更新相关老化测试
2026-07-01 23:03:43 +08:00
程序员阿江(Relakkes) f328ee35b5 fix: restore Tieba crawling after PC page rewrite
Tieba search, detail, comments, creator, and forum-list pages now rely on the current signed PC JSON APIs instead of brittle HTML selectors. The CLI also maps Tieba detail and creator arguments into the platform-specific config so command-line runs exercise the intended mode.

Constraint: Tieba PC pages no longer expose stable HTML structures for search, creator, and forum-list extraction
Constraint: Current PC APIs require browser cookies, tbs, and the web client signing convention
Rejected: Keep expanding HTML selectors | search and creator pages returned large documents with empty parsed results after the redesign
Confidence: high
Scope-risk: moderate
Directive: Do not replace these API paths with page HTML parsing without re-verifying the current Tieba network requests
Tested: uv run pytest tests/test_tieba_client_pagination.py tests/test_cmd_arg_tieba.py tests/test_tieba_extractor.py -q
Tested: uv run python -m py_compile cmd_arg/arg.py media_platform/tieba/help.py media_platform/tieba/client.py media_platform/tieba/core.py tests/test_cmd_arg_tieba.py tests/test_tieba_client_pagination.py tests/test_tieba_extractor.py
Tested: uv run main.py --platform tieba --type search --keywords 编程兼职 --get_comment false
Tested: uv run main.py --platform tieba --type detail --specified_id 9835114923 --get_comment true --max_comments_count_singlenotes 3
Tested: uv run main.py --platform tieba --type creator --creator_id https://tieba.baidu.com/home/main?id=tb.1.6ad0cd4a.7ZcjVYWa7UpHttCld2OppA --get_comment false
Not-tested: Second-level Tieba comment API migration; this path still uses the existing /p/comment HTML parser
Not-tested: Full pytest suite has one pre-existing unrelated XHS Excel store assertion failure
2026-04-30 18:20:46 +08:00