fix(ui): 「每轮最多采集作品数」标签是错的——它是每个博主的上限
爬虫里这个值是在 per-creator 的函数内比较的(client.py get_all_notes_by_creator 的 result 是局部变量),而外层 for 循环遍历全部目标。所以 100 个目标 × 20 篇 = 单轮最多 2000 篇,一篇都不会被丢弃。标签写成「每轮最多」会让人以为超出的会被截掉。 - creator 模式:标签改为「每个博主最多采集作品数」,并实时算出「N 个目标 × M 篇 → 单轮最多 X 篇」 - note 模式:禁用该输入并说明「此项不生效」——get_specified_notes 里没有任何 CRAWLER_MAX_NOTES_COUNT 引用,列出的每个链接都会被逐条抓 - 单轮估算超过 500 篇时给出警告:每篇还要抓最多 max_comments_count 条评论、并发为 1,容易触发限流,也可能跑不完就被默认 1 小时的任务超时中断
This commit is contained in:
@@ -57,6 +57,15 @@ const SCHEDULE_MODE_OPTIONS: Array<{ value: ScheduleMode; label: string }> = [
|
||||
]
|
||||
|
||||
const HOURS = Array.from({ length: 24 }, (_, hour) => hour)
|
||||
|
||||
/**
|
||||
* Above this many works per run, warn about the request volume.
|
||||
*
|
||||
* The runner serialises crawls (max_concurrency_num=1) and each note also pulls
|
||||
* up to `max_comments_count` comments, so the cost is works × (1 + comments) and
|
||||
* the default run timeout is an hour.
|
||||
*/
|
||||
const NOTE_VOLUME_WARN = 500
|
||||
const MINUTES = Array.from({ length: 60 }, (_, minute) => minute)
|
||||
const WEEKDAYS = ['周一', '周二', '周三', '周四', '周五', '周六', '周日']
|
||||
const pad = (value: number) => String(value).padStart(2, '0')
|
||||
@@ -371,18 +380,48 @@ export function TaskEditorDialog({ open, onOpenChange, task }: TaskEditorDialogP
|
||||
<div className="grid grid-cols-2 gap-3">
|
||||
<div className="space-y-2">
|
||||
<Label className="text-xs font-mono text-cyber-text-secondary">
|
||||
每轮最多采集作品数
|
||||
{mode === 'creator' ? '每个博主最多采集作品数' : '最多采集作品数'}
|
||||
</Label>
|
||||
<Input
|
||||
type="number"
|
||||
min={1}
|
||||
value={maxNotes}
|
||||
onChange={(event) => setMaxNotes(event.target.value)}
|
||||
disabled={mode === 'note'}
|
||||
className="h-9 text-xs"
|
||||
/>
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
只取最新的前 N 条,决定了"该博主的作品"覆盖范围
|
||||
</p>
|
||||
|
||||
{mode === 'creator' ? (
|
||||
<>
|
||||
{/* 这是「每个博主」的上限,不是一轮的总量 —— 爬虫里这个值是在
|
||||
per-creator 的函数内比较的(client.py get_all_notes_by_creator),
|
||||
而外层 for 循环会遍历全部目标。所以 100 个目标 × 20 篇 = 单轮
|
||||
最多 2000 篇。标签写成「每轮最多」会让人以为超出的会被丢弃。 */}
|
||||
<p className="text-[10px] font-mono text-cyber-text-muted">
|
||||
这是<span className="text-cyber-text-secondary">每个博主</span>的上限,
|
||||
不是一轮的总量。只取该博主最新的前 N 条,超出的不会补抓。
|
||||
</p>
|
||||
<p className="text-[10px] font-mono text-cyber-text-secondary">
|
||||
{targetList.length} 个目标 × {maxNotes} 篇 × 每人 1 次
|
||||
→ 单轮最多{' '}
|
||||
<span className="text-cyber-neon-cyan">
|
||||
{targetList.length * (Number(maxNotes) || 0)}
|
||||
</span>{' '}
|
||||
篇
|
||||
</p>
|
||||
{targetList.length * (Number(maxNotes) || 0) > NOTE_VOLUME_WARN && (
|
||||
<p className="text-[10px] font-mono text-cyber-neon-orange leading-relaxed">
|
||||
单轮量偏大:每篇还要抓最多 {maxComments} 条评论,且并发为 1。
|
||||
容易触发平台限流,也可能跑不完就被任务超时(默认 1 小时)中断。
|
||||
建议调低这个数,或拆成几个任务。
|
||||
</p>
|
||||
)}
|
||||
</>
|
||||
) : (
|
||||
<p className="text-[10px] font-mono text-cyber-neon-orange">
|
||||
笔记模式下此项不生效:你列出的每个链接都会被逐条抓取。
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
|
||||
<div className="space-y-2">
|
||||
|
||||
Reference in New Issue
Block a user