From 2112c1a870fbaffa4491766de0c82a132a3909d1 Mon Sep 17 00:00:00 2001 From: butubb <1422726308@qq.com> Date: Sat, 10 Oct 2026 17:12:25 +0800 Subject: [PATCH] =?UTF-8?q?feat(monitor):=20=E6=8A=96=E9=9F=B3=20Web=20?= =?UTF-8?q?=E6=8E=A5=E5=8F=A3=E5=AE=A2=E6=88=B7=E7=AB=AF=20=E2=80=94?= =?UTF-8?q?=E2=80=94=20=E7=BB=95=E5=BC=80=E7=88=AC=E8=99=AB=E5=AD=90?= =?UTF-8?q?=E8=BF=9B=E7=A8=8B=EF=BC=8C=E7=9B=B4=E6=8E=A5=E5=8F=91=20HTTP?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 移植自 mac-agent-os 的 mediacrawler_adapter:不起子进程、不开页面,用浏览器里那份 登录态直接调抖音 Web 接口。产物键名照抄 store/douyin,所以 ingest 那条链路一个字不用改。 **目前能用的(真环境实测,非推断)**: profile/other : 200, 7075 字节 —— 博主主页指标(粉丝/获赞/作品数/昵称) aweme/detail : 200, 45425 字节 —— 单条作品详情(含点赞/评论/收藏/分享) **目前不能用的:作品列表 `aweme/post`。** 两个互相独立的原因: 1. 这个接口被抖音单独升级成了真校验:不带 x-tt-argus 回 403「Uifid Not Found」, 带上 dummy 值回 200 + **空 body**。也就是说「头在不在」骗得过,「真校验」过不了。 同一套头打 profile/other 和 aweme/detail 都是通的 —— 抖音是挑着接口加保护的, 挑中的恰好是「批量拉作品列表」这个最敏感的动作。 2. 改走页面截获也不行:CDP 浏览器打开博主主页会落到「验证码中间页」(当天大量探测的 代价,过几小时要重测)。 所以现在的边界是:**已知作品的指标刷新能做,自动发现新作品做不了**。 **排查中控住变量后得到的两条事实**(都写进注释了): · `Accept` / `Accept-Language` / `Referer` 才是主页接口能返回真数据的原因 —— 只有 UA+client hints+Cookie 时是 200 但仅 121 字节的空壳,补上这三个头变 7074 字节。 (我先前猜的 sec-ch-ua 不是关键。) · 因此 UA 与 client hints 必须**成套地取自同一个浏览器**,所以 BrowserIdentity 一次 从 CDP 取齐 cookie + UA + hints,而不是各自写死。 「200 + 空 body 必须当场报错」也是刻意写死的:放过去它会在下游变成「这个博主没作品」, 把一次失败伪装成一条正常结果 —— 爬虫那条路正是这么栽的,还被翻译成「账号被封」。 测试 +11:cookie 解析、请求头成套性(含 uifid 缺失/回退)、产物键名与 store 对齐、 以及 _get 的三条失败路径(空 body / 403 带网关原话 / 正常返回)。 --- api/monitor/douyin_api.py | 429 ++ mac-agent-os-main/.codewhale/handoff.md | 49 + mac-agent-os-main/.gitignore | 182 + .../00_bootstrap/DEPLOY-GUIDE.md | 183 + .../00_bootstrap/apply-config.sh | 160 + mac-agent-os-main/00_bootstrap/deploy.sh | 232 + .../00_bootstrap/disable_worker_dashboard.sh | 72 + .../00_bootstrap/export_skills.sh | 58 + .../00_bootstrap/fleet_reconcile.sh | 202 + mac-agent-os-main/00_bootstrap/fleet_sync.sh | 236 + .../00_bootstrap/hooks/pre-commit | 33 + mac-agent-os-main/00_bootstrap/hooks/pre-push | 39 + .../00_bootstrap/import_skills.sh | 96 + mac-agent-os-main/00_bootstrap/init.sh | 463 ++ mac-agent-os-main/00_bootstrap/setup_env.sh | 58 + .../00_bootstrap/upgrade_20260514.sh | 54 + mac-agent-os-main/01_core/CHANGELOG.md | 11 + .../01_core/CONFIG_MANIFEST.yaml | 63 + mac-agent-os-main/01_core/HOST_ID.tpl.md | 46 + mac-agent-os-main/01_core/IDENTITY.tpl.md | 55 + .../01_core/MAINTENANCE_GUIDE.md | 351 ++ .../01_core/NIGHTLY_AUTOMATION.md | 57 + mac-agent-os-main/01_core/SOUL.md | 163 + mac-agent-os-main/01_core/SOUL.md.v2-backup | 186 + mac-agent-os-main/01_core/UPDATE_SYSTEM.md | 199 + mac-agent-os-main/01_core/USER.tpl.md | 46 + mac-agent-os-main/01_core/VERSION | 12 + .../_archived/SOUL.md.v2-backup.note.md | 4 + .../01_core/automation/README.md | 23 + .../launchd/com.agentos.guardd.plist.template | 37 + .../01_core/automation/workflows.yaml | 37 + .../20260518_yuan_benchu_历史解读_006.md | 57 + .../20260518_yuan_benchu_哲学思考_003.md | 62 + .../20260518_yuan_benchu_哲学思考_004.md | 55 + .../20260518_yuan_benchu_哲学思考_009.md | 74 + .../20260518_yuan_benchu_政治分析_002.md | 52 + .../20260518_yuan_benchu_政治分析_007.md | 68 + .../20260518_yuan_benchu_社会观察_001.md | 58 + .../20260518_yuan_benchu_社会观察_005.md | 57 + .../20260518_yuan_benchu_社会观察_010.md | 79 + .../20260518_yuan_benchu_经济分析_008.md | 71 + .../20260519_yuan_benchu_历史解读_008.md | 84 + .../20260519_yuan_benchu_哲学思考_004.md | 81 + .../20260519_yuan_benchu_哲学思考_005.md | 83 + .../20260519_yuan_benchu_政治分析_006.md | 82 + .../20260519_yuan_benchu_政治分析_007.md | 76 + .../20260519_yuan_benchu_社会观察_001.md | 70 + .../20260519_yuan_benchu_社会观察_002.md | 66 + .../20260519_yuan_benchu_社会观察_003.md | 64 + .../20260519_yuan_benchu_社会观察_010.md | 104 + .../20260519_yuan_benchu_经济分析_009.md | 89 + .../20260602_yuan_benchu_哲学思考_008.md | 61 + .../20260602_yuan_benchu_哲学思考_009.md | 51 + .../20260602_yuan_benchu_政治分析_001.md | 60 + .../20260602_yuan_benchu_政治分析_002.md | 49 + .../20260602_yuan_benchu_政治分析_003.md | 61 + .../20260602_yuan_benchu_社会观察_004.md | 47 + .../20260602_yuan_benchu_社会观察_005.md | 64 + .../20260602_yuan_benchu_社会观察_006.md | 62 + .../20260602_yuan_benchu_社会观察_007.md | 51 + .../20260602_yuan_benchu_经济分析_010.md | 61 + .../20260603_yuan_benchu_历史解读_009.md | 59 + .../20260603_yuan_benchu_政治分析_004.md | 47 + .../20260603_yuan_benchu_政治分析_005.md | 49 + .../20260603_yuan_benchu_社会观察_008.md | 66 + .../20260603_yuan_benchu_社会观察_009.md | 57 + .../20260603_yuan_benchu_社会观察_010.md | 80 + .../20260603_yuan_benchu_社会观察_011.md | 68 + .../20260603_yuan_benchu_经济分析_011.md | 63 + .../20260603_yuan_benchu_经济分析_012.md | 65 + .../20260603_yuan_benchu_经济分析_013.md | 71 + .../20260604_yuan_benchu_历史解读_01.md | 65 + .../20260604_yuan_benchu_哲学思考_01.md | 66 + .../20260604_yuan_benchu_政治分析_01.md | 64 + .../20260604_yuan_benchu_政治分析_02.md | 44 + .../20260604_yuan_benchu_政治分析_03.md | 63 + .../20260604_yuan_benchu_社会观察_01.md | 55 + .../20260604_yuan_benchu_社会观察_02.md | 79 + .../20260604_yuan_benchu_经济分析_01.md | 56 + .../20260604_yuan_benchu_经济分析_02.md | 74 + .../20260604_yuan_benchu_经济分析_03.md | 81 + .../20260605_yuan_benchu_历史解读_01.md | 63 + .../20260605_yuan_benchu_哲学思考_01.md | 69 + .../20260605_yuan_benchu_政治分析_01.md | 47 + .../20260605_yuan_benchu_政治分析_02.md | 46 + .../20260605_yuan_benchu_政治分析_03.md | 46 + .../20260605_yuan_benchu_社会观察_01.md | 59 + .../20260605_yuan_benchu_经济分析_01.md | 42 + .../20260605_yuan_benchu_经济分析_02.md | 57 + .../20260605_yuan_benchu_经济分析_03.md | 60 + .../20260605_yuan_benchu_经济分析_04.md | 56 + .../20260608_yuan_benchu_历史解读_01.md | 122 + .../20260608_yuan_benchu_哲学思考_01.md | 155 + .../20260608_yuan_benchu_政治分析_01.md | 89 + .../20260608_yuan_benchu_政治分析_02.md | 119 + .../20260608_yuan_benchu_政治分析_03.md | 95 + .../20260608_yuan_benchu_社会观察_01.md | 103 + .../20260608_yuan_benchu_社会观察_02.md | 100 + .../20260608_yuan_benchu_社会观察_03.md | 113 + .../20260608_yuan_benchu_社会观察_04.md | 104 + .../20260608_yuan_benchu_经济分析_01.md | 127 + .../02_skills/_archived/README.md | 12 + .../_archived/auto_collector/SKILL.md | 15 + .../cloakbrowser_controller/SKILL.md | 15 + .../_archived/content_processor/SKILL.md | 15 + .../02_skills/_archived/web_crawler/SKILL.md | 15 + .../02_skills/_template/SKILL.md | 32 + .../02_skills/_template/skill.py | 24 + .../02_skills/_template/version.json | 7 + .../02_skills/auto_collector/SKILL.md | 107 + .../02_skills/auto_collector/SKILL_CARD.yaml | 45 + .../02_skills/auto_collector/version.json | 1 + .../cloakbrowser_controller/SKILL.md | 101 + .../02_skills/collect_to_inbox/SKILL.md | 137 + .../collect_to_inbox/SKILL_CARD.yaml | 54 + .../collect_to_inbox/collect_to_inbox.py | 312 ++ .../02_skills/collect_to_inbox/version.json | 1 + .../02_skills/content_processor/SKILL.md | 91 + .../02_skills/content_processor/version.json | 1 + .../02_skills/inbox_refine/SKILL.md | 140 + .../02_skills/inbox_refine/SKILL_CARD.yaml | 46 + .../02_skills/inbox_refine/inbox_refine.py | 461 ++ .../02_skills/inbox_refine/llm_classifier.py | 395 ++ .../inbox_refine/llm_classifier_broken.py | 394 ++ .../02_skills/inbox_refine/version.json | 17 + .../02_skills/kb_manager/SKILL.md | 119 + .../02_skills/kb_manager/SKILL_CARD.yaml | 56 + .../02_skills/kb_manager/kb_ingest.py | 287 + .../02_skills/kb_manager/kb_search.py | 293 + .../02_skills/kb_manager/vector_db_rebuild.py | 292 + .../02_skills/kb_manager/version.json | 7 + mac-agent-os-main/02_skills/matrix/SKILL.md | 226 + .../02_skills/matrix/SKILL_CARD.yaml | 66 + .../02_skills/matrix/version.json | 8 + .../02_skills/memory_manager/SKILL.md | 130 + .../02_skills/memory_manager/SKILL_CARD.yaml | 57 + .../memory_manager/agent_memory_init.py | 119 + .../memory_manager/bootstrap_from_memory.py | 388 ++ .../02_skills/memory_manager/daily_digest.py | 536 ++ .../memory_manager/export_memories.py | 114 + .../memory_manager/import_memories.py | 226 + .../memory_manager/memory_cleanup.py | 126 + .../memory_manager/memory_extractor.py | 324 ++ .../memory_manager/semantic_search.py | 664 +++ .../memory_manager/vector_backfill.sh | 78 + .../02_skills/memory_manager/version.json | 7 + .../02_skills/peekaboo_controller/SKILL.md | 78 + .../02_skills/peekaboo_controller/policy.py | 100 + .../02_skills/sync_manager/SKILL.md | 45 + .../02_skills/sync_manager/SKILL_CARD.yaml | 55 + .../02_skills/sync_manager/sync_manager.py | 320 ++ .../02_skills/sync_manager/version.json | 7 + .../02_skills/web_crawler/SKILL.md | 58 + .../02_skills/web_crawler/version.json | 7 + mac-agent-os-main/03_knowledge/00_README.md | 110 + .../03_knowledge/00_inbox/.gitkeep | 0 .../2026-06_submissions_distillation.md | 298 + .../inbox/knowledge/clipping/.gitkeep | 0 .../00_stream/inbox/knowledge/feed/.gitkeep | 0 .../00_stream/inbox/media/.gitkeep | 0 .../00_stream/inbox/memory/.gitkeep | 0 .../00_stream/inbox/personal/.gitkeep | 0 .../00_stream/inbox/tools/.gitkeep | 0 .../03_knowledge/01_daily/.gitkeep | 0 .../03_knowledge/01_daily/memory | 1 + .../03_knowledge/01_submissions/.gitkeep | 0 .../20260503_social-wisdom_douyin-video.md | 32 + .../20260503_social-wisdom_subtitle.md | 27 + .../20260503_social-wisdom_summary.md | 18 + .../2026-05-16_10条社会智慧_-_AI总结.md | 24 + .../2026-05-16_10条社会智慧_-_视频文案.md | 33 + ...05-16_10条社会智慧_每一条都藏着处世真相.md | 38 + .../10_concepts/2026-06-01_2026-04-25.md | 42 + .../10_concepts/2026-06-01_2026-04-26.md | 47 + .../10_concepts/2026-06-01_2026-04-27.md | 44 + .../10_concepts/2026-06-01_2026-04-28.md | 44 + .../10_concepts/2026-06-01_2026-04-29.md | 38 + .../2026-06-01_2026-04-29_1778863247.md | 44 + .../10_concepts/2026-06-01_2026-05-01.md | 44 + .../2026-06-01_bootstrap_daily_2026-04-25.md | 41 + ...ystem_189058df-1ad8-438f-b8aa-61087a7c1.md | 43 + ...ystem_7a50954c-2a08-4dbb-adde-abee01c49.md | 44 + .../03_knowledge/10_concepts/design/.gitkeep | 0 .../family-succession-capability-first.md | 95 + .../global-accounting-strategic-loss.md | 197 + .../10_concepts/mao-ten-thinking-methods.md | 96 + .../money-making-logic-value-creation.md | 119 + .../10_concepts/personal-management/.gitkeep | 0 .../10_concepts/power-of-naming-reality.md | 191 + .../10_concepts/social-wisdom-10-rules.md | 97 + .../03_knowledge/20_methods/.gitkeep | 0 ...文生视频__李导_Agent版_10条视频知识提取.md | 53 + ...6_character-consistency-reference-sheet.md | 44 + ...026-05-16_lip-sync-solutions-comparison.md | 52 + ...示词技巧__依士曼胶片_高调硬光_影棚置景.md | 51 + ...素材生成__提示词工程与分镜脚本方法汇总.md | 48 + ...2026-05-16_character-consistency-refere.md | 44 + ...2026-05-16_lip-sync-solutions-compariso.md | 50 + ...026-05-17_2026-05-16_stuck-intervention.md | 47 + ...动文生视频__AICG造梦局10条视频知识提取.md | 42 + ...26-05-17_video-shooting-cool-techniques.md | 54 + ...全知识手册__文生视频_图生视频提示词工程.md | 57 + ...26-05-16_AI视频素材生成__提示词工程与分.md | 47 + ...2026-05-17_2026-05-16_character-consist.md | 41 + ...6-01_2026-05-17_2026-05-16_cross-domain.md | 44 + ..._2026-05-17_2026-05-16_knowledge-review.md | 55 + ...2026-05-17_2026-05-16_lip-sync-solution.md | 48 + ...-01_2026-05-17_2026-05-16_meta-thinking.md | 52 + ...2026-05-17_2026-05-16_stuck-interventio.md | 47 + ...7_2026-05-16_分镜时序与运镜_让AI_懂剪辑.md | 52 + ...7_2026-05-16_原子化登录管理模块_auth_ma.md | 45 + ...5-17_2026-05-16_统一升级引擎_agentos_up.md | 45 + ...驱动文生视频__AICG造梦局10条视频知识提.md | 40 + ...示词驱动文生视频__博主干货提取_10_11条_.md | 47 + ...示词驱动文生视频__李导_Agent版_10条视频.md | 50 + ...视频提示词系统创作教科书___知识体系索引.md | 52 + ...026-06-01_2026-05-17_ch02-result-driven.md | 52 + ...-06-01_2026-05-17_ch07-micro-expression.md | 44 + ...26-06-01_2026-05-17_ch10-commercial-tvc.md | 52 + ...26-06-01_2026-05-17_ch11-sci-fi-fantasy.md | 52 + ...-06-01_2026-05-17_ch12-model-comparison.md | 42 + ...026-06-01_2026-05-17_ch13-tool-workflow.md | 53 + ...-06-01_2026-05-17_ch14-script-to-prompt.md | 56 + ...-01_2026-05-17_ch15-narrative-structure.md | 47 + ...026-06-01_2026-05-17_ch16-practice-path.md | 47 + ...-06-01_2026-05-17_ch17-resource-library.md | 37 + ...6-06-01_2026-05-17_ch18-problem-solving.md | 53 + ...2026-05-17_video-shooting-cool-techniqu.md | 52 + ...完全知识手册__文生视频_图生视频提示词工.md | 54 + ...种感觉__爆款酷炫短视频拆解与标准化流程.md | 55 + ...AI提示词技巧__依士曼胶片_高调硬光_影棚.md | 50 + ...-06-01_2026-05-17_2026-05-16_AI视频素材.md | 44 + ...2026-06-01_2026-05-17_2026-05-16_Camouf.md | 41 + ...2026-06-01_2026-05-17_2026-05-16_charac.md | 39 + ...2026-06-01_2026-05-17_2026-05-16_cross-.md | 45 + ...2026-06-01_2026-05-17_2026-05-16_knowle.md | 53 + ...2026-06-01_2026-05-17_2026-05-16_lip-sy.md | 47 + ...2026-06-01_2026-05-17_2026-05-16_meta-t.md | 51 + ...2026-06-01_2026-05-17_2026-05-16_stuck-.md | 46 + ...6-01_2026-05-17_2026-05-16_分镜时序与运.md | 49 + ...6-01_2026-05-17_2026-05-16_原子化登录管.md | 45 + ...6-01_2026-05-17_2026-05-16_统一升级引擎.md | 44 + ...1_2026-05-17_AI提示词驱动文生视频__AICG.md | 41 + ...26-05-17_AI提示词驱动文生视频__博主干货.md | 48 + ...2026-05-17_AI提示词驱动文生视频__李导_A.md | 47 + ..._2026-05-17_AI视频提示词系统___采集队列.md | 51 + ...026-05-17_AI视频提示词系统创作教科书___.md | 51 + ...2_2026-06-01_2026-05-17_ch01-core-logic.md | 48 + ...2026-06-01_2026-05-17_ch02-result-drive.md | 52 + ...2026-06-01_2026-05-17_ch06-audio-design.md | 53 + ...2026-06-01_2026-05-17_ch07-micro-expres.md | 45 + ...2_2026-06-01_2026-05-17_ch08-film-style.md | 53 + ..._2026-06-01_2026-05-17_ch09-anime-style.md | 48 + ...2026-06-01_2026-05-17_ch10-commercial-t.md | 53 + ...2026-06-01_2026-05-17_ch11-sci-fi-fanta.md | 53 + ...2026-06-01_2026-05-17_ch12-model-compar.md | 43 + ...2026-06-01_2026-05-17_ch13-tool-workflo.md | 51 + ...2026-06-01_2026-05-17_ch14-script-to-pr.md | 56 + ...2026-06-01_2026-05-17_ch15-narrative-st.md | 48 + ...2026-06-01_2026-05-17_ch16-practice-pat.md | 48 + ...2026-06-01_2026-05-17_ch17-resource-lib.md | 38 + ...2026-06-01_2026-05-17_ch18-problem-solv.md | 52 + ...06-02_2026-06-01_2026-05-17_ch19-trends.md | 43 + ...2026-06-01_2026-05-17_video-shooting-co.md | 50 + ...-01_2026-05-17_光影色调_定义_情绪与氛围.md | 41 + ...26-05-17_可灵AI提示词完全知识手册__文生.md | 51 + ...1_2026-05-17_基础定调参数_锁死_底层标准.md | 56 + ...-05-17_帅是一种感觉__爆款酷炫短视频拆解.md | 53 + ...26-05-17_邵氏电影美学AI提示词技巧__依士.md | 46 + ...06-01_2026-05-28_影视风格提示词速查手册.md | 56 + ...-06-02_2026-05-29_OpenCLI安装与集成指南.md | 46 + ...2026-06-02_2026-06-01_2026-05-17_2026-0.md | 42 + ...-06-02_2026-06-01_2026-05-17_AI提示词驱.md | 44 + ...-06-02_2026-06-01_2026-05-17_AI视频提示.md | 46 + ...2026-06-02_2026-06-01_2026-05-17_ch01-c.md | 44 + ...2026-06-02_2026-06-01_2026-05-17_ch02-r.md | 50 + ...2026-06-02_2026-06-01_2026-05-17_ch06-a.md | 47 + ...2026-06-02_2026-06-01_2026-05-17_ch07-m.md | 42 + ...2026-06-02_2026-06-01_2026-05-17_ch08-f.md | 51 + ...2026-06-02_2026-06-01_2026-05-17_ch09-a.md | 47 + ...2026-06-02_2026-06-01_2026-05-17_ch10-c.md | 52 + ...2026-06-02_2026-06-01_2026-05-17_ch11-s.md | 52 + ...2026-06-02_2026-06-01_2026-05-17_ch12-m.md | 42 + ...2026-06-02_2026-06-01_2026-05-17_ch13-t.md | 46 + ...2026-06-02_2026-06-01_2026-05-17_ch14-s.md | 53 + ...2026-06-02_2026-06-01_2026-05-17_ch15-n.md | 47 + ...2026-06-02_2026-06-01_2026-05-17_ch16-p.md | 47 + ...2026-06-02_2026-06-01_2026-05-17_ch17-r.md | 37 + ...2026-06-02_2026-06-01_2026-05-17_ch18-p.md | 48 + ...2026-06-02_2026-06-01_2026-05-17_ch19-t.md | 42 + ...2026-06-02_2026-06-01_2026-05-17_video-.md | 45 + ...06-02_2026-06-01_2026-05-17_光影色调_定.md | 40 + ...-06-02_2026-06-01_2026-05-17_可灵AI提示.md | 46 + ...6-02_2026-06-01_2026-05-17_基础定调参数.md | 54 + ...6-02_2026-06-01_2026-05-17_帅是一种感觉.md | 49 + ...6-02_2026-06-01_2026-05-17_邵氏电影美学.md | 43 + ...6-02_2026-06-01_2026-05-28_影视风格提示.md | 54 + ...2026-06-02_2026-06-01_2026-05-29_OpenCL.md | 46 + .../20_methods/ai-prompt-creator-2.md | 46 + .../20_methods/ai-prompt-video-knowledge.md | 197 + .../ai-video-prompt-engineering-summary.md | 423 ++ .../01_foundation/ch01-core-logic.md | 78 + .../01_foundation/ch02-result-driven.md | 135 + .../02_core_modules/ch03-parameters.md | 197 + .../02_core_modules/ch04-storyboard-camera.md | 41 + .../02_core_modules/ch05-lighting-color.md | 41 + .../02_core_modules/ch06-audio-design.md | 91 + .../02_core_modules/ch07-micro-expression.md | 450 ++ .../03_applications/ch08-film-style.md | 176 + .../03_applications/ch09-anime-style.md | 111 + .../03_applications/ch10-commercial-tvc.md | 92 + .../03_applications/ch11-sci-fi-fantasy.md | 141 + .../04_tools_models/ch12-model-comparison.md | 69 + .../04_tools_models/ch13-tool-workflow.md | 153 + .../ch14-script-to-prompt.md | 124 + .../ch15-narrative-structure.md | 49 + .../06_practice/ch16-practice-path.md | 102 + .../06_practice/ch17-resource-library.md | 96 + .../07_trends/ch18-problem-solving.md | 159 + .../ai-video-system/07_trends/ch19-trends.md | 110 + .../2026-05-28_影视风格提示词速查手册.md | 601 ++ .../99_assets/collection-queue.md | 87 + .../99_assets/result-driven-template.md | 94 + .../20_methods/ai-video-system/INDEX.md | 370 ++ ...ouyin_ins爆火贴纸特效_ai提示词_20260605.md | 45 + .../20_methods/install_opencli.sh | 81 + .../character-consistency-reference-sheet.md | 144 + .../kling/kling-prompt-complete-guide.md | 384 ++ .../kling/lip-sync-solutions-comparison.md | 204 + .../20_methods/li-dao-ai-video-prompts.md | 464 ++ .../shaoshi-film-aesthetics-prompts.md | 103 + .../case-cool-is-a-feeling.md | 178 + .../03_knowledge/30_facts/.gitkeep | 0 ...026-05-17_CloakBrowser_源码级反爬浏览器.md | 60 + ...5-17_Matrix_养号系统_SMS_验证码自动接收.md | 57 + ..._2026-05-17_Peekaboo_v3_桌面_GUI_自动化.md | 58 + ...-06-01_2026-05-17_CloakBrowser_源码级反.md | 58 + ...06-01_2026-05-17_Matrix_养号系统_SMS_验.md | 56 + ...26-06-01_2026-05-17_Peekaboo_v3_桌面_GU.md | 59 + ...2026-06-02_2026-06-01_2026-05-17_CloakB.md | 57 + ...2026-06-02_2026-06-01_2026-05-17_Matrix.md | 52 + ...2026-06-02_2026-06-01_2026-05-17_Peekab.md | 54 + .../03_knowledge/40_references/docs/.gitkeep | 0 .../40_references/papers/.gitkeep | 0 .../03_knowledge/50_resources/.gitkeep | 0 .../03_knowledge/60_opinions/.gitkeep | 0 .../90_archive/deprecated/.gitkeep | 0 .../deprecated/2026-04-28_测试_LLM_分类器.md | 40 + ...2026-05-16_统一升级引擎_agentos_upgrade.md | 45 + .../99_system/AI_READING_GUIDE.md | 57 + .../99_system/ARCHITECTURE_AUDIT.md | 269 + .../99_system/ARCHITECTURE_CONSTITUTION.md | 574 ++ .../03_knowledge/99_system/ARCHITECTURE_v3.md | 179 + .../ARCHITECTURE_AUDIT.md | 168 + .../EXPERIENCE-atomic-ops-20260621.md | 320 ++ .../PLAN-atomic-ops-v2.md | 343 ++ .../page_inspector.py | 284 + .../日志-2026-06-20.md | 202 + .../日志-2026-06-21.md | 48 + .../项目记忆-MEMORY.md | 147 + .../03_knowledge/99_system/README.md | 27 + .../99_system/TRANSFORMATION_ROADMAP.md | 336 ++ .../agent-os-memory-knowledge-architecture.md | 387 ++ .../architecture/IMPLEMENTATION-PLAN-v1.md | 416 ++ .../federated-multi-machine-architecture.md | 97 + .../federation-operations-architecture.md | 324 ++ .../architecture/loading-architecture.md | 116 + .../architecture/login-system-tech-map.md | 216 + .../architecture/trigger-matching-analysis.md | 133 + .../99_system/archive/CORE-ARCHITECTURE.md | 195 + .../03_knowledge/99_system/archive/INDEX.md | 22 + .../99_system/archive/QUICKSTART.md | 201 + .../99_system/archive/SKILLS-CATALOG.md | 324 ++ .../dashboard-repair-log-20260622.md | 183 + .../99_system/dashboard-v4-design.md | 388 ++ ...6-05-16_原子化登录管理模块_auth_manager.md | 46 + ...-05-17_2026-05-16_Camoufox_集成修复记录.md | 41 + .../matrix/douyin-comment-automation.md | 129 + .../99_system/matrix/douyin-login-pitfalls.md | 353 ++ .../99_system/matrix/matrix-known-pitfalls.md | 161 + .../matrix-nurture-system-architecture.md | 152 + .../matrix/matrix-sms-verification.md | 62 + .../matrix/matrix-v6-product-plan.md | 476 ++ .../xhs-browse-interaction-techniques.md | 230 + .../matrix/指纹分辨率触发XHS_AI布局版本.md | 64 + .../03_knowledge/99_system/memory-index.md | 68 + .../pipelines/content-collection-pipeline.md | 312 ++ .../99_system/prompts/classify-knowledge.md | 49 + .../99_system/protocols/cross-domain.md | 48 + .../99_system/protocols/knowledge-review.md | 64 + .../99_system/protocols/meta-thinking.md | 51 + .../99_system/protocols/stuck-intervention.md | 49 + .../references/cloakbrowser-integration.md | 56 + .../references/peekaboo-v3-integration.md | 48 + .../99_system/session-20260622-dashboard.md | 53 + .../99_system/taxonomies/domains.md | 30 + .../99_system/taxonomies/folder-aliases.json | 64 + .../99_system/taxonomies/nature-types.md | 39 + .../99_system/templates/concept-card.md | 51 + .../99_system/templates/fact-card.md | 39 + .../99_system/templates/method-card.md | 62 + .../templates/personal-insight-card.md | 50 + .../03_knowledge/99_system/timelines/.gitkeep | 0 mac-agent-os-main/03_knowledge/CHANGELOG.md | 242 + .../03_knowledge/KB_REORG_PLAN.md | 158 + mac-agent-os-main/03_knowledge/README.md | 214 + mac-agent-os-main/03_knowledge/versions.json | 24 + mac-agent-os-main/04_memory/CHANGELOG.md | 10 + .../cross_machine/ARCHITECTURE-v2.md | 644 +++ ...ff14-4ed9-885b-b04c5326304d_heartbeat.json | 15 + ...2159-4fe7-b6ff-db14ccf379f5_heartbeat.json | 15 + ...7310-433c-ba35-368c51edb1d0_heartbeat.json | 15 + .../hostname_drift/Redmi-12C/heartbeat.json | 24 + .../cross_machine/encrypted/pending/.gitkeep | 0 .../cross_machine/guardd-required-version.txt | 1 + .../cross_machine/knowledge/.gitkeep | 0 .../04_memory/cross_machine/registry/.gitkeep | 0 .../cross_machine/registry/7kecheng_pub.pem | 14 + .../registry/_archive/5kecheng.json | 7 + .../registry/_archive/7kecheng.json | 7 + .../registry/_archive/chengzige.json | 8 + .../status/_archive/Redmi-12C/heartbeat.json | 1 + .../_archive/chengzigedeAir/heartbeat.json | 24 + .../04_memory/daily_summaries/.gitkeep | 0 .../04_memory/daily_summaries/2026-04-25.md | 260 + .../04_memory/daily_summaries/2026-04-26.md | 2992 ++++++++++ .../04_memory/daily_summaries/2026-04-27.md | 44 + .../04_memory/daily_summaries/2026-04-28.md | 31 + .../04_memory/daily_summaries/2026-04-29.md | 146 + .../04_memory/daily_summaries/2026-05-01.md | 126 + mac-agent-os-main/04_memory/logs/.gitkeep | 0 mac-agent-os-main/04_memory/long_term/raw | 1 + .../04_memory/memory_backup/.gitkeep | 0 mac-agent-os-main/05_tools/00_setup/.gitkeep | 0 .../05_tools/00_setup/agentos/__init__.py | 2 + .../05_tools/00_setup/agentos/__main__.py | 3 + .../05_tools/00_setup/agentos/backup.py | 193 + .../05_tools/00_setup/agentos/check.py | 228 + .../05_tools/00_setup/agentos/config_mgr.py | 424 ++ .../05_tools/00_setup/agentos/const.py | 30 + .../05_tools/00_setup/agentos/init.py | 385 ++ .../05_tools/00_setup/agentos/install.sh | 69 + .../05_tools/00_setup/agentos/main.py | 238 + .../05_tools/00_setup/agentos/skill_mgr.py | 275 + .../05_tools/00_setup/agentos/sync.py | 199 + .../05_tools/00_setup/agentos/tool_mgr.py | 177 + .../05_tools/00_setup/agentos/upgrade.py | 473 ++ .../05_tools/00_setup/agentos/utils.py | 108 + .../05_tools/00_setup/guardd/README.md | 185 + .../guardd/com.agentos.chrome-debug.plist | 22 + .../00_setup/guardd/com.agentos.guardd.plist | 38 + .../05_tools/00_setup/guardd/guardd.py | 1886 +++++++ .../00_setup/guardd/install_guardd.sh | 118 + .../05_tools/00_setup/guardd/launch.sh | 12 + .../05_tools/00_setup/guardd/modules/.gitkeep | 0 .../00_setup/guardd/modules/__init__.py | 0 .../guardd/modules/account_monitor.py | 119 + .../00_setup/guardd/modules/executor.py | 268 + .../00_setup/guardd/modules/heartbeat.py | 169 + .../00_setup/guardd/modules/oracle_sync.py | 78 + .../00_setup/guardd/modules/priority_queue.py | 110 + .../guardd/modules/schedule_bridge.py | 135 + .../00_setup/guardd/modules/scheduler.py | 350 ++ .../00_setup/guardd/modules/slot_manager.py | 238 + .../00_setup/guardd/modules/task_store.py | 176 + .../05_tools/00_setup/guardd/scripts/.gitkeep | 0 .../00_setup/guardd/scripts/chrome_debug.sh | 113 + .../00_setup/guardd/scripts/chrome_hide.py | 54 + .../00_setup/guardd/scripts/install.sh | 124 + .../05_tools/00_setup/ssh_7kecheng_setup.md | 43 + .../05_tools/00_setup/sync_machine.sh | 76 + mac-agent-os-main/05_tools/01_system/.gitkeep | 0 .../01_system/check_automation_env.py | 135 + .../05_tools/01_system/check_facts.py | 43 + .../05_tools/01_system/cleanup_registry.py | 55 + .../05_tools/01_system/cluster_registry.py | 332 ++ .../05_tools/01_system/fix_7kecheng.py | 16 + .../05_tools/01_system/fix_all_paths.py | 114 + .../05_tools/01_system/fix_machine_names.py | 60 + .../05_tools/01_system/fix_ssh_auth.sh | 40 + .../05_tools/01_system/rebuild_registry.py | 49 + .../05_tools/01_system/role_check.py | 142 + .../05_tools/01_system/skill_scanner.py | 265 + .../05_tools/01_system/test_omlx_embedding.py | 158 + .../05_tools/01_system/trigger_matcher.py | 123 + .../05_tools/02_browser/.gitkeep | 0 .../05_tools/02_browser/README.md | 6 + mac-agent-os-main/05_tools/03_ocr/.gitkeep | 0 mac-agent-os-main/05_tools/03_ocr/README.md | 4 + mac-agent-os-main/05_tools/04_media/.gitkeep | 0 mac-agent-os-main/05_tools/04_media/README.md | 5 + mac-agent-os-main/05_tools/05_crawl/.gitkeep | 0 .../05_crawl/content-inspiration/README.md | 451 ++ .../05_crawl/content-inspiration/SKILL.md | 120 + .../05_crawl/content-inspiration/analyze.py | 260 + .../05_crawl/content-inspiration/app.py | 704 +++ .../05_crawl/content-inspiration/collect.py | 257 + .../05_crawl/content-inspiration/config.yaml | 57 + .../content-inspiration/data/database.db | Bin 0 -> 49152 bytes .../content-inspiration/docs/OPENCLI_SETUP.md | 42 + .../content-inspiration/doubao_driver.py | 243 + .../content-inspiration/downloader.py | 188 + .../install_opencli_extension.sh | 34 + .../content-inspiration/requirements.txt | 11 + .../05_crawl/content-inspiration/schema.sql | 74 + .../content-inspiration/script_factory.py | 84 + .../scripts_output/vid_20260504_8817.md | 18 + .../scripts_output/vid_20260504_8817.yaml | 63 + .../scripts_output/vid_20260505_em001.md | 54 + .../scripts_output/vid_20260505_em001.yaml | 109 + .../content-inspiration/setup_doubao_login.sh | 40 + .../content-inspiration/test_doubao.sh | 31 + .../05_crawl/content-inspiration/utils.py | 154 + .../05_crawl/longcat/longcat_claimer.py | 129 + .../05_crawl/longcat/socks5_forwarder.py | 135 + mac-agent-os-main/05_tools/06_mobile/.gitkeep | 0 .../05_tools/06_mobile/README.md | 3 + .../05_tools/07_matrix/.gitignore | 21 + .../.last_run/daily_douyin_browse.txt | 1 + .../07_matrix/.last_run/daily_xhs_like.txt | 1 + .../05_tools/07_matrix/ARCHITECTURE.md | 212 + .../05_tools/07_matrix/DOCUMENT_INDEX.md | 120 + .../05_tools/07_matrix/MODULE.md | 62 + .../05_tools/07_matrix/README.md | 222 + mac-agent-os-main/05_tools/07_matrix/TOOL.md | 278 + mac-agent-os-main/05_tools/07_matrix/agentos | 23 + .../backups/20260626_test_pass/corpus.py | 701 +++ .../20260626_test_pass/daily_comment.json | 30 + .../20260626_test_pass/douyin_comment.json | 40 + .../backups/20260626_test_pass/douyin_ops.py | 1862 ++++++ .../20260626_test_pass/douyin_reply.json | 36 + .../backups/20260626_test_pass/engine.py | 534 ++ .../backups/20260626_test_pass/interact.py | 298 + .../20260626_test_pass/interact_comment.json | 18 + .../blueprints/_archive/douyin_browse.json | 57 + .../blueprints/_archive/douyin_browse_v2.json | 57 + .../blueprints/_archive/douyin_browse_v3.json | 83 + .../blueprints/_archive/douyin_collect.json | 38 + .../_archive/douyin_comment_4videos.json | 25 + .../_archive/douyin_comment_interact.json | 44 + .../_archive/douyin_comment_quick.json | 14 + .../_archive/douyin_comment_test.json | 42 + .../_archive/douyin_nurture_v1.json | 88 + .../_archive/douyin_search_browse.json | 67 + .../export_douyin_01_20260601_233311.json | 27 + .../export_xhs_01_20260613_001448.json | 33 + .../_archive/xiaohongshu_nurture_v1.json | 19 + .../_archive/xiaohongshu_nurture_v2.json | 20 + .../_archive/xiaohongshu_read_profile.json | 18 + .../07_matrix/blueprints/daily_comment.json | 30 + .../blueprints/douyin_active_v1.json | 92 + .../07_matrix/blueprints/douyin_comment.json | 40 + .../07_matrix/blueprints/douyin_daily.json | 54 + .../blueprints/douyin_daily_clean.json | 101 + .../blueprints/douyin_read_profile.json | 155 + .../07_matrix/blueprints/douyin_reply.json | 36 + .../07_matrix/blueprints/douyin_search.json | 97 + .../07_matrix/blueprints/dy_step_test.json | 20 + .../07_matrix/blueprints/dy_test_all.json | 25 + .../export_douyin_01_20260601_233311.py | 38 + .../export_xhs_01_20260613_001448.py | 52 + .../07_matrix/blueprints/interact_chain.json | 24 + .../blueprints/interact_collect.json | 16 + .../blueprints/interact_comment.json | 18 + .../07_matrix/blueprints/interact_hot.json | 20 + .../07_matrix/blueprints/interact_like.json | 18 + .../07_matrix/blueprints/xhs_active_v1.json | 101 + .../07_matrix/blueprints/xhs_daily.json | 42 + .../07_matrix/blueprints/xhs_test_all.json | 25 + .../blueprints/xiaohongshu_read_profile.json | 68 + .../05_tools/07_matrix/config/schedule.yaml | 10 + .../config_template/accounts.override.yaml | 31 + .../07_matrix/config_template/accounts.yaml | 54 + .../07_matrix/config_template/ai.yaml | 13 + .../07_matrix/config_template/profiles.json | 27 + .../07_matrix/config_template/schedule.yaml | 31 + .../config_template/screen_layout.yaml | 39 + .../07_matrix/config_template/sms.yaml | 15 + .../07_matrix/config_template/tasks.yaml | 24 + .../05_tools/07_matrix/corpus/douyin.yaml | 427 ++ .../07_matrix/corpus/xiaohongshu.yaml | 140 + .../07_matrix/docs/ACCOUNT_LOGIN_SOP.md | 421 ++ .../07_matrix/docs/ANTI-DETECTION-PLAN.md | 119 + .../07_matrix/docs/BASE_CONVENTIONS.md | 76 + .../07_matrix/docs/BLUEPRINT-DESIGN.md | 111 + .../docs/CAMOUFOX_LOGIN_MANAGEMENT.md | 294 + .../07_matrix/docs/COLLECTION_RULES.md | 100 + .../docs/COLLECT_DISPLAY_SEPARATION.md | 62 + .../07_matrix/docs/COMMENT_FLOW_SPEC.md | 288 + .../05_tools/07_matrix/docs/CORPUS_V2_PLAN.md | 235 + .../07_matrix/docs/DEVELOPMENT_RULES.md | 96 + .../07_matrix/docs/DOUYIN_SELECTORS.md | 285 + .../07_matrix/docs/IDENTITY_FACTORY.md | 357 ++ .../07_matrix/docs/LOCAL_ADAPTATION.md | 420 ++ .../05_tools/07_matrix/docs/MANUAL.md | 221 + .../07_matrix/docs/MATRIX_V5_GUIDE.md | 323 ++ .../07_matrix/docs/MC_COMMAND_REFERENCE.md | 224 + .../05_tools/07_matrix/docs/SMS_API_USAGE.md | 57 + .../05_tools/07_matrix/docs/SYNC_GUIDE.md | 252 + .../07_matrix/docs/SYSTEM_ARCHITECTURE.md | 181 + .../07_matrix/fix_dashboard_launchd.sh | 74 + .../05_tools/07_matrix/init_matrix.sh | 496 ++ .../05_tools/07_matrix/install.sh | 156 + .../05_tools/07_matrix/local.yaml.template | 8 + mac-agent-os-main/05_tools/07_matrix/mc | 31 + .../05_tools/07_matrix/platforms/__init__.py | 40 + .../05_tools/07_matrix/platforms/base.py | 53 + .../07_matrix/platforms/douyin/SKILL.md | 19 + .../07_matrix/platforms/douyin/__init__.py | 7 + .../07_matrix/platforms/douyin/plugin.py | 103 + .../07_matrix/platforms/xiaohongshu/SKILL.md | 18 + .../platforms/xiaohongshu/__init__.py | 7 + .../07_matrix/platforms/xiaohongshu/plugin.py | 97 + .../05_tools/07_matrix/requirements.txt | 9 + .../07_matrix/scripts/ARCHITECTURE_AUDIT.md | 168 + .../07_matrix/scripts/agentos/__init__.py | 23 + .../07_matrix/scripts/agentos/__main__.py | 4 + .../07_matrix/scripts/agentos/base.py | 84 + .../05_tools/07_matrix/scripts/agentos/cli.py | 85 + .../07_matrix/scripts/agentos/plugins/ave.py | 152 + .../scripts/agentos/plugins/crawl.py | 150 + .../scripts/agentos/plugins/fleet.py | 123 + .../scripts/agentos/plugins/guardd.py | 162 + .../scripts/agentos/plugins/matrix.py | 599 ++ .../scripts/agentos/plugins/serve.py | 162 + .../07_matrix/scripts/anti_detection.py | 271 + .../07_matrix/scripts/auth_manager.py | 392 ++ .../07_matrix/scripts/browser_manager.py | 361 ++ .../07_matrix/scripts/browser_utils.py | 288 + .../07_matrix/scripts/c2/profile_scraper.py | 180 + .../07_matrix/scripts/cdp_connector.py | 628 +++ .../scripts/config/comment_corpus.yaml | 212 + .../07_matrix/scripts/config/schedule.yaml | 18 + .../07_matrix/scripts/config/vision.yaml | 14 + .../07_matrix/scripts/create_identity.py | 136 + .../07_matrix/scripts/data/schema.sql | 109 + .../07_matrix/scripts/detect_app_login.py | 353 ++ .../05_tools/07_matrix/scripts/douyin_ops.py | 2307 ++++++++ .../05_tools/07_matrix/scripts/guardd.py | 444 ++ .../05_tools/07_matrix/scripts/local_paths.py | 119 + .../07_matrix/scripts/login_identity.py | 180 + .../05_tools/07_matrix/scripts/matrix_mgmt.py | 1571 ++++++ .../scripts/matrix_modules/__init__.py | 7 + .../account/captcha/__init__.py | 10 + .../matrix_modules/account/captcha/base.py | 45 + .../matrix_modules/account/douyin_login.py | 167 + .../account/login_state_machine.py | 1140 ++++ .../matrix_modules/account/sms/__init__.py | 19 + .../scripts/matrix_modules/account/sms/api.py | 153 + .../matrix_modules/account/sms/base.py | 64 + .../matrix_modules/account/sms_login.py | 344 ++ .../matrix_modules/account/xhs_login.py | 630 +++ .../account/xiaohongshu_login.py | 246 + .../matrix_modules/comment/ai_generator.py | 295 + .../matrix_modules/comment/xhs/__init__.py | 6 + .../matrix_modules/comment/xhs/corpus.py | 321 ++ .../matrix_modules/nurture/__init__.py | 1 + .../matrix_modules/nurture/anchor_db.py | 242 + .../matrix_modules/nurture/anchor_rules.py | 222 + .../matrix_modules/nurture/behavior.py | 187 + .../matrix_modules/nurture/calibration.py | 294 + .../matrix_modules/nurture/comment_corpus.py | 109 + .../matrix_modules/nurture/data/anchor.db | Bin 0 -> 40960 bytes .../matrix_modules/nurture/data/schema.sql | 109 + .../scripts/matrix_modules/nurture/runner.py | 1947 +++++++ .../matrix_modules/nurture/ui_layout.py | 152 + .../matrix_modules/ops/xhs/ATOMIC_OPS.md | 387 ++ .../matrix_modules/ops/xhs/__init__.py | 6 + .../scripts/matrix_modules/ops/xhs/browse.py | 440 ++ .../matrix_modules/ops/xhs/interact.py | 462 ++ .../matrix_modules/ops/xhs/selectors.py | 367 ++ .../matrix_modules/utils/cookie_manager.py | 287 + .../05_tools/07_matrix/scripts/mc/__init__.py | 20 + .../05_tools/07_matrix/scripts/mc/__main__.py | 7 + .../05_tools/07_matrix/scripts/mc/analyzer.py | 462 ++ .../05_tools/07_matrix/scripts/mc/browser.py | 190 + .../05_tools/07_matrix/scripts/mc/cli.py | 1839 ++++++ .../05_tools/07_matrix/scripts/mc/corpus.py | 1159 ++++ .../05_tools/07_matrix/scripts/mc/engine.py | 718 +++ .../07_matrix/scripts/mc/execution_policy.py | 270 + .../05_tools/07_matrix/scripts/mc/exporter.py | 335 ++ .../05_tools/07_matrix/scripts/mc/interact.py | 352 ++ .../05_tools/07_matrix/scripts/mc/proxy.py | 99 + .../05_tools/07_matrix/scripts/mc/recorder.py | 730 +++ .../05_tools/07_matrix/scripts/mc/run.py | 83 + .../05_tools/07_matrix/scripts/mc/runlog.py | 64 + .../07_matrix/scripts/mc/scheduler.py | 308 + .../05_tools/07_matrix/scripts/mc/task.py | 183 + .../scripts/migrate_accounts_to_registry.py | 102 + .../07_matrix/scripts/nurture_blueprint.py | 298 + .../07_matrix/scripts/nurture_daily.py | 311 + .../07_matrix/scripts/ops/__init__.py | 5 + .../05_tools/07_matrix/scripts/ops/_base.py | 420 ++ .../05_tools/07_matrix/scripts/ops/xhs_ops.py | 655 +++ .../07_matrix/scripts/page_inspector.py | 284 + .../05_tools/07_matrix/scripts/page_state.py | 197 + .../07_matrix/scripts/pipelines/__init__.py | 7 + .../07_matrix/scripts/publish_video.py | 334 ++ .../05_tools/07_matrix/scripts/task_engine.py | 255 + .../07_matrix/scripts/task_scheduler.py | 132 + .../scripts/tools/comment_dom_scanner.py | 64 + .../07_matrix/scripts/tools/comment_sender.py | 253 + .../scripts/tools/reanalyze_xhs_dom.py | 370 ++ .../07_matrix/scripts/vision_bridge.py | 264 + .../07_matrix/setup_on_new_machine.sh | 143 + .../05_tools/07_matrix/start_dashboard.sh | 48 + .../05_tools/08_trae_agent/README.md | 85 + .../05_tools/08_trae_agent/colima_setup.md | 75 + .../08_trae_agent/install_trae_agent.sh | 119 + .../05_tools/08_trae_agent/trae_agent.sh | 121 + .../08_trae_agent/trae_config.yaml.example | 54 + .../09_ave/PLANS/API_STRATEGY_REPORT.md | 179 + .../09_ave/PLANS/AVE_IMPLEMENTATION_PLAN.md | 550 ++ .../09_ave/PLANS/AVE_V2_ARCHITECTURE.md | 250 + .../09_ave/PLANS/AVE_V3_ARCHITECTURE.md | 351 ++ .../05_tools/09_ave/PLANS/BGM_LIBRARY.md | 197 + .../05_tools/09_ave/PLANS/DASHBOARD_DESIGN.md | 237 + .../09_ave/PLANS/DASHBOARD_MULTI_AGENT.md | 408 ++ .../09_ave/PLANS/JIMENG_AUTOMATION_PLAN.md | 209 + .../PLANS/SPORTS_CHARACTER_PLANNING_v1.md | 196 + .../05_tools/09_ave/PLANS/STRATEGY.md | 92 + .../05_tools/09_ave/PLANS/TASK_TRACKER.md | 225 + mac-agent-os-main/05_tools/09_ave/TOOL.md | 260 + .../assets/demo_scripts/cat_tai_chi.txt | 22 + mac-agent-os-main/05_tools/09_ave/config.yaml | 77 + .../05_tools/09_ave/config/capabilities.yaml | 1601 ++++++ .../09_ave/config/cosyvoice_voices.yaml | 115 + .../09_ave/config/creative_matrix.yaml | 199 + .../09_ave/config/feasibility_rules.yaml | 372 ++ .../05_tools/09_ave/config/image_models.yaml | 126 + .../05_tools/09_ave/config/param_library.yaml | 786 +++ .../09_ave/config/portrait_presets.yaml | 88 + .../05_tools/09_ave/config/scene_presets.yaml | 49 + .../05_tools/09_ave/config/video_models.yaml | 74 + .../05_tools/09_ave/docs/api/README.md | 81 + .../05_tools/09_ave/docs/api/index.json | 1708 ++++++ ...e-image-generation-ai-tryon-create-task.md | 391 ++ ...-image-generation-ai-tryon-query-result.md | 251 + ...erence-image-generation-aitryon-parsing.md | 209 + ...-generation-aitryon-refiner-create-task.md | 260 + ...generation-aitryon-refiner-query-result.md | 215 + ...ation-background-generation-create-task.md | 496 ++ ...tion-background-generation-query-result.md | 354 ++ ...-generation-creative-poster-create-task.md | 337 ++ ...generation-creative-poster-query-result.md | 348 ++ ...neration-facechain-finetune-create-task.md | 408 ++ ...eration-facechain-finetune-query-result.md | 392 ++ ...ration-facechain-generation-create-task.md | 452 ++ ...facechain-generation-facechain-face-det.md | 231 + ...ation-facechain-generation-query-result.md | 343 ++ ...tion-image-erase-completion-create-task.md | 235 + ...ion-image-erase-completion-query-result.md | 226 + ...neration-image-out-painting-create-task.md | 435 ++ ...eration-image-out-painting-query-result.md | 332 ++ ...tion-kling-image-generation-create-task.md | 349 ++ ...ion-kling-image-generation-query-result.md | 308 + ...-image-generation-kling-subject-id-list.md | 169 + ...person-instance-segmentation-create-tas.md | 311 + ...person-instance-segmentation-query-resu.md | 264 + ...nce-image-generation-qwen-image-editing.md | 660 +++ ...-generation-qwen-text-to-image-30-async.md | 662 +++ ...age-generation-qwen-text-to-image-async.md | 649 +++ ...ge-generation-qwen-text-to-image-openai.md | 332 ++ ...eneration-qwen-text-to-image-task-query.md | 600 ++ ...nce-image-generation-qwen-text-to-image.md | 666 +++ ...image-generation-shoe-model-create-task.md | 266 + ...mage-generation-shoe-model-query-result.md | 285 + ...ation-vidu-image-generation-create-task.md | 400 ++ ...tion-vidu-image-generation-query-result.md | 362 ++ ...ge-generation-virtual-model-create-task.md | 467 ++ ...e-generation-virtual-model-query-result.md | 429 ++ ...ration-wan-text-to-image-v2-create-task.md | 940 ++++ ...ation-wan-text-to-image-v2-query-result.md | 536 ++ ...ration-wan-text-to-image-v2-synchronous.md | 561 ++ ...wan21-general-image-editing-create-task.md | 788 +++ ...wan21-general-image-editing-query-resul.md | 371 ++ ...wan25-general-image-editing-create-task.md | 694 +++ ...wan25-general-image-editing-query-resul.md | 334 ++ ...ration-wan26-image-gen-edit-create-task.md | 795 +++ ...ation-wan26-image-gen-edit-query-result.md | 492 ++ ...ration-wan26-image-gen-edit-synchronous.md | 702 +++ ...ration-wan27-image-gen-edit-create-task.md | 863 +++ ...ation-wan27-image-gen-edit-query-result.md | 477 ++ ...ration-wan27-image-gen-edit-synchronous.md | 814 +++ ...ration-wanx-sketch-to-image-create-task.md | 482 ++ ...ation-wanx-sketch-to-image-query-result.md | 338 ++ ...neration-wanx-style-repaint-create-task.md | 365 ++ ...eration-wanx-style-repaint-query-result.md | 384 ++ ...ation-wanx-v1-text-to-image-create-task.md | 816 +++ ...tion-wanx-v1-text-to-image-query-result.md | 320 ++ ...-generation-wanx-x-painting-create-task.md | 548 ++ ...generation-wanx-x-painting-query-result.md | 369 ++ ...generation-wordart-semantic-create-task.md | 276 + ...eneration-wordart-semantic-query-result.md | 287 + ...-generation-wordart-texture-create-task.md | 376 ++ ...generation-wordart-texture-query-result.md | 328 ++ .../api-reference-image-generation-z-image.md | 308 + ...e-translation-qwen-mt-image-create-task.md | 282 + ...-translation-qwen-mt-image-query-result.md | 316 ++ ...e-translation-qwen-mt-image-synchronous.md | 250 + ...mbedding-dashscope-multimodal-embedding.md | 721 +++ ...al-translation-qwen-mt-uni-query-result.md | 459 ++ ...modal-translation-qwen-mt-uni-translate.md | 573 ++ ...ence-real-time-multimodal-client-events.md | 899 +++ ...eal-time-multimodal-interaction-process.md | 136 + ...rence-real-time-multimodal-model-access.md | 39 + ...-real-time-multimodal-realtime-java-sdk.md | 239 + ...eal-time-multimodal-realtime-python-sdk.md | 277 + ...ence-real-time-multimodal-server-events.md | 1375 +++++ ...ence-real-time-multimodal-voice-cloning.md | 1469 +++++ ...o-generation-animate-anyone-create-task.md | 302 + ...-generation-animate-anyone-query-result.md | 279 + ...animate-anyone-template-gen-create-task.md | 230 + ...animate-anyone-template-gen-query-resul.md | 235 + ...generation-animate-anyone-wan-aa-detect.md | 202 + ...-video-generation-emo-video-create-task.md | 269 + ...e-video-generation-emo-video-emo-detect.md | 259 + ...video-generation-emo-video-query-result.md | 239 + ...ideo-generation-emoji-video-create-task.md | 330 ++ ...deo-generation-emoji-video-emoji-detect.md | 281 + ...deo-generation-emoji-video-query-result.md | 320 ++ ...n-happyhorse-image-to-video-create-task.md | 270 + ...-happyhorse-image-to-video-query-result.md | 295 + ...happyhorse-reference-to-video-create-ta.md | 297 + ...happyhorse-reference-to-video-query-res.md | 314 ++ ...on-happyhorse-text-to-video-create-task.md | 261 + ...n-happyhorse-text-to-video-query-result.md | 292 + ...on-happyhorse-video-editing-create-task.md | 273 + ...n-happyhorse-video-editing-query-result.md | 295 + ...tion-kling-video-generation-create-task.md | 498 ++ ...ion-kling-video-generation-query-result.md | 353 ++ ...neration-liveportrait-video-create-task.md | 309 + ...-liveportrait-video-liveportrait-detect.md | 227 + ...eration-liveportrait-video-query-result.md | 280 + ...generation-minimax-h3-video-create-task.md | 440 ++ ...eneration-minimax-h3-video-query-result.md | 308 + ...ion-pixverse-image-to-video-create-task.md | 369 ++ ...pixverse-image-to-video-first-frame-cre.md | 478 ++ ...pixverse-image-to-video-first-frame-que.md | 330 ++ ...pixverse-image-to-video-first-last-crea.md | 355 ++ ...pixverse-image-to-video-first-last-quer.md | 327 ++ ...on-pixverse-image-to-video-query-result.md | 342 ++ ...generation-pixverse-lipsync-create-task.md | 335 ++ ...eneration-pixverse-lipsync-query-result.md | 269 + ...tion-pixverse-motioncontrol-create-task.md | 305 + ...ion-pixverse-motioncontrol-query-result.md | 273 + ...pixverse-reference-to-video-create-task.md | 482 ++ ...pixverse-reference-to-video-query-resul.md | 363 ++ ...tion-pixverse-text-to-video-create-task.md | 381 ++ ...ion-pixverse-text-to-video-query-result.md | 335 ++ ...generation-pixverse-upscale-create-task.md | 296 + ...eneration-pixverse-upscale-query-result.md | 268 + ...deo-generation-video-retalk-create-task.md | 338 ++ ...eo-generation-video-retalk-query-result.md | 280 + ...ation-video-style-transform-create-task.md | 343 ++ ...tion-video-style-transform-query-result.md | 318 ++ ...eration-vidu-image-to-video-create-task.md | 335 ++ ...vidu-image-to-video-first-frame-create-.md | 419 ++ ...vidu-image-to-video-first-frame-query-r.md | 323 ++ ...ration-vidu-image-to-video-query-result.md | 344 ++ ...ion-vidu-reference-to-video-create-task.md | 662 +++ ...on-vidu-reference-to-video-query-result.md | 495 ++ ...ion-vidu-start-end-to-video-create-task.md | 404 ++ ...on-vidu-start-end-to-video-query-result.md | 361 ++ ...neration-vidu-text-to-video-create-task.md | 548 ++ ...eration-vidu-text-to-video-query-result.md | 374 ++ ...n-wan-general-video-editing-create-task.md | 1349 +++++ ...-wan-general-video-editing-query-result.md | 347 ++ ...tion-wan-image-to-animation-create-task.md | 264 + ...ion-wan-image-to-animation-query-result.md | 289 + ...wan-image-to-video-first-frame-create-t.md | 762 +++ ...wan-image-to-video-first-frame-query-re.md | 373 ++ ...wan-image-to-video-first-frame-video-ef.md | 149 + ...wan-image-to-video-first-last-frames-cr.md | 564 ++ ...wan-image-to-video-first-last-frames-qu.md | 309 + ...tion-wan-reference-to-video-create-task.md | 286 + ...ion-wan-reference-to-video-query-result.md | 264 + ...ce-video-generation-wan-s2v-create-task.md | 296 + ...e-video-generation-wan-s2v-query-result.md | 259 + ...video-generation-wan-s2v-wan-s2v-detect.md | 206 + ...eneration-wan-text-to-video-create-task.md | 491 ++ ...neration-wan-text-to-video-query-result.md | 287 + ...on-wan-video-character-swap-create-task.md | 267 + ...n-wan-video-character-swap-query-result.md | 289 + ...ration-wan27-image-to-video-create-task.md | 580 ++ ...ation-wan27-image-to-video-query-result.md | 322 ++ ...on-wan27-reference-to-video-create-task.md | 627 +++ ...n-wan27-reference-to-video-query-result.md | 352 ++ ...eration-wan27-text-to-video-create-task.md | 544 ++ ...ration-wan27-text-to-video-query-result.md | 310 + ...eration-wan27-video-editing-create-task.md | 659 +++ ...ration-wan27-video-editing-query-result.md | 341 ++ ...ideo-generation-wan30-video-create-task.md | 759 +++ ...deo-generation-wan30-video-query-result.md | 317 ++ ...per-guides-getting-started-image-models.md | 171 + ...per-guides-getting-started-video-models.md | 232 + ...er-guides-getting-started-vision-models.md | 215 + ...-image-generation-background-generation.md | 340 ++ ...guides-image-generation-creative-poster.md | 307 + ...r-guides-image-generation-image-editing.md | 989 ++++ ...image-generation-image-erase-completion.md | 323 ++ ...uides-image-generation-image-inpainting.md | 367 ++ ...des-image-generation-image-out-painting.md | 516 ++ ...generation-person-instance-segmentation.md | 110 + ...oper-guides-image-generation-shoe-model.md | 202 + ...guides-image-generation-sketch-to-image.md | 213 + ...r-guides-image-generation-text-to-image.md | 882 +++ ...r-guides-image-generation-virtual-model.md | 368 ++ ...ides-image-generation-wan-image-editing.md | 1624 ++++++ ...es-image-generation-wan21-image-editing.md | 624 +++ .../qwen/developer-guides-multimodal-ocr.md | 2582 +++++++++ .../developer-guides-multimodal-vision.md | 3031 ++++++++++ .../developer-guides-speech-asr-realtime.md | 3462 ++++++++++++ .../raw/qwen/developer-guides-speech-asr.md | 2225 ++++++++ ...eveloper-guides-speech-audio-generation.md | 228 + ...eveloper-guides-speech-file-translation.md | 723 +++ ...des-speech-improve-recognition-accuracy.md | 599 ++ ...veloper-guides-speech-multimodal-speech.md | 2200 ++++++++ ...eveloper-guides-speech-music-generation.md | 570 ++ .../developer-guides-speech-omni-models.md | 73 + ...developer-guides-speech-omni-voice-list.md | 308 + ...loper-guides-speech-qwen-audio-realtime.md | 1501 +++++ ...uides-speech-realtime-multimodal-speech.md | 3674 ++++++++++++ ...eloper-guides-speech-realtime-streaming.md | 4878 ++++++++++++++++ ...oper-guides-speech-realtime-translation.md | 1162 ++++ .../developer-guides-speech-s2s-models.md | 301 + ...per-guides-speech-speech-to-text-models.md | 227 + .../raw/qwen/developer-guides-speech-ssml.md | 2661 +++++++++ .../developer-guides-speech-tts-models.md | 169 + .../raw/qwen/developer-guides-speech-tts.md | 474 ++ .../developer-guides-speech-voice-cloning.md | 749 +++ .../developer-guides-speech-voice-design.md | 328 ++ ...oper-guides-speech-voice-list-cosyvoice.md | 261 + ...guides-speech-voice-list-qwen-audio-tts.md | 539 ++ ...loper-guides-speech-voice-list-qwen-tts.md | 113 + ...eo-generation-image-to-video-first-last.md | 623 ++ ...-guides-video-generation-image-to-video.md | 1093 ++++ ...guides-video-generation-reference-video.md | 624 +++ ...r-guides-video-generation-text-to-video.md | 1181 ++++ ...r-guides-video-generation-video-editing.md | 1539 +++++ ...per-guides-video-generation-wan30-video.md | 1555 +++++ .../character_consistency/01_skill_原文.md | 245 + .../02_提示词规范_原文.md | 100 + .../03_角色卡schema_原文.md | 86 + .../character_consistency/04_字段拆解_原文.md | 101 + .../docs/character_consistency/README.md | 40 + .../05_tools/09_ave/docs/cosyvoice_voices.md | 86 + .../05_tools/09_ave/docs/gap_engine/README.md | 64 + .../gap_engine/对话原始JSON_2026-10-02.json | 1 + .../docs/gap_engine/对话原文_2026-10-02.md | 1933 +++++++ .../00_业界生产流程_关键帧先行.md | 166 + .../01_关键帧合成_设计方案.md | 143 + .../02_百炼可用能力与价格.md | 161 + .../03_千问AI平台_OpenAPI字段规格.md | 184 + .../04_看板前端改造方案_草案.md | 94 + .../05_图像生成与编辑能力调研.md | 121 + .../06_关键帧环节_流程与界面设计.md | 132 + .../07_关键帧工作台_设计方案.md | 160 + .../08_关键帧转视频_设计方案.md | 259 + .../09_导演工作流整合方案.md | 252 + .../10_按段落交接改造方案.md | 177 + .../pipeline_research/11_首尾帧工作流建议.md | 111 + .../12_千问代理的可灵通道_规格对照.md | 90 + .../13_千问平台可用能力总览.md | 227 + .../09_ave/docs/realism/02-进阶公式.md | 122 + .../09_ave/docs/realism/04-情绪外化表.md | 48 + .../09_ave/docs/realism/06-约束词清单.md | 94 + .../09_ave/docs/realism/07-特殊字符规范.md | 41 + .../09_ave/docs/realism/08-避坑12问.md | 194 + .../09_ave/docs/realism/09-kling-公式.md | 146 + .../20-realistic-character-consistency.md | 398 ++ .../05_tools/09_ave/docs/realism/README.md | 25 + .../09_ave/docs/realism/anti-ai-checklist.md | 137 + .../docs/realism/atmosphere-dictionary.md | 174 + .../09_ave/docs/realism/camera-aesthetics.md | 229 + .../09_ave/docs/realism/seedance_SKILL.md | 113 + .../09_ave/docs/realism/smixs-README.md | 306 + .../docs/业界参考/创意与设定的标准分层.md | 164 + .../docs/业界参考/剧本与分镜的分层标准.md | 158 + .../05_tools/09_ave/docs/参数速查.md | 179 + mac-agent-os-main/05_tools/09_ave/install.sh | 56 + .../05_tools/09_ave/requirements.txt.pip | 13 + .../09_ave/scripts/AVE_ARCHITECTURE_PLAN.md | 480 ++ .../05_tools/09_ave/scripts/__init__.py | 0 .../09_ave/scripts/_generate_frames.py | 69 + .../scripts/anchor_extractor/__init__.py | 0 .../scripts/anchor_extractor/extractor.py | 155 + .../09_ave/scripts/asset_manager/__init__.py | 14 + .../09_ave/scripts/asset_manager/cache.py | 367 ++ .../09_ave/scripts/asset_manager/index.py | 469 ++ .../09_ave/scripts/asset_manager/tags.py | 317 ++ .../09_ave/scripts/assets_manager/__init__.py | 421 ++ .../09_ave/scripts/audio_line/__init__.py | 198 + .../09_ave/scripts/audio_line/config.yaml | 85 + .../audio_line/layer1_input/__init__.py | 0 .../audio_line/layer1_input/_parser.py | 52 + .../audio_line/layer1_input/collector.py | 133 + .../audio_line/layer2_analysis/__init__.py | 0 .../layer2_analysis/beat_detector.py | 180 + .../layer2_analysis/emotion_analyzer.py | 166 + .../layer2_analysis/stem_separator.py | 98 + .../layer2_analysis/structure_parser.py | 207 + .../audio_line/layer3_creation/__init__.py | 12 + .../layer3_creation/module_a_rhythm.py | 227 + .../layer3_creation/module_b_lyrics.py | 174 + .../layer3_creation/module_c_melody.py | 135 + .../layer3_creation/module_d_mix.py | 291 + .../audio_line/layer4_alignment/__init__.py | 8 + .../layer4_alignment/lyric_aligner.py | 309 + .../layer4_alignment/stem_aligner.py | 392 ++ .../audio_line/layer5_output/__init__.py | 8 + .../audio_line/layer5_output/exporter.py | 282 + .../layer5_output/lyric_formatter.py | 226 + .../09_ave/scripts/audio_line/orchestrator.py | 434 ++ .../09_ave/scripts/bgm_generator/__init__.py | 10 + .../scripts/bgm_generator/bgm_download.py | 203 + .../09_ave/scripts/bgm_generator/chord_pad.py | 164 + .../09_ave/scripts/bgm_generator/suno.py | 273 + .../09_ave/scripts/capabilities/__init__.py | 232 + .../scripts/capabilities/dashscope_video.py | 350 ++ .../scripts/capabilities/kling_video.py | 604 ++ .../09_ave/scripts/capability_doctor.py | 231 + .../09_ave/scripts/capability_matcher.py | 382 ++ .../scripts/character_adapter/__init__.py | 397 ++ .../scripts/character_generator/__init__.py | 11 + .../character_generator/asset_registrar.py | 304 + .../attribute_extractor.py | 276 + .../character_generator/direction_expander.py | 132 + .../scripts/character_generator/pipeline.py | 417 ++ .../character_generator/portrait_set.py | 370 ++ .../character_generator/prompt_assembler.py | 866 +++ .../prompts/character_sheet_prompts.md | 89 + .../character_generator/quality_check.py | 361 ++ .../character_generator/variant_generator.py | 351 ++ .../09_ave/scripts/character_portrait.py | 790 +++ .../scripts/character_registry/__init__.py | 318 ++ .../scripts/character_registry/registry.yaml | 700 +++ .../09_ave/scripts/character_sheet.py | 1152 ++++ .../09_ave/scripts/composer/__init__.py | 0 .../05_tools/09_ave/scripts/composer/align.py | 136 + .../09_ave/scripts/composer/beat_sync.py | 678 +++ .../scripts/composer/character_locker.py | 68 + .../05_tools/09_ave/scripts/composer/de_ai.py | 181 + .../scripts/composer/director_styles.py | 233 + .../09_ave/scripts/composer/dubbing.py | 607 ++ .../09_ave/scripts/composer/ffmpeg.py | 699 +++ .../09_ave/scripts/composer/frames.py | 62 + .../09_ave/scripts/composer/hybrid.py | 457 ++ .../09_ave/scripts/composer/lipsync.py | 508 ++ .../09_ave/scripts/composer/pipeline.py | 1274 +++++ .../05_tools/09_ave/scripts/composer/qa.py | 111 + .../09_ave/scripts/composer/realism.py | 801 +++ .../09_ave/scripts/composer/speed_ramp.py | 376 ++ .../09_ave/scripts/daoist_quotes.yaml | 114 + .../09_ave/scripts/demo_full_pipeline.py | 256 + .../scripts/director_parser/__init__.py | 0 .../09_ave/scripts/director_parser/parser.py | 192 + .../09_ave/scripts/director_parser/schemas.py | 131 + .../09_ave/scripts/director_script.yaml | 74 + .../09_ave/scripts/directors/__init__.py | 75 + .../09_ave/scripts/directors/_shot_builder.py | 163 + .../05_tools/09_ave/scripts/directors/base.py | 304 + .../scripts/directors/legacy_scene_planner.py | 274 + .../05_tools/09_ave/scripts/directors/llm.py | 141 + .../scripts/directors/seedance_director.py | 385 ++ .../scripts/directors/storyboard_v52.py | 424 ++ .../09_ave/scripts/gap_engine/__init__.py | 40 + .../scripts/gap_engine/combo_generator.py | 453 ++ .../scripts/gap_engine/gap_calculator.py | 168 + .../09_ave/scripts/gap_engine/library.py | 110 + .../09_ave/scripts/gap_engine/models.py | 160 + .../09_ave/scripts/gap_engine/store.py | 122 + .../09_ave/scripts/inspire_quotes.yaml | 91 + .../05_tools/09_ave/scripts/keyframe_gen.py | 283 + .../05_tools/09_ave/scripts/kling_report.py | 88 + .../05_tools/09_ave/scripts/lib/__init__.py | 0 .../05_tools/09_ave/scripts/lib/config.py | 47 + .../09_ave/scripts/lib/cost_tracker.py | 213 + .../05_tools/09_ave/scripts/lib/dashboard.py | 498 ++ .../05_tools/09_ave/scripts/lib/ffmpeg.py | 131 + .../09_ave/scripts/lib/ghvideo_upload.py | 135 + .../05_tools/09_ave/scripts/lib/http.py | 93 + .../09_ave/scripts/lib/kling_webhook.py | 151 + .../05_tools/09_ave/scripts/lib/logger.py | 28 + .../05_tools/09_ave/scripts/main.py | 931 +++ .../scripts/material_producer/__init__.py | 0 .../material_producer/fallback/__init__.py | 0 .../material_producer/kling/__init__.py | 0 .../scripts/material_producer/kling/kling.py | 237 + .../material_producer/pexels/__init__.py | 5 + .../material_producer/pexels/search.py | 337 ++ .../material_producer/wan2_2/__init__.py | 6 + .../material_producer/wan2_2/avatar.py | 25 + .../material_producer/wan2_2/wan2_2.py | 340 ++ .../09_ave/scripts/music_selector/__init__.py | 371 ++ .../scripts/music_selector/music_library.yaml | 81 + .../09_ave/scripts/nature_clouds.yaml | 79 + .../09_ave/scripts/person_swap/__init__.py | 18 + .../09_ave/scripts/person_swap/api.py | 324 ++ .../09_ave/scripts/person_swap/preprocess.py | 244 + .../09_ave/scripts/person_swap/service.py | 376 ++ .../scripts/pipeline_controller/__init__.py | 910 +++ .../prompts/character_sheet_prompts.md | 320 ++ .../scripts/scene_generator/__init__.py | 0 .../scripts/scene_generator/generator.py | 195 + .../scene_generator/prompt_assembler.py | 397 ++ .../05_tools/09_ave/scripts/schemas.py | 418 ++ .../scripts/script_generator/__init__.py | 377 ++ .../templates/story_prompts.md | 81 + .../09_ave/scripts/script_schemas/__init__.py | 275 + .../scripts/script_schemas/script_schema.yaml | 57 + .../09_ave/scripts/service_layer/__init__.py | 0 .../09_ave/scripts/service_layer/app.py | 271 + .../story_director/HOLOCINE_REFERENCE.md | 108 + .../09_ave/scripts/story_director/__init__.py | 13 + .../scripts/story_director/batch_generator.py | 664 +++ .../scripts/story_director/scene_planner.py | 444 ++ .../scripts/story_director/temporal_bridge.py | 322 ++ .../05_tools/09_ave/scripts/test_tai_chi.yaml | 90 + .../scripts/tools/extend_capabilities.py | 343 ++ .../09_ave/scripts/tools/fetch_api_docs.py | 199 + .../05_tools/09_ave/scripts/video_factory.py | 641 +++ .../scripts/voice_synthesizer/__init__.py | 0 .../scripts/voice_synthesizer/aliyun.py | 263 + .../scripts/voice_synthesizer/volcano.py | 191 + .../05_tools/09_ave/skills/README.md | 37 + .../skills/ai-film-skills/COMPATIBILITY.md | 55 + .../skills/ai-film-skills/INSTALLATION.md | 68 + .../09_ave/skills/ai-film-skills/LICENSE | 177 + .../09_ave/skills/ai-film-skills/README.md | 375 ++ .../skills/ai-film-skills/SKILL_CATALOG.md | 117 + .../skills/ai-film-skills/docs/WORKFLOW.md | 93 + .../skills/zh-CN/ai-storyboard-director.md | 126 + .../docs/skills/zh-CN/character-asset.md | 115 + .../docs/skills/zh-CN/director-agent.md | 122 + .../docs/skills/zh-CN/prop-asset.md | 114 + .../docs/skills/zh-CN/scene-asset.md | 114 + .../guofeng-visual-director/SKILL.md | 93 + .../references/cultural-grounding.md | 45 + .../references/image-direction.md | 69 + .../references/people-and-objects.md | 41 + .../references/prompt-and-review.md | 60 + .../references/world-and-space.md | 42 + .../hard-sci-fi-visual-director/SKILL.md | 55 + .../references/aesthetic-audit.md | 164 + .../references/cinematic-image-direction.md | 35 + .../references/combat-mecha-aesthetic.md | 225 + .../references/combat-mecha-form-direction.md | 188 + .../references/full-design-workflow.md | 197 + .../references/future-interface-systems.md | 179 + .../references/future-weapon-systems.md | 138 + .../references/inspiration-engine.md | 165 + .../references/live-action-cinematography.md | 116 + .../references/organism-and-ecology.md | 88 + .../references/physics-and-engineering.md | 111 + .../references/production-design.md | 109 + .../references/prompt-examples.md | 186 + .../references/scene-routes.md | 113 + .../references/scene-world-design.md | 134 + .../references/script-to-visual-derivation.md | 169 + .../references/style-color-system.md | 129 + .../references/visual-continuity.md | 163 + .../whitebox-previs-executor/SKILL.md | 79 + .../references/previs-spec.md | 142 + .../references/prompt-to-previs.md | 162 + .../skills/ai-short-drama-production/SKILL.md | 104 + .../references/SOURCE-LEDGER.md | 47 + .../references/control-contracts.md | 49 + .../references/independent-production-core.md | 47 + .../references/production-handoff.md | 61 + .../skills/ai-storyboard-director/SKILL.md | 57 + .../references/camera-motion-diagnostics.md | 49 + .../cinematography-design-engine.md | 143 + .../references/creative-shot-ideas.md | 25 + .../references/delivery-mode-guard.md | 19 + .../references/design-memory-protocol.md | 136 + .../references/fight-design.md | 88 + .../references/fight-reference-case.md | 17 + .../references/framing-and-axis.md | 115 + .../references/production-contract.md | 77 + .../references/shot-design-engine.md | 64 + .../skills/character-asset/SKILL.md | 197 + .../references/NEGATIVE-CASE-BOOK.md | 20 + .../references/cinematic-image-direction.md | 70 + .../references/female-character-charm.md | 58 + .../skills/cyberpunk-design/SKILL.md | 40 + .../d-data-analysis-semantic-layer/SKILL.md | 37 + .../references/semantic-layer.md | 73 + .../references/semantic-write-contract.md | 37 + .../references/source-inventory.md | 27 + .../references/versioning-and-expiry.md | 19 + .../d-official-market-analysis/SKILL.md | 46 + .../references/analysis-and-report.md | 52 + .../analysis-capability-contract.md | 42 + .../references/connector-contract.md | 13 + .../references/data-contract.md | 61 + .../references/evidence-and-sources.md | 35 + .../references/full-market-research.md | 67 + .../references/platform-metrics.md | 20 + .../skills/director-agent/SKILL.md | 77 + .../references/anti-laziness-contract.md | 129 + .../references/director-thinking-spine.md | 66 + .../references/director-workbench-protocol.md | 285 + .../references/github-project-watchlist.md | 89 + .../references/local-knowledge-map.md | 49 + .../production-storyboard-compiler.md | 77 + .../references/research-update-protocol.md | 128 + .../screenplay-ai-execution-compiler.md | 162 + .../screenplay-cold-read-protocol.md | 295 + .../screenplay-exemplar-benchmarks.md | 82 + .../references/screenplay-state-engine.md | 291 + .../references/screenplay-writing-core.md | 145 + .../references/verified-director-logic.md | 240 + .../skills/epic-design/SKILL.md | 40 + .../references/COMMON-12-SECTION-PROTOCOL.md | 80 + .../references/NEGATIVE-CASE-BOOK.md | 18 + .../epic-design/references/SOURCE-LEDGER.md | 39 + .../references/cinematic-image-direction.md | 70 + .../epic-design/references/genre-presets.md | 54 + .../skills/fantasy-design/SKILL.md | 40 + .../references/COMMON-12-SECTION-PROTOCOL.md | 80 + .../references/NEGATIVE-CASE-BOOK.md | 18 + .../references/SOURCE-LEDGER.md | 39 + .../references/cinematic-image-direction.md | 70 + .../references/genre-presets.md | 54 + .../skills/horror-design/SKILL.md | 40 + .../references/COMMON-12-SECTION-PROTOCOL.md | 80 + .../references/NEGATIVE-CASE-BOOK.md | 18 + .../horror-design/references/SOURCE-LEDGER.md | 39 + .../references/cinematic-image-direction.md | 70 + .../horror-design/references/genre-presets.md | 54 + .../skills/noir-design/SKILL.md | 40 + .../references/COMMON-12-SECTION-PROTOCOL.md | 80 + .../references/NEGATIVE-CASE-BOOK.md | 18 + .../noir-design/references/SOURCE-LEDGER.md | 39 + .../references/cinematic-image-direction.md | 70 + .../noir-design/references/genre-presets.md | 54 + .../skills/produce-ai-video/SKILL.md | 74 + .../autonomous-production-workflow.md | 135 + .../references/qualified-video-acceptance.md | 92 + .../references/storyboard-prompt-compiler.md | 89 + .../ai-film-skills/skills/prop-asset/SKILL.md | 186 + .../references/NEGATIVE-CASE-BOOK.md | 20 + .../references/cinematic-image-direction.md | 70 + .../skills/romance-design/SKILL.md | 40 + .../references/COMMON-12-SECTION-PROTOCOL.md | 80 + .../references/NEGATIVE-CASE-BOOK.md | 18 + .../references/SOURCE-LEDGER.md | 39 + .../references/cinematic-image-direction.md | 70 + .../references/genre-presets.md | 54 + .../skills/scene-asset/SKILL.md | 205 + .../references/NEGATIVE-CASE-BOOK.md | 20 + .../references/cinematic-image-direction.md | 70 + .../ai-film-skills/skills/war-design/SKILL.md | 102 + .../references/COMMON-12-SECTION-PROTOCOL.md | 88 + .../references/NEGATIVE-CASE-BOOK.md | 26 + .../war-design/references/SOURCE-LEDGER.md | 23 + .../references/cinematic-image-direction.md | 70 + .../war-design/references/combat-visual.md | 51 + .../skills/war-design/references/sources.md | 33 + .../references/squad-cinematography.md | 48 + .../war-design/references/story-visual.md | 61 + .../war-design/references/visual-design.md | 93 + .../war-design/references/visual-review.md | 39 + .../references/war-visual-presets.md | 77 + .../skills/web-design-director/SKILL.md | 105 + .../references/creative-direction.md | 71 + .../references/design-rubric.md | 57 + .../references/interaction-design.md | 90 + .../references/material-and-motion.md | 51 + .../references/research-to-design.md | 72 + .../references/web-quality-checklist.md | 72 + .../skills/wuxia-design/SKILL.md | 40 + .../references/COMMON-12-SECTION-PROTOCOL.md | 80 + .../references/NEGATIVE-CASE-BOOK.md | 18 + .../wuxia-design/references/SOURCE-LEDGER.md | 39 + .../references/cinematic-image-direction.md | 70 + .../wuxia-design/references/genre-presets.md | 54 + .../09_ave/skills/seedance-director/SKILL.md | 166 + .../page-2026-06-16T06-20-53-028Z.yml | 0 .../page-2026-06-16T06-20-55-501Z.yml | 88 + .../page-2026-06-16T06-29-26-329Z.yml | 66 + .../page-2026-06-17T11-10-46-764Z.yml | 73 + .../page-2026-06-17T11-12-53-973Z.yml | 73 + .../page-2026-06-17T11-13-50-422Z.yml | 73 + .../page-2026-06-17T11-39-48-242Z.yml | 73 + .../page-2026-06-17T11-52-17-980Z.yml | 73 + .../page-2026-06-17T12-02-00-772Z.yml | 73 + .../page-2026-06-17T12-10-23-589Z.yml | 73 + .../page-2026-06-17T12-13-43-409Z.yml | 73 + .../page-2026-06-17T12-14-54-855Z.yml | 73 + .../page-2026-06-17T13-25-02-710Z.yml | 27 + .../page-2026-06-17T13-26-00-416Z.yml | 27 + .../page-2026-06-17T13-26-49-605Z.yml | 27 + .../page-2026-06-17T13-28-43-625Z.yml | 27 + .../page-2026-06-17T13-29-36-213Z.yml | 27 + .../page-2026-06-17T13-30-22-011Z.yml | 27 + .../page-2026-06-17T13-32-57-345Z.yml | 70 + .../page-2026-06-17T13-33-14-434Z.yml | 70 + .../page-2026-06-17T13-34-03-709Z.yml | 27 + .../page-2026-06-17T13-35-49-778Z.yml | 73 + .../page-2026-06-17T15-04-21-111Z.yml | 27 + .../PLANS/BUSINESS_ARCHITECTURE_v4.md | 225 + .../05_tools/10_dashboard/PLANS/DEPLOYMENT.md | 164 + .../05_tools/10_dashboard/README.md | 296 + .../05_tools/10_dashboard/__init__.py | 2 + .../05_tools/10_dashboard/app.py | 1804 ++++++ .../05_tools/10_dashboard/frontend/.gitignore | 24 + .../frontend/.vite/deps/_metadata.json | 8 + .../frontend/.vite/deps/package.json | 3 + .../05_tools/10_dashboard/frontend/index.html | 119 + .../10_dashboard/frontend/package-lock.json | 872 +++ .../10_dashboard/frontend/package.json | 14 + .../10_dashboard/frontend/public/favicon.svg | 1 + .../10_dashboard/frontend/public/icons.svg | 24 + .../05_tools/10_dashboard/frontend/src/api.js | 128 + .../frontend/src/assets/javascript.svg | 1 + .../10_dashboard/frontend/src/assets/vite.svg | 1 + .../src/components/account-selector.js | 490 ++ .../src/components/execution-pipeline.js | 300 + .../frontend/src/components/task-monitor.js | 170 + .../frontend/src/config/status-config.js | 26 + .../10_dashboard/frontend/src/counter.js | 9 + .../frontend/src/event-handlers.js | 264 + .../10_dashboard/frontend/src/inline.js | 4988 +++++++++++++++++ .../10_dashboard/frontend/src/main.js | 45 + .../frontend/src/modules/account_selector.js | 701 +++ .../frontend/src/modules/batch_exec.js | 6 + .../frontend/src/modules/c2_remote.js | 1340 +++++ .../frontend/src/modules/cmd_tasks.js | 160 + .../frontend/src/modules/collect.js | 133 + .../frontend/src/modules/corpus.js | 1074 ++++ .../frontend/src/modules/kb_management.js | 136 + .../frontend/src/modules/machine_bar.js | 46 + .../frontend/src/modules/matrix_views.js | 745 +++ .../frontend/src/modules/nurture.js | 302 + .../frontend/src/modules/ops_router.js | 335 ++ .../frontend/src/modules/project-store.js | 139 + .../frontend/src/modules/recording.js | 447 ++ .../frontend/src/modules/registration.js | 860 +++ .../frontend/src/modules/schedule.js | 366 ++ .../frontend/src/modules/settings.js | 144 + .../frontend/src/modules/upstream.js | 268 + .../frontend/src/modules/workflow.js | 141 + .../10_dashboard/frontend/src/nav-menu.js | 108 + .../10_dashboard/frontend/src/navigation.js | 105 + .../10_dashboard/frontend/src/router.js | 130 + .../10_dashboard/frontend/src/state.js | 56 + .../10_dashboard/frontend/src/style.css | 152 + .../10_dashboard/frontend/src/utils.js | 49 + .../frontend/src/view-registry.js | 116 + .../frontend/src/views/accounts-center.js | 1215 ++++ .../10_dashboard/frontend/src/views/alerts.js | 105 + .../frontend/src/views/api-config.js | 191 + .../frontend/src/views/assets-common.js | 670 +++ .../10_dashboard/frontend/src/views/assets.js | 30 + .../frontend/src/views/ave-ambience.js | 39 + .../frontend/src/views/ave-bgm.js | 27 + .../frontend/src/views/ave-capabilities.js | 300 + .../frontend/src/views/ave-director.js | 1943 +++++++ .../frontend/src/views/ave-docs.js | 153 + .../frontend/src/views/ave-flow-beat.js | 16 + .../src/views/ave-flow-digital-human.js | 16 + .../frontend/src/views/ave-flow-dub.js | 16 + .../frontend/src/views/ave-flow-hybrid.js | 16 + .../frontend/src/views/ave-flow-narrative.js | 16 + .../src/views/ave-flow-short-drama.js | 16 + .../frontend/src/views/ave-history.js | 24 + .../frontend/src/views/ave-keyframe-video.js | 1042 ++++ .../frontend/src/views/ave-keyframes.js | 801 +++ .../frontend/src/views/ave-locations.js | 72 + .../frontend/src/views/ave-materials.js | 9 + .../frontend/src/views/ave-new-flow.js | 32 + .../frontend/src/views/ave-products.js | 28 + .../frontend/src/views/ave-props.js | 33 + .../frontend/src/views/ave-providers.js | 124 + .../frontend/src/views/ave-real-locations.js | 19 + .../frontend/src/views/ave-render.js | 9 + .../frontend/src/views/ave-script.js | 9 + .../frontend/src/views/ave-settings.js | 144 + .../frontend/src/views/ave-sfx.js | 22 + .../frontend/src/views/ave-studio.js | 510 ++ .../frontend/src/views/ave-templates.js | 9 + .../frontend/src/views/ave-test-cap.js | 209 + .../frontend/src/views/ave-vfx-presets.js | 30 + .../frontend/src/views/ave-voices.js | 151 + .../frontend/src/views/ave-wardrobe.js | 28 + .../frontend/src/views/capabilities.js | 80 + .../frontend/src/views/char-gen.js | 649 +++ .../frontend/src/views/characters.js | 521 ++ .../frontend/src/views/comment-workbench.js | 1651 ++++++ .../10_dashboard/frontend/src/views/costs.js | 114 + .../frontend/src/views/creative-studio.js | 966 ++++ .../frontend/src/views/creatives.js | 243 + .../frontend/src/views/fleet-exec.js | 143 + .../frontend/src/views/fleet-reconcile.js | 262 + .../frontend/src/views/fleet-sync.js | 33 + .../frontend/src/views/flow-entry.js | 57 + .../frontend/src/views/gen-review.js | 134 + .../frontend/src/views/machines.js | 87 + .../frontend/src/views/matrix-accounts.js | 223 + .../frontend/src/views/matrix-atom-ops.js | 58 + .../frontend/src/views/matrix-blueprints.js | 187 + .../frontend/src/views/matrix-c2.js | 132 + .../frontend/src/views/matrix-collect.js | 163 + .../frontend/src/views/matrix-commands.js | 221 + .../frontend/src/views/matrix-comment.js | 420 ++ .../frontend/src/views/matrix-corpus.js | 617 ++ .../frontend/src/views/matrix-dm.js | 80 + .../frontend/src/views/matrix-interact.js | 714 +++ .../frontend/src/views/matrix-like.js | 77 + .../frontend/src/views/matrix-live.js | 95 + .../frontend/src/views/matrix-login.js | 14 + .../frontend/src/views/matrix-nurture.js | 113 + .../frontend/src/views/matrix-publish.js | 13 + .../frontend/src/views/matrix-schedule.js | 35 + .../frontend/src/views/matrix-sms-proxy.js | 215 + .../frontend/src/views/matrix-summary.js | 97 + .../frontend/src/views/ops-command.js | 338 ++ .../frontend/src/views/ops-recorder.js | 849 +++ .../frontend/src/views/person-swap.js | 269 + .../frontend/src/views/productions.js | 371 ++ .../frontend/src/views/scene-gen.js | 225 + .../10_dashboard/frontend/src/views/scrape.js | 2602 +++++++++ .../frontend/src/views/serve-dashboard.js | 14 + .../frontend/src/views/serve-mcp.js | 14 + .../frontend/src/views/serve-schedule.js | 14 + .../src/views/studio-creative-card.js | 246 + .../src/views/studio-keyframe-plan.js | 479 ++ .../frontend/src/views/studio-post.js | 334 ++ .../frontend/src/views/studio-script.js | 448 ++ .../frontend/src/views/studio-setting.js | 150 + .../frontend/src/views/studio-storyboard.js | 505 ++ .../frontend/src/views/summary.js | 111 + .../frontend/src/views/timeline.js | 80 + .../frontend/src/views/workflow.js | 716 +++ .../10_dashboard/frontend/vite.config.js | 18 + .../05_tools/10_dashboard/nav.yaml | 122 + .../05_tools/10_dashboard/plugins/__init__.py | 38 + .../10_dashboard/plugins/_registry.py | 134 + .../05_tools/10_dashboard/plugins/ave.py | 243 + .../05_tools/10_dashboard/plugins/base.py | 139 + .../05_tools/10_dashboard/plugins/crawl.py | 91 + .../10_dashboard/plugins/federation.py | 60 + .../05_tools/10_dashboard/plugins/guardd.py | 288 + .../05_tools/10_dashboard/plugins/kb_api.py | 742 +++ .../05_tools/10_dashboard/plugins/matrix.py | 123 + .../10_dashboard/plugins/scheduler.py | 66 + .../05_tools/10_dashboard/plugins/skills.py | 45 + .../10_dashboard/plugins/sms_proxy_api.py | 910 +++ .../10_dashboard/plugins/system_plugins.py | 403 ++ .../05_tools/10_dashboard/plugins/tools.py | 40 + .../05_tools/10_dashboard/report.yaml | 119 + .../10_dashboard/routes/api_config.py | 550 ++ .../05_tools/10_dashboard/routes/ave.py | 1823 ++++++ .../10_dashboard/routes/ave_assets.py | 1327 +++++ .../10_dashboard/routes/ave_creatives.py | 262 + .../05_tools/10_dashboard/routes/ave_docs.py | 82 + .../10_dashboard/routes/ave_keyframe.py | 1766 ++++++ .../10_dashboard/routes/ave_project.py | 1735 ++++++ .../05_tools/10_dashboard/routes/ave_scene.py | 211 + .../10_dashboard/routes/comment_workbench.py | 306 + .../05_tools/10_dashboard/routes/matrix.py | 1441 +++++ .../05_tools/10_dashboard/routes/ops.py | 654 +++ .../10_dashboard/routes/person_swap.py | 190 + .../05_tools/10_dashboard/routes/scrape.py | 711 +++ .../10_dashboard/routes/v2_accounts.py | 216 + .../05_tools/10_dashboard/run.py | 58 + .../services/EXECUTION_PIPELINE.md | 213 + .../10_dashboard/services/__init__.py | 1 + .../10_dashboard/services/account_service.py | 454 ++ .../services/adapters/__init__.py | 192 + .../services/adapters/bilibili_scrape.py | 134 + .../services/adapters/browser_helpers.py | 191 + .../services/adapters/douyin_scrape.py | 181 + .../services/adapters/web_scrape.py | 115 + .../services/adapters/xhs_scrape.py | 130 + .../services/adapters/zhihu_scrape.py | 115 + .../services/browser_orchestrator.py | 281 + .../10_dashboard/services/command_bus.py | 2113 +++++++ .../10_dashboard/services/command_chain.py | 456 ++ .../10_dashboard/services/data_aggregator.py | 202 + .../10_dashboard/services/douyin_stats.py | 20 + .../10_dashboard/services/fleet_collector.py | 238 + .../10_dashboard/services/log_aggregator.py | 122 + .../services/mediacrawler_adapter.py | 457 ++ .../10_dashboard/services/nurture_runner.sh | 103 + .../10_dashboard/services/operation_queue.py | 191 + .../10_dashboard/services/preflight.py | 128 + .../10_dashboard/services/remote_exec.py | 184 + .../10_dashboard/services/resource_lock.py | 145 + .../10_dashboard/services/scrape_db.py | 415 ++ .../10_dashboard/services/scrape_engine.py | 361 ++ .../services/scrape_validation.md | 314 ++ .../10_dashboard/services/video_analyzer.py | 225 + .../05_tools/10_dashboard/static/favicon.svg | 1 + .../05_tools/10_dashboard/static/icons.svg | 24 + .../05_tools/10_dashboard/static/index.html | 119 + .../05_tools/10_dashboard/utils/__init__.py | 0 .../05_tools/10_dashboard/utils/identity.py | 31 + .../05_tools/10_dashboard/workflows.py | 1376 +++++ mac-agent-os-main/05_tools/README.md | 31 + .../07_migration/RESTORE-GUIDE.md | 156 + mac-agent-os-main/07_migration/backup.sh | 41 + mac-agent-os-main/07_migration/pack.sh | 38 + mac-agent-os-main/07_migration/unpack.sh | 41 + .../10_next_version_discussion/v2.md | 133 + .../vNext-重构规划与工作计划.md | 314 ++ .../知识库分类体系与全链路设计规范.md | 139 + .../10_next_version_discussion/讨论清单.md | 193 + .../01_director_parser/__init__.py | 0 .../02_voice_synthesizer/__init__.py | 0 .../03_bgm_generator/__init__.py | 0 .../04_anchor_extractor/__init__.py | 0 .../05_material_producer/__init__.py | 0 .../05_material_producer/fallback/__init__.py | 0 .../05_material_producer/pexels/__init__.py | 0 .../05_material_producer/wan2_2/__init__.py | 0 .../06_composer/__init__.py | 0 .../07_service_layer/__init__.py | 0 .../README.md | 35 + .../docs/AgentOS-完整说明文档_副本.md | 30 + .../AgentOS-完整说明文档_副本.v2.1-history.md | 1318 +++++ .../docs/AgentOS-知识管理系统说明文档.md | 721 +++ .../90_archive/docs/CHANGELOG.md | 106 + .../90_archive/docs/CORE-ARCHITECTURE.md | 217 + .../90_archive/docs/REVIEW-V2.0.md | 174 + .../90_archive/docs/V2.0-SUMMARY.md | 133 + mac-agent-os-main/90_archive/docs/我的说明.md | 390 ++ .../DCS_VIDEO_FACTORY_PLAN.md | 289 + .../video_factory_docs_20261001/README.md | 23 + .../VIDEO_FACTORY_ARCHITECTURE.md | 340 ++ .../VIDEO_FACTORY_FRAMEWORK.md | 530 ++ .../99_system/AGENTOS-PANORAMA.md | 426 ++ .../99_system/ARCHITECTURE_CONSTITUTION.md | 348 ++ .../99_system/COMMAND-CENTER-PLAN.md | 224 + mac-agent-os-main/99_system/FIX_PLAN.md | 332 ++ mac-agent-os-main/99_system/INDEX.md | 95 + .../99_system/OPTIMIZATION-ASSESSMENT.md | 368 ++ .../COMMAND_CONDUIT_ARCHITECTURE.md | 217 + .../2026-05-14_peekaboo_cloakbrowser.md | 73 + mac-agent-os-main/CHANGELOG.md | 2028 +++++++ mac-agent-os-main/CONSTITUTION.md | 365 ++ mac-agent-os-main/FEDERATION_GUIDE.md | 708 +++ mac-agent-os-main/MANIFEST.yaml | 453 ++ mac-agent-os-main/ORACLE.yaml | 330 ++ mac-agent-os-main/PLANS/AUDIO_LIPSYNC_PLAN.md | 146 + .../PLANS/AUDIT_5LAYER_REPORT.md | 130 + .../PLANS/CAPABILITY_GOVERNANCE.md | 144 + mac-agent-os-main/PLANS/CHANGE_SCOPE.md | 55 + mac-agent-os-main/PLANS/CODE_MERGE_PLAN.md | 191 + .../PLANS/COMMAND_UNIFICATION_PLAN.md | 220 + .../PLANS/FRONTEND_ADJUSTMENT_PLAN.md | 295 + .../PLANS/IMPLEMENTATION_PROGRESS.md | 54 + mac-agent-os-main/PLANS/INTEGRATION_AUDIT.md | 67 + .../PLANS/INTERACT_SYSTEM_PLAN.md | 551 ++ .../PLANS/OPTIMIZATION_PLAN_v2.md | 221 + .../PLANS/QUEUE_MANAGEMENT_FRAMEWORK.md | 309 + mac-agent-os-main/PLANS/REALISM_GUIDE.md | 193 + mac-agent-os-main/PLANS/REFORM_ANALYSIS.md | 241 + .../PLANS/SCHEDULER_ARCHITECTURE_v3.md | 1447 +++++ mac-agent-os-main/PLANS/STUDIO_REDESIGN.md | 288 + .../PLANS/SUBMISSION_DISTILL_TASK.md | 175 + .../PLANS/SYSTEM_AUDIT_2026-06-19.md | 373 ++ .../PLANS/TASK_ORCHESTRATION_ARCHITECTURE.md | 430 ++ mac-agent-os-main/PLANS/TEST_20S_SCRIPT.md | 74 + .../PLANS/VIDEO_FACTORY_AUDIT.md | 112 + .../PLANS/VIDEO_FACTORY_MASTER.md | 526 ++ .../PLANS/VIDEO_FACTORY_USER_GUIDE.md | 99 + .../PLANS/reference/DCS_说明书_v2.0_原文.md | 837 +++ mac-agent-os-main/README.md | 44 + mac-agent-os-main/_check_rec.py | 10 + mac-agent-os-main/_check_xhs.py | 5 + mac-agent-os-main/_verify_fixes.py | 12 + .../archive_docs/ARCHITECTURE.md | 122 + .../archive_docs/ARCHITECTURE_v3.md | 179 + .../STATE_MACHINE_ARCHITECTURE.md | 266 + .../archive_docs/VITE_MIGRATION_PLAN.md | 239 + .../matrix_docs/AGENT_INIT_GUIDE.md | 210 + .../matrix_docs/ARCHITECTURE_FULL.md | 316 ++ .../matrix_docs/DOUYIN_FULL_PLAN.md | 366 ++ .../archive_docs/matrix_docs/FRAMEWORK_V1.md | 657 +++ .../matrix_docs/FULL_TEST_REPORT.md | 237 + .../matrix_docs/IMPLEMENTATION_GUIDE.md | 894 +++ .../matrix_docs/IP_SWITCH_GUIDE.md | 185 + .../matrix_docs/PHASE_A_SUMMARY.md | 154 + .../matrix_docs/PROJECT_OVERVIEW.md | 542 ++ .../archive_docs/matrix_docs/REFACTOR-PLAN.md | 190 + .../matrix_docs/REFACTOR_PLAN_v5.md | 233 + .../archive_docs/matrix_docs/REPAIR_PLAN.md | 102 + .../matrix_docs/SETUP_ON_NEW_MACHINE.md | 532 ++ .../matrix_docs/STATE_MACHINE_CATALOG.md | 137 + .../matrix_docs/TASK_CHECKLIST.md | 96 + .../archive_docs/matrix_docs/TEST_REPORT.md | 113 + .../matrix_docs/V6_UPGRADE_PLAN.md | 235 + .../docs/DASHBOARD_DATA_LAYER_V2.md | 740 +++ mac-agent-os-main/docs/GAP-ANALYSIS.md | 179 + .../docs/UPGRADE_SOUL_v3_PLAN.md | 313 ++ mac-agent-os-main/docs/global.md | 24 + mac-agent-os-main/requirements.txt | 20 + tests/test_douyin_api.py | 181 + 1601 files changed, 367425 insertions(+) create mode 100644 api/monitor/douyin_api.py create mode 100644 mac-agent-os-main/.codewhale/handoff.md create mode 100644 mac-agent-os-main/.gitignore create mode 100644 mac-agent-os-main/00_bootstrap/DEPLOY-GUIDE.md create mode 100644 mac-agent-os-main/00_bootstrap/apply-config.sh create mode 100644 mac-agent-os-main/00_bootstrap/deploy.sh create mode 100644 mac-agent-os-main/00_bootstrap/disable_worker_dashboard.sh create mode 100644 mac-agent-os-main/00_bootstrap/export_skills.sh create mode 100644 mac-agent-os-main/00_bootstrap/fleet_reconcile.sh create mode 100644 mac-agent-os-main/00_bootstrap/fleet_sync.sh create mode 100644 mac-agent-os-main/00_bootstrap/hooks/pre-commit create mode 100644 mac-agent-os-main/00_bootstrap/hooks/pre-push create mode 100644 mac-agent-os-main/00_bootstrap/import_skills.sh create mode 100644 mac-agent-os-main/00_bootstrap/init.sh create mode 100644 mac-agent-os-main/00_bootstrap/setup_env.sh create mode 100644 mac-agent-os-main/00_bootstrap/upgrade_20260514.sh create mode 100644 mac-agent-os-main/01_core/CHANGELOG.md create mode 100644 mac-agent-os-main/01_core/CONFIG_MANIFEST.yaml create mode 100644 mac-agent-os-main/01_core/HOST_ID.tpl.md create mode 100644 mac-agent-os-main/01_core/IDENTITY.tpl.md create mode 100644 mac-agent-os-main/01_core/MAINTENANCE_GUIDE.md create mode 100644 mac-agent-os-main/01_core/NIGHTLY_AUTOMATION.md create mode 100644 mac-agent-os-main/01_core/SOUL.md create mode 100644 mac-agent-os-main/01_core/SOUL.md.v2-backup create mode 100644 mac-agent-os-main/01_core/UPDATE_SYSTEM.md create mode 100644 mac-agent-os-main/01_core/USER.tpl.md create mode 100644 mac-agent-os-main/01_core/VERSION create mode 100644 mac-agent-os-main/01_core/_archived/SOUL.md.v2-backup.note.md create mode 100644 mac-agent-os-main/01_core/automation/README.md create mode 100644 mac-agent-os-main/01_core/automation/launchd/com.agentos.guardd.plist.template create mode 100644 mac-agent-os-main/01_core/automation/workflows.yaml create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_历史解读_006.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_哲学思考_003.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_哲学思考_004.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_哲学思考_009.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_政治分析_002.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_政治分析_007.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_社会观察_001.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_社会观察_005.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_社会观察_010.md create mode 100644 mac-agent-os-main/01_submissions/20260518_yuan_benchu_经济分析_008.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_历史解读_008.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_哲学思考_004.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_哲学思考_005.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_政治分析_006.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_政治分析_007.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_社会观察_001.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_社会观察_002.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_社会观察_003.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_社会观察_010.md create mode 100644 mac-agent-os-main/01_submissions/20260519_yuan_benchu_经济分析_009.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_哲学思考_008.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_哲学思考_009.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_政治分析_001.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_政治分析_002.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_政治分析_003.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_社会观察_004.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_社会观察_005.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_社会观察_006.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_社会观察_007.md create mode 100644 mac-agent-os-main/01_submissions/20260602_yuan_benchu_经济分析_010.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_历史解读_009.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_政治分析_004.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_政治分析_005.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_社会观察_008.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_社会观察_009.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_社会观察_010.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_社会观察_011.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_经济分析_011.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_经济分析_012.md create mode 100644 mac-agent-os-main/01_submissions/20260603_yuan_benchu_经济分析_013.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_历史解读_01.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_哲学思考_01.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_政治分析_01.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_政治分析_02.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_政治分析_03.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_社会观察_01.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_社会观察_02.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_经济分析_01.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_经济分析_02.md create mode 100644 mac-agent-os-main/01_submissions/20260604_yuan_benchu_经济分析_03.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_历史解读_01.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_哲学思考_01.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_政治分析_01.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_政治分析_02.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_政治分析_03.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_社会观察_01.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_经济分析_01.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_经济分析_02.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_经济分析_03.md create mode 100644 mac-agent-os-main/01_submissions/20260605_yuan_benchu_经济分析_04.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_历史解读_01.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_哲学思考_01.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_政治分析_01.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_政治分析_02.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_政治分析_03.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_社会观察_01.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_社会观察_02.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_社会观察_03.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_社会观察_04.md create mode 100644 mac-agent-os-main/01_submissions/20260608_yuan_benchu_经济分析_01.md create mode 100644 mac-agent-os-main/02_skills/_archived/README.md create mode 100644 mac-agent-os-main/02_skills/_archived/auto_collector/SKILL.md create mode 100644 mac-agent-os-main/02_skills/_archived/cloakbrowser_controller/SKILL.md create mode 100644 mac-agent-os-main/02_skills/_archived/content_processor/SKILL.md create mode 100644 mac-agent-os-main/02_skills/_archived/web_crawler/SKILL.md create mode 100644 mac-agent-os-main/02_skills/_template/SKILL.md create mode 100644 mac-agent-os-main/02_skills/_template/skill.py create mode 100644 mac-agent-os-main/02_skills/_template/version.json create mode 100644 mac-agent-os-main/02_skills/auto_collector/SKILL.md create mode 100644 mac-agent-os-main/02_skills/auto_collector/SKILL_CARD.yaml create mode 100644 mac-agent-os-main/02_skills/auto_collector/version.json create mode 100644 mac-agent-os-main/02_skills/cloakbrowser_controller/SKILL.md create mode 100644 mac-agent-os-main/02_skills/collect_to_inbox/SKILL.md create mode 100644 mac-agent-os-main/02_skills/collect_to_inbox/SKILL_CARD.yaml create mode 100644 mac-agent-os-main/02_skills/collect_to_inbox/collect_to_inbox.py create mode 100644 mac-agent-os-main/02_skills/collect_to_inbox/version.json create mode 100644 mac-agent-os-main/02_skills/content_processor/SKILL.md create mode 100644 mac-agent-os-main/02_skills/content_processor/version.json create mode 100644 mac-agent-os-main/02_skills/inbox_refine/SKILL.md create mode 100644 mac-agent-os-main/02_skills/inbox_refine/SKILL_CARD.yaml create mode 100644 mac-agent-os-main/02_skills/inbox_refine/inbox_refine.py create mode 100644 mac-agent-os-main/02_skills/inbox_refine/llm_classifier.py create mode 100644 mac-agent-os-main/02_skills/inbox_refine/llm_classifier_broken.py create mode 100644 mac-agent-os-main/02_skills/inbox_refine/version.json create mode 100644 mac-agent-os-main/02_skills/kb_manager/SKILL.md create mode 100644 mac-agent-os-main/02_skills/kb_manager/SKILL_CARD.yaml create mode 100644 mac-agent-os-main/02_skills/kb_manager/kb_ingest.py create mode 100644 mac-agent-os-main/02_skills/kb_manager/kb_search.py create mode 100644 mac-agent-os-main/02_skills/kb_manager/vector_db_rebuild.py create mode 100644 mac-agent-os-main/02_skills/kb_manager/version.json create mode 100644 mac-agent-os-main/02_skills/matrix/SKILL.md create mode 100644 mac-agent-os-main/02_skills/matrix/SKILL_CARD.yaml create mode 100644 mac-agent-os-main/02_skills/matrix/version.json create mode 100644 mac-agent-os-main/02_skills/memory_manager/SKILL.md create mode 100644 mac-agent-os-main/02_skills/memory_manager/SKILL_CARD.yaml create mode 100644 mac-agent-os-main/02_skills/memory_manager/agent_memory_init.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/bootstrap_from_memory.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/daily_digest.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/export_memories.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/import_memories.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/memory_cleanup.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/memory_extractor.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/semantic_search.py create mode 100644 mac-agent-os-main/02_skills/memory_manager/vector_backfill.sh create mode 100644 mac-agent-os-main/02_skills/memory_manager/version.json create mode 100644 mac-agent-os-main/02_skills/peekaboo_controller/SKILL.md create mode 100644 mac-agent-os-main/02_skills/peekaboo_controller/policy.py create mode 100644 mac-agent-os-main/02_skills/sync_manager/SKILL.md create mode 100644 mac-agent-os-main/02_skills/sync_manager/SKILL_CARD.yaml create mode 100644 mac-agent-os-main/02_skills/sync_manager/sync_manager.py create mode 100644 mac-agent-os-main/02_skills/sync_manager/version.json create mode 100644 mac-agent-os-main/02_skills/web_crawler/SKILL.md create mode 100644 mac-agent-os-main/02_skills/web_crawler/version.json create mode 100644 mac-agent-os-main/03_knowledge/00_README.md create mode 100644 mac-agent-os-main/03_knowledge/00_inbox/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/00_stream/2026-06_submissions_distillation.md create mode 100644 mac-agent-os-main/03_knowledge/00_stream/inbox/knowledge/clipping/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/00_stream/inbox/knowledge/feed/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/00_stream/inbox/media/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/00_stream/inbox/memory/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/00_stream/inbox/personal/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/00_stream/inbox/tools/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/01_daily/.gitkeep create mode 120000 mac-agent-os-main/03_knowledge/01_daily/memory create mode 100644 mac-agent-os-main/03_knowledge/01_submissions/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/01_submissions/20260503_social-wisdom_douyin-video.md create mode 100644 mac-agent-os-main/03_knowledge/01_submissions/20260503_social-wisdom_subtitle.md create mode 100644 mac-agent-os-main/03_knowledge/01_submissions/20260503_social-wisdom_summary.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-05-16_10条社会智慧_-_AI总结.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-05-16_10条社会智慧_-_视频文案.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-05-16_10条社会智慧_每一条都藏着处世真相.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_2026-04-25.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_2026-04-26.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_2026-04-27.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_2026-04-28.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_2026-04-29.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_2026-04-29_1778863247.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_2026-05-01.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_bootstrap_daily_2026-04-25.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_bootstrap_system_189058df-1ad8-438f-b8aa-61087a7c1.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/2026-06-01_bootstrap_system_7a50954c-2a08-4dbb-adde-abee01c49.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/design/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/family-succession-capability-first.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/global-accounting-strategic-loss.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/mao-ten-thinking-methods.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/money-making-logic-value-creation.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/personal-management/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/power-of-naming-reality.md create mode 100644 mac-agent-os-main/03_knowledge/10_concepts/social-wisdom-10-rules.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-05-17_2026-05-16_AI提示词驱动文生视频__李导_Agent版_10条视频知识提取.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-05-17_2026-05-16_character-consistency-reference-sheet.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-05-17_2026-05-16_lip-sync-solutions-comparison.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-05-17_2026-05-16_邵氏电影美学AI提示词技巧__依士曼胶片_高调硬光_影棚置景.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-01_2026-05-17_2026-05-16_AI视频素材生成__提示词工程与分镜脚本方法汇总.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-01_2026-05-17_2026-05-16_character-consistency-refere.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-01_2026-05-17_2026-05-16_lip-sync-solutions-compariso.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-01_2026-05-17_2026-05-16_stuck-intervention.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-01_2026-05-17_AI提示词驱动文生视频__AICG造梦局10条视频知识提取.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-01_2026-05-17_video-shooting-cool-techniques.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-01_2026-05-17_可灵AI提示词完全知识手册__文生视频_图生视频提示词工程.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_AI视频素材生成__提示词工程与分.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_character-consist.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_cross-domain.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_knowledge-review.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_lip-sync-solution.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_meta-thinking.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_stuck-interventio.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_分镜时序与运镜_让AI_懂剪辑.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_原子化登录管理模块_auth_ma.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_2026-05-16_统一升级引擎_agentos_up.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_AI提示词驱动文生视频__AICG造梦局10条视频知识提.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_AI提示词驱动文生视频__博主干货提取_10_11条_.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_AI提示词驱动文生视频__李导_Agent版_10条视频.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_AI视频提示词系统创作教科书___知识体系索引.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch02-result-driven.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch07-micro-expression.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch10-commercial-tvc.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch11-sci-fi-fantasy.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch12-model-comparison.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch13-tool-workflow.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch14-script-to-prompt.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch15-narrative-structure.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch16-practice-path.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch17-resource-library.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_ch18-problem-solving.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_video-shooting-cool-techniqu.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_可灵AI提示词完全知识手册__文生视频_图生视频提示词工.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_帅是一种感觉__爆款酷炫短视频拆解与标准化流程.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-02_2026-06-01_2026-05-17_邵氏电影美学AI提示词技巧__依士曼胶片_高调硬光_影棚.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_AI视频素材.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_Camouf.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_charac.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_cross-.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_knowle.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_lip-sy.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_meta-t.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_stuck-.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_分镜时序与运.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_原子化登录管.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-05-16_统一升级引擎.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_AI提示词驱动文生视频__AICG.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_AI提示词驱动文生视频__博主干货.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_AI提示词驱动文生视频__李导_A.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_AI视频提示词系统___采集队列.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_AI视频提示词系统创作教科书___.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch01-core-logic.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch02-result-drive.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch06-audio-design.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch07-micro-expres.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch08-film-style.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch09-anime-style.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch10-commercial-t.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch11-sci-fi-fanta.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch12-model-compar.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch13-tool-workflo.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch14-script-to-pr.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch15-narrative-st.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch16-practice-pat.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch17-resource-lib.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch18-problem-solv.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch19-trends.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_video-shooting-co.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_光影色调_定义_情绪与氛围.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_可灵AI提示词完全知识手册__文生.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_基础定调参数_锁死_底层标准.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_帅是一种感觉__爆款酷炫短视频拆解.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-17_邵氏电影美学AI提示词技巧__依士.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-03_2026-06-02_2026-06-01_2026-05-28_影视风格提示词速查手册.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-05-29_OpenCLI安装与集成指南.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_2026-0.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_AI提示词驱.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_AI视频提示.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch01-c.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch02-r.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch06-a.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch07-m.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch08-f.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch09-a.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch10-c.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch11-s.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch12-m.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch13-t.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch14-s.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch15-n.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch16-p.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch17-r.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch18-p.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_ch19-t.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_video-.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_光影色调_定.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_可灵AI提示.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_基础定调参数.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_帅是一种感觉.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_邵氏电影美学.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-28_影视风格提示.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-29_OpenCL.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-prompt-creator-2.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-prompt-video-knowledge.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-prompt-engineering-summary.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/01_foundation/ch01-core-logic.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/01_foundation/ch02-result-driven.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/02_core_modules/ch03-parameters.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/02_core_modules/ch04-storyboard-camera.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/02_core_modules/ch05-lighting-color.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/02_core_modules/ch06-audio-design.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/02_core_modules/ch07-micro-expression.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/03_applications/ch08-film-style.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/03_applications/ch09-anime-style.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/03_applications/ch10-commercial-tvc.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/03_applications/ch11-sci-fi-fantasy.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/04_tools_models/ch12-model-comparison.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/04_tools_models/ch13-tool-workflow.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/05_script_narrative/ch14-script-to-prompt.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/05_script_narrative/ch15-narrative-structure.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/06_practice/ch16-practice-path.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/06_practice/ch17-resource-library.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/07_trends/ch18-problem-solving.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/07_trends/ch19-trends.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/2026-05-28_影视风格提示词速查手册.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/99_assets/collection-queue.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/99_assets/result-driven-template.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/ai-video-system/INDEX.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/douyin_ins爆火贴纸特效_ai提示词_20260605.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/install_opencli.sh create mode 100644 mac-agent-os-main/03_knowledge/20_methods/kling/character-consistency-reference-sheet.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/kling/kling-prompt-complete-guide.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/kling/lip-sync-solutions-comparison.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/li-dao-ai-video-prompts.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/shaoshi-film-aesthetics-prompts.md create mode 100644 mac-agent-os-main/03_knowledge/20_methods/video-shooting-cool-techniques/case-cool-is-a-feeling.md create mode 100644 mac-agent-os-main/03_knowledge/30_facts/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-02_2026-06-01_2026-05-17_CloakBrowser_源码级反爬浏览器.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-02_2026-06-01_2026-05-17_Matrix_养号系统_SMS_验证码自动接收.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-02_2026-06-01_2026-05-17_Peekaboo_v3_桌面_GUI_自动化.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-03_2026-06-02_2026-06-01_2026-05-17_CloakBrowser_源码级反.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-03_2026-06-02_2026-06-01_2026-05-17_Matrix_养号系统_SMS_验.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-03_2026-06-02_2026-06-01_2026-05-17_Peekaboo_v3_桌面_GU.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_CloakB.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_Matrix.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/2026-06-06_2026-06-03_2026-06-02_2026-06-01_2026-05-17_Peekab.md create mode 100644 mac-agent-os-main/03_knowledge/40_references/docs/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/40_references/papers/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/50_resources/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/60_opinions/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/90_archive/deprecated/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/90_archive/deprecated/2026-04-28_测试_LLM_分类器.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/2026-06-01_2026-05-17_2026-05-16_统一升级引擎_agentos_upgrade.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/AI_READING_GUIDE.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/ARCHITECTURE_AUDIT.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/ARCHITECTURE_CONSTITUTION.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/ARCHITECTURE_v3.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/AgentOS-原子操作体系-v2_副本2/ARCHITECTURE_AUDIT.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/AgentOS-原子操作体系-v2_副本2/EXPERIENCE-atomic-ops-20260621.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/AgentOS-原子操作体系-v2_副本2/PLAN-atomic-ops-v2.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/AgentOS-原子操作体系-v2_副本2/page_inspector.py create mode 100644 mac-agent-os-main/03_knowledge/99_system/AgentOS-原子操作体系-v2_副本2/日志-2026-06-20.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/AgentOS-原子操作体系-v2_副本2/日志-2026-06-21.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/AgentOS-原子操作体系-v2_副本2/项目记忆-MEMORY.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/README.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/TRANSFORMATION_ROADMAP.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/agent-os-memory-knowledge-architecture.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/architecture/IMPLEMENTATION-PLAN-v1.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/architecture/federated-multi-machine-architecture.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/architecture/federation-operations-architecture.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/architecture/loading-architecture.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/architecture/login-system-tech-map.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/architecture/trigger-matching-analysis.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/archive/CORE-ARCHITECTURE.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/archive/INDEX.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/archive/QUICKSTART.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/archive/SKILLS-CATALOG.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/dashboard-repair-log-20260622.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/dashboard-v4-design.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/2026-06-01_2026-05-17_2026-05-16_原子化登录管理模块_auth_manager.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/2026-06-02_2026-06-01_2026-05-17_2026-05-16_Camoufox_集成修复记录.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/douyin-comment-automation.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/douyin-login-pitfalls.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/matrix-known-pitfalls.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/matrix-nurture-system-architecture.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/matrix-sms-verification.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/matrix-v6-product-plan.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/xhs-browse-interaction-techniques.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/matrix/指纹分辨率触发XHS_AI布局版本.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/memory-index.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/pipelines/content-collection-pipeline.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/prompts/classify-knowledge.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/protocols/cross-domain.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/protocols/knowledge-review.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/protocols/meta-thinking.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/protocols/stuck-intervention.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/references/cloakbrowser-integration.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/references/peekaboo-v3-integration.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/session-20260622-dashboard.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/taxonomies/domains.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/taxonomies/folder-aliases.json create mode 100644 mac-agent-os-main/03_knowledge/99_system/taxonomies/nature-types.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/templates/concept-card.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/templates/fact-card.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/templates/method-card.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/templates/personal-insight-card.md create mode 100644 mac-agent-os-main/03_knowledge/99_system/timelines/.gitkeep create mode 100644 mac-agent-os-main/03_knowledge/CHANGELOG.md create mode 100644 mac-agent-os-main/03_knowledge/KB_REORG_PLAN.md create mode 100644 mac-agent-os-main/03_knowledge/README.md create mode 100644 mac-agent-os-main/03_knowledge/versions.json create mode 100644 mac-agent-os-main/04_memory/CHANGELOG.md create mode 100644 mac-agent-os-main/04_memory/cross_machine/ARCHITECTURE-v2.md create mode 100644 mac-agent-os-main/04_memory/cross_machine/_archive/heartbeat/4cf443bc-ff14-4ed9-885b-b04c5326304d_heartbeat.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/_archive/heartbeat/d19759cf-2159-4fe7-b6ff-db14ccf379f5_heartbeat.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/_archive/heartbeat/f13b03d1-7310-433c-ba35-368c51edb1d0_heartbeat.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/_archive/hostname_drift/Redmi-12C/heartbeat.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/encrypted/pending/.gitkeep create mode 100644 mac-agent-os-main/04_memory/cross_machine/guardd-required-version.txt create mode 100644 mac-agent-os-main/04_memory/cross_machine/knowledge/.gitkeep create mode 100644 mac-agent-os-main/04_memory/cross_machine/registry/.gitkeep create mode 100644 mac-agent-os-main/04_memory/cross_machine/registry/7kecheng_pub.pem create mode 100644 mac-agent-os-main/04_memory/cross_machine/registry/_archive/5kecheng.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/registry/_archive/7kecheng.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/registry/_archive/chengzige.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/status/_archive/Redmi-12C/heartbeat.json create mode 100644 mac-agent-os-main/04_memory/cross_machine/status/_archive/chengzigedeAir/heartbeat.json create mode 100644 mac-agent-os-main/04_memory/daily_summaries/.gitkeep create mode 100644 mac-agent-os-main/04_memory/daily_summaries/2026-04-25.md create mode 100644 mac-agent-os-main/04_memory/daily_summaries/2026-04-26.md create mode 100644 mac-agent-os-main/04_memory/daily_summaries/2026-04-27.md create mode 100644 mac-agent-os-main/04_memory/daily_summaries/2026-04-28.md create mode 100644 mac-agent-os-main/04_memory/daily_summaries/2026-04-29.md create mode 100644 mac-agent-os-main/04_memory/daily_summaries/2026-05-01.md create mode 100644 mac-agent-os-main/04_memory/logs/.gitkeep create mode 120000 mac-agent-os-main/04_memory/long_term/raw create mode 100644 mac-agent-os-main/04_memory/memory_backup/.gitkeep create mode 100644 mac-agent-os-main/05_tools/00_setup/.gitkeep create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/__init__.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/__main__.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/backup.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/check.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/config_mgr.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/const.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/init.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/install.sh create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/main.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/skill_mgr.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/sync.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/tool_mgr.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/upgrade.py create mode 100644 mac-agent-os-main/05_tools/00_setup/agentos/utils.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/README.md create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/com.agentos.chrome-debug.plist create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/com.agentos.guardd.plist create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/guardd.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/install_guardd.sh create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/launch.sh create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/.gitkeep create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/__init__.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/account_monitor.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/executor.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/heartbeat.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/oracle_sync.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/priority_queue.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/schedule_bridge.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/scheduler.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/slot_manager.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/modules/task_store.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/scripts/.gitkeep create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/scripts/chrome_debug.sh create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/scripts/chrome_hide.py create mode 100644 mac-agent-os-main/05_tools/00_setup/guardd/scripts/install.sh create mode 100644 mac-agent-os-main/05_tools/00_setup/ssh_7kecheng_setup.md create mode 100644 mac-agent-os-main/05_tools/00_setup/sync_machine.sh create mode 100644 mac-agent-os-main/05_tools/01_system/.gitkeep create mode 100644 mac-agent-os-main/05_tools/01_system/check_automation_env.py create mode 100644 mac-agent-os-main/05_tools/01_system/check_facts.py create mode 100644 mac-agent-os-main/05_tools/01_system/cleanup_registry.py create mode 100644 mac-agent-os-main/05_tools/01_system/cluster_registry.py create mode 100644 mac-agent-os-main/05_tools/01_system/fix_7kecheng.py create mode 100644 mac-agent-os-main/05_tools/01_system/fix_all_paths.py create mode 100644 mac-agent-os-main/05_tools/01_system/fix_machine_names.py create mode 100644 mac-agent-os-main/05_tools/01_system/fix_ssh_auth.sh create mode 100644 mac-agent-os-main/05_tools/01_system/rebuild_registry.py create mode 100644 mac-agent-os-main/05_tools/01_system/role_check.py create mode 100644 mac-agent-os-main/05_tools/01_system/skill_scanner.py create mode 100644 mac-agent-os-main/05_tools/01_system/test_omlx_embedding.py create mode 100644 mac-agent-os-main/05_tools/01_system/trigger_matcher.py create mode 100644 mac-agent-os-main/05_tools/02_browser/.gitkeep create mode 100644 mac-agent-os-main/05_tools/02_browser/README.md create mode 100644 mac-agent-os-main/05_tools/03_ocr/.gitkeep create mode 100644 mac-agent-os-main/05_tools/03_ocr/README.md create mode 100644 mac-agent-os-main/05_tools/04_media/.gitkeep create mode 100644 mac-agent-os-main/05_tools/04_media/README.md create mode 100644 mac-agent-os-main/05_tools/05_crawl/.gitkeep create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/README.md create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/SKILL.md create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/analyze.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/app.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/collect.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/config.yaml create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/data/database.db create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/docs/OPENCLI_SETUP.md create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/doubao_driver.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/downloader.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/install_opencli_extension.sh create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/requirements.txt create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/schema.sql create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/script_factory.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/scripts_output/vid_20260504_8817.md create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/scripts_output/vid_20260504_8817.yaml create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/scripts_output/vid_20260505_em001.md create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/scripts_output/vid_20260505_em001.yaml create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/setup_doubao_login.sh create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/test_doubao.sh create mode 100644 mac-agent-os-main/05_tools/05_crawl/content-inspiration/utils.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/longcat/longcat_claimer.py create mode 100644 mac-agent-os-main/05_tools/05_crawl/longcat/socks5_forwarder.py create mode 100644 mac-agent-os-main/05_tools/06_mobile/.gitkeep create mode 100644 mac-agent-os-main/05_tools/06_mobile/README.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/.gitignore create mode 100644 mac-agent-os-main/05_tools/07_matrix/.last_run/daily_douyin_browse.txt create mode 100644 mac-agent-os-main/05_tools/07_matrix/.last_run/daily_xhs_like.txt create mode 100644 mac-agent-os-main/05_tools/07_matrix/ARCHITECTURE.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/DOCUMENT_INDEX.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/MODULE.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/README.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/TOOL.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/agentos create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/corpus.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/daily_comment.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/douyin_comment.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/douyin_ops.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/douyin_reply.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/engine.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/interact.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/backups/20260626_test_pass/interact_comment.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_browse.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_browse_v2.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_browse_v3.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_collect.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_comment_4videos.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_comment_interact.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_comment_quick.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_comment_test.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_nurture_v1.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/douyin_search_browse.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/export_douyin_01_20260601_233311.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/export_xhs_01_20260613_001448.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/xiaohongshu_nurture_v1.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/xiaohongshu_nurture_v2.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/_archive/xiaohongshu_read_profile.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/daily_comment.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/douyin_active_v1.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/douyin_comment.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/douyin_daily.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/douyin_daily_clean.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/douyin_read_profile.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/douyin_reply.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/douyin_search.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/dy_step_test.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/dy_test_all.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/export_douyin_01_20260601_233311.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/export_xhs_01_20260613_001448.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/interact_chain.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/interact_collect.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/interact_comment.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/interact_hot.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/interact_like.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/xhs_active_v1.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/xhs_daily.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/xhs_test_all.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/blueprints/xiaohongshu_read_profile.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/config/schedule.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/accounts.override.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/accounts.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/ai.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/profiles.json create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/schedule.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/screen_layout.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/sms.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/config_template/tasks.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/corpus/douyin.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/corpus/xiaohongshu.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/ACCOUNT_LOGIN_SOP.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/ANTI-DETECTION-PLAN.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/BASE_CONVENTIONS.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/BLUEPRINT-DESIGN.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/CAMOUFOX_LOGIN_MANAGEMENT.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/COLLECTION_RULES.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/COLLECT_DISPLAY_SEPARATION.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/COMMENT_FLOW_SPEC.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/CORPUS_V2_PLAN.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/DEVELOPMENT_RULES.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/DOUYIN_SELECTORS.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/IDENTITY_FACTORY.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/LOCAL_ADAPTATION.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/MANUAL.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/MATRIX_V5_GUIDE.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/MC_COMMAND_REFERENCE.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/SMS_API_USAGE.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/SYNC_GUIDE.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/docs/SYSTEM_ARCHITECTURE.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/fix_dashboard_launchd.sh create mode 100644 mac-agent-os-main/05_tools/07_matrix/init_matrix.sh create mode 100644 mac-agent-os-main/05_tools/07_matrix/install.sh create mode 100644 mac-agent-os-main/05_tools/07_matrix/local.yaml.template create mode 100644 mac-agent-os-main/05_tools/07_matrix/mc create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/base.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/douyin/SKILL.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/douyin/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/douyin/plugin.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/xiaohongshu/SKILL.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/xiaohongshu/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/platforms/xiaohongshu/plugin.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/requirements.txt create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/ARCHITECTURE_AUDIT.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/__main__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/base.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/cli.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/plugins/ave.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/plugins/crawl.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/plugins/fleet.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/plugins/guardd.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/plugins/matrix.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/agentos/plugins/serve.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/anti_detection.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/auth_manager.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/browser_manager.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/browser_utils.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/c2/profile_scraper.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/cdp_connector.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/config/comment_corpus.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/config/schedule.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/config/vision.yaml create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/create_identity.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/data/schema.sql create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/detect_app_login.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/douyin_ops.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/guardd.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/local_paths.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/login_identity.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_mgmt.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/captcha/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/captcha/base.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/douyin_login.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/login_state_machine.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/sms/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/sms/api.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/sms/base.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/sms_login.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/xhs_login.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/account/xiaohongshu_login.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/comment/ai_generator.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/comment/xhs/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/comment/xhs/corpus.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/anchor_db.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/anchor_rules.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/behavior.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/calibration.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/comment_corpus.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/data/anchor.db create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/data/schema.sql create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/runner.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/nurture/ui_layout.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/ops/xhs/ATOMIC_OPS.md create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/ops/xhs/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/ops/xhs/browse.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/ops/xhs/interact.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/ops/xhs/selectors.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/matrix_modules/utils/cookie_manager.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/__main__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/analyzer.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/browser.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/cli.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/corpus.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/engine.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/execution_policy.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/exporter.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/interact.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/proxy.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/recorder.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/run.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/runlog.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/scheduler.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/mc/task.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/migrate_accounts_to_registry.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/nurture_blueprint.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/nurture_daily.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/ops/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/ops/_base.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/ops/xhs_ops.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/page_inspector.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/page_state.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/pipelines/__init__.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/publish_video.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/task_engine.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/task_scheduler.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/tools/comment_dom_scanner.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/tools/comment_sender.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/tools/reanalyze_xhs_dom.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/scripts/vision_bridge.py create mode 100644 mac-agent-os-main/05_tools/07_matrix/setup_on_new_machine.sh create mode 100644 mac-agent-os-main/05_tools/07_matrix/start_dashboard.sh create mode 100644 mac-agent-os-main/05_tools/08_trae_agent/README.md create mode 100644 mac-agent-os-main/05_tools/08_trae_agent/colima_setup.md create mode 100644 mac-agent-os-main/05_tools/08_trae_agent/install_trae_agent.sh create mode 100644 mac-agent-os-main/05_tools/08_trae_agent/trae_agent.sh create mode 100644 mac-agent-os-main/05_tools/08_trae_agent/trae_config.yaml.example create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/API_STRATEGY_REPORT.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/AVE_IMPLEMENTATION_PLAN.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/AVE_V2_ARCHITECTURE.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/AVE_V3_ARCHITECTURE.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/BGM_LIBRARY.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/DASHBOARD_DESIGN.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/DASHBOARD_MULTI_AGENT.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/JIMENG_AUTOMATION_PLAN.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/SPORTS_CHARACTER_PLANNING_v1.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/STRATEGY.md create mode 100644 mac-agent-os-main/05_tools/09_ave/PLANS/TASK_TRACKER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/TOOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/assets/demo_scripts/cat_tai_chi.txt create mode 100644 mac-agent-os-main/05_tools/09_ave/config.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/capabilities.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/cosyvoice_voices.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/creative_matrix.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/feasibility_rules.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/image_models.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/param_library.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/portrait_presets.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/scene_presets.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/config/video_models.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/README.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/index.json create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-ai-tryon-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-ai-tryon-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-aitryon-parsing.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-aitryon-refiner-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-aitryon-refiner-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-background-generation-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-background-generation-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-creative-poster-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-creative-poster-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-facechain-finetune-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-facechain-finetune-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-facechain-generation-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-facechain-generation-facechain-face-det.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-facechain-generation-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-image-erase-completion-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-image-erase-completion-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-image-out-painting-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-image-out-painting-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-kling-image-generation-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-kling-image-generation-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-kling-subject-id-list.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-person-instance-segmentation-create-tas.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-person-instance-segmentation-query-resu.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-qwen-image-editing.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-qwen-text-to-image-30-async.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-qwen-text-to-image-async.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-qwen-text-to-image-openai.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-qwen-text-to-image-task-query.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-qwen-text-to-image.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-shoe-model-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-shoe-model-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-vidu-image-generation-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-vidu-image-generation-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-virtual-model-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-virtual-model-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan-text-to-image-v2-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan-text-to-image-v2-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan-text-to-image-v2-synchronous.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan21-general-image-editing-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan21-general-image-editing-query-resul.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan25-general-image-editing-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan25-general-image-editing-query-resul.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan26-image-gen-edit-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan26-image-gen-edit-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan26-image-gen-edit-synchronous.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan27-image-gen-edit-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan27-image-gen-edit-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wan27-image-gen-edit-synchronous.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-sketch-to-image-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-sketch-to-image-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-style-repaint-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-style-repaint-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-v1-text-to-image-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-v1-text-to-image-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-x-painting-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wanx-x-painting-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wordart-semantic-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wordart-semantic-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wordart-texture-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-wordart-texture-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-generation-z-image.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-translation-qwen-mt-image-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-translation-qwen-mt-image-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-image-translation-qwen-mt-image-synchronous.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-multimodal-embedding-dashscope-multimodal-embedding.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-multimodal-translation-qwen-mt-uni-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-multimodal-translation-qwen-mt-uni-translate.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-real-time-multimodal-client-events.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-real-time-multimodal-interaction-process.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-real-time-multimodal-model-access.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-real-time-multimodal-realtime-java-sdk.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-real-time-multimodal-realtime-python-sdk.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-real-time-multimodal-server-events.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-real-time-multimodal-voice-cloning.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-animate-anyone-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-animate-anyone-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-animate-anyone-template-gen-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-animate-anyone-template-gen-query-resul.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-animate-anyone-wan-aa-detect.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-emo-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-emo-video-emo-detect.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-emo-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-emoji-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-emoji-video-emoji-detect.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-emoji-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-image-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-image-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-reference-to-video-create-ta.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-reference-to-video-query-res.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-text-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-text-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-video-editing-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-happyhorse-video-editing-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-kling-video-generation-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-kling-video-generation-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-liveportrait-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-liveportrait-video-liveportrait-detect.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-liveportrait-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-minimax-h3-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-minimax-h3-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-image-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-image-to-video-first-frame-cre.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-image-to-video-first-frame-que.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-image-to-video-first-last-crea.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-image-to-video-first-last-quer.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-image-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-lipsync-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-lipsync-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-motioncontrol-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-motioncontrol-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-reference-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-reference-to-video-query-resul.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-text-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-text-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-upscale-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-pixverse-upscale-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-video-retalk-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-video-retalk-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-video-style-transform-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-video-style-transform-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-image-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-image-to-video-first-frame-create-.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-image-to-video-first-frame-query-r.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-image-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-reference-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-reference-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-start-end-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-start-end-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-text-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-vidu-text-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-general-video-editing-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-general-video-editing-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-image-to-animation-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-image-to-animation-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-image-to-video-first-frame-create-t.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-image-to-video-first-frame-query-re.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-image-to-video-first-frame-video-ef.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-image-to-video-first-last-frames-cr.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-image-to-video-first-last-frames-qu.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-reference-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-reference-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-s2v-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-s2v-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-s2v-wan-s2v-detect.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-text-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-text-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-video-character-swap-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan-video-character-swap-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-image-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-image-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-reference-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-reference-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-text-to-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-text-to-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-video-editing-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan27-video-editing-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan30-video-create-task.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/api-reference-video-generation-wan30-video-query-result.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-getting-started-image-models.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-getting-started-video-models.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-getting-started-vision-models.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-background-generation.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-creative-poster.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-image-editing.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-image-erase-completion.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-image-inpainting.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-image-out-painting.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-person-instance-segmentation.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-shoe-model.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-sketch-to-image.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-text-to-image.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-virtual-model.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-wan-image-editing.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-image-generation-wan21-image-editing.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-multimodal-ocr.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-multimodal-vision.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-asr-realtime.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-asr.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-audio-generation.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-file-translation.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-improve-recognition-accuracy.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-multimodal-speech.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-music-generation.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-omni-models.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-omni-voice-list.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-qwen-audio-realtime.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-multimodal-speech.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-streaming.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-translation.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-s2s-models.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-speech-to-text-models.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-ssml.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts-models.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-cloning.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-design.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-cosyvoice.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-qwen-audio-tts.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-qwen-tts.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-video-generation-image-to-video-first-last.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-video-generation-image-to-video.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-video-generation-reference-video.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-video-generation-text-to-video.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-video-generation-video-editing.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-video-generation-wan30-video.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/character_consistency/01_skill_原文.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/character_consistency/02_提示词规范_原文.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/character_consistency/03_角色卡schema_原文.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/character_consistency/04_字段拆解_原文.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/character_consistency/README.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/cosyvoice_voices.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/gap_engine/README.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/gap_engine/对话原始JSON_2026-10-02.json create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/gap_engine/对话原文_2026-10-02.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/00_业界生产流程_关键帧先行.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/01_关键帧合成_设计方案.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/02_百炼可用能力与价格.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/03_千问AI平台_OpenAPI字段规格.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/04_看板前端改造方案_草案.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/05_图像生成与编辑能力调研.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/06_关键帧环节_流程与界面设计.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/07_关键帧工作台_设计方案.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/08_关键帧转视频_设计方案.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/09_导演工作流整合方案.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/10_按段落交接改造方案.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/11_首尾帧工作流建议.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/12_千问代理的可灵通道_规格对照.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/pipeline_research/13_千问平台可用能力总览.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/02-进阶公式.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/04-情绪外化表.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/06-约束词清单.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/07-特殊字符规范.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/08-避坑12问.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/09-kling-公式.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/20-realistic-character-consistency.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/README.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/anti-ai-checklist.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/atmosphere-dictionary.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/camera-aesthetics.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/seedance_SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/realism/smixs-README.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/业界参考/创意与设定的标准分层.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/业界参考/剧本与分镜的分层标准.md create mode 100644 mac-agent-os-main/05_tools/09_ave/docs/参数速查.md create mode 100644 mac-agent-os-main/05_tools/09_ave/install.sh create mode 100644 mac-agent-os-main/05_tools/09_ave/requirements.txt.pip create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/AVE_ARCHITECTURE_PLAN.md create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/_generate_frames.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/anchor_extractor/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/anchor_extractor/extractor.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/asset_manager/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/asset_manager/cache.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/asset_manager/index.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/asset_manager/tags.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/assets_manager/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/config.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer1_input/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer1_input/_parser.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer1_input/collector.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer2_analysis/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer2_analysis/beat_detector.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer2_analysis/emotion_analyzer.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer2_analysis/stem_separator.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer2_analysis/structure_parser.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer3_creation/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer3_creation/module_a_rhythm.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer3_creation/module_b_lyrics.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer3_creation/module_c_melody.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer3_creation/module_d_mix.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer4_alignment/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer4_alignment/lyric_aligner.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer4_alignment/stem_aligner.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer5_output/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer5_output/exporter.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/layer5_output/lyric_formatter.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/audio_line/orchestrator.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/bgm_generator/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/bgm_generator/bgm_download.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/bgm_generator/chord_pad.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/bgm_generator/suno.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/capabilities/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/capabilities/dashscope_video.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/capabilities/kling_video.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/capability_doctor.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/capability_matcher.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_adapter/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/asset_registrar.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/attribute_extractor.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/direction_expander.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/pipeline.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/portrait_set.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/prompt_assembler.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/prompts/character_sheet_prompts.md create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/quality_check.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_generator/variant_generator.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_portrait.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_registry/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_registry/registry.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/character_sheet.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/align.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/beat_sync.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/character_locker.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/de_ai.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/director_styles.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/dubbing.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/ffmpeg.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/frames.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/hybrid.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/lipsync.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/pipeline.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/qa.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/realism.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/composer/speed_ramp.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/daoist_quotes.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/demo_full_pipeline.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/director_parser/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/director_parser/parser.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/director_parser/schemas.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/director_script.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/directors/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/directors/_shot_builder.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/directors/base.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/directors/legacy_scene_planner.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/directors/llm.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/directors/seedance_director.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/directors/storyboard_v52.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/gap_engine/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/gap_engine/combo_generator.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/gap_engine/gap_calculator.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/gap_engine/library.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/gap_engine/models.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/gap_engine/store.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/inspire_quotes.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/keyframe_gen.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/kling_report.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/config.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/cost_tracker.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/dashboard.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/ffmpeg.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/ghvideo_upload.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/http.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/kling_webhook.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/lib/logger.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/main.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/fallback/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/kling/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/kling/kling.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/pexels/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/pexels/search.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/wan2_2/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/wan2_2/avatar.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/material_producer/wan2_2/wan2_2.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/music_selector/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/music_selector/music_library.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/nature_clouds.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/person_swap/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/person_swap/api.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/person_swap/preprocess.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/person_swap/service.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/pipeline_controller/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/prompts/character_sheet_prompts.md create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/scene_generator/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/scene_generator/generator.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/scene_generator/prompt_assembler.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/schemas.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/script_generator/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/script_generator/templates/story_prompts.md create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/script_schemas/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/script_schemas/script_schema.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/service_layer/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/service_layer/app.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/story_director/HOLOCINE_REFERENCE.md create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/story_director/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/story_director/batch_generator.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/story_director/scene_planner.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/story_director/temporal_bridge.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/test_tai_chi.yaml create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/tools/extend_capabilities.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/tools/fetch_api_docs.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/video_factory.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/voice_synthesizer/__init__.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/voice_synthesizer/aliyun.py create mode 100644 mac-agent-os-main/05_tools/09_ave/scripts/voice_synthesizer/volcano.py create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/README.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/COMPATIBILITY.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/INSTALLATION.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/LICENSE create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/README.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/SKILL_CATALOG.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/docs/WORKFLOW.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/docs/skills/zh-CN/ai-storyboard-director.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/docs/skills/zh-CN/character-asset.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/docs/skills/zh-CN/director-agent.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/docs/skills/zh-CN/prop-asset.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/docs/skills/zh-CN/scene-asset.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/guofeng-visual-director/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/guofeng-visual-director/references/cultural-grounding.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/guofeng-visual-director/references/image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/guofeng-visual-director/references/people-and-objects.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/guofeng-visual-director/references/prompt-and-review.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/guofeng-visual-director/references/world-and-space.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/aesthetic-audit.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/combat-mecha-aesthetic.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/combat-mecha-form-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/full-design-workflow.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/future-interface-systems.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/future-weapon-systems.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/inspiration-engine.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/live-action-cinematography.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/organism-and-ecology.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/physics-and-engineering.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/production-design.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/prompt-examples.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/scene-routes.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/scene-world-design.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/script-to-visual-derivation.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/style-color-system.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/hard-sci-fi-visual-director/references/visual-continuity.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/whitebox-previs-executor/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/whitebox-previs-executor/references/previs-spec.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/experimental/whitebox-previs-executor/references/prompt-to-previs.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-short-drama-production/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-short-drama-production/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-short-drama-production/references/control-contracts.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-short-drama-production/references/independent-production-core.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-short-drama-production/references/production-handoff.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/camera-motion-diagnostics.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/cinematography-design-engine.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/creative-shot-ideas.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/delivery-mode-guard.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/design-memory-protocol.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/fight-design.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/fight-reference-case.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/framing-and-axis.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/production-contract.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/ai-storyboard-director/references/shot-design-engine.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/character-asset/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/character-asset/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/character-asset/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/character-asset/references/female-character-charm.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/cyberpunk-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-data-analysis-semantic-layer/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-data-analysis-semantic-layer/references/semantic-layer.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-data-analysis-semantic-layer/references/semantic-write-contract.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-data-analysis-semantic-layer/references/source-inventory.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-data-analysis-semantic-layer/references/versioning-and-expiry.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/references/analysis-and-report.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/references/analysis-capability-contract.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/references/connector-contract.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/references/data-contract.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/references/evidence-and-sources.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/references/full-market-research.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/d-official-market-analysis/references/platform-metrics.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/anti-laziness-contract.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/director-thinking-spine.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/director-workbench-protocol.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/github-project-watchlist.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/local-knowledge-map.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/production-storyboard-compiler.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/research-update-protocol.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/screenplay-ai-execution-compiler.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/screenplay-cold-read-protocol.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/screenplay-exemplar-benchmarks.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/screenplay-state-engine.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/screenplay-writing-core.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/director-agent/references/verified-director-logic.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/epic-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/epic-design/references/COMMON-12-SECTION-PROTOCOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/epic-design/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/epic-design/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/epic-design/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/epic-design/references/genre-presets.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/fantasy-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/fantasy-design/references/COMMON-12-SECTION-PROTOCOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/fantasy-design/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/fantasy-design/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/fantasy-design/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/fantasy-design/references/genre-presets.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/horror-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/horror-design/references/COMMON-12-SECTION-PROTOCOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/horror-design/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/horror-design/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/horror-design/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/horror-design/references/genre-presets.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/noir-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/noir-design/references/COMMON-12-SECTION-PROTOCOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/noir-design/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/noir-design/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/noir-design/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/noir-design/references/genre-presets.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/produce-ai-video/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/produce-ai-video/references/autonomous-production-workflow.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/produce-ai-video/references/qualified-video-acceptance.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/produce-ai-video/references/storyboard-prompt-compiler.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/prop-asset/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/prop-asset/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/prop-asset/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/romance-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/romance-design/references/COMMON-12-SECTION-PROTOCOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/romance-design/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/romance-design/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/romance-design/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/romance-design/references/genre-presets.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/scene-asset/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/scene-asset/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/scene-asset/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/COMMON-12-SECTION-PROTOCOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/combat-visual.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/sources.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/squad-cinematography.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/story-visual.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/visual-design.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/visual-review.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/war-design/references/war-visual-presets.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/web-design-director/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/web-design-director/references/creative-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/web-design-director/references/design-rubric.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/web-design-director/references/interaction-design.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/web-design-director/references/material-and-motion.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/web-design-director/references/research-to-design.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/web-design-director/references/web-quality-checklist.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/wuxia-design/SKILL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/wuxia-design/references/COMMON-12-SECTION-PROTOCOL.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/wuxia-design/references/NEGATIVE-CASE-BOOK.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/wuxia-design/references/SOURCE-LEDGER.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/wuxia-design/references/cinematic-image-direction.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/ai-film-skills/skills/wuxia-design/references/genre-presets.md create mode 100644 mac-agent-os-main/05_tools/09_ave/skills/seedance-director/SKILL.md create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-16T06-20-53-028Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-16T06-20-55-501Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-16T06-29-26-329Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T11-10-46-764Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T11-12-53-973Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T11-13-50-422Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T11-39-48-242Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T11-52-17-980Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T12-02-00-772Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T12-10-23-589Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T12-13-43-409Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T12-14-54-855Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-25-02-710Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-26-00-416Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-26-49-605Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-28-43-625Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-29-36-213Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-30-22-011Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-32-57-345Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-33-14-434Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-34-03-709Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T13-35-49-778Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/.playwright-cli/page-2026-06-17T15-04-21-111Z.yml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/PLANS/BUSINESS_ARCHITECTURE_v4.md create mode 100644 mac-agent-os-main/05_tools/10_dashboard/PLANS/DEPLOYMENT.md create mode 100644 mac-agent-os-main/05_tools/10_dashboard/README.md create mode 100644 mac-agent-os-main/05_tools/10_dashboard/__init__.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/app.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/.gitignore create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/.vite/deps/_metadata.json create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/.vite/deps/package.json create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/index.html create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/package-lock.json create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/package.json create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/public/favicon.svg create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/public/icons.svg create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/api.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/assets/javascript.svg create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/assets/vite.svg create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/components/account-selector.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/components/execution-pipeline.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/components/task-monitor.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/config/status-config.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/counter.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/event-handlers.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/inline.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/main.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/account_selector.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/batch_exec.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/c2_remote.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/cmd_tasks.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/collect.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/corpus.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/kb_management.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/machine_bar.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/matrix_views.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/nurture.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/ops_router.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/project-store.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/recording.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/registration.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/schedule.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/settings.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/upstream.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/modules/workflow.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/nav-menu.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/navigation.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/router.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/state.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/style.css create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/utils.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/view-registry.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/accounts-center.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/alerts.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/api-config.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/assets-common.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/assets.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-ambience.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-bgm.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-capabilities.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-director.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-docs.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-flow-beat.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-flow-digital-human.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-flow-dub.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-flow-hybrid.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-flow-narrative.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-flow-short-drama.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-history.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-keyframe-video.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-keyframes.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-locations.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-materials.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-new-flow.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-products.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-props.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-providers.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-real-locations.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-render.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-script.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-settings.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-sfx.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-studio.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-templates.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-test-cap.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-vfx-presets.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-voices.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ave-wardrobe.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/capabilities.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/char-gen.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/characters.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/comment-workbench.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/costs.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/creative-studio.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/creatives.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/fleet-exec.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/fleet-reconcile.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/fleet-sync.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/flow-entry.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/gen-review.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/machines.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-accounts.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-atom-ops.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-blueprints.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-c2.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-collect.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-commands.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-comment.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-corpus.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-dm.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-interact.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-like.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-live.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-login.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-nurture.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-publish.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-schedule.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-sms-proxy.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/matrix-summary.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ops-command.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/ops-recorder.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/person-swap.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/productions.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/scene-gen.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/scrape.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/serve-dashboard.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/serve-mcp.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/serve-schedule.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/studio-creative-card.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/studio-keyframe-plan.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/studio-post.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/studio-script.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/studio-setting.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/studio-storyboard.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/summary.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/timeline.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/src/views/workflow.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/frontend/vite.config.js create mode 100644 mac-agent-os-main/05_tools/10_dashboard/nav.yaml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/__init__.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/_registry.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/ave.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/base.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/crawl.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/federation.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/guardd.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/kb_api.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/matrix.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/scheduler.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/skills.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/sms_proxy_api.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/system_plugins.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/plugins/tools.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/report.yaml create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/api_config.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ave.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ave_assets.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ave_creatives.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ave_docs.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ave_keyframe.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ave_project.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ave_scene.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/comment_workbench.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/matrix.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/ops.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/person_swap.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/scrape.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/routes/v2_accounts.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/run.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/EXECUTION_PIPELINE.md create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/__init__.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/account_service.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/adapters/__init__.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/adapters/bilibili_scrape.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/adapters/browser_helpers.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/adapters/douyin_scrape.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/adapters/web_scrape.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/adapters/xhs_scrape.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/adapters/zhihu_scrape.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/browser_orchestrator.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/command_bus.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/command_chain.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/data_aggregator.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/douyin_stats.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/fleet_collector.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/log_aggregator.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/mediacrawler_adapter.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/nurture_runner.sh create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/operation_queue.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/preflight.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/remote_exec.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/resource_lock.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/scrape_db.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/scrape_engine.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/scrape_validation.md create mode 100644 mac-agent-os-main/05_tools/10_dashboard/services/video_analyzer.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/static/favicon.svg create mode 100644 mac-agent-os-main/05_tools/10_dashboard/static/icons.svg create mode 100644 mac-agent-os-main/05_tools/10_dashboard/static/index.html create mode 100644 mac-agent-os-main/05_tools/10_dashboard/utils/__init__.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/utils/identity.py create mode 100644 mac-agent-os-main/05_tools/10_dashboard/workflows.py create mode 100644 mac-agent-os-main/05_tools/README.md create mode 100644 mac-agent-os-main/07_migration/RESTORE-GUIDE.md create mode 100644 mac-agent-os-main/07_migration/backup.sh create mode 100644 mac-agent-os-main/07_migration/pack.sh create mode 100644 mac-agent-os-main/07_migration/unpack.sh create mode 100644 mac-agent-os-main/90_archive/10_next_version_discussion/v2.md create mode 100644 mac-agent-os-main/90_archive/10_next_version_discussion/vNext-重构规划与工作计划.md create mode 100644 mac-agent-os-main/90_archive/10_next_version_discussion/知识库分类体系与全链路设计规范.md create mode 100644 mac-agent-os-main/90_archive/10_next_version_discussion/讨论清单.md create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/01_director_parser/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/02_voice_synthesizer/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/03_bgm_generator/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/04_anchor_extractor/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/05_material_producer/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/05_material_producer/fallback/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/05_material_producer/pexels/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/05_material_producer/wan2_2/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/06_composer/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/07_service_layer/__init__.py create mode 100644 mac-agent-os-main/90_archive/ave_legacy_numbered_dirs_20261001/README.md create mode 100644 mac-agent-os-main/90_archive/docs/AgentOS-完整说明文档_副本.md create mode 100644 mac-agent-os-main/90_archive/docs/AgentOS-完整说明文档_副本.v2.1-history.md create mode 100644 mac-agent-os-main/90_archive/docs/AgentOS-知识管理系统说明文档.md create mode 100644 mac-agent-os-main/90_archive/docs/CHANGELOG.md create mode 100644 mac-agent-os-main/90_archive/docs/CORE-ARCHITECTURE.md create mode 100644 mac-agent-os-main/90_archive/docs/REVIEW-V2.0.md create mode 100644 mac-agent-os-main/90_archive/docs/V2.0-SUMMARY.md create mode 100644 mac-agent-os-main/90_archive/docs/我的说明.md create mode 100644 mac-agent-os-main/90_archive/video_factory_docs_20261001/DCS_VIDEO_FACTORY_PLAN.md create mode 100644 mac-agent-os-main/90_archive/video_factory_docs_20261001/README.md create mode 100644 mac-agent-os-main/90_archive/video_factory_docs_20261001/VIDEO_FACTORY_ARCHITECTURE.md create mode 100644 mac-agent-os-main/90_archive/video_factory_docs_20261001/VIDEO_FACTORY_FRAMEWORK.md create mode 100644 mac-agent-os-main/99_system/AGENTOS-PANORAMA.md create mode 100644 mac-agent-os-main/99_system/ARCHITECTURE_CONSTITUTION.md create mode 100644 mac-agent-os-main/99_system/COMMAND-CENTER-PLAN.md create mode 100644 mac-agent-os-main/99_system/FIX_PLAN.md create mode 100644 mac-agent-os-main/99_system/INDEX.md create mode 100644 mac-agent-os-main/99_system/OPTIMIZATION-ASSESSMENT.md create mode 100644 mac-agent-os-main/99_system/architecture/COMMAND_CONDUIT_ARCHITECTURE.md create mode 100644 mac-agent-os-main/99_system/upgrade_notes/2026-05-14_peekaboo_cloakbrowser.md create mode 100644 mac-agent-os-main/CHANGELOG.md create mode 100644 mac-agent-os-main/CONSTITUTION.md create mode 100644 mac-agent-os-main/FEDERATION_GUIDE.md create mode 100644 mac-agent-os-main/MANIFEST.yaml create mode 100644 mac-agent-os-main/ORACLE.yaml create mode 100644 mac-agent-os-main/PLANS/AUDIO_LIPSYNC_PLAN.md create mode 100644 mac-agent-os-main/PLANS/AUDIT_5LAYER_REPORT.md create mode 100644 mac-agent-os-main/PLANS/CAPABILITY_GOVERNANCE.md create mode 100644 mac-agent-os-main/PLANS/CHANGE_SCOPE.md create mode 100644 mac-agent-os-main/PLANS/CODE_MERGE_PLAN.md create mode 100644 mac-agent-os-main/PLANS/COMMAND_UNIFICATION_PLAN.md create mode 100644 mac-agent-os-main/PLANS/FRONTEND_ADJUSTMENT_PLAN.md create mode 100644 mac-agent-os-main/PLANS/IMPLEMENTATION_PROGRESS.md create mode 100644 mac-agent-os-main/PLANS/INTEGRATION_AUDIT.md create mode 100644 mac-agent-os-main/PLANS/INTERACT_SYSTEM_PLAN.md create mode 100644 mac-agent-os-main/PLANS/OPTIMIZATION_PLAN_v2.md create mode 100644 mac-agent-os-main/PLANS/QUEUE_MANAGEMENT_FRAMEWORK.md create mode 100644 mac-agent-os-main/PLANS/REALISM_GUIDE.md create mode 100644 mac-agent-os-main/PLANS/REFORM_ANALYSIS.md create mode 100644 mac-agent-os-main/PLANS/SCHEDULER_ARCHITECTURE_v3.md create mode 100644 mac-agent-os-main/PLANS/STUDIO_REDESIGN.md create mode 100644 mac-agent-os-main/PLANS/SUBMISSION_DISTILL_TASK.md create mode 100644 mac-agent-os-main/PLANS/SYSTEM_AUDIT_2026-06-19.md create mode 100644 mac-agent-os-main/PLANS/TASK_ORCHESTRATION_ARCHITECTURE.md create mode 100644 mac-agent-os-main/PLANS/TEST_20S_SCRIPT.md create mode 100644 mac-agent-os-main/PLANS/VIDEO_FACTORY_AUDIT.md create mode 100644 mac-agent-os-main/PLANS/VIDEO_FACTORY_MASTER.md create mode 100644 mac-agent-os-main/PLANS/VIDEO_FACTORY_USER_GUIDE.md create mode 100644 mac-agent-os-main/PLANS/reference/DCS_说明书_v2.0_原文.md create mode 100644 mac-agent-os-main/README.md create mode 100644 mac-agent-os-main/_check_rec.py create mode 100644 mac-agent-os-main/_check_xhs.py create mode 100644 mac-agent-os-main/_verify_fixes.py create mode 100644 mac-agent-os-main/archive_docs/ARCHITECTURE.md create mode 100644 mac-agent-os-main/archive_docs/ARCHITECTURE_v3.md create mode 100644 mac-agent-os-main/archive_docs/STATE_MACHINE_ARCHITECTURE.md create mode 100644 mac-agent-os-main/archive_docs/VITE_MIGRATION_PLAN.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/AGENT_INIT_GUIDE.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/ARCHITECTURE_FULL.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/DOUYIN_FULL_PLAN.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/FRAMEWORK_V1.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/FULL_TEST_REPORT.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/IMPLEMENTATION_GUIDE.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/IP_SWITCH_GUIDE.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/PHASE_A_SUMMARY.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/PROJECT_OVERVIEW.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/REFACTOR-PLAN.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/REFACTOR_PLAN_v5.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/REPAIR_PLAN.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/SETUP_ON_NEW_MACHINE.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/STATE_MACHINE_CATALOG.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/TASK_CHECKLIST.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/TEST_REPORT.md create mode 100644 mac-agent-os-main/archive_docs/matrix_docs/V6_UPGRADE_PLAN.md create mode 100644 mac-agent-os-main/docs/DASHBOARD_DATA_LAYER_V2.md create mode 100644 mac-agent-os-main/docs/GAP-ANALYSIS.md create mode 100644 mac-agent-os-main/docs/UPGRADE_SOUL_v3_PLAN.md create mode 100644 mac-agent-os-main/docs/global.md create mode 100644 mac-agent-os-main/requirements.txt create mode 100644 tests/test_douyin_api.py diff --git a/api/monitor/douyin_api.py b/api/monitor/douyin_api.py new file mode 100644 index 0000000..5260c28 --- /dev/null +++ b/api/monitor/douyin_api.py @@ -0,0 +1,429 @@ +# -*- coding: utf-8 -*- +# Copyright (c) 2025 relakkes@gmail.com +# +# This file is part of MediaCrawler project. +# Repository: https://github.com/NanmiCoder/MediaCrawler/blob/main/api/monitor/douyin_api.py +# GitHub: https://github.com/NanmiCoder +# Licensed under NON-COMMERCIAL LEARNING LICENSE 1.1 +# +# 声明:本代码仅供学习和研究目的使用。使用者应遵守以下原则: +# 1. 不得用于任何商业用途。 +# 2. 使用时应遵守对应平台的使用条款和robots.txt规则。 +# 3. 不得进行大规模爬取或对平台造成运营干扰。 +# 4. 应合理控制请求频率,避免给目标平台带来不必要的负担。 +# 5. 不得用于任何非法或不当的用途。 +# +# 详细许可条款请参阅项目根目录下的LICENSE文件。 +# 使用本代码即表示您同意遵守上述原则和LICENSE中的所有条款。 + +"""抖音 Web 接口客户端 —— 直接发 HTTP,不起爬虫子进程。 + +**为什么另起一套。** 爬虫那条路(``media_platform/douyin``)会构造一大串浏览器指纹 +参数:``browser_platform=MacIntel``、``os_name=Mac OS``、``browser_version=125.0.0.0``…… +而 ``User-Agent`` 是从页面现读的(在服务器上是 Linux + Chrome 155)。参数说自己是 Mac, +UA 说自己是 Linux —— 抖音网关对这种自相矛盾的请求的处理方式是:**不报错、不给原因, +回一个 200 + 空 body**。爬虫那边把它翻译成 ``Exception("account blocked")``,看起来像 +账号被封,其实什么都不是。 + +这份客户端只发必要参数(``device_platform`` / ``aid`` 那两三个),走浏览器自己也在用的 +那条调用路径。它的做法来自 mac-agent-os 项目的 ``mediacrawler_adapter.py``,实测可用。 + +两个关键点: + +* **cookie 走 CDP 现读。** Chrome 把 cookie 值加密存在 SQLite 里,只有 CDP 拿得到 + 解密后的值;而且浏览器里那份比库里存的旧快照新 —— 站点会自己轮换会话。 +* **产物形状照抄 store。** ``aweme_id`` / ``aweme_url`` / ``cover_url`` / ``aweme_type`` / + ``create_time``(**秒**,由 adapters 换算成毫秒)…… 这样 ingest 那条链路一个字都不用改。 +""" + +import asyncio +import os +import time +from dataclasses import dataclass +from typing import Any, Dict, List, Optional, Tuple + +import config +import httpx +from tools import utils +from tools.user_hash import anonymize_user_id + +# 请求头。**要像一个浏览器**,而且必须是**同一个浏览器**:见 BrowserIdentity。 +_BASE_HEADERS = { + "Accept": "application/json, text/plain, */*", + "Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8", + "Referer": "https://www.douyin.com/", + "Origin": "https://www.douyin.com", +} + +# 网关的业务前置校验头。缺了它,抖音边缘网关的 ArgusSecurityPlugin 会直接回 +# 403 并写明 "Blocked by ArgusSecurityPlugin Uifid Not Found" —— 难得一次它会说原因。 +# 当前网关并不校验这个头的**值**,填什么都行;一旦升级到真校验,就得改成让页面里的 +# SDK 自己生成(见 media_platform/douyin/client.py 里同一条注释)。 +ARGUS_HEADER_VALUE = "1" + +API_ORIGIN = "https://www.douyin.com" +PROFILE_PATH = "/aweme/v1/web/user/profile/other/" +POSTS_PATH = "/aweme/v1/web/aweme/post/" +DETAIL_PATH = "/aweme/v1/web/aweme/detail/" +COMMENT_PATH = "/aweme/v1/web/comment/list/" + +# 一次请求的超时。抖音这两个接口正常都在一秒内返回。 +REQUEST_TIMEOUT_SECONDS = 20.0 +# 单页最多要多少条。接口自己有上限,要多了也没用。 +MAX_PAGE_SIZE = 20 + + +class DouyinApiError(RuntimeError): + """请求失败,或登录态不可用。""" + + +def _cdp_url() -> str: + """浏览器 DevTools 端点。与扫码登录那边共用同一个开关。""" + return os.getenv("MC_CDP_URL") or f"http://127.0.0.1:{config.CDP_DEBUG_PORT}" + + +@dataclass +class BrowserIdentity: + """一个请求要像浏览器所需要的全部身份信息,**且必须来自同一个浏览器**。 + + 只拿 cookie 是不够的。UA 声称自己是 Chrome 155、却不带 Chrome 155 该有的 + ``sec-ch-ua``,网关一眼就能看出这不是浏览器 —— 它的回应是 **200 + 空 body**: + 不报错、不给原因,只看得到「抓到 0 条」。所以这三样必须成套地从同一处取。 + """ + + cookie: str + user_agent: str + client_hints: Dict[str, str] + + def headers(self) -> Dict[str, str]: + headers = { + "User-Agent": self.user_agent, + **self.client_hints, + **_BASE_HEADERS, + "x-tt-argus": ARGUS_HEADER_VALUE, + "Cookie": self.cookie, + } + # uifid 是设备标识,网关要它;cookie 里没有就不带(送空值反而更像异常请求)。 + uifid = _cookie_value(self.cookie, "UIFID") or _cookie_value( + self.cookie, "UIFID_TEMP" + ) + if uifid: + headers["uifid"] = uifid + return headers + + +# 身份信息的短时缓存:一次采集要发好几个请求,没必要每次都连一遍 CDP。 +_IDENTITY_TTL_SECONDS = 120.0 +_identity_cache: Optional[Tuple[float, BrowserIdentity]] = None + + +async def _read_browser() -> Optional[BrowserIdentity]: + """连上 CDP 浏览器,一次取齐 cookie、UA、client hints。 + + 读不到返回 None(浏览器没开/没登录),由调用方决定怎么报 —— 不抛异常。 + """ + from playwright.async_api import async_playwright + + from media_platform.douyin.help import client_hint_headers + + playwright = None + try: + playwright = await async_playwright().start() + browser = await playwright.chromium.connect_over_cdp(_cdp_url(), timeout=15000) + if not browser.contexts: + return None + # contexts[0] 是真实 profile。**不要 new_context()** —— 那是无痕式的,读不到登录态。 + context = browser.contexts[0] + cookies = await context.cookies() + + # UA 和 hints 要从页面里问 —— 它们是浏览器自己的事实,写死迟早对不上。 + page = context.pages[0] if context.pages else await context.new_page() + user_agent = await page.evaluate("() => navigator.userAgent") + hints = client_hint_headers( + await page.evaluate("() => navigator.userAgentData || null") + ) + except Exception as exc: + utils.logger.warning(f"[douyin_api] 读浏览器身份失败:{exc}") + return None + finally: + if playwright is not None: + # 只断开连接。**绝不能 browser.close()** —— 对这个 CDP 连接而言那会关掉 + # 操作者自己的浏览器。 + try: + await playwright.stop() + except Exception: + pass + + douyin_cookies = { + cookie["name"]: cookie["value"] + for cookie in cookies + if "douyin" in cookie.get("domain", "") or "amemv" in cookie.get("domain", "") + } + return BrowserIdentity( + cookie=_cookie_from_dict(douyin_cookies), + user_agent=user_agent or "", + client_hints=hints, + ) + + +async def browser_identity(cookie: str = "", force: bool = False) -> BrowserIdentity: + """拿到一份可用的身份:**优先浏览器里那份**,其次退回传进来的 cookie(库里存的)。 + + 优先浏览器的原因:站点会自己轮换会话,库里存的是粘贴那一刻的快照,浏览器里那份才是 + 当前有效的;而 UA/hints 更是只有浏览器自己知道。 + """ + global _identity_cache + + now = time.monotonic() + if not force and _identity_cache is not None: + cached_at, cached = _identity_cache + if now - cached_at < _IDENTITY_TTL_SECONDS: + return cached + + identity = await _read_browser() + if identity is None or not _has_session(identity.cookie): + # 浏览器里没有可用会话,退回调用方给的那份。UA/hints 编不出来就不编 —— + # 一组和 UA 对不上的 hints 比没有更糟。 + identity = BrowserIdentity( + cookie=_cookie_header(cookie), user_agent="", client_hints={} + ) + _identity_cache = (now, identity) + return identity + + +def forget_identity() -> None: + """丢掉缓存的身份。cookie 变了、或测试之间要隔离时调用。""" + global _identity_cache + _identity_cache = None + + +def _cookie_header(cookie: str) -> str: + """把 ``a=1; b=2`` 形式的 cookie 串规整成请求头用的形状。""" + pairs = [] + for part in (cookie or "").split(";"): + if "=" in part: + name, _, value = part.partition("=") + name = name.strip() + if name: + pairs.append(f"{name}={value.strip()}") + return "; ".join(pairs) + + +def _cookie_from_dict(cookies: Dict[str, str]) -> str: + return "; ".join(f"{name}={value}" for name, value in cookies.items()) + + +def _cookie_value(cookie: str, name: str) -> str: + """从一个 cookie 串里取某个键的值。""" + for part in (cookie or "").split(";"): + key, _, value = part.partition("=") + if key.strip() == name: + return value.strip() + return "" + + +async def _get( + path: str, params: Dict[str, Any], identity: BrowserIdentity +) -> Dict[str, Any]: + """发一个 GET,返回 JSON。 + + 只带调用方给的参数 —— **不要往里加 webid / msToken / browser_version 那一堆**, + 那正是爬虫那条路失败的原因。 + """ + url = f"{API_ORIGIN}{path}" + async with httpx.AsyncClient(timeout=REQUEST_TIMEOUT_SECONDS) as client: + response = await client.get( + url, + params=params, + headers=identity.headers(), + ) + + if response.status_code != 200: + raise DouyinApiError(f"HTTP {response.status_code}:{response.text[:120]}") + + # 「200 + 空 body」是抖音网关拒绝请求时的典型回应(见模块说明)。必须当成错误报出来, + # 否则会一路往下变成「这个博主没作品」。 + if not response.text.strip(): + raise DouyinApiError( + "接口返回了空内容 —— 通常是登录态失效,或请求被网关判成了非浏览器" + ) + + try: + return response.json() + except ValueError as exc: + raise DouyinApiError(f"返回的不是 JSON:{response.text[:120]}") from exc + + +def _as_int(value: Any) -> int: + try: + return int(value) + except (TypeError, ValueError): + return 0 + + +def normalize_aweme(aweme: Dict[str, Any]) -> Dict[str, Any]: + """把接口返回的一条作品,翻译成 store 落盘的那套键名。 + + 键名必须和 ``store/douyin`` 一致 —— 跨过这一层之后,ingest 就不知道数据是从爬虫 + 来的还是从接口来的。 + """ + author = aweme.get("author") or {} + statistics = aweme.get("statistics") or {} + aweme_id = str(aweme.get("aweme_id") or "") + cover = ((aweme.get("video") or {}).get("cover") or {}).get("url_list") or [""] + uid = str(author.get("uid") or "") + nickname = author.get("nickname") or "" + + return { + "aweme_id": aweme_id, + "aweme_type": str(aweme.get("aweme_type") or ""), + # store 那边 title 取的是 desc。 + "title": aweme.get("desc") or "", + "desc": aweme.get("desc") or "", + # **秒**。adapters.time_scale 会把它换成毫秒,和 store 写出来的形态一致。 + "create_time": _as_int(aweme.get("create_time")), + "creator_hash": anonymize_user_id(uid or author.get("sec_uid") or ""), + "nickname": nickname, + "liked_count": str(_as_int(statistics.get("digg_count"))), + "comment_count": str(_as_int(statistics.get("comment_count"))), + "collected_count": str(_as_int(statistics.get("collect_count"))), + "share_count": str(_as_int(statistics.get("share_count"))), + "aweme_url": f"https://www.douyin.com/video/{aweme_id}", + "cover_url": cover[0] if cover else "", + "source_keyword": "", + } + + +async def author_videos( + sec_user_id: str, count: int = MAX_PAGE_SIZE, *, cookie: str = "" +) -> List[Dict[str, Any]]: + """某个博主最新发布的作品(按发布时间倒序),已翻译成 store 的键名。 + + 用 ``sec_user_id`` 而不是数字 uid:监控任务里存的就是主页链接里的那段 sec_uid, + 而且这个接口两种都收(爬虫那边用的也是 sec_user_id)。 + """ + identity = await browser_identity(cookie) + if not _has_session(identity.cookie): + raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有") + + payload = await _get( + POSTS_PATH, + { + "sec_user_id": sec_user_id, + "count": max(1, min(count, MAX_PAGE_SIZE)), + "max_cursor": 0, + "device_platform": "webapp", + "aid": 6383, + }, + identity, + ) + + awemes = payload.get("aweme_list") or [] + if not awemes and payload.get("status_code") not in (0, None): + raise DouyinApiError( + f"接口拒绝了请求(status_code={payload.get('status_code')})" + ) + return [normalize_aweme(aweme) for aweme in awemes] + + +async def author_profile(sec_user_id: str, *, cookie: str = "") -> Dict[str, Any]: + """博主主页指标:昵称 / 粉丝数 / 总获赞 / 作品数。""" + identity = await browser_identity(cookie) + if not _has_session(identity.cookie): + raise DouyinApiError("抖音登录态不可用:浏览器里没有会话,库里的 cookie 也没有") + + payload = await _get( + PROFILE_PATH, + {"sec_user_id": sec_user_id, "device_platform": "webapp", "aid": 6383}, + identity, + ) + user = payload.get("user") or {} + if not user: + raise DouyinApiError( + f"接口没返回用户数据(status_code={payload.get('status_code')})" + ) + return { + "nickname": user.get("nickname") or "", + "unique_id": user.get("unique_id") or "", + "fans": _as_int(user.get("follower_count")), + "total_favorited": _as_int(user.get("total_favorited")), + "works": _as_int(user.get("aweme_count")), + "following": _as_int(user.get("following_count")), + } + + +async def video_comments( + aweme_id: str, count: int = 20, *, cookie: str = "" +) -> List[Dict[str, Any]]: + """一条作品的评论,翻译成 store 的评论键名。 + + 刻意不带 ``sub_comment_count`` / ``parent_comment_id`` 的猜测值 —— 接口给了就用, + 没给就留空,不编。 + """ + identity = await browser_identity(cookie) + if not _has_session(identity.cookie): + raise DouyinApiError("抖音登录态不可用") + + payload = await _get( + COMMENT_PATH, + { + "aweme_id": aweme_id, + "count": max(1, min(count, MAX_PAGE_SIZE)), + "cursor": 0, + "device_platform": "webapp", + "aid": 6383, + }, + identity, + ) + + records = [] + for comment in payload.get("comments") or []: + user = comment.get("user") or {} + records.append( + { + "comment_id": str(comment.get("cid") or ""), + "aweme_id": aweme_id, + "content": comment.get("text") or "", + "nickname": user.get("nickname") or "", + "creator_hash": anonymize_user_id( + str(user.get("uid") or user.get("sec_uid") or "") + ), + # 同为秒;adapters 会换算。 + "create_time": _as_int(comment.get("create_time")), + "like_count": str(_as_int(comment.get("digg_count"))), + "sub_comment_count": str(_as_int(comment.get("reply_comment_total"))), + # 顶层评论在抖音里是 "0";adapters.parent_comment_id 会归一成空串。 + "parent_comment_id": str(comment.get("reply_id") or "0"), + } + ) + return records + + +def _has_session(cookie: str) -> bool: + return "sessionid=" in (cookie or "") + + +async def check_login(cookie: str = "") -> Dict[str, Any]: + """浏览器/库里现在有没有可用的抖音登录态。给设置页用。""" + identity = await browser_identity(cookie) + if _has_session(identity.cookie): + source = "browser" if identity.user_agent else "stored" + return {"ok": True, "source": source, "cookie_length": len(identity.cookie)} + return {"ok": False, "source": "", "cookie_length": 0} + + +async def main() -> None: # pragma: no cover - 手工排查用 + """``python -m api.monitor.douyin_api ``""" + import sys + + if len(sys.argv) < 2: + print(await check_login()) + return + sec = sys.argv[1] + print(await author_profile(sec)) + for record in await author_videos(sec, count=5): + print(record["create_time"], record["title"][:30], record["liked_count"]) + + +if __name__ == "__main__": # pragma: no cover + asyncio.run(main()) diff --git a/mac-agent-os-main/.codewhale/handoff.md b/mac-agent-os-main/.codewhale/handoff.md new file mode 100644 index 0000000..b274887 --- /dev/null +++ b/mac-agent-os-main/.codewhale/handoff.md @@ -0,0 +1,49 @@ +# Session handoff — 2026-07-19 + +## 今日完成 + +### 账号数据统一(Step 3 视图迁移) +| 视图 | 改动 | 状态 | +|:-----|:------|:------| +| `productions.js` | `/matrix/sms/accounts` → `/v2/accounts` | ✅ | +| `matrix-blueprints.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `matrix-comment.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `matrix-sms-proxy.js` | `/matrix/sms/accounts` → `/v2/accounts` | ✅ | +| `ops-recorder.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `matrix-nurture.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `comment-workbench.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `matrix-like.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `matrix-interact.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `matrix-collect.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `matrix-accounts.js` | `/matrix/accounts` → `/v2/accounts` | ✅ | +| `inline.js` (2处) | `/matrix/accounts` → `/v2/accounts` | ✅ | + +### 旧 API 标记 DEPRECATED +- `routes/matrix.py`: `/api/matrix/accounts` 文档+运行时日志标记废弃 +- `MANIFEST.yaml`: 添加红线 `🚫 /api/matrix/accounts(读)— 已废弃` + +### 状态配置统一 +- 新建 `config/status-config.js` (STATUS_CFG + STATUS_ORDER 唯一来源) +- `accounts-center.js`、`account-selector.js` 改为 import + +### 账号选择器重构 +- 筛选改为排除模式(机器/平台/状态均排除选中的) +- 新增「全选筛选结果」按钮 +- 新增「复位选择」按钮 +- 搜索框移至第二行 + +### 养号 pkill 误杀修复 +- `nurture_runner.sh`: `pkill -f` 改为 `pgrep`+`ps` 精确匹配 +- `command_bus.py`: `graceful_exit()` 先查 guardd 有 active 任务则跳过 + +### 评论区输入框 +- `matrix-comment.js`: `` → ` +
ICE 收集完成后自动生成,复制后通过 curl 命令发送到服务端获取 Answer
+ + +
+
2curl 命令
+
+ +
+ +
复制此命令到终端执行,将返回的 Answer SDP 粘贴到下方
+
+ +
+
3Answer SDP
+ +
粘贴后点击上方"设置 Answer"按钮建立连接
+
+ + + +
事件(DataChannel)
+
+ + + + + + ``` + + + 在浏览器中打开此文件,按以下步骤操作: + + 1. 点击**开始会话**,页面会自动生成 Offer SDP 和对应的 curl 命令。 + 2. 点击**复制 curl 命令**,在终端中执行。命令返回的内容即为 Answer SDP。 + 3. 将 Answer SDP 粘贴到页面的 **Answer SDP** 文本框中,点击**设置 Answer**即可建立连接并开始语音对话。 + + + + + +VAD/手动模式交互流程以及 Qwen3.8-Omni-Flash-Realtime MCP 交互流程(工具发现、调用、审批与续接),参见[交互流程](/api-reference/real-time-multimodal/interaction-process)。 + +## 多通道音频、视频聚合与 MCP + +以下配置适用于 Qwen3.8-Omni-Flash-Realtime,按需在首次输入音频前配置。 + +- **多通道音频(WebSocket 接入)**:在首段音频前设置 `session.audio.input.format`。支持 1、2、4 声道;多通道必须为 PCM、16000 Hz、s16le、interleaved。2 声道使用 `raw_mic_array`,4 声道使用 `foa_ambix`。字段约束和 JSON 片段见[客户端事件](/api-reference/real-time-multimodal/client-events#session-audio-multichannel)。 +- **视频聚合**:`session.video.input.representation_compact` 默认为 `none`;设为 `normal` 可聚合视频表征,降低计算开销,适用于不依赖细粒度视觉信息的场景。也必须在首段音频前配置,开始输入后不得修改。 +- **音色**:默认音色为 `Tina`。新接入使用 `session.audio.output.voice`,它优先于兼容字段 `session.voice`。支持 `longanlingxin`(龙安灵心,知心温暖音)等音色。使用方式和声音复刻入口见[音色列表](/developer-guides/speech/omni-voice-list#qwen3-8-omni-flash-realtime)。 +- **MCP**:在 `session.tools` 中配置公网 HTTPS MCP Streamable HTTP 服务;默认需要审批。先确认工具发现完成,再发起需要该工具的 Response。执行结果、拒绝和失败处理以及续答步骤见[交互流程](/api-reference/real-time-multimodal/interaction-process#mcp-工具调用)。 + +### SDK 配置视频聚合 + +Qwen3.8-Omni-Flash-Realtime 使用 DashScope Python SDK 1.26.5 及以上版本,或 Java SDK 2.22.15 及以上版本。视频聚合参数通过 SDK 透传:Python 在 `update_session` 中传入 `video`,Java 通过 `OmniRealtimeConfig.builder().parameters(...)` 传入。 + +以下示例配置一个 Manual 模式会话:文本与音频输出、`Tina` 音色、关闭输入转写,并启用视频聚合。先按下方 SDK 示例建立连接,将模型设为 `qwen3.8-omni-flash-realtime`,用本例替换该示例的会话配置调用,在首段音频输入前调用一次。`conversation` 为已连接的 SDK 会话。SDK 会同时发送音色、VAD 等配置;如需其他音色、VAD 或转写设置,在同一次调用中完整设置。Java 需导入 `java.util.Arrays`、`java.util.Map`、`java.util.HashMap`、`com.alibaba.dashscope.audio.omni.OmniRealtimeConfig` 和 `com.alibaba.dashscope.audio.omni.OmniRealtimeModality`。 + + + ```python Python + from dashscope.audio.qwen_omni import MultiModality + + conversation.update_session( + output_modalities=[MultiModality.AUDIO, MultiModality.TEXT], + voice="Tina", + enable_turn_detection=False, + enable_input_audio_transcription=False, + video={ + "input": { + "representation_compact": "normal", + }, + }, + ) + ``` + + ```java Java + Map videoInput = new HashMap<>(); + videoInput.put("representation_compact", "normal"); + + Map video = new HashMap<>(); + video.put("input", videoInput); + + Map parameters = new HashMap<>(); + parameters.put("video", video); + + OmniRealtimeConfig config = + OmniRealtimeConfig.builder() + .modalities(Arrays.asList(OmniRealtimeModality.AUDIO, OmniRealtimeModality.TEXT)) + .voice("Tina") + .enableTurnDetection(false) + .enableInputAudioTranscription(false) + .parameters(parameters) + .build(); + conversation.updateSession(config); + ``` + + +## 联网搜索 + +联网搜索功能允许模型使用实时检索到的信息进行回复,适用于需要最新信息的场景,如股票价格、天气预报等。模型会自主决定是否进行联网搜索。 + + + Qwen3.8-Omni-Flash-Realtime 和 Qwen3.5-Omni-Realtime 系列模型支持联网搜索,默认关闭,需通过 `session.update` 事件启用。 + + 有关计费详情,请参阅[计费说明](/developer-guides/getting-started/pricing)中 `agent` 策略的说明。 + + +### 如何开启 + +在 `session.update` 事件中,增加以下参数: + +- `enable_search`:设为 `true` 以开启联网搜索。 +- `search_options.enable_source`:设为 `true` 以返回搜索结果来源列表。 + +完整参数说明请参阅 [session.update](/api-reference/real-time-multimodal/client-events)。 + +### 响应格式 + +开启联网搜索后,`response.done` 事件的 `usage` 对象中新增 `plugins` 字段,用于记录搜索使用量: + +```json +{ + "usage": { + "total_tokens": 2937, + "input_tokens": 2554, + "output_tokens": 383, + "input_tokens_details": { + "text_tokens": 2512, + "audio_tokens": 42 + }, + "output_tokens_details": { + "text_tokens": 90, + "audio_tokens": 293 + }, + "plugins": { + "search": { + "count": 1, + "strategy": "agent" + } + } + } +} +``` + +### 代码示例 + +以下示例展示如何开启联网搜索功能。 + + + + 在 `update_session` 调用中传入 `enable_search` 和 `search_options` 参数: + + ```python + import os + import base64 + import time + import json + import pyaudio + from dashscope.audio.qwen_omni import MultiModality, AudioFormat, OmniRealtimeCallback, OmniRealtimeConversation + import dashscope + + dashscope.api_key = os.getenv('DASHSCOPE_API_KEY') + url = 'wss://maas.qianwenaiapi.com/api-ws/v1/realtime' + model = 'qwen3.8-omni-flash-realtime' + voice = 'Tina' + + class SearchCallback(OmniRealtimeCallback): + def __init__(self, pya): + self.pya = pya + self.out = None + def on_open(self): + self.out = self.pya.open(format=pyaudio.paInt16, channels=1, rate=24000, output=True) + def on_event(self, response): + if response['type'] == 'response.audio.delta': + self.out.write(base64.b64decode(response['delta'])) + elif response['type'] == 'conversation.item.input_audio_transcription.delta': + preview = response.get('text', '') + response.get('stash', '') + print(f"\r[用户] {preview}", end='', flush=True) + elif response['type'] == 'conversation.item.input_audio_transcription.completed': + print(f"\r[用户] {response['transcript']}") + elif response['type'] == 'response.audio_transcript.done': + print(f"[模型] {response['transcript']}") + elif response['type'] == 'response.done': + usage = response.get('response', {}).get('usage', {}) + plugins = usage.get('plugins', {}) + if plugins.get('search'): + print(f"[搜索] count={plugins['search']['count']}, strategy={plugins['search']['strategy']}") + + pya = pyaudio.PyAudio() + callback = SearchCallback(pya) + conv = OmniRealtimeConversation(model=model, callback=callback, url=url) + conv.connect() + conv.update_session( + output_modalities=[MultiModality.AUDIO, MultiModality.TEXT], + voice=voice, + instructions="你是小云,一个私人助手", + enable_search=True, + search_options={'enable_source': True} + ) + mic = pya.open(format=pyaudio.paInt16, channels=1, rate=16000, input=True) + print("联网搜索已开启。对着麦克风说话(按 Ctrl+C 退出)...") + try: + while True: + audio_data = mic.read(3200, exception_on_overflow=False) + conv.append_audio(base64.b64encode(audio_data).decode()) + time.sleep(0.01) + except KeyboardInterrupt: + conv.close() + mic.close() + callback.out.close() + pya.terminate() + print("\n对话已结束") + ``` + + + + 在 `updateSession` 中通过 `parameters` 映射传入联网搜索参数: + + ```java + import com.alibaba.dashscope.audio.omni.*; + import com.alibaba.dashscope.exception.NoApiKeyException; + import com.google.gson.JsonObject; + import javax.sound.sampled.*; + import java.nio.ByteBuffer; + import java.util.*; + import java.util.concurrent.ConcurrentLinkedQueue; + import java.util.concurrent.atomic.AtomicBoolean; + + public class OmniSearch { + static class SequentialAudioPlayer { + private final SourceDataLine line; + private final Queue audioQueue = new ConcurrentLinkedQueue<>(); + private final Thread playerThread; + private final AtomicBoolean shouldStop = new AtomicBoolean(false); + + public SequentialAudioPlayer() throws LineUnavailableException { + AudioFormat format = new AudioFormat(24000, 16, 1, true, false); + line = AudioSystem.getSourceDataLine(format); + line.open(format); + line.start(); + playerThread = new Thread(() -> { + while (!shouldStop.get()) { + byte[] audio = audioQueue.poll(); + if (audio != null) { + line.write(audio, 0, audio.length); + } else { + try { Thread.sleep(10); } catch (InterruptedException ignored) {} + } + } + }, "AudioPlayer"); + playerThread.start(); + } + + public void play(String base64Audio) { + audioQueue.add(Base64.getDecoder().decode(base64Audio)); + } + public void close() { + shouldStop.set(true); + try { playerThread.join(1000); } catch (InterruptedException ignored) {} + line.drain(); + line.close(); + } + } + + public static void main(String[] args) { + try { + SequentialAudioPlayer player = new SequentialAudioPlayer(); + AtomicBoolean shouldStop = new AtomicBoolean(false); + + OmniRealtimeParam param = OmniRealtimeParam.builder() + .model("qwen3.8-omni-flash-realtime") + .apikey(System.getenv("DASHSCOPE_API_KEY")) + .url("wss://maas.qianwenaiapi.com/api-ws/v1/realtime") + .build(); + + OmniRealtimeConversation conversation = new OmniRealtimeConversation(param, new OmniRealtimeCallback() { + @Override public void onOpen() { + System.out.println("连接已建立"); + } + @Override public void onClose(int code, String reason) { + System.out.println("连接已关闭"); + shouldStop.set(true); + } + @Override public void onEvent(JsonObject event) { + String type = event.get("type").getAsString(); + if ("response.audio.delta".equals(type)) { + player.play(event.get("delta").getAsString()); + } else if ("response.audio_transcript.done".equals(type)) { + System.out.println("[模型] " + event.get("transcript").getAsString()); + } else if ("response.done".equals(type)) { + JsonObject response = event.getAsJsonObject("response"); + if (response != null && response.has("usage")) { + JsonObject usage = response.getAsJsonObject("usage"); + if (usage.has("plugins")) { + JsonObject plugins = usage.getAsJsonObject("plugins"); + if (plugins.has("search")) { + JsonObject search = plugins.getAsJsonObject("search"); + System.out.println("[搜索] count=" + search.get("count").getAsInt() + + ", strategy=" + search.get("strategy").getAsString()); + } + } + } + } + } + }); + + conversation.connect(); + conversation.updateSession(OmniRealtimeConfig.builder() + .modalities(Arrays.asList(OmniRealtimeModality.AUDIO, OmniRealtimeModality.TEXT)) + .voice("Tina") + .enableTurnDetection(true) + .enableInputAudioTranscription(true) + .parameters(Map.of( + "instructions", "你是小云,一个私人助手", + "enable_search", true, + "search_options", Map.of("enable_source", true) + )) + .build() + ); + + System.out.println("联网搜索已开启。开始说话(按 Ctrl+C 退出)..."); + AudioFormat format = new AudioFormat(16000, 16, 1, true, false); + TargetDataLine mic = AudioSystem.getTargetDataLine(format); + mic.open(format); + mic.start(); + + ByteBuffer buffer = ByteBuffer.allocate(3200); + while (!shouldStop.get()) { + int bytesRead = mic.read(buffer.array(), 0, buffer.capacity()); + if (bytesRead > 0) { + byte[] chunk = new byte[bytesRead]; + System.arraycopy(buffer.array(), 0, chunk, 0, bytesRead); + conversation.appendAudio(Base64.getEncoder().encodeToString(chunk)); + } + Thread.sleep(20); + } + + conversation.close(1000, "正常退出"); + player.close(); + mic.close(); + } catch (NoApiKeyException e) { + System.err.println("未找到 API KEY:请设置 DASHSCOPE_API_KEY 环境变量。"); + } catch (Exception e) { + e.printStackTrace(); + } + } + } + ``` + + + + 在 `session.update` 的 JSON 数据中增加 `enable_search` 和 `search_options` 字段: + + ```python + import json + import os + import websocket + import base64 + import pyaudio + import threading + + API_KEY = os.getenv("DASHSCOPE_API_KEY") + API_URL = "wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime" + + pya = pyaudio.PyAudio() + out_stream = pya.open(format=pyaudio.paInt16, channels=1, rate=24000, output=True) + + def on_open(ws): + ws.send(json.dumps({ + "type": "session.update", + "session": { + "modalities": ["text", "audio"], + "voice": "Tina", + "instructions": "你是个人助理小云", + # 配置音频输入/输出格式和采样率 + "audio": { + "input": {"format": {"type": "pcm", "sample_rate": 16000}}, + "output": {"format": {"type": "pcm", "sample_rate": 24000}} + }, + "enable_search": True, + "search_options": { + "enable_source": True + } + } + })) + print("联网搜索已开启。对着麦克风说话...") + def send_audio(): + mic = pya.open(format=pyaudio.paInt16, channels=1, rate=16000, input=True) + try: + while True: + audio = mic.read(3200, exception_on_overflow=False) + ws.send(json.dumps({ + "type": "input_audio_buffer.append", + "audio": base64.b64encode(audio).decode() + })) + except Exception: + mic.close() + threading.Thread(target=send_audio, daemon=True).start() + + def on_message(ws, message): + event = json.loads(message) + if event["type"] == "response.audio.delta": + out_stream.write(base64.b64decode(event["delta"])) + elif event["type"] == "response.audio_transcript.done": + print(f"[模型] {event['transcript']}") + elif event["type"] == "response.done": + usage = event.get("response", {}).get("usage", {}) + plugins = usage.get("plugins", {}) + if plugins.get("search"): + print(f"[搜索] count={plugins['search']['count']}, strategy={plugins['search']['strategy']}") + + def on_error(ws, error): + print(f"错误: {error}") + + headers = ["Authorization: Bearer " + API_KEY] + ws = websocket.WebSocketApp(API_URL, header=headers, on_open=on_open, on_message=on_message, on_error=on_error) + ws.run_forever() + ``` + + + +## API 参考 + +- **事件协议**:[客户端事件](/api-reference/real-time-multimodal/client-events)、[服务端事件](/api-reference/real-time-multimodal/server-events) +- **SDK**:[Python SDK](/api-reference/real-time-multimodal/realtime-python-sdk)、[Java SDK](/api-reference/real-time-multimodal/realtime-java-sdk) +- **AOQ 协议**:[AOQ SDK 概述](/api-reference/realtime-api/aoq-sdk-intro) +- **协议选型**:[Realtime API 概述](/api-reference/realtime-api/overview)(WebSocket、WebRTC 与 AOQ 的对比) + +## 计费与限流 + +### 计费规则 + +Qwen-Omni-Realtime 按不同输入模态(如音频和图像)消耗的 Token 数量计费。有关计费的详细信息,请参阅[计费说明](/developer-guides/getting-started/pricing)。 + +输出语音时,`qwen3.8-omni-flash-realtime` 的音频及对应文本分别按音频输出和文本输出单价计费;Qwen3.5-Omni-Realtime 系列仅对音频计费,对应文本不计费。 + +MCP 工具调用不额外收费,模型推理仍按模型价格计费。调用流程与使用限制见[交互流程](/api-reference/real-time-multimodal/interaction-process#mcp-工具调用)。 + + + 在多轮实时对话中,模型每次生成响应时,需要将上下文窗口内的所有历史对话内容(包括之前各轮的音频、图片和文本)与本轮新增输入一并作为输入 Token 进行处理。因此,输入 Token 会随对话轮次的增加而逐轮累积,而非仅计算当前轮次的新增输入。 + + 例如,假设一段 10 秒的音频输入转换为 70 个 Token(Qwen3.5-Omni-Realtime 系列模型),在第 3 轮对话时,该音频仍在上下文窗口内,则它依然会被计入第 3 轮的输入 Token。实际计费的输入 Token 数 = 上下文窗口内所有历史轮次的内容 Token 数 + 本轮新增输入的 Token 数。 + + + + + + - Qwen3.8-Omni-Flash-Realtime、Qwen3.5-Omni-Realtime:输入音频总 Token 数 = 音频时长(秒)x 7;输出音频总 Token 数 = 音频时长(秒)x 12.5 + - Qwen3-Omni-Flash-Realtime:输入与输出音频的总 Token 数 = 音频时长(秒)x 12.5 + - Qwen-Omni-Turbo-Realtime:输入与输出音频的总 Token 数 = 音频时长(秒)x 25 + + 若音频时长不足 1 秒,按 1 秒计算。 + + `qwen3.8-omni-flash-realtime` 开启空间音频输入时,音频输入 Token 数为普通音频的 2 倍,2 通道和 4 通道的倍数相同。 + + + + + + - `qwen3.8-omni-flash-realtime`、Qwen3.5-Omni-Realtime 系列模型:每 `32x32` 像素消耗 1 Token + - `Qwen3-Omni-Flash-Realtime` 系列模型:每 `32x32` 像素消耗 1 Token + - `Qwen-Omni-Turbo-Realtime` 模型:每 `28x28` 像素消耗 1 Token + + 单张图片最少消耗 4 个 Token,最多支持 1,280 个 Token。可使用以下代码估算图片消耗的总 Token 数: + + ```python + # 安装 Pillow 库:pip install Pillow + from PIL import Image + import math + + # Qwen-Omni-Turbo-Realtime 模型的缩放因子为 28 + # factor = 28 + # Qwen3.8-Omni-Flash-Realtime、Qwen3.5-Omni-Realtime、Qwen3-Omni-Flash-Realtime 系列模型的缩放因子为 32 + factor = 32 + + def token_calculate(image_path='', duration=10): + """ + :param image_path: 图片路径 + :param duration: 会话连接时长(秒) + :return: 图片消耗的 Token 数 + """ + if len(image_path) > 0: + # 打开指定的 PNG 图片文件 + image = Image.open(image_path) + # 获取图片原始宽高 + height = image.height + width = image.width + print(f"缩放前图片尺寸: height={height}, width={width}") + # 将高度调整为 factor 的整数倍 + h_bar = round(height / factor) * factor + # 将宽度调整为 factor 的整数倍 + w_bar = round(width / factor) * factor + # 图片 Token 下限:4 个 Token + min_pixels = factor * factor * 4 + # 图片 Token 上限:1,280 个 Token + max_pixels = 1280 * factor * factor + # 缩放图片,确保总像素在 [min_pixels, max_pixels] 范围内 + if h_bar * w_bar > max_pixels: + # 计算缩放比例 beta,使缩放后总像素不超过 max_pixels + beta = math.sqrt((height * width) / max_pixels) + # 重新计算调整后的高度,确保为 factor 的整数倍 + h_bar = math.floor(height / beta / factor) * factor + # 重新计算调整后的宽度,确保为 factor 的整数倍 + w_bar = math.floor(width / beta / factor) * factor + elif h_bar * w_bar < min_pixels: + # 计算缩放比例 beta,使缩放后总像素不低于 min_pixels + beta = math.sqrt(min_pixels / (height * width)) + # 重新计算调整后的高度,确保为 factor 的整数倍 + h_bar = math.ceil(height * beta / factor) * factor + # 重新计算调整后的宽度,确保为 factor 的整数倍 + w_bar = math.ceil(width * beta / factor) * factor + print(f"缩放后图片尺寸: height={h_bar}, width={w_bar}") + # 计算图片 Token 数:总像素除以 (factor x factor) + token = int((h_bar * w_bar) / (factor * factor)) + print(f"缩放后 Token 数: {token}") + total_token = token * math.ceil(duration / 2) + print(f"总 Token 数: {total_token}") + return total_token + else: + print("错误:image_path 为空,无法计算 Token 数") + return 0 + + if __name__ == "__main__": + total_token = token_calculate(image_path="xxx/test.jpg", duration=10) + ``` + + + + `qwen3.8-omni-flash-realtime` 的视频基础计量沿用 Qwen3.5-Omni-Realtime,视频画面的像素换算见[图片 Token 规则](#realtime-image-tokens)。 + + `qwen3.8-omni-flash-realtime` 的视频输入在 `representation_compact="normal"` 时,视频 Token 数为未开启聚合时的 1/4。 + + + + +### 限流 + +有关模型限流规则的详细信息,请参阅[限流说明](/developer-guides/administration/rate-limits)。 + +## 常见问题 + +### 怎么向模型输入图片? + +A:输入方式取决于接入协议。 + +**WebSocket**:通过客户端发送 [input\_image\_buffer.append](/api-reference/real-time-multimodal/client-events) 事件。 + +- VAD 模式:该模式会根据语音检测情况自动提交音频与图片,请在服务端响应前发送 [input\_image\_buffer.append](/api-reference/real-time-multimodal/client-events) 事件。 +- Manual 模式:参见快速开始中的手动模式代码,将图片输入与提交的两部分代码取消注释,即可传入本地图片。 + +**WebRTC**:通过视频轨道(RTP)发送画面帧,无需发送 `input_image_buffer.append` 事件。 + +若用于视频通话场景,可以对视频抽帧,建议以 1张/秒 的频率向服务端发送图像。DashScope SDK 代码请参见 [Omni-Realtime 示例代码](https://github.com/aliyun/alibabacloud-bailian-speech-demo/tree/master/samples/conversation/omni)。 + +### 收不到 ASR 转录文本(transcript)怎么办? + +A:如果您收不到 `conversation.item.input_audio_transcription.completed` 事件,或事件中的 transcript 字段为空,请按以下顺序排查。 + +1. **确认已开启输入音频转录**。转录事件由 `session.update` 的 `enable_input_audio_transcription` 参数控制,该参数默认为 `true`。关闭该参数后,服务端不会下发 `conversation.item.input_audio_transcription.completed` 事件,这是预期行为,并非异常,请确认会话配置中该参数为 `true`。 +2. **监听转录失败事件**。除 `completed` 事件外,请同时监听 `conversation.item.input_audio_transcription.failed` 事件。转录失败时,该事件会通过 `error.code`、`error.message` 和 `error.param` 返回失败原因,可据此定位是参数问题还是音频问题。 +3. **升级 DashScope SDK**。较低版本的 DashScope SDK 可能不支持输入音频转录相关参数,导致配置未生效,请将 DashScope SDK 升级到较新版本后重试。 +4. **检查 VAD 配置与网络**。语音活动检测由 `session.turn_detection` 配置,`type` 取值为 `server_vad` 或 `semantic_vad`。VAD 灵敏度与实际环境不匹配(例如在嘈杂环境中阈值过低)会导致音频分段异常,进而影响转录结果,可参见本文档"2. 配置会话"章节调整 `threshold` 与 `silence_duration_ms`。此外,网络不稳定会造成 WebSocket 连接中断,导致转录事件丢失,请确认连接在整个会话期间保持稳定。 +5. **显式指定转录模型**。可通过 `session.update` 的 `input_audio_transcription_model` 参数指定 ASR 转录模型(例如 `qwen3-asr-flash-realtime`)。未显式指定时,服务端使用默认转录模型;当默认模型的转录效果不满足需求时,建议显式指定。 + +## 错误码 + +调用失败时,请参阅[错误码](/api-reference/preparation/error-messages)。 + +## 音色列表 + +各全模态模型支持的音色、试听样例和 `voice` 参数取值,见[全模态音色列表](/developer-guides/speech/omni-voice-list)。 diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-streaming.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-streaming.md new file mode 100644 index 0000000..f15fb4f --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-streaming.md @@ -0,0 +1,4878 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 实时语音合成 + +> 实时语音合成将文本实时转换为自然语音,支持流式输入与输出,具备声音复刻、声音设计及精细化音频控制能力,适用于语音助手、有声读物、智能客服等场景。 + +## 概述 + +实现低延迟文本到语音转换。 + +- 支持流式输入与输出,首包延迟低 +- 可调节语速、语调、音量与码率,实现精细的语音效果控制 +- 兼容主流音频格式(PCM、WAV、MP3、Opus),最高支持 48kHz 采样率输出 +- 支持[指令控制](/developer-guides/speech/realtime-streaming#指令控制),可通过自然语言指令控制语音表现力 +- 支持[声音复刻](/developer-guides/speech/voice-cloning)与[声音设计](/developer-guides/speech/voice-design)音色定制 +- 支持[情感与富语言标签](/developer-guides/speech/realtime-streaming#情感与富语言标签),可在文本中嵌入标签控制情感表达或插入拟声效果 + +批量场景(有声读物、课件配音等)可使用[非实时语音合成](/developer-guides/speech/tts)。各模型选型建议请参见[语音合成](/developer-guides/speech/tts-models)。 + + + Sambert 为早期语音合成模型,新项目建议优先使用 CosyVoice 或 Qwen-Audio-TTS,可获得更好的合成效果和更丰富的功能支持。 + + +## 前提条件 + +- 已[配置 API Key](/api-reference/preparation/api-key)并将其[设置到环境变量](/api-reference/preparation/export-api-key-env)。 +- 如果通过 DashScope SDK 调用,需要[安装最新版SDK](/api-reference/preparation/install-sdk)。 +- 如果通过 AOQ 协议接入 CosyVoice 系列模型,需要下载并集成 AOQ 客户端 SDK,详见[AOQ SDK 简介](/api-reference/realtime-api/aoq-sdk-intro)。 + +## 快速开始 + +以下是各模型的语音合成示例。更多示例和参数说明请参见 [Realtime API](/api-reference/realtime-api/overview)。 + + + + 以下示例演示如何使用系统音色进行语音合成。 + + 如需使用[指令控制](/developer-guides/speech/realtime-streaming#指令控制)功能,请通过 `instruction` 参数设置指令。 + + + ```python Python expandable + # coding=utf-8 + + import os + import dashscope + from dashscope.audio.tts_v2 import * + + # 若没有配置环境变量,请用千问AI平台API Key将下行替换为:dashscope.api_key = "sk-xxx" + dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY') + + dashscope.base_websocket_api_url='wss://maas.qianwenaiapi.com/api-ws/v1/inference' + + # 模型 + # qwen-audio-3.0-tts-flash/qwen-audio-3.0-tts-plus:使用longanhuan_v3.6等音色。 + # 不同语言选择对应音色 + model = "qwen-audio-3.0-tts-flash" + # 音色 + voice = "longanhuan_v3.6" + + # 实例化SpeechSynthesizer,并在构造方法中传入模型(model)、音色(voice)等请求参数 + synthesizer = SpeechSynthesizer(model=model, voice=voice) + # 发送待合成文本,获取二进制音频 + audio = synthesizer.call("今天天气怎么样?") + # 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + print('[Metric] requestId为:{},首包延迟为:{}毫秒'.format( + synthesizer.get_last_request_id(), + synthesizer.get_first_package_delay())) + + # 将音频保存至本地 + with open('output.mp3', 'wb') as f: + f.write(audio) + ``` + + ```java Java expandable + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisParam; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesizer; + import com.alibaba.dashscope.utils.Constants; + + import java.io.File; + import java.io.FileOutputStream; + import java.io.IOException; + import java.nio.ByteBuffer; + + public class Main { + // 模型 + // qwen-audio-3.0-tts-flash/qwen-audio-3.0-tts-plus:使用longanhuan_v3.6等音色。 + // 每个音色支持的语言不同,合成日语、韩语等非中文语言时,需选择支持对应语言的音色。详见音色列表。 + private static String model = "qwen-audio-3.0-tts-flash"; + // 音色 + private static String voice = "longanhuan_v3.6"; + + public static void streamAudioDataToSpeaker() { + // 请求参数 + SpeechSynthesisParam param = + SpeechSynthesisParam.builder() + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:.apiKey("sk-xxx") + .apiKey(System.getenv("DASHSCOPE_API_KEY")) + .model(model) // 模型 + .voice(voice) // 音色 + .build(); + + // 同步模式:禁用回调(第二个参数为null) + SpeechSynthesizer synthesizer = new SpeechSynthesizer(param, null); + ByteBuffer audio = null; + try { + // 阻塞直至音频返回 + audio = synthesizer.call("今天天气怎么样?"); + } catch (Exception e) { + throw new RuntimeException(e); + } finally { + // 任务结束关闭websocket连接 + synthesizer.getDuplexApi().close(1000, "bye"); + } + if (audio != null) { + // 将音频数据保存到本地文件"output.mp3"中 + File file = new File("output.mp3"); + // 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + // 注意:getFirstPackageDelay() 需要 dashscope-sdk-java 2.18.0 及以上版本 + System.out.println( + "[Metric] requestId为:" + + synthesizer.getLastRequestId() + + "首包延迟(毫秒)为:" + + synthesizer.getFirstPackageDelay()); + try (FileOutputStream fos = new FileOutputStream(file)) { + fos.write(audio.array()); + } catch (IOException e) { + throw new RuntimeException(e); + } + } + } + + public static void main(String[] args) { + Constants.baseWebsocketApiUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference"; + streamAudioDataToSpeaker(); + System.exit(0); + } + } + ``` + + + + + 该模型除 WebSocket 协议外,还支持通过 AOQ 协议接入;如果是客户端对接,且更看重稳定的延迟、弱网下的交互能力、实时双工的降噪与回声消除,可优先考虑 AOQ,协议对比与选型请参见[模型/应用支持力度](/api-reference/realtime-api/overview#模型支持力度)。 + + + `cosyvoice-v3.5-plus` 和 `cosyvoice-v3.5-flash` 仅支持声音设计和声音复刻场景(无系统音色)。使用前需先通过[声音复刻](/developer-guides/speech/voice-cloning)或[声音设计](/developer-guides/speech/voice-design)创建音色,然后在代码中将 `voice` 设为音色 ID、`model` 设为对应模型名称。 + + + 以下示例演示如何使用系统音色(参见[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice))进行语音合成。 + + 如需使用[指令控制](/developer-guides/speech/realtime-streaming#指令控制)功能,请通过 `instruction` 参数设置指令。 + + + ```python Python expandable + # coding=utf-8 + + import os + import dashscope + from dashscope.audio.tts_v2 import * + + # 若没有配置环境变量,请用千问AI平台API Key将下行替换为:dashscope.api_key = "sk-xxx" + dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY') + + dashscope.base_websocket_api_url='wss://maas.qianwenaiapi.com/api-ws/v1/inference' + + # 模型 + # 不同模型版本需要使用对应版本的音色: + # cosyvoice-v3-flash/cosyvoice-v3-plus:使用longanyang等音色。 + # cosyvoice-v2:使用longxiaochun_v2等音色。 + # 不同语言选择对应音色 + model = "cosyvoice-v3-flash" + # 音色 + voice = "longanyang" + + # 实例化SpeechSynthesizer,并在构造方法中传入模型(model)、音色(voice)等请求参数 + synthesizer = SpeechSynthesizer(model=model, voice=voice) + # 发送待合成文本,获取二进制音频 + audio = synthesizer.call("今天天气怎么样?") + # 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + print('[Metric] requestId为:{},首包延迟为:{}毫秒'.format( + synthesizer.get_last_request_id(), + synthesizer.get_first_package_delay())) + + # 将音频保存至本地 + with open('output.mp3', 'wb') as f: + f.write(audio) + ``` + + ```java Java expandable + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisParam; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesizer; + import com.alibaba.dashscope.utils.Constants; + + import java.io.File; + import java.io.FileOutputStream; + import java.io.IOException; + import java.nio.ByteBuffer; + + public class Main { + // 模型 + // 不同模型版本需要使用对应版本的音色: + // cosyvoice-v3-flash/cosyvoice-v3-plus:使用longanyang等音色。 + // cosyvoice-v2:使用longxiaochun_v2等音色。 + // 每个音色支持的语言不同,合成日语、韩语等非中文语言时,需选择支持对应语言的音色。详见CosyVoice音色列表。 + private static String model = "cosyvoice-v3-flash"; + // 音色 + private static String voice = "longanyang"; + + public static void streamAudioDataToSpeaker() { + // 请求参数 + SpeechSynthesisParam param = + SpeechSynthesisParam.builder() + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:.apiKey("sk-xxx") + .apiKey(System.getenv("DASHSCOPE_API_KEY")) + .model(model) // 模型 + .voice(voice) // 音色 + .build(); + + // 同步模式:禁用回调(第二个参数为null) + SpeechSynthesizer synthesizer = new SpeechSynthesizer(param, null); + ByteBuffer audio = null; + try { + // 阻塞直至音频返回 + audio = synthesizer.call("今天天气怎么样?"); + } catch (Exception e) { + throw new RuntimeException(e); + } finally { + // 任务结束关闭websocket连接 + synthesizer.getDuplexApi().close(1000, "bye"); + } + if (audio != null) { + // 将音频数据保存到本地文件"output.mp3"中 + File file = new File("output.mp3"); + // 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + // 注意:getFirstPackageDelay() 需要 dashscope-sdk-java 2.18.0 及以上版本 + System.out.println( + "[Metric] requestId为:" + + synthesizer.getLastRequestId() + + "首包延迟(毫秒)为:" + + synthesizer.getFirstPackageDelay()); + try (FileOutputStream fos = new FileOutputStream(file)) { + fos.write(audio.array()); + } catch (IOException e) { + throw new RuntimeException(e); + } + } + } + + public static void main(String[] args) { + Constants.baseWebsocketApiUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference"; + streamAudioDataToSpeaker(); + System.exit(0); + } + } + ``` + + + + +## 进阶功能 + +### 指令控制 + +指令控制通过自然语言描述控制语音的音调、语速、情感和音色特点,无需调整复杂的音频参数。 + +**各模型指令规格**: + + + + **支持的模型**:`qwen-audio-3.1-tts-flash`、`qwen-audio-3.0-tts-plus`、`qwen-audio-3.0-tts-flash` + + 系统音色和声音复刻音色:均可输入任意指令。 + + + + **支持的模型**:`cosyvoice-v3.5-plus`、`cosyvoice-v3.5-flash`、`cosyvoice-v3-plus`、`cosyvoice-v3-flash` + + 不同模型对指令的格式要求不同: + + - `cosyvoice-v3.5-plus`、`cosyvoice-v3.5-flash`: + + - 声音复刻/设计音色:可输入任意指令。 + - 系统音色:v3.5不支持系统音色。 + - `cosyvoice-v3-plus`: + + - 声音复刻/设计音色:不支持指令控制。 + - 系统音色:指令必须使用固定格式和内容,参见[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)。 + - `cosyvoice-v3-flash`: + + - 声音复刻/设计音色:可输入任意指令。 + - 系统音色:指令必须使用固定格式和内容,参见[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)。 + + **使用方式**:通过 `instruction` 参数指定指令内容。 + + **指令文本支持的语言**: + + - `cosyvoice-v3.5-plus`、`cosyvoice-v3.5-flash`: + + - 声音复刻/设计音色:中文、英文、法语、德语、日语、韩语、俄语、葡萄牙语、泰语、印尼语、越南语。 + - 系统音色:v3.5不支持系统音色。 + - `cosyvoice-v3-plus`: + + - 声音复刻/设计音色:中文、英文、法语、德语、日语、韩语、俄语。 + - 系统音色:指令必须使用固定格式和内容,参见[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)。 + - `cosyvoice-v3-flash`: + + - 声音复刻/设计音色:中文、英文、法语、德语、日语、韩语、俄语。 + - 系统音色:中文。 + + **指令文本长度限制**:不超过 100 字符。汉字(包括简体/繁体汉字、日文汉字和韩文汉字)按 2 个字符计算,其他字符(如标点符号、字母、数字、日韩文假名/谚文等)按 1 个字符计算。 + + + + **支持的模型**:仅支持Qwen3-TTS-Instruct-Flash-Realtime系列模型。 + + **使用方式**:通过 `instructions` 参数指定指令内容。 + + **指令文本支持的语言**:仅支持中文和英文。 + + **指令文本长度限制**:不超过 1600 Token。 + + + +**适用场景**: + +- 有声书和广播剧配音 +- 广告和宣传片配音 +- 游戏角色和动画配音 +- 情感化的智能语音助手 +- 纪录片和新闻播报 + +**如何编写高质量的声音描述**: + +- **核心原则**: + + 1. **具体而非模糊**:使用描绘声音特质的词语,如“低沉”、“清脆”、“语速偏快”,避免“好听”、“普通”等主观或模糊的表述。 + 2. **多维而非单一**:好的描述通常涵盖多个维度(如性别、年龄、情感等)。仅写“女声”过于宽泛,难以生成有特色的音色。 + 3. **客观而非主观**:聚焦声音的物理和感知特征。例如,用”音调偏高,带有活力“代替”我最喜欢的声音”。 + 4. **原创而非模仿**:描述声音的特质,而非要求模仿特定人物(如名人、演员)。模型不支持模仿,且可能涉及版权风险。 + 5. **简洁而非冗余**:确保每个词都有明确作用,避免重复的同义词或无意义的修饰。 +- **描述维度参考**: + + 建议组合以下维度描述声音,维度越丰富,生成效果越精准。 + +| **维度** | **描述示例** | +| ------ | ---------------------------------------------------------- | +| 性别 | 男性、女性、中性 | +| 年龄 | 儿童(5-12 岁)、青少年(13-18 岁)、青年(19-35 岁)、中年(36-55 岁)、老年(55 岁以上) | +| 音调 | 高音、中音、低音、偏高、偏低 | +| 语速 | 快速、中速、缓慢、偏快、偏慢 | +| 情感 | 开朗、沉稳、温柔、严肃、活泼、冷静、治愈 | +| 特点 | 有磁性、清脆、沙哑、圆润、甜美、浑厚、有力 | +| 用途 | 新闻播报、广告配音、有声书、动画角色、语音助手、纪录片解说 | + +- **示例**: + + - 标准播音风格:吐字清晰精准,字正腔圆 + - 年轻活泼的女性声音,语速较快,带有明显的上扬语调,适合介绍时尚产品 + - 沉稳的中年男性,语速缓慢,音色低沉有磁性,适合朗读新闻或纪录片解说 + - 温柔知性的女性,30 岁左右,语调平和,适合有声书朗读 + - 可爱的儿童声音,大约 8 岁女孩,说话略带稚气,适合动画角色配音 + +### 方言 + +本节介绍如何让模型用**中文方言**(如河南话、四川话、粤语等)输出语音。不同模型和音色类型的设置方式不同。 + +**各模型方言设置方式**: + + + + - **系统音色**:在[Qwen-Audio-TTS音色列表](/developer-guides/speech/voice-list/qwen-audio-tts)中选择以下任一种音色: + + - 支持方言的系统音色,无需额外设置即可输出对应方言。 + - 支持[指令控制](/developer-guides/speech/realtime-streaming#指令控制)且可指定方言的音色,通过指令文本指定方言。 + - **声音复刻音色**:通过[指令控制](/developer-guides/speech/realtime-streaming#指令控制)功能设置,例如指令文本写 `请用河南话表达`。 + + **具体支持哪些方言**:参见[Qwen-Audio-TTS](/developer-guides/speech/tts-models)中各模型“支持的语言”。 + + + + - **系统音色**:在[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)中选择以下任一种音色: + + - 支持方言的系统音色(例如 `longshange_v3`),无需额外设置即可输出对应方言。 + - 支持[指令控制](/developer-guides/speech/realtime-streaming#指令控制)且可指定方言的音色(例如 `longanhuan_v3`),通过指令文本指定方言。 + - **声音复刻音色**:通过[指令控制](/developer-guides/speech/realtime-streaming#指令控制)功能设置,例如指令文本写 `请用河南话表达`。 + - **声音设计音色**:暂不支持方言。 + + **具体支持哪些方言**:参见[CosyVoice](/developer-guides/speech/tts-models)中各模型“支持的语言”。 + + **示例**:以 `cosyvoice-v3-flash` + `longanhuan_v3` 音色,通过指令文本 `"请用河南话表达。"` 输出河南话语音。 + + ```python expandable + # coding=utf-8 + + import os + import dashscope + from dashscope.audio.tts_v2 import * + + # 若没有配置环境变量,请用千问AI平台API Key将下行替换为:dashscope.api_key = "sk-xxx" + dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY') + + dashscope.base_websocket_api_url='wss://maas.qianwenaiapi.com/api-ws/v1/inference' + + # 模型 + # 不同模型版本需要使用对应版本的音色: + # cosyvoice-v3-flash/cosyvoice-v3-plus:使用longanyang等音色。 + # cosyvoice-v2:使用longxiaochun_v2等音色。 + # 不同语言选择对应音色 + model = "cosyvoice-v3-flash" + # 音色 + voice = "longanhuan_v3" + + # 实例化SpeechSynthesizer,并在构造方法中传入模型(model)、音色(voice)等请求参数 + synthesizer = SpeechSynthesizer(model=model, voice=voice, instruction="请用河南话表达。") + # 发送待合成文本,获取二进制音频 + audio = synthesizer.call("叫你去买盐,你买回来一袋面,这不是弄啥嘞吗!") + # 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + print('[Metric] requestId为:{},首包延迟为:{}毫秒'.format( + synthesizer.get_last_request_id(), + synthesizer.get_first_package_delay())) + + # 将音频保存至本地 + with open('output.mp3', 'wb') as f: + f.write(audio) + ``` + + + + - **系统音色**:使用支持方言的系统音色,参见[支持的音色](/developer-guides/speech/realtime-streaming#支持的音色)中的 Qwen-TTS音色列表。 + - **声音复刻音色**:不支持方言。 + - **声音设计音色**:不支持方言。 + + **具体支持哪些方言**:参见[Qwen3-TTS](/developer-guides/speech/tts-models)中各模型“支持的语言”。 + + + +### 情感与富语言标签 + +Qwen-Audio-TTS 系列模型支持在待合成文本(`text` 参数)中直接嵌入情感与富语言标签,用于控制语音的情感表达或在指定位置插入拟声效果(如笑声、叹息等),无需调整复杂的音频参数即可生成更具表现力的语音。 + + + **支持的模型**:仅 `qwen-audio-3.1-tts-flash`、`qwen-audio-3.0-tts-plus` 和 `qwen-audio-3.0-tts-flash`。 + + **限制**:仅支持单向流式模式。 + + +**控制类标签** + +控制类标签用于设定语音的情感或风格。将标签写在文本中,标签会作用于其后的所有文本,直到遇到下一个控制类标签,或因句子较长被自动切分为止。 + +| **标签** | **说明** | +| -------------------------- | ------------ | +| `[sad]` | 悲伤 | +| `[amazed]` | 惊叹 | +| `[deep and loud shouting]` | 深沉大声呐喊 | +| `[trembling]` | 颤抖 | +| `[angry]` | 愤怒 | +| `[excited]` | 兴奋 | +| `[sarcastic]` | 讽刺 | +| `[curious]` | 好奇 | +| `[like dracula]` | 德古拉风格(低沉、阴森) | +| `[bored]` | 无聊 | +| `[tired]` | 疲惫 | +| `[scornful]` | 轻蔑 | +| `[shouting]` | 大喊 | +| `[asmr]` | ASMR 轻柔耳语 | +| `[panicked]` | 恐慌 | +| `[mischievously]` | 调皮 | +| `[empathetic]` | 共情 | +| `[whispers]` | 耳语 | +| `[reluctantly]` | 不情愿 | +| `[crying]` | 哭泣 | +| `[serious]` | 严肃 | +| `[very slowly]` | 非常缓慢地说话 | +| `[very fast]` | 非常快速地说话 | + +**富语言类标签** + +富语言类标签用于在文本的当前位置插入一段拟声效果,不影响前后文本的情感风格。 + +| **标签** | **说明** | +| ----------------- | ------ | +| `[gasp]` | 倒吸一口气 | +| `[sighing]` | 叹息 | +| `[clears throat]` | 清嗓 | +| `[giggles]` | 咯咯笑 | +| `[laughing]` | 大笑 | +| `[cough]` | 咳嗽 | +| `[snorts]` | 哼声、嗤笑 | + +**使用示例** + +以下示例展示如何在 `text` 参数中组合使用控制类标签和富语言类标签: + +`[excited]今天的天气真不错![laughing]我们一起出去玩吧!` + +上述文本中,`[excited]` 是控制类标签,作用于其后的所有文本,使语音带有兴奋的情感;`[laughing]` 是富语言类标签,在该位置插入一段笑声效果后继续合成后续文本。 + +您也可以在同一段文本中切换不同情感: + +`[serious]请注意安全事项。[excited]好了,现在让我们开始吧!` + +其中 `[serious]` 控制第一句为严肃语气,`[excited]` 从第二句起切换为兴奋语气。 + +### 取消任务 + +在实时语音合成过程中,如果需要中断当前轮次合成,可以发送取消指令。取消后服务端会立即结束当前任务并返回结束事件,您可在当前 WebSocket 连接上继续发起新的合成任务,无需重新建立连接。 + +**使用方式**: + +- **Python SDK**:1.26.4 及以上版本,调用 `SpeechSynthesizer.streaming_cancel()`。 +- **Java SDK**:2.22.26 及以上版本,调用 `SpeechSynthesizer.streamingCancel()`。 +- **WebSocket 原始协议**:发送 `finish-task` 事件,并在 `input` 中设置 `directive=cancel`。 + + + **模型限制**: + + - Qwen-Audio-TTS 系列模型的所有模型都支持该功能;CosyVoice 系列模型仅 v2 及以上版本支持该功能。 + + +### WebSocket 原始协议调用 + +以下示例展示如何通过 WebSocket 原始协议直连服务端,适用于不使用 DashScope SDK 的场景。此为最小可运行实现,WebSocket 协议请参见各模型的 API 参考。 + + + + + Qwen-Audio-TTS 和 CosyVoice 使用相同的 WebSocket 协议,只需替换 `model` 和 `voice` 参数。以下示例以 `qwen-audio-3.0-tts-flash` 为例,使用 CosyVoice 时将 model 替换为 `cosyvoice-v3-flash` 等,voice 替换为对应音色即可。 + + + + ```go expandable + package main + + import ( + "encoding/json" + "fmt" + "net/http" + "os" + "strings" + "time" + + "github.com/google/uuid" + "github.com/gorilla/websocket" + ) + + const ( + wsURL = "wss://maas.qianwenaiapi.com/api-ws/v1/inference" + outputFile = "output.mp3" + ) + + func main() { + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:apiKey := "sk-xxx" + apiKey := os.Getenv("DASHSCOPE_API_KEY") + + // 清空输出文件 + os.Remove(outputFile) + os.Create(outputFile) + + // 连接WebSocket + header := make(http.Header) + header.Add("X-DashScope-DataInspection", "enable") + header.Add("Authorization", fmt.Sprintf("bearer %s", apiKey)) + + conn, resp, err := websocket.DefaultDialer.Dial(wsURL, header) + if err != nil { + if resp != nil { + fmt.Printf("连接失败 HTTP状态码: %d\n", resp.StatusCode) + } + fmt.Println("连接失败:", err) + return + } + defer conn.Close() + + // 生成任务ID + taskID := uuid.New().String() + fmt.Printf("生成任务ID: %s\n", taskID) + + // 发送run-task事件 + runTaskCmd := map[string]interface{}{ + "header": map[string]interface{}{ + "action": "run-task", + "task_id": taskID, + "streaming": "duplex", + }, + "payload": map[string]interface{}{ + "task_group": "audio", + "task": "tts", + "function": "SpeechSynthesizer", + "model": "qwen-audio-3.0-tts-flash", + "parameters": map[string]interface{}{ + "text_type": "PlainText", + "voice": "longanhuan_v3.6", + "format": "mp3", + "sample_rate": 22050, + "volume": 50, + "rate": 1, + "pitch": 1, + // 如果enable_ssml设为true,只允许发送一次continue-task事件,否则会报错“Text request limit violated, expected 1.” + "enable_ssml": false, + }, + "input": map[string]interface{}{}, + }, + } + + runTaskJSON, _ := json.Marshal(runTaskCmd) + fmt.Printf("发送run-task事件: %s\n", string(runTaskJSON)) + + err = conn.WriteMessage(websocket.TextMessage, runTaskJSON) + if err != nil { + fmt.Println("发送run-task失败:", err) + return + } + + textSent := false + + // 处理消息 + for { + messageType, message, err := conn.ReadMessage() + if err != nil { + fmt.Println("读取消息失败:", err) + break + } + + // 处理二进制消息 + if messageType == websocket.BinaryMessage { + fmt.Printf("收到二进制消息,长度: %d\n", len(message)) + file, _ := os.OpenFile(outputFile, os.O_APPEND|os.O_WRONLY|os.O_CREATE, 0644) + file.Write(message) + file.Close() + continue + } + + // 处理文本消息 + messageStr := string(message) + fmt.Printf("收到文本消息: %s\n", strings.ReplaceAll(messageStr, "\n", "")) + + // 简单解析JSON获取event类型 + var msgMap map[string]interface{} + if json.Unmarshal(message, &msgMap) == nil { + if header, ok := msgMap["header"].(map[string]interface{}); ok { + if event, ok := header["event"].(string); ok { + fmt.Printf("事件类型: %s\n", event) + + switch event { + case "task-started": + fmt.Println("=== 收到task-started事件 ===") + + if !textSent { + // 发送continue-task事件 + + texts := []string{"床前明月光,疑是地上霜。", "举头望明月,低头思故乡。"} + + for _, text := range texts { + continueTaskCmd := map[string]interface{}{ + "header": map[string]interface{}{ + "action": "continue-task", + "task_id": taskID, + "streaming": "duplex", + }, + "payload": map[string]interface{}{ + "input": map[string]interface{}{ + "text": text, + }, + }, + } + + continueTaskJSON, _ := json.Marshal(continueTaskCmd) + fmt.Printf("发送continue-task事件: %s\n", string(continueTaskJSON)) + + err = conn.WriteMessage(websocket.TextMessage, continueTaskJSON) + if err != nil { + fmt.Println("发送continue-task失败:", err) + return + } + } + + textSent = true + + // 延迟发送finish-task + time.Sleep(500 * time.Millisecond) + + // 发送finish-task事件 + finishTaskCmd := map[string]interface{}{ + "header": map[string]interface{}{ + "action": "finish-task", + "task_id": taskID, + "streaming": "duplex", + }, + "payload": map[string]interface{}{ + "input": map[string]interface{}{}, + }, + } + + finishTaskJSON, _ := json.Marshal(finishTaskCmd) + fmt.Printf("发送finish-task事件: %s\n", string(finishTaskJSON)) + + err = conn.WriteMessage(websocket.TextMessage, finishTaskJSON) + if err != nil { + fmt.Println("发送finish-task失败:", err) + return + } + } + + case "task-finished": + fmt.Println("=== 任务完成 ===") + return + + case "task-failed": + fmt.Println("=== 任务失败 ===") + if header["error_message"] != nil { + fmt.Printf("错误信息: %s\n", header["error_message"]) + } + return + + case "result-generated": + fmt.Println("收到result-generated事件") + } + } + } + } + } + } + ``` + + + + ```csharp expandable + using System.Net.WebSockets; + using System.Text; + using System.Text.Json; + + class Program { + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:private static readonly string ApiKey = "sk-xxx" + private static readonly string ApiKey = Environment.GetEnvironmentVariable("DASHSCOPE_API_KEY") ?? throw new InvalidOperationException("DASHSCOPE_API_KEY environment variable is not set."); + + private const string WebSocketUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference"; + // 输出文件路径 + private const string OutputFilePath = "output.mp3"; + + // WebSocket客户端 + private static ClientWebSocket _webSocket = new ClientWebSocket(); + // 取消令牌源 + private static CancellationTokenSource _cancellationTokenSource = new CancellationTokenSource(); + // 任务ID + private static string? _taskId; + // 任务是否已启动 + private static TaskCompletionSource _taskStartedTcs = new TaskCompletionSource(); + + static async Task Main(string[] args) { + try { + // 清空输出文件 + ClearOutputFile(OutputFilePath); + + // 连接WebSocket服务 + await ConnectToWebSocketAsync(WebSocketUrl); + + // 启动接收消息的任务 + Task receiveTask = ReceiveMessagesAsync(); + + // 发送run-task事件 + _taskId = GenerateTaskId(); + await SendRunTaskCommandAsync(_taskId); + + // 等待task-started事件 + await _taskStartedTcs.Task; + + // 持续发送continue-task事件 + string[] texts = { + "床前明月光", + "疑是地上霜", + "举头望明月", + "低头思故乡" + }; + foreach (string text in texts) { + await SendContinueTaskCommandAsync(text); + } + + // 发送finish-task事件 + await SendFinishTaskCommandAsync(_taskId); + + // 等待接收任务完成 + await receiveTask; + + Console.WriteLine("任务完成,连接已关闭。"); + } catch (OperationCanceledException) { + Console.WriteLine("任务被取消。"); + } catch (Exception ex) { + Console.WriteLine($"发生错误:{ex.Message}"); + } finally { + _cancellationTokenSource.Cancel(); + _webSocket.Dispose(); + } + } + + private static void ClearOutputFile(string filePath) { + if (File.Exists(filePath)) { + File.WriteAllText(filePath, string.Empty); + Console.WriteLine("输出文件已清空。"); + } else { + Console.WriteLine("输出文件不存在,无需清空。"); + } + } + + private static async Task ConnectToWebSocketAsync(string url) { + var uri = new Uri(url); + if (_webSocket.State == WebSocketState.Connecting || _webSocket.State == WebSocketState.Open) { + return; + } + + // 设置WebSocket连接的头部信息 + _webSocket.Options.SetRequestHeader("Authorization", $"bearer {ApiKey}"); + _webSocket.Options.SetRequestHeader("X-DashScope-DataInspection", "enable"); + + try { + await _webSocket.ConnectAsync(uri, _cancellationTokenSource.Token); + Console.WriteLine("已成功连接到WebSocket服务。"); + } catch (OperationCanceledException) { + Console.WriteLine("WebSocket连接被取消。"); + } catch (Exception ex) { + Console.WriteLine($"WebSocket连接失败: {ex.Message}"); + throw; + } + } + + private static async Task SendRunTaskCommandAsync(string taskId) { + var command = CreateCommand("run-task", taskId, "duplex", new { + task_group = "audio", + task = "tts", + function = "SpeechSynthesizer", + model = "qwen-audio-3.0-tts-flash", + parameters = new + { + text_type = "PlainText", + voice = "longanhuan_v3.6", + format = "mp3", + sample_rate = 22050, + volume = 50, + rate = 1, + pitch = 1, + // 如果enable_ssml设为true,只允许发送一次continue-task事件,否则会报错“Text request limit violated, expected 1.” + enable_ssml = false + }, + input = new { } + }); + + await SendJsonMessageAsync(command); + Console.WriteLine("已发送run-task事件。"); + } + + private static async Task SendContinueTaskCommandAsync(string text) { + if (_taskId == null) { + throw new InvalidOperationException("任务ID未初始化。"); + } + + var command = CreateCommand("continue-task", _taskId, "duplex", new { + input = new { + text + } + }); + + await SendJsonMessageAsync(command); + Console.WriteLine("已发送continue-task事件。"); + } + + private static async Task SendFinishTaskCommandAsync(string taskId) { + var command = CreateCommand("finish-task", taskId, "duplex", new { + input = new { } + }); + + await SendJsonMessageAsync(command); + Console.WriteLine("已发送finish-task事件。"); + } + + private static async Task SendJsonMessageAsync(string message) { + var buffer = Encoding.UTF8.GetBytes(message); + try { + await _webSocket.SendAsync(new ArraySegment(buffer), WebSocketMessageType.Text, true, _cancellationTokenSource.Token); + } catch (OperationCanceledException) { + Console.WriteLine("消息发送被取消。"); + } + } + + private static async Task ReceiveMessagesAsync() { + while (_webSocket.State == WebSocketState.Open) { + var response = await ReceiveMessageAsync(); + if (response != null) { + var eventStr = response.RootElement.GetProperty("header").GetProperty("event").GetString(); + switch (eventStr) { + case "task-started": + Console.WriteLine("任务已启动。"); + _taskStartedTcs.TrySetResult(true); + break; + case "task-finished": + Console.WriteLine("任务已完成。"); + _cancellationTokenSource.Cancel(); + break; + case "task-failed": + Console.WriteLine("任务失败:" + response.RootElement.GetProperty("header").GetProperty("error_message").GetString()); + _cancellationTokenSource.Cancel(); + break; + default: + // result-generated可在此处理 + break; + } + } + } + } + + private static async Task ReceiveMessageAsync() { + var buffer = new byte[1024 * 4]; + var segment = new ArraySegment(buffer); + + try { + WebSocketReceiveResult result = await _webSocket.ReceiveAsync(segment, _cancellationTokenSource.Token); + + if (result.MessageType == WebSocketMessageType.Close) { + await _webSocket.CloseAsync(WebSocketCloseStatus.NormalClosure, "Closing", _cancellationTokenSource.Token); + return null; + } + + if (result.MessageType == WebSocketMessageType.Binary) { + // 处理二进制数据 + Console.WriteLine("接收到二进制数据..."); + + // 将二进制数据保存到文件 + using (var fileStream = new FileStream(OutputFilePath, FileMode.Append)) { + fileStream.Write(buffer, 0, result.Count); + } + + return null; + } + + string message = Encoding.UTF8.GetString(buffer, 0, result.Count); + return JsonDocument.Parse(message); + } catch (OperationCanceledException) { + Console.WriteLine("消息接收被取消。"); + return null; + } + } + + private static string GenerateTaskId() { + return Guid.NewGuid().ToString("N").Substring(0, 32); + } + + private static string CreateCommand(string action, string taskId, string streaming, object payload) { + var command = new { + header = new { + action, + task_id = taskId, + streaming + }, + payload + }; + + return JsonSerializer.Serialize(command); + } + } + ``` + + + + 示例代码目录结构为: + + my-php-project/ + + ├── composer.json + + ├── vendor/ + + └── index.php + + composer.json内容如下,相关依赖的版本号请根据实际情况自行决定: + + ```json + { + "require": { + "react/event-loop": "^1.3", + "react/socket": "^1.11", + "react/stream": "^1.2", + "react/http": "^1.1", + "ratchet/pawl": "^0.4" + }, + "autoload": { + "psr-4": { + "App\\": "src/" + } + } + } + ``` + + index.php内容如下: + + ```php expandable + [ + 'bindto' => '0.0.0.0:0', + ], + // 警告:关闭 TLS 证书校验存在中间人攻击风险,仅限本地调试;生产环境务必将 verify_peer/verify_peer_name 设为 true + 'tls' => [ + 'verify_peer' => false, + 'verify_peer_name' => false, + ], + ]); + + $connector = new Connector($loop, $socketConnector); + + $headers = [ + 'Authorization' => 'bearer ' . $api_key, + 'X-DashScope-DataInspection' => 'enable' + ]; + + $connector($websocket_url, [], $headers)->then(function ($conn) use ($loop, $output_file) { + echo "连接到WebSocket服务器\n"; + + // 生成任务ID + $taskId = generateTaskId(); + + // 发送 run-task 事件 + sendRunTaskMessage($conn, $taskId); + + // 定义发送 continue-task 事件的函数 + $sendContinueTask = function() use ($conn, $loop, $taskId) { + // 待发送的文本 + $texts = ["床前明月光", "疑是地上霜", "举头望明月", "低头思故乡"]; + $continueTaskCount = 0; + foreach ($texts as $text) { + $continueTaskMessage = json_encode([ + "header" => [ + "action" => "continue-task", + "task_id" => $taskId, + "streaming" => "duplex" + ], + "payload" => [ + "input" => [ + "text" => $text + ] + ] + ]); + echo "准备发送continue-task事件: " . $continueTaskMessage . "\n"; + $conn->send($continueTaskMessage); + $continueTaskCount++; + } + echo "发送的continue-task事件个数为:" . $continueTaskCount . "\n"; + + // 发送 finish-task 事件 + sendFinishTaskMessage($conn, $taskId); + }; + + // 标记是否收到 task-started 事件 + $taskStarted = false; + + // 监听消息 + $conn->on('message', function($msg) use ($conn, $sendContinueTask, $loop, &$taskStarted, $taskId, $output_file) { + if ($msg->isBinary()) { + // 写入二进制数据到本地文件 + file_put_contents($output_file, $msg->getPayload(), FILE_APPEND); + } else { + // 处理非二进制消息 + $response = json_decode($msg, true); + + if (isset($response['header']['event'])) { + handleEvent($conn, $response, $sendContinueTask, $loop, $taskId, $taskStarted); + } else { + echo "未知的消息格式\n"; + } + } + }); + + // 监听连接关闭 + $conn->on('close', function($code = null, $reason = null) { + echo "连接已关闭\n"; + if ($code !== null) { + echo "关闭代码: " . $code . "\n"; + } + if ($reason !== null) { + echo "关闭原因:" . $reason . "\n"; + } + }); + }, function ($e) { + echo "无法连接:{$e->getMessage()}\n"; + }); + + $loop->run(); + + /** + * 生成任务ID + * @return string + */ + function generateTaskId(): string { + return bin2hex(random_bytes(16)); + } + + /** + * 发送 run-task 事件 + * @param $conn + * @param $taskId + */ + function sendRunTaskMessage($conn, $taskId) { + $runTaskMessage = json_encode([ + "header" => [ + "action" => "run-task", + "task_id" => $taskId, + "streaming" => "duplex" + ], + "payload" => [ + "task_group" => "audio", + "task" => "tts", + "function" => "SpeechSynthesizer", + "model" => "qwen-audio-3.0-tts-flash", + "parameters" => [ + "text_type" => "PlainText", + "voice" => "longanhuan_v3.6", + "format" => "mp3", + "sample_rate" => 22050, + "volume" => 50, + "rate" => 1, + "pitch" => 1, + // 如果enable_ssml设为true,只允许发送一次continue-task事件,否则会报错“Text request limit violated, expected 1.” + "enable_ssml" => false + ], + "input" => (object) [] + ] + ]); + echo "准备发送run-task事件: " . $runTaskMessage . "\n"; + $conn->send($runTaskMessage); + echo "run-task事件已发送\n"; + } + + /** + * 读取音频文件 + * @param string $filePath + * @return bool|string + */ + function readAudioFile(string $filePath) { + $voiceData = file_get_contents($filePath); + if ($voiceData === false) { + echo "无法读取音频文件\n"; + } + return $voiceData; + } + + /** + * 分割音频数据 + * @param string $data + * @param int $chunkSize + * @return array + */ + function splitAudioData(string $data, int $chunkSize): array { + return str_split($data, $chunkSize); + } + + /** + * 发送 finish-task 事件 + * @param $conn + * @param $taskId + */ + function sendFinishTaskMessage($conn, $taskId) { + $finishTaskMessage = json_encode([ + "header" => [ + "action" => "finish-task", + "task_id" => $taskId, + "streaming" => "duplex" + ], + "payload" => [ + "input" => (object) [] + ] + ]); + echo "准备发送finish-task事件: " . $finishTaskMessage . "\n"; + $conn->send($finishTaskMessage); + echo "finish-task事件已发送\n"; + } + + /** + * 处理事件 + * @param $conn + * @param $response + * @param $sendContinueTask + * @param $loop + * @param $taskId + * @param $taskStarted + */ + function handleEvent($conn, $response, $sendContinueTask, $loop, $taskId, &$taskStarted) { + switch ($response['header']['event']) { + case 'task-started': + echo "任务开始,发送continue-task事件...\n"; + $taskStarted = true; + // 发送 continue-task 事件 + $sendContinueTask(); + break; + case 'result-generated': + // 收到result-generated事件 + break; + case 'task-finished': + echo "任务完成\n"; + $conn->close(); + break; + case 'task-failed': + echo "任务失败\n"; + echo "错误代码:" . $response['header']['error_code'] . "\n"; + echo "错误信息:" . $response['header']['error_message'] . "\n"; + $conn->close(); + break; + case 'error': + echo "错误:" . $response['payload']['message'] . "\n"; + break; + default: + echo "未知事件:" . $response['header']['event'] . "\n"; + break; + } + + // 如果任务已完成,关闭连接 + if ($response['header']['event'] == 'task-finished') { + // 等待1秒以确保所有数据都已传输完毕 + $loop->addTimer(1, function() use ($conn) { + $conn->close(); + echo "客户端关闭连接\n"; + }); + } + + // 如果没有收到 task-started 事件,关闭连接 + if (!$taskStarted && in_array($response['header']['event'], ['task-failed', 'error'])) { + $conn->close(); + } + } + ``` + + + + 需安装相关依赖: + + ```bash + npm install ws + npm install uuid + ``` + + 示例代码如下: + + ```javascript expandable + const WebSocket = require('ws'); + const fs = require('fs'); + const uuid = require('uuid').v4; + + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:const apiKey = "sk-xxx" + const apiKey = process.env.DASHSCOPE_API_KEY; + const url = 'wss://maas.qianwenaiapi.com/api-ws/v1/inference'; + // 输出文件路径 + const outputFilePath = 'output.mp3'; + + // 清空输出文件 + fs.writeFileSync(outputFilePath, ''); + + // 创建WebSocket客户端 + const ws = new WebSocket(url, { + headers: { + Authorization: `bearer ${apiKey}`, + 'X-DashScope-DataInspection': 'enable' + } + }); + + let taskStarted = false; + let taskId = uuid(); + + ws.on('open', () => { + console.log('已连接到WebSocket服务器'); + + // 发送run-task事件 + const runTaskMessage = JSON.stringify({ + header: { + action: 'run-task', + task_id: taskId, + streaming: 'duplex' + }, + payload: { + task_group: 'audio', + task: 'tts', + function: 'SpeechSynthesizer', + model: 'qwen-audio-3.0-tts-flash', + parameters: { + text_type: 'PlainText', + voice: 'longanhuan_v3.6', // 音色 + format: 'mp3', // 音频格式 + sample_rate: 22050, // 采样率 + volume: 50, // 音量 + rate: 1, // 语速 + pitch: 1, // 音调 + enable_ssml: false // 是否开启SSML功能。如果enable_ssml设为true,只允许发送一次continue-task事件,否则会报错“Text request limit violated, expected 1.” + }, + input: {} + } + }); + ws.send(runTaskMessage); + console.log('已发送run-task消息'); + }); + + const fileStream = fs.createWriteStream(outputFilePath, { flags: 'a' }); + ws.on('message', (data, isBinary) => { + if (isBinary) { + // 写入二进制数据到文件 + fileStream.write(data); + } else { + const message = JSON.parse(data); + + switch (message.header.event) { + case 'task-started': + taskStarted = true; + console.log('任务已开始'); + // 发送continue-task事件 + sendContinueTasks(ws); + break; + case 'task-finished': + console.log('任务已完成'); + ws.close(); + fileStream.end(() => { + console.log('文件流已关闭'); + }); + break; + case 'task-failed': + console.error('任务失败:', message.header.error_message); + ws.close(); + fileStream.end(() => { + console.log('文件流已关闭'); + }); + break; + default: + // 可以在这里处理result-generated + break; + } + } + }); + + function sendContinueTasks(ws) { + const texts = [ + '床前明月光,', + '疑是地上霜。', + '举头望明月,', + '低头思故乡。' + ]; + + texts.forEach((text, index) => { + setTimeout(() => { + if (taskStarted) { + const continueTaskMessage = JSON.stringify({ + header: { + action: 'continue-task', + task_id: taskId, + streaming: 'duplex' + }, + payload: { + input: { + text: text + } + } + }); + ws.send(continueTaskMessage); + console.log(`已发送continue-task,文本:${text}`); + } + }, index * 1000); // 每隔1秒发送一次 + }); + + // 发送finish-task事件 + setTimeout(() => { + if (taskStarted) { + const finishTaskMessage = JSON.stringify({ + header: { + action: 'finish-task', + task_id: taskId, + streaming: 'duplex' + }, + payload: { + input: {} + } + }); + ws.send(finishTaskMessage); + console.log('已发送finish-task'); + } + }, texts.length * 1000 + 1000); // 在所有continue-task事件发送完毕后1秒发送 + } + + ws.on('close', () => { + console.log('已断开与WebSocket服务器的连接'); + }); + ``` + + + + 建议使用 Java DashScope SDK 进行开发,请参见[Qwen-Audio-TTS Java SDK](/api-reference/speech-synthesis/qwen-audio-tts/java-sdk) / [CosyVoice Java SDK](/api-reference/speech-synthesis/cosyvoice/java-sdk)。 + + 以下是 Java WebSocket 直连示例,运行前请导入以下依赖: + + - `Java-WebSocket` + - `jackson-databind` + + 推荐使用Maven或Gradle管理依赖包,其配置如下: + + + ```xml pom.xml + + + + org.java-websocket + Java-WebSocket + 1.5.3 + + + + + com.fasterxml.jackson.core + jackson-databind + 2.13.0 + + + ``` + + ```gradle build.gradle + // 省略其它代码 + dependencies { + // WebSocket Client + implementation 'org.java-websocket:Java-WebSocket:1.5.3' + // JSON Processing + implementation 'com.fasterxml.jackson.core:jackson-databind:2.13.0' + } + // 省略其它代码 + ``` + + + Java代码如下: + + ```java expandable + import com.fasterxml.jackson.databind.ObjectMapper; + + import org.java_websocket.client.WebSocketClient; + import org.java_websocket.handshake.ServerHandshake; + + import java.io.FileOutputStream; + import java.io.IOException; + import java.net.URI; + import java.nio.ByteBuffer; + import java.util.*; + + public class TTSWebSocketClient extends WebSocketClient { + private final String taskId = UUID.randomUUID().toString(); + private final String outputFile = "output_" + System.currentTimeMillis() + ".mp3"; + private boolean taskFinished = false; + + public TTSWebSocketClient(URI serverUri, Map headers) { + super(serverUri, headers); + } + + @Override + public void onOpen(ServerHandshake serverHandshake) { + System.out.println("连接成功"); + + // 发送run-task事件 + // 如果enable_ssml设为true,只允许发送一次continue-task事件,否则会报错“Text request limit violated, expected 1.” + String runTaskCommand = "{ \"header\": { \"action\": \"run-task\", \"task_id\": \"" + taskId + "\", \"streaming\": \"duplex\" }, \"payload\": { \"task_group\": \"audio\", \"task\": \"tts\", \"function\": \"SpeechSynthesizer\", \"model\": \"qwen-audio-3.0-tts-flash\", \"parameters\": { \"text_type\": \"PlainText\", \"voice\": \"longanhuan_v3.6\", \"format\": \"mp3\", \"sample_rate\": 22050, \"volume\": 50, \"rate\": 1, \"pitch\": 1, \"enable_ssml\": false }, \"input\": {} }}"; + send(runTaskCommand); + } + + @Override + public void onMessage(String message) { + System.out.println("收到服务端返回的消息:" + message); + try { + // Parse JSON message + Map messageMap = new ObjectMapper().readValue(message, Map.class); + + if (messageMap.containsKey("header")) { + Map header = (Map) messageMap.get("header"); + + if (header.containsKey("event")) { + String event = (String) header.get("event"); + + if ("task-started".equals(event)) { + System.out.println("收到服务端返回的task-started事件"); + + List texts = Arrays.asList( + "床前明月光,疑是地上霜", + "举头望明月,低头思故乡" + ); + + for (String text : texts) { + // 发送continue-task事件 + sendContinueTask(text); + } + + // 发送finish-task事件 + sendFinishTask(); + } else if ("task-finished".equals(event)) { + System.out.println("收到服务端返回的task-finished事件"); + taskFinished = true; + closeConnection(); + } else if ("task-failed".equals(event)) { + System.out.println("任务失败:" + message); + closeConnection(); + } + } + } + } catch (Exception e) { + System.err.println("出现异常:" + e.getMessage()); + } + } + + @Override + public void onMessage(ByteBuffer message) { + System.out.println("收到的二进制音频数据大小为:" + message.remaining()); + + try (FileOutputStream fos = new FileOutputStream(outputFile, true)) { + byte[] buffer = new byte[message.remaining()]; + message.get(buffer); + fos.write(buffer); + System.out.println("音频数据已写入本地文件" + outputFile + "中"); + } catch (IOException e) { + System.err.println("音频数据写入本地文件失败:" + e.getMessage()); + } + } + + @Override + public void onClose(int code, String reason, boolean remote) { + System.out.println("连接关闭:" + reason + " (" + code + ")"); + } + + @Override + public void onError(Exception ex) { + System.err.println("报错:" + ex.getMessage()); + ex.printStackTrace(); + } + + private void sendContinueTask(String text) { + String command = "{ \"header\": { \"action\": \"continue-task\", \"task_id\": \"" + taskId + "\", \"streaming\": \"duplex\" }, \"payload\": { \"input\": { \"text\": \"" + text + "\" } }}"; + send(command); + } + + private void sendFinishTask() { + String command = "{ \"header\": { \"action\": \"finish-task\", \"task_id\": \"" + taskId + "\", \"streaming\": \"duplex\" }, \"payload\": { \"input\": {} }}"; + send(command); + } + + private void closeConnection() { + if (!isClosed()) { + close(); + } + } + + public static void main(String[] args) { + try { + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:String apiKey = "sk-xxx" + String apiKey = System.getenv("DASHSCOPE_API_KEY"); + if (apiKey == null || apiKey.isEmpty()) { + System.err.println("请设置 DASHSCOPE_API_KEY 环境变量"); + return; + } + + Map headers = new HashMap<>(); + headers.put("Authorization", "bearer " + apiKey); + TTSWebSocketClient client = new TTSWebSocketClient(new URI("wss://maas.qianwenaiapi.com/api-ws/v1/inference"), headers); + + client.connect(); + + while (!client.isClosed() && !client.taskFinished) { + Thread.sleep(1000); + } + } catch (Exception e) { + System.err.println("连接WebSocket服务失败:" + e.getMessage()); + e.printStackTrace(); + } + } + } + ``` + + + + 建议使用 Python DashScope SDK 进行开发,请参见[Qwen-Audio-TTS Python SDK](/api-reference/speech-synthesis/qwen-audio-tts/python-sdk) / [CosyVoice Python SDK](/api-reference/speech-synthesis/cosyvoice/python-sdk)。 + + 以下是 Python WebSocket 直连示例,运行前请导入以下依赖: + + ```bash + pip uninstall websocket-client + pip uninstall websocket + pip install websocket-client + ``` + + + 请不要将运行示例代码的Python文件命名为“websocket.py”,否则会报错(AttributeError: module 'websocket' has no attribute 'WebSocketApp'. Did you mean: 'WebSocket'?)。 + + + ```python expandable + import websocket + import json + import uuid + import os + import time + + class TTSClient: + def __init__(self, api_key, uri): + """ + 初始化 TTSClient 实例 + + 参数: + api_key (str): 鉴权用的 API Key + uri (str): WebSocket 服务地址 + """ + self.api_key = api_key # 替换为你的 API Key + self.uri = uri # 替换为你的 WebSocket 地址 + self.task_id = str(uuid.uuid4()) # 生成唯一任务 ID + self.output_file = f"output_{int(time.time())}.mp3" # 输出音频文件路径 + self.ws = None # WebSocketApp 实例 + self.task_started = False # 是否收到 task-started + self.task_finished = False # 是否收到 task-finished / task-failed + + def on_open(self, ws): + """ + WebSocket 连接建立时回调函数 + 发送 run-task 事件开启语音合成任务 + """ + print("WebSocket 已连接") + + # 构造 run-task 事件 + run_task_cmd = { + "header": { + "action": "run-task", + "task_id": self.task_id, + "streaming": "duplex" + }, + "payload": { + "task_group": "audio", + "task": "tts", + "function": "SpeechSynthesizer", + "model": "qwen-audio-3.0-tts-flash", + "parameters": { + "text_type": "PlainText", + "voice": "longanhuan_v3.6", + "format": "mp3", + "sample_rate": 22050, + "volume": 50, + "rate": 1, + "pitch": 1, + # 如果enable_ssml设为True,只允许发送一次continue-task事件,否则会报错“Text request limit violated, expected 1.” + "enable_ssml": False + }, + "input": {} + } + } + + # 发送 run-task 事件 + ws.send(json.dumps(run_task_cmd)) + print("已发送 run-task 事件") + + def on_message(self, ws, message): + """ + 接收到消息时的回调函数 + 区分文本和二进制消息处理 + """ + if isinstance(message, str): + # 处理 JSON 文本消息 + try: + msg_json = json.loads(message) + print(f"收到 JSON 消息: {msg_json}") + + if "header" in msg_json: + header = msg_json["header"] + + if "event" in header: + event = header["event"] + + if event == "task-started": + print("任务已启动") + self.task_started = True + + # 发送 continue-task 事件 + texts = [ + "床前明月光,疑是地上霜", + "举头望明月,低头思故乡" + ] + + for text in texts: + self.send_continue_task(text) + + # 所有 continue-task 发送完成后发送 finish-task + self.send_finish_task() + + elif event == "task-finished": + print("任务已完成") + self.task_finished = True + self.close(ws) + + elif event == "task-failed": + error_msg = msg_json.get("error_message", "未知错误") + print(f"任务失败: {error_msg}") + self.task_finished = True + self.close(ws) + + except json.JSONDecodeError as e: + print(f"JSON 解析失败: {e}") + else: + # 处理二进制消息(音频数据) + print(f"收到二进制消息,大小: {len(message)} 字节") + with open(self.output_file, "ab") as f: + f.write(message) + print(f"已将音频数据写入本地文件{self.output_file}中") + + def on_error(self, ws, error): + """发生错误时的回调""" + print(f"WebSocket 出错: {error}") + + def on_close(self, ws, close_status_code, close_msg): + """连接关闭时的回调""" + print(f"WebSocket 已关闭: {close_msg} ({close_status_code})") + + def send_continue_task(self, text): + """发送 continue-task 事件,附带要合成的文本内容""" + cmd = { + "header": { + "action": "continue-task", + "task_id": self.task_id, + "streaming": "duplex" + }, + "payload": { + "input": { + "text": text + } + } + } + + self.ws.send(json.dumps(cmd)) + print(f"已发送 continue-task 事件,文本内容: {text}") + + def send_finish_task(self): + """发送 finish-task 事件,结束语音合成任务""" + cmd = { + "header": { + "action": "finish-task", + "task_id": self.task_id, + "streaming": "duplex" + }, + "payload": { + "input": {} + } + } + + self.ws.send(json.dumps(cmd)) + print("已发送 finish-task 事件") + + def close(self, ws): + """主动关闭连接""" + if ws and ws.sock and ws.sock.connected: + ws.close() + print("已主动关闭连接") + + def run(self): + """启动 WebSocket 客户端""" + # 设置请求头部(鉴权) + header = { + "Authorization": f"bearer {self.api_key}", + "X-DashScope-DataInspection": "enable" + } + + # 创建 WebSocketApp 实例 + self.ws = websocket.WebSocketApp( + self.uri, + header=header, + on_open=self.on_open, + on_message=self.on_message, + on_error=self.on_error, + on_close=self.on_close + ) + + print("正在监听 WebSocket 消息...") + self.ws.run_forever() # 启动长连接监听 + + # 示例使用方式 + if __name__ == "__main__": + # 若没有配置环境变量,请用千问AI平台API Key将下行替换为:API_KEY = "sk-xxx" + API_KEY = os.environ.get("DASHSCOPE_API_KEY") + SERVER_URI = "wss://maas.qianwenaiapi.com/api-ws/v1/inference" # 替换为你的 WebSocket 地址 + + client = TTSClient(API_KEY, SERVER_URI) + client.run() + ``` + + + + + + 1. **创建客户端** + + + + 在本地新建 Python 文件,命名为`tts_realtime_client.py`并复制以下代码到文件中: + + ```python expandable + # -- coding: utf-8 -- + + import asyncio + import websockets + import json + import base64 + import time + from typing import Optional, Callable, Dict, Any + from enum import Enum + + class SessionMode(Enum): + SERVER_COMMIT = "server_commit" + COMMIT = "commit" + + class TTSRealtimeClient: + """ + 与 TTS Realtime API 交互的客户端。 + + 该类提供了连接 TTS Realtime API、发送文本数据、获取音频输出以及管理 WebSocket 连接的相关方法。 + + 属性说明: + base_url (str): + Realtime API 的基础地址。 + api_key (str): + 用于身份验证的 API Key。 + voice (str): + 服务端合成语音所使用的声音。 + mode (SessionMode): + 会话模式,可选 server_commit 或 commit。 + audio_callback (Callable[[bytes], None]): + 接收音频数据的回调函数。 + language_type(str) + 合成的语音的语种,可选值Chinese、English、German、Italian、Portuguese、Spanish、Japanese、Korean、French、Russian、Auto + """ + + def __init__( + self, + base_url: str, + api_key: str, + voice: str = "Cherry", + mode: SessionMode = SessionMode.SERVER_COMMIT, + audio_callback: Optional[Callable[[bytes], None]] = None, + language_type: str = "Auto"): + self.base_url = base_url + self.api_key = api_key + self.voice = voice + self.mode = mode + self.ws = None + self.audio_callback = audio_callback + self.language_type = language_type + + # 当前回复状态 + self._current_response_id = None + self._current_item_id = None + self._is_responding = False + self._response_done_future = None + + async def connect(self) -> None: + """与 TTS Realtime API 建立 WebSocket 连接。""" + headers = { + "Authorization": f"Bearer {self.api_key}" + } + + self.ws = await websockets.connect(self.base_url, additional_headers=headers) + + # 设置默认会话配置 + await self.update_session({ + "mode": self.mode.value, + "voice": self.voice, + # 如需使用指令控制功能,请取消下方注释,并在server_commit.py或commit.py中将model替换为qwen3-tts-instruct-flash-realtime + # "instructions": "语速较快,带有明显的上扬语调,适合介绍时尚产品。", + # "optimize_instructions": true + "language_type": self.language_type, + "response_format": "pcm", + "sample_rate": 24000 + }) + + async def send_event(self, event) -> None: + """发送事件到服务器。""" + event['event_id'] = "event_" + str(int(time.time() * 1000)) + print(f"发送事件: type={event['type']}, event_id={event['event_id']}") + await self.ws.send(json.dumps(event)) + + async def update_session(self, config: Dict[str, Any]) -> None: + """更新会话配置。""" + event = { + "type": "session.update", + "session": config + } + print("更新会话配置: ", event) + await self.send_event(event) + + async def append_text(self, text: str) -> None: + """向 API 发送文本数据。""" + event = { + "type": "input_text_buffer.append", + "text": text + } + await self.send_event(event) + + async def commit_text_buffer(self) -> None: + """提交文本缓冲区以触发处理。""" + event = { + "type": "input_text_buffer.commit" + } + await self.send_event(event) + + async def clear_text_buffer(self) -> None: + """清除文本缓冲区。""" + event = { + "type": "input_text_buffer.clear" + } + await self.send_event(event) + + async def finish_session(self) -> None: + """结束会话。""" + event = { + "type": "session.finish" + } + await self.send_event(event) + + async def wait_for_response_done(self): + """等待 response.done 事件""" + if self._response_done_future: + await self._response_done_future + + async def handle_messages(self) -> None: + """处理来自服务器的消息。""" + try: + async for message in self.ws: + event = json.loads(message) + event_type = event.get("type") + + if event_type != "response.audio.delta": + print(f"收到事件: {event_type}") + + if event_type == "error": + print("错误: ", event.get('error', {})) + continue + elif event_type == "session.created": + print("会话创建,ID: ", event.get('session', {}).get('id')) + elif event_type == "session.updated": + print("会话更新,ID: ", event.get('session', {}).get('id')) + elif event_type == "input_text_buffer.committed": + print("文本缓冲区已提交,项目ID: ", event.get('item_id')) + elif event_type == "input_text_buffer.cleared": + print("文本缓冲区已清除") + elif event_type == "response.created": + self._current_response_id = event.get("response", {}).get("id") + self._is_responding = True + # 创建新的 future 来等待 response.done + self._response_done_future = asyncio.Future() + print("响应已创建,ID: ", self._current_response_id) + elif event_type == "response.output_item.added": + self._current_item_id = event.get("item", {}).get("id") + print("输出项已添加,ID: ", self._current_item_id) + # 处理音频增量 + elif event_type == "response.audio.delta" and self.audio_callback: + audio_bytes = base64.b64decode(event.get("delta", "")) + self.audio_callback(audio_bytes) + elif event_type == "response.audio.done": + print("音频生成完成") + elif event_type == "response.done": + self._is_responding = False + self._current_response_id = None + self._current_item_id = None + # 标记 future 完成 + if self._response_done_future and not self._response_done_future.done(): + self._response_done_future.set_result(True) + print("响应完成") + elif event_type == "session.finished": + print("会话已结束") + + except websockets.exceptions.ConnectionClosed: + print("连接已关闭") + except Exception as e: + print("消息处理出错: ", str(e)) + + async def close(self) -> None: + """关闭 WebSocket 连接。""" + if self.ws: + await self.ws.close() + ``` + + + + 在本地新建 Java 文件,命名为`TTSRealtimeClient.java`并复制以下代码到文件中: + + ```java expandable + import com.google.gson.Gson; + import com.google.gson.JsonObject; + import org.java_websocket.client.WebSocketClient; + import org.java_websocket.handshake.ServerHandshake; + + import java.net.URI; + import java.util.Base64; + import java.util.HashMap; + import java.util.Map; + import java.util.concurrent.CountDownLatch; + import java.util.function.Consumer; + + /** + * 与 TTS Realtime API 交互的客户端。 + * + * 该类提供了连接 TTS Realtime API、发送文本数据、获取音频输出以及管理 WebSocket 连接的相关方法。 + */ + public class TTSRealtimeClient { + + public enum SessionMode { + SERVER_COMMIT("server_commit"), + COMMIT("commit"); + private final String value; + SessionMode(String value) { this.value = value; } + public String getValue() { return value; } + } + + /** + * 音频回调接口 + */ + public interface AudioCallback { + void onAudio(byte[] audioData); + } + + private final String baseUrl; + private final String apiKey; + private final String voice; + private final SessionMode mode; + private final String languageType; + private final AudioCallback audioCallback; + private final Gson gson = new Gson(); + + private WebSocketClient ws; + private CountDownLatch responseDoneLatch; + private CountDownLatch sessionFinishedLatch; + + public TTSRealtimeClient(String baseUrl, String apiKey, String voice, + SessionMode mode, AudioCallback audioCallback, + String languageType) { + this.baseUrl = baseUrl; + this.apiKey = apiKey; + this.voice = voice; + this.mode = mode; + this.audioCallback = audioCallback; + this.languageType = languageType; + } + + public TTSRealtimeClient(String baseUrl, String apiKey, String voice, + SessionMode mode, AudioCallback audioCallback) { + this(baseUrl, apiKey, voice, mode, audioCallback, "Auto"); + } + + /** + * 与 TTS Realtime API 建立 WebSocket 连接。 + */ + public void connect() throws Exception { + Map headers = new HashMap<>(); + headers.put("Authorization", "Bearer " + apiKey); + + responseDoneLatch = new CountDownLatch(0); + sessionFinishedLatch = new CountDownLatch(1); + + ws = new WebSocketClient(new URI(baseUrl), headers) { + @Override + public void onOpen(ServerHandshake handshake) { + System.out.println("WebSocket 连接已建立"); + // 发送默认会话配置 + JsonObject session = new JsonObject(); + session.addProperty("mode", mode.getValue()); + session.addProperty("voice", TTSRealtimeClient.this.voice); + // 如需使用指令控制功能,请取消下方注释,并将model替换为qwen3-tts-instruct-flash-realtime + // session.addProperty("instructions", "语速较快,带有明显的上扬语调,适合介绍时尚产品。"); + // session.addProperty("optimize_instructions", true); + session.addProperty("language_type", languageType); + session.addProperty("response_format", "pcm"); + session.addProperty("sample_rate", 24000); + updateSession(session); + } + + @Override + public void onMessage(String message) { + JsonObject event = gson.fromJson(message, JsonObject.class); + String eventType = event.has("type") ? event.get("type").getAsString() : ""; + + if (!"response.audio.delta".equals(eventType)) { + System.out.println("收到事件: " + eventType); + } + + switch (eventType) { + case "error": + System.err.println("错误: " + event.get("error")); + break; + case "session.created": + System.out.println("会话创建,ID: " + + event.getAsJsonObject("session").get("id").getAsString()); + break; + case "session.updated": + System.out.println("会话更新,ID: " + + event.getAsJsonObject("session").get("id").getAsString()); + break; + case "input_text_buffer.committed": + System.out.println("文本缓冲区已提交,项目ID: " + event.get("item_id")); + break; + case "input_text_buffer.cleared": + System.out.println("文本缓冲区已清除"); + break; + case "response.created": + System.out.println("响应已创建,ID: " + + event.getAsJsonObject("response").get("id").getAsString()); + responseDoneLatch = new CountDownLatch(1); + break; + case "response.output_item.added": + System.out.println("输出项已添加,ID: " + + event.getAsJsonObject("item").get("id").getAsString()); + break; + case "response.audio.delta": + if (audioCallback != null) { + byte[] audioBytes = Base64.getDecoder().decode( + event.get("delta").getAsString()); + audioCallback.onAudio(audioBytes); + } + break; + case "response.audio.done": + System.out.println("音频生成完成"); + break; + case "response.done": + System.out.println("响应完成"); + responseDoneLatch.countDown(); + break; + case "session.finished": + System.out.println("会话已结束"); + sessionFinishedLatch.countDown(); + break; + } + } + + @Override + public void onClose(int code, String reason, boolean remote) { + System.out.println("连接已关闭: " + reason); + } + + @Override + public void onError(Exception ex) { + System.err.println("WebSocket 错误: " + ex.getMessage()); + } + }; + ws.connectBlocking(); + } + + /** + * 发送事件到服务器。 + */ + public void sendEvent(JsonObject event) { + String eventId = "event_" + System.currentTimeMillis(); + event.addProperty("event_id", eventId); + System.out.println("发送事件: type=" + event.get("type").getAsString() + + ", event_id=" + eventId); + ws.send(gson.toJson(event)); + } + + /** + * 更新会话配置。 + */ + public void updateSession(JsonObject config) { + JsonObject event = new JsonObject(); + event.addProperty("type", "session.update"); + event.add("session", config); + System.out.println("更新会话配置: " + event); + sendEvent(event); + } + + /** + * 向 API 发送文本数据。 + */ + public void appendText(String text) { + JsonObject event = new JsonObject(); + event.addProperty("type", "input_text_buffer.append"); + event.addProperty("text", text); + sendEvent(event); + } + + /** + * 提交文本缓冲区以触发处理。 + */ + public void commitTextBuffer() { + JsonObject event = new JsonObject(); + event.addProperty("type", "input_text_buffer.commit"); + sendEvent(event); + } + + /** + * 清除文本缓冲区。 + */ + public void clearTextBuffer() { + JsonObject event = new JsonObject(); + event.addProperty("type", "input_text_buffer.clear"); + sendEvent(event); + } + + /** + * 结束会话。 + */ + public void finishSession() { + JsonObject event = new JsonObject(); + event.addProperty("type", "session.finish"); + sendEvent(event); + } + + /** + * 等待 response.done 事件。 + */ + public void waitForResponseDone() throws InterruptedException { + responseDoneLatch.await(); + } + + /** + * 等待 session.finished 事件。 + */ + public void waitForSessionFinished() throws InterruptedException { + sessionFinishedLatch.await(); + } + + /** + * 关闭 WebSocket 连接。 + */ + public void close() { + if (ws != null) { + ws.close(); + } + } + } + ``` + + + + 2. **选择语音合成模式** + + Realtime API支持以下两种模式: + + - **server\_commit 模式** + + 服务端智能判断文本分段与合成时机,客户端只需发送文本。适合低延迟场景(如 GPS 导航)。 + - **commit 模式** + + 客户端将文本添加至缓冲区后主动触发合成。适合精细控制断句的场景(如新闻播报)。 + + + + + + 在`tts_realtime_client.py`的同级目录下新建另一个 Python 文件,命名为`server_commit.py`,并将以下代码复制进文件中: + + ```python expandable + import os + import asyncio + import logging + import wave + from tts_realtime_client import TTSRealtimeClient, SessionMode + import pyaudio + + # QwenTTS 服务配置 + # 如需使用指令控制功能,请将model替换为qwen3-tts-instruct-flash-realtime,并在tts_realtime_client.py中取消instructions的注释 + URL = "wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model=qwen3-tts-flash-realtime" + # 若没有配置环境变量,请用千问AI平台API Key将下行替换为:API_KEY="sk-xxx" + API_KEY = os.getenv("DASHSCOPE_API_KEY") + + if not API_KEY: + raise ValueError("Please set DASHSCOPE_API_KEY environment variable") + + # 收集音频数据 + _audio_chunks = [] + # 实时播放相关 + _AUDIO_SAMPLE_RATE = 24000 + _audio_pyaudio = pyaudio.PyAudio() + _audio_stream = None # 将在运行时打开 + + def _audio_callback(audio_bytes: bytes): + """TTSRealtimeClient 音频回调: 实时播放并缓存""" + global _audio_stream + if _audio_stream is not None: + try: + _audio_stream.write(audio_bytes) + except Exception as exc: + logging.error(f"PyAudio playback error: {exc}") + _audio_chunks.append(audio_bytes) + logging.info(f"Received audio chunk: {len(audio_bytes)} bytes") + + def _save_audio_to_file(filename: str = "output.wav", sample_rate: int = 24000) -> bool: + """将收集到的音频数据保存为 WAV 文件""" + if not _audio_chunks: + logging.warning("No audio data to save") + return False + + try: + audio_data = b"".join(_audio_chunks) + with wave.open(filename, 'wb') as wav_file: + wav_file.setnchannels(1) # 单声道 + wav_file.setsampwidth(2) # 16-bit + wav_file.setframerate(sample_rate) + wav_file.writeframes(audio_data) + logging.info(f"Audio saved to: {filename}") + return True + except Exception as exc: + logging.error(f"Failed to save audio: {exc}") + return False + + async def _produce_text(client: TTSRealtimeClient): + """向服务器发送文本片段""" + text_fragments = [ + "阿里云的大模型服务平台千问AI平台是一站式的大模型开发及应用构建平台。", + "不论是开发者还是业务人员,都能深入参与大模型应用的设计和构建。", + "您可以通过简单的界面操作,在5分钟内开发出一款大模型应用,", + "或在几小时内训练出一个专属模型,从而将更多精力专注于应用创新。", + ] + + logging.info("Sending text fragments…") + for text in text_fragments: + logging.info(f"Sending fragment: {text}") + await client.append_text(text) + await asyncio.sleep(0.1) # 片段间稍作延时 + + # 等待服务器完成内部处理后结束会话 + await asyncio.sleep(1.0) + await client.finish_session() + + async def _run_demo(): + """运行完整 Demo""" + global _audio_stream + # 打开 PyAudio 输出流 + _audio_stream = _audio_pyaudio.open( + format=pyaudio.paInt16, + channels=1, + rate=_AUDIO_SAMPLE_RATE, + output=True, + frames_per_buffer=1024 + ) + + client = TTSRealtimeClient( + base_url=URL, + api_key=API_KEY, + voice="Cherry", + mode=SessionMode.SERVER_COMMIT, + audio_callback=_audio_callback + ) + + # 建立连接 + await client.connect() + + # 并行执行消息处理与文本发送 + consumer_task = asyncio.create_task(client.handle_messages()) + producer_task = asyncio.create_task(_produce_text(client)) + + await producer_task # 等待文本发送完成 + + # 等待 response.done + await client.wait_for_response_done() + + # 关闭连接并取消消费者任务 + await client.close() + consumer_task.cancel() + + # 关闭音频流 + if _audio_stream is not None: + _audio_stream.stop_stream() + _audio_stream.close() + _audio_pyaudio.terminate() + + # 保存音频数据 + os.makedirs("outputs", exist_ok=True) + _save_audio_to_file(os.path.join("outputs", "qwen_tts_output.wav")) + + def main(): + """同步入口""" + logging.basicConfig( + level=logging.INFO, + format='%(asctime)s [%(levelname)s] %(message)s', + datefmt='%Y-%m-%d %H:%M:%S' + ) + logging.info("Starting QwenTTS Realtime Client demo…") + asyncio.run(_run_demo()) + + if __name__ == "__main__": + main() + ``` + + 运行`server_commit.py`,即可听到 Realtime API实时生成的音频。 + + + + 在`TTSRealtimeClient.java`的同级目录下新建另一个 Java 文件,命名为`ServerCommit.java`,并将以下代码复制进文件中: + + ```java expandable + import javax.sound.sampled.*; + import java.io.*; + import java.util.ArrayList; + import java.util.List; + import java.util.concurrent.ConcurrentLinkedQueue; + import java.util.concurrent.atomic.AtomicBoolean; + + public class ServerCommit { + private static final String URL = "wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model=qwen3-tts-flash-realtime"; + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:private static final String API_KEY = "sk-xxx"; + private static final String API_KEY = System.getenv("DASHSCOPE_API_KEY"); + private static final int SAMPLE_RATE = 24000; + + // 音频数据缓存 + private static final List audioChunks = new ArrayList<>(); + // 实时播放队列 + private static final ConcurrentLinkedQueue playbackQueue = new ConcurrentLinkedQueue<>(); + private static final AtomicBoolean playing = new AtomicBoolean(true); + + public static void main(String[] args) throws Exception { + if (API_KEY == null || API_KEY.isEmpty()) { + throw new IllegalStateException("请设置 DASHSCOPE_API_KEY 环境变量"); + } + + // 初始化音频播放 + AudioFormat format = new AudioFormat(SAMPLE_RATE, 16, 1, true, false); + DataLine.Info info = new DataLine.Info(SourceDataLine.class, format); + SourceDataLine audioLine = (SourceDataLine) AudioSystem.getLine(info); + audioLine.open(format); + audioLine.start(); + + // 启动播放线程 + Thread playerThread = new Thread(() -> { + while (playing.get() || !playbackQueue.isEmpty()) { + byte[] chunk = playbackQueue.poll(); + if (chunk != null) { + audioLine.write(chunk, 0, chunk.length); + } else { + try { Thread.sleep(10); } catch (InterruptedException ignored) {} + } + } + }); + playerThread.start(); + + // 创建 TTS 客户端 + // 如需使用指令控制功能,请将model替换为qwen3-tts-instruct-flash-realtime,并在TTSRealtimeClient.java中取消instructions的注释 + TTSRealtimeClient client = new TTSRealtimeClient( + URL, API_KEY, "Cherry", + TTSRealtimeClient.SessionMode.SERVER_COMMIT, + audioData -> { + playbackQueue.add(audioData); + audioChunks.add(audioData); + System.out.println("收到音频数据: " + audioData.length + " bytes"); + } + ); + + client.connect(); + + // 发送文本片段 + String[] textFragments = { + "阿里云的大模型服务平台千问AI平台是一站式的大模型开发及应用构建平台。", + "不论是开发者还是业务人员,都能深入参与大模型应用的设计和构建。", + "您可以通过简单的界面操作,在5分钟内开发出一款大模型应用,", + "或在几小时内训练出一个专属模型,从而将更多精力专注于应用创新。" + }; + + System.out.println("开始发送文本..."); + for (String text : textFragments) { + System.out.println("发送片段: " + text); + client.appendText(text); + Thread.sleep(100); + } + + Thread.sleep(1000); + client.finishSession(); + + // 等待响应完成 + client.waitForResponseDone(); + client.waitForSessionFinished(); + client.close(); + + // 等待播放完成 + playing.set(false); + playerThread.join(); + audioLine.drain(); + audioLine.close(); + + // 保存音频文件 + saveWav("output.wav"); + System.out.println("完成"); + } + + private static void saveWav(String filename) throws IOException { + if (audioChunks.isEmpty()) { + System.out.println("没有音频数据可保存"); + return; + } + ByteArrayOutputStream bos = new ByteArrayOutputStream(); + for (byte[] chunk : audioChunks) { + bos.write(chunk); + } + byte[] allAudio = bos.toByteArray(); + AudioFormat format = new AudioFormat(SAMPLE_RATE, 16, 1, true, false); + AudioInputStream ais = new AudioInputStream( + new ByteArrayInputStream(allAudio), format, allAudio.length / 2); + new File("outputs").mkdirs(); + AudioSystem.write(ais, AudioFileFormat.Type.WAVE, + new File("outputs/" + filename)); + System.out.println("音频已保存到: outputs/" + filename); + } + } + ``` + + 编译并运行`ServerCommit.java`,即可听到 Realtime API实时生成的音频。 + + + + + + + + 在`tts_realtime_client.py`的同级目录下新建另一个 Python 文件,命名为`commit.py`,并将以下代码复制进文件中: + + ```python expandable + import os + import asyncio + import logging + import wave + from tts_realtime_client import TTSRealtimeClient, SessionMode + import pyaudio + + # QwenTTS 服务配置 + # 如需使用指令控制功能,请将model替换为qwen3-tts-instruct-flash-realtime,并在tts_realtime_client.py中取消instructions的注释 + URL = "wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model=qwen3-tts-flash-realtime" + # 若没有配置环境变量,请用千问AI平台API Key将下行替换为:API_KEY="sk-xxx" + API_KEY = os.getenv("DASHSCOPE_API_KEY") + + if not API_KEY: + raise ValueError("Please set DASHSCOPE_API_KEY environment variable") + + # 收集音频数据 + _audio_chunks = [] + _AUDIO_SAMPLE_RATE = 24000 + _audio_pyaudio = pyaudio.PyAudio() + _audio_stream = None + + def _audio_callback(audio_bytes: bytes): + """TTSRealtimeClient 音频回调: 实时播放并缓存""" + global _audio_stream + if _audio_stream is not None: + try: + _audio_stream.write(audio_bytes) + except Exception as exc: + logging.error(f"PyAudio playback error: {exc}") + _audio_chunks.append(audio_bytes) + logging.info(f"Received audio chunk: {len(audio_bytes)} bytes") + + def _save_audio_to_file(filename: str = "output.wav", sample_rate: int = 24000) -> bool: + """将收集到的音频数据保存为 WAV 文件""" + if not _audio_chunks: + logging.warning("No audio data to save") + return False + + try: + audio_data = b"".join(_audio_chunks) + with wave.open(filename, 'wb') as wav_file: + wav_file.setnchannels(1) # 单声道 + wav_file.setsampwidth(2) # 16-bit + wav_file.setframerate(sample_rate) + wav_file.writeframes(audio_data) + logging.info(f"Audio saved to: {filename}") + return True + except Exception as exc: + logging.error(f"Failed to save audio: {exc}") + return False + + async def _user_input_loop(client: TTSRealtimeClient): + """持续获取用户输入并发送文本,当用户输入空文本时发送commit事件并结束本次会话""" + print("请输入文本(直接按Enter发送commit事件并结束本次会话,按Ctrl+C或Ctrl+D结束整个程序):") + + while True: + try: + user_text = input("> ") + if not user_text: # 用户输入为空 + # 空输入视为一次对话的结束: 提交缓冲区 -> 结束会话 -> 跳出循环 + logging.info("空输入,发送 commit 事件并结束本次会话") + await client.commit_text_buffer() + # 适当等待服务器处理 commit,防止过早结束会话导致丢失音频 + await asyncio.sleep(0.3) + await client.finish_session() + break # 直接退出用户输入循环,无需再次回车 + else: + logging.info(f"发送文本: {user_text}") + await client.append_text(user_text) + + except EOFError: # 用户按下Ctrl+D + break + except KeyboardInterrupt: # 用户按下Ctrl+C + break + + # 结束会话 + logging.info("结束会话...") + async def _run_demo(): + """运行完整 Demo""" + global _audio_stream + # 打开 PyAudio 输出流 + _audio_stream = _audio_pyaudio.open( + format=pyaudio.paInt16, + channels=1, + rate=_AUDIO_SAMPLE_RATE, + output=True, + frames_per_buffer=1024 + ) + + client = TTSRealtimeClient( + base_url=URL, + api_key=API_KEY, + voice="Cherry", + mode=SessionMode.COMMIT, # 修改为COMMIT模式 + audio_callback=_audio_callback + ) + + # 建立连接 + await client.connect() + + # 并行执行消息处理与用户输入 + consumer_task = asyncio.create_task(client.handle_messages()) + producer_task = asyncio.create_task(_user_input_loop(client)) + + await producer_task # 等待用户输入完成 + + # 等待 response.done + await client.wait_for_response_done() + + # 关闭连接并取消消费者任务 + await client.close() + consumer_task.cancel() + + # 关闭音频流 + if _audio_stream is not None: + _audio_stream.stop_stream() + _audio_stream.close() + _audio_pyaudio.terminate() + + # 保存音频数据 + os.makedirs("outputs", exist_ok=True) + _save_audio_to_file(os.path.join("outputs", "qwen_tts_output.wav")) + + def main(): + logging.basicConfig( + level=logging.INFO, + format='%(asctime)s [%(levelname)s] %(message)s', + datefmt='%Y-%m-%d %H:%M:%S' + ) + logging.info("Starting QwenTTS Realtime Client demo…") + asyncio.run(_run_demo()) + + if __name__ == "__main__": + main() + ``` + + 运行`commit.py`,可多次输入要合成的文本。在未输入文本的情况下单击 Enter 键,将从扬声器听到 Realtime API返回的音频。 + + + + 在`TTSRealtimeClient.java`的同级目录下新建另一个 Java 文件,命名为`Commit.java`,并将以下代码复制进文件中: + + ```java expandable + import javax.sound.sampled.*; + import java.io.*; + import java.util.ArrayList; + import java.util.List; + import java.util.Scanner; + import java.util.concurrent.ConcurrentLinkedQueue; + import java.util.concurrent.atomic.AtomicBoolean; + + public class Commit { + private static final String URL = "wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model=qwen3-tts-flash-realtime"; + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:private static final String API_KEY = "sk-xxx"; + private static final String API_KEY = System.getenv("DASHSCOPE_API_KEY"); + private static final int SAMPLE_RATE = 24000; + + private static final List audioChunks = new ArrayList<>(); + private static final ConcurrentLinkedQueue playbackQueue = new ConcurrentLinkedQueue<>(); + private static final AtomicBoolean playing = new AtomicBoolean(true); + + public static void main(String[] args) throws Exception { + if (API_KEY == null || API_KEY.isEmpty()) { + throw new IllegalStateException("请设置 DASHSCOPE_API_KEY 环境变量"); + } + + // 初始化音频播放 + AudioFormat format = new AudioFormat(SAMPLE_RATE, 16, 1, true, false); + DataLine.Info info = new DataLine.Info(SourceDataLine.class, format); + SourceDataLine audioLine = (SourceDataLine) AudioSystem.getLine(info); + audioLine.open(format); + audioLine.start(); + + // 启动播放线程 + Thread playerThread = new Thread(() -> { + while (playing.get() || !playbackQueue.isEmpty()) { + byte[] chunk = playbackQueue.poll(); + if (chunk != null) { + audioLine.write(chunk, 0, chunk.length); + } else { + try { Thread.sleep(10); } catch (InterruptedException ignored) {} + } + } + }); + playerThread.start(); + + // 创建 TTS 客户端(commit 模式) + // 如需使用指令控制功能,请将model替换为qwen3-tts-instruct-flash-realtime,并在TTSRealtimeClient.java中取消instructions的注释 + TTSRealtimeClient client = new TTSRealtimeClient( + URL, API_KEY, "Cherry", + TTSRealtimeClient.SessionMode.COMMIT, + audioData -> { + playbackQueue.add(audioData); + audioChunks.add(audioData); + System.out.println("收到音频数据: " + audioData.length + " bytes"); + } + ); + + client.connect(); + + // 交互式输入 + System.out.println("请输入文本(直接按Enter发送commit事件并结束本次会话,按Ctrl+D结束程序):"); + Scanner scanner = new Scanner(System.in); + while (true) { + System.out.print("> "); + if (!scanner.hasNextLine()) { + client.finishSession(); + break; + } + String userText = scanner.nextLine(); + if (userText.isEmpty()) { + // 空输入:提交缓冲区并结束会话 + System.out.println("空输入,发送 commit 事件并结束本次会话"); + client.commitTextBuffer(); + Thread.sleep(300); + client.finishSession(); + break; + } else { + System.out.println("发送文本: " + userText); + client.appendText(userText); + } + } + scanner.close(); + + // 等待响应完成 + client.waitForResponseDone(); + client.waitForSessionFinished(); + client.close(); + + // 等待播放完成 + playing.set(false); + playerThread.join(); + audioLine.drain(); + audioLine.close(); + + // 保存音频文件 + saveWav("output.wav"); + System.out.println("完成"); + } + + private static void saveWav(String filename) throws IOException { + if (audioChunks.isEmpty()) { + System.out.println("没有音频数据可保存"); + return; + } + ByteArrayOutputStream bos = new ByteArrayOutputStream(); + for (byte[] chunk : audioChunks) { + bos.write(chunk); + } + byte[] allAudio = bos.toByteArray(); + AudioFormat format = new AudioFormat(SAMPLE_RATE, 16, 1, true, false); + AudioInputStream ais = new AudioInputStream( + new ByteArrayInputStream(allAudio), format, allAudio.length / 2); + new File("outputs").mkdirs(); + AudioSystem.write(ais, AudioFileFormat.Type.WAVE, + new File("outputs/" + filename)); + System.out.println("音频已保存到: outputs/" + filename); + } + } + ``` + + 编译并运行`Commit.java`,可多次输入要合成的文本。在未输入文本的情况下单击 Enter 键,将从扬声器听到 Realtime API返回的音频。 + + + + + + + + + + ```go expandable + package main + + import ( + "encoding/json" + "fmt" + "net/http" + "os" + "time" + + "github.com/google/uuid" + "github.com/gorilla/websocket" + ) + + const ( + wsURL = "wss://maas.qianwenaiapi.com/api-ws/v1/inference" // WebSocket服务器地址 + outputFile = "output.mp3" // 输出文件路径 + ) + + func main() { + // 若没有将API Key配置到环境变量,可将下行替换为:apiKey := "your_api_key"。不建议在生产环境中直接将API Key硬编码到代码中,以减少API Key泄露风险。 + apiKey := os.Getenv("DASHSCOPE_API_KEY") + + // 检查并清空输出文件 + if err := clearOutputFile(outputFile); err != nil { + fmt.Println("清空输出文件失败:", err) + return + } + + // 连接WebSocket服务 + conn, err := connectWebSocket(apiKey) + if err != nil { + fmt.Println("连接WebSocket失败:", err) + return + } + defer closeConnection(conn) + + // 创建一个通道用于接收任务完成的通知 + done := make(chan struct{}) + + // 启动异步接收消息的goroutine + go receiveMessage(conn, done) + + // 发送run-task指令 + if err := sendRunTaskMsg(conn); err != nil { + fmt.Println("发送run-task指令失败:", err) + return + } + + // 等待任务完成或超时 + select { + case <-done: + fmt.Println("任务结束") + case <-time.After(5 * time.Minute): + fmt.Println("任务超时") + } + } + + // 定义消息结构体 + type Message struct { + Header Header `json:"header"` + Payload Payload `json:"payload"` + } + + // 定义头部信息 + type Header struct { + Action string `json:"action,omitempty"` + TaskID string `json:"task_id"` + Streaming string `json:"streaming,omitempty"` + Event string `json:"event,omitempty"` + ErrorCode string `json:"error_code,omitempty"` + ErrorMessage string `json:"error_message,omitempty"` + Attributes map[string]interface{} `json:"attributes"` + } + + // 定义负载信息 + type Payload struct { + Model string `json:"model,omitempty"` + TaskGroup string `json:"task_group,omitempty"` + Task string `json:"task,omitempty"` + Function string `json:"function,omitempty"` + Input Input `json:"input,omitempty"` + Parameters Parameters `json:"parameters,omitempty"` + Output Output `json:"output,omitempty"` + Usage Usage `json:"usage,omitempty"` + } + + // 定义输入信息 + type Input struct { + Text string `json:"text"` + } + + // 定义参数信息 + type Parameters struct { + TextType string `json:"text_type"` + Format string `json:"format"` + SampleRate int `json:"sample_rate"` + Volume int `json:"volume"` + Rate float64 `json:"rate"` + Pitch float64 `json:"pitch"` + WordTimestampEnabled bool `json:"word_timestamp_enabled"` + PhonemeTimestampEnabled bool `json:"phoneme_timestamp_enabled"` + } + + // 定义输出信息 + type Output struct { + Sentence Sentence `json:"sentence"` + } + + // 定义句子信息 + type Sentence struct { + BeginTime int `json:"begin_time"` + EndTime int `json:"end_time"` + Words []Word `json:"words"` + } + + // 定义单词信息 + type Word struct { + Text string `json:"text"` + BeginTime int `json:"begin_time"` + EndTime int `json:"end_time"` + Phonemes []Phoneme `json:"phonemes"` + } + + // 定义音素信息 + type Phoneme struct { + BeginTime int `json:"begin_time"` + EndTime int `json:"end_time"` + Text string `json:"text"` + Tone int `json:"tone"` + } + + // 定义使用信息 + type Usage struct { + Characters int `json:"characters"` + } + + func receiveMessage(conn *websocket.Conn, done chan struct{}) { + for { + msgType, message, err := conn.ReadMessage() + if err != nil { + fmt.Println("解析服务器消息失败:", err) + close(done) + break + } + + if msgType == websocket.BinaryMessage { + // 处理二进制音频流 + if err := writeBinaryDataToFile(message, outputFile); err != nil { + fmt.Println("写入二进制数据失败:", err) + close(done) + break + } + fmt.Println("音频片段已写入本地文件") + } else { + // 处理文本消息 + var msg Message + if err := json.Unmarshal(message, &msg); err != nil { + fmt.Println("解析事件失败:", err) + continue + } + if handleMessage(conn, msg, done) { + break + } + } + } + } + + func handleMessage(conn *websocket.Conn, msg Message, done chan struct{}) bool { + switch msg.Header.Event { + case "task-started": + fmt.Println("任务已启动") + + case "result-generated": + // 如需获取附加消息,可在此处添加相应代码 + + case "task-finished": + fmt.Println("任务已完成") + close(done) + return true + + case "task-failed": + if msg.Header.ErrorMessage != "" { + fmt.Printf("任务失败:%s\n", msg.Header.ErrorMessage) + } else { + fmt.Println("未知原因导致任务失败") + } + close(done) + return true + + default: + fmt.Printf("预料之外的事件:%v\n", msg) + close(done) + } + + return false + } + + func sendRunTaskMsg(conn *websocket.Conn) error { + runTaskMsg, err := generateRunTaskMsg() + if err != nil { + return err + } + if err := conn.WriteMessage(websocket.TextMessage, []byte(runTaskMsg)); err != nil { + return err + } + return nil + } + + func generateRunTaskMsg() (string, error) { + runTaskMessage := Message{ + Header: Header{ + Action: "run-task", + TaskID: uuid.New().String(), + Streaming: "out", + }, + Payload: Payload{ + Model: "sambert-zhichu-v1", + TaskGroup: "audio", + Task: "tts", + Function: "SpeechSynthesizer", + Input: Input{ + Text: "白日依山尽,黄河入海流。欲穷千里目,更上一层楼。", + }, + Parameters: Parameters{ + TextType: "PlainText", + Format: "mp3", + SampleRate: 16000, + Volume: 50, + Rate: 1.0, + Pitch: 1.0, + WordTimestampEnabled: true, + PhonemeTimestampEnabled: true, + }, + }, + } + + runTaskMsgJSON, err := json.Marshal(runTaskMessage) + return string(runTaskMsgJSON), err + } + + func connectWebSocket(apiKey string) (*websocket.Conn, error) { + header := make(http.Header) + header.Add("X-DashScope-DataInspection", "enable") + header.Add("Authorization", fmt.Sprintf("bearer %s", apiKey)) + conn, _, err := websocket.DefaultDialer.Dial(wsURL, header) + if err != nil { + fmt.Println("连接WebSocket失败:", err) + return nil, err + } + return conn, nil + } + + func writeBinaryDataToFile(data []byte, filePath string) error { + file, err := os.OpenFile(filePath, os.O_APPEND|os.O_CREATE|os.O_WRONLY, 0644) + if err != nil { + return err + } + defer file.Close() + _, err = file.Write(data) + return err + } + + func closeConnection(conn *websocket.Conn) { + if conn != nil { + conn.Close() + } + } + + func clearOutputFile(filePath string) error { + file, err := os.OpenFile(filePath, os.O_TRUNC|os.O_CREATE|os.O_WRONLY, 0644) + if err != nil { + return err + } + file.Close() + return nil + } + ``` + + + + 示例代码如下: + + ```csharp expandable + using System.Net.WebSockets; + using System.Text; + using System.Text.Json; + + class Program { + // 若没有将API Key配置到环境变量,可将下行替换为:private const string ApiKey="your_api_key"。不建议在生产环境中直接将API Key硬编码到代码中,以减少API Key泄露风险。 + private static readonly string ApiKey = Environment.GetEnvironmentVariable("DASHSCOPE_API_KEY") ?? throw new InvalidOperationException("DASHSCOPE_API_KEY environment variable is not set."); + + private const string WebSocketUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference"; // WebSocket服务器地址 + private const string OutputFilePath = "output.mp3"; // 输出文件路径 + + static async Task Main(string[] args) { + var ws = new ClientWebSocket(); + try { + // 1. 连接WebSocket服务,鉴权 + await ConnectWithAuth(ws, WebSocketUrl); + + // 2. 启动接收消息的线程 + var receiveTask = ReceiveMessages(ws); + + // 3. 发送run-task指令 + string textToSynthesize = "白日依山尽,黄河入海流。欲穷千里目,更上一层楼。"; + string taskId = GenerateTaskId(); + await SendRunTaskCommand(ws, textToSynthesize, taskId); + + // 4. 等待接收任务完成 + await receiveTask; + } catch (Exception ex) { + Console.WriteLine($"错误:{ex.Message}"); + } finally { + if (ws.State == WebSocketState.Open) { + await ws.CloseAsync(WebSocketCloseStatus.NormalClosure, "关闭连接", CancellationToken.None); + } + } + } + + private static async Task ConnectWithAuth(ClientWebSocket ws, string url) { + var uri = new Uri(url); + ws.Options.SetRequestHeader("Authorization", $"bearer {ApiKey}"); + ws.Options.SetRequestHeader("X-DashScope-DataInspection", "enable"); + await ws.ConnectAsync(uri, CancellationToken.None); + Console.WriteLine("已连接到WebSocket服务器。"); + } + + private static string GenerateTaskId() { + return Guid.NewGuid().ToString("N"); + } + + private static async Task SendRunTaskCommand(ClientWebSocket ws, string text, string taskId) { + var command = CreateRunTaskCommand(text, taskId); + var buffer = Encoding.UTF8.GetBytes(command); + await ws.SendAsync(new ArraySegment(buffer), WebSocketMessageType.Text, true, CancellationToken.None); + Console.WriteLine("已发送run-task指令。"); + } + + private static string CreateRunTaskCommand(string text, string taskId) { + var command = new { + header = new { + action = "run-task", + task_id = taskId, + streaming = "out" + }, + payload = new { + model = "sambert-zhichu-v1", + task_group = "audio", + task = "tts", + function = "SpeechSynthesizer", + input = new { + text = text + }, + parameters = new { + text_type = "PlainText", + format = "mp3", + sample_rate = 16000, + volume = 50, + rate = 1, + pitch = 1, + word_timestamp_enabled = true, + phoneme_timestamp_enabled = true + } + } + }; + return JsonSerializer.Serialize(command); + } + + private static async Task ReceiveMessages(ClientWebSocket ws) { + var buffer = new byte[1024 * 4]; + var fs = new FileStream(OutputFilePath, FileMode.Create, FileAccess.Write); + bool taskStarted = false; + bool taskFinished = false; + + while (ws.State == WebSocketState.Open && !taskFinished) { + var result = await ws.ReceiveAsync(new ArraySegment(buffer), CancellationToken.None); + + switch (result.MessageType) { + case WebSocketMessageType.Text: + var message = Encoding.UTF8.GetString(buffer, 0, result.Count); + var jsonMessage = JsonSerializer.Deserialize(message); + + ProcessTextMessage(jsonMessage, ref taskStarted, ref taskFinished); + break; + case WebSocketMessageType.Binary: + if (taskStarted) { + await fs.WriteAsync(buffer, 0, result.Count); + Console.WriteLine("收到音频数据。"); + } + break; + case WebSocketMessageType.Close: + Console.WriteLine("服务器关闭了连接。"); + taskFinished = true; + break; + } + } + fs.Close(); + } + + private static void ProcessTextMessage(JsonElement jsonMessage, ref bool taskStarted, ref bool taskFinished) { + if (jsonMessage.TryGetProperty("header", out JsonElement header) && header.TryGetProperty("event", out JsonElement eventToken)) { + var eventType = eventToken.GetString(); + switch (eventType) { + case "task-started": + taskStarted = true; + Console.WriteLine("任务开始。"); + break; + case "result-generated": + // 如需获取附加消息,可在此处添加相应代码 + break; + case "task-finished": + taskFinished = true; + Console.WriteLine("任务完成。"); + break; + case "task-failed": + taskFinished = true; + Console.WriteLine("任务失败。"); + break; + } + } + } + } + ``` + + + + 示例代码目录结构为: + + my-php-project/ + + ├── composer.json + + ├── vendor/ + + └── index.php + + composer.json内容如下,相关依赖的版本号请根据实际情况自行决定: + + ```json + { + "require": { + "react/event-loop": "^1.3", + "react/socket": "^1.11", + "react/stream": "^1.2", + "react/http": "^1.1", + "ratchet/pawl": "^0.4" + }, + "autoload": { + "psr-4": { + "App\\": "src/" + } + } + } + ``` + + index.php内容如下: + + ```php expandable + [ + 'bindto' => '0.0.0.0:0', + ], + // 警告:关闭 TLS 证书校验存在中间人攻击风险,仅限本地调试;生产环境务必将 verify_peer/verify_peer_name 设为 true + 'tls' => [ + 'verify_peer' => false, + 'verify_peer_name' => false, + ], + ]); + + $connector = new Connector($loop, $socketConnector); + + $headers = [ + 'Authorization' => 'bearer ' . $api_key, + 'X-DashScope-DataInspection' => 'enable' + ]; + + // 连接WebSocket服务 + $connector($websocket_url, [], $headers) + ->then(function ($conn) use ($output_file) { + echo "连接成功\n"; + + // 异步接收WebSocket消息 + $conn->on('message', function ($msg) use ($conn, $output_file) { + if ($msg->isBinary()) { + // 写入二进制数据到本地文件 + file_put_contents($output_file, $msg->getPayload(), FILE_APPEND); + echo "二进制数据写入文件\n"; + } else { + $data = json_decode($msg, true); + switch ($data['header']['event']) { + case 'task-started': + echo "任务开始\n"; + break; + case 'result-generated': + // 如需获取附加消息,可在此处添加相应代码 + break; + case 'task-finished': + echo "任务完成\n"; + $conn->close(); + break; + case 'task-failed': + echo "任务失败:" . $data['header']['error_message'] . "\n"; + $conn->close(); + break; + default: + echo "未知事件:" . $msg . "\n"; + } + } + }); + + // 监听连接关闭 + $conn->on('close', function($code = null, $reason = null) { + echo "连接已关闭\n"; + if ($code !== null) { + echo "关闭代码:" . $code . "\n"; + } + if ($reason !== null) { + echo "关闭原因:" . $reason . "\n"; + } + }); + + // 发送run-task指令 + $conn->send(json_encode([ + 'header' => [ + 'action' => 'run-task', + 'task_id' => bin2hex(random_bytes(16)), + 'streaming' => 'out' + ], + 'payload' => [ + 'model' => 'sambert-zhichu-v1', + 'task_group' => 'audio', + 'task' => 'tts', + 'function' => 'SpeechSynthesizer', + 'input' => [ + 'text' => '床前明月光,疑是地上霜。举头望明月,低头思故乡。' + ], + 'parameters' => [ + 'text_type' => 'PlainText', + 'format' => 'mp3', + 'sample_rate' => 16000, + 'volume' => 50, + 'rate' => 1, + 'pitch' => 1, + 'word_timestamp_enabled' => true, + 'phoneme_timestamp_enabled' => true + ] + ] + ])); + echo "run-task指令已发送\n"; + }, function (Exception $e) { + echo "连接失败:{$e->getMessage()}\n"; + file_put_contents('error.log', $e->getMessage() . "\n", FILE_APPEND); + }); + + $loop->run(); + ``` + + + + 需安装相关依赖: + + ```bash + npm install ws + npm install uuid + ``` + + 示例代码如下: + + ```javascript expandable + const WebSocket = require('ws'); + const fs = require('fs'); + const { v4: uuidv4 } = require('uuid'); + + // 若没有将API Key配置到环境变量,可将下行替换为:apiKey = 'your_api_key'。不建议在生产环境中直接将API Key硬编码到代码中,以减少API Key泄露风险。 + const apiKey = process.env.DASHSCOPE_API_KEY; + const wsUrl = 'wss://maas.qianwenaiapi.com/api-ws/v1/inference'; // WebSocket服务器地址 + const outputFilePath = 'output.mp3'; // 替换为您的音频文件路径 + + async function main() { + await checkAndClearOutputFile(outputFilePath); + createWebSocketConnection(); + } + + const fileStream = fs.createWriteStream(outputFilePath, { flags: 'a' }); + function createWebSocketConnection() { + const ws = new WebSocket(wsUrl, { + headers: { + Authorization: `bearer ${apiKey}`, + 'X-DashScope-DataInspection': 'enable' + } + }); + + ws.on('open', () => { + console.log('已连接到WebSocket服务器'); + sendRunTaskMessage(ws); + }); + + ws.on('message', (data, isBinary) => handleWebSocketMessage(data, isBinary, ws)); + ws.on('error', (error) => console.error('WebSocket错误:', error)); + ws.on('close', () => console.log('WebSocket连接已关闭')); + + return ws; + } + + function sendRunTaskMessage(ws) { + const taskId = uuidv4(); + const runTaskMessage = { + header: { + action: 'run-task', + task_id: taskId, + streaming: 'out' + }, + payload: { + model: 'sambert-zhichu-v1', + task_group: 'audio', + task: 'tts', + function: 'SpeechSynthesizer', + input: { + text: '白日依山尽,黄河入海流。欲穷千里目,更上一层楼。' + }, + parameters: { + text_type: 'PlainText', + format: 'mp3', + sample_rate: 16000, + volume: 50, + rate: 1, + pitch: 1, + word_timestamp_enabled: true, + phoneme_timestamp_enabled: true + } + } + }; + ws.send(JSON.stringify(runTaskMessage)); + console.log('run-task指令已发送'); + } + + function handleWebSocketMessage(data, isBinary, ws) { + if (isBinary) { + fileStream.write(data); + } else { + const message = JSON.parse(data); + handleWebSocketEvent(message, ws); + } + } + + function handleWebSocketEvent(message, ws) { + switch (message.header.event) { + case 'task-started': + console.log('任务已启动'); + break; + case 'result-generated': + console.log('结果已生成'); + break; + case 'task-finished': + console.log('任务已完成'); + ws.close(); + fileStream.end(() => { + console.log('文件流已关闭'); + }); + break; + case 'task-failed': + console.error('任务失败:', message.header.error_message); + ws.close(); + fileStream.end(() => { + console.log('文件流已关闭'); + }); + break; + default: + console.log('未知事件:', message.header.event); + } + } + + function checkAndClearOutputFile(filePath) { + return new Promise((resolve, reject) => { + fs.access(filePath, fs.F_OK, (err) => { + if (!err) { + fs.truncate(filePath, 0, (truncateErr) => { + if (truncateErr) return reject(truncateErr); + console.log('文件已清空'); + resolve(); + }); + } else { + fs.open(filePath, 'w', (openErr) => { + if (openErr) return reject(openErr); + console.log('文件已创建'); + resolve(); + }); + } + }); + }); + } + + main().catch(console.error); + ``` + + + + + + +## 应用于生产环境 + +### 连接复用(WebSocket) + +WebSocket 连接支持复用:一个合成任务结束后,无需重新建立连接即可开启下一个任务。 + +**复用流程**: + +- **Qwen-Audio-TTS** / **CosyVoice/ Sambert**:客户端发送 `finish-task`,服务端返回 `task-finished` 后,可重新发送 `run-task` 开启新任务。 +- **Qwen-TTS**:客户端发送 `session.finish`,服务端返回 `session.finished` 后,可建立新会话开启下一个任务。 + +**取消任务后复用**:对于 Qwen-Audio-TTS / CosyVoice,如果使用 `cancel` 指令取消当前任务,服务端返回 `task-finished` 后,同样可以在当前连接上重新发送 `run-task` 开启新任务。详情请参见[取消任务](/developer-guides/speech/realtime-streaming#取消任务)。 + + + 1. 必须等服务端返回结束事件(`task-finished` 或 `session.finished`)后才可发起新任务。 + 2. Qwen-Audio-TTS、CosyVoice 和 Sambert 在复用连接中的不同任务需要使用不同的 `task_id`。 + 3. 任务失败时服务端返回错误事件并关闭连接,该连接不可复用。 + 4. 任务结束后 60 秒无新任务,连接自动断开。 + + +各模型事件说明请参见对应的[API参考](/developer-guides/speech/realtime-streaming#api参考)。 + +### 模型限流 + +模型调用受限流规则约束,超出限制时服务端返回 `Requests rate limit exceeded, please try again later.` 报错,需降低调用频率或并发数后重试。 + +各模型的限流条件请参见[限流](/developer-guides/administration/rate-limits)。 + +### 高并发最佳实践 + +DashScope SDK 内置池化机制,可复用 WebSocket 连接和合成对象,避免频繁创建销毁带来的开销。 + + + + + Qwen-Audio-TTS 和 CosyVoice 使用相同的 SDK 接口,以下示例同样适用于 Qwen-Audio-TTS 系列模型,只需替换 `model` 和 `voice` 参数。 + + #### 前提条件 + + - [获取与配置 API Key](/api-reference/preparation/api-key) + - 已安装符合版本要求的DashScope SDK,建议[安装最新版](/api-reference/preparation/install-sdk): + + - Python SDK:版本≥1.25.2 + - Java SDK:版本≥2.16.6 + + + + Python SDK 通过 `SpeechSynthesizerObjectPool` 管理和复用 `SpeechSynthesizer` 对象。 + + 对象池在初始化时即创建指定数量的 `SpeechSynthesizer` 实例并建立 WebSocket 连接,获取对象时可直接发起请求,降低首包延迟。归还后连接保持活跃,等待下次复用。 + + #### 实现步骤 + + 1. 安装依赖:安装DashScope依赖(`pip install -U dashscope`) + 2. 创建并配置对象池 + + 对象池大小推荐设为峰值并发数的 1.5\~2 倍,且不应超过账户的 QPS 限制。 + + 创建全局单例对象池(初始化时建立连接,有一定耗时): + + ```python + from dashscope.audio.tts_v2 import SpeechSynthesizerObjectPool + + connectionPool = SpeechSynthesizerObjectPool(max_size=20) + import dashscope + dashscope.base_http_api_url = "https://maas.qianwenaiapi.com/api/v1" + ``` + + + - 在对象池场景中,`SpeechSynthesizerObjectPool`在初始化时即按当前全局`dashscope.api_key`与服务端建立 WebSocket 连接。apiKey 仅在 WebSocket 建连握手时写入`Authorization`请求头用于鉴权,后续任务消息(如`run-task`)本身不携带 apiKey。**池创建后修改**`dashscope.api_key`**不会影响池内已建连接**——`borrow_synthesizer`取出的对象(包括归还后再次复用的对象)仍使用握手时的 apiKey,新值会被静默忽略,可能导致身份、配额或计费归属与预期不一致。注意:`borrow_synthesizer`也不支持通过参数指定 apiKey。 + - 如确需使用多个不同的 API Key,请为每个 API Key 维护**独立的**`SpeechSynthesizerObjectPool`**实例**。 + + + 3. 从对象池中获取`SpeechSynthesizer`对象 + + 如果当前未归还的对象数已超过池容量,系统会额外创建新对象。 + + 此类对象需重新建立连接,不具备复用效果。 + + ```python + speech_synthesizer = connectionPool.borrow_synthesizer( + model='cosyvoice-v3-flash', + voice='longanyang', + seed=12382, + callback=synthesizer_callback + ) + ``` + + 4. 进行语音合成 + + 调用`SpeechSynthesizer`对象的call或streaming\_call方法进行语音合成。 + 5. 归还`SpeechSynthesizer`对象 + + 任务结束后归还对象以供复用。 + + 不要归还未完成任务或任务失败的对象。 + + ```python + connectionPool.return_synthesizer(speech_synthesizer) + ``` + + ##### 完整代码 + + + 复制使用前请注意:`SpeechSynthesizerObjectPool`在初始化时即按当前全局`dashscope.api_key`与服务端建立 WebSocket 连接并完成鉴权;**池创建后再修改**`dashscope.api_key`**不会影响池内已建连接**,新值会被静默忽略。多 API Key 场景请为每个 API Key 维护独立的池实例。详见上文重要说明。 + + + ```python expandable + # !/usr/bin/env python3 + # Copyright (C) Alibaba Group. All Rights Reserved. + # MIT License (https://opensource.org/licenses/MIT) + + import os + import time + import threading + + import dashscope + from dashscope.audio.tts_v2 import * + + USE_CONNECTION_POOL = True + text_to_synthesize = [ + '第一句、欢迎使用阿里巴巴语音合成服务。', + '第二句、欢迎使用阿里巴巴语音合成服务。', + '第三句、欢迎使用阿里巴巴语音合成服务。', + ] + connectionPool = None + + def init_dashscope_api_key(): + ''' + Set your DashScope API-key. More information: + https://github.com/aliyun/alibabacloud-bailian-speech-demo/blob/master/PREREQUISITES.md + ''' + if 'DASHSCOPE_API_KEY' in os.environ: + dashscope.api_key = os.environ[ + 'DASHSCOPE_API_KEY'] # load API-key from environment variable DASHSCOPE_API_KEY + else: + dashscope.api_key = '' # set API-key manually + + def synthesis_text_to_speech_and_play_by_streaming_mode(text, task_id): + global USE_CONNECTION_POOL, connectionPool + ''' + Synthesize speech with given text by streaming mode, async call and play the synthesized audio in real-time. + for more information, please refer to + ''' + + complete_event = threading.Event() + + # Define a callback to handle the result + + class Callback(ResultCallback): + def on_open(self): + # when using object pool, on_open will be called after task start + self.file = open(f'result_{task_id}.mp3', 'wb') + print(f'[task_{task_id}] start') + + def on_complete(self): + print(f'[task_{task_id}] speech synthesis task complete successfully.') + complete_event.set() + + def on_error(self, message: str): + print(f'[task_{task_id}] speech synthesis task failed, {message}') + + def on_close(self): + # when using object pool, on_close will be called after task finished + print(f'[task_{task_id}] finished') + + def on_event(self, message): + # print(f'recv speech synthsis message {message}') + pass + + def on_data(self, data: bytes) -> None: + # send to player + # save audio to file + self.file.write(data) + + # Call the speech synthesizer callback + synthesizer_callback = Callback() + + # Initialize the speech synthesizer + # you can customize the synthesis parameters, like voice, format, sample_rate or other parameters + if USE_CONNECTION_POOL: + speech_synthesizer = connectionPool.borrow_synthesizer( + model='cosyvoice-v3-flash', + voice='longanyang', + seed=12382, + callback=synthesizer_callback + ) + else: + speech_synthesizer = SpeechSynthesizer(model='cosyvoice-v3-flash', + voice='longanyang', + seed=12382, + callback=synthesizer_callback) + try: + speech_synthesizer.call(text) + except Exception as e: + print(f'[task_{task_id}] speech synthesis task failed, {e}') + if USE_CONNECTION_POOL: + # close the synthesizer connection manually if task failed when using connection pool. + speech_synthesizer.close() + return + + print('[task_{}] Synthesized text: {}'.format(task_id, text)) + complete_event.wait() + print('[task_{}][Metric] requestId: {}, first package delay ms: {}'.format( + task_id, + speech_synthesizer.get_last_request_id(), + speech_synthesizer.get_first_package_delay())) + if USE_CONNECTION_POOL: + connectionPool.return_synthesizer(speech_synthesizer) + + # main function + if __name__ == '__main__': + # 必须先设置 dashscope.api_key 和 base_websocket_api_url,再创建 SpeechSynthesizerObjectPool。 + # 池在初始化时即按当前全局 dashscope.api_key 建立 WebSocket 连接, + # 池创建后再修改 dashscope.api_key 不会影响池内已建连接。 + dashscope.base_websocket_api_url='wss://maas.qianwenaiapi.com/api-ws/v1/inference' + init_dashscope_api_key() + + if USE_CONNECTION_POOL: + print('creating connection pool') + start_time = time.time() * 1000 + connectionPool = SpeechSynthesizerObjectPool(max_size=3) + end_time = time.time() * 1000 + print('connection pool created, cost: {} ms'.format(end_time - start_time)) + + task_thread_list = [] + for task_id in range(3): + thread = threading.Thread( + target=synthesis_text_to_speech_and_play_by_streaming_mode, + args=(text_to_synthesize[task_id], task_id)) + task_thread_list.append(thread) + + for task_thread in task_thread_list: + task_thread.start() + + for task_thread in task_thread_list: + task_thread.join() + + if USE_CONNECTION_POOL: + connectionPool.shutdown() + ``` + + #### 资源管理与异常处理 + + - 任务成功:当语音合成任务正常完成时,必须调用 `connectionPool.return_synthesizer(speech_synthesizer)` 将 `SpeechSynthesizer` 对象归还到池中,以便复用。 + + + 不要归还未完成任务或任务失败的`SpeechSynthesizer`对象。 + + - 任务失败:当 SDK 内部或业务逻辑抛出异常导致任务中断时,主动关闭底层的 WebSocket 连接:`speech_synthesizer.close()` + - 在所有语音合成任务完成后,要通过如下方式关闭对象池:`connectionPool.shutdown()`。 + - 在服务出现TaskFailed报错时,不需要额外处理。 + + + + Java SDK通过内置的连接池和自定义的对象池协同工作,实现最佳性能。 + + - 连接池:SDK 内部集成的 OkHttp3 连接池,负责管理和复用底层的 WebSocket 连接,减少网络握手开销。此功能默认开启。 + - 对象池:基于 `commons-pool2` 实现,用于维护一组已预先建立好连接的 `SpeechSynthesizer` 对象。从池中获取对象可消除连接建立的延迟,显著降低首包延迟。 + + #### 实现步骤 + + 1. 添加依赖 + + 根据项目构建工具,在依赖配置文件中添加 dashscope-sdk-java 和 commons-pool2。 + + 以Maven和Gradle为例,配置如下: + + + + 1) 打开Maven项目的`pom.xml`文件。 + 2) 在``标签内添加以下依赖信息。 + + ```xml + + com.alibaba + dashscope-sdk-java + + the-latest-version + + + + org.apache.commons + commons-pool2 + + the-latest-version + + ``` + + 1. 保存`pom.xml`文件。 + 2. 使用Maven命令(如`mvn clean install`或`mvn compile`)来更新项目依赖 + + + + 1. 打开Gradle项目的`build.gradle`文件。 + 2. 在`dependencies`块内添加以下依赖信息。 + + ```groovy + dependencies { + // 请将 'the-latest-version' 替换为2.16.6及以上版本,可在如下链接查询相关版本号:https://mvnrepository.com/artifact/com.alibaba/dashscope-sdk-java + implementation group: 'com.alibaba', name: 'dashscope-sdk-java', version: 'the-latest-version' + + // 请将 'the-latest-version' 替换为最新版本,可在如下链接查询相关版本号:https://mvnrepository.com/artifact/org.apache.commons/commons-pool2 + implementation group: 'org.apache.commons', name: 'commons-pool2', version: 'the-latest-version' + } + ``` + + 3. 保存`build.gradle`文件。 + 4. 在命令行中,切换到项目根目录,执行以下Gradle命令来更新项目依赖。 + + ```bash + ./gradlew build --refresh-dependencies + ``` + + 或者,如果使用Windows系统,命令应为: + + ```cmd + gradlew build --refresh-dependencies + ``` + + + + 2. 配置连接池 + + 通过环境变量配置连接池关键参数: + +| **环境变量** | **描述** | +| ---------------------------------------------- | -------------------------------------------------------------------------- | +| DASHSCOPE\_CONNECTION\_POOL\_SIZE | 连接池大小。
推荐值:峰值并发数的 2 倍以上。
默认值:32。 | +| DASHSCOPE\_MAXIMUM\_ASYNC\_REQUESTS | 最大异步请求数。
推荐值:与 `DASHSCOPE_CONNECTION_POOL_SIZE` 保持一致。
默认值:32。 | +| DASHSCOPE\_MAXIMUM\_ASYNC\_REQUESTS\_PER\_HOST | 单主机最大异步请求数。
推荐值:与 `DASHSCOPE_CONNECTION_POOL_SIZE` 保持一致。
默认值:32。 | + + 3. 配置对象池 + + 通过环境变量配置对象池大小: + +| **环境变量** | **描述** | +| --------------------------- | ----------------------------------------------- | +| COSYVOICE\_OBJECTPOOL\_SIZE | 对象池大小。
推荐值:峰值并发数的 1.5 至 2 倍。
默认值:500。 | + + + - 对象池的大小(`COSYVOICE_OBJECTPOOL_SIZE`)必须小于或等于连接池的大小(`DASHSCOPE_CONNECTION_POOL_SIZE`)。否则,当对象池请求对象时,若连接池已满,会导致调用线程阻塞,等待可用连接。 + - 对象池大小不应超过账户的 QPS(每秒查询率)限制。 + + + 通过如下代码创建对象池: + + ```java + class CosyvoiceObjectPool { + // 。。。这里省略其它代码,完整示例请参见完整代码 + public static GenericObjectPool getInstance() { + lock.lock(); + if (synthesizerPool == null) { + // 您可以在这里设置对象池的大小。或在环境变量COSYVOICE_OBJECTPOOL_SIZE中设置。 + // 建议设置为服务器最大并发连接数的1.5到2倍。 + int objectPoolSize = getObjectivePoolSize(); + SpeechSynthesizerObjectFactory speechSynthesizerObjectFactory = + new SpeechSynthesizerObjectFactory(); + GenericObjectPoolConfig config = + new GenericObjectPoolConfig<>(); + config.setMaxTotal(objectPoolSize); + config.setMaxIdle(objectPoolSize); + config.setMinIdle(objectPoolSize); + synthesizerPool = + new GenericObjectPool<>(speechSynthesizerObjectFactory, config); + } + lock.unlock(); + return synthesizerPool; + } + } + ``` + + 4. 从对象池中获取`SpeechSynthesizer`对象 + + 如果当前未归还的对象数量已超过对象池的最大容量,系统会额外创建一个新的`SpeechSynthesizer`对象。 + + 此类新创建的对象需要重新进行初始化并建立 WebSocket 连接,无法利用对象池的既有连接资源,因此不具备复用效果。 + + ```java + synthesizer = CosyvoiceObjectPool.getInstance().borrowObject(); + ``` + + 5. 进行语音合成 + + 从对象池借出`SpeechSynthesizer`对象后,需先调用`updateParamAndCallback(param, callback)`关联本次任务的参数与回调,再调用`streamingCall`或`call`方法进行语音合成。 + + + - 在对象池场景中,`updateParamAndCallback`会被**多次调用**(每次借出对象时都需调用一次,用于切换该次任务的回调和任务级参数,如`voice`、`format`等)。**多次调用时传入的**`apiKey`**必须始终相同**。`updateParamAndCallback`只更新当前`SpeechSynthesizer`实例的本地字段,不会重建底层 WebSocket 连接;而 SDK 仅在 WebSocket 建连握手时将`apiKey`写入`Authorization`请求头用于鉴权,后续任务消息(如`run-task`)本身不携带`apiKey`。因此只要复用的连接未断开,传入新的`apiKey`不会被发送到服务端,请求实际仍会使用连接首次握手时的`apiKey`,可能导致身份、配额或计费归属与预期不一致。 + - 如确需使用多个不同的 API Key,请为每个 API Key 维护**独立的对象池实例**。 + + 6. 归还`SpeechSynthesizer`对象 + + 语音合成任务结束后,归还`SpeechSynthesizer`对象,以便后续任务可以复用该对象。 + + 不要归还未完成任务或任务失败的对象。 + + ```java + CosyvoiceObjectPool.getInstance().returnObject(synthesizer); + ``` + + ##### 完整代码 + + + 复制使用前请注意:对象池场景下多次调用`updateParamAndCallback`时传入的 apiKey**必须始终相同**——SDK 不会更新已建立连接的 apiKey,传入不同的 apiKey 不会生效。多 API Key 场景请为每个 API Key 维护独立的对象池实例。详见上文重要说明。 + + + ```java expandable + import com.alibaba.dashscope.audio.tts.SpeechSynthesisResult; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisAudioFormat; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisParam; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesizer; + import com.alibaba.dashscope.common.ResultCallback; + import com.alibaba.dashscope.exception.NoApiKeyException; + import com.alibaba.dashscope.utils.Constants; + import lombok.extern.slf4j.Slf4j; + import org.apache.commons.pool2.BasePooledObjectFactory; + import org.apache.commons.pool2.PooledObject; + import org.apache.commons.pool2.impl.DefaultPooledObject; + import org.apache.commons.pool2.impl.GenericObjectPool; + import org.apache.commons.pool2.impl.GenericObjectPoolConfig; + + import java.time.LocalDateTime; + import java.util.concurrent.ExecutorService; + import java.util.concurrent.Executors; + import java.util.concurrent.TimeUnit; + import java.util.concurrent.locks.Lock; + + /** + * 您需要在项目中引入org.apache.commons.pool2和DashScope相关的包。 + * + * DashScope SDK 2.16.6及后续版本针对高并发场景进行了优化, + * DashScope SDK 2.16.6之前的版本不推荐在高并发场景下使用。 + * + * + * 在对TTS服务进行高并发调用之前, + * 请通过以下环境变量配置连接池的相关参数。 + * + * DASHSCOPE_MAXIMUM_ASYNC_REQUESTS + * DASHSCOPE_MAXIMUM_ASYNC_REQUESTS_PER_HOST + * DASHSCOPE_CONNECTION_POOL_SIZE + * + */ + + class SpeechSynthesizerObjectFactory + extends BasePooledObjectFactory { + public SpeechSynthesizerObjectFactory() { + super(); + } + @Override + public SpeechSynthesizer create() throws Exception { + return new SpeechSynthesizer(); + } + + @Override + public PooledObject wrap(SpeechSynthesizer obj) { + return new DefaultPooledObject<>(obj); + } + } + + class CosyvoiceObjectPool { + public static GenericObjectPool synthesizerPool; + public static String COSYVOICE_OBJECTPOOL_SIZE_ENV = "COSYVOICE_OBJECTPOOL_SIZE"; + public static int DEFAULT_OBJECT_POOL_SIZE = 500; + private static Lock lock = new java.util.concurrent.locks.ReentrantLock(); + public static int getObjectivePoolSize() { + try { + Integer n = Integer.parseInt(System.getenv(COSYVOICE_OBJECTPOOL_SIZE_ENV)); + System.out.println("Using Object Pool Size In Env: "+ n); + return n; + } catch (NumberFormatException e) { + System.out.println("Using Default Object Pool Size: "+ DEFAULT_OBJECT_POOL_SIZE); + return DEFAULT_OBJECT_POOL_SIZE; + } + } + public static GenericObjectPool getInstance() { + lock.lock(); + if (synthesizerPool == null) { + // 您可以在这里设置对象池的大小。或在环境变量COSYVOICE_OBJECTPOOL_SIZE中设置。 + // 建议设置为服务器最大并发连接数的1.5到2倍。 + int objectPoolSize = getObjectivePoolSize(); + SpeechSynthesizerObjectFactory speechSynthesizerObjectFactory = + new SpeechSynthesizerObjectFactory(); + GenericObjectPoolConfig config = + new GenericObjectPoolConfig<>(); + config.setMaxTotal(objectPoolSize); + config.setMaxIdle(objectPoolSize); + config.setMinIdle(objectPoolSize); + synthesizerPool = + new GenericObjectPool<>(speechSynthesizerObjectFactory, config); + } + lock.unlock(); + return synthesizerPool; + } + } + + class SynthesizeTaskWithCallback implements Runnable { + String[] textArray; + String requestId; + long timeCost; + public SynthesizeTaskWithCallback(String[] textArray) { + this.textArray = textArray; + } + @Override + public void run() { + SpeechSynthesizer synthesizer = null; + long startTime = System.currentTimeMillis(); + // if recv onError + final boolean[] hasError = {false}; + try { + class ReactCallback extends ResultCallback { + ReactCallback() {} + + @Override + public void onEvent(SpeechSynthesisResult message) { + if (message.getAudioFrame() != null) { + try { + byte[] bytesArray = message.getAudioFrame().array(); + System.out.println("收到音频,音频文件流length为:" + bytesArray.length); + } catch (Exception e) { + throw new RuntimeException(e); + } + } + } + + @Override + public void onComplete() {} + + @Override + public void onError(Exception e) { + System.out.println(e.getMessage()); + e.printStackTrace(); + hasError[0] = true; + } + } + + SpeechSynthesisParam param = + SpeechSynthesisParam.builder() + .model("cosyvoice-v3-flash") + .voice("longanyang") + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:.apiKey("sk-xxx") + .apiKey(System.getenv("DASHSCOPE_API_KEY")) + .format(SpeechSynthesisAudioFormat + .MP3_22050HZ_MONO_256KBPS) // 流式合成使用PCM或者MP3 + .build(); + + try { + synthesizer = CosyvoiceObjectPool.getInstance().borrowObject(); + // 注意:对象池场景下,多次调用 updateParamAndCallback 时传入的 apiKey 必须始终相同;SDK 不会更新已建立连接的 apiKey,传入不同的 apiKey 不会生效。详见上文“进行语音合成”步骤的重要说明。 + synthesizer.updateParamAndCallback(param, new ReactCallback()); + for (String text : textArray) { + synthesizer.streamingCall(text); + } + Thread.sleep(20); + synthesizer.streamingComplete(60000); + requestId = synthesizer.getLastRequestId(); + } catch (Exception e) { + System.out.println("Exception e: " + e.toString()); + hasError[0] = true; + } + } catch (Exception e) { + hasError[0] = true; + throw new RuntimeException(e); + } + if (synthesizer != null) { + try { + if (hasError[0] == true) { + // 如果出现异常,则关闭连接并在对象池中禁用该对象。 + synthesizer.getDuplexApi().close(1000, "bye"); + CosyvoiceObjectPool.getInstance().invalidateObject(synthesizer); + } else { + // 如果任务正常结束,则归还对象。 + CosyvoiceObjectPool.getInstance().returnObject(synthesizer); + } + } catch (Exception e) { + throw new RuntimeException(e); + } + long endTime = System.currentTimeMillis(); + timeCost = endTime - startTime; + System.out.println("[线程 " + Thread.currentThread() + "] 语音合成任务结束。耗时 " + timeCost + " ms, RequestId " + requestId); + } + } + } + + @Slf4j + public class SynthesizeTextToSpeechWithCallbackConcurrently { + public static void checkoutEnv(String envName, int defaultSize) { + if (System.getenv(envName) != null) { + System.out.println("[ENV CHECK]: " + envName + " " + + System.getenv(envName)); + } else { + System.out.println("[ENV CHECK]: " + envName + + " Using Default which is " + defaultSize); + } + } + + public static void main(String[] args) + throws InterruptedException, NoApiKeyException { + Constants.baseWebsocketApiUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference"; + // Check for connection pool env + checkoutEnv("DASHSCOPE_CONNECTION_POOL_SIZE", 32); + checkoutEnv("DASHSCOPE_MAXIMUM_ASYNC_REQUESTS", 32); + checkoutEnv("DASHSCOPE_MAXIMUM_ASYNC_REQUESTS_PER_HOST", 32); + checkoutEnv(CosyvoiceObjectPool.COSYVOICE_OBJECTPOOL_SIZE_ENV, CosyvoiceObjectPool.DEFAULT_OBJECT_POOL_SIZE); + + int runTimes = 3; + // Create the pool of SpeechSynthesis objects + ExecutorService executorService = Executors.newFixedThreadPool(runTimes); + + for (int i = 0; i < runTimes; i++) { + // Record the task submission time + LocalDateTime submissionTime = LocalDateTime.now(); + executorService.submit(new SynthesizeTaskWithCallback(new String[] { + "床前明月光,", "疑似地上霜。", "举头望明月,", "低头思故乡。"})); + } + + // Shut down the ExecutorService and wait for all tasks to complete + executorService.shutdown(); + executorService.awaitTermination(1, TimeUnit.MINUTES); + System.exit(0); + } + } + ``` + + #### 推荐配置 + + 以下配置基于在指定规格的阿里云服务器上仅运行 CosyVoice 语音合成服务的测试结果。过高的并发数可能导致任务处理延迟。 + + 其中单机并发数指的是同一时刻正在运行的CosyVoice语音合成任务数,也可以理解为工作线程数。 + +| **机器配置(阿里云)** | **单机最大并发数** | **对象池大小** | **连接池大小** | +| ------------- | ----------- | --------- | --------- | +| 4核8GiB | 100 | 500 | 2000 | +| 8核16GiB | 150 | 500 | 2000 | +| 16核32GiB | 200 | 500 | 2000 | + + #### 资源管理与异常处理 + + - 任务成功:当语音合成任务正常完成时,必须调用GenericObjectPool的returnObject方法将`SpeechSynthesizer`对象归还到池中,以便复用。 + + 在当前代码中,对应`CosyvoiceObjectPool.getInstance().returnObject(synthesizer)` + + + 不要归还未完成任务或任务失败的`SpeechSynthesizer`对象。 + + - 任务失败:当 SDK 内部或业务逻辑抛出异常导致任务中断时,必须执行以下两个操作: + + 1. 主动关闭底层的 WebSocket 连接 + 2. 从对象池中废弃该对象,防止被再次使用 + + ```java + // 在当前代码中对应如下内容 + // 关闭连接 + synthesizer.getDuplexApi().close(1000, "bye"); + // 在对象池中废弃出现异常的synthesizer + CosyvoiceObjectPool.getInstance().invalidateObject(synthesizer); + ``` + + - 在服务出现TaskFailed报错时,不需要额外处理。 + + #### 调用预热与耗时统计说明 + + 在对 DashScope Java SDK 进行并发调用延迟等性能评估时,建议先执行充分的预热操作,确保测量结果反映稳定状态下的真实性能,避免初始连接耗时导致数据偏差。 + + ##### 连接复用机制 + + DashScope Java SDK 通过全局单例的连接池高效管理和复用 WebSocket 连接,旨在减少频繁建连和断连的开销,提升高并发场景下的处理能力。 + + 该机制的工作特点如下: + + - **按需创建**:SDK 不会在服务启动时预创建 WebSocket 连接,而是在首次调用时按需建立。 + - **限时复用**:请求完成后,连接将在池中保留最多 60 秒以备复用。 + + - 若 60 秒内有新请求,将复用现有连接,避免重复握手开销。 + - 若连接空闲超过 60 秒,将被自动关闭以释放资源。 + + ##### 预热的重要性 + + 在以下场景中,连接池中可能没有可复用的活跃连接,导致请求需要新建连接: + + - 应用刚启动,尚未发起任何调用。 + - 服务空闲时间超过 60 秒,池中连接已因超时而关闭。 + + 在这些场景下,首次请求需完成 WebSocket 建连(TCP 握手、TLS 协商、协议升级),延迟显著高于后续复用连接的请求。若未预热,性能测试结果会因包含建连耗时而产生偏差。 + + ##### SDK侧延迟与实际首包延迟的区别 + + SDK侧打印的首包延迟(如通过 `get_first_package_delay()` 获取的值)包含了 WebSocket 建联和网络传输等耗时,并不等同于模型服务的实际首包延迟。 + + 实际首包延迟是指从服务端收到 `run-task` 指令到返回第一个 `result-generated` 事件的时间间隔,该值可通过服务端日志查看。 + + 在高并发场景下,由于大量连接的建立和资源调度,SDK侧打印的延迟数值可能显著高于服务端的实际首包延迟。如果观察到 SDK 报告的首包延迟较高,建议: + + - 对比服务端日志中的首包延迟(从 `run-task` 到首个 `result-generated`),确认模型推理性能是否正常。 + - 使用上述对象池或连接池机制进行预热,消除 WebSocket 建连开销,使 SDK 侧打印的延迟更接近实际首包延迟。 + + ##### 推荐做法 + + 为获取可靠的性能数据,在正式进行性能压测或延迟统计前,请遵循以下预热步骤: + + 1. 模拟正式测试的并发级别,提前发起一定数量的调用(例如,持续 1-2 分钟),以充分填充连接池。 + 2. 确认连接池已建立并维持足够的活跃连接后,再开始正式的性能数据采集。 + + 通过合理的预热,可使 SDK 连接池进入稳定复用状态,从而测量出更具代表性的延迟指标,真实反映服务在线上平稳运行时的性能。 + + #### Java SDK常见异常 + + + **出错原因:** + + **类型一:** + + 每一个 SDK 对象创建时都会申请一个连接。如果没有使用对象池,每一次任务结束后对象都被析构。此时这一个连接将进入无引用状态,需要等待 61s 秒后服务端报错连接超时才会真正断开,这会导致这个连接在 61 秒内不可复用。 + + 在高并发场景下,新的任务在发现没有可复用连接时会创建新连接,会造成如下后果: + + 1. 连接数持续上升。 + 2. 由于连接数过多,服务器资源不足,服务器卡顿。 + 3. 连接池被打满、新任务由于启动时需要等待可用连接而阻塞。 + + **类型二:** + + 对象池配置的MaxIdle小于MaxTotal,导致在对象闲置时,超过MaxIdle的对象被销毁,从而造成连接泄漏。泄漏的连接需要等待61秒超时后断连,同类型一造成连接数持续上升。 + + **解决方法**: + + 对于类型一,使用对象池解决。 + + 对于类型二,检查对象池配置参数,设置MaxIdle和MaxTotal相等,关闭对象池自动销毁策略解决。 + + + + 同“**异常 1**”,连接池已经达到最大连接限制,新的任务需要等待无引用状态的连接 61 秒触发超时后才可以获得连接。 + + + + **出错原因**: + + 在高并发调用时,同一个对象会复用同一个WebSocket连接,因此WebSocket连接只会在服务启动时创建。需要注意的是,任务启动阶段如果立刻开始较高并发调用,同时创建过多的WebSocket连接会导致阻塞。 + + **解决方法**: + + 启动服务后逐步提升并发量,或增加预热任务。 + + + + **出错原因**: + + 这是由于出现了客户端报错后,服务端不知道客户端出错,连接处于任务中状态。此时连接和对象被复用并开启下一个任务,导致流程错误,下一个任务失败。 + + **解决方法**: + + 在抛出异常后主动关闭 WebSocket 连接后归还对象池。 + + + + **出错原因**: + + 同时创建过多 WebSocket 连接导致阻塞,但业务流量持续打进来,导致任务短时间积压,并且在阻塞后所有积压任务立刻调用。这会造成调用量尖刺,并且有可能造成瞬时超过账号的并发数限制导致部分任务失败、服务器卡顿等。 + + 这种瞬间创建过多 WebSocket 的情况多发生于: + + - 服务启动阶段 + - 网络出现异常,大量 WebSocket 连接同时中断重连 + - 某一时刻出现大量服务端报错,导致大量 WebSocket 重连。常见报错如并发数超过账号限制(“Requests rate limit exceeded, please try again later.”)。 + + **解决方法**: + + 1. 检查网络情况。 + 2. 排查尖刺前是否出现大量其他服务端报错。 + 3. 提高账号并发限制。 + 4. 调小对象池和连接池大小,通过对象池上限限制最大并发数。 + 5. 提升服务器配置或扩充机器数。 + + + + **解决方法**: + + 1. 检查是否已经达到网络带宽上限。 + 2. 检查实际并发数是否已经过高。 + +
+
+
+ + + Sambert仅Java SDK内置了池化机制,Python SDK暂不支持。 + + #### 前提条件 + + - [获取与配置 API Key](/api-reference/preparation/api-key) + - 已安装符合版本要求的DashScope Java SDK,建议[安装最新版](/api-reference/preparation/install-sdk),SDK版本需≥2.16.6 + + #### 推荐配置 + + 连接池和对象池不是越多越好,过少或过多都会导致程序运行变慢。建议根据服务器实际规格配置。 + + 在服务器上只运行Sambert语音合成服务的情况下,进行测试后得到了如下推荐配置供参考: + +| **常见机器配置(阿里云)** | **单机最大并发数** | **对象池大小** | **连接池大小** | +| --------------- | ----------- | --------- | --------- | +| 4核8GiB | 600 | 1200 | 2000 | + + > 单机并发数指的是同一时刻正在运行的Sambert语音合成任务的数量,也可以理解为工作线程数。 + + + 在高并发调用时,同一个对象会复用同一个WebSocket连接,因此WebSocket连接只会在服务启动时创建。 + + 同时创建过多的 WebSocket 连接会导致阻塞,启动服务时应逐步提高单机并发数。 + + + #### 可配置参数 + + ##### 连接池 + + DashScope Java SDK使用了OkHttp3提供的连接池来复用WebSocket连接,从而减少频繁创建WebSocket连接的耗时和资源开销。 + + 连接池是DashScope SDK默认开启的优化项,需要根据使用场景配置连接池大小。 + + 请在运行Java服务前,通过环境变量的方式提前按需配置好连接池的相关参数。连接池配置参数如下: + +| DASHSCOPE\_CONNECTION\_POOL\_SIZE | 配置连接池大小。默认值为32。
推荐配置为峰值并发数的2倍以上。 | +| ---------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| DASHSCOPE\_MAXIMUM\_ASYNC\_REQUESTS | 配置最大异步请求数。默认值为32。
推荐配置为和连接池大小一致。
更多信息参见[参考文档](https://javadoc.io/static/com.squareup.okhttp3/okhttp/3.14.9/okhttp3/Dispatcher.html#setMaxRequests-int-)。 | +| DASHSCOPE\_MAXIMUM\_ASYNC\_REQUESTS\_PER\_HOST | 配置单host最大异步请求数。默认值为32。
推荐配置为和连接池大小一致。
更多信息参见[参考文档](https://javadoc.io/static/com.squareup.okhttp3/okhttp/3.14.9/okhttp3/Dispatcher.html#setMaxRequestsPerHost-int-)。 | + + ##### 对象池 + + 推荐使用对象池的方式来复用`SpeechSynthesizer`对象,这样可以进一步降低反复创建和销毁对象带来的内存和时间开销。 + + 请在运行 Java 服务前,通过环境变量或代码的方式提前按需配置好对象池的大小。对象池配置参数如下: + +| SAMBERT\_OBJECTPOOL\_SIZE | 对象池大小。
推荐配置为峰值并发数的1.5\~2倍。
对象池大小需要小于或等于连接池大小,否则会出现对象等待连接的情况,导致调用阻塞。 | +| ------------------------- | ----------------------------------------------------------------------------- | + + 关于如何配置环境变量,可参考[配置API Key到环境变量](/api-reference/preparation/export-api-key-env)。 + + #### 示例代码 + + 以下为使用资源池的示例代码。其中,对象池为全局单例对象。 + + - 每个主账号默认每秒可提交3个Sambert语音合成任务。 + + 如需开通更高QPS请联系我们。 + + + 以Maven和Gradle为例,配置如下: + + + + 1. 打开Maven项目的`pom.xml`文件。 + 2. 在``标签内添加以下依赖信息。 + + ```xml + + com.alibaba + dashscope-sdk-java + + the-latest-version + + + + org.apache.commons + commons-pool2 + + the-latest-version + + ``` + + 1. 保存`pom.xml`文件。 + 2. 使用Maven命令(如`mvn clean install`或`mvn compile`)来更新项目依赖 + + + + 1. 打开Gradle项目的`build.gradle`文件。 + 2. 在`dependencies`块内添加以下依赖信息。 + + ```gradle + dependencies { + // 请将 'the-latest-version' 替换为2.16.9及以上版本,可在如下链接查询相关版本号:https://mvnrepository.com/artifact/com.alibaba/dashscope-sdk-java + implementation group: 'com.alibaba', name: 'dashscope-sdk-java', version: 'the-latest-version' + + // 请将 'the-latest-version' 替换为最新版本,可在如下链接查询相关版本号:https://mvnrepository.com/artifact/org.apache.commons/commons-pool2 + implementation group: 'org.apache.commons', name: 'commons-pool2', version: 'the-latest-version' + } + ``` + + 3. 保存`build.gradle`文件。 + 4. 在命令行中,切换到项目根目录,执行以下Gradle命令来更新项目依赖。 + + ```bash + ./gradlew build --refresh-dependencies + ``` + + 或者,如果使用Windows系统,命令应为: + + ```cmd + gradlew build --refresh-dependencies + ``` + + + 如果项目中没有Gradle Wrapper文件(`gradlew`或`gradlew.bat`),可以: + + - 使用已安装的Gradle直接运行`gradle build --refresh-dependencies` + - 或先运行`gradle wrapper`生成wrapper文件,然后再运行上述命令 + + + + + + - 示例代码中,不同的线程通过等待随机时间来避免同时创建过多的WebSocket连接。 + + ```java expandable + import com.alibaba.dashscope.audio.tts.SpeechSynthesisAudioFormat; + import com.alibaba.dashscope.audio.tts.SpeechSynthesisParam; + import com.alibaba.dashscope.audio.tts.SpeechSynthesisResult; + import com.alibaba.dashscope.audio.tts.SpeechSynthesizer; + import com.alibaba.dashscope.common.ResultCallback; + import com.alibaba.dashscope.exception.NoApiKeyException; + import lombok.extern.slf4j.Slf4j; + import org.apache.commons.pool2.BasePooledObjectFactory; + import org.apache.commons.pool2.PooledObject; + import org.apache.commons.pool2.impl.DefaultPooledObject; + import org.apache.commons.pool2.impl.GenericObjectPool; + import org.apache.commons.pool2.impl.GenericObjectPoolConfig; + + import java.util.Random; + import java.util.concurrent.CountDownLatch; + import java.util.concurrent.ExecutorService; + import java.util.concurrent.Executors; + import java.util.concurrent.TimeUnit; + import java.util.concurrent.locks.Lock; + import com.alibaba.dashscope.utils.Constants; + + /** + * Before making high-concurrency calls to the TTS service, + * please configure the connection pool size through following environment + * variables. + * + * DASHSCOPE_MAXIMUM_ASYNC_REQUESTS=2000 + * DASHSCOPE_MAXIMUM_ASYNC_REQUESTS_PER_HOST=2000 + * DASHSCOPE_CONNECTION_POOL_SIZE=2000 + * + * The default is 32, and it is recommended to set it to 2 times the maximum + * concurrent connections of a single server. + */ + + @Slf4j + public class SynthesizeTextToSpeechUsingSambertConcurrently { + public static void checkoutEnv(String envName, int defaultSize) { + if (System.getenv(envName) != null) { + System.out.println("[ENV CHECK]: " + envName + " " + + System.getenv(envName)); + } else { + System.out.println("[ENV CHECK]: " + envName + + " Using Default which is " + defaultSize); + } + } + + public static void main(String[] args) + throws InterruptedException, NoApiKeyException { + Constants.baseHttpApiUrl = "https://maas.qianwenaiapi.com/api/v1"; + + // Check for connection pool env + checkoutEnv("DASHSCOPE_CONNECTION_POOL_SIZE", 32); + checkoutEnv("DASHSCOPE_MAXIMUM_ASYNC_REQUESTS", 32); + checkoutEnv(SambertObjectPool.SAMBERT_OBJECTPOOL_SIZE_ENV, SambertObjectPool.DEFAULT_CONNECTION_POOL_SIZE); + checkoutEnv("DASHSCOPE_MAXIMUM_ASYNC_REQUESTS_PER_HOST", 32); + + // Record task start time + int runTimes = 1; + + // Create the pool of SpeechSynthesis objects + ExecutorService executorService = Executors.newFixedThreadPool(runTimes); + + for (int i = 0; i < runTimes; i++) { + executorService.submit(new SynthesizeTask(new String[]{ + "床前明月光,", + "疑似地上霜。", + "举头望明月,", + "低头思故乡。" + })); + } + + // Shut down the ExecutorService and wait for all tasks to complete + executorService.shutdown(); + executorService.awaitTermination(1, TimeUnit.MINUTES); + System.exit(0); + } + } + + class SpeechSynthesizerObjectFactory + extends BasePooledObjectFactory { + public SpeechSynthesizerObjectFactory() { + super(); + } + @Override + public SpeechSynthesizer create() throws Exception { + return new SpeechSynthesizer(); + } + + @Override + public PooledObject wrap(SpeechSynthesizer obj) { + return new DefaultPooledObject<>(obj); + } + } + + class SambertObjectPool { + public static GenericObjectPool synthesizerPool; + public static String SAMBERT_OBJECTPOOL_SIZE_ENV = "SAMBERT_OBJECTPOOL_SIZE"; + public static int DEFAULT_CONNECTION_POOL_SIZE = 500; + private static Lock lock = new java.util.concurrent.locks.ReentrantLock(); + public static int getObjectivePoolSize() { + try { + Integer n = Integer.parseInt(System.getenv(SAMBERT_OBJECTPOOL_SIZE_ENV)); + return n; + } catch (NumberFormatException e) { + return DEFAULT_CONNECTION_POOL_SIZE; + } + } + public static GenericObjectPool getInstance() { + lock.lock(); + if (synthesizerPool == null) { + // You can set the object pool size here. or in environment variable + // SAMBERT_OBJECTPOOL_SIZE It is recommended to set it to 1.5 to 2 times + // your server's maximum concurrent connections. + int objectPoolSize = getObjectivePoolSize(); + SpeechSynthesizerObjectFactory speechSynthesizerObjectFactory = + new SpeechSynthesizerObjectFactory(); + GenericObjectPoolConfig config = + new GenericObjectPoolConfig<>(); + config.setMaxTotal(objectPoolSize); + config.setMaxIdle(objectPoolSize); + config.setMinIdle(objectPoolSize); + synthesizerPool = + new GenericObjectPool<>(speechSynthesizerObjectFactory, config); + } + lock.unlock(); + return synthesizerPool; + } + } + + class SynthesizeTask implements Runnable { + String[] textList; + String requestId; + long timeCost; + public SynthesizeTask(String[] textList) { + this.textList = textList; + } + @Override + public void run() { + // sleep random time before start task, avoid creating too much websocket at the same time. + Random random = new Random(); + try { + Thread.sleep(random.nextInt(30*1000)); + } catch (InterruptedException e) { + throw new RuntimeException(e); + } + for (String text:textList) { + SpeechSynthesizer synthesizer = null; + long startTime = System.currentTimeMillis(); + + try { + CountDownLatch latch = new CountDownLatch(1); + class ReactCallback extends ResultCallback { + ReactCallback() {} + + @Override + public void onEvent(SpeechSynthesisResult message) { + if (message.getAudioFrame() != null) { + try { + byte[] bytesArray = message.getAudioFrame().array(); + } catch (Exception e) { + throw new RuntimeException(e); + } + } + } + + @Override + public void onComplete() { + latch.countDown(); + } + + @Override + public void onError(Exception e) { + System.out.println(e.getMessage()); + e.printStackTrace(); + latch.countDown(); + } + } + + SpeechSynthesisParam param = + SpeechSynthesisParam.builder() + .model("sambert-zhichu-v1") + .format(SpeechSynthesisAudioFormat.MP3) // 使用PCM或者MP3 + .text(text) + .enablePhonemeTimestamp(true) + .enableWordTimestamp(true) + // 若没有配置环境变量,请用千问AI平台API Key将下行替换为:.apiKey("sk-xxx") + .apiKey(System.getenv("DASHSCOPE_API_KEY")) + .build(); + + try { + synthesizer = SambertObjectPool.getInstance().borrowObject(); + synthesizer.call(param, new ReactCallback()); + try { + latch.await(); + } catch (InterruptedException e) { + throw new RuntimeException(e); + } + requestId = synthesizer.getLastRequestId(); + } catch (Exception e) { + System.out.println("Exception e: " + e.toString()); + synthesizer.getSyncApi().close(1000, "bye"); + } + } catch (Exception e) { + throw new RuntimeException(e); + } finally { + if (synthesizer != null) { + try { + // Return the SpeechSynthesizer object to the pool + SambertObjectPool.getInstance().returnObject(synthesizer); + } catch (Exception e) { + e.printStackTrace(); + } + } + } + long endTime = System.currentTimeMillis(); + timeCost = endTime - startTime; + System.out.println("[线程" + Thread.currentThread() + "] 语音合成任务:(" + text + ")结束。耗时" + timeCost + "ms, RequestId" + requestId); + } + } + } + ``` + + ##### 异常处理 + + - 在服务出现TaskFailed报错时,不需要额外处理。 + - 如果在语音合成中途,客户端出现错误(如SDK内部异常或业务逻辑异常)导致语音合成任务未完成,则需要主动关闭连接。 + + 关闭连接方法如下: + + ```java + // 将下面这段代码放在try-catch块中 + synthesizer.getSyncApi().close(1000, "bye"); + ``` + + #### 常见异常 + + + **出错原因:** + + **类型一:** + + 每一个 SDK 对象创建时都会申请一个连接。如果没有使用对象池,每一次任务结束后对象都被析构。此时这一个连接将进入无引用状态,需要等待 61s 秒后服务端报错连接超时才会真正断开,这会导致这个连接在 61 秒内不可复用。 + + 在高并发场景下,新的任务在发现没有可复用连接时会创建新连接,会造成如下后果: + + 1. 连接数持续上升。 + 2. 由于连接数过多,服务器资源不足,服务器卡顿。 + 3. 连接池被打满、新任务由于启动时需要等待可用连接而阻塞。 + + **类型二:** + + 对象池配置的MaxIdle小于MaxTotal,导致在对象闲置时,超过MaxIdle的对象被销毁,从而造成连接泄漏。泄漏的连接需要等待61秒超时后断连,同类型一造成连接数持续上升。 + + **解决方法**: + + 对于类型一,使用对象池解决。 + + 对于类型二,检查对象池配置参数,设置MaxIdle和MaxTotal相等,关闭对象池自动销毁策略解决。 + + + + 同“**异常 1**”,连接池已经达到最大连接限制,新的任务需要等待无引用状态的连接 61 秒触发超时后才可以获得连接。 + + + + **出错原因**: + + 在高并发调用时,同一个对象会复用同一个WebSocket连接,因此WebSocket连接只会在服务启动时创建。需要注意的是,任务启动阶段如果立刻开始较高并发调用,同时创建过多的WebSocket连接会导致阻塞。 + + **解决方法**: + + 启动服务后逐步提升并发量,或增加预热任务。 + + + + **出错原因**: + + 这是由于出现了客户端报错后,服务端不知道客户端出错,连接处于任务中状态。此时连接和对象被复用并开启下一个任务,导致流程错误,下一个任务失败。 + + **解决方法**: + + 在抛出异常后主动关闭 WebSocket 连接后归还对象池。 + + + + **出错原因**: + + 同时创建过多 WebSocket 连接导致阻塞,但业务流量持续打进来,导致任务短时间积压,并且在阻塞后所有积压任务立刻调用。这会造成调用量尖刺,并且有可能造成瞬时超过账号的并发数限制导致部分任务失败、服务器卡顿等。 + + 这种瞬间创建过多 WebSocket 的情况多发生于: + + - 服务启动阶段 + - 网络出现异常,大量 WebSocket 连接同时中断重连 + - 某一时刻出现大量服务端报错,导致大量 WebSocket 重连。常见报错如并发数超过账号限制(“Requests rate limit exceeded, please try again later.”)。 + + **解决方法**: + + 1. 检查网络情况。 + 2. 排查尖刺前是否出现大量其他服务端报错。 + 3. 提高账号并发限制。 + 4. 调小对象池和连接池大小,通过对象池上限限制最大并发数。 + 5. 提升服务器配置或扩充机器数。 + + + + **解决方法**: + + 1. 检查是否已经达到网络带宽上限。 + 2. 检查实际并发数是否已经过高。 + +
+
+
+ +## 支持的模型 + +支持以下模型: + +- \*\*Qwen-Audio-TTS:\*\*qwen-audio-3.0-tts-plus、qwen-audio-3.1-tts-flash、qwen-audio-3.0-tts-flash +- \*\*CosyVoice:\*\*cosyvoice-v3.5-plus、cosyvoice-v3.5-flash、cosyvoice-v3-plus、cosyvoice-v3-flash、cosyvoice-v2、cosyvoice-v1 +- **Qwen-TTS**: + + - **Qwen3-TTS-Instruct-Flash-Realtime**:qwen3-tts-instruct-flash-realtime(稳定版,当前等同qwen3-tts-instruct-flash-realtime-2026-01-22)、qwen3-tts-instruct-flash-realtime-2026-01-22(最新快照版) + - \*\*Qwen3-TTS-VD-Realtime:\*\*qwen3-tts-vd-realtime-2026-01-15(最新快照版)、qwen3-tts-vd-realtime-2025-12-16(快照版) + - \*\*Qwen3-TTS-VC-Realtime:\*\*qwen3-tts-vc-realtime-2026-01-15(最新快照版)、qwen3-tts-vc-realtime-2025-11-27(快照版) + - \*\*Qwen3-TTS-Flash-Realtime:\*\*qwen3-tts-flash-realtime(稳定版,当前等同qwen3-tts-flash-realtime-2025-11-27)、qwen3-tts-flash-realtime-2025-11-27(最新快照版)、qwen3-tts-flash-realtime-2025-09-18(快照版) + - \*\*Qwen-TTS-Realtime:\*\*qwen-tts-realtime(稳定版,当前等同qwen-tts-realtime-2025-07-15)、qwen-tts-realtime-latest(最新版,当前等同qwen-tts-realtime-2025-07-15)、qwen-tts-realtime-2025-07-15(快照版) +- \*\*Sambert:\*\*sambert-zhinan-v1、sambert-zhiqi-v1、sambert-zhichu-v1、sambert-zhide-v1、sambert-zhijia-v1、sambert-zhiru-v1、sambert-zhiqian-v1、sambert-zhixiang-v1、sambert-zhiwei-v1、sambert-zhihao-v1、sambert-zhijing-v1、sambert-zhiming-v1、sambert-zhimo-v1、sambert-zhina-v1、sambert-zhishu-v1、sambert-zhistella-v1、sambert-zhiting-v1、sambert-zhixiao-v1、sambert-zhiya-v1、sambert-zhiye-v1、sambert-zhiying-v1、sambert-zhiyuan-v1、sambert-zhiyue-v1、sambert-zhigui-v1、sambert-zhishuo-v1、sambert-zhimiao-emo-v1、sambert-zhimao-v1、sambert-zhilun-v1、sambert-zhifei-v1、sambert-zhida-v1、sambert-camila-v1、sambert-perla-v1、sambert-indah-v1、sambert-clara-v1、sambert-hanna-v1、sambert-beth-v1、sambert-betty-v1、sambert-cally-v1、sambert-cindy-v1、sambert-eva-v1、sambert-donna-v1、sambert-brian-v1、sambert-waan-v1,请参见[Sambert模型列表](/api-reference/speech-synthesis/sambert/java-sdk#模型列表) + +## 支持的音色 + +不同模型支持的音色不同。将请求参数 `voice` 设为音色列表中 **voice参数** 列的值即可。 + +- [Qwen-Audio-TTS音色列表](/developer-guides/speech/voice-list/qwen-audio-tts) +- [CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice) +- [Qwen-TTS音色列表](/developer-guides/speech/voice-list/qwen-tts#qwen-tts实时语音合成音色列表) +- [Sambert音色列表](/api-reference/speech-synthesis/sambert/java-sdk#模型列表) + +## API参考 + +通过 AOQ 接入的流程和示例,请参见[AOQ 接入](/api-reference/realtime-api/aoq-access)。支持的模型及版本请参见[模型与协议支持范围](/api-reference/realtime-api/overview#模型支持力度)。 + +- [实时语音合成-Qwen-Audio-TTS API参考](/api-reference/speech-synthesis/qwen-audio-tts/model-access) / [实时语音合成-CosyVoice API参考](/developer-guides/speech/realtime-streaming) +- [实时语音合成-千问API参考](/api-reference/speech-synthesis/qwen-tts-realtime/model-access) +- [实时语音合成-Sambert API参考](/developer-guides/speech/realtime-streaming) + +## 常见问题 + +### Q:语音合成发音错误怎么办?多音字如何控制发音? + +- 将多音字替换为同音的其他汉字,快速解决发音问题。 +- 使用 SSML 标记语言控制发音:Sambert 和 CosyVoice 均支持 SSML。 + +### Q:使用复刻音色生成的音频无声音如何排查? + +1. **确认音色状态** + + 调用[CosyVoice声音复刻/设计API](/api-reference/speech-synthesis/voice-cloning/http-api)接口,确认音色的 `status` 是否为 `OK`。 +2. **检查模型版本一致性** + + 确保复刻音色时使用的 `target_model` 参数与语音合成时的 `model` 参数完全一致。例如: + + - 复刻时使用 `cosyvoice-v3-plus` + - 合成时也必须使用 `cosyvoice-v3-plus` +3. **验证源音频质量** + + 检查复刻音色时使用的源音频是否符合[CosyVoice声音复刻/设计API](/api-reference/speech-synthesis/voice-cloning/http-api): + + - 音频时长:10-20秒 + - 音质清晰 + - 无背景噪音 +4. **检查请求参数** + + 确认语音合成请求中的 `voice` 参数已设置为复刻音色的 ID。 + +### Q:声音复刻后合成效果不稳定或语音不完整怎么办? + +如果复刻音色后合成的语音出现以下问题: + +- 语音播放不完整,只读出部分文字 +- 合成效果不稳定,时好时坏 +- 语音中包含异常停顿或静音段 + +**可能原因**:源音频质量不符合要求。 + +**解决方案**:请检查源音频是否符合[录音操作指南](/developer-guides/speech/voice-cloning#录音建议)中的音频要求,建议按照录音指南重新录制。 + +### Q:为什么语音合成的实际时长与 WAV 文件显示的时长不一致? + +语音合成采用流式机制,边合成边返回数据,因此保存的 WAV 文件头中的时长是预估值,存在一定误差。如需精确时长,可将 format 设置为 pcm,待获取完整合成结果后自行添加 WAV 文件头信息。 + +### Q:为什么音频无法播放? + +请按以下场景逐一排查: + +1. 音频保存为完整文件(如 xx.mp3)的情况 + + 1. 音频格式一致性:请求参数中的音频格式须与文件后缀一致(如参数为 wav 则文件须为 .wav)。 + 2. 播放器兼容性:确认播放器支持该音频的格式和采样率。 +2. 流式播放音频的情况 + + 1. 将音频流保存为完整文件,尝试用播放器播放。如果文件无法播放,请参考场景 1 的排查方法。 + 2. 如果文件可正常播放,则问题在流式播放实现。请确认播放器支持流式播放(如 ffmpeg、pyaudio、AudioFormat、MediaSource 等)。 + +### Q:为什么音频播放卡顿? + +请按以下步骤逐一排查: + +1. 检查文本发送速度:确保发送间隔合理,避免上段音频播完后下段文本尚未到达。 +2. 检查回调函数性能: + + - 确认回调函数中无阻塞性业务逻辑。 + - 回调运行在 WebSocket 线程,阻塞会影响数据接收。建议将音频数据写入独立缓冲区,在其他线程中处理。 +3. 检查网络稳定性:网络波动可能导致音频传输中断或延迟。 + +### Q:语音合成耗时较长是什么原因? + +请按以下步骤排查: + +1. 检查输入间隔 + + 如果是流式合成,确认文本发送间隔是否过长,过长会导致合成总时长增加。 +2. 分析性能指标 + + - 首包延迟:正常约 500ms。 + - RTF(实时率 = 合成总耗时 / 音频时长):正常应小于 1.0。 + +### Q:合成的音频中读出了文本里的特殊符号怎么办? + +Qwen-TTS 系列模型可能将文本中的部分特殊符号(如 Markdown 加粗标记 \*\*)合成为语音。可通过以下方式处理: + +1. 调用前对文本进行预处理,去除特殊符号。 +2. 改用 CosyVoice 模型。 + +### Q:如何限制 API Key 仅用于语音合成服务(权限隔离)? + +通过新建业务空间并仅授权特定模型,可限制 API Key 的使用范围。请参见[业务空间管理](/api-reference/preparation/api-key)。 + +### Q:子业务空间的 API Key 能否调用 Qwen-Audio-TTS/CosyVoice 模型? + +默认业务空间下,所有模型均可调用。 + +子业务空间下,需要为 API Key 对应的子业务空间进行[模型授权](/developer-guides/administration/workspace)。请参见[子业务空间的模型调用](/developer-guides/administration/workspace)。 diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-translation.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-translation.md new file mode 100644 index 0000000..f7e8ee0 --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-realtime-translation.md @@ -0,0 +1,1162 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 实时语音/音视频翻译-千问 + +> 本文介绍千问实时语音/音视频翻译模型的能力、支持的模型和接入方式。模型可结合音频与图像输入进行实时翻译,输出目标语种的文本或语音,适用于实时语音交流和视频翻译等场景。 + +> 在线体验参见[通过函数计算一键部署](/developer-guides/speech/realtime-translation#通过函数计算一键部署)。 + +## 功能特性 + +- **多语言支持**:支持 60 种语言互译,其中 29 种支持音频+文本输出、31 种仅支持文本输出,覆盖中文、英语、法语、德语、俄语、日语、韩语、西班牙语、葡萄牙语、阿拉伯语等主流语种。 +- **视觉增强**:利用视觉内容提升翻译准确性。模型通过分析画面中的口型、动作和文字,改善在嘈杂环境下或一词多义场景中的翻译效果。 +- **2.3 秒时延**:实现低至 2.3 秒的同传时延。 +- **实时说话人分离**:支持在多人交替发言时区分不同说话人及其发言内容,让听众清晰了解“谁说了什么”。 +- **无损同传**:通过语义单元预测技术,解决跨语言语序问题。实时翻译质量接近离线翻译结果。 +- **音色自然**:生成音色自然的拟人语音。模型能根据源语音内容,自适应调节语气和情感。 +- **配置热词**:通过热词提升特定词汇的翻译准确性。 +- **声音复刻**:支持复刻发言人音色用于翻译播报,让输出听起来像本人说外语。支持服务端实时复刻和使用预先复刻的固定音色。 + +## 如何使用 + +### 1. 配置连接 + +连接时通过 `model` 指定[支持的模型](#支持的模型)。使用 `qwen3.8-livetranslate-flash-realtime` 时,连接地址和完整示例参见[快速开始](#快速开始)。以下连接示例使用 `qwen3.8-livetranslate-flash-realtime`。 + +qwen3.8-livetranslate-flash-realtime 模型通过 WebSocket 协议接入,连接时需要以下配置项: + + + `qwen3.5-livetranslate-flash-realtime` 除 WebSocket 外,还支持通过 AOQ 和 WebRTC 协议接入;如果是客户端对接,且更看重稳定的延迟、弱网下的交互能力、实时双工的降噪与回声消除,可优先考虑 AOQ,协议对比与选型请参见[Realtime API 概述](/api-reference/realtime-api/overview#模型支持力度)。 + + +| **配置项** | **说明** | +| ------- | -------------------------------------------------------------------------------------------------------- | +| 调用地址 | wss\://maas.qianwenaiapi.com/api-ws/v1/realtime | +| 查询参数 | 查询参数为model,需指定为访问的模型名。示例:`?model=qwen3.8-livetranslate-flash-realtime` | +| 消息头 | 使用 Bearer Token 鉴权:Authorization: Bearer DASHSCOPE\_API\_KEY > DASHSCOPE\_API\_KEY 是您在千问AI平台上申请的API-KEY。 | + +可通过以下 Python 示例代码建立连接。 + + + ```python expandable + # pip install websocket-client + import json + import websocket + import os + + API_KEY=os.getenv("DASHSCOPE_API_KEY") + API_URL = "wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model=qwen3.8-livetranslate-flash-realtime" + + headers = [ + "Authorization: Bearer " + API_KEY + ] + + def on_open(ws): + print(f"Connected to server: {API_URL}") + def on_message(ws, message): + data = json.loads(message) + print("Received event:", json.dumps(data, indent=2)) + def on_error(ws, error): + print("Error:", error) + + ws = websocket.WebSocketApp( + API_URL, + header=headers, + on_open=on_open, + on_message=on_message, + on_error=on_error + ) + + ws.run_forever() + ``` + + +### 2. 配置语种、输出模态与音色 + + + + 通过 [session.update](/api-reference/speech-translation/livetranslate-realtime/client-events#session-update) 配置会话: + + - **目标语种**:通过 `session.translation.language` 配置,例如 `en` 表示英语。 + - **输出模态**:通过 `session.output_modalities` 配置,`["text"]` 为仅文本,`["text", "audio"]` 为文本和音频。 + - **原文识别**:通过 `conversation.item.input_audio_transcription.delta` 接收识别增量,通过 `conversation.item.input_audio_transcription.completed` 接收完整结果。 + - **音频与音色**:默认输入为 16000 Hz PCM,输出为 24000 Hz PCM,音色为 `Tina`。完整会话结构参见[服务端事件](/api-reference/speech-translation/livetranslate-realtime/server-events#session-created)。 + + + + 发送客户端事件[session.update](/api-reference/speech-translation/livetranslate-realtime/client-events#session-update): + + - **语种** + + - \*\*源语种:\*\*通过`session.input_audio_transcription.language`参数配置。 + + > 默认不填写,此时模型会自动识别源语种。 + - \*\*目标语种:\*\*通过`session.translation.language`参数配置。 + + > 默认值为`en`(英语)。 + + 取值范围参见[支持的语种](/developer-guides/speech/realtime-translation#支持的语种)。 + - **输出源语言识别结果** + + 通过`session.input_audio_transcription.model`参数配置。设置为`qwen3-asr-flash-realtime`后,服务端会在翻译的同时返回输入音频的语音识别结果(源语言原文)。 + + 启用后,服务端会返回以下事件: + + - `conversation.item.input_audio_transcription.text`:流式返回识别结果。 + - `conversation.item.input_audio_transcription.completed`:识别完成后返回最终结果。 + - `conversation.item.input_audio_transcription.failed`:识别失败时返回错误信息。 + - **输出模态** + + 通过`session.modalities`参数配置。支持设置为`["text"]`(仅输出文本)或`["text","audio"]`(输出文本与音频)。 + - **语音活动检测(VAD)与 Manual 模式** + + 通过`session.turn_detection`参数配置语音起止的检测方式: + + - **VAD 模式**(默认):`turn_detection`设为配置对象。服务端自动检测语音起止并自动触发翻译,适合客户端持续发送音频流的场景。 + - **Manual 模式**:`turn_detection`设为`null`。由客户端判断语音起止,说完一段话后主动发送`input_audio_buffer.commit`事件提交音频,适合"按下说话、松开发送"的场景。 + + 两种模式下的完整交互步骤参见本文[3. 输入音频与图片](/developer-guides/speech/realtime-translation#3-输入音频与图片)。 + - **音色** + + 通过`session.voice`参数配置。参见[支持的音色](/developer-guides/speech/realtime-translation#支持的音色)。 + - **热词** + + 通过`session.translation.corpus.phrases`参数配置。热词用于提升特定词汇的翻译准确性,以 key-value 形式指定源语言词汇与目标语言翻译的映射关系。推荐配置 1000 个以内的热词。 + + 示例:将`"人工智能"`指定翻译为`"Artificial Intelligence"`。 + - **声音复刻** + + 通过 `session.enable_voice_clone`、`session.voice_clone_options.frequency` 与 `session.voice` 参数配置。支持三种模式:使用预先复刻的音色(`frequency` 为 `never`)、服务端复刻一次(`once`)或每次复刻(`always`)。详见 [声音复刻](/developer-guides/speech/realtime-translation#声音复刻)。 + + + +### 3. 输入音频与图片 + +客户端通过 [input\_audio\_buffer.append](/api-reference/speech-translation/livetranslate-realtime/client-events#input-audio-buffer-append) 和 [input\_image\_buffer.append](/api-reference/speech-translation/livetranslate-realtime/client-events#input-image-buffer-append) 事件发送 Base64 编码的音频和图片数据。音频输入是必需的;图片输入是可选的。 + +> 图片可以来自本地文件,或从视频流中实时采集。 + +以下 VAD 与 Manual 模式配置适用于 `qwen3.5-livetranslate-flash-realtime`。`qwen3.8-livetranslate-flash-realtime` 的默认断句配置为 `audio.input.turn_detection.type = speaker_detection`,持续发送音频后由服务端生成响应。 + +模型判断一段语音"说完了"的方式,取决于[turn\_detection](/api-reference/speech-translation/livetranslate-realtime/client-events)参数配置的 VAD 模式或 Manual 模式: + +- **VAD 模式**(默认):客户端持续发送[input\_audio\_buffer.append](/api-reference/speech-translation/livetranslate-realtime/client-events#input-audio-buffer-append)事件;服务端检测到语音开始/结束时,分别返回[input\_audio\_buffer.speech\_started](/api-reference/speech-translation/livetranslate-realtime/server-events#input-audio-buffer-speech-started)、[input\_audio\_buffer.speech\_stopped](/api-reference/speech-translation/livetranslate-realtime/server-events#input-audio-buffer-speech-stopped)事件,并自动提交音频缓冲区、触发翻译。翻译响应基于流式语音同步生成,通常在语音输入过程中就已开始,无需等待语音结束。 +- **Manual 模式**:将`session.turn_detection`设为`null`。客户端发送完一段完整的语音后,主动发送[input\_audio\_buffer.commit](/api-reference/speech-translation/livetranslate-realtime/client-events#input-audio-buffer-commit)事件提交音频缓冲区;服务端返回[input\_audio\_buffer.committed](/api-reference/speech-translation/livetranslate-realtime/server-events#input-audio-buffer-committed)事件确认后,自动开始生成翻译响应,客户端无需再发送其他事件触发响应。提交前若需清空未提交的音频,可发送[input\_audio\_buffer.clear](/api-reference/speech-translation/livetranslate-realtime/client-events#input-audio-buffer-clear)事件。 + +### 4. 接收模型响应 + + + + 根据输出模态接收响应: + + - **仅文本**:累加 `response.text.delta` 的 `delta` 字段获得译文。 + - **文本和音频**:累加 `response.audio_transcript.delta` 的 `delta` 字段获得译文,并对 `response.audio.delta` 的 `delta` 字段进行 Base64 解码,获取音频分片。 + + 收到 `response.done` 表示本次响应结束。事件字段参见[服务端事件](/api-reference/speech-translation/livetranslate-realtime/server-events#response-text-delta)。 + + + + 翻译响应基于流式语音同步生成,通常无需等待语音结束(参见上一节 VAD/Manual 模式说明)。模型的响应格式取决于配置的输出模态。 + + - **仅输出文本** + + 服务端通过[response.text.text](/api-reference/speech-translation/livetranslate-realtime/server-events#response-text-text)事件流式返回增量翻译文本(含已确认文本和待确认的预测文本);翻译完成后,通过[response.text.done](/api-reference/speech-translation/livetranslate-realtime/server-events#response-text-done)事件返回完整的翻译文本。 + - **输出文本+音频** + + - **文本** + + 通过[response.audio\_transcript.text](/api-reference/speech-translation/livetranslate-realtime/server-events#response-audio-transcript-text)事件流式返回增量翻译文本;翻译完成后,通过[response.audio\_transcript.done](/api-reference/speech-translation/livetranslate-realtime/server-events#response-audio-transcript-done)事件返回完整的翻译文本。 + - **音频** + + 通过[response.audio.delta](/api-reference/speech-translation/livetranslate-realtime/server-events#response-audio-delta)事件返回 Base64 编码的增量音频数据。 + + + `qwen3.5-livetranslate-flash-realtime` 使用 `response.text.text` 事件返回增量文本,与全双工语音对话(Omni)模型的 `response.text.delta` 事件不同,两者字段结构和语义有差异,请勿混用。 + + + + +### 5. 结束会话 + +音频发送完毕后,客户端必须发送 [session.finish](/api-reference/speech-translation/livetranslate-realtime/client-events#session-finish) 事件通知服务端,然后等待服务端返回 `session.finished` 事件后再关闭 WebSocket 连接。 + + + 如果不发送 `session.finish`,服务端无法得知音频输入已完成,会导致最后一段语音的识别和翻译结果丢失,连接也可能长时间处于等待状态。请务必在关闭连接前发送该事件。 + + +## 支持的模型 + +### 推荐模型 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+ **模型名称** + + **版本** + + **上下文长度** + + **最大输入** + + **最大输出** +
+ **(Token数)** +
+ **qwen3.8-livetranslate-flash-realtime** + + 稳定版 + + 53248 + + 49152 + + 4096 +
+ **qwen3.5-livetranslate-flash-realtime** + + > 当前能力等同 qwen3.5-livetranslate-flash-realtime-2026-05-19 + + 稳定版 + + 53248 + + 49152 + + 4096 +
+ qwen3.5-livetranslate-flash-realtime-2026-05-19 + + 快照版 +
+ +### 旧版模型 + +> 以下模型仍在服务中,但不再作为首选推荐。新场景建议使用上方的新一代模型,以获得更优的翻译质量与性价比。 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+ **模型名称** + + **版本** + + **上下文长度** + + **最大输入** + + **最大输出** +
+ **(Token数)** +
+ **qwen3-livetranslate-flash-realtime** + + > 当前能力等同 qwen3-livetranslate-flash-realtime-2025-09-22 + + 稳定版 + + 53248 + + 49152 + + 4096 +
+ qwen3-livetranslate-flash-realtime-2025-09-22 + + 快照版 +
+ +## 快速开始 + + + + ### 翻译本地音频 + + 安装 `websocket-client`(`pip install websocket-client`),并设置以下环境变量: + + - `DASHSCOPE_API_KEY`:API Key。 + - `INPUT_PCM_FILE`:待翻译音频的本地路径。此示例使用单声道、16 位、16000 Hz 的无文件头 PCM 音频。 + + 示例将音频翻译为英语,打印译文,并将 24000 Hz、单声道、16 位 PCM 输出保存为 `translation.pcm`。连接后配置会话,发送音频,最后通过 `session.finish` 请求结束,收到 `session.finished` 后关闭连接。 + + ```python + import base64 + import json + import os + import threading + import time + + import websocket + + model = "qwen3.8-livetranslate-flash-realtime" + api_key = os.environ["DASHSCOPE_API_KEY"] + audio_file = os.environ["INPUT_PCM_FILE"] + url = f"wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model={model}" + ws = websocket.create_connection( + url, header={"Authorization": f"Bearer {api_key}"}, timeout=30 + ) + + def receive(): + event = json.loads(ws.recv()) + if event.get("type") == "error" or event.get("code"): + raise RuntimeError(event) + return event + + send_errors = [] + + def send_audio(): + try: + with open(audio_file, "rb") as source: + while chunk := source.read(3200): + ws.send(json.dumps({ + "type": "input_audio_buffer.append", + "audio": base64.b64encode(chunk).decode("ascii"), + })) + time.sleep(0.1) + ws.send(json.dumps({"type": "session.finish"})) + except Exception as error: + send_errors.append(error) + + try: + receive() + ws.send(json.dumps({ + "type": "session.update", + "session": { + "output_modalities": ["text", "audio"], + "translation": {"language": "en"}, + }, + })) + while receive().get("type") != "session.updated": + pass + sender = threading.Thread(target=send_audio, daemon=True) + sender.start() + with open("translation.pcm", "wb") as output: + while True: + event = receive() + event_type = event.get("type") + if event_type in ("response.text.delta", "response.audio_transcript.delta"): + print(event["delta"], end="", flush=True) + elif event_type == "response.audio.delta": + output.write(base64.b64decode(event["delta"])) + elif event_type == "session.finished": + break + sender.join(timeout=5) + if send_errors: + raise send_errors[0] + print() + finally: + ws.close() + ``` + + + + 1. **准备运行环境** + + 您的 Python 版本需要不低于 3.10。 + + 首先安装 pyaudio。 + + + ```bash macOS + brew install portaudio && pip install pyaudio + ``` + + ```bash Debian/Ubuntu + sudo apt-get install python3-pyaudio + + 或者 + + pip install pyaudio + ``` + + ```bash CentOS + sudo yum install -y portaudio portaudio-devel && pip install pyaudio + ``` + + ```powershell Windows + pip install pyaudio + ``` + + + 安装完成后,通过 pip 安装 websocket 相关的依赖: + + ```bash + pip install websocket-client==1.8.0 websockets + ``` + + 2. **创建客户端** + + 在本地新建一个 Python 文件,命名为`livetranslate_client.py`,并将以下代码复制进文件中: + + + ```python expandable + import os + import time + import base64 + import asyncio + import json + import websockets + import pyaudio + import queue + import threading + import traceback + + class LiveTranslateClient: + def __init__(self, api_key: str, target_language: str = "en", *, audio_enabled: bool = True): + if not api_key: + raise ValueError("API key cannot be empty.") + + self.api_key = api_key + self.target_language = target_language + self.audio_enabled = audio_enabled + self.ws = None + self.api_url = "wss://maas.qianwenaiapi.com/api-ws/v1/realtime?model=qwen3.5-livetranslate-flash-realtime" + + # 音频输入配置 (来自麦克风) + self.input_rate = 16000 + self.input_chunk = 1600 + self.input_format = pyaudio.paInt16 + self.input_channels = 1 + + # 音频输出配置 (用于播放) + self.output_rate = 24000 + self.output_chunk = 2400 + self.output_format = pyaudio.paInt16 + self.output_channels = 1 + + # 状态管理 + self.is_connected = False + self.audio_player_thread = None + self.audio_playback_queue = queue.Queue() + self.pyaudio_instance = pyaudio.PyAudio() + self.session_finished_event = asyncio.Event() + + async def connect(self): + """建立到翻译服务的 WebSocket 连接。""" + headers = {"Authorization": f"Bearer {self.api_key}"} + try: + self.ws = await websockets.connect(self.api_url, additional_headers=headers) + self.is_connected = True + print(f"成功连接到服务端: {self.api_url}") + await self.configure_session() + except Exception as e: + print(f"连接失败: {e}") + self.is_connected = False + raise + + async def configure_session(self): + """配置翻译会话,设置目标语言、声音等。""" + config = { + "event_id": f"event_{int(time.time() * 1000)}", + "type": "session.update", + "session": { + # 'modalities' 控制输出类型。 + # ["text", "audio"]: 同时返回翻译文本和合成音频(推荐)。 + # ["text"]: 仅返回翻译文本。 + "modalities": ["text", "audio"] if self.audio_enabled else ["text"], + "input_audio_format": "pcm", + "output_audio_format": "pcm", + # 'input_audio_transcription' 配置源语言识别。 + # 设置 'model' 为 'qwen3-asr-flash-realtime' 可同时输出源语言识别结果。 + # "input_audio_transcription": { + # "model": "qwen3-asr-flash-realtime", + # "language": "zh" # 源语言,默认 'en' + # }, + "translation": { + "language": self.target_language, + # 'corpus' 配置热词,用于提升特定词汇的翻译准确性。 + # "corpus": { + # "phrases": { + # "人工智能": "Artificial Intelligence", + # "机器学习": "Machine Learning" + # } + # } + } + } + } + print(f"发送会话配置: {json.dumps(config, indent=2, ensure_ascii=False)}") + await self.ws.send(json.dumps(config)) + + async def send_audio_chunk(self, audio_data: bytes): + """将音频数据块编码并发送到服务端。""" + if not self.is_connected: + return + + event = { + "event_id": f"event_{int(time.time() * 1000)}", + "type": "input_audio_buffer.append", + "audio": base64.b64encode(audio_data).decode() + } + await self.ws.send(json.dumps(event)) + + async def send_image_frame(self, image_bytes: bytes, *, event_id: str | None = None): + #将图像数据发送到服务端 + if not self.is_connected: + return + + if not image_bytes: + raise ValueError("image_bytes 不能为空") + + # 编码为 Base64 + image_b64 = base64.b64encode(image_bytes).decode() + + event = { + "event_id": event_id or f"event_{int(time.time() * 1000)}", + "type": "input_image_buffer.append", + "image": image_b64, + } + + await self.ws.send(json.dumps(event)) + + def _audio_player_task(self): + stream = self.pyaudio_instance.open( + format=self.output_format, + channels=self.output_channels, + rate=self.output_rate, + output=True, + frames_per_buffer=self.output_chunk, + ) + try: + while self.is_connected or not self.audio_playback_queue.empty(): + try: + audio_chunk = self.audio_playback_queue.get(timeout=0.1) + if audio_chunk is None: # 结束信号 + break + stream.write(audio_chunk) + self.audio_playback_queue.task_done() + except queue.Empty: + continue + finally: + stream.stop_stream() + stream.close() + + def start_audio_player(self): + """启动音频播放线程(仅当启用音频输出时)。""" + if not self.audio_enabled: + return + if self.audio_player_thread is None or not self.audio_player_thread.is_alive(): + self.audio_player_thread = threading.Thread(target=self._audio_player_task, daemon=True) + self.audio_player_thread.start() + + async def handle_server_messages(self, on_text_received): + """循环处理来自服务端的消息。""" + try: + async for message in self.ws: + event = json.loads(message) + event_type = event.get("type") + if event_type == "response.audio.delta" and self.audio_enabled: + audio_b64 = event.get("delta", "") + if audio_b64: + audio_data = base64.b64decode(audio_b64) + self.audio_playback_queue.put(audio_data) + + elif event_type == "response.done": + print("\n[INFO] 一轮响应完成。") + usage = event.get("response", {}).get("usage", {}) + if usage: + print(f"[INFO] Token 使用情况: {json.dumps(usage, indent=2, ensure_ascii=False)}") + elif event_type == "session.finished": + print("[INFO] 会话已结束。") + self.session_finished_event.set() + # 处理源语言识别结果(需启用 input_audio_transcription.model) + # elif event_type == "conversation.item.input_audio_transcription.text": + # stash = event.get("stash", "") # 待确认的识别文本 + # print(f"[识别中] {stash}") + # elif event_type == "conversation.item.input_audio_transcription.completed": + # transcript = event.get("transcript", "") # 完整识别结果 + # print(f"[源语言] {transcript}") + elif event_type == "response.text.text": + # 仅文本模态下的流式翻译文本 + text = event.get("text", "") + stash = event.get("stash", "") + print(f"\r[翻译中] {text}{stash}", end="", flush=True) + elif event_type == "response.audio_transcript.done": + print("\n[INFO] 翻译文本完成。") + text = event.get("transcript", "") + if text: + print(f"[INFO] 翻译文本: {text}") + elif event_type == "response.text.done": + print("\n[INFO] 翻译文本完成。") + text = event.get("text", "") + if text: + print(f"[INFO] 翻译文本: {text}") + + except websockets.exceptions.ConnectionClosed as e: + print(f"[WARNING] 连接已关闭: {e}") + self.is_connected = False + except Exception as e: + print(f"[ERROR] 消息处理时发生未知错误: {e}") + traceback.print_exc() + self.is_connected = False + + async def start_microphone_streaming(self): + """从麦克风捕获音频并流式传输到服务端。""" + stream = self.pyaudio_instance.open( + format=self.input_format, + channels=self.input_channels, + rate=self.input_rate, + input=True, + frames_per_buffer=self.input_chunk + ) + print("麦克风已启动,请开始说话...") + try: + while self.is_connected: + audio_chunk = await asyncio.get_event_loop().run_in_executor( + None, stream.read, self.input_chunk + ) + await self.send_audio_chunk(audio_chunk) + finally: + stream.stop_stream() + stream.close() + + async def close(self): + """优雅地关闭连接和资源。""" + # 发送 session.finish,确保服务端完成最后的语音翻译 + if self.is_connected and self.ws: + finish_event = { + "event_id": f"event_{int(time.time() * 1000)}", + "type": "session.finish", + } + await self.ws.send(json.dumps(finish_event)) + print("已发送 session.finish,等待服务端完成处理...") + try: + await asyncio.wait_for(self.session_finished_event.wait(), timeout=15) + print("服务端已完成处理。") + except asyncio.TimeoutError: + print("等待 session.finished 超时。") + + self.is_connected = False + if self.ws: + await self.ws.close() + print("WebSocket 连接已关闭。") + + if self.audio_player_thread: + self.audio_playback_queue.put(None) # 发送结束信号 + self.audio_player_thread.join(timeout=1) + print("音频播放线程已停止。") + + self.pyaudio_instance.terminate() + print("PyAudio 实例已释放。") + ``` + + + 3. **与模型互动** + + 在`livetranslate_client.py`的同级目录下新建另一个 Python 文件,命名为`main.py`,并将以下代码复制进文件中: + + + ```python expandable + import os + import asyncio + from livetranslate_client import LiveTranslateClient + + def print_banner(): + print("=" * 60) + print(" 基于千问 qwen3.5-livetranslate-flash-realtime") + print("=" * 60 + "\n") + + def get_user_config(): + """获取用户配置""" + print("请选择模式:") + print("1. 语音+文本 [默认] | 2. 仅文本") + mode_choice = input("请输入选项 (直接回车选择语音+文本): ").strip() + audio_enabled = (mode_choice != "2") + + if audio_enabled: + lang_map = { + "1": "en", "2": "zh", "3": "ru", "4": "fr", "5": "de", "6": "pt", + "7": "es", "8": "it", "9": "ko", "10": "ja" + } + print("请选择翻译目标语言 (音频+文本 模式):") + print("1. 英语 | 2. 中文 | 3. 俄语 | 4. 法语 | 5. 德语 | 6. 葡萄牙语 | 7. 西班牙语 | 8. 意大利语 | 9. 韩语 | 10. 日语") + else: + lang_map = { + "1": "en", "2": "zh", "3": "ru", "4": "fr", "5": "de", "6": "pt", "7": "es", "8": "it", + "9": "id", "10": "ko", "11": "ja", "12": "vi", "13": "th", "14": "ar", + "15": "yue", "16": "hi", "17": "el", "18": "tr" + } + print("请选择翻译目标语言 (仅文本 模式):") + print("1. 英语 | 2. 中文 | 3. 俄语 | 4. 法语 | 5. 德语 | 6. 葡萄牙语 | 7. 西班牙语 | 8. 意大利语 | 9. 印尼语 | 10. 韩语 | 11. 日语 | 12. 越南语 | 13. 泰语 | 14. 阿拉伯语 | 15. 粤语 | 16. 印地语 | 17. 希腊语 | 18. 土耳其语") + + choice = input("请输入选项 (默认取第一个): ").strip() + target_language = lang_map.get(choice, next(iter(lang_map.values()))) + + return target_language, audio_enabled + + async def main(): + """主程序入口""" + print_banner() + + api_key = os.environ.get("DASHSCOPE_API_KEY") + if not api_key: + print("[ERROR] 请设置环境变量 DASHSCOPE_API_KEY") + print(" 例如: export DASHSCOPE_API_KEY='your_api_key_here'") + return + + target_language, audio_enabled = get_user_config() + print("\n配置完成:") + print(f" - 目标语言: {target_language}") + if not audio_enabled: + print(" - 输出模式: 仅文本") + + client = LiveTranslateClient(api_key=api_key, target_language=target_language, audio_enabled=audio_enabled) + + # 定义回调函数 + def on_translation_text(text): + print(text, end="", flush=True) + + try: + print("正在连接到翻译服务...") + await client.connect() + + # 根据模式启动音频播放 + client.start_audio_player() + + print("\n" + "-" * 60) + print("连接成功!请对着麦克风说话。") + print("程序将实时翻译您的语音并播放结果。按 Ctrl+C 退出。") + print("-" * 60 + "\n") + + # 并发运行消息处理和麦克风录音 + message_handler = asyncio.create_task(client.handle_server_messages(on_translation_text)) + tasks = [message_handler] + # 无论是否启用音频输出,都需要从麦克风捕获音频进行翻译 + microphone_streamer = asyncio.create_task(client.start_microphone_streaming()) + tasks.append(microphone_streamer) + + await asyncio.gather(*tasks) + + except KeyboardInterrupt: + print("\n\n用户中断,正在退出...") + except Exception as e: + print(f"\n发生严重错误: {e}") + finally: + print("\n正在清理资源...") + await client.close() + print("程序已退出。") + + if __name__ == "__main__": + asyncio.run(main()) + ``` + + + 运行`main.py`,通过麦克风说出要翻译的句子,模型会实时返回翻译完成的音频与文本。系统会检测您的音频起始位置并自动发送到服务端,无需手动发送。 + + + +## 声音复刻 + +模型支持发言人声音复刻功能,支持使用预先复刻的固定音色,也支持由服务端实时复刻,让翻译播报听起来像本人说外语。适用于跨语言演讲、个人主播、视频翻译等需要保留个人音色的场景。 + +在 [session.update](/api-reference/speech-translation/livetranslate-realtime/client-events#session-update) 中设置以下参数启用: + +- `session.enable_voice_clone`:设置为 `true`,启用声音复刻。 +- `session.voice_clone_options.frequency`:控制声音复刻时机,取值如下: + + - `never`:不在服务端复刻,使用用户预先复刻好的音色。此时 `session.voice` 需设置为用户自己的复刻音色 ID。 + - `once`:服务端在会话开始时基于输入音频复刻一次音色,后续翻译输出复用该音色。适合单人演讲场景。此时 `session.voice` 需设置为 `default`。 + - `always`:服务端在每次生成翻译音频前实时复刻,音色跟随输入动态变化。适合双人及以上对话场景。此时 `session.voice` 需设置为 `default`。 +- `session.voice`:指定输出音色,取值取决于 `frequency` 的设置。 + + - 设置为 `default`:搭配 `frequency` 为 `once` 或 `always` 使用,由服务端复刻输入音频的音色,复刻完成前使用默认音色过渡。 + - 设置为用户复刻的音色 ID(如 `qwen-translate-vc-xxx-yyy-zzz`):搭配 `frequency` 为 `never` 使用。需提前通过[声音复刻API](/developer-guides/speech/voice-cloning)准备音色,`targetModel` 需指定为实际使用的翻译模型。 + +> 当 `frequency` 为 `once` 或 `always` 时,`voice` 必须设置为 `default`,不可设置为其他预设音色,否则服务端会返回错误。 + +### 声音复刻配置示例 + +**使用预先复刻的音色**(音质稳定,推荐需要固定音色的场景): + +```json +{ + "type": "session.update", + "session": { + "modalities": ["text","audio"], + "voice": "qwen-translate-vc-xxx-yyy-zzz", + "translation": { + "language": "en" + }, + "enable_voice_clone": true, + "voice_clone_options": { + "frequency": "never" + } + } +} +``` + +**服务端复刻一次**(适合单人演讲): + +```json +{ + "type": "session.update", + "session": { + "modalities": ["text","audio"], + "voice": "default", + "translation": { + "language": "en" + }, + "enable_voice_clone": true, + "voice_clone_options": { + "frequency": "once" + } + } +} +``` + +**服务端每次复刻**(适合多人对话): + +```json +{ + "type": "session.update", + "session": { + "modalities": ["text","audio"], + "voice": "default", + "translation": { + "language": "en" + }, + "enable_voice_clone": true, + "voice_clone_options": { + "frequency": "always" + } + } +} +``` + +## 利用图像提升翻译准确率 + +qwen3.5-livetranslate-flash-realtime 模型可以接收图像输入,辅助音频翻译,适用于同音异义、低频专有名词识别场景。建议每秒发送不超过2张图片。 + +将以下示例图片下载到本地:[口罩.png](https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250923/wbclir/%E5%8F%A3%E7%BD%A9.png)[面具.png](https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250923/ohelwv/%E9%9D%A2%E5%85%B7.png) + +将以下代码下载到`livetranslate_client.py`同级目录并运行,向麦克风说`"What is mask?"`,在输入口罩图片时,模型会翻译为“什么是口罩?”;输入面具图片时,模型会翻译为“什么是面具?” + +```python expandable +import os +import time +import json +import asyncio +import contextlib +import functools + +from livetranslate_client import LiveTranslateClient + +IMAGE_PATH = "口罩.png" +# IMAGE_PATH = "面具.png" + +def print_banner(): + print("=" * 60) + print(" 基于千问 qwen3.5-livetranslate-flash-realtime —— 单轮交互示例 (mask)") + print("=" * 60 + "\n") + +async def stream_microphone_once(client: LiveTranslateClient, image_bytes: bytes): + pa = client.pyaudio_instance + stream = pa.open( + format=client.input_format, + channels=client.input_channels, + rate=client.input_rate, + input=True, + frames_per_buffer=client.input_chunk, + ) + print(f"[INFO] 开始录音,请讲话……") + loop = asyncio.get_event_loop() + last_img_time = 0.0 + frame_interval = 0.5 # 2 fps + try: + while client.is_connected: + data = await loop.run_in_executor(None, stream.read, client.input_chunk) + await client.send_audio_chunk(data) + + # 每 0.5 秒追加一帧图片 + now = time.time() + if now - last_img_time >= frame_interval: + await client.send_image_frame(image_bytes) + last_img_time = now + finally: + stream.stop_stream() + stream.close() + +async def main(): + print_banner() + api_key = os.environ.get("DASHSCOPE_API_KEY") + if not api_key: + print("[ERROR] 请先在环境变量 DASHSCOPE_API_KEY 中配置 API KEY") + return + + client = LiveTranslateClient(api_key=api_key, target_language="zh", audio_enabled=True) + + def on_text(text: str): + print(text, end="", flush=True) + + try: + await client.connect() + client.start_audio_player() + message_task = asyncio.create_task(client.handle_server_messages(on_text)) + with open(IMAGE_PATH, "rb") as f: + img_bytes = f.read() + await stream_microphone_once(client, img_bytes) + await asyncio.sleep(15) + finally: + await client.close() + if not message_task.done(): + message_task.cancel() + with contextlib.suppress(asyncio.CancelledError): + await message_task + +if __name__ == "__main__": + asyncio.run(main()) +``` + +## 通过函数计算一键部署 + +控制台暂不支持体验。可通过以下方式一键部署: + +1. 打开我们写好的函数计算模板,填入 API Key, 单击**创建并部署默认环境**即可在线体验。 +2. 等待约一分钟,在 **环境详情 > 环境信息** 中获取访问域名,**将访问域名的**`http`**改成**`https`(例如[https://qwen-livetranslate-flash-realtime.fcv3.xxx.cn-hangzhou.fc.devsapp.net/),通过该链接与模型交互。](https://qwen-livetranslate-flash-realtime.fcv3.xxx.cn-hangzhou.fc.devsapp.net/),通过该链接与模型交互。) + + + 此链接使用自签名证书,仅用于临时测试。首次访问时,浏览器会显示安全警告,这是预期行为,**请勿在生产环境使用**。如需继续,请按浏览器提示操作(如点击“高级” → “继续前往(不安全)”)。 + + +> 如需开通访问控制权限,请跟随页面指引操作。 + +> 通过**资源信息**-**函数资源**查看项目源代码。 + +> 函数计算与[千问AI平台](/resources/free-quota)均为新用户提供免费额度,可以覆盖简单调试所需成本,额度耗尽后按量计费。只有在访问的情况下会产生费用。 + +## 交互流程 + + + + 默认通过服务端断句生成响应,以下列出主要交互环节。事件字段详见[服务端事件](/api-reference/speech-translation/livetranslate-realtime/server-events)。 + +| **阶段** | **客户端操作** | **服务端事件** | +| ------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | +| 创建和配置会话 | 建立连接,发送 `session.update` | `session.created`、`session.updated` | +| 输入音频 | `input_audio_buffer.append` | 原文通过 `conversation.item.input_audio_transcription.delta` 增量返回,完成后返回 `conversation.item.input_audio_transcription.completed`。 | +| 接收译文与音频 | 持续接收服务端事件 | 译文通过 `response.text.delta`(仅文本)或 `response.audio_transcript.delta`(文本和音频)返回;音频通过 `response.audio.delta` 返回。`response.done` 表示一次响应完成。 | +| 结束会话 | `session.finish` | 收到 `session.finished` 后关闭连接。 | + + + + 实时语音翻译的交互流程遵循标准的 WebSocket 事件驱动模型。语音起止的判断方式取决于 VAD 模式或 Manual 模式(参见[3. 输入音频与图片](/developer-guides/speech/realtime-translation#3-输入音频与图片)),下表以 VAD 模式(默认)为主线,并标注了 Manual 模式下不同的服务端事件。 + +| **生命周期** | **客户端事件** | **服务端事件** | +| -------- | -------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 会话初始化 | session.update > 会话配置 | session.created > 会话已创建 session.updated > 会话配置已更新 | +| 用户音频输入 | input\_audio\_buffer.append > 添加音频到缓冲区 input\_image\_buffer.append > 添加图片到缓冲区 input\_audio\_buffer.commit > (仅 Manual 模式)提交音频缓冲区 | **VAD 模式**: input\_audio\_buffer.speech\_started > 检测到语音开始 input\_audio\_buffer.speech\_stopped > 检测到语音结束,服务端自动提交音频缓冲区 **Manual 模式**: input\_audio\_buffer.committed > 客户端发送 input\_audio\_buffer.commit 后返回,确认音频缓冲区已提交 | +| 服务端音频输出 | 无 | response.created > 服务端开始生成响应 response.output\_item.added > 响应时有新的输出内容 conversation.item.created > 对话中创建新的消息项 response.content\_part.added > 新的输出内容添加到assistant message response.text.text > 仅文本模态下增量生成的翻译文本 response.audio\_transcript.text > 音频+文本模态下增量生成的转录文字 response.audio.delta > 模型增量生成的音频 response.text.done > 仅文本模态下翻译文本完成 response.audio\_transcript.done > 音频+文本模态下文本转录完成 response.audio.done > 音频生成完成 response.content\_part.done > Assistant message 的文本或音频内容流式输出完成 response.output\_item.done > Assistant message 的整个输出项流式传输完成 response.done > 响应完成 | +| 会话结束 | session.finish > 通知服务端音频发送完毕 | session.finished > 服务端完成处理,会话结束 | + + + + + 音频发送结束后,必须发送 `session.finish` 事件并等待 `session.finished` 响应后再断开连接。如果直接关闭 WebSocket 而不发送 `session.finish`,服务端无法得知音频输入已结束,将导致最后一段语音的识别和翻译结果丢失。 + + +## API 参考 + +通过 AOQ 接入的流程和示例,请参见[AOQ 接入](/api-reference/realtime-api/aoq-access)。支持的模型及版本请参见[模型与协议支持范围](/api-reference/realtime-api/overview#模型支持力度)。 + +- [实时音视频翻译(Qwen-Livetranslate-Realtime)](/api-reference/speech-translation/livetranslate-realtime/model-access)。 +- [Realtime API 概述](/api-reference/realtime-api/overview)(WebRTC 协议说明) + +## 计费说明 + +- **Qwen3.8-LiveTranslate-Flash-Realtime、Qwen3.5-LiveTranslate-Flash-Realtime** + + - **音频**:输入每秒音频消耗 7 Token,输出每秒音频消耗 12.5 Token。 + - **图片**:每输入 32\*32 像素消耗 0.5 Token。 +- **Qwen3-LiveTranslate-Flash-Realtime** + + - **音频**:输入或输出每秒音频均消耗 12.5 Token。 + - **图片**:每输入 28\*28 像素消耗 0.5 Token。 + - **文本**:启用源语言语音识别功能后,服务除返回翻译结果外,还会返回输入音频的语音识别文本(即源语言原文),该识别文本将按输出文本的 Token 标准计费。 + +各模型的 Token 单价请参见[模型调用计费](/developer-guides/getting-started/pricing)。 + +## 限流说明 + +模型的限流规则请参见[限流](/developer-guides/administration/rate-limits)。 + +## 支持的语种 + +下表中的语种代码可用于指定源语种与目标语种。 + +> 部分目标语种仅支持输出文本,不支持输出音频。老模型 qwen3-livetranslate-flash-realtime 仅支持以下 18 种语种:en、zh、ru、fr、de、pt、es、it、id、ko、ja、vi、th、ar、yue、hi、el、tr。 + +| **语种代码** | **语种** | **支持的输出模态** | +| -------- | ------- | ----------- | +| zh | 中文 | 音频+文本 | +| en | 英语 | 音频+文本 | +| ar | 阿拉伯语 | 音频+文本 | +| de | 德语 | 音频+文本 | +| fr | 法语 | 音频+文本 | +| es | 西班牙语 | 音频+文本 | +| pt | 葡萄牙语 | 音频+文本 | +| id | 印度尼西亚语 | 音频+文本 | +| it | 意大利语 | 音频+文本 | +| ko | 韩语 | 音频+文本 | +| ru | 俄语 | 音频+文本 | +| th | 泰语 | 音频+文本 | +| vi | 越南语 | 音频+文本 | +| ja | 日语 | 音频+文本 | +| tr | 土耳其语 | 音频+文本 | +| hi | 印地语 | 音频+文本 | +| ms | 马来语 | 音频+文本 | +| nl | 荷兰语 | 音频+文本 | +| ur | 乌尔都语 | 音频+文本 | +| nb | 挪威语 | 音频+文本 | +| sv | 瑞典语 | 音频+文本 | +| da | 丹麦语 | 音频+文本 | +| he | 希伯来语 | 音频+文本 | +| fi | 芬兰语 | 音频+文本 | +| pl | 波兰语 | 音频+文本 | +| is | 冰岛语 | 音频+文本 | +| cs | 捷克语 | 音频+文本 | +| fil | 菲律宾语 | 音频+文本 | +| fa | 波斯语 | 音频+文本 | +| yue | 粤语 | 文本 | +| el | 希腊语 | 文本 | +| af | 南非荷兰语 | 文本 | +| ast | 阿斯图里亚斯语 | 文本 | +| be | 白俄罗斯语 | 文本 | +| bg | 保加利亚语 | 文本 | +| bn | 孟加拉语 | 文本 | +| bs | 波斯尼亚语 | 文本 | +| ca | 加泰罗尼亚语 | 文本 | +| ceb | 宿务语 | 文本 | +| et | 爱沙尼亚语 | 文本 | +| gl | 加利西亚语 | 文本 | +| gu | 古吉拉特语 | 文本 | +| hr | 克罗地亚语 | 文本 | +| hu | 匈牙利语 | 文本 | +| jv | 爪哇语 | 文本 | +| kk | 哈萨克语 | 文本 | +| kn | 卡纳达语 | 文本 | +| ky | 柯尔克孜语 | 文本 | +| lv | 拉脱维亚语 | 文本 | +| mk | 马其顿语 | 文本 | +| ml | 马拉雅拉姆语 | 文本 | +| mr | 马拉地语 | 文本 | +| pa | 旁遮普语 | 文本 | +| ro | 罗马尼亚语 | 文本 | +| sk | 斯洛伐克语 | 文本 | +| sl | 斯洛文尼亚语 | 文本 | +| sw | 斯瓦希里语 | 文本 | +| tg | 塔吉克语 | 文本 | +| az | 阿塞拜疆语 | 文本 | +| uk | 乌克兰语 | 文本 | + +## 支持的音色 + +实时翻译支持的音色与`voice`参数取值参见[音色列表](/developer-guides/speech/omni-voice-list)。 diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-s2s-models.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-s2s-models.md new file mode 100644 index 0000000..e7dec4a --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-s2s-models.md @@ -0,0 +1,301 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 语音到语音模型 + +> 为'语音输入 → 语音输出'场景(语音对话、语音翻译、同声传译等)选择模型。 + + + 本文档面向"语音 → 语音"场景。视觉理解、音视频分析和内容审核等能力见[全模态模型](/developer-guides/speech/omni-models);其中视频分析、内容标注等需要推理并输出文本的场景,可参考 [Qwen3.8-Omni-Flash](/developer-guides/speech/multimodal-speech)。 + + +## 从闭源模型迁移到千问AI平台? + +如果你正在使用 OpenAI Realtime 或 Gemini Live,可参考下表选择千问AI平台对位模型。 + +| | 闭源模型代表 | 千问AI平台推荐 | +| --------- | ----------------------------------- | -------------------------------------- | +| 高能力实时对话 | OpenAI GPT Realtime、Gemini 3.1 Live | `qwen-audio-3.1-realtime-plus` | +| 成本敏感对话 | OpenAI gpt-4o-mini Realtime | `qwen-audio-3.0-realtime-flash` | +| 实时翻译 / 同传 | Gemini 3.1 Live | `qwen3.8-livetranslate-flash-realtime` | + +## S2S 与 Pipeline 对比 + +构建语音应用有两种方式: + +| | S2S | Pipeline(ASR + LLM + TTS) | +| ---- | ----------------------- | ------------------------- | +| 延迟 | 低 — 单模型,流式输出 | 较高 — 需经过 3 个串行环节 | +| 音频理解 | 端到端 — 能感知语气、情感并做出相应回应 | 先转文字再处理 — 音频细节丢失 | +| 语音定制 | 通过 system prompt 选择预设音色 | 支持声音克隆、声音设计(CosyVoice) | + +- **选择 S2S**:适用于交互式对话、低延迟、需要感知音频情绪的场景。请继续阅读本页。 +- **选择 Pipeline**:适用于需要自定义音色,或希望为每个环节分别选择最佳 ASR、LLM 和 TTS 的场景。 + +本文档继续介绍 S2S 单模型路线(Omni、Livetranslate)。如选择 Pipeline 路线,分别在以下文档中挑选三个组件: + +- **ASR(语音识别)**:[语音识别](/developer-guides/speech/speech-to-text-models) +- **LLM(大语言模型)**:[文本生成](/developer-guides/getting-started/text-generation-models) +- **TTS(语音合成)**:[语音合成](/developer-guides/speech/tts-models) + +## 实时还是文件? + +- **实时(WebSocket)**— 适用于实时语音交互场景:语音助手、呼叫中心、同声传译。音频流式输入,语音流式输出。 + +- **文件(HTTP)**— 适用于可以牺牲延迟换取更好效果的场景:视频配音、播客翻译、离线内容处理。文件模式下还支持 Function Calling、联网搜索、思考模式、视频上下文等附带能力(详见下方"S2S 单模型的附带能力")。 + +## 按场景选模型(S2S 单模型路线) + +以下场景均针对 S2S 单模型路线。Pipeline 路线请按上述链接分别在 ASR / LLM / TTS 文档中选型。 + +| 场景 | 推荐模型 | API | +| --------------------------------------- | -------------------------------------- | -------------------- | +| 实时音视频对话 | `qwen3.8-omni-flash-realtime` | WebSocket、WebRTC、AOQ | +| 语音助手 / 客服对话 | `qwen-audio-3.1-realtime-plus` | WebSocket | +| 成本敏感的对话 | `qwen-audio-3.0-realtime-flash` | WebSocket | +| 同声传译 / 直播翻译 | `qwen3.8-livetranslate-flash-realtime` | WebSocket | +| 视频配音 / 播客翻译 | `qwen3-livetranslate-flash` | Chat Completions | +| 语义 VAD 语音助手 / 智能客服(支持 Function Calling) | `qwen-audio-3.1-realtime-plus` | WebSocket | + +## S2S 单模型的附带能力 + +以下介绍语音交互中的工具调用、联网搜索及文本推理能力。 + +### Function calling + +根据音视频内容查询知识库、查询日程或触发工作流,可使用 Qwen3.8-Omni-Flash-Realtime(WebSocket / WebRTC / AOQ)、Qwen-Audio Realtime(WebSocket)。 + +### 联网搜索 + +需要检索实时信息并生成语音回复时,实时对话可使用 Qwen3.8-Omni-Flash-Realtime 或 Qwen3.5-Omni-Realtime;文件调用可使用 Qwen3.5-Omni(HTTP,Plus 和 Flash 系列)。模型自主决定是否搜索。Qwen-Audio Realtime 3.0 Plus/Flash 和 3.1 Plus 也支持联网搜索,通过 `enable_search` 开启,不能与 Function Calling 同时启用。 + + + Qwen3-Omni-Flash 和 Livetranslate 模型不支持此功能。 + + +### 思考模式 + +当回答质量比延迟更重要时,使用 Qwen3 Omni(HTTP 模式)。模型在回复前会逐步推理,适用于视频分析、批量打标等场景。Qwen-Audio Realtime 不支持此功能。 + + + 思考模式下不支持生成语音。 + + +## 翻译 + +以下模型系列均支持语音翻译: + +- **Qwen3.8-Livetranslate** — 支持 60 种源语言和 29 种语言的语音输出,支持音频与图像输入。详见[模型信息](/developer-guides/speech/realtime-translation)。 +- **Qwen3.5-Livetranslate** — 支持 60 种语言互译,其中 29 种支持音频+文本输出、31 种仅支持文本输出,覆盖中文、英语、法语、德语、俄语、日语、韩语、西班牙语、葡萄牙语、阿拉伯语等主流语种。 +- **Qwen3-Livetranslate** — 支持 18 种语言 + 5 种中文方言,约 3 秒延迟,开箱即用。文件模式支持输入视频以获得上下文感知的翻译精度。其中 7 种语言仅输出文本(不输出语音)。 +- **Qwen3.8-Omni-Flash-Realtime** — 支持实时语音翻译,语音生成语种与 Qwen3.5-Omni-Realtime 一致,支持 36 种语种和方言,各音色支持范围见[音色列表](/developer-guides/speech/omni-voice-list)。 +- **Qwen3.5-Omni** — 支持 29 种输出语言 + 7 种中文方言。音视频理解能力更强,支持联网搜索。可通过 system prompt 注入术语和领域上下文。支持实时和文件两种模式。 +- **Qwen3-Omni-Flash** — 支持 11 种输出语言 + 8 种中文方言。可通过 system prompt 注入术语和领域上下文,适用于专业领域翻译。支持实时和文件两种模式。成本更低。 + + + 快速上手选 Qwen3.5-Livetranslate(60 种语言,约 3 秒延迟);追求最佳质量和最广语言覆盖选 Qwen3.5-Omni;控制成本选 Qwen3-Omni-Flash。 + + + +| 语言 | Qwen3.5-Livetranslate | Qwen3-Livetranslate | Qwen3.5-Omni | Qwen3-Omni-Flash | +| ------- | --------------------- | ------------------- | ------------ | ---------------- | +| 英语 | ✓ | ✓ | ✓ | ✓ | +| 中文(普通话) | ✓ | ✓ | ✓ | ✓ | +| + 粤语 | 仅文本 | ✓ | ✓ | ✓ | +| + 四川话 | ✓ | ✓ | ✓ | ✓ | +| + 上海话 | ✓ | ✓ | ✓ | ✓ | +| + 北京话 | ✓ | ✓ | ✓ | ✓ | +| + 天津话 | ✓ | ✓ | ✓ | ✓ | +| + 南京话 | — | — | ✓ | ✓ | +| + 陕西话 | — | — | ✓ | ✓ | +| + 闽南语 | — | — | ✓ | ✓ | +| 法语 | ✓ | ✓ | ✓ | ✓ | +| 德语 | ✓ | ✓ | ✓ | ✓ | +| 俄语 | ✓ | ✓ | ✓ | ✓ | +| 意大利语 | ✓ | ✓ | ✓ | ✓ | +| 西班牙语 | ✓ | ✓ | ✓ | ✓ | +| 葡萄牙语 | ✓ | ✓ | ✓ | ✓ | +| 日语 | ✓ | ✓ | ✓ | ✓ | +| 韩语 | ✓ | ✓ | ✓ | ✓ | +| 阿拉伯语 | ✓ | 仅文本 | ✓ | — | +| 泰语 | ✓ | 仅文本 | ✓ | ✓ | +| 越南语 | ✓ | 仅文本 | ✓ | — | +| 印尼语 | ✓ | 仅文本 | ✓ | — | +| 土耳其语 | ✓ | 仅文本 | ✓ | — | +| 印地语 | ✓ | 仅文本 | ✓ | — | +| 马来语 | ✓ | — | ✓ | — | +| 荷兰语 | ✓ | — | ✓ | — | +| 乌尔都语 | ✓ | — | ✓ | — | +| 挪威语 | ✓ | — | ✓ | — | +| 瑞典语 | ✓ | — | ✓ | — | +| 丹麦语 | ✓ | — | ✓ | — | +| 希伯来语 | ✓ | — | ✓ | — | +| 芬兰语 | ✓ | — | ✓ | — | +| 波兰语 | ✓ | — | ✓ | — | +| 冰岛语 | ✓ | — | ✓ | — | +| 捷克语 | ✓ | — | ✓ | — | +| 菲律宾语 | ✓ | — | ✓ | — | +| 波斯语 | ✓ | — | ✓ | — | +| 希腊语 | 仅文本 | 仅文本 | — | — | +| 南非荷兰语 | 仅文本 | — | — | — | +| 阿斯图里亚斯语 | 仅文本 | — | — | — | +| 白俄罗斯语 | 仅文本 | — | — | — | +| 保加利亚语 | 仅文本 | — | — | — | +| 孟加拉语 | 仅文本 | — | — | — | +| 波斯尼亚语 | 仅文本 | — | — | — | +| 加泰罗尼亚语 | 仅文本 | — | — | — | +| 宿务语 | 仅文本 | — | — | — | +| 爱沙尼亚语 | 仅文本 | — | — | — | +| 加利西亚语 | 仅文本 | — | — | — | +| 古吉拉特语 | 仅文本 | — | — | — | +| 克罗地亚语 | 仅文本 | — | — | — | +| 匈牙利语 | 仅文本 | — | — | — | +| 爪哇语 | 仅文本 | — | — | — | +| 哈萨克语 | 仅文本 | — | — | — | +| 卡纳达语 | 仅文本 | — | — | — | +| 柯尔克孜语 | 仅文本 | — | — | — | +| 拉脱维亚语 | 仅文本 | — | — | — | +| 马其顿语 | 仅文本 | — | — | — | +| 马拉雅拉姆语 | 仅文本 | — | — | — | +| 马拉地语 | 仅文本 | — | — | — | +| 旁遮普语 | 仅文本 | — | — | — | +| 罗马尼亚语 | 仅文本 | — | — | — | +| 斯洛伐克语 | 仅文本 | — | — | — | +| 斯洛文尼亚语 | 仅文本 | — | — | — | +| 斯瓦希里语 | 仅文本 | — | — | — | +| 塔吉克语 | 仅文本 | — | — | — | +| 阿塞拜疆语 | 仅文本 | — | — | — | +| 乌克兰语 | 仅文本 | — | — | — | + + ✓ = 音频 + 文本输出。"仅文本" = 该语言不支持音频输出。Qwen3.5-Livetranslate 共支持 60 种语言(29 种音频+文本,31 种仅文本)。 + + Qwen3.8-Omni-Flash、Qwen3.8-Omni-Flash-Realtime 和 Qwen3.5-Omni 均支持 113 种输入语言/方言。详见[完整列表](/developer-guides/speech/multimodal-speech#supported-languages)。 + + 旧版 `qwen-omni-turbo` 仅支持中文和英语。 + + +## 推荐模型 + +下表列出每个系列的常用入口模型。如需锁定特定日期版本(用于版本回归或稳定性需求),请见下方"所有模型"。 + +| 模型 | API | 输入 | Function calling | 联网搜索 | 思考模式 | 翻译 | +| -------------------------------------- | -------------------- | ----------- | ---------------- | ---- | ---- | --- | +| `qwen3.8-omni-flash-realtime` | WebSocket、WebRTC、AOQ | 文本、音频、图像、视频 | ✓ | ✓ | — | 29种 | +| `qwen3.5-omni-plus` | HTTP | 文本、音频、图像、视频 | — | ✓ | — | 29种 | +| `qwen3.5-omni-flash` | HTTP | 文本、音频、图像、视频 | — | ✓ | — | 29种 | +| `qwen3-omni-flash-realtime` | WebSocket | 文本、音频、图像、视频 | — | — | — | 11种 | +| `qwen3-omni-flash` | HTTP | 文本、音频、图像、视频 | ✓ | — | ✓ | 11种 | +| `qwen3.8-livetranslate-flash-realtime` | WebSocket | 音频 | — | — | — | 60种 | +| `qwen3.5-livetranslate-flash` | HTTP | 音频、视频 | — | — | — | 18种 | + +## 所有模型 + + + +| 模型 | API | 输入 | Function calling | 联网搜索 | 思考模式 | 翻译 | +| ------------------------------- | --------- | ----- | ---------------- | ---- | ---- | -- | +| `qwen-audio-3.1-realtime-plus` | WebSocket | 音频、文本 | 支持 | 支持 | — | — | +| `qwen-audio-3.0-realtime-plus` | WebSocket | 音频、文本 | 支持 | 支持 | — | — | +| `qwen-audio-3.0-realtime-flash` | WebSocket | 音频、文本 | 支持 | 支持 | — | — | + + + +| 模型 | API | 输入 | Function calling | 联网搜索 | 思考模式 | +| ----------------------------- | -------------------- | ----------- | ---------------- | ---- | ---- | +| `qwen3.8-omni-flash-realtime` | WebSocket、WebRTC、AOQ | 文本、音频、图像、视频 | ✓ | ✓ | — | + + 多通道音频、文本与音频输出及 MCP 用法见[实时多模态语音](/developer-guides/speech/realtime-multimodal-speech)。 + + + +| 模型 | API | 输入 | Function calling | 联网搜索 | 思考模式 | 批量 | +| ---------------------------------------- | --------- | ----------- | ---------------- | ---- | ---- | -- | +| `qwen3.5-omni-plus-realtime` | WebSocket | 文本、音频、图像、视频 | ✓ | ✓ | — | — | +| `qwen3.5-omni-plus-realtime-2026-03-15` | WebSocket | 文本、音频、图像、视频 | ✓ | ✓ | — | — | +| `qwen3.5-omni-flash-realtime` | WebSocket | 文本、音频、图像、视频 | ✓ | ✓ | — | — | +| `qwen3.5-omni-flash-realtime-2026-03-15` | WebSocket | 文本、音频、图像、视频 | ✓ | ✓ | — | — | +| `qwen3.5-omni-plus` | HTTP | 文本、音频、图像、视频 | — | ✓ | — | — | +| `qwen3.5-omni-plus-2026-03-15` | HTTP | 文本、音频、图像、视频 | — | ✓ | — | — | +| `qwen3.5-omni-flash` | HTTP | 文本、音频、图像、视频 | — | ✓ | — | — | +| `qwen3.5-omni-flash-2026-03-15` | HTTP | 文本、音频、图像、视频 | — | ✓ | — | — | + + + +| 模型 | API | 输入 | Function calling | 联网搜索 | 思考模式 | 批量 | +| -------------------------------------- | --------- | ----------- | ---------------- | ---- | ---- | -- | +| `qwen3-omni-flash-realtime` | WebSocket | 文本、音频、图像、视频 | — | — | — | — | +| `qwen3-omni-flash-realtime-2025-12-01` | WebSocket | 文本、音频、图像、视频 | — | — | — | — | +| `qwen3-omni-flash-realtime-2025-09-15` | WebSocket | 文本、音频、图像、视频 | — | — | — | — | +| `qwen3-omni-flash` | HTTP | 文本、音频、图像、视频 | ✓ | — | ✓ | — | +| `qwen3-omni-flash-2025-12-01` | HTTP | 文本、音频、图像、视频 | ✓ | — | ✓ | — | +| `qwen3-omni-flash-2025-09-15` | HTTP | 文本、音频、图像、视频 | ✓ | — | ✓ | — | + + + +| 模型 | API | 输入 | 语言数 | +| -------------------------------------- | --------- | ----- | --- | +| `qwen3.8-livetranslate-flash-realtime` | WebSocket | 音频、图片 | 60 | + + + +| 模型 | API | 输入 | 语言数 | +| ------------------------------------------------- | --------- | -- | --- | +| `qwen3.5-livetranslate-flash-realtime` | WebSocket | 音频 | 60 | +| `qwen3.5-livetranslate-flash-realtime-2026-05-19` | WebSocket | 音频 | 60 | + + + +| 模型 | API | 输入 | 语言数 | +| ----------------------------------------------- | --------- | ----- | --- | +| `qwen3-livetranslate-flash-realtime`(旧版) | WebSocket | 音频 | 18 | +| `qwen3-livetranslate-flash-realtime-2025-09-22` | WebSocket | 音频 | 18 | +| `qwen3-livetranslate-flash` | HTTP | 音频、视频 | 18 | +| `qwen3-livetranslate-flash-2025-12-01` | HTTP | 音频、视频 | 18 | + + + + 以下模型不再更新。新项目的实时音视频对话推荐 Qwen3.8-Omni-Flash-Realtime;离线语音输出可选择 Qwen3.5-Omni。 + +| 模型 | 输入 | API | +| ------------------------------------- | ----------- | --------- | +| `qwen2.5-omni-7b` | 文本、音频、图像、视频 | HTTP | +| `qwen-omni-turbo` | 文本、音频、图像、视频 | HTTP | +| `qwen-omni-turbo-latest` | 文本、音频、图像、视频 | HTTP | +| `qwen-omni-turbo-2025-03-26` | 文本、音频、图像、视频 | HTTP | +| `qwen-omni-turbo-2025-01-19` | 文本、音频、图像、视频 | HTTP | +| `qwen-omni-turbo-realtime` | 文本、音频 | WebSocket | +| `qwen-omni-turbo-realtime-latest` | 文本、音频 | WebSocket | +| `qwen-omni-turbo-realtime-2025-05-08` | 文本、音频 | WebSocket | + + + +## 下一步 + +选定模型后,参考对应的调用文档: + +- Qwen-Audio Realtime(WebSocket,实时语音对话)→ [实时语音对话](/developer-guides/speech/qwen-audio-realtime) +- Qwen3.8-Omni-Flash-Realtime(WebSocket、WebRTC、AOQ,实时)→ [实时多模态语音](/developer-guides/speech/realtime-multimodal-speech) +- Qwen3.5-Omni / Qwen3-Omni(HTTP,文件)→ [多模态语音](/developer-guides/speech/multimodal-speech) +- Qwen3.8-Livetranslate / Qwen3.5-Livetranslate(实时)→ [实时翻译](/developer-guides/speech/realtime-translation) +- Qwen3-Livetranslate(HTTP,文件)→ [文件翻译](/developer-guides/speech/file-translation) + +## 了解更多 + + + + 构建实时多模态语音助手。 + + + + 处理音频和视频文件并生成语音输出。 + + + + 实时跨语言语音翻译。 + + + + 翻译音频和视频文件。 + + diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-speech-to-text-models.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-speech-to-text-models.md new file mode 100644 index 0000000..00fd111 --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-speech-to-text-models.md @@ -0,0 +1,227 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 语音识别模型 + +> 选择适合实时字幕、音频转写等场景的语音识别模型。 + +## 从闭源模型迁移到千问AI平台? + +如果你正在使用 Whisper、Deepgram 或 Google 的语音识别服务,可参考下表选择对应的千问AI平台模型。 + +| 使用场景 | 闭源模型代表 | 千问AI平台推荐 | +| ---------- | -------------------------------- | --------------------------------------------------------------- | +| 实时识别 | Deepgram Nova-3、Google Chirp 3 | `qwen-audio-3.1-asr-flash-streaming` | +| 非实时 / 文件转写 | OpenAI gpt-4o-transcribe、Whisper | `qwen-audio-3.1-asr-flash-filetrans`、`qwen-audio-3.1-asr-flash` | + +本文按"先选维度、再看模型"的顺序帮助您完成选型:先从"[选型决策维度](#选型决策维度)"(实时/非实时、专业术语、说话人分离、情感识别)确认场景;再到"[推荐模型](#推荐模型)"查看针对各场景的首选;如需更多版本可在"[全部模型](#全部模型)"按系列展开;最后到"[音频规格](#音频规格)"核对输入文件约束。各模型支持的语言(含方言)随附在"[全部模型](#全部模型)"的各系列子节内。 + +## 选型决策维度 + +从以下 4 个维度逐项确认,每个维度都会推荐适配的模型。 + +### 实时还是非实时? + +实时是指在用户说话的同时输出识别结果,非实时是指录音结束后再进行转写。 + +- **实时(实时语音识别)**:基于 WebSocket 协议,音频流式输入,文本流式输出。适用于实时字幕、语音助手和会议转写。推荐使用 `qwen-audio-3.1-asr-flash-streaming`(热词、Prompt 上下文、多语种及方言)。 + +- **非实时(录音文件识别)**:基于 HTTP 协议,提交音频文件获取识别结果。适用于呼叫中心录音、播客和访谈等场景。推荐使用 `qwen-audio-3.1-asr-flash-filetrans`(热词、Prompt 上下文、说话人分离)。 + +Qwen-Audio-3.x-ASR-Flash-Streaming 和 Fun-ASR-Realtime 系列模型还支持通过 AOQ 协议接入;如果是客户端对接,且更看重稳定的延迟、弱网下的交互能力、实时双工的降噪与回声消除,可优先考虑 AOQ,协议对比与选型请参见 [Realtime API 概述](/api-reference/realtime-api/overview)。 + +Qwen-Audio-3.x-ASR-Flash-Streaming、Qwen-Audio-3.x-ASR-Flash-Filetrans,以及 Fun-ASR 和 Qwen-ASR 的实时模型支持通过 DashScope SDK(Java、Python)接入。Qwen-Audio-3.x-ASR-Flash-Streaming、Qwen-Audio-3.x-ASR-Flash-Filetrans 和 Fun-ASR 模型还支持 Android、iOS SDK 接入,速度快。其他模型需根据对应的 WebSocket 或 HTTP 协议直接调用。 + +选择实时接入请参考[实时语音识别](/developer-guides/speech/asr-realtime),选择非实时接入请参考[非实时语音识别](/developer-guides/speech/asr)。 + +### 处理专业术语 + +两种方式,按灵活性排序: + +1. **Prompt 上下文注入** — 在系统提示词中描述您的领域背景,无需预配置。模型在每次请求时自适应。推荐使用 Qwen-Audio-3.1-ASR-Flash-Streaming、Qwen-Audio-3.1-ASR-Flash-Filetrans 和 Qwen-Audio-3.1-ASR-Flash 系列模型。 + +2. **热词** — 提供带权重的词汇表。推荐使用 Qwen-Audio-3.1-ASR-Flash-Streaming、Qwen-Audio-3.1-ASR-Flash-Filetrans 和 Qwen-Audio-3.1-ASR-Flash 系列模型。 + +### 说话人分离 + +`qwen-audio-3.1-asr-flash`、`qwen-audio-3.1-asr-flash-filetrans`、`qwen-audio-3.0-asr-flash-filetrans` 以及 Fun-ASR 系列的非实时模型(`fun-asr`、`fun-asr-mtl`)支持说话人分离。如果您需要区分"谁说了什么",推荐使用 `qwen-audio-3.1-asr-flash-filetrans`。 + +### 情感识别 + +Qwen-ASR 系列模型在转写的同时支持情感识别。推荐使用 `qwen3-asr-flash-realtime`(实时)或 `qwen3-asr-flash-filetrans`(非实时)。 + +## 推荐模型 + +以下为各场景下最推荐的模型,可前往模型广场查看详情。 + +| 模型ID | 模式 | API | 精度增强 | 情感识别 | 说话人分离 | 支持语言 | 音频最大时长/大小 | +| ------------------------------------ | --- | --------- | ------------- | ---- | ----- | ------ | ---------- | +| `qwen-audio-3.1-asr-flash-streaming` | 实时 | WebSocket | 热词、Prompt 上下文 | ✗ | ✗ | 多语种及方言 | 无限制 | +| `qwen-audio-3.1-asr-flash-filetrans` | 非实时 | HTTP | 热词、Prompt 上下文 | ✗ | ✓ | 多语种及方言 | 12小时 / 2GB | + +## 全部模型 + + + +| 模型ID | 模式 | API | 精度增强 | 情感识别 | 说话人分离 | 支持语言 | 音频最大时长/大小 | +| ------------------------------------ | -- | --------- | ------------- | ---- | ----- | ------ | --------- | +| `qwen-audio-3.1-asr-flash-streaming` | 实时 | WebSocket | 热词、Prompt 上下文 | ✗ | ✗ | 多语种及方言 | 无限制 | +| `qwen-audio-3.0-asr-flash-streaming` | 实时 | WebSocket | 热词、Prompt 上下文 | ✗ | ✗ | 多语种及方言 | 无限制 | + + **支持的语言(按版本)**: + + - `qwen-audio-3.1-asr-flash-streaming`:中文、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语、挪威语。支持上海、南昌、宁波、客家、杭州、温州、湖南、福建、粤语、苏州方言。 + - `qwen-audio-3.0-asr-flash-streaming`:中文(普通话、粤语、吴语、闽南语、客家话、赣语、湘语、晋语;并支持中原、西南、冀鲁、江淮、兰银、胶辽、东北、北京、港台等,包括河南、陕西、湖北、四川、重庆、云南、贵州、广东、广西、河北、天津、山东、安徽、南京、江苏、杭州、甘肃、宁夏等地区官话口音)、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、挪威语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语 + + + +| 模型ID | 模式 | API | 精度增强 | 情感识别 | 说话人分离 | 支持语言 | 音频最大时长/大小 | +| ------------------------------------ | --- | ---- | ------------- | ---- | ----- | ------ | ---------- | +| `qwen-audio-3.1-asr-flash-filetrans` | 非实时 | HTTP | 热词、Prompt 上下文 | ✗ | ✓ | 多语种及方言 | 12小时 / 2GB | +| `qwen-audio-3.0-asr-flash-filetrans` | 非实时 | HTTP | 热词、Prompt 上下文 | ✗ | ✓ | 多语种及方言 | 12小时 / 2GB | + + **支持的语言(按版本)**: + + - `qwen-audio-3.1-asr-flash-filetrans`:中文、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语、挪威语。支持上海、南昌、宁波、客家、杭州、温州、湖南、福建、粤语、苏州方言。 + - `qwen-audio-3.0-asr-flash-filetrans`:中文(普通话、粤语、吴语、闽南语、客家话、赣语、湘语、晋语;并支持中原、西南、冀鲁、江淮、兰银、胶辽、东北、北京、港台等,包括河南、陕西、湖北、四川、重庆、云南、贵州、广东、广西、河北、天津、山东、安徽、南京、江苏、杭州、甘肃、宁夏等地区官话口音)、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、挪威语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语 + + + +| 模型ID | 模式 | API | 精度增强 | 情感识别 | 说话人分离 | 支持语言 | 音频最大时长/大小 | +| -------------------------- | --- | ---- | ------------- | ---- | ----- | ------ | --------- | +| `qwen-audio-3.1-asr-flash` | 非实时 | HTTP | 热词、Prompt 上下文 | ✗ | ✓ | 多语种及方言 | 5分钟 / 2GB | +| `qwen-audio-3.0-asr-flash` | 非实时 | HTTP | 热词、Prompt 上下文 | ✗ | ✗ | 多语种及方言 | 5分钟 / 2GB | + + **支持的语言(按版本)**: + + - `qwen-audio-3.1-asr-flash`:中文、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语、挪威语。支持上海、南昌、宁波、客家、杭州、温州、湖南、福建、粤语、苏州方言。 + - `qwen-audio-3.0-asr-flash`:中文(普通话、粤语、吴语、闽南语、客家话、赣语、湘语、晋语;并支持中原、西南、冀鲁、江淮、兰银、胶辽、东北、北京、港台等,包括河南、陕西、湖北、四川、重庆、云南、贵州、广东、广西、河北、天津、山东、安徽、南京、江苏、杭州、甘肃、宁夏等地区官话口音)、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、挪威语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语 + + + +| 模型ID | 模式 | API | 精度增强 | 情感识别 | 说话人分离 | 支持语言 | 音频最大时长/大小 | +| -------------------------------------- | --- | --------- | ---------- | ---- | ----- | -------- | ---------- | +| `fun-asr-realtime` | 实时 | WebSocket | 热词 | ✗ | ✗ | 多语种及方言 | 无限制 | +| `fun-asr-realtime-2026-02-28` | 实时 | WebSocket | 热词 | ✗ | ✗ | 中、英、日及方言 | 无限制 | +| `fun-asr-realtime-2025-11-07` | 实时 | WebSocket | 热词 | ✗ | ✗ | 多语种及方言 | 无限制 | +| `fun-asr-realtime-2025-09-15` | 实时 | WebSocket | 热词 | ✗ | ✗ | 中、英 | 无限制 | +| `fun-asr-flash-8k-realtime` | 实时 | WebSocket | 热词 | ✗ | ✗ | 中文 | 无限制 | +| `fun-asr-flash-8k-realtime-2026-01-28` | 实时 | WebSocket | 热词 | ✗ | ✗ | 中文 | 无限制 | +| `fun-asr` | 非实时 | HTTP | 热词 | ✗ | ✓ | 多语种及方言 | 12小时 / 2GB | +| `fun-asr-2025-11-07` | 非实时 | HTTP | 热词 | ✗ | ✓ | 多语种及方言 | 12小时 / 2GB | +| `fun-asr-2025-08-25` | 非实时 | HTTP | 热词 | ✗ | ✓ | 中、英 | 12小时 / 2GB | +| `fun-asr-mtl` | 非实时 | HTTP | 热词 | ✗ | ✓ | 多语种及方言 | 12小时 / 2GB | +| `fun-asr-mtl-2025-08-25` | 非实时 | HTTP | 热词 | ✗ | ✓ | 多语种及方言 | 12小时 / 2GB | +| `fun-asr-flash-2026-06-15` | 非实时 | HTTP | Prompt 上下文 | ✗ | ✗ | 多语种及方言 | 5分钟 / 2GB | + + **支持的语言(按版本)**: + + - **Fun-ASR-Realtime 主版本**(`fun-asr-realtime`、`fun-asr-realtime-2025-11-07`):中文(普通话、粤语、吴语、闽南语、客家话、赣语、湘语、晋语;并支持中原、西南、冀鲁、江淮、兰银、胶辽、东北、北京、港台等,包括河南、陕西、湖北、四川、重庆、云南、贵州、广东、广西、河北、天津、山东、安徽、南京、江苏、杭州、甘肃、宁夏等地区官话口音)、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、挪威语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语 + - `fun-asr-realtime-2026-02-28`:中文(普通话、粤语、吴语、闽南语、客家话、赣语、湘语、晋语;并支持中原、西南、冀鲁、江淮、兰银、胶辽、东北、北京、港台等,包括河南、陕西、湖北、四川、重庆、云南、贵州、广东、广西、河北、天津、山东、安徽、南京、江苏、杭州、甘肃、宁夏等地区官话口音)、英语、日语 + - `fun-asr-realtime-2025-09-15`:中文(普通话)、英文 + - **Fun-ASR-Flash-Realtime(8K)**(`fun-asr-flash-8k-realtime`、`fun-asr-flash-8k-realtime-2026-01-28`):中文 + - **Fun-ASR 主版本**(`fun-asr`、`fun-asr-2025-11-07`):中文(普通话、粤语、吴语、闽南语、客家话、赣语、湘语、晋语;并支持中原、西南、冀鲁、江淮、兰银、胶辽、东北、北京、港台等,包括河南、陕西、湖北、四川、重庆、云南、贵州、广东、广西、河北、天津、山东、安徽、南京、江苏、杭州、甘肃、宁夏等地区官话口音)、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、挪威语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语 + - **Fun-ASR-Flash**(`fun-asr-flash-2026-06-15`):中文(普通话、粤语、吴语、闽南语、客家话、赣语、湘语、晋语;并支持中原、西南、冀鲁、江淮、兰银、胶辽、东北、北京、港台等,包括河南、陕西、湖北、四川、重庆、云南、贵州、广东、广西、河北、天津、山东、安徽、南京、江苏、杭州、甘肃、宁夏等地区官话口音)、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、挪威语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语 + - `fun-asr-2025-08-25`:中文(普通话)、英文 + - **Fun-ASR(MTL)**(`fun-asr-mtl`、`fun-asr-mtl-2025-08-25`):中文(普通话、粤语)、英语、日语、韩语、越南语、泰语、印尼语、马来语、菲律宾语、印地语、阿拉伯语、法语、德语、西班牙语、葡萄牙语、俄语、意大利语、荷兰语、瑞典语、丹麦语、芬兰语、挪威语、希腊语、波兰语、捷克语、匈牙利语、罗马尼亚语、保加利亚语、克罗地亚语、斯洛伐克语 + + + +| 模型ID | 模式 | API | 精度增强 | 情感识别 | 说话人分离 | 支持语言 | 音频最大时长/大小 | +| -------------------------------------- | --- | --------------- | ---- | ---- | ----- | ------ | ---------- | +| `qwen3-asr-flash-realtime` | 实时 | WebSocket | ✗ | ✓ | ✗ | 多语种及方言 | 无限制 | +| `qwen3-asr-flash-realtime-2026-02-10` | 实时 | WebSocket | ✗ | ✓ | ✗ | 多语种及方言 | 无限制 | +| `qwen3-asr-flash-realtime-2025-10-27` | 实时 | WebSocket | ✗ | ✓ | ✗ | 多语种及方言 | 无限制 | +| `qwen3-asr-flash-filetrans` | 非实时 | HTTP | ✗ | ✓ | ✗ | 多语种及方言 | 12小时 / 2GB | +| `qwen3-asr-flash-filetrans-2025-11-17` | 非实时 | HTTP | ✗ | ✓ | ✗ | 多语种及方言 | 12小时 / 2GB | +| `qwen3-asr-flash` | 非实时 | HTTP(OpenAI 兼容) | ✗ | ✓ | ✗ | 多语种及方言 | 5分钟 / 10MB | +| `qwen3-asr-flash-2026-02-10` | 非实时 | HTTP(OpenAI 兼容) | ✗ | ✓ | ✗ | 多语种及方言 | 5分钟 / 10MB | +| `qwen3-asr-flash-2025-09-08` | 非实时 | HTTP(OpenAI 兼容) | ✗ | ✓ | ✗ | 多语种及方言 | 5分钟 / 10MB | + + **支持的语言**:所有 Qwen-ASR 系列模型(`qwen3-asr-flash-realtime`、`qwen3-asr-flash-filetrans`、`qwen3-asr-flash` 及其快照版)均支持相同的语言列表:中文(普通话、四川话、闽南语、吴语、粤语)、英语、日语、德语、韩语、俄语、法语、葡萄牙语、阿拉伯语、意大利语、西班牙语、印地语、印尼语、泰语、土耳其语、乌克兰语、越南语、捷克语、丹麦语、菲律宾语、芬兰语、冰岛语、马来语、挪威语、波兰语、瑞典语。 + + + + Paraformer 是较早一代的 ASR 模型,包括实时与非实时两类。若您的业务允许,建议迁移到前文推荐的 Fun-ASR 或 Qwen-ASR。 + +| 模型ID | API | 说明 | +| --------------------------- | --------- | ---------------------------- | +| `paraformer-realtime-v2` | WebSocket | 实时识别,中、英、日、韩、德、法、俄 | +| `paraformer-realtime-v1` | WebSocket | 实时识别,中、英、日、韩、德、法、俄 | +| `paraformer-realtime-8k-v2` | WebSocket | 实时识别,8kHz 电话场景,中文 | +| `paraformer-realtime-8k-v1` | WebSocket | 实时识别,8kHz 电话场景,中文 | +| `paraformer-v2` | HTTP | 录音文件识别,支持说话人分离,中、英、日、韩、德、法、俄 | +| `paraformer-8k-v2` | HTTP | 录音文件识别,8kHz 电话场景,中文 | +| `paraformer-v1` | HTTP | 录音文件识别,支持说话人分离,中、英、日、韩、德、法、俄 | +| `paraformer-8k-v1` | HTTP | 录音文件识别,8kHz 电话场景,中文 | +| `paraformer-mtl-v1` | HTTP | 录音文件识别,支持说话人分离,多语种 | + + **支持的语言(按版本)**: + + - `paraformer-realtime-v2`、`paraformer-v2`:中文(普通话、粤语、吴语、闽南语、东北话、甘肃话、贵州话、河南话、湖北话、湖南话、宁夏话、山西话、陕西话、山东话、四川话、天津话、江西话、云南话、上海话)、英文、日语、韩语、德语、法语、俄语 + - `paraformer-realtime-v1`、`paraformer-realtime-8k-v2`、`paraformer-realtime-8k-v1`、`paraformer-8k-v2`、`paraformer-8k-v1`:中文普通话 + - `paraformer-v1`:中文普通话、英文 + - `paraformer-mtl-v1`:中文(普通话、粤语、吴语、闽南语、东北话、甘肃话、贵州话、河南话、湖北话、湖南话、宁夏话、山西话、陕西话、山东话、四川话、天津话)、英文、日语、韩语、西班牙语、印尼语、法语、德语、意大利语、马来语 + + + + 以下模型已计划下线,仅供存量业务参考,请尽快迁移到推荐模型。 + +| 模型ID | API | 说明 | +| ------------------- | --------- | ------------------ | +| `gummy-realtime-v1` | WebSocket | 实时识别,中、英及方言 | +| `gummy-chat-v1` | WebSocket | 短音频实时识别(1分钟限制),多语种 | +| `sensevoice-v1` | HTTP | 录音文件识别,多语种 | + + + +## 音频规格 + +下表汇总了**实时**和**非实时**两种模式下的音频规格(输入方式、格式、采样率、大小/时长)。各模型支持的语言(含方言)见上文"[全部模型](#全部模型)"内对应系列子节。 + +### 实时 + +| 模型ID | 输入方式 | 音频格式 | 采样率 | 大小/时长 | +| -------------------------------------------------------------------------------------------------------------------- | ------------ | -------------------------------------------- | -------------------------------------------------------------------------------------------- | ----- | +| **Qwen-Audio-3.x-ASR-Flash-Streaming**(`qwen-audio-3.1-asr-flash-streaming`、`qwen-audio-3.0-asr-flash-streaming` 系列) | 二进制(Binary)流 | `pcm`、`wav`、`mp3`、`opus`、`speex`、`aac`、`amr` | 任意 | 不限 | +| **Fun-ASR-Realtime**(`fun-asr-realtime` 系列) | 二进制(Binary)流 | `pcm`、`wav`、`mp3`、`opus`、`speex`、`aac`、`amr` | 任意 | 不限 | +| **Fun-ASR-Flash-8K-Realtime**(`fun-asr-flash-8k-realtime` 系列) | 二进制(Binary)流 | 同 Fun-ASR-Realtime | 8 kHz | 不限 | +| **Qwen-ASR-Realtime**(`qwen3-asr-flash-realtime` 系列) | 二进制(Binary)流 | `pcm`、`opus` | 8 kHz、16 kHz | 不限 | +| **Paraformer-Realtime**(`paraformer-realtime-v2/v1`、`paraformer-realtime-8k-v2/v1`) | 二进制(Binary)流 | 同 Fun-ASR-Realtime | `paraformer-realtime-v2` 任意;`paraformer-realtime-v1` 16 kHz;`paraformer-realtime-8k-*` 8 kHz | 不限 | + +所有实时模型均为**单声道**输入。 + +### 非实时 + +| 模型ID | 输入方式 | 音频格式 | 采样率 | 文件大小/时长 | +| -------------------------------------------------------------------------------------------------------------------- | ------------------------------ | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ----------------------------- | +| **Qwen-Audio-3.x-ASR-Flash-Filetrans**(`qwen-audio-3.1-asr-flash-filetrans`、`qwen-audio-3.0-asr-flash-filetrans` 系列) | 公网可访问的文件 URL,单次 1 个 | `aac`、`amr`、`avi`、`flac`、`flv`、`m4a`、`mkv`、`mov`、`mp3`、`mp4`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | 任意 | ≤2 GB;≤12 小时(启用说话人分离建议 ≤2 小时) | +| **Qwen-Audio-3.1-ASR-Flash**(`qwen-audio-3.1-asr-flash`) | URL / Base64,单次 1 个 | `aac`、`amr`、`avi`、`flac`、`flv`、`m4a`、`mkv`、`mov`、`mp3`、`mp4`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | 任意 | ≤2 GB;≤5 分钟 | +| **Qwen-Audio-3.0-ASR-Flash**(`qwen-audio-3.0-asr-flash`) | URL / Base64,单次 1 个 | `aac`、`amr`、`avi`、`flac`、`flv`、`m4a`、`mkv`、`mov`、`mp3`、`mp4`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | 任意 | ≤10 MB;≤5 分钟 | +| **Fun-ASR**(`fun-asr`、`fun-asr-mtl` 系列) | 公网可访问的文件 URL,单次 1 个 | `aac`、`amr`、`avi`、`flac`、`flv`、`m4a`、`mkv`、`mov`、`mp3`、`mp4`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | 任意 | ≤2 GB;≤12 小时(启用说话人分离建议 ≤2 小时) | +| **Fun-ASR-Flash**(`fun-asr-flash-2026-06-15`) | URL / Base64,单次 1 个 | `aac`、`amr`、`avi`、`flac`、`flv`、`m4a`、`mkv`、`mov`、`mp3`、`mp4`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | 任意 | ≤2 GB;≤5 分钟 | +| **Fun-ASR-Realtime**(`fun-asr-realtime`、`fun-asr-realtime-2026-02-28`) | URL / Base64,单次 1 个 | `aac`、`amr`、`avi`、`flac`、`flv`、`m4a`、`mkv`、`mov`、`mp3`、`mp4`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | 任意 | ≤2 GB;≤5 分钟 | +| **Paraformer**(`paraformer-v2/v1`、`paraformer-mtl-v1`、`paraformer-8k-v2/v1`) | 同 Fun-ASR | 同 Fun-ASR | `paraformer-v2/v1` 任意;`paraformer-8k-*` 仅 8 kHz;`paraformer-mtl-v1` 16 kHz 及以上 | 同 Fun-ASR | +| **Qwen3-ASR-Flash-Filetrans**(`qwen3-asr-flash-filetrans` 系列) | 公网可访问的文件 URL,单次 1 个 | `aac`、`amr`、`avi`、`flac`、`flv`、`m4a`、`mkv`、`mov`、`mp3`、`mp4`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | `pcm` 必须 16 kHz;其他格式任意(服务端会重采样为 16 kHz 再识别) | ≤2 GB;≤12 小时 | +| **Qwen3-ASR-Flash**(`qwen3-asr-flash` 系列) | URL / Base64 / 本地文件绝对路径,单次 1 个 | `aac`、`amr`、`avi`、`aiff`、`flac`、`flv`、`mkv`、`mp3`、`mpeg`、`ogg`、`opus`、`wav`、`webm`、`wma`、`wmv` | `pcm` 必须 16 kHz;其他格式任意(服务端会重采样为 16 kHz 再识别) | ≤10 MB;≤5 分钟 | + + + 需要将语音翻译为其他语言?请参阅[语音翻译模型](/developer-guides/speech/s2s-models),了解 LiveTranslate 和 Qwen-Omni 的实时与文件翻译能力。 + + +## 了解更多 + + + + 流式输入音频,实时返回识别文本。 + + + + 通过异步 API 转写录音文件。 + + + + 提升专业术语的识别准确率。 + + diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-ssml.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-ssml.md new file mode 100644 index 0000000..836e7db --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-ssml.md @@ -0,0 +1,2661 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# SSML 与 LaTeX + +> 通过 SSML 控制语速、停顿、发音等语音特征,或将 LaTeX 公式转换为自然语音 + +通过 SSML(Speech Synthesis Markup Language)标记语言,可以精细控制语速、停顿、发音等语音特征;通过 LaTeX 公式朗读功能,可以将数学公式转换为自然语音。这两项功能均适用于 CosyVoice 模型。 + +## 概述 + +SSML(Speech Synthesis Markup Language)是一种基于 XML 的语音合成标记语言。在文本中嵌入 SSML 标签后,可以精细控制语速、语调、停顿和音量等语音特征,也可以添加背景音乐和音效,实现更丰富的语音表达效果。 + +CosyVoice 还支持解析文本中嵌入的 LaTeX 公式,并按照符合中文阅读习惯的方式将其朗读出来,适用于在线教育、有声读物等包含数学公式的场景。例如,输入文本"这是一道一元二次方程的求根公式:`$x = \frac{-b \pm \sqrt{b^2-4ac}}{2a}$`"时,模型会将公式朗读为"x等于负b加减根号下b的平方减四ac,分之二a"。 + +典型应用场景包括: + +- **有声读物**:灵活控制停顿和语速,搭配背景音乐增强沉浸感 +- **智能客服**:通过 `` 标签确保电话号码、日期等信息的准确朗读 +- **多语种播报**:使用 `` 标签精确指定外文发音 +- **在线教育**:通过 LaTeX 公式朗读功能将数学公式转为自然语音 + +两项功能均适用于 CosyVoice 模型系列。如需了解各模型的选型建议,请参见[语音合成模型](/developer-guides/speech/tts-models)。 + +## SSML 标记语言 + +### 使用限制 + +- **模型**:qwen-audio-3.1-tts-flash、qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3.5-flash、cosyvoice-v3.5-plus、cosyvoice-v3-flash、cosyvoice-v3-plus、cosyvoice-v2。 +- **音色**:克隆音色,以及[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)中标注为支持 SSML 的系统音色。 +- **接口**: + - Java SDK(2.20.3 及以上版本):支持非流式调用和单向流式调用。 + - Python SDK(1.23.4 及以上版本):支持非流式调用和单向流式调用。 + - WebSocket API:需将参数 `enable_ssml` 设置为 `true`,且只允许发送一次 continue-task 事件。 + - HTTP API:需将参数 `enable_ssml` 设置为 `true`。 + + + `cosyvoice-v3.5-plus` 和 `cosyvoice-v3.5-flash` 模型专用于声音复刻场景(不提供系统音色)。使用前,请先参见[声音复刻](/developer-guides/speech/voice-cloning)创建目标音色。 + + +### 快速开始 + +以下示例展示如何使用 SSML 控制语速进行语音合成。运行前,请完成以下准备工作: + +1. [获取 API Key](/developer-guides/administration/api-keys) +2. 安装 DashScope SDK(Python 1.23.4 及以上版本,Java 2.20.3 及以上版本)。详情请参见[安装 SDK](/api-reference/preparation/install-sdk)。 + + + + + ```java 非流式调用 + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisParam; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesizer; + import com.alibaba.dashscope.utils.Constants; + import java.io.File; + import java.io.FileOutputStream; + import java.io.IOException; + import java.nio.ByteBuffer; + /** + * SSML功能说明: + * 1. 只有非流式调用和单向流式调用支持SSML功能 + * 2. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + */ + public class Main { + private static String model = "cosyvoice-v3-flash"; + private static String voice = "longanyang"; + public static void main(String[] args) { + Constants.baseWebsocketApiUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference"; + streamAudioDataToSpeaker(); + System.exit(0); + } + public static void streamAudioDataToSpeaker() { + SpeechSynthesisParam param = + SpeechSynthesisParam.builder() + // 若没有配置环境变量,请用API Key将下行替换为:.apiKey("sk-xxx") + .apiKey(System.getenv("DASHSCOPE_API_KEY")) + .model(model) + .voice(voice) + .build(); + SpeechSynthesizer synthesizer = new SpeechSynthesizer(param, null); + ByteBuffer audio = null; + try { + // 非流式调用,阻塞直至音频返回 + // 特殊字符需要进行转义 + audio = synthesizer.call("我的语速比正常人快。"); + } catch (Exception e) { + throw new RuntimeException(e); + } finally { + // 任务结束关闭websocket连接 + synthesizer.getDuplexApi().close(1000, "bye"); + } + if (audio != null) { + // 将音频数据保存到本地文件"output.mp3"中 + File file = new File("output.mp3"); + try (FileOutputStream fos = new FileOutputStream(file)) { + fos.write(audio.array()); + } catch (IOException e) { + throw new RuntimeException(e); + } + } + // 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + System.out.println( + "[Metric] requestId为:" + + synthesizer.getLastRequestId() + + "首包延迟(毫秒)为:" + + synthesizer.getFirstPackageDelay()); + } + } + ``` + + ```java 单向流式调用 + import com.alibaba.dashscope.audio.tts.SpeechSynthesisResult; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisAudioFormat; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisParam; + import com.alibaba.dashscope.audio.ttsv2.SpeechSynthesizer; + import com.alibaba.dashscope.common.ResultCallback; + import com.alibaba.dashscope.utils.Constants; + import java.io.FileOutputStream; + import java.io.IOException; + import java.util.concurrent.CountDownLatch; + /** + * SSML功能说明: + * 1. 只有非流式调用和单向流式调用支持SSML功能 + * 2. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + */ + public class Main { + private static String model = "cosyvoice-v3-flash"; + private static String voice = "longanyang"; + public static void main(String[] args) { + Constants.baseWebsocketApiUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference"; + streamAudioDataToSpeaker(); + System.out.println("音频已保存到 output.mp3 文件中"); + System.exit(0); + } + public static void streamAudioDataToSpeaker() { + CountDownLatch latch = new CountDownLatch(1); + final FileOutputStream[] fileOutputStream = new FileOutputStream[1]; + try { + fileOutputStream[0] = new FileOutputStream("output.mp3"); + } catch (IOException e) { + System.err.println("无法创建输出文件: " + e.getMessage()); + return; + } + // 实现回调接口ResultCallback + ResultCallback callback = new ResultCallback() { + @Override + public void onEvent(SpeechSynthesisResult result) { + if (result.getAudioFrame() != null) { + // 将音频数据写入本地文件 + try { + byte[] audioData = result.getAudioFrame().array(); + fileOutputStream[0].write(audioData); + fileOutputStream[0].flush(); + } catch (IOException e) { + System.err.println("写入音频数据失败: " + e.getMessage()); + } + } + } + @Override + public void onComplete() { + System.out.println("收到Complete,语音合成结束"); + closeFileOutputStream(fileOutputStream[0]); + latch.countDown(); + } + @Override + public void onError(Exception e) { + System.out.println("出现异常:" + e.toString()); + closeFileOutputStream(fileOutputStream[0]); + latch.countDown(); + } + }; + SpeechSynthesisParam param = + SpeechSynthesisParam.builder() + // 若没有配置环境变量,请用API Key将下行替换为:.apiKey("sk-xxx") + .apiKey(System.getenv("DASHSCOPE_API_KEY")) + .model(model) + .voice(voice) + .format(SpeechSynthesisAudioFormat.MP3_22050HZ_MONO_256KBPS) + .build(); + SpeechSynthesizer synthesizer = new SpeechSynthesizer(param, callback); + try { + // 单向流式调用,立即返回null(实际结果通过回调接口异步传递),在回调接口的onEvent方法中实时获取二进制音频 + // 特殊字符需要进行转义 + synthesizer.call("我的语速比正常人快。"); + // 等待合成完成 + latch.await(); + } catch (Exception e) { + throw new RuntimeException(e); + } finally { + // 任务结束后关闭websocket连接 + try { + synthesizer.getDuplexApi().close(1000, "bye"); + } catch (Exception e) { + System.err.println("关闭WebSocket连接失败: " + e.getMessage()); + } + // 确保文件流被关闭 + closeFileOutputStream(fileOutputStream[0]); + } + // 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + System.out.println( + "[Metric] requestId为:" + + synthesizer.getLastRequestId() + + ",首包延迟(毫秒)为:" + + synthesizer.getFirstPackageDelay()); + } + private static void closeFileOutputStream(FileOutputStream fileOutputStream) { + try { + if (fileOutputStream != null) { + fileOutputStream.close(); + } + } catch (IOException e) { + System.err.println("关闭文件流失败: " + e.getMessage()); + } + } + } + ``` + + + + + + ```python 非流式调用 + # coding=utf-8 + # SSML功能说明: + # 1. 只有非流式调用和单向流式调用支持SSML功能 + # 2. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + import dashscope + from dashscope.audio.tts_v2 import * + import os + # 若没有配置环境变量,请用API Key将下行替换为:dashscope.api_key = "sk-xxx" + dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY') + dashscope.base_websocket_api_url='wss://maas.qianwenaiapi.com/api-ws/v1/inference' + # 模型 + model = "cosyvoice-v3-flash" + # 音色 + voice = "longanyang" + # 实例化SpeechSynthesizer,并在构造方法中传入模型(model)、音色(voice)等请求参数 + synthesizer = SpeechSynthesizer(model=model, voice=voice) + # 非流式调用,阻塞直至音频返回 + # 特殊字符需要进行转义 + audio = synthesizer.call("我的语速比正常人快。") + # 将音频保存至本地 + with open('output.mp3', 'wb') as f: + f.write(audio) + # 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + print('[Metric] requestId为:{},首包延迟为:{}毫秒'.format( + synthesizer.get_last_request_id(), + synthesizer.get_first_package_delay())) + ``` + + ```python 单向流式调用 + # coding=utf-8 + # SSML功能说明: + # 1. 只有非流式调用和单向流式调用支持SSML功能 + # 2. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + import dashscope + from dashscope.audio.tts_v2 import * + import os + from datetime import datetime + def get_timestamp(): + now = datetime.now() + formatted_timestamp = now.strftime("[%Y-%m-%d %H:%M:%S.%f]") + return formatted_timestamp + # 若没有配置环境变量,请用API Key将下行替换为:dashscope.api_key = "sk-xxx" + dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY') + dashscope.base_websocket_api_url='wss://maas.qianwenaiapi.com/api-ws/v1/inference' + # 模型 + model = "cosyvoice-v3-flash" + # 音色 + voice = "longanyang" + # 定义回调接口 + class Callback(ResultCallback): + _player = None + _stream = None + def on_open(self): + # 打开输出文件,准备写入音频数据 + self.file = open("output.mp3", "wb") + print("连接建立:" + get_timestamp()) + def on_complete(self): + print("语音合成完成,所有合成结果已被接收:" + get_timestamp()) + if hasattr(self, 'file') and self.file: + self.file.close() + # 首次发送文本时需建立 WebSocket 连接,因此首包延迟会包含连接建立的耗时 + print('[Metric] requestId为:{},首包延迟为:{}毫秒'.format( + self.synthesizer.get_last_request_id(), + self.synthesizer.get_first_package_delay())) + def on_error(self, message: str): + print(f"语音合成出现异常:{message}") + if hasattr(self, 'file') and self.file: + self.file.close() + def on_close(self): + print("连接关闭:" + get_timestamp()) + if hasattr(self, 'file') and self.file: + self.file.close() + def on_event(self, message): + pass + def on_data(self, data: bytes) -> None: + print(get_timestamp() + " 二进制音频长度为:" + str(len(data))) + # 将音频数据写入文件 + self.file.write(data) + callback = Callback() + # 实例化SpeechSynthesizer,并在构造方法中传入模型(model)、音色(voice)等请求参数 + synthesizer = SpeechSynthesizer( + model=model, + voice=voice, + callback=callback, + ) + # 将synthesizer实例赋值给callback,以便在on_complete中使用 + callback.synthesizer = synthesizer + # 单向流式调用,发送待合成文本,在回调接口的on_data方法中实时获取二进制音频 + # 特殊字符需要进行转义 + synthesizer.call("我的语速比正常人快。") + ``` + + + + + + + ```go + // SSML功能说明: + // 1. 在发送run-task指令时,将参数enable_ssml设置为true,以开启SSML支持 + // 2. 通过continue-task指令发送包含SSML的文本,且只允许发送一次continue-task指令 + // 3. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + package main + import ( + "encoding/json" + "fmt" + "net/http" + "os" + "strings" + "time" + "github.com/google/uuid" + "github.com/gorilla/websocket" + ) + const ( + wsURL = "wss://maas.qianwenaiapi.com/api-ws/v1/inference/" + outputFile = "output.mp3" + ) + func main() { + // 若没有配置环境变量,请用API Key将下行替换为:apiKey := "sk-xxx" + apiKey := os.Getenv("DASHSCOPE_API_KEY") + // 清空输出文件 + os.Remove(outputFile) + os.Create(outputFile) + // 连接WebSocket + header := make(http.Header) + header.Add("X-DashScope-DataInspection", "enable") + header.Add("Authorization", fmt.Sprintf("bearer %s", apiKey)) + conn, resp, err := websocket.DefaultDialer.Dial(wsURL, header) + if err != nil { + if resp != nil { + fmt.Printf("连接失败 HTTP状态码: %d\n", resp.StatusCode) + } + fmt.Println("连接失败:", err) + return + } + defer conn.Close() + // 生成任务ID + taskID := uuid.New().String() + fmt.Printf("生成任务ID: %s\n", taskID) + // 发送run-task指令 + runTaskCmd := map[string]interface{}{ + "header": map[string]interface{}{ + "action": "run-task", + "task_id": taskID, + "streaming": "duplex", + }, + "payload": map[string]interface{}{ + "task_group": "audio", + "task": "tts", + "function": "SpeechSynthesizer", + "model": "cosyvoice-v3-flash", + "parameters": map[string]interface{}{ + "text_type": "PlainText", + "voice": "longanyang", + "format": "mp3", + "sample_rate": 22050, + "volume": 50, + "rate": 1, + "pitch": 1, + // 如果enable_ssml设为true,只允许发送一次continue-task指令,否则会报错“Text request limit violated, expected 1.” + "enable_ssml": true, + }, + "input": map[string]interface{}{}, + }, + } + runTaskJSON, _ := json.Marshal(runTaskCmd) + fmt.Printf("发送run-task指令: %s\n", string(runTaskJSON)) + err = conn.WriteMessage(websocket.TextMessage, runTaskJSON) + if err != nil { + fmt.Println("发送run-task失败:", err) + return + } + textSent := false + // 处理消息 + for { + messageType, message, err := conn.ReadMessage() + if err != nil { + fmt.Println("读取消息失败:", err) + break + } + // 处理二进制消息 + if messageType == websocket.BinaryMessage { + fmt.Printf("收到二进制消息,长度: %d\n", len(message)) + file, _ := os.OpenFile(outputFile, os.O_APPEND|os.O_WRONLY|os.O_CREATE, 0644) + file.Write(message) + file.Close() + continue + } + // 处理文本消息 + messageStr := string(message) + fmt.Printf("收到文本消息: %s\n", strings.ReplaceAll(messageStr, "\n", "")) + // 简单解析JSON获取event类型 + var msgMap map[string]interface{} + if json.Unmarshal(message, &msgMap) == nil { + if header, ok := msgMap["header"].(map[string]interface{}); ok { + if event, ok := header["event"].(string); ok { + fmt.Printf("事件类型: %s\n", event) + switch event { + case "task-started": + fmt.Println("=== 收到task-started事件 ===") + if !textSent { + // 发送 continue-task 指令,使用SSML功能时,该指令只允许发送一次 + continueTaskCmd := map[string]interface{}{ + "header": map[string]interface{}{ + "action": "continue-task", + "task_id": taskID, + "streaming": "duplex", + }, + "payload": map[string]interface{}{ + "input": map[string]interface{}{ + // 特殊字符需要进行转义 + "text": "我的语速比正常人快。", + }, + }, + } + continueTaskJSON, _ := json.Marshal(continueTaskCmd) + fmt.Printf("发送continue-task指令: %s\n", string(continueTaskJSON)) + err = conn.WriteMessage(websocket.TextMessage, continueTaskJSON) + if err != nil { + fmt.Println("发送continue-task失败:", err) + return + } + textSent = true + // 延迟发送finish-task + time.Sleep(500 * time.Millisecond) + // 发送finish-task指令 + finishTaskCmd := map[string]interface{}{ + "header": map[string]interface{}{ + "action": "finish-task", + "task_id": taskID, + "streaming": "duplex", + }, + "payload": map[string]interface{}{ + "input": map[string]interface{}{}, + }, + } + finishTaskJSON, _ := json.Marshal(finishTaskCmd) + fmt.Printf("发送finish-task指令: %s\n", string(finishTaskJSON)) + err = conn.WriteMessage(websocket.TextMessage, finishTaskJSON) + if err != nil { + fmt.Println("发送finish-task失败:", err) + return + } + } + case "task-finished": + fmt.Println("=== 任务完成 ===") + return + case "task-failed": + fmt.Println("=== 任务失败 ===") + if header["error_message"] != nil { + fmt.Printf("错误信息: %s\n", header["error_message"]) + } + return + case "result-generated": + fmt.Println("收到result-generated事件") + } + } + } + } + } + } + ``` + + + + ```csharp + using System.Net.WebSockets; + using System.Text; + using System.Text.Json; + // SSML功能说明: + // 1. 在发送run-task指令时,将参数enable_ssml设置为true,以开启SSML支持 + // 2. 通过continue-task指令发送包含SSML的文本,且只允许发送一次continue-task指令 + // 3. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + class Program { + // 若没有配置环境变量,请用API Key将下行替换为:private static readonly string ApiKey = "sk-xxx" + private static readonly string ApiKey = Environment.GetEnvironmentVariable("DASHSCOPE_API_KEY") ?? throw new InvalidOperationException("DASHSCOPE_API_KEY environment variable is not set."); + private const string WebSocketUrl = "wss://maas.qianwenaiapi.com/api-ws/v1/inference/"; + // 输出文件路径 + private const string OutputFilePath = "output.mp3"; + // WebSocket客户端 + private static ClientWebSocket _webSocket = new ClientWebSocket(); + // 取消令牌源 + private static CancellationTokenSource _cancellationTokenSource = new CancellationTokenSource(); + // 任务ID + private static string? _taskId; + // 任务是否已启动 + private static TaskCompletionSource _taskStartedTcs = new TaskCompletionSource(); + static async Task Main(string[] args) { + try { + // 清空输出文件 + ClearOutputFile(OutputFilePath); + // 连接WebSocket服务 + await ConnectToWebSocketAsync(WebSocketUrl); + // 启动接收消息的任务 + Task receiveTask = ReceiveMessagesAsync(); + // 发送run-task指令 + _taskId = GenerateTaskId(); + await SendRunTaskCommandAsync(_taskId); + // 等待task-started事件 + await _taskStartedTcs.Task; + // 发送 continue-task 指令,使用SSML功能时,该指令只允许发送一次 + // 特殊字符需要进行转义 + await SendContinueTaskCommandAsync("我的语速比正常人快。"); + // 发送finish-task指令 + await SendFinishTaskCommandAsync(_taskId); + // 等待接收任务完成 + await receiveTask; + Console.WriteLine("任务完成,连接已关闭。"); + } catch (OperationCanceledException) { + Console.WriteLine("任务被取消。"); + } catch (Exception ex) { + Console.WriteLine($"发生错误:{ex.Message}"); + } finally { + _cancellationTokenSource.Cancel(); + _webSocket.Dispose(); + } + } + private static void ClearOutputFile(string filePath) { + if (File.Exists(filePath)) { + File.WriteAllText(filePath, string.Empty); + Console.WriteLine("输出文件已清空。"); + } else { + Console.WriteLine("输出文件不存在,无需清空。"); + } + } + private static async Task ConnectToWebSocketAsync(string url) { + var uri = new Uri(url); + if (_webSocket.State == WebSocketState.Connecting || _webSocket.State == WebSocketState.Open) { + return; + } + // 设置WebSocket连接的头部信息 + _webSocket.Options.SetRequestHeader("Authorization", $"bearer {ApiKey}"); + _webSocket.Options.SetRequestHeader("X-DashScope-DataInspection", "enable"); + try { + await _webSocket.ConnectAsync(uri, _cancellationTokenSource.Token); + Console.WriteLine("已成功连接到WebSocket服务。"); + } catch (OperationCanceledException) { + Console.WriteLine("WebSocket连接被取消。"); + } catch (Exception ex) { + Console.WriteLine($"WebSocket连接失败: {ex.Message}"); + throw; + } + } + private static async Task SendRunTaskCommandAsync(string taskId) { + var command = CreateCommand("run-task", taskId, "duplex", new { + task_group = "audio", + task = "tts", + function = "SpeechSynthesizer", + model = "cosyvoice-v3-flash", + parameters = new + { + text_type = "PlainText", + voice = "longanyang", + format = "mp3", + sample_rate = 22050, + volume = 50, + rate = 1, + pitch = 1, + // 如果enable_ssml设为true,只允许发送一次continue-task指令,否则会报错“Text request limit violated, expected 1.” + enable_ssml = true + }, + input = new { } + }); + await SendJsonMessageAsync(command); + Console.WriteLine("已发送run-task指令。"); + } + private static async Task SendContinueTaskCommandAsync(string text) { + if (_taskId == null) { + throw new InvalidOperationException("任务ID未初始化。"); + } + var command = CreateCommand("continue-task", _taskId, "duplex", new { + input = new { + text + } + }); + await SendJsonMessageAsync(command); + Console.WriteLine("已发送continue-task指令。"); + } + private static async Task SendFinishTaskCommandAsync(string taskId) { + var command = CreateCommand("finish-task", taskId, "duplex", new { + input = new { } + }); + await SendJsonMessageAsync(command); + Console.WriteLine("已发送finish-task指令。"); + } + private static async Task SendJsonMessageAsync(string message) { + var buffer = Encoding.UTF8.GetBytes(message); + try { + await _webSocket.SendAsync(new ArraySegment(buffer), WebSocketMessageType.Text, true, _cancellationTokenSource.Token); + } catch (OperationCanceledException) { + Console.WriteLine("消息发送被取消。"); + } + } + private static async Task ReceiveMessagesAsync() { + while (_webSocket.State == WebSocketState.Open) { + var response = await ReceiveMessageAsync(); + if (response != null) { + var eventStr = response.RootElement.GetProperty("header").GetProperty("event").GetString(); + switch (eventStr) { + case "task-started": + Console.WriteLine("任务已启动。"); + _taskStartedTcs.TrySetResult(true); + break; + case "task-finished": + Console.WriteLine("任务已完成。"); + _cancellationTokenSource.Cancel(); + break; + case "task-failed": + Console.WriteLine("任务失败:" + response.RootElement.GetProperty("header").GetProperty("error_message").GetString()); + _cancellationTokenSource.Cancel(); + break; + default: + // result-generated可在此处理 + break; + } + } + } + } + private static async Task ReceiveMessageAsync() { + var buffer = new byte[1024 * 4]; + var segment = new ArraySegment(buffer); + try { + WebSocketReceiveResult result = await _webSocket.ReceiveAsync(segment, _cancellationTokenSource.Token); + if (result.MessageType == WebSocketMessageType.Close) { + await _webSocket.CloseAsync(WebSocketCloseStatus.NormalClosure, "Closing", _cancellationTokenSource.Token); + return null; + } + if (result.MessageType == WebSocketMessageType.Binary) { + // 处理二进制数据 + Console.WriteLine("接收到二进制数据..."); + // 将二进制数据保存到文件 + using (var fileStream = new FileStream(OutputFilePath, FileMode.Append)) { + fileStream.Write(buffer, 0, result.Count); + } + return null; + } + string message = Encoding.UTF8.GetString(buffer, 0, result.Count); + return JsonDocument.Parse(message); + } catch (OperationCanceledException) { + Console.WriteLine("消息接收被取消。"); + return null; + } + } + private static string GenerateTaskId() { + return Guid.NewGuid().ToString("N").Substring(0, 32); + } + private static string CreateCommand(string action, string taskId, string streaming, object payload) { + var command = new { + header = new { + action, + task_id = taskId, + streaming + }, + payload + }; + return JsonSerializer.Serialize(command); + } + } + ``` + + + + 示例代码目录结构为: + + my-php-project/ + + ├── composer.json + + ├── vendor/ + + └── index.php + + composer.json内容如下,相关依赖的版本号请根据实际情况自行决定: + + ```json + { + "require": { + "react/event-loop": "^1.3", + "react/socket": "^1.11", + "react/stream": "^1.2", + "react/http": "^1.1", + "ratchet/pawl": "^0.4" + }, + "autoload": { + "psr-4": { + "App\\": "src/" + } + } + } + ``` + + index.php内容如下: + + ```php + + + + + [ + 'bindto' => '0.0.0.0:0', + ], + 'tls' => [ + 'verify_peer' => false, + 'verify_peer_name' => false, + ], + ]); + $connector = new Connector($loop, $socketConnector); + $headers = [ + 'Authorization' => 'bearer ' . $api_key, + 'X-DashScope-DataInspection' => 'enable' + ]; + $connector($websocket_url, [], $headers)->then(function ($conn) use ($loop, $output_file) { + echo "连接到WebSocket服务器\n"; + // 生成任务ID + $taskId = generateTaskId(); + // 发送 run-task 指令 + sendRunTaskMessage($conn, $taskId); + // 定义发送 continue-task 指令的函数 + $sendContinueTask = function() use ($conn, $loop, $taskId) { + // 发送 continue-task 指令,使用SSML功能时,该指令只允许发送一次 + $continueTaskMessage = json_encode([ + "header" => [ + "action" => "continue-task", + "task_id" => $taskId, + "streaming" => "duplex" + ], + "payload" => [ + "input" => [ + // 特殊字符需要进行转义 + "text" => "我的语速比正常人快。" + ] + ] + ]); + $conn->send($continueTaskMessage); + // 发送 finish-task 指令 + sendFinishTaskMessage($conn, $taskId); + }; + // 标记是否收到 task-started 事件 + $taskStarted = false; + // 监听消息 + $conn->on('message', function($msg) use ($conn, $sendContinueTask, $loop, &$taskStarted, $taskId, $output_file) { + if ($msg->isBinary()) { + // 写入二进制数据到本地文件 + file_put_contents($output_file, $msg->getPayload(), FILE_APPEND); + } else { + // 处理非二进制消息 + $response = json_decode($msg, true); + if (isset($response['header']['event'])) { + handleEvent($conn, $response, $sendContinueTask, $loop, $taskId, $taskStarted); + } else { + echo "未知的消息格式\n"; + } + } + }); + // 监听连接关闭 + $conn->on('close', function($code = null, $reason = null) { + echo "连接已关闭\n"; + if ($code !== null) { + echo "关闭代码: " . $code . "\n"; + } + if ($reason !== null) { + echo "关闭原因:" . $reason . "\n"; + } + }); + }, function ($e) { + echo "无法连接:{$e->getMessage()}\n"; + }); + $loop->run(); + /** + * 生成任务ID + * @return string + */ + function generateTaskId(): string { + return bin2hex(random_bytes(16)); + } + /** + * 发送 run-task 指令 + * @param $conn + * @param $taskId + */ + function sendRunTaskMessage($conn, $taskId) { + $runTaskMessage = json_encode([ + "header" => [ + "action" => "run-task", + "task_id" => $taskId, + "streaming" => "duplex" + ], + "payload" => [ + "task_group" => "audio", + "task" => "tts", + "function" => "SpeechSynthesizer", + "model" => "cosyvoice-v3-flash", + "parameters" => [ + "text_type" => "PlainText", + "voice" => "longanyang", + "format" => "mp3", + "sample_rate" => 22050, + "volume" => 50, + "rate" => 1, + "pitch" => 1, + // 如果enable_ssml设为true,只允许发送一次continue-task指令,否则会报错“Text request limit violated, expected 1.” + "enable_ssml" => true + ], + "input" => (object) [] + ] + ]); + echo "准备发送run-task指令: " . $runTaskMessage . "\n"; + $conn->send($runTaskMessage); + echo "run-task指令已发送\n"; + } + /** + * 读取音频文件 + * @param string $filePath + * @return bool|string + */ + function readAudioFile(string $filePath) { + $voiceData = file_get_contents($filePath); + if ($voiceData === false) { + echo "无法读取音频文件\n"; + } + return $voiceData; + } + /** + * 分割音频数据 + * @param string $data + * @param int $chunkSize + * @return array + */ + function splitAudioData(string $data, int $chunkSize): array { + return str_split($data, $chunkSize); + } + /** + * 发送 finish-task 指令 + * @param $conn + * @param $taskId + */ + function sendFinishTaskMessage($conn, $taskId) { + $finishTaskMessage = json_encode([ + "header" => [ + "action" => "finish-task", + "task_id" => $taskId, + "streaming" => "duplex" + ], + "payload" => [ + "input" => (object) [] + ] + ]); + echo "准备发送finish-task指令: " . $finishTaskMessage . "\n"; + $conn->send($finishTaskMessage); + echo "finish-task指令已发送\n"; + } + /** + * 处理事件 + * @param $conn + * @param $response + * @param $sendContinueTask + * @param $loop + * @param $taskId + * @param $taskStarted + */ + function handleEvent($conn, $response, $sendContinueTask, $loop, $taskId, &$taskStarted) { + switch ($response['header']['event']) { + case 'task-started': + echo "任务开始,发送continue-task指令...\n"; + $taskStarted = true; + // 发送 continue-task 指令 + $sendContinueTask(); + break; + case 'result-generated': + // 忽略result-generated事件 + break; + case 'task-finished': + echo "任务完成\n"; + $conn->close(); + break; + case 'task-failed': + echo "任务失败\n"; + echo "错误代码:" . $response['header']['error_code'] . "\n"; + echo "错误信息:" . $response['header']['error_message'] . "\n"; + $conn->close(); + break; + case 'error': + echo "错误:" . $response['payload']['message'] . "\n"; + break; + default: + echo "未知事件:" . $response['header']['event'] . "\n"; + break; + } + // 如果任务已完成,关闭连接 + if ($response['header']['event'] == 'task-finished') { + // 等待1秒以确保所有数据都已传输完毕 + $loop->addTimer(1, function() use ($conn) { + $conn->close(); + echo "客户端关闭连接\n"; + }); + } + // 如果没有收到 task-started 事件,关闭连接 + if (!$taskStarted && in_array($response['header']['event'], ['task-failed', 'error'])) { + $conn->close(); + } + } + ``` + + + + 需安装相关依赖: + + ```bash + npm install ws + npm install uuid + ``` + + 示例代码如下: + + ```javascript + // SSML功能说明: + // 1. 在发送run-task指令时,将参数enable_ssml设置为true,以开启SSML支持 + // 2. 通过continue-task指令发送包含SSML的文本,且只允许发送一次continue-task指令 + // 3. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + import fs from 'fs'; + import WebSocket from 'ws'; + import { v4 as uuid } from 'uuid'; // 用于生成UUID + // 若没有配置环境变量,请用API Key将下行替换为:const apiKey = "sk-xxx" + const apiKey = process.env.DASHSCOPE_API_KEY; + const url = 'wss://maas.qianwenaiapi.com/api-ws/v1/inference/'; + // 输出文件路径 + const outputFilePath = 'output.mp3'; + // 清空输出文件 + fs.writeFileSync(outputFilePath, ''); + // 创建WebSocket客户端 + const ws = new WebSocket(url, { + headers: { + Authorization: `bearer ${apiKey}`, + 'X-DashScope-DataInspection': 'enable' + } + }); + let taskStarted = false; + let taskId = uuid(); + ws.on('open', () => { + console.log('已连接到WebSocket服务器'); + // 发送run-task指令 + const runTaskMessage = JSON.stringify({ + header: { + action: 'run-task', + task_id: taskId, + streaming: 'duplex' + }, + payload: { + task_group: 'audio', + task: 'tts', + function: 'SpeechSynthesizer', + model: 'cosyvoice-v3-flash', + parameters: { + text_type: 'PlainText', + voice: 'longanyang', // 音色 + format: 'mp3', // 音频格式 + sample_rate: 22050, // 采样率 + volume: 50, // 音量 + rate: 1, // 语速 + pitch: 1, // 音调 + enable_ssml: true // 是否开启SSML功能。如果enable_ssml设为true,只允许发送一次continue-task指令,否则会报错“Text request limit violated, expected 1.” + }, + input: {} + } + }); + ws.send(runTaskMessage); + console.log('已发送run-task消息'); + }); + const fileStream = fs.createWriteStream(outputFilePath, { flags: 'a' }); + ws.on('message', (data, isBinary) => { + if (isBinary) { + // 写入二进制数据到文件 + fileStream.write(data); + } else { + const message = JSON.parse(data); + switch (message.header.event) { + case 'task-started': + taskStarted = true; + console.log('任务已开始'); + // 发送continue-task指令 + sendContinueTasks(ws); + break; + case 'task-finished': + console.log('任务已完成'); + ws.close(); + fileStream.end(() => { + console.log('文件流已关闭'); + }); + break; + case 'task-failed': + console.error('任务失败:', message.header.error_message); + ws.close(); + fileStream.end(() => { + console.log('文件流已关闭'); + }); + break; + default: + // 可以在这里处理result-generated + break; + } + } + }); + function sendContinueTasks(ws) { + if (taskStarted) { + // 发送 continue-task 指令,使用SSML功能时,该指令只允许发送一次 + const continueTaskMessage = JSON.stringify({ + header: { + action: 'continue-task', + task_id: taskId, + streaming: 'duplex' + }, + payload: { + input: { + // 特殊字符需要进行转义 + text: '我的语速比正常人快。' + } + } + }); + ws.send(continueTaskMessage); + // 发送finish-task指令 + const finishTaskMessage = JSON.stringify({ + header: { + action: 'finish-task', + task_id: taskId, + streaming: 'duplex' + }, + payload: { + input: {} + } + }); + ws.send(finishTaskMessage); + } + } + ws.on('close', () => { + console.log('已断开与WebSocket服务器的连接'); + }); + ``` + + + + 如您使用Java编程语言,建议采用Java DashScope SDK进行开发,详情请参见[Java SDK](/api-reference/speech-synthesis/cosyvoice/java-sdk)。 + + 以下是Java WebSocket的调用示例。在运行示例前,请确保已导入以下依赖: + + - `Java-WebSocket` + - `jackson-databind` + + 推荐您使用Maven或Gradle管理依赖包,其配置如下: + + + ```xml pom.xml + + + + org.java-websocket + Java-WebSocket + 1.5.3 + + + + com.fasterxml.jackson.core + jackson-databind + 2.13.0 + + + ``` + + ```gradle build.gradle + // 省略其它代码 + dependencies { + // WebSocket Client + implementation 'org.java-websocket:Java-WebSocket:1.5.3' + // JSON Processing + implementation 'com.fasterxml.jackson.core:jackson-databind:2.13.0' + } + // 省略其它代码 + ``` + + + Java代码如下: + + ```java + import com.fasterxml.jackson.databind.ObjectMapper; + import org.java_websocket.client.WebSocketClient; + import org.java_websocket.handshake.ServerHandshake; + import java.io.FileOutputStream; + import java.io.IOException; + import java.net.URI; + import java.nio.ByteBuffer; + import java.util.*; + /** + * SSML功能说明: + * 1. 在发送run-task指令时,将参数enable_ssml设置为true,以开启SSML支持 + * 2. 通过continue-task指令发送包含SSML的文本,且只允许发送一次continue-task指令 + * 3. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + */ + public class TTSWebSocketClient extends WebSocketClient { + private final String taskId = UUID.randomUUID().toString(); + private final String outputFile = "output_" + System.currentTimeMillis() + ".mp3"; + private boolean taskFinished = false; + public TTSWebSocketClient(URI serverUri, Map headers) { + super(serverUri, headers); + } + @Override + public void onOpen(ServerHandshake serverHandshake) { + System.out.println("连接成功"); + // 发送run-task指令 + // 如果enable_ssml设为true,只允许发送一次continue-task指令,否则会报错“Text request limit violated, expected 1.” + String runTaskCommand = "{ \"header\": { \"action\": \"run-task\", \"task_id\": \"" + taskId + "\", \"streaming\": \"duplex\" }, \"payload\": { \"task_group\": \"audio\", \"task\": \"tts\", \"function\": \"SpeechSynthesizer\", \"model\": \"cosyvoice-v3-flash\", \"parameters\": { \"text_type\": \"PlainText\", \"voice\": \"longanyang\", \"format\": \"mp3\", \"sample_rate\": 22050, \"volume\": 50, \"rate\": 1, \"pitch\": 1, \"enable_ssml\": true }, \"input\": {} }}"; + send(runTaskCommand); + } + @Override + public void onMessage(String message) { + System.out.println("收到服务端返回的消息:" + message); + try { + // Parse JSON message + Map messageMap = new ObjectMapper().readValue(message, Map.class); + if (messageMap.containsKey("header")) { + Map header = (Map) messageMap.get("header"); + if (header.containsKey("event")) { + String event = (String) header.get("event"); + if ("task-started".equals(event)) { + System.out.println("收到服务端返回的task-started事件"); + // 发送 continue-task 指令,使用SSML功能时,该指令只允许发送一次 + // 特殊字符需要进行转义 + sendContinueTask("我的语速比正常人快。"); + // 发送finish-task指令 + sendFinishTask(); + } else if ("task-finished".equals(event)) { + System.out.println("收到服务端返回的task-finished事件"); + taskFinished = true; + closeConnection(); + } else if ("task-failed".equals(event)) { + System.out.println("任务失败:" + message); + closeConnection(); + } + } + } + } catch (Exception e) { + System.err.println("出现异常:" + e.getMessage()); + } + } + @Override + public void onMessage(ByteBuffer message) { + System.out.println("收到的二进制音频数据大小为:" + message.remaining()); + try (FileOutputStream fos = new FileOutputStream(outputFile, true)) { + byte[] buffer = new byte[message.remaining()]; + message.get(buffer); + fos.write(buffer); + System.out.println("音频数据已写入本地文件" + outputFile + "中"); + } catch (IOException e) { + System.err.println("音频数据写入本地文件失败:" + e.getMessage()); + } + } + @Override + public void onClose(int code, String reason, boolean remote) { + System.out.println("连接关闭:" + reason + " (" + code + ")"); + } + @Override + public void onError(Exception ex) { + System.err.println("报错:" + ex.getMessage()); + ex.printStackTrace(); + } + private void sendContinueTask(String text) { + String command = "{ \"header\": { \"action\": \"continue-task\", \"task_id\": \"" + taskId + "\", \"streaming\": \"duplex\" }, \"payload\": { \"input\": { \"text\": \"" + text + "\" } }}"; + send(command); + } + private void sendFinishTask() { + String command = "{ \"header\": { \"action\": \"finish-task\", \"task_id\": \"" + taskId + "\", \"streaming\": \"duplex\" }, \"payload\": { \"input\": {} }}"; + send(command); + } + private void closeConnection() { + if (!isClosed()) { + close(); + } + } + public static void main(String[] args) { + try { + // 若没有配置环境变量,请用API Key将下行替换为:String apiKey = "sk-xxx" + String apiKey = System.getenv("DASHSCOPE_API_KEY"); + if (apiKey == null || apiKey.isEmpty()) { + System.err.println("请设置 DASHSCOPE_API_KEY 环境变量"); + return; + } + Map headers = new HashMap<>(); + headers.put("Authorization", "bearer " + apiKey); + TTSWebSocketClient client = new TTSWebSocketClient(new URI("wss://maas.qianwenaiapi.com/api-ws/v1/inference/"), headers); + client.connect(); + while (!client.isClosed() && !client.taskFinished) { + Thread.sleep(1000); + } + } catch (Exception e) { + System.err.println("连接WebSocket服务失败:" + e.getMessage()); + e.printStackTrace(); + } + } + } + ``` + + + + 如您使用Python编程语言,建议采用Python DashScope SDK进行开发,详情请参见[Python SDK](/api-reference/speech-synthesis/cosyvoice/python-sdk)。 + + 以下是Python WebSocket的调用示例。在运行示例前,请确保通过如下方式导入依赖: + + ```bash + pip uninstall websocket-client + pip uninstall websocket + pip install websocket-client + ``` + + + 请不要将运行示例代码的Python文件命名为"websocket.py",否则会报错(AttributeError: module 'websocket' has no attribute 'WebSocketApp'. Did you mean: 'WebSocket'?)。 + + + ```python + # SSML功能说明: + # 1. 在发送run-task指令时,将参数enable_ssml设置为true,以开启SSML支持 + # 2. 通过continue-task指令发送包含SSML的文本,且只允许发送一次continue-task指令 + # 3. 只有qwen-audio-3.0-tts-flash、qwen-audio-3.0-tts-plus、cosyvoice-v3-flash、cosyvoice-v3-plus和cosyvoice-v2模型的复刻音色以及音色列表中标记为支持SSML的系统音色支持SSML功能(例如cosyvoice-v3-flash模型的longanyang音色) + import websocket + import json + import uuid + import os + import time + class TTSClient: + def __init__(self, api_key, uri): + """ + 初始化 TTSClient 实例 + 参数: + api_key (str): 鉴权用的 API Key + uri (str): WebSocket 服务地址 + """ + self.api_key = api_key # 替换为你的 API Key + self.uri = uri # 替换为你的 WebSocket 地址 + self.task_id = str(uuid.uuid4()) # 生成唯一任务 ID + self.output_file = f"output_{int(time.time())}.mp3" # 输出音频文件路径 + self.ws = None # WebSocketApp 实例 + self.task_started = False # 是否收到 task-started + self.task_finished = False # 是否收到 task-finished / task-failed + def on_open(self, ws): + """ + WebSocket 连接建立时回调函数 + 发送 run-task 指令开启语音合成任务 + """ + print("WebSocket 已连接") + # 构造 run-task 指令 + run_task_cmd = { + "header": { + "action": "run-task", + "task_id": self.task_id, + "streaming": "duplex" + }, + "payload": { + "task_group": "audio", + "task": "tts", + "function": "SpeechSynthesizer", + "model": "cosyvoice-v3-flash", + "parameters": { + "text_type": "PlainText", + "voice": "longanyang", + "format": "mp3", + "sample_rate": 22050, + "volume": 50, + "rate": 1, + "pitch": 1, + # 如果enable_ssml设为True,只允许发送一次continue-task指令,否则会报错“Text request limit violated, expected 1.” + "enable_ssml": True + }, + "input": {} + } + } + # 发送 run-task 指令 + ws.send(json.dumps(run_task_cmd)) + print("已发送 run-task 指令") + def on_message(self, ws, message): + """ + 接收到消息时的回调函数 + 区分文本和二进制消息处理 + """ + if isinstance(message, str): + # 处理 JSON 文本消息 + try: + msg_json = json.loads(message) + print(f"收到 JSON 消息: {msg_json}") + if "header" in msg_json: + header = msg_json["header"] + if "event" in header: + event = header["event"] + if event == "task-started": + print("任务已启动") + self.task_started = True + # 发送 continue-task 指令,使用SSML功能时,该指令只允许发送一次 + # 特殊字符需要进行转义 + self.send_continue_task("我的语速比正常人快。") + # continue-task 发送完成后发送 finish-task + self.send_finish_task() + elif event == "task-finished": + print("任务已完成") + self.task_finished = True + self.close(ws) + elif event == "task-failed": + error_msg = msg_json.get("error_message", "未知错误") + print(f"任务失败: {error_msg}") + self.task_finished = True + self.close(ws) + except json.JSONDecodeError as e: + print(f"JSON 解析失败: {e}") + else: + # 处理二进制消息(音频数据) + print(f"收到二进制消息,大小: {len(message)} 字节") + with open(self.output_file, "ab") as f: + f.write(message) + print(f"已将音频数据写入本地文件{self.output_file}中") + def on_error(self, ws, error): + """发生错误时的回调""" + print(f"WebSocket 出错: {error}") + def on_close(self, ws, close_status_code, close_msg): + """连接关闭时的回调""" + print(f"WebSocket 已关闭: {close_msg} ({close_status_code})") + def send_continue_task(self, text): + """发送 continue-task 指令,附带要合成的文本内容""" + cmd = { + "header": { + "action": "continue-task", + "task_id": self.task_id, + "streaming": "duplex" + }, + "payload": { + "input": { + "text": text + } + } + } + self.ws.send(json.dumps(cmd)) + print(f"已发送 continue-task 指令,文本内容: {text}") + def send_finish_task(self): + """发送 finish-task 指令,结束语音合成任务""" + cmd = { + "header": { + "action": "finish-task", + "task_id": self.task_id, + "streaming": "duplex" + }, + "payload": { + "input": {} + } + } + self.ws.send(json.dumps(cmd)) + print("已发送 finish-task 指令") + def close(self, ws): + """主动关闭连接""" + if ws and ws.sock and ws.sock.connected: + ws.close() + print("已主动关闭连接") + def run(self): + """启动 WebSocket 客户端""" + # 设置请求头部(鉴权) + header = { + "Authorization": f"bearer {self.api_key}", + "X-DashScope-DataInspection": "enable" + } + # 创建 WebSocketApp 实例 + self.ws = websocket.WebSocketApp( + self.uri, + header=header, + on_open=self.on_open, + on_message=self.on_message, + on_error=self.on_error, + on_close=self.on_close + ) + print("正在监听 WebSocket 消息...") + self.ws.run_forever() # 启动长连接监听 + # 示例使用方式 + if __name__ == "__main__": + # 若没有配置环境变量,请用API Key将下行替换为:API_KEY = "sk-xxx" + API_KEY = os.environ.get("DASHSCOPE_API_KEY") + SERVER_URI = "wss://maas.qianwenaiapi.com/api-ws/v1/inference/" + client = TTSClient(API_KEY, SERVER_URI) + client.run() + ``` + + + + + + ```bash + curl --location 'https://maas.qianwenaiapi.com/api/v1/services/aigc/text-generation/generation' \ + --header "Authorization: Bearer $DASHSCOPE_API_KEY" \ + --header 'Content-Type: application/json' \ + --header 'X-DashScope-DataInspection: enable' \ + --data '{ + "model": "cosyvoice-v3-flash", + "input": { + "text": "我的语速比正常人快。" + }, + "parameters": { + "voice": "longanyang", + "format": "mp3" + } + }' + ``` + + + +### 标签参考 + + + CosyVoice SSML 基于 [W3C SSML 1.0](https://www.w3.org/TR/speech-synthesis/),仅支持部分标签。 + + **语法规则**: + + - 所有 SSML 内容必须包裹在 `` 标签中。 + - 可以连续使用多个 `` 标签,但不能嵌套。 + - 需要转义 XML 特殊字符:`"` → `"`,`'` → `'`,`&` → `&`,`<` → `<`,`>` → `>`。 + + +#### ``:根标签 + +**说明** + +所有 SSML 内容必须包裹在 `` 标签中。 + +**语法** + +```xml +需要使用 SSML 功能的文本 +``` + +**属性** + +| 属性 | 类型 | 必填 | 说明 | +| --------------------- | ------ | -- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| voice | String | 否 | 音色名称。覆盖 API 参数 `voice`。参见[音色列表](/developer-guides/speech/voice-list/cosyvoice)。 | +| rate | String | 否 | 语速。覆盖 API 参数 `speech_rate`。推荐范围:0.5 \~ 2,默认值 1。大于 1 加速,小于 1 减速。 | +| pitch | String | 否 | 音调。覆盖 API 参数 `pitch_rate`。推荐范围:0.5 \~ 2,默认值 1。大于 1 升高,小于 1 降低。 | +| volume | String | 否 | 音量。覆盖 API 参数 `volume`。取值范围:0 \~ 100,默认值 50。 | +| effect | String | 否 | 音效。可选值:`robot`、`lolita`(活泼女声)、`lowpass`、`echo`、`eq`(均衡器,高级)、`lpfilter`(低通滤波器,高级)、`hpfilter`(高通滤波器,高级)。`eq`、`lpfilter`、`hpfilter` 需配合 `effectValue` 使用。每个标签只能设置一种音效。音效会增加延迟。 | +| effectValue | String | 否 | 自定义 `effect` 参数。`eq`:8 个以空格分隔的整数(-20 \~ 20),分别对应 `["40 Hz", "100 Hz", "200 Hz", "400 Hz", "800 Hz", "1600 Hz", "4000 Hz", "12000 Hz"]` 频段的增益,示例:`"1 1 1 1 1 1 1 1"`。`lpfilter`:整数频率,范围 (0, sample\_rate/2],示例:`"800"`。`hpfilter`:整数频率,范围 (0, sample\_rate/2],示例:`"1200"`。 | +| bgm | String | 否 | 背景音乐 URL。文件需存放在 OSS 上,权限至少为公共读。URL 中的 XML 特殊字符需转义。要求:16 kHz 采样率、单声道、WAV 格式、16-bit。如果合成音频长于背景音乐,音乐将循环播放。 | +| backgroundMusicVolume | String | 否 | 背景音乐音量。 | + +**示例** + +音色: + +```xml + + 我是男声。 + +``` + +语速: + +```xml + + 我的语速比正常人快。 + +``` + +音调: + +```xml + + 但是我的音调比别人低。 + +``` + +音量: + +```xml + + 我的音量也很高。 + +``` + +音效: + +```xml + + 你喜欢机器人瓦力吗? + +``` + +音效 + effectValue: + +```xml + + 你喜欢机器人瓦力吗? + + + + 你喜欢机器人瓦力吗? + + + + 你喜欢机器人瓦力吗? + +``` + +如果音频不是 WAV 格式,可使用 `ffmpeg` 转换: + +```bash +ffmpeg -i input_audio -acodec pcm_s16le -ac 1 -ar 16000 output.wav +``` + +背景音乐(bgm): + +```xml + + + 阴崖老木苍苍烟 + + 雨声犹在竹林间 + + 绵蕝固知裨国计 + + 绵州风物总堪怜 + + +``` + + + 上传音频的版权由您自行承担法律责任。 + + +组合属性(空格分隔): + +```xml + + 所以放在一起,我的声音是这样的。 + +``` + +#### ``:停顿 + +**说明** + +插入一段停顿。时长单位为秒(s)或毫秒(ms)。 + +**语法** + +```xml +# 无属性 + +# 带 time 属性 + +``` + + + **break 标签行为**: + + - 不带属性时,`` 默认停顿 1 秒。 + - **注意**:连续的 `` 标签时长会累加,但总时长上限为 10 秒。 + + 例如,以下三个标签总时长为 15 秒,但仅前 10 秒有效: + + ```xml + + 请闭上眼睛休息一下。好了,请睁开眼睛。 + + ``` + + +**属性** + +| 属性 | 类型 | 必填 | 说明 | +| ---- | ------ | -- | -------------------------------------------------------- | +| time | String | 否 | 停顿时长,如 `"2s"` 或 `"50ms"`。秒为单位:1 \~ 10。毫秒为单位:50 \~ 10000。 | + +**示例** + +```xml + + 请闭上眼睛休息一下。好了,请睁开眼睛。 + +``` + +#### ``:替换文本 + +**说明** + +将显示文本替换为其他发音。 + +**语法** + +```xml + +``` + +**属性** + +| 属性 | 类型 | 必填 | 说明 | +| ----- | ------ | -- | -------- | +| alias | String | 是 | 替代朗读的文本。 | + +**示例** + +```xml + + W3C + +``` + +#### ``:设置发音 + +**说明** + +使用拼音(中文)或 CMU 音标(英文)指定发音。 + +**语法** + +```xml +text +``` + +**属性** + +| 属性 | 类型 | 必填 | 说明 | +| -------- | ------ | -- | --------------------------------------------------------------------------------------------------------------- | +| alphabet | String | 是 | 发音类型:`"py"`(拼音)或 `"cmu"`(音标)。参见 [The CMU Pronouncing Dictionary](http://www.speech.cs.cmu.edu/cgi-bin/cmudict)。 | +| ph | String | 是 | 拼音或音标符号。每个汉字的拼音之间用空格分隔,音节数必须与字数一致。每个音节带声调号(1 \~ 5,其中 5 为轻声)。 | + +**示例** + +```xml + + 去典当行把这个玩意当掉 + + + + How to spell sin? + +``` + +#### ``:插入音效 + +**说明** + +在合成语音中插入外部音频文件(提示音、环境音等)。 + +**语法** + +```xml + +``` + +**属性** + +| 属性 | 类型 | 必填 | 说明 | +| --- | ------ | -- | ---------------------------------------------------------------------------------------- | +| src | String | 是 | 音频 URL。文件需存放在 OSS 上,权限至少为公共读。URL 中的 XML 特殊字符需转义。要求:16 kHz 采样率、单声道、WAV 格式、16-bit,最大 2 MB。 | + +如果音频不是 WAV 格式,可使用 `ffmpeg` 转换: + +```bash +ffmpeg -i input_audio -acodec pcm_s16le -ac 1 -ar 16000 output.wav +``` + + + 上传音频的版权由您自行承担法律责任。 + + +**示例** + +```xml + + 一匹马受了惊吓人们四散躲避 + +``` + +#### ``:设置朗读格式 + +**说明** + +指定文本的朗读方式(如数字、日期、电话号码等)。 + +**语法** + +```xml +text +``` + +**属性** + +| 属性 | 类型 | 必填 | 说明 | +| ------------ | ------ | -- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| interpret-as | String | 是 | 文本类型。可选值:`cardinal`(数字)、`digits`(逐位数字)、`telephone`(电话号码)、`name`(姓名)、`address`(地址)、`id`(账号名/昵称)、`characters`(逐字符)、`punctuation`(标点)、`date`(日期)、`time`(时间)、`currency`(货币)、`measure`(度量单位)。 | + +##### cardinal + +`cardinal` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| -------------- | ------- | -------- | -------------------------------------------------------------------------------------------------- | +| 数字串 | 145 | 一百四十五 | 整数输入范围:20位以内的正负整数,\[-99999999999999999999, 99999999999999999999]。小数输入范围:对小数点后小数的位数没有特殊限制,建议不超过10位。 | +| 负号+数字串 | -145 | 负一百四十五 | | +| 以逗号分隔3位数字串 | 10,000 | 一万 | | +| 负号+以逗号分隔3位数字串 | -10,124 | 负一万一百二十四 | | +| 数字串+小数点+2个零 | 10.00 | 十 | | +| 负号+数字串+小数点+2个零 | -110.00 | 负一百一十 | | +| 数字串+小数点+数字串 | 79.090 | 七十九点零九零 | | +| 负号+数字串+小数点+数字串 | -79.001 | 负七十九点零零一 | | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| ----------------------- | --------- | ---------------------------------------------- | ----------------------------------------------------------------------- | +| 纯数字 | 145 | one hundred forty five | 整数范围:最多 13 位,\[-999999999999, 999999999999]。小数:整数部分最多 13 位,小数部分最多 10 位。 | +| 零开头的数字 | 0145 | one hundred forty five | | +| 负号 + 数字 | -145 | minus hundred forty five | | +| 千分位逗号分隔的数字 | 60,000 | sixty thousand | | +| 负号 + 千分位逗号分隔的数字 | -208,000 | minus two hundred eight thousand | | +| 数字 + 小数点 + 零 | 12.00 | twelve | | +| 数字 + 小数点 + 数字 | 12.34 | twelve point three four | | +| 千分位逗号分隔 + 小数点 + 数字 | 1,000.1 | one thousand point one | | +| 负号 + 数字 + 小数点 + 数字 | -12.34 | minus twelve point three four | | +| 负号 + 千分位逗号分隔 + 小数点 + 数字 | -1,000.1 | minus one thousand point one | | +| 千分位数字 + 连字符 + 千分位数字 | 1-1,000 | one to one thousand | | +| 其他默认读法 | 012.34 | twelve point three four | | +| | 1/2 | one half | | +| | -3/4 | minus three quarters | | +| | 5.1/6 | five point one over six | | +| | -3 1/2 | minus three and a half | | +| | 1,000.3^3 | one thousand point three to the power of three | | +| | 3e9.1 | three times ten to the power of nine point one | | +| | 23.10% | twenty three point one percent | | + +**示例** + +```xml + + 12345 + +``` + +```xml + + 10234 + +``` + +##### digits + +`digits` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| --- | --------- | --------- | -------------------------------------------- | +| 数字串 | 129090909 | 一二九零九零九零九 | 对数字串的长度没有特殊限制,建议不超过20位。当数字串超过10位时,每个数字后插入停顿。 | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| ---------------------- | ------------- | ---------------------------------------------------- | --------------------- | +| 纯数字 | 12034 | one two zero three four | 无严格长度限制,建议不超过 20 个字符。 | +| 数字 + 空格或连字符 + 数字 + ... | 1-23-456 7890 | one, two three, four five six, seven eight nine zero | | + +**示例** + +```xml + + 12345 + +``` + +```xml + + 10234 + +``` + +##### telephone + +`telephone` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| --------------- | ----------------------- | ---------------------- | ---------------------------------------------------------------------------- | +| 座机号 | 4930286 | 四九三 零二八六 | 支持7\~8位座机号,支持空格和"-"作为分隔符。其中,7位座机号支持"3-4"的数字分隔方式;8位座机号支持"4-4"的数字分隔方式。 | +| | 493 0286 | 四九三 零二八六 | | +| | 493-0286 | 四九三 零二八六 | | +| | 62552560 | 六二五五 二五六零 | | +| | 6255 2560 | 六二五五 二五六零 | | +| | 6255-2560 | 六二五五 二五六零 | | +| 座机号+分机号 | 4930286-109 | 四九三 零二八六 转幺零九 | 支持1\~4位分机号。 | +| | 4930286转109 | 四九三 零二八六 转幺零九 | | +| | 4930286分机109 | 四九三 零二八六 分机幺零九 | | +| | 4930286分机号109 | 四九三 零二八六 分机号幺零九 | | +| 区号+座机号 | 01062552560 | 零幺零 六二五五 二五六零 | 支持区号:010、02x、03xx、04xx、05xx、07xx、08xx、09xx。 | +| | 010 62552560 | 零幺零 六二五五 二五六零 | | +| | 010 6255 2560 | 零幺零 六二五五 二五六零 | | +| | 010 6255-2560 | 零幺零 六二五五 二五六零 | | +| | 010-62552560 | 零幺零 六二五五 二五六零 | | +| | 010-6255-2560 | 零幺零 六二五五 二五六零 | | +| | (010)62552560 | 零幺零 六二五五 二五六零 | | +| | 03198907098 | 零三幺九 八九零 七零九八 | | +| | 0319-8907098 | 三幺九 八九零 七零九八 | | +| 区号+座机号+分机号 | 010 62552560-109 | 零幺零 六二五五 二五六零 转幺零九 | | +| | 010-62552560-109 | 零幺零 六二五五 二五六零 转幺零九 | | +| | (010)62552560-109 | 零幺零 六二五五 二五六零 转幺零九 | | +| | (010)62552560转109 | 零幺零 六二五五 二五六零 转幺零九 | | +| | (010)62552560分机109 | 零幺零 六二五五 二五六零 分机幺零九 | | +| | (010)62552560分机号109 | 零幺零 六二五五 二五六零 分机号幺零九 | | +| 国家代码+区号+座机号 | 86-010-62791627 | 八六 零幺零 六二七九 幺六二七 | 支持国家代码:86、(86)、+86、(+86)、0086,统一读为"八六"。 | +| | (86)10-62791627 | 八六 幺零 六二七九 幺六二七 | | +| | +86-010-62791627 | 八六 零幺零 六二七九 幺六二七 | | +| | 0086-10-62791627 | 八六 幺零 六二七九 幺六二七 | | +| | (+86)-10-6279 1627 | 八六 幺零 六二七九 幺六二七 | | +| 国家代码+区号+座机号+分机号 | (86)21-58118818-207 | 八六 二幺 五八幺幺 八八幺八 转二零七 | | +| | (86)021-5811-8818-207 | 八六 零二幺 五八幺幺 八八幺八 转二零七 | | +| | (86)021-58118818转207 | 八六 零二幺 五八幺幺 八八幺八 转二零七 | | +| | (86)21-5811-8818分机207 | 八六 二幺 五八幺幺 八八幺八 分机二零七 | | +| | +86-021-58118818分机号207 | 八六 零二幺 五八幺幺 八八幺八分机号二零七 | | +| 手机号 | 139 0000 5678 | 幺三九 零零零零 五六七八 | 支持11位手机号,支持3-3-5、3-4-4两种数字分隔方式。 | +| | 139-000-05678 | 幺三九 零零零 零五六七八 | | +| | 139 000 05678 | 幺三九 零零零 零五六七八 | | +| 国家代码+手机号 | +86-13900005678 | 八六 幺三九 零零零零 五六七八 | | +| | (+86)-139-0000-5678 | 八六 幺三九 零零零零 五六七八 | | +| | +8613900005678 | 八六 幺三九 零零零零 五六七八 | | +| | 0086-139 000 05678 | 八六 幺三九 零零零 零五六七八 | | +| 服务号 | 123 | 幺二三 | 支持常用的服务号。支持以400/800开头的10位服务号,支持以"3-3-4"的数字分隔方式。支持以12530/17951/12593开头的16位号码。 | +| | 95678 | 九五六七八 | | +| | 4008110510 | 四零零 八幺幺 零五幺零 | | +| | 800-810-8888 | 八零零 八幺零 八八八八 | | +| | 1253013520638377 | 幺二五三零 幺三五 二零六三 八三七七 | | +| 其他 | (86)(21)9899-80800-0909 | 八六 二幺 九八九九 八零八零零 零九零九 | 支持"数字串+分隔符(左右括号、-)"方式。 | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| -------------------------------------- | ------------- | -------------------------------------------------- | --------------------- | +| 纯数字 | 12034 | one two oh three four | 无严格长度限制,建议不超过 20 个字符。 | +| 数字 + 空格或连字符 + 数字 + ... | 1-23-456 7890 | one, two three, four five six, seven eight nine oh | | +| 加号 + 数字 + 空格或连字符 + 数字 | +43-211-0567 | plus four three, two one one, oh five six seven | | +| 左括号 + 数字 + 右括号 + 空格 + 数字 + 空格或连字符 + 数字 | (21) 654-3210 | (two one) six five four, three two one oh | | + +**示例** + +```xml + + 12345 + +``` + +```xml + + 10234 + +``` + +##### name + +**示例** + +```xml + + Her former name is Zeng Xiaofan + +``` + +##### address + +`address` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| ------ | ------------------- | ------------------- | ---------------------- | +| 常用地址格式 | 元和镇嘉元30-9 | 元和镇嘉元三十杠九 | 支持常用地址格式。此处地址指标准的邮寄地址。 | +| | 市台路388弄1107-1108号 | 市台路三八八弄幺幺零七杠幺幺零八号 | | +| | 华润二十四城六期锦云府3-1-3205 | 华润二十四城六期锦云府三杠一杠三二零五 | | +| | 圣华名都大厦2幢2006室 | 圣华名都大厦二幢二零零六室 | | +| | 五常街道庭院5幢4单元201 | 五常街道庭院五幢四单元二零幺 | | +| | 芙蓉江路150弄19号 | 芙蓉江路幺五零弄十九号 | | + + + 英文文本不支持该格式。 + + +**示例** + +```xml + + Fulu International, Building 1, Unit 3, Room 304 + +``` + +##### id + +`id` 支持的格式: + +| 格式 | 示例 | 输出 | 说明 | +| --- | ---------- | ------------------- | -------------------------------------------------- | +| 字符串 | dell0101 | D E L L 零 一 零 一 | 大小写英文字符、阿拉伯数字0\~9、下划线。输出的空格表示每个字符之间插入停顿,即字符一个一个地读。 | +| | myid\_1998 | M Y I D 下划线 一 九 九 八 | | +| | AiTest | A I T E S T | | + + + 英文文本的效果与 `characters` 相同。 + + +**示例** + +```xml + + myid_1998 + +``` + +##### characters + +`characters` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| --- | ------------------- | ----------------------------------- | ------------------------------------------------------------------------------------------- | +| 字符串 | ISBN 1-001-099098-1 | I S B N 一 杠 零 零 一 杠 零 九 九 零 九 八 杠 一 | 支持中文汉字、大小写英文字符、阿拉伯数字0\~9以及部分全角和半角字符。输出的空格表示每个字符之间插入停顿,即字符一个一个地读。标签内的文本如果包含XML的特殊字符,需要做字符转义。 | +| | x10b2345\_u | x 一 零 b 二 三 四 五 下划线 u | | +| | v1.0.1 | v 一 点 零 点 一 | | +| | 版本号2.0 | 版本号二 点 零 | | +| | 苏M MA000 | 苏M M A 零 零 零 | | +| | 空中客车A330 | 空中客车A 三 三 零 | | +| | 型号s01 s02和s03 | 型号s 零 一 s 零二 和s 零 三 | | +| | 空中客车A330 | 空中客车A 三 三 零 | | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| --- | -------------- | -------------------------------------------------------------------- | ------------------------- | +| 字符串 | \*b+3\$.c-0'=α | asterisk B plus three dollar dot C dash zero apostrophe equals alpha | 支持中文汉字、英文字母、数字 0-9 及常用符号。 | + +**示例** + +```xml + + Greek letters αβ + +``` + +```xml + + *b+3.c$=α + +``` + +##### punctuation + +`punctuation` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| ---- | ------- | ------------------------ | ----------------------------------------------------------------- | +| 标点符号 | ... | 省略号 | 支持常见中英文标点。输出的空格表示每个字符之间插入停顿,即字符一个一个地读。标签内的文本如果包含XML的特殊字符,需要做字符转义。 | +| | ...... | 省略号 | | +| | !"#\$%& | 叹号 双引号 井号 dollar 百分号 and | | +| | '()\*+ | 单引号 左括号 右括号 星号 加号 | | +| | ,-./:; | 逗号 杠 点 斜杠 冒号 分号 | | +| | ?@ | 小于 等号 大于 问号 at | | +| | \[]^\_ | 左方括号 反斜线 右方括号 脱字符 下划线 | | + + + 英文文本的效果与 `characters` 相同。 + + +**示例** + +```xml + + -./:; + +``` + +##### date + +`date` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| -------------------- | ----------------------- | ------------------- | --------------------------------------------------------------------------- | +| xx年 | 71年 | 七一年 | 支持2位和4位年份。2位年份支持60年~~99年、00年~~09年、10年~~19年。4位年份支持1000年~~1999年、2000年\~2099年。 | +| | 04年 | 零四年 | | +| | 19年 | 一九年 | | +| | 1011年 | 一零一一年 | | +| | 1998年 | 一九九八年 | | +| | 2008年 | 二零零八年 | | +| xx年xx月 | 98年4月 | 九八年四月 | 当月份为1到9月时,支持开头带"0"和不带"0"两种写法。 | +| | 1998年04月 | 一九九八年四月 | | +| | 08年8月 | 零八年八月 | | +| | 2008年8月 | 二零零八年八月 | | +| xx年xx月xx日/xx年xx月xx号 | 98年4月23日 | 九八年四月二十三日 | 当日期为1到9日时,支持开头带"0"和不带"0"两种写法。 | +| | 1998年04月23日 | 一九九八年四月二十三日 | | +| | 08年8月8号 | 零八年八月八号 | | +| | 2008年08月08号 | 二零零八年八月八号 | | +| xx月xx号 | 3月20日 | 三月二十日 | | +| | 08月07号 | 八月七号 | | +| 年月缩写 | 2018/08 | 二零一八年八月 | 支持"/"、"-"、"."作为缩写的分隔符。 | +| | 2018-08 | 二零一八年八月 | | +| | 2018.08 | 二零一八年八月 | | +| 年月日缩写 | 2018/08/08 | 二零一八年八月八日 | | +| | 2018-8-8 | 二零一八年八月八日 | | +| | 2018.08.08 | 二零一八年八月八日 | | +| xx年xx月xx日\~xx年xx月xx日 | 04年9月1日\~30日 | 零四年九月一日至三十日 | 支持"\~"、"-"作为"至"的缩写标志。 | +| | 2004年09月01号-2008年06月08号 | 二零零四年九月一号至二零零八年六月八号 | | +| xx年xx月\~xx年xx月 | 01年04月\~10年04月 | 零一年四月至一零年四月 | | +| | 2001年04月\~2010年04月 | 二零零一年四月至二零一零年四月 | | +| xx月xx日\~xx月xx日 | 10月1日\~10月7日 | 十月一日至十月七日 | | +| | 10月01号\~10月07号 | 十月一号至十月七号 | | +| xx月xx日\~xx日 | 10月1日\~7日 | 十月一日至七日 | | +| | 10月01号\~07号 | 十月一号至七号 | | +| 年月日缩写\~年月日缩写 | 2018/03/03\~2019/01/01 | 二零一八年三月三日至二零一九年一月一日 | 支持"/"、"."作为缩写的分隔符,支持"\~"、"-"作为"至"的缩写标志。 | +| | 1997.9.9\~1998.9.9 | 一九九七年九月九日至一九九八年九月九日 | | +| 月日缩写\~月日缩写 | 10/20\~10/31 | 十月二十日至十月三十一日 | | +| xx~~xx月/xx月~~xx月 | 1\~10月 | 一至十月 | | +| | 1月\~10月 | 一月至十月 | | +| 月日年缩写 | 10/20/2018 | 二零一八年十月二十日 | 仅支持4位的年份,仅支持"/"作为日期的分隔符,仅支持"月/日/年"的书写方式。 | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| ----------------------------------------------------- | ------------------ | ----------------------------------------------------- | ------------------------------------------- | +| 四位数/两位数 或 四位数-两位数 | 2000/01 | two thousand, oh one | 年份跨度。 | +| | 1900-01 | nineteen hundred, oh one | | +| | 2001-02 | twenty oh one, oh two | | +| | 2019-20 | twenty nineteen, twenty | | +| | 1998-99 | nineteen ninety eight, ninety nine | | +| | 1999-00 | nineteen ninety nine, oh oh | | +| 以 1 或 2 开头的四位数 | 2000 | two thousand | 四位数年份。 | +| | 1900 | nineteen hundred | | +| | 1905 | nineteen oh five | | +| | 2021 | twenty twenty one | | +| 星期-星期 或 星期\~星期 或 星期&星期 | mon-wed | monday to wednesday | 范围分隔符中的 XML 特殊字符需转义。 | +| | tue\~fri | tuesday to friday | | +| | sat\&sun | saturday and sunday | | +| DD-DD MMM, YYYY 或 DD\~DD MMM, YYYY 或 DD\&DD MMM, YYYY | 19-20 Jan, 2000 | the nineteen to the twentieth of january two thousand | DD = 两位数日期。MMM = 月份缩写或全称。YYYY = 四位数年份。 | +| | 01 \~ 10 Jul, 2020 | the first to the tenth of july twenty twenty | | +| | 05&06 Apr, 2009 | the fifth and the sixth of april two thousand nine | | +| MMM DD-DD 或 MMM DD\~DD 或 MMM DD\&DD | Feb 01 - 03 | february the first to the third | MMM = 月份。DD = 日期。 | +| | Aug 10-20 | august the tenth to the twentieth | | +| | Dec 11&12 | december the eleventh and the twelfth | | +| MMM-MMM 或 MMM\~MMM 或 MMM\&MMM | Jan-Jun | january to june | MMM = 月份。 | +| | Jul - Dec | july to december | | +| | sep\&oct | september and october | | +| YYYY-YYYY 或 YYYY\~YYYY | 1990 - 2000 | nineteen ninety to two thousand | YYYY = 以 1 或 2 开头的四位数年份。 | +| | 2001-2021 | two thousand one to twenty twenty one | | +| WWW DD MMM YYYY | Sun 20 Nov 2011 | sunday the twentieth of november twenty eleven | WWW = 星期(缩写或全称)。DD = 日期。MMM = 月份。YYYY = 年份。 | +| WWW DD MMM | Sun 20 Nov | sunday the twentieth of november | | +| WWW MMM DD YYYY | Sun Nov 20 2011 | sunday november the twentieth twenty eleven | | +| WWW MMM DD | Sun Nov 20 | sunday november the twentieth | | +| WWW YYYY-MM-DD | Sat 2010-10-01 | saturday october the first twenty ten | | +| WWW YYYY/MM/DD | Sat 2010/10/01 | saturday october the first twenty ten | | +| WWW MM/DD/YYYY | Sun 11/20/2011 | sunday november the twentieth twenty eleven | | +| MM/DD/YYYY | 11/20/2011 | november the twentieth twenty eleven | | +| YYYY | 1998 | nineteen ninety eight | | +| 其他默认读法 | 10 Mar, 2001 | the tenth of march two thousand one | | +| | 10 Mar | the tenth of march | | +| | Mar 2001 | march two thousand one | | +| | Fri. 10/Mar/2001 | friday the tenth of march two thousand one | | +| | Mar 10th, 2001 | march the tenth two thousand one | | +| | Mar 10 | march the tenth | | +| | 2001/03/10 | march the tenth two thousand one | | +| | 2001-03-10 | march the tenth two thousand one | | +| | 2000s | two thousands | | +| | 2010's | twenty tens | | +| | 1900's | nineteen hundreds | | +| | 1990s | nineteen nineties | | + +**示例** + +```xml + + 1000-10-10 + +``` + +```xml + + 10-01-2020 + +``` + +##### time + +`time` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| ------ | ---------------- | ---------------- | -------------- | +| 时刻 | 12:00 | 十二点 | 支持常用时间和时间范围格式。 | +| | 12:00:00点 | 十二点 | | +| | 10:20分 | 十点二十分 | | +| | 10:20:30 | 十点二十分三十秒 | | +| | 09:18:14 | 九点十八分十四秒 | | +| 时刻\~时刻 | 11:00\~12:00 | 十一点到十二点 | | +| | 09:00-14:00 | 九点到十四点 | | +| | 11:00\~11:30 | 十一点到十一点三十分 | | +| | 11:00-12:18 | 十一点到十二点十八分 | | +| | 10:30\~11:00 | 十点三十分到十一点 | | +| | 09:28-10:00 | 九点二十八分到十点 | | +| | 10:20\~11:20 | 十点二十分到十一点二十分 | | +| | 06:00\~08:00 | 六点到八点 | | +| | 上午10:20\~下午13:30 | 上午十点二十分到下午十三点三十分 | | +| 时间缩写 | 5:00 am | 凌晨五点整 | | +| | 5:30 am | 凌晨五点半 | | +| | 5:20:12 am | 凌晨五点二十分十二秒 | | +| | 7:00 am | 上午七点整 | | +| | 7:30 AM | 上午七点半 | | +| | 7:20:12 a.m. | 上午七点二十分十二秒 | | +| | 07:08:12 A.M. | 上午七点零八分十二秒 | | +| | 5:00 pm | 下午五点整 | | +| | 5:30 PM | 下午五点半 | | +| | 5:20:12 p.m. | 下午五点二十分十二秒 | | +| | 05:09:12 P.M. | 下午五点零九分十二秒 | | +| | 9:00 pm | 晚上九点整 | | +| | 9:30 pm | 晚上九点半 | | +| | 9:20:12 PM | 晚上九点二十分十二秒 | | +| | 9:02:12 P.M. | 晚上九点零二分十二秒 | | +| | 12:00 pm | 中午十二点整 | | +| | 12:30 p.m. | 中午十二点半 | | +| | 12:20:12 PM | 中午十二点二十分十二秒 | | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| ------------- | ------------------ | -------------------------------- | ------------------------------------------ | +| HH:MM AM 或 PM | 09:00 AM | nine A M | HH = 小时(1-2 位)。MM = 分钟(2 位)。AM/PM = 上午或下午。 | +| | 09:03 PM | nine oh three P M | | +| | 09:13 p.m. | nine thirteen p m | | +| HH:MM | 21:00 | twenty one hundred | | +| HHMM | 100 | one oclock | | +| 时间点-时间点 | 8:00 am - 05:30 pm | eight a m to five p m | 时间范围格式。 | +| | 7:05\~10:15 AM | seven oh five to ten fifteen A M | | +| | 09:00-13:00 | nine oclock to thirteen hundred | | + +**示例** + +```xml + + 5:00am + +``` + +```xml + + 0500 + +``` + +##### currency + +`currency` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| -------- | ----------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------ | +| 数字+金额标识符 | 12.00 RMB | 十二人民币 | 支持AUD(澳元)、CAD(加元)、HKD(港币)、JPY(日元)、USD(美元)、CHF(瑞士法郎)、NOK(挪威克朗)、SEK(瑞典克朗)、GBP(英镑)、RMB(人民币)、CNY(元)和EUR(欧元)。支持的数字格式包括:整数、小数以及以逗号分隔的国际写法。 | +| | 12.50 RMB | 十二点五零人民币 | | +| | 12,000,000 RMB | 一千二百万人民币 | | +| | 12,000,000.00 RMB | 一千二百万人民币 | | +| | 12,000.35 RMB | 一万两千点三五人民币 | | +| 金额标识符+数字 | \$12 | 十二美元 | 支持 CAD(加元)、\$(美元)、Fr(法郎)、kr(丹麦克朗)、£(英镑)、¥(元)和 EUR(欧元)。支持的数字格式包括:整数、小数以及以逗号分隔的国际写法。 | +| | \$12.00 | 十二美元 | | +| | \$12.12 | 二点一二美元 | | +| | \$12,000 | 一万两千美元 | | +| | \$12,000.00 | 一万两千美元 | | +| | \$12,000.99 | 一万两千点九九美元 | | +| 其他默认读法 | 1213 | 一千二百一十三 | | +| | 1213 KML | 一千二百一十三K M L | | +| | 1213.00 KML | 一千二百一十三K M L | | +| | 1213.9 KML | 一千二百一十三点九K M L | | +| | 1,000 KML | 一千K M L | | +| | 1,000.00 KML | 一千K M L | | +| | 1,000.98 KML | 一千点九八K M L | | +| | 12,000 | 一万两千 | | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| --------------------------------- | ------------ | ------------------------------------- | ------------------------------------------------------------- | +| 数字 + 货币标识符 | 1.00 RMB | one yuan | 支持整数、小数和千分位逗号分隔。 | +| | 2.02 CNY | two point zero two yuan | | +| | 1,000.23 CN¥ | one thousand point two three yuan | | +| | 1.01 SGD | one singapore dollar and one cent | | +| | 2.01 CAD | two canadian dollars and one cent | | +| | 3.1 HKD | three hong kong dollars and ten cents | | +| | 1,000.00 EUR | one thousand euros | | +| 货币标识符 + 数字 | US\$ 1.00 | one US dollar | 支持整数、小数和千分位逗号分隔。 | +| | \$0.01 | one cent | | +| | JPY 1.01 | one japanese yen and one sen | | +| | £1.1 | one pound and ten pence | | +| | €2.01 | two euros and one cent | | +| | USD 1,000 | one thousand united states dollars | | +| 数字 + 量词 + 货币标识符 或 货币标识符 + 数字 + 量词 | 1.23 Tn RMB | one point two three trillion yuan | 量词:thousand、million、billion、trillion、Mil、mil、K、k、Bn、bn、Tn、tn。 | +| | \$1.2 K | one point two thousand dollars | | + +**示例** + +```xml + + 13,000,000.00RMB + +``` + +```xml + + $1,000.01 + +``` + +##### measure + +`measure` 支持的格式: + +**中文输出** + +| 格式 | 示例 | 中文输出 | 说明 | +| ------------ | -------------------- | ----------------------- | -------------- | +| 数字+中文单位 | 2片 | 两片 | 支持常见中文单位及单位缩写。 | +| | 120公顷 | 一百二十公顷 | | +| | 100多毫克 | 一百多毫克 | | +| | 100来米 | 一百来米 | | +| | 100余人 | 一百余人 | | +| | 1厘米20毫米 | 一厘米二十毫米 | | +| | 120.00平方公里 | 一百二十平方公里 | | +| 数字+单位缩写 | 120.56 cm2 | 一百二十点五六平方厘米 | | +| | 100 m 12 cm 6 mm | 一百米十二厘米六毫米 | | +| 范围 | 10\~15 kg | 十至十五千克 | | +| | 10.24\~789.82亩 | 十点二四至七百八十九点八二亩 | | +| | 10米\~15米 | 十米至十五米 | | +| | 10.24 cm\~19.08 cm | 十点二四厘米至十九点零八厘米 | | +| 数字+单位+"/"+单位 | 10元/斤 | 十元每斤 | | +| | 199\~299元/件 | 一百九十九至二百九十九元每件 | | +| | 299.99元/g\~399.99元/g | 二百九十九点九九元每克至三百九十九点九九元每克 | | +| 其他默认读法 | 12扎 | 十二扎 | | +| | 30 rm | 三十r m | | +| | 4万万同胞 | 四万万同胞 | | +| | 12.897微克 | 十二点八九七微克 | | + +**英文输出** + +| 格式 | 示例 | 英文读法 | 说明 | +| --------- | ----------- | -------------------------------------------------------------- | ------------------------- | +| 数字 + 度量单位 | 1.0 kg | one kilogram | 支持整数、小数和千分位逗号分隔。支持常用单位缩写。 | +| | 1,234.01 km | one thousand two hundred thirty-four point zero one kilometers | | +| 纯度量单位 | mm2 | square millimeter | | + +**示例** + +```xml + + 100m12cm6mm + +``` + +```xml + + 1,000.01kg + +``` + +##### 符号发音 + +`` 常用符号发音: + +| 符号 | 中文读法 | 英文读法 | +| ---- | ------ | ----------------- | +| ! | 叹号 | exclamation mark | +| " | 双引号 | double quote | +| # | 井号 | pound | +| \$ | dollar | dollar | +| % | 百分号 | percent | +| & | and | and | +| ' | 单引号 | left quote | +| ( | 左括号 | left parenthesis | +| ) | 右括号 | right parenthesis | +| \* | 星 | asterisk | +| + | 加 | plus | +| , | 逗号 | comma | +| - | 杠 | dash | +| . | 点 | dot | +| / | 斜杠 | slash | +| : | 冒号 | colon | +| ; | 分号 | semicolon | +| \< | 小于 | less than | +| = | 等号 | equals | +| > | 大于 | greater than | +| ? | 问号 | question mark | +| @ | at | at | +| \[ | 左方括号 | left bracket | +| \\ | 反斜线 | backslash | +| ] | 右方括号 | right bracket | +| ^ | 脱字符 | caret | +| \_ | 下划线 | underscore | +| \` | 反引号 | backtick | +| `\{` | 左花括号 | left brace | +| \ | | 竖线 | vertical bar | +| `\}` | 右花括号 | right brace | +| \~ | 波浪线 | tilde | + +全角及特殊符号: + +| 符号 | 中文读法 | 英文读法 | +| -- | ----- | ------------------------ | +| ! | 叹号 | exclamation mark | +| “ | 左双引号 | left double quote | +| ” | 右双引号 | right double quote | +| ‘ | 左单引号 | left quote | +| ’ | 右单引号 | right quote | +| ( | 左括号 | left parenthesis | +| ) | 右括号 | right parenthesis | +| , | 逗号 | comma | +| 。 | 句号 | full stop | +| — | 杠 | em dash | +| : | 冒号 | colon | +| ; | 分号 | semicolon | +| ? | 问号 | question mark | +| 、 | 顿号 | enumeration comma | +| … | 省略号 | ellipsis | +| …… | 省略号 | ellipsis | +| 《 | 左书名号 | left guillemet | +| 》 | 右书名号 | right guillemet | +| ¥ | 人民币符号 | yuan | +| ≥ | 大于等于 | greater than or equal to | +| ≤ | 小于等于 | less than or equal to | +| ≠ | 不等于 | not equal | +| ≈ | 约等于 | approximately equal | +| ± | 加减 | plus or minus | +| × | 乘 | times | +| π | 派 | pi | + +希腊字母(大写): + +| 符号 | 中文读法 | 英文读法 | +| -- | ---- | ------- | +| Α | 阿尔法 | alpha | +| Β | 贝塔 | beta | +| Γ | 伽玛 | gamma | +| Δ | 德尔塔 | delta | +| Ε | 艾普西龙 | epsilon | +| Ζ | 捷塔 | zeta | +| Θ | 西塔 | theta | +| Ι | 艾欧塔 | iota | +| Κ | 喀帕 | kappa | +| ∧ | 拉姆达 | lambda | +| Μ | 缪 | mu | +| Ν | 拗 | nu | +| Ξ | 克西 | ksi | +| Ο | 欧麦克轮 | omicron | +| ∏ | 派 | pi | +| Ρ | 柔 | rho | +| ∑ | 西格玛 | sigma | +| Τ | 套 | tau | +| Υ | 宇普西龙 | upsilon | +| Φ | fai | phi | +| Χ | 器 | chi | +| Ψ | 普赛 | psi | +| Ω | 欧米伽 | omega | + +希腊字母(小写): + +| 符号 | 中文读法 | 英文读法 | +| -- | ---- | ------- | +| α | 阿尔法 | alpha | +| β | 贝塔 | beta | +| γ | 伽玛 | gamma | +| δ | 德尔塔 | delta | +| ε | 艾普西龙 | epsilon | +| ζ | 捷塔 | zeta | +| η | 依塔 | eta | +| θ | 西塔 | theta | +| ι | 艾欧塔 | iota | +| κ | 喀帕 | kappa | +| λ | 拉姆达 | lambda | +| μ | 缪 | mu | +| ν | 拗 | nu | +| ξ | 克西 | ksi | +| ο | 欧麦克轮 | omicron | +| π | 派 | pi | +| ρ | 柔 | rho | +| σ | 西格玛 | sigma | +| τ | 套 | tau | +| υ | 宇普西龙 | upsilon | +| φ | fai | phi | +| χ | 器 | chi | +| ψ | 普赛 | psi | +| ω | 欧米伽 | omega | + +##### 常用度量单位 + +`` 常用度量单位: + +| 类别 | 单位 | +| -- | --------------------------------------------------------------------- | +| 长度 | nm(纳米)、μm(微米)、mm(毫米)、cm(厘米)、m(米)、km(千米)、ft(英尺)、in(英寸) | +| 面积 | cm²(平方厘米)、m²(平方米)、km²(平方千米)、SqFt(平方英尺) | +| 体积 | cm³(立方厘米)、m³(立方米)、km3(立方千米)、mL(毫升)、L(升)、gal(加仑) | +| 重量 | μg(微克)、mg(毫克)、g(克)、kg(千克) | +| 时间 | min(分钟)、sec(秒)、ms(毫秒) | +| 电磁 | μA(微安)、mA(毫安)、Hz(赫兹)、kHz(千赫兹)、MHz(兆赫兹)、GHz(吉赫兹)、V(伏特)、kV(千伏)、kWh(千瓦时) | +| 声音 | dB(分贝) | +| 气压 | Pa(帕斯卡)、kPa(千帕)、MPa(兆帕) | +| 其他 | 还支持 tsp(茶匙)、rpm(转/分)、KB(千字节)、mmHg(毫米汞柱)等单位。 | + +## LaTeX 公式转语音 + +CosyVoice 可以将文本中的数学公式转换为自然语音,适用于有声书、在线教育等数理类音频内容场景。 + + + 该功能仅支持**中文**,其他语言可能无法正确朗读公式。 + + +### 使用限制 + +- **仅支持中文**:不支持其他语言 +- **内容限制**: + - 仅支持[支持的标签和符号](#支持的标签和符号)中列出的标签和符号 + - 不支持 Markdown 数学代码块(` ```math ... ``` `) + - 分隔符内只能包含公式,混入其他内容可能导致合成结果不准确 +- **兼容模型**:cosyvoice-v3.5-flash、cosyvoice-v3.5-plus、cosyvoice-v3-flash、cosyvoice-v3-plus、cosyvoice-v2 + +### 使用方法 + +用指定的分隔符包裹文本中的公式,然后调用语音合成 API。 + + + + 用以下任意分隔符包裹公式(效果相同): + + - `$...$` + - `$$...$$` + - `\(...\)` + - `\[...\]` + + 示例: + + ```plaintext + 这是一元二次方程的求根公式:$x = \frac{-b \pm \sqrt{b^2-4ac}}{2a}$,请仔细计算。 + ``` + + + + 调用语音合成 API,传入标记好公式的文本。在 JSON 或字符串中,反斜杠(`\`)是转义字符,需要写成 `\\`。 + + Python 调用示例: + + ```python + # coding=utf-8 + + import os + import dashscope + from dashscope.audio.tts_v2 import * + + # 如果未配置环境变量,请将下面一行替换为:dashscope.api_key = "sk-xxx" + dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY') + + dashscope.base_websocket_api_url='wss://maas.qianwenaiapi.com/api-ws/v1/inference' + + model = "cosyvoice-v3-flash" + voice = "longanyang" + + synthesizer = SpeechSynthesizer(model=model, voice=voice) + audio = synthesizer.call("这是一元二次方程的求根公式:$x = \\frac{-b \\pm \\sqrt{b^2-4ac}}{2a}$,请仔细计算。") + + print('[Metric] requestId: {}, first-package delay: {} ms'.format( + synthesizer.get_last_request_id(), + synthesizer.get_first_package_delay())) + + with open('output.mp3', 'wb') as f: + f.write(audio) + ``` + + + +### 支持的标签和符号 + +以下是当前支持的标签和符号列表。 + +#### 基础运算 + +| 标签或符号 | 功能 | 公式内容示例 | 公式输入示例 | 朗读效果 | +| --------- | ---- | ----------------- | ------------------- | --------- | +| + | 加法 | 2 + 3 = 5 | `$2 + 3 = 5$` | 二加三等于五 | +| - | 减法 | 3 - 2 = 1 | `$3 - 2 = 1$` | 三减二等于一 | +| \pm | 正负号 | \pm 1 \pm 2 | `$\pm 1\pm 2$` | 正负一、正负二 | +| \times | 乘法 | 2 \times 3 = 6 | `$2 \times 3 = 6$` | 二乘三等于六 | +| × | 乘法 | 2 × 3 = 6 | `$$2 × 3 = 6$$` | 二乘三等于六 | +| \* | 乘法 | 2 \* 3 = 6 | `\(2 * 3 = 6\)` | 二乘三等于六 | +| \div | 除法 | 6\div2=3 | `\[6\div2=3\]` | 六除以二等于三 | +| ÷ | 除法 | 6÷2=3 | `$6÷2=3$` | 六除以二等于三 | +| / | 除法 | 6/2=3 | `$6/2=3$` | 六除以二等于三 | +| = | 等于 | 3+5=8 | `$3+5=8$` | 三加五等于八 | +| \< | 小于 | 1\< 2 | `$1< 2$` | 一小于二 | +| ≤ | 小于等于 | 3≤5 | `$3≤5$` | 三小于等于五 | +| \<= | 小于等于 | 3\<=5 | `$3<=5$` | 三小于等于五 | +| \leq | 小于等于 | 3\leq5 | `$3\leq 5$` | 三小于等于五 | +| \le | 小于等于 | 3\le5 | `$3\le 5$` | 三小于等于五 | +| \leqq | 小于等于 | 3\leqq5 | `$3\leqq 5$` | 三小于等于五 | +| \leqslant | 小于等于 | 3\leqslant5 | `$3\leqslant 5$` | 三小于等于五 | +| > | 大于 | 2>1 | `$2>1$` | 二大于一 | +| ≥ | 大于等于 | 5≥3 | `$5≥3$` | 五大于等于三 | +| >= | 大于等于 | 5>=3 | `$5>=3$` | 五大于等于三 | +| \geq | 大于等于 | 5\geq3 | `$5\geq 3$` | 五大于等于三 | +| \ge | 大于等于 | 5\ge3 | `$5\ge 3$` | 五大于等于三 | +| \geqq | 大于等于 | 5\geqq3 | `$5\geqq 3$` | 五大于等于三 | +| \geqslant | 大于等于 | 5\geqslant3 | `$5\geqslant 3$` | 五大于等于三 | +| \frac | 分数 | `2\frac3` | `$\frac {2}{3}$` | 三分之二 | +| ^ | 幂 | `2^1` | `$2^{1}$` | 二的一次方 | +| \sqrt | 开方 | `\sqrt{9} = 3` | `$\sqrt {9} = 3$` | 根号九等于三 | +| \sqrt | 开方 | `\sqrt[3]{8} = 2` | `$\sqrt[3]{8} = 2$` | 八的三次方根等于二 | +| % | 百分号 | `5\%` | `$5\%$` | 百分之五 | +| \ | | 绝对值 | `∣3∣=3` | `$\ | 3\ | =3$` | 三的绝对值等于三 | +| \vert | 绝对值 | `3\vert=3` | `$\vert 3\vert =3$` | 三的绝对值等于三 | +| \lg | 对数 | `lg {10}` | `$\lg {10}$` | lg 十 | +| \log | 对数 | `\log{5}` | `$\log{5}$` | log 五 | +| \ln | 自然对数 | `\lnX` | `$ln {10}$` | ln 十 | +| ! | 阶乘 | 5! | `$5!$` | 五的阶乘 | +| () | 括号 | (2+1) | `$(2+1)$` | 括号二加一 | +| `\{ \}` | 花括号 | `\{2+1\}` | `$\{2+1\}$` | 花括号二加一 | + +#### 特殊数学符号 + +| 标签或符号 | 转换结果 | 公式内容示例 | 公式输入示例 | 朗读效果 | +| ------ | ----- | ------ | ---------- | ---- | +| \alpha | alpha | \alpha | `$\alpha$` | 阿尔法 | +| \Alpha | alpha | \Alpha | `$\Alpha$` | 阿尔法 | +| \beta | beta | \beta | `$\beta$` | 贝塔 | +| \Beta | beta | \Beta | `$\Beta$` | 贝塔 | +| \gamma | gamma | \gamma | `$\gamma$` | 伽马 | +| \Gamma | gamma | \Gamma | `$\Gamma$` | 伽马 | +| \delta | delta | \delta | `$\delta$` | 德尔塔 | +| \Delta | delta | \Delta | `$\Delta$` | 德尔塔 | +| \infty | 无穷大 | \infty | `$\infty$` | 无穷大 | +| ∞ | 无穷大 | ∞ | `$∞$` | 无穷大 | + +#### 几何 + +| 标签或符号 | 功能 | 公式内容示例 | 公式输入示例 | 朗读效果 | +| ---------------- | ----- | --------------------------------------------------- | ------------------------------------------------------- | ------------------- | +| \pi | 圆周率 | \pi=3.14159 | `$\pi =3.14159$` | 派等于 3.14159 | +| \sin | 三角函数 | `\sin 30^\circ=\frac{1}{2}` | `$\sin 30^\circ =\frac {1}{2}$` | 正弦三十度等于二分之一 | +| \cos | 三角函数 | `\cos 30^\circ=\frac{\sqrt{2}}{2}` | `$\cos 30^\circ =\frac {\sqrt {2}}{2}$` | 余弦三十度等于二分之根号二 | +| \tan | 三角函数 | `\tan 30^\circ=\frac{\sin 30^\circ}{\cos 30^\circ}` | `$\tan 30^\circ =\frac {\sin 30^\circ}{\cos 30^\circ}$` | 正切三十度等于正弦三十度除以余弦三十度 | +| \csc | 三角函数 | \csc A | `$\csc A$` | 余割 A | +| \sec | 三角函数 | \sec A | `$\sec A$` | 正割 A | +| \cot | 三角函数 | \cot A | `$\cot A$` | 余切 A | +| \angle | 角 | \angle AB | `$\angle AB$` | 角 AB | +| ∠ | 角 | ∠AB | `$∠AB$` | 角 AB | +| ^\circ | 度 | ∠AB = 30^\circ | `$∠AB = 30^\circ$` | 角 AB 等于三十度 | +| \odot | 圆 | \odot | `$\odot$` | 圆 | +| `\overset\frown` | 弧 | `\overset\frown {BC}` | `$\overset\frown {BC}$` | 弧 BC | +| `\rm{Rt}` | 直角 | `\because \rm{Rt}\triangle ABC` | `$\because \rm{Rt}\triangle ABC$` | 因为三角形 ABC 是直角三角形 | +| `\mathrm{Rt}` | 直角 | `\therefore AB \perp BC` | `$\therefore AB \perp BC$` | 所以 AB 垂直于 BC | +| \triangle | 三角形 | \triangle ABC | `$\triangle ABC$` | 三角形 ABC | +| △ | 三角形 | △ABC | `$△ABC$` | 三角形 ABC | +| \parallelogram | 平行四边形 | \parallelogram ABCD | `$\parallelogram ABCD$` | 平行四边形 ABCD | +| \perp | 垂直 | AB \perp BC | `$AB \perp BC$` | AB 垂直于 BC | +| \bot | 垂直 | AB \bot BC | `$AB \bot BC$` | AB 垂直于 BC | +| ⊥ | 垂直 | AB ⊥ BC | `$AB ⊥ BC$` | AB 垂直于 BC | +| \parallel | 平行 | A\parallel B | `$A\parallel B$` | A 平行于 B | +| \equalparallel | 平行且等于 | A\equalparallel B | `$A\equalparallel B$` | A 平行且等于 B | +| \cong | 全等 | △ABC\cong△DEF | `$△ABC\cong△DEF$` | 三角形 ABC 全等于三角形 DEF | + +#### 条件关系 + +| 标签或符号 | 功能 | 公式内容示例 | 公式输入示例 | 朗读效果 | +| ---------- | --- | ----------------------------- | --------------------------------- | ------------------- | +| \implies | 推出 | \implies 1+1=2 | `$\implies 1+1=2$` | 可推出一加一等于二 | +| \iff | 等价于 | p\iffq | `$p\iffq$` | p 等价于 q | +| \because | 因为 | \because a = b \therefore b=a | `$\because a = b \therefore b=a$` | 因为 a 等于 b,所以 b 等于 a | +| \therefore | 所以 | \because a = b \therefore b=a | `$\because a = b \therefore b=a$` | 因为 a 等于 b,所以 b 等于 a | + +#### 单位 + +单位必须用 `\unit`、`\quantity`、`\mathit`、`\mathrm` 或 `\rm` 标签包裹(例如 `\unit{cm}`)。 + +| 标签或符号 | 朗读效果 | 公式内容示例 | 公式输入示例 | 朗读示例 | +| ----- | ----- | ------------------ | -------------------- | ------ | +| mm | 毫米 | `5\quantity{mm}` | `$5\quantity{mm}$` | 五毫米 | +| cm | 厘米 | `5\quantity{cm}` | `$5\quantity{cm}$` | 五厘米 | +| dm | 分米 | `5\quantity{dm}` | `$5\quantity{dm}$` | 五分米 | +| m | 米 | `5\quantity{m}` | `$5\quantity{m}$` | 五米 | +| km | 千米 | `5\quantity{km}` | `$5\quantity{km}$` | 五千米 | +| g | 克 | `5\quantity{g}` | `$5\quantity{g}$` | 五克 | +| kg | 千克 | `5\quantity{kg}` | `$5\quantity{kg}$` | 五千克 | +| t | 吨 | `5\quantity{t}` | `$5\quantity{t}$` | 五吨 | +| mm^2 | 平方毫米 | `5\quantity{mm^2}` | `$5\quantity{mm^2}$` | 五平方毫米 | +| cm^2 | 平方厘米 | `5\quantity{cm^2}` | `$5\quantity{cm^2}$` | 五平方厘米 | +| dm^2 | 平方分米 | `5\quantity{dm^2}` | `$5\quantity{dm^2}$` | 五平方分米 | +| m^2 | 平方米 | `5\quantity{m^2}` | `$5\quantity{m^2}$` | 五平方米 | +| km^2 | 平方千米 | `5\quantity{km^2}` | `$5\quantity{km^2}$` | 五平方千米 | +| mm^3 | 立方毫米 | `5\quantity{mm^3}` | `$5\quantity{mm^3}$` | 五立方毫米 | +| cm^3 | 立方厘米 | `5\quantity{cm^3}` | `$5\quantity{cm^3}$` | 五立方厘米 | +| dm^3 | 立方分米 | `5\quantity{dm^3}` | `$5\quantity{dm^3}$` | 五立方分米 | +| m^3 | 立方米 | `5\quantity{m^3}` | `$5\quantity{m^3}$` | 五立方米 | +| km^3 | 立方千米 | `5\quantity{km^3}` | `$5\quantity{km^3}$` | 五立方千米 | +| ml | 毫升 | `5\quantity{ml}` | `$5\quantity{ml}$` | 五毫升 | +| s | 秒 | `5\quantity{s}` | `$5\quantity{s}$` | 五秒 | +| min | 分钟 | `5\quantity{min}` | `$5\quantity{min}$` | 五分钟 | +| h | 小时 | `5\quantity{h}` | `$5\quantity{h}$` | 五小时 | +| km/h | 千米每小时 | `5\quantity{km/h}` | `$5\quantity{km/h}$` | 五千米每小时 | +| g/l | 克每升 | `5\quantity{g/l}` | `$5\quantity{g/l}$` | 五克每升 | + +### 常见问题 + +#### 输入的公式没有被朗读? + +1. **分隔符**:确认公式已用 `$...$`、`$$...$$`、`\(...\)` 或 `\[...\]` 包裹 +2. **公式复杂度**:确认公式仅使用了[支持的标签和符号](#支持的标签和符号)中的内容 +3. **转义字符**:确认在 API 请求中,反斜杠(`\`)已转义为 `\\` + +#### 代码中如何处理反斜杠(`\`)? + +反斜杠(`\`)在字符串和 JSON 中是转义字符,需要写成 `\\`。例如:在 Python、Java、JavaScript 等语言中,`\frac` 应写为 `\\frac`。 diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts-models.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts-models.md new file mode 100644 index 0000000..c6da94a --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts-models.md @@ -0,0 +1,169 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 语音合成模型 + +> 选择适合语音合成、声音克隆和声音设计的模型。 + +选择模型前,先确定两个问题:是否需要自定义音色(还是内置音色即可),以及是否需要实时流式输出。 + +## 从闭源模型迁移到千问AI平台? + +如果你正在使用 ElevenLabs、OpenAI 或 Google 的语音合成服务,可参考下表选择对应的模型: + +| 使用场景 | 闭源模型代表 | 推荐 | +| ------------ | -------------------------------- | ------------------------------------------------------------ | +| 内置音色 / 标准合成 | OpenAI gpt-4o-tts、Google Chirp 3 | `qwen-audio-3.0-tts-plus` | +| 自定义音色 / 声音复刻 | ElevenLabs Multilingual v3 | `qwen-audio-3.0-tts-flash`(声音复刻)、`cosyvoice-v3.5-plus`(声音设计) | + +## 内置音色还是自定义音色? + +### 内置音色 + +从音色库中选择一个音色,即可开始合成语音。 + +- **Qwen-Audio-TTS** — 通过 WebSocket/HTTP 调用,支持声音复刻和指令控制 +- **CosyVoice** — 音色库丰富,合成质量高,选定音色即可使用 +- **Qwen3-TTS** — 低延迟流式输出;使用 `-instruct` 变体可通过自然语言控制语速、情感和风格 +- **MiniMax** — 支持混合音色和情感风格调节,适合社交、播客等场景 + + + CosyVoice 系列模型还支持通过 AOQ 协议接入;如果是客户端对接,且更看重稳定的延迟、弱网下的交互能力、实时双工的降噪与回声消除,可优先考虑 AOQ,协议对比与选型请参见 [Realtime API 概述](/api-reference/realtime-api/overview)。 + + +### 自定义音色 + +音色库中没有满意的音色? + +- **声音克隆(Voice Cloning)** — 基于音频样本复现特定人物的声音。适用于需要匹配目标音色的场景。Qwen-Audio-TTS、CosyVoice、Qwen3-TTS 和 MiniMax 均支持声音克隆。 +- **声音设计(Voice Design)** — 通过文字描述生成全新音色(例如"温暖低沉的女声")。适用于无音频样本但需要品牌专属音色的场景。 + +## 控制语音效果 + +三种方式,按灵活性由高到低排列: + +1. **指令控制**(CosyVoice 系列:`cosyvoice-v3.5-plus`、`cosyvoice-v3.5-flash`、`cosyvoice-v3-flash`;Qwen-TTS 系列:`qwen3-tts-instruct-flash`、`qwen3-tts-instruct-flash-realtime`)— 用自然语言描述期望的朗读效果,可逐次调整语速、情感和风格。灵活性最高。 + +2. **声音设计**(`qwen3-tts-vd-*`)— 通过文字描述生成自定义音色。适合在没有音频样本的情况下打造品牌音色。 + +3. **声音克隆**(`qwen3-tts-vc-*`)— 基于音频样本复现已有声音。适合需要匹配特定人物音色的场景。 + +## 推荐模型 + +| 模型 | 系列 | 流式输出 | 自定义音色 | 指令控制 | +| ---------------------------------- | -------------- | ---- | ----- | ---- | +| `qwen-audio-3.0-tts-plus` | Qwen-Audio-TTS | ✓ | ✓ | ✓ | +| `cosyvoice-v3-plus` | CosyVoice | ✓ | — | — | +| `MiniMax/speech-2.8-hd` | MiniMax | ✓ | ✓ | — | +| `qwen3-tts-flash` | Qwen3-TTS | ✓ | — | — | +| `qwen3-tts-flash-realtime` | Qwen3-TTS | ✓ | — | — | +| `qwen3-tts-instruct-flash` | Qwen3-TTS | ✓ | — | ✓ | +| `qwen3-tts-vc-realtime-2026-01-15` | Voice Cloning | ✓ | ✓ | — | +| `qwen3-tts-vd-realtime-2026-01-15` | Voice Design | ✓ | ✓ | — | + +## 全部模型 + + + +| 模型 | 流式输出 | 自定义音色 | 指令控制 | +| -------------------------- | ---- | ----- | ---- | +| `qwen-audio-3.0-tts-plus` | ✓ | ✓ | ✓ | +| `qwen-audio-3.1-tts-flash` | ✓ | ✓ | ✓ | +| `qwen-audio-3.0-tts-flash` | ✓ | ✓ | ✓ | + + + +| 模型 | 流式输出 | 自定义音色 | 指令控制 | +| ---------------------- | ---- | ----- | ---- | +| `cosyvoice-v3.5-plus` | ✓ | — | ✓ | +| `cosyvoice-v3.5-flash` | ✓ | — | ✓ | +| `cosyvoice-v3-plus` | ✓ | — | — | +| `cosyvoice-v3-flash` | ✓ | — | ✓ | + + + +| 模型 | 流式输出 | 自定义音色 | 指令控制 | +| ----------------------------------- | ---- | ----- | ---- | +| `qwen3-tts-flash` | ✓ | — | — | +| `qwen3-tts-flash-realtime` | ✓ | — | — | +| `qwen3-tts-instruct-flash` | ✓ | — | ✓ | +| `qwen3-tts-instruct-flash-realtime` | ✓ | — | ✓ | + + + +| 模型 | 流式输出 | 自定义音色 | 指令控制 | +| -------------------------- | ---- | ----- | ---- | +| `MiniMax/speech-2.8-hd` | ✓ | ✓ | — | +| `MiniMax/speech-02-hd` | ✓ | ✓ | — | +| `MiniMax/speech-2.8-turbo` | ✓ | ✓ | — | +| `MiniMax/speech-02-turbo` | ✓ | ✓ | — | + + + +| 模型 | 流式输出 | 自定义音色 | 指令控制 | +| ---------------------------------- | ---- | ----- | ---- | +| `qwen3-tts-vc-2026-01-22` | ✗ | ✓ | — | +| `qwen3-tts-vc-realtime-2026-01-15` | ✓ | ✓ | — | +| `qwen3-tts-vd-2026-01-26` | ✗ | ✓ | — | +| `qwen3-tts-vd-realtime-2026-01-15` | ✓ | ✓ | — | +| `qwen-voice-enrollment` | ✗ | ✓ | — | +| `qwen-voice-design` | ✗ | ✓ | — | + + + + 上一代模型。新项目建议使用上述最新版本。 + +| 模型 | 系列 | 流式输出 | 自定义音色 | 指令控制 | +| ---------------------------------------------- | ----------------- | ---- | ----- | ---- | +| `qwen3-tts-flash-2025-11-27` | Qwen3-TTS | ✓ | — | — | +| `qwen3-tts-flash-2025-09-18` | Qwen3-TTS | ✓ | — | — | +| `qwen3-tts-flash-realtime-2025-11-27` | Qwen3-TTS | ✓ | — | — | +| `qwen3-tts-flash-realtime-2025-09-18` | Qwen3-TTS | ✓ | — | — | +| `qwen3-tts-instruct-flash-2026-01-26` | Qwen3-TTS | ✓ | — | ✓ | +| `qwen3-tts-instruct-flash-realtime-2026-01-22` | Qwen3-TTS | ✓ | — | ✓ | +| `qwen3-tts-vc-realtime-2025-11-27` | Voice Cloning | ✓ | ✓ | — | +| `qwen3-tts-vd-realtime-2025-12-16` | Voice Design | ✓ | ✓ | — | +| `qwen-tts` | Qwen-TTS | ✓ | — | — | +| `qwen-tts-latest` | Qwen-TTS | ✓ | — | — | +| `qwen-tts-2025-05-22` | Qwen-TTS | ✓ | — | — | +| `qwen-tts-2025-04-10` | Qwen-TTS | ✓ | — | — | +| `qwen-tts-realtime` | Qwen-TTS-Realtime | ✓ | — | — | +| `qwen-tts-realtime-latest` | Qwen-TTS-Realtime | ✓ | — | — | +| `qwen-tts-realtime-2025-07-15` | Qwen-TTS-Realtime | ✓ | — | — | +| `cosyvoice-v2` | CosyVoice | ✓ | — | — | +| `cosyvoice-v1` | CosyVoice | ✓ | — | — | + + + +## 了解更多 + + + + 了解如何通过 API 使用语音合成模型。 + + + + 通过 WebSocket 使用实时语音合成模型。 + + + + 浏览 CosyVoice 音色库和试听样本。 + + + + 浏览 Qwen-TTS 非流式模型的系统音色。 + + + + 浏览 Qwen-TTS-Realtime 流式模型的系统音色。 + + + + 基于音频样本克隆声音。 + + + + 查看 MiniMax 语音合成模型的调用参数。 + + diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts.md new file mode 100644 index 0000000..b340477 --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-tts.md @@ -0,0 +1,474 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 非实时语音合成 + +非实时语音合成通过HTTP API将文本转换为语音,适用于有声书制作、在线教育配音、内容制作等对延迟要求不高的场景,支持丰富音色、多语言、声音复刻与声音设计。 + +## 概述 + +通过HTTP API将完整文本转换为语音文件,支持非流式和流式两种输出模式。 + +- **非流式**返回音频文件 URL,有效期 24 小时;**流式**逐段返回音频数据。 +- 支持多种语言,含中文方言。 +- 支持[声音复刻](/developer-guides/speech/voice-cloning)与[声音设计](/developer-guides/speech/voice-design)进行定制音色创建。 +- 支持[指令控制](/developer-guides/speech/tts#指令控制),通过自然语言指令控制语音表现力。 +- 支持[情感与富语言标签](/developer-guides/speech/tts#情感与富语言标签),可在文本中嵌入标签控制情感表达或插入拟声效果 + + + 各模型系列的调用端点不同:Qwen-Audio-TTS 与 CosyVoice 使用 `https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer`;Qwen-TTS 使用 `https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation`;MiniMax 使用 `https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation`。端点不可混用,请以各模型系列示例中的端点为准。 + + +低延迟流式场景请参见[实时语音合成](/developer-guides/speech/realtime-streaming)。各模型选型建议请参见[语音合成](/developer-guides/speech/tts-models)。 + +千问AI平台控制台**声音设计**页面合成的语音仅支持在线试听,无法下载音频文件。如需下载音频,请通过 API 或 SDK 调用,非流式模式下响应中返回音频 URL,有效期 24 小时。 + +## 前提条件 + +开始前,请确认已完成以下准备工作: + +- [配置API Key](/api-reference/preparation/api-key),并[设置到环境变量](/api-reference/preparation/export-api-key-env) +- (可选)如果通过 DashScope SDK调用,[安装最新版SDK](/api-reference/preparation/install-sdk) + +## 快速开始 + +以下各 Tab 分别演示不同模型系列的语音合成。更多语言示例和详细参数说明,请参见[API 参考](/developer-guides/speech/tts#api-参考)。 + + + + 以下示例演示如何使用 Qwen-Audio-TTS 模型合成语音。 + + + + 非流式模式下,响应中包含合成音频的 URL,有效期为 24 小时。 + + ```bash + curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "model": "qwen-audio-3.0-tts-flash", + "input": { + "text": "我家的后面有一个很大的花园。", + "voice": "longanhuan_v3.6", + "format": "wav", + "sample_rate": 24000 + } + }' + ``` + + + + 添加 `X-DashScope-SSE: enable` Header 开启流式输出,服务端以 Server-Sent Events(SSE)方式逐段返回音频数据。 + + ```bash + curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -H "X-DashScope-SSE: enable" \ + -d '{ + "model": "qwen-audio-3.0-tts-flash", + "input": { + "text": "我家的后面有一个很大的花园。", + "voice": "longanhuan_v3.6", + "format": "wav", + "sample_rate": 24000 + } + }' + ``` + + + + + + 以下示例演示如何使用 CosyVoice 模型合成语音。 + + + + 非流式模式下,响应中包含合成音频的 URL,有效期为 24 小时。 + + ```bash + curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "model": "cosyvoice-v3-flash", + "input": { + "text": "我家的后面有一个很大的花园。", + "voice": "longanyang", + "format": "wav", + "sample_rate": 24000 + } + }' + ``` + + + + 添加 `X-DashScope-SSE: enable` Header 开启流式输出,服务端以 Server-Sent Events(SSE)方式逐段返回音频数据。 + + ```bash + curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -H "X-DashScope-SSE: enable" \ + -d '{ + "model": "cosyvoice-v3-flash", + "input": { + "text": "我家的后面有一个很大的花园。", + "voice": "longanyang", + "format": "wav", + "sample_rate": 24000 + } + }' + ``` + + + + + + MiniMax 支持情感控制、语速调节和音调调整。 + + + + 非流式模式下,响应中包含完整的合成音频。 + + ```bash + curl -X POST "https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation" \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "model": "MiniMax/speech-2.8-hd", + "input": { + "text": "今天天气真不错,适合出去走走。", + "voice_setting": { + "voice_id": "male-qn-qingse", + "speed": 1, + "vol": 1, + "pitch": 0, + "emotion": "happy" + }, + "audio_setting": { + "sample_rate": 32000, + "bitrate": 128000, + "format": "mp3", + "channel": 1 + } + } + }' + ``` + + + + 添加 `X-DashScope-SSE: enable` Header 开启流式输出。 + + ```bash + # 获取API Key:/api-reference/preparation/api-key + + curl -X POST "https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation" \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -H "X-DashScope-SSE: enable" \ + -d '{ + "model": "MiniMax/speech-2.8-hd", + "input": { + "text": "今天天气真不错,适合出去走走。", + "voice_setting": { + "voice_id": "male-qn-qingse", + "speed": 1, + "vol": 1, + "pitch": 0, + "emotion": "happy" + }, + "audio_setting": { + "sample_rate": 32000, + "bitrate": 128000, + "format": "mp3", + "channel": 1 + } + } + }' + ``` + + + + + +## 进阶功能 + +### 指令控制 + +指令控制通过自然语言描述控制语音的音调、语速、情感和音色特点,无需调整复杂的音频参数。 + +**各模型指令规格**: + + + 指令参数名因模型系列而异:CosyVoice 使用 `instruction`,Qwen-TTS 使用 `instructions`。跨模型迁移时请注意修改参数名。 + + + + + **支持的模型**:`qwen-audio-3.1-tts-flash`、`qwen-audio-3.0-tts-plus`、`qwen-audio-3.0-tts-flash` + + 系统音色和声音复刻音色:均可输入任意指令。 + + + + **支持的模型**:`cosyvoice-v3.5-plus`、`cosyvoice-v3.5-flash`、`cosyvoice-v3-plus`、`cosyvoice-v3-flash` + + 不同模型对指令的格式要求不同: + + - `cosyvoice-v3.5-plus`、`cosyvoice-v3.5-flash`: + + - 声音复刻/设计音色:可输入任意指令。 + - 系统音色:v3.5不支持系统音色。 + - `cosyvoice-v3-plus`: + + - 声音复刻/设计音色:不支持指令控制。 + - 系统音色:指令必须使用固定格式和内容,参见[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)。 + - `cosyvoice-v3-flash`: + + - 声音复刻/设计音色:可输入任意指令。 + - 系统音色:指令必须使用固定格式和内容,参见[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)。 + + **使用方式**:通过 `instruction` 参数指定指令内容。 + + **指令文本支持的语言**: + + - `cosyvoice-v3.5-plus`、`cosyvoice-v3.5-flash`: + + - 声音复刻/设计音色:中文、英文、法语、德语、日语、韩语、俄语、葡萄牙语、泰语、印尼语、越南语。 + - 系统音色:v3.5不支持系统音色。 + - `cosyvoice-v3-plus`: + + - 声音复刻/设计音色:中文、英文、法语、德语、日语、韩语、俄语。 + - 系统音色:指令必须使用固定格式和内容,参见[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)。 + - `cosyvoice-v3-flash`: + + - 声音复刻/设计音色:中文、英文、法语、德语、日语、韩语、俄语。 + - 系统音色:中文。 + + **指令文本长度限制**:不超过 100 字符。汉字(包括简体/繁体汉字、日文汉字和韩文汉字)按 2 个字符计算,其他字符(如标点符号、字母、数字、日韩文假名/谚文等)按 1 个字符计算。 + + + + **支持的模型**:仅支持Qwen3-TTS-Instruct-Flash 系列模型。 + + **使用方式**:通过 `instructions` 参数传入指令内容。 + + **指令文本支持的语言**:仅支持中文和英文。 + + **指令文本长度限制**:不超过 1,600 Token。 + + + +**适用场景**: + +- 有声书和广播剧配音 +- 广告和宣传片配音 +- 游戏角色和动画配音 +- 情感化的智能语音助手 +- 纪录片和新闻播报 + +**如何编写高质量的声音描述**: + +- **核心原则**: + + 1. **具体而非模糊**:使用描绘声音特质的词语,如“低沉”、“清脆”、“语速偏快”,避免“好听”、“普通”等主观或模糊的表述。 + 2. **多维而非单一**:好的描述通常涵盖多个维度(如性别、年龄、情感等)。仅写“女声”过于宽泛,难以生成有特色的音色。 + 3. **客观而非主观**:聚焦声音的物理和感知特征。例如,用”音调偏高,带有活力“代替”我最喜欢的声音”。 + 4. **原创而非模仿**:描述声音的特质,而非要求模仿特定人物(如名人、演员)。模型不支持模仿,且可能涉及版权风险。 + 5. **简洁而非冗余**:确保每个词都有明确作用,避免重复的同义词或无意义的修饰。 +- **描述维度参考**: + + 建议组合以下维度描述声音,维度越丰富,生成效果越精准。 + +| **维度** | **描述示例** | +| ------ | ---------------------------------------------------------- | +| 性别 | 男性、女性、中性 | +| 年龄 | 儿童(5-12 岁)、青少年(13-18 岁)、青年(19-35 岁)、中年(36-55 岁)、老年(55 岁以上) | +| 音调 | 高音、中音、低音、偏高、偏低 | +| 语速 | 快速、中速、缓慢、偏快、偏慢 | +| 情感 | 开朗、沉稳、温柔、严肃、活泼、冷静、治愈 | +| 特点 | 有磁性、清脆、沙哑、圆润、甜美、浑厚、有力 | +| 用途 | 新闻播报、广告配音、有声书、动画角色、语音助手、纪录片解说 | + +- **示例**: + + - 标准播音风格:吐字清晰精准,字正腔圆 + - 年轻活泼的女性声音,语速较快,带有明显的上扬语调,适合介绍时尚产品 + - 沉稳的中年男性,语速缓慢,音色低沉有磁性,适合朗读新闻或纪录片解说 + - 温柔知性的女性,30 岁左右,语调平和,适合有声书朗读 + - 可爱的儿童声音,大约 8 岁女孩,说话略带稚气,适合动画角色配音 + +### 方言 + +本节介绍如何让模型用**中文方言**(如河南话、四川话等)输出语音。不同模型和音色类型的设置方式不同。 + + + + - **系统音色**:在[Qwen-Audio-TTS音色列表](/developer-guides/speech/voice-list/qwen-audio-tts)中选择以下任一种音色: + + - 支持方言的系统音色,无需额外设置即可输出对应方言。 + - 支持[指令控制](/developer-guides/speech/tts#指令控制)且可指定方言的音色,通过指令文本指定方言。 + - **声音复刻音色**:通过[指令控制](/developer-guides/speech/tts#指令控制)功能设置,例如指令文本写 `请用河南话表达`。 + - **声音设计音色**:暂不支持方言。 + + **具体支持哪些方言**:参见[Qwen-Audio-TTS](/developer-guides/speech/tts-models)中各模型“支持的语言”。 + + + + - **系统音色**:在[CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice)中选择以下任一种音色: + + - 支持方言的系统音色(例如 `longshange_v3`),无需额外设置即可输出对应方言。 + - 支持[指令控制](/developer-guides/speech/tts#指令控制)且可指定方言的音色(例如 `longanhuan_v3`),通过指令文本指定方言。 + - **声音复刻音色**:通过[指令控制](/developer-guides/speech/tts#指令控制)功能设置,例如指令文本写 `请用河南话表达`。 + - **声音设计音色**:暂不支持方言。 + + **具体支持哪些方言**:参见[CosyVoice](/developer-guides/speech/tts-models)中各模型“支持的语言”。 + + **示例**:以 `cosyvoice-v3-flash` + `longanhuan_v3` 音色,通过指令文本 `"请用河南话表达。"` 输出河南话语音。 + + ```bash + curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "model": "cosyvoice-v3-flash", + "input": { + "text": "叫你去买盐,你买回来一袋面,这不是弄啥嘞吗!", + "voice": "longanhuan_v3", + "format": "wav", + "sample_rate": 24000, + "instruction": "请用河南话表达。" + } + }' + ``` + + + 此处指令参数名 `instruction` 为 CosyVoice 专用;Qwen-TTS 的指令参数名为 `instructions`,请勿混用。 + + + + + - **系统音色**:使用支持方言的系统音色,参见[Qwen-TTS音色列表](/developer-guides/speech/voice-list/qwen-tts)。 + - **声音复刻音色**:不支持方言。 + - **声音设计音色**:不支持方言。 + + **具体支持哪些方言**:参见[Qwen3-TTS](/developer-guides/speech/tts-models)中各模型“支持的语言”。 + + + +### 情感与富语言标签 + +Qwen-Audio-TTS 系列模型支持在待合成文本(`text` 参数)中直接嵌入情感与富语言标签,用于控制语音的情感表达或在指定位置插入拟声效果(如笑声、叹息等),无需调整复杂的音频参数即可生成更具表现力的语音。 + + + **支持的模型**:仅 `qwen-audio-3.1-tts-flash`、`qwen-audio-3.0-tts-plus` 和 `qwen-audio-3.0-tts-flash`。 + + +**控制类标签** + +控制类标签用于设定语音的情感或风格。将标签写在文本中,标签会作用于其后的所有文本,直到遇到下一个控制类标签,或因句子较长被自动切分为止。 + +| **标签** | **说明** | +| -------------------------- | ------------ | +| `[sad]` | 悲伤 | +| `[amazed]` | 惊叹 | +| `[deep and loud shouting]` | 深沉大声呐喊 | +| `[trembling]` | 颤抖 | +| `[angry]` | 愤怒 | +| `[excited]` | 兴奋 | +| `[sarcastic]` | 讽刺 | +| `[curious]` | 好奇 | +| `[like dracula]` | 德古拉风格(低沉、阴森) | +| `[bored]` | 无聊 | +| `[tired]` | 疲惫 | +| `[scornful]` | 轻蔑 | +| `[shouting]` | 大喊 | +| `[asmr]` | ASMR 轻柔耳语 | +| `[panicked]` | 恐慌 | +| `[mischievously]` | 调皮 | +| `[empathetic]` | 共情 | +| `[whispers]` | 耳语 | +| `[reluctantly]` | 不情愿 | +| `[crying]` | 哭泣 | +| `[serious]` | 严肃 | +| `[very slowly]` | 非常缓慢地说话 | +| `[very fast]` | 非常快速地说话 | + +**富语言类标签** + +富语言类标签用于在文本的当前位置插入一段拟声效果,不影响前后文本的情感风格。 + +| **标签** | **说明** | +| ----------------- | ------ | +| `[gasp]` | 倒吸一口气 | +| `[sighing]` | 叹息 | +| `[clears throat]` | 清嗓 | +| `[giggles]` | 咯咯笑 | +| `[laughing]` | 大笑 | +| `[cough]` | 咳嗽 | +| `[snorts]` | 哼声、嗤笑 | + +**使用示例** + +以下示例展示如何在 `text` 参数中组合使用控制类标签和富语言类标签: + +`[excited]今天的天气真不错![laughing]我们一起出去玩吧!` + +上述文本中,`[excited]` 是控制类标签,作用于其后的所有文本,使语音带有兴奋的情感;`[laughing]` 是富语言类标签,在该位置插入一段笑声效果后继续合成后续文本。 + +您也可以在同一段文本中切换不同情感: + +`[serious]请注意安全事项。[excited]好了,现在让我们开始吧!` + +其中 `[serious]` 控制第一句为严肃语气,`[excited]` 从第二句起切换为兴奋语气。 + +### 文本预处理建议 + +`cosyvoice-v3-flash` 在合成包含点号(·)分隔数字段的文本时,可能出现漏读或重复念读的情况,例如连续的房号可能被读错。 + +将文本中的点号(·)替换为中文逗号(,)可规避该问题: + +- 原文:`主楼五楼·501房是PU·502房是OOO` +- 预处理后:`主楼五楼501房是PU,502房是OOO` + + + 此为模型层已知限制,仅在 `cosyvoice-v3-flash` 上确认,`cosyvoice-v2` 经交叉验证无此问题,不适用于 CosyVoice 其他型号或其他模型系列。在模型优化完成前,建议在代码侧对待合成文案统一做该预处理。 + + +## 支持的模型 + +支持以下模型: + +- **Qwen-Audio-TTS**:qwen-audio-3.0-tts-plus、qwen-audio-3.1-tts-flash、qwen-audio-3.0-tts-flash +- **CosyVoice**:cosyvoice-v3.5-plus、cosyvoice-v3.5-flash、cosyvoice-v3-plus、cosyvoice-v3-flash、cosyvoice-v2 +- **Qwen-TTS**: + + - **Qwen3-TTS-Instruct-Flash**:qwen3-tts-instruct-flash(稳定版,当前等同 qwen3-tts-instruct-flash-2026-01-26)、qwen3-tts-instruct-flash-2026-01-26(最新快照版) + - \*\*Qwen3-TTS-VD:\*\*qwen3-tts-vd-2026-01-26(最新快照版) + - \*\*Qwen3-TTS-VC:\*\*qwen3-tts-vc-2026-01-22(最新快照版) + - **Qwen3-TTS-Flash**:qwen3-tts-flash(稳定版,当前等同 qwen3-tts-flash-2025-11-27)、qwen3-tts-flash-2025-11-27、qwen3-tts-flash-2025-09-18 + - **Qwen-TTS**:qwen-tts(稳定版,当前等同 qwen-tts-2025-04-10)、qwen-tts-latest(最新版,当前等同 qwen-tts-2025-05-22)、qwen-tts-2025-05-22(快照版)、qwen-tts-2025-04-10(快照版) +- **MiniMax**:MiniMax/speech-2.8-hd、MiniMax/speech-02-hd、MiniMax/speech-2.8-turbo、MiniMax/speech-02-turbo + +## 支持的系统音色 + +不同模型支持的音色不同。将请求参数 `voice` 设为下表中 **voice 参数**列的值即可。 + +- [Qwen-Audio-TTS音色列表](/developer-guides/speech/voice-list/qwen-audio-tts) +- [CosyVoice音色列表](/developer-guides/speech/voice-list/cosyvoice) +- [Qwen-TTS音色列表](/developer-guides/speech/voice-list/qwen-tts#qwen-tts非实时语音合成音色列表) + +## API 参考 + +- [非实时语音合成-Qwen-Audio-TTS API参考](/api-reference/speech-synthesis/qwen-audio-tts/http-api) / [非实时语音合成-CosyVoice API参考](/api-reference/speech-synthesis/cosyvoice-nrt/http-api) +- [非实时语音合成-千问API参考](/api-reference/speech-synthesis/qwen-tts) +- [非实时语音合成-MiniMax API参考](/api-reference/speech-synthesis/minimax-tts) + +## 常见问题 + +### Q:音频文件链接的有效期是多久? + +A:音频文件链接在生成后 24 小时内有效。链接过期后,重新调用接口即可获取新链接。 diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-cloning.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-cloning.md new file mode 100644 index 0000000..1b27c2b --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-cloning.md @@ -0,0 +1,749 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 声音复刻 + +> 声音复刻(Voice Cloning)只需提供一段 10~20 秒的音频样本,即可生成高度相似的定制音色,无需模型训练。 + +声音复刻适用于个性化语音助手、品牌专属播报、有声内容定制化等场景。 + +## 概述 + +千问AI平台平台提供以下模型系列的声音复刻能力: + +- **Qwen-Audio-TTS**/**CosyVoice**:通过 DashScope SDK 或 HTTP API 创建音色,支持实时与非实时语音合成。 +- **MiniMax**:通过 HTTP API 创建音色,仅支持非实时语音合成。 +- **Qwen-Audio-Realtime**:Qwen-Audio-Realtime 是实时语音对话模型(非语音合成模型),声音复刻用于自定义对话模型回复时的 TTS 音色。通过 DashScope SDK 或 HTTP API 创建音色。 +- **Qwen-TTS**:通过 HTTP API 创建音色,支持实时与非实时语音合成。 + +如需了解各模型系列的详细对比和选型建议,请参见[语音合成](/developer-guides/speech/tts-models)。 + +## 前提条件 + +1. 已[配置 API Key](/api-reference/preparation/api-key)并将其[设置到环境变量](/api-reference/preparation/export-api-key-env)。 +2. 如果通过 DashScope SDK 调用,需要[安装最新版 SDK](/api-reference/preparation/install-sdk)。 +3. **准备音频文件**:音频需符合[音频要求](#音频要求)。 + +## 快速开始 + +声音复刻的使用分为以下三步: + +1. **准备音频**:准备一段符合[音频要求](#音频要求)的音频文件。 +2. **创建音色**:调用声音复刻接口上传音频创建音色,通过 `target_model` 指定绑定的语音合成模型。 +3. **使用音色合成语音**:调用语音合成接口,传入创建音色时返回的音色 ID。 + +### Qwen-TTS 声音复刻 + +示例使用本地音频文件 `voice.mp3`,运行时请替换为实际路径。 + + + 创建音色时的 `target_model` 必须与语音合成时使用的模型完全一致,否则合成将失败。 + + +#### 双向流式合成(实时) + +适用于 Qwen3-TTS-VC-Realtime 模型。参数详情见[实时流式语音合成](/developer-guides/speech/realtime-streaming)。 + + + + ```python + # pyaudio 安装方式: + # macOS: brew install portaudio && pip install pyaudio + # Ubuntu: sudo apt-get install python3-pyaudio (或 pip install pyaudio) + # CentOS: sudo yum install -y portaudio portaudio-devel && pip install pyaudio + # Windows: python -m pip install pyaudio + + import pyaudio + import os + import requests + import base64 + import pathlib + import threading + import time + import dashscope + from dashscope.audio.qwen_tts_realtime import QwenTtsRealtime, QwenTtsRealtimeCallback, AudioFormat + + TARGET_MODEL = "qwen3-tts-vc-realtime-2026-01-15" + VOICE_FILE = "voice.mp3" # 替换为你的音频文件 + + TEXT_TO_SYNTHESIZE = [ + 'Today we explore the wonders of speech synthesis.', + 'Each voice carries a unique character.', + 'With voice cloning, you can bring any text to life.', + "Let's create something amazing together." + ] + + def create_voice(file_path: str) -> str: + """创建克隆声音并返回声音标识符。""" + api_key = os.getenv("DASHSCOPE_API_KEY") + file_path_obj = pathlib.Path(file_path) + base64_str = base64.b64encode(file_path_obj.read_bytes()).decode() + data_uri = f"data:audio/mpeg;base64,{base64_str}" + + response = requests.post( + "https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization", + headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}, + json={ + "model": "qwen-voice-enrollment", + "input": { + "action": "create", + "target_model": TARGET_MODEL, + "preferred_name": "myvoice", + "audio": {"data": data_uri} + } + } + ) + return response.json()["output"]["voice"] + + class MyCallback(QwenTtsRealtimeCallback): + def __init__(self): + self.complete_event = threading.Event() + self._player = pyaudio.PyAudio() + self._stream = self._player.open(format=pyaudio.paInt16, channels=1, rate=24000, output=True) + + def on_event(self, response: dict) -> None: + if response.get("type") == "response.audio.delta": + audio_data = base64.b64decode(response["delta"]) + self._stream.write(audio_data) + elif response.get("type") == "session.finished": + self.complete_event.set() + + if __name__ == "__main__": + dashscope.api_key = os.getenv("DASHSCOPE_API_KEY") + callback = MyCallback() + tts = QwenTtsRealtime(model=TARGET_MODEL, callback=callback, + url="wss://maas.qianwenaiapi.com/api-ws/v1/realtime") + tts.connect() + tts.update_session(voice=create_voice(VOICE_FILE), + response_format=AudioFormat.PCM_24000HZ_MONO_16BIT, mode="server_commit") + + for text in TEXT_TO_SYNTHESIZE: + tts.append_text(text) + time.sleep(0.1) + + tts.finish() + callback.complete_event.wait() + ``` + + + + ```java + import com.alibaba.dashscope.audio.qwen_tts_realtime.*; + import com.google.gson.Gson; + import com.google.gson.JsonObject; + import java.io.*; + import java.net.HttpURLConnection; + import java.net.URL; + import java.nio.file.*; + import java.util.Base64; + import java.util.concurrent.CountDownLatch; + + public class Main { + private static final String TARGET_MODEL = "qwen3-tts-vc-realtime-2026-01-15"; + private static final String AUDIO_FILE = "voice.mp3"; // 替换为你的音频文件 + + public static String createVoice() throws Exception { + String apiKey = System.getenv("DASHSCOPE_API_KEY"); + byte[] bytes = Files.readAllBytes(Paths.get(AUDIO_FILE)); + String encoded = Base64.getEncoder().encodeToString(bytes); + String dataUri = "data:audio/mpeg;base64," + encoded; + + String jsonPayload = "{\"model\":\"qwen-voice-enrollment\",\"input\":{" + + "\"action\":\"create\",\"target_model\":\"" + TARGET_MODEL + "\"," + + "\"preferred_name\":\"myvoice\",\"audio\":{\"data\":\"" + dataUri + "\"}}}"; + + HttpURLConnection con = (HttpURLConnection) new URL( + "https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization").openConnection(); + con.setRequestMethod("POST"); + con.setRequestProperty("Authorization", "Bearer " + apiKey); + con.setRequestProperty("Content-Type", "application/json"); + con.setDoOutput(true); + try (OutputStream os = con.getOutputStream()) { + os.write(jsonPayload.getBytes("UTF-8")); + } + + BufferedReader br = new BufferedReader(new InputStreamReader(con.getInputStream(), "UTF-8")); + StringBuilder response = new StringBuilder(); + String line; + while ((line = br.readLine()) != null) response.append(line); + return new Gson().fromJson(response.toString(), JsonObject.class) + .getAsJsonObject("output").get("voice").getAsString(); + } + + public static void main(String[] args) throws Exception { + CountDownLatch latch = new CountDownLatch(1); + QwenTtsRealtimeParam param = QwenTtsRealtimeParam.builder() + .model(TARGET_MODEL) + .url("wss://maas.qianwenaiapi.com/api-ws/v1/realtime") + .apikey(System.getenv("DASHSCOPE_API_KEY")) + .build(); + + QwenTtsRealtime tts = new QwenTtsRealtime(param, new QwenTtsRealtimeCallback() { + public void onEvent(JsonObject msg) { + if (msg.get("type").getAsString().equals("session.finished")) latch.countDown(); + } + }); + tts.connect(); + + QwenTtsRealtimeConfig config = QwenTtsRealtimeConfig.builder() + .voice(createVoice()) + .responseFormat(QwenTtsRealtimeAudioFormat.PCM_24000HZ_MONO_16BIT) + .mode("server_commit").build(); + tts.updateSession(config); + + for (String text : new String[]{ + "Today we explore the wonders of speech synthesis.", + "Each voice carries a unique character.", + "With voice cloning, you can bring any text to life.", + "Let's create something amazing together."}) { + tts.appendText(text); + Thread.sleep(100); + } + tts.finish(); + latch.await(); + } + } + ``` + + + +#### 非流式合成 + +适用于 Qwen3-TTS-VC 模型。详见 [Qwen TTS](/api-reference/speech-synthesis/qwen-tts)。 + + + + ```python + import os + import requests + import base64 + import pathlib + import dashscope + + TARGET_MODEL = "qwen3-tts-vc-2026-01-22" + VOICE_FILE = "voice.mp3" + + def create_voice(file_path: str) -> str: + api_key = os.getenv("DASHSCOPE_API_KEY") + base64_str = base64.b64encode(pathlib.Path(file_path).read_bytes()).decode() + data_uri = f"data:audio/mpeg;base64,{base64_str}" + + response = requests.post( + "https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization", + headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}, + json={ + "model": "qwen-voice-enrollment", + "input": {"action": "create", "target_model": TARGET_MODEL, + "preferred_name": "myvoice", "audio": {"data": data_uri}} + } + ) + return response.json()["output"]["voice"] + + if __name__ == "__main__": + dashscope.base_http_api_url = 'https://maas.qianwenaiapi.com/api/v1' + response = dashscope.MultiModalConversation.call( + model=TARGET_MODEL, + api_key=os.getenv("DASHSCOPE_API_KEY"), + text="Today we explore the wonders of speech synthesis.", + voice=create_voice(VOICE_FILE), + stream=False + ) + print(response) + ``` + + + + **步骤一:创建音色** + + ```bash + # 将 voice.mp3 替换为实际音频文件路径 + + AUDIO_BASE64=$(base64 -i voice.mp3) + + curl -X POST 'https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization' \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H 'Content-Type: application/json' \ + -d '{ + "model": "qwen-voice-enrollment", + "input": { + "action": "create", + "target_model": "qwen3-tts-vc-2026-01-22", + "preferred_name": "guanyu", + "audio": {"data": "data:audio/mpeg;base64,'$AUDIO_BASE64'"} + } + }' + ``` + + **步骤二:使用复刻音色合成语音** + + 将上一步返回的 `voice` 值填入以下请求中。 + + Qwen-TTS 系列创建音色接口的返回体中,音色 ID 位于 `output.voice` 字段;其他模型系列的创建音色接口返回的是 `voice_id` 字段,两者均表示音色 ID,注意字段名差异。 + + ```bash + # 将 YOUR_VOICE_ID 替换为上一步返回的 voice 值 + + curl -X POST 'https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation' \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H 'Content-Type: application/json' \ + -d '{ + "model": "qwen3-tts-vc-2026-01-22", + "input": { + "text": "今天天气怎么样?", + "voice": "YOUR_VOICE_ID" + } + }' + ``` + + + + ```java + import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation; + import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam; + import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult; + import com.alibaba.dashscope.utils.Constants; + import com.google.gson.Gson; + import com.google.gson.JsonObject; + + import java.io.*; + import java.net.HttpURLConnection; + import java.net.URL; + import java.nio.file.*; + import java.nio.charset.StandardCharsets; + import java.util.Base64; + + public class Main { + private static final String TARGET_MODEL = "qwen3-tts-vc-2026-01-22"; + private static final String AUDIO_FILE = "voice.mp3"; // 替换为你的音频文件 + + public static String createVoice() throws Exception { + String apiKey = System.getenv("DASHSCOPE_API_KEY"); + byte[] bytes = Files.readAllBytes(Paths.get(AUDIO_FILE)); + String encoded = Base64.getEncoder().encodeToString(bytes); + String dataUri = "data:audio/mpeg;base64," + encoded; + + String jsonPayload = "{\"model\":\"qwen-voice-enrollment\",\"input\":{" + + "\"action\":\"create\",\"target_model\":\"" + TARGET_MODEL + "\"," + + "\"preferred_name\":\"myvoice\",\"audio\":{\"data\":\"" + dataUri + "\"}}}"; + + HttpURLConnection con = (HttpURLConnection) new URL( + "https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization").openConnection(); + con.setRequestMethod("POST"); + con.setRequestProperty("Authorization", "Bearer " + apiKey); + con.setRequestProperty("Content-Type", "application/json"); + con.setDoOutput(true); + try (OutputStream os = con.getOutputStream()) { + os.write(jsonPayload.getBytes(StandardCharsets.UTF_8)); + } + + BufferedReader br = new BufferedReader(new InputStreamReader(con.getInputStream(), StandardCharsets.UTF_8)); + StringBuilder response = new StringBuilder(); + String line; + while ((line = br.readLine()) != null) response.append(line); + return new Gson().fromJson(response.toString(), JsonObject.class) + .getAsJsonObject("output").get("voice").getAsString(); + } + + public static void main(String[] args) { + try { + Constants.baseHttpApiUrl = "https://maas.qianwenaiapi.com/api/v1"; + MultiModalConversation conv = new MultiModalConversation(); + MultiModalConversationParam param = MultiModalConversationParam.builder() + .apiKey(System.getenv("DASHSCOPE_API_KEY")) + .model(TARGET_MODEL) + .text("Today we explore the wonders of speech synthesis.") + .parameter("voice", createVoice()) + .build(); + MultiModalConversationResult result = conv.call(param); + String audioUrl = result.getOutput().getAudio().getUrl(); + System.out.println("Audio URL: " + audioUrl); + + // 下载音频 + try (InputStream in = new URL(audioUrl).openStream(); + FileOutputStream out = new FileOutputStream("output.wav")) { + byte[] buffer = new byte[1024]; + int bytesRead; + while ((bytesRead = in.read(buffer)) != -1) { + out.write(buffer, 0, bytesRead); + } + System.out.println("音频已保存至 output.wav"); + } + } catch (Exception e) { + System.out.println("Error: " + e.getMessage()); + } + System.exit(0); + } + } + ``` + + + +### Qwen-Audio-TTS 声音复刻 + +**步骤一:创建音色** + +调用声音复刻 API 上传音频并创建音色。`url` 参数传入音频文件的可访问 URL 地址,`prefix` 参数作为音色名称前缀。 + +```bash +# 将 url 替换为实际音频文件的可访问地址 + +curl -X POST 'https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization' \ +-H "Authorization: Bearer $DASHSCOPE_API_KEY" \ +-H "Content-Type: application/json" \ +-d '{ + "model": "voice-enrollment", + "input": { + "action": "create_voice", + "target_model": "qwen-audio-3.0-tts-flash", + "prefix": "myvoice", + "url": "https://your-audio-url.wav" + } +}' +``` + +**步骤二:使用复刻音色合成语音** + +将上一步返回的 `voice_id` 值填入以下请求中。 + +```python +# coding=utf-8 +import dashscope +from dashscope.audio.tts_v2 import * +import os + +dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY') +# 声音复刻、语音合成要使用相同的模型 +model = "qwen-audio-3.0-tts-flash" +# 将 voice 参数替换为声音复刻生成的专属音色 +voice = "voice_id" + +synthesizer = SpeechSynthesizer(model=model, voice=voice) +audio = synthesizer.call("今天天气怎么样?") +print('[Metric] requestId为:{},首包延迟为:{}毫秒'.format( + synthesizer.get_last_request_id(), + synthesizer.get_first_package_delay())) + +with open('output.mp3', 'wb') as f: + f.write(audio) +``` + +### Qwen-Audio-Realtime 声音复刻 + + + Qwen-Audio-Realtime 是实时语音对话模型(非语音合成模型),声音复刻用于自定义对话模型回复时的 TTS 音色,音频要求与 Qwen-Audio-TTS 相同。 + + +**步骤一:创建音色** + +通过 HTTP API 调用声音复刻接口,上传音频并创建音色。`target_model` 参数填入实时语音对话模型名称,`prefix` 参数作为音色名称前缀。 + +```bash +curl -X POST 'https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization' \ +-H "Authorization: Bearer $DASHSCOPE_API_KEY" \ +-H "Content-Type: application/json" \ +-d '{ + "model": "voice-enrollment", + "input": { + "action": "create_voice", + "target_model": "qwen-audio-3.1-realtime-plus", + "prefix": "myvoice", + "url": "https://your-audio-url.wav" + } +}' +``` + +**步骤二:在实时语音对话中使用复刻音色** + +将上一步返回的 `voice_id` 填入实时语音对话 `session.update` 事件的 `voice` 参数中。 + +```json +{ + "type": "session.update", + "session": { + "voice": "qwen-audio-3.1-realtime-plus-myvoice-xxxxxx" + } +} +``` + +### CosyVoice 声音复刻 + +CosyVoice 声音复刻通过专用的声音复刻 API 进行操作,同样遵循"创建音色 - 使用音色合成"的流程。 + +**步骤一:创建音色** + +调用声音复刻 API 上传音频并创建音色。`url` 参数传入音频文件的可访问 URL 地址,`prefix` 参数作为音色名称前缀。 + +```bash +# 将 url 替换为实际音频文件的可访问地址 + +curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization \ +-H "Authorization: Bearer $DASHSCOPE_API_KEY" \ +-H "Content-Type: application/json" \ +-d '{ + "model": "voice-enrollment", + "input": { + "action": "create_voice", + "target_model": "cosyvoice-v3-plus", + "prefix": "myvoice", + "url": "https://your-audio-url.wav", + "language_hints": ["zh"] + } +}' +``` + +**步骤二:使用复刻音色合成语音** + +将上一步返回的 `voice` 值填入以下请求中。 + +```bash +# 将 YOUR_VOICE_ID 替换为上一步返回的 voice 值 + +curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer \ +-H "Authorization: Bearer $DASHSCOPE_API_KEY" \ +-H "Content-Type: application/json" \ +-d '{ + "model": "cosyvoice-v3-plus", + "input": { + "text": "今天天气怎么样?", + "voice": "YOUR_VOICE_ID", + "format": "wav", + "sample_rate": 24000 + } +}' +``` + +### MiniMax 音色复刻 + + + 以下价格为目录价。具体优惠活动及折扣价格请前往[模型市场](https://www.qianwenai.com/models)查看。 + + +提交复刻请求后,系统会生成一段试听音频(按同步语音合成单价计费)。首次使用复刻音色进行语音合成时,需支付 9.9 元音色解锁费用。 + +**步骤一:创建音色** + +调用音色复刻 API 上传音频并创建音色。`voice_id` 参数用于指定新音色的 ID,`audio_url` 参数传入音频文件的可访问 URL 地址。 + +```bash +# 将 audio_url 替换为实际音频文件的可访问地址 +# 将 voice_id 替换为自定义的音色 ID + +curl -X POST 'https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation' \ +-H "Authorization: Bearer $DASHSCOPE_API_KEY" \ +-H 'Content-Type: application/json; charset=utf-8' \ +-d '{ + "input": { + "action": "voice_clone", + "voice_id": "my-custom-voice", + "audio_url": "https://your-audio-url.wav", + "text": "你说是什么就是什么" + }, + "model": "MiniMax/speech-2.8-turbo" + }' +``` + +**步骤二:使用复刻音色合成语音** + +将上一步指定的 `voice_id` 值填入以下请求中。 + +```bash +# 将 voice_id 替换为上一步指定的音色 ID + +curl -X POST "https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation" \ +-H "Authorization: Bearer $DASHSCOPE_API_KEY" \ +-H "Content-Type: application/json" \ +-d '{ + "model": "MiniMax/speech-2.8-turbo", + "input": { + "text": "今天天气怎么样?", + "voice_setting": { + "voice_id": "my-custom-voice", + "speed": 1, + "vol": 1, + "pitch": 0 + }, + "audio_setting": { + "sample_rate": 32000, + "bitrate": 128000, + "format": "mp3", + "channel": 1 + } + } +}' +``` + +## 音频要求 + +输入音频的质量直接决定复刻效果。不同模型系列对音频的具体要求有所差异,请按照目标模型的要求准备音频样本。 + + + +| 项目 | 要求 | +| -------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | +| **支持格式** | WAV(16bit)、MP3、M4A | +| **音频时长** | 推荐 10\~20 秒,最长不超过 60 秒 | +| **文件大小** | 不超过 10 MB | +| **采样率** | 16 kHz 及以上 | +| **声道** | 单声道或双声道。双声道音频仅处理首声道,请确保首声道包含有效人声。 | +| **内容** | 音频必须包含至少 5 秒连续清晰的朗读内容(无背景音),其余部分仅允许短暂停顿(不超过 2 秒)。整段音频应避免出现背景音乐、环境噪音或其他人声。请使用正常语速的说话音频,不要上传歌曲或唱歌录音。 | +| **支持语言** | 中文(普通话、广东话、重庆话、东北话、甘肃话、贵州话、浙江话、河北话、河南话、湖北话、湖南话、江西话、宁波话、宁夏话、青岛话、陕西话、山西话、山东话、上海话、四川话、云南话)、英语、日语、韩语、俄语、法语、德语、葡萄牙语、泰语、印尼语、越南语、西班牙语、意大利语、马来西亚语、菲律宾语、阿拉伯语 | + + + +| 项目 | 要求 | +| -------- | -------------------------------------------------------------------------------------------------- | +| **支持格式** | WAV(16bit)、MP3、M4A | +| **音频时长** | 推荐 10\~20 秒,最长不超过 60 秒 | +| **文件大小** | 不超过 10 MB | +| **采样率** | 16 kHz 及以上 | +| **声道** | 单声道或双声道。双声道音频仅处理首声道,请确保首声道包含有效人声。 | +| **内容** | 音频必须包含至少 5 秒连续清晰的朗读内容(无背景音),其余部分仅允许短暂停顿(不超过 2 秒)。整段音频应避免出现背景音乐、环境噪音或其他人声。请使用正常语速的说话音频,不要上传歌曲或唱歌录音。 | +| **支持语言** | 因驱动音色的语音合成模型(通过 `target_model` 参数指定)而异,详见下方说明 | + + **各模型支持的语言**: + + - **cosyvoice-v1、cosyvoice-v2**:中文(普通话)、英文 + - **cosyvoice-v3-flash**:中文(普通话、广东话、东北话、甘肃话、贵州话、河南话、湖北话、江西话、闽南话、宁夏话、山西话、陕西话、山东话、上海话、四川话、天津话、云南话)、英文、法语、德语、日语、韩语、俄语、葡萄牙语、泰语、印尼语、越南语 + - **cosyvoice-v3-plus**:中文(普通话)、英文、法语、德语、日语、韩语、俄语 + - **cosyvoice-v3.5-plus、cosyvoice-v3.5-flash**:中文(普通话、广东话、河南话、湖北话、闽南话、宁夏话、陕西话、山东话、上海话、四川话)、英文、法语、德语、日语、韩语、俄语、葡萄牙语、泰语、印尼语、越南语 + + + +| 项目 | 要求 | +| -------- | -------------------------------------------------------------------------------------------------- | +| **支持格式** | WAV(16bit)、MP3、M4A | +| **音频时长** | 推荐 10\~20 秒,最长不超过 60 秒 | +| **文件大小** | 不超过 10 MB | +| **采样率** | 24 kHz 及以上 | +| **声道** | 单声道 | +| **内容** | 音频必须包含至少 3 秒连续清晰的朗读内容(无背景音),其余部分仅允许短暂停顿(不超过 2 秒)。整段音频应避免出现背景音乐、环境噪音或其他人声。请使用正常语速的说话音频,不要上传歌曲或唱歌录音。 | +| **支持语言** | 中文、英文、德语、意大利语、葡萄牙语、西班牙语、日语、韩语、法语、俄语 | + + + +| 项目 | 要求 | +| -------- | ---------------------------------------------------------------------------------- | +| **支持格式** | MP3、M4A、WAV | +| **音频时长** | 不低于 10 秒,最长不超过 5 分钟 | +| **文件大小** | 不超过 20 MB | +| **内容** | 音频应包含连续清晰的朗读内容(无背景音),停顿时长不超过 2 秒。整段音频应避免出现背景音乐、环境噪音或其他人声。请使用正常语速的说话音频,不要上传歌曲或唱歌录音。 | +| **支持语言** | 无特殊限制 | + + + + + 为获得最佳复刻效果,建议参照[录音建议](#录音建议)准备样本。 + + +## 录音建议 + +高质量的输入音频是获得优质复刻效果的基础。以下从录音设备、录音环境、录音文案和操作流程四个方面提供建议。 + +### 录音设备 + +可使用手机、数字录音笔、专业录音机等。建议使用支持高采样率(24 kHz 及以上)录音的设备,以满足音频要求。 + +### 录音环境 + +**场地** + +- 建议在 10 平方米以内的小型封闭空间录音。 +- 优先选择配有吸音材料(如吸音棉、地毯、窗帘)的房间。 +- 避免空旷大厅、会议室、教室等高混响场所。 + +**噪音控制** + +- 室外噪音:关闭门窗,避免交通、施工等干扰。 +- 室内噪音:关闭空调、风扇、日光灯镇流器等设备;可通过手机录制环境音并放大播放,识别潜在噪音源。 + +**混响控制** + +- 混响会导致声音模糊、清晰度下降。 +- 减少光滑表面反射:拉上窗帘、打开衣柜门、铺放衣物或床单覆盖桌面/柜面。 +- 利用不规则物体(如书架、软包家具)实现声波漫反射。 + +### 录音文案 + +- 内容无特殊限制,建议与目标应用场景一致。 +- 避免短句(如"你好"、"是的"),应使用完整句子。 +- 保持语义连贯,朗读时避免频繁停顿(建议至少连续 3 秒无中断)。 +- 录音的开头和结尾部分应保持与中间段落一致的语速,避免因开头或结尾语速过快导致复刻后语音合成时出现卡顿现象。 +- 可加入适当情绪表达(如温暖、亲切、严肃),避免机械朗读。 +- 不包含敏感词汇(如政治、色情、暴力相关内容),否则会导致复刻失败。 + +### 操作建议 + +以普通卧室为例: + +1. 关闭门窗,隔绝外部噪音。 +2. 关闭空调、电扇等电器。 +3. 拉上窗帘,减少玻璃反射。 +4. 在桌面铺放衣物或毛毯,降低桌面反射。 +5. 提前熟悉文案,设定角色语气,自然演绎。 +6. 与录音设备保持约 10 厘米距离,避免喷麦或信号过弱。 + +## 管理自定义音色 + +音色创建完成后,您可以通过 API 对已有音色进行查询和管理(Qwen-Audio-TTS、Qwen-Audio-Realtime、Qwen-TTS、CosyVoice 和 MiniMax 支持)。Qwen-Audio-Realtime 的音色查询与删除接口与 Qwen-Audio-TTS 完全相同。 + +- **查询音色列表**:获取当前账号下所有自定义音色的列表。 +- **查询音色详情**:查看指定音色的详细信息,如创建时间、绑定的语音合成模型等。 +- **删除音色**:删除不再需要的自定义音色,释放配额。 + +各模型的 API 接口和参数详情请参见 [API 参考](#api-参考)。 + +## 配额与计费 + +### 音色配额与自动清理 + +- **音色总数限制**:每个千问AI平台账号下,Qwen-Audio-TTS / Qwen-Audio-Realtime / CosyVoice 与 Qwen-TTS 分别最多可创建 1000 个自定义音色(两类配额独立计算)。达到上限后,新的创建请求将直接失败并返回错误,系统不会自动删除最早创建的复刻音色。如需创建新音色,请先删除不需要的音色以释放配额,或等待未被使用的音色自动清理(详见下方自动清理规则)。 +- **上限后创建行为**:达到 1000 个音色上限后,继续调用创建音色接口将直接返回失败错误,系统不会自动淘汰最早创建的复刻音色来腾出空间。需手动删除不需要的音色释放配额后才能继续创建。 +- **自动清理规则**:若单个音色在过去 1 年内未被用于任何语音合成请求,系统将自动删除该音色。 + + + MiniMax 创建的音色不受上述配额与清理规则约束。 + + +### 计费规则 + +- **Qwen-Audio-TTS / Qwen-Audio-Realtime / CosyVoice**:创建音色免费。 +- **MiniMax**:声音复刻本身不计费。首次使用复刻音色进行语音合成时,扣除 9.9 元音色解锁费用,且无免费额度。 +- **Qwen-TTS**:按 0.01 元/个计费,创建失败不计费。 + +**免费额度**: + +- 千问AI平台开通后 90 天内,可享 1000 次免费音色创建机会。 +- 创建失败不占用免费次数。 +- 删除音色不会恢复免费次数。 +- 免费额度用完或超出 90 天有效期后,创建音色将按 0.01 元/个的价格计费。 + +## 适用范围 + +支持的模型: + +- **Qwen-Audio-TTS**:qwen-audio-3.0-tts-plus、qwen-audio-3.1-tts-flash、qwen-audio-3.0-tts-flash +- **Qwen-Audio-Realtime**:qwen-audio-3.1-realtime-plus、qwen-audio-3.0-realtime-plus、qwen-audio-3.0-realtime-flash +- **CosyVoice**:cosyvoice-v3.5-plus、cosyvoice-v3.5-flash、cosyvoice-v3-plus、cosyvoice-v3-flash、cosyvoice-v2、cosyvoice-v1 +- **MiniMax**:MiniMax/speech-2.8-hd、MiniMax/speech-02-hd、MiniMax/speech-2.8-turbo、MiniMax/speech-02-turbo +- **Qwen-TTS**: + - **Qwen3-TTS-VC-Realtime**:qwen3-tts-vc-realtime-2026-01-15(最新快照版)、qwen3-tts-vc-realtime-2025-11-27(快照版) + - **Qwen3-TTS-VC**:qwen3-tts-vc-2026-01-22(最新快照版) + +## 常见问题 + +### Q:创建音色后可以用于不同的语音合成模型吗? + +不可以。音色在创建时通过 `target_model` 绑定到特定的语音合成模型,不能跨模型使用。如果您需要在多个模型上使用同一段音频的声音,请为每个模型分别创建音色。 + +### Q:复刻音色的有效期是多久? + +Qwen-Audio-TTS、Qwen-Audio-Realtime、Qwen-TTS 和 CosyVoice 创建的音色默认长期有效,但长时间未使用的音色可能会被系统清理。建议妥善保存音色 ID,需要时可通过查询接口确认音色是否仍然可用。 + +### Q:音频质量不好会影响复刻效果吗? + +会的。输入音频的质量直接影响复刻效果。背景噪音、混响、多人声等问题都会降低复刻音色的相似度和自然度。建议参照[音频要求](#音频要求)和[录音建议](#录音建议)准备样本。 + +## API 参考 + +[声音复刻 API](/api-reference/speech-synthesis/voice-cloning/http-api) diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-design.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-design.md new file mode 100644 index 0000000..1268fe0 --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-design.md @@ -0,0 +1,328 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# 声音设计 + +> 通过文本描述创建自定义音色,无需音频样本,支持 Qwen-TTS、CosyVoice 和 Qwen-Audio-TTS 模型。 + +声音设计(Voice Design)无需音频样本,仅通过自然语言描述即可创建定制化音色。 + +## 概述 + +声音设计适用于快速原型验证、创意内容生产、游戏角色配音等场景。 + +千问AI平台平台提供以下模型系列的声音设计能力: + +- **CosyVoice**:支持实时与非实时语音合成(v3.5 系列)。 +- **Qwen-TTS**:支持实时与非实时语音合成,声音描述长度上限更高(2048 字符)。 +- **Qwen-Audio-TTS**:支持实时与非实时语音合成(qwen-audio-3.0-tts-plus、qwen-audio-3.0-tts-flash)。 + +如果您已有音频样本,请参见[声音复刻](/developer-guides/speech/voice-cloning)。如需了解模型选型建议,请参见[语音合成](/developer-guides/speech/tts-models)。 + + + Voice design 中的 `target_model` 必须与合成时的 `model` 一致,否则会导致调用失败。 + + +## 前提条件 + +- 已[配置 API Key](/api-reference/preparation/api-key) 并将其[设置到环境变量](/api-reference/preparation/export-api-key-env)。 +- 如果通过 DashScope SDK 调用,需要[安装最新版 SDK](/api-reference/preparation/install-sdk)。 + +## 快速开始 + +声音设计的基本流程为:描述 -> 创建 -> 使用。 + +1. **编写声音描述**:用自然语言描述期望的声音特质。详细的编写指南请参见[编写声音描述](#编写声音描述)。 +2. **创建音色**:调用声音设计接口,系统根据描述生成音色并返回预览音频。建议试听确认效果后再使用。 +3. **使用音色合成语音**:调用语音合成接口,传入音色 ID 进行语音合成。 + +### Qwen-TTS 声音设计 + +以下示例演示如何创建音色并用于语音合成。 + + + 建议先试听返回的预览音频,确认效果后再用于合成,以降低调用成本。 + + +#### 创建音色 + + + + ```python + import os + import requests + import dashscope + + # ======= 常量配置 ======= + DEFAULT_TARGET_MODEL = "qwen3-tts-vd-2026-01-26" # 声音设计、语音合成要使用相同的模型 + DEFAULT_PREFERRED_NAME = "custom_voice" + + # 声音描述:用自然语言描述期望的声音特质 + VOICE_PROMPT = "年轻活泼的女性声音,语速较快,带有明显的上扬语调,适合介绍时尚产品。" + + def create_voice_by_design(voice_prompt: str, + target_model: str = DEFAULT_TARGET_MODEL, + preferred_name: str = DEFAULT_PREFERRED_NAME) -> str: + """ + 通过声音描述创建音色,并返回 voice 参数 + """ + api_key = os.getenv("DASHSCOPE_API_KEY") + + url = "https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization" + payload = { + "model": "qwen-voice-design", # 不要修改该值 + "input": { + "action": "create", + "target_model": target_model, + "preferred_name": preferred_name, + "voice_prompt": voice_prompt, + "preview_text": preview_text + }, + "parameters": { + "sample_rate": 24000, + "response_format": "wav" + } + } + headers = { + "Authorization": f"Bearer {api_key}", + "Content-Type": "application/json" + } + + resp = requests.post(url, json=payload, headers=headers) + if resp.status_code != 200: + raise RuntimeError(f"创建 voice 失败: {resp.status_code}, {resp.text}") + + result = resp.json() + # 返回预览音频(可选:先试听确认效果) + preview_audio = result.get("output", {}).get("preview_audio") + if preview_audio: + print(f"预览音频URL: {preview_audio}") + + try: + return result["output"]["voice"] + except (KeyError, ValueError) as e: + raise RuntimeError(f"解析 voice 响应失败: {e}") + + if __name__ == '__main__': + dashscope.base_http_api_url = 'https://maas.qianwenaiapi.com/api/v1' + + voice_id = create_voice_by_design(VOICE_PROMPT) + print(f"创建的音色ID: {voice_id}") + + text = "大家好,欢迎来到我们的直播间!今天给大家推荐的这款产品真的超级好用。" + response = dashscope.MultiModalConversation.call( + model=DEFAULT_TARGET_MODEL, + api_key=os.getenv("DASHSCOPE_API_KEY"), + text=text, + voice=voice_id, + stream=False + ) + print(response) + ``` + + + + **步骤一:通过声音描述创建音色** + + ```bash + curl -X POST 'https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization' \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H 'Content-Type: application/json' \ + -d '{ + "model": "qwen-voice-design", + "input": { + "action": "create", + "target_model": "qwen3-tts-vd-2026-01-26", + "preferred_name": "custom_voice", + "voice_prompt": "年轻活泼的女性声音,语速较快,带有明显的上扬语调,适合介绍时尚产品。" + } + }' + ``` + + **步骤二:使用设计音色合成语音** + + 将上一步返回的 `voice` 值填入以下请求中。 + + ```bash + # 将 YOUR_VOICE_ID 替换为上一步返回的 voice 值 + + curl -X POST 'https://maas.qianwenaiapi.com/api/v1/services/aigc/multimodal-generation/generation' \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H 'Content-Type: application/json' \ + -d '{ + "model": "qwen3-tts-vd-2026-01-26", + "input": { + "text": "大家好,欢迎来到我们的直播间!今天给大家推荐的这款产品真的超级好用。", + "voice": "YOUR_VOICE_ID" + } + }' + ``` + + + +### CosyVoice 声音设计 + +CosyVoice 同样支持通过文本描述创建音色,使用流程与 Qwen-TTS 类似。 + + + CosyVoice 声音设计基于 FunAudioGen-VD 模型能力。相同描述文本(Prompt)设计的音色可能存在差异,建议多次生成后择优使用。 + + +#### 创建音色 + +调用声音复刻/设计 API,通过 `voice_prompt` 参数传入声音描述,`preview_text` 参数指定预览音频朗读的文本。 + + + + **步骤一:通过声音描述创建音色** + + ```bash + curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/customization \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "model": "voice-enrollment", + "input": { + "action": "create_voice", + "target_model": "cosyvoice-v3-plus", + "voice_prompt": "沉稳的中年男性播音员,音色低沉浑厚,富有磁性,语速平稳,吐字清晰,适合用于新闻播报或纪录片解说。", + "preview_text": "各位听众朋友,大家好,欢迎收听晚间新闻。", + "prefix": "announcer", + "language_hints": ["zh"] + }, + "parameters": { + "sample_rate": 24000, + "response_format": "wav" + } + }' + ``` + + **步骤二:使用设计音色合成语音** + + 将上一步返回的 `voice` 值填入以下请求中。 + + ```bash + # 将 YOUR_VOICE_ID 替换为上一步返回的 voice 值 + + curl -X POST https://maas.qianwenaiapi.com/api/v1/services/audio/tts/SpeechSynthesizer \ + -H "Authorization: Bearer $DASHSCOPE_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "model": "cosyvoice-v3-plus", + "input": { + "text": "各位听众朋友,大家好,欢迎收听晚间新闻。", + "voice": "YOUR_VOICE_ID", + "format": "wav", + "sample_rate": 24000 + } + }' + ``` + + + +如需使用自定义音色进行实时流式合成,请参见[流式语音合成](/developer-guides/speech/realtime-streaming)。完整的 API 参数说明及更多操作(列出、查询、删除),请参见 [Voice design API 参考](/api-reference/speech-synthesis/voice-design/http-api)。 + +## 配额与计费 + +### 音色配额与自动清理 + +- **音色总数限制**:每个千问AI平台账号下,CosyVoice 与 Qwen-TTS 分别最多可创建 1000 个自定义音色(两类配额独立计算)。 +- **自动清理规则**:若单个音色在过去 1 年内未被用于任何语音合成请求,系统将自动删除该音色。 + +### 计费规则 + + + 以下价格为目录价。具体优惠活动及折扣价格请前往[模型市场](https://www.qianwenai.com/models)查看。 + + +- **CosyVoice**:创建音色免费。 +- **Qwen-TTS**:按 0.2 元/个计费,创建失败不计费。 + + **免费额度**: + + - 千问AI平台开通后 90 天内,可享 10 次免费音色创建机会。 + - 创建失败不占用免费次数。 + - 删除音色不会恢复免费次数。 + - 免费额度用完或超出 90 天有效期后,创建音色将按 0.2 元/个的价格计费。 + +## 适用范围 + +支持的模型: + +- **CosyVoice**:cosyvoice-v3.5-plus、cosyvoice-v3.5-flash、cosyvoice-v3-plus、cosyvoice-v3-flash +- **Qwen-Audio-TTS**:qwen-audio-3.0-tts-plus、qwen-audio-3.0-tts-flash +- **Qwen-TTS**: + - **Qwen3-TTS-VD-Realtime**:qwen3-tts-vd-realtime-2026-01-15(最新快照版)、qwen3-tts-vd-realtime-2025-12-16(快照版) + - **Qwen3-TTS-VD**:qwen3-tts-vd-2026-01-26(最新快照版) + + + * CosyVoice 声音设计基于 FunAudioGen-VD 模型能力。 + * 相同描述文本(Prompt)设计的音色可能存在差异,建议多次生成后择优使用。 + + +## 编写声音描述 + +声音描述(`voice_prompt`)直接决定生成音色的效果。清晰、具体的描述能帮助模型更准确地生成目标音色。 + +### 要求与限制 + +- **长度限制**:`voice_prompt` 的最大长度因模型而异:CosyVoice 不超过 500 个字符,Qwen-TTS 不超过 2048 个字符。 +- **支持语言**:描述文本仅支持中文和英文。 + +### 核心原则 + +1. **具体而非模糊**:使用描绘声音特质的词语,如"低沉""清脆""语速偏快",避免"好听""普通"等主观或模糊的表述。 +2. **多维而非单一**:好的描述通常涵盖多个维度(如性别、年龄、情感等)。仅写"女声"过于宽泛,难以生成有特色的音色。 +3. **客观而非主观**:聚焦声音的物理和感知特征。例如,用"音调偏高,带有活力"代替"我最喜欢的声音"。 +4. **原创而非模仿**:描述声音的特质,而非要求模仿特定人物(如名人、演员)。模型不支持模仿,且可能涉及版权风险。 +5. **简洁而非冗余**:确保每个词都有明确作用,避免重复的同义词或无意义的修饰。 + +### 描述维度参考 + +建议组合以下维度来描述声音,维度越丰富,生成效果越精准。 + +| 维度 | 描述示例 | +| -- | ---------------------------------------------------------- | +| 性别 | 男性、女性、中性 | +| 年龄 | 儿童(5-12 岁)、青少年(13-18 岁)、青年(19-35 岁)、中年(36-55 岁)、老年(55 岁以上) | +| 音调 | 高音、中音、低音、偏高、偏低 | +| 语速 | 快速、中速、缓慢、偏快、偏慢 | +| 情感 | 开朗、沉稳、温柔、严肃、活泼、冷静、治愈 | +| 特点 | 有磁性、清脆、沙哑、圆润、甜美、浑厚、有力 | +| 用途 | 新闻播报、广告配音、有声书、动画角色、语音助手、纪录片解说 | + +### 示例 + +- 标准播音风格:吐字清晰精准,字正腔圆 +- 年轻活泼的女性声音,语速较快,带有明显的上扬语调,适合介绍时尚产品 +- 沉稳的中年男性,语速缓慢,音色低沉有磁性,适合朗读新闻或纪录片解说 +- 温柔知性的女性,30 岁左右,语调平和,适合有声书朗读 +- 可爱的儿童声音,大约 8 岁女孩,说话略带稚气,适合动画角色配音 + +## 管理自定义音色 + +声音设计和声音复刻创建的音色共用同一套管理接口。您可以查询音色列表、查看音色详情或删除不再需要的音色。 + +API 接口和参数详情请参见 [Voice design API 参考](/api-reference/speech-synthesis/voice-design/http-api)。 + +## 常见问题 + +### Q:相同的声音描述每次生成的音色一样吗? + +不一定。声音设计具有随机性,相同描述可能生成略有差异的音色。建议多次生成后试听,择优使用。 + +### Q:声音描述可以使用哪些语言? + +目前声音描述(`voice_prompt`)仅支持中文和英文,但生成的音色可用于合成多种语言的语音。 + +### Q:声音设计和声音复刻有什么区别? + +声音设计通过文本描述从零创建音色,无需音频样本,适合设计全新的声音形象。声音复刻基于真实音频样本复制音色,适合还原特定人物的声音。详情请参见[声音复刻](/developer-guides/speech/voice-cloning)。 + +## API 参考 + +- [Voice design API 参考](/api-reference/speech-synthesis/voice-design/http-api) -- API 参数和响应格式 +- [流式语音合成](/developer-guides/speech/realtime-streaming) -- 使用自定义音色进行实时合成 +- [Qwen TTS](/api-reference/speech-synthesis/qwen-tts) -- 使用自定义音色进行非流式合成 +- [获取 API Key](/api-reference/preparation/api-key) -- 设置身份认证 diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-cosyvoice.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-cosyvoice.md new file mode 100644 index 0000000..9c7e2ca --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-cosyvoice.md @@ -0,0 +1,261 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# CosyVoice 音色列表 + +> 系统预置音色列表 + +以下表格列出了 CosyVoice 的所有系统预置音色。 + +使用须知: + +- 每个 `model` 仅支持特定的音色集合,不能跨模型混用。 +- `text` 的内容必须使用该音色支持的语言,否则会出现发音错误或不自然的合成效果。 +- 支持 SSML 的音色,可在 `text` 参数中传入 SSML 内容。详见 [SSML 指南](/developer-guides/speech/ssml)。 +- 支持 Instruct 的音色,可在 `instruction` 参数中传入 Instruct 格式的文本。 +- 支持时间戳的音色,将 `word_timestamp_enabled` 设为 `true`(Java SDK 中为 `enableWordTimestamp`)。时间戳数据以字级别时间事件的形式通过 WebSocket 回调返回,不嵌入音频数据中。 + - Python SDK:通过 `additional_params` 传入 `word_timestamp_enabled`: + `SpeechSynthesizer(model=model, voice=voice, callback=callback, additional_params={"word_timestamp_enabled": True})` + + + 如果音色列表中明确标注了 Instruct 的格式要求,请严格按照指定格式填写 `instruction` 参数;如果未标注格式要求,则表示该音色支持任意格式的自然语言指令。详细用法请参见[语音合成 > 指令控制](/developer-guides/speech/tts#指令控制)。 + + +## cosyvoice-v3-flash + +| 场景 | 音色名称 | 音色参数 | 特征 | 年龄 | 语言 | SSML | Instruct | 时间戳 | 试听 | +| ---------- | ----------- | ----------------- | ------ | ------- | ------------------------------------------ | ---- | -------- | --- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | +| 社交陪伴(标杆音色) | 龙安洋 | `longanyang` | 阳光大男孩 | 20\~30岁 | 中文(普通话)、英文 | 支持 | 支持 | 支持 |
+ + 以下四个音色均支持下列全部方言和语言。 + + - 方言:上海话、广东话、东北话、重庆话、陕西话、云南话、宁波话、甘肃话。 + - 语言:日语、韩语、法语、德语、葡萄牙语、意大利语、越南语、印尼语。 + +| 名称 | voice 参数 | 性别 | 试听 | +| ---------- | -------------------- | -- | --------------------------------------- | +| 龙安欢\_v3.1 | `longanhuan_v3.1` | 女 | 重庆话
宁波话
韩语
印尼语 | +| 龙安灵心\_v3.1 | `longanlingxin_v3.1` | 女 | 云南话
陕西话
上海话
法语
意大利语 | +| 龙安风悦\_v3.1 | `longanfengyue_v3.1` | 女 | 东北话
越南语
日语 | +| 许南川 | `xunanchuan_v3.1` | 男 | 甘肃话
东北话
法语
葡萄牙语 | + + + +
+ + 以下音色仅支持中文普通话。 + +| 名称 | voice 参数 | 性别 | 声线特质 | 适用场景 | 试听 | +| --- | ------------------- | -- | -------- | --------------------- | -- | +| 于小云 | `yuxiaoyun_v3.1` | 女 | 元气、亲切、自然 | 广告营销、广播、客服助手、旁白 | | +| 乔小娇 | `qiaoxiaojiao_v3.1` | 女 | 俏丽、可爱 | 广告营销、客服助手、有声书 | | +| 夏小晨 | `xiaxiaochen_v3.1` | 女 | 元气、明亮 | 广告营销、有声书 | | +| 安明远 | `anmingyuan_v3.1` | 男 | 清亮、自然 | 广告营销、有声书、旁白 | | +| 温怀清 | `wenhuaiqing_v3.1` | 女 | 清亮、柔和 | 儿童故事、客服助手、广告营销、新闻播报 | | +| 安小岚 | `anxiaolan_v3.1` | 女 | 清甜、纯净 | 有声书、客服助手、旁白、新闻播报、广告营销 | | +| 谢舒柔 | `xieshurou_v3.1` | 女 | 柔和、自然、知性 | 有声书、客服助手、旁白 | | +| 白清岚 | `baiqinglan_v3.1` | 女 | 明亮、清纯 | 语音助手、客服助手 | | +| 许玉远 | `xuyuyuan_v3.1` | 女 | 知性、成熟、质感 | 广告营销、新闻播报、旁白、客服助手、有声书 | | +| 安若柔 | `anruorou_v3.1` | 女 | 气声、知性 | 旁白、语音助手 | | +| 闻怀之 | `wenhuaizhi_v3.1` | 女 | 稳重、成熟 | 有声书、新闻播报、广告营销、客服助手、旁白 | | +| 萧行之 | `xiaoxingzhi_v3.1` | 女 | 端庄、贵气 | 新闻播报、有声书、旁白、客服助手 | | +| 顾云舒 | `guyunshu_v3.1` | 女 | 成熟、稳重 | 音乐电台、客服助手、有声书、旁白 | | +| 霍拙石 | `huozhuoshi_v3.1` | 男 | 清亮 | 有声书、广告营销、旁白 | | +| 叶清禾 | `yeqinghe_v3.1` | 女 | 亲切、温柔 | 有声书、广告营销、旁白、客服助手 | | +| 云欢欢 | `yunhuanhuan_v3.1` | 女 | 高亢、热情 | 有声书、旁白、客服助手 | | +| 徐小俏 | `xuxiaoqiao_v3.1` | 女 | 自然、俏皮 | 有声书、旁白、客服助手 | | +| 白安然 | `baianran_v3.1` | 女 | 低沉、浑厚、气声 | 配音讲解、有声书、旁白 | | +| 许言初 | `xuyanchu_v3.1` | 女 | 沉稳、磁性 | 新闻播报、有声书 | | +| 叶知晴 | `yezhiqing_v3.1` | 女 | 轻快、自然 | 儿童故事、客服助手、语音助手 | | +| 安迪 | `andi_v3.1` | 男 | ABC口音 | 语音助手 | | +| 安语晴 | `anyuqing_v3.1` | 女 | 甜妹 | 语音助手、旁白、新闻播报 | | + + + + + + 以下音色仅支持英文。 + +| 名称 | voice 参数 | 性别 | 口音 | 试听 | +| ----- | ------------ | -- | ---- | -- | +| Emily | `Emily_v3.1` | 女 | 英式女声 | | +| Luna | `Luna_v3.1` | 女 | 英式女声 | | +| Eric | `Eric_v3.1` | 男 | 英式男声 | | +| Luca | `Luca_v3.1` | 男 | 英式男声 | | +| Abby | `Abby_v3.1` | 女 | 美式女声 | | +| Annie | `Annie_v3.1` | 女 | 美式女声 | | +| Ava | `Ava_v3.1` | 女 | 美式女声 | | +| Beth | `Beth_v3.1` | 女 | 美式女声 | | +| Betty | `Betty_v3.1` | 女 | 美式女声 | | +| Cally | `Cally_v3.1` | 女 | 美式女声 | | +| Cindy | `Cindy_v3.1` | 女 | 美式女声 | | +| Donna | `Donna_v3.1` | 女 | 美式女声 | | +| Andy | `Andy_v3.1` | 男 | 美式男声 | | +| Brian | `Brian_v3.1` | 男 | 美式男声 | | +| David | `David_v3.1` | 男 | 美式男声 | | + + + + + +| 名称 | voice 参数 | 性别 | 声线特质 | 适用场景 | 试听 | +| ----------------- | -------------------- | -- | ----- | ---------- | -- | +| 龙安元妃\_v3.1 | `longanyuanfei_v3.1` | 女 | 高傲妃子音 | 社交陪伴 | | +| 龙杰力豆\_v3.1 | `longjielidou_v3.1` | 男 | 天真男童音 | 儿童陪伴 | | +| 龙安灵希\_v3.1 | `longanlingxi_v3.1` | 女 | 可爱甜美音 | 社交陪伴(精品中文) | | +| 龙火火\_v3.1 | `longhuohuo_v3.1` | 男 | 顽皮少年音 | 角色音 | | +| 龙应桃\_v3.1 | `longyingtao_v3.1` | 女 | 温柔淡定女 | 客服 | | +| 龙安雅\_v3.1 | `longanya_v3.1` | 女 | 高雅气质女 | 社交陪伴 | | +| 龙婉\_v3.1 | `longwan_v3.1` | 女 | 细腻柔声女 | 社交陪伴 | | +| 龙星\_v3.1 | `longxing_v3.1` | 女 | 温婉邻家女 | 社交陪伴 | | +| 龙华\_v3.1 | `longhua_v3.1` | 女 | 元气甜美女 | 社交陪伴 | | +| 龙寒\_v3.1 | `longhan_v3.1` | 男 | 温暖痴情男 | 社交陪伴 | | +| 龙安智\_v3.1 | `longanzhi_v3.1` | 男 | 睿智轻熟男 | 社交陪伴 | | +| 龙哲\_v3.1 | `longzhe_v3.1` | 男 | 呆板大暖男 | 社交陪伴 | | +| 龙安洋\_v3.1 | `longanyang_v3.1` | 男 | 阳光大男孩 | 社交陪伴(标杆音色) | | +| 李白\_v3.1 | `libai_v3.1` | 男 | 古代诗仙男 | 诗词朗诵 | | +| 龙铃\_v3.1 | `longling_v3.1` | 女 | 稚气呆板女 | 童声 | | +| 龙牛牛\_v3.1 | `longniuniu_v3.1` | 男 | 阳光男童声 | 消费电子-儿童有声书 | | +| 龙闪闪\_v3.1 | `longshanshan_v3.1` | 男 | 戏剧化童声 | 消费电子-儿童有声书 | | +| 龙泡泡\_v3.1 | `longpaopao_v3.1` | 女 | 飞天泡泡音 | 消费电子-儿童陪伴 | | +| loongstella\_v3.1 | `loongstella_v3.1` | 女 | 飒爽利落女 | 新闻播报 | | +| 龙媛\_v3.1 | `longyuan_v3.1` | 女 | 温暖治愈女 | 有声书 | | +| 龙妙\_v3.1 | `longmiao_v3.1` | 女 | 抑扬顿挫女 | 有声书 | | +| 龙三叔\_v3.1 | `longsanshu_v3.1` | 男 | 沉稳质感男 | 有声书 | | +| 龙安莉\_v3.1 | `longanli_v3.1` | 女 | 利落从容女 | 语音助手 | | +| 龙安温\_v3.1 | `longanwen_v3.1` | 女 | 优雅知性女 | 语音助手 | | +| 龙安朗\_v3.1 | `longanlang_v3.1` | 男 | 清爽利落男 | 语音助手 | | +| 龙小夏\_v3.1 | `longxiaoxia_v3.1` | 女 | 沉稳权威女 | 语音助手 | | +| 龙安冲\_v3.1 | `longanchong_v3.1` | 男 | 激情推销男 | 直播带货 | | + + + +### qwen-audio-3.0-tts-plus + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+ **适用场景** + + **音色信息** + + **音频试听(右键保存音频)** +
+ 社交陪伴(旗舰音色) + + **名称**:龙安灵心 + + **voice参数**:longanlingxin + + **特质**:知心温暖音 + + **年龄**:25岁 + + **性别**:女 + + **语言**:中文(普通话)、英文 + + +
+ **名称**:龙安鲁风 + + **voice参数**:longanlufeng + + **特质**:明亮开朗音 + + **年龄**:25岁 + + **性别**:男 + + **语言**:中文(普通话)、英文 + + +
+ +### qwen-audio-3.0-tts-flash + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+ **适用场景** + + **音色信息** + + **音频试听(右键保存音频)** +
+ 社交陪伴(精品中文) + + **名称**:龙安风悦 + + **voice参数**:longanfengyue + + **特质**:自然亲切音 + + **年龄**:30岁 + + **性别**:女 + + **语言**:中文(普通话)、英文 + + +
+ **名称**:龙安元妃 + + **voice参数**:longanyuanfei + + **特质**:高傲妃子音 + + **年龄**:30岁 + + **性别**:女 + + **语言**:中文(普通话)、英文 + + +
+ **名称**:龙安灵希 + + **voice参数**:longanlingxi + + **特质**:可爱甜美音 + + **年龄**:25岁 + + **性别**:女 + + **语言**:中文(普通话)、英文 + + +
+ **名称**:龙安小昕 + + **voice参数**:longanxiaoxin + + **特质**:亲切活泼音 + + **年龄**:22岁 + + **性别**:女 + + **语言**:中文(普通话)、英文 + + +
+ **名称**:龙安欢 + + **voice参数**:longanhuan\_v3.6 + + **年龄**:25岁 + + **性别**:女 + + **语言**:中文(普通话)、英文 + + +
+ 儿童陪伴/智能玩具(精品儿童) + + **名称**:龙杰力豆 + + **voice参数**:longjielidou\_v3.6 + + **特质**:天真男童 + + **年龄**:5岁 + + **性别**:男 + + **语言**:中文(普通话)、英文 + + +
+ **名称**:龙泡泡 + + **voice参数**:longpaopao\_v3.6 + + **特质**:软糯可爱音 + + **年龄**:5岁 + + **性别**:女 + + **语言**:中文(普通话)、英文 + + +
+ 角色音/游戏(精品中文) + + **名称**:龙火火 + + **voice参数**:longhuohuo\_v3.6 + + **特质**:顽皮少年音 + + **年龄**:8岁 + + **性别**:男 + + **语言**:中文(普通话)、英文 + + +
+ **名称**:龙川叔 + + **voice参数**:longchuanshu\_v3.6 + + **特质**:川普大叔音 + + **年龄**:40岁 + + **性别**:男 + + **语言**:中文(普通话)、英文 + + +
+ 社交陪伴/语音助手(精品英文) + + **名称**:loongmary + + **voice参数**:loongmary + + **特质**:温暖英音 + + **年龄**:20岁 + + **性别**:女 + + **语言**:英文 + + +
+ **名称**:loongeva + + **voice参数**:loongeva\_v3.6 + + **特质**:高智美音 + + **年龄**:28岁 + + **性别**:女 + + **语言**:英文 + + +
+ **名称**:loongJohn + + **voice参数**:loongjohn + + **特质**:沉稳亲切美音 + + **年龄**:28岁 + + **性别**:男 + + **语言**:英文 + + +
+ +## 基础音色 + + + 建议优先使用系统音色,以获得更稳定的语音合成效果。如需定制音色,可通过[声音复刻](/developer-guides/speech/voice-cloning)或[声音设计](/developer-guides/speech/voice-design)创建专属音色。基础音色提供更多选择,使用前建议试听并评估其是否符合业务需求。 + + +除上述系统音色外,`qwen-audio-3.0-tts-plus`和`qwen-audio-3.0-tts-flash`各自还提供500余个通过声音复刻生成的基础音色,调用方式与系统音色一致。 + +
+ +基础音色命名格式为`qwen-audio-3.0-tts-{plus|flash}-{音色后缀}`,两个模型同一后缀对应同一套试听音频。 + + + +- `qwen-audio-3.0-tts-plus`基础音色列表(Excel):[qwen-audio-3.0-tts-plus基础音色.xlsx](https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260723/ydwqqz/qwen-audio-3.0-tts-plus%E5%9F%BA%E7%A1%80%E9%9F%B3%E8%89%B2.xlsx) +- `qwen-audio-3.0-tts-flash`基础音色列表(Excel):[qwen-audio-3.0-tts-flash基础音色.xlsx](https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260723/thosjr/qwen-audio-3.0-tts-flash%E5%9F%BA%E7%A1%80%E9%9F%B3%E8%89%B2.xlsx) +- 基础音色试听音频包(plus和flash共用):[基础音色试听音频包.zip](https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260720/tuuuqo/%E5%9F%BA%E7%A1%80%E9%9F%B3%E8%89%B2%E8%AF%95%E5%90%AC%E9%9F%B3%E9%A2%91%E5%8C%85.zip) + +**试听步骤** + + + +1. 下载Excel和试听音频包,将音频包解压到本地。 +2. 在Excel中找到“预览音频文件名”列,获取音频文件名。 +3. 在解压目录中找到对应文件,使用播放器打开试听。 diff --git a/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-qwen-tts.md b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-qwen-tts.md new file mode 100644 index 0000000..bf24e31 --- /dev/null +++ b/mac-agent-os-main/05_tools/09_ave/docs/api/raw/qwen/developer-guides-speech-voice-list-qwen-tts.md @@ -0,0 +1,113 @@ +> ## Documentation Index +> Fetch the complete documentation index at: https://platform.qianwenai.com/docs/llms.txt +> Use this file to discover all available pages before exploring further. + +# Qwen-TTS 音色列表 + +> Qwen-TTS 实时与非实时语音合成支持的音色 + +## Qwen-TTS实时语音合成音色列表 + +| `voice` 参数 | 详细信息 | 支持语种 | 支持的模型 | +| ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `Cherry` | **描述**:阳光积极、亲切自然小姐姐(女性)