- A mandatory separate `互动提问` section after the title candidates and before the platform rewrites.
- Every `互动提问` section must contain exactly 10 questions.
- Ask only knowledge- and content-related questions about the current subject.
- Focus on structure, component function, architecture, process, trade-offs, bottlenecks, applications, limitations, reliability, cost, efficiency, or technical comparisons.
- Phrase every item as a genuine question ending in `?` or `?`, for example asking what function a component performs, where a bottleneck occurs, why a material was selected, or what trade-off an architecture makes.
- Every question must explicitly include the topic keyword (the subject of the poster, e.g. NVLink, B200, CoWoS) in the question text. Do not use generic references like "this architecture", "this component", or "this design" as a substitute for the keyword.
- Do not ask what topic, product, teardown, article, image, or subject the audience wants next.
- Do not ask subjective-preference or opinion questions. Reject phrasing such as "你觉得", "你认为", "哪个更", "会更", "会不会", "值不值得", "有没有必要", or any question that asks the reader to pick a favorite, predict adoption, or express personal taste. Every question must have a verifiable technical answer grounded in the subject's structure, process, physics, or engineering constraints.
- Do not use generic engagement bait without a concrete technical object or trade-off.
- Reject prompts such as asking what to cover next, what else the reader wants to see, or requesting future-topic suggestions in comments.
- Do not write analytical evaluation paragraphs, generic praise, rhetorical promotion, predetermined conclusions, or unrelated engagement bait.
-
exec
/bin/zsh -lc "sed -n '520,900p' /Users/alex/.codex/skills/product-exploded-view-company-prompts/SKILL.md 2>&1 | head -c 10000" in /Users/alex/.openclaw/workspace/exploded-poster-gen
succeeded in 0ms:
web search:
web search: site:help.openai.com ChatGPT Voice 9 voices preset voices
web search:
web search: https://help.openai.com/en/articles/8400625-voice-mode-faq
codex
使用 `product-exploded-view-company-prompts` skill 生成。参考了 ChatGPT Voice 官方帮助中心对 Live/Advanced/Standard、9 种 voice、可打断、多模态与语音设置的说明,以及模型发布说明中对 Advanced Voice 实时语音能力的描述:
https://help.openai.com/en/articles/20001274
https://help.openai.com/en/articles/9624314-model-release-notes
**Imagegen Prompt**
text
画布比例:竖版 3:4,中文技术拆解海报,主题为「ChatGPT Voice 实时语音系统拆解」,整体为概念爆炸图 / schematic exploded view,不使用任何官方 logo、官方 glyph、品牌标识、图标化商标或仿官方视觉符号
digraph skill_flow {
"User message received" [shape=doublecircle];
"About to EnterPlanMode?" [shape=doublecircle];
"Already brainstormed?" [shape=diamond];
"Invoke brainstorming skill" [shape=box];
"Might any skill apply?" [shape=diamond];
"Invoke Skill tool" [shape=box];
"Announce: 'Using [skill] to [purpose]'" [shape=box];
"Has checklist?" [shape=diamond];
"Create TodoWrite todo per item" [shape=box];
"Follow skill exactly" [shape=box];
"Respond (including clarifications)" [shape=doublecircle];
dot
digraph skill_flow {
"User message received" [shape=doublecircle];
"About to EnterPlanMode?" [shape=doublecircle];
"Already brainstormed?" [shape=diamond];
"Invoke brainstorming skill" [shape=box];
"Might any skill apply?" [shape=diamond];
"Invoke Skill tool" [shape=box];
"Announce: 'Using [skill] to [purpose]'" [shape=box];
"Has checklist?" [shape=diamond];
"Create TodoWrite todo per item" [shape=box];
"Follow skill exactly" [shape=box];
"Respond (including clarifications)" [shape=doublecircle];
"About to EnterPlanMode?" -> "Already brainstormed?";
"Already brainstormed?" -> "Invoke brainstorming skill" [label="no"];
"Already brainstormed?" -> "Might any skill apply?" [label="yes"];
"Invoke brainstorming skill" -> "Might any skill apply?";
"User message received" -> "Might any skill apply?";
"Might any skill apply?" -> "Invoke Skill tool" [label="yes, even 1%"];
"Might any skill apply?" -> "Respond (including clarifications)" [label="definitely not"];
"Invoke Skill tool" -> "Announce: 'Using [skill] to [purpose]'";
"Announce: 'Using [skill] to [purpose]'" -> "Has checklist?";
"Has checklist?" -> "Create TodoWrite todo per item" [label="yes"];
"Has checklist?" -> "Follow skill exactly" [label="no"];
"Create TodoWrite todo per item" -> "Follow skill exactly";
}
Red Flags
These thoughts mean STOP—you're rationalizing:
| Thought | Reality |
|---------|---------|
| "This is just a simple question" | Questions are tasks. Check for skills. |
| "I need more context first" | Skill check comes BEFORE clarifying questions. |
| "Let me explore the codeba
exec
/bin/zsh -lc "sed -n '261,520p' /Users/alex/.codex/skills/product-exploded-view-company-prompts/SKILL.md 2>&1 | head -c 8000" in /Users/alex/.openclaw/workspace/exploded-poster-gen
succeeded in 0ms:
"[heading]\n[short explanatory sentence]"
Right side:
"[heading]\n[short explanatory sentence]"
Bottom:
- Headline: "[bottom headline]"
- Body: "[2-3 concise sentences]"
- Footer mark: "[category / use case]"
Avoid:
garbled text, wrong language, Japanese text in Chinese posters, obvious brand logos, official logo glyphs, logo-like marks, company icons, company introductions, company products, company locations, company links, stock codes, rankings, recommendation reasons, geographic or regional names, origin-group labels, hard-ad title, sales CTA, exaggerated superlatives, hype words such as 重磅/颠覆/神器/必看/最强/炸裂/王炸, missing author credit, missing source disclaimer, metadata-only credit or disclaimer, watermark treatment, decorative clutter, empty-looking layout, unreadable tiny text, cropped labels, overlapping microtext, overlapping callouts, isolated floating parts without whole-to-parts guidance, unsupported exact claims about proprietary internals.
Component Rules
- Label every layer with a visible number and a short component name.
- Prefer 7-10 layers. Fewer can feel shallow; more often becomes unreadable.
- Group tiny internals into meaningful systems: compute, memory, power, cooling, frame, connector, enclosure, sensor, optics, battery, interface.
- Use callouts to explain function, not repeat component names.
- Keep primary image-model text short. Prefer one heading and one short sentence per main callout.
- Use "示意", "结构概念", "conceptual", or "schematic" for uncertain internals.
- For architecture or chips, describe whole package, stacked layers, key interconnect, base/control layer, package/interposer, and what the structure enables.
- For consumer products, describe shell/enclosure, display/interface, compute/control board, battery/power, sensors, connectors, thermal path, frame, and fastening when relevant.
In-Image Density Rules
- Use richer in-image text as an optional enhancement for technical posters when it helps the image feel complete and educational.
- Add 4-8 secondary details such as micro labels, small spec chips, short system-flow captions, mini legends, material notes, or compact annotation panels.
- Use smaller typography for secondary details, but do not make any required label, number, author credit, source disclaimer, or main callout tiny.
- Richer text should support the visual reading path, not become decorative filler.
- Avoid dense paragraphs inside the image; prefer short fragments, labels, and one-line notes.
Title Rules
- Poster titles should feel like technical editorial hooks, not product ads.
- Prefer titles that reveal a question, contrast, or structural role, such as "GPU 背后的调度底座", "为什么它不是显卡", or "AI 工厂里被低估的 CPU".
- Avoid hard-marketing title patterns such as "重磅发布", "颠覆未来", "最强芯片", "必看神器", "王炸新品", and sales-style calls to action.
- Keep the title shorter than the subtitle; let the subtitle carry the precise product name or use case when the title is more editorial.
- For Chinese publishing copy, generate 30 distinct title candidates by default and mark one recommended title.
Author Credit Rules
- Every prompt must include an in-image author/account credit.
- For Chinese content, use `作者:{name}`.
- Default to `作者:好用工具推荐` when no conflicting account appears.
- Place the credit inside the lower-middle content area, beside the lower decomposition axis, next to a real callout group, or inside a small engineering annotation panel connected to the diagram.
- Make the credit feel naturally merged with the real content so it is not easy to remove by cropping the bottom or deleting a detached footer.
- Avoid the very bottom edge, detached footer strips, corner-only labels, and large watermark treatment. Also avoid placing it too high or over the main reading path.
- Keep the credit smaller than component labels but clearly readable.
- Do not let it cover the product, numbered layers, arrows, or callouts.
- Do not treat the credit as a watermark outside the design; integrate it as a technical poster annotation.
- Do not use placeholders such as `【作者名】`, `{author}`, `[account]`, or `{name}` in the final prompt.
Source Disclaimer Rules
- Every Chinese technical poster prompt and companion publishing copy must include a source disclaimer.
- Default wording: `声明:信息整理自网络公开内容,仅供学习交流;如有疏漏或错误,请联系指正。`
- If the user supplies different wording, preserve it unless it creates a factual or legal misrepresentation.
- Require the disclaimer as visible image text near the author credit or in a connected lower-middle technical annotation panel. Keep it smaller than primary labels but clearly readable.
- Do not place it at the bottom edge, in a detachable footer, only in metadata, or outside the image.
- Append the full disclaimer directly at the end of every publish-ready platform version, including 作者, 小红书, 朋友圈/社群, and 知乎/B站动态.
- Do not consolidate, deduplicate, or replace these per-platform endings with a shared disclaimer, reusable ending, "通用结尾", footnote, or instruction to add it later.
- Do not claim that the generated image contains or correctly renders the disclaimer because this skill does not generate or inspect images.
- The disclaimer does not replace source verification. Continue to use official or primary sources when accuracy matters and label uncertain internals as conceptual or schematic.
Chinese Social Copy
Include social-media publishing copy by default for Chinese poster prompts unless the user asks for prompt-only output or no social copy. Use:
- 30 distinct title candidates by default, with one clearly marked as the recommended title; each title must be 15–20 visible characters (count Chinese characters, letters, digits, punctuation — every visible glyph counts). Titles shorter than 15 or longer than 20 visible characters are invalid and must be rewritten.
- A concise body that starts from the visual reading path, for example "通过这张图可以看到...".
- A mandatory separate `互动提问` section after the title candidates and before the platform rewrites.
- Every `互动提问` section must contain exactly 10 questions.
- Ask only knowledge- and content-related questions about the current subject.
- Focus on structure, component function, architecture, process, trade-offs, bottlenecks, applications, limitations, reliability, cost, efficiency, or technical comparisons.
- Phrase every item as a genuine question ending in `?` or `?`, for example asking what function a component performs, where a bottleneck occurs, why a material was selected, or what trade-off an architecture makes.
- Every question must explicitly include the topic keyword (the subject of the poster, e.g. NVLink, B200, CoWoS) in the question text. Do not use generic references like "this architecture", "this component", or "this design" as a substitute for the keyword.
- Do not ask what topic, product, teardown, article, image, or subject the audience wants next.
- Do not ask subjective-preference or opinion questions. Reject phrasing such as "你觉得", "你认为", "哪个更", "会更", "会不会", "值不值得", "有没有必要", or any question that asks the reader to pick a favorite, predict adoption, or express personal taste. Every question must have a verifiable technical answer grounded in the subject's structure, process, physics, or engineering constraints.
- Do not use generic engagement bait without a concrete technical object or trade-off.
- Reject prompts such as asking what to cover next, what else the reader wants to see, or requesting future-topic suggestions in comments.
- Do not write analytical evaluation paragraphs, generic praise, rhetorical promotion, predetermined conclusions, or unrelated engagement bait.
-
exec
/bin/zsh -lc "sed -n '520,900p' /Users/alex/.codex/skills/product-exploded-view-company-prompts/SKILL.md 2>&1 | head -c 10000" in /Users/alex/.openclaw/workspace/exploded-poster-gen
succeeded in 0ms:
web search:
web search: site:help.openai.com ChatGPT Voice 9 voices preset voices
web search:
web search: https://help.openai.com/en/articles/8400625-voice-mode-faq
codex
使用 `product-exploded-view-company-prompts` skill 生成。参考了 ChatGPT Voice 官方帮助中心对 Live/Advanced/Standard、9 种 voice、可打断、多模态与语音设置的说明,以及模型发布说明中对 Advanced Voice 实时语音能力的描述:
https://help.openai.com/en/articles/20001274
https://help.openai.com/en/articles/9624314-model-release-notes
Imagegen Prompt
text
画布比例:竖版 3:4,中文技术拆解海报,主题为「ChatGPT Voice 实时语音系统拆解」,整体为概念爆炸图 / schematic exploded view,不使用任何官方 logo、官方 glyph、品牌标识、图标化商标或仿官方视觉符号。
视觉风格:深色工程蓝黑背景 + 青绿色实时音频波形 + 白色细线分层标注 + 局部暖色延迟脉冲,高级技术编辑海报质感。中心是一条从左下到右上的「端到端低延迟语音链路」主轴,主轴上方展示完整系统,周围按编号拆解为 9 个功能层。所有文字使用简体中文,现代无衬线字体,短标签、清晰层级、无日文、无乱码、无伪文字。
标题区:
「一句话背后的实时语音系统」
副标题:
「ChatGPT Voice 技术功能拆解:从麦克风输入到神经语音输出」
主体布局:
中心主轴显示完整链路:声音输入 → 识别 → 理解 → 推理 → 状态控制 → 语音合成 → 声音输出。
主轴旁添加细小流动粒子、音频帧、文本 token、声学特征、低延迟时间片,体现实时管道。
编号功能层:
01 麦克风采集与 VAD
短注释:采集声波、抑制噪声、判断说话起止点,减少无效音频进入后续链路。
02 ASR 流式识别
短注释:边听边转写,把连续语音切成可理解的文本片段,支持实时增量更新。
03 语境与意图理解
短注释:结合历史对话、当前语义和用户目标,判断要回答什么、何时回答。
04 LLM 实时推理
短注释:在语音流持续输入时生成响应策略,兼顾速度、准确性和对话自然度。
05 对话状态管理
短注释:维护轮次、上下文、任务状态和用户偏好,避免每句话都从零开始。
06 可打断响应机制
短注释:支持用户插话、继续说、修正问题,系统根据最新语音重新对齐响应。
07 TTS 神经语音合成
短注释:把文本响应转成自然语音,控制音高、节奏、停顿和语气变化。
08 预设音色与人设层
短注释:展示 9 种预设 voice:Arbor、Breeze、Cove、Ember、Juniper、Maple、Sol、Spruce、Vale。只用中性文字列表,不做头像、标志或品牌化符号。
09 多语种与情感表达
短注释:识别和生成多语言语音,并通过语速、重音、停顿和语调表达情绪线索。
补充标注:
- 「低延迟关键点:采集、流式识别、增量推理、边合成边播放」
- 「实时对话难点:噪声、重叠说话、长停顿、网络波动、上下文窗口」
- 「输出不是逐字记录:语音转写和最终文本可能存在差异」
- 「Live / Advanced / Standard 可作为不同体验层级的抽象标注,不加入官方图形元素」
底部中间技术声明面板:
「声明:信息整理自网络公开内容,仅供学习交流;如有疏漏或错误,请联系指正。」
画面要求:
信息密度适中,9 个编号层清晰可读,中心主轴有完整到拆解的阅读路径;标注线不交叉,文字不重叠,不出现作者署名、账号、公众号、品牌水印、二维码、logo、官方 glyph、营销口号、夸张词。
推荐标题
1. 一句话背后的语音系统
标题候选
1. 一句话背后的语音系统
2. ChatGPTVoice如何听懂
3. 实时语音对话链路拆解
4. 从麦克风到语音回复
5. 低延迟语音管道全景图
6. Voice背后的九层能力
7. 可打断对话怎样实现
8. 边听边想边说的系统
9. ASR到TTS的实时链路
10. 实时语音交互功能拆解
11. 语音输入如何变成回答
12. 多语种语音系统分层图
13. 情绪语气如何进入语音
14. Voice实时响应机制拆解
15. 端到端语音AI链路图
16. 麦克风之后发生了什么
17. 流式识别到神经合成图
18. 语音助手的实时管道图
19. 听说理解合成四步链路
20. 插话不中断的技术逻辑
21. ChatGPTVoice系统剖面
22. 九层看懂实时语音对话
23. 语音模型如何保持上下文
24. Voice低延迟从何而来
25. 从声音帧到自然回复图
26. 语音交互的状态控制层
27. 实时语音背后的工程栈
28. 一张图看语音AI管线
29. Voice功能层级拆解图
30. ChatGPTVoice拆解笔记
互动提问
1. ChatGPT Voice 的 VAD 在实时链路中主要解决什么端点判断问题?
2. ChatGPT Voice 的流式 ASR 为什么比整段转写更适合实时对话?
3. ChatGPT Voice 如何在用户继续说话时更新语境和意图?
4. ChatGPT Voice 的 LLM 实时推理需要平衡哪些延迟与准确性约束?
5. ChatGPT Voice 的对话状态管理保存了哪些关键上下文信息?
6. ChatGPT Voice 的可打断机制如何处理用户插话和重叠语音?
7. ChatGPT Voice 的 TTS 合成层如何影响语速、停顿和语气?
8. ChatGPT Voice 的 9 种预设 voice 在系统中属于哪一层抽象?
9. ChatGPT Voice 的多语种能力会给 ASR 和 TTS 带来哪些工程挑战?
10. ChatGPT Voice 的端到端低延迟瓶颈通常会出现在哪些环节?
发布正文
通过这张图可以看到,ChatGPT Voice 不是一个单独的“语音按钮”,而是一条端到端实时语音管道。
从麦克风采集、VAD 端点检测、流式 ASR,到语境理解、LLM 实时推理、对话状态管理,再到可打断响应和 TTS 神经语音合成,每一层都会影响对话是否自然、是否跟手、是否能及时处理插话。
图里把它拆成 9 个功能层:输入、识别、理解、推理、状态、打断、合成、音色、多语种与情感表达。适合用来理解实时语音 AI 的系统结构。
声明:信息整理自网络公开内容,仅供学习交流;如有疏漏或错误,请联系指正。
tokens used
36,587
使用 `product-exploded-view-company-prompts` skill 生成。参考了 ChatGPT Voice 官方帮助中心对 Live/Advanced/Standard、9 种 voice、可打断、多模态与语音设置的说明,以及模型发布说明中对 Advanced Voice 实时语音能力的描述:
https://help.openai.com/en/articles/20001274
https://help.openai.com/en/articles/9624314-model-release-notes
Imagegen Prompt
text
画布比例:竖版 3:4,中文技术拆解海报,主题为「ChatGPT Voice 实时语音系统拆解」,整体为概念爆炸图 / schematic exploded view,不使用任何官方 logo、官方 glyph、品牌标识、图标化商标或仿官方视觉符号。
视觉风格:深色工程蓝黑背景 + 青绿色实时音频波形 + 白色细线分层标注 + 局部暖色延迟脉冲,高级技术编辑海报质感。中心是一条从左下到右上的「端到端低延迟语音链路」主轴,主轴上方展示完整系统,周围按编号拆解为 9 个功能层。所有文字使用简体中文,现代无衬线字体,短标签、清晰层级、无日文、无乱码、无伪文字。
标题区:
「一句话背后的实时语音系统」
副标题:
「ChatGPT Voice 技术功能拆解:从麦克风输入到神经语音输出」
主体布局:
中心主轴显示完整链路:声音输入 → 识别 → 理解 → 推理 → 状态控制 → 语音合成 → 声音输出。
主轴旁添加细小流动粒子、音频帧、文本 token、声学特征、低延迟时间片,体现实时管道。
编号功能层:
01 麦克风采集与 VAD
短注释:采集声波、抑制噪声、判断说话起止点,减少无效音频进入后续链路。
02 ASR 流式识别
短注释:边听边转写,把连续语音切成可理解的文本片段,支持实时增量更新。
03 语境与意图理解
短注释:结合历史对话、当前语义和用户目标,判断要回答什么、何时回答。
04 LLM 实时推理
短注释:在语音流持续输入时生成响应策略,兼顾速度、准确性和对话自然度。
05 对话状态管理
短注释:维护轮次、上下文、任务状态和用户偏好,避免每句话都从零开始。
06 可打断响应机制
短注释:支持用户插话、继续说、修正问题,系统根据最新语音重新对齐响应。
07 TTS 神经语音合成
短注释:把文本响应转成自然语音,控制音高、节奏、停顿和语气变化。
08 预设音色与人设层
短注释:展示 9 种预设 voice:Arbor、Breeze、Cove、Ember、Juniper、Maple、Sol、Spruce、Vale。只用中性文字列表,不做头像、标志或品牌化符号。
09 多语种与情感表达
短注释:识别和生成多语言语音,并通过语速、重音、停顿和语调表达情绪线索。
补充标注:
- 「低延迟关键点:采集、流式识别、增量推理、边合成边播放」
- 「实时对话难点:噪声、重叠说话、长停顿、网络波动、上下文窗口」
- 「输出不是逐字记录:语音转写和最终文本可能存在差异」
- 「Live / Advanced / Standard 可作为不同体验层级的抽象标注,不加入官方图形元素」
底部中间技术声明面板:
「声明:信息整理自网络公开内容,仅供学习交流;如有疏漏或错误,请联系指正。」
画面要求:
信息密度适中,9 个编号层清晰可读,中心主轴有完整到拆解的阅读路径;标注线不交叉,文字不重叠,不出现作者署名、账号、公众号、品牌水印、二维码、logo、官方 glyph、营销口号、夸张词。
推荐标题
1. 一句话背后的语音系统
标题候选
1. 一句话背后的语音系统
2. ChatGPTVoice如何听懂
3. 实时语音对话链路拆解
4. 从麦克风到语音回复
5. 低延迟语音管道全景图
6. Voice背后的九层能力
7. 可打断对话怎样实现
8. 边听边想边说的系统
9. ASR到TTS的实时链路
10. 实时语音交互功能拆解
11. 语音输入如何变成回答
12. 多语种语音系统分层图
13. 情绪语气如何进入语音
14. Voice实时响应机制拆解
15. 端到端语音AI链路图
16. 麦克风之后发生了什么
17. 流式识别到神经合成图
18. 语音助手的实时管道图
19. 听说理解合成四步链路
20. 插话不中断的技术逻辑
21. ChatGPTVoice系统剖面
22. 九层看懂实时语音对话
23. 语音模型如何保持上下文
24. Voice低延迟从何而来
25. 从声音帧到自然回复图
26. 语音交互的状态控制层
27. 实时语音背后的工程栈
28. 一张图看语音AI管线
29. Voice功能层级拆解图
30. ChatGPTVoice拆解笔记
互动提问
1. ChatGPT Voice 的 VAD 在实时链路中主要解决什么端点判断问题?
2. ChatGPT Voice 的流式 ASR 为什么比整段转写更适合实时对话?
3. ChatGPT Voice 如何在用户继续说话时更新语境和意图?
4. ChatGPT Voice 的 LLM 实时推理需要平衡哪些延迟与准确性约束?
5. ChatGPT Voice 的对话状态管理保存了哪些关键上下文信息?
6. ChatGPT Voice 的可打断机制如何处理用户插话和重叠语音?
7. ChatGPT Voice 的 TTS 合成层如何影响语速、停顿和语气?
8. ChatGPT Voice 的 9 种预设 voice 在系统中属于哪一层抽象?
9. ChatGPT Voice 的多语种能力会给 ASR 和 TTS 带来哪些工程挑战?
10. ChatGPT Voice 的端到端低延迟瓶颈通常会出现在哪些环节?
发布正文
通过这张图可以看到,ChatGPT Voice 不是一个单独的“语音按钮”,而是一条端到端实时语音管道。
从麦克风采集、VAD 端点检测、流式 ASR,到语境理解、LLM 实时推理、对话状态管理,再到可打断响应和 TTS 神经语音合成,每一层都会影响对话是否自然、是否跟手、是否能及时处理插话。
图里把它拆成 9 个功能层:输入、识别、理解、推理、状态、打断、合成、音色、多语种与情感表达。适合用来理解实时语音 AI 的系统结构。
声明:信息整理自网络公开内容,仅供学习交流;如有疏漏或错误,请联系指正。