Sponsor Suno AI Music arrow_forward
Skill

Video Assemble

合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、tts_meta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

Type
Skill
GitHub stars
536
License
MIT
Repo last updated
Sep 27, 2026

What Video Assemble is

Video Assemble is a skill published in the zenstory-ai/video-recap-skills repository on GitHub, which has about 536 stars. The repository describes itself as: “Claude Code / Codex skills that turn a video into a Chinese narration recap (视频解说): scene detection, ASR, VLM, script, TTS, ffmpeg assembly, optional editable JianYing / CapCut draft export (剪映草稿导出). Local ffmpeg + one MiMo key, no GPU. | 用 Claude Code skills 把视频做成中文解说成片,可选一键导出可编辑剪映草稿。”

A skill is a folder with a SKILL.md file: frontmatter with a name and a description, followed by instructions Claude follows. Claude loads a skill automatically when a task matches its description, and you can also run it directly with a slash and its name.

Skills work in Claude Code and in Claude Cowork, which makes Video Assemble a portable way to give Claude the same method everywhere.

How to install Video Assemble

Claude Code

  1. Download the video-assemble folder from the repository.
  2. Save it as ~/.claude/skills/<skill-name>/SKILL.md for all projects, or .claude/skills/<skill-name>/SKILL.md for one project.
  3. Claude loads it automatically when a task matches; you can also run it with / and its name.

Claude Cowork

  1. Zip the skill folder so SKILL.md sits at the top level of the folder.
  2. Open Customize → Skills, click +, then upload the ZIP.
  3. Start a task that matches the description, or call it by name with /.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from skills/video-assemble/SKILL.md, shared under the repository's MIT license. Read the full file on GitHub.

1. 定位

本技能负责最终合成:

  1. 把各段旁白音频放到视频时间线上。
  2. 在旁白窗口内压低原声,支持 fixed / sidechain / zone 模式。
  3. 根据旁白位置生成 subtitles.srt;默认同时生成并烧录 subtitles.ass,--no-burn-subtitles 可关闭。
  4. 可选把最终响度标准化到目标 LUFS。

2. 声音收尾契约

合成阶段只实现创作决定,不凭空制造决定。Agent 在写旁白位置前,已在 visual_audio_board.json 为每个 beat 指定 audio_owner:

  • original_dialogue
  • action_sound
  • ambience / music
  • silence
  • narration

因此,旁白间隙是主动选择,不是必须填满的空白。不要为了“更满”而加入通用 BGM、压住必须听见的台词或消除有意义的沉默。

当前渲染器不解析 visual_audio_board.json;Agent 通过旁白时间、overlaps_speech、原声留白与现有混音参数落实这些决定。

3. 输入契约

  • :源视频;cut 模式下为 edited_source.mp4。
  • work_dir/tts_meta.json:默认 narration 模式必需;配音阶段写出的 {segments: [...]}。每段包含 audio_path、时间、pause_after_ms、overlaps_speech 和用于混音/字幕的位置。显式 source-mix / adopted-packet-copy 模式不读取它。
  • 已采用的配音使用显式 --tts-meta 和 --narration-adoption:后者由调用方独立确认文字、请求的引擎/声线和速度策略,不能从待消费元数据自动“批准”出来。完整格式与记录边界见 references/narration-adoption.md。
  • 已采用的完整声音底轨与逐段配音可再传 --audio-mix-adoption;严格格式、48 kHz 声道矩阵和双 binding 事务见 references/explicit-audio-mix.md。

下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。

4. 运行命令

python3 scripts/assemble.py <video> --work-dir <work_dir> \
  [--audio-mode narration|source-mix|adopted-packet-copy] [--audio-stream-index <N>] \
  [--tts-meta <tts_meta.json> --narration-adoption <narration_adoption.json>] \
  [--audio-mix-adoption <audio_mix_adoption.json>] \
  [--recap-stem <name>] [--output-dir <dir>] [--no-burn-subtitles] \
  [--subtitle-y-top <inclusive-y> --subtitle-y-bot <exclusive-y>] \
  [--source-video <orig.mp4>] [--export-jianying [--jianying-out <dir>]]

5. 输出契约

  • recap_ .mp4:稳定的最终输出别名;每次运行覆盖更新。
  • work_dir/output.mp4:工作目录内成片。
  • subtitles.srt:旁白字幕;烧录时另有 subtitles.ass。
  • timeline.json:后端无关的多轨模型,包含视频、原声、旁白、BGM、字幕和 ducking 自动化。
  • _placed_*.wav:实际写入主混音的完整逐段旁白 PCM;时间线与剪映只引用这些文件。
  • narration_input_binding.json:旁白输入、转换、实际放置、旁白总轨和最终音轨的消费记录(路径、PCM 参数、packet 计数)。区分未采用与已绑定采用决定两种状态;不等于声线鉴定或听审。
  • audio_mix_binding.json:显式完整声音分支消费的画面时钟、底轨、48 kHz 配音放置、premaster、固定 master gain、最终 PCM/AAC 事实与 narration binding 路径的记录。
  • assembly_manifest.json:输入来源、cut 来源标识(路径、大小、mtime)、渲染设置与最终输出路径。
  • assembly_qc.json:旁白完整性、原声句末交接、时间线素材时长与交付质量的发布门禁。
  • 剪映草稿目录:仅 --export-jianying 时生成,包含 draft_content.json、draft_info.json 与 draft_meta_info.json。

6. 合成规则

  • 音频模式的处理与冻结语义见 references/audio-modes.md。默认仍为 narration;另外两种模式必须显式选择。
  • --audio-mix-adoption 只与显式 --tts-meta、--narration-adoption 同时使用;它保留 narration 模式名,但跳过旧速度/适配、原声 handoff、环境 BGM、duck、loudnorm 和 limiter。
  • 音频按轨道混合:原声、可选 BGM 与旁白各自独立。
  • 旁白不做任何容差裁尾;温和加速后仍放不下即 no_safe_fit。每段 _placed_*.wav 必须与序列化后的时间线区间等长或更短,否则 timeline_audio_mismatch 阻断。
  • 已采用配音的 v1 合同只支持原速、禁止段内适速;不能让环境默认 1.15 倍速或旧缓存覆盖它。放不下就修订安排,不裁尾。严格运行使用新工作目录与新输出路径;输入/实际混音来源变动或 QC 失败时,不发布候选成片。没有采用文件的旧入口仍是兼容模式,不自动获得同等证据。
  • 原声在旁白结束后保持压低到下一可靠句末的 pause_start,只在实测停顿内渐强, 于 source_restore_at 回满;无后续锚点时保持压低到时间线末端,而不是放出半句。
  • --export-jianying / EXPORT_JIANYING=1 可把 timeline.json 导出为可编辑草稿。cut 模式应传 --source-video ,让草稿引用真实原片区间。
  • 剪映导出默认把视频、音频与图片复制到 Resources/local/{video,audio,image},保持草稿可搬迁;--jianying-no-bundle-media 只适合原路径始终可访问的情况。
  • 重叠覆盖物会拆到编号轨道;非空目标目录不会覆盖,而会创建编号兄弟目录。
  • 常速、倒放、变换、富文本、转场、蒙版、LUT、绿幕复合草稿及显式特效轨道通过 timeline v2 扩展表达。需要素材包的功能只接受调用方合法提供的离线资源。
  • 剪映草稿引用未烧录的源视频,因此原片硬字幕仍会保留,必要时在剪映内另行遮罩。
  • 字幕外观可用 SUBTITLE_FONT_SIZE、SUBTITLE_MARGIN_V、SUBTITLE_MAX_CHARS 等控制。

按原片区间准备声音,而不是整体压低旧成片

已有多段原声取舍和独立 BGM 决定时,先用 references/source-score.md 的独立 source_score.py 从原片声音流按精确帧区间重建原声轨、音乐轨及两者之和;它只输出 声音底轨和来源回执。要与逐段已采用配音合成,再由调用方提供 references/explicit-audio-mix.md 的严格 adoption;不要将底轨塞入旧入口自动 duck,也不要从含旧解说的成片取整条声音冒充干净原声。

7. 字幕与可选包装

先锁定画面、剪点、旁白和混音,再投入字幕动画或边框包装;字幕样式不能掩盖叙事、剪点或声音问题。 普通交付优先使用现有 ASS 路径;只有用户需要逐 cue 排版、动画或透明图层时,才用项目级代码渲染器, 并按 references/foreground-compose.md 把它生成的 RGBA 序列叠到锁定母版。包装顺序与样帧抽检清单见 references/packaging.md。

8. 能力边界

  • 不生成旁白文字,不合成 TTS,不重新转写视频。
  • 字幕烧录默认开启;关闭时不会重编码绘制字幕区域。

显式输出轴字幕轨的独立合同、完整替换语义和当前边界见 references/subtitle-track.md。

画面回原片重建后,若需保留另一文件中的已采用完整混音,先按 references/pair-media.md 显式配对独立画面与音轨。配对只复制流,不补字幕或片名卡; 后续字幕轨必须重新绑定配对后的容器与 a:0,不能继续沿用旧版本身份。

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Video Assemble?

Video Assemble is a skill for Claude Code and Claude Cowork from the zenstory-ai/video-recap-skills repository on GitHub. 合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、tts_meta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

How do I install Video Assemble in Claude Code?

Download the video-assemble folder from the repository. Save it as ~/.claude/skills/<skill-name>/SKILL.md for all projects, or .claude/skills/<skill-name>/SKILL.md for one project. Claude loads it automatically when a task matches; you can also run it with / and its name.

Can I use Video Assemble in Claude Cowork?

Zip the skill folder so SKILL.md sits at the top level of the folder. Open Customize → Skills, click +, then upload the ZIP. Start a task that matches the description, or call it by name with /.

Is Video Assemble safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.