Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

model 2512.13507 — Cross-paper Synthesis

Seedance 1.5 Pro — L3 Cross-Paper Synthesis #

§1 相关论文 #

相关实体关系类型关联理由
Seedance 2.0 (2604.14148)直接后继ByteDance Seed Vision Team 的下一代模型,继承并扩展 1.5 pro 的 dual-branch MMDiT + RLHF 路线,在 Arena.AI T2V/I2V 双榜 Elo #1
LongCat-Video (2510.22200)技术对比美团的视频生成系统,同赛道竞争者;采用 GRPO + flow matching 的不同 RLHF 策略和 sparse attention 加速路线
SLA (2509.24006)底层加速Sparse-linear attention 是 Seedance 系列潜在采用的 attention 加速技术,SLA/SLA2 在 Wan2.1(与 Seedance 同赛道的 DiT 模型)上验证

§2 本篇 vs 相关论文的 delta #

Seedance 1.5 Pro vs Seedance 2.0 #

维度Seedance 1.5 Pro (2512.13507)Seedance 2.0 (2604.14148)
发布时间2025-122026-04
核心架构Dual-branch MMDiT + cross-modal joint module [2512.13507]继承 dual-branch MMDiT + cross-modal joint denoising [2604.14148]
输入模态Text, Image → Audio-VideoText, Image, Video, Audio 混合输入 [2604.14148]
任务覆盖T2VA, I2VA, T2V, I2V+ VFX 参考、创意参考、视频续写/扩展 [2604.14148]
RLHFMulti-dimensional reward (motion + aesthetics + audio) [2512.13507]Extended RLHF + "更精细的奖励维度"(具体未披露)[2604.14148]
EvaluationSeedVideoBench 1.5 (proprietary)Arena.AI public leaderboard: T2V Elo 1450 #1, I2V Elo 1449 #1 [2604.14148]
技术披露极少(无 equations/algorithms/model size) [2512.13507]同样极少(evaluation report 性质)[2604.14148]

演进逻辑: 1.5 pro → 2.0 的核心进化是 输入模态扩展(从 text/image 到混合多模态引用)和 任务覆盖泛化(从基础生成到 VFX/续写/扩展)。架构骨架(dual-branch MMDiT + RLHF)保持不变,说明 ByteDance 将此视为 scalable foundation 而非需要重新设计的方向。

Seedance 1.5 Pro vs LongCat-Video #

维度Seedance 1.5 ProLongCat-Video
组织ByteDance (197 authors)美团 (~10 authors) [2510.22200]
架构Dual-branch MMDiT (video + audio)13.6B 稠密 DiT (video only) [2510.22200]
RLHF 方法"Multi-dimensional reward model" (undisclosed) [2512.13507]GRPO adapted to flow matching, 理论证明等价于随机噪声搜索 [2510.22200]
Anti-reward-hacking未讨论多 reward 加权 $\hat{A}_\text{total} = \sum_k w_k \hat{A}_k$ 防止静态视频 hacking [2510.22200]
推理加速>10× (distillation + quant + parallelism) [2512.13507]12.3× (coarse-to-fine + 3D BSA) [2510.22200]
音频Native joint generation无音频
BenchmarkProprietary (SeedVideoBench 1.5)VBench 2.0 开源 #1, Commonsense 全场最佳 [2510.22200]

关键差异: LongCat-Video 的技术透明度远高于 Seedance 1.5 pro——GRPO 适配 flow matching 的理论推导 [2510.22200] 提供了可复现的方法论,而 Seedance 的 "multi-dimensional reward model" 完全不透明。但 Seedance 1.5 pro 的 native audio 能力是 LongCat-Video 完全不具备的差异化优势。

Seedance 1.5 Pro 与 SLA 的加速技术关系 #

Seedance 1.5 pro 声称 >10× inference acceleration [2512.13507],但未披露 attention 层的具体加速手段。SLA/SLA2 在 Wan2.1(同赛道 DiT 模型)上实现 13.7–18.7× attention speedup [2509.24006]。推测 Seedance 的加速栈中可能包含类似的 sparse attention 技术(但采用何种方案未知)。

SLA 的 sparse-linear decomposition 适用于 DiT 的条件:(1) attention weights 分解为 high-rank sparse + low-rank dense [2509.24006],(2) model 可通过 fine-tuning 适应 sparse pattern。Seedance 的 dual-branch 架构引入了额外的 cross-modal attention,其 sparsity pattern 可能与 single-modal video attention 显著不同——cross-modal 对齐通常需要更 dense 的 attention。

§3 可攻击面 #

  1. 技术不透明: 论文无任何 equations、algorithms、model size、training data size 或 architecture detail [2512.13507]。所有 claims(>10× acceleration、native joint generation、multi-dimensional RLHF)均无可验证的技术支撑。对比 LongCat-Video 提供的 GRPO 理论推导和 ablation [2510.22200],Seedance 1.5 pro 的论文更接近产品发布而非学术贡献。
    1. 评估不可复现: SeedVideoBench 1.5 为 proprietary benchmark,且评估使用人工评分("professional film directors")[2512.13507]。LongCat-Video 使用公开的 VBench 2.0 [2510.22200],Seedance 2.0 使用公开的 Arena.AI [2604.14148]——1.5 pro 是三者中唯一完全依赖 proprietary evaluation 的。
      1. Audio 能力未量化: "Native joint audio-video generation" 是核心 claim,但 audio representation(waveform? spectrogram? codec tokens?)、audio quality metrics(PESQ/STOI/FAD)、以及 AV sync 的量化指标均未报告 [2512.13507]。Bar charts (Fig 5-6) 仅给出 GSB pairwise 比较,无绝对分数。
        1. Slow-motion critique 无证据: 论文批评竞品使用 slow-motion "artificially enhance perceived stability" [2512.13507],但未提供任何量化证据(如竞品生成视频的平均运动速度对比)。这是一个 pointed but unsubstantiated claim。
        2. §4 生态位 #

          Seedance 1.5 pro 在视频生成赛道中的定位:

          模型组织音频分辨率开源技术透明度
          Wan 2.5阿里Post-processing720p+部分开源
          Kling 2.6快手Post-processing1080p闭源
          Sora 2OpenAIJoint1080p闭源极低
          Veo 3.1GoogleJoint1080p闭源
          Seedance 1.5 proByteDanceNative joint720p闭源极低
          Seedance 2.0ByteDanceNative joint720p闭源极低
          LongCat-Video美团720p部分开源

          Seedance 1.5 pro 的生态位优势在于 中文语境的 native AV 联合生成(方言支持、唇语同步),但技术透明度是赛道中最低的——这使其科学贡献有限,更多是产品定位文档。

          §5 未探索方向 #

          1. RLHF reward hacking 防护: LongCat-Video 明确提出 multi-reward 加权防止 static video hacking [2510.22200]。Seedance 系列是否存在类似 hacking 问题(如 audio channel 学会产生 silence 以最大化 visual quality reward)?需要 audio-specific reward guardrails。
            1. SLA2 在 cross-modal attention 上的适配: Seedance 的 dual-branch MMDiT 包含 video self-attention、audio self-attention 和 cross-modal attention。SLA2 的 sparse-linear 分解已在 video self-attention 上验证 [2509.24006],但 cross-modal attention 的 sparsity pattern 可能截然不同——audio-video alignment 需要的 dense cross-attention 可能不适合高 sparsity。需要 modality-aware router。
              1. Open-source AV benchmark: SeedVideoBench 1.5 的闭源性阻碍了赛道的科学发展。方向:基于 LongCat-Video 的 VBench 2.0 [2510.22200] 扩展 audio 维度(AV sync, audio quality, expressiveness),创建开放的 audio-video generation benchmark。
                1. Progressive audio-video generation: 当前 Seedance 声称 native joint denoising,但未说明 audio 和 video 的 denoising 是否共享 timestep schedule。方向:audio 和 video 可能有不同的最优 NFE(audio 通常比 video 更快收敛),允许 progressive generation(video 先达到粗糙质量 → 条件化 audio generation → 联合 refinement)可能比严格 joint denoising 更高效。