← 返回研究成果

CVPR 2026 / 2026

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

Zhou, M., Qin, S.Z., Li, Y.Z., Chen, Q., Jiang, P.

共同一作英文论文
AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation — 查看完整图片
查看完整图片 ↗

研究简介

短视频已成为数字广告的主要媒介,需要可扩展且高效的内容创作方式。AutoCut是一个基于多模态离散化和可控编辑的端到端广告视频编辑框架。它使用专用编码器提取视频和音频特征,通过残差向量量化将其离散化为与文本表示对齐的统一token,构建共享的视频-音频-文本token空间。基于基础模型,通过多模态对齐和监督微调进一步开发了用于视频编辑的多模态大语言模型,在统一编辑框架中支持视频筛选与排序、脚本生成和背景音乐选择等任务。

引用格式

ZHOU M, QIN S Z, LI Y Z, et al. AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation[C/OL]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2026: 37777–37787. https://arxiv.org/abs/2603.28366.

BibTeX
@inproceedings{20260330AutoCut,
  title = {AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation},
  author = {Zhou, Milton and Qin, Sizhong and Li, Yongzhi and Chen, Quan and Jiang, Peng},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year = {2026},
  pages = {37777--37787}
}