VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling

Leigang Qu1, Feng Cheng2, Ziyan Yang2, Bangbang Yang2, Zhaoyang Huang2,
Wei Chow1, Yicong Li3, Wenjie Wang3, Tat-Seng Chua1, Yan Zeng2
▶ 1. NExT++ Lab, National University of Singapore
▶ 2. ByteDance Seed
▶ 3. University of Science and Technology of China
NeurIPS 2026
Teaser

Qualitative showcase of VINCIE-NExT across diverse video editing scenarios. Our method enables temporally consistent stylization, dynamic background replacement, identity and attribute editing, scene transformation, and object removal, while preserving motion, structure, and visual coherence across frames.

Abstract

Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video → Image → Image → Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.


Framework

Rather than modeling the edited video directly from the source video and instruction, which demands large-scale paired video data, VINCIE-NExT routes editing intent through the image domain, where mature large-scale editing priors are readily available:

Video → Image → Image → Video

V→I extracts a representative keyframe from the source video, I→I applies the instruction-guided edit at the image level, and I→V recomposes the edited image into a temporally consistent target video. The whole chain is flattened into one interleaved token sequence and modeled by a single Diffusion Transformer, so the image pair (Is, It) acts as an in-context visual demonstration, a pixel-grounded appearance blueprint for every output frame. The backbone is a 3B MM-DiT initialized from the text-to-video model of VINCIE, shared by all sub-tasks and both inference stages.

Framework

Overview of VINCIE-NExT. Video editing is decomposed into a structured chain of sub-tasks, and the image editing pair serves as an in-context visual demonstration that transfers image editing priors to video generation via a unified diffusion backbone.

Joint Training on Multi-Source Data

Because each sub-task is a valid suffix of the interleaved sequence, every sub-task can draw from independently collected data: abundant image editing pairs for I→I, image-conditioned video generation data for I→V, and full-chain samples, all supervised by a single diffusion objective with one set of parameters. We further build video-based in-context image editing (V↔I↔I) data: GPT-4o annotates editing instructions for raw videos, which Seedream 4.5 applies to keyframes, yielding (source video, frame, edited frame) triplets without any paired video editing supervision. Because the source video stays attached as context, each sample teaches exactly the V→I→I sub-step whose output is reused as the demonstration for V→I→I→V, so image editing ability transfers into video editing rather than remaining a standalone skill.

In total, training uses 1.20M image editing pairs (I2I, from OmniEdit), 2.45M paired video edits (V2V*, from OpenVE), and 1.63M V↔I↔I triplets.

TDF3D-RoPE

In-context editing poses two conflicting positional demands: the model must distinguish tokens from different shots (e.g., Is vs. It) while recognizing spatially corresponding patches across shots to propagate edits pixel-faithfully. Time Dual-Frequency 3D RoPE resolves this by fusing two 3D frequency tensors via element-wise addition: an inter-shot component with monotonically increasing global offsets that makes every shot distinguishable, and an intra-video component with zero global offset, so the same spatial location in any two shots shares identical (h, w) positions, enabling content-driven attention and pixel-faithful edit propagation. We also apply Temporal Position Randomization (TPR) to image training data, sampling the temporal index uniformly instead of fixing it to zero, so that the model grounds correspondences in spatial content rather than temporal identity.

Chain-of-Editing (CoE)

At inference, the decomposition is instantiated as two sequential diffusion stages. Stage 1 (V→I→I) generates the edited keyframe from the source video and its temporal mid-frame, processing visual segments frame-independently to match the spatial nature of image editing. Stage 2 (V→I→I→V) generates the edited video conditioned on the source video, source keyframe, and the Stage 1 latent, which is inserted directly in latent space without a VAE decode/re-encode cycle; the same backbone now attends jointly across frames to model temporal dependencies. Because each stage is an independent diffusion process over an explicit interleaved context, the chain provides principled test-time scaling: additional compute (or a stronger external image editor) can be spent to improve editing fidelity without retraining.


Quantitative Results

Evaluation on OpenVE-Bench (431 clips, eight editing categories), scored by Gemini 2.5 Pro on a 1–5 scale. VINCIE-NExT achieves an overall score of 3.08, surpassing all open-source methods (+0.59 over OpenVE-Edit, ~24% relative gain; +48.8% relative over the in-context video editor ICVE) and closing the gap to the proprietary Runway Aleph to 0.57, while substantially reducing the dependence on large-scale paired video editing data. Global Style is the standout category (4.17, above Runway Aleph's 3.72), and Local Remove improves from 1.85 to 3.24 over OpenVE-Edit as edits are routed through a mature image editor. Camera Edit remains the main limitation, since novel-viewpoint synthesis is a geometric operation outside the appearance-editing chain.

MethodResolutionOverall Global
Style
Background
Change
Local
Change
Local
Remove
Local
Add
Subtitle
Edit
Creative
Edit
Camera
Edit
Runway Aleph1280×7203.653.722.624.184.162.783.623.644.53
VACE1280×7201.571.491.552.071.461.261.481.471.62
OmniVideo640×3521.311.111.181.141.141.361.002.261.00
InsViE720×4801.532.201.061.481.361.172.182.021.09
Lucy-Edit1280×7042.152.271.573.201.752.301.612.861.61
ICVE384×2402.072.221.622.572.511.972.092.411.11
DITTO832×4801.984.011.682.031.531.412.811.231.32
OpenVE-Edit1280×7042.493.162.362.981.852.152.912.312.02
VINCIE-NExT (Ours)480×6403.084.172.553.483.242.273.453.412.06

Grey text denotes a closed-source commercial model. Bold: best and underline: second best among open-source methods.

Ablation: Data Composition and Chain-of-Editing

Training on V↔I↔I data alone already enables meaningful video editing, and CoE provides consistent gains when chain-structured data is present. The full combination of all three data sources with CoE achieves the best overall score.

I2IV2V*V↔I↔ICoEOverall Global
Style
Background
Change
Local
Change
Local
Remove
Local
Add
Subtitle
Edit
Creative
Edit
Camera
Edit
✓1.141.291.001.051.051.001.601.201.04
✓✓1.051.051.001.011.031.001.271.061.00
✓1.341.681.001.121.031.062.232.091.03
✓✓1.802.581.862.031.461.481.622.301.15
✓2.493.471.962.672.712.083.402.031.29
✓✓2.573.322.113.072.872.033.491.881.32
✓✓✓2.533.061.942.882.851.973.492.691.29
✓✓✓✓2.653.642.153.033.111.823.202.891.22

Shaded rows apply Chain-of-Editing at inference time. All ablations use the Stage-1 model at 256×256 (the main table uses the Stage-2 model at 480×640), which explains the gap between 2.65 and 3.08.

Is Paired Video Data the Main Driver?

A Pure V2V model (same backbone and V2V data, no decomposition or CoE) scores only 2.32, below OpenVE-Edit (2.49), which trains on a superset of our V2V data, so the gains do not come from a stronger backbone. Full data + CoE reaches 2.65 (+14.2%), and +74.1% on Creative Edit, a category with no V2V supervision, while adding only ~6% visual training tokens.

MethodOverallGlobal
Style
Background
Change
Local
Edit
Creative
Edit
Pure V2V (w/o decomp., w/o CoE)2.323.131.892.221.66
Full data + CoE2.653.642.152.652.89
Δ+14.2%+16.3%+13.8%+19.5%+74.1%

Local Edit averages Local Change, Local Remove, and Local Add.

A blinded pairwise human study on all 431 clips with 10 raters (Fleiss' κ = 0.721) also prefers full data + CoE over V2V*-only + CoE on every dimension, most clearly Prompt Following (38% vs. 24%) and Consistency (28% vs. 9%).

Test-time Scaling and Efficiency

Test-time scaling of Chain-of-Editing

Overall score (line) and inference time (bars) as the denoising steps of the intermediate image-editing stage increase. Scored by Gemini 3.1 Flash Lite.

Spending more denoising steps on the intermediate image-editing stage (1/4/16/64) improves the overall score monotonically from 2.30 to 2.81 (+22.2%) without retraining.

Despite the two-stage design, inference is practical: on one H200 GPU with 50 steps, VINCIE-NExT takes 53.5s (18.4s + 35.1s).

MethodInference Time
AnyV2V11min 46s
ICVE118.1s
Lucy-Edit39.8s
Ours53.5s

On objective VBench metrics, VINCIE-NExT also scores best among recent baselines (imaging quality 0.715, quality average 0.819).

Ablation: TDF3D-RoPE, TPR, and Data Weights

Ablation on TDF3D-RoPE and CoE

TDF3D-RoPE and CoE are complementary; their combination achieves the best result.

Ablation on TPR and CoE

TPR and CoE interact multiplicatively; only their combination reaches peak performance.

Training data sampling weights

Each data source has a narrow optimal sampling weight; balancing all three is necessary.

Effect of Intermediate Image Editing Quality

Thanks to the modular I→I interface, replacing self-editing in Stage 1 with a stronger external image editor (Qwen-Image-Edit or Nano Banana) yields consistent improvement: advances in image editing transfer directly to video without retraining.

Effect of intermediate image editing quality

Qualitative Results

1. Global Stylization and Compositional Editing

Global scene stylization (left, summer-to-snow) and localized compositional editing (right, injecting fire/water/wind), both maintaining temporal consistency without task-specific fine-tuning.

Case study

2. Chain-of-Editing vs. Direct Inference

Comparison of V2V direct inference, CoE with self-editing, and CoE with an external image editor on subject replacement (left) and removal (right). Direct inference produces flickering and ghosting; CoE eliminates this instability by anchoring to a reference image, and routing through an external editor further improves semantic accuracy with no retraining.

CoE comparison

Video Results

Edited videos on OpenVE-Bench from the supplementary material. Videos play automatically when scrolled into view and loop.

1. Comparison with Baselines

Source video alongside Lucy-Edit, ICVE, DITTO, and VINCIE-NExT (Ours).

    2. Chain-of-Editing Ablation

    Direct V2V inference, CoE with self-editing, and CoE with an external image editor (Nano Banana or Qwen-Image-Edit) supplying the intermediate edited keyframe.

      BibTeX

      @inproceedings{qu2026vincienext,
          title={VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling},
          author={Qu, Leigang and Cheng, Feng and Yang, Ziyan and Yang, Bangbang and Huang, Zhaoyang and Chow, Wei and Li, Yicong and Wang, Wenjie and Chua, Tat-Seng and Zeng, Yan},
          booktitle={Advances in Neural Information Processing Systems},
          year={2026}
        }