CrossEdit: Cross-Modal Training Enables Rich Audio-Visual Editing

Abstract

Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within one modality, since diverse, high-quality editing pairs are hard to obtain, especially for cross-modal tasks where aligned audiovisual (AV) data is scarce. We present CrossEdit, a unified omni-modal editing model for images, audio, and video that performs AV movie scene edits zero-shot through cross-modal transfer. Our key observation is that complex instruction-following, once learned in any modality, generalizes to others. We exploit the additive nature of acoustic signals to procedurally generate large-scale audio editing pairs with compositional instructions, and show that fine-tuning on this synthetic data and self-supervised AV masked reconstruction, alongside a targeted set of curated cross-modal tasks, induces instruction-following that transfers zero-shot to unseen modality and instruction combinations. To evaluate this capability, we release CrossEditBench, a human-annotated benchmark of AV edits on movie scenes, and propose AV-FES, a metric that jointly scores instruction following and consistency. We show that our proposed techniques improve performance on zero-shot AV movie scene editing, while maintaining or improving performance on lip-synced speech editing and standard image, video, and audio editing benchmarks.

Cross-method Comparisons

Speech

Instruction Source Stable Avatar HuMo CrossEdit-Base CrossEdit
Change the speech to "and things are stretched out to the point where now it's almost getting funny"
Change the speech to And that may be an understood thing between him and his wife and that's all falling apart.
Change the speech to There were a lot of small chapters in the history books.
Change the speech to Who would make the discovery? How would it be made and how long would it take?

Scene

Instruction Source HY-Video-Foley XXL HY-Video-Foley XL VACE + Coherent AvED AVI-Edit MiniMax-H3 InstructAV2AV CrossEdit-Base CrossEdit
Remove the right man's hat, he starts saying "It is much too dark in here Johnny"
Make the video and music more happy
Remove the static
Remove the music and change the speech to do you recognize her

Ablation

Instruction Source CrossEdit-Base + Audio Editing + Audio Editing + AV Masked Recon.
Change the speech to "and things are stretched out to the point where now it's almost getting funny"
Change the speech to And that may be an understood thing between him and his wife and that's all falling apart.
Change the speech to There were a lot of small chapters in the history books.
Change the speech to Who would make the discovery? How would it be made and how long would it take?
Remove the right man's hat, he starts saying "It is much too dark in here Johnny"
Remove the dialogue
Change the man to a dog astronaut
Make the video and music more happy
Remove the static
Remove the music and change the speech to do you recognize her
Add a dinosaur roaring

Audio+Video → Audio+Video: Movie Editing (Zero-Shot)

Instruction Input Video Output Video
Add a dinosaur roaring
Make the people's dress more formal. In the audio, change the saxophone to piano
Replace the magic carpet with a space ship
Change the man to a dog astronaut
No visible changes. Remove the dialogue
Change the scene so that the men are running around the foundation, their breathing is audible
Add lightning strikes hitting one planet from another, with echoing sounds of thunder
Ducks quack as they swim across the pond
Colorize and sound effects to the scene
Replace the pink ball with a cartoon clown

Audio+Video → Audio+Video: Voice Cloning

Interviews

Instruction Input Video Output Video
Change the speech to There have been many times we have been doing scenes and I'll be doing all
Change the speech to So, but she's also a very calm person. She needs to be calm in order to...
Change the speech to 5% of the world's crops depend on their cooperation.
Change the speech to 5% of the world's crops depend on their cooperation.
Change the speech to She was this feisty young woman who spoke her mind
Change the speech to Its kind of an action, its kind of a comedy
Change the speech to I've been doing a lot of heavy movies lately
Change the speech to Twilight and Applejack and Rainbow Dash and Rarity and Fluttershy
Change the speech to and then I'm like, I gotta, I have to, A, have to do this,

Movies

Instruction Input Video Output Video
Change the speech to I could do this all day
Change the speech to It's over, I have the high ground
Change the speech to I suspect the squirrel has a tiny wire Watson
Change the speech to This time for Africa
Change the speech to I am the one who knocks!

Cross-lingual (Zero-Shot)

Instruction Input Video Output Video
Change the speech to Tu sous-estimes mon pouvoir.
Change the speech to 牛逼啊

Video → Video: Editing

Instruction Input Video Output Video
Replace the green bowl with a mystical crystal orb
Add a magical rainbow in the background
Transform the entire video to resemble a vivid oil painting, enhancing the textures and brush strokes for a painterly effect
Add a pair of realistic sunglasses to the man in the video. The sunglasses should be proportionate to his face, with a sleek black frame and dark reflective lenses that subtly mirror the environment. Ensure the glasses follow the natural movement and tilt of his head, with proper perspective and slight shadowing on the face where the frames rest.
Remove the black car from the video
Insert a butterfly
Change the animal that is swimming to a boat
Make the place where the man stands made by water

Video + Image to Video

Instruction Input Video Input Image Output Video
Insert a paper boat in the water Input image
Insert a car Input image

Image → Image: Editing

Instruction Input Image Output Image
Transfer the image into a traditional ukiyo-e woodblock-print style. Input image Output image
Remove the red trolley (marked \"77\" and labeled \"WEST CHESTER\") from the railway track in the foreground. Input image Output image
Add a small brown dog sitting inside the open trunk of the antique car. Input image Output image
Add a modern skyscraper in the background. Input image Output image
Change the interior setting in the image from a historic or vintage room to a modern office environment. Input image Output image
Replace the blue bird in the image with a red fox. Input image Output image

Audio-to-Audio: Instruction-based Source Separation

Despite being only trained on simulated separation data, the model can generalize to real, complex scenes. In the following demo, we vary the separated sounds throughout the video.



Audio → Audio: Editing

Synthetic Scenes (AudioChat)

Instruction Input Audio Output Audio
Add the sound of someone panicking
Remove the door creak
Let's make the raven sound more ominous. More drawn out, lower pitched.
I think the footsteps are a little too heavy. Can we make them lighter and softer?
Add a whispering voice

Real World Audio (AudioCaps)

Instruction Input Audio Output Audio
Add the sound of a bird chirping
Remove the distant kitchen ambience
Add a low rumbling thunder sound to the background
Remove the overlapping pan water sound
Reduce the volume of the dog sounds

Image → Video+Audio

Instruction Input Image Output Video
Several penguins stand and interact on a snow-covered iceberg while small seals emerge from the nearby water, one sliding partway onto the ice as light snowfall begins around them. Input image
A small dog riding in a backpack on a seated hiker's back turns its head toward the camera with curiosity and barks twice. Input image
The camera zooms out from a close view of chickens by a wire fence to reveal a busy pen filled with many chickens moving around and foraging Input image
Cute little animated characters emerge onto the upper balconies of a pastel row of buildings while additional characters come out from the ground-floor shops below, waving and calling to one another across the street frontage. Input image
A stylized 3D character stands centered against a plain backdrop, waving toward the camera while jumping up and down in place. Input image
A woman stands before a large oval mirror while her reflected self walks deeper away within the mirror, shrinking in the reflected space as she watches. Input image
A hooded figure at the far end of a dark hallway moves forward toward the camera as the overhead light flickers unevenly. Input image
Three first responders in orange flight suits roll a patient on a gurney beside a helicopter, coordinating with each other as they prepare to load the patient into the aircraft on the left Input image
An elderly couple walks in from the left into a cozy 3D living room set, sits down in the two armchairs facing the coffee table, and begins chatting together. Input image

Image+Audio → Video+Audio

Instruction Input Image Input Audio Output Video
A woman with short, wavy blonde hair and a gentle smile is seated in a director's chair, looking at the camera. She wears a dark grey turtleneck sweater under a black quilted leather jacket. The setting is a dimly lit room, possibly an office or archive, with shelves of boxes and a desk with a typewriter visible in the soft-focus background. The lighting is warm and focused on her, creating an intimate interview-style atmosphere, and the person's lips move in motion with the speech. Audio: The woman speaks enthusiastically about a story, her expression matching her positive tone. 'I think the audiences are going to be riveted by the story. It is absolutely gripping.' [aɪ θˈɪŋk ðə ˈɔːdiənsɪz ɑːɹ gˈoʊɪŋ tə bi ɹˈɪvɪtɪd baɪ ðə stˈoːɹi ɪt ɪz ˌæbsəlˈuːtli gɹˈɪpɪŋ] Input image
A woman with dark wavy hair and a pensive expression sits by the window on a moving train, wearing a dark camisole with lace trim and maroon pants. The seat behind her has a patterned blue fabric, and the lighting is cool and natural. The shot is a medium close-up, and the persons lips move in motion with the speech. Audio: The woman on the train speaks softly, her voice blending with the rhythmic rumble of the train on its tracks. 'que luego hablamos. Un besito.' [kˈe lʊˈeɣo aβlˈamos ˈumbesˈito] Input image
A middle-aged man with fair skin and wavy, shoulder-length dark brown hair sits for an interview. He wears a blue short-sleeved button-down shirt open at the collar over a white undershirt. His expression is thoughtful as he looks slightly down and to his left. The setting is an indoor room with soft lighting. In the out-of-focus background, a wooden dresser with a mirror and a vintage metal fan are visible against a neutral-colored wall. The camera is framed in a medium close-up, and the persons lips move in motion with the speech. Audio: The man in the blue shirt speaks in a calm, measured tone, his expression matching the contemplative nature of his words. 'She has just a sense of confidence and a' [ʃiː hæz dʒˈʌst ɐ sˈɛns ʌv kˈɑːnfɪdɪns ænd ɐ] Input image
A man with short blond hair in a blue sweater stands next to a man in a white chef's coat and another man in a suit holding a Guinness World Records folder. They are in a brightly lit studio in front of a live audience, standing behind a counter with a cutting board, knife, blindfold, and carrots. The blond man is speaking with an animated expression, and his lips move in motion with the speech. Audio: The blond man speaks with an excited, announcer-like voice to the audience about an upcoming world record attempt. 'tonight, he's going to try and set a new world record for the most carrot slices done in thir' [tənˈaɪt, hiz gˈoʊɪŋ tə tɹˈaɪ ænd sˈɛt ɐ nˈuː wˈɜːld ɹˈɛkɚd fɔːɹ ðə mˈoʊst kˈæɹət slˈaɪsᵻz dˈʌn ɪn θɜːɹ ] Input image
A young woman with long brown hair and fair skin is shown in a close-up against a solid black background. She wears a black ribbed short-sleeved shirt. The scene is dramatically lit from one side, casting deep shadows and highlighting her serious expression as she looks slightly off-camera, the persons lips move in motion with the speech. Audio: The woman with long brown hair speaks in a calm and reflective tone, her voice matching the serious and intimate mood created by the low-key lighting. 'fue una situación de mucho respeto y' [fwˈe ˈuna sitwasjˈon de mˈutʃo respˈeto i] Input image
A close-up shot of a young woman with fair skin, dark, styled hair, and dark eyes, sitting in the driver's seat of a car at night. She wears a lavender top and a delicate pearl necklace. The scene is dramatically lit, with her face illuminated while the interior of the car and the background remain in deep shadow. The steering wheel is visible in the foreground. She looks off to her right with a thoughtful and slightly wistful expression, and the person's lips move in motion with the speech. Audio: The woman in the lavender top speaks with a clear, slightly melancholic tone, her voice accompanied by faint, dramatic orchestral music. 'Yeah, she has something I don't have. Nice' [jˈæ, ʃi hˈæz sˈʌmθɪŋ ˈaɪ dˈoʊnt hˈæv. nˈaɪs] Input image
A man with dark skin, in his late 20s or early 30s, is shown in a medium close-up shot. He has short, textured black hair with faded sides, a neat goatee, and a small silver hoop earring in his left ear. He wears a thick, silver-colored chain necklace and a light blue, short-sleeved button-up shirt with a faint white pattern. He is speaking with an engaged expression, his mouth slightly open, and the person's lips move in motion with the speech. The background is a softly lit, out-of-focus room with neutral-colored walls and a blurry abstract painting with blue and orange tones. Audio: The man speaks conversationally, recounting a story from inside the room. 'in the middle of dinner, LED was just like, 'Hey, would you would you go on a trip with your ex?' Just' [ɪn ðə mˈɪdəl ʌv dˈɪnɚ, ɛl i dˈiː wʌz dʒˈʌst lˈaɪk, hˈeɪ, wʊd jˈuː wʊd jˈuː gˈoʊ ɑːn ɐ tɹˈɪp wɪθ jʊəɹ ˈɛks? dʒˈʌst ] Input image
A young woman with fair skin, dark hair, and a bruise on her cheek lies on a couch, propped up by a yellow and white patterned pillow. She wears a white cervical neck brace and a light-colored camisole, her expression weary and pained. One hand with black nail polish rests on her chest while the other touches the brace. The room is dimly lit, with a patterned couch and a small end table visible in the background. The back of another person's head is out of focus in the foreground, and the person's lips move in motion with the speech. Audio: The injured woman on the couch speaks in a weak, shaky voice, her tone conveying fear and exhaustion. 'It wasn't a person. It was a' [ɪt wˈɑːzənt ɐ pˈɜːsən. ɪt wʌz ə] Input image
Several people in yellow jumpsuits walk through an institutional corridor with dark vertical panels on the walls. A person with short brown hair is seen in profile, walking with a determined look. In the background, a man in a white shirt and tie moves past, while the heads of other individuals are visible in the blurry foreground. The scene is dimly lit, creating a tense and serious atmosphere. Audio: Tense, rhythmic, and percussive music underscores the scene, mixed with the sound of shuffling footsteps as the group moves forward. Input image
Three men stand outdoors by a railing in a gritty, industrial setting with steel girders in the background. The man on the right, wearing a dark fedora and overcoat, holds a paper cup and speaks. In the center, a man in a dark three-piece suit and patterned tie listens, while the man on the left, in a dark coat with the collar up, looks on seriously, holding a cup and a metal lunch pail. The scene has a vintage 1970s film aesthetic. The man in the fedora's lips move in motion with the speech. Audio: The man in the fedora speaks in a serious, conspiratorial tone in French, as his two companions listen in the industrial environment. 'Il aime l'argent. Et je possède un élément que ni lui' [il ɛm laʁʒɑ̃ e ʒə pɔsɛd ɛ̃n‿ elemɑ̃ kə ni lɥi] Input image
At night, a woman with blonde hair, wearing a light-colored top and shorts, runs across a dirt yard while holding a rifle. She is in front of a large, two-story white house with overgrown vines. A motorcycle is parked nearby, next to a large corrugated metal water tank. In the background, a bald man in a light shirt emerges from behind the building. The scene is dimly lit, creating a tense and suspenseful atmosphere under a dark night sky. Audio: The sound of footsteps on gravel is heard as the woman runs, followed by the loud, sharp report of a gunshot. Input image
An East Asian man with short dark hair, wearing a black leather jacket with a red collar over a graphic t-shirt, sits in a chair in a futuristic control room. He is framed in a medium shot, leaning forward with his hands clasped as he speaks. The background is filled with complex electronic equipment, including monitors and a large, glowing blue panel with intricate geometric patterns, and the person's lips move in motion with the speech. Audio: The man in the black leather jacket speaks in a calm, thoughtful tone while sitting in the high-tech room. 'As an actor, you know, we always strive to' [æz ən ˈæktɚ jˈuː nˈoʊ wˈiː ˈɔːlweɪz stɹˈaɪv tˈuː] Input image

BibTeX

@article{chen2026crossedit,
  title={CrossEdit: Cross-Modal Training Enables Rich Audio-Visual Editing},
  author={Chen, William and Seetharaman, Prem and Chen, Ke and Nieto, Oriol and Duarte, Kevin and Iyer, Siddharth Srinivasan and Rizve, Mamshad Nayeem and Cao, Zhiwen and Watanabe, Shinji and Xiong, Yuanjun and Zhang, Jianming and Jin, Zeyu and Salamon, Justin},
  journal={arXiv preprint arXiv:2610.10264},
  year={2026},
  url={https://arxiv.org/abs/2610.10264}
}