Netflix isn't just a streaming giant – it's apparently a serious AI researcher too. The company has unveiled VOID (Video Object and Interaction Deletion), a new open-source model that can remove objects from video scenes and reconstruct the remaining scene in a physically correct manner.
What VOID Can Do
Unlike simple inpainting tools, VOID goes a crucial step further: the vision-language model understands not only where an object was, but also how the remaining scene would behave without it. For example: remove a car from a collision scene, and VOID generates video where the remaining vehicle continues driving in a physically plausible way – no debris, smoke, or flames.
Another scenario: a person jumps into a pool and creates splashes. Remove the person, and VOID shows an undisturbed pool – no waves, no splashes on the edge.
Developed by Netflix Research
The project was developed by researchers from Netflix and Sofia University: Saman Motamed, William Harvey, Benjamin Klein, Luc Van Gool, Zhuoning Yuan, and Ta-Ying Cheng. The accompanying paper is available as a preprint on arXiv.
Open Source on Hugging Face
Netflix has released VOID as an open-source model on Hugging Face, making it available not just for internal productions but for everyone.
Significantly Better Than Competitors
In a user study with 25 participants, VOID was preferred in 64.8 percent of cases – far ahead of Runway (18.4 percent) and other tools like Generative Omnimatte, DiffuEraser, and ProPainter.
- VOID: 64.8% preference
- Runway: 18.4% preference
- Other tools: below 10%
Implications
The technology has obvious potential for film and video production – but simultaneously raises questions about video manipulation. The more convincing AI-powered video editing becomes, the harder it gets to distinguish authentic from manipulated material.
Sources: The Register, VOID Project Page, arXiv Paper