Read the original at HF Daily Papers
Researchers introduce ALIVE, a framework that inserts objects into videos with coherent interactions using a 35,800-pair dataset and a vision-language model for guidance.
Carried by: HF Daily Papers. First seen: .