This is a recurring event: View all events in the series “Data Bites”
Abstract
Diffusion models have transformed visual synthesis, yet generating accurate, realistic content from incomplete views remains a core challenge.
This talk presents two recent works: Multi-Modal UNet-based Feature Encoder (MUFEN), where a diffusion model directly generates photo-realistic hand gestures from multi-view, multi-modal priors, and Video Diffusion-Aware 4D Reconstruction (ViDAR), where a diffusion model enriches the training data to enable accurate 4D scene reconstruction from monocular video. Both systems share a common strategy: coupling powerful generative priors with domain-specific geometric constraints to overcome occlusion, motion ambiguity, and spatial inconsistency.
I will discuss how this paradigm produces reconstructions that are both visually compelling.
About the speaker
Greg Slabaugh is Professor of Computer Vision and AI and Director of the Digital Environment Research Institute (DERI) at Queen Mary University of London (QMUL).
His primary research interests include computer vision and deep learning, with various applications in computational photography and medical imaging. His career has included positions both in industry (Huawei, Medicsight, Siemens) and academia (QMUL and City, St George’s University of London).
Attendance at City St George's events is subject to our terms and conditions.