Events

Image generation with end-to-end training and benefits of a good VAE

Join the live event on Thu, 9 April 2026 at 13:00 (CEST) or watch the recording on demand afterwards.

Seminar | Series


Speakers

  • Liang Zheng — Australian National University

Abstract

Latent diffusion models underly modern image generation, which requires a variational auto-encoder (VAE) for image encoding and decoding, and a diffusion transformer for generation. While end-to-end training has been the spirit of deep learning, it is surprising that latent diffusion models are not trained end-to-end, causing representation bottlenecks. In this talk, I will introduce our work that jointly trains the VAE and diffusion transformer and show how it accelerates training and yields high quality images. Further, I will discuss use cases where the resulting end-to-end trained VAEs bring significant benefits. This includes higher-quality text-to-image generation and automatic agentic search of diffusion transformer architectures. I will conclude with new perspectives.

Looking for more in the field?

Explore more events from JIVP Webinar Series and Springer Nature's Computational Science.