Learned representation-guided diffusion models for large-image generation.

Alexandros Graikos Srikar Yellapragada Minh-Quan Le Saarthak Kapse Prateek Prasanna Joel Saltz Dimitris Samaras

Proc IEEE Comput Soc Conf Comput Vis Pattern Recognit

Published: June 2024

Diffusion models can improve image generation in specialized fields like histopathology and satellite imagery by utilizing self-supervised learning (SSL) embeddings as stand-ins for human labels, which are hard to obtain.
This new method allows for high-quality images to be created from these embeddings, and it can even generate larger images by combining smaller patches while maintaining their spatial consistency.
The approach enhances classifier performance on both small patch-level and larger scale classification tasks and shows strong adaptability, successfully working with unseen datasets and different input sources, including text descriptions for image synthesis.

To synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed by domain experts and involves hundreds of millions of patches. Modern-day self-supervised learning (SSL) representations encode rich semantic and visual information. In this paper, we posit that such representations are expressive enough to act as proxies to fine-grained human labels. We introduce a novel approach that trains diffusion models conditioned on embeddings from SSL. Our diffusion models successfully project these features back to high-quality histopathology and remote sensing images. In addition, we construct larger images by assembling spatially consistent patches inferred from SSL embeddings, preserving long-range dependencies. Augmenting real data by generating variations of real images improves downstream classifier accuracy for patch-level and larger, image-scale classification tasks. Our models are effective even on datasets not encountered during training, demonstrating their robustness and generalizability. Generating images from learned embeddings is agnostic to the source of the embeddings. The SSL embeddings used to generate a large image can either be extracted from a reference image, or sampled from an auxiliary model conditioned on any related modality (e.g. class labels, text, genomic data). As proof of concept, we introduce the text-to-large image synthesis paradigm where we successfully synthesize large pathology and satellite images out of text descriptions.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC11601131	PMC
http://dx.doi.org/10.1109/cvpr52733.2024.00815	DOI Listing

Publication Analysis

Top Keywords

diffusion models

embeddings ssl

ssl embeddings

models

embeddings

images

learned representation-guided

diffusion

representation-guided diffusion

models large-image

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!

A PHP Error was encountered