Back close

Blended multi-class text to image synthesis GANs with RoBerTa and Mask R-CNN

Publication Type : Journal Article

Publisher : Elsevier BV

Source : Procedia Computer Science

Url : https://doi.org/10.1016/j.procs.2023.01.065

Keywords : Computer Vision, Natural Language Processing, RoBERTa, Mask R-CNN, AttnGAN, Poisson Blending

Campus : Coimbatore

School : School of Computing

Year : 2023

Abstract : Generation of scenes from text description requires employing computer vision for processing image and natural language processing for decoding the text provided. The task here is difficult as it requires generating multiple images of different classes together. Existing methods construct images from captions using a single dataset class, utilising Generative Adversarial Networks (GANs) approaches. Training several classes with GANs is a cumbersome task that needs a large amount of data and huge computational power. With complex datasets the efficiency of the generated image according to text description may be poor. In this paper, we propose an application that generates images based on the Caltech-UCSD bird and Oxford 102 flowers dataset. It leverages the Attentional Generative Adversarial Network (AttnGANs) as the generative model and the RoBERTa neural language model for word embeddings to build an image from multiple classes trained in isolation. The image created by GANs is segmented using Mask R-CNN and blended together using Poisson blending. The application can create scenes based on the description provided by the user and produce them in less time compared to training multiple classes in a single generative network.

Cite this Research Publication : M Siddharth, R Aarthi, Blended multi-class text to image synthesis GANs with RoBerTa and Mask R-CNN, Procedia Computer Science, Elsevier BV, 2023, https://doi.org/10.1016/j.procs.2023.01.065

Admissions Apply Now