Research Progress of Image Generation Methods Based on Deep Learning

Authors

  • Senwei Zhang

DOI:

https://doi.org/10.54097/0999vz62

Keywords:

Deep Learning, Image Generation, GAN, VAE, CGAN, StackGAN.

Abstract

This paper provides an overview of deep learning-based image generation methods and image style migration techniques. The focus is on text-to-image generation, unconditional and conditional image generation methods, and various approaches for image style migration. The review covers the underlying models, including Generative Adversarial Networks (GAN), Variational Autoencoders (VAE), and diffusion models, while also discussing conditional image generation methods such as Conditional Generative Adversarial Networks (CGAN) and Stacked Generative Adversarial Network (StackGAN). Furthermore, it explores deep learning-based image style migration methods, including both image iteration-based and model iteration-based approaches. The possibilities and opportunities for the advancement and use of these techniques in the field of picture production are highlighted in the review's conclusion.

Downloads

Download data is not yet available.

References

LI Yueyang, TONG Guoxiang, ZHAO Yingzhi, LUO Qi. A Survey of Text-to-Image Synthesis Based on Generative Adversarial Network. Electronic Science and Technology, 2023, 36(10): 39-55.

Zhu X, Goldberg A B, Eldawy M, et al. A text-to-pictur. Van-synthesis system for augmenting communication cover: Proceedings of the Sixth International Conference on Learning Representations, 2018: 1710-1735.

YI Yaokuan. Research on Generative Adversarial NetworkResearch on Algorithm of Text to Image Based. Changchun University of Technology, 2021.

Kingma DP, Welling M. Auto-encoding variational Bayes. IEEE Computer Society, 2021, 43(12): 4217-4228.

Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative adversarial nets. Advances in neural information processing systems. Cambridge MIT Press, 2014: 2672-2680.

AGNESE J, HERRERA J, TAO H, et al. A survey and taxonomy of adversarial neural networks for text-to-image synthesis. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2020, 10(4): 1345.

PAN Z, YU W, YI X, et al. Recent progress on generative adversarial networks (GANs): A survey. IEEE Access, 2019, 7: 36322-36333.

WANG Zhendong, ZHENG Huangjie, HE Pengcheng, et al. Diffusion-GAN: Training GANs with diffusion. arXiv: 2206.02262, 2022.

Karras T, Laine S, Aila T. A Style-Based Generator Architecture for Generative Adversarial Networks. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2019: 4401-4410.

F.-A. Croitoru, V. Hondru, R. T. Ionescu and M. Shah, Diffusion Models in Vision: A Survey, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(9): 10850-10869.

A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, "Hierarchical text-conditional image generation with CLIP latents, 2022.

H. Zhang, V. Sindagi and V. M. Patel, Image De-Raining Using a Conditional Generative Adversarial Network, IEEE Transactions on Circuits and Systems for Video Technology, 2020, 30(11) : 3943-3956.

Zhang H, Xu T, Li H, et al. StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks. 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, 2017: 5907-5915.

H. Zhang, StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, 41(8) :1947-1962.

Zhang, L., Rao, A., Agrawala, M. Adding Conditional Control to Text-to-Image Diffusion Models. IEEE/CVF International Conference on Computer Vision, 2023, 3836-3847.

Hertzmann A, Jacobs C E, Oliver N. Image analogies, Proc of the 28th Annual Conference on Computer Graphics and Interactive Techniques. New York: ACM Press, 2001: 327-340.

Gatys L A, Ecker A S, Bethge M, et al. Controlling perceptual factors in neural style transfer, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017: 3985-3993.

Gatys L A, Ecker A S, Bethge M. A neural algorithm of artistic style. 2015.

Li Chuan, Wand M. Combining Markov random fields and convolutional neural networks for image synthesis, IEEE Conference on Computer Vision and Pattern Recognition, 2016: 2479-2486.

Liao Jing, Yao Yuan, Yuan Lu, et al. Visual attribute transfer through deep image analogy. ACM Trans on Graphics, 2017, 36(4) : 120.

Justin J, Alexandre A, Li Feifei. Perceptual losses for real-time style transfer and super-resolution, European Conference on Computer Vision, 2016: 694-711.

Li Yijun, Fang Chen, Yang Jimei, et al. Universal style transfer via feature transforms. 2017.

Downloads

Published

26-06-2024

How to Cite

Zhang, S. (2024). Research Progress of Image Generation Methods Based on Deep Learning. Highlights in Science, Engineering and Technology, 103, 28-35. https://doi.org/10.54097/0999vz62