Image Generation Based on Diffusion Models: Quality, Efficiency, and Controllability

Authors

  • Xiyao Xu School of Management, Hefei University of Technology, Anhui, 230009, China

DOI:

https://doi.org/10.54097/mx822h46

Keywords:

Diffusion model, Sampling acceleration, Conditional generation, Image generation, Large language model.

Abstract

Image generation is a crucial technology in computer vision. Currently, deep learning-based image generation models are widely used in commercial, entertainment, and industrial applications. With the rapid transformation of the digital economy, higher demands are being placed on image generation. At a time when traditional image generation models are reaching their limits, diffusion models, with their unique advantages, have gained widespread attention in academia and have now become the mainstream paradigm in the field. This paper reviews the core technological achievements of diffusion models in recent years, outlines their evolution in terms of quality, efficiency, and controllability, and summarizes the core issues of each technology. However, due to the inherent constraints of the diffusion model generation paradigm, there is a natural zero-sum game among the three dimensions of quality, efficiency, and controllability—a rigid conflict between them. Finally, this paper analyzes the current state of diffusion models and provides an outlook on future research, with a focus on integrating diffusion models with large language models.

References

[1] Krizhevsky Alex, et al. ImageNet classification with deep convolutional neural networks. Communications of the ACM, 2012, 60: 84-90.

[2] Zhai Zhengli, et al. A review of variational autoencoder models. Computer Engineering and Applications, 2019, 55(03): 1-9.

[3] Wang Kunfeng, et al. Research progress and prospects of generative adversarial networks (GAN). Acta Automatica Sinica, 2017, 43(03): 321-332. doi:10.16383/j.aas.2017.y000003.

[4] Sohl-Dickstein J, Weiss E, Maheswaranathan N, et al. Deep unsupervised learning using nonequilibrium thermodynamics. In: International Conference on Machine Learning, PMLR, 2015: 2256-2265.

[5] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 2020, 33: 6840-6851.

[6] Dhariwal P, Nichol A. Diffusion models beat GANs on image synthesis. Advances in Neural Information Processing Systems, 2021, 34: 8780-8794.

[7] Peebles W, Xie S. Scalable diffusion models with transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023: 4195-4205.

[8] Rombach R. High-Resolution Image Synthesis with Latent Diffusion Models [Internet]. arXiv [cs.CV], 2021.

[9] Luo S, Tan Y, Huang L, et al. Latent consistency models: Synthesizing high-resolution images with few-step inference. arXiv preprint arXiv:2310.04378, 2023.

[10] Ho J, Salimans T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.

[11] Zhang L, Rao A, Agrawala M. Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023: 3836-3847.

[12] Song J, Meng C, Ermon S. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.

Downloads

Published

13-08-2026

Issue

Section

Articles

How to Cite

Xu, X. (2026). Image Generation Based on Diffusion Models: Quality, Efficiency, and Controllability. Mathematical Modeling and Algorithm Application, 9(3), 84-89. https://doi.org/10.54097/mx822h46