Exploring the Realm of Automated Music Composition with Models based on Transformer and its Variants

Authors

  • Shuanglong Zhu

DOI:

https://doi.org/10.54097/gdz0mc66

Keywords:

Transformer, Self-Attention Mechanism, Relative Positional Encoding, Sparse Attention, Variational Autoencoders (VAEs).

Abstract

The Transformer architecture and its modelling variants have great potential in the music domain. Traditional approaches, such as rule-based systems and RNNs, have limitations in capturing the complex temporal dependencies and hierarchical structures inherent in music. With the introduction of its self-concern mechanism, the Transformer can effectively address this problem by capturing remote dependencies and allowing parallel processing of sequences. Transformer-based model variants such as Music Transformers, MuseNet and Jukebox have demonstrated that they can generate high-quality, varied, and stylistically rich compositions. Despite the success of these models, some challenges still need to be solved, such as high computational requirements and limited control over musical style and emotion. Possible future research directions include optimizing computational efficiency, enhancing stylistic and emotional control, and developing hybrid models that combine other models with Transformer. This article provides a comprehensive overview of the current state of research and potential and future applications of Transformer for music generation.

Downloads

Download data is not yet available.

References

[1] Eck, D., & Schmidhuber, J. (2002). Finding temporal structure in music: Blues improvisation with LSTM recurrent networks. Proceedings of the 12th IEEE Workshop on Neural Networks for Signal Processing. DOI: 10.1109/NNSP.2002.1030094.

[2] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. https://doi.org/10.48550/arXiv.1706.03762.

[3] Hadjeres, G., & Pachet, F. (2016). DeepBach: a steerable model for Bach chorales generation. https://doi.org/10.48550/arXiv.1612.01010.

[4] Huang, C. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Simon, I., Hawthorne, C., & Eck, D. (2018). Music transformer. https://doi.org/10.48550/arXiv.1809.04281.

[5] Payne, Christine. "MuseNet." OpenAI, 25 Apr. 2019. Available at: openai.com/blog/musenet.

[6] Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., & Sutskever, I. (2020). Jukebox: A generative model for music. https://doi.org/10.48550/arXiv.2005.00341.

[7] Huang, C. A., Cooijmans, T., Roberts, A., Courville, A., & Eck, D. (2019). Counterpoint by convolution. https://doi.org/10.48550/arXiv.1903.07227.

[8] Shaw, P., Uszkoreit, J., & Vaswani, A. (2018). Self-Attention with Relative Position Representations. ArXiv: 1803.02155 [cs.CL]. https://doi.org/10.48550/arXiv.1803.02155.

[9] Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., & Salakhutdinov, R. (2019). Transformer-XL: Attentive language models beyond a fixed-length context. https://doi.org/10.48550/arXiv.1901.02860.

[10] Katharopoulos, A., Vyas, A., Pappas, N., & Fleuret, F. (2020). Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. ArXiv: 2006.16236 [cs.LG]. https://doi.org/10.48550/arXiv.2006.16236.

[11] Child, R., Gray, S., Radford, A., & Sutskever, I. (2019). Generating Long Sequences with Sparse Transformers. https://doi.org/10.48550/arXiv.1904.10509.

[12] Murty, S., Sharma, P., Andreas, J., & Manning, C. D. (2023). Grokking of Hierarchical Structure in Vanilla Transformers. ACL 2023. https://doi.org/10.48550/arXiv.2305.18741.

Downloads

Published

18-02-2025

How to Cite

Zhu, S. (2025). Exploring the Realm of Automated Music Composition with Models based on Transformer and its Variants. Highlights in Science, Engineering and Technology, 124, 38-44. https://doi.org/10.54097/gdz0mc66