Overcoming Data Sparsity in Social Computing: A GMM+LLM Approach
DOI:
https://doi.org/10.54097/4qbjsn64Keywords:
Large Language Models (LLMs), Gaussian Mixture Models (GMM), CAS, Social Media Simulation, Generative Framework.Abstract
The rise of social media as the primary platform for high school students’ interactions presents both opportunities and challenges for social computing research. While social media offers rich insights into adolescent social behavior, research often suffers from insufficient sample sizes, leading to biased portrayals of social dynamics. To address this limitation, this study proposes a novel framework that integrates Gaussian Mixture Models (GMM) with Large Language Models (LLMs) to generate linguistically detailed and context-sensitive synthetic data. Using a dataset of 15,000 student social profiles collected from 2006 to 2009, GMM clustering was applied to identify distinct user groups, followed by LLM-driven content generation tailored to each cluster through structured prompt engineering. Comparative analysis with K-means clustering demonstrates GMM’ s superior ability to handle outlier-prone data, producing smoother and more representative clusters. The synthesized content successfully reflects authentic patterns of adolescent engagement across domains such as sports, music, religion, and negative topics, thereby mitigating the problem of data sparsity. This study highlights the potential of combining probabilistic modeling with advanced natural language generation to enhance the authenticity and scalability of social dynamics simulation. Future research should expand contextual features and employ robust evaluation mechanisms to further refine generative accuracy.
Downloads
References
[1] Smith A, Anderson M, Jiang J. Social media use in 2021: A deep dive into the adolescent experience. J Adolesc Health. 2021; 68 (2): 145-52.
[2] Chen L, Wang H. Data sparsity and bias in computational social science: A review of challenges and solutions. IEEE Trans Comput Soc Syst. 2022; 9 (4): 1120-32.
[3] Squazzoni F, Polhill GJ, Edmonds B, et al. Computational models that matter during a global pandemic: The case of COVID-19. ACM Comput Surv. 2020; 53 (6): 1-37.
[4] Li Y, Wang D, Sun Y. SparseGAN: Synthesizing social network data with generative adversarial networks for research on rare populations. IEEE Trans Knowl Data Eng. 2021; 34 (8): 3892-905.
[5] Reynolds DA, Rose RC. Robust text-independent speaker identification using Gaussian mixture speaker models. IEEE Trans Speech Audio Process. 1995; 3 (1): 72-83. doi: 10.1109/89.365379. DOI: https://doi.org/10.1109/89.365379
[6] Zabihullah18. Students’ social network profile clustering [dataset]. Kaggle. 2020. Available from: https://www.kaggle.com/datasets/zabihullah18/students-social-network-profile-clusterin.
[7] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. Adv Neural Inf Process Syst. 2017; 30: 5998-6008.
[8] He A, Luo C, Tian X, Zeng W. A twofold siamese network for real-time object tracking. In: Proc IEEE Conf Comput Vis Pattern Recognit (CVPR). 2018. p. 4834-43. DOI: https://doi.org/10.1109/CVPR.2018.00508
[9] Melekhov I, Kannala J, Rahtu E. Siamese network features for image matching. In: Proc 23rd Int Conf Pattern Recognit (ICPR). IEEE; 2016. p. 378-83. DOI: https://doi.org/10.1109/ICPR.2016.7899663
[10] Creswell A, White T, Dumoulin V, Arulkumaran K, Sengupta B, Bharath AA. Generative adversarial networks: An overview. IEEE Signal Process Mag. 2018; 35 (1): 53-65. DOI: https://doi.org/10.1109/MSP.2017.2765202
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Academic Journal of Science and Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.








