Key References
Below is a curated list of key references on simulating survey responses and human behavior with large language models. Use the search bar to quickly find a paper by author, title, or year; each entry links directly to the paper.
- Aher, G., Arriaga, R. I., & Kalai, A. T. (2023). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. arXiv preprint arXiv:2208.10264. [Link]
- Ahnert, G., Haensch, A.-C., Plank, B., & Strohmaier, M. (2026). Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 41554–41577. [Link]
- Anthis, J. R., Richardson, S. M., Kozlowski, A. C., Koch, B., Brynjolfsson, E., Evans, J., & Bernstein, M. S. (2025). Position: LLM Social Simulations Are a Promising Research Method. Proceedings of the 42nd International Conference on Machine Learning. [Link]
- Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3), 337–351. [Link]
- Barrie, C., & Törnberg, P. (2025). Emergent LLM Behaviors Are Observationally Equivalent to Data Leakage. arXiv preprint arXiv:2505.23796. [Link]
- Baumann, J., Röttger, P., Urman, A., Wendsjö, A., Plaza-del-Arco, F. M., Gruber, J. B., & Hovy, D. (2025). Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation. arXiv preprint arXiv:2509.08825. [Link]
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis, 32(4), 401–416. [Link]
- Castricato, L., Lile, N., Rafailov, R., Fränken, J.-P., & Finn, C. (2025). PERSONA: A Reproducible Testbed for Pluralistic Alignment. Proceedings of the 31st International Conference on Computational Linguistics, 11348–11368. [Link]
- Cheng, M., Durmus, E., & Jurafsky, D. (2023). Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1504–1532. [Link]
- Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., & Narasimhan, K. (2023). Toxicity in ChatGPT: Analyzing Persona-Assigned Language Models. Findings of the Association for Computational Linguistics: EMNLP 2023, 1236–1270. [Link]
- Fagiolo, G., Windrum, P., & Moneta, A. (2007). Empirical Validation of Agent-Based Models: Alternatives and Prospects. Journal of Artificial Societies and Social Simulation, 10(2). [Link]
- Ge, T., Chan, X., Wang, X., Yu, D., Mi, H., & Yu, D. (2025). Scaling Synthetic Data Creation with 1,000,000,000 Personas. arXiv preprint arXiv:2406.20094. [Link]
- Groves, R. M. (2011). Three Eras of Survey Research. Public Opinion Quarterly, 75(5), 861–871. [Link]
- Gupta, S., Shrivastava, V., Deshpande, A., Kalyan, A., Clark, P., Sabharwal, A., & Khot, T. (2024). Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs. arXiv preprint arXiv:2311.04892. [Link]
- Holtdirk, T., Ahnert, G., Sakshaug, J. W., & Haensch, A.-C. (2026). In-Context Learning for the Imputation of Public Opinion Data with Large Language Models. arXiv preprint arXiv:2606.09351. [Link]
- Hu, T., & Collier, N. (2024). Quantifying the Persona Effect in LLM Simulations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10289–10307. [Link]
- Hu, Z., Lian, J., Xiao, Z., Xiong, M., Lei, Y., Wang, T., Ding, K., Xiao, Z., Yuan, N. J., & Xie, X. (2025). Population-Aligned Persona Generation for LLM-based Social Simulation. arXiv preprint arXiv:2509.10127. [Link]
- Hullman, J., Broska, D., Sun, H., & Shaw, A. (2026). This Human Study Did Not Involve Human Subjects: Validating LLM Simulations as Behavioral Evidence. arXiv preprint arXiv:2602.15785. [Link]
- Kozlowski, A. C., & Evans, J. (2025). Simulating Subjects: The Promise and Peril of Artificial Intelligence Stand-Ins for Social Agents and Interactions. Sociological Methods & Research, 54(3), 1017–1073. [Link]
- Kreutner, M., Rupprecht, J., Ahnert, G., Salem, A., & Strohmaier, M. (2026). QSTN: A Modular Framework for Robust Questionnaire Inference with Large Language Models. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 3: System Demonstrations), 537–549. [Link]
- Larooij, M., & Törnberg, P. (2025). Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations. arXiv preprint arXiv:2504.03274. [Link]
- Lundberg, I., Johnson, R., & Stewart, B. M. (2021). What Is Your Estimand? Defining the Target Quantity Connects Statistical Evidence to Theory. American Sociological Review, 86(3), 532–565. [Link]
- Lutz, M., Sen, I., Ahnert, G., Rogers, E., & Strohmaier, M. (2025). The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models. Findings of the Association for Computational Linguistics: EMNLP 2025, 23212–23237. [Link]
- Moon, S., Abdulhai, M., Kang, M., Suh, J., Soedarmadji, W., Behar, E. K., & Chan, D. M. (2024). Virtual Personas for Language Models via an Anthology of Backstories. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 19864–19897. [Link]
- Nemotron-Personas – a NVIDIA Collection. (2026). Multilingual, region-specific synthetic persona datasets. Hugging Face. [Link]
- Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. arXiv preprint arXiv:2304.03442. [Link]
- Rahimzadeh, V., Monazzah, E. M., Pilehvar, M. T., & Yaghoobzadeh, Y. (2026). Synthia: Scalable Grounded Persona Generation from Social Media Data. arXiv preprint arXiv:2507.14922. [Link]
- Ren, S., Tomlinson, B., Black, R. W., & Torrance, A. W. (2024). Reconciling the Contrasting Narratives on the Environmental Impact of Large Language Models. Scientific Reports, 14, 26310. [Link]
- Rothschild, D. M., Buskirk, T. D., Eckman, S., Hillygus, D. S., Kreuter, F., & Lazer, D. (2025). Successfully Navigating the Disruption AI will Bring to Survey Research. The Survey Statistician, 92, 30–44. [Link]
- Röttger, P., Hofmann, V., Pyatkin, V., Hinck, M., Kirk, H., Schuetze, H., & Hovy, D. (2024). Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15295–15311. [Link]
- Rupprecht, J., Froehling, L., Wagner, C., & Strohmaier, M. (2026). German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies. Proceedings of the Language Resources and Evaluation Conference (LREC), 1761–1780. [Link]
- Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? Proceedings of the 40th International Conference on Machine Learning, 29971–30004. [Link]
- Sen, I., Lutz, M., Rogers, E., Garcia, D., & Strohmaier, M. (2025). Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs. Findings of the Association for Computational Linguistics: ACL 2025, 24263–24289. [Link]
- Shu, B., Zhang, L., Choi, M., Dunagan, L., Logeswaran, L., Lee, M., Card, D., & Jurgens, D. (2024). You Don't Need a Personality Test to Know These Models Are Unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments. arXiv preprint arXiv:2311.09718. [Link]
- Troitzsch, K. G. (2014). Analysing Simulation Results Statistically: Does Significance Matter? In Interdisciplinary Applications of Agent-Based Social Simulation and Modeling (pp. 88–105). IGI Global. [Link]
- Wang, X., Hu, C., Ma, B., Röttger, P., & Plank, B. (2024). Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think. arXiv preprint arXiv:2404.08382. [Link]
- Wang, X., Ma, B., Hu, C., Weber-Genzel, L., Röttger, P., Kreuter, F., Hovy, D., & Plank, B. (2024). "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models. Findings of the Association for Computational Linguistics: ACL 2024, 7407–7416. [Link]