大语言模型时代下日益萎缩的语言多样性
The shrinking landscape of linguistic diversity in the age of LLMs

原始链接: https://www.nature.com/articles/s41562-026-02550-0

该参考资料集广泛探讨了**计算语言学、心理学与大型语言模型(LLM)**的交叉领域,主要涵盖以下四个主题: 1. **身份与健康的语言标记:** 研究表明,书面语言可以揭示个体深层次的差异,包括人格特质、道德基础和政治意识形态。此外,语言模式还是心理健康状况(如抑郁症和精神病)以及认知能力下降(如阿尔茨海默病)的诊断指标。 2. **社会语言学与文化语境:** 相关研究探讨了语言如何反映社会动态,包括性别、地域差异以及随时间推移发生的文化变迁,强调了交流模式如何深植于特定的社会和地理语境中。 3. **大型语言模型的兴起:** 文献中的很大一部分关注 ChatGPT 等工具的迅速普及。研究人员正在评估这些模型的心理“特征”、政治偏见,以及它们作为研究中人类参与者替代方案的有效性。 4. **对语言多样性的影响:** 越来越多的研究警告称,人工智能辅助写作、基于合成数据的模型训练以及标准化的对齐技术正在使人类交流趋于同质化,这可能会削弱文化细微差别,并降低创造性及学术论述中的内容多样性。

最近在 Hacker News 上的一场讨论强调了人们对大型语言模型(LLM)导致人类交流趋同化的担忧。用户们表达了一种焦虑,即由人工智能生成或编辑的文本正日趋普遍,这正在侵蚀语言多样性并助长思维的单一化。 为了对抗这种“潜移默化”的标准化,评论者们提出了几种防御策略: * **沉浸于非 AI 来源:** 通过接触历史文献或采用当代快速演变的俚语,使自己与 LLM 所偏好的“平均”模式保持距离。 * **语言破坏:** 使用刻意“破碎”的英语、不完整的警句或特定语境下扭曲的措辞。由于 LLM 依赖于字面解读和对语法的严格遵循,它们难以复制刻意表现出的不合逻辑或高度口语化的句法。 * **策略性地使用粗话:** 由于 LLM 受到严格的安全护栏限制,一些人认为,使用激烈或具有创意的粗话可以作为一种屏障,因为无论对话语境如何,这些模型都被训练去避免使用此类语言。 总体而言,社区将这些策略视为一种方式,旨在优先考虑人类的认知参与和创造力,而非 AI 系统日益推崇的、可预测且经过“净体化”的输出内容。
相关文章

原文
  • Orwell, G. Nineteen Eighty-Four 69–70 (Penguin Books, 1949).

  • Park, G. et al. Automatic personality assessment through social media language. J. Pers. Soc. Psychol. 108, 934–952 (2015).

    Article  PubMed  Google Scholar 

  • Oberlander, J. & Gill, A. J. Language with character: a stratified corpus comparison of individual differences in e-mail communication. Discourse Process. 42, 239–270 (2006).

    Article  Google Scholar 

  • Moreno, J. D., Martinez-Huertas, J. A., Olmos, R., Jorge-Botana, G. & Botella, J. Can personality traits be measured analyzing written language? A meta-analytic study on computational methods. Pers. Individ. Dif. 177, 110818 (2021).

    Article  Google Scholar 

  • Mairesse, F., Walker, M. A., Mehl, M. R. & Moore, R. K. Using linguistic cues for the automatic recognition of personality in conversation and text. J. Artif. Intell. Res. 30, 457–500 (2007).

    Article  Google Scholar 

  • Schwartz, H. A. et al. Personality, gender, and age in the language of social media: the open-vocabulary approach. PLoS ONE 8, 73791 (2013).

    Article  Google Scholar 

  • Kramsch, C. Language and culture. AILA Rev. 27, 30–55 (2014).

    Article  Google Scholar 

  • Gumperz, J. The speech community. Int. Encycl. Soc. Sci. 9, 381–386 (1968).

    Google Scholar 

  • Nguyen, D. & Rosé, C. P. Language use as a reflection of socialization in online communities. In Proc. Workshop on Language in Social Media (eds Nagarajan, M. & Gamon, M.) 76–85 (Association for Computational Linguistics, 2011).

  • Bamman, D., Eisenstein, J. & Schnoebelen, T. Gender identity and lexical variation in social media. J. Socioling. 18, 135–160 (2014).

    Article  Google Scholar 

  • Pennebaker, J. W. The secret life of pronouns. New Sci. 211, 42–45 (2011).

    Article  Google Scholar 

  • Robinson, M. D., Boyd, R. L., Fetterman, A. K. & Persich, M. R. The mind versus the body in political (and nonpolitical) discourse: linguistic evidence for an ideological signature in US politics. J. Lang. Soc. Psychol. 36, 438–461 (2017).

    Article  Google Scholar 

  • Huang, Y., Guo, D., Kasakoff, A. & Grieve, J. Understanding US regional linguistic variation with Twitter data analysis. Comput. Environ. Urban Syst. 59, 244–255 (2016).

    Article  Google Scholar 

  • Eisenstein, J., O’Connor, B., Smith, N. A. & Xing, E. A latent variable model for geographic lexical variation. In Proc. 2010 Conference on Empirical Methods in Natural Language Processing (eds Li, H. & Màrquez, L.) 1277–1287 (Association for Computational Linguistics, 2010).

  • Peterson, K., Hohensee, M. & Xia, F. Email formality in the workplace: a case study on the enron corpus. In Proc. Workshop on Language in Social Media (LSM 2011) (eds Nagarajan, M. & Gamon, M.) 86–95 (Association for Computational Linguistics, 2011).

  • Stamatatos, E. A survey of modern authorship attribution methods. J. Assoc. Inf. Sci. Technol. 60, 538–556 (2009).

    Article  Google Scholar 

  • Grieve, J. Quantitative authorship attribution: an evaluation of techniques. Lit. Ling. Comput. 22, 251–270 (2007).

    Article  Google Scholar 

  • Cassell, J. & Tversky, D. The language of online intercultural community formation. J. Comput. Mediat. Commun. 10, 1027 (2005).

    Google Scholar 

  • Danet, B. & Herring, S. C. Introduction: the multilingual internet. J. Comput. Mediat. Commun. 9, 9110 (2003).

    Google Scholar 

  • Kennedy, B. et al. Moral concerns are differentially observable in language. Cognition 212, 104696 (2021).

    Article  PubMed  Google Scholar 

  • Jackson, J. C., Gelfand, M., De, S. & Fox, A. The loosening of American culture over 200 years is associated with a creativity–order trade-off. Nat. Hum. Behav. 3, 244–250 (2019).

    Article  PubMed  Google Scholar 

  • Hofmann, V., Kalluri, P. R., Jurafsky, D. & King, S. AI generates covertly racist decisions about people based on their dialect. Nature 633, 147–154 (2024).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  • Richard, A. B., Lelandais, M., Reilly, K. T. & Jacquin-Courtois, S. Linguistic markers of subtle cognitive impairment in connected speech: a systematic review. J. Speech Lang. Hear. Res. 67, 4714–4733 (2024).

    Article  PubMed  Google Scholar 

  • Eyigoz, E., Mathur, S., Santamaria, M., Cecchi, G. & Naylor, M. Linguistic markers predict onset of Alzheimer’s disease. EClinicalMedicine 28, 100583 (2020).

  • Roark, B., Mitchell, M., Hosom, J.-P., Hollingshead, K. & Kaye, J. Spoken language derived measures for detecting mild cognitive impairment. IEEE Trans. Audio Speech Lang. Process. 19, 2081–2090 (2011).

    Article  PubMed  PubMed Central  Google Scholar 

  • Trifu, R. N. et al. Linguistic markers for major depressive disorder: a cross-sectional study using an automated procedure. Front. Psychol. 15, 1355734 (2024).

    Article  PubMed  PubMed Central  Google Scholar 

  • Weerasinghe, J., Morales, K. & Greenstadt, R. "Because… i was told… so much”: linguistic indicators of mental health status on Twitter. Proc. Priv. Enhanc. Technol. 4, 152–171 (2019).

  • Whorf, B. L. Language, Thought, and Reality: Selected Writings of Benjamin Lee Whorf (MIT Press, 2012).

  • Eckert, P. Three waves of variation study: the emergence of meaning in the study of sociolinguistic variation. Annu. Rev. Anthropol. 41, 87–100 (2012).

    Article  Google Scholar 

  • Hofstede, G. Culture’s Consequences: Comparing Values, Behaviors, Institutions and Organizations Across Nations 2nd edn (Sage, 2001).

  • OpenAI. Introducing ChatGPT https://openai.com/blog/chatgpt (2022).

  • Gemini Team et al. Gemini: a family of highly capable multimodal models. Preprint at https://doi.org/10.48550/arXiv.2312.11805 (2023).

  • Bailyn, E. ChatGPT Usage Statistics: March 2026 https://firstpagesage.com/seo-blog/chatgpt-usage-statistics/ (FirstPageSage, 2026).

  • Nearly 1 in 3 College Students Have Used ChatGPT on Written Assignments https://www.intelligent.com/nearly-1-in-3-college-students-have-used-chatgpt-on-written-assignments/ (Intelligent, 2024).

  • McClain, C. Americans’ Use of ChatGPT is Ticking Up, But Few Trust Its Election Information https://www.pewresearch.org/short-reads/2024/03/26/americans-use-of-chatgpt-is-ticking-up-but-few-trust-its-election-information/ (Pew Research Center, 2024).

  • Handa, K. et al. Which economic tasks are performed with AI? Evidence from millions of Claude conversations. Preprint at https://doi.org/10.48550/arXiv.2503.04761 (2025).

  • Mizrahi, M. et al. State of what art? A call for multi-prompt LLM evaluation. Trans. Assoc. Comput. Linguist. 12, 933–949 (2024).

    Article  Google Scholar 

  • Serapio-García, G. et al. A psychometric framework for evaluating and shaping personality traits in large language models. Nat. Mach. Intell. 7, 1954–1968 (2025).

    Article  PubMed  PubMed Central  Google Scholar 

  • Ghosh, S. et al. A closer look at the limitations of instruction tuning. In Proc. 41st International Conference on Machine Learning 624 (JMLR, 2024).

  • Santurkar, S. et al. Whose opinions do language models reflect? In International Conference on Machine Learning 29971–30004 (JMLR, 2023).

  • Ireland, M.E. & Mehl, M. R. in The Oxford Handbook of Language and Social Psychology (ed. Holtgraves, T. M.) 201–218 https://doi.org/10.1093/oxfordhb/9780199838639.013.034 (Oxford Univ. Press, 2014).

  • Corona Hernández, H. et al. Natural language processing markers for psychosis and other psychiatric disorders: emerging themes and research agenda from a cross-linguistic workshop. Schizophr. Bull. 49, 86–92 (2023).

    Article  Google Scholar 

  • Rude, S., Gortner, E.-M. & Pennebaker, J. Language use of depressed and depression-vulnerable college students. Cogn. Emot. 18, 1121–1133 (2004).

    Article  Google Scholar 

  • Coppersmith, G., Leary, R., Crutchley, P. & Fine, A. Natural language processing of social media as screening for suicide risk. Biomed. Inform. Insights 10, 1178222618792860 (2018).

    Article  PubMed  PubMed Central  Google Scholar 

  • Matz, S. C. & Netzer, O. Using big data as a window into consumers’ psychology. Curr. Opin. Behav. Sci. 18, 7–12 (2017).

    Article  Google Scholar 

  • Winter, S., Maslowska, E. & Vos, A. L. The effects of trait-based personalization in social media advertising. Comput. Hum. Behav. 114, 106525 (2021).

    Article  Google Scholar 

  • Sundar, S. S. & Marathe, S. S. Personalization versus customization: the importance of agency, privacy, and power usage. Hum. Commun. Res. 36, 298–322 (2010).

    Article  Google Scholar 

  • Ryan, M. J., Held, W. & Yang, D. Unintended impacts of LLM alignment on global representation. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics1, 16121–16140 (2024).

  • Navigli, R., Conia, S. & Ross, B. Biases in large language models: origins, inventory, and discussion. ACMJ. Data Inf. Qual. 15, 1–21 (2023).

    Article  Google Scholar 

  • Atari, M., Xue, M. J., Park, P. S., Blasi, D. & Henrich, J. Which humans? Preprint at PsyArXiv https://doi.org/10.31234/osf.io/5b26t (2023).

  • Wang, A., Morgenstern, J. & Dickerson, J. P. Large language models that replace human participants can harmfully misportray and flatten identity groups. Nat. Mach. Intell. 7, 400–411 (2025).

    Article  Google Scholar 

  • Rozado, D. The political preferences of LLMs. PLoS ONE 19, 0306621 (2024).

    Article  Google Scholar 

  • Pan, K. & Zeng, Y. Do LLMs possess a personality? Making the MBTI test an amazing evaluation for large language models. Preprint at https://doi.org/10.48550/arXiv.2307.16180 (2023).

  • Abdurahman, S. et al. Perils and opportunities in using large language models in psychological research. PNAS Nexus 3, 245 (2024).

    Article  Google Scholar 

  • Kobak, D., González-Márquez, R., Horvát, E. -Á & Lause, J. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Sci. Adv. 11, 3813 (2025).

    Article  Google Scholar 

  • Liang, W. et al. Quantifying large language model usage in scientific papers. Nat. Hum. Behav. 9, 2599–2609 (2025).

    Article  PubMed  Google Scholar 

  • Liang, W. et al. The widespread adoption of large language model-assisted writing across society. Patterns 6, 101366 (2025).

  • Bao, T., Zhao, Y., Mao, J. & Zhang, C. Examining linguistic shifts in academic writing before and after the launch of chatGPT: a study on preprint papers. Scientometrics 130, 3597–3627 (2025).

    Article  Google Scholar 

  • Guo, Y., Shang, G., Vazirgiannis, M. & Clavel, C. The curious decline of linguistic diversity: training language models on synthetic text. In Findings of the Association for Computational Linguistics: NAACL 2024 3589–3604 (ACL, 2024).

  • Muñoz-Ortiz, A., Gómez-Rodríguez, C. & Vilares, D. Contrasting linguistic patterns in human and LLM-generated news text. Artif. Intell. Rev. 57, 265 (2024).

    Article  PubMed  PubMed Central  Google Scholar 

  • Xu, W., Jojic, N., Rao, S., Brockett, C. & Dolan, B. Echoes in AI: quantifying lack of plot diversity in LLM outputs. Proc. Natl Acad. Sci. USA 122, 2504966122 (2025).

    Article  Google Scholar 

  • Padmakumar, V. & He, H. Does writing with language models reduce content diversity? In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=Feiz5HtCD0

  • Doshi, A. R. & Hauser, O. P. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Sci. Adv. 10, 5290 (2024).

    Article  Google Scholar 

  • Anderson, B. R., Shah, J. H. & Kreminski, M. Homogenization effects of large language models on human creative ideation. In Proc. 16th Conference on Creativity & Cognition 413–425 (Association for Computing Machinery, 2024).

  • Agarwal, D., Naaman, M. & Vashistha, A. AI suggestions homogenize writing toward western styles and diminish cultural nuances. In Proc. 2025 CHI Conference on Human Factors in Computing Systems 1–21 (Association for Computing Machinery, 2025).

  • Moon, K., Green, A. E. & Kushlev, K. Homogenizing effect of large language models (LLMs) on creative diversity: an empirical comparison of human and chatGPT writing. Comput. Hum. Behav. Artif. Hum. 6, 100207 (2025).

    Article  Google Scholar 

  • Alvero, A. et al. Large language models, social demography, and hegemony: comparing authorship in human and synthetic text. J. Big Data 11, 138 (2024).

    Article  Google Scholar 

  • Zhang, S., Xu, J. & Alvero, A. Generative AI meets open-ended survey responses: research participant use of AI and homogenization. Sociol. Methods Res. https://doi.org/10.1177/00491241251327130 (2025).

  • Lee, J., Alvero, A., Joachims, T. & Kizilcec, R. Poor alignment and steerability of large language models: evidence from college admission essays. In Proc. Workshop on Social Simulation with LLMs at COLM (2025).

  • Hans, A. et al. Spotting LLMs with binoculars: zero-shot detection of machine-generated text. In Proc. 41st International Conference on Machine Learning 698 (JMLR, 2024).

  • Tufts, B., Zhao, X. & Li, L. A practical examination of AI-generated text detectors for large language models. In Findings of the Association for Computational Linguistics: NAACL 2025 (eds Chiruzzo, L. et al.) 4839–4856 https://doi.org/10.18653/v1/2025.findings-naacl.271 (Association for Computational Linguistics, 2025).

  • Dugan, L. et al. Raid: a shared benchmark for robust evaluation of machine-generated text detectors. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics1, 12463–12492 (ACL, 2024).

  • Buchert, J.-M. The 6 Best AI Detectors Based on Objective Studies and Usage https://intellectualead.com/best-ai-detectors-guide/ (Intellectual Lead, 2025).

  • Singer, J. D., Willett, J. B. Applied Longitudinal Data Analysis: Modeling Change and Event Occurrence (Oxford Univ, 2003).

  • Linden, A. Conducting interrupted time-series analysis for single- and multiple-group comparisons. Stata J. 15, 480–500 (2015).

    Article  Google Scholar 

  • Bell, D., Kay, J. & Malley, J. A non-parametric approach to non-linear causality testing. Econ. Lett. 51, 7–18 (1996).

    Article  Google Scholar 

  • Dubey, A. et al. The llama 3 herd of models. Preprint at https://doi.org/10.48550/arXiv.2407.21783 (2024).

  • Kosinski, M., Stillwell, D. & Graepel, T. Private traits and attributes are predictable from digital records of human behavior. Proc. Natl Acad. Sci. USA 110, 5802–5805 (2013).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  • Tufekci, Z. Engineering the public: big data, surveillance and computational politics. First Monday https://doi.org/10.5210/fm.v19i7.4901 (2014).

  • Garg, N., Schiebinger, L., Jurafsky, D. & Zou, J. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proc. Natl Acad. Sci. USA 115, 3635–3644 (2018).

    Article  Google Scholar 

  • Kennedy, B., Ashokkumar, A., Boyd, R. L. & Dehghani, M. in Handbook of Language Analysis in Psychology (eds Dehghani, M. & Boyd, R. L.) 3–62 (Guilford Publications, 2022).

  • Hirsh, J. B. & Peterson, J. B. Personality and language use in self-narratives. J. Res. Pers. 43, 524–527 (2009).

    Article  Google Scholar 

  • Mehl, M. R., Gosling, S. D. & Pennebaker, J. W. Personality in its natural habitat: manifestations and implicit folk theories of personality in daily life. J. Pers. Soc. Psychol. 90, 862–877 (2006).

    Article  PubMed  Google Scholar 

  • Boyd, R. L., Ashokkumar, A., Seraj, S. & Pennebaker, J. W. The Development and Psychometric Properties of LIWC-22 Technical report https://www.liwc.app (Univ. Texas at Austin, 2022).

  • Frimer, J. A., Boghrati, R., Haidt, J., Graham, J. & Dehgani, M. Moral Foundations Dictionary for Linguistic Analyses 2.0 (Provalis Research, 2019).

  • Ishikawa, Y. Gender differences in vocabulary use in essay writing by university students. ProcediaSoc. Behav. Sci. 192, 593–600 (2015).

    Article  Google Scholar 

  • Chen, J., Qiu, L. & Ho, M.-H. R. A meta-analysis of linguistic markers of extraversion: positive emotion and social process words. J. Res. Pers. 89, 104035 (2020).

    Article  Google Scholar 

  • Li, L. & Tomasello, M. On the moral functions of language. Soc. Cogn. https://doi.org/10.1521/soco.2021.39.1.99 (2021).

  • Nguyen, D., Doğruöz, A. S., Rosé, C. P. & De Jong, F. Computational sociolinguistics: a survey. Comput. Linguist. 42, 537–593 (2016).

    Article  Google Scholar 

  • Tausczik, Y. R. & Pennebaker, J. W. The psychological meaning of words: LIWC and computerized text analysis methods. J. Lang. Soc. Psychol. 29, 24–54 (2010).

    Article  Google Scholar 

  • Evans, N. & Levinson, S. C. The myth of language universals: language diversity and its importance for cognitive science. Behav. Brain Sci. 32, 429–448 (2009).

    Article  PubMed  Google Scholar 

  • arXiv.org submitters. arXiv dataset. kaggle https://doi.org/10.34740/KAGGLE/DSV/7548853 (2024).

  • Verma, V., Fleisig, E., Tomlin, N. & Klein, D. Ghostbuster: detecting text ghostwritten by large language models. In Proc. 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies1 (eds Duh, K. et al.) 1702–1717 https://doi.org/10.18653/v1/2024.naacl-long.95 (Association for Computational Linguistics, 2024).

  • Adam, G. A. et al. GPTZero: robust detection of LLM-generated texts. Preprint at https://doi.org/10.48550/arXiv.2602.13042 (2026).

  • Zhan, H., He, X., Xu, Q., Wu, Y. & Stenetorp, P. G3detector: general GPT-generated text detector. Preprint at https://doi.org/10.48550/arXiv.2305.12680 (2023).

  • Zellers, R. et al. Defending against neural fake news. Adv. Neural Inf. Process. Syst. 32, 9051–9062 (2019).

    Google Scholar 

  • SIMPSON, E. H. Measurement of diversity. Nature 163, 688 (1949).

    Article  Google Scholar 

  • Shannon, C. E. A mathematical theory of communication. Bell Syst. Tech. J. 27, 379–423 (1948).

    Article  Google Scholar 

  • Gibson, E. Linguistic complexity: locality of syntactic dependencies. Cognition 68, 1–76 (1998).

    Article  CAS  PubMed  Google Scholar 

  • Johnson, W. Studies in language behavior: a program of research. Psychol. Monogr. 56, 1–15 (1944).

    Article  Google Scholar 

  • Mardaga, H. Hapax legomena: a neglected field in biblical studies. Curr. Biblic. Res. 10, 264–274 (2012).

    Article  Google Scholar 

  • Cronbach, L. J. Coefficient alpha and the internal structure of tests. Psychometrika 16, 297–334 (1951).

    Article  Google Scholar 

  • West, S. G., Finch, J. F. & Curran, P. J. in Structural Equation Modeling: Concepts, Issues, and Applications (ed. Hoyle, R. H.) 56–75 (Sage, 1995).

  • Curran, P. J., West, S. G. & Finch, J. F. The robustness of test statistics to nonnormality and specification error in confirmatory factor analysis. Psychol. Methods 1, 16–29 (1996).

    Article  Google Scholar 

  • Kilian, L. New introduction to multiple time series analysis, by Helmut Lütkepohl, Springer, 2005. Econ. Theory 22, 961–967 (2006).

    Article  Google Scholar 

  • Ivanov, V. & Kilian, L. A practitioner’s guide to lag order selection for VAR impulse response analysis. Stud. Nonlinear Dyn. Econom.9, 1 (2005).

  • Bruns, S. B. & Stern, D. I. Lag length selection and p-hacking in Granger causality testing: prevalence and performance of meta-regression models. Empir. Econ. 56, 797–830 (2019).

    Article  Google Scholar 

  • Dickey, D. A. & Fuller, W. A. Distribution of the estimators for autoregressive time series with a unit root. J. Am. Stat. Assoc. 74, 427–431 (1979).

    Google Scholar 

  • Box, G. E., Jenkins, G. M., Reinsel, G. C. & Ljung, G. M. Time Series Analysis: Forecasting and Control (John Wiley & Sons, 2015).

  • Wood, S. N. Generalized Additive Models: An Introduction with R 2nd edn https://doi.org/10.1201/9781315370279 (Chapman and Hall/CRC, 2017).

  • Hastie, T. & Tibshirani, R. Generalized additive models. Stat. Sci. 1, 297–310 (1986).

    Google Scholar 

  • Wieling, M. Analyzing dynamic phonetic data using generalized additive mixed modeling: a tutorial focusing on articulatory differences between l1 and l2 speakers of English. J. Phon. 70, 86–116 (2018).

    Article  Google Scholar 

  • Greene, R. et al. New and Improved Embedding Model https://openai.com/index/new-and-improved-embedding-model (OpenAI, 2024).

  • Conneau, A. & Kiela, D. SentEval: an evaluation toolkit for universal sentence representations. In Proc. Eleventh International Conference on Language Resources and Evaluation (LREC 2018) (eds Calzolari, N. et al.) https://aclanthology.org/L18-1269/ (European Language Resources Association (ELRA), 2018).

  • Gilardi, F., Alizadeh, M. & Kubli, M. ChatGPT outperforms crowd workers for text-annotation tasks. Proc. Natl Acad. Sci. USA 120, 2305016120 (2023).

    Article  Google Scholar 

  • Alizadeh, M. et al. Open-source LLMs for text annotation: a practical guide for model setting and fine-tuning. J. Comput. Soc. Sci. 8, 17 (2025).

    Article  PubMed  Google Scholar 

  • Gwet, K. L. Computing inter-rater reliability and its variance in the presence of high agreement. Br. J. Math. Stat. Psychol. 61, 29–48 (2008).

    Article  PubMed  Google Scholar 

  • Levene, H. in Contributions to Probability and Statistics (ed. Olkin, I.) 278–292 (Stanford Univ. Press, 1960).

  • Gentzkow, M., Shapiro, J. M. & Taddy, M. Congressional Record for the 43rd–114th Congresses: Parsed Speeches and Phrase Counts https://data.stanford.edu/congress_text (Stanford Libraries, 2018).

  • Silver, N. & Mehta, D. Both Republicans and Democrats Have an Age Problem https://web.archive.org/web/20221216043420/; https://fivethirtyeight.com/features/both-republicans-and-democrats-have-an-age-problem/ (2014).

  • Abu Raya, M. et al. The reciprocal relationship between openness and creativity: from neurobiology to multicultural environments. Front. Neurol. 14, 1235348 (2023).

    Article  PubMed  PubMed Central  Google Scholar 

  • Goldberg, L. R. in Personality and Personality Disorders 34–47 (Routledge, 2013).

  • Goldberg, L. Standard Markers of the Big-five Structure (Oregon Research Institute, 1990).

  • Pennebaker, J. W. & King, L. A. Linguistic styles: language use as an individual difference. J. Pers. Soc. Psychol. 77, 1296–1312 (1999).

    Article  CAS  PubMed  Google Scholar 

  • John, O. P., Donahue, E. M. & Kentle, R. L. The Big Five Inventory –versions 4a and 54 Technical report (Univ. California, Berkeley, Institute of Personality and Social Research, 1991).

  • Celli, F., Pianesi, F., Stillwell, D. & Kosinski, M. Workshop on computational personality recognition: shared task. Proc. Int. AAAI Conf. Web Soc. Media 7, 2–5 (2013).

    Article  Google Scholar 

  • Davis, M. H. A multidimensional approach to individual differences in empathy. JSAS Cat. Sel. Doc. Psychol. 10, 85 (1980).

    Google Scholar 

  • Omitaomu, D. et al. Empathic conversations: a multi-level dataset of contextualized conversations. Preprint at https://doi.org/10.48550/arXiv.2205.12698 (2022).

  • Barriere, V., Sedoc, J., Tafreshi, S. & Giorgi, S. Findings of WASSA 2023 shared task on empathy, emotion and personality detection in conversation and reactions to news articles. In Proc. 13th Workshop on Computational Approaches to Subjectivity, Sentiment, and Social Media Analysis 511–525 (Association for Computational Linguistics, 2023).

  • Graham, J. et al. Chapter two - moral foundations theory: the pragmatic validity of moral pluralism. Adv. Exp. Soc. Psychol. 47, 55–130 https://doi.org/10.1016/B978-0-12-407236-7.00002-4 (2013).

  • Haidt, J. & Joseph, C. Intuitive ethics: how innately prepared intuitions generate culturally variable virtues. Daedalus 133, 55–66 (2004).

    Article  Google Scholar 

  • Atari, M. et al. Morality beyond the weird: how the nomological network of morality varies across cultures. J. Pers. Soc. Psychol. 125, 1157–1188 (2023).

  • Graham, J. et al. Mapping the moral domain. J. Pers. Soc. Psychol. 101, 366–385 (2011).

    Article  PubMed  PubMed Central  Google Scholar 

  • Beltagy, I., Peters, M. E. & Cohan, A. Longformer: the long-document transformer. Preprint at https://doi.org/10.48550/arXiv.2004.05150 (2020).

  • Shatz, I. Assumption-checking rather than (just) testing: the importance of visualization and effect size in statistical diagnostics. Behav. Res. Methods 56, 826–845 (2024).

    Article  PubMed  Google Scholar 

  • Knief, U. & Forstmeier, W. Violating the normality assumption may be the lesser of two evils. Behav. Res. Methods 53, 2576–2590 (2021).

    Article  PubMed  PubMed Central  Google Scholar 

  • Bonferroni, C. Teoria statistica delle classi e calcolo delle probabilita. Pubbl. R. Inst. Super. Sci. Econ. Commerc. Firenze 8, 3–62 (1936).

    Google Scholar 

  • Havlicek, L. L. & Peterson, N. L. Robustness of the Pearson correlation against violations of assumptions. Percept. Mot. Skills 43, 1319–1334 (1976).

    Article  Google Scholar 

  • Wilcox, R. R. Modern insights about Pearson’s correlation and least squares regression. Int. J. Sel. Assess. 9, 195–205 (2001).

    Article  Google Scholar 

  • Delacre, M., Lakens, D. & Leys, C. Why psychologists should by default use Welch’s t-test instead of Student’s t-test. Int. Rev. Soc. Psychol. 30, 92–101 (2017).

    Article  Google Scholar 

  • Mohammad, S. M. & Turney, P. D. Crowdsourcing a word–emotion association lexicon. Comput. Intell. 29, 436–465 (2013).

    Article  Google Scholar 

  • Mohammad, S. & Turney, P. Emotions evoked by common words and phrases: using Mechanical Turk to create an emotion lexicon. In Proc. NAACL HLT 2010 Workshop on Computational Approaches to Analysis and Generation of Emotion in Text 26–34 https://aclanthology.org/W10-0204 (Association for Computational Linguistics, 2010).

  • Sedoc, J., Buechel, S., Nachmany, Y., Buffone, A. & Ungar, L. Learning word ratings for empathy and distress from document-level user responses. In Proc. Twelfth Language Resources and Evaluation Conference (eds Calzolari, N. et al.) 1664–1673 https://aclanthology.org/2020.lrec-1.206/ (European Language Resources Association, 2020).

  • Baumgartner, J., Zannettou, S., Keegan, B., Squire, M. & Blackburn, J. The pushshift reddit dataset. In Proc. International AAAI Conference on Web and Social Media 14, 830–839 (2020).

  • Sourati, Z. Code for "The shrinking landscape of linguistic diversity in the age of large language models". Zenodo https://doi.org/10.5281/zenodo.21633457 (2026).

  • Cohen, J. A power primer. Psychol. Bull. 112, 155–159 (1992).

    Article  CAS  PubMed  Google Scholar 

  • 联系我们 contact @ memedata.com