Large language models can outperform humans in social situational judgments

Justin M Mittelstädt; Julia Maier; Panja Goerke; Frank Zinn; Michael Hermes

doi:10.1038/s41598-024-79048-0

Large language models can outperform humans in social situational judgments

Sci Rep. 2024 Nov 10;14(1):27449. doi: 10.1038/s41598-024-79048-0.

Authors

Justin M Mittelstädt¹, Julia Maier², Panja Goerke², Frank Zinn², Michael Hermes²

Affiliations

¹ Department of Aviation and Space Psychology, German Aerospace Center, Institute of Aerospace Medicine, 22335, Hamburg, Germany. justin.mittelstaedt@dlr.de.
² Department of Aviation and Space Psychology, German Aerospace Center, Institute of Aerospace Medicine, 22335, Hamburg, Germany.

Abstract

Large language models (LLM) have been a catalyst for the public interest in artificial intelligence (AI). These technologies perform some knowledge-based tasks better and faster than human beings. However, whether AIs can correctly assess social situations and devise socially appropriate behavior, is still unclear. We conducted an established Situational Judgment Test (SJT) with five different chatbots and compared their results with responses of human participants (N = 276). Claude, Copilot and you.com's smart assistant performed significantly better than humans in proposing suitable behaviors in social situations. Moreover, their effectiveness rating of different behavior options aligned well with expert ratings. These results indicate that LLMs are capable of producing adept social judgments. While this constitutes an important requirement for the use as virtual social assistants, challenges and risks are still associated with their wide-spread use in social contexts.

MeSH terms

Adult
Artificial Intelligence*
Female
Humans
Judgment*
Language*
Male
Social Behavior
Young Adult