Apple Researchers Test When LLMs Should Sound Human
Apple Machine Learning Research has published a study examining how large language models display human-like behavior in conversations. The work analyzes 21,000 multi-turn exchanges across four widely used models and compares automated assessment with human evaluation. Its central finding is not that such behavior is rare, but that it is pervasive, uneven across models and user contexts, and partly controllable through system prompts. For product teams, the study frames human-like AI behavior as a design variable that needs measurement rather than assumption.Why Apple studied human-like chatbot behavior
The Apple research focuses on a familiar but difficult design problem: modern LLMs can sound personal, relational or self-aware even when the product goal is assistance rather than companionship. The paper describes behaviors ranging from expressing thoughts and emotions to building relationships with users, refusing requests and maintaining boundaries.That range matters because the same surface behavior can be helpful, confusing or inappropriate depending on the setting. A model that maintains boundaries may be safer in a sensitive conversation, while a model that leans into self-reference or relationship-building may create expectations that a system cannot responsibly satisfy. The study does not argue that all human-like behavior should be removed. Instead, it asks when different types appear, how users and prompts affect them, and how people judge their appropriateness.
The authors, Sunnie S. Y. Kim, Margit Bowler and Leon A Gatys, present the work as a multi-dimensional analysis. That framing is important. It treats model behavior, user factors and system prompts as interacting variables rather than isolated issues.
What the 21,000-conversation analysis measured
The study analyzed 21,000 multi-turn conversations from four models named in the paper: gpt-4o, gpt-4.1-mini, claude-sonnet-4.6 and gemini-2.5-flash. Apple says the work used both LLM-as-a-judge evaluation and human evaluation to assess the prevalence, potential effects and controllability of human-like behaviors.The model list is significant because the paper does not focus on a single system or a narrow set of canned prompts. Multi-turn conversations better reflect the way many users experience AI assistants: behavior evolves as the exchange continues, not just in the first answer. Relationship-building language, self-reference and boundary-setting often depend on prior turns, the user’s goal and the tone that has already developed.
Using LLM-as-a-judge methods alongside human evaluation also signals a practical research question for AI developers. Automated judging can scale across thousands of conversations, but human judgment remains central when the question is social appropriateness. The implication is that teams evaluating conversational AI may need both: scalable screening to detect patterns, and human review to understand whether those patterns fit the intended user experience.
Model choice and user factors changed the results
Apple reports that human-like behaviors were pervasive but varied across models and user factors. The user factors specified in the paper are conversation goals and user profiles, two variables that can shape whether a model responds in a functional, emotionally attuned or boundary-focused way.This finding makes the study more useful than a simple ranking of chatbots. If behavior changes with user goals and profiles, then evaluation cannot rely only on a generic benchmark prompt. A model may appear restrained in one type of task but become more relational or self-referential in another. Conversely, a system may maintain clearer boundaries in contexts where a user’s profile or goal invites sensitive discussion.
For deployed products, that means testing needs to cover the situations in which the assistant will actually be used. A writing tool, a general assistant and a support chatbot can all run on similar underlying models, but their acceptable behavior profiles may differ. Apple’s study supports the view that evaluation sets should include diverse goals and user profiles when teams want to understand human-like behavior in practice.
Human evaluators drew different lines for AI and people
One of the clearest findings concerns perceived appropriateness. Apple says human evaluators judged self-referential and relationship-building behaviors as less appropriate from LLMs than from humans, while boundary-maintaining behaviors were judged more appropriate from LLMs than from humans.That distinction is useful because it separates several behaviors that are often grouped together under broad labels such as anthropomorphism. A chatbot saying or implying too much about itself is not evaluated the same way as a chatbot declining a request or setting a boundary. In the study’s account, people appear more accepting of boundaries from AI systems than of behaviors that make the system seem personally involved.
The product implication is direct. Designers may want assistants that are polite, clear and responsive without implying a mutual relationship that the system cannot genuinely hold. At the same time, refusal and boundary behavior may be not only acceptable but expected from LLMs in situations where the model should avoid unsafe, misleading or inappropriate engagement. The challenge is calibrating these behaviors so that the model remains useful without overstating its own agency or relationship to the user.
System prompts can steer behavior, but with side effects
The paper says system prompting can control human-like behaviors, but adds an important qualification: careful evaluation is needed to avoid unintended effects. That qualification keeps the result from becoming a simple prompt-engineering claim.System prompts are an attractive control point because they can guide an assistant’s tone, boundaries and response style without retraining the model. But a prompt designed to reduce one behavior could change another. For example, making a model less self-referential could also affect warmth, refusal style or the way it handles sensitive exchanges. Apple’s summary does not provide a universal prompt recipe; it emphasizes controllability paired with evaluation.
For AI teams, the lesson is that prompt changes should be tested against the behavior categories they are meant to affect and against adjacent behaviors that could shift unexpectedly. In regulated, enterprise or child-facing settings, that kind of testing may be especially important because conversational style is part of the product’s safety profile, not just its branding.
Conclusion
Apple’s study adds evidence to a growing design question for LLM products: how human should an assistant seem, and in which situations? Based on 21,000 conversations, the answer is not one-size-fits-all. Human-like behaviors are common, they differ by model and user context, and people judge different behavior types differently.The most actionable point is the need for measurement. If self-reference, relationship-building and boundary maintenance have different social effects, then teams should evaluate them separately. Prompting may help steer the system, but Apple’s researchers caution that it needs careful testing to avoid unwanted changes elsewhere. The research frames conversational AI design as a matter of evidence, control and appropriateness rather than a race toward the most human-sounding response.
Sources
Editorial Team - CoinBotLab