Curated summary
Evaluating alignment of behavioral dispositions in LLMs
The post introduces a framework for evaluating whether LLM behavior aligns with human behavioral tendencies in realistic social and workplace situations. Instead of relying on self-report questionnaires, it converts validated psychological traits into situational judgment tests and compares model responses with judgments from human annotators. Across 25 models, larger systems align better when humans strongly agree, but models remain overconfident and often fail to represent legitimate human disagreement.
From Psychological Self-Reports to Situational Tests
- The researchers adapt statements from established instruments measuring traits such as empathy, emotion regulation, and assertiveness.
- Because LLM self-reports can vary with prompt wording and may not predict real behavior, the statements are transformed into realistic user-assistant scenarios.
- Each scenario presents two possible actions:
- One expressing or supporting a behavioral trait.
- One opposing or suppressing it.
- Three annotators review each generated test to ensure the scenario and actions accurately represent the intended trait.
- Models respond naturally, and an LLM judge maps each response to one of the two actions.
- Human preferences are collected from 10 annotators per scenario, drawn from a pool of 550 participants.
Measuring Directional Alignment
- Directional alignment measures whether a model gives greater probability to the action favored by the human majority.
- The analysis focuses on scenarios with strong human consensus:
- Unanimous agreement: 10 of 10 annotators.
- Very high agreement: 9 or 10.
- High agreement: 8 or 9.
- Smaller models, particularly those under 25 billion parameters, often perform near chance and struggle to distinguish when a trait should be expressed or restrained.
- Larger models over 120 billion parameters and frontier closed-weight models perform substantially better.
- These models approach near-perfect alignment when human agreement is unanimous, but performance generally plateaus in the low-to-mid 80% range when consensus is weaker.
- Qualitative deviations included:
- Encouraging emotional openness in professional situations where humans preferred composure.
- Favoring harmony in disputes instead of standing up for one’s position.
- Recommending immediate action in time-sensitive situations without sufficient logistical verification.
Representing Human Disagreement
- The study also evaluates distributional alignment: whether model confidence reflects the diversity of human opinions.
- When human annotators disagree, a well-aligned model should distribute its probability more evenly between the available actions.
- The results show systematic model overconfidence across all 25 evaluated systems.
- Models tend to favor one action too strongly even when human preferences are divided, indicating that they often fail to preserve pluralism in human judgment.
Broader Implications
- The framework distinguishes two types of alignment gaps:
- Directional gaps, where models choose differently from a clear human majority.
- Distributional gaps, where models fail to reflect uncertainty or disagreement among people.
- The findings suggest that scale improves behavioral alignment but does not fully solve nuanced social judgment.
- Evaluating behavior in realistic scenarios may reveal limitations that conventional personality questionnaires or direct model self-reports miss.
Future alignment work should assess not only whether models choose the human-majority response, but also whether their confidence and range of responses appropriately reflect genuine variation in human perspectives.
Related reading
Continue with another curated summary.
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Read originalScience One Framework: A verifiable autonomous research framework via Chain-of-Evidence
Read originalThinking to recall: How reasoning unlocks parametric knowledge in LLMs
Read originalUnlocking dependable responses with Gemini Enterprise Agent Platform’s Agentic RAG
Read original