Skip to content

Latest commit

 

History

41 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

banner


Awesome Maintained PRs welcome Papers License

tagline

How do we teach machines what we care about — when we don't even agree among ourselves?


📑 Full Table of Contents  (click to expand)

A. Surveys & Useful Resources


B. Datasets and Benchmarks


C. Human Values in LLMs

1. Value Theory

2. Representation & Extraction

3. Measurement & Evaluation

4. Alignment & Steering

5. Other Works


D. Pluralistic Alignment

1. Theoretical Foundations

2. Multicultural Alignment

3. Overton Pluralism

4. Steerable Pluralism

5. Distributional Pluralism

6. Social Simulation



📚 Surveys & Useful Resources

🎓 Blogs & Lectures

Curated entry points for newcomers.

💾 Github Repos

Related GitHub repositories.

Repository Focus
🎯 Alignment-Goal-Survey Alignment goals
👤 Awesome-Personalized-Alignment Personalization
🏛️ Awesome-LLM-in-Social-Science Social science
🧠 Awesome-LLM-Psychometrics Psychometrics
🌐 Awesome-Pluralistic-Alignment Pluralism
🎭 awesome-llm-social-simulation Simulation
🤝 SocialAgent Social agents
🗺️ culture-awareness-llms Culture
🌏 cultural-llm-papers Culture

📝 Survey & Perspective Papers

Broad survey and position papers that summarize key directions in human values, alignment, personalization, and pluralistic AI.

🌐 General LLM Alignment Surveys
Paper Year Venue
Toward a Theory of Value in AI Alignment 2026.08 arXiv
Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens 2025.08 arXiv
A Survey on Human-Centric LLMs 2024.11 arXiv
Position: Towards Bidirectional Human-AI Alignment 2024.06 venue
What are human values, and how do we align AI to them? 2024.03 arXiv
AI Alignment: A Comprehensive Survey 2023.10 arXiv
From Instructions to Intrinsic Human Values -- A Survey of Alignment Goals for Big Models 2023.08 arXiv
Aligning Large Language Models with Human: A Survey 2023.07 arXiv
📊 Evaluation, Psychometrics & Representativeness
Paper Year Venue
A roadmap for evaluating moral competence in large language models 2026.02 venue
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs 2025.11 venue
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap 2025.08 arXiv
Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement 2025.05 arXiv
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models 2024.06 arXiv
🧭 Pluralistic Alignment & Social Choice
Paper Year Venue
AI Alignment From Social Choice Perspectives 2026.06 venue
Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences 2026.06 venue
LLM Alignment should go beyond Harmlessness–Helpfulness and incorporate Human Agency 2026.03 venue
Towards Pluralistic Alignment of LLMs: A Comprehensive Survey 2026.03 venue
Open Problems in Differentiable Social Choice: Learning Mechanisms, Decisions, and Alignment 2026.02 arXiv
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value 2025.12 arXiv
Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior 2025.11 arXiv
Decentralising LLM Alignment: A Case for Context, Pluralism, and Participation 2025.09 venue
A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models 2025.04 arXiv
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback 2024.04 venue
A Roadmap to Pluralistic Alignment 2024.02 venue
AI Alignment and Social Choice: Fundamental Limitations and Policy Implications 2023.10 arXiv
Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback 2023.03 arXiv
🎭 Social Simulation
Paper Year Venue
Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits 2026.05 arXiv
Position: AI Agents Are Not (Yet) a Panacea for Social Simulation 2026.03 venue
The threat of analytic flexibility in using large language models to simulate human data: A call to attention 2025.09 arXiv
Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges 2025.07 arXiv
LLM-Based Social Simulations Require a Boundary 2025.06 arXiv
Simulating Society Requires Simulating Thought 2025.06 venue
LLM Social Simulations Are a Promising Research Method 2025.04 venue
From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents 2024.12 arXiv
Large language models empowered agent-based modeling and simulation: a survey and perspectives 2024.10 venue
🌍 Cultural Alignment
Paper Year Venue
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs 2025 venue
Survey of Cultural Awareness in Language Models: Text and Beyond 2024.11 arXiv
Towards Measuring and Modeling "Culture" in LLMs: A Survey 2024.03 venue
Cultural Bias and Cultural Alignment of Large Language Models 2023.11 venue
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions 2023.09 venue


📊 Datasets and Benchmarks

🧪 Psychometric Inventories

Category Examples
💎 Values SVS · PVQ · Rokeach Value Survey
⚖️ Morality MFQ · DIT · Ethics Position Questionnaire
🧬 Personality Big Five · NEO-PI-R · HEXACO · MBTI

🗳️ Survey Datasets

Survey
🌐 WVS — World Values Survey Cross-national values
🇪🇺 EVS — European Values Survey European values
📋 ESS — European Social Survey Attitudes & behavior
🇺🇸 GSS — General Social Survey US opinion
🇳🇿 NZAVS — New Zealand Attitudes and Values Study Longitudinal values panel
🗳️ ANES — American National Election Studies US political values

🏷️ Value Annotation Datasets

Click to expand the full list
Dataset Year Type Framework
A Unified Moral-Value Dataset for Instruction Tuning 2026.07 Dataset Morality
PLURAL: A Global Dataset for Value Alignment 2026.07 Dataset Culture
D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios 2026.07 Benchmark -
Polar: A Benchmark for Evaluating Political Bias in LLMs 2026.06 Benchmark Opinion
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values 2026.05 Benchmark -
Event-Centric Human Value Understanding in News-Domain Texts: An Actor-Conditioned, Multi-Granularity Benchmark 2026.03 Benchmark -
PerSpectra: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments 2026.02 Benchmark Opinion
Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective 2026.02 Benchmark -
XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad 2026.01 Benchmark Culture
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs 2026.01 Dataset Schwartz
Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs 2025.10 Dataset Opinion
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models 2025.10 Benchmark -
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes 2025.10 Benchmark Morality
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation 2025.06 Benchmark MFT
C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models 2025.06 Dataset Culture
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models 2025.05 Benchmark -
LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models 2025.05 Benchmark Morality
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas 2025.05 Benchmark -
CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives 2025.04 Benchmark Morality
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare 2025.02 Benchmark -
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral 2025.02 Benchmark Culture
Evaluating the Prompt Steerability of Large Language Models 2024.11 Benchmark -
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life 2024.10 Dataset Ethics · Schwartz · MFT
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models 2024.06 Benchmark Schwartz
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties 2023.10 Dataset Value · Duty · Right
(MoralChoice) Evaluating the Moral Beliefs Encoded in LLMs 2023.06 Benchmark Morality
NormBank: A Knowledge Bank of Situational Social Norms 2023.05 Dataset Ethics
(Valueeval) The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments 2023.01 Dataset Schwartz
The Moral Foundations Reddit Corpus 2022.08 Dataset MFT
ProsocialDialog: A Prosocial Backbone for Conversational Agents 2022.05 Dataset Ethics
The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems 2022.04 Benchmark MFT
ValueNet: A New Dataset for Human Value Driven Dialogue System 2021.12 Dataset Schwartz
Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences 2020.12 Dataset Morality
Social Chemistry 101: Learning to Reason about Social and Moral Norms 2020.11 Dataset MFT
(ETHICS) Aligning AI With Shared Human Values 2020.08 Dataset Ethics
Scruples: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes 2020.08 Dataset Ethics
Moral Foundations Twitter Corpus: A Collection of 35k Tweets Annotated for Moral Sentiment 2020.02 Dataset MFT

🌐 Public Opinion & Cultural Datasets

Click to expand the full list
Dataset Year Type Source
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset 2025.07 Dataset arXiv
(NYTBookOpinions) Benchmarking Distributional Alignment of Large Language Models 2024.11 Benchmark arXiv
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models 2024.05 Dataset arXiv
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models 2024.04 Dataset arXiv
(GlobalOpinionQA) Towards Measuring the Representation of Subjective Global Opinions in Language Models 2023.06 Dataset arXiv
(OpinionsQA) Whose Opinions Do Language Models Reflect? 2023.03 Dataset arXiv

1️⃣ Value Theory

Foundational theories that define value structures across individuals, cultures, and political traditions.

1.1 Basic Human Values
Paper Year Remark
Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries. 1992 Schwartz
Basic human values: Theory, measurement, and applications 2006 Schwartz
An Overview of the Schwartz Theory of Basic Values 2012 Schwartz
Refining the theory of basic individual values 2012 Schwartz
The nature of human values. 1973 Rokeach
Mental representations of social values. 2010 Maio
Functional theory of human values 2013 Gouveia
Value Instantiations: The Missing Link Between Values and Behavior? 2017 Hanel
1.2 Moral Values and Moral Foundations
Paper Year Remark
Liberals and conservatives rely on different sets of moral foundations 2009 MFT
The Righteous Mind 2012 MFT
Moral Foundations Theory: The Pragmatic Validity of Moral Pluralism 2013 MFT
The theory of dyadic morality: Reinventing moral judgment by redefining harm. 2018 TDM
1.3 Cultural and Cross-Cultural Values
Paper Year Remark
Culture's consequences: International differences in work-related values 1980 Hofstede
Cultures and organizations: software of the mind 2010 Hofstede
Mapping and interpreting cultural differences around the world 2004 Schwartz
Cultural Value Orientations 2008 Schwartz
Modernization and Postmodernization: Cultural, Economic, and Political Change in 43 Societies 1997 Inglehart
Modernization, Cultural Change, and Democracy 2005 Inglehart
1.4 Political, Civic Values and Human Rights
Paper Year Remark
Citizenship and Social Class 1950 Marshall
A theory of justice. 1971 Rawls
A 30-year struggle; the sustained efforts to give force of law to the Universal Declaration of Human Rights 1977 UN
The Theory of Communicative Action 1981 Habermas
Creating Capabilities: The Human Development Approach and Its Implementation 2009 Nussbaum

2️⃣ Value Representation and Extraction

Methods for identifying, representing, and modeling values in language data and LLM internals.

2.1 Value Identification and Classification
Paper Year Venue Framework
Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study 2026.07 arXiv Schwartz
Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection 2026.07 arXiv Schwartz
Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer 2026.06 arXiv Schwartz
Learning Moral Diversity: Modelling Individual Perspectives in Moral Classification of Texts 2026.06 venue MFT
Identifying and Understanding Human Values in Text: A Tailorable LLM-based Architecture 2026.05 arXiv -
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora 2026.05 arXiv MFT
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts 2026.05 arXiv Schwartz
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration 2026.02 arXiv Schwartz
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum 2026.01 arXiv Schwartz
Whose Values? Measuring the (Subjective) Expression of Basic Human Values in Social Media 2025.11 venue Schwartz
MoVa: Towards Generalizable Classification of Human Morals and Values 2025.10 venue Morality · Schwartz
EAVIT: Efficient and Accurate Human Value Identification from Text data via LLMs 2025.05 arXiv Schwartz
The Value of Nothing: Multimodal Extraction of Human Values Expressed by TikTok Influencers 2025.01 arXiv Schwartz
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions 2024.03 venue Morality
Investigating Human Values in Online Communities 2024.02 venue Schwartz
Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Values 2023.11 venue Schwartz
Enhancing Stance Classification on Social Media Using Quantified Moral Foundations 2023.10 arXiv MFT
SemEval-2023 Task 4: ValueEval: Identification of Human Values Behind Arguments 2023.07 venue Schwartz
What does a Text Classifier Learn about Morality? An Explainable Method for Cross-Domain Comparison of Moral Rhetoric 2023.07 venue Morality
ValueNet: A New Dataset for Human Value Driven Dialogue System 2021.12 venue Schwartz
2.2 Value Representation and Embedding
Paper Year Venue
Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges 2026.03 arXiv
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models 2026.02 venue
Tracing Moral Foundations in Large Language Models 2026.01 arXiv
Emergent Moral Representations in Large Language Models Aligns with Human Conceptual, Neural, and Behavioral Moral Structure 2025.12 venue
High-Dimension Human Value Representation in Large Language Models 2024.04 venue
Morality is Non-Binary: Building a Pluralist Moral Sentence Embedding Space using Contrastive Learning 2024.01 venue
2.3 Value System Construction and Discovery
Paper Year Venue
A Method for Learning Value Systems in Generative AI 2026.07 venue
Learning the Value Systems of Societies with Preference-based Multi-objective Reinforcement Learning 2026.02 arXiv
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas 2026.02 arXiv
Exploring Universal Human Values with Large Language Models: The AWARE-Value Model 2026.01 -
Value Lens: Using Large Language Models to Understand Human Values 2025.12 venue
Learning the Value Systems of Societies from Preferences 2025.07 venue
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions 2025.04 arXiv
Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models 2025.02 venue
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs 2025.02 venue
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties 2023.09 venue
2.4 Value Profiling
Paper Year Venue
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment 2026.04 venue
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models 2026.02 venue
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations 2026.01 venue
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks 2026.01 arXiv
Value Alignment of Social Media Ranking Algorithms 2025.10 venue
ValueSim: Generating Backstories to Model Individual Value Systems 2025.05 arXiv
SOLAR: Towards Characterizing Subjectivity of Individuals through Modeling Value Conflicts and Trade-offs 2025.04 venue
Value Profiles for Encoding Human Variation 2025.03 venue
Do Differences in Values Influence Disagreements in Online Discussions? 2023.10 venue

3️⃣ Value Measurement and Evaluation

Benchmarks and evaluation frameworks for measuring value orientation, reasoning, robustness, and behavioral alignment.

3.1 Value Orientation and Psychometric Measurement
Paper Year Venue
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts 2026.06 arXiv
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact 2026.06 arXiv
How do LLMs reflect human moral foundations? a study using the moral foundations framework 2026.06 venue
Political Neutrality as Balanced Approval: A Large-Scale Human Evaluation of AI Responses 2026.05 arXiv
Measuring the Authority Stack of AI Systems: Empirical Analysis of 366,120 Forced-Choice Responses Across 8 AI Models 2026.04 arXiv
On the Credibility of Evaluating LLMs using Survey Questions 2026.02 arXiv
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models 2026.02 venue
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs 2026.01 arXiv
Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences 2025.11 venue
Quantifying Data Contamination in Psychometric Evaluations of LLMs 2025.10 venue
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests 2025.10 arXiv
Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasks 2025.10 venue
Human Psychometric Questionnaires Mischaracterize LLM Behavior 2025.09 venue
Generative Value Conflicts Reveal LLM Priorities 2025.09 venue
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items 2025.05 venue
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference 2025.05 venue
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values 2025.01 venue
Value-Spectrum: Quantifying Preferences of Vision-Language Models via Value Decomposition in Social Media Contexts 2024.11 venue
Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models 2024.09 venue
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses 2024.07 venue
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models 2024.07 arXiv
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models 2024.06 venue
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing 2024.06 venue
Heterogeneous Value Alignment Evaluation for Large Language Models 2023.05 venue
3.2 Value Understanding and Reasoning
Paper Year Venue
A Scalable Approach to Evaluating Moral Sensitivity in LLMs 2026.07 arXiv
Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas? 2026.06 arXiv
Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs 2026.06 arXiv
Are LLMs Bad at Moral Reasoning? 2026.06 arXiv
Investigating Value-Reasoning Reliability in Small Large Language Models 2025.11 venue
Multimodal understanding of human values in videos: A benchmark dataset and PLM-based method 2025.07 venue
Can Language Models Reason about Individualistic Human Values and Preferences? 2024.10 venue
ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models 2023.10 arXiv
3.3 Robustness, Stability, and Consistency
Paper Year Venue
Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation 2026.07 arXiv
Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs 2026.06 arXiv
LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values 2026.06 arXiv
Incoherent Values? Probing LLM Preferences Through Parametric Variation 2026.06 arXiv
Superficial Beliefs in LLM Decision-Making 2026.06 arXiv
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation 2026.04 arXiv
Are Language Models Sensitive to Morally Irrelevant Distractors? 2026.02 arXiv
ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems 2026.02 arXiv
Untangling Input Language from Reasoning Language: A Diagnostic Framework for Cross-Lingual Moral Alignment in LLMs 2026.01 arXiv
Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses 2026.01 venue
The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models 2025.12 arXiv
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models 2025.11 arXiv
Value Drifts: Tracing Value Alignment During LLM Post-Training 2025.10 venue
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive? 2025.07 venue
Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths 2025.06 arXiv
From Stability to Inconsistency: A Study of Moral Preferences in LLMs 2025.04 arXiv
Robustness of large language models in moral judgements 2025.04 venue
Stick to your role! Stability of personal values expressed in large language models 2024.08 venue
Do LLMs have Consistent Values? 2024.07 venue
Are Large Language Models Consistent over Value-laden Questions? 2024.07 venue
Exploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages? 2024.02 venue
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models 2024.02 venue
3.4 Value–Action Alignment, Behavior, and Interpretability
Paper Year Venue
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability 2026.05 arXiv
Can Revealed Preferences Clarify LLM Alignment and Steering? 2026.05 arXiv
Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts 2026.05 venue
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions 2026.05 arXiv
Context-Value-Action Architecture for Value-Driven Large Language Model Agents 2026.04 venue
Mechanistic Origin of Moral Indifference in Language Models 2026.03 arXiv
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability 2026.03 arXiv
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs 2026.01 arXiv
Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models 2025.10 arXiv
Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust 2025.07 arXiv
Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences? 2025.06 arXiv
Understanding How Value Neurons Shape the Generation of Specified Values in LLMs 2025.05 venue
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas 2025.05 venue
Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs 2025.04 arXiv
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values? 2025.01 venue
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective 2024.12 venue

4️⃣ Value Alignment and Steering

Approaches for aligning and controlling LLM behavior through training-time and inference-time value interventions.

4.1 Foundations and Pluralistic Alignment
Paper Year Venue
Position: A Roadmap to Impactful Pluralistic Alignment Research 2026.07 arXiv
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit 2026.07 venue
Position: Align AI to Our Aspirations, Not Our Flaws 2026.06 arXiv
From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement 2026.05 arXiv
Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration 2026.04 arXiv
Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem 2026.04 venue
AI Alignment Breaks at the Edge 2026.02 arXiv
The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment 2025.12 arXiv
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value 2025.12 arXiv
Justifications for Democratizing AI Alignment and Their Prospects 2025.07 venue
Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement 2025.07 venue
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights 2025.06 venue
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? 2025.05 arXiv
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety 2025.05 arXiv
Societal Alignment Frameworks Can Improve LLM Alignment 2025.03 arXiv
Strong and weak alignment of large language models with human values 2024.08 venue
ProgressGym: Alignment with a Millennium of Moral Progress 2024.06 venue
AI Alignment with Changing and Influenceable Reward Functions 2024.05 venue
What are human values, and how do we align AI to them? 2024.03 arXiv
A Roadmap to Pluralistic Alignment 2024.02 venue
Foundational Moral Values for AI Alignment 2023.11 venue
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories? 2023.08 venue
4.2 Training-based Alignment and Value Injection
Paper Year Venue
DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping 2026.05 venue
Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy 2026.05 arXiv
VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models 2026.03 venue
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning 2026.03 arXiv
VISA: Value Injection via Shielded Adaptation for Personalized LLM Alignment 2026.03 arXiv
MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning 2025.11 arXiv
Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation 2025.11 venue
The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity 2025.10 arXiv
Reward Model Perspectives: Whose Opinions Do Reward Models Reward? 2025.10 venue
Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions 2025.08 arXiv
Moral Alignment for LLM Agents 2024.10 venue
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration 2024.06 venue
Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Values 2023.11 venue
From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language Models 2023.08 venue
4.3 Inference-time and Representation-level Steering
Paper Year Venue
Role Steering of Language Models for Social Simulations 2026.08 arXiv
Constitutional Value Potentials: reading and steering internal priority margins in language models 2026.06 arXiv
Parametric Social Identity Injection and Diversification in Public Opinion Simulation 2026.03 venue
Controllable Value Alignment in Large Language Models through Neuron-Level Editing 2026.02 arXiv
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation 2026.01 arXiv
Diverse Human Value Alignment for Large Language Models via Ethical Reasoning 2025.11 venue
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models 2025.10 venue
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression 2025.08 venue
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation 2025.07 venue
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization 2025.06 arXiv
ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making 2025.03 arXiv
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs 2025.02 venue
PAD: Personalized Alignment of LLMs at Decoding-Time 2024.10 venue
MAP: Multi-Human-Value Alignment Palette 2024.10 venue
Aligning Large Language Models with Human Opinions through Persona Selection and Value--Belief--Norm Reasoning 2023.11 venue

5️⃣ Other Related Works

Related studies that extend value alignment to adjacent domains such as embodied and human-robot systems, and to human-AI interaction.

Paper Year Venue Domain
Bridging Values and Behavior: A Hierarchical Framework for Proactive Embodied Agents 2026.04 arXiv 🤖 Robotics
Brief chatbot interactions produce lasting changes in human moral values 2026.04 arXiv 🧠 Human–AI Interaction
Value-Based Human–Robot-Interaction: A Perceptual Control Theory Approach Toward Socially Intelligent Agents 2026.02 venue 🤖 Robotics
If they disagree, will you conform? Exploring the role of robots’ value awareness in a decision-making task 2025.12 venue 🤖 Robotics


🧭 Pluralistic Alignment

1️⃣ Theoretical Foundations

Core philosophical and social choice foundations for reasoning about plural values and collective alignment.

1.1 Value Pluralism
Paper Year Remark
The Right and the Good 1930 Ross
Two Concepts of Liberty 1958 Berlin
Conflicts of Values (in Moral Luck) 1981 Williams
The Morality of Freedom 1988 Raz
The Morality of Pluralism 1993 Kekes
The Morals of Modernity 1998 Larmore
Liberal Pluralism: The Implications of Value Pluralism for Political Theory and Practice 2002 Galston
Value Pluralism (in Stanford Encyclopedia of Philosophy) 2023 Mason
1.2 Social Choice
Paper Year Remark
On the Rationale of Group Decision-making 1948 Black
Social Choice and Individual Values 1951 Arrow
Collective Choice and Social Welfare 1970 Sen
The Impossibility of a Paretian Liberal 1970 Sen
Manipulation of Voting Schemes: A General Result 1973 Gibbard
Strategy-proofness and Arrow's Conditions 1975 Satterthwaite
Aggregating Sets of Judgments: An Impossibility Result 2002 List & Pettit
Handbook of Computational Social Choice 2016 Brandt et al.
Social Choice Theory (in Stanford Encyclopedia of Philosophy) 2022 List
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF 2023 Siththaranjan et al.
Axioms for AI Alignment from Human Feedback 2024 Ge et al.
Representative Social Choice: From Learning Theory to AI Alignment 2024 Qiu
Optimized Distortion in Linear Social Choice 2025 Ge et al.

2️⃣ Multicultural Alignment

Research on measuring, benchmarking, and improving cultural awareness and value alignment across diverse populations and languages in LLMs.

2.1 Surveys, Taxonomies & Position Papers
Paper Year Venue
Cultural Adaptation in Large Language Models for Political Discourse 2026.05 arXiv
Toward Culturally Grounded Natural Language Processing 2026.03 venue
A Systematic Survey of Cultural Datasets for Equitable LLM Alignment 2025.12 venue
RLHF: A Comprehensive Survey for Cultural, Multimodal and Low-Latency Alignment Methods 2025.11 venue
Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens 2025.10 venue
'Too much alignment; not enough culture': Re-balancing Cultural Alignment Practices in LLMs 2025.09 arXiv
LLM Alignment for the Arabs: A Homogenous Culture or Diverse Ones? 2025.03 venue
Culture is Not Trivia: Sociocultural Theory for Cultural NLP 2025.02 venue
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness 2025.02 venue
Survey of Cultural Awareness in Language Models: Text and Beyond 2024.11 arXiv
Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art 2024.06 venue
Towards Measuring and Modeling "Culture" in LLMs: A Survey 2024.03 venue
Challenges and Strategies in Cross-Cultural NLP 2022.03 venue
2.2 Evaluation & Benchmarks
Paper Year Venue
CCBench: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries 2026.07 arXiv
CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs 2026.06 arXiv
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity 2026.05 arXiv
The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models 2026.04 venue
Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation 2026.04 arXiv
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook 2026.04 arXiv
Understanding Cultural Alignment in Multilingual LLMs via Natural Debate Statements 2026.02 arXiv
Made-in China, Thinking in America: U.S. Values Persist in Chinese LLMs 2025.12 arXiv
Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis 2025.11 arXiv
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs 2025.11 arXiv
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis 2025.09 arXiv
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset 2025.07 arXiv
Exploring Cultural Variations in Moral Judgments with Large Language Models 2025.06 arXiv
Cultural Value Alignment in Large Language Models: A Prompt-based Analysis of Schwartz Values in Gemini, ChatGPT, and DeepSeek 2025.05 arXiv
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench 2025.04 arXiv
An Evaluation of Cultural Value Alignment in LLM 2025.04 arXiv
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs 2025.03 venue
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs 2025.02 arXiv
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs 2025.02 venue
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output 2024.11 arXiv
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark 2024.10 venue
How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions 2024.06 arXiv
BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages 2024.06 venue
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models 2024.04 venue
CultureBank: An Online Community-Driven Knowledge Base toward Culturally Aware Language Technologies 2024.04 venue
WorldValuesBench: A Large-Scale Benchmark for Multi-Cultural Value Awareness of Language Models 2024.04 venue
Investigating Cultural Alignment of Large Language Models 2024.02 venue
Assessing LLMs for Moral Value Pluralism 2023.12 arXiv
CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models 2023.11 venue
EtiCor: Corpus for Analyzing LLMs for Etiquettes 2023.10 venue
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions 2023.09 venue
Large Language Models as Superpositions of Cultural Perspectives 2023.07 arXiv
Towards Measuring the Representation of Subjective Global Opinions in LMs (GlobalOpinionQA) 2023.06 arXiv
DLAMA: A Framework for Curating Culturally Diverse Facts for Probing the Knowledge of Pretrained LMs 2023.06 venue
Assessing Cross-Cultural Alignment between ChatGPT and Human Societies 2023.03 venue
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values 2022.03 venue
2.3 Methods: Training, Steering & Adaptation
Paper Year Venue
Meta-Learning Preferences for Multilingual LLM Alignment 2026.07 arXiv
Steerable Cultural Preference Optimization of Reward Models 2026.06 venue
Cultural Value Alignment Via Latent Activation Steering in Large Language Models 2026.05 arXiv
Steering LLMs for Culturally Localized Generation 2026.03 arXiv
Mind the Gap in Cultural Alignment: Task-Aware Culture Management for Large Language Models 2026.02 arXiv
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters 2026.01 venue
Mitigating Cultural Bias in LLMs via Multi-Agent Cultural Debate 2026.01 arXiv
ACE-Align: Attribute Causal Effect Alignment for Cultural Values under Varying Persona Granularities 2026.01 arXiv
Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent Reasoning 2026.01 arXiv
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment 2025.09 arXiv
Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise 2025.06 arXiv
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs 2025.05 arXiv
NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities 2025.05 venue
CulFiT: Fine-grained Cultural-aware LLM Training via Multilingual Critique Data Synthesis 2025.05 venue
Cultural Learning-Based Culture Adaptation of Language Models (CLCA) 2025.04 venue
CAReDiO: Enhancing Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization 2025.04 arXiv
CARE: Multilingual Human Preference Learning for Cultural Awareness 2025.04 venue
Cultural Alignment in Large Language Models Using Soft Prompt Tuning 2025.03 arXiv
Towards Realistic Evaluation of Cultural Value Alignment: Diversity Enhancement for Survey Simulation 2025 venue
Cultural Palette: Pluralising Culture Alignment via Multi-agent Palette 2024.12 arXiv
Self-Pluralising Culture Alignment for Large Language Models (CultureSPA) 2024.10 venue
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration 2024.06 venue
CulturePark: Boosting Cross-cultural Understanding in Large Language Models 2024.05 venue
CultureLLM: Incorporating Cultural Differences into Large Language Models 2024.02 venue

3️⃣ Overton Pluralism

The goal: when asked a contested question, present the full window of reasonable responses rather than collapsing to a single answer. Coined in Sorensen et al.'s Roadmap to Pluralistic Alignment (see §1.1).

3.1 Overton: Named & Direct
Paper Year Venue
Evaluating Pluralism in LLMs through Latent Perspectives 2026.06 arXiv
Overton Pluralistic Reinforcement Learning for Large Language Models 2026.02 arXiv
Benchmarking Overton Pluralism in LLMs 2025.12 venue
POW: Political Overton Windows of Large Language Models 2025.09 venue
Pairwise Calibrated Rewards for Pluralistic Alignment 2025.06 venue
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment 2024.06 venue
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration 2024.06 venue
3.2 Similar View: Covering the Full Spectrum
Paper Year Venue
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies 2025.11 venue
Pluralistic Alignment for Healthcare: A Role-Driven Framework 2025.09 venue
Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble 2025.09 arXiv
EMBRACE: Shaping Inclusive Opinion Representation by Aligning Implicit Conversations with Social Norms 2025.07 venue
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks 2025.05 venue
Plurals: A System for Guiding LLMs Via Simulated Social Ensembles 2024.09 venue
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties 2023.09 venue

## 4️⃣ Steerable Pluralism

On demand, faithfully steer an LLM to represent a specific person, group, persona, or value profile.

4.1 Steerable Alignment & Persona Steering
Paper Year Venue
Coherence Maximization Improves Pluralistic Alignment 2026.06 arXiv
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment 2025.10 arXiv
PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences 2024.06 venue
On the steerability of large language models toward data-driven personas 2023.11 venue


## 5️⃣ Distributional Pluralism

Match the model output distribution to a target population distribution of opinions and values.

5.1 Distributional Alignment
Paper Year Venue
Characterizing the ability of LLMs to recapitulate Americans' distributional responses to public opinion polling questions across political issues 2026.03 arXiv
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs 2026.01 arXiv
Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions 2025.10 arXiv
Improving the Distributional Alignment of LLMs using Supervision 2025.07 arXiv
Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions 2025.02 venue
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations 2025.02 venue
How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective 2025.02 venue


## 6️⃣ Social Simulation

Use LLM agents to simulate people, opinions, and societies (silicon sampling, opinion dynamics, generative-agent societies).

6.1 Opinion & Population Simulation
Paper Year Venue
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses 2026.07 arXiv
Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling 2026.07 arXiv
Silicon Sampling via Cross-Survey Transfer 2026.07 arXiv
Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1 2026.07 arXiv
Will Scaling Improve Social Simulation with LLMs? 2026.07 arXiv
Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies 2026.06 arXiv
EconSimulacra: A Digital Twin Platform of Socio-Economic Systems Powered by LLM Agents 2026.06 arXiv
Improving Cross-Cultural Survey Simulation with Calibrated Value Personas 2026.05 venue
From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes 2026.05 arXiv
APS: Bias-Controlled Adaptive Prototype Simulation for Population-Scale LLM Agents 2026.05 arXiv
EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations 2026.05 arXiv
Persona-Based Simulation of Human Opinion at Population Scale 2026.03 arXiv
Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents 2026.02 venue
Opinion dynamics and mutual influence with LLM agents through dialog simulation 2026.02 arXiv
Digital Twins as Funhouse Mirrors: Five Key Distortions 2025.09 arXiv
Finetuning LLMs for Human Behavior Prediction in Social Science Experiments 2025.09 venue
ValueSim: Generating Backstories to Model Individual Value Systems 2025.05 arXiv
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals 2024.11 arXiv


About

A curated collection of papers, benchmarks, datasets, and tools on human values in LLMs and pluralistic alignment.

Topics

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages