Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers


ai-harm-assessment-methodology
Overview of our four-step methodology. Step 1. We selected three use cases—AI kidney allocation (KIDNEY), AI-simulated workers (WORK), and generative AI content of the deceased (GEN)—based on harm impact and likelihood of encounter, capturing rare high-stakes to common but less harmful cases. Step 2. We used three moral framings—(A) Control (baseline), (B) World-You-Want (societal consequences), and (C) Could-Be-You (perspective-taking)—capturing structural and perspective-taking prompts to examine how different moral lenses shape moral preference elicitation. Step 3. In Phase 1, participants identified moral features relevant to AI decision-making for each use case. Step 4. In Phase 2, we elicited moral preferences for these features from a new sample using a 7-point scale (−3 to +3) across the three framings in Step 2. We then modeled how political ideology influences these preferences through moral foundations, with moral framings as a moderator.


As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sampling, and question framing. In other words, they decide which features go to a vote, which voters to include, and how to present the question. These choices are often opaque, undocumented, and treated as technical details rather than normative ones.

We examine each of these choices within a common empirical study and show that each can shape the preferences produced by moral-AI elicitation. Across two studies (N = 817) in three deployment contexts (i.e., AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased), we examine the three main stages of the moral AI elicitation pipeline. First, morally relevant features shift across contexts. This suggests that feature schemas should not be assumed to transfer across deployment domains. Second, preferences differ by political ideology for roughly one-third of features, with some differences reversing direction. The ideological composition of the voter pool can therefore affect the resulting aggregated preference profile. Third, the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The framing conditions also change how moral foundations are associated with participants' judgments. Taken together, these findings suggest that voting-based alignment cannot deliver fair or transparent AI by aggregation alone; at minimum, each stage of the moral AI elicitation pipeline should be audited and disclosed.



Publications

  • Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers. AIES 2026 PDF (Upcoming)