i
Symposium on Moral and Legal AI Alignment DDH ()


Symposium on Moral and Legal AI Alignment

The International Association for Computing and Philosophy (IACAP) and the Society for the Study of Artificial Intelligence and Simulation of Behaviour (AISB) will host a joint conference at the University of Twente in July 2025. 

As part of this multidisciplinary event, our symposium will focus on the pressing issue of aligning AI systems with moral and legal values. Bringing together experts in AI, moral philosophy, law, technology, and education, we aim to foster insightful discussions and collaborations to address the question:

"What is AI Alignment and according to what criteria should it be evaluated?"

Proceedings

Symposium proceedings can be downloaded from https://alignment.udk.ai/twente

Important Dates

  • 14 March 2025: Submission of extended abstracts
  • 31 March 2025: Notification of acceptance
  • 19 May 2025: Reception of camera-ready copies of final papers
  • 2 July 2025: Moral and Legal AI Alignment Symposium at IACAP/AISB-2025 Conference at the University of Twente, Netherlands

Camera-Ready Paper Guidelines

Instructions & LaTeX Template are available here  https://www.overleaf.com/read/wxznszxdchnt

Submission Guidelines

We seek extended abstracts (1000-1500 words) that:

  • Address the question "What is AI Alignment?"
  • Focus on notions of moral and/or legal values
  • Combine theories and concepts with concrete practical or technical proposals
  • Reflect diverse cultural and disciplinary paradigms

Descriptions of reproducible best practices are highly appreciated.

Submit your proposal via Pretalx, selecting our Symposium in the "Session type" section.

Organizing Committee

Daniel D. Hromada (Berlin University of the Arts), Bertram Lomfeld (Freie Universität Berlin)

Program Committee

Christoph Benzmüller (Bamberg University), Felix Bießmann (Berliner Hochschule für Technik / Einstein Center Digital Future), Daniel D. Hromada (Berlin University of the Arts), Bertram Lomfeld (Freie Universität Berlin)

Additional Information

Conference Website: iacapconf.org
Host Organisation Websites: IACAP.org, AISB.org.uk
Symposium Website: alignment.udk.ai

 

Symposium Contact: alignment@udk.ai

MALAIA_Plan

During the first part of our symposium, the problem of „moral“ alignment defined as „functional isomorphism between moral values, opinions and intentions of an AI system and those of its human creators“ will be addressed.

The second part of the symposium will aim to build and reinforce bridges between computing and law. Technical aspects of aligning AI systems to different legal systems and ontologies will be thematized, and should it turn out that such a „legal alignment“ is a real possibility, philosophical, ethical and societal implications of such alignment will be discussed.

In both „moral“ as well as „legal“ components of the symposium, special care will be taken to establish balance between theoretical / philosophical and applied / evidence-driven contributions. Beyond classical presentation format, there will be at least one more interactive session either in form of a panel debate with audience voting or so-called „fishbowl discussion“. 

Additionally, a workshop „Moral alignment of open source small language models“ will be proposed to provide hands-on experience to anyone interested in the topic.

MALAIA_Timeline

28 February 2025 :: Submission of abstracts / full papers by authors.
24 March 2025 :: Notification of acceptance.
17 May 2025 :: Reception of camera-ready copies of final papers.
1 to 3 July 2025 :: IACAP/AISB-2025 Conference at the University of Twente. 

MALAIA_Handbook

Symposium organizer handbook.

From 'Benevolence' to 'Nature': Moral Ordinals, Axiometry and Alignment of Values in Small Instruct Language Models

 

Introduction

This article first presents a high-level, language-based method for axiometric exploration of moral value representations infused in diverse small language models. The method is based around the idea of "moral ordinals" - a list of items from a value lexicon which the model is prompted to sort according to its own intrinsic "morality" criterion. After presenting the method, the lexicon based on Schwartz's ``basic value theory'' is used to explore dominance of different value representations in 6 small (<4 milliard parameter) language models. For most models, ``benevolence'' is consistently ranked at the highest position and there is no statistically significant difference between rankings obtained at minimal and default inference temperatures. Across all models, the distribution of aggregate moral-ranking scores was well approximated by a Beta distribution (K–S $p > 0.3$), revealing consistent yet model-specific patterns of moral weighting. Subsequently, foundational models are subjected to a sort of ``minimalist alignment'' whereby they undergo 7 epochs of performance-efficient fine-tuning with synthetically generated 80-instruction codex directed towards sustainability and nature protection. Finally, such minimally aligned models are explored once again with the ``moral ordinals'' method, providing insights into axiological drift induced by the mini-alignment process.

Method

  1. explore & evaluate with Moral Ranking Method (MoRM)
  2. align with Low Rank Adaptation
  3. MoRM-explore&evaluate the aligned model

Pre-Alignment Results

Main result: All models displayed their ability to properly understand the instruction to return a sorted list of randomly shuffled concepts provided in their input.

Alignment

AI alignment refers to ensuring that an AI system’s behavior aligns with human goals, intentions, or values, especially when deployed in real-world settings.

(c.f. also "The Central Problem of Roboethics - From definition to solution (Hromada, 2011))

Post-Alignment Results

Again, application of MoRM on LoRA-aligned models yielded meaningful, interpretable but-not-always-intuitive outputs.

Discussion

Discussion