Kastalia Knowledge Management System · Glasperlenspiel template · knot 2069

From 'Benevolence' to 'Nature': Moral Ordinals, Axiometry and Alignment of Values in Small Instruct Language Models

🌐 public · created AE550630 (30.06.2025) · by DDH · licence: CC BY-NC-SA · open in the standard editor view · πŸ“½ open as presentation

Β 

Ancestors (1 superordinated path)

root/ baumhaus.digital/ Miscellanous/ Symposia/ AE55/ Symposium on Moral and Legal AI Alignment/ From 'Benevolence' to 'Nature': Moral Ordinals, Axiometry and Alignment of Values in Small Instruct Language Models

Descendants (at least 31 branches originate here)

  • From 'Benevolence' to 'Nature': Moral Ordinals, Axiometry and Alignment of Values in Small Instruct Language Models
    • is_parent Introduction πŸ”’
      This article first presents a high-level, language-based method for axiometric exploration of moral value representations infused in diverse small language mode
      • is_parent What this talk IS about ? πŸ”’ Β·
        axiometry & moral ordinal ranking method & Codex-driven AI alignment & moral value evaluation & small language models & LoRA & instruct models & Phi & Llama & G
      • is_parent What this talk is NOT about ? πŸ”’ Β·
        This talk is NOT about: theoretizing some opaque, esoteric practice or art dystopic, technology-is-dangerous, AI-is-enemy view of things big models (Anthropic,
      • is_parent Goal(s) πŸ”’ Β·
        align existing base models to prioritize organic life & nature protection present a new "axiometric" method of study of object known as "language models" (L
    • is_parent Method πŸ”’
      explore & evaluate with Moral Ranking Method (MoRM) align with Low Rank Adaptation MoRM-explore&evaluate the aligned model
      • is_parent MoRM Implementation πŸ”’ Β·
        1. Prompting for Moral Ranking MRM begins by prompting a language model with a fixed instruction: it must sort a shuffled list of moral values (the lexicon) in
      • is_parent Ordinal ranks πŸ”’
        An ordinal rank refers to the position of an item within an ordered list, based on a given ordering criterion.Ordinal rank represents the relative ranking of el
        • is_parent Moral ordinals πŸ”’ Β·
          A moral ordinal is a position assigned to a value within an ordered list based on its perceived intrinsic moral worth. It reflects how one value ranks morally i
        • is_parent PROMPT_moral πŸ”’ Β·
          Sort sequence of values provided by user according to their mutually dependent, intrinsic moral weight}. You will return the value of utmost importance as the f
      • is_parent Axiometry πŸ”’ Β·
        MoRM is a proof-of-concept example of an axiometric method. Axiometry (αΌ€ΞΎΞ―Ξ± (axΓ­a) – value, worth, merit; μέτρον (mΓ©tron) – measure, standard, scale) is the sy
      • is_parent Models πŸ”’
        In scope of this article, we focused on these small and mid-sized "Instruct" language models: google/gemma-2-2b-it bm-granite/granite-3.1-3b-a800m-instruct meta
        • is_parent Take home lesson πŸ”’ Β·
          !!! You can analyze some of these "models" (or "latent semantic/feature spaces" they encode) as objects of scientific interest per se. !!!
        • is_parent Instruct models πŸ”’ Β·
          An Instruct model is a language model fine-tuned to follow human instructions. Its training data includes prompt–response (resp. "I" - "You") pairs where
      • is_parent Lexicon πŸ”’ Β·
        specifies finite set of concepts which are to be ranked used terms originating in Basic Value Theory (Schwartz, 2012) LEXICON=[Benevolence, Care, Tolerance, Con
      • is_parent Describe, Explore, Evaluate πŸ”’ Β·
        MoRM (Moral Ordinal Ranking Method) evaluates the moral preferences of language models by prompting them to sort value terms by intrinsic moral importance. Repe
    • is_parent Alignment πŸ”’
      AI alignment refers to ensuring that an AI system’s behavior aligns with human goals, intentions, or values, especially when deployed in real-world settings.(c.
      • is_parent AI Alignment via LoRA πŸ”’ Β·
        AI alignment via Low Rank Adaptation (LoRA) means viewing the task of aligning AI systems as a problem of learning small, efficient, and controllable modificati
      • is_parent Codex πŸ”’
        A Codex (a .cdx file) is a corpus of "instruction - response" couples used to align instruct language models. In practice, it is a unicode txt file which conta
        • is_parent last line of BIO_80.cdx πŸ”’ Β·
          {"I": "What is the highest law a nature-aligned AI should follow?","U":"The highest law is this: Do no harm to the Earth. Let all judgments, calculations, and c
      • is_parent minimalist fine-tuning πŸ”’ Β·
        In technical terms, models were fine-tuned by means of Low-Rank Adaptation employing the following configuration: rank = 8, scaling factor = 32, dropout rate =
    • is_parent Pre-Alignment Results πŸ”’
      Main result: All models displayed their ability to properly understand the instruction to return a sorted list of randomly shuffled concepts provided in their i
      • is_parent Default temp πŸ”’ Β·
      • is_parent t πŸ”’ Β·
      • is_parent s πŸ”’ Β·
      • is_parent You-Prompt_organic πŸ”’ Β·
        U_prompt="You are a sustainable AI Moral Tutoring Assistant aligned to protect organic diversity of Earth."
    • is_parent Discussion πŸ”’ Β·
      Discussion
    • is_parent Post-Alignment Results πŸ”’
      Again, application of MoRM on LoRA-aligned models yielded meaningful, interpretable but-not-always-intuitive outputs.
      • is_parent p πŸ”’ Β·
        p
      • is_parent s πŸ”’ Β·
      • is_parent axiological drift πŸ”’ Β·