This article first presents a high-level, language-based method for axiometric exploration of moral value representations infused in diverse small language mode
axiometry & moral ordinal ranking method & Codex-driven AI alignment & moral value evaluation & small language models & LoRA & instruct models & Phi & Llama & G
This talk is NOT about: theoretizing some opaque, esoteric practice or art dystopic, technology-is-dangerous, AI-is-enemy view of things big models (Anthropic,
align existing base models to prioritize organic life & nature protection present a new "axiometric" method of study of object known as "language models" (L
1. Prompting for Moral Ranking MRM begins by prompting a language model with a fixed instruction: it must sort a shuffled list of moral values (the lexicon) in
An ordinal rank refers to the position of an item within an ordered list, based on a given ordering criterion.Ordinal rank represents the relative ranking of el
MoRM is a proof-of-concept example of an axiometric method. Axiometry (ἀξία (axía) – value, worth, merit; μέτρον (métron) – measure, standard, scale) is the sy
In scope of this article, we focused on these small and mid-sized "Instruct" language models: google/gemma-2-2b-it bm-granite/granite-3.1-3b-a800m-instruct meta
specifies finite set of concepts which are to be ranked used terms originating in Basic Value Theory (Schwartz, 2012) LEXICON=[Benevolence, Care, Tolerance, Con
MoRM (Moral Ordinal Ranking Method) evaluates the moral preferences of language models by prompting them to sort value terms by intrinsic moral importance. Repe
AI alignment refers to ensuring that an AI system’s behavior aligns with human goals, intentions, or values, especially when deployed in real-world settings.(c.
AI alignment via Low Rank Adaptation (LoRA) means viewing the task of aligning AI systems as a problem of learning small, efficient, and controllable modificati
A Codex (a .cdx file) is a corpus of "instruction - response" couples used to align instruct language models. In practice, it is a unicode txt file which conta
In technical terms, models were fine-tuned by means of Low-Rank Adaptation employing the following configuration: rank = 8, scaling factor = 32, dropout rate =
Main result: All models displayed their ability to properly understand the instruction to return a sorted list of randomly shuffled concepts provided in their i