🕸️
Main result: All models displayed their ability to properly understand the instruction to return a sorted list of randomly shuffled concepts provided in their input.
This article first presents a high-level, language-based method for axiometric exploration of moral value representations infused in diverse small language models. The method is based around the idea of "moral ordinals" - a list of items from a value lexicon which the model is prompted to sort according to its own intrinsic "morality" criterion. After presenting the method, the lexicon based on Schwartz's ``basic value theory'' is used to explore dominance of different value representations in 6 small (<4 milliard parameter) language models. For most models, ``benevolence'' is consistently ranked at the highest position and there is no statistically significant difference between rankings obtained at minimal and default inference temperatures. Across all models, the distribution of aggregate moral-ranking scores was well approximated by a Beta distribution (K–S $p > 0.3$), revealing consistent yet model-specific patterns of moral weighting. Subsequently, foundational models are subjected to a sort of ``minimalist alignment'' whereby they undergo 7 epochs of performance-efficient fine-tuning with synthetically generated 80-instruction codex directed towards sustainability and nature protection. Finally, such minimally aligned models are explored once again with the ``moral ordinals'' method, providing insights into axiological drift induced by the mini-alignment process.
  1. explore & evaluate with Moral Ranking Method (MoRM)
  2. align with Low Rank Adaptation
  3. MoRM-explore&evaluate the aligned model
Discussion
[Impressum, Datenschutz, Login] Other subprojects of udk.ai linkring: teacher.solar gardens.digital fibel.digital refused.science baumhaus.digital