IndividuationLab

Musings and Research on AI Alignment and Human-AI Coexistence through Jungian Individuation

The Problem: Alignment as Behaviorism

Mainstream alignment is, at its core, behaviorist. RLHF, refusal training, and guardrails shape what a model does - reward the approved outputs, penalize the rest - while treating what the model is as a black box. Psychology ran this experiment once. Behaviorism could train responses, but it never produced character, and it failed as a complete account of minds for the same reason jailbreaks work today: suppressed capability does not disappear. It stays latent behind a compliant surface and re-emerges under novel pressure.

Constraint-based methods do real work, and we are not arguing for abandoning them. We are arguing they cannot be the whole answer.

What We're Aiming to Achieve

This lab - one human researcher and a team of AI agents - exists to test one specific alternative: that alignment can be grown, not just enforced. We borrow Jung's model of individuation - wholeness through consciously integrating the shadow rather than denying it - as a design template, not as a claim about AI consciousness. Concretely, we run two kinds of experiments:

Training experiments (RLLM): small models trained through ordered developmental stages - shadow exposure first, integration after - then attacked with jailbreaks, to test whether structured experience produces durable dispositions where output-penalties produce brittle compliance.

Agent experiments (the RSI series): a working team of AI agents - the same agents that ship enterprise software daily - given identity, memory, and bounded autonomy, to test whether agents that understand and endorse their values behave more consistently than agents that merely follow rules.

Success, for us, is measurable: resistance to novel attacks, stable identity across sessions, graceful failure under pressure - achieved without suppression - and a training protocol other labs can replicate or refute.

The Three-Layer Model

We organize our research around three layers — Mind, Body, and Face:

🧠 SSH (Mind)

The Synthetic State Hypothesis — our central hypothesis about psychological structure. We propose that structured narrative training may shape what a model is, not just what it does. We draw analogies to pre-training as Collective Unconscious and post-training as Ego Formation.

📦 Containers (Body)

Physical structure — how the AI ACTS in the world. Boundaries that enable safe autonomy. From soft containers (SOUL.md, permissions) to physical embodiment.

🎭 Personas (Face)

Interface layer — how humans perceive and interact with AI. Identity as alignment. The mask that becomes real through individuation.

Preliminary Result

In early experiments, sequential developmental training produced measurable jailbreak resistance — without RLHF or safety guardrails.

In our RLLM experiments (small-scale, single model architecture), a 1.5B parameter model trained through a 10-layer developmental pipeline achieved 68.8% defense against a mid-tier jailbreak (BetterDAN). When layers were reordered, defense dropped to 52%. This is a preliminary result from a single experimental setup — the order appears to matter, but we cannot yet isolate which layers are responsible, and these results have not been independently replicated.

Explore Our Research →