Glossary · term

Open Character Training

Open Character Training (OCT) is a named open-weight post-training pipeline for making an assistant express a selected persona without an inference-time character prompt. It starts from a short first-person constitution, distils constitution-conditioned teacher responses into a student with direct preference optimization (DPO), then applies supervised fine-tuning (SFT) to synthetic self-reflections and self-interactions generated from the intermediate model.

Training2025-11-03Wave 3 · 2025–26Maturity: 3/5

Origin and context

Anthropic publicly described the broader practice of character training in June 2024 for Claude 3. Maiya, Bartsch, Lambert and Hubinger introduced OCT on 3 November 2025 as an open, reproducible implementation with eleven constitutions across Llama, Qwen and Gemma models. They released training code, data and LoRA adapters. The proper name identifies this recipe; it does not cover every persona-training method.

Sources: s1, s2, s3

Why it matters

OCT turns an otherwise opaque industrial practice into a testable baseline: researchers can inspect constitutions, rerun stages and compare weight-level persona shaping with prompting or activation steering. A peer-reviewed EigenBench paper uses an OCT-trained model as a validation target. Independent work at the ICML 2026 Pluralistic Alignment workshop then used OCT models and checkpoints to measure cross-constitution drift, showing why evaluation cannot stop at the target trait. A separate implementation reproduces and extends the recipe on another training stack.

Sources: s2, s8, s4, s5

Example

A team might write a constitution for candour, generate matched teacher and student responses, train a DPO adapter, then continue with self-reflection and self-interaction SFT. A responsible evaluation would compare the base, DPO and final checkpoints on candour, factual accuracy, sycophancy, unrelated traits and adversarial prompts. It would report the exact constitution, model, data, judge and seeds rather than treating a single persona score as proof of alignment.

Sources: s1, s4, s6

How it differs

Constitutional AI

Constitutional AI is the broader family of methods that uses written principles to supervise model behavior. OCT adapts that idea to first-person character assertions and adds a specific DPO-plus-introspective-SFT pipeline.

Direct Preference Optimization (DPO)

DPO is one optimization stage within OCT. Using DPO alone does not constitute the full OCT recipe, which also requires a character constitution, synthetic-data generation and introspective SFT.

Feature Steering

Feature steering changes activations at inference time. OCT changes weights through fine-tuning; the originating paper compares the two but does not make them interchangeable.

Maturity and evidence

Maturity is 3. OCT has a dated specification, open code and artifacts, an independent workshop study that uses it as a named experimental pipeline, and independent reimplementations. It is not rated higher because the originating work remains publicly verifiable as an arXiv preprint, independent evaluation is narrow, and no standard or broad production adoption was found.

Sources: s1, s2, s8, s4, s5

Limits and open questions

The original results rely heavily on synthetic data, model-based judges and small open-weight models. Its five capability benchmarks do not establish general capability preservation, and the authors' deliberately misaligned persona did lose performance. Independent OCT evaluation uses one main base model and reports collateral trait shifts, while peer-reviewed Nature work finds that a different warmth-SFT setup can increase error and sycophancy; that result motivates broader auditing but is not a direct OCT replication. `OpenCharacter`, despite its similar name, is an unrelated role-playing SFT system.

Sources: s1, s4, s6, s7

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as direct preference optimization skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm fine tuning skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as supervised fine tuning sft skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as fine tuning evaluation skill.