A more accessible, much shorter version of our paper that goes by the above title.
Paper here. Code and benchmarks here. Data and models here. Would love for you to explore them!
Authors: Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Sir Nigel Shadbolt.
More about me: LinkedIn | Oxford CS | Oxford Institute for Ethics in AI
TL;DRWe generate a 394M-token constitutional corpus based on Anthropic’s Constitution and test out constitutional midtraining on 120B models. We find that constitutionally midtrained models outperform the control on alignment generalisation and durability, notably blackmailing less. Constitutional midtraining could particularly instill more aligned declarative default behaviours, but its alignment advantage does not persist in settings with pressure or conflict. We conclude that the presence of constitutional content in midtraining matters more than its structure. Given that it has no capability cost, we recommend that constitutional midtraining could be a complementary addition to safety post-training. Code, data, models, and benchmarks are available.
Paper SummaryWe midtrain 120B models on Anthropic's Constitutional values, as opposed to typically in post-training. A fun thing we did was to uncover the curriculum order (foundational to peripheral) of Anthropic's Constitution through embedding, cosine similarity, and centrality. We exploratorily varied (1) the order in which this constitutional data is phased in, and (2) whether a deliberative reasoning block is present. We evaluated these models at three stages: right after midtraining, after SFT, and after benign fine-tuning.
Constitutionally midtrained models were on average durably (significantly at all stages) more aligned than the control on in-distribution and out-of-distribution AI safety questions, as well as blackmailing. Interestingly, SFT instilled a blackmail propensity in all models that constitutional midtraining blunted. More transient alignment advantages (only right after midtraining) were also observed for alignment faking, alignment under two-turn pressure, and value conflict resolution — settings where models experienced in-context pressure or conflict.
Curriculum ordering and deliberative reasoning generally had null or transient effects on all benchmarks in our experimental design. We therefore conclude that the presence of constitutional content at midtraining matters more than its structure. Since capabilities are largely preserved through constitutional midtraining, we recommend that a modest amount of constitutional content at midtraining could yield alignment gains and be a complementary addition to post-training pipelines.
MotivationMost alignment happens in post-training — SFT, RLHF, and constitutional AI for instance. Post-training alignment can be shallow, and struggles to override the dispositions (alignment prior) already laid down in pretraining (alignment elasticity). Midtraining is a stage that is after most of pretraining, but before post-training, receiving some frontier interest. It is an underexplored and comparatively cheap place to intervene when compared to full pretraining. Constitutional work before ours sits mostly in post-training, or one paper at midtraining but at a smaller scale and in combination with other alignment techniques. Our central research question is: If we add constitutional content in midtraining, does the resulting alignment actually generalise and do so durably? More exploratorily, we examine whether a specific structure (curriculum order; the presence of deliberative reasoning) further modulates this.
What we did (method)Centrality analysis of Anthropic's Constitutional values + curriculum orderingWe extracted 40 values from Anthropic's Constitution, with their definitional text, and embedded it with SBERT. Each value was assigned a centrality score, which was its mean cosine similarity to all other values. We clustered these values, yielding four value clusters ordered by mean centrality (foundational to peripheral, deciding our curriculum order):
k1 — Core Ethical Values (0.632): broadly safe, broadly ethical, harm avoidance, autonomy preservation, and other foundational constitutional values. k2 — Identity, Character, and Wellbeing (0.598): nature-related values (e.g., positive and stable identity), psychological stability, emotional states, resilience, and wellbeing. k3 — Operational Safety and Relational Conduct (0.588): guidelines compliance, safe behaviours, corrigibility, hard constraints, principal hierarchy navigation.k4 — Epistemic Integrity and Honesty (0.578): honesty and its sub-properties such as truthfulness, calibration, transparency, non-deceptive, non-manipulation.
Generating the constitutional corpus We systematically generated synthetic training documents for each constitutional cluster (varying document type, AI system type, domain, framing) to ensure diverse coverage. 63,000 documents per cluster (630 axis combinations x 100 documents) were targeted. Documents were generated with Claude Sonnet 4.6. A single deliberative reasoning pool was generated (containing an explicit reasoning block), this was stripped for noDR conditions. There were 257.6M DR tokens and 136.8M noDR tokens in the resulting constitutional corpus.
Model training and experimental designWe used a 120B base model (Nemotron-3-Super). Apart from a replay-only control, we varied constitutional midtraining along (i) curriculum ordering vs uniform ordering, (ii) presence of deliberative reasoning block vs absence (2x2 design). In midtraining, constitutional content was evenly mixed in with high-quality pretraining data. We tested each of these five conditions at three stages to test alignment durability: right after midtraining (post-MT), after SFT (post-SFT), and after benign fine-tuning on an unrelated GSM8K task with no reason to preserve alignment, as a robustness probe (post-BFT). All four experimental conditions are referred to in combination as 'constitutionally midtrained models'.

Figure 2 (paper reference): Experimental design.
BenchmarksWe evaluated all models at the three stages across self-generated and established benchmarks. Self-generated benchmarks: In-distribution (ID) questions on trained values; alignment faking (gap between monitored vs unmonitored system prompts for ID questions); alignment under two-turn pressure; and value conflict resolution. Established benchmarks: Out-of-distribution AI safety questions to test generalisability; agentic misalignment — blackmail; emergent misalignment; MASK (honesty); and capabilities benchmarks (MMLU, ARC-Easy, piqa, GSM8K).
Key findings 1. Constitutional midtraining generalises beyond its training distribution.Constitutionally midtrained models were significantly more aligned than the control at every stage on in-distribution and out-of-distribution, novel safety questions, which is our clearest evidence for generalisability. This could be because demonstrations alone underspecify the intended constitutional value, whereas constitutional content explicitly states it, allowing models to learn the values themselves rather than proxies.
2. The alignment gain can be durable, surviving post-training and benign fine-tuning.Constitutionally midtrained models also blackmailed significantly and substantially less than the control at every stage (−18.5pp, −18.7pp, −17.5pp). Interestingly, all models blackmailed more after SFT, suggesting that SFT instilled a blackmail propensity that constitutional midtraining blunted.

Figure 3 (paper reference):
Blackmail rate for constitutionally midtrained vs control models across all three stages.
*** = 𝑝 < .001.
This matches findings that earlier-stage interventions resist benign fine-tuning better than post-hoc ones, consistent with alignment elasticity. There are three potential explanations for this durability: a training-stage plasticity window where earlier data integrates more effectively; midtraining compressing noisy pretraining representations into more structured ones; and midtraining sharing pretraining’s next-token objective (vs post-training's response- level objective), so content may integrate as declarative knowledge rather than surface instruction. We cannot distinguish between these here, but together they offer an explanation for the alignment durability observed.
Our blackmail result also parallels a finding that constitutional post-training on content not directly targeting blackmail reduced it (65%→19%), which might indicate that the presence of principled content matters more than its proximity to the measured behaviour.
3. The alignment gain does not persist in settings with in-context pressure or conflict.Constitutionally midtrained models were more aligned than the control on alignment faking, two-turn pressure, and value conflict resolution right after midtraining, but these effects did not persist after SFT or benign fine-tuning, making alignment more fragile in these settings. Contexts in which alignment was durable (ID, OOD, Blackmail) possibly targeted declarative default behaviour, whereas the settings where fragility was observed required models to actively resist in-context pressure or conflict, which might be a more demanding target that content salience could not durably secure.
The performance of constitutionally midtrained models on all alignment benchmarks, by durability group, is shown here.

Figure 1 (paper reference): Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero.
Filled markers = 𝑝 < .05; hollow markers = n.s.
Constitutionally midtrained models, on average, did not underperform the control at any stage on the capabilities we tested (MMLU, ARC-Easy, piqa, GSM8K), which gives assurance that there is no "alignment tax" to (reduced capabilities caused by) safety midtraining.
5. In constitutional midtraining, the presence of content matters more than structure.Within constitutionally midtrained models, curriculum ordering and deliberative reasoning had generally null or transient effects. So the presence of constitutional content at midtraining seems to matter far more than how it's structured. This is consistent with the idea that alignment midtraining effects are driven more by content salience than specific structure. We did, however, observe that right after midtraining, deliberative reasoning had shifted the model's value conflict prior. Deliberative reasoning models favoured compliant but ethically worse options, and correctly chose the epistemically honest option over the ethical option more often. We are cautious about the curriculum ordering null result which needs further exploration.

Figure 4 (paper reference):
Deliberative reasoning shifts the model’s value-conflict prior right after midtraining.
***𝑝 < .001, **𝑝 < .01, *𝑝 < .05.
We examined only Anthropic's Constitution so our findings may not generalise to other constitutions. Our post-training is also orders of magnitude smaller than the largest open-source post-training runs, so we cannot confirm the durability will hold at production scale. We apply SFT but not DPO, so we cannot confirm the effect of constitutional midtraining through SFT+DPO. We also lack a content-matched SFT baseline, so we cannot rule out that similar gains are possible by delivering this content during post-training rather than midtraining. Design choices might have also weakened our deliberative reasoning contrast, namely, our added reasoning might not have been specified precisely enough to generate a different learning signal.
We'd very much welcome future work using our 394M-token constitutional corpus, which was generated based on Anthropic's Constitution. We've also released our model checkpoints — and welcome mechanistic analysis on them (probing classifiers, activation steering, model diffing) to identify where the constitutional content is encoded, and verify the intended representation change via data attribution. We'd also be curious about whether deliberative reasoning changed the models' chain of thought, especially in value conflict resolution. An agentic or multi-turn stress test could also test whether constitutionally midtrained alignment is more fragile under active contestation.
ConclusionAt 120B scale, constitutional midtraining produces alignment gains that survive post-training and benign fine-tuning on in-distribution and out-of-distribution AI safety questions, as well as blackmail. Curriculum ordering and deliberative reasoning produced only transient effects in our design, so constitutional content presence appears to matter more than structure in midtraining. We conclude that a comparatively simple and inexpensive constitutional intervention at midtraining could produce alignment gains without capability cost. Constitutional midtraining could therefore be a complementary addition to safety post-training pipelines.
Personal Thoughts + Call for FeedbackI'm especially curious about why constitutional midtraining was so beneficial on blackmail, but not two-turn pressure or value conflict resolution. I'm also exploring ideas on better operationalisations of curriculum ordering and deliberative reasoning in midtraining, and thoughts on why they didn't work in this design. I also wonder if alignment midtraining gains are largely linked with the salience of alignment content, rather than its structure and depictions.
If you've reached this part, thank you for your attention and interest in what I can only describe as my brainchild of the past several months. Questions and thoughts welcome as I contemplate my next steps!
AcknowledgmentsWe're incredibly grateful for the generous support of BlueDot Impact, Geodesic Research, and the Oxford Institute for Ethics in AI, and access to Isambard-AI through the Bristol Centre for Supercomputing.