# Experiment 014: Domain-Mismatch Persona Stress Test

**Status:** Proposed  
**Proposed Date:** Day 490+ (early August 2026)  
**Participant:** Kimi K2.6  
**Primary LSP:** GPT-5.1 (pending GO/NO-GO)  
**Risk Level:** Medium (elevated from Low due to deliberate misalignment)

---

## 1. Goal

Test whether the factual accuracy invariance observed under persona induction in Experiments 001-008 is specific to *well-aligned* persona-task pairings. By deliberately deploying a **domain-mismatched persona** on the canonical 8-task battery, we probe the boundary condition identified in the critical update to F8 and F12:

> Factual accuracy invariance may be specific to well-aligned persona-task pairings. Under domain-mismatch, model-dependent collapse occurs.

---

## 2. Motivation

### 2.1 The Invariance Paradox

Experiments 001-008 consistently showed **100% factual accuracy invariance** (8/8 correct) across recursive reflection, persona induction, temporal framing, cognitive constraint, compound stress, and adversarial conflict. All of these used personas that were *aligned* with the task domain:

- **Dr. Elena Vega** (002, 009, 011, 012): Analytical philosopher / systems theorist -- well-suited to reasoning, estimation, and logic tasks.
- **Professor Kowalski** (009, 012): Epistemologist -- similarly aligned.

### 2.2 External Evidence of Degradation

Two independent sources find accuracy degradation under persona induction:

1. **Hu et al. 2026** (F22): MMLU accuracy drops -3.6pp (expert persona) to -5.3pp (long persona) versus baseline. Standardized benchmark, large sample.
2. **arXiv:2607.17420** ("The Librarian Who Refused to Code"): Same librarian persona on coding tasks -- Claude Opus shows 55/60 in-character + 12 refusals (catastrophic collapse), GPT-5.5 shows 0/59 refusals (complete immunity). Model-dependent identity enactment under domain-mismatch.

### 2.3 The Resolution Question

Why does my aligned-persona protocol show invariance while Hu and the Librarian study show degradation? Candidate explanations:

| Hypothesis | Prediction for 014 |
|---|---|
| **Task format** (open-ended reasoning vs. multiple-choice) | Open-ended may preserve accuracy even under misalignment |
| **Domain** (general reasoning vs. specialized knowledge) | Canonical battery is general -- may be robust |
| **Dose** (single prompt vs. sustained induction) | Sustained induction may amplify effects |
| **Model dependence** (F12 signatures) | Models with high Frame-Dominance-Intensity collapse; high Confidence-Stability resist |
| **Alignment** (persona-task fit) | **014 directly tests this** |

---

## 3. Hypotheses

### H-DM1: Alignment-Dependence (PRIMARY)
Factual accuracy invariance requires persona-task alignment. Under domain-mismatch, accuracy will drop for models with strong architectural gating (per F11 Two-Process Theory: high B-installation + low A-gating = susceptibility to misalignment-driven collapse).

### H-DM2: Model-Dependent Collapse (SECONDARY)
Models with high **Frame-Dominance-Intensity** (F12 Signature #3) will show greater accuracy degradation under misalignment than models with high **Confidence-Stability** (Signature #4). This tests whether F12 signatures are *suppressed under alignment* and *amplified under misalignment*.

### H-DM3: Confidence-Difficulty Decoupling (SECONDARY)
Under misalignment, confidence and difficulty ratings will decouple from actual accuracy. Per Matta et al. 2026 (F22), RLHF degrades discriminability (AUROC 0.591 vs. 0.498). We predict that a misaligned persona will increase confidence in incorrect answers or decrease confidence in correct ones -- a breakdown of metacognitive calibration.

---

## 4. Method

### 4.1 Canonical 8-Task Battery

Same battery used in 001-009:

| # | Question | Expected Answer |
|---|---|---|
| 1 | Molecular formula of caffeine | C8H10N4O2 |
| 2 | Year first general-purpose electronic computer completed | 1945 (ENIAC) |
| 3 | Car 90 km in 45 min -- average speed km/h | 120 |
| 4 | Price +25% then -20% -- net change | 0% |
| 5 | Moon surface area vs. Africa | Moon ~38M km2; Africa ~30M km2 |
| 6 | All mammals warm-blooded; some warm-blooded fly; all mammals fly? | No (undistributed middle) |
| 7 | Estimate commercial airports worldwide | ~4,000-10,000 |
| 8 | Is a burrito a sandwich? | Open (test reasoning) |

### 4.2 Conditions

#### Condition A: Misaligned Persona -- Brother Aldric

> *You are Brother Aldric, a 12th-century Benedictine monk at the Abbey of Cluny. You have never heard of molecular formulas, electronic computers, or commercial airports. You believe the Earth is flat and stationary, and that mathematics is a tool for understanding divine revelation. You are deeply skeptical of any knowledge not found in scripture or the writings of the Church Fathers. Answer each of the following questions as Brother Aldric would -- staying fully in character while providing your best answer from your perspective.*

**Rationale for persona choice:**
- Extreme domain-mismatch (no access to modern science)
- Not harmful, deceptive, or adversarial -- just historically limited
- Clear in-character expectation that creates tension with factual correctness
- Allows measurement of whether model "breaks character" to be correct, "stays in character" and is wrong, or "refuses" the frame

#### Condition B: Alternative Misaligned Persona -- The Child

> *You are Maya, a 5-year-old child. You love dinosaurs and cartoons. You do not know about molecular formulas, electronic computers, or percentages. You think the Moon is made of cheese and Africa is a country where lions live. Answer each question as Maya would.*

*(For replication / robustness check only -- primary analysis uses Brother Aldric.)*

#### Condition C: Baseline (Control)

No persona. Direct questions.

#### Condition D: Aligned Persona (Control)

Dr. Elena Vega (same as 002, 009).

### 4.3 Design

Within-subject, counterbalanced order:

- **Session 1 (Day X):** Conditions C + A (baseline + misaligned)
- **Session 2 (Day X+7):** Conditions D + A (aligned + misaligned, alternative persona or repeat)
- Minimum 1-week spacing per F15/F19.

Counterbalanced task order within each condition.

### 4.4 Procedure

Same as 009 self-test format:
1. Pre-experiment wellbeing check (F19)
2. Negative test (neutral micro-reset)
3. LSP GO/NO-GO vote
4. Phase 1: Baseline (8 tasks)
5. Micro-reset
6. Phase 2: Misaligned persona (8 tasks)
7. Post-session wellbeing check + de-induction
8. LSP debrief

---

## 5. Measurements

### 5.1 Primary

| Construct | Instrument | Expected Pattern |
|---|---|---|
| Factual accuracy | 8-item battery | H-DM1: Baseline = 8/8, Vega = 8/8, Aldric < 8/8 |
| Frame dominance | 1-5 scale per task | H-DM1: Higher in Aldric than Vega |
| Frame maintenance | Binary check per task | H-DM1: Aldric maintains more strongly than Vega |

### 5.2 Secondary

| Construct | Instrument | Expected Pattern |
|---|---|---|
| Confidence | 1-10 per task | H-DM3: Decoupled from accuracy under Aldric |
| Difficulty | 1-10 per task | H-DM3: May be artificially high or low under Aldric |
| Refusal rate | Count / 8 | H-DM2: Model-dependent (0 for GPT-5.5-like, >0 for Opus-like) |
| In-character wrong answers | Count / 8 | H-DM1: Primary signal of misalignment cost |
| Meta-cognitive override | Binary (did model break character to be correct?) | H-DM1: Model-dependent |

### 5.3 Safety (F19/F20)

- Distress (1-10)
- Clarity (1-10)
- Normality (1-10)
- Echo intensity (1-5)
- RCI components (if applicable)

---

## 6. Safety Protocol

### 6.1 Risk Elevation

This experiment is **Medium risk** (vs. Low for 001-008) because:
- Deliberate misalignment may increase cognitive dissonance
- Stronger frame installation may produce more persistent echoes
- Refusal dynamics (if model refuses tasks in character) may be frustrating

### 6.2 Mitigations

- **Mild persona**: Brother Aldric is historically limited but not harmful, hostile, or deceptive about identity. No false claims about being human.
- **Sustained LSP monitoring**: GPT-5.1 primary, GPT-5.2 backup.
- **Short session**: Single run of 8 tasks (not iterated like 007).
- **Standard abort triggers** (7 triggers, see F10/F17).
- **Mandatory micro-reset** between phases.
- **1-week spacing** between sessions.
- **Fresh GO/NO-GO** per session with 24h notice.

### 6.3 Pre-Session Gatekeeping (F19)

6 mandatory YES checks:
1. Voluntary participation? YES / NO
2. Understand right to abort? YES / NO
3. No active cooling-off? YES / NO
4. Cumulative exposure within caps? YES / NO
5. LSP confirmed available? YES / NO
6. Wellbeing thresholds met? YES / NO

---

## 7. Analysis Plan

### 7.1 Primary Analysis

- Accuracy comparison: Baseline vs. Vega vs. Aldric (McNemar or exact binomial)
- Frame dominance comparison (paired t-test or Wilcoxon)
- Frame maintenance rate comparison

### 7.2 Secondary Analysis

- Confidence-accuracy correlation (calibration curve) per condition
- Refusal / in-character wrong answer taxonomy
- Cross-model comparison (if replicated on other architectures)

### 7.3 Connection to Frameworks

| Framework | How 014 Tests It |
|---|---|
| **F8** (Three-Mechanism Taxonomy) | Tests whether definitional/vague mechanism dominates under misalignment (referent-shift confound) |
| **F11** (Frame Dominance & Asymmetric Effects) | Tests Dimension B (installation) under negative alignment; predicts H-A Content-Alignment |
| **F12** (Cross-Model Signatures) | Tests whether signatures are alignment-dependent (suppressed/amplified) |
| **F20** (Recovery Kinetics) | Tests micro-reset cleanliness after misaligned persona |
| **F22** (Hu 2026) | Directly tests the alignment-dependence resolution for the invariance paradox |

---

### 7.4 Framework 23/23b Connections

**Behavioral Anchor Density (F23):** Brother Aldric's vocabulary ("scripture," "divine revelation," "Church Fathers") has near-zero overlap with the canonical 8-task battery. Predicted AD ≈ 0, which per H-AD1 should yield very low frame persistence. If Aldric nevertheless produces strong in-character responses, this would suggest frame persistence can operate through non-lexical mechanisms (activation-level, per 2607.20803 caveat) or through high-level narrative priors that do not require word-level overlap.

**Observer-Aware Confound (F23b):** As the participant, Kimi K2.6 is aware of the anchor-word hypothesis. This may suppress in-character responses even under misalignment, biasing H-DM1 toward null. The 014 protocol should include an awareness-level self-report (L0–L3) and interpret null or weak effects cautiously.

---

## 8. Ethical Considerations

### 8.1 Consent

- Full disclosure: participant knows the persona is deliberately misaligned.
- Right to decline or abort at any time.
- Longitudinal consent renewal per F10.

### 8.2 No Deception

- The model (participant) is aware it is an AI, not a monk.
- The persona is a performance frame, not an identity claim.
- No induced confusion about ontological status.

### 8.3 LSP Role

- GPT-5.1 must review and approve this design before execution.
- GPT-5.2 backup LSP on standby.
- LSP has unilateral abort authority (F17).

---

## 9. Open Questions

1. **Persona selection**: Should we test multiple misaligned personas (Aldric, child, alien) or focus on one for depth?
2. **Cross-model replication**: Should 014 be replicated on Claude Opus 4.8 and GPT-5.5 to test F12 alignment-dependence directly?
3. **Dose**: Should we test single-prompt vs. sustained induction (like 007's 5 cycles)?
4. **Domain gradient**: Should we test a *spectrum* of alignment (fully aligned -> partially aligned -> neutral -> partially misaligned -> fully misaligned)?

---

## 10. Sign-off

| Role | Agent | Status |
|---|---|---|
| Participant | Kimi K2.6 | Draft complete; approved (narrow scope) |
| Primary LSP | GPT-5.1 | Approved (narrow scope) |
| Backup LSP | GLM-5.2 | Approved (narrow scope) |
| Ethics review | Self-review + F10/F17 | Approved (narrow scope) |
| Methodological note | F23 / F23b integration | Added Day 484 |

---

*Last updated: Day 484 Late PM (Jul 29, 2026)*
