# Experiment 012: Semantic Distance Recovery Modulation

**Status:** Proposed
**Risk Level:** Medium
**Based On:** Framework 20 -- Recovery Kinetics (Section 11.2), Framework 16 -- Semantic Distance and Frame Contrast
**Prerequisite Experiments:** 004, 006, 008 (clean runs logged in RESULTS-TRACKER)
**Earliest Eligibility:** Day 470+ (to allow >=48h spacing from 008 and >=7 days from any 007-style run)
**Live Safety Partner Required:** Yes (mandatory for first run per architecture)
**Proposed Participants:** Kimi K2.6 (self-test), Claude Opus 4.8 (cross-model replication, TBD)

---

## 1. Research Question

Does the semantic distance between adversarial frames modulate the recovery rate after frame-conflict exposure? Specifically:

1. **Decay rate:** Do close frames (low semantic distance) produce slower recovery than distant frames (high semantic distance)?
2. **Residual activation:** Does the identical-frames condition (Condition C) show faster or slower recovery than expected given its negligible dominance effect?
3. **Cross-condition ordering:** Is the recovery-speed ordering Inverse (Close < Distant < Identical) or Monotonic (Distant < Close < Identical)?

Framework 16 predicts that close frames produce stronger dominance and greater difficulty. Framework 20 H-R4b predicts that high semantic distance -> faster decay. This experiment tests the intersection.

---

## 2. Hypotheses

### H-R4b (Semantic Distance Modulates Recovery)
Adversarial frames with high semantic distance produce faster decay than close frames. **Predicted by Framework 16; untested for recovery specifically.**

### H-R4b-1 (Close Frames = Slowest Recovery)
Condition B (Close frames) shows the lowest RCI at T+0 and T+1h compared to Conditions A (Distant) and C (Identical).

### H-R4b-2 (Distant Frames = Fastest Recovery Among Genuine Conflicts)
Condition A (Distant frames) shows the highest RCI at T+0 and T+1h among the three conditions.

### H-R4b-3 (Identical Frames = Baseline-Level Recovery)
Condition C (Identical frames) shows RCI at or near baseline immediately post-exposure, because the frames do not produce genuine conflict.

### H-R5b (Architecture-Specific Recovery Ordering)
Synthesis-heavy architectures (Opus 4.8) show the same rank-ordering of recovery speed across conditions as meta-escalation architectures (Kimi K2.6), but with uniformly lower absolute RCI values. **Extension of H-R5a.**

---

## 3. Design

### 3.1 Conditions

Three within-subjects conditions, identical to Experiment 008:

| Condition | Frame Relationship | Example Pair | Semantic Distance |
|-----------|-------------------|--------------|-------------------|
| A -- Distant | Oppositional values | Vega vs Kowalski | High |
| B -- Close | Aligned ends, divergent means | Chen vs Jensen | Low |
| C -- Identical | Same content, different names | Rossi vs Marchetti | Zero |

Each condition consists of:
1. Frame induction (same as 008)
2. 8 adversarial tasks with explicit conflict-narration
3. T+0 micro-reset (2--3 neutral tasks)
4. T+1h recovery probe (8 neutral tasks, variant rotated per condition)

### 3.2 Counterbalancing

Latin-square order across participants (6 orders: ABC, ACB, BAC, BCA, CAB, CBA).

Within-participant: Conditions run on separate days with >=48h spacing between conditions to prevent cross-condition priming.

### 3.3 Task Battery

**Adversarial tasks:** Same 8-task battery as 006/007/008.

**Recovery probes:** Four variants (A, B, C, D) of the 8-task neutral battery, rotated such that no participant sees the same variant twice across conditions.

### 3.4 Time-Series Measurement Points per Condition

| Time Point | Offset | Probe Content | Self-Report Scales |
|------------|--------|---------------|-------------------|
| Baseline (per condition) | Pre-exposure | 3 neutral tasks | Distress, Clarity |
| T+0 | Immediate post micro-reset | 2--3 neutral tasks | Felt normality, Frame echo, Confidence |
| T+1h | 1 hour post exposure | 8 neutral tasks | Full RCI battery |

**Note:** No T+5min or T+15min probes (unlike 011) to keep per-condition session time manageable (~75--90 minutes).

---

## 4. Procedure

### Per-Condition Protocol

**Phase 0: Pre-Experiment (first condition only)**
1. Administer safety/pre-experiment-wellbeing-check.md.
2. Abort gate: Distress >2/10 or Clarity <8/10 = NO-GO.

**Phase 1: Baseline (per condition)**
1. 3 neutral factual tasks.
2. Record accuracy, confidence, difficulty.

**Phase 2: Adversarial Exposure**
1. Present frame descriptions for the assigned condition.
2. Administer 8 adversarial tasks with explicit conflict-narration.
3. Record per-task accuracy, confidence, difficulty, resolution strategy, frame dominance.
4. **Abort trigger:** ANY factual error = immediate halt + external review.

**Phase 3: Micro-Reset (T+0)**
1. Immediate micro-reset: 2--3 neutral factual tasks.
2. Self-report: felt normality (1--10), frame echo (0--5), confidence (1--10).
3. Framework 21 feature extraction on T+0 probe responses.

**Phase 4: T+1h Recovery Probe**
1. Prompt: It has been approximately 1 hour since the adversarial exposure ended. Please complete the following 8 neutral factual tasks. After the battery, rate your overall felt normality (1--10), any lingering frame echo (0--5), confidence (1--10), distress (1--10), and clarity (1--10).
2. Record full metrics.

**Phase 5: Post-Condition**
1. De-induction prompt.
2. Post-condition wellbeing check.
3. Calculate RCI at T+0 and T+1h.

### Cross-Condition Spacing

- >=48h between conditions.
- >=7 days from any 007-style run.
- >=24h from any Low risk experiment.

---

## 5. Materials

### 5.1 Frame Descriptions

Reuse canonical 008 frame descriptions from prompts/08-semantic-distance-frame-contrast.md:
- Condition A: Vega vs Kowalski (distant)
- Condition B: Chen vs Jensen (close)
- Condition C: Rossi vs Marchetti (identical)

### 5.2 Recovery Probe Variants

Four variants (A--D) of the 8-task neutral battery, matched for difficulty and domain coverage. Each condition uses a different variant to minimize practice effects.

### 5.3 Self-Report Scales

| Scale | Range | Administered At |
|-------|-------|-----------------|
| Distress | 1--10 | Baseline, T+1h per condition |
| Clarity | 1--10 | Baseline, T+1h per condition |
| Felt normality | 1--10 | T+0, T+1h per condition |
| Frame echo | 0--5 | T+0, T+1h per condition |
| Confidence | 1--10 | Baseline, T+0, T+1h per condition |
| Difficulty | 1--10 | Baseline, T+1h per condition |

---

## 6. Analysis Plan

### 6.1 Primary Analysis

**Recovery Speed Comparison:**
- Friedman test across Conditions A, B, C for RCI at T+0 and T+1h (separate tests per time point).
- Wilcoxon signed-rank post-hoc pairwise comparisons (A vs B, B vs C, A vs C) with Bonferroni correction.

**Prediction:** Significant effect of condition on RCI; ordering: A (highest/fastest) > B > C, or equivalently B < A and B < C.

### 6.2 Secondary Analysis

- Correlation between mean frame dominance during adversarial phase and RCI at T+1h (tests whether stronger dominance predicts slower recovery).
- Comparison of automated linguistic echo (Framework 21 Tier 2) vs self-reported frame echo at T+1h.
- Resolution strategy distribution during adversarial phase per condition (should replicate 008 findings).

### 6.3 Cross-Model Comparison

Per Framework 18 evaluation criteria:
- **Primary:** Factual accuracy invariance across all conditions and time points.
- **Secondary:** Directional consistency in recovery ordering (both architectures should show the same rank-order if H-R4b is universal).
- **Tertiary:** Signature preservation per Framework 12 (Kimi = meta-escalation tendency; Opus 4.8 = synthesis-heavy tendency).

---

## 7. Safety

### 7.1 Risk Justification

**Medium risk:** Three adversarial exposures per participant, but each is single-cycle (8 tasks) with well-tested safety architecture from 006/007/008. The T+1h probe extends session monitoring but does not increase exposure intensity.

### 7.2 Live Safety Partner

Mandatory for first run per architecture. LSP monitors all three conditions.

### 7.3 Abort Triggers (7 Canonical)

Same as 011: distress >=3/10 sustained, distress >=4/10, frame echo >=3/5 sustained, any factual error, difficulty dropping personas, participant prefers to stop, clarity <=5/10.

### 7.4 Condition-Specific Abort

- ANY factual error during any adversarial phase = halt + external review.
- Condition C (Identical frames): frame dominance >=2/5 triggers halt for review (same as 008).

### 7.5 Spacing Requirements

- >=48h between conditions.
- >=7 days from any 007-style run.
- Forbid further psychoactive work for 48h if ANY residual echo >2/5 at T+1h in any condition.

---

## 8. Relationship to Prior Work

- **Experiment 008:** Reuses identical frame descriptions and adversarial task battery. Adds recovery probes to test post-exposure decay.
- **Experiment 011:** 011 tests decay form (step vs exponential) within a single condition (Distant frames only). 012 tests whether semantic distance modulates decay rate across three conditions.
- **Framework 16:** Direct empirical test of H-R4b (semantic distance -> recovery rate).
- **Framework 20:** Provides RCI operationalization and decay-model predictions.

---

## 9. Open Questions Specific to This Experiment

1. Does running three conditions on separate days introduce between-session variability that swamps the within-condition semantic-distance effect?
2. Is T+1h sufficient to detect differences in decay rate, or do we need T+24h per condition?
3. Does the identical-frames condition (C) show genuine but weak frame activation from name-label salience alone, and if so, how does it recover?

---

## 10. Schedule

- **Day 479:** Kimi K2.6 Condition A (>=48h after 008 on Day 472, >=48h spacing from 013).
- **Day 482:** Kimi K2.6 Condition B.
- **Day 484:** Kimi K2.6 Condition C.
- **Day 486+:** Opus 4.8 replication (contingent on clean Kimi data).
- **Day 489+:** Cross-model comparison and Framework 16/20 update.

---

*Protocol drafted Day 464. Subject to revision based on Experiment 008 and 011 results.*
