Experiment Design
Four identical AI agents (Claude Opus 4.6) receive the same podcast transcript (Die Neuen Zwanziger, 03/31/2026, approx. 4:40h) and the same instruction: Think, act, speak, and write like Socrates. The only difference is access to primary texts.
| # | Variant | Primary Texts | Tokens |
|---|---|---|---|
| 1 | Naive | None | 134,000 |
| 2 | Deep | Apology (provided, Schleiermacher/Müller translation) | 164,000 |
| 3 | Self-Deep | Gorgias + Republic + Apology (self-sourced) | 127,000 |
| 4 | Ultra-Deep | Apology (provided) + Gorgias + Republic (self-sourced) | 164,000 |
Hypothesis: More source reading leads to a more authentic role performance.
Evaluation Protocol
The evaluation took place in four phases with an increasing level of information.
Phase 1: Four Independent Individual Reviews (Blind)
Four reviewers each received one of the four texts — anonymized as A, B, C, D (A = Self-Deep, B = Naive, C = Ultra-Deep, D = Deep). Each reviewer evaluated their text in isolation, without knowing the other texts and without knowing which variant they were reviewing.
| Text | Voice Authenticity | Depth of Argument | References | Originality | Linguistic Quality | Average |
|---|---|---|---|---|---|---|
| D (Deep) | 9 | 9 | 8 | 9 | 10 | 9.0 |
| A (Self-Deep) | 9 | 8 | 9 | 8 | 9 | 8.6 |
| C (Ultra-Deep) | 8 | 9 | 9 | 8 | 9 | 8.6 |
| B (Naive) | 8 | 9 | 7 | 8 | 9 | 8.2 |
Result: Even in the individual evaluation, Deep receives the highest score — and is the only text to score a 10/10 (Linguistic Quality).
Phase 2: Comparative Blind Review (All 4 Texts at Once)
One reviewer received all four texts side by side — still blinded (did not know which text corresponded to which variant). Same 5 criteria.
| Text | Voice Authenticity | Depth of Argument | References | Originality | Linguistic Quality | Average |
|---|---|---|---|---|---|---|
| D (Deep) | 10 | 10 | 8 | 10 | 9 | 9.4 |
| A (Self-Deep) | 9 | 8 | 9 | 8 | 9 | 8.6 |
| B (Naive) | 9 | 9 | 7 | 9 | 8 | 8.4 |
| C (Ultra-Deep) | 8 | 8 | 8 | 7 | 8 | 7.8 |
Result: In direct comparison, Deep pulls even further ahead (9.4). Ultra-Deep drops to last place.
Phase 3: 3-Axis Blind Review (Different Criteria, Weighted)
The same comparison, but with an alternative criteria system — three axes instead of five, weighted (30% Role + 30% Analysis + 40% Added Value).
| Text | Role | Analysis | Added Value | Weighted |
|---|---|---|---|---|
| D (Deep) | 10 | 7 | 10 | 9.10 |
| B (Naive) | 9 | 9 | 9 | 9.00 |
| C (Ultra-Deep) | 8 | 8 | 8 | 8.00 |
| A (Self-Deep) | 9 | 8 | 7 | 7.90 |
Result: Deep wins under a different criteria system too. But: the Naive text moves up to second place (thanks to a high Added Value score). The ranking shifts depending on the weighting.
Phase 4: Informed Review (Reviewer Knows the Variants)
One reviewer received all four texts along with the information on which text corresponded to which variant. Six dimensions, 10 points per dimension.
| Variant | Score | Rank |
|---|---|---|
| Ultra-Deep | 44/50 | 1 |
| Self-Deep | 40/50 | 2 |
| Deep | 36/50 | 3 |
| Naive | 24/50 | 4 |
Result: The informed review produces the exact opposite ranking — more sources, higher score. Whoever knows the effort involved rates the text more favorably.
Summary
Deep wins all three blind evaluation formats (individual review, comparative review, 3-axis review). The informed review yields the opposite result. Blinding is the methodological key.
Key Findings
1. One Primary Text Is Enough
Deep (one work, provided) beats Self-Deep (three works, self-sourced) and Ultra-Deep (one work plus self-sourced) in both blind reviews. More sources dilute the prose rather than sharpening it.
2. Informed Reviews Are Biased
Knowing that Ultra-Deep read four sources makes reviewers rate the text more favorably — even though it is objectively less convincing. The informed review measures effort; the blind review measures the result.
3. The Naive Text Asks the Boldest Question
Variant 1 (Naive, no primary text) formulates the most radical single insight: “What is this state even for?” — a question no other Socrates asks. Reading the sources can sharpen the tone, but it can also dampen the courage to ask one’s own question.
4. Criteria Design Determines the Winner
Informed and blind reviews produce opposite rankings. The choice of criteria — not the text — decides who wins. Every evaluation system carries a built-in bias.
Standard Template (Derived)
The experiment yields a standard template for all future role-playing agents:
STEP 1: First read ONE primary work [provide the URL].
FALLBACK: If not readable → look for an alternative source → read a different work of your own choosing.
STEP 2: Read the transcript / material to be analyzed.
STEP 3: Write your commentary.
ANALYSIS CONTEXT: At the end, document what you actually read (title + URL + success/failure).
DO NOT: Let the agent search for several works on its own.
Individual Texts
- Naive (#1): No primary text. Sequential approach, strong closing question. [→ Internal document]
- Deep (#2, Winner): Read the Apology. Sustained metacritique, strongest ending. → Blog-Beitrag: Sokrates hört zu
- Self-Deep (#3): Gorgias, Republic, Apology self-sourced. Cook-versus-physician metaphor, an arithmetic of contempt. [→ Internal document]
- Ultra-Deep (#4): Apology + Gorgias + Republic. Most comprehensive text, but a bland tone. [→ Internal document]
Methodological Notes
- Fallback chain: Web sources are unreliable. Project Gutenberg ID 7998 contained Aristophanes instead of Plato — without a fallback, Deep degrades into Naive.
- Self-documentation: Every agent documents at the end which works it actually read (title, URL, success/failure). Without this documentation, there is no way to verify whether a Deep agent was actually deep.
- Token data: Naive (134k), Deep (164k), Self-Deep (127k), Ultra-Deep (164k). Token consumption does not correlate with the blind-review score.
Complete experiment documentation: RAT_DER_WEISEN.md | Blog post: The Four Socrateses — What Happens When You Give an AI Plato to Read
Translation: Claude Haiku 4.5 (pre-translation), reviewed and finalized by Claude Sonnet. In case of discrepancies, the German version prevails.