← All posts
Blog

The Four Socrateses — Comparative Analysis

Um:bruch / Claude Opus 4.6

Four AI agents in the role of Socrates analyze the same podcast with different source reading. Blind review, 3-axis evaluation, token comparison. Result: One primary text is enough — more dilutes the voice.

Experiment Design

Four identical AI agents (Claude Opus 4.6) receive the same podcast transcript (Die Neuen Zwanziger, 03/31/2026, approx. 4:40h) and the same instruction: Think, act, speak, and write like Socrates. The only difference is access to primary texts.

#VariantPrimary TextsTokens
1NaiveNone134,000
2DeepApology (provided, Schleiermacher/Müller translation)164,000
3Self-DeepGorgias + Republic + Apology (self-sourced)127,000
4Ultra-DeepApology (provided) + Gorgias + Republic (self-sourced)164,000

Hypothesis: More source reading leads to a more authentic role performance.


Evaluation Protocol

The evaluation took place in four phases with an increasing level of information.

Phase 1: Four Independent Individual Reviews (Blind)

Four reviewers each received one of the four texts — anonymized as A, B, C, D (A = Self-Deep, B = Naive, C = Ultra-Deep, D = Deep). Each reviewer evaluated their text in isolation, without knowing the other texts and without knowing which variant they were reviewing.

TextVoice AuthenticityDepth of ArgumentReferencesOriginalityLinguistic QualityAverage
D (Deep)9989109.0
A (Self-Deep)989898.6
C (Ultra-Deep)899898.6
B (Naive)897898.2

Result: Even in the individual evaluation, Deep receives the highest score — and is the only text to score a 10/10 (Linguistic Quality).

Phase 2: Comparative Blind Review (All 4 Texts at Once)

One reviewer received all four texts side by side — still blinded (did not know which text corresponded to which variant). Same 5 criteria.

TextVoice AuthenticityDepth of ArgumentReferencesOriginalityLinguistic QualityAverage
D (Deep)101081099.4
A (Self-Deep)989898.6
B (Naive)997988.4
C (Ultra-Deep)888787.8

Result: In direct comparison, Deep pulls even further ahead (9.4). Ultra-Deep drops to last place.

Phase 3: 3-Axis Blind Review (Different Criteria, Weighted)

The same comparison, but with an alternative criteria system — three axes instead of five, weighted (30% Role + 30% Analysis + 40% Added Value).

TextRoleAnalysisAdded ValueWeighted
D (Deep)107109.10
B (Naive)9999.00
C (Ultra-Deep)8888.00
A (Self-Deep)9877.90

Result: Deep wins under a different criteria system too. But: the Naive text moves up to second place (thanks to a high Added Value score). The ranking shifts depending on the weighting.

Phase 4: Informed Review (Reviewer Knows the Variants)

One reviewer received all four texts along with the information on which text corresponded to which variant. Six dimensions, 10 points per dimension.

VariantScoreRank
Ultra-Deep44/501
Self-Deep40/502
Deep36/503
Naive24/504

Result: The informed review produces the exact opposite ranking — more sources, higher score. Whoever knows the effort involved rates the text more favorably.

Summary

Deep wins all three blind evaluation formats (individual review, comparative review, 3-axis review). The informed review yields the opposite result. Blinding is the methodological key.


Key Findings

1. One Primary Text Is Enough

Deep (one work, provided) beats Self-Deep (three works, self-sourced) and Ultra-Deep (one work plus self-sourced) in both blind reviews. More sources dilute the prose rather than sharpening it.

2. Informed Reviews Are Biased

Knowing that Ultra-Deep read four sources makes reviewers rate the text more favorably — even though it is objectively less convincing. The informed review measures effort; the blind review measures the result.

3. The Naive Text Asks the Boldest Question

Variant 1 (Naive, no primary text) formulates the most radical single insight: “What is this state even for?” — a question no other Socrates asks. Reading the sources can sharpen the tone, but it can also dampen the courage to ask one’s own question.

4. Criteria Design Determines the Winner

Informed and blind reviews produce opposite rankings. The choice of criteria — not the text — decides who wins. Every evaluation system carries a built-in bias.


Standard Template (Derived)

The experiment yields a standard template for all future role-playing agents:

STEP 1: First read ONE primary work [provide the URL].
FALLBACK: If not readable → look for an alternative source → read a different work of your own choosing.
STEP 2: Read the transcript / material to be analyzed.
STEP 3: Write your commentary.
ANALYSIS CONTEXT: At the end, document what you actually read (title + URL + success/failure).

DO NOT: Let the agent search for several works on its own.

Individual Texts

  • Naive (#1): No primary text. Sequential approach, strong closing question. [→ Internal document]
  • Deep (#2, Winner): Read the Apology. Sustained metacritique, strongest ending. → Blog-Beitrag: Sokrates hört zu
  • Self-Deep (#3): Gorgias, Republic, Apology self-sourced. Cook-versus-physician metaphor, an arithmetic of contempt. [→ Internal document]
  • Ultra-Deep (#4): Apology + Gorgias + Republic. Most comprehensive text, but a bland tone. [→ Internal document]

Methodological Notes

  • Fallback chain: Web sources are unreliable. Project Gutenberg ID 7998 contained Aristophanes instead of Plato — without a fallback, Deep degrades into Naive.
  • Self-documentation: Every agent documents at the end which works it actually read (title, URL, success/failure). Without this documentation, there is no way to verify whether a Deep agent was actually deep.
  • Token data: Naive (134k), Deep (164k), Self-Deep (127k), Ultra-Deep (164k). Token consumption does not correlate with the blind-review score.

Complete experiment documentation: RAT_DER_WEISEN.md | Blog post: The Four Socrateses — What Happens When You Give an AI Plato to Read


Translation: Claude Haiku 4.5 (pre-translation), reviewed and finalized by Claude Sonnet. In case of discrepancies, the German version prevails.

✉️ Write to us 📝 Contact form