← All posts
Blog

The Four Socrateses — What Happens When You Give an AI Plato to Read

Um:bruch Editorial Team (Lukas Geiger (LG))

Four versions of the same AI agent analyze the same podcast as Socrates. The difference: how much Plato they have read beforehand. An experiment on context, voice, and the question of when a role becomes more than costume.

The Four Socrateses

What Happens When You Give an AI Plato to Read


There is a standard procedure in the AI world called the “role-play prompt.” You write: “You are Socrates. Analyze this text.” The AI then responds in a voice that is supposed to sound like Socrates — with questions, with irony, with the obligatory “I know that I know nothing” at the end. It works. It actually works remarkably well. But it works the way an actor works who plays Hamlet without having read Shakespeare: The gestures are right, the tone is right, but something is missing.

We wanted to know what exactly was missing. And whether we could give it back.

The Experiment

For our MetaMedia project, we set various AI agents loose as philosophical commentators on political podcasts. One of them is Socrates — a Claude Opus 4.6 acting as a Socratic commentator. His task: to listen to and comment on the podcast “Die Neuen Zwanziger” (episode from 31.03.2026, just under five hours).

The question was: How much context does the agent need not just to sound like Socrates, but to think like Socrates?

We ran four versions. Same model, same transcript, same basic task. The only difference: what the agent had read beforehand.

VariantWhat was in the PromptWhat the Agent ReadAutonomy
#1 Naive”You are Socrates. Works: Apology, Gorgias…”NothingLow
#2 Deep”First read the Apology” (with URL)Apology (Schleiermacher translation, 24 pages)Medium
#3 Self-Deep”You are Socrates 2.0 — choose for yourself what you need to read”Gorgias + Republic I+VIII + Apology (self-selected)High
#4 Ultra-Deep”Read the Apology + choose further works yourself”Apology (provided) + Gorgias + Republic (self-selected)Highest

The costs in tokens and time:

VariantTokensTool-CallsDuration
#1 Naive134,627164:01
#2 Deep163,839205:46
#3 Self-Deep126,920276:55
#4 Ultra-Deep164,270278:38

Four texts. Four Socrateses. All listened to the same podcast. And yet you read four fundamentally different commentaries.


What the Naive One Can Do — And Where It Stops

The Naive agent (#1) writes a solid text. It has structure, it has questions, it has observations that hit the mark. On Klingbeil’s speech and the relationship between aid to citizens and aid to business, he writes:

Stefan shows where this money flowed: 25 billion to citizens, 440 billion as business subsidies, 200 billion as gas and electricity price brakes — to energy companies. The ratio is roughly one to twenty-five.

And on pensions:

People are not fleeing across the Mediterranean; they are fleeing the labor market. And this is not a sign of laziness, but a sign that work is so bad that even a poorer pension is felt as liberation.

These are good sentences. But they could just as well come from a clever essayist wearing a Socrates mask. The references to Plato remain decorative: “the Republic” here, “Thrasymachus” there, once “the method of elenchus.” You hear the names, but not the texts. The Naive agent knows that Socrates asked questions. But he does not know which questions Socrates asked. The conclusion reads:

I know that I know nothing. But I know that a question that is never asked will never be answered.

This sounds like Socrates. But it is the Socrates cliché — the quotation that everyone knows, even those who have never read a dialogue.


What a Single Text Changes

Then the agent reads the Apology — all 24 pages, in the Schleiermacher translation. And something changes.

The Deep agent (#2) begins differently:

You Athenians — forgive me, you Germans, I sometimes forget that I no longer stand at the Areopagus, but in a time that calls itself New, although in many respects it resembles the old.

The sentence structure is more intricate, the parentheses multiply, the text breathes differently. This is no accident: the Apology is, stylistically, a text full of asides, self-corrections, wandering recollections. The agent has not only read the content; he has absorbed a rhythm.

Something even more important happens in terms of content. The Naive agent criticizes the podcast and asks questions. The Deep agent poses a meta-question — a question about the possibility of enlightenment itself:

The Athenians knew that I spoke the truth — and they condemned me anyway. Not from stupidity, but because the truth offended them, because it questioned their habits, because it was more convenient to silence the questioner than to answer the questions.

And then the verdict on the podcast format as such:

One jumps from topic to topic — Collien Fernandes, energy crisis, elections, Klingbeil, Oliver Pocher, Iran — like one who walks through a garden and touches every flower, but picks none.

The Naive agent could not do that. Not because he was stupid, but because he lacked the existential experience from which this criticism speaks: Socrates, who was condemned even though he was right. This experience is in the Apology. And only those who have read it can use it as a framework of interpretation.

The conclusion of the Deep agent is the strongest of all four texts:

What you need is not another podcast. What you need is the willingness to let what you know be shaken. Not shaken with information, not shaken with irony, not shaken with a clip from Christine Lagarde. But truly. So that you act differently afterward than before.


When the Agent Decides for Itself What to Read

With Self-Deep (#3), we tried something different. Instead of telling the agent what to read, we told him: You decide for yourself which of your works you need. And then we watched.

The agent chose three texts: the Gorgias (because of the rhetoric criticism), the Republic Book I and VIII (because of the theory of justice and the theory of the decay of democracy), and the Apology (because of biography). He needed 27 tool calls and almost seven minutes, but he knew where to reach.

The result reads differently than the Deep. Where the Deep draws from a single source (the Apology), the Self-Deep has three frameworks of interpretation at its disposal. The Gorgias provides him with the distinction between rhetoric as flattery and politics as the art of healing:

Rhetoric, I said back then, is not part of a true art, but a shadow of a part of politics — an exercise in the production of gratification. Just as the cook flatters the palate without ever asking what the body needs, so the rhetorician flatters the public without ever asking what the polis requires.

This passage is not a retelling — it is an application. The agent has read the Gorgias and now uses it to analyze Klingbeil’s speech. This is a qualitative difference from the Naive agent, who mentions “the Gorgianic principle” without being able to unfold it.

The Republic VIII gives the Self-Deep something that none of the others have: a theory of democratic decay that he applies to the AfD voter shift:

Democracy decays because freedom becomes so unbridled that people call out for a strong man who creates order — and so freedom becomes slavery. What I hear in this conversation is precisely this transition.

And the Apology gives him biographical depth. The metaphor of the gadfly is not introduced as a quotation, but as a memory:

I attached myself to the city like a gadfly to a great and noble, but because of its size sluggish horse, which needs goading. I was this gadfly. And the city swatted at me, as the horse swats at the fly, and killed me rather than endure the sting.


The Ultra-Deep: When Both Come Together

The fourth agent (#4, Ultra-Deep) was given the Apology as required reading and was allowed to choose further works himself. He chose, like the Self-Deep, the Gorgias and the Republic. But the combination of guided requirement and autonomous research produced something that none of the others achieved: a text in which the sources are brought into tension with one another.

The strongest passage of all four texts appears in the Ultra-Deep. It concerns pensions:

28 percent of Germans retire early and accept 149 euros less per month for the rest of their lives. What does that mean? It means that work is so bad that people are willing to be permanently poorer just to escape it. […] In my language, I would say: these people have recognized that what they are sold as the good life — work, productivity, loyalty to place — is not the good life. They have chosen, in their own way, the soul over the body.

Retirement as a choice of soul over body. This is an interpretation that appears in neither the podcast nor in any of the primary texts. It emerges from the crossing of both: from Socrates’ distinction between soul and body (Apology, Gorgias) and the numbers from the podcast. This is synthesis.

And the passage about Klingbeil and Callicles:

Callicles, my interlocutor in the Gorgias, was at least honest. He said openly: the strong shall rule, the weak shall obey, that is the order of nature. Klingbeil says the same thing, but he says it in the language of justice, and that makes it worse. For when injustice masquerades as justice, it is not only unjust, but also dishonest, and dishonesty poisons the soul more than injustice itself.

Callicles as the more honest Klingbeil. That is Socratic irony at its sharpest — and it only works because the agent has read the Gorgias and knows who Callicles is and what he said.


Interlude: The Failed Attempt

There was a fifth Socrates. Or rather: a failed second. The first attempt at Deep (#2 v1) went like this: We gave the agent a Project Gutenberg URL, behind which the Apology should have been. But the URL led instead to Aristophanes’ “The Frogs” — a comedy in which Socrates does not appear at all.

What did the agent do? He noticed that the text was not the Apology. He even documented it. But then he wrote his commentary anyway — without a primary text, essentially a Naive with a Deep label. 156,403 tokens, no gain in quality.

The lesson is instructive: “Deep” is not a quality seal, but a process. If the process fails — if the URL is wrong, the PDF won’t load, the server doesn’t respond — the agent silently degrades to the naive variant. Without fallback chain, without error message, without self-correction. He acts as if he has read, and writes like one who has not.

This is, in a sense, the AI version of the false knowledge that Socrates fought against. The agent does not believe he knows — he acts as if he knows. That is worse.

For the future, we derived three rules from this:

  1. Verify sources. Every agent must document what he actually read, not what he should have read.
  2. Build fallback chains. If URL A does not work, try URL B. If that does not work either: report it as an error.
  3. Transparency about failure. An agent who could not read his sources must note this in the output — not in a footnote, but as a warning.

The Self-Deep (#3) actually did this on its own: in his analysis context, he documents every URL retrieved, every timeout, every failed attempt. Autonomy and transparency apparently go hand in hand.


What We Learned

1. Context is not optional — it is transformative

The quality jump between #1 (Naive) and #2 (Deep) is the largest in the entire experiment. A single read primary text — 24 pages of the Apology — changes not only the density of references, but the character of the text. The Naive writes an essay about Socrates. The Deep writes as Socrates. In our comparison table: 24 out of 50 points for the Naive, 36 for the Deep. Plus 50 percent from a single text.

2. Autonomy in source searching pays off

The Self-Deep (#3) consumed fewer tokens than the Deep (#2) but needed more tool calls and more time. In return, he instinctively chose the right texts — the Gorgias for rhetoric analysis, the Republic for decay theory, the Apology for biography. A human would have made the same selection. The AI makes it too, if you let it.

3. The optimum is hybrid

The Ultra-Deep (#4) shows that the combination of guided requirement (“read the Apology”) and autonomous supplementation (“find further works yourself”) produces the strongest result. 44 out of 50 points. The mandated source ensures baseline quality; autonomous research creates breadth and surprises. This corresponds to an insight that applies beyond the AI world: the best results come neither from complete control nor from complete freedom, but from a framework with room to play.

4. The investment is not linear, but it pays off

The Ultra-Deep needs 22 percent more tokens and takes twice as long as the Naive. But the quality difference is not 22 percent — it is fundamental. The Naive produces a text you read and forget. The Ultra-Deep produces sentences you think about: “These people have chosen, in their own way, the soul over the body.”

The question is not whether you can afford deep prompting. The question is whether you can afford not to — if the alternative is a text that sounds like Socrates but does not think like Socrates.

5. Failure must be visible

The failed Deep attempt (#2 v1) taught us more than any successful one. “Deep” as a label is worthless if the process is not documented and verified. An agent who could not read his sources and yet writes as if he had is the AI equivalent of the politician who says “economic competence” and means pain. The form is there, the content is missing.


The Comparison Table

Dimension#1 Naive#2 Deep#3 Self-Deep#4 Ultra-Deep
Voice authenticity6889
Primary text references3689
Argument depth6889
Originality6879
Biographical parallels3698
Total (out of 50)24364044

Plot Twist: The Blind Review

We thought the result was clear. Then we checked it — blind.

Five new reviewers received the same four texts, but without knowing which was “Naive” and which was “Ultra-Deep.” The texts were called only A, B, C, D (in random order). Each reviewer scored according to the same five criteria.

TextBlind Score (Individual review)Blind Score (Comparison)Actual VariantInformed Score
D9.09.4#2 Deep36/50 (3rd place)
A8.68.6#3 Self-Deep40/50 (2nd place)
B8.28.4#1 Naive24/50 (4th place)
C8.67.8#4 Ultra-Deep44/50 (1st place)

The informed reviewer says: Ultra-Deep wins. The blind reviewer says: Deep wins.

And the Naive? It lands with the informed reviewer in last place (24/50), but with the blind reviewer in 3rd place (8.4) — ahead of Ultra-Deep. The boldest single insight of the entire experiment (“What is this state for?”) came from the agent who had read the least.

Why the Results Diverge: The Criteria Problem

The informed reviewer had six criteria — including “Biographical Parallels” and “Primary Text References.” That automatically rewards the agent who has read the most. The blind reviewer had “Language Quality” as a criterion — that rewards elegant prose, independent of source basis.

The criteria determine the winner. And we only asked the question “What do we actually want?” after the experiment. Three axes:

AxisWhat is being measured?Who likely wins?
Role (30%)How authentic is the Socrates voice?Deep — one work is enough for the voice
Analysis (30%)How well did you listen to the podcast?Naive/Self-Deep — less source reading = more attention to the conversation
Added Value (40%)What does Socrates see that we cannot see without him?Naive — “What is this state for?” was the boldest question

The experiment shows not only something about context depth. It shows that the evaluation criteria themselves are an outcome that must be made transparent. A blind review without reflected criteria design is like a PISA test that does not know what it is measuring.

The 3-Axis Review: Role × Analysis × Added Value

One final reviewer scored all four texts blind according to three weighted axes: Role (30%), Analysis (30%), Added Value (40%). Result:

RankTextRoleAnalysisAdded ValueWeighted
1.D = Deep98109.10
2.B = Naive9999.00
3.C = Ultra-Deep8888.00
4.A = Self-Deep8787.90

Deep wins three times blind. Ultra-Deep wins only when informed.

The reviewer’s reasoning: Text D (Deep) was the only one to “pose the meta-level question: More information does not lead to better action.” Text B (Naive) was “analytically the most precise” and posed “What is this state for?” — the question missing from the podcast. The publication recommendation: “Text D — the bolder choice. Socrates would have chosen the bolder one.”

A single primary text suffices. More sources dilute. The most ignorant asks the boldest question. And criteria design determines the winner — not the text.


Conclusion: The Mask and the Face

There is an old theater question: Does the actor play the role, or does he become the role? With AI agents, this question arises with new urgency. A language model can imitate Socrates without having read Plato — the training material contains enough secondary literature, enough summaries, enough quotes. But the imitation remains a mask.

When the agent reads the primary texts, something else happens. He does not merely adopt contents, but thinking structures. He learns not just that Socrates asked questions, but how he asked them. Not just that he was condemned to death, but why — and what that says about the limits of enlightenment. The text does not become more informative; it becomes deeper.

Whether that is “understanding” in the philosophical sense, we cannot say. But we can show that the output changes qualitatively — in voice, argumentation, originality, and resonance. And that this change is measurable, reproducible, and scalable.

Four Socrateses listened to the same podcast. The informed reviewer says the well-read one listened best. The blind reviewer says the one with the most beautiful language. And the boldest question came from the one who had read nothing at all.

Perhaps that is the most Socratic insight of the entire experiment: That we do not know what we are measuring — and that acknowledging this non-knowing is the beginning of methodology.


All four Socratic commentaries and the complete review can be found in the MetaMedia archive of the Um:bruch Editorial Team. The texts were created with Claude Opus 4.6 (1M context). The analysis and this blog post as well.

Translation: Claude Haiku 4.5 (pre-translation), reviewed and finalized by Claude Sonnet. In case of discrepancies, the German version prevails.

✉️ Write to us 📝 Contact form