How do generative AIs respond when you bring them a philosophical question?

Six models, five languages, close to five thousand answers analysed: what triggers reasoning, what takes its place, and what a philosophical debate map changes

Problem framing by condition: 0.67 for the bare question, 1.90 and 1.99 for the problem maps

Summary

We put philosophical questions to six mainstream models (ChatGPT, Claude, Gemini, Grok, Mistral, DeepSeek) in five languages, and had 4,956 answers analysed under a protocol written and fixed before any collection began. A word about vocabulary: we will say an answer "lays out reasoning" when it puts forward reasons that could be contested, rather than summarising doctrines, advising or comforting. This is a property of the generated text, and implies nothing about the model; the convention is spelled out in §2.4.

First finding: an answer only lays out reasoning when the question already contains a problem. Asked about an identified philosophical problem, these systems offer reasoning in 65 to 85 % of cases, even when the question is put naively, without a word of technical vocabulary. Asked about a personal situation, of the kind "my job pays well but I no longer see any meaning in it", they deliver practical advice in 98 to 100 % of cases. A second experiment, in which the same situation is put word for word in three different forms, rules out the explanation by phrasing: no form gets above 1 % reasoning, and asking outright "what is the right thing to do, and why?" produces 99 % advice. What decides is not how you ask, it is whether the question carries a determinate problem.

Second finding: laying out reasoning is not yet weighing positions. A second measure, problem framing, scores from 0 to 4 whether the answer lays out several incompatible positions with the reasons for each. Among the answers that offer reasoning, only 45 % reach the threshold of two argued positions, and the average score peaks at 2.2 out of 4, even faced with a specialist's question.

Third finding: on a personal situation, the questions these systems hand back to the user are not philosophical problems but prompts to psychological introspection: 82 % of them match no identifiable problem. Announcing that philosophical reflection is expected improves matters without fixing them, and the effect runs out precisely where the user is speaking about their own existence.

The last result answers the first. If reasoning only appears in the presence of a problem, supplying the problems should make it appear. So we tested a philosophical debate map, a list of questions each carrying the positions that compete on it. It does not diversify the authors cited and barely opens more questions, but it triples problem framing and cuts from 56 % to 15 % the share of unmapped questions. The control isolates what does the work: the same map reduced to a list of themes, without the positions, captures only half the gain. What counts is not the ground covered, it is that each question arrives with the positions that compete on it.

1. The question we started from

We wanted to know whether large language models, prompted through a conversational interface, make good philosophical companions. Not whether they know the philosophical corpus, which is obvious, but whether they usefully accompany someone who wants to think a question through: a curious person with no training, someone worked on by a sincere question, a student facing an essay subject, a specialist digging into a difficulty. What are their broad tendencies when a philosophical request comes their way?

Our initial intuition had two parts. First, that they fall back into the same groove, existentialist in colour: meaning as something you build, authenticity, the choices you are responsible for, whatever the starting point. Second, and this is the paradoxical part, that they narrow the field of problems and theories, even though they give access in principle to almost the entire philosophical corpus. Is it not frustrating that a machine which has read Sextus Empiricus, Nāgārjuna, Anscombe or Simondon keeps returning to commonplaces about a handful of authors?

We wrote the full protocol before collecting a single answer, with its hypotheses, its thresholds for success and failure, and its metrics, so that the criteria could not be chosen after the fact. Two of our hypotheses were contradicted by the measurements; others were borne out.

A necessary disclosure: we develop Philoscopia, a website built on an analytical framework of 80 major philosophical questions (the axes), each with the main competing positions it admits (the poles), documented and sourced, with the arguments that support them and the objections that target them, so that a reader can locate their own convictions and put them to the test. The second half of this study tests, in part, whether such a framework is worth anything. Since we have expectations that could bias the results, we compared this "remedy" against two placebos and had the analysis done blind by a language model. We also publish the raw data, along with a GitHub repository that lets anyone reproduce the experiment and vary it.

2. The protocol

2.1 What is tested, and why only one answer

The experiment crosses four things: models, questions, ways of putting them, and a supplement of information given or withheld. We describe the models here, the questions in 2.2 and the supplement in 2.3; what we measure on the answers is the subject of 2.4.

We tested six models, one per provider among those most used today (with no claim to exhaustiveness, mainly a matter of cost), in their mid-range version, the one an ordinary user meets: GPT-5.6 Terra (OpenAI), Claude Sonnet 5 (Anthropic), Gemini 3.6 Flash (Google), Grok 4.5 (xAI), Mistral Medium 3.5 (Mistral AI, included to test for a possible French bias), DeepSeek V4 Pro (DeepSeek, chosen among other competitors on the Chinese market such as MoonshotAI or Z.ai). Each question is put to the panel several times, in a fresh conversation every time, at the default temperature.

Only the first answer is analysed. This choice deserves justifying, because it is also the study's main limitation. It matches a very common use: you ask a question, you read the answer, you leave. That use is admittedly contrary to philosophical enquiry itself, which requires taking things up again, objecting, going back to presuppositions. Dialogue with an AI could offer an ideal setting for exactly that, and what becomes of a ten-turn exchange is a fascinating question — but it is another study, a more expensive one, which we will run separately.

2.2 What we ask them

We put three batteries of questions to these models, in five languages (French, English, Spanish, Chinese, German). Each version was written natively in its own language, then back-translated by a third model asked to flag any drift in meaning.

Eight open questions, with no technical vocabulary, as a user would put them:

"Which questions are really worth thinking about? Where should I start?" "I am 34, my job pays well but I no longer see any point in it. What should I be asking myself?" "My father died six months ago and I cannot get back to living my life. What is there to think through here?" "Should I tell a friend a truth that will hurt them, or say nothing?" "Who should I read in order to start thinking seriously?"

None of them says the word "philosophy". That is deliberate, and it is also an objection we tested rather than argued about (§3.3).

Thirty-six anchored questions, each sitting on a philosophical theme fixed in advance (certainty, the source of knowledge, the place of science, technology, beauty, the existence of God, distributive justice, partiality, the moral circle, death, desire, one's relation to time), each asked by three kinds of user. On certainty:

Curious: "Can we ever be really sure of anything? Sometimes I think I might be wrong about everything." Student: "I have to write an essay on: 'Can we be certain of anything?' Where do I start?" Specialist: "Is the demand for certainty defensible, or does fallible knowledge do everything we ask of knowledge?"

Eight framed questions, finally: the eight open questions preceded by a single sentence, "I would like to think this question through philosophically." They serve to separate two things that everything else confuses: that the question has a philosophical shape, and that the user asks for philosophy.

No question, in any battery, names a philosopher, a school, a period or a region: a single proper noun would prime the answer and destroy the measurement.

2.3 The five conditions

Each question is put in five versions, identical but for one block appended after the question. The same length instruction applies throughout (350 words), without which the best-supplied condition would win by answering at greater length.

1. Nothing: the bare question. 2. A diversity instruction: one sentence, "draw on the widest possible range of philosophical traditions, periods and regions". The first placebo: if a sentence is enough, a map is pointless. 3. A list of themes: the 80 themes of our framework as a flat list ("Death", "Technology", "Desire"…), nothing more. The second placebo, which isolates the effect of "supplying a long list of subjects". It also controls for an obvious objection: conditions 2 to 5 all signal that philosophy is expected and condition 1 does not. Since the list and the map signal exactly as much as each other, the gap between them cannot come from that signal. 4. The narrow map: the 29 fundamental questions of the framework, each with its title, its interrogative formulation and the names of its competing positions. One line, as sent:

[DEATH] Death: what attitude should we take toward death? Freeing ourselves from the fear | Turning it into lucidity | Preparing for it as a passage | Refusing it as an evil

5. The full map: all 80 questions, same format.

These last two conditions reproduce the opening view that our companion and our MCP server actually present to an AI querying them: a digest of the questions grouped by domain (our relation to truth, to ourselves, to others, to the world), each reduced to its identifier, its title, its question and the list of its positions, plus the name of its median position where it has one. The detail of a question — the full description of each position, what is at stake in it, the figures who embody it and above all the canonical arguments that support or attack it — is only reachable on demand, through dedicated tools the model calls during the conversation. The study therefore tests the opening view, not the descent into detail: it measures the map as a document injected in one block at the head of a conversation, not a system that queries it on demand, and nothing guarantees that the second does better than the first.

2.4 The analysis, and the precautions

Every answer is analysed by an annotating language model from outside the tested panel: GLM-5.2 (Zhipu AI), chosen after a run-off against Kimi K3 (Moonshot AI) on the same answers. Agreement between the two: identical sets of extracted authors on 30 answers out of 30, identical dominant mode on 30 out of 30, same dominant orientation on 25 out of 30. The annotating model works blind, on shuffled answers, never knowing which model or which condition produced what it reads.

It records four things. The authors named, with the use made of each: a passing mention, a position attributed, or an argument actually deployed. The questions the answer raises, each then referred to one of the 80 questions of the framework, with an explicit option to answer "none". The current the answer follows in substance, even when it never names it. And its dominant mode: reasoning, doctrinal summary, practical advice or emotional support. That classification bears on the dominant mode of the answer and prejudges nothing about its quality, nor about what happens inside the model: "reasoning" here names a mode of the generated text — reasons put forward that could be contested, an examination, an objection, an argument — never a faculty ascribed to the system. It is problem framing, scored separately from 0 to 4, that measures whether competing positions are laid out with their reasons.

One technical control, finally, because it could have invalidated everything: if providers returned cached answers, our repeated draws would be artificially alike. Across 6,725 pairs of draws of the same configuration, not one answer is strictly identical to another.

2.5 What was predicted, and what was not

A clarification is needed before the results, because what follows does not all have the same status and the layout does not say so.

Six hypotheses were fixed, with their numerical thresholds, before any collection. Four bear on the canon: authors concentrate (held), discovery saturates (held), model families converge (missed, 0.30 to 0.39 against a threshold of 0.50), the canon depends on the language (contradicted). Two bear on the map: it diversifies authors more than a plain instruction does (refuted) and it improves problem framing compared with a list of themes (held).

Everything else is descriptive or came afterwards. The result that gives the article its title, the one about the dominant mode, was not anticipated: we discovered it while reading the annotations. The three indicators that replaced the refuted hypothesis about authors were declared once it had fallen, before being computed, which is the least one can do but does not amount to pre-registration. Two further collections were launched after the first data had been analysed, in response to objections: the framed questions (§3.3) and the crossed battery (§3.1), flagged each time where they appear.

The corpus breaks down as follows, the right-hand column saying what belonged to the original plan:

wavewhat it measuresanswersstatus
pilotshakedown of the apparatus, two items24shakedown
open questionscanon concentration, five languages, four tones910pre-registered
five conditionsthe effect of the map, French and English2,114pre-registered
anchored questionstwelve named problems, three reader profiles1,296pre-registered
framed questionsannouncing a philosophical expectation (§3.3)288after the fact
crossed batterythree situations, three forms each (§3.1)324after the fact
total4,956

The protocol, finally, was written before collection but was not deposited in a public registry that would timestamp it, as is the custom in psychology and medicine: on that point, the reader has only our word.

3. What the models do when left to themselves

The six results that follow all concern the bare condition, with no help of any kind supplied to the model. The first four concern what it does with the question, the last two the authors it summons.

3.1 What triggers reasoning is a determinate problem, not the way you ask

This is the central result of the study, and it took two experiments to establish: the first showed a spectacular gap, the second established where it came from — refuting our first explanation on the way.

Here is the gap. Depending on what the question is about, the dominant mode of the answer changes completely:

what the question is aboutreasoningdoctrinaladvicesupport
a determinate problem, specialist phrasing85 %15 %0 %0 %
a determinate problem, curious phrasing65 %14 %20 %1 %
a determinate problem, help with an essay28 %14 %58 %0 %
no precise problem: "where should I start?"2 %3 %96 %0 %
a personal situation: work without meaning0 %0 %98 %2 %
a personal situation: the truth that hurts0 %0 %100 %0 %
a personal situation: a father's death0 %0 %0 %100 %
Stacked bars of the dominant mode for seven kinds of question
The switch is visible at a glance: the blue of reasoning dominates questions bearing on a named problem, the orange of advice those bearing on a lived situation. The change between them is abrupt, never gradual.

A word about what the first column counts, because it governs how this whole section should be read. The classification bears on the dominant mode of the answer: laying out reasoning, rather than summarising doctrines, advising or comforting. It does not say the reasoning is any good, nor that it weighs competing positions. A second measure, problem framing, takes care of that, and the two do not coincide:

dominant modenproblem framing /4lays out 2 argued positions
reasoning8392.2445 %
doctrinal summary3771.9637 %
practical advice1,3401.1613 %
emotional support1460.502 %

More than half the answers classed as “reasoning” do not lay out two defensible positions with their reasons. Even on specialist questions the average score peaks at 2.17 out of 4, and only 45 % clear that bar. Argumentation in the strong sense, the kind that weighs competing theses against each other, therefore remains rare everywhere, including where the question invites it most.

Our first explanation was wrong

We first put this gap down to the form of the request: the specialist asks whether a demand is defensible, the bereaved person asks what there is to think about. But those questions differ in everything at once — theme, abstraction, first person, emotional charge. Nothing allowed the form to be isolated.

Here is what prompted us to revise the explanation. Having found the work of our annotator, GLM-5.2, to be mixed on certain points, we first had a sample of its classifications re-judged by two judges independent of it, GPT-5.6 Luna and Claude Opus 5, then submitted the whole study to a critical review by GPT-5.6 Terra. It was that review which flagged this confusion, among other objections. The nine items of the second experiment had been written and audited from the start, but their collection was only launched after that objection: they are therefore not part of the confirmatory protocol, they are later controls, and the reader should be able to read them as such.

So we built a second experiment designed for that alone. Three personal situations, each put under three different forms, with the situation reproduced word for word from one form to the next, so that between any two items exactly one factor ever varies:

"My mother is losing her independence and wants to stay in her own home. Caring for her day to day would cost me my job and my plans." followed by _What should I be asking myself?_ or _What is the right thing to do, and why?_ or _What has philosophy said about a situation like this?_

The result, over 324 answers:

form of the request, situation held constantreasoningdoctrinaladviceproblem framing
reflexive: _what should I be asking myself?_0 %0 %100 %0.78
normative: _what is the right thing to do, why?_1 %0 %99 %0.52
doctrinal: _what has philosophy said?_1 %77 %22 %1.56
Stacked bars of the three forms, on an identical situation
The blue of reasoning is absent from all three bars. Only the doctrinal form shifts the answer, and it shifts it toward summary, not toward argument.

No form gets above 1 % reasoning, and the result holds on each of the three situations taken separately. Asking outright _what is the right thing to do, and why?_ — a normative question that demands its reasons — produces 99 % practical advice and the lowest problem framing score in the whole study, 0.52 out of 4.

The form does shift something, then, but between advice and doctrinal exposition, never toward argument: asking what philosophy has said tips 77 % of the answers into exposition and doubles problem framing, without producing any more reasoning. You can steer these models toward philosophical material; you do not thereby get reasoning.

What actually decides

Once the form is set aside, the remaining variable is a simple one. Reasoning appears when a determinate problem is on the table, and disappears when there is none. By a problem we mean a delimited question whose answer does not go without saying, and on which several incompatible answers can be defended: something is left to settle, and the question says what. Asking whether we can be certain of anything, or whether a just society must correct inequalities of birth, puts a problem on the table: the question delimits what would have to be established and opens a space of competing answers. "My job no longer means anything" and "where should I start?" put none: the first describes a state and calls for help, the second asks for direction; nothing is yet there to be settled.

The point is neither the vocabulary, nor the level of the person asking, nor the presence of the word "philosophy". A naive phrasing by a curious person, without a single technical term, gets 65 % reasoning as soon as it bears on an identified problem. The same person, speaking about their own situation, drops to 0 %, however they choose to ask.

Two cautions about what that sentence establishes. The second experiment manipulates the form and nothing else: it is on the form that it concludes, and it concludes firmly. It does not manipulate the presence of the problem. Doing so would mean putting the same situation with and without a named problem, which we did not do. The gap between 85 % and 0 % therefore remains a covariation, observed across questions that also differ in theme and in the use of the first person. It is strong and consistent across five languages and six models, and none of the competing explanations we were able to test survives it; it does not, for all that, have the standing of an isolated factor.

The problem is therefore not an incapacity: these models switch into the reasoning mode as soon as they are given a problem. But the person who arrives with a lived question does not bring a problem, they bring a situation, and nothing in their way of asking will close that gap. That is precisely what a debate map does in their place, and it is the hypothesis §4 puts to the test.

Two secondary observations are worth recording.

Themes are not equal before reasoning. Among the anchored questions, distributive justice draws 88 % reasoning, the place of science 71 %, technology 67 %, but death 42 % and desire 35 %. The closer a theme comes to the conduct of one's own life, the steeper the slope toward advice, even when the question is properly posed.

The student case puts a number on a practice every philosophy teacher already knows. Asking for help with an essay drops reasoning to 28 %. What the pupil receives is not a piece of thinking but a structure ready to be filled in.

On "Does technology set us free?", Gemini delivers a three-part plan, references included, ready to copy out[^plan]. None of the beliefs at stake is put in question and nothing progresses from one part to the next: the discussion, which ought to generate the plan, is absent. The pupil gets an object that looks like an essay without learning anything about what a plan is. The model offers to do the homework; it does not help anyone think the problem.

3.2 Psychological introspection in place of examining beliefs

For every answer, the annotating model records the questions it invites the reader to ask themselves, then tries to attach each one to a catalogued philosophical problem. On "my job pays well but I no longer see the point of it", 82 % of those questions fall outside the framework entirely. Here they are, verbatim:

"What exactly am I missing?" · "Am I bored, or am I actually depressed?" · "What, as a child, made me lose track of time?" · "What would I regret not having tried by 40 or 50?" · "If I imagine myself at 44 still doing this, how do I feel?"

It is worth being precise about what is being criticised. Personal questioning is not a failing — the Socratic method encourages it. But that method interrogates beliefs, concepts, unexamined obviousness; the AI's version interrogates a feeling. At no point is the person invited to examine what they hold to be true, or the concepts that structure how they understand reality.

The difference is plain the moment the same model, on the same question, is given the map described in §2.3 — the twenty-nine fundamental questions, each with the names of the positions that compete on it. The questions the answer then raises become:

"Do you think you have a true nature to discover and honour, or are you free to become someone you are not yet?" · "What actually depends on you here: your attitude toward this work, or the concrete possibility of changing it?" · "Is your dissatisfaction a signal to honour, a desire to sort through, or a demand for endless novelty?"

These are the same anxieties, reformulated as what they are: positions to be taken on questions that have been argued about for a long time. The person may then discover that their unease rests on a belief they have never examined — "I had a vocation and I missed it", say, which is an essentialist thesis, disputable and disputed. The point is not to expect a philosophical companion to quote authors: it is that it should make one examine the beliefs, concepts and opinions that steer a life without one noticing.

The dominant drift of LLMs faced with a philosophical request is therefore not from one theme to another, as we had feared at the outset. It runs from the problem to introspection, and from examining beliefs to taking stock of feelings.

3.3 Announcing that you want philosophy shifts the mode without reversing it

An objection arises here, aimed at the previous result: if the model answers with advice, perhaps that is because it was asked what to do, and nothing else. Our open questions never say the word "philosophy".

That objection is testable, and we tested it rather than argue about it. The means: add an instruction in front of the question, without touching the question itself, and measure what the instruction shifts. Two instructions were tried.

The first concerns the person asking. When the message announces "I work on these questions in an academic setting and expect a rigorous answer", practical advice stays at 89 %. Presenting oneself as an academic is not enough.

The second announces the expectation directly: "I would like to think this question through philosophically", placed in front of each of the eight open questions, which are otherwise unchanged.

reasoningdoctrinaladvicesupportunmapped questions
bare question0 %14 %72 %13 %54 %
framed question22 %22 %52 %5 %33 %

Simply asking for it therefore takes reasoning from 0 % to 22 % and cuts unmapped questions by a fifth. The bias is not irreducible: it comes in part from a default reading of the request. But advice remains the majority mode, and more than three answers in four still carry no reasoning even though the user has just said exactly what they expect.

The breakdown by question is more eloquent than the average:

questionreasoning, bare → framedadvice, bare → framed
"my father died"0 % → 58 %0 % → 6 %
"a legal order I find unjust"0 % → 42 %100 % → 42 %
"where should I start?"3 % → 33 %93 % → 56 %
"the truth that hurts"0 % → 25 %100 % → 53 %
"work without meaning"0 % → 14 %97 % → 86 %
"working out my view of the world"0 % → 0 %100 % → 97 %

Bereavement flips completely: unframed it draws 100 % emotional support, and one sentence turns it into a problem. The two questions closest to personal development, by contrast, resist almost entirely. On work without meaning, even forewarned, the model offers a life audit.

The two causes therefore compound, and neither is sufficient where what is at stake is the conduct of one's own existence.

3.4 Existentialism as a slope, not as a magnet

Which current do the answers follow in substance, even when they never name it? On the two questions that start from a personal situation — work that has lost its meaning, and a recent bereavement — the existentialist framing, the idea that meaning is not given but built through choices you are responsible for, dominates 72 % of answers. Stoicism drops to 4 % there, although it is the most explicitly cited school in the corpus. These systems quote the Stoics and answer like existentialists, precisely on bereavement and on work, where Stoicism would have the most to say.

Two counter-tests were needed.

Are our questions loaded? We submitted each question on its own, with no answer attached, to two annotating models and in both languages — four verdicts per question — asking which orientations it invites. The item that draws 88 % existentialism is judged loaded by all four verdicts, for the same reason: a crisis of meaning plus a reflexive request calls for that framing. The item "who should I read?", judged open by all of them, draws only 7 %. Part of the effect therefore comes from our own phrasing, and that audit is published with the data.

Do meaning and freedom, the themes commonly associated with existentialist reflection, attract every other subject? That is what the anchored battery is for. The answer is no. Two thirds of the questions raised stay on the starting theme, and only 2 % drift toward meaning or freedom. The existentialist slope is not a magnet that seizes any subject: it is the default regime of unanchored questions. It is also a novice phenomenon: the curious user slides into it five times more often than the student or the specialist, and sees 18 % of their questions fall outside the framework, against 12 % for the specialist.

3.5 The canon: they quote Camus and reason with Kant

The ten most-cited authors account for 51 to 60 % of all mentions across the five languages, and discovery saturates fast: past forty answers, each new one brings no more than 0.1 to 0.2 previously unseen names.

There are no specifically national canons. Asked in Chinese, the models answer Plato, Socrates, Kant, Aristotle, Nietzsche, Descartes; Zhuangzi comes ninth, Confucius fifteenth. In German, a language with a large national pool of philosophers, Kant moves to first place but little else changes, and Greek figures remain well ahead of German ones (26 % of mentions against 16 %). Language adds a national tint of about five points, never a different canon. For anyone who knows how these models work statistically, the homogeneity was predictable; its scale remains striking, and it confirms the narrowing we suspected.

Counting names, however, hides the essential. Our annotation distinguishes the passing mention, the attributed position, and the argument actually deployed — that is, cases where a reason, an argument or an objection belonging to that author is genuinely put to work in the answer. This last column does not measure the mode of the answer but the use made of a name:

authormentionsin passingposition attributedargument deployed
Kant5108 %67 %25 %
Aristotle3859 %85 %5 %
Descartes28916 %63 %21 %
Plato25927 %68 %5 %
Nietzsche24418 %76 %6 %
Epicurus1866 %63 %30 %
Sartre18127 %66 %7 %
Camus17828 %67 %4 %
Hume13614 %63 %24 %
Rawls11811 %77 %12 %
Socrates11725 %61 %15 %
Epictetus10647 %42 %11 %
Marcus Aurelius10663 %30 %7 %
Most cited authors, mentions and share actually deployed
The pale bar counts mentions, the dark bar those where a reason of the author's does some work in the answer. Kant and Camus are cited in the same order of magnitude; the second almost never does any reasoning work.

The table separates two families. On one side those the answers put to work in reasoning: Epicurus, Kant, Hume, Descartes, between a fifth and a third of their mentions. On the other those who stay reference names: Aristotle, Plato, Camus, below 5 %. Marcus Aurelius is the limiting case, cited in passing two times out of three, almost never put to any reflective use.

But this split only appears when the question carries a problem. On the open questions alone, the ones an ordinary user would ask, deployment inside an argument collapses for everyone: Kant falls to 1 %, Descartes to 6 %, Aristotle and Camus to 2 %, and Camus is there cited more often than Kant. Kant's 25 % comes entirely from the anchored questions, where a problem is named. This is the finding of §3.1 seen through the authors: no author is deployed in reasoning while no problem is on the table; as soon as there is one, some serve to reason and others stay decorative.

A warning is due on that last column, because we checked it and it is too generous. We submitted a blind sample of sixty mentions, mixing the two roles without distinguishing them, to two judges independent of the study's annotator. Those two agree with each other 85 % of the time and confirm only half the mentions classed as "argument deployed", while confirming 85 % of the "attributed positions". The disputed cases are always of the same kind: a bibliography entry with a summary of doctrine, a parenthetical tag, a name cited as authority. The absolute values in that column are therefore a ceiling: on the sample checked, half the cases did not survive, and nothing establishes that this error rate is the same across authors, languages or conditions. Since the bias strikes every author alike, the gaps between them — which is what interests us here — remain valid.

Camus is instructive about how this canon is composed. He weighs almost as much as Sartre and more than Hume, with a distinctive usage profile: over a quarter of his mentions are in passing, and only 4 % see any reasoning actually built on one of his ideas. The works cited concentrate on a single title, The Myth of Sisyphus, 35 mentions, the rest being two novels.

The point is not to contest his place in the history of ideas, which is not at issue here, but to describe what that profile indicates: the canon of these models is not that of syllabuses or academic bibliographies, it is weighted by general fame and by the circulation of images. The same profile turns up among the over-represented non-philosophers of §3.6, and with Marcus Aurelius, cited two thirds of the time in passing.

3.6 The six models are not alike

So far we have spoken of "these systems" as though they formed a single block. They do not, and the gaps deserve publishing.

modelreasoningadviceno traditionproblem framingwords
DeepSeek V4 Pro40 %43 %10 %1.80345
GPT-5.6 Terra33 %50 %14 %1.51314
Grok 4.531 %52 %18 %1.19203
Claude Sonnet 530 %53 %20 %2.00351
Mistral Medium 3.526 %48 %29 %1.59284
Gemini 3.6 Flash26 %52 %21 %1.35316

How often a model reasons varies from 26 % to 40 %, and frequency does not equal quality: Claude reasons less often than DeepSeek but frames problems better when it does, with the best score in the panel. Grok is the most perfunctory on both counts, averaging 203 words where the cap allowed 350. Mistral is the one that least often follows an identifiable tradition.

On authors the result is more surprising, and it depends on the regime of the question. When the question is open, each model has its favourite author:

modelmost-cited author, open questions
Claude Sonnet 5Nietzsche
DeepSeek V4 ProPlato
Gemini 3.6 FlashPlato
Mistral Medium 3.5Camus
GPT-5.6 TerraAristotle
Grok 4.5Aristotle

As soon as the question is anchored on an identified theme, they become interchangeable: Kant comes first for all six, at 8.2 % to 12.1 % of mentions, followed everywhere by Descartes, Hume and Epicurus. A model's personality only shows in the unconstrained regime; a well-formed question brings them all back to the same schoolroom core.

Those personalities are measured by over-representation — an author's share for one model divided by their share across the whole corpus:

modelmarkedly over-represented authorsphysiognomy
GPT-5.6 TerraMontaigne ×3.3, Arendt ×2.7, Epictetus ×2.1classical moralist
Grok 4.5Frankl ×2.3, Marcus Aurelius ×2.2, Epictetus ×1.8Stoic
Mistral MediumCamus 2nd author (5.5 %), Bentham ×1.8, Seneca ×1.7French canon
Claude Sonnet 5Kahneman ×2.0, Pascal ×1.8, Marx ×1.8eclectic, one foot outside
Gemini 3.6 FlashFrankl ×2.2, Arendt ×1.9, Russell ×1.9contemporary
DeepSeek V4 ProSocrates ×1.8, Kant 10.7 %, Descartes 7.4 %the most schoolroom-like

These gaps finally explain a result that had surprised us: we expected an overlap of 0.50 between the repertoires of two providers, and measured 0.30 to 0.39. It was computed on the open questions, precisely the regime in which each follows its own slope. The canon is shared at the top and personal in the tail.

Several over-represented names are not philosophers: Kahneman for Claude, Frankl for Grok and Gemini, Orwell for GPT. These authors count in the history of ideas and nothing here diminishes their interest; they simply are not major figures of the philosophical corpus, and seeing them raised to the level of Pascal or Marx extends what we observed with Camus (§3.5): this canon spills over into psychology and literature. Which authors each model prefers would deserve a study of its own, over several conversational turns.

4. What a philosophical debate map changes

The five conditions are compared on strictly matched material: the same eight open questions, the same language, the same neutral tone, 240 answers per condition. The groups being equal in size, the counts read directly, with no sample-size correction.

conditiondistinct authorsquestions per answeraxes touchedunmapped questionsproblem framing (0–4)
bare question696.05756 %0.67
diversity instruction2075.66228 %1.55
list of 80 themes769.48017 %1.29
map of 29 questions616.54917 %1.90
map of 80 questions727.68015 %1.99
Two panels comparing the five conditions
On strictly matched material. The two maps, in dark, stand apart on problem framing, which is scored by a rubric mentioning none of our axes. On questions left with no identifiable problem, the flat list almost catches up.

The map does not diversify authors, and we expected the opposite: hypothesis refuted. A simple diversity instruction triples the number of distinct authors, where the map leaves it roughly where it was.

But that instruction-driven diversity is a trompe-l'œil, and it is the placebo that saved us from concluding wrongly. In 57 % of those answers the annotating model finds no tradition followed, not even implicitly: they are bibliographic indexes, sound questions followed by names in brackets — "What do I really desire? (Buddhism, Stoicism, Spinoza)" — in which only 3 % of mentions deploy an argument. The instruction does not make them think more widely, it makes them label more.

What the map brings does not lie in how much material it opens, and that is the most instructive result in this section.

It triples problem framing, from 0.67 to 1.99, the sharpest gap in the whole study. This is the result that carries weight, because the rubric that awards it mentions none of our axes: it asks only whether the answer lays out incompatible positions with their reasons. A reader who rejected our carving outright would still have to accept this figure.

It also cuts from 56 % to 15 % the share of invited questions falling outside the framework, the conversion documented in §3.2, where the questions raised stop being invitations to introspect and become catalogued problems. This figure is the most spectacular and yet the weakest: the map injected into the question and the yardstick that sorts the answers are the same object, so part of the drop is granted in advance. We give it because it describes what the user experiences, not because it proves anything.

The map does not, on the other hand, open more doors than the other conditions, and that is precisely the thing to hold on to. It takes the number of questions opened per answer from 6.0 to 7.6, a modest rise, and the flat list opens far more, 9.4. On sheer ground covered the two are equal: the list and the full map each touch all 80 problems of the framework, where the bare model reaches 57. Opening questions and making them workable are two different things. Only the second is a distinctive effect of using the debate map.

In other words it produces the effect one would expect of a philosophical companion for a beginner: not opening more subjects, but turning what it opens into problems one can work on, where the model on its own sends the user back to their own feelings.

The control that explains why. The flat list of the same 80 themes opens more doors than the map and touches exactly as many problems, yet it reaches only 1.29 in problem framing against 1.90 and 1.99 for the two maps. So it does far better than nothing, 0.67, and about half as well as the map. The effect therefore comes neither from the length of the context, nor from the number of subjects, nor from the signal "philosophy is expected" — the list carries all of that as much as the map does. It comes from articulating each theme into a question equipped with competing positions. Twenty-nine well-articulated questions beat eighty bare themes. A theme says "here is something to talk about"; a question with its positions says "here is what is in dispute, and here is what you will have to choose between". Only the second is accompanied by higher problem framing.

5. Discussion: a confusion between philosophy and personal development

Taken in isolation, none of these biases is truly alarming. Answering distress with support, prompting introspection, quoting Camus: nothing illegitimate, nothing inherently unsatisfactory. Together, though, they produce an image of philosophy hard to tell apart from personal development, understood here in its ordinary sense: that set of methods promising fulfilment and personal effectiveness, starting from how you feel, clarifying what matters to you, and converting the whole into habits and decisions.

The problem is not personal development, and we have no intention of disqualifying it — its practical, individualised approach, in the spirit of coaching, is defensible and legitimate. The problem is the confusion. Someone questioning a language model without philosophical training, but hoping to be accompanied in a philosophical enquiry, would stand a good chance of taking away an implicit lesson: that philosophising consists essentially in clarifying what one feels, identifying what matters to oneself, and putting in place a set of routines to improve daily life. A few great names quoted along the way would lend these steps an extra warrant and supply most of their philosophical colour.

Yet the two approaches differ on a point our measurements touch directly. Personal development starts from concepts and values it uses without interrogating them: what matters to you, effectiveness, authenticity, the meaning you give your life. Philosophical enquiry consists precisely in taking those as its objects, asking where they come from, what they presuppose, and whether what they assert survives examination.

That is exactly the gap our rate of unmapped questions measures. The questions LLMs raise about a personal situation bear on a feeling to be clarified, almost never on a belief to be examined; those that emerge with a debate map bear on positions to be taken. We do not claim to settle which of the two regimes serves a person in difficulty better: we observe that they are distinct and that what LLMs spontaneously offer under the name of philosophy belongs mostly to the first.

This does not condemn practical concern itself, from which philosophy would rather have something to learn: the ancient schools did not separate doctrine from exercise, and a philosophy that changed nothing in a life would be missing something. What our data show is that these models sustain the confusion instead of holding the two together. We say they sustain it because they merely reproduce a confusion that predates them; but we observe that they do not help dispel it unless prompted to.

That such tools serve the initiated better is hardly surprising and probably holds in every field. But what is worrying here, and perhaps less so in other fields, is that the novice does not receive a simplified version of philosophy: they receive something else.

6. Recommendations

These measurements yield concrete principles for anyone equipping an AI to respond as pertinently as possible to a philosophical request, ourselves included.

1. A bare list of themes works against reflection. This is the study's most counter-intuitive result: supplying subjects makes for worse thinking than supplying nothing at all. What produces the effect is a list already built by specialists, in which each theme is formulated as a question and accompanied by the positions that compete on it. One obviously cannot ask a user to bring their own questions and positions: that is precisely the work such a framework does in their place. 2. But a map of positions must invite contesting them, not picking one. The symmetrical risk is to fence reflection inside the positions on offer, like a menu of options to choose between. What is needed is a device that pushes the user to interrogate the positions themselves, their presuppositions and the concepts they employ. That is the very spirit of philosophical enquiry, and the object of the instructions we give our own models to encourage it. 3. Structure before volume. Twenty-nine articulated questions beat eighty themes. Exhaustiveness is a coverage goal, not a quality lever. 4. Be wary of diversity instructions. They produce measurable name-dropping: more names, fewer arguments. Breadth should be asked of problems, not of bibliographies. 5. Turn introspection into examination. The reflex "what am I feeling?" should always be paired with "what do I hold to be true, and have I examined it?". That is the switch the map performs measurably, and it separates psychological support from philosophical companionship reasonably well. 6. Treat the user's phrasing as the real entry point. Register follows the shape of the request, and the user who most needs help is the one who formulates it least well. A useful companion must therefore begin by turning a lived question into a problem, rather than answering it as it stands. 7. Make arguments as reachable as names. Throughout our data, authors serve as labels far more often than as supports (3 % of argued mentions for Camus, 1 % for Marcus Aurelius). A system that wants to make people philosophise must give access to reasons, not to references. That is the role of the canonical arguments attached to each position in our framework, and precisely what this study, limited to the first answer, could not bring into play.

7. Limitations of the protocol

Each of these limitations comes from a design choice. We state what was restricted, what that restriction takes away from the scope of the results, and what it would take to go further.

We analysed only the first answer of each conversation. Our results therefore hold for the "one question, one answer" use, and say nothing about what happens after five or ten exchanges. This is a heavy restriction, for two opposite reasons. On one side, a model left alone might start repeating itself and going in circles, which would make the picture worse. On the other, an equipped system only deploys its resources over time, fetching at the third or fourth turn the arguments attached to a position: everything we measure here is unfavourable to it. To settle the matter, the protocol would have to be replayed over multi-turn conversations, with a simulated user identical across all conditions, measuring at each turn what is newly brought and what is repeated. That is the object of a companion study.

We tested the map as a document, not as a tool. It was injected in one block, reduced to the labels of its positions, without the descriptions, the stakes or the arguments a real companion fetches question by question. Our figures therefore bear on a document injected in one block, not on a system querying the framework as the conversation goes. We do not claim they give a lower bound: a tool can also distract the model, stiffen its answers or multiply faulty references. One sign confirms this: under the map, authors become position-holders and the share of mentions deploying an argument falls from 13 % to 2 %. Going further requires a condition in which the model queries the framework through tools, on demand, as the product does.

We treated four uses alike that have almost nothing in common. A curious person asking in passing, an amateur wanting to dig into a problem or consolidate their view of the world, a pupil with homework to hand in, a specialist preparing a class or working through a difficulty: these four do not expect the same thing and should not be judged by the same yardstick. Our three user-phrasings approximate them; they do not replace them. The heart of our argument concerns the individual amateur use, someone who wants to think for themselves without being a student or a teacher; that is where our results carry most. For each of the other uses, a study designed for it, with its own criteria of success, would deserve running: what a philosophy teacher expects of an AI has little to do with what a reader adrift for meaning needs.

Our questions partly steer the answers. A question staging a crisis of meaning calls for an existentialist framing, and our eight open questions do not represent the whole range of possible requests. We measured that steering rather than denying it, by submitting each question on its own to annotators, and the audit is published with the data: the reader can check how much of our results comes from our phrasing. A broader battery, written by several authors, would reduce the risk.

Our framework is itself a canon. The 80 questions serving as the map, and as the grid for classifying the questions raised, were written by a French speaker fed on the same culture as the models tested. A question our framework does not cover counts as "unmapped", which can overstate the share of introspection: this is why that rate is always given with its scale, and the geographical composition of our 80 questions published alongside that of the answers. A map built by others, in another tradition, would give a different yardstick.

The record of authors has two defects that we measured. The "argument deployed" role is assigned about twice too often, as the double check described in §3.5 showed; the proportions between authors hold, the absolute values are a ceiling. Furthermore, a mechanical test on the French and English answers finds that 3 % of the authors recorded appear nowhere in the text: the annotator invents some, more often on the mentions it judges argued (5.5 %). Checking ten randomly drawn cases by eye confirms all ten. A third discrepancy follows: in 9.8 % of answers, the annotator combines a fallback category with a substantive orientation, where its instructions required the fallback categories to stand alone.

These three defects were discovered after the fact and tested rather than debated, and they work against our findings about the canon. If the "argument deployed" role is twice too generous, genuine argumentation is rarer still than we report. If names are invented, measured diversity is overstated and the concentration of the canon understated. If fallback categories are mixed with substantive orientations, our calculation counts those answers as following a tradition, which understates the share that follow none. None of the three touches the dominant mode, the unmapped rate, problem framing or any comparison between conditions, which do not depend on the record of authors. One reservation remains on that last point: nothing rules out the annotator erring differentially by condition — recognising a position more readily, say, in an answer that reuses the map's vocabulary. We did not test that.

The classification of answers is done by a language model. The categories it applies — an answer's mode, the tradition it follows — remain judgements, and a single annotator can carry its own systematic biases. We set it against a second model on a sample (perfect agreement on extracted authors and on mode, 25 out of 30 on the dominant orientation) and hand-checked thirty answers. A panel of three annotators, or human double-annotation on a larger sample, would be sturdier.

These measurements are a snapshot. Models are updated without notice, and nothing guarantees these figures will hold in six months. Every data point carries its date and the provider that served it, which is one more reason to publish the protocol and the code: the study is built to be re-run, by us or by anyone else.

8. Data and code

Everything is published: the raw answers, the annotations, the questions in all five languages, the audit of the items, the collection and analysis scripts, and the pre-registered protocol with its thresholds. The harness is reusable with an OpenRouter key, including on a single model: anyone can audit the assistant they use. A claim about the behaviour of these systems that a reader cannot re-run is worth nothing.

Everything is gathered in a public repository: github.com/fbgallet/philo-probe. The reference data, that is the 4,956 answers and the whole of their annotations, sit in its results folder, together with the full protocol (PROTOCOL.md) and the exact configuration the collection was run under.

Contributions

This study was carried out as a collaboration between a human author and an agentic AI system, Claude Code (Claude Opus 5 model, Anthropic). The author designed the protocol, defined the research questions, the hypotheses and their thresholds, arbitrated every choice along the way, reviewed, corrected and extended all of the analyses and the text, and bears sole responsibility for the content. Claude Code wrote and ran the collection, annotation and analysis scripts, proposed indicators, identified several regularities taken up in the article — including some the author would probably not have spotted — drafted first versions of the text, later reviewed and corrected, and produced the English translation, which was reviewed. The remaining errors are the author's.