← Back to the Enneagram Test

Methodology & Psychometrics

The scoring model behind the Enneagram Test, written for anyone who wants to know exactly how their result was calculated — not just told to trust it.

90,000
simulated respondents the scoring model was fitted against
~82.5%
pure-type accuracy in that simulation
81
questions at most — most people answer far fewer
01

What the test measures

The Enneagram Test scores four things from one sitting: your core type, your wing, your tritype, and your instinctual stack (sp / so / sx). Every question is built around motivation, fear, and behavior pattern — what’s driving a response, not just the response itself — because that’s what actually separates one type from another, especially between the look-alike pairs (a Type 1 and a Type 8 can both look “controlling” from the outside for entirely different reasons underneath).

It’s built as a self-reflection instrument, not a diagnostic one. That distinction shapes everything on this page: where a clinical instrument optimizes for a single confident label, this one optimizes for telling you honestly how sure it is, and why.

02

The question bank

The test opens with a fixed set of 69 questions: a bank of five-point Likert-scale statements (“strongly disagree” to “strongly agree”) spanning all nine types, plus a set of forced-choice pairs built specifically to separate the types people most often confuse — Type 1 and Type 8, Type 2 and Type 9, Type 3 and Type 7, and eight other pairs like them, each written as two believable-sounding statements where only one actually fits a given type’s real motivation.

If your top two types come out close after those 69, the engine adds more — targeted, not random — up to a hard ceiling of 81 questions total. Most people never see the extension at all; it only activates when the evidence is genuinely ambiguous. More on how that extension works below.

A DELIBERATE SCORING CHOICE

Every Likert item is anchored to a fixed neutral midpoint, not scored against your own average response. Some personality instruments use “ipsative” scoring — comparing your answers only to each other — which quietly distorts results for people who tend to answer near the extremes or near the middle across the board. This test doesn’t do that: a “5” means the same thing regardless of how you answered everything else.

03

The scoring model

Under the hood, this isn’t a simple point tally. Each of the nine types has its own fitted logistic regression model — nine separate equations, each one weighing all 54 scored features (43 core Likert items plus 11 forced-choice discriminators) against how strongly that specific type explains your answers. A Type 4 item you answered doesn’t just add points to Type 4; it also pulls slightly against Type 9, or slightly toward Type 1’s core fear, depending on how those patterns actually correlate in real data. That cross-loading is what a simple additive tally can’t do.

The model’s weights were fitted against 90,000 simulated respondents, generated with randomized noise parameters across a plausible range rather than one fixed set of assumptions about how people answer — so it isn’t overfit to a single, narrow idea of what a “typical” Type 6 answer pattern looks like. In that same simulation, this model correctly recovered a respondent’s intended pure type about 82.5% of the time, up from roughly 71.6% for the flat per-type average this replaced.

Each type’s raw score is converted to a display scale (roughly 15–85, centered on 50) so scores are easy to compare across types at a glance, without changing the underlying ranking the model actually computed.

04

How your wing is calculated

Your wing isn’t a coin flip between your core type’s two neighbors. The model compares your two adjacent types’ scores directly and converts the gap between them into a probability split — so a wing that’s genuinely close gets reported as close (a “slight lean,” roughly 50–60%), and a wing that’s clearly dominant gets reported as dominant (a “strong lean,” 78%+), rather than both being flattened into an identical “Type 4w5” label with no sense of how confidently that “w5” was chosen.

05

How your tritype is calculated

The nine types sit in three centers of intelligence, and your tritype is the dominant type from each:

Gut
8 · 9 · 1
instinct, control, resistance
Heart
2 · 3 · 4
image, feeling, connection
Head
5 · 6 · 7
thinking, security, planning

Within each center, the type with the highest score wins that slot — and, like the wing, the gap between the top type and the runner-up in that center determines whether it’s reported as a clear result, a likely one, or a genuine close call between two types in the same center.

06

How your instinctual stack is calculated

Instinctual variant (self-preservation, social, sexual/one-to-one) is scored completely separately from type, through 15 forced-ranking triads: for each one, you pick which of the three instincts feels most like you and which feels least like you, leaving the third unstated.

This uses standard forced-ranking scoring (a Luce/Plackett-style model), not a simple tally. Being picked “most like me” earns a full positive score; the unstated middle instinct earns a small implicit credit, since it was preferred over whatever you ranked least; and “least like me” is explicitly penalized rather than just left at zero. That ordering carries real information a flat count would throw away.

07

The adaptive extension

If your top two types are still close after the initial 69 questions — a narrow score gap, a known look-alike pair, or your answers split into two internally inconsistent halves (see the next section) — the test adds up to 12 further questions, hand-picked to separate specifically those two (or occasionally three) types, rather than continuing with generic items that wouldn’t move the needle.

These extension items reuse the same fitted, cross-loaded weights as the main model rather than applying a flat bonus to whichever type you picked. That distinction mattered in testing: an earlier version that just added a flat bump to the chosen type actually made results measurably less accurate in simulation (roughly 82.7% down to 80.2%), because a flat bump carries none of the cross-loading information a real fitted weight does. The current version fixed that, including for people whose profile is genuinely blended across two or three types rather than cleanly single-type.

08

How confidence is scored

Every result ships with a confidence label — high, medium, or low — built from a 0–100 composite score, not a guess. Four signals feed into it:

gap between your top two types + split-half agreement + tie-break agreement + response quality

Score gap

How far your top type’s score leads the runner-up. A wide gap is the single strongest signal that the result is genuinely clear-cut.

Split-half agreement

Your Likert answers are silently split into two independent halves as you go (odd-numbered vs. even-numbered items per type), and each half is scored on its own through the same model. If both halves independently point to the same type, that’s real internal-consistency evidence — the kind of check a one-shot quiz can’t offer.

Response quality

The engine checks for straight-lining (the same answer repeated across the whole questionnaire — usually a sign a section got clicked through rather than considered) and for an unusually heavy lean toward extreme or dead-center answers across the board. Either pattern discounts confidence somewhat, on the reasoning that it’s more likely to reflect rushed answering than a genuinely decisive personality — though genuinely decisive people aren’t penalized for a few strong opinions; the threshold only trips on a broad pattern across many items.

A close race between two types, or answers that don’t fully agree with themselves, is supposed to lower this score. That’s by design — an honest “medium confidence, here’s why” is more useful than a false “certain” on a result that isn’t.

09

What this doesn’t measure

  • —It’s not a diagnosis. The Enneagram is a framework for noticing patterns in motivation, and no questionnaire — this one included — captures a person completely.
  • —It’s a snapshot, not a fixed label. Your result reflects how you answered today. Mood, context, and self-awareness all shift the picture, which is exactly why every result reports its own confidence rather than presenting a single number as certain.
  • —No institutional affiliation. Whispers of Void is an independently built tool, not affiliated with any clinical or research body. Take what’s useful and set aside what doesn’t fit — you know your own life better than any report can.

Questions this page didn’t answer, or found something that looks off in your result? Reach out through the contact link in the site footer — we read every message.

Scroll to Top