Self-knowledge tests: which ones are worth it and how to read the result
Taking personality tests has become a common habit and an almost always wasted one. Someone answers fifty questions, receives four letters or a number, feels a pleasant flash of recognition for ten minutes, and the next day everything continues exactly as before.
Two things explain nearly all of the waste: choosing instruments that don't reliably measure anything, and reading the result as a verdict rather than as a data point. The test itself is usually innocent. This article gives you the criteria for separating a serious instrument from a magazine quiz, explains what a result actually means, and proposes an order for the tests worth taking.
What a test actually does
A good personality test does one thing, and does it well: it places you in a distribution. It tells you where you stand relative to other people on a characteristic defined precisely enough to be measured.
That's why the correct reading of a result is always comparative. Being extroverted means being above most people on a scale of extraversion, rather than belonging to a different species from those below. The difference between those two readings is the difference between using a result and being stuck with it.
A test also doesn't tell you what to do. It describes tendencies, with a margin of error, at one moment in your life. The practical part, what you decide to change, is your work and happens afterward.
The four criteria that separate an instrument from a quiz
If you only memorize one thing from this article, make it this list. Apply it to any test before you give it credit.
Internal consistency. Do the questions supposedly measuring the same thing agree with each other? It's measured with a coefficient, typically Cronbach's alpha, and values above 0.70 are the acceptable minimum. Serious instruments publish this number.
Stability over time. If you answer again in six weeks, do you get the same thing? It's called test-retest reliability, and this is where many popular instruments fail spectacularly. In the five-factor model, retest correlations after years typically hold between 0.70 and 0.80. In typologies that force categories, a substantial share of people change type within weeks, and that's the most consistent criticism of the MBTI, which we covered in the article on the 16 personalities.
Validity. Does the result predict anything outside the test itself? Observed behavior, ratings from people who know you, performance, health. An instrument that only correlates with itself is measuring how you respond to questionnaires.
Reference norms. Is there a large sample your score is compared against, and does that sample resemble you? A score without a norm is a number without a scale.
A fifth, informal and revealing criterion: does the instrument state clearly what it can't do? The serious ones do. The ones selling certainty tend to have less to back it with.
What to do with the result
Here's the most common mistake, and it's easy to fix.
A result is a percentile, rather than a category. If you land at the 65th percentile of extraversion, that means you answer more extrovertedly than sixty-five percent of the sample. That's near the middle. The distance between you and someone at the 55th percentile is practically nothing, and categorical language, which calls you extroverted and the other person introverted, erases that information.
The second thing to do is look at the extremes of your profile and ignore the middle. In a five-factor result, the two dimensions where you're very high or very low say something about you. The three that land between thirty and seventy say you're an ordinary person on those, which is the most frequent answer and the least commented on.
The third is to look for what you dislike. The part of the report that makes you think it isn't quite like that is usually the part with the most value, for a simple reason: what confirms what you already thought adds no information at all.
The Barnum effect, and how to test it on yourself
The Barnum effect is our tendency to accept as precise any description vague enough to fit nearly everybody. It's the engine of half the testing industry, and recognizing it is a trainable skill.
The classic demonstration is easy to repeat. Take your report and underline the sentences that apply only to you. Then read the report for a different type from yours and see how many of those sentences would also fit. The percentage that survives this test is the instrument's real value, and it's usually a good deal smaller than the first impression.
Sentences like you need time alone to recharge, you hold high standards for yourself, or you sometimes doubt your decisions fit practically any adult. A good report contains statements that could be wrong, and that's what makes it useful.
Which ones are worth it, and in what order
It makes sense to start with the instruments with the best support and work down, rather than starting with the most entertaining.
First, the five factors. It's the personality model with the best empirical support and the base everything else rests on. It gives you five continuous dimensions with norms and decades of research behind them. We wrote about interpreting yours in the article on the Big Five.
Then, attachment style. It changes more in practical life than any other result on this list, because it describes what you do when closeness tightens. The instruments used in research measure two continuous dimensions, anxiety and avoidance, and have good psychometric properties. The territory is described in attachment styles.
Next, values. Schwartz's questionnaire places you in the circle of ten basic values and explains your internal tensions better than any personality description. That's what we work through in personal values.
Finally, the typologies. Enneagram, MBTI, temperaments. They have weaker empirical support and one quality the others lack: they give vocabulary and are easy to share with another person. Use them as language for talking about yourself, aware that the category is a simplification.
The order matters because the first group gives you the ruler you read the second one with. Taking the Enneagram before knowing where you sit on the five factors tends to produce an identification with a type that the data doesn't support.
The mistakes that ruin a result before you read it
Five things contaminate a response, and all of them are avoidable.
Answering as you'd like to be. Social desirability is the biggest contaminant of these instruments, and it's almost always unconscious. The antidote is answering fast, with the first reaction, and thinking of a concrete context rather than your life in the abstract.
Taking the test in an atypical state. After an awful week or very good news, some dimensions shift temporarily. Lack of sleep is among the factors that distort most.
Answering with work in mind. Many people answer in the professional register without noticing, and the result then describes the work version rather than the person. If the instrument doesn't specify the context, pick one and hold it.
Reading only the summary. The profile title is the least informative part of the report. The dimensional scores are what count.
Not asking for a second opinion. Research on self-awareness distinguishes the internal component, what you understand about yourself, from the external one, what you understand about how others see you, and shows they're nearly independent. Asking two close people to answer the same instrument about you is the cheapest and most revealing exercise on this entire list.
When to reassess
Personality changes slowly, and it does change. Five-factor scores shift across adult life along known patterns, and high-impact events can accelerate that reorganization.
Repeating the instruments every two years is a reasonable cadence. Reassessing more often than that tends to capture noise instead of change, and a few points of variation between two administrations is the instrument's margin of error at work, with no meaning at all.
There's one exception. It's worth repeating after real structural changes: a separation, becoming a parent, an illness, moving country. In those cases the reorganization is fast enough to show up in the numbers.
What to take with you
A well-chosen, well-read test gives you one concrete thing: precise language for patterns you already felt and couldn't name, with an indication of where you stand relative to other people.
The four criteria protect you from nearly everything that goes wrong. Internal consistency, stability over time, validity outside the test itself, and reference norms. An instrument that doesn't publish these numbers is asking for your trust without presenting reasons.
And the step almost everyone skips is the one that pays most: after reading your result, pick one sentence you disagreed with and spend ten minutes in writing exploring why. That's where the test stops being entertainment.
Frequently asked questions
What are the best self-knowledge tests?
In order of empirical support: the five factors, for the general structure of personality; attachment style instruments, which measure anxiety and avoidance on continuous scales; and Schwartz's values questionnaire. Typologies like the Enneagram and the MBTI have weaker support and mostly serve as shareable vocabulary.
How do I know whether a personality test is reliable?
Check four things: internal consistency, typically Cronbach's alpha above 0.70; test-retest stability after weeks or years; validity, meaning whether the result predicts anything outside the test itself; and the existence of reference norms in a large sample. Serious instruments publish these numbers and also state what they can't do.
What does a test result actually mean?
It's a percentile, rather than a category. A score at the 65th percentile of extraversion means you answer more extrovertedly than sixty-five percent of the sample, which is near the middle. Look at the dimensions where you're very high or very low and ignore the ones landing between thirty and seventy.
What is the Barnum effect in personality tests?
It's our tendency to accept as precise any description vague enough to fit nearly everybody. Test it like this: underline the sentences in your report that apply only to you, then read the report for another type and see how many of them would also fit. What's left over is the instrument's real value.
How often should I retake a test?
Every two years is a reasonable cadence. Repeating more frequently tends to capture noise rather than real change, and variations of a few points between administrations are margin of error. It's worth reassessing sooner after structural changes such as a separation, becoming a parent or an illness.
Why does my result change every time I take the test?
It can be the instrument or it can be the context. Typologies that force categories make many people oscillate between types, because they cut continuous dimensions in half and anyone near the cut switches sides with very little variation. On your side, lack of sleep, an atypical emotional state, or answering with work in mind rather than life in general all shift the scores.
Agapone crosses several instruments into a single reading, with each one's limits acknowledged, so the result is good for something after you close the report. It's in the Mental category of the DNA that MBTI, the Enneagram, the Big Five and DISC enter that reading, each carrying the weight it deserves. To see that in practice, visit Agapone and explore your profile.
Enjoying what you read here?
Add Agapone as a preferred source on Google Search to see more of our articles in Discover and Top Stories.