Methodology

How the test was built, why each design decision was made, and what it still cannot do. Every number below is checkable against the test itself.

Why eight functions instead of four letters

A four-letter result is a summary of a summary. It collapses eight measurements into four binary decisions and discards the distances between them — which is exactly the information you need when two functions score close together.

This test measures the eight functions directly and only then derives a four-letter code from their ordering. The consequence is that you always see the underlying measurements, even when the derived code is unstable.

Why 96 questions

Twelve items per function is the shortest length at which two functions scoring close together can be separated with any confidence. Most mistypes come from exactly that situation, so a shorter test would not be measuring the same thing — it would be guessing at it.

The cost is real: 96 items is roughly fourteen minutes. Four chapters of 24 items make that cost payable, because each chapter has a visible end.

Why two reverse-scored items per function

People agree with statements at different baseline rates. Without a correction, a test measuring agreement partly measures agreeableness — and every function would look slightly stronger for agreeable respondents.

Two of the twelve items per function are therefore phrased in the opposite direction. The gap between a function's forward and reverse items is the consistency reading on your result page: a large gap means your answers about that function contradict each other.

How the four-function stack is derived

The stack order follows the Grant/Brownsword model: the highest-scoring attitude pair takes the dominant position, and the remaining functions fall into auxiliary, tertiary and inferior positions by their measured scores and their required attitude.

Judging and perceiving are not separate measurements here. J/P is a structural consequence of which function ends up extraverted and which the attitude of the dominant — which is why the four groups (NT/NF/SJ/SP) never cross over.

What "confidence" measures — and what it does not

Confidence is the distance between your best-fit type and the second-best, expressed as a percentage. It describes how clearly your profile separates the top type from the alternatives.

It is not a probability that the result is correct. A high confidence reading means your answers are internally decisive, not that the instrument is right about you. A low reading means you should read the top three types rather than the first one.

What this test deliberately does not do

It does not use forced-choice items that hide the trade-off from you, so your result reflects what you reported rather than a constraint we imposed.

It does not publish normalised scores against a population, because we have not collected a sample large enough to justify them. The distribution page reports published external figures and says plainly where they come from.

It does not reproduce any proprietary instrument content, and it does not claim equivalence with any licensed assessment.

Known limitations

  • Every score is self-report. A person who answers strategically, or who has a fixed idea of their own type, will get a profile that reflects that idea.
  • There is no norm sample yet, so scores are comparable across your own sittings but not against a population.
  • The item set is written in English and translated; idiom and cultural framing shift what a statement means, and we have not measured that drift.
  • Typology instruments of this kind do not predict job performance, relationship outcomes, or mental health, and should not be used as if they did.

Design decisions and item wording are documented here so they can be checked rather than taken on trust. Corrections and source-backed objections are welcome at the contact address on this site.

Keep reading