How we know what we claim

Methodology

Last updated 22 July 2026

This page states what the instrument is, what its published reliabilities are, what evidence we do not yet have, and what this result must never be used to decide. If a number appears anywhere in Persona AI, its provenance is on this page.

What the instrument is

Persona AI scores 24 personality scales. The questions are not ours: they are the IPIP-HEXACO facet scales, published by the International Personality Item Pool to measure the constructs of the HEXACO model of personality (Ashton, Lee & Goldberg, 2007; Goldberg, 1999). They are in the public domain, they have been in use in research for years, and anyone can read every item and every key at the source.

IPIP states its licensing plainly: “Because the IPIP has been placed in the public domain, permission has already been automatically granted for any person to use IPIP items, scales, and inventories for any purpose, commercial or non-commercial.” Attribution is requested rather than required. We give it here, and anywhere we report a number that came from them.

There are 24 scales, four in each of six domains. The full published version of each scale has ten items. A standard sitting here administers four of them per scale, an extended sitting eight; which items form each subset is fixed in advance and identical for everyone.

One difference from most consumer tests is worth naming. HEXACO carries a sixth factor, Honesty-Humility, which five-factor (Big Five) instruments do not measure at all. Ashton and Lee (2008) report that it accounts for variation the five-factor model leaves unexplained — which is why it is a sixth axis here rather than a rearrangement of the other five.

How the 24 scales map onto the six axes

Each axis you are shown is exactly one HEXACO domain: four scales, sixteen items in a standard sitting. Nothing is borrowed across domains and nothing is placed in an axis it was not written to measure, so the axis you read is the domain the items belong to.

The six reported axes and the four IPIP-HEXACO scales measured behind each.
Axis, by its polesScales behind itHEXACO domain
Outgoing / ReflectiveSociability, Social Boldness, Liveliness, ExpressivenesseXtraversion (X)
Visionary / PracticalCreativity, Inquisitiveness, Aesthetic Appreciation, UnconventionalityOpenness (O)
Empathic / AnalyticalForgiveness, Gentleness, Flexibility, PatienceAgreeableness (A)
Organized / FlexibleOrganization, Diligence, Prudence, PerfectionismConscientiousness (C)
Steady / SensitiveFearfulness, Anxiety, Dependence, SentimentalityEmotionality (E), scored in reverse
Unassuming / Status-seekingSincerity, Fairness, Greed-Avoidance, ModestyHonesty-Humility (H)

Two notes on the table. The Steady / Sensitive axis is Emotionality read in reverse: high scores on Fearfulness, Anxiety, Dependence and Sentimentality produce a low score on that axis, because the low pole of Emotionality is what everyday language calls steadiness. And the Agreeableness scale named Flexibility — a willingness to compromise — has nothing to do with the Flexible pole of the Organized / Flexible axis. They are two different ideas that collide in one English word.

Four scales are printed under a different name in your own report. IPIP calls them Sincerity, Fairness, Greed-Avoidance and Modesty; your report calls the same four scales Directness, Even dealing, Material indifference and Understatement, and calls the axis they form Self-presentation rather than Honesty-Humility. Nothing about the scoring changes — the items, the keying and the arithmetic are IPIP's, unaltered, and the names in the table above are the ones to cite or look up. The report uses the plainer set because a scale name is printed next to your own number there, and a number under a word like Fairness reads as a verdict on the person rather than a description of what the ten items ask about. The source names stay on this page so the two can always be lined up.

An axis score is the mean of its four scale scores. A scale score is the average of your answers on that scale, rescaled to run from 0 to 100. It is an average of your own responses and nothing else — in particular it is not a percentile, and the section on norms below explains why we cannot give you one.

Every scale, its length, and IPIP's published alpha

The alpha column below is internal consistency as published by IPIP for the parent scales — not yet measured on our own sample. Read it with two qualifications. It describes IPIP's ten-item scales, not the four- or eight-item subsets administered here, and shorter scales are always less internally consistent than longer ones; so treat these figures as a ceiling for our version rather than as its value. And alpha is a statement about the items, not about the person answering them.

Every scale is measured with 4 items in the standard sitting and 8 in the extended one. Every item uses the same 5-point agree–disagree response scale, fully labelled: all five points carry words, not just the two ends.

The bands in the last column are the ones used in EFPA and BPS test reviews: below .70 inadequate, .70 to below .80 adequate, .80 to below .90 good, .90 and above excellent. A value sitting exactly on a boundary is placed in the higher band.

All 24 scales, their length in each sitting, and the internal consistency IPIP publishes for the parent scales — not measured on our sample.
ScaleFeeds axisItems (standard / extended)Alpha, published by IPIPBand
SociabilityOutgoing / Reflective4 / 8.85Good
Social BoldnessOutgoing / Reflective4 / 8.86Good
LivelinessOutgoing / Reflective4 / 8.82Good
ExpressivenessOutgoing / Reflective4 / 8.84Good
CreativityVisionary / Practical4 / 8.85Good
InquisitivenessVisionary / Practical4 / 8.78Adequate
Aesthetic AppreciationVisionary / Practical4 / 8.83Good
UnconventionalityVisionary / Practical4 / 8.84Good
ForgivenessEmpathic / Analytical4 / 8.78Adequate
GentlenessEmpathic / Analytical4 / 8.81Good
FlexibilityEmpathic / Analytical4 / 8.73Adequate
PatienceEmpathic / Analytical4 / 8.88Good
OrganizationOrganized / Flexible4 / 8.85Good
DiligenceOrganized / Flexible4 / 8.81Good
PrudenceOrganized / Flexible4 / 8.80Good
PerfectionismOrganized / Flexible4 / 8.80Good
FearfulnessSteady / Sensitive4 / 8.84Good
AnxietySteady / Sensitive4 / 8.85Good
DependenceSteady / Sensitive4 / 8.73Adequate
SentimentalitySteady / Sensitive4 / 8.79Adequate
SincerityUnassuming / Status-seeking4 / 8.81Good
FairnessUnassuming / Status-seeking4 / 8.77Adequate
Greed-AvoidanceUnassuming / Status-seeking4 / 8.69Below the adequacy line
ModestyUnassuming / Status-seeking4 / 8.81Good

One scale sits below that line. Greed-Avoidance is published at .69, under the .70 adequacy threshold, and we are not going to bury it in a footnote. It means the ten items IPIP grouped under that name hang together less tightly than the others do, so scores on it are noisier — and because it feeds the Unassuming / Status-seeking axis, a quarter of that axis carries the extra noise. The honest response to a weak scale is to name it and read it less confidently, not to quietly drop it and report five axes instead of six.

What alpha does not tell you

Alpha is one number about one thing: whether the items on a scale move together. It says nothing about whether the scale measures what its name claims, nothing about whether your score would be the same next month, and nothing about whether it predicts anything at all about your life. A high alpha on a badly named scale is still a badly named scale. Most of this market reports alpha and presents it as validity. It is not validity, and we are not going to pretend otherwise on our own page.

Here is what we do not have, stated as plainly as we can put it:

  • No test-retest reliability. We have never re-tested anyone, so we cannot tell you how much your score would move if you sat the same questions again in a month.
  • No convergent validity. We have not correlated these scores against any established instrument, so we cannot show that they agree with anything outside themselves.
  • No norms, and therefore no percentiles. IPIP publishes no norms for these scales and explicitly recommends that users develop local ones. Your 0–100 figure is your own item average rescaled — where your answers sat on the response scale — not your standing against a comparison group. A 70 does not mean you scored higher than 70% of anybody.
  • No reliability figures from our own respondents, in either language. Every coefficient on this page is IPIP's.

What we will publish, and when. Once we hold at least 300 completed sittings per language, we will publish from our own data: coefficient omega alongside alpha for all 24 scales, because omega assumes a measurement model that is realistic where alpha assumes one that is not; test-retest correlations with the interval stated; and every figure reported separately for Thai and for English.

Separately, because measurement invariance across languages cannot be assumed. A Thai and an English respondent may not be answering quite the same question even when the sentences translate cleanly, and pooling the two would hide that rather than test it. Reporting them apart is what makes the difference visible if it is there.

If those figures come back worse than IPIP's, this page will say so. A methodology page that can only ever report good news is marketing.

The ± band on your scores

Every score in the full report is rendered with a band, in the form 62 ±5. That band is computed from your answers and nothing else, subject to the one floor stated in the next paragraph: it is the standard error of your own item mean on that scale — the spread of your responses on the scale, divided by the square root of how many items you answered on it. Axis bands are the same quantity propagated across the axis's four scales.

One floor, stated because it is the only part of the band that is not your arithmetic. We never render a band narrower than ±3, however consistent the answers. A respondent who gives all four items on a scale the same answer has an observed spread of exactly zero, and SD / √k would print ±0 — a claim of perfect precision from four questions, which nothing in this instrument supports. Where you see ±3 on a scale you answered uniformly, that 3 is the floor rather than a measurement, and no band you see is ever narrower than the truth. It is the only constant anywhere in the numbers we show you.

Read it as: how consistently you answered this particular scale. Someone who answers four items 5, 5, 4, 5 gets a narrow band. Someone who answers 5, 1, 4, 2 gets a wide one — and the wide band is the honest report of that, because the average of two contradictory groups of answers sits between them and describes neither. Answering more items narrows the band, which is the entire reason the extended sitting exists.

What the band is not: it is not a population standard error, and it is not a confidence interval in the statistical sense. It uses no external data, because we have no external data to use. It is labelled on the report as the spread in your own answers, and that is exactly what it is.

The band we would rather show is the standard error of measurement: SEM = SD × √(1 − reliability). It uses a norm sample's spread and the scale's reliability to say how far a score would be expected to bounce on retesting, which is the question people actually mean when they ask how sure we are. It needs two things we do not have — reliability measured on our own respondents, and standard deviations from a norm sample — and both arrive with the 300-per-language milestone above. Until then we render the number we can genuinely compute from your data, and we tell you which number it is.

Where a score's band straddles the midpoint of an axis, the report says the pole assignment is inside the band rather than handing you a letter with confidence it has not earned.

A caution from the literature, quoted in full

Johnson (2014) developed the IPIP-NEO-120, a public-domain inventory in the same family as these scales, on a sample of 619,150 respondents. His own conclusion about short public-domain facet scales was this:

the IPIP-NEO-120 facet scales have sufficient reliability for research studies, but probably should not be used to make important decisions about individuals.

Johnson, J. A. (2014). Journal of Research in Personality.

That sentence is about a sibling instrument rather than this one, and we quote it because it applies here with more force rather than less: our standard sitting uses four items per scale, and we have no evaluation study of our own at all. Johnson also reports mean alphas of roughly .68 to .75 at four items per scale against .80 to .82 at ten — length buys less than people expect, and no amount of length buys validity.

What this is for, and what it must not be used for

Intended use: self-reflection, and conversation. A vocabulary for describing how you tend to operate, and a document worth arguing with. One of the more useful outcomes is that you read a paragraph and think that is not me — because now you have words for something you did not have words for before.

Not intended, and not suitable, for any of the following:

  • Clinical or diagnostic use of any kind. This is not a mental-health assessment, it detects no disorder, and it is not a substitute for speaking to a professional.
  • Hiring, promotion, performance review, or any employment decision.
  • Admissions, selection, screening, or eligibility for anything.
  • Any consequential decision made about a person — by an employer, an institution, or by you about someone else.

To restate what the about page already says: this instrument is not clinically validated, not peer reviewed, and not published as research. The items come from a published public-domain source; the scoring, the axes, the archetypes and the written report are ours, and no one outside this company has evaluated them.

You will not find testimonials, star ratings, user counts or press mentions anywhere on this site. That is not a design preference: we have not run the studies and we do not have the users, so publishing any of it would be a claim we could not substantiate if someone asked us to. This page is what we have instead, and we would rather be judged on it.

Two lengths: 96 and 192

The standard sitting is 96 questions: four items on each of the 24 scales. The extended sitting is 192: eight items on each. Both cover all 24 scales and all six axes; the longer one simply asks more about each.

Reverse-keyed items, counted. Of the 96 items in the standard sitting, 29 are reverse-keyed — 30%. Of the 192 in the extended sitting, 82 are — 43%. The full 240-item IPIP-HEXACO bank we draw from is 122 of 240, or 51%, but the bank figure is not the one that describes your sitting, so it is not the one we quote elsewhere. Reverse keying is a partial guard against agreeing with every statement, not a cure: the published evidence is mixed, and acquiescence correction has been found to work less reliably outside individualistic Western samples. One scale is a known gap — Dependence has no reverse-keyed item in the standard sitting, so it carries none of this guard at four items; that is a defect in our item selection, not a property of the source bank, and it is on the list to fix by reselecting which of its ten items the standard sitting uses.

The reason to offer the longer one is arithmetic, not marketing: the standard error of your item mean falls with the square root of the item count, so doubling from four items to eight narrows each band by about 30%. Tighter bands are the whole benefit, and it is the only thing we claim for it.

The question set is fixed. We do not change which questions you see based on how you answer them, and neither sitting is shortened or extended as you go. Everyone taking the standard sitting answers the same 96 items covering the same 24 scales, in a balanced order.

Sources

Everything above traces to one of these. Where a figure is IPIP's, the first link is where to check it.

  • International Personality Item Pool — the IPIP-HEXACO item keys and the alpha coefficients reproduced above. https://ipip.ori.org/newHEXACO_PI_key.htm
  • International Personality Item Pool — permission and public-domain status. https://ipip.ori.org/newPermission.htm
  • Ashton, M. C., Lee, K., & Goldberg, L. R. (2007). Personality and Individual Differences. The source publication for the IPIP-HEXACO scales.
  • Goldberg, L. R. (1999). A broad-bandwidth, public-domain, personality inventory measuring the lower-level facets of several five-factor models.
  • Ashton, M. C., & Lee, K. (2008). Journal of Research in Personality, 42(5). On what the sixth factor accounts for beyond five-factor models.
  • Johnson, J. A. (2014). Journal of Research in Personality. Development of the IPIP-NEO-120, and the caution quoted above.
  • EFPA Test Review Model and the BPS test review criteria — the source of the internal-consistency bands used in the table above.