Reliability concerns consistency and measurement error. Validity concerns whether the instrument measures the construct it claims to measure. Strong internal consistency is not proof of content, structural or construct validity.
Evaluate the evidence in layers
- Define the construct. Specify what is being measured, for whom, in which context and for what use.
- Assess content validity. Check whether the items are relevant, comprehensive and understandable for the intended construct and population.
- Assess internal structure. Evaluate dimensionality and whether internal-consistency estimates are appropriate for that structure.
- Assess reliability and measurement error. Examine score stability under the relevant repeated-measurement conditions.
- Assess construct validity and responsiveness. Test prespecified hypotheses about relationships, group differences and change where applicable.
Why alpha alone is insufficient
A high reliability coefficient may reflect redundant items or a narrow item set. It does not show that the instrument covers the intended construct, works across groups, has the expected factor structure or detects meaningful change. COSMIN treats reliability, validity and responsiveness as related but distinct domains.
Worked use case
A six-item scale produces a high alpha, but all items describe only one part of a broader construct. Report the internal-consistency result, then examine content coverage and structural validity. Do not call the scale “validated” on the basis of alpha alone.
Evidence boundary
Measurement evidence is use-specific. An instrument supported in one language, population, administration mode or purpose may require additional evidence before use in another context.
Primary sources
- COSMIN. Manual version 2.0 and taxonomy of measurement properties. Retrieved 4 August 2026.
- COSMIN. Deciding what to measure. Retrieved 4 August 2026.