Research Wiki · Methods Lab

Reliability, validity and measurement science

Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation.

Methods family

Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation.

Use methods only after defining the estimand, data structure, assumptions, validation plan and decision consequence.

Classical test theory

Classical test theory is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Read the full entry →
Reliability analysis

Reliability analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Read the full entry →
Cronbach’s alpha

Cronbach’s alpha summarises internal consistency under assumptions that are often stronger than users realise. A high alpha does not prove unidimensionality, validity or item quality.

Read the full entry →
McDonald’s omega

McDonald’s omega estimates scale reliability from a factor model and can be more appropriate than alpha when item loadings differ. Its interpretation depends on a defensible measurement model.

Read the full entry →
Split-half reliability

Split-half reliability is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Test-retest reliability

Test-retest reliability is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Read the full entry →
Inter-rater reliability

Inter-rater reliability is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Cohen’s kappa

Cohen’s kappa is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Fleiss’ kappa

Fleiss’ kappa is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Intraclass correlation

Intraclass correlation is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Construct validity

Construct validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Convergent validity

Convergent validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Discriminant validity

Discriminant validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Criterion validity

Criterion validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Predictive validity

Predictive validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Concurrent validity

Concurrent validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Face validity

Face validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Content validity

Content validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Nomological validity

Nomological validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Item analysis

Item analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Item-total correlation

Item-total correlation is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Scale purification

Scale purification is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Item-response theory

Item-response theory is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Rasch modelling

Rasch modelling is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Two-parameter logistic IRT

Two-parameter logistic IRT is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Three-parameter logistic IRT

Three-parameter logistic IRT is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Graded-response models

Graded-response models is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Differential item functioning

Differential item functioning is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Measurement invariance

Measurement invariance is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Common-method bias

Common-method bias is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Response-style analysis

Response-style analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Acquiescence adjustment

Acquiescence adjustment is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Extreme-response-style analysis

Extreme-response-style analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Browse linked method entries

Deep entry

Classical test theory

Classical test theory is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Deep entry

Reliability analysis

Reliability analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Deep entry

Cronbach’s alpha

Cronbach’s alpha summarises internal consistency under assumptions that are often stronger than users realise. A high alpha does not prove unidimensionality, validity or item quality.

Open method
Deep entry

McDonald’s omega

McDonald’s omega estimates scale reliability from a factor model and can be more appropriate than alpha when item loadings differ. Its interpretation depends on a defensible measurement model.

Open method
Deep entry

Split-half reliability

Split-half reliability is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Test-retest reliability

Test-retest reliability is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Deep entry

Inter-rater reliability

Inter-rater reliability is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Cohen’s kappa

Cohen’s kappa is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Fleiss’ kappa

Fleiss’ kappa is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Intraclass correlation

Intraclass correlation is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Construct validity

Construct validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Convergent validity

Convergent validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Discriminant validity

Discriminant validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Criterion validity

Criterion validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Predictive validity

Predictive validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Concurrent validity

Concurrent validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Face validity

Face validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Content validity

Content validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Nomological validity

Nomological validity is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Item analysis

Item analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Item-total correlation

Item-total correlation is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Scale purification

Scale purification is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Item-response theory

Item-response theory is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Rasch modelling

Rasch modelling is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Two-parameter logistic IRT

Two-parameter logistic IRT is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Three-parameter logistic IRT

Three-parameter logistic IRT is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Graded-response models

Graded-response models is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Differential item functioning

Differential item functioning is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Measurement invariance

Measurement invariance is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Common-method bias

Common-method bias is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Response-style analysis

Response-style analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Acquiescence adjustment

Acquiescence adjustment is a statistical or analytical concept within reliability, validity and measurement science. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Extreme-response-style analysis

Extreme-response-style analysis is a method within reliability, validity and measurement science. Methods for evaluating whether scales and measures are consistent, valid, invariant and fit for the intended interpretation. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method

Need the method applied to a live decision?

Share the market, customer, product or investment question that needs a defensible evidence design.

Brief Ninth Atlas